SĂĽeda Asil
Corporate
- Thread Author
- #1
🏠What Robots Learn: The Data-Model Divide
Contracts for companies using robots in warehouses or factories typically grant them ownership of operational data, while giving vendors unrestricted rights to train their models with this data. However, there is no commercially accepted way to measure how much a model has learned from your facility. Even if vendors remove your influence from the model, it doesn't give you anything usable back. Current legal regulations and standard industrial automation contracts do not reach the model layer; this is only possible through a negotiated contract.
⚡ Lessons from Energy Trading: The Difference Between Data and Learning
Before physical AI, I managed a fund that traded electricity and gas. The distinction that the industrial automation sector overlooks is something every energy trader knows. Gas is a commodity that can be stored, delivered, divided, and its ownership moves with the molecule. But when you burn gas to generate electricity, that electricity mixes into the grid and cannot be retrieved. No meter will give you your own electrons back.
Data behaves like gas. The trained model, which is everything that makes the robot fleet get better each quarter, behaves like grid electricity.
đź§ Learning Cannot Be Disaggregated
When an industrial operator's data is used to train a common model, the impact of that data on the model becomes inseparable. It becomes part of the model's behavior, blended with contributions from all other sites. There is no copy to give back to you as "your share."
Extracting an operator's data from a trained model usually means retraining without that data, which is rarely practical at fleet scale. Complete "unlearning" only works under specific architectures. In any case, extraction is the wrong test: even if a system successfully "unlearns" a customer's data, it cannot return that customer's contribution as an independent entity.
Measurement is no better. There is no commercially accepted method that can tell how much an industrial site's data contributed to fleet performance. A before-and-after comparison shows that training helped, not which data helped. Data is divisible; learning, once mixed into a common model, is not, unless the system was built for extraction before training began.
Even the intuitive compromise ("Your fine-tune is yours.") runs into the same limitation. A site fine-tune is typically a small set of adjustments that adapt the vendor's general model to a site. Unlike shared learning, it can be stored and delivered as a separate file. But it only works with a compatible base model and runtime. Technical separability does not make it useful on its own, and no contract changes that. What a contract can do is license the customer to continue running it after exit.
⚖️ The Law Is Missing a Layer
Let's re-read the legal records with these physical realities in mind.
The proposed FTC John Deere settlement would be a landmark outcome on repair access. Access to repair resources "need not include ownership or possession of Deere’s intellectual property rights." Farmers gained access to diagnostic tools, not the model. The ruling never reaches the model behind Deere’s machine learning spraying system, See & Spray; that model remains with Deere by silence, not by exception.
The EU Data Act, effective September 2025, grants distributors a legal right to the data their machines generate. It reaches raw and pre-processed data. However, Article 15 excludes information derived through proprietary complex algorithms "unless otherwise agreed." Those three words return the entire model layer to contract, and in practice, to the vendor's paper.
Two instruments, two jurisdictions, one line: the law leaves out the model. While distributors gain access to the raw material, the composite derivative remains outside of every tool they've gained.
đź’° Large Buyers Get It
Until recently, model rights were rarely explicitly addressed. But at the top of the market, this has changed.
The largest industrial automation buyers are now writing the model layer directly into their physical AI deal drafts. I've seen the drafts and negotiated against them. Their demands ascend like a ladder: ownership of raw data, then derived data, then customer-specific weights, then perpetual licenses to any base model required to run those weights. The final rung is the architecture itself: royalty triggers for systems that only resemble what was co-built. Buyers have understood that a fine-tune only works on the vendor's base, so the demands ascend until they reach the platform.
Two details stand out. First, the same documents demanding perpetual economic rights to everything that benefits from a robotic deployment treat the hardware as a cost-plus item. When the most sophisticated automation buyers allocate negotiation capital, they issue a depreciation schedule. Hardware depreciates. Learning compounds.
Second, the most fiercely contested clause is rarely data ownership. It's pool membership: whether a customer's live production data can train the fleet's common brain. If you spend a quarter teaching the fleet how to handle your most fragile SKU, that skill can be shipped to a competitor working with the same vendor in the next update. While a single site's deployment data is a small fraction of the total training data, these edge cases in production have disproportionate value.
Today, fleet models stubbornly remain site-specific, and a fine-tune from one facility often breaks in the next building. But the perpetual licenses being signed now will govern the models five years from now, if cross-site transfer works. The asset is being allocated before it fully exists, which is when it is cheapest to acquire by default.
Automation buyers are demanding approval gates for training, and this gate is technically real: site-specific fine-tuning can increasingly run on a single in-house GPU, with operational data never leaving the facility. But commingling is not inevitable. It is a pipeline decision made upstream, in the contract. Vendors argue for base models trained on anonymized data across all sites, improving all customers. Some drafts now point to the honest solution: a priced data license for training rights. Others impose model provenance obligations, documented inputs, and decision logic years before any regulator.
Meanwhile, the same buyers are demanding source code: warehouses, embedded engineering teams, escrow with trigger release. Access to source code is an incomplete solution; it transfers a snapshot, not the learning loop that developed it.


















