大家都在押注更大的模型。但對于物理人工智能而言智能本身并非倍增器學習才是。彼得·路德維希2026年7月28日A billion machines will become autonomous or intelligent over the next ten years. Cars, trucks, tractors, mining haulers, defense systems, warehouse robots, humanoids—the physical economy will be rebuilt around software that perceives, decides, and acts.The prevailing assumption about how we get there goes something like this: models keep improving, world models mature, foundation models for robotics arrive, and autonomy falls out the other end. Intelligence is the whole game; scale the intelligence and the machines will follow.I’ve spent a decade helping Applied Intuition build the software infrastructure behind many of the world’s most ambitious physical AI programs, from software-defined vehicles and autonomous trucks to construction equipment, mining systems, defense platforms, and robotics. That vantage point has given me a front-row seat to where the industry is accelerating—and where it continues to slow itself down. I’m as bullish on the intelligence as anyone, but the prevailing assumption gets the math wrong. Deployed physical AI is a product of two variables: the capability of the models, and the capacity of the engineering system around them—i.e. how requirements become software, how software gets validated, and how validated systems get deployed, monitored, and improved. The industry has largely poured everything into the first variable while the second sits roughly where it was a decade ago, built for quarterly releases and hundred-person integration teams. Frontier intelligence running on a legacy engineering system doesn’t produce frontier outcomes, because the old engineering system is the limiting factor.The contrarian bet, then, isn’t against intelligence. It’s that the next order of magnitude in physical AI comes from making the engineering system as intelligent as the models it carries. The industry’s roadmap, however, rests on a handful of assumptions that once made sense but no longer match where physical AI is headed.Smarter Models Don’t Create Deployed MachinesThe gap between a capable model and a certified, operating machine is enormous, and model quality alone doesn’t close it. A model that’s 20% better in benchmark terms still has to be integrated with many other software components, tested across millions of scenario variations, traced against safety requirements, validated on hardware, rolled out to a fleet, and monitored in the field. In a typical program, that pipeline — not the model — sets the tempo. Teams take delivery of a meaningfully better model and then spend two quarters proving it’s safe to ship.This is why world models, as remarkable as they are, won’t get us to a fully autonomous future on their own. They are advancing faster than engineering organizations can keep up. Every leap in model capability relocates the bottleneck rather than eliminating it. The constraint moves downstream, from “can the machine perceive the world?” to “can we validate, integrate, and operate what the machine can now do?” A team whose validation cycle takes months is, in effect, throttling frontier AI down to the speed of its own process.Here’s the implication the industry hasn’t priced in: as models commoditize toward the frontier, two companies with access to the same intelligence will have wildly different outcomes. The difference will be determined by how fast their engineering systems can absorb what the models can do. Intelligence is becoming ubiquitous. The ability to operationalize it isn’t.Digital AI Doesn’t Transfer to Physical AIThe second assumption is subtler: that the agentic revolution happening in digital work will naturally extend to physical AI. Direct the coding agents and copilots at the autonomy stack, and the same productivity gains will follow.They won’t, because most digital AI stops at documents, conversations, and code. Physical AI work doesn’t live there. It lives in drive logs and sensor data, in simulation runs and hardware-in-the-loop test rigs, and in requirements databases and validation reports. It lives in fleet telemetry streaming from real vehicles on real roads and job sites. An agent that has never seen a disengagement, doesn’t know why a perception regression matters, and can’t trace a requirement to a test case, is not a productivity tool in this domain. It’s a liability with a friendly interface.Making agents genuinely capable in deploying physical AI itself is a frontier intelligence problem, one that is entirely different than training a bigger model. The agents need access to the actual data and tools of the trade—simulators, data pipelines, validation systems—through interfaces hardened enough to trust. They need embedded domain judgment and the accumulated knowledge of what “validated” actually means when the artifact ships into a multi-ton machine. And they need evaluation and governance built in, because in this domain a plausible answer and a correct answer can be separated by fatal consequences.When I say physical AI needs its own agentic platform, I’m talking about a platform that combines state-of-the-art models grounded in the data layer, tooling, and domain expertise of physical systems, with evaluation and governance native to the platform. It is not a chatbot bolted onto engineering tools, and it is not something you get by fine-tuning a general-purpose agent. It’s a different architecture, and it demands as much AI innovation as the models themselves.The future of physical AI wont be determined by the smartest models alone, but by the engineering systems that turn intelligence into autonomous machines at scale.Speed Doesn’t Compromise SafetyWhen discussing agentic capabilities in physical AI development, the reflexive objection is that agents have no place in safety-critical engineering. Automation means moving fast and perhaps getting a few things wrong. It’s a reasonable position when the thing in question weighs several tons.However, in safety-critical systems, the speed of your feedback loopisa safety mechanism. When validation takes weeks, teams test at irregular milestones. When it takes minutes, they test on every change. Problems surface earlier, when they’re cheap to fix. Coverage expands to orders of magnitude more scenarios. Requirements get implemented more accurately and verified more often. The slow, careful-looking process isn’t by its nature the safest one. It’s often the one where defects age quietly for months before anyone notices.What actually makes agents safe in this domain isn’t slowing them down. It’s drawing the line correctly. Automate development, validation, and operations workflows, so safety-critical issues get resolved faster, and refuse to automate certification, regulatory sign-off, and final engineering judgment. High-stakes agents propose; humans decide. Anything touching production systems runs behind approval gates. This isn’t a temporary concession while the models improve. It’s the correct permanent architecture for physical AI, in the same way that a well-designed autonomous vehicle has a defined operational domain rather than unlimited authority.When agents help build physical AI faster, that speed can make the end product safer.Models Don’t Compound. Systems Do.When you understand these assumptions are holding physical AI back, the logical next step is combining better intelligence with faster learning, creating an agentic flywheel where frontier models and frontier engineering systems feed each other.A continuously turning flywheel looks like this. A machine underperforms in the field. Agents mine the operational data to find out where and why, or synthetically generate the scenarios that expose the gap. Findings become requirements, and requirements become test cases. Test cases become validated software, which is deployed to the fleet. The fleet generates new data, which makes the models better, which makes the agents better, which speeds up the next turn of the loop. Every turn used to take months and a room full of specialists. Each stage that agents accelerate doesn’t just save time; it increases the number of turns, and each turn compounds both the speed and the intelligence. Smarter models turn the loop faster. A faster loop makes the models smarter. That’s the compounding effect the industry is leaving on the table when it treats intelligence as the whole game.We know this flywheel is real because we’ve been running it on ourselves. Applied Intuition has spent a decade at the frontier of physical intelligence — perception, simulation, validation, vehicle software across automotive, trucking, mining, agriculture, and defense. Over the past year we built an agentic platform grounded in that infrastructure. It’s called Dana, and our engineers have built more than a thousand internal apps and agents on it. With Dana, development cycles have become roughly 20x faster, with higher output quality. Deployments went from once every few weeks to multiple times a day. Applications that took months to build now take days or hours. And, tellingly, applications that would never have justified months of effort now get built regularly. When the cost of building drops by an order of magnitude, the set of things worth building expands by more than an order of magnitude.The intelligence made the engineering system possible; the engineering system made the intelligence matter. Neither alone gets you there.What This Means for the Next DecadeIf the industry keeps betting on intelligence alone, the next decade of physical AI looks like a slow one: dazzling demos, decade-long programs, and a widening gap between what machines can do in a lab and what’s actually operating in the world. The models will be extraordinary and the deployment curve will stay stubbornly flat, because every improvement will queue up behind engineering organizations that absorb change at last year’s speeds.Pair frontier intelligence with agentic engineering systems and the curve bends. Physical AI starts reaching its full potential. Farms that stabilize output through labor and climate shocks. Mines with continuous, safer extraction. Freight networks that self-route around disruption. Defense systems that hold under degraded conditions. Multi-hundred-billion-dollar markets converging on the same stack, with the learning loop at the center of all of them.Software ate the world by making it cheap to build applications for the digital economy. Physical AI will do the same for the physical one, but only if building intelligent machines becomes as fast and iterative as building software, without compromising the discipline safety-critical systems demand. That takes the best modelsanda reinvention of how we engineer, and the second half is the one almost nobody is building.Twenty years from now, we won’t remember which company had the best world model in 2027. We’ll remember which company figured out how to continuously turn intelligence into deployed systems. That’s the problem we’ve been working on.未來十年將有十億臺機器實現自主運行或具備智能。汽車、卡車、拖拉機、礦用運輸車、國防系統、倉庫機器人、人形機器人——實體經濟將圍繞能夠感知、決策和行動的軟件進行重建。目前普遍認為實現這一目標的過程大致如下模型不斷改進世界模型日趨成熟機器人技術的基礎模型最終形成自主性也隨之而來。智能是關鍵所在提升智能規模機器自然會隨之發展。過去十年我一直幫助Applied Intuition公司構建支撐眾多全球最具雄心的物理人工智能項目的軟件基礎設施涵蓋軟件定義車輛、自動駕駛卡車、建筑設備、采礦系統、國防平臺和機器人等領域。這段經歷讓我得以近距離觀察行業的發展趨勢以及它自身發展停滯不前的原因。我對人工智能的未來充滿信心但目前普遍的假設存在邏輯錯誤。物理人工智能的部署取決于兩個變量模型的能力以及圍繞模型構建的工程系統的容量——即需求如何轉化為軟件、軟件如何得到驗證以及經過驗證的系統如何部署、監控和改進。行業目前幾乎將所有資源都投入到了第一個變量上而第二個變量卻仍然停留在十年前的水平仍然沿用著季度發布和百人集成團隊的模式。在老舊的工程系統上運行的前沿智能無法產生前沿成果因為老舊的工程系統本身就是制約因素。因此這種反主流觀點并非反對智能本身而是認為物理人工智能的下一個飛躍階段將來自于使工程系統本身達到與其承載的模型相同的智能水平。然而行業路線圖卻建立在一些曾經合理但如今已不再符合物理人工智能發展方向的假設之上。更智能的模型不會創建已部署的機器一個性能優異的模型與一臺經過認證、可實際運行的機器之間存在著巨大的差距單靠模型質量的提升并不能彌合這一差距。即使一個模型在基準測試中提升了20%它仍然需要與其他眾多軟件組件集成在數百萬種場景變化中進行測試對照安全要求進行驗證在硬件上進行驗證部署到整個機隊并在現場進行監控。在一個典型的項目中決定項目進度的是整個流程而不是模型本身。團隊拿到一個性能顯著提升的模型后還需要花費兩個季度的時間來證明其安全性才能最終交付使用。這就是為什么世界模型雖然卓越但僅靠它們本身無法帶我們走向完全自主的未來。它們的進步速度遠遠超過了工程組織的響應速度。模型能力的每一次飛躍都只是將瓶頸轉移到了其他地方而不是消除它。限制因素向下游轉移從“機器能否感知世界”變成了“我們能否驗證、集成并運行機器現在能夠做到的事情”一個驗證周期需要數月之久的團隊實際上是在將前沿人工智能的速度限制在自身流程的速度之內。行業尚未充分考慮以下影響隨著模型向前沿領域商品化兩家擁有相同情報資源的公司最終會得出截然不同的結果。這種差異取決于它們的工程系統吸收模型功能的速度。情報正變得無處不在但將其轉化為實際應用的能力卻遠未普及。數字人工智能無法轉化為物理人工智能第二個假設更為微妙數字工作中正在發生的智能體革命自然會擴展到物理人工智能領域。引導編碼智能體和副駕駛關注自主系統架構同樣的生產力提升也將隨之而來。他們不會因為大多數數字人工智能都止步于文檔、對話和代碼。而物理人工智能的工作并非如此。它存在于駕駛日志和傳感器數據中存在于模擬運行和硬件在環測試平臺中存在于需求數據庫和驗證報告中。它存在于來自真實道路和作業現場真實車輛的車隊遙測數據流中。一個從未經歷過用戶脫離、不了解感知回歸為何重要、也無法將需求追溯到測試用例的智能體在這個領域并非生產力工具。它只是一個界面友好的累贅。讓智能體真正具備部署物理人工智能的能力是前沿智能領域的一大難題這與訓練一個更大的模型截然不同。智能體需要通過足夠可靠的接口訪問實際數據和工具——模擬器、數據管道、驗證系統等等。它們需要具備嵌入式領域判斷能力以及在將產品交付給一臺重達數噸的機器時“驗證”的真正含義。此外它們還需要內置評估和治理機制因為在這個領域一個看似合理的答案和一個正確的答案之間可能存在致命的后果。我所說的物理人工智能需要其自身的代理平臺指的是一個將基于物理系統數據層、工具和領域專業知識的尖端模型與平臺原生評估和治理功能相結合的平臺。它并非簡單地將聊天機器人附加到工程工具上也不是通過微調通用代理就能實現的。它是一種不同的架構并且對人工智能創新提出了與模型本身同等的要求。物理人工智能的未來不僅僅取決于最智能的模型還取決于將智能大規模轉化為自主機器的工程系統。速度并不影響安全性在討論物理人工智能開發中的智能體能力時人們往往會反駁說智能體在安全攸關的工程領域沒有立足之地。自動化意味著快速行動但也可能導致一些錯誤。當目標物體重達數噸時這種觀點不無道理。然而在安全關鍵型系統中反饋循環的速度本身就是一種安全機制。當驗證需要數周時間時團隊只能在不規則的里程碑節點進行測試。而當驗證只需幾分鐘時他們就能對每一次變更進行測試。問題會更早地被發現從而降低修復成本。測試覆蓋范圍也因此擴展到更多場景。需求能夠得到更準確的實現并被更頻繁地驗證。緩慢而謹慎的流程本質上并非最安全的。在這種流程中缺陷往往會悄無聲息地存在數月之久直到有人發現。在這個領域真正確保智能體安全的并非降低其運行速度而是正確劃定界限。自動化開發、驗證和運維工作流程以便更快地解決安全關鍵問題同時拒絕自動化認證、監管審批和最終工程判斷。高風險智能體提出方案由人類做出決定。任何涉及生產系統的事項都必須經過審批流程。這并非模型改進期間的臨時妥協而是物理人工智能的正確永久架構正如設計良好的自動駕駛汽車擁有明確的運行范圍而非無限的權限一樣。當智能體幫助更快地構建物理人工智能時這種速度可以使最終產品更安全。模型不會產生復合效應系統才會。當你理解這些假設阻礙了物理人工智能的發展時合乎邏輯的下一步就是將更智能的技能與更快的學習速度結合起來創建一個智能飛輪使前沿模型和前沿工程系統相互促進。一個持續運轉的飛輪看起來是這樣的一臺機器在實際應用中性能不佳。智能體會挖掘運行數據找出問題所在及原因或者合成場景來暴露差距。發現的問題轉化為需求需求轉化為測試用例。測試用例轉化為經過驗證的軟件并部署到整個系統中。系統生成新的數據從而改進模型改進智能體進而加快循環的下一輪。過去每一輪循環都需要數月時間并且需要一屋子的專家參與。智能體加速的每一個階段不僅僅節省了時間它增加了循環次數而每一次循環都會同時提升速度和智能水平。更智能的模型能夠更快地完成循環。更快的循環速度又使模型更加智能。這就是行業將智能視為全部時所忽略的復合效應。我們深知這種飛輪效應真實存在因為我們一直在自身實踐中驗證它。Applied Intuition 十年來一直致力于物理智能領域的前沿研究涵蓋感知、仿真、驗證以及汽車、卡車、采礦、農業和國防等行業的車輛軟件。過去一年我們基于這一基礎架構構建了一個智能體平臺名為 Dana。我們的工程師已在其上開發了超過一千個內部應用程序和智能體。借助 Dana開發周期縮短了約 20 倍同時輸出質量也顯著提高。部署頻率從幾周一次提升到每天多次。過去需要數月才能構建的應用程序現在只需幾天甚至幾小時即可完成。更重要的是那些過去根本不值得花費數月時間開發的應用程序現在也開始定期構建。當構建成本降低一個數量級時值得構建的項目數量級也會相應增加。智慧使工程系統成為可能工程系統使智慧發揮作用。兩者缺一不可。這對未來十年意味著什么如果業界繼續只依賴智能那么未來十年物理人工智能的發展將會十分緩慢令人眼花繚亂的演示、長達十年的項目以及機器在實驗室中的表現與實際應用之間日益擴大的差距。模型固然會非常出色但部署曲線卻會始終保持平緩因為每一項改進都將排在那些以去年速度吸收變革的工程團隊之后。將前沿智能與智能工程系統相結合曲線將發生轉變。物理人工智能開始充分發揮其潛力。農場能夠應對勞動力和氣候沖擊穩定產量。礦山能夠持續、安全地開采。貨運網絡能夠自動繞過中斷。防御系統能夠在惡劣條件下保持有效運行。數千億美元的市場將匯聚到同一技術棧上而學習循環則是所有這些市場的核心。軟件通過降低構建數字經濟應用的成本徹底改變了世界。物理人工智能也將對物理世界產生同樣的影響但這只有在構建智能機器的速度和迭代性能夠與構建軟件一樣快并且不損害安全關鍵系統所要求的嚴謹性時才能實現。這需要最佳模型和對工程方式的徹底革新而后半部分幾乎無人涉足。二十年后我們不會記得哪家公司在2027年擁有最佳的世界模型。我們會記住哪家公司找到了將情報持續轉化為可部署系統的方法。這正是我們一直在努力解決的問題。