Vision Part 1: Cameras, LiDAR, Depth Sensors: The $38B Format War
Cameras, LiDAR, and depth sensors are fighting for the same spot. The winner locks in a supply chain for a decade.
Eyes of the Machine: A 3-Part Series
1. Cameras, LiDAR, Depth Sensors: The $38B Format War
2. Software Eats the Sensors
3. Sony Won the Last Camera War. Who Wins This One?
A camera-based perception stack costs $120 to $400. A LiDAR-supplemented stack costs $500 to $4,000. Same spot on the robot’s head. Same job description. The price gap is 5 to 10x, and the humanoid robotics industry is splitting into camps over which sensor technology earns the right to sit there. The winner determines which supply chain captures value for the next decade. We have seen this movie before.
If this analysis is useful, subscribe for free to follow the full Eyes of the Machine series.
The Format War Precedent
Toshiba surrendered on February 19, 2008. The Blu-ray versus HD DVD format war lasted two years, burned billions, and ended with a single winner. The lesson was not about which technology was “better.” HD DVD was cheaper. Blu-ray had more storage. The lesson was about adoption. Sony put Blu-ray in every PlayStation 3. Toshiba could not match that installed base. The supply chain locked in around Blu-ray. It still exists today, 18 years later.
The perception stack format war is happening right now. Once enough OEMs commit to a sensor architecture, the supply chain crystallizes. Component manufacturers optimize for the winning format. Software stacks mature around it. Switching costs compound. The technology that wins does not just win a spot on the robot. It wins a decade of supply chain lock-in.
The question is which camp has the PlayStation 3.
The Three Contenders
Every humanoid robot needs to answer one question: what is around me, and how far away is it? Each sensor answers that question differently, with a different cost structure and a different weakness.
Cameras are the cheapest option. A typical humanoid carries six to eight modules covering 360 degrees. Total cost: $90 to $400. Sony, Samsung, and OmniVision dominate the CMOS image sensor market. The weakness is compute: cameras need light, and extracting depth from 2D images requires GPU-intensive AI models. That compute cost shows up in the power draw and thermal envelope of the robot’s head.
LiDAR shoots laser pulses and measures return time to build a precise 3D point cloud. It works in complete darkness and delivers millimeter-level depth accuracy. The weakness is cost. A single LiDAR unit can cost more than the entire rest of a camera-based stack. For indoor humanoid applications, 200-meter range is overkill. The supply chain is thinner: Hesai, LIVOX, Luminar, and Ouster are the primary suppliers. The landscape is consolidating.
Depth sensors (structured light, time-of-flight, stereo) fill the gap. Better depth than a camera alone, cheaper than LiDAR, optimized for manipulation-range tasks at 0.5 to 5 meters. Intel’s RealSense is the most widely deployed. Orbbec and Stereolabs (now owned by Ouster) compete in the same space. The weakness is range. A depth sensor that works for picking up a coffee cup does not work for navigating a warehouse.
The cost spread is not subtle: a 5 to 10x gap that shows up directly in the bill of materials.
The camera camp argues the gap proves cameras win on economics. The LiDAR camp argues the gap is closing. Both are right. The question is timing.
The OEM Bets
The camps are forming. Five of the most visible humanoid robotics companies have made their sensor choices.
Tesla is the loudest voice in the camera-first camp. Optimus uses cameras exclusively. No LiDAR. No dedicated depth sensors. Tesla’s bet is that its AI vision stack, trained on billions of miles of driving data, can extract sufficient 3D understanding from 2D feeds. The cost argument is devastating: Tesla’s perception stack is the cheapest in the industry by a factor of five. The risk is that “sufficient” 3D understanding is not the same as “reliable” 3D understanding for manipulation tasks where millimeter precision matters.
Figure AI takes the pragmatic middle ground. Figure 02 pairs cameras with Intel RealSense depth sensors. No LiDAR. Better manipulation-range perception than cameras alone, at a fraction of LiDAR’s cost. Figure’s stack lands in the $200 to $500 range.
Unitree surprises. The G1 and H1 carry a LIVOX-MID360 LiDAR unit plus depth cameras at roughly $800 to $1,200. Unitree absorbed that cost because the robot’s navigation and SLAM capabilities required it. They are buying all three.
1X Technologies uses camera pairs to derive depth through stereo vision, a technique that is computationally expensive but avoids dedicated depth hardware. Their bet is that compute gets cheaper faster than sensors get cheaper.
Boston Dynamics uses LiDAR plus cameras on Atlas. When a company that has spent 30 years building robots says LiDAR is necessary, the market listens.
Two things stand out. The camera-first camp (Tesla, Figure, 1X) has more adherents arriving through different engineering philosophies. The LiDAR camp (Unitree, Boston Dynamics) includes the companies with the most real-world deployment experience. The companies that have shipped robots into messy environments chose LiDAR. The companies building for scale chose cameras.
The Convergence
LiDAR costs are falling. Fast.
In 2025, a solid-state LiDAR unit suitable for indoor humanoid applications costs $500 to $2,000. By 2028, industry roadmaps target $200 to $500. Hesai, Luminar, and Ouster have all published product roadmaps with sub-$500 units for robotics applications by 2027 to 2028. The cost reduction is driven by silicon photonics integration (moving from discrete optics to chip-scale LiDAR) and volume scaling from the automotive market.
The camera stack is not standing still, but the absolute cost floor is already low. A camera module costs $15 to $50. There is not much room to compress. The total camera-based stack sits at $120 to $400. That number will not drop below $100 without a fundamental change in camera technology.
If LiDAR reaches $200 to $500 by 2028, the cost argument for camera-first perception evaporates. The 5 to 10x gap collapses to a 1.5 to 3x gap. At that differential, the capability advantages of LiDAR (works in darkness, native 3D, no compute-intensive depth estimation) become the dominant decision variable.
There is a second-order effect. When LiDAR costs fall below $500, the “LiDAR plus cameras plus depth sensors” stack that Unitree currently ships at $1,000 to $2,500 drops to $500 to $1,200. That is within striking distance of Figure’s camera-plus-depth stack ($200 to $500). A LiDAR-supplemented stack that costs 20% more than a camera-only stack but delivers 50% better depth accuracy becomes an easy engineering call.
The camera camp’s strongest argument is cost. That argument has a shelf life of 24 months.
If you want to understand where the supply chain value actually accrues, subscribe for free. Part 2 covers the fusion layer.
The Fusion Layer
Regardless of which sensor dominates the robot’s head, the data has to be combined into a single, coherent model of the world. That is the fusion layer. And the companies building it do not need to pick sides.
Ouster acquired StereoLabs in February 2026. On the surface, this looks like a LiDAR company buying a camera company. Underneath, it is a fusion play. Ouster now owns both the LiDAR hardware and the camera-based depth perception software. They can sell a combined stack to any OEM, regardless of camp. The acquisition signals that the smart money is not betting on one sensor. It is betting on the ability to combine all of them.
Bosch is building multi-sensor perception modules that integrate cameras, radar, and LiDAR into a single package. Bosch is an automotive Tier 1 supplier bringing automotive-grade sensor fusion to robotics. The modules are not cheap, but they are certified, tested, and available at volume. For OEMs that do not want to build their own fusion stack, Bosch offers a turnkey solution.
NVIDIA’s Isaac platform is positioning as the operating system for robot perception. Isaac does not care whether the input comes from cameras, LiDAR, or depth sensors. It processes all three. NVIDIA’s bet is that the value in robotics is not in the sensors (which are commoditizing) or in the actuators (which are commoditizing) but in the compute layer that makes sense of the sensor data.
The fusion layer is where the real supply chain value accrues. The sensor format war determines which hardware suppliers win. The fusion layer determines which platform suppliers win. And platform suppliers tend to capture more margin than hardware suppliers. In automotive, the Tier 1 suppliers (Bosch, Continental, Denso) capture more margin than the component manufacturers below them. In robotics, the same pattern will emerge.
Ouster’s software-attached bookings doubled in 2025. Software now ships with more than 15% of sensors. The fusion layer is already monetizing before the humanoid market even arrives.
The Bear Case
Two risks. Both real.
The numbers are small.
Humanoid-specific sensor revenue is approximately $59 million today. That is a rounding error in a $12 to $13 billion total perception TAM. The humanoid-specific slice requires shipments to grow from roughly 13,000 units in 2025 to 1.4 million units in a decade. That is a 100x increase. The robotics market has burned forecasters before. Bull cases that assume hockey-stick adoption curves in hardware markets have a poor track record.
AI might make the debate irrelevant.
If AI vision models advance to the point where a camera-only stack can extract reliable 3D depth from 2D images in real time, the entire LiDAR value proposition collapses. Not gradually. Suddenly. Tesla is already betting on this. Their FSD stack extracts depth from monocular cameras using neural networks trained on billions of labeled frames. If that approach transfers to manipulation tasks, cameras win by default. No format war. No convergence. Just cameras and compute. The counter-argument is that AI depth estimation is probabilistic, not deterministic. LiDAR measures distance with millimeter precision. AI estimates distance with confidence intervals. For navigation, estimation is good enough. For manipulation, where a robot hand needs to grasp a 10mm bolt, estimation might not be. The question is whether AI closes that precision gap before LiDAR closes the cost gap. We think the AI risk is the more dangerous one for the LiDAR supply chain. The convergence argument assumes LiDAR costs fall while camera capabilities stay flat. If AI vision models improve faster than LiDAR costs decline, the camera camp wins on both cost and capability. And the LiDAR supply chain becomes stranded.
Key Takeaways
The format war is real. Once enough OEMs commit to a sensor architecture, the supply chain locks in. Switching costs compound. The winner captures a decade of value.
The camera camp wins on cost today. Tesla, Figure, and 1X all arrive at camera-first through different engineering philosophies. The 5 to 10x cost gap is the camera camp’s strongest argument.
The cost argument has a 24-month shelf life. LiDAR is targeting $200 to $500 by 2028. At that price, capability advantages (darkness, native 3D, no compute-intensive depth estimation) become the dominant decision variable.
The fusion layer wins regardless. Ouster, Bosch, and NVIDIA are building platforms that process whatever sensors the OEM chooses. Platform suppliers capture more margin than hardware suppliers. This pattern held in automotive. It will hold in robotics.
The real risk is AI, not cost. If AI depth estimation reaches sub-5mm precision at manipulation range, the LiDAR supply chain becomes stranded. That is the bear case that matters.
Format war timelines
VHS vs Betamax, 5 years. Blu-ray vs HD DVD, 22 months. The perception stack format war started in 2025. The clock is running.
We think the camera-first camp wins the next 24 months on cost. We think the LiDAR-supplemented camp wins the decade on capability. And we think the fusion layer captures the margin either way.
Part 2 examines the fusion layer in detail: who owns it, who is building it, and where the margin actually accrues. Subscribe for free to follow the series.
Frequently Asked Questions
Why can’t robots just use cameras like self-driving cars?
Self-driving cars operate outdoors in well-lit conditions with structured roads and known map data. Humanoid robots operate indoors, in variable lighting, manipulating objects at close range. The perception requirements, depth accuracy under 5mm, object recognition at arm’s length, operation in low-light rooms, are different in kind. What works for a car at 60 mph on a highway does not translate to a robot reaching for a door handle in a dim hallway.
Is LiDAR going to become as cheap as cameras?
LiDAR costs have fallen from $75,000 in 2015 to $500–2,000 today, with a $200–500 target by 2028. At that price, LiDAR would still cost 4–13x more than a camera module ($15–50). The gap narrows but does not close. The real question is whether LiDAR becomes cheap enough that the capability advantage, deterministic 3D point clouds that work in darkness, outweighs the remaining price premium. We think the answer is yes, at roughly the $300 unit price threshold.
What happens if AI makes cameras “good enough” for 3D perception?
This is the wildcard. Vision-language-action models are learning to extract depth from 2D images with improving accuracy. If that capability matures to sub-5mm precision at manipulation range, camera-first perception wins by default. No format war, just cameras and compute. But we are not there yet. Current AI depth estimation is probabilistic, not deterministic. For navigation, estimation is good enough. For grasping a 10mm bolt from a parts bin, it is not.
Which companies benefit regardless of which sensor wins?
The fusion layer, companies building software that combines data from multiple sensor types, wins either way. Ouster (via its StereoLabs acquisition), Bosch (via its NEURA Robotics partnership), and NVIDIA (via the Isaac platform) all sell to both camps. The fusion layer does not need to pick a side. It processes whatever sensors the OEM chooses and extracts margin at the platform level.
Eyes of the Machine: A 3-Part Series
1. Cameras, LiDAR, Depth Sensors: The $38B Format War
2. Software Eats the Sensors
3. Sony Won the Last Camera War. Who Wins This One?








