Last week we trained a camera to recognise products in a simulated tote, got a perfect score, and said we did not trust it. The scene was too tidy and nothing was ever picked up.
So this week we picked things up.
What the robot does now
A simulated robot arm, the same model as a real Universal Robots UR5e, stands next to a tote. We drop three to six products into it, tins, cartons and cleaner jugs, and let physics pile them up however they land. Some lean on the wall. Some lie on top of each other.
An order asks for one of them. The robot looks down through a camera, works out where to put its suction cup, reaches in, lifts, carries the item over the rim, turns it so it fits, and lowers it into the order box. Then we check what actually landed in the box.
Watch one pick (12 seconds)
The numbers
We ran each version on the same 120 piles, so the comparison is fair.
- 78% when the simulator tells the robot exactly where each item is. This is the ceiling for our arm, cup and motion settings.
- 59% when the robot works it out from the camera alone, using a model that outlines each item.
- 28% with our first camera approach, which drew rough boxes around items instead of outlines. Boxes overlap the neighbours, so the robot kept aiming at the wrong thing.
Each successful pick takes about 10 seconds, which works out to roughly 360 an hour if nothing fails. Once the camera-guided robot has an item on the cup, it gets it to the box as reliably as the version that knows the answers. All of the camera's shortfall is in deciding where to grab.
Most of what is left of that gap turned out to be about tins. The camera sees the lid sitting a few millimetres below the rim, but our physics model treats the tin as a solid flat-topped cylinder. The robot aims for a surface the physics does not have. That is a flaw in how we set up the simulation, not in the camera, and it is next on the list.
Three things we got wrong
This is the part we would want to read if someone else had done the work.
We blamed leverage. Items kept falling off the cup the instant it started lifting. We assumed the cup was grabbing too far from the item's centre, so the weight pried it loose, and changed the robot to grab nearer the middle. Success fell from 61% to 26%. Wrong idea, and we would not have known without running it.
We blamed suction strength. Next we made the cup effectively unbreakable. Items still fell. So that was not it either.
The real cause was small. The cup was stopping four millimetres above the surface instead of touching it. Pressing it three millimetres in took success from 61% to 82% in a quick test. Real suction cups have to squash a little to seal, so this is the right behaviour, not a trick.
And one was plainly our bug. For a while the camera-guided robot was terrible at cartons. The model had scored 99% in testing, so something did not add up. The software passing camera images to the model had red and blue swapped. The pink jugs and red-and-white cartons were arriving looking like different objects. Every camera result before the fix went in the bin.
What this is and is not
Every number here comes from a simulation. Simulated suction and friction are approximations, so 59% is not a promise about a real robot. What a simulation is good at is comparing choices, a cup length, a grasp rule, a camera model, cheaply and on identical piles, and catching mistakes before they are made in metal.
It already changed our plans. The two-finger gripper in our hardware spec opens to 98 mm, and the tin is 114 mm wide. We found that on paper. The simulation is how we will find the ones we cannot see on paper.
Next: fix the tin, try a two-finger gripper against suction, and photograph real totes, because the only number that finally matters is the one from a real camera.