a ~5 MB neural net (COCO-SSD) running inside this tab on your GPU via TensorFlow.js. it knows exactly 80 words β person, cup, laptop, dogβ¦ β and shouts them dozens of times a second. no server, no api key, no video upload. this is what shipped-to-every-browser computer vision looks like.
reflexes are fast but shallow. press ask the expert and one snapshot goes to Claude, which actually understands the scene β the specific tool on your bench, the state it's in, the thing the 80 words can't say. reflexes are cheap and everywhere; judgment is one api call away. that's the whole thesis: one expert, a whole team.
this took one working session to build β no ML degree, no GPU cluster, no training run. the reflex tier is a free download; the judgment tier is an api. any engineer can wire senses like these into their own tools, inspections, and workflows. the hard part stopped being the AI.
the reflex net only knows the 80 words below. it will sometimes call your dog a cat, miss small things, and label a 3d-printed bracket "scissors". that gap is the demo β watch what happens when the expert looks instead.