β—‰ AGENT EYES
STANDBY
AI agents for engineers β†—
your browser can see. live object detection with no install, no upload, no account β€” plus an expert layer that knows what things actually are.
point it at your world.
a neural net is about to run inside this tab and name what it sees, live. pick a source β€” the video never leaves your device.
πŸ“± you're inside an app's browser β€” the live camera usually needs Safari or Chrome. tap β‹― β†’ open in browser for the full demo, or try a ✨ sample / πŸ–ΌοΈ image right here.
MODEL RUNS LOCALLY Β· COCO-SSD VIA TENSORFLOW.JS Β· ~5 MB, LOADS ONCE
AGENT-EYES-TAU.VERCEL.APP
conf β‰₯ 45%
C or ESC to exit cinema mode Β· E asks the expert
WHAT IS ACTUALLY HAPPENING HERE
01 Β· THE REFLEX LAYER

a ~5 MB neural net (COCO-SSD) running inside this tab on your GPU via TensorFlow.js. it knows exactly 80 words β€” person, cup, laptop, dog… β€” and shouts them dozens of times a second. no server, no api key, no video upload. this is what shipped-to-every-browser computer vision looks like.

02 Β· THE EXPERT LAYER

reflexes are fast but shallow. press ask the expert and one snapshot goes to Claude, which actually understands the scene β€” the specific tool on your bench, the state it's in, the thing the 80 words can't say. reflexes are cheap and everywhere; judgment is one api call away. that's the whole thesis: one expert, a whole team.

03 Β· WHY IT MATTERS

this took one working session to build β€” no ML degree, no GPU cluster, no training run. the reflex tier is a free download; the judgment tier is an api. any engineer can wire senses like these into their own tools, inspections, and workflows. the hard part stopped being the AI.

HONEST LIMITS

the reflex net only knows the 80 words below. it will sometimes call your dog a cat, miss small things, and label a 3d-printed bracket "scissors". that gap is the demo β€” watch what happens when the expert looks instead.