Vision-based interfaces
What
Interfaces that take what a person is already doing — tilting their head, looking, holding something — as the input, instead of asking them to touch the screen.
Why
Some places leave the hands busy or the screen far away: a moving car, a patrol control room, an exhibition floor. There the body is the input. But reading the body brings misreadings and privacy with it, so the first decision is how much to read at all.
Process
- A head tilt – Glide – AirPods head-motion sensors and CoreMotion read the passenger's tilt, and a haptic arrives 2–7 seconds before the car moves. No face video is stored, and health-data access was dropped even at the cost of accuracy.
- A gaze – Physix AI – eye-tracking checked the patrol-robot console. Gaze was used as evidence about the design, not as an input device.
- A hand – Distance of Thinking – the handle a visitor holds is the control. A sensor was chosen over a camera, which is why it works in a dark gallery.
Result
- The more of the body you read, the more you must explain. When a system moves first, people ask why — which is why Glide always says what it noticed.
- Some limits come before accuracy. No stored face video, no health data.
- A camera is often the wrong answer. Using one where a sensor would do brings lighting, privacy and computation along with it.