Real-world intelligence: how ViroReact lets AR understand the space around your users

Most AR apps still treat the camera feed as wallpaper. A model floats in front of it, the phone tracks its own motion well enough to keep the model roughly still, and that is the whole trick. It demos well and it falls apart the moment the content has to mean something: sit on the actual floor, hide behind the actual sofa, be found again tomorrow in the same corner of the same room.
The difference between that and something useful is understanding. Not the app understanding the user, the app understanding the room. ViroReact, our open-source renderer for React Native, gives a phone three kinds of that understanding, and it gives them to iPhone and Android from one codebase. This post is about what they are, how they fit together, and what people are building on them.
Three questions an AR app could not ask before
Each capability answers a question. Put the answers together and an app knows where it is, what the space around it looks like, and what is in it.
Where am I? Visual positioning (VPS)
A phone's motion tracking is relative. It knows it has moved a metre to the left since the session started, and nothing more. Close the app and that frame of reference is gone, which is why AR content has traditionally lived for one session only.
Visual positioning fixes content to the real world rather than to a session. In ViroReact this comes through ReactVision Platform in two forms. Cloud anchors capture the visual features of a place so another device, or the same device a week later, can recognise it and resolve the same anchor in the same spot. Geospatial anchors pin content to latitude, longitude and altitude, using the phone's location and a visual match against the surroundings to get far closer than GPS alone.
In practice this is the capability that turns a one-off demo into a product. A survey marked on Tuesday is still there on Thursday. A route through a building is the same route for every visitor. A game level built in a park is the same level for everyone who turns up to play it. If you are choosing between the two, see cloud anchors vs geospatial anchors.
What shape is the space? Depth
Depth is the sense of distance: how far away every pixel in the camera image is. On an iPhone Pro that comes from the LiDAR scanner, and ARKit turns it into a mesh of the room that grows as you walk. On Android it comes from the ARCore Depth API. Where neither is available, ViroReact can run a neural depth estimator on the camera image instead.
ViroReact collects this into a world mesh: a live 3D surface of the floors, walls and furniture around the user. Once you have that mesh, a long list of things become straightforward. Virtual objects can be hidden behind real ones (occlusion). A ball can roll off a real table and land on the real floor (physics against the mesh). A tap on the screen can land on a real wall and return a point in metres, which is all you need to measure a room, size a repair or place a piece of furniture at the right scale. And the mesh can be saved, so a scanned room becomes a 3D model you can view, measure and share.
This is what that understanding looks like from the inside. Every triangle is a surface the phone has measured, and the renderer keeps adding to it for as long as the user keeps walking. Nothing in the picture was modelled by hand.

What is in view? Object detection
Depth tells you there is a surface. It does not tell you the surface is a sofa, a boiler, a door or a pallet. Object detection does. ViroReact's object detector runs ONNX models on the device, frame by frame, and hands your React Native code a list of what it sees with a bounding box and a confidence for each.
Because it runs on the phone, it works offline and the camera frames never leave the device. Because the model is yours to choose, it can be a general model for a quick start or a model trained on the things your product cares about: damage types, fixtures, parts, plants. The detector gives an app the vocabulary to react to the world rather than just to a tap: label the objects in a room, count what is on a shelf, or confirm that the thing in frame is the thing the user was asked to find.
Why the combination matters
Any one of these is a feature. Together they are something closer to a sense of place, and that is what lets people act on the physical world through the screen.
Take a surveyor sizing water damage on a ceiling. Depth gives the size of the stain in real units from a photo. Object detection notes the light fitting and the extractor in the same frame. VPS pins the report to that ceiling, so the follow-up visit opens on the same patch with the earlier outline still drawn on it. No single capability does that job. The three of them do it in one app, and the same app runs on whichever phone the surveyor carries.
That is the pattern we see again and again: the interesting products sit where the three overlap.
Who is building on it
Property. Room scans that become floor plans, measurements and 3D models. Damage and defect reports sized from a photo rather than a tape measure. Inspections where the notes stay attached to the fabric of the building. Furnishing and renovation previews that respect the real walls and the real light.
Robotics. A phone with ViroReact is a cheap, well-understood sensor package. Teams use it to build and label maps a machine can work from, to teach a robot where things are by pointing a phone at them, and to overlay what the robot thinks it knows onto what the operator can see.
Games. Depth and detection let a game use the real room as the level: enemies come round real corners, loot hides behind real furniture, and the floor is the floor. VPS makes that level persistent and shared, so a group can play in the same space on the same terms.
These sectors have very different buyers, but they are asking the renderer the same three questions, which is why one renderer can serve all of them.
Trying it: Viro Intelligence
We keep an internal test-bed app called Viro Intelligence that exercises each of these capabilities on real phones, and it is a good way to see what the understanding actually looks like. It runs a Depth Check on first launch to find out what the phone can sense, scans rooms into meshes you can turn, measure and view as a floor plan, tape-measures objects in view, sizes an area from a single photo, and lists what the object detector saw while the camera was open. Every screen is ordinary React Native over ViroReact.
It's worth noting depth quality varies between phones, detection is only as good as the model you load, and visual positioning needs the place to look roughly as it did when it was captured. All gaps that we continue to work on closing.
Getting started
All of this is in ViroReact today. Add the renderer to a React Native app, turn on the world mesh and hit testing for depth, mount the object detector with a model, and connect ReactVision Platform for cloud and geospatial anchors. If you would rather not write the first version by hand, connect the ViroReact MCP server to your agent and let the agent do the heavy lifting (its validation and Render testing will help ensure what it produces works properly).
AR that acts on the world has to understand the world first. That understanding is what ViroReact is for.
Docs for depth, VPS and object detection in ViroReact
Depth and the world mesh
- ViroReact API reference and guides:
ViroARSceneNavigatorwithdepthEnabledandworldMeshEnabled,worldMeshConfig,getWorldMeshStats(),snapshotWorldMeshToFile()andloadWorldMeshFromFile(), plusperformARHitTestWithPoint()onViroARScenefor depth, plane and feature-point hits. - Installation guide: the platform permissions and build settings depth needs on iOS (ARKit, LiDAR) and Android (ARCore Depth API).
Visual positioning: cloud anchors, geospatial anchors and VPS
- Cloud Anchors guide: host an anchor from one session and resolve it from another device or another day.
- Geospatial Anchors guide: pin content to latitude, longitude and altitude with a visual match against the surroundings.
- ReactVision Studio: a free account includes ReactVision Platform access, the
rvApiKeyandrvProjectIdthat turn these on.
Object detection
@reactvision/react-viro-onnx(source and README on GitHub): the on-device ONNX Runtime provider behindViroObjectDetector, with install steps for Expo and React Native CLI, how to bundle a model, and the model's input and output contract.- ViroReact API reference: the
ViroObjectDetectorcomponent, itsmodel,confidenceThreshold,maxFPSandmaxDetectionsprops and theonDetectionevent.
Starting a project
- ViroReact MCP server: connect it to your coding agent for the most current guidance, with scene validation and rendering.
- Expo and TypeScript starter kit and the ViroReact repository.
See it in action
Viro Intelligence running on ViroReact: room scanning, depth and on-device object detection. Watch the demo on X.
Oliver Edis (@OliEdis) October 6, 2026