Oliver Edis

Real-world intelligence: how ViroReact lets AR understand the space around your users

ViroReact real-world intelligence: visual positioning, depth and object detection for AR apps on iOS and Android

Most AR apps still treat the camera feed as wallpaper. A model floats in front of it, the phone tracks its own motion well enough to keep the model roughly still, and that is the whole trick. It demos well and it falls apart the moment the content has to mean something: sit on the actual floor, hide behind the actual sofa, be found again tomorrow in the same corner of the same room.

The difference between that and something useful is understanding. Not the app understanding the user, the app understanding the room. ViroReact, our open-source renderer for React Native, gives a phone three kinds of that understanding, and it gives them to iPhone and Android from one codebase. This post is about what they are, how they fit together, and what people are building on them.

Three questions an AR app could not ask before

Each capability answers a question. Put the answers together and an app knows where it is, what the space around it looks like, and what is in it.

Where am I? Visual positioning (VPS)

A phone's motion tracking is relative. It knows it has moved a metre to the left since the session started, and nothing more. Close the app and that frame of reference is gone, which is why AR content has traditionally lived for one session only.

Visual positioning fixes content to the real world rather than to a session. In ViroReact this comes through ReactVision Platform in two forms. Cloud anchors capture the visual features of a place so another device, or the same device a week later, can recognise it and resolve the same anchor in the same spot. Geospatial anchors pin content to latitude, longitude and altitude, using the phone's location and a visual match against the surroundings to get far closer than GPS alone.

In practice this is the capability that turns a one-off demo into a product. A survey marked on Tuesday is still there on Thursday. A route through a building is the same route for every visitor. A game level built in a park is the same level for everyone who turns up to play it. If you are choosing between the two, see cloud anchors vs geospatial anchors.

What shape is the space? Depth

Depth is the sense of distance: how far away every pixel in the camera image is. On an iPhone Pro that comes from the LiDAR scanner, and ARKit turns it into a mesh of the room that grows as you walk. On Android it comes from the ARCore Depth API. Where neither is available, ViroReact can run a neural depth estimator on the camera image instead.

ViroReact collects this into a world mesh: a live 3D surface of the floors, walls and furniture around the user. Once you have that mesh, a long list of things become straightforward. Virtual objects can be hidden behind real ones (occlusion). A ball can roll off a real table and land on the real floor (physics against the mesh). A tap on the screen can land on a real wall and return a point in metres, which is all you need to measure a room, size a repair or place a piece of furniture at the right scale. And the mesh can be saved, so a scanned room becomes a 3D model you can view, measure and share.

This is what that understanding looks like from the inside. Every triangle is a surface the phone has measured, and the renderer keeps adding to it for as long as the user keeps walking. Nothing in the picture was modelled by hand.

A room scanned with Viro Intelligence on an iPhone, drawn as ViroReact's world mesh. The floor, walls, sofa and shelves are all there as triangles, and the holes are simply the places the camera never looked.

What is in view? Object detection

Depth tells you there is a surface. It does not tell you the surface is a sofa, a boiler, a door or a pallet. Object detection does. ViroReact's object detector runs ONNX models on the device, frame by frame, and hands your React Native code a list of what it sees with a bounding box and a confidence for each.

Because it runs on the phone, it works offline and the camera frames never leave the device. Because the model is yours to choose, it can be a general model for a quick start or a model trained on the things your product cares about: damage types, fixtures, parts, plants. The detector gives an app the vocabulary to react to the world rather than just to a tap: label the objects in a room, count what is on a shelf, or confirm that the thing in frame is the thing the user was asked to find.

Why the combination matters

Any one of these is a feature. Together they are something closer to a sense of place, and that is what lets people act on the physical world through the screen.

Take a surveyor sizing water damage on a ceiling. Depth gives the size of the stain in real units from a photo. Object detection notes the light fitting and the extractor in the same frame. VPS pins the report to that ceiling, so the follow-up visit opens on the same patch with the earlier outline still drawn on it. No single capability does that job. The three of them do it in one app, and the same app runs on whichever phone the surveyor carries.

That is the pattern we see again and again: the interesting products sit where the three overlap.

Who is building on it

Property. Room scans that become floor plans, measurements and 3D models. Damage and defect reports sized from a photo rather than a tape measure. Inspections where the notes stay attached to the fabric of the building. Furnishing and renovation previews that respect the real walls and the real light.

Robotics. A phone with ViroReact is a cheap, well-understood sensor package. Teams use it to build and label maps a machine can work from, to teach a robot where things are by pointing a phone at them, and to overlay what the robot thinks it knows onto what the operator can see.

Games. Depth and detection let a game use the real room as the level: enemies come round real corners, loot hides behind real furniture, and the floor is the floor. VPS makes that level persistent and shared, so a group can play in the same space on the same terms.

These sectors have very different buyers, but they are asking the renderer the same three questions, which is why one renderer can serve all of them.

Trying it: Viro Intelligence

We keep an internal test-bed app called Viro Intelligence that exercises each of these capabilities on real phones, and it is a good way to see what the understanding actually looks like. It runs a Depth Check on first launch to find out what the phone can sense, scans rooms into meshes you can turn, measure and view as a floor plan, tape-measures objects in view, sizes an area from a single photo, and lists what the object detector saw while the camera was open. Every screen is ordinary React Native over ViroReact.

It's worth noting depth quality varies between phones, detection is only as good as the model you load, and visual positioning needs the place to look roughly as it did when it was captured. All gaps that we continue to work on closing.

Getting started

All of this is in ViroReact today. Add the renderer to a React Native app, turn on the world mesh and hit testing for depth, mount the object detector with a model, and connect ReactVision Platform for cloud and geospatial anchors. If you would rather not write the first version by hand, connect the ViroReact MCP server to your agent and let the agent do the heavy lifting (its validation and Render testing will help ensure what it produces works properly).

AR that acts on the world has to understand the world first. That understanding is what ViroReact is for.

Docs for depth, VPS and object detection in ViroReact

Depth and the world mesh

  • ViroReact API reference and guides: ViroARSceneNavigator with depthEnabled and worldMeshEnabled, worldMeshConfig, getWorldMeshStats(), snapshotWorldMeshToFile() and loadWorldMeshFromFile(), plus performARHitTestWithPoint() on ViroARScene for depth, plane and feature-point hits.
  • Installation guide: the platform permissions and build settings depth needs on iOS (ARKit, LiDAR) and Android (ARCore Depth API).

Visual positioning: cloud anchors, geospatial anchors and VPS

  • Cloud Anchors guide: host an anchor from one session and resolve it from another device or another day.
  • Geospatial Anchors guide: pin content to latitude, longitude and altitude with a visual match against the surroundings.
  • ReactVision Studio: a free account includes ReactVision Platform access, the rvApiKey and rvProjectId that turn these on.

Object detection

Starting a project

See it in action

Frequently asked questions

What is real-world intelligence in AR?
It is an AR app's understanding of the physical space around the user: where the phone is (visual positioning), what shape the space is (depth and a world mesh) and what is in view (object detection). ViroReact provides all three to React Native apps on iOS and Android from one codebase.
How does ViroReact get depth on phones without LiDAR?
On an iPhone Pro, depth comes from the LiDAR scanner through ARKit. On Android it comes from the ARCore Depth API. Where neither is available, ViroReact can run a neural depth estimator on the camera image instead.
What is the difference between cloud anchors and geospatial anchors?
Cloud anchors capture the visual features of a place so another device, or the same device later, can resolve the same anchor in the same spot. Geospatial anchors pin content to latitude, longitude and altitude, using the phone's location and a visual match against the surroundings. Both come through ReactVision Platform.
Does ViroReact object detection run on the device?
Yes. ViroReact's object detector runs ONNX models on the phone, frame by frame, so it works offline and camera frames never leave the device. You can load a general model or one trained on the objects your product cares about.
How do I start building with depth, VPS and object detection in ViroReact?
Add ViroReact to a React Native app, turn on the world mesh and hit testing for depth, mount the object detector with an ONNX model, and connect ReactVision Platform (a free ReactVision Studio account) for cloud and geospatial anchors. The ViroReact MCP server can help a coding agent write and validate the first version.