Visual SLAM and photogrammetry both use images to recover information about a space, but they are designed around different problems. Visual simultaneous localization and mapping estimates camera motion while building a map that supports localization. Photogrammetry uses overlapping images to reconstruct scene geometry and appearance, usually through a processing workflow after capture. For indoor mapping, the right choice depends on whether the priority is navigation, documentation, geometric reconstruction, or a combination.
What visual SLAM does
A visual SLAM system identifies and tracks image features as a camera moves. It estimates the camera trajectory, adds observations to a map, and attempts to recognize previously visited areas so accumulated drift can be corrected. Some systems combine cameras with inertial measurements or other sensors. The ORB-SLAM3 paper is a useful technical reference for understanding visual, visual-inertial, and multi-map SLAM concepts.
The map maintained by a SLAM system is often optimized for localization rather than for a polished survey deliverable. Depending on the implementation, it may contain sparse landmarks, keyframes, a trajectory, or denser spatial information. Teams should inspect the actual output instead of assuming every SLAM product creates the same kind of model.
What photogrammetry does
Photogrammetry reconstructs spatial information from images captured with sufficient overlap and viewpoint diversity. Processing typically estimates camera positions, identifies corresponding features, and produces outputs such as point clouds, meshes, textured models, or orthographic views. Autodesk's explanation of photogrammetry for construction provides context for its use in site documentation, mapping, and model creation.
Indoor photogrammetry can be demanding because rooms may contain blank walls, repeated patterns, reflective finishes, changing light, narrow circulation areas, and moving people. A successful workflow therefore depends on deliberate capture rather than simply recording more images.
Visual SLAM and photogrammetry compared
| Factor | Visual SLAM | Photogrammetry |
|---|---|---|
| Primary objective | Estimate motion and maintain a map for localization | Reconstruct scene geometry and appearance from images |
| Processing pattern | Often incremental during or shortly after movement | Often processed as a collected image set |
| Typical outputs | Trajectory, landmarks, keyframes, or system-specific maps | Point cloud, mesh, texture, orthographic output, or derived model |
| Scale | May require inertial data, known dimensions, depth, or another scale source | Requires control, known dimensions, camera information, or another scale strategy |
| Common indoor risks | Feature loss, drift, rapid motion, lighting changes, and poor loop closure | Weak overlap, blur, reflective surfaces, limited texture, and incomplete coverage |
Choose based on the operating task
Navigation and live tracking
Visual SLAM is the more natural foundation when a device or operator must understand movement through a building as capture happens. Robotics, augmented-reality positioning, and guided inspection can depend on an evolving pose estimate. The essential tests are whether tracking survives the actual indoor environment, whether relocalization works after interruptions, and whether the map can be exported in a form the wider project can use.
Detailed visual reconstruction
Photogrammetry is often the clearer choice when the goal is a reviewable model derived from a planned image set. It allows teams to prioritize coverage and reconstruction quality rather than continuous localization. It is useful for documenting rooms, facades, equipment areas, and visible conditions, provided the scene offers enough stable visual information and the team establishes scale and verification.
A combined workflow
The methods can complement each other. A SLAM estimate may help organize image trajectories or provide initial camera positions, while photogrammetric processing refines a reconstruction. Other sensors can contribute scale or depth. A hybrid pipeline should still identify which system controls the final coordinate frame, how transformations are recorded, and which observations are used for acceptance.
Indoor capture conditions to assess
- Texture: blank or repetitive surfaces may provide too few distinctive features for reliable matching.
- Lighting: dark areas, strong windows, flicker, and exposure changes can interrupt consistent observations.
- Motion: rapid turns and blur reduce the quality of feature tracking and image matching.
- Dynamic activity: people, doors, equipment, and temporary materials can introduce inconsistent scene content.
- Geometry: long similar corridors and repeated rooms can make place recognition more difficult.
- Scale and coordinates: define known dimensions, control, or sensor inputs before processing rather than adding scale informally later.
Build quality control into the workflow
Begin with a small representative area containing the hardest expected conditions. Record a repeatable route, capture revisits that support loop recognition, and avoid ending a sequence in an unconnected area. For photogrammetry, maintain intentional image overlap, vary viewpoints without abrupt movement, and capture transitions between rooms. Keep original imagery and metadata so results can be traced back to observations.
Review more than visual appearance. Check trajectory continuity, disconnected components, scale, coordinate orientation, missing surfaces, duplicate geometry, and distortions around reflective or low-texture areas. Compare selected dimensions or control observations that were not used to build the model. Document excluded spaces and any manual corrections.
Design the handoff, not just the map
An indoor map becomes useful when downstream teams understand its coordinate frame, date, coverage, object meaning, and limitations. A viewer, point cloud, mesh, trajectory, and BIM model represent different levels of interpretation. NIST's work on building digitization and semantic interoperability is relevant because spatial data must ultimately connect with building concepts and other project systems.
Define whether the deliverable is evidence, a localization map, a geometric reference, or an authored building model. Assign responsibility for converting raw reconstruction into named spaces, systems, and model elements. This prevents a visually convincing output from being used beyond its verified purpose.
Selection checklist
- Choose visual SLAM when continuous pose estimation or relocalization is central.
- Choose photogrammetry when a processed visual reconstruction is the primary deliverable.
- Use a hybrid approach only when the integration and coordinate strategy are explicit.
- Test the most difficult lighting, texture, and circulation conditions before scaling capture.
- Validate scale, coverage, and downstream compatibility against the intended decision.
Preimage supports image-led spatial workflows that connect capture, reconstruction, review, and field-to-model handoff. Teams evaluating indoor mapping methods can book a focused workflow discussion with Preimage to review their environment, intended outputs, and verification plan.











.webp)