A 3D mesh model uses co-visible points to select normals for accurate surface reconstruction.
Per-pixel alpha masks filter specular highlights and locomotion errors during gradient-based depth map refinement.
Neural rendering synthesizes high-fidelity images from sparse views, eliminating dense camera requirements while resolving occlusion artifacts.