Methods and systems for executing a
Gaussian online splatting model for simultaneously localizing and mapping a surrounding 3D space are disclosed. The model is configured to receive an image-based data sample representing a first
field of view of the 3D space and, using a
Gaussian 3D map of the model, to render a new image-based data sample representing a new
field of view distinct from the first, as well as to render the corresponding speech features. By incorporating a hierarchical
encoder and a contrastive speech-image pre-training model (CLIP model) into the architecture of the
Gaussian online splatting model, the overall architecture is configured to operate in near real-time.