A real-
time system for the multimodal generation of
virtual reality scenes based on
artificial intelligence for the creation of immersive three-dimensional environments from
natural language narratives, consisting of: a speech capture module configured to continuously
record a user's spoken narrative via one or more directional microphones, preprocesses the captured
signal by
noise reduction and temporal alignment, and outputs a digital speech
stream; A speech-to-
text processing unit that is operationally coupled to the speech capture module and configured for real-time
speech recognition using a continuous neural
transformer model. The unit is trained to transcribe
natural language utterances into
structured text data while maintaining contextual continuity throughout the evolving narrative. a
semantic interpretation processing unit that is communicatively linked to the
speech recognition unit and configured to perform
natural language understanding techniques to extract contextual entities, spatial references, temporal relationships, and object attributes from the transcribed narrative; the engine includes a large
language model that is fine-tuned for spatial reasoning tasks; a
scene graph generation module configured to transform the interpreted semantic data into a structured, hierarchical representation that defines nodes for identified entities and edges for corresponding relationships, with each node associated with
metadata describing geometry, position, orientation, texture, and linking attributes between objects; a
multimodal image-
language model processor coupled with the
scene graph generation module, wherein the processor is configured to retrieve, adapt, or synthesize appropriate three-dimensional elements from a pre-trained visual-lexical embedding space and align these elements with their semantic and spatial definitions derived from the
scene graph; a scene
assembly and rendering controller configured to create a cohesive virtual scene from the aligned assets, perform real-time rendering using a GPU-accelerated
ray tracing pipeline, and produce a stereoscopic visual output that corresponds to the evolving narrative; A head-mounted
virtual reality visualization device connected to the rendering engine and configured to display the generated immersive environment to the user in real time. The device features
motion sensors and inside-out tracking cameras to detect head and body movements, dynamically updating viewing angles and perspective within the rendered scene; and a bidirectional feedback module integrated into the head-mounted device and connected to the
semantic interpretation processing unit; the module is configured to interpret corrective commands, gestures, or supplementary comments from the user to refine or modify specific scene elements without interrupting the real-time
visualization; The
system continuously updates the virtual scene as the narrative develops, ensuring temporal synchronization between
speech input and rendered output below a defined latency threshold, thus enabling a natural, dialogic construction of complex three-dimensional virtual environments.