Immersive psychotherapy VR system based on dialogue real-time generation

By generating dynamic VR scenes in real time through a modular system architecture, the problem of insufficient immersion and difficulty in empathy in traditional narrative therapy is solved, thereby improving the immersive memory experience and therapeutic effect of psychotherapy.

CN120878082APending Publication Date: 2025-10-31GUANGZHOU COLLEGE OF COMMERCE

Patent Information

Application Number
CN202510965839.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional narrative therapy lacks immersion and emotional empathy, and existing VR technology cannot generate personalized dynamic scenes in real time, making it difficult to meet the needs of immediate interaction in psychotherapy.

Method used

Adopting a modular system architecture, and combining speech recognition, natural language processing, scene generation, and graphics rendering technologies, the system converts patients' language content into dynamic VR scenes in real time, enhancing immersion and empathy.

Benefits of technology

It improved patients' immersion and concentration, enhanced therapists' empathy, and improved the effectiveness and efficiency of psychotherapy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120878082A_ABST
    Figure CN120878082A_ABST
Patent Text Reader

Abstract

The invention discloses an immersive psychotherapy VR system based on dialogue real-time generation, and the system comprises a voice recognition module which is used for collecting dialogue voice in a consultation process, and transcribing the dialogue voice into a text in real time; the natural language processing module is used for performing semantic understanding on the transcribed dialogue text and extracting scene elements; the scene generation module is used for dynamically generating a corresponding three-dimensional virtual scene according to the scene elements and adjusting scene details in real time along with updating of the dialogue content; the graphic rendering module is used for rendering the three-dimensional virtual scene in real time and then outputting the three-dimensional virtual scene to the VR equipment; and the user interaction module is used for realizing interaction operation among the patient, the therapist and the system. According to the invention, a modular system architecture is adopted, technologies such as speech recognition, large language model understanding, three-dimensional scene generation and graphic rendering are organically combined, and real-time conversion from dialogue content to a virtual scene is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of virtual reality, and in particular relates to an immersive psychotherapy VR system based on real-time dialogue generation. Background Technology

[0002] Narrative therapy is a commonly used method in psychotherapy, where therapists guide patients to recount their personal experiences to help them process emotions and trauma. However, traditional narrative therapy relies primarily on verbal communication and lacks intuitive situational support, resulting in the following shortcomings:

[0003] Lack of immersion: Patients find it difficult to fully immerse themselves in their memories in a purely verbal communication environment. When asked to recall traumatic events or past experiences, patients often struggle to concentrate due to the monotonous surroundings, making it difficult to recall details.

[0004] Difficulty in emotional empathy: Psychotherapists can only understand patients' experiences and emotions based on their verbal descriptions, which is an indirect and unintuitive approach. Therapists struggle to promptly and accurately perceive the patient's situation and emotions, limiting the effectiveness of empathy.

[0005] Current technology lacks support for dynamic scenarios: While there have been attempts to use VR in psychotherapy (such as exposure therapy), these typically employ pre-made, fixed scenes, failing to generate corresponding scenarios in real time based on each patient's unique narrative. This lack of personalization and real-time capability makes it difficult to fully meet the immediate interactive needs of narrative therapy.

[0006] Given the above shortcomings, there is an urgent need for a new technology that enables patients to have a more immersive memory experience in narrative therapy and allows therapists to visually observe the context of the patient's narrative, thereby enhancing the visual interaction effect. Summary of the Invention

[0007] The purpose of this invention is to provide a VR system for psychotherapy that generates immersive scenes in real time based on dialogue, in order to address the problem of insufficient immersion in existing narrative therapy. By converting the patient's language content into dynamic VR scenes in real time during psychological counseling dialogues, the system enhances the patient's immersion, helping them to recall and express their personal experiences more attentively and vividly. Simultaneously, it enables psychologists to "immerse themselves" in the patient's memory context and emotional changes, strengthening empathy between doctor and patient and improving the visual interaction effect.

[0008] To achieve the above objectives, the present invention provides an immersive psychotherapy VR system based on real-time dialogue generation, comprising:

[0009] The speech recognition module is used to collect the dialogue voice during the consultation process and transcribe it into text in real time;

[0010] The natural language processing module, connected to the speech recognition module, is used to perform semantic understanding on the transcribed dialogue text and extract contextual elements;

[0011] The scene generation module, connected to the natural language processing module, is used to dynamically generate corresponding three-dimensional virtual scenes based on the contextual elements, and adjust scene details in real time as the dialogue content is updated.

[0012] The graphics rendering module, connected to the scene generation module, is used to render the three-dimensional virtual scene in real time and output it to the VR device.

[0013] The user interaction module, connected to the graphics rendering module, is used to enable interactive operations between patients, therapists, and the system.

[0014] Preferably, the speech recognition module includes a microphone array and a speech transcription subsystem;

[0015] The microphone array is used to pick up conversational audio and perform noise reduction and speech separation processing;

[0016] The speech-to-text subsystem outputs text recognition results in real time in a streaming manner using a deep neural network acoustic model and a large language model.

[0017] Preferably, the deep neural network acoustic model includes a Conformer encoder employing end-to-end automatic speech recognition technology, a Conformer Block fusing convolutional neural networks and Transformer self-attention mechanisms, for simultaneously capturing local detail features and global dependencies of the speech signal;

[0018] The large language model includes an explicit language model and a decoder;

[0019] The decoder performs the search based on the beam search principle.

[0020] Preferably, the natural language processing module stores the aforementioned elements in the dialogue state through the context preservation function, and associates new information with the existing context when further describing details to obtain contextual elements;

[0021] The contextual elements include the time, place, characters, emotional states, and environmental description of the event.

[0022] Preferably, the natural language processing module includes a keyword and event extraction unit, a sentiment analysis unit, an entity recognition and relation extraction unit, and a contextual modeling unit;

[0023] The keyword and event extraction unit is used to extract keywords that express core semantics from each dialogue sentence, as well as sentence components that describe important events;

[0024] The sentiment analysis unit is used to analyze the sentiment tendency of dialogue statements and determine the speaker's current emotion.

[0025] The entity recognition and relationship extraction unit is used to identify entities from the dialogue and further extract the relationships between entities;

[0026] The entities include names of people, locations, and organizations;

[0027] The contextual modeling unit is used to understand multi-turn dialogues as a whole context through long sequence processing and to splice the dialogue history into a long text; it is also used to detect dialogue turns and contextual changes.

[0028] Preferably, the scene generation module includes an environment generation unit and a character modeling unit;

[0029] The environment generation unit is used to render the corresponding scene background according to time, location and environmental elements;

[0030] The character modeling unit is used to create virtual dolls or characters based on character elements, and to adjust the atmosphere of the scene and the character's expression in combination with emotional elements.

[0031] Preferably, the scene generation module constructs a dynamic three-dimensional virtual scene based on a generative 3D model;

[0032] The generative 3D model integrates NeRF, Gaussian Splatting, and DreamFusion generative modeling technologies, and uses extracted text elements to drive scene generation and evolution.

[0033] Preferably, the graphics rendering module includes a real-time ray tracing rendering unit, a foveated rendering unit, a neural rendering and AI optimization unit, a multimodal feedback optimization unit, and a performance optimization unit;

[0034] The ray tracing rendering unit is used to simulate the propagation, reflection and refraction of light in the scene using real-time ray tracing, to simulate indirect diffuse reflection caused by multiple light bounces using a global illumination algorithm, and to accelerate global illumination calculation using a neural radiation cache.

[0035] The gaze point rendering unit is used to render the user's gaze area and the surrounding area with different precisions during rendering using an eye-tracking device;

[0036] The neural rendering and AI optimization unit is used to perform neural rendering by combining neural shading with materials and neural rendering of complex characters.

[0037] The multimodal feedback optimization unit is used to continuously optimize rendering quality through an AI feedback loop;

[0038] The performance optimization unit is used for global performance monitoring and adaptive optimization through frame rate monitoring and dynamic adjustment, multi-threading and asynchronous pipelines, and resource load balancing.

[0039] Preferably, the user interaction module includes a patient interaction unit and a therapist interaction unit;

[0040] The patient interaction unit is used to experience the generated scene and interact naturally through VR devices;

[0041] The therapist interaction unit is used to observe the scene from the patient's perspective through the control interface and control the scene presentation or the rhythm of the dialogue.

[0042] Compared with the prior art, the present invention has the following advantages and technical effects:

[0043] This invention adopts a modular system architecture, organically combining technologies such as speech recognition, large language model understanding, 3D scene generation and graphics rendering to achieve real-time conversion of dialogue content into virtual scenes.

[0044] This invention significantly enhances the patient's immersive experience during treatment by visualizing the patient's narrative in real-time into an immersive VR scene. The patient feels as if they are back in the scene, making it easier to concentrate and recall the details of past events.

[0045] The immersive narrative environment created by the system of this invention helps to stimulate and release the patient's emotions. The realistically recreated scenes allow the patient to express their emotions more naturally, helping them to more fully express and process their inner trauma, and improving the clarity and completeness of recalled memories.

[0046] The immersive narrative therapy of this invention can enhance patients' engagement in treatment, shorten the time required to build trust and a sense of security, and thus is expected to improve the overall effectiveness of psychotherapy. Patients process their emotions in a highly realistic setting, which helps to consolidate and deepen the therapeutic effect.

[0047] This invention allows doctors to synchronously observe the patient's virtual environment using VR, essentially "witnessing" the patient's experience. This intuitive approach enables doctors to more accurately understand the context and emotions of the events described by the patient, significantly enhancing their empathy. Simultaneously, doctors can more efficiently capture and analyze key details in the patient's narrative, improving the accuracy of consultation decisions and subsequent interventions. Attached Figure Description

[0048] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0049] Figure 1 This is a schematic diagram of the system structure according to an embodiment of the present invention. Detailed Implementation

[0050] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0051] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0052] like Figure 1 As shown, this embodiment provides an immersive psychotherapy VR system based on real-time dialogue generation, including: a speech recognition module, a natural language processing module, a scene generation module, a graphics rendering module, and a user interaction module.

[0053] Patients and therapists wear corresponding audio acquisition devices and VR display devices. At the start of the consultation, the patient's voice is first processed by the speech recognition module. The generated text stream is then processed by the natural language processing module to extract semantic elements, and subsequently passed to the scene generation module to construct a virtual scene. The graphics rendering module then renders the scene and outputs it to the patient's VR headset. The therapist, through the user interaction module's interface, can observe the scene from the patient's perspective and influence the scene presentation or the pace of the conversation through the control interface. All modules are connected via high-speed data interfaces, enabling the system to operate in a real-time closed-loop manner: on the one hand, instantly transforming the patient's narrative into a visual experience; on the other hand, allowing the therapist to intervene and guide the process. The entire system architecture adopts a modular + parallel stream processing design, ensuring real-time performance and stability.

[0054] Furthermore, the speech recognition module is used to collect the dialogue during the consultation process and uses speech recognition technology to transcribe the patient's and therapist's speech into text in real time. This module supports streaming recognition of continuous speech, ensuring that the dialogue content can be synchronously provided to subsequent processing units.

[0055] In this embodiment, the speech recognition module includes a high-sensitivity microphone array and a speech-to-text subsystem. The microphone array picks up the dialogue between the patient and therapist in the consultation room, and after noise reduction and speech separation processing, it is sent to the speech recognition engine. The speech-to-text subsystem uses a deep neural network acoustic model and a large language model, enabling it to output text recognition results in real-time streaming. For example, when the patient says, "I had an accident at school when I was a child," the speech recognition module quickly outputs the corresponding text, "I had an accident at school when I was a child," and labels the speaker (patient). This module provides an accurate and rapid text foundation for subsequent semantic analysis. To reduce latency, the system uses a segmented output strategy for long sentences, that is, the words are output as the patient speaks, so that the patient's narration is captured by the system almost synchronously. Speech recognition errors are automatically corrected by the subsequent language model correction module, improving the overall accuracy.

[0056] Furthermore, the deep neural network acoustic model in this embodiment adopts end-to-end automatic speech recognition technology, specifically including an acoustic model with the Conformer (Convolution-augmented Transformer) architecture as its core, and combining a language model and a decoder to achieve high-precision speech transcription.

[0057] Specifically, the acoustic model with the Conformer (Convolution-augmented Transformer) architecture at its core is the Conformer encoder. The Conformer encoder in this embodiment combines the advantages of convolutional neural networks (CNN) and the Transformer self-attention mechanism, and can simultaneously capture the local detailed features and global dependencies of the speech signal.

[0058] First, the input speech is convolutionally downsampled to reduce the temporal length and preliminary features are extracted. Then, it is encoded by stacking multiple Conformer Blocks.

[0059] Each Conformer Block consists of four sequential parts: a feedforward module, a multi-head self-attention module, a convolutional module, and a feedforward module. A Macaron structure is used, with the two feedforward modules located at the beginning and end of the block, respectively. This structure integrates convolutional units into the self-attention module, enhancing the ability to model local neighborhood patterns. Residual connections and layer normalization ensure the stability of deep network training. Relative position encoding is introduced into the self-attention sublayer to adapt to variable-length speech sequences, while the convolutional sublayer includes gated linear units (GLUs) and depthwise separable convolutions to effectively extract temporal local features.

[0060] Thanks to the aforementioned architectural design, the Conformer acoustic model in this embodiment significantly outperforms previous pure Transformer or CNN models on public speech recognition benchmarks, achieving the lowest word error rate and other metrics currently available in the industry. For example, training the Conformer model on speech data in this embodiment revealed that its error rate in recognizing real-world noisy environments is reduced by up to 43% compared to previous mainstream ASR models. This demonstrates that the Conformer architecture, as the acoustic model of this system, can provide excellent recognition accuracy and robustness.

[0061] Specifically, this system integrates an explicit language model (LM) and an efficient decoder to further improve recognition accuracy and stability. The language model employs a Transformer language model trained on a large-scale text corpus (e.g., at the character or word level). During the decoding stage, it combines this model with acoustic model probabilities through shallow fusion to supplement the syntactic and contextual information of speech recognition. Specifically, the decoder performs a search based on the beam search principle, adding the acoustic probability output by the Transformer encoder to the external language model score for evaluation at each hypothesis expansion, thereby selecting the text path with the highest probability.

[0062] To improve decoding efficiency, the system can adopt end-to-end criteria such as connectionist temporal classification (CTC) or RNN-Transducer (RNNT). For example, when combined with CTC, a weighted finite state machine (WFST) decoding graph containing acoustic and language models can be pre-constructed to achieve efficient search of candidate strings. When combined with RNNT, a prediction network (equivalent to an internal language model) and a joint network are added on the basis of the Conformer encoder to achieve streaming output.

[0063] During decoding, speech activity detection can be applied to skip silent segments, and dictionary constraints can be used to ensure that the output is a valid word sequence. Through these multi-layered processing steps, the transcribed text output by the speech recognition module achieves industry-leading accuracy and fluency, providing reliable input for subsequent natural language processing modules.

[0064] This embodiment's speech recognition module utilizes the Conformer end-to-end architecture for recognizing psychotherapy dialogues, fully leveraging its robustness and high accuracy in noisy environments. Compared to traditional HMM / DNN cascaded models, this embodiment uses a deep learning model to integrate acoustic and linguistic information, and achieves accurate transcription of long dialogues by fusing external language models and improved decoding algorithms. This architecture simplifies system complexity while maintaining accuracy.

[0065] Furthermore, the Natural Language Processing module is used to perform semantic understanding on the real-time transcribed dialogue text through a large language model, extracting key information elements, including the time, location, characters, emotional states, and environmental descriptions of events. This module analyzes and summarizes the dialogue content to form structured "contextual element" data, providing a basis for scene generation.

[0066] In this embodiment, the natural language processing module immediately performs semantic understanding and element extraction after acquiring the text. The module's built-in large language model has the capabilities for character relationship extraction, sentiment analysis, and scene description understanding. Taking the patient's description "I experienced an accident at school when I was a child" as an example, the natural language processing module will extract: time element = "childhood" (corresponding to past childhood), location element = "school," event element = "an accident," character element = "I (the patient)," emotion element = inferred as "fear / horror" (inferred from the context of "accident"), and environment element = "daytime (assuming school time based on common sense)" and "classroom / campus." The extracted elements are represented in a structured form, for example:

[0067] {

[0068] "Time": "Childhood"

[0069] Location: Campus

[0070] "Person":["Patient (Childhood)"],

[0071] "Event": "Accident"

[0072] "Emotion": "Fear"

[0073] }

[0074] During the ongoing dialogue, the natural language processing module also features context preservation, storing the aforementioned elements in the dialogue state. This allows the system to connect new information with the existing context when the patient provides further details (e.g., describing the specific details of an accident), preventing inconsistencies in the generated scene. This semantic extraction and context management enables the system to progressively enrich the scene content: for example, if the patient continues by saying, "It was raining heavily, and everyone in the classroom was panicked," the system will update the environmental elements (weather = "heavy rain") and the character elements (adding "classmates," "teachers," etc.), and the emotional element may be strengthened to "panic." The high accuracy and ability to extract implicit information from the natural language processing module ensure that the scene generation module obtains the most complete possible contextual information.

[0075] The natural language processing module in this embodiment is based on a large language model and is responsible for extracting contextual elements from the transcribed dialogue text. This module integrates various natural language processing functions such as keyword extraction, event extraction, sentiment analysis, entity recognition, relation extraction, and context modeling, all supported by a unified large model framework, thereby enabling deep semantic understanding and information extraction from continuous dialogues.

[0076] More specifically, the natural language processing module includes a keyword and event extraction unit, a sentiment analysis unit, an entity recognition and relation extraction unit, and a contextual modeling unit;

[0077] The keyword and event extraction unit first extracts keywords expressing core semantics and sentence components describing important events from each dialogue sentence. The model unit identifies the most informative words in the current context (e.g., nouns and verbs related to traumatic experiences) as keywords by distributing attention weights across the input sentences. Simultaneously, the model unit leverages its powerful semantic understanding capabilities to analyze syntactic structure and semantic roles, capturing the narrative event flow from the dialogue. For example, for a patient's description of a traffic accident, the model will identify the event of the accident and related actions and objects (e.g., "car crash," "injury"), among other key information fragments. Compared to traditional rule-based or shallow machine learning methods, large-model-based event extraction can perform deep semantic analysis by comprehensively considering the context, accurately capturing event trigger words and their arguments implicit in the sentence. The extracted keywords and event elements will be used to determine the thematic elements appearing in the scene and their dynamic evolution.

[0078] Next, the entity recognition and relation extraction unit identifies various entities mentioned in the dialogue (such as names, locations, and organizations) and further extracts the relationships between these entities. The model unit, pre-trained with massive amounts of knowledge, possesses strong Named Entity Recognition (NER) capabilities, accurately identifying names, locations, and time expressions in the dialogue. For example, it detects person entities (the patient, relatives, etc.) and location entities (the location of the accident, the hospital name, etc.) in the patient's narrative. Based on entity recognition, the model also uses contextual understanding to infer relationships between entities, such as determining that someone is another person's doctor, or that a location is the location of an event. Through relation extraction, the system can construct a knowledge graph about the current dialogue content (e.g., "Patient A - Encountered -> Accident X" or "Event X - Occurred -> Location Y"), clarifying the connections between various elements. This joint entity and relation extraction method, centered on a large model, continuously improves the scene element network while the dialogue is ongoing, providing the subsequent scene generation module with structured knowledge input (characters, scene locations, and their relationships).

[0079] In psychotherapy settings, capturing the client's emotional state is crucial. The emotional analysis unit analyzes the emotional tendency of dialogue statements to determine the speaker's current emotions (such as sadness, anger, anxiety, fear, etc.). This is achieved by identifying emotional vocabulary and tone in the dialogue (e.g., words like "pain" and "fear"), and by inferring implicit emotions from the context. To improve accuracy, this embodiment fine-tunes the emotional analysis unit for emotional classification tasks, or adds an emotional discriminator layer to its output, assigning emotional labels and intensity scores to each segment. For example, the model unit can identify the level of sadness or anxiety hidden when a patient recounts a memory. The emotional analysis results serve as an important basis for scene rendering: the system can adjust the scene atmosphere (such as lighting, color tone, and weather effects) based on the user's emotional state to synchronously reflect or appropriately guide the client's emotional changes. This deep emotional understanding, supported by a large model, is more intelligent and sensitive than traditional approaches based on emotional dictionaries, capable of capturing subtle emotional cues.

[0080] Since psychological counseling dialogues are often continuous and closely related to the context, this embodiment places particular emphasis on the overall modeling of the dialogue context. The long-sequence processing capability of the context modeling unit enables it to understand multi-turn dialogues as a whole context.

[0081] Specifically, the dialogue history (e.g., recent sentences, or even the entire dialogue) is concatenated into a long text and input into the model. The model's internal attention mechanism tracks information such as character references, event development, and emotional evolution. The model forms an implicit contextual representation of the current dialogue state, encoding the time, place, people, events, and emotions involved. For example, if the model remembers that an event mentioned by a patient occurred "last winter," it can correctly relate it when mentioned again without confusing it with other time information. Furthermore, contextual modeling includes detecting changes in dialogue turns and context, such as determining whether the current topic has shifted from event details to emotional expression. This continuous contextual tracking is implicitly performed by the model units; in this embodiment, the representation vectors of the model's intermediate layers can also be extracted as the dialogue state. When needed, the system can request the model to output a summary of the current contextual elements, for example, by prompting the model to list the elements such as "time, place, people, events, and emotions" up to the present, thereby explicitly obtaining structured contextual elements.

[0082] As an optional implementation, the large language model in this embodiment can be connected to DeepSeek, which divides the massive neural network into multiple "expert" sub-models. Each expert is good at handling different types of input features. During inference, only a small number of experts most relevant to the current input are activated, thereby significantly reducing computational overhead and improving the model's computational efficiency and generalization ability, enabling it to run with limited computing power while maintaining a huge capacity.

[0083] Furthermore, this embodiment of the large language model introduces a Multi-Head Latent Attention (MLA) mechanism, which is a modification of Transformer self-attention. MLA reduces storage and computational overhead as sequence length increases by compressing the "key" and "value" vectors in the attention mechanism into low-rank latent vector representations. This mechanism enables the model to efficiently handle extremely long dialogue contexts, minimizing memory usage while maintaining attention effectiveness, thus supporting global modeling of the entire dialogue. Additionally, the model employs a Multi-Token Prediction (MTP) strategy during training, allowing the model to predict multiple future tokens simultaneously to enhance its grasp of long-term context. This strategy improves the coherence of the model's generated long texts, which is highly beneficial for tasks requiring the summarization of contextual information in dialogues.

[0084] In addition to the above structural innovations, this embodiment also employs optimizations such as sparse attention, which further improves inference speed and scalability.

[0085] Furthermore, the scene generation module dynamically generates corresponding 3D virtual scenes based on the contextual elements extracted by the natural language processing module. This module comprises two parts: environment generation and character modeling. It renders the corresponding scene background (e.g., indoor rooms, school playgrounds, accident scenes, etc.) based on "time," "location," and "environment" elements; creates virtual dolls or characters based on "character" elements; and adjusts the atmosphere of the scene and the expressions of the characters based on "emotion" elements. Scene generation is continuous and real-time, constantly updating or changing relevant details as new information from the dialogue emerges, allowing the virtual environment to evolve synchronously with the narrative.

[0086] In this embodiment, the scene generation module receives contextual element data from the natural language processing module and dynamically constructs a virtual scene using a 3D engine. The module consists of an environment rendering subunit and a character generation subunit. The environment rendering subunit maps "time" and "location" elements to a predefined environment model library or generation algorithm. For example, when location = "school" and time = "childhood," a 3D model scene template of a primary school classroom is selected, and rain effects and corresponding lighting changes are added based on weather or environmental descriptions (such as "it's raining heavily"). If the environmental elements are very specific (such as "family living room" or "accident scene on the highway"), the corresponding scene template is called or a parametric scene generation algorithm is used to construct the corresponding environment. The character generation subunit creates virtual human characters based on "person" elements: such as the patient's childhood self-image, a teacher or classmate who might be present, etc. The system selects a 3D character model matching the person's identity from the character model library, or dynamically generates a simple humanoid (a general virtual doll representative can be used). Simultaneously, based on the "emotion" element, the character is given corresponding facial expressions and body language. For example, if the emotion is "fear / panic", then the character model's facial expression will show terror, and the posture may show retreating or defensive actions.

[0087] The scene generation module updates existing scenes in real time upon receiving new narrative information. It employs an incremental generation strategy, avoiding a complete scene reconstruction each time and instead adjusting only the changed parts. For example, in the aforementioned school accident scene, when the patient shifts from describing the rain to describing the accident (e.g., "the window suddenly shattered"), the system dynamically adds "window shattering" effects and audio-visual feedback to the classroom scene, enhancing the presentation of the event. Through this series of processes, the VR scene generated by the module closely matches the patient's narrative and evolves as the story progresses. The generated scene data structure is maintained in memory, including environment objects, character objects, and their states, for rendering purposes.

[0088] More specifically, the scene generation module in this embodiment is based on a generative 3D model to construct a dynamic 3D scene. The generative 3D model integrates generative model technologies such as NeRF (Neural Radiance Field), Gaussian Splatting, and DreamFusion, and combines extracted text elements to drive scene generation and evolution.

[0089] The NeRF neural radiation field technique described in this embodiment is an implicit 3D representation method that takes coordinates as input. It uses a multilayer perceptron (MLP) to model the color and density distribution of each point in 3D space. The NeRF model can reconstruct realistic 3D scenes from a set of 2D images with a limited number of viewpoints and achieve continuous rendering from any viewpoint, making it possible to synthesize realistic 3D from sparse data. However, the training and rendering overhead of classic NeRF is relatively large, making it difficult to apply in real time to interactive systems.

[0090] To address this, this embodiment introduces the latest acceleration technology, Gaussian Splatting, to improve the efficiency of NeRF. The Gaussian Splatting method uses a series of 3D Gaussian points distributed in space to represent the scene, approximating the continuous volumetric radiation field with a Gaussian distribution. During the training phase, a batch of sparse points is first obtained by back-projecting the multi-view images from the camera, and then expanded into an initial Gaussian distribution set. Next, an alternating optimization strategy is employed to continuously adjust the parameters of each Gaussian point (including the anisotropic covariance matrix) to more accurately fit the volumetric density distribution of the real scene. During the rendering phase, a special visibility-optimized Gaussian rasterization algorithm is used, rendering only Gaussians near the viewing direction and performing transparency synthesis. This maintains the photographic visual quality of the NeRF model while achieving true real-time rendering, reaching a new perspective synthesis speed of over 100 frames per second at 1080p resolution. Therefore, the system in this embodiment can render realistic scenes that update with changes in the user's viewing angle at a refresh rate of 90Hz or even higher in VR, significantly enhancing the immersive experience.

[0091] In addition to the implicit representation methods mentioned above, this embodiment also employs DreamFusion, a generative deep learning model that directly synthesizes 3D content from text. It utilizes a pre-trained text-to-image diffusion model (such as Imagen or StableDiffusion) to guide the generation of 3D scenes. Specifically, DreamFusion uses a randomly initialized NeRF model as the 3D representation, then continuously renders 2D images from randomly selected viewpoints within the NeRF. These rendered images are then fed into a pre-trained diffusion model to determine their matching degree with the input text description. DreamFusion proposes a probability density distillation loss, treating the 2D diffusion model as a scorer, optimizing the NeRF parameters to continuously reduce the loss of each view rendering under the diffusion model's judgment. Through this self-approximation optimization process (similar to DeepDream), the final NeRF model can generate a 3D scene that matches the input text. This method cleverly bypasses the need for large amounts of labeled 3D data, fully utilizing the prior knowledge of the 2D generative model to achieve the synthesis of high-fidelity 3D content from plain text.

[0092] In this system, once the NLP module provides a complete scene text description, the scene generation module can invoke a pipeline similar to DreamFusion to generate a corresponding 3D scene representation. For example, based on the scene elements described as "a tranquil beach at sunset, with the sky gradually darkening," the system will generate a panoramic image that matches the description using a diffusion model, and then optimize the corresponding NeRF radiation field model. To improve efficiency, this embodiment can use only low-resolution cascading of the diffusion model for coarse model generation, and then obtain a high-resolution result through NeRF's own refinement training. After generation, the implicit NeRF representation can also be exported as an explicit triangular mesh and texture (extracted through Marching Cubes, etc.) to facilitate subsequent loading and rendering in the engine. In summary, this embodiment, leveraging text-guided generative 3D technology, can construct a virtual scene from scratch that matches the semantics of the dialogue, with highly realistic content details and visual effects.

[0093] More specifically, as an additional implementation method, the scene generation module of this embodiment includes an environment background generation unit, a character and entity generation unit, an event and plot reproduction unit, and an emotional atmosphere rendering unit.

[0094] The environment background generation unit determines the scene's environmental background based on extracted location and time elements. For example, if the dialogue mentions an accident occurring on a "highway at night," scene generation will use "highway at night" as the text prompt and generate the base environment using a diffusion model or existing environmental assets in the library. Similarly, "childhood bedroom" will trigger the construction of an indoor bedroom scene. For specific locations, this embodiment can prioritize retrieving predefined 3D materials (such as common city streets, parks, etc.) and then fine-tune them using a generation model to match the description. If the location is very unique, the scene is generated entirely based on the text. Time elements (such as day / night, season, etc.) will determine the scene's lighting and weather conditions, which can be reflected in the generated text prompts (such as "night," "snowy streets in winter," etc.) or achieved by modifying the ambient light map after generation.

[0095] The human entities identified by the character and entity generation unit are mapped to virtual characters in the scene. For key characters (such as the patient or significant others), pre-built controllable digital human models can be used, or corresponding character images can be created through generative methods. Deep learning text-to-head / body models (e.g., GAN-based character generation) can generate the character's face and body based on the description, and then bind motion skeletons for subsequent animation. If real-life reference images are available, neural human rendering techniques can also be used to generate realistic digital avatars. In addition to human characters, other entities mentioned in the dialogue (such as vehicles, animals, and specific items) will also be instantiated as 3D objects and inserted into the scene through generative models. For example, if the event involves a "car accident," a model of a damaged car is generated and placed in the corresponding position in the scene; if "house" is mentioned, a house model of the corresponding style is generated. These objects can be generated individually using the DreamFusion method (generating NeRF objects from given category and attribute text, then extracting them as meshes), or similar models can be called from an asset library and their appearance adjusted using neural texturing techniques to match the description.

[0096] The event and plot reenactment unit recreates events occurring in a scene using animation and special effects. Based on the provided event descriptions (such as "collision" or "fire"), the objects in the scene are driven to dynamically evolve according to the narrative. For example, in a "car accident" scene, two car models will move along a preset trajectory and trigger a collision physics simulation, producing particle sparks and debris flying effects to recreate the moment of the accident. If the event is an "argument," the corresponding virtual character model will trigger argumentative actions and expressions. To achieve these dynamic events, this embodiment uses a game engine's physics engine and animation system. The system sets key event frames on the timeline based on the event script (which can be automatically generated from the large model based on the description): for example, T = 1 second for the two cars to collide, T = 1.5 seconds for an explosion effect. The physics engine calculates the motion of objects after the collision based on realistic mechanics, and the particle system generates realistic flames, smoke, and other effects to achieve an immersive dynamic demonstration. The entire process is automatically arranged by the scene generation module, eliminating the need for manual animation script creation, demonstrating a high degree of intelligence and creative efficiency.

[0097] User emotions are also a crucial driving factor in scene generation. The emotional atmosphere rendering unit in this embodiment can adjust the atmosphere of the virtual environment based on the visitor's emotional state, thereby achieving emotional resonance or guidance. The sentiment analysis results will influence the scene's visual style parameters: when a visitor is detected to be depressed (sad, traumatized), the scene may use cooler tones, dim lighting, and rainy weather to echo their inner feelings; if it is necessary to guide them out of their gloom, the scene atmosphere will be gradually changed at appropriate times, such as the clouds dissipating and sunlight appearing, to symbolize hope.

[0098] This embodiment adjusts the overall color saturation and brightness of the scene during the rendering stage using neural style transfer or image filters, and changes the weather and day / night cycle by switching ambient light maps. These changes can be smoothly transitioned to ensure that users do not abruptly perceive scene changes and be pulled out of the immersive experience. For example, when the system detects that the user's mood has eased, the overcast sky can gradually clear, increasing the scene's brightness and color. In addition, background music or sound effects can also be generated or switched according to mood changes, enhancing the immersive experience through multimodal means.

[0099] In summary, the scene generation module in this embodiment uses extracted scenario elements as a blueprint to drive the real-time construction and updating of the VR environment using a generative 3D model. This process is highly automated and intelligent: unlike traditional VR content that requires pre-design by artists, this system can instantly create corresponding scenes based on the dialogue content and adjust and evolve the scenes as the dialogue continues, truly achieving "what is said is what is seen." For example, in psychotherapy, when a client first recounts a traumatic event, the system generates a scene depicting the event; if the client's emotions gradually calm down and they re-examine the event under the guidance of the therapist, the system's scene also undergoes positive changes simultaneously (such as the sky brightening and the surrounding environment gradually returning to calm from chaos), thus providing the client with an intuitive healing experience. This dialogue-driven dynamic scene generation technically combines multiple cutting-edge AI models (large language model + generative 3D + physics engine), opening up a completely new interactive content generation mode in the field of psychotherapy.

[0100] Furthermore, the graphics rendering module is used to render the 3D scene output by the scene generation module in real time to adapt to the display requirements of virtual reality devices. This module uses a high-performance graphics engine to convert dynamically generated content into immersive images that can be displayed by VR headsets, while ensuring low latency and high frame rate, making patients feel as if they are in the situation they are describing.

[0101] In this embodiment, the graphics rendering module uses a real-time rendering engine such as Unity3D or Unreal Engine to convert the scene data provided by the scene generation module into VR visual images. The module optimizes rendering for VR binocular displays to ensure immersion and realism. For example, in a classroom accident scene, the rendering module loads a classroom model, sets up lighting to simulate the dim lighting of a rainy day, renders raindrops and broken glass shards, and performs animation rendering for each virtual character (such as a panicked expression animation). Since the patient wears a head-mounted display (HMD), the rendering module needs to output images to both eyes at a rate of at least 90 frames per second to prevent dizziness and update the viewing angle based on the patient's head movements (achieving six degrees of freedom head tracking). This module has the highest requirements for system real-time performance; therefore, it employs multi-threading and GPU acceleration to keep the latency from voice input to image output within an acceptable range (e.g., around 100 milliseconds), ensuring that scene changes are almost synchronized with dialogue content. The graphics rendering module is also responsible for pushing the final image stream to the patient's VR device and sending the synchronized image to the therapist's monitoring end (e.g., displaying it on the therapist's computer screen or auxiliary HMD). Through an efficient rendering pipeline, the VR scene seen by the patient can change in sync with their own narrative voice, creating a seamless immersive experience.

[0102] More specifically, the graphics rendering module in this embodiment is responsible for presenting the 3D content output by the scene generation module to the user in real time with high-fidelity image quality. Considering the extremely high requirements of VR immersive experience for rendering effects and performance, this module integrates several cutting-edge rendering technologies, including real-time ray tracing, foveated rendering based on eye tracking, and neural rendering. With the support of advanced hardware, through optimized algorithms and performance assurance mechanisms, the system ensures that while maintaining cinematic image quality, it achieves the high frame rate and low latency rendering required for VR.

[0103] The graphics rendering module in this embodiment includes a real-time ray tracing rendering unit, a foveated rendering unit, a neural rendering and AI optimization unit, a multimodal feedback optimization unit, and a performance optimization unit;

[0104] Real-time ray tracing rendering units simulate the propagation, reflection, and refraction of light within a scene to produce realistic lighting effects. However, traditional ray tracing requires significant computation and has primarily been used for offline movie rendering. In recent years, with advancements in GPU architecture (such as the NVIDIA RTX series graphics cards), real-time ray tracing is gradually becoming a reality. The dedicated RT Cores units introduced in RTX hardware accelerate the traversal of the bounding volume hierarchy (BVH) and the intersection testing of rays and triangles in the scene, offloading these two time-consuming operations from general-purpose shaders. Simultaneously, Tensor Cores units execute AI-driven denoising algorithms, reconstructing near-noise-free images with fewer ray samples. Utilizing these technologies, a single consumer-grade GPU can complete a frame of ray-traced rendering for complex scenes in milliseconds.

[0105] This system fully utilizes real-time ray tracing to enhance scene realism: for key objects and environmental surfaces, real-time reflection and refraction calculations are enabled, allowing users to see realistic reflections of people in mirrors and scenes refracted through glass; a global illumination algorithm is used to simulate indirect diffuse reflection caused by multiple light bounces, resulting in soft and natural lighting in indoor scenes. In special event scenarios such as vehicle collisions, this embodiment also enables dynamic shadows and volumetric lighting effects through ray tracing, such as the light pillar effect formed by smoke and dust under car headlights. These ray tracing effects are provided by the underlying rendering engine, and this embodiment allows for custom adjustments through shader programming to match the visual style required by the psychological scene. Considering that real-time ray tracing is still expensive, this embodiment selectively applies it within the limits of performance: for example, the scene background may use a hybrid scheme of rasterization and ambient occlusion (SSAO), and ray tracing is only enabled for detailed rendering of local areas of interest to the user, thereby balancing performance and quality.

[0106] It is worth mentioning that this embodiment also introduces neural network-accelerated ray tracing optimization. For example, a neural radiance cache is used to accelerate global illumination calculation. This technique learns the distribution of multiple indirect illuminations in a scene through a small neural network, replacing the calculations required for multi-hop ray propagation in traditional path tracing with AI inference. Specifically, after the initial ricochet calculation, subsequent higher-order illuminations are quickly predicted by the trained radiance cache network based on scene features, thereby significantly reducing the overhead of complete path tracing.

[0107] Experiments show that this method accelerates global illumination computation and reduces noise with almost no loss of visual accuracy. Similarly, this embodiment uses AI super-resolution and AI denoising techniques (such as DLSS deep learning supersampling) to process the ray-traced output. Rendering is performed at a slightly lower resolution, then upscaled to the target resolution by the super-resolution network, increasing the frame rate while maintaining sharp details. For residual image noise from ray tracing, a denoising neural network trained extensively offline is used to filter each frame pixel by pixel, resulting in a smooth and clean image. These neural rendering aids, closely integrated with hardware ray tracing, significantly enhance the practicality of real-time ray tracing in this system, enabling this embodiment to present near-offline path-tracing quality images in VR, achieving a truly convincing immersive experience.

[0108] Gaze-based rendering unit: In VR rendering, simultaneously achieving high resolution, a large field of view, and a high frame rate is a significant challenge. The physiological characteristics of the human eye provide a key optimization strategy for this embodiment: the human eye has the highest visual acuity only when focusing on the focal area, while the resolution and perception of the peripheral vision are significantly reduced. Gaze-based rendering based on eye tracking is based on this principle, using different levels of precision for the user's gaze area and the surrounding areas during rendering, thus saving substantial computational resources. This system is equipped with a high-precision eye-tracking device capable of continuously obtaining the user's current gaze direction with a latency of less than ~5 milliseconds. The rendering module divides the field of view into concentric regions based on the eye's gaze point: a high-resolution central region (corresponding to the macular field of view), a surrounding medium-resolution region, and a low-resolution edge region. In actual rendering, this embodiment uses the original full-resolution rendering for the central region, while reducing the rendering resolution or level of detail for the peripheral regions. For example, the central 10° field of view is rendered with 1.0 times the pixel density, the middle ring with 0.5 times the density, and the edges with 0.25 times or lower density, with appropriate blur filtering applied to avoid hard-switching boundaries. Since the human eye is not sensitive to blurred edges, this approach of reducing peripheral resolution has almost no impact on perceptual quality, but it can significantly reduce pixel shading and ray tracing computations. Using dynamic foveated rendering at the default resolution can save approximately 60-70% of the GPU shading load. In actual testing, this embodiment also observed a significant increase in frame rate and more stable frame time after enabling foveated rendering; the frame rate easily dropped below 90Hz when disabled, but remained consistently above 90Hz after enabling it. This is crucial for avoiding VR motion sickness and ensuring a good user experience.

[0109] This unit prioritizes accuracy and low latency in foveated rendering. Eye-tracking data is transmitted to the rendering engine via a dedicated high-speed interface. This embodiment uses a prediction algorithm to make micro-predictions of eye movements to compensate for system latency. For example, a first-order exponential prediction model is used to estimate the user's gaze point position upon completion of rendering, preventing "misalignment." The rendering thread and eye-tracking acquisition thread operate in high parallelism, ensuring that the latest gaze point is obtained for each frame. This embodiment also corrects for VR lens distortion, deforming the shape of the gaze area on the rendering plane accordingly to match the actual circular area range in the field of vision after wearing the headset. For the ray tracing pipeline, this embodiment implements foveated adaptive ray allocation: more sampled rays are projected at the gaze center, while sampling is reduced or even completely approximated in the peripheral areas, thereby further saving ray tracing costs. In summary, through a combination of hardware and software foveated rendering, this system achieves the goal of "rendering exactly where the human eye is looking." While ensuring uncompromised image quality at the user's gaze point, it significantly reduces unnecessary peripheral rendering calculations, allowing for the most efficient use of GPU resources.

[0110] Neural Rendering and AI Optimization Unit: This unit improves performance or image quality by introducing neural networks into various stages of the traditional graphics pipeline. In addition to the previously mentioned AI denoising, super-resolution, and radiosity caching, this embodiment also applies various neural rendering techniques to the system, including:

[0111] Neural Shading and Materials: This embodiment delegates some complex lighting and material calculations to small neural network agents. For example, for multi-layered complex materials that would otherwise be rendered offline (such as automotive paint with multiple coatings), by training a neural network to approximate its shading function, near-offline quality can be achieved at runtime with extremely low cost. Furthermore, large amounts of high-resolution textures are compressed and encoded using neural textures, and then decoded and restored by AI during loading. These methods save video memory, accelerate rendering calculations, and make film-quality material details possible in real-time environments.

[0112] Neural Rendering for Complex Characters: For highly complex and realistic objects such as virtual faces, this embodiment introduces generative AI-assisted rendering. For example, Neural Faces technology is used to render realistic facial expressions. Specifically, the character's face is first coarsely rendered using traditional methods (containing only basic textures and simple lighting), and then a pre-trained generative model is used to produce a high-fidelity facial image replacement based on the input expression parameters and lighting information. This generative model is trained on a large amount of real-world facial data, including samples from different angles, lighting, and expressions. Therefore, at runtime, it can "complete" the coarse input into a photorealistic facial image.

[0113] This embodiment applies this technology to facial rendering of key characters (such as patient stand-ins), allowing subtle skin texture and facial expression changes to be seen even at close range, without the need for expensive cinematic facial capture equipment. This incredibly realistic digital human rendering is highly beneficial for psychotherapy scenarios, as the lifelike expressions of the characters enhance user immersion and trust. In addition to faces, this embodiment also explores neural human animation (e.g., rendering clothing wrinkles, hair strands, etc. using generative models) and neural ambient lighting (automatically adjusting light source color temperature and brightness based on scene context using AI). All these neural rendering techniques share the commonality of utilizing the prior knowledge of pre-trained models to achieve real-time real-time real-world effects that are difficult to achieve using traditional methods at extremely low cost, thus pushing the system's graphics fidelity to the extreme.

[0114] The multimodal feedback optimization unit continuously optimizes rendering quality through an AI feedback loop. Using a neural network similar to a GAN discriminator, this embodiment evaluates the difference between the rendered output and the real-world image in real time. If unrealistic artifacts (such as ray tracing noise or unnatural animation) are detected in certain areas of some frames, the feedback network locates these problems and notifies the rendering module to adjust parameters or trigger stronger AI correction. This mechanism is similar to an "AI quality inspector" constantly monitoring each frame, allowing the rendering effect to continuously self-correct and gradually approach physical realism. Furthermore, this embodiment uses reinforcement learning to train an intelligent LOD (Level of Detail) scheduler, where AI automatically determines the trade-offs between model detail levels based on scene complexity and hardware load, balancing frame rate and quality.

[0115] Performance Optimization Unit: To ensure that VR real-time requirements are met even in complex scenes and with a large amount of dynamic content, this module, in addition to employing the aforementioned cutting-edge technologies to improve single-frame rendering efficiency, also implements a global performance monitoring and adaptive optimization mechanism.

[0116] Frame rate monitoring and dynamic adjustment: The system sets a target frame rate (e.g., 90 FPS) and a maximum allowable latency. A monitoring module embedded in the rendering thread samples the GPU time and CPU overhead for each frame in real time. When a frame time approaches a threshold (e.g., >11ms), the system triggers a degradation strategy, such as temporarily reducing the rendering resolution, simplifying post-processing effects (disabling depth of field / motion blur, etc.), reducing the global illumination sampling rate, and even, in extreme cases, reducing concurrently active secondary scene objects. Image quality details are gradually restored after the load recovers. This dynamic resolution and effects adjustment ensures a stable frame rate, preventing severe frame drops due to individual complex scenes.

[0117] Multithreading and Asynchronous Pipeline: The rendering module fully utilizes multi-core CPUs and heterogeneous computing resources to parallelize the workflow. For example, the main rendering thread and the physics simulation thread are separated; physics calculations are performed in an independent thread, and the results are predicted several frames in advance to avoid blocking during the rendering phase. The scene generation module may perform NeRF optimization on the background GPU, which is also decoupled from the foreground rendering, and scene data is handed over through a double-buffer mechanism. This embodiment also employs asynchronous time warp technology: if a frame cannot be fully rendered in time, the depth buffer of the previous frame and the latest posture of the head-mounted display are used to perform a lightweight reprojection of the previous frame image to synthesize a temporary frame to fill in the gaps, thereby avoiding screen stuttering. These parallel and asynchronous measures effectively reduce the long-tail latency of the rendering pipeline, ensuring a continuous and smooth VR experience.

[0118] Resource load balancing: In VR applications, synchronously coordinating CPU, GPU, memory, and bandwidth is crucial. This unit introduces an intelligent resource manager to dynamically monitor GPU utilization and temperature, as well as CPU thread usage. If the GPU becomes a bottleneck while the CPU has spare capacity, it attempts to shift some computation to the CPU or reduce the GPU workload (e.g., by reducing the view distance clipping range). Conversely, the same applies. Through this cross-device load balancing, the overall system performance is fully utilized. For video memory, this embodiment employs on-demand loading and streaming detail technology to ensure that only essential high-precision resources within the field of view reside in video memory. Other resources are loaded or released in the background with low priority, avoiding swapping overhead caused by video memory overflow.

[0119] In summary, the graphics rendering module in this embodiment significantly improves the realism and rendering efficiency of VR scenes, from hardware-accelerated ray tracing to biologically inspired foveated rendering and AI-driven neural rendering. More importantly, this embodiment constructs a comprehensive performance optimization and assurance system, enabling the system to run stably at high frame rates and low latency under various complex conditions, meeting the requirements of immersive therapy. No matter how complex the scene or how frequent the dialogue-driven changes, the user always sees a smooth and realistic virtual world. This breakthrough in performance and image quality opens up new possibilities for immersive psychotherapy: counselors and clients can explore emotions and cognition in a highly realistic virtual context without being disturbed by technical lag or rough graphics.

[0120] Furthermore, the user interaction module provides an interface for patients and therapists to interact with the system. On one hand, patients experience the generated scene from a first-person perspective through the VR device they wear, and can react to or perform simple controls on the virtual environment through natural interactions (such as gazing, gesture control handles, etc.). On the other hand, therapists can observe the scene seen by the patient through the interactive interface, and pause / play scene generation, mark important events, or insert guiding content as needed, thereby participating in guiding the treatment process.

[0121] In this embodiment, the user interaction module includes two parts: patient interaction and therapist interaction. For patients, the module integrates VR interaction devices (such as controllers or motion-sensing devices), allowing patients to interact to a certain extent in the virtual environment. For example, patients can trigger the system to record a marker by looking at a virtual object and pressing a button on the controller, indicating that the element has special meaning to them; or patients can freely look around and move their position in the scene, re-examining the event from different angles. These interactions help patients exercise their agency and improve the treatment effect. For therapists, the interaction module provides a control interface: therapists can view the VR screen seen by the patient in real time through a tablet or PC interface, as well as the summary of scene elements extracted by the system (such as currently detected emotions, scene keywords, and other auxiliary information). The therapist interface allows for several operations, such as pausing / resuming automatic scene generation (pausing screen changes when a detail needs to be explored in depth), manually adjusting environmental parameters (for example, if the patient's narration is inaccurate, the therapist can manually correct erroneous elements in the scene), or adding therapist-defined guidance content to the scene (for example, introducing a virtual guide when the patient is confused). The user interaction module ensures flexibility in human-computer interaction through these functions, enabling the system to operate automatically and receive human guidance when needed, thereby serving the treatment process safely and efficiently.

[0122] As an additional implementation method, a patient recalled an accident that occurred during their childhood at school. At the start of the treatment, the patient wore a VR headset, and the therapist activated the system of this invention. When the patient recounted, "It was a rainy afternoon, I was in fifth grade, in class," the system sequentially performed speech recognition and semantic analysis: extracting the time as "afternoon / rainy day / fifth grade (approximately 11 years old)," the location as "elementary school classroom," the characters as "the patient (11-year-old student) and classmates / teacher," the emotion as "nervous" (inferred from tone and the rainy situation), and the environment as "rainy weather, dimly lit indoors." The scene generation module immediately constructed a classroom scene recreated on a rainy day: the sky outside was dim with the sound of raindrops, the classroom was furnished with desks and a blackboard appropriate for the era, a virtual student character (representing the patient) of similar age sat in the classroom, surrounded by other students and teachers, some of whom appeared tense. The patient saw this recreated classroom scene through the VR headset and heard environmental sound effects such as rain and thunder, as if transported back to that afternoon in their memory. As the patient continued to recount, "Suddenly there was a clap of thunder, and then the window shattered, everyone was terrified," the system's natural language processing module updated the event element to "the window was struck by lightning and shattered," escalating the emotion to "panic." The scene generation module instantly added the window-shattering effect and sound to the classroom scene, with virtual teacher and student characters reacting with animations of dodging and screaming. The scene the patient saw was perfectly synchronized with their own narrative: the window shattered, shards flew, and the surrounding classmates screamed in chaos. The patient seemed to relive the scene, their repressed fear being triggered and released and processed under the therapist's guidance. Throughout the process, the therapist observed the virtual scene through a monitoring interface, gaining a direct understanding of the patient's environment and emotions, enabling more targeted reassurance and questioning. For example, the therapist could have the patient point out the location or object they were most afraid of in the VR, allowing for a deeper exploration of related details. This embodiment fully demonstrates the value of this invention in psychotherapy: through a real-time generated immersive VR scene, patients can safely revisit traumatic memories and express their emotions, while therapists gain an unprecedented intuitive perspective to understand the patient, significantly improving therapeutic efficacy.

[0123] It should be noted that the application of this embodiment is not limited to the above-described scenarios. For other narrative therapy scenarios, such as a patient recalling a warm or traumatic scene in their childhood family living room, the scene of an accident in adulthood, or even descriptions of inner dreams, this system can extract key elements to generate corresponding VR scenes, making the treatment process more vivid and effective.

[0124] In this embodiment, the various modules of the system communicate and collaborate via a system bus or network: speech flows through the speech recognition module to generate text, the text is sent to the language processing module to extract contextual elements, then the scene generation module generates a 3D scene in real time, the rendering module displays it on the VR device, and the interaction module ensures human-computer interaction and the therapist's control. The system architecture is as follows: Figure 1 As shown, the modules are closely integrated, enabling the automatic conversion of dialogue content into immersive VR scenes.

[0125] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An immersive psychotherapy VR system based on real-time dialogue generation, characterized in that, include: The speech recognition module is used to collect the dialogue voice during the consultation process and transcribe it into text in real time; The natural language processing module, connected to the speech recognition module, is used to perform semantic understanding on the transcribed dialogue text and extract contextual elements; The scene generation module, connected to the natural language processing module, is used to dynamically generate corresponding three-dimensional virtual scenes based on the contextual elements, and adjust scene details in real time as the dialogue content is updated. The graphics rendering module, connected to the scene generation module, is used to render the three-dimensional virtual scene in real time and output it to the VR device. The user interaction module, connected to the graphics rendering module, is used to enable interactive operations between patients, therapists, and the system.

2. The system according to claim 1, characterized in that, The speech recognition module includes a microphone array and a speech transcription subsystem; The microphone array is used to pick up conversational audio and perform noise reduction and speech separation processing; The speech-to-text subsystem outputs text recognition results in real time in a streaming manner using a deep neural network acoustic model and a large language model.

3. The system according to claim 2, characterized in that, The deep neural network acoustic model includes a Conformer encoder employing end-to-end automatic speech recognition technology, a Conformer Block that integrates convolutional neural networks and Transformer self-attention mechanisms, and is used to simultaneously capture local detail features and global dependencies of speech signals. The large language model includes an explicit language model and a decoder; The decoder performs the search based on the beam search principle.

4. The system according to claim 1, characterized in that, The natural language processing module stores the aforementioned elements in the dialogue state through the context preservation function, and associates new information with the existing context when further describing details to obtain contextual elements; The contextual elements include the time, place, characters, emotional states, and environmental description of the event.

5. The system according to claim 1, characterized in that, The natural language processing module includes a keyword and event extraction unit, a sentiment analysis unit, an entity recognition and relation extraction unit, and a contextual modeling unit; The keyword and event extraction unit is used to extract keywords that express core semantics from each dialogue sentence, as well as sentence components that describe important events; The sentiment analysis unit is used to analyze the sentiment tendency of dialogue statements and determine the speaker's current emotion. The entity recognition and relationship extraction unit is used to identify entities from the dialogue and further extract the relationships between entities; The entities include names of people, locations, and organizations; The contextual modeling unit is used to understand multi-turn dialogues as a whole context through long sequence processing and to splice the dialogue history into a long text; it is also used to detect dialogue turns and contextual changes.

6. The system according to claim 1, characterized in that, The scene generation module includes an environment generation unit and a character modeling unit; The environment generation unit is used to render the corresponding scene background according to time, location and environmental elements; The character modeling unit is used to create virtual dolls or characters based on character elements, and to adjust the atmosphere of the scene and the character's expression in combination with emotional elements.

7. The system according to claim 1, characterized in that, The scene generation module constructs dynamic three-dimensional virtual scenes based on generative 3D models; The generative 3D model integrates NeRF, Gaussian Splatting, and DreamFusion generative modeling technologies, and uses extracted text elements to drive scene generation and evolution.

8. The system according to claim 1, characterized in that, The graphics rendering module includes a real-time ray tracing rendering unit, a foveated rendering unit, a neural rendering and AI optimization unit, a multimodal feedback optimization unit, and a performance optimization unit; The ray tracing rendering unit is used to simulate the propagation, reflection and refraction of light in the scene using real-time ray tracing, to simulate indirect diffuse reflection caused by multiple light bounces using a global illumination algorithm, and to accelerate global illumination calculation using a neural radiation cache. The gaze point rendering unit is used to render the user's gaze area and the surrounding area with different precisions during rendering using an eye-tracking device; The neural rendering and AI optimization unit is used to perform neural rendering by combining neural shading with materials and neural rendering of complex characters. The multimodal feedback optimization unit is used to continuously optimize rendering quality through an AI feedback loop; The performance optimization unit is used for global performance monitoring and adaptive optimization through frame rate monitoring and dynamic adjustment, multi-threading and asynchronous pipelines, and resource load balancing.

9. The system according to claim 1, characterized in that, The user interaction module includes a patient interaction unit and a therapist interaction unit; The patient interaction unit is used to experience the generated scene and interact naturally through VR devices; The therapist interaction unit is used to observe the scene from the patient's perspective through the control interface and control the scene presentation or the rhythm of the dialogue.

Citation Information

Patent Citations

  • Psychological counseling virtual interaction scene construction system based on voice recognition

    CN115206488A

  • Real-time ray tracing method and system for VR large-space immersive touring

    CN120298566A

Cited By

  • Video AR virtual reality interactive communication method and system based on AI model

    CN121788771A