Mental health assessment method and system based on VR interaction and multi-modal data

By combining voice and behavioral data in a virtual reality environment, a deep learning model has been developed to address the issues of insufficient feature recognition and device dependence in the assessment of depression and anxiety disorders in existing technologies. This has enabled efficient and low-cost mental health assessment, improving the accuracy and universality of early screening.

CN121905500APending Publication Date: 2026-04-21BEIJING JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING JIAOTONG UNIV
Filing Date
2025-12-15
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

Existing technologies for assessing depression and anxiety have problems such as insufficient sensitivity in feature recognition, reliance on specialized equipment or personnel, and limited design of VR emotion-stimulating scenarios, resulting in low accuracy, high cost, poor universality, and low data ecosystem validity in early screening.

Method used

This study employs a mental health assessment method based on VR interaction and multimodal data. By completing interactive question-and-answer and recording tasks in a virtual environment, voice signals and interactive behavior data are obtained. A pre-trained GRU-based deep learning model is used for feature extraction and fusion analysis. Complementary features are extracted by combining VGGish and VLAD features. Interactive tasks that are both fun and guided are designed to stimulate high-quality emotional response data.

Benefits of technology

It improves the sensitivity and accuracy of identifying subtle and dynamic physiological and behavioral patterns in people with sub-optimal mental health, enables low-cost and easily promoted early screening of mental states, provides an immersive and highly interactive assessment solution, and protects user privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121905500A_ABST
    Figure CN121905500A_ABST
Patent Text Reader

Abstract

The invention provides a psychological health assessment method and system based on VR interaction and multi-modal data, and belongs to the technical field of psychological emotion detection based on virtual reality, and the method comprises the steps: carrying out the feature extraction and fusion analysis of obtained voice signals and interaction behavior data through a pre-trained GRU-based deep learning model; obtaining a psychological state assessment result of the user; wherein the GRU-based deep learning model utilizes a GRU branch to deeply mine time sequence dynamic information in the audio, and the time sequence dynamic information is fused with behavior characteristics. According to the psychological assessment method, weak and dynamic physiological behavior modes of psychological sub-health people can be effectively captured, and by adopting the time sequence deep learning model, the recognition sensitivity and accuracy of early psychological state abnormity are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual reality-based psychological and emotional detection technology, specifically to a psychological health assessment method and system based on VR interaction and multimodal data. Background Technology

[0002] Traditional neuropsychological tests face challenges in achieving ecological validity when assessing cognitive deficits in depression and anxiety. These tests rely heavily on patient subjective descriptions, clinical observations, and feedback from family members. The test tasks fail to accurately simulate real-life situations, leading to potentially unreliable and concealable results. Furthermore, their time-consuming nature, complex procedures, and reliance on professional involvement hinder the early and rapid diagnosis of conditions such as depression and anxiety.

[0003] Extensive research has focused on using speech, facial expressions, and body gestures for emotion recognition. However, collecting higher-quality objective physiological signals to provide more detailed and complex emotional state information remains a challenge. In recent years, many studies have conducted emotion recognition based on physiological signals using virtual reality (VR). The application potential of VR in the field of mental health prediction is increasingly prominent, mainly due to its two core advantages: firstly, VR environments can effectively stimulate user emotions through multi-sensory stimulation, providing a high degree of realism and immersion, thus creating conditions for collecting high-quality, high-volume behavioral and reaction data; secondly, all data is generated in the virtual world, offering a natural advantage in data privacy protection. Against this backdrop, a series of studies have been dedicated to verifying the effectiveness of VR technology in emotion stimulation and recognition. Hidaka et al. (2017) established a series of VR scenarios, providing emotion-stimulating experiments to subjects using head-mounted displays (HMDs), verifying the usability and effectiveness of emotion-stimulating VR scenarios using HMDs. However, most studies, similar to Hidaka et al.'s, stimulate emotions by playing music, videos, or digital content. However, research on emotion recognition using interactive VR scenarios is scarce. Péron et al. (2011) found that patients with major depressive disorder had significantly reduced responses to positive emotions and significantly increased responses to negative emotions, and this correlation was significantly higher than the severity of the condition. Lauren S. Hallion et al. (2018) found that depression and anxiety were associated with impaired attention. However, past studies have generally suffered from small sample sizes because data from individuals with mood disorders such as depression and anxiety is difficult to collect. VR technology is being used more extensively for mental health prediction and assessment. Using VR data to predict mental health status emphasizes that the immersive and engaging nature of the VR environment can improve the quality and quantity of data, and VR data does not involve sensitive information, offering privacy advantages.

[0004] However, existing technical solutions still have the following main technical drawbacks:

[0005] (1) Insufficient sensitivity of feature recognition leads to low accuracy of early screening: Existing methods based on voice or behavior analysis (Wang et al. (2019), Xin Yu (2023)) mostly rely on static statistical features or simple feature fusion. Their models are not sensitive to the weak, nonlinear and dynamic temporal feature patterns exhibited by people with sub-optimal mental health, making it difficult to achieve effective early identification and warning.

[0006] (2) The assessment system relies on professional equipment or personnel and has poor universality: Some existing solutions (Wan Chunting (2023), Fu Yixiao et al. (2018)) require professional physiological signal acquisition equipment such as EEG, or must be operated by psychological professionals, resulting in high system cost and complex process, making it difficult to achieve low-cost, large-scale universal screening in community, family and other scenarios.

[0007] (3) The VR emotion-stimulating scenario design is monotonous and the data ecosystem validity is low: Existing VR psychological assessment schemes (Fu Yixiao et al. (2018), Shu Lin (2024)) mostly involve passive observation of static materials or are designed as high-pressure, low-interest tasks. Such settings with insufficient interactivity or poor experience make it difficult to stimulate users' real, natural, and high-quality emotional response data, thereby affecting the intrinsic quality of the collected data and the effectiveness of the assessment. Summary of the Invention

[0008] The purpose of this invention is to provide a mental health assessment method and system based on VR interaction and multimodal data, so as to solve at least one of the technical problems existing in the background art.

[0009] To achieve the above objectives, the present invention adopts the following technical solution:

[0010] In a first aspect, the present invention provides a mental health assessment method based on VR interaction and multimodal data, comprising:

[0011] Users to be evaluated interactively complete quiz and recording tasks in a virtual environment;

[0012] Acquire voice signals and interactive behavior data of the user to be evaluated in the interactive task;

[0013] A pre-trained GRU-based deep learning model is used to extract and fuse features from acquired speech signals and interactive behavior data to obtain the user's psychological state assessment results. The GRU-based deep learning model uses a GRU branch to deeply mine the temporal dynamic information in the audio and fuse it with behavioral features.

[0014] As a further limitation of the first aspect of the present invention, the GRU-based deep learning model includes: an audio temporal modeling branch, used to input the generated comprehensive audio feature vector sequence into a gated recurrent unit neural network to learn the dynamic evolution pattern of the emotional state reflected by the user's speech features in continuous tasks; the hidden state of the last time step of the last layer of the gated recurrent unit neural network is used as a high-level temporal feature representation of the entire audio sequence; a behavioral feature concatenation and fusion network, used to horizontally concatenate the normalized user reaction time feature vector with the high-level audio temporal feature representation output by the gated recurrent unit neural network in the feature dimension to form a unified feature vector that integrates complex temporal speech patterns and simple static behavioral indicators; and a classification output network, used to pass the concatenated fused feature vector to a single-layer fully connected network, and finally output the final psychological state classification result through a Softmax classifier.

[0015] As a further limitation of the first aspect of the present invention, for voice data, silent segments and invalid audio shorter than 1 second are automatically removed; the silent portions at the beginning and end of each recording are cut off, and valid voice is retained.

[0016] As a further limitation of the first aspect of the present invention, a dual-path parallel architecture is adopted to extract complementary features from the audio, including: VGGish embedding features: the preprocessed audio is input into a pre-trained VGGish model, which outputs a fixed-dimensional embedding feature vector to represent the high-level semantic content of the audio; VLAD features: the audio is converted into a Mel spectrogram using the Librosa library, and then encoded using the NetVLAD model to generate a VLAD feature vector to capture the local statistical characteristics of the audio; feature fusion and serialization: the VGGish feature vector and the VLAD feature vector of each audio sample are concatenated to form a comprehensive audio feature; the comprehensive audio feature sequence and the behavioral data sequence are aligned according to the task order to form multimodal temporal data.

[0017] As a further limitation of the first aspect of the present invention, the construction of the virtual environment scene includes: first, building a teaching interface and multiple task scenes according to psychological paradigms and assessment needs; the entrance music before entering the scene relaxes the subjects and returns them to a relatively calm and focused state, so as to better prepare for the experience; multiple interactive areas are distributed in the scene, and emotional stimulation materials that have been screened and labeled by psychological experts are placed in the area, including paintings and text fragments with three emotional tendencies: positive, neutral and negative.

[0018] As a further definition of the first aspect of the present invention, the user's interactive tasks in the virtual scene include: after entering the interactive area in the virtual scene, the user completes different types of interactive tasks; the question-and-answer task requires the user to observe the painting and answer questions about the details of the painting; the recording task requires the user to observe the painting and record an audio recording of a text describing the painting; after each task is completed, the data is packaged immediately and includes the user ID, task ID and timestamp; when all preset tasks are completed, or the user actively chooses to exit, the current data collection session ends.

[0019] Secondly, the present invention provides a mental health assessment system based on VR interaction and multimodal data, comprising:

[0020] The virtual scene interaction module is used by the user to be evaluated to complete the task of answering questions and recording audio in a virtual scene.

[0021] The acquisition module is used to acquire the voice signals and interactive behavior data of the user to be evaluated in the interactive task.

[0022] The evaluation module is used to perform feature extraction and fusion analysis on the acquired speech signals and interactive behavior data using a pre-trained GRU-based deep learning model to obtain the user's psychological state evaluation results; wherein, the GRU-based deep learning model uses a GRU branch to deeply mine the temporal dynamic information in the audio and fuse it with behavioral features.

[0023] Thirdly, the present invention provides a non-transitory computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the mental health assessment method based on VR interaction and multimodal data as described in the first aspect.

[0024] Fourthly, the present invention provides a computer device including a memory and a processor, wherein the processor and the memory communicate with each other, the memory stores program instructions executable by the processor, and the processor invokes the program instructions to execute the mental health assessment method based on VR interaction and multimodal data as described in the first aspect.

[0025] Fifthly, the present invention provides an electronic device, comprising: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions for implementing the mental health assessment method based on VR interaction and multimodal data as described in the first aspect.

[0026] Terminology Explanation: Virtual Reality (VR): A computer-generated simulated environment that allows users to experience an immersive environment through head-mounted devices and interactive devices. Multimodal: In this project, it refers to the comprehensive analysis of different types of data (such as sound and behavior). Ecological Validity: Ecological validity is also an indicator of the degree to which the results of psychological theories or experimental research can be generalized to real-life situations. In psychological research, high ecological validity usually means that the research materials, environment, and tasks are real, rather than hypothetical. GRU: An improved recurrent neural network based on gated recurrent units, used to process sequential data. Emotionally Arousing Material: Content used to evoke or stimulate emotions, such as paintings, music, and text.

[0027] This invention offers several advantages. First, it provides a psychological assessment method that effectively captures subtle and dynamic physiological and behavioral patterns in individuals with sub-optimal mental health. By employing a temporal deep learning model, it improves the sensitivity and accuracy of identifying early psychological abnormalities. Second, it provides a non-invasive, low-cost, and easily deployable early screening solution for mental health. This solution utilizes consumer-grade hardware such as VR headsets and microphones to collect voice and behavioral data, eliminating reliance on specialized equipment and personnel and enabling convenient deployment and large-scale application. Third, it provides a highly immersive, interactive, and user-friendly VR emotion-stimulating scenario. By designing engaging and guided interactive tasks, it stimulates more realistic and high-quality behavioral and voice data while protecting user privacy, laying the foundation for accurate assessment.

[0028] The advantages of additional aspects of the invention will be set forth more clearly in the following description or will be learned by practice of the invention. Attached Figure Description

[0029] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0030] Figure 1 This is a flowchart illustrating the process of scenario interaction and emotion assessment for test subjects as described in this embodiment of the invention.

[0031] Figure 2 This is a design architecture diagram of the emotional stimulation scenario and material placement for the psychological paradigm described in this embodiment of the invention.

[0032] Figure 3 This is a diagram illustrating the positive emotion-stimulating question-answering task described in an embodiment of the present invention.

[0033] Figure 4This is a diagram illustrating the recording task for stimulating positive emotions as described in an embodiment of the present invention.

[0034] Figure 5 This is a diagram illustrating the question-and-answer task that induces neutral emotions as described in an embodiment of the present invention.

[0035] Figure 6 This is a diagram illustrating the recording task for stimulating neutral emotions as described in an embodiment of the present invention.

[0036] Figure 7 This is a diagram illustrating the question-and-answer task for stimulating negative emotions as described in an embodiment of the present invention.

[0037] Figure 8 This is a diagram illustrating the recording task for stimulating negative emotions as described in an embodiment of the present invention.

[0038] Figure 9 This is a flowchart illustrating the interactive tasks of the test subjects in a virtual scene as described in an embodiment of the present invention.

[0039] Figure 10 This is the core flowchart of the MGRU model (binary classification) described in the embodiments of the present invention.

[0040] Figure 11 This is a diagram showing the login page of the data platform according to an embodiment of the present invention.

[0041] Figure 12 This is a schematic diagram illustrating the distribution of personal data displayed on the login homepage in a population, as described in an embodiment of the present invention.

[0042] Figure 13 This is a population distribution map of the mental health assessment dataset described in this embodiment of the invention.

[0043] Figure 14 This is a schematic diagram showing the population distribution of scale scores according to an embodiment of the present invention.

[0044] Figure 15 This is a schematic diagram illustrating the distribution of anxiety levels based on sleep quality adjustment, as described in an embodiment of the present invention.

[0045] Figure 16 This is a schematic diagram illustrating the distribution of the comprehensive level of depressive mood based on sleep quality adjustment as described in an embodiment of the present invention.

[0046] Figure 17 This diagram illustrates the distribution of individuals with different levels of depressive mood across three voice indicators, as described in this embodiment of the invention. In the diagram, the horizontal axis represents 0 as no depressive mood, 1 as mild depressive mood, and 2 as moderate depressive mood. Detailed Implementation

[0047] Embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0048] It will be understood by those skilled in the art that, unless otherwise defined, all terms used herein (including technical and scientific terms) have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0049] It should also be understood that terms such as those defined in general dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art, and should not be interpreted in an idealized or overly formal sense unless defined as described here.

[0050] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, and / or groups thereof.

[0051] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.

[0052] To facilitate understanding of the present invention, the present invention will be further explained and described below with reference to the accompanying drawings and specific embodiments. However, the specific embodiments do not constitute a limitation on the embodiments of the present invention.

[0053] Those skilled in the art should understand that the accompanying drawings are merely schematic diagrams of embodiments, and the components in the drawings are not necessarily essential for implementing the present invention.

[0054] This invention integrates VR immersive interactive scenarios, emotional stimulation, and machine learning: it effectively stimulates the user's sensory system and brain neural network through VR interactive tasks that conform to psychological paradigms, while simultaneously collecting and intelligently analyzing their speech and behavioral response data. Data shows that while the acoustic characteristics of sub-healthy individuals are consistent with the trends observed in clinical patients, the signal strength is weak, highlighting the need for highly sensitive analysis technology. Therefore, this invention specifically designs a machine learning-based method to identify and evaluate these weak signals, overcoming the limitations of existing methods in early screening.

[0055] Example 1

[0056] In this embodiment 1, a mental health assessment system based on VR interaction and multimodal data is first provided, including: a virtual scene interaction module, used for the user to be assessed to interactively complete answering and recording tasks in a virtual scene; an acquisition module, used to acquire the user's voice signals and interactive behavior data in the interactive tasks; and an assessment module, used to use a pre-trained GRU-based deep learning model to perform feature extraction and fusion analysis on the acquired voice signals and interactive behavior data to obtain the user's mental state assessment results; wherein, the GRU-based deep learning model uses a GRU branch to deeply mine the temporal dynamic information in the audio and fuse it with behavioral features.

[0057] In this embodiment, the above-described system can be used to implement a mental health assessment method based on VR interaction and multimodal data, including: using a virtual scene interaction module to allow the user to be assessed to interactively complete answering and recording tasks in a virtual scene; using an acquisition module to acquire the user's voice signals and interactive behavior data during the interactive tasks; and using an assessment module to perform feature extraction and fusion analysis on the acquired voice signals and interactive behavior data using a pre-trained GRU-based deep learning model to obtain the user's mental state assessment result; wherein, the GRU-based deep learning model uses a GRU branch to deeply mine the temporal dynamic information in the audio and fuse it with behavioral features.

[0058] The GRU-based deep learning model includes: an audio temporal modeling branch, which inputs the generated comprehensive audio feature vector sequence into a gated recurrent unit neural network to learn the dynamic evolution pattern of the user's emotional state reflected by their speech features in continuous tasks; the hidden state of the last time step of the last layer of the gated recurrent unit neural network is used as a high-level temporal feature representation of the entire audio sequence; a behavioral feature concatenation and fusion network, which horizontally concatenates the normalized user reaction time feature vector with the high-level audio temporal feature representation output by the gated recurrent unit neural network in the feature dimension to form a unified feature vector that integrates complex temporal speech patterns and simple static behavioral indicators; and a classification output network, which passes the concatenated fused feature vector to a single-layer fully connected network, and finally outputs the final psychological state classification result through a Softmax classifier.

[0059] For voice data, silent segments and invalid audio shorter than 1 second are automatically removed; the silent portions at the beginning and end of each recording are cut off, and valid voice is retained.

[0060] A dual-path parallel architecture is employed to extract complementary features from the audio, including: VGGish embedding features: preprocessed audio is input into a pre-trained VGGish model, outputting a fixed-dimensional embedding feature vector to represent the high-level semantic content of the audio; VLAD features: the audio is converted into a Mel spectrogram using the Librosa library, and then encoded using the NetVLAD model to generate a VLAD feature vector to capture the local statistical characteristics of the audio; feature fusion and serialization: the VGGish feature vector and VLAD feature vector of each audio sample are concatenated to form a comprehensive audio feature; the comprehensive audio feature sequence is aligned with the behavioral data sequence according to the task order to form multimodal temporal data.

[0061] The construction of the virtual environment scenario includes: First, building a teaching interface and multiple task scenarios based on psychological paradigms and assessment needs; the entrance music before entering the scenario helps the subjects relax and return to a relatively calm and focused state, better preparing them for the experience; multiple interactive areas are distributed in the scenario, and emotionally stimulating materials selected and labeled by psychology experts are placed in the areas, including paintings and text fragments with three emotional tendencies: positive, neutral, and negative.

[0062] User interaction tasks in the virtual scene include: after entering the interaction area in the virtual scene, users complete different types of interaction tasks; the question-and-answer task requires users to observe the painting and answer questions about the details of the painting; the recording task requires users to observe the painting and record an audio recording of a text describing the painting; after each task is completed, the data is packaged immediately and includes the user ID, task ID and timestamp; when all preset tasks are completed, or the user actively chooses to exit, the data collection session ends.

[0063] Specifically, the system and method described in this embodiment will be described in detail below with reference to the accompanying drawings.

[0064] The overall system architecture of this embodiment mainly includes three core components: a VR interaction and data acquisition terminal, an intelligent data management platform, and a multimodal emotion recognition and analysis module. These three components work together to complete the entire process from emotion stimulation and data acquisition to intelligent analysis and result feedback.

[0065] (1) VR Interaction and Data Acquisition Terminal: It consists of a VR head-mounted display device, an operating handle and a built-in microphone. It is responsible for presenting an immersive virtual scene to the user, receiving user interaction commands, and simultaneously collecting the user's voice signals and interactive behavior data.

[0066] (2) Multimodal emotion recognition and analysis module: This is the core computing unit of the present invention. It obtains data from the platform, uses a deep learning model based on GRU to perform feature extraction and fusion analysis, and finally outputs the psychological state assessment results and returns the results to the platform for display.

[0067] (3) Intelligent data management platform: responsible for storing and managing all raw data and user information from the terminal, and providing a data visualization interface.

[0068] The workflow of the system (e.g.) Figure 1 As shown): After the user completes the experience on the VR terminal, the collected data is uploaded to the platform; the experiment administrator can start the analysis module through the platform; after the analysis module completes the calculation, the evaluation results are returned to the platform and can be queried by the administrator and the user (with authorization).

[0069] The core of VR interaction and data acquisition terminals lies in using original VR scenes that conform to psychological paradigms to stimulate user emotions through proactive interaction and simultaneously collect multimodal data.

[0070] For the development of virtual environment scenarios, firstly, a teaching interface and various task scenarios are built based on psychological paradigms and assessment needs. The entrance music before entering the scenario helps participants relax and return to a relatively calm and focused state, better preparing them for the experience. Multiple interactive areas are distributed within the scenario, containing emotionally stimulating materials selected and labeled by psychological experts, including paintings and text fragments with positive, neutral, and negative emotional tendencies. In this embodiment, single images, videos, or audio are placed in a three-dimensional environment using VR technology, making them more realistic and three-dimensional, constructing more vivid and diverse emotionally stimulating scenarios (such as...). Figure 2 (As shown).

[0071] In this embodiment, the original scenarios considered both the stimulating effect and the duration of the experience. Each scenario included 22 paintings designed to stimulate the participants' emotions. The overall material library was manually selected and labeled with different emotion types, including positive, neutral, and negative, with the assistance of psychology experts. Each emotion type was used in two types of emotion-stimulating tasks: a question-and-answer task and a recording task.

[0072] The scene contains 7 tasks to stimulate positive emotions: (1) such as Figure 3 As shown, the positive emotion-stimulating question-answering task interface presents a painting labeled as having a positive emotion (e.g., a blooming colorful bouquet), along with a descriptive text hinting at the painting's positive tendency. A multiple-choice interactive area is located at the bottom of the interface, where questions drive the user to be evaluated to carefully observe and experience the painting's emotion (e.g., 'What color of flower is in this painting?'). After the user selects an answer using the handle, the system records the selection result and response time as behavioral data under positive emotion stimulation. (2) As Figure 4 As shown, the interface for the positive emotion-inducing recording task displays a painting labeled as evoking positive emotions (such as a bright coastal landscape), accompanied by a positive text description of the painting's overall atmosphere and detailed depiction. A recording prompt is displayed on the right side of the interface; the user presses the button on the controller to begin recording. The system simultaneously collects voice data and records the task duration for analysis of internal emotional characteristics.

[0073] The scene contains 7 neutral emotion-evoking tasks: (1) such as Figure 5 As shown, the neutral emotion-stimulating question-answering task interface presents a painting marked as neutral emotion (e.g., a geometric shape with progressively darkening colors) and a neutral, objective description of the painting. A multiple-choice question interaction area is located at the bottom of the interface, guiding the user to view and objectively analyze the painting (e.g., 'How does the color change from the periphery to the center of this painting?'). After the user selects an answer using the handle, the system records the selection result and response time as behavioral data under neutral emotion stimulation. (2) As Figure 6As shown, the interface for the neutral emotion-inducing recording task displays a painting labeled as neutral in emotion (e.g., a combination of geometric shapes) along with a text description explaining the objective content of the painting. A recording prompt is displayed on the right side of the interface; the user presses the button on the controller to begin recording. The system simultaneously collects voice data and records the task duration for analysis of internal emotional characteristics.

[0074] The scene contains 8 tasks to evoke negative emotions: (1) such as Figure 7 As shown, the negative emotion-stimulating question-and-answer task interface presents a painting labeled as having a negative emotion (e.g., a disturbing lion statue and puppet), along with a textual analysis of the painting with a negative tendency. A multiple-choice interactive area is located at the bottom of the interface, where questions drive the user to be evaluated to carefully observe and experience the emotions in the painting (e.g., 'What is in the center of this painting?'). After the user selects an answer using the handle, the system records the selection result and response time as behavioral data under negative emotion stimulation. (2) As Figure 8 As shown, the interface for the negative emotion-inducing recording task displays an artwork evoking negative emotions (such as a negotiating president, death, a bloodstained white sheet, and recording equipment), accompanied by a description of the overall expression and details of the artwork. Recording prompts are displayed on the right side of the interface; the user presses the button on the controller to begin recording. The system simultaneously collects voice data and records the task duration for analysis of underlying emotional characteristics.

[0075] To ensure immersion and standardization, the scene design follows these key technical points:

[0076] (1) Field of vision and free exploration: The scene layout is open, and users can move and rotate the view freely through the handle, minimizing the dependence on the experimenter and simulating the exploration behavior in a real environment.

[0077] (2) Interaction area triggering mechanism: To avoid cluttered interface, all tasks are triggered by "entry". When the user (through their virtual avatar) enters the preset interaction area, the corresponding task UI (user interface) will be activated and displayed.

[0078] like Figure 9 As shown, the user experience and data collection process in the scenario are as follows:

[0079] (1) Initialization and learning: Users first enter the initial page and learn the basic operation of the handle through guidance. This greatly reduces the subject's dependence on the experimenter and external interference. (2) Free exploration and task triggering: Users freely explore the scene. After entering the interaction area, they complete different types of interactive tasks. The question-and-answer task requires users to observe the painting and answer questions about the details of the painting. The recording task requires users to observe the painting and record an audio recording of a text describing the painting. (3) Data encapsulation and uploading: After each task is completed, the data (behavioral data, options, audio files) is packaged in real time and automatically uploaded with information such as user ID, task ID, and timestamp. (4) Termination mechanism: When all preset tasks are completed, or the user actively chooses to exit, the data collection session ends.

[0080] The multimodal emotion recognition and analysis module is responsible for feature extraction and pattern recognition of the collected data. Its core process is as follows: Figure 10 As shown.

[0081] Data preprocessing includes: (1) Speech data: Automatically remove silent segments and invalid audio shorter than 1 second. Remove the silent parts at the beginning and end of each recording and retain the valid speech. All audio is finally manually verified to ensure quality. This embodiment adopts a dual-path parallel architecture to extract complementary features from the audio. VGGish embedding features: Input the preprocessed audio into the pre-trained VGGish model and output a fixed-dimensional embedding feature vector to represent the high-level semantic content of the audio. VLAD features: Use the Librosa library to convert the audio into a Mel spectrogram, and then encode it through the NetVLAD model to generate a VLAD feature vector to capture the local statistical characteristics of the audio. (2) Behavioral data: Mainly use reaction time. Normalize the reaction time of all answer tasks and form a behavioral sequence. (3) Feature fusion and serialization: Concatenate the VGGish feature vector and VLAD feature vector of each audio sample to form a comprehensive audio feature. Align the comprehensive audio feature sequence and the behavioral data sequence according to the task order to form multimodal temporal data.

[0082] For the MGRU model, an efficient hybrid architecture is adopted, such as... Figure 10 As shown. The specific steps for training the MGRU model are as follows:

[0083] (1) Training set construction

[0084] The training set originates from user multimodal data collected by the VR interaction and data acquisition terminal, including audio feature sequences, behavioral feature vectors, and psychological state labels. Audio feature sequences: Raw speech signals synchronously collected when users complete various recording tasks in the VR scene. Each user's audio data is organized by emotion type and task number. Emotion types include three categories: Positive, Neutral, and Negative. Each category contains three task audio files, totaling nine files. Behavioral feature vectors: Reaction time sequences recorded when users complete various question-and-answer tasks in the VR scene, pre-standardized. These two types of data are combined according to the above data preprocessing steps to form multimodal time-series data. Psychological state labels: Serving as the ground truth for supervised model learning, they are calculated using one of two methods. The binary classification method refers to labeling users as "no depressive mood" (label 0) or "having depressive mood" (label 1) based on the calculation results of the user's Beck Depression Scale Version 2. The five-category approach refers to dividing users' psychological states into five levels based on the combined calculation results of the Beck Depression Rating Scale 2 and the Pittsburgh Sleep Quality Index: "Healthy", "Mild Distress", "Mild Depressive Mood", "Moderate Depressive Mood", and "Severe Depressive Mood", corresponding to labels 0 to 4 respectively.

[0085] Then, data augmentation and balancing are performed: the number of samples in each category is counted, and data augmentation is performed on categories with fewer samples. Augmentation methods include adding random noise or time offsets to balance the distribution of the training set.

[0086] Finally, batch storage is performed: the feature sequences and corresponding labels of all users are stored as NumPy array files for subsequent model training.

[0087] (2) Model architecture and data processing flow

[0088] For an input batch of samples x, each sample contains an audio feature sequence and a behavioral feature vector.

[0089] Audio temporal modeling: Layer normalization is applied to the input audio feature sequence to stabilize training and accelerate convergence. The normalized sequence is fed into a 4-layer unidirectional GRU network. The GRU is computed at each time step t according to the following formula:

[0090]

[0091]

[0092]

[0093]

[0094] in: The input feature vector at time step t, This is the hidden state from the previous time step. It is the Sigmoid activation function. These are the reset door and the update door, respectively. The result was obtained by resetting the door. The candidate vector for time, For trainable weight matrix, The updated state is determined by the value of the update gate.

[0095] The temporal features output by the GRU are summed along the time dimension to obtain a compact representation of the entire audio sequence. The pooled vector is then input into a fully connected network to obtain abstract audio features. One-dimensional adaptive average pooling is used to adjust the length to 128 dimensions, yielding the final audio feature representation. The behavioral features are directly used as the feature output of the reaction time modality. .

[0096] The processed audio features and behavioral features are concatenated along the feature dimension to obtain a 141-dimensional fused feature vector. Subsequently, H_fused is fed into a single-layer fully connected network, and the final mental state classification probability distribution is output through the Softmax function:

[0097]

[0098] in, This is the weight matrix. This is the bias vector. In this embodiment, during the binary classification task... Five-category task .

[0099] (3) Model training and loss function: total loss of the model It is the sum of the cross-entropy loss of the audio mode and the reaction time mode:

[0100] (1)

[0101] in, Indicates the mode used (audio or response time rt); It is modal The representation vector for audio modal is For reaction time modes ; y represents the weights corresponding to mode m in a fully connected network; y represents the true label of the sample.

[0102] Loss for each mode The loss is calculated using the standard classification cross-entropy loss function:

[0103] (2)

[0104] in, Number of categories (for binary classification) Five-part time ), One-hot encoding of the real label. The model predicts the first Class probability.

[0105] (4) Training process:

[0106] Forward propagation: computation and intermediate predictions for each mode .

[0107] Loss Calculation: Calculate the total loss according to formulas (1) and (2). .

[0108] Backpropagation and optimization: Minimizing using the Adam optimizer Updates include All model parameters, including those in the model.

[0109] Iterative training: Optimize model parameters through multiple iterations of training.

[0110] This architecture utilizes a GRU branch to deeply mine temporal dynamic information in audio and fuses it with behavioral features, as detailed below:

[0111] (1) Audio temporal modeling branch: The generated sequence of comprehensive audio feature vectors (VGGish + VLAD) is input into a gated recurrent unit (GRU) neural network. The role of this GRU network is to learn the dynamic evolution pattern of the emotional state reflected by the user's speech features in continuous tasks. The hidden state of the last time step of the last layer of the GRU is used as a high-level temporal feature representation of the entire audio sequence.

[0112] (2) Behavioral feature splicing and fusion: The normalized user response time feature vector is horizontally spliced ​​with the high-level audio temporal feature representation output by the GRU above in the feature dimension to form a unified feature vector that integrates complex temporal speech patterns and simple static behavioral indicators.

[0113] (3) Classification output: The concatenated fused feature vector is passed to a single-layer fully connected network, and finally the final psychological state classification result is output through a Softmax classifier.

[0114] In this embodiment, for a mental health assessment dataset, the machine learning model of this invention was used to conduct experimental analysis on multimodal data of audio and behavioral data to obtain the model's performance in the abnormal emotion assessment and detection task, including performance on four indicators: Accuracy, Precision, Recall, and F1. The MGRU model (binary classification) only distinguishes between the presence and absence of depressive mood. The five-category label of the MGRU optimized model integrates the results of the three scales mentioned above (sleep, depression, and anxiety), and the indicators are weighted averages. The experimental results are shown in Table 1.

[0115] Table 1 Model Performance Comparison

[0116]

[0117] For intelligent data management platforms, such as Figure 11 , 12 As shown, this platform is a B / S (Browser / Server) architecture web application, and its main functions include:

[0118] (1) User and Data Management: Provides user registration, login and permission management. The experiment administrator has the highest privileges and can view all data and perform model analysis.

[0119] (2) Data visualization: Provide users and individuals with data dashboards to visually display the distribution of their various indicators in the norm population.

[0120] (3) Results feedback: The evaluation results will be safely returned to the platform, and users will be reminded to check them via in-site messages or notifications.

[0121] The VR scene, built using emotional arousal materials based on psychological paradigms, is responsible for collecting data and exporting it to the data platform. The platform can manage and display answer data, audio data, user questionnaire data, etc. The experiment administrator uniformly calls the MGRU model for prediction and evaluation, and the evaluation results are fed back to the participants through the data platform.

[0122] like Figure 13 As shown in the figure, this embodiment generated a Chinese psychological state assessment dataset during the experiment, including data from 103 participants. The data collection scope was a group aged 18-27, including 60 females and 43 males. Among them, there were 53 undergraduate students, 41 graduate students, and 9 doctoral students.

[0123] Each sample was provided with three labels: anxiety level calculated using the SAS Self-Rating Anxiety Scale, depression level calculated using the Beck Depression Rating Scale, Version 2, and sleep index calculated using the Pittsburgh Sleep Quality Index. Response details, including reaction time and outcome, were provided, along with three audio files each for positive, neutral, and negative responses. Statistical analysis showed that female participants performed better in the model's predicted results.

[0124] Figure 14 This is a distribution chart of scale scores. Considering that this invention is an innovative solution addressing the limitations of the scale, we obtained a new comprehensive distribution, which comprehensively considers the Pittsburgh Sleep Scale scores and adjusts the distribution of anxiety and depression levels, such as... Figure 15 , Figure 16 As shown.

[0125] In the Chinese mental health assessment dataset constructed in this embodiment, we observed that among individuals ranging from those without depression to those with moderate depressive symptoms, some acoustic features (such as dimensions 5 and 7 of the MFCC) exhibited a decreasing trend consistent with existing studies (Wang et al. (2019)). However, in sub-healthy individuals, the change in these features was extremely slight (e.g., Figure 17 If left untreated, this could develop into depression or anxiety. This finding has a dual significance: it confirms the feasibility of using acoustic features for emotion assessment, and it also directly reveals the core challenge of insufficient sensitivity when directly applying clinical conclusions to early screening due to weak signals. Therefore, the solution developed in this embodiment, which integrates interactive VR and deep learning models, is precisely designed to accurately capture these early signals that are overlooked by traditional methods.

[0126] Example 2

[0127] This embodiment 2 provides a non-transitory computer-readable storage medium for storing computer instructions. When executed by a processor, the computer instructions implement the mental health assessment method based on VR interaction and multimodal data as described above. The method includes:

[0128] Users to be evaluated interactively complete quiz and recording tasks in a virtual environment;

[0129] Acquire voice signals and interactive behavior data of the user to be evaluated in the interactive task;

[0130] A pre-trained GRU-based deep learning model is used to extract and fuse features from acquired speech signals and interactive behavior data to obtain the user's psychological state assessment results. The GRU-based deep learning model uses a GRU branch to deeply mine the temporal dynamic information in the audio and fuse it with behavioral features.

[0131] Example 3

[0132] This embodiment 3 provides a computer device, including a memory and a processor, wherein the processor and the memory communicate with each other, and the memory stores program instructions that can be executed by the processor. The processor calls the program instructions to execute the mental health assessment method based on VR interaction and multimodal data as described above, the method including:

[0133] Users to be evaluated interactively complete quiz and recording tasks in a virtual environment;

[0134] Acquire voice signals and interactive behavior data of the user to be evaluated in the interactive task;

[0135] A pre-trained GRU-based deep learning model is used to extract and fuse features from acquired speech signals and interactive behavior data to obtain the user's psychological state assessment results. The GRU-based deep learning model uses a GRU branch to deeply mine the temporal dynamic information in the audio and fuse it with behavioral features.

[0136] Example 4

[0137] This embodiment 4 provides an electronic device, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, and the computer program is stored in the memory. When the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions to implement the mental health assessment method based on VR interaction and multimodal data as described above. The method includes:

[0138] Users to be evaluated interactively complete quiz and recording tasks in a virtual environment;

[0139] Acquire voice signals and interactive behavior data of the user to be evaluated in the interactive task;

[0140] A pre-trained GRU-based deep learning model is used to extract and fuse features from acquired speech signals and interactive behavior data to obtain the user's psychological state assessment results. The GRU-based deep learning model uses a GRU branch to deeply mine the temporal dynamic information in the audio and fuse it with behavioral features.

[0141] In summary, the mental health assessment method based on VR interaction and multimodal data described in this embodiment of the invention is as follows: (1) Original scene design conforming to psychological paradigms: The virtual reality scene in this invention, through design conforming to psychological paradigms and combined with multimodal emotional arousal materials (positive, negative, neutral paintings, music, text, etc.), effectively stimulates the user's emotional response to improve the accuracy of emotion recognition. This design provides a reliable experimental environment for emotion detection. (2) Emotion stimulation and data collection scheme based on interactive VR tasks: Protect the design logic of the VR scene itself, that is, through a preset process that includes active interactive tasks (such as the selection, operation, or problem solving of specific goals), to synchronously stimulate the user's emotions and collect multimodal behavioral data. This is different from the passive scene of simply playing videos. (3) Propose an emotion recognition model architecture for multimodal temporal data: Protect the specific model design, that is, use a gated recurrent unit (GRU) network as the core to perform feature-level fusion and temporal modeling of speech signals (extracting features such as MFCC) and user behavior data from VR tasks. This emphasizes the application of GRU in processing such specific data streams. This invention introduces a GRU (Gated Recurrent Unit) neural network model, specifically designed to process multimodal data of users in virtual scenarios, including reaction time and speech. The final MGRU model achieves real-time identification of user emotional states and accurate detection of depressive moods through the analysis of sequence data. (4) A psychological assessment system and data construction method integrating VR acquisition and data analysis: protecting the entire system architecture and data pipeline. This includes a collaborative working architecture of VR terminal devices (used to present scenes and collect data), a data intelligent management platform (used to store, manage data and run analysis models), and a method for constructing a dedicated mental health dataset containing VR interactive behavior data through this system. The platform developed in this invention supports multi-level permission management and data visualization, and can perform emotional state analysis on the collected data through the data analysis module. The data within the platform forms a new Chinese mental health assessment dataset, including audio recordings and behavioral data, providing rich basic data for subsequent emotion analysis and mental health assessment.

[0142] This invention effectively overcomes the experiential problems associated with traditional passive viewing or high-pressure tasks by designing immersive and interactive VR emotion-stimulating scenarios. This design not only enhances user engagement and data authenticity but also, through carefully crafted emotional stimulation materials, reliably elicits assessable psychophysiological responses, thus providing a higher-quality data foundation for subsequent analysis. Addressing the challenge of weak and difficult-to-capture signals in individuals with sub-optimal mental health, this invention employs the MGRU temporal deep learning model. This model can automatically learn potential dynamic change patterns from users' speech and behavioral response sequence data, exhibiting higher sensitivity and accuracy in identifying early, mild abnormal psychological states compared to existing methods that rely on static feature analysis. This invention constructs a complete solution from data collection, management, to analysis. The system utilizes consumer-grade VR devices and non-invasive microphones for data collection and integrates and analyzes the data through a developed data platform, eliminating reliance on specialized equipment and personnel. Experiments show that this approach can effectively collect data with few adverse reactions during a 30-minute experience. Its assessment results are highly correlated with traditional scale scores, demonstrating its great potential as a low-cost, easy-to-promote, and highly effective early mental state screening tool.

[0143] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0144] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0145] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0146] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment, whereby a series of operational steps are performed to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0147] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that, based on the technical solutions disclosed in the present invention, various modifications or variations that can be made by those skilled in the art without creative effort should be included within the scope of protection of the present invention.

Claims

1. A mental health assessment method based on VR interaction and multimodal data, characterized in that, include: Users to be evaluated interactively complete quiz and recording tasks in a virtual environment; Acquire voice signals and interactive behavior data of the user to be evaluated in the interactive task; A pre-trained GRU-based deep learning model is used to extract and fuse features from acquired speech signals and interactive behavior data to obtain the user's psychological state assessment results. The GRU-based deep learning model uses a GRU branch to deeply mine the temporal dynamic information in the audio and fuse it with behavioral features.

2. The mental health assessment method based on VR interaction and multimodal data according to claim 1, characterized in that, The GRU-based deep learning model includes: an audio temporal modeling branch, which inputs the generated comprehensive audio feature vector sequence into a gated recurrent unit neural network to learn the dynamic evolution pattern of the user's emotional state reflected by their speech features in continuous tasks; the hidden state of the last time step of the last layer of the gated recurrent unit neural network is used as a high-level temporal feature representation of the entire audio sequence; a behavioral feature concatenation and fusion network, which horizontally concatenates the normalized user reaction time feature vector with the high-level audio temporal feature representation output by the gated recurrent unit neural network in the feature dimension to form a unified feature vector that integrates complex temporal speech patterns and simple static behavioral indicators; and a classification output network, which passes the concatenated fused feature vector to a single-layer fully connected network, and finally outputs the final psychological state classification result through a Softmax classifier.

3. The mental health assessment method based on VR interaction and multimodal data according to claim 1, characterized in that, For voice data, silent segments and invalid audio shorter than 1 second are automatically removed; the silent portions at the beginning and end of each recording are cut off, and valid voice is retained.

4. The mental health assessment method based on VR interaction and multimodal data according to claim 1, characterized in that, A dual-path parallel architecture is employed to extract complementary features from the audio, including: VGGish embedding features: preprocessed audio is input into a pre-trained VGGish model, outputting a fixed-dimensional embedding feature vector to represent the high-level semantic content of the audio; VLAD features: the audio is converted into a Mel spectrogram using the Librosa library, and then encoded using the NetVLAD model to generate a VLAD feature vector to capture the local statistical characteristics of the audio; feature fusion and serialization: the VGGish feature vector and VLAD feature vector of each audio sample are concatenated to form a comprehensive audio feature; the comprehensive audio feature sequence is aligned with the behavioral data sequence according to the task order to form multimodal temporal data.

5. The mental health assessment method based on VR interaction and multimodal data according to claim 1, characterized in that, The construction of the virtual environment scenario includes: First, building a teaching interface and multiple task scenarios based on psychological paradigms and assessment needs; the entrance music before entering the scenario helps the subjects relax and return to a relatively calm and focused state, better preparing them for the experience; multiple interactive areas are distributed in the scenario, and emotionally stimulating materials selected and labeled by psychology experts are placed in the areas, including paintings and text fragments with three emotional tendencies: positive, neutral, and negative.

6. The mental health assessment method based on VR interaction and multimodal data according to claim 1, characterized in that, User interaction tasks in the virtual scene include: after entering the interaction area in the virtual scene, users complete different types of interaction tasks; the question-and-answer task requires users to observe the painting and answer questions about the details of the painting; the recording task requires users to observe the painting and record an audio recording of a text describing the painting; after each task is completed, the data is packaged immediately and includes the user ID, task ID and timestamp; when all preset tasks are completed, or the user actively chooses to exit, the data collection session ends.

7. A mental health assessment system based on VR interaction and multimodal data, characterized in that, include: The virtual scene interaction module is used by the user to be evaluated to complete the task of answering questions and recording audio in a virtual scene. The acquisition module is used to acquire the voice signals and interactive behavior data of the user to be evaluated in the interactive task. The evaluation module is used to extract and fuse features from the acquired speech signals and interactive behavior data using a pre-trained GRU-based deep learning model to obtain the user's psychological state evaluation results; wherein, the GRU-based deep learning model uses a GRU branch to deeply mine the temporal dynamic information in the audio and fuse it with behavioral features.

8. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, which, when executed by a processor, implement the mental health assessment method based on VR interaction and multimodal data as described in any one of claims 1-6.

9. A computer device, characterized in that, The device includes a memory and a processor, the processor and the memory communicating with each other, the memory storing program instructions executable by the processor, and the processor calling the program instructions to execute the mental health assessment method based on VR interaction and multimodal data as described in any one of claims 1-6.

10. An electronic device, characterized in that, include: The device includes a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to cause the electronic device to execute instructions for implementing the mental health assessment method based on VR interaction and multimodal data as described in any one of claims 1-6.