Teenager psychology detection method and system based on VR and multi-modal data fusion
By synchronously collecting multimodal data in a virtual reality system, generating contribution weight distribution and hierarchical compression processing, using a dynamic topology fusion network for cross-modal fusion, generating psychological state decision-making features, and finally performing psychological risk assessment and dynamic intervention strategy generation, the problem of low efficiency in multimodal feature collaborative analysis and dynamic decision-making in existing VR psychological detection technology is solved, and more efficient multimodal feature perception and psychological risk assessment are achieved.
Patent Information
- Application Number
- CN202510981993.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-16
- Publication Date
- 2025-09-09
AI Technical Summary
Existing VR psychological detection technology has problems in the collaborative analysis of multimodal features and dynamic decision-making, such as single-modal detection is difficult to capture multi-dimensional behavioral characteristics, multimodal fusion has inefficient high-dimensional redundant processing and rigid static weight mechanism, and risk assessment mechanism lacks closed-loop optimization capabilities, resulting in cross-modal correlation breakdown and disconnection from intervention strategies.
Through the virtual reality interactive system, eye movement trajectory streams, EEG signal streams, speech spectrum streams and gesture trajectory streams are synchronously collected. The contribution weight distribution is generated based on the real-time psychological state classification results. Layered compression processing is performed to generate a compressed feature vector set, and cross-modal fusion is performed through a dynamic topology fusion network to generate psychological state decision-making features. Finally, psychological risk assessment and dynamic intervention strategy generation are carried out.
It improves the accuracy of multimodal feature perception, optimizes the efficiency of dynamic decision-making, enhances the adaptability of psychological risk assessment, and solves the problems of cross-modal correlation breakdown and disconnected intervention strategies in existing technologies.
Smart Images

Figure CN120605015A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimodal data analysis and psychological assessment, and in particular to a method and system for adolescent psychological testing based on the fusion of VR and multimodal data. Background Art
[0002] With the deep integration of virtual reality and multimodal sensing technologies, psychological state monitoring is playing an increasingly critical role in adolescent mental health intervention. Accurate psychological risk assessment requires not only capturing subtle physiological and behavioral characteristics but also collaborative analysis and dynamic decision-making across multi-source heterogeneous data. This has become a core technical challenge in modern psychological diagnosis and treatment systems.
[0003] However, related VR psychological detection technologies have the following problems in the collaborative analysis of multimodal features and dynamic decision-making: single-modal detection is difficult to capture multidimensional behavioral characteristics such as anxiety and depression; multimodal fusion has the defects of inefficient high-dimensional redundant processing and rigid static weight mechanism; the risk assessment mechanism lacks closed-loop optimization capabilities, resulting in cross-modal correlation breaks and disconnection from intervention strategies. Summary of the Invention
[0004] Based on this, it is necessary to provide a youth psychology detection method and system based on the fusion of VR and multimodal data to address the above technical problems, so as to achieve the technical effects of improving the accuracy of multimodal feature perception, optimizing dynamic decision-making efficiency, and enhancing the adaptability of risk assessment.
[0005] In the first aspect, the present application provides a method for adolescent psychological testing based on the fusion of VR and multimodal data, the method comprising:
[0006] Through the virtual reality interactive system, eye movement trajectory stream, EEG signal stream, voice spectrum stream and gesture trajectory stream are synchronously collected;
[0007] Based on the real-time mental state classification results, the contribution weight distribution is generated by processing the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream through the dimension importance evaluation model to generate the contribution weight distribution;
[0008] Perform hierarchical compression processing on the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream according to the contribution weight distribution to generate a compressed feature vector set;
[0009] The compressed feature vector set is fused across modalities through a dynamic topology fusion network to generate psychological state decision features.
[0010] Psychological risk assessment is conducted based on the psychological state decision-making characteristics to generate psychological risk level reports and dynamic intervention strategies.
[0011] Furthermore, the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream are hierarchically compressed according to the contribution weight distribution to generate a compressed feature vector set, including:
[0012] Based on the contribution weight distribution, the core dimension data in the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream are losslessly preserved to generate lossless preserved data;
[0013] Based on the contribution weight distribution, the auxiliary dimension data in the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream are processed in the time domain to generate reduced dimension data;
[0014] Based on the contribution weight distribution, feature aggregation processing is performed on the basic dimension data in the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream to generate aggregated data;
[0015] Multi-dimensional integration processing is performed on lossless preserved data, dimensionality reduced data and aggregated data to generate a compressed feature vector set.
[0016] Furthermore, based on the contribution weight distribution, the core dimension data in the eye movement trajectory stream, EEG signal stream, speech spectrum stream, and gesture trajectory stream are losslessly preserved to generate lossless preserved data, including:
[0017] Based on the weight threshold in the contribution weight distribution, the core dimension identification processing is performed on the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream and the gesture trajectory stream to generate a core dimension identification set;
[0018] According to the core dimension identification set, the gaze coordinate sequence dimension in the eye movement trajectory stream is processed with full resolution time series preservation to obtain full resolution time series preservation data;
[0019] Performing waveform integrity preservation processing on the frequency band energy dimension in the EEG signal stream to obtain waveform integrity preservation data;
[0020] Performing harmonic structure preservation processing on the fundamental frequency envelope dimension in the speech spectrum stream to obtain harmonic structure preservation data;
[0021] Perform complete kinematic feature extraction on the joint angular velocity dimension in the gesture trajectory stream to obtain complete kinematic feature data;
[0022] Integrate full-resolution timing preservation data, waveform integrity preservation data, harmonic structure preservation data, and kinematic feature integrity data to generate lossless preservation data.
[0023] Furthermore, based on the contribution weight distribution, the auxiliary dimension data in the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream are subjected to time domain dimensionality reduction processing to generate dimensionality reduction data, including:
[0024] Based on the modal correlation parameters in the contribution weight distribution, key frame extraction processing is performed on the scanning path dimension in the eye movement trajectory stream to obtain key frame extraction data;
[0025] Perform envelope feature-preserving downsampling on the alpha band oscillation dimension in the EEG signal stream to obtain envelope feature-preserving data;
[0026] Perform piecewise linear approximation processing on the dimension of the formant trajectory in the speech spectrum stream to obtain piecewise linear approximation data;
[0027] Perform motion trend preservation sampling processing on the displacement velocity dimension in the gesture trajectory stream to obtain motion trend preservation data;
[0028] The key frame extraction data, envelope feature preservation data, piecewise linear approximation data and motion trend preservation data are processed synchronously in the time domain through a cross-modal time aligner to generate dimensionality reduction data.
[0029] Furthermore, based on the real-time mental state classification results, the contribution weight distribution of the eye movement trajectory stream, EEG signal stream, speech spectrum stream, and gesture trajectory stream is processed through the dimension importance evaluation model to generate the contribution weight distribution, including:
[0030] Based on the real-time mental state classification results, the dominant modality identification process is performed through the state-modality mapping rule engine to generate the dominant modality identifier;
[0031] According to the dominant modality identification, the eye movement trajectory stream is analyzed and processed with time-space feature correlation to generate the eye movement modality weight coefficient;
[0032] Perform frequency band-emotion response matching processing on the EEG signal stream to generate EEG modality weight coefficients;
[0033] Performing rhythm-psychological state correlation calculation on the speech spectrum stream to generate speech modal weight coefficients;
[0034] Perform motion mode-pressure level mapping on the gesture trajectory stream to generate gesture modality weight coefficients;
[0035] The eye movement modality weight coefficient, EEG modality weight coefficient, speech modality weight coefficient and gesture modality weight coefficient are multimodally integrated through a dynamic weighted fusion to generate a contribution weight distribution.
[0036] Furthermore, according to the dominant modality identification, the eye movement trajectory stream is subjected to time-space feature correlation analysis to generate the eye movement modality weight coefficient, including:
[0037] Based on the dominant modality identification, the hotspot area identification and processing of the spatial distribution of gaze points in the eye movement trajectory stream are performed to generate visual attention hotspots;
[0038] Perform pattern segmentation on the time series of saccadic paths in the eye movement trajectory stream to generate saccadic motion rhythms;
[0039] The spatiotemporal correlation mapper is used to couple the visual attention hotspots with the saccadic rhythm to generate the eye movement feature correlation.
[0040] The psychological state response intensity is quantified based on the correlation of eye movement features to generate eye movement modality weight coefficients.
[0041] Furthermore, the compressed feature vector set is cross-modally fused through a dynamic topology fusion network to generate psychological state decision features, including:
[0042] In the dynamic topology fusion network, the modal correlation of the compressed feature vector set is analyzed to generate dynamic topology connection weights;
[0043] Based on dynamic topological connection weights, a cross-modal feature interaction channel is constructed;
[0044] Through the cross-modal feature interaction channel, multiple rounds of iterative feature transfer and aggregation are performed on the compressed feature vector set;
[0045] Extract the fused feature representation after iterative aggregation to generate the mental state decision feature.
[0046] Secondly, this application also provides a youth psychology detection system based on VR and multimodal data fusion, which includes:
[0047] Multimodal synchronous acquisition module, used to synchronously acquire eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream through the virtual reality interaction system;
[0048] The dynamic weight evaluation module is used to generate contribution weight distribution based on the real-time mental state classification results and the dimension importance evaluation model for the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream.
[0049] A hierarchical compression processing module is used to perform hierarchical compression processing on the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream, and the gesture trajectory stream according to the contribution weight distribution to generate a compressed feature vector set;
[0050] A topological dynamic fusion module is used to perform cross-modal fusion processing on the compressed feature vector set through a dynamic topological fusion network to generate psychological state decision features;
[0051] The closed-loop decision output module is used to perform psychological risk assessment based on psychological state decision-making characteristics, generate psychological risk level reports and dynamic intervention strategies.
[0052] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any method in the first aspect of the present application are implemented.
[0053] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any method in the first aspect of the present application when the computer program is executed by a processor.
[0054] This application provides a method and system for adolescent psychological detection based on VR and multimodal data fusion, the method comprising: synchronously collecting eye movement trajectory streams, EEG signal streams, voice spectrum streams and gesture trajectory streams through a virtual reality interactive system; based on the real-time psychological state classification results, performing contribution weight distribution processing on the eye movement trajectory streams, EEG signal streams, voice spectrum streams and gesture trajectory streams through a dimension importance evaluation model to generate a contribution weight distribution; performing hierarchical compression processing on the eye movement trajectory streams, EEG signal streams, voice spectrum streams and gesture trajectory streams according to the contribution weight distribution to generate a compressed feature vector set; performing cross-modal fusion processing on the compressed feature vector set through a dynamic topology fusion network to generate psychological state decision features; performing psychological risk assessment processing based on the psychological state decision features to generate a psychological risk level report and a dynamic intervention strategy, so as to achieve the technical effects of improving the perception accuracy of multimodal features, optimizing the efficiency of dynamic decision-making, and enhancing the adaptability of risk assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0056] Figure 1 This is a flow chart of a method for adolescent psychology detection based on VR and multimodal data fusion in one embodiment of the present invention;
[0057] Figure 2 This is a flowchart of performing temporal-spatial feature correlation analysis on an eye movement trajectory stream according to a dominant modality identifier to generate an eye movement modality weight coefficient in one embodiment of the present invention;
[0058] Figure 3This is a structural diagram of a youth psychology detection system based on VR and multimodal data fusion in one embodiment of the present invention. DETAILED DESCRIPTION
[0059] In order to make the above-mentioned purposes, features and advantages of the present application more clearly understood, the specific implementation methods of the present application are described in detail below with reference to the accompanying drawings. Many specific details are set forth in the following description to fully understand the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar improvements without violating the connotation of the application. Therefore, the present application is not limited to the specific embodiments disclosed below.
[0060] like Figure 1 As shown, this application provides a method for adolescent psychological testing based on the fusion of VR and multimodal data, the method comprising:
[0061] S101: Synchronously collect eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream through the virtual reality interaction system.
[0062] Specifically, an immersive virtual scene is constructed through a virtual reality interactive system to guide the tester to interact with the scene. During this process, the eye tracking unit, EEG signal acquisition unit, voice spectrum capture unit, and gesture trajectory recording unit are started synchronously. Among them, the eye tracking unit captures the tester's eye movements through the built-in sensors of the VR device and generates an eye trajectory stream, which contains information such as the coordinates of the gaze point and the scanning path; the EEG signal acquisition unit uses the EEG cap integrated with the VR glasses to record the EEG signals generated by the tester's cerebral cortex activity and generate an EEG signal stream, including brain wave data of different frequency bands; the voice spectrum capture unit collects the tester's voice information during the interaction process through a microphone and converts it into a voice spectrum stream, which contains features such as fundamental frequency and resonance peak; the gesture trajectory recording unit uses relevant sensors to capture the tester's body movements and generates a gesture trajectory stream, which contains data such as joint movement angle and speed.
[0063] S102: Based on the real-time mental state classification results, the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream, and the gesture trajectory stream are processed through the dimension importance evaluation model to generate a contribution weight distribution to generate a contribution weight distribution.
[0064] Specifically, based on the real-time mental state classification results, the state-modality mapping rule engine identifies the current dominant modality. For the eye movement trajectory stream, the visual attention hotspots generated by the spatial distribution of the fixation points and the saccadic rhythms segmented from the saccadic path time series are analyzed. Through coupled analysis, the correlation between eye movement features is determined and quantified as the eye movement modality weight coefficient.
[0065] For EEG signal streams, the EEG modal weight coefficients are generated by matching frequency bands with emotional responses. For speech spectrum streams, the correlation between rhythm and psychological state is calculated to obtain speech modal weight coefficients. For gesture trajectory streams, movement patterns and stress levels are mapped to determine gesture modal weight coefficients. These four modal weight coefficients are integrated using a dynamic weighted fusion to generate a contribution weight distribution.
[0066] S103: Performing hierarchical compression processing on the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream, and the gesture trajectory stream according to the contribution weight distribution to generate a compressed feature vector set.
[0067] Specifically, according to the contribution weight distribution, the core dimension data in the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream are processed: the core dimensions are identified according to the weight threshold, the full-resolution time series of the eye movement gaze coordinate sequence is retained, the waveform of the EEG frequency band energy is kept intact, the harmonic structure of the fundamental frequency envelope of the speech is retained, and the kinematic features of the joint angular velocity of the gesture are fully extracted, and the data are integrated to generate lossless retained data.
[0068] Reprocess auxiliary dimensional data: Based on modal correlation parameters, extract key frames from the eye movement scanning path, perform envelope feature-preserving downsampling on the EEG Alpha band oscillation, perform piecewise linear approximation on the speech formant trajectory, maintain motion trend sampling on the gesture displacement speed, and generate dimensionality reduction data synchronously in the time domain.
[0069] After that, the basic dimensional data is processed and aggregated data is generated through feature aggregation. The lossless data, dimensionality-reduced data, and aggregated data are multi-dimensionally integrated to generate a compressed feature vector set.
[0070] S104: Perform cross-modal fusion processing on the compressed feature vector set through a dynamic topology fusion network to generate psychological state decision features.
[0071] Specifically, in a dynamic topological fusion network, the inherent correlations among the various modal data—eye movement, EEG, speech, and gesture—in the compressed feature vector set are analyzed, and dynamic topological connection weights are generated based on the strength of the correlations. Based on these weights, a cross-modal feature interaction channel is constructed, enabling features from different modalities to be iteratively transferred and aggregated within the channel over multiple rounds. For example, emotion-related frequency band features in EEG signals are cross-mapped and information complemented with visual attention features in eye movement trajectories, emotional rhythmic features in speech spectra, and motor tension features in gesture trajectories. Through multiple rounds of iterative optimization, a fused feature representation that integrates key multimodal information is extracted, generating psychological state decision-making features.
[0072] S105: Conduct psychological risk assessment based on the psychological state decision-making characteristics to generate a psychological risk level report and dynamic intervention strategy.
[0073] Specifically, based on the psychological state decision-making characteristics and pre-set psychological risk assessment criteria, a comprehensive analysis of multimodal fusion information, including EEG, eye movement, speech, and gestures, is conducted to determine the individual's risk level for depression, anxiety, phobia, and other conditions. Based on the assessed risk level, a corresponding psychological risk level report is generated to clarify the risk category and severity. Furthermore, based on the risk level and the specific psychological state reflected in the decision-making characteristics, a pre-set library of intervention scenarios and strategy templates is invoked. For example, dynamic intervention strategies are generated by matching anxiety-prone individuals with relaxing virtual scenarios and interactive activities.
[0074] An embodiment of the present application provides a method for adolescent psychological detection based on VR and multimodal data fusion, including: synchronously collecting eye movement trajectory streams, EEG signal streams, voice spectrum streams and gesture trajectory streams through a virtual reality interactive system; based on the real-time psychological state classification results, performing contribution weight distribution generation processing on the eye movement trajectory streams, EEG signal streams, voice spectrum streams and gesture trajectory streams through a dimension importance evaluation model to generate a contribution weight distribution; performing hierarchical compression processing on the eye movement trajectory streams, EEG signal streams, voice spectrum streams and gesture trajectory streams according to the contribution weight distribution to generate a compressed feature vector set; performing cross-modal fusion processing on the compressed feature vector set through a dynamic topology fusion network to generate psychological state decision features; performing psychological risk assessment processing based on the psychological state decision features to generate a psychological risk level report and a dynamic intervention strategy, so as to achieve the technical effects of improving the perception accuracy of multimodal features, optimizing the efficiency of dynamic decision-making, and enhancing the adaptability of risk assessment.
[0075] Furthermore, the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream are hierarchically compressed according to the contribution weight distribution to generate a compressed feature vector set, including:
[0076] Based on the contribution weight distribution, the core dimension data in the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream are losslessly preserved to generate lossless preserved data;
[0077] Based on the contribution weight distribution, the auxiliary dimension data in the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream are processed in the time domain to generate reduced dimension data;
[0078] Based on the contribution weight distribution, feature aggregation processing is performed on the basic dimension data in the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream to generate aggregated data;
[0079] Multi-dimensional integration processing is performed on lossless preserved data, dimensionality reduced data and aggregated data to generate a compressed feature vector set.
[0080] Specifically, the core dimension data is losslessly preserved: the core dimensions of each modality are locked according to the weight threshold, and the gaze coordinate sequence reflecting visual focus in the eye movement trajectory stream is recorded in full resolution time series to completely preserve its dynamic changes; the waveform integrity of the frequency band energy dimension closely related to emotions in the EEG signal stream is maintained to ensure that the original characteristics of the signal are not lost; the harmonic structure of the fundamental frequency envelope dimension reflecting the intonation characteristics in the speech spectrum stream is retained to maintain the integrity of its acoustic characteristics; the kinematic features of the joint angular velocity dimension reflecting the strength of the movement in the gesture trajectory stream are comprehensively extracted to fully capture the details of the limb movement, and the above data are integrated to generate losslessly preserved data.
[0081] Time domain dimensionality reduction processing is performed on auxiliary dimension data: based on modal correlation parameters, key frames are extracted from the scanning path dimension reflecting the movement of the gaze in the eye movement trajectory stream to retain its main motion trajectory; the envelope feature-preserving downsampling is performed on the specific band oscillation dimension related to the alert state in the EEG signal stream, compressing the data volume while retaining the core fluctuation characteristics; piecewise linear approximation is performed on the resonance peak trajectory dimension reflecting the sound resonance in the speech spectrum stream to simplify the data while maintaining the overall change trend; the motion trend-preserving sampling is performed on the displacement velocity dimension reflecting the displacement characteristics in the gesture trajectory stream to retain the main motion posture, and then the above data are synchronously calibrated in the time dimension through a cross-modal time aligner to generate reduced dimensionality data.
[0082] We perform feature aggregation on basic dimensional data: We statistically integrate and merge low-weight but still valuable basic dimensional information from each modality, such as eye movement blink frequency, EEG background noise, voice volume changes, and gesture amplitude, to generate aggregated data. Finally, we perform a multi-dimensional structured integration of the lossless data, dimensionality-reduced data, and aggregated data to generate a compressed feature vector set that combines key information integrity with data simplicity.
[0083] Furthermore, based on the contribution weight distribution, the core dimension data in the eye movement trajectory stream, EEG signal stream, speech spectrum stream, and gesture trajectory stream are losslessly preserved to generate lossless preserved data, including:
[0084] Based on the weight threshold in the contribution weight distribution, the core dimension identification processing is performed on the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream and the gesture trajectory stream to generate a core dimension identification set;
[0085] According to the core dimension identification set, the gaze coordinate sequence dimension in the eye movement trajectory stream is processed with full resolution time series preservation to obtain full resolution time series preservation data;
[0086] Performing waveform integrity preservation processing on the frequency band energy dimension in the EEG signal stream to obtain waveform integrity preservation data;
[0087] Performing harmonic structure preservation processing on the fundamental frequency envelope dimension in the speech spectrum stream to obtain harmonic structure preservation data;
[0088] Perform complete kinematic feature extraction on the joint angular velocity dimension in the gesture trajectory stream to obtain complete kinematic feature data;
[0089] Integrate full-resolution timing preservation data, waveform integrity preservation data, harmonic structure preservation data, and kinematic feature integrity data to generate lossless preservation data.
[0090] Specifically, based on the weight threshold in the contribution weight distribution, the core dimensions of the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream are identified: the dimensions with the highest correlation with the psychological state in each modality are screened out, and the gaze coordinate sequence of the eye movement trajectory stream, the frequency band energy of the EEG signal stream, the fundamental frequency envelope of the speech spectrum stream, and the joint angular velocity of the gesture trajectory stream are identified as core dimensions to generate a core dimension identification set.
[0091] Based on the above-mentioned set of identifiers, the gaze coordinate sequence dimension of the eye movement trajectory stream is preserved in full-resolution time series to completely record its coordinate information that changes over time in the virtual scene, and obtain full-resolution time series preserved data; the frequency band energy dimension of the EEG signal stream is maintained in waveform integrity to ensure that the original waveform characteristics of different emotion-related frequency bands are not destroyed, and obtain waveform integrity preserved data.
[0092] The fundamental frequency envelope dimension of the speech spectrum stream is harmonically preserved, maintaining the harmonic characteristics of the voice that reflect emotion, resulting in harmonically preserved data. Kinematic features are fully extracted from the joint angular velocity dimension of the gesture trajectory stream, comprehensively capturing the velocity variations of limb joint rotations and generating kinematically complete data. Finally, the full-resolution time-series data, waveform integrity-preserved data, harmonically preserved data, and kinematically complete data are integrated to generate lossless preserved data.
[0093] Furthermore, based on the contribution weight distribution, the auxiliary dimension data in the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream are subjected to time domain dimensionality reduction processing to generate dimensionality reduction data, including:
[0094] Based on the modal correlation parameters in the contribution weight distribution, key frame extraction processing is performed on the scanning path dimension in the eye movement trajectory stream to obtain key frame extraction data;
[0095] Perform envelope feature-preserving downsampling on the alpha band oscillation dimension in the EEG signal stream to obtain envelope feature-preserving data;
[0096] Perform piecewise linear approximation processing on the dimension of the formant trajectory in the speech spectrum stream to obtain piecewise linear approximation data;
[0097] Perform motion trend preservation sampling processing on the displacement velocity dimension in the gesture trajectory stream to obtain motion trend preservation data;
[0098] The key frame extraction data, envelope feature preservation data, piecewise linear approximation data and motion trend preservation data are processed synchronously in the time domain through a cross-modal time aligner to generate dimensionality reduction data.
[0099] Specifically, based on the modal association parameters in the contribution weight distribution, key frame extraction processing is performed on the scanning path dimension in the eye movement trajectory stream: frame information that can reflect the key turning points of rapid gaze movement is screened out to retain the core nodes of its path changes and obtain key frame extraction data.
[0100] The oscillation dimension of the Alpha band in the EEG signal stream is subjected to envelope feature-preserving downsampling processing: while reducing the data sampling density, the waveform contour characteristics of this band related to the relaxation state are maintained to obtain envelope feature-preserving data.
[0101] The dimension of the formant trajectory in the speech spectrum stream is piecewise linearly approximated: the trajectory reflecting the change of sound resonance is divided into several segments according to the trend, and the overall direction of each segment is approximated by a straight line to obtain piecewise linear approximation data.
[0102] The displacement velocity dimension in the gesture trajectory stream is sampled to maintain the motion trend: representative sample points that can reflect the increase or decrease trend of limb movement speed are selected, and their dynamic change characteristics are retained to obtain motion trend maintenance data.
[0103] Through the cross-modal time aligner, the key frame extraction data, envelope feature preservation data, piecewise linear approximation data and motion trend preservation data are matched on the time axis to ensure that the time nodes of different modal data correspond to each other and generate dimensionality reduction data.
[0104] Furthermore, based on the real-time mental state classification results, the contribution weight distribution of the eye movement trajectory stream, EEG signal stream, speech spectrum stream, and gesture trajectory stream is processed through the dimension importance evaluation model to generate the contribution weight distribution, including:
[0105] Based on the real-time mental state classification results, the dominant modality identification process is performed through the state-modality mapping rule engine to generate the dominant modality identifier;
[0106] According to the dominant modality identification, the eye movement trajectory stream is analyzed and processed with time-space feature correlation to generate the eye movement modality weight coefficient;
[0107] Perform frequency band-emotion response matching processing on the EEG signal stream to generate EEG modality weight coefficients;
[0108] Performing rhythm-psychological state correlation calculation on the speech spectrum stream to generate speech modal weight coefficients;
[0109] Perform motion mode-pressure level mapping on the gesture trajectory stream to generate gesture modality weight coefficients;
[0110] The eye movement modality weight coefficient, EEG modality weight coefficient, speech modality weight coefficient and gesture modality weight coefficient are multimodally integrated through a dynamic weighted fusion to generate a contribution weight distribution.
[0111] Specifically, based on the real-time mental state classification results, the state-modality mapping rule engine identifies the dominant modality that most significantly influences the current mental state, generating a dominant modality identifier. Based on this identifier, the eye movement trajectory stream undergoes temporal-spatial feature correlation analysis: identifying visual attention hotspots formed by the spatial distribution of fixations, segmenting the time series of saccade paths to obtain saccade rhythms, and coupling analysis to analyze the correlation between the two, quantifying them as eye movement modality weight coefficients.
[0112] Frequency-band-emotional response matching is performed on the EEG signal stream: EEG waves of different frequency bands are compared with corresponding emotional response patterns to determine the weight of each frequency band's influence on the current psychological state and generate EEG modal weight coefficients. Prosody-psychological state correlation is calculated on the speech spectrum stream: The correlation between speech prosodic features such as pitch and rhythm and psychological state is analyzed and converted into speech modal weight coefficients.
[0113] Motion mode-pressure level mapping is performed on the gesture trajectory stream: the amplitude, speed, and other patterns of limb movement are mapped to pressure levels to obtain gesture modal weight coefficients. A dynamic weighted fuser adjusts the fusion weights of each coefficient based on the dominant modality identifier, integrating the weight coefficients of the four modalities (eye movement, EEG, speech, and gesture) to generate a contribution weight distribution.
[0114] like Figure 2 As shown in the figure, according to the dominant modality identification, the eye movement trajectory stream is subjected to time-space feature correlation analysis to generate the eye movement modality weight coefficient, including:
[0115] S201: Based on the dominant modality identifier, perform hotspot area identification processing on the spatial distribution of gaze points in the eye movement trajectory stream to generate visual attention hotspots;
[0116] S202: performing pattern segmentation processing on the saccadic path time series in the eye movement trajectory stream to generate saccadic motion rhythm;
[0117] S203: performing coupling analysis and processing on the visual attention hotspot and the saccadic movement rhythm through a spatiotemporal correlation mapper to generate an eye movement feature correlation degree;
[0118] S204: Quantify the intensity of the psychological state response based on the eye movement feature correlation to generate an eye movement modality weight coefficient.
[0119] Specifically, based on the dominant modality identifier, the spatial distribution of fixations in the eye movement stream is processed for hotspot identification. By analyzing the density and duration of the subject's fixations in the virtual scene, the core areas of visual attention are located and visual hotspots are generated. Pattern segmentation is performed on the time series of saccade paths in the eye movement stream. Based on characteristics such as saccade speed, direction, and interval, continuous saccade paths are divided into segments with similar motion characteristics to generate saccade motion rhythms.
[0120] The spatiotemporal correlation mapper is used to couple the visual attention hotspots with the saccade rhythm. The correlation between the position and range of the hotspots and the start timing and movement trajectory of the saccade segments is analyzed to clarify the degree of matching between the two in time and space dimensions, and generate the eye movement feature correlation.
[0121] The intensity of psychological state response is quantified based on the correlation of eye movement features: combined with the psychological state corresponding to the dominant mode (such as visual vigilance in anxiety and distraction in depression), the correlation is converted into a numerical value reflecting the degree of influence of the eye movement mode on the current psychological state, generating the eye movement mode weight coefficient.
[0122] Furthermore, the compressed feature vector set is cross-modally fused through a dynamic topology fusion network to generate psychological state decision features, including:
[0123] In the dynamic topology fusion network, the modal correlation of the compressed feature vector set is analyzed to generate dynamic topology connection weights;
[0124] Based on dynamic topological connection weights, a cross-modal feature interaction channel is constructed;
[0125] Through the cross-modal feature interaction channel, multiple rounds of iterative feature transfer and aggregation are performed on the compressed feature vector set;
[0126] Extract the fused feature representation after iterative aggregation to generate the mental state decision feature.
[0127] Specifically, within the dynamic topological fusion network, the inherent connections between the various modal data sets—eye movement, EEG, speech, and gesture—are analyzed within the compressed feature vector set. The network analyzes the interactions and influence of different modal features in reflecting psychological states, such as the strength of the correlation between emotion-related EEG frequency bands and eye movement visual attention areas, and the emotional rhythm of speech and the motor intensity of gestures. Dynamic topological connection weights are generated based on these weights. Cross-modal feature interaction channels are constructed based on these weights, enabling targeted information transfer between modal features based on the closeness of their connections.
[0128] Through a cross-modal feature interaction channel, the compressed feature vector set undergoes multiple rounds of iterative feature transfer and aggregation. In each iteration, the key features of one modality are transferred to other closely related modalities, fused with the corresponding modal features, and then fed back to the original modality. After multiple rounds of optimization, the features of different modalities are deeply complementary. Finally, the fused feature representation formed after iterative aggregation, which integrates the core information of multiple modalities, is extracted to generate the psychological state decision feature.
[0129] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0130] In one embodiment, if Figure 3 As shown, the present application also provides a youth psychology detection system 300 based on VR and multimodal data fusion, which includes:
[0131] A multimodal synchronous acquisition module 301 is used to synchronously acquire eye movement trajectory streams, EEG signal streams, speech spectrum streams, and gesture trajectory streams through a virtual reality interaction system;
[0132] Dynamic weight evaluation module 302 is used to generate contribution weight distribution based on the real-time mental state classification results by using the dimension importance evaluation model to generate contribution weight distribution for the eye movement trajectory stream, EEG signal stream, speech spectrum stream, and gesture trajectory stream;
[0133] A layered compression processing module 303 is used to perform layered compression processing on the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream, and the gesture trajectory stream according to the contribution weight distribution to generate a compressed feature vector set;
[0134] A topological dynamic fusion module 304 is configured to perform cross-modal fusion processing on the compressed feature vector set through a dynamic topological fusion network to generate a psychological state decision feature;
[0135] The closed-loop decision output module 305 is used to perform psychological risk assessment based on the psychological state decision-making characteristics and generate a psychological risk level report and dynamic intervention strategy.
[0136] Specifically, the multimodal synchronous acquisition module 301 constructs an immersive scene through a virtual reality interactive system, guides the tester to interact, and simultaneously performs eye tracking, EEG acquisition, voice capture and gesture recording functions to obtain eye movement trajectory stream, EEG signal stream, voice spectrum stream and gesture trajectory stream respectively.
[0137] The dynamic weight evaluation module 302 uses the dimension importance evaluation model to identify the dominant mode based on the real-time psychological state classification results, analyzes the correlation between each mode and the current psychological state, generates modal weight coefficients for eye movement, EEG, speech, and gestures, and obtains the contribution weight distribution through integration.
[0138] The layered compression processing module 303 performs lossless preservation of the core dimension data, performs time domain dimensionality reduction on the auxiliary dimension data, performs feature aggregation on the basic dimension data, and then integrates the above three types of data to generate a compressed feature vector set according to the contribution weight distribution.
[0139] The topological dynamic fusion module 304 analyzes the modal association of the compressed feature vector set through the dynamic topological fusion network, generates dynamic topological connection weights, constructs a cross-modal interaction channel, extracts the fusion feature representation through multiple rounds of iterative transmission and aggregation, and generates the psychological state decision feature.
[0140] The closed-loop decision output module 305 determines the risk level based on the psychological state decision characteristics and refers to the psychological risk assessment standards, generates a psychological risk level report, matches targeted virtual intervention scenarios, and generates a dynamic intervention strategy.
[0141] The layered compression processing module 303 is further configured to:
[0142] Based on the contribution weight distribution, the core dimension data in the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream are losslessly preserved to generate lossless preserved data;
[0143] Based on the contribution weight distribution, the auxiliary dimension data in the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream are processed in the time domain to generate reduced dimension data;
[0144] Based on the contribution weight distribution, feature aggregation processing is performed on the basic dimension data in the eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream to generate aggregated data;
[0145] Multi-dimensional integration processing is performed on lossless preserved data, dimensionality reduced data and aggregated data to generate a compressed feature vector set.
[0146] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any method in the first aspect of the present application are implemented.
[0147] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps of any method in the first aspect of the present application when the computer program is executed by a processor.
[0148] In one embodiment, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0149] In one embodiment, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the above-mentioned method embodiments when the computer program is executed by a processor.
[0150] In one embodiment, in order to achieve more accurate detection of EEG and eye movement signals, the overall design of the device fully considers the user's wearing experience, and innovatively designs the device. The EEG cap and VR glasses are deeply integrated. Through precise structural design, the electrode distribution and wearing arc of the EEG cap meet ergonomic requirements. At the same time, an eye movement signal acquisition module is integrated in the VR glasses to ensure that the two cooperate with each other in spatial layout and do not interfere with each other.
[0151] In terms of virtual reality interaction design and psychological data analysis systems, we have initially completed virtual scene modeling, the construction of a large multimodal data analysis model, and the establishment of auxiliary treatment scenarios. The entire equipment set can be used for preliminary diagnosis of psychological disorders such as depression, bipolar disorder, and phobias.
[0152] ① Virtual reality of psychological scales:
[0153] (1) Environmental stress mapping: Using multimodal stimuli such as visual and auditory stimulation in virtual scenes, we simulate anxiety-triggering situations in reality and activate the subject’s autonomic nervous system response.
[0154] (2) Emotional linking: Using task designs such as time pressure and multitasking to induce cognitive patterns such as catastrophic thinking and excessive worrying that are related to anxiety.
[0155] (3) Implicit assessment: The purpose of the test is concealed through gamified narratives, which reduces the user’s psychological defenses. For example, in response to the question “Do you often feel nervous, anxious, or ‘restless’?” in the GAB scale, users are asked to play the role of emergency nurses, dealing with multiple patients who are rushing in at the same time: first, cases are dynamically generated, and the severity of the patient’s symptoms is randomly selected as mild or critical; second, a 30-second countdown is given, and 10 patients need to be distinguished, allowing users to experience anxiety in a specific scenario.
[0156] Scenario-based theory of psychological scales:
[0157] (1) Task mechanism: transform psychological symptoms into interactive challenges, use the Unity physics engine to simulate pressure feedback, and capture micro-manipulation capabilities through the PICO handle.
[0158] (2) Environmental metaphor: Using visual symbols to imply psychological states, such as using the wild growth of vines to map the accumulation of anxiety, using Shader Graph to dynamically adjust the plant morphology, and the environmental color temperature changes with speed parameters such as event triggering frequency.
[0159] (3) Dynamic adjustment: Adjust the scene difficulty based on real-time data and dynamically control the event triggering frequency through algorithms.
[0160] (4) Temporal dimension: Assess the persistence of symptoms through gradual changes such as scene fading and task cycle accumulation.
[0161] ②Bioelectric signal detection technology:
[0162] The EEG software system architecture, timing alignment, and data fusion process are as follows:
[0163] 1. Full-cycle baseline records
[0164] The global clock starts, and the master clock is automatically activated (accuracy ±0.1ms) the moment the user begins detection. The three sensors of EEG, eye movement, and VR behavior are simultaneously turned on for continuous recording; environmental baseline collection is performed, and resting physiological data is recorded in the first 30 seconds to establish a personalized benchmark. At the same time, interference parameters such as environmental noise and light intensity are simultaneously collected.
[0165] 2. Task Node Enhancement
[0166] When a VR scene triggers preset tasks such as timed puzzle solving and frightening events, a three-level tagging system is generated, including tag level, trigger conditions, and processing priority. Red tags correspond to core assessment tasks such as emotional stimulation, and the data cache must be frozen immediately; yellow tags increase the sampling rate by 50% for auxiliary observation tasks such as attention tests; green tags are associated with common events such as scene switching, and standard recording is performed. In addition, the enhanced mode will be activated during the red tag period, with the EEG sampling rate increased from 256Hz to 512Hz, the eye tracking frequency increased from 60Hz to 120Hz, and an independent storage channel will be allocated to prevent data overwriting.
[0167] 3. Layered Data Fusion
[0168] Prioritizing task node alignment, we constructed task time windows from 5 seconds before to 10 seconds after an event. Raw data was first determined to identify task nodes. If so, μs-level interpolation alignment was performed; otherwise, ms-level sliding window alignment was used. Using different alignment strategies, we compared the conventional accuracy of data types like EEG signals, eye movement coordinates, and behavioral events with the accuracy of task nodes, promoting fusion through incremental data synthesis. The fusion process was divided into three rounds. The first round processed only red-marked data to generate a core event analysis report. The second round incorporated yellow-marked data to refine the "Behavioral Pattern Evolution Map." The final round integrated all data cycles to generate the "Comprehensive Mental State Timeline."
[0169] 4. Abnormal Change Detection
[0170] It includes dual baseline comparison (comparison of task node data with the initial resting baseline, and historical data of similar tasks with current performance) and intelligent difference identification (preset threshold alarms, dynamic pattern recognition, and cross-modal contradiction detection); the technical advantages are reflected in resource optimization, with a small proportion of key data but more effective information retained, improved computing efficiency, enhanced accuracy (small time error of key events), and enhanced interpretability (generation of a fused timeline with task tags).
[0171] Attention is a cognitive function that involves the entire brain. In-depth neuroimaging research has revealed that attention-related functions are mediated by a network of specific neural regions. Based on brain network imaging studies, attention involves three core functions: alertness, which concerns achieving and maintaining a state of alertness; orienting, which focuses sensory resources on specific stimuli; and execution, which involves resolving conflicts between neural systems.
[0172] While the cognitive process of attention has been shown to involve a wide range of brain regions, including the primary sensory cortex responsible for basic information perception, the limbic system closely linked to emotion and memory, and the motor cortex that controls physical movement, the neural mechanisms that activate these areas actually originate from a relatively small set of neural networks involving three cortical regions of the attention network. These include the alerting network, which encompasses thalamic and cortical sites associated with the brain's norepinephrine system; the orienting network, centered in the parietal lobe; and the executive network, which encompasses the anterior cingulate gyrus and other frontal lobe regions.
[0173] In terms of EEG evoked potentials, specific pictures are used to induce EEG during the virtual scene test. For example, the International Affective Picture System (IAPS) is used as a source of emotional stimulation, and a series of emotional and neutral images are carefully selected. A total of 72 pictures of four types were screened from the IAPS database, including 36 neutral pictures and 36 pictures with obvious emotional expressions, with 12 each of sadness, threat, and positivity. The standardized emotional stimuli provided by IAPS have been widely used in emotional research on psychopathologies such as depression, anxiety, and bipolar disorder. In the virtual scene, one of the emotional types of pictures, sadness, threat, neutrality, and positivity, is displayed. Each trial starts with a 1000-ms fixation cross, followed by a 6-second emotional or neutral picture. Each participant views a total of 9 blocks of pictures in a pseudo-random order, with each block containing four emotional pictures and four neutral pictures. To ensure consistency and standardization of the experiment, all images used maintained the same size and resolution of 1024*768 pixels. By comparing the waveforms in different scenarios, it was determined whether the subjects had a mental illness.
[0174] In terms of eye movement technology, by installing eye movement sensors and infrared sensors in VR glasses, the heat map of the subjects' responses to international emotional pictures was detected. It was found that patients with depression, anxiety, and bipolar disorder had different levels of attention to different pictures, which can be used to determine whether there are mental health problems.
[0175] The auxiliary therapy technology has designed different relaxing scenes such as forests and cherry blossoms, and set up interactions in the scenes. Users can perform interactive operations such as rowing and meditation, while playing relaxing music to soothe their emotions.
[0176] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separated, and the parts displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the disclosed solution. A person of ordinary skill in the art can understand and implement it without expending creative work.
[0177] The above-described embodiments merely represent several implementation methods of the embodiments of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the concept of the embodiments of the present application, and these modifications and improvements fall within the scope of protection of the embodiments of the present application.
Claims
1. A youth psychology detection method based on VR and multimodal data fusion, characterized by: The method comprises: Through the virtual reality interactive system, eye movement trajectory stream, EEG signal stream, voice spectrum stream and gesture trajectory stream are synchronously collected; Based on the real-time mental state classification result, the contribution weight distribution generation process is performed on the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream, and the gesture trajectory stream using a dimension importance evaluation model to generate a contribution weight distribution; Performing hierarchical compression processing on the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream, and the gesture trajectory stream according to the contribution weight distribution to generate a compressed feature vector set; Performing cross-modal fusion processing on the compressed feature vector set through a dynamic topology fusion network to generate a psychological state decision feature; A psychological risk assessment is performed based on the psychological state decision-making characteristics to generate a psychological risk level report and a dynamic intervention strategy.
2. The adolescent psychology detection method based on VR and multimodal data fusion according to claim 1 is characterized in that: The step of performing hierarchical compression processing on the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream, and the gesture trajectory stream according to the contribution weight distribution to generate a compressed feature vector set includes: Based on the contribution weight distribution, losslessly retaining the core dimension data in the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream, and the gesture trajectory stream is processed to generate losslessly retained data; Based on the contribution weight distribution, performing time-domain dimensionality reduction processing on the auxiliary dimension data in the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream, and the gesture trajectory stream to generate reduced-dimensionality data; Based on the contribution weight distribution, feature aggregation processing is performed on the basic dimension data in the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream, and the gesture trajectory stream to generate aggregated data; Multi-dimensional integration processing is performed on the lossless retained data, the dimensionality reduced data, and the aggregated data to generate the compressed feature vector set.
3. The adolescent psychology detection method based on VR and multimodal data fusion according to claim 2 is characterized in that: The step of performing lossless preservation processing on the core dimension data in the eye movement trajectory stream, the EEG signal stream, the voice spectrum stream, and the gesture trajectory stream based on the contribution weight distribution to generate lossless preserved data includes: Based on a weight threshold in the contribution weight distribution, performing core dimension recognition processing on the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream, and the gesture trajectory stream to generate a core dimension identification set; According to the core dimension identifier set, performing full-resolution time-series preservation processing on the gaze coordinate sequence dimension in the eye movement trajectory stream to obtain full-resolution time-series preservation data; Performing waveform integrity preservation processing on the frequency band energy dimension in the EEG signal stream to obtain waveform integrity preservation data; Performing harmonic structure preservation processing on the fundamental frequency envelope dimension in the speech spectrum stream to obtain harmonic structure preservation data; Performing complete kinematic feature extraction processing on the joint angular velocity dimension in the gesture trajectory stream to obtain complete kinematic feature data; The full-resolution time-series preservation data, the waveform integrity preservation data, the harmonic structure preservation data, and the kinematic feature integrity data are integrated to generate the lossless preservation data.
4. The adolescent psychology detection method based on VR and multimodal data fusion according to claim 2 is characterized in that: The step of performing time-domain dimensionality reduction processing on the auxiliary dimension data in the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream, and the gesture trajectory stream based on the contribution weight distribution to generate reduced-dimensionality data includes: Based on the modal association parameters in the contribution weight distribution, performing key frame extraction processing on the scanning path dimension in the eye movement trajectory stream to obtain key frame extraction data; Performing envelope feature-preserving downsampling processing on the Alpha band oscillation dimension in the EEG signal stream to obtain envelope feature-preserving data; Performing piecewise linear approximation processing on the formant trajectory dimension in the speech spectrum stream to obtain piecewise linear approximation data; Performing motion trend preservation sampling processing on the displacement velocity dimension in the gesture trajectory stream to obtain motion trend preservation data; The key frame extraction data, the envelope feature preservation data, the piecewise linear approximation data and the motion trend preservation data are subjected to time domain synchronization processing by a cross-modal time aligner to generate the dimensionality reduction data.
5. The adolescent psychology detection method based on VR and multimodal data fusion according to claim 1 is characterized in that: The step of performing contribution weight distribution generation processing on the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream, and the gesture trajectory stream based on the real-time mental state classification result through a dimension importance evaluation model to generate a contribution weight distribution includes: Based on the real-time mental state classification result, a dominant modality identification process is performed through a state-modality mapping rule engine to generate a dominant modality identifier; According to the dominant modality identifier, performing time-space feature correlation analysis on the eye movement trajectory stream to generate an eye movement modality weight coefficient; Performing frequency band-emotion response matching processing on the EEG signal stream to generate an EEG modality weight coefficient; Performing prosody-mental state correlation calculation on the speech spectrum stream to generate a speech modality weight coefficient; Performing motion mode-pressure level mapping processing on the gesture trajectory stream to generate a gesture modality weight coefficient; The eye movement modality weight coefficient, the EEG modality weight coefficient, the speech modality weight coefficient and the gesture modality weight coefficient are subjected to multimodal integration processing by a dynamic weighted fusion device to generate the contribution weight distribution.
6. The adolescent psychology detection method based on VR and multimodal data fusion according to claim 5 is characterized in that: The step of performing time-space feature correlation analysis on the eye movement trajectory stream according to the dominant modality identifier to generate an eye movement modality weight coefficient includes: Based on the dominant modality identifier, performing hotspot area identification processing on the spatial distribution of gaze points in the eye movement trajectory stream to generate visual attention hotspots; performing pattern segmentation processing on the saccadic path time series in the eye movement trajectory stream to generate saccadic motion rhythm; Performing coupling analysis and processing on the visual attention hotspot and the saccadic motion rhythm by a spatiotemporal correlation mapper to generate an eye movement feature correlation; The psychological state response intensity is quantified based on the eye movement feature correlation to generate the eye movement modality weight coefficient.
7. The adolescent psychology detection method based on VR and multimodal data fusion according to claim 1 is characterized in that: The cross-modal fusion processing of the compressed feature vector set by the dynamic topology fusion network to generate the mental state decision feature includes: In the dynamic topology fusion network, analyzing the modal relevance of the compressed feature vector set to generate dynamic topology connection weights; Building a cross-modal feature interaction channel based on the dynamic topological connection weights; Performing multiple rounds of iterative feature transfer and aggregation on the compressed feature vector set through the cross-modal feature interaction channel; Extract the fused feature representation after iterative aggregation to generate the mental state decision feature.
8. A youth psychology detection system based on VR and multimodal data fusion is characterized by: The system comprises: Multimodal synchronous acquisition module, used to synchronously acquire eye movement trajectory stream, EEG signal stream, speech spectrum stream and gesture trajectory stream through the virtual reality interaction system; a dynamic weight evaluation module, configured to generate a contribution weight distribution based on the real-time mental state classification result and the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream, and the gesture trajectory stream using a dimension importance evaluation model; a hierarchical compression processing module, configured to perform hierarchical compression processing on the eye movement trajectory stream, the EEG signal stream, the speech spectrum stream, and the gesture trajectory stream according to the contribution weight distribution to generate a compressed feature vector set; A topological dynamic fusion module, configured to perform cross-modal fusion processing on the compressed feature vector set through a dynamic topological fusion network to generate a psychological state decision feature; The closed-loop decision output module is used to perform psychological risk assessment based on the psychological state decision characteristics, and generate a psychological risk level report and dynamic intervention strategy.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the adolescent psychology detection method based on VR and multimodal data fusion described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the adolescent psychology detection method based on VR and multimodal data fusion according to any one of claims 1 to 7 are implemented.