An emotion-guided virtual reality motion sickness intervention system

CN122557902APending Publication Date: 2026-08-14SHANDONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-19
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

现阶段针对虚拟现实晕动症的干预方式,多采用单一生理指标进行状态检测,仅依托脑电图或头部姿态单一维度数据开展晕动情况判别,数据采集维度较为局限,评估依据较为单薄

Benefits of technology

整合脑电图信号、心电图信号、皮肤电反应信号、眼动追踪参数及头部姿态数据形成完整生理数据集合,多类生理信号从脑电活动、心脏律动、皮肤应激、眼部运动、头部姿态多个维度同步捕捉人体状态,各类数据之间形成互补关联,能够细致捕捉虚拟现实体验过程中人体细微的生理波动变化,生理状态表征覆盖范围更广,可依托全方位的生理数据完成晕动症状态的连续量化刻画。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122557902A_ABST
    Figure CN122557902A_ABST
Patent Text Reader

Abstract

This invention relates to the field of virtual reality intervention technology, specifically disclosing a virtual reality motion sickness intervention system based on emotion guidance. The system includes a data acquisition module that acquires a multimodal physiological response data set comprising electroencephalogram (EEG), electrocardiogram (ECG), skin conductance response, eye-tracking parameters, and head posture. This data is then input into a pre-trained motion sickness assessment model to generate a real-time motion sickness index sequence. An emotion analysis module simultaneously acquires facial expression image sequences and speech signal fragments, extracting the user's emotional tendency characteristics. When the real-time motion sickness index exceeds a trigger threshold, a target emotion guidance strategy is matched based on the emotional tendency characteristics. This involves jointly adjusting the virtual scene's visual presentation, audio playback, and narrative parameters to generate a new experience flow. This system, relying on a multimodal physiological assessment and emotion-linked control mode, can dynamically intervene according to the user's physiological and emotional changes, improving motion sickness and discomfort in virtual reality experiences.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of virtual reality intervention technology, and more specifically to a virtual reality motion sickness intervention system based on emotion guidance. Background Technology

[0002] Virtual reality technology has been widely used in many scenarios such as entertainment simulation and virtual training. When users wear virtual reality devices for immersive experiences, they are very likely to experience motion sickness and other discomfort reactions such as dizziness and nausea. At present, intervention methods for virtual reality motion sickness mostly use a single physiological indicator for state detection, relying solely on electroencephalogram (EEG) or head posture data to determine the degree of motion sickness. The data collection dimensions are relatively limited, and the assessment basis is relatively weak.

[0003] Current conventional intervention methods lack a standardized multimodal physiological data fusion and evaluation system. A single physiological signal cannot fully characterize the comprehensive physiological stress state of the human body during virtual reality experiences, leading to discrepancies between motion sickness severity assessments and actual physical sensations. Furthermore, traditional intervention models only adjust single-dimensional parameters of the virtual scene's visuals or sound effects, failing to consider the user's real-time emotional changes. This makes it impossible to adapt intervention methods to individual emotional differences, and the fixed and rigid scene adjustment dimensions are ill-suited to accommodate varying levels of physical tolerance among different users.

[0004] It is necessary to build an assessment framework for the collaborative acquisition of multiple types of physiological signals, improve the real-time quantitative assessment method for motion sickness, and combine the emotional representations of users' facial and voice levels with the guidance method that matches and adapts to emotional features to achieve the collaborative adjustment of multi-dimensional parameters in virtual reality scenes, thereby making up for the shortcomings of existing technologies such as insufficient data dimensions and lack of emotional adaptability in intervention methods. Summary of the Invention

[0005] The purpose of this invention is to provide an emotion-guided virtual reality motion sickness intervention system to address the problems mentioned above.

[0006] The objective of this invention can be achieved through the following technical solutions: An emotion-guided virtual reality motion sickness intervention system includes: The data acquisition module acquires a set of multimodal physiological response data of the user of the virtual reality device. The set of multimodal physiological response data includes electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, skin conductance response signals, eye tracking parameters, and head posture data. The motion sickness assessment module inputs the multimodal physiological response data set into a pre-trained motion sickness assessment model to generate a real-time motion sickness index sequence; The emotion analysis module simultaneously collects the user's facial expression image sequence and voice signal segment, and extracts emotional tendency features based on the facial expression image sequence and the voice signal segment; The strategy matching module selects a target guidance strategy from multiple emotion guidance strategies based on the emotion tendency characteristics when the current index value in the real-time motion sickness index sequence exceeds a preset intervention trigger threshold. The scene control module performs joint adjustments on the visual presentation parameters, audio playback parameters, and narrative parameters of the virtual reality scene according to the target guidance strategy, generating an adjusted virtual reality experience stream.

[0007] As a further aspect of the present invention, obtaining a set of multimodal physiological response data of virtual reality device users specifically includes: The user's electroencephalogram (EEG), electrocardiogram (ECG), and skin conductance signals are collected via dry electrodes integrated into the padding of the virtual reality headset. The eye-tracking module built into the virtual reality headset collects data on the user's pupil diameter, gaze point offset, and eyelid flicker frequency. The virtual reality headset's built-in posture sensing module collects data on the user's head center of gravity shift trajectory and head sway frequency. The electroencephalogram (EEG) signal, electrocardiogram (ECG) signal, skin conductance response signal, pupil diameter data, fixation point offset data, eyelid flicker frequency data, head center of gravity offset trajectory data, and head sway frequency data are time-stamped and aligned to form the multimodal physiological response data set.

[0008] As a further aspect of the present invention, the multimodal physiological response data set is input into a pre-trained motion sickness assessment model to generate a real-time motion sickness index sequence, specifically including: Baseline drift correction and power frequency notch filtering are performed on each signal channel in the multimodal physiological response dataset to obtain a denoised multimodal physiological feature matrix; The denoised multimodal physiological feature matrix is ​​segmented using a sliding window to generate multiple temporal physiological feature fragments; Each temporal physiological feature segment is sequentially input into a pre-trained motion sickness assessment model, which includes a cascaded convolutional feature extraction layer and a long short-term memory temporal coding layer. The spatial feature map of each temporal physiological feature segment is extracted by the convolutional feature extraction layer, and then the temporal dependency relationship of the spatial feature map is modeled by the long short-term memory temporal coding layer to output the motion sickness index prediction value corresponding to each temporal physiological feature segment. The motion sickness index prediction values ​​corresponding to all time-series physiological feature segments are arranged in chronological order to generate the real-time motion sickness index sequence.

[0009] As a further aspect of the present invention, the user's facial expression image sequence and speech signal segment are acquired simultaneously, and emotional tendency features are extracted based on the facial expression image sequence and the speech signal segment, specifically including: The user's facial region image sequence is captured by a downward-facing camera built into the virtual reality headset, and the facial region image sequence covers the eye, eyebrow and mouth areas; The user's voice signal segments are collected via the microphone array built into the virtual reality headset. The voice signal segments include spontaneous sighs, hums and short phrases produced by the user. For each frame of the facial region image sequence, facial key point localization processing is performed to extract features of eyebrow curvature, eye opening, and mouth corner angle. The speech signal segment is subjected to fundamental frequency extraction processing and acoustic energy envelope analysis processing to obtain speech pitch fluctuation characteristics and speech energy attenuation characteristics; The features of eyebrow curvature, eye opening, mouth corner angle, voice pitch fluctuation, and voice energy attenuation are concatenated to generate the emotion tendency feature.

[0010] As a further aspect of the present invention, when the current index value in the real-time motion sickness index sequence exceeds a preset intervention trigger threshold, a target guidance strategy is selected from multiple emotion guidance strategies based on the emotional tendency characteristics, specifically including: Detect whether the current index value in the real-time motion sickness index sequence is greater than a preset intervention trigger threshold; If the current index value is greater than the intervention trigger threshold, the emotional tendency feature is input into the pre-trained emotion classifier to obtain the current emotion category label. The strategy mapping table is used to query the guidance strategy identifier that has the highest matching degree with the current emotion category label. The strategy mapping table stores the correspondence between various emotion category labels and various guidance strategy identifiers in advance. Based on the guidance strategy identifier, at least one guidance strategy is selected from the visual guidance strategy library, the audio guidance strategy library, and the narrative guidance strategy library as the target guidance strategy.

[0011] As a further aspect of the present invention, at least one guidance strategy is selected as the target guidance strategy from the visual guidance strategy library, the audio guidance strategy library, and the narrative guidance strategy library based on the guidance strategy identifier, specifically including: When the guidance strategy identifier points to a visual guidance type, a color temperature adjustment strategy and a light and shadow softening strategy are selected from the visual guidance strategy library as the target guidance strategy; When the guidance strategy identifier points to an audio guidance type, the pink noise superposition strategy and the neural entrainment beat strategy are selected from the audio guidance strategy library as the target guidance strategy; When the guidance strategy identifier points to the narrative guidance type, a plot branch switching strategy and a movement speed decay strategy are selected from the narrative guidance strategy library as the target guidance strategy; When the guidance strategy identifier points to a composite guidance type, a guidance strategy is selected from the visual guidance strategy library, the audio guidance strategy library, and the narrative guidance strategy library, and the three selected guidance strategies are combined into the target guidance strategy.

[0012] As a further aspect of the present invention, the visual presentation parameters, audio playback parameters, and narrative parameters of the virtual reality scene are jointly adjusted according to the target guidance strategy to generate an adjusted virtual reality experience stream, specifically including: The visual adjustment instructions contained in the target guidance strategy are analyzed, and the current color temperature value of the virtual reality scene is gradually adjusted to the warm color temperature range according to the visual adjustment instructions, and the sharpness of the light and shadow edges in the virtual reality scene is reduced to below the preset softening coefficient. The audio adjustment instructions contained in the target guidance strategy are analyzed, and a pink noise layer is superimposed on the original audio track of the virtual reality scene according to the audio adjustment instructions. The volume of the neural bandage beat audio is adjusted to a modulation depth that is synchronized with the user's heart rate. The narrative adjustment instructions contained in the target guidance strategy are analyzed, and the current task target in the virtual reality scene is switched from a high-speed movement task to a low-speed exploration task according to the narrative adjustment instructions, and the movement speed of the virtual character is limited to a preset comfortable speed threshold. The virtual reality scene, after color temperature adjustment, light and shadow softening, audio overlay, and narrative switching, is rendered into a continuous stream of the adjusted virtual reality experience.

[0013] As a further aspect of the present invention, after jointly adjusting the visual presentation parameters, audio playback parameters, and narrative parameters of the virtual reality scene according to the stated target guidance strategy, the method further includes performing: During the playback of the adjusted virtual reality experience stream, the user's updated multimodal physiological response data set, updated facial expression image sequence, and updated speech signal segment are continuously collected; The updated multimodal physiological response dataset is input into the motion sickness assessment model to generate an updated real-time motion sickness index sequence; The updated facial expression image sequence and the updated speech signal segment are input into the emotion classifier to generate updated emotion tendency features; Compare the relationship between the index values ​​in the updated real-time motion sickness index sequence and the intervention trigger threshold, and determine the degree of mood improvement based on the updated mood tendency characteristics; Based on the relationship between the index value and the intervention trigger threshold, as well as the degree of emotional improvement, the intensity of the target guidance strategy is incrementally adjusted.

[0014] As a further aspect of the present invention, based on the relationship between the index value and the intervention trigger threshold and the degree of emotional improvement, the intensity of the targeted guidance strategy is incrementally adjusted, specifically including: The difference between the index value in the updated real-time motion sickness index sequence and the intervention trigger threshold is calculated to obtain the motion sickness reduction value; Calculate the feature distance between the updated sentiment tendency feature and the sentiment tendency feature to obtain the sentiment improvement magnitude value; When the decrease in motion sickness is positive and the improvement in mood exceeds the preset effective improvement threshold, the current application intensity of the target guidance strategy remains unchanged. When the decrease in motion sickness is positive but the improvement in mood does not exceed the effective threshold for improvement, the adjustment step size of the visual adjustment instruction and the superposition gain of the audio adjustment instruction in the target guidance strategy are each increased by one gradient unit. When the motion sickness reduction value is negative, the narrative adjustment instruction in the target guidance strategy is switched to a higher intensity scene replacement instruction, and the target color temperature range in the visual adjustment instruction is shifted towards a longer wavelength color temperature direction.

[0015] As a further aspect of the present invention, after generating the adjusted virtual reality experience stream, the following steps are also included: Record the real-time motion sickness index sequence, the emotional tendency characteristics, the identification information of the target guidance strategy, and the rendering parameters of the adjusted virtual reality experience stream to form a single intervention event record; The single intervention event record is stored in the user's personalized configuration profile, which contains the user's historical intervention event record set; The next time the user starts the virtual reality device, the set of historical intervention event records in the personalized configuration profile is loaded; Based on the frequency of occurrence of the target guidance strategy identifier in the historical intervention event record set, the baseline value of the intervention trigger threshold of the motion sickness assessment model is pre-adjusted; Based on the distribution center of the rendering parameters of the adjusted virtual reality experience stream in the set of historical intervention event records, the initial values ​​of the visual presentation parameters, the audio playback parameters, and the narrative parameters are preset.

[0016] The beneficial effects of this invention are: By integrating EEG signals, ECG signals, skin conductance signals, eye tracking parameters, and head posture data, a complete set of physiological data is formed. Multiple physiological signals simultaneously capture the human body's state from multiple dimensions, including brain activity, heart rhythm, skin stress, eye movement, and head posture. The various data forms complementary correlations, which can meticulously capture subtle physiological fluctuations and changes in the human body during virtual reality experiences. The physiological state representation has a wider coverage and can complete the continuous quantitative characterization of motion sickness based on comprehensive physiological data.

[0017] Simultaneously acquiring facial expression image sequences and voice signal fragments completes the extraction of emotional tendency features, capturing the user's inner emotional changes from two dimensions: visual expression and voice tone. The sources of emotional feature extraction are more diverse, accurately capturing subtle fluctuations in the user's emotions during the immersive experience. Based on the extracted emotional tendency features, suitable emotional guidance strategies are selected, and simultaneously, three types of parameters—visual presentation of the virtual reality scene, audio playback, and narrative—are adjusted in sync. Scene control is no longer limited to single-dimensional changes; multi-parameter coordinated adjustment can conform to the linkage between human physiological discomfort and emotional changes, making the adjustment and changes in the virtual experience flow more in line with human sensory acceptance habits. Attached Figure Description

[0018] The invention will now be further described with reference to the accompanying drawings.

[0019] Figure 1 This is a timing diagram of the virtual reality motion sickness intervention system based on emotion guidance described in this invention; Figure 2 This is a flowchart illustrating the operation of the multimodal physiological response data acquisition method; Figure 3 This is a flowchart illustrating the process of the emotion tendency feature extraction method. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] See Figure 1 This invention is an emotion-guided virtual reality motion sickness intervention system, comprising: The system comprises a data acquisition module, a motion sickness assessment module, an emotion analysis module, a strategy matching module, and a scene control module. The data acquisition module acquires a multimodal physiological response data set from the virtual reality device user, including EEG signals, ECG signals, skin conductance signals, eye-tracking parameters, and head posture data. The motion sickness assessment module inputs this data set into a pre-trained motion sickness assessment model to generate a real-time motion sickness index sequence. The emotion analysis module simultaneously acquires facial expression image sequences and speech signal fragments from the user, extracting emotional tendency features based on these sequences. When the current index value in the real-time motion sickness index sequence exceeds a preset intervention trigger threshold, the strategy matching module selects a target guidance strategy from multiple emotion guidance strategies based on the emotional tendency features. The scene control module performs joint adjustments to the visual presentation parameters, audio playback parameters, and narrative parameters of the virtual reality scene according to the target guidance strategy, generating an adjusted virtual reality experience stream.

[0022] In one embodiment of the present invention, see [reference] Figure 2 The data acquisition module collects the user's electroencephalogram (EEG), electrocardiogram (ECG), and skin conductance response signals via dry electrodes integrated into the virtual reality headset pad; it collects the user's pupil diameter, gaze point offset, and eyelid flicker frequency data via the eye-tracking module built into the virtual reality headset; and it collects the user's head center of gravity offset trajectory data and head sway frequency data via the posture sensing module built into the virtual reality headset. The EEG, ECG, skin conductance response signals, pupil diameter, gaze point offset, eyelid flicker frequency, head center of gravity offset trajectory, and head sway frequency data are time-stamped and aligned to form a multimodal physiological response data set. The motion sickness assessment module performs baseline drift correction and power frequency notch filtering on each signal channel in the multimodal physiological response dataset to obtain a denoised multimodal physiological feature matrix. The denoised multimodal physiological feature matrix is ​​then segmented using a sliding window to generate multiple temporal physiological feature segments. Each temporal physiological feature segment is sequentially input into a pre-trained motion sickness assessment model, which includes cascaded convolutional feature extraction layers and long short-term memory temporal coding layers. The convolutional feature extraction layers extract spatial feature maps for each temporal physiological feature segment, and the long short-term memory temporal coding layer models the temporal dependencies of these spatial feature maps, outputting a predicted motion sickness index value for each temporal physiological feature segment. Finally, the predicted motion sickness index values ​​for all temporal physiological feature segments are arranged in chronological order to generate a real-time motion sickness index sequence.

[0023] In practice, once the user activates the virtual reality roller coaster experience application, the data acquisition module begins to work. Dry electrodes integrated into the forehead and behind the ears of the virtual reality headset pads come into contact with the user's skin, continuously acquiring EEG signals, ECG signals, and skin conductance response signals. The eye-tracking module built into the virtual reality headset synchronously acquires pupil diameter data, fixation point offset data, and eyelid flicker frequency data through an infrared light source and a camera. The posture sensing module built into the virtual reality headset acquires head center of gravity offset trajectory data and head sway frequency data through an inertial measurement unit. The data acquisition module performs timestamp alignment processing on the EEG signals, ECG signals, skin conductance response signals, pupil diameter data, fixation point offset data, eyelid flicker frequency data, head center of gravity offset trajectory data, and head sway frequency data to form a multimodal physiological response data set.

[0024] In some embodiments, the timestamp alignment process is as follows: the data acquisition module compares the timestamps attached to the EEG signal, ECG signal, skin conductance response signal, pupil diameter data, fixation point offset data, eyelid flicker frequency data, head center of gravity offset trajectory data, and head shaking frequency data according to a unified system clock reference. The signal with different sampling rates is resampled to the same 100Hz target frequency using a linear interpolation method. The data of each channel is fine-tuned so that each sampling point corresponds to the same time point, resulting in a synchronized multimodal physiological response data set.

[0025] In practical implementation, after obtaining the multimodal physiological response data set, the motion sickness assessment module performs baseline drift correction and power frequency notch filtering on the following channels: EEG signal channel, ECG signal channel, skin conductance response signal channel, pupil diameter data channel, fixation point offset data channel, eyelid flicker frequency data channel, head center of gravity offset trajectory data channel, and head sway frequency data channel. Baseline drift correction is achieved by subtracting the median baseline calculated through a 1-second sliding window from each channel signal. Power frequency notch filtering applies a band-stop filter to each channel signal after baseline drift correction to eliminate 50Hz power line interference. The transfer function of the band-stop filter is: in: This represents the system function of a band-stop filter. This represents the normalized notch angle frequency. This represents the extreme contraction factor and its value is between 0.9 and 0.99. The Z-transform complex variable is represented by the multimodal physiological feature matrix obtained after the signals of each channel pass through the band-stop filter.

[0026] In some embodiments, when performing sliding window segmentation on the denoised multimodal physiological feature matrix, the time width of the sliding window is set to 2 seconds, the sliding step size between adjacent windows is 0.5 seconds, and starting from the initial time of the denoised multimodal physiological feature matrix, a multidimensional signal segment with a duration of 2 seconds is extracted every 0.5 seconds as a temporal physiological feature segment. The temporal physiological feature segments overlap each other on the time axis, covering the complete virtual reality experience process.

[0027] In practice, each temporal physiological feature segment is sequentially input into a pre-trained motion sickness assessment model. The motion sickness assessment model consists of cascaded convolutional feature extraction layers and long short-term memory temporal coding layers. The convolutional feature extraction layer applies multiple one-dimensional convolutional kernels to the input temporal physiological feature segments to perform convolution operations along the time axis, generating multiple feature maps. After batch normalization and linear rectified activation, max pooling is performed in the time dimension to obtain spatial feature maps. The long short-term memory temporal coding layer takes the spatial feature maps extracted from the same channel on consecutive temporal physiological feature segments as a time series. At each time step, the hidden state and cell state are updated through forget gate, input gate, and output gate, and the latent vector corresponding to the current temporal physiological feature segment is output. This vector is mapped to a motion sickness index prediction value through a fully connected layer. The motion sickness index prediction values ​​corresponding to all temporal physiological feature segments are arranged sequentially according to the segment start timestamp to generate a real-time motion sickness index sequence.

[0028] It is understandable that convolutional kernels of different lengths in the convolutional feature extraction layer can capture the local time-frequency patterns of EEG signals, ECG signals, skin conductance response signals, pupil diameter data, fixation point shift data, eyelid flicker frequency data, head center of gravity shift trajectory data, and head sway frequency data within a 2-second segment. The spatial feature map encodes the spatial interaction relationships between multimodal physiological response data within a short time window. Similarly, the Long Short-Term Memory (LSTM) temporal coding layer utilizes memory units to maintain motion sickness-related physiological changes over a long period. When a user's motion sickness progresses from mild to moderate, the previously accumulated EEG signal spectral shift features and pupil diameter data expansion trends can be transferred to the current temporal physiological feature segment, enabling the motion sickness index prediction value to reflect the gradual evolution of motion sickness severity.

[0029] Optionally, when performing baseline drift correction on each signal channel, the median baseline calculation window for the EEG signal channel can be extended to 5 seconds to retain more low-frequency EEG components, while the median baseline calculation window for the skin conductance response signal channel can be shortened to 0.5 seconds to accommodate the rapid drift characteristics of skin conductance. Optionally, when arranging the motion sickness index prediction values ​​corresponding to all temporal physiological feature segments to generate a real-time motion sickness index sequence, the motion sickness index prediction values ​​of the overlapping parts of adjacent temporal physiological feature segments are Gaussian weighted smoothed and fused. The value at the center of the overlapping region is taken using the weighted average of the prediction values ​​of the two segments, with the weight decreasing from the center to the edge of the segment according to a Gaussian distribution to eliminate the index jump caused by the window boundary effect.

[0030] In one embodiment of the present invention, the emotion analysis module acquires a sequence of facial region images of the user via a downward-facing camera built into the virtual reality headset. The facial region image sequence covers the eye, eyebrow, and mouth areas; see reference. Figure 3 The system collects user speech signal segments via the microphone array built into the virtual reality headset. These segments include spontaneous sighs, hums, and short phrases. Facial key point localization is performed on each frame of the facial image sequence to extract features such as eyebrow curvature, eye opening, and mouth corner angle. Fundamental frequency extraction and acoustic energy envelope analysis are performed on the speech signal segments to obtain speech pitch fluctuation features and speech energy attenuation features. These features are then concatenated to generate emotional tendency features.

[0031] In practice, the user wears a virtual reality headset and launches a marine exploration virtual reality application. The emotion analysis module begins to simultaneously collect the user's facial expression image sequence and voice signal fragments. The virtual reality headset's built-in downward-facing camera captures the user's facial area image sequence at a sampling rate of 60 frames per second, covering the eye, eyebrow, and mouth areas. The virtual reality headset's built-in microphone array continuously collects the user's voice signal fragments at a sampling rate of 16kHz. The voice signal fragments include the user's spontaneous sighs, hums, and short phrases. During the virtual reality experience, the user's natural voice reactions caused by scene changes are completely recorded.

[0032] In some embodiments, a downward-facing camera is installed inside the shell of the virtual reality headset near the bridge of the nose, with the lens optical axis pointing towards the lower half of the user's face. Infrared supplementary lighting is used to eliminate the influence of light leakage from the internal screen of the virtual reality headset on the facial image sequence. Each frame of the image captured by the downward-facing camera includes the user's facial area from the upper edge of the eyebrows to the lower edge of the lower lip. The range of eyebrow movement, eyelid opening and closing, and the shape of the corners of the mouth can all be presented completely without obstruction. In its implementation, the emotion analysis module performs facial key point localization processing on each frame of the facial region image sequence. It locates the eyebrow contour points, eyelid edge points, and mouth corner endpoints using a pre-trained 68-point facial feature point detection model. It extracts features of eyebrow curvature, eye opening, and mouth corner angle. The eyebrow curvature feature is obtained by calculating the maximum vertical distance between the line connecting the inner and outer endpoints of the eyebrow and the fitted curve of the eyebrow contour. The eye opening feature is obtained by calculating the ratio of the vertical pixel distance between the upper and lower eyelid edge points to the user's individual calibration baseline. The mouth corner angle feature is obtained by calculating the angle between the line connecting the left and right mouth corner endpoints and the horizontal reference line.

[0033] In some embodiments, the emotion analysis module performs fundamental frequency extraction and acoustic energy envelope analysis on the speech signal segments. Fundamental frequency extraction uses the autocorrelation function method to estimate the short-time fundamental frequency of the speech signal segments with a frame length of 25 milliseconds and a frame shift of 10 milliseconds, obtaining speech pitch fluctuation characteristics. Acoustic energy envelope analysis calculates the short-time root mean square energy of each frame of the speech signal segments with the same frame length and frame shift, extracts the upper envelope curve of the frame energy sequence, and calculates the decay rate of the envelope curve as the speech energy decay characteristic. It can be understood that when a user feels tense in a virtual reality scene, the eyebrow curvature is characterized by eyebrows converging towards the center of the brow and an increased degree of eyebrow curvature; the eye opening is characterized by eyelid tension and contraction leading to a decreased degree of eye opening; the mouth corner angle is characterized by a downward tilt of the mouth corner; simultaneously, the speech pitch fluctuation characteristics show an increase in the mean fundamental frequency and an increase in the fundamental frequency jitter amplitude; and the speech energy decay characteristics show a slower energy envelope decay and even a prolonged tail sound. It is understandable that when users feel happy and relaxed in a virtual reality scene, the characteristics of eyebrow curvature are manifested as eyebrows naturally relaxed, the characteristics of eye opening are manifested as eyelid relaxation causing the degree of eye opening to be close to the calibration reference value, the characteristics of mouth corner angle are manifested as an increased upward angle of the corner of the mouth, the characteristics of voice pitch fluctuation are manifested as a smooth change in fundamental frequency, and the characteristics of voice energy attenuation are manifested as a natural attenuation of energy envelope without obvious prolongation of the tail sound.

[0034] In its implementation, the emotion analysis module concatenates features such as eyebrow curvature, eye opening, mouth corner angle, voice pitch fluctuation, and voice energy attenuation to generate emotion tendency features. The concatenation method involves stitching the feature vectors for eyebrow curvature, eye opening, mouth corner angle, voice pitch fluctuation, and voice energy attenuation in a preset order. When the frame rate of the facial region image sequence is 60 frames per second and the fundamental frequency extraction frame rate of the voice signal segment is 100 frames per second, the voice pitch fluctuation and voice energy attenuation feature vectors are downsampled to align with the frame rate of the facial region image sequence, ensuring that all feature vectors have the same number of sampling points in the time dimension. Optionally, the emotion analysis module introduces an independent evaluation mechanism for the left and right eyebrows when extracting eyebrow curvature features, calculating the curvature of the left and right eyebrows separately as independent feature dimensions to capture the asymmetrical emotional signal expressed by the user's unilateral eyebrow elevation. Optionally, when the sentiment analysis module performs fundamental frequency extraction on speech signal segments, it distinguishes between unvoiced and voiced frames, estimating the fundamental frequency only for voiced frames and setting the fundamental frequency value to zero for unvoiced frames. When generating speech pitch fluctuation features, it removes zero-value segments from consecutive unvoiced frame regions to avoid interference from silent segments on the fundamental frequency fluctuation statistics. Optionally, during feature concatenation, the features of eyebrow curvature, eye opening, mouth corner angle, speech pitch fluctuation, and speech energy attenuation are respectively subjected to Z-score standardization. The standardization formula is: in: This represents the original feature value of the m-th feature channel. This represents the mean of the m-th feature channel during the individual baseline acquisition phase. This represents the standard deviation of the m-th feature channel during the individual baseline acquisition phase. This represents the standardized eigenvalue of the m-th feature channel. The standardization process ensures that all feature channels have the same dimensional scale before feature cascading.

[0035] In one embodiment of the present invention, the strategy matching module detects whether the current index value in the real-time motion sickness index sequence is greater than a preset intervention trigger threshold; if the current index value is greater than the intervention trigger threshold, the emotional tendency feature is input into a pre-trained emotion classifier to obtain the current emotion category label; the guidance strategy identifier with the highest matching degree with the current emotion category label is queried from the strategy mapping table, the strategy mapping table pre-stores the correspondence between multiple emotion category labels and multiple guidance strategy identifiers; at least one guidance strategy is selected as the target guidance strategy from the visual guidance strategy library, the audio guidance strategy library and the narrative guidance strategy library according to the guidance strategy identifier. The method for selecting a target guidance strategy from the visual guidance strategy library, audio guidance strategy library, and narrative guidance strategy library based on the guidance strategy identifier is as follows: When the guidance strategy identifier points to the visual guidance type, the color temperature adjustment strategy and the light and shadow softening strategy are selected from the visual guidance strategy library as the target guidance strategy; when the guidance strategy identifier points to the audio guidance type, the pink noise superposition strategy and the neural entrainment beat strategy are selected from the audio guidance strategy library as the target guidance strategy; when the guidance strategy identifier points to the narrative guidance type, the plot branch switching strategy and the movement speed decay strategy are selected from the narrative guidance strategy library as the target guidance strategy; when the guidance strategy identifier points to the composite guidance type, one guidance strategy is selected from each of the visual guidance strategy library, audio guidance strategy library, and narrative guidance strategy library, and the three selected guidance strategies are combined into the target guidance strategy.

[0036] In practice, users wear virtual reality headsets and run ocean exploration virtual reality applications. When the experience scene switches to the deep-sea canyon crossing segment, the current index value in the real-time motion sickness index sequence continuously climbs from 3.2 to 7.5. The strategy matching module detects that the current index value is greater than the preset intervention trigger threshold of 6.0, triggering the intervention process. The strategy matching module obtains the emotional tendency features generated by the emotion analysis module within the same time window. The emotional tendency features are manifested as the degree of eyebrow curvature indicating that the eyebrows are gathered, the degree of eye opening indicating that the eyelids are tense and contracted, the angle of the mouth corners indicating that the corners of the mouth are pressed down, the voice tone fluctuation features indicating that the fundamental frequency is increased and the shaking is intensified, and the voice energy attenuation features indicating that the energy attenuation is slow. The strategy matching module inputs the emotional tendency features into the pre-trained emotion classifier and obtains the current emotion category label as "tension-anxiety".

[0037] In some embodiments, the pre-trained emotion classifier employs a fully connected neural network structure. The number of nodes in the input layer equals the dimension of the emotion tendency feature. The hidden layer contains two fully connected layers, each with 128 neurons, and each layer is followed by a wired rectified activation function and a random dropout operation with a dropout rate of 0.3. The output layer contains 6 neurons corresponding to six emotion category labels: "calm," "tension-anxiety," "pleasure-excitement," "frustration-disappointment," "disgust-resistance," and "surprise-curiosity." The output layer uses a Softmax activation function. The emotion classifier calculates the probability that the input emotion tendency feature vector f belongs to the j-th emotion category label: in: The feature vector f representing the emotion tendency belongs to the emotion category label. The posterior probability, This represents the weight vector corresponding to the j-th neuron in the output layer. This represents the bias term corresponding to the j-th neuron in the output layer. and Let represent the weight vector and bias term corresponding to the k-th emotion category label, respectively. The denominator term sums the output values ​​of all six emotion category labels. The emotion classifier selects the emotion category label corresponding to the maximum posterior probability as the current emotion category label output.

[0038] In practice, after the strategy matching module obtains the current emotion category label, it queries the strategy mapping table for the guidance strategy identifier that has the highest matching degree with "tension-anxiety". The strategy mapping table stores the correspondence between various emotion category labels and various guidance strategy identifiers in advance. For the contents of the strategy mapping table, please refer to Table 1.

[0039] Table 1: Strategy Mapping Table The strategy matching module determines the guidance strategy identifier as "VA-001" based on the lookup table results. This identifier points to a composite guidance type. The strategy matching module simultaneously selects one guidance strategy from each of the visual guidance strategy library, audio guidance strategy library, and narrative guidance strategy library. From the visual guidance strategy library, it selects the color temperature adjustment strategy and the light and shadow softening strategy. From the audio guidance strategy library, it selects the pink noise superposition strategy and the neural entrainment beat strategy. From the narrative guidance strategy library, it selects the plot branch switching strategy and the movement speed decay strategy. The selected three guidance strategies are combined into the target guidance strategy.

[0040] It is understandable that when the guidance strategy identifier points to the visual guidance type, the strategy matching module selects only the color temperature adjustment strategy and the light and shadow softening strategy from the visual guidance strategy library as the target guidance strategy. The color temperature adjustment strategy in the visual guidance strategy library defines an adjustment curve that gradually moves from the current color temperature value to the warm-toned target color temperature range. The light and shadow softening strategy defines a sequence of filter parameters that gradually reduces the sharpness of the lighting edges in the scene to below the softening coefficient. Each adjustment strategy in the visual guidance strategy library contains a corresponding adjustment parameter template and application timing. Similarly, it is understandable that when the guidance strategy identifier points to the audio guidance type, the strategy matching module selects only the pink noise superposition strategy and the neural entrainment beat strategy from the audio guidance strategy library as the target guidance strategy. The pink noise superposition strategy in the audio guidance strategy library defines the spectral envelope and superposition gain curve of pink noise. The neural entrainment beat strategy defines the synchronous mapping relationship between the beat frequency and the user's heart rate, as well as the modulation depth envelope parameters. The audio guidance strategy library stores the executable adjustment instructions for each strategy in the form of an audio processing plugin chain.

[0041] In some embodiments, when the guidance strategy identifier points to a narrative guidance type, the strategy matching module selects a plot branch switching strategy and a movement speed decay strategy from the narrative guidance strategy library as the target guidance strategy. The plot branch switching strategy in the narrative guidance strategy library stores alternative plot branches for each scene node of the virtual reality application, and the movement speed decay strategy stores a table of speed decay coefficients for different virtual character types under different degrees of motion sickness. The plot branch switching strategy changes the task objective by modifying the activation state of scene graph nodes, and the movement speed decay strategy takes effect by modifying the movement speed parameters of the virtual character animation controller.

[0042] In practical implementation, after selecting the target guidance strategy, the strategy matching module outputs the structured description data of the target guidance strategy to the scene control module. This structured description data includes identifiers of all adjustment instructions in the target guidance strategy, the range of adjustment target parameters, the execution sequence, and the collaborative constraints between instructions. The scene control module parses each adjustment instruction based on the structured description data and executes subsequent joint adjustment operations. Optionally, whenever a new correspondence between an emotion category label and a guidance strategy identifier is added to the strategy mapping table, the strategy matching module records the timestamp and source of the addition. When multiple guidance strategy identifiers corresponding to the same emotion category label show frequency differences, the strategy matching module dynamically adjusts the default value of the guidance strategy identifier corresponding to the highest matching degree based on the frequency. Optionally, when the difference between the maximum posterior probability and the second-largest posterior probability output by the emotion classifier is less than 0.15, the strategy matching module determines that the current emotion category label is in a state of boundary ambiguity, uses the emotion category label corresponding to the second-largest posterior probability as a candidate label, and simultaneously queries the guidance strategy identifiers corresponding to the current emotion category label and the candidate label from the strategy mapping table, selecting the one with the lower intervention intensity as the identifier of the target guidance strategy.

[0043] In one embodiment of the present invention, the scene control module parses the visual adjustment instructions contained in the target guidance strategy, and gradually adjusts the current color temperature value of the virtual reality scene to the warm color temperature range according to the visual adjustment instructions, and reduces the sharpness of the light and shadow edges in the virtual reality scene to below a preset softening coefficient; it parses the audio adjustment instructions contained in the target guidance strategy, and superimposes a pink noise layer on the original audio track of the virtual reality scene according to the audio adjustment instructions, and adjusts the volume of the neural bandage beat audio to a modulation depth synchronized with the user's heart rate; it parses the narrative adjustment instructions contained in the target guidance strategy, and switches the current task target in the virtual reality scene from a high-speed movement task to a low-speed exploration task according to the narrative adjustment instructions, and limits the movement speed of the virtual character to within a preset comfortable speed threshold; and renders the virtual reality scene after color temperature adjustment, light and shadow softening, audio superposition, and narrative switching into a continuous adjusted virtual reality experience stream.

[0044] In practice, users wear virtual reality headsets and run ocean exploration virtual reality applications. When the experience enters the deep-sea canyon crossing stage, the strategy matching module selects a composite guidance strategy based on the real-time motion sickness index sequence exceeding the intervention trigger threshold and the emotional tendency characteristics being classified as "tension-anxiety". This target guidance strategy includes simultaneously activated visual adjustment instructions, audio adjustment instructions, and narrative adjustment instructions. The scene control module receives the structured description data of the target guidance strategy and initiates a joint adjustment process, synchronously modifying the visual presentation parameters, audio playback parameters, and narrative parameters of the virtual reality scene.

[0045] In practical implementation, the scene control module parses the visual adjustment instructions contained in the target guidance strategy, and extracts the color temperature adjustment parameter set and the lighting softening parameter set from the data structure of the visual adjustment instructions. The color temperature adjustment parameter set defines the initial color temperature value, the warm-toned target color temperature range, the color temperature adjustment time constant, and the color temperature increment per frame. Based on the color temperature adjustment parameter set, the scene control module gradually adjusts the current color temperature value of the virtual reality scene from 6500K to the warm-toned target color temperature range of 3500K to 4200K. Before rendering each frame, the global illumination color temperature parameter of the virtual reality scene is incremented by an increment until it enters the warm-toned target color temperature range. The color temperature adjustment process follows the following time change relationship: in: This represents the real-time color temperature value of the virtual reality scene after time t since the start of self-adjustment. This indicates the current color temperature value at the start of the adjustment. This represents the center value of the target color temperature range for warm tones. This represents the time constant for color temperature adjustment, with time t increasing from zero in seconds.

[0046] In practice, the scene control module reduces the sharpness of light and shadow edges in the virtual reality scene to below a preset softening coefficient of 0.35 based on the light and shadow softening parameter set. The light and shadow softening parameter set defines the softening target coefficient, the softening filter kernel size, and the kernel size increment per frame. In the post-processing stage of the virtual reality scene rendering pipeline, the scene control module performs a Gaussian blur operation on the scene lighting buffer. The standard deviation of the Gaussian blur kernel increases frame by frame from the initial value of 0.5 until it reaches the standard deviation of 8.0 corresponding to the softening target coefficient, thus achieving a smooth transition from sharp edges to soft edges.

[0047] In some embodiments, the scene control module parses the audio adjustment instructions contained in the target guidance strategy. These instructions include a pink noise superposition parameter set and a neural entrainment beat parameter set. The pink noise superposition parameter set defines the normalized spectral envelope, initial superposition gain, and gain decay time curve of the pink noise. The scene control module superimposes a layer of pink noise onto the original audio track of the virtual reality scene via an audio mixing bus. The spectrum of the pink noise is generated according to the 1 / f characteristic, and the superposition gain gradually decays from -18dBFS at a rate decreasing by 1dB every 5 seconds, creating an auditory effect that is not... The eye-catching background overlay effect, the neural-encased beat parameter set defines the synchronization mapping function of the beat fundamental frequency, beat frequency and the user's real-time heart rate and the modulation depth. The scene control module obtains the user's real-time heart rate value from the motion sickness assessment module, locks the beat frequency of the neural-encased beat audio to the real-time heart rate value minus 6 beats per minute, and adjusts the volume of the neural-encased beat audio to a modulation depth of 0.4. A modulation depth of 0.4 means that the ratio of the peak value of the amplitude envelope of the neural-encased beat audio to the peak value of the pink noise superposition layer is 0.4, so that the beat signal is embedded in the audio stream with a subthreshold perceived intensity.

[0048] In some embodiments, the scene control module parses the narrative adjustment instructions contained in the target guidance strategy. The narrative adjustment instructions include a plot branch switching identifier and a movement speed decay parameter set. The plot branch switching identifier points to an alternative narrative branch, "slow exploration of the reef area," in the deep-sea canyon crossing scene. The scene control module changes the currently active scene node from "high-speed channel crossing" to "slow exploration of the reef area" by modifying the scene node state of the virtual reality application. The current task objective in the virtual reality scene immediately changes from high-speed crossing along the canyon channel to slow walking and observing marine life in the reef area. The movement speed decay parameter set defines the upper limit of the virtual character's movement speed and the acceleration decay coefficient. The scene control module limits the maximum movement speed of the virtual character to within a preset comfortable speed threshold of 2.0 meters per second and decays the movement acceleration of the virtual character from the default value of 8.0 meters per square second to 2.5 meters per square second. The movement instructions input by the user through the controller are mapped to the limited movement speed range.

[0049] It is understandable that when the scene control module performs joint adjustments to visual presentation parameters, audio playback parameters, and narrative parameters, the adjustments of the three types of parameters work together on the timeline. The process of color temperature shifting to warm tones and light and shadow softening lasts for about 12 seconds. The addition of pink noise and the introduction of neural-driven beats are initiated simultaneously with the visual adjustments. The narrative switching is completed within 2 seconds after the color temperature adjustment begins to avoid continuous exposure for users in a high-speed movement state. The combined effect of the three types of adjustments causes the overall stimulation intensity of the virtual reality scene to decrease in an orderly manner in a short period of time.

[0050] It is understandable that when the scene control module renders the virtual reality scene, after color temperature adjustment, lighting softening, audio overlay, and narrative switching, into a continuous virtual reality experience stream, the rendering pipeline comprehensively applies the current color temperature offset value, Gaussian blur kernel parameters, audio mixing buffer state, and virtual character movement speed limits for each frame. This ensures smooth parameter interpolation between consecutive frames of the virtual reality experience stream, allowing users to experience natural and seamless scene transitions in the virtual reality headset. A summary of the visual presentation parameters, audio playback parameters, and narrative parameters involved in the joint adjustment process is shown in Table 2.

[0051] Table 2: Detailed List of Joint Adjustment Parameters for Virtual Reality Scenes Optionally, when adjusting the color temperature, the scene control module applies differentiated color temperature offsets to different light source types in the virtual reality scene. The main directional light's color temperature is offset to the upper limit of the warm-toned target color temperature range (4200K), and the ambient light's color temperature is offset to the lower limit of the warm-toned target color temperature range (3500K). This separation of color temperature levels between the main light source and the ambient light enhances the visual comfort of the scene. Optionally, when reducing the sharpness of light and shadow edges, the scene control module applies Gaussian blur with different kernel sizes to the distant and near areas of the virtual reality scene. The standard deviation of the Gaussian blur kernel in the distant area is 1.5 times that of the near area, allowing the visual focus area to naturally fall on the near-field activity area, helping to reduce the user's visual vestibular conflict.

[0052] In one embodiment of the present invention, during the playback of the adjusted virtual reality experience stream, the user's updated multimodal physiological response data set, updated facial expression image sequence, and updated speech signal segment are continuously collected; the updated multimodal physiological response data set is input into a motion sickness assessment model to generate an updated real-time motion sickness index sequence; the updated facial expression image sequence and updated speech signal segment are input into an emotion classifier to generate updated emotion tendency features; the relationship between the index value in the updated real-time motion sickness index sequence and the intervention trigger threshold is compared, and the degree of emotion improvement is determined based on the updated emotion tendency features; based on the relationship between the index value and the intervention trigger threshold and the degree of emotion improvement, the intensity of the target guidance strategy is incrementally adjusted. The incremental correction method is as follows: calculate the difference between the index value in the updated real-time motion sickness index sequence and the intervention trigger threshold to obtain the motion sickness reduction value; calculate the feature distance between the updated emotion tendency feature and the original emotion tendency feature to obtain the emotion improvement value; when the motion sickness reduction value is positive and the emotion improvement value exceeds the preset effective improvement threshold, maintain the current application intensity of the target guidance strategy unchanged; when the motion sickness reduction value is positive but the emotion improvement value does not exceed the effective improvement threshold, increase the adjustment step size of the visual adjustment instruction and the superposition gain of the audio adjustment instruction in the target guidance strategy by one gradient unit; when the motion sickness reduction value is negative, switch the narrative adjustment instruction in the target guidance strategy to a higher intensity scene replacement instruction, and shift the target color temperature range in the visual adjustment instruction towards a longer wavelength color temperature direction. In addition, after generating the adjusted virtual reality experience stream, the real-time motion sickness index sequence, emotional tendency characteristics, target guidance strategy identification information, and rendering parameters of the adjusted virtual reality experience stream are recorded to form a single intervention event record. The single intervention event record is stored in the user's personalized configuration file, which contains the user's historical intervention event record set. When the user starts the virtual reality device again, the historical intervention event record set in the personalized configuration file is loaded. Based on the frequency of occurrence of target guidance strategy identification in the historical intervention event record set, the baseline value of the intervention trigger threshold of the motion sickness assessment model is pre-adjusted. Based on the distribution center of the rendering parameters of the adjusted virtual reality experience stream in the historical intervention event record set, the initial values ​​of visual presentation parameters, audio playback parameters, and narrative parameters are pre-set.

[0053] In practice, the user wears a virtual reality headset and runs an ocean exploration virtual reality application, triggering a composite guidance strategy. The adjusted virtual reality experience stream continuously plays in a narrative branch where the deep-sea canyon is replaced by a slow exploration of a reef area. The data acquisition module and the emotion analysis module run continuously during the playback of the adjusted virtual reality experience stream. The data acquisition module synchronously collects the user's updated multimodal physiological response data set, which includes updated EEG signals, ECG signals, skin conductance response signals, pupil diameter data, fixation point offset data, eyelid flicker frequency data, head center of gravity offset trajectory data, and head shaking frequency data. The emotion analysis module synchronously collects the user's updated facial expression image sequence and updated speech signal fragments. The updated facial expression image sequence covers the eye, eyebrow, and mouth areas, and the updated speech signal fragments include short phrases and sighs spontaneously generated by the user during the slow exploration.

[0054] In practice, the motion sickness assessment module inputs the updated multimodal physiological response data set into the pre-trained motion sickness assessment model. After baseline drift correction, power frequency notch filtering, sliding window segmentation, convolutional feature extraction, and long short-term memory temporal coding, an updated real-time motion sickness index sequence is generated. The emotion analysis module inputs the updated facial expression image sequence and the updated speech signal segment into the pre-trained emotion classifier. After facial key point localization, eyebrow curvature feature extraction, eye opening feature extraction, mouth corner angle feature extraction, speech fundamental frequency extraction, and acoustic energy envelope analysis, the features are cascaded and input into the emotion classifier to generate updated emotion tendency features. The strategy matching module obtains the updated real-time motion sickness index sequence from the motion sickness assessment module and the updated emotion tendency features from the emotion analysis module. It compares the relationship between the index values ​​in the updated real-time motion sickness index sequence and the intervention trigger threshold, and determines the degree of emotion improvement based on the updated emotion tendency features.

[0055] In some embodiments, when the strategy matching module determines the degree of emotion improvement, it calculates the difference between the mean index value within the latest 5-second window in the updated real-time motion sickness index sequence and the intervention trigger threshold to obtain the motion sickness reduction value. A positive motion sickness reduction value indicates that the degree of motion sickness has decreased to below the intervention trigger threshold, while a negative motion sickness reduction value indicates that the degree of motion sickness is still higher than the intervention trigger threshold. At the same time, it calculates the feature distance between the updated emotion tendency feature and the original emotion tendency feature at the time of triggering the intervention to obtain the emotion improvement value. The feature distance characterizes the offset of the emotion state in the feature space before and after the intervention. The larger the emotion improvement value, the more obvious the positive change in emotion.

[0056] In practice, the strategy matching module calculates the Euclidean distance between the updated emotion tendency feature vector and the original emotion tendency feature vector stored at the time of triggering intervention as the magnitude of emotion improvement. The formula for calculating the Euclidean distance is as follows: in: This represents the updated sentiment tendency feature vector. Compared with the original sentiment tendency feature vector The characteristic distance between them The total number of dimensions representing emotional tendency characteristics. Representing vectors The One portion, Representing vectors The Each component, summation index Iterate from 1 to the total number of dimensions. The square root operation yields the Euclidean distance value, which represents the magnitude of mood improvement.

[0057] In some embodiments, the strategy matching module performs incremental adjustments to the intensity of the target guidance strategy based on the decrease in motion sickness and the improvement in mood. When the decrease in motion sickness is positive and the improvement in mood exceeds the preset effective threshold of 0.35, it indicates that the current intervention combination has effectively relieved motion sickness symptoms and significantly improved mood. The strategy matching module maintains the current intensity of the target guidance strategy unchanged and continues to regulate the virtual reality scene according to the original color temperature adjustment step size, the incremental increase of the light and shadow softening parameters, the pink noise superposition gain attenuation curve, and the neural entrainment beat modulation depth.

[0058] In practice, when the decrease in motion sickness symptoms is positive but the improvement in mood does not exceed the effective improvement threshold of 0.35, it indicates that although the motion sickness symptoms have been relieved, the user's emotional tension has not been fully relieved. The strategy matching module increases the adjustment step size of the visual adjustment command and the superposition gain of the audio adjustment command in the target guidance strategy by one gradient unit. The color temperature adjustment step size is increased from 0.8K per frame to 1.6K per frame to accelerate convergence to the warm color temperature range of the target color temperature. The standard deviation increment of the Gaussian blur kernel for light and shadow softening is increased from 0.3 per frame to 0.6 per frame. The decay rate of the pink noise superposition gain is adjusted from decreasing by 1dB every 5 seconds to decreasing by 0.5dB every 5 seconds to extend the pink noise coverage period. The modulation depth of the neurally-entrained beat audio is increased from 0.4 to 0.55 to enhance the auditory soothing intensity of the beat signal.

[0059] In practice, when the decrease in motion sickness is negative, it indicates that the index value in the updated real-time motion sickness index sequence is still higher than the intervention trigger threshold and the intervention effect is insufficient. The strategy matching module switches the narrative adjustment instruction in the target guidance strategy to a higher intensity scene replacement instruction. It selects the "static rest cabin" scene from the narrative guidance strategy library to replace the current "slow exploration of the reef area" scene, and switches the virtual reality scene as a whole to a fixed rest cabin environment floating on the virtual sea surface. The virtual character stops moving completely, and the target color temperature range in the visual adjustment instruction is shifted to a longer wavelength color temperature direction. The warm color target color temperature range is further shifted from 3500K to 4200K to the orange-red light band of 2800K to 3500K. At the same time, the light and shadow edge softening coefficient is further reduced from 0.35 to 0.15 to eliminate visual detail stimulation.

[0060] It is understandable that when incremental adjustments are made to the intensity of the target guidance strategy, the adjustment operation is to add to the adjustment parameters that are currently being executed, rather than restarting the entire joint adjustment process. The step size change of the visual adjustment command and the gain change of the audio adjustment command take effect from the next frame rendering cycle after the adjustment command is issued. The scene replacement command of the narrative adjustment command will trigger a scene switching transition animation to ensure a smooth transition. All adjustments are seamlessly connected in the adjusted virtual reality experience stream.

[0061] It is understandable that the user's personalized configuration profile plays the role of adjusting the baseline in the incremental correction process. After a user has experienced multiple intervention events, the historical intervention event record set accumulated in the personalized configuration profile will provide a reference for the branch decision of incremental correction. For users who frequently trigger scene replacement commands in the historical records, the effective improvement threshold is adjusted from 0.35 to 0.25 to adapt to their individual differences.

[0062] In practice, after a complete virtual reality experience ends, the system records the real-time motion sickness index sequence, emotional tendency characteristics, the target guidance strategy identifier "VA-001", and the rendering parameters of the adjusted virtual reality experience stream, including the final color temperature value and softening coefficient of the visual presentation parameters, the final gain of pink noise and the final modulation depth of the neural banding beat of the audio playback parameters, and the scene switching identifier and movement speed limit value of the plot narrative parameters. This data is packaged into a single intervention event record and stored in the user's personalized configuration file. The personalized configuration file exists in the device's local storage in the form of a structured file indexed by the user's identity identifier, and contains the user's historical intervention event record set.

[0063] In practice, when a user launches the virtual reality headset and enters any virtual reality application, the system automatically loads the historical intervention event record set from the personalized configuration profile. Based on the frequency of occurrence of the target guidance strategy identifier in the historical intervention event record set, the baseline value of the intervention trigger threshold of the motion sickness assessment model is pre-adjusted. The system also calculates the proportion of occurrences of the composite guidance type identifier "VA-001" in the historical intervention event record set to the total number of interventions. If the proportion exceeds 60%, the baseline value of the intervention trigger threshold is lowered from the default 6.0 to 5.2 to enable earlier intervention to address the user's motion sickness. Based on the distribution center of the rendering parameters of the adjusted virtual reality experience stream in the historical intervention event record set, the initial values ​​of visual presentation parameters, audio playback parameters, and narrative parameters are preset. The initial color temperature value of the visual presentation parameter is set to the weighted average center of the historical color temperature final value, 3840K. The initial pink noise superposition gain of the audio playback parameter is set to the average of the historical final value -18dBFS. The initial movement speed limit of the narrative parameter is set to the historical average of 2.1 meters per second. Users can enjoy personalized comfort parameter settings from the first frame of the virtual reality experience.

[0064] Optionally, when calculating the magnitude of emotion improvement, the strategy matching module uses weighted Euclidean distance instead of standard Euclidean distance. It assigns a 1.5-fold weight to features such as eyebrow curvature and mouth corner angle, and a 1.0-fold weight to features such as eye opening, voice tone fluctuation, and voice energy attenuation. The weighting coefficients are set based on the explicitness of facial expressions in emotion expression, making the magnitude of emotion improvement more sensitive to the degree of facial muscle relaxation. Optionally, a time decay factor is added to each individual intervention event record in the personalized configuration file. The longer the intervention event record is from the current time, the lower the time decay factor. When determining the distribution center of the rendering parameters for the historical intervention event record set, the time decay factor is used as a weighting coefficient for exponential weighted averaging, making the initial values ​​of the pre-set visual presentation parameters, audio playback parameters, and narrative parameters closer to the user's recent physiological state.

[0065] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. A virtual reality motion sickness intervention system based on emotion guidance, characterized in that, The system includes: The data acquisition module acquires a set of multimodal physiological response data of the user of the virtual reality device. The set of multimodal physiological response data includes electroencephalogram (EEG) signals, electrocardiogram (ECG) signals, skin conductance response signals, eye tracking parameters, and head posture data. The motion sickness assessment module inputs the multimodal physiological response data set into a pre-trained motion sickness assessment model to generate a real-time motion sickness index sequence; The emotion analysis module simultaneously collects the user's facial expression image sequence and voice signal segment, and extracts emotional tendency features based on the facial expression image sequence and the voice signal segment; The strategy matching module selects a target guidance strategy from multiple emotion guidance strategies based on the emotion tendency characteristics when the current index value in the real-time motion sickness index sequence exceeds a preset intervention trigger threshold. The scene control module performs joint adjustments on the visual presentation parameters, audio playback parameters, and narrative parameters of the virtual reality scene according to the target guidance strategy, generating an adjusted virtual reality experience stream.

2. The virtual reality motion sickness intervention system based on emotion guidance according to claim 1, characterized in that, Acquire a set of multimodal physiological response data from users of virtual reality devices, specifically including: The user's electroencephalogram (EEG), electrocardiogram (ECG), and skin conductance signals are collected via dry electrodes integrated into the padding of the virtual reality headset. The eye-tracking module built into the virtual reality headset collects data on the user's pupil diameter, gaze point offset, and eyelid flicker frequency. The virtual reality headset's built-in posture sensing module collects data on the user's head center of gravity shift trajectory and head sway frequency. The electroencephalogram (EEG) signal, electrocardiogram (ECG) signal, skin conductance response signal, pupil diameter data, fixation point offset data, eyelid flicker frequency data, head center of gravity offset trajectory data, and head sway frequency data are time-stamped and aligned to form the multimodal physiological response data set.

3. The virtual reality motion sickness intervention system based on emotion guidance according to claim 1, characterized in that, The multimodal physiological response dataset is input into a pre-trained motion sickness assessment model to generate a real-time motion sickness index sequence, specifically including: Baseline drift correction and power frequency notch filtering are performed on each signal channel in the multimodal physiological response dataset to obtain a denoised multimodal physiological feature matrix; The denoised multimodal physiological feature matrix is ​​segmented using a sliding window to generate multiple temporal physiological feature fragments; Each temporal physiological feature segment is sequentially input into a pre-trained motion sickness assessment model, which includes a cascaded convolutional feature extraction layer and a long short-term memory temporal coding layer. The spatial feature map of each temporal physiological feature segment is extracted by the convolutional feature extraction layer, and then the temporal dependency relationship of the spatial feature map is modeled by the long short-term memory temporal coding layer to output the motion sickness index prediction value corresponding to each temporal physiological feature segment. The motion sickness index prediction values ​​corresponding to all time-series physiological feature segments are arranged in chronological order to generate the real-time motion sickness index sequence.

4. The virtual reality motion sickness intervention system based on emotion guidance according to claim 1, characterized in that, The system simultaneously acquires sequences of facial expression images and audio signal segments from the user, and extracts emotional tendency features based on the facial expression image sequences and audio signal segments, specifically including: The user's facial region image sequence is captured by a downward-facing camera built into the virtual reality headset, and the facial region image sequence covers the eye, eyebrow and mouth areas; The user's voice signal segments are collected via the microphone array built into the virtual reality headset. The voice signal segments include spontaneous sighs, hums and short phrases produced by the user. For each frame of the facial region image sequence, facial key point localization processing is performed to extract features of eyebrow curvature, eye opening, and mouth corner angle. The speech signal segment is subjected to fundamental frequency extraction processing and acoustic energy envelope analysis processing to obtain speech pitch fluctuation characteristics and speech energy attenuation characteristics; The features of eyebrow curvature, eye opening, mouth corner angle, voice pitch fluctuation, and voice energy attenuation are concatenated to generate the emotion tendency feature.

5. The virtual reality motion sickness intervention system based on emotion guidance according to claim 1, characterized in that, When the current index value in the real-time motion sickness index sequence exceeds a preset intervention trigger threshold, a target guidance strategy is selected from multiple emotion guidance strategies based on the emotion tendency characteristics, specifically including: Detect whether the current index value in the real-time motion sickness index sequence is greater than a preset intervention trigger threshold; If the current index value is greater than the intervention trigger threshold, the emotional tendency feature is input into the pre-trained emotion classifier to obtain the current emotion category label. The strategy mapping table is used to query the guidance strategy identifier that has the highest matching degree with the current emotion category label. The strategy mapping table stores the correspondence between various emotion category labels and various guidance strategy identifiers in advance. Based on the guidance strategy identifier, at least one guidance strategy is selected from the visual guidance strategy library, the audio guidance strategy library, and the narrative guidance strategy library as the target guidance strategy.

6. The virtual reality motion sickness intervention system based on emotion guidance according to claim 5, characterized in that, Based on the guidance strategy identifier, at least one guidance strategy is selected as the target guidance strategy from the visual guidance strategy library, the audio guidance strategy library, and the narrative guidance strategy library. Specifically, this includes: When the guidance strategy identifier points to a visual guidance type, a color temperature adjustment strategy and a light and shadow softening strategy are selected from the visual guidance strategy library as the target guidance strategy; When the guidance strategy identifier points to an audio guidance type, the pink noise superposition strategy and the neural entrainment beat strategy are selected from the audio guidance strategy library as the target guidance strategy; When the guidance strategy identifier points to the narrative guidance type, a plot branch switching strategy and a movement speed decay strategy are selected from the narrative guidance strategy library as the target guidance strategy; When the guidance strategy identifier points to a composite guidance type, a guidance strategy is selected from the visual guidance strategy library, the audio guidance strategy library, and the narrative guidance strategy library, and the three selected guidance strategies are combined into the target guidance strategy.

7. The virtual reality motion sickness intervention system based on emotion guidance according to claim 5, characterized in that, According to the aforementioned target guidance strategy, the visual presentation parameters, audio playback parameters, and narrative parameters of the virtual reality scene are jointly adjusted to generate an adjusted virtual reality experience stream, specifically including: The visual adjustment instructions contained in the target guidance strategy are analyzed, and the current color temperature value of the virtual reality scene is gradually adjusted to the warm color temperature range according to the visual adjustment instructions, and the sharpness of the light and shadow edges in the virtual reality scene is reduced to below the preset softening coefficient. The audio adjustment instructions contained in the target guidance strategy are analyzed, and a pink noise layer is superimposed on the original audio track of the virtual reality scene according to the audio adjustment instructions. The volume of the neural bandage beat audio is adjusted to a modulation depth that is synchronized with the user's heart rate. The narrative adjustment instructions contained in the target guidance strategy are analyzed, and the current task target in the virtual reality scene is switched from a high-speed movement task to a low-speed exploration task according to the narrative adjustment instructions, and the movement speed of the virtual character is limited to a preset comfortable speed threshold. The virtual reality scene, after color temperature adjustment, light and shadow softening, audio overlay, and narrative switching, is rendered into a continuous stream of the adjusted virtual reality experience.

8. The virtual reality motion sickness intervention system based on emotion guidance according to claim 7, characterized in that, After jointly adjusting the visual presentation parameters, audio playback parameters, and narrative parameters of the virtual reality scene according to the stated target guidance strategy, the process also includes executing: During the playback of the adjusted virtual reality experience stream, the user's updated multimodal physiological response data set, updated facial expression image sequence, and updated speech signal segment are continuously collected; The updated multimodal physiological response dataset is input into the motion sickness assessment model to generate an updated real-time motion sickness index sequence; The updated facial expression image sequence and the updated speech signal segment are input into the emotion classifier to generate updated emotion tendency features; Compare the relationship between the index values ​​in the updated real-time motion sickness index sequence and the intervention trigger threshold, and determine the degree of mood improvement based on the updated mood tendency characteristics; Based on the relationship between the index value and the intervention trigger threshold, as well as the degree of emotional improvement, the intensity of the target guidance strategy is incrementally adjusted.

9. The virtual reality motion sickness intervention system based on emotion guidance according to claim 8, characterized in that, Based on the relationship between the index value and the intervention trigger threshold, and the degree of emotional improvement, the intensity of the targeted guidance strategy is incrementally adjusted, specifically including: The difference between the index value in the updated real-time motion sickness index sequence and the intervention trigger threshold is calculated to obtain the motion sickness reduction value; Calculate the feature distance between the updated sentiment tendency feature and the sentiment tendency feature to obtain the sentiment improvement magnitude value; When the decrease in motion sickness is positive and the improvement in mood exceeds the preset effective improvement threshold, the current application intensity of the target guidance strategy remains unchanged. When the decrease in motion sickness is positive but the improvement in mood does not exceed the effective threshold for improvement, the adjustment step size of the visual adjustment instruction and the superposition gain of the audio adjustment instruction in the target guidance strategy are each increased by one gradient unit. When the motion sickness reduction value is negative, the narrative adjustment instruction in the target guidance strategy is switched to a higher intensity scene replacement instruction, and the target color temperature range in the visual adjustment instruction is shifted towards a longer wavelength color temperature direction.

10. The virtual reality motion sickness intervention system based on emotion guidance according to claim 7, characterized in that, After generating the adjusted virtual reality experience stream, the process also includes executing: Record the real-time motion sickness index sequence, the emotional tendency characteristics, the identification information of the target guidance strategy, and the rendering parameters of the adjusted virtual reality experience stream to form a single intervention event record; The single intervention event record is stored in the user's personalized configuration profile, which contains the user's historical intervention event record set; The next time the user starts the virtual reality device, the set of historical intervention event records in the personalized configuration profile is loaded; Based on the frequency of occurrence of the target guidance strategy identifier in the historical intervention event record set, the baseline value of the intervention trigger threshold of the motion sickness assessment model is pre-adjusted; Based on the distribution center of the rendering parameters of the adjusted virtual reality experience stream in the set of historical intervention event records, the initial values ​​of the visual presentation parameters, the audio playback parameters, and the narrative parameters are preset.