Facial expression information acquisition and reconstruction and emotional physiological response labeling method and system under VR head-mounted display shielding condition

By setting up multimodal sensors on VR headsets to acquire multi-source facial expression data, and using deep neural networks and convolutional neural networks to reconstruct the full-face expression state, the problem of incomplete acquisition of facial expression information and ineffective correlation between facial expression and emotional state and physiological representation in VR environment is solved, achieving high-precision physiological emotion labeling and maintaining the continuity of immersive experience.

CN121910329APending Publication Date: 2026-04-24UNIV OF ELECTRONICS SCI & TECH OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
UNIV OF ELECTRONICS SCI & TECH OF CHINA
Filing Date
2026-01-14
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

In VR environments, existing technologies suffer from compatibility conflicts between emotional state feedback methods and immersive experiences. Facial expression state reconstruction is poor, and there is a lack of effective correlation between facial expression and emotional state and physiological representation, resulting in insufficient accuracy and temporal consistency of emotion labeling.

Method used

Multimodal sensors are used to acquire multi-source facial expression data on VR headsets, including visual dynamic features of unobstructed areas, eye data of occluded areas, and electromyographic signals and contact pressure data of the contact interface. The user's full facial expression state is reconstructed by fusing deep neural networks and convolutional neural networks, and physiological emotion or cognitive state labels are generated using emotion recognition models.

Benefits of technology

It achieves high-precision, non-disruptive facial expression acquisition and physiological response annotation under VR headset occlusion conditions, maintains the continuity of immersive experience, improves the temporal alignment accuracy of facial expression and physiological response data, and reduces the uncertainty and deviation caused by occlusion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121910329A_ABST
    Figure CN121910329A_ABST
Patent Text Reader

Abstract

The invention discloses a facial expression information acquisition and reconstruction and emotional physiological response labeling method and system under a VR head-mounted display shielding condition, solves the problems of compatibility conflict and the like between emotional state feedback and immersive experience in the prior art, and belongs to the crossing field of artificial intelligence, human-computer interaction, virtual reality, psychology and acknowledgement science. According to the method, a multi-modal sensor is arranged on a VR helmet, and full facial expression data under the condition that the VR helmet is partially shielded are obtained, namely, facial expression multi-source data are obtained, and the facial expression multi-source data comprise eye data of a physical shielding area, electromyographic signals and contact pressure data captured by a contact interface and visual dynamic feature data of a non-shielding area; and fusing the visual dynamic feature data of the non-occlusion area, the eye data in the physical occlusion boundary, and the electromyographic signal and the contact pressure data captured by the contact interface, and constructing a full face expression state of the user. The method is used for realizing emotional physiological response labeling according to facial expressions under the condition of VR helmet shielding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] A method and system for acquiring and reconstructing facial expression information and annotating emotional physiological responses under VR headset occlusion conditions is provided. This method is used to annotate emotional physiological responses based on facial expressions under VR headset occlusion conditions. It belongs to the interdisciplinary field of artificial intelligence, human-computer interaction, virtual reality, psychology and cognitive science. Background Technology

[0002] Virtual reality (VR) technology, with its unique immersion, interactivity, and imagination, has become an important platform for research in emotion recognition, cognitive science, and human-computer interaction. VR head-mounted displays (HMDs) can construct high-dimensional immersive experimental scenarios through multi-sensory channels and panoramic displays, effectively shielding external interference and giving subjects a strong sense of presence, thereby inducing more realistic, continuous, and profound emotional experiences. Furthermore, VR environments possess high ecological validity and controllability, providing repeatable, low-cost, and safe experimental conditions, offering an ideal new environment for quantifying emotion recognition.

[0003] In existing research on emotion or cognitive state recognition, methods based on physiological representations (including physiological signals such as EEG, physiological indicators such as heart rate, and eye movement information) have become mainstream technologies in the field of emotion and cognitive computing due to their core advantages such as objectively mapping users' intrinsic physiological characteristics, strong temporal continuity, and independence from subjective feedback. EEG, GSR, heart rate (HR), skin temperature, and eye movement indicators can effectively represent users' emotional fluctuations and cognitive states, and are widely applicable to diverse human-computer interaction scenarios. Against this backdrop, applying physiological representations to emotion and cognitive recognition in VR environments has become an emerging and promising research direction with important applications in health, medicine, and rehabilitation. However, the effectiveness of physiological modeling of emotion or cognition still highly depends on the support of large-scale, high-confidence physiological datasets and emotional or cognitive response state labels / annotations (hereinafter referred to as "emotion labels").

[0004] Existing emotion labeling methods are mainly divided into two categories: "self-assessment labeling" and "third-party labeling." Traditional self-assessment labeling typically requires subjects to conduct subjective evaluations during or after the experiment using interactive media such as buttons (e.g., mouse, keyboard), scales, or questionnaires. While this method has advantages such as ease of operation and low deployment costs, its effectiveness heavily relies on the subject's subjective awareness and retrospective evaluation, making it highly susceptible to interference from social expectation effects, self-presentation bias, and memory decay of emotional / cognitive states. This often leads to significant discrepancies between the labeling results and physiological data, making it difficult to support the objective, real-time data requirements of high-precision emotion modeling. For example, a subject's actual emotion might be "sadness," but they subjectively believe "happiness" is more reasonable, so they label it "happiness," resulting in a complete mismatch between physiological data and label data.

[0005] In contrast, third-party annotation aims to circumvent the subjective cognitive filtering of subjects through an external observational perspective. This method utilizes real-time judgments by experts or experimenters on subjects' overt behaviors (such as facial expressions, body movements, and speech / voice) for annotation, largely overcoming reporting bias caused by memory decay. However, third-party annotation is not absolutely objective: on the one hand, subjects may engage in behavioral masquerading under monitored conditions, such as feigning composure during stress tests; on the other hand, the annotation results are highly dependent on the observer's subjective inferences.

[0006] Furthermore, in a VR environment, the reliability of this "external inference" (or facial expression recognition) is further weakened: due to the physical occlusion of the periocular face by head-mounted devices (HMDs), the most expressive upper facial features of the subjects are blocked (missed), resulting in a severe lack of information for third-party annotators. This makes observational methods originally intended to reduce bias difficult to implement in VR, and instead increases the randomness and uncertainty of annotation due to insufficient evidence, making it impossible to provide high-reliability temporal alignment labels for physiological signal modeling.

[0007] In addition, facial expression recognition technology based on computer vision is also widely used for the recognition of physiological emotions or cognitive states. The recognition results can also be used to label physiological data of physiological emotions or cognitive responses, but it also requires the subject's face to be clearly visible, so it is not suitable for VR environments.

[0008] Existing technologies attempt to directly transfer annotation methods from planar environments to VR environments, such as spatially mapping traditional planar annotation tools (e.g., SAM scale, sliders) to adapt them to virtual reality scenes. However, this method often pops up the annotation interface during the critical window of emotion induction, forcibly interrupting the subject's immersive experience. This interference not only disrupts the continuity of emotion generation but may even trigger the subject's reverse emotional resistance, causing the annotation values ​​to deviate from the actual induction state. This results in poor temporal consistency and repeatability of the acquired label data, limiting its application in real-time emotion monitoring systems.

[0009] Some technologies, by introducing 3D face modeling, attention mechanisms, and image inpainting techniques, achieve feature compensation and accuracy restoration for occluded areas. In VR environments, they utilize limited visible areas to drive 3D virtual avatars, maintaining visual consistency during interaction. However, this approach primarily focuses on facial reconstruction in the computer vision (CV) dimension, with its technological endpoint limited to "expression reconstruction" or "visual classification." Essentially, it still serves virtual social interaction or film / television-driven applications, not expression recognition. It doesn't involve expression recognition based on facial reconstruction, nor does it establish a connection between reconstructed expression information and underlying physiological representations (such as EEG and electrodermatology) (emotional / cognitive physiological data annotation). Therefore, it cannot provide effective physiological data annotation for high-precision physiological emotion models. Furthermore, these methods heavily rely on facial geometric topology (such as facial meshes and muscle deformation models). Because feature extraction originates from physical deformation information, it is insufficient for depicting deep, subtle emotional cues (such as micro-expression differences and the fineness of emotional arousal), making it difficult to meet the needs of affective computing for high-dimensional emotional state mining.

[0010] In summary, the existing technology has the following technical problems: The “compatibility conflict” between emotional state feedback methods and immersive experiences: In the current VR environment, the annotation of emotional state data only relies on the “explicit interaction” feedback method. That is, in the HMD display interface, the subject uses the on-screen floating UI prompts (SAM table or scale) and the interactive device (mouse, gamepad, etc.) to provide feedback or annotation on the induced emotional / cognitive feelings (e.g., filling in the scale). This actually leads to a “task switching” state. At this time, the subject will temporarily stop watching the induced content and also pause the identification of emotional state, and focus on the emotional feedback or filling in the scale. This will cause the subject’s temporary interruption of the induced response, causing interference with the emotional / cognitive state. 2. The reconstruction of full-face expression under partial occlusion conditions is poor, making it difficult to fully capture complete facial expression and emotional state, and unable to achieve high-precision temporal alignment with physiological response data: Current visual information-based 3D facial reconstruction heavily relies on vision (camera) to capture facial geometric topological deformations and construct facial meshes and facial muscle deformation models. When key facial information (upper face and periorbital area) is missing, this method is difficult to reconstruct or may fail. In addition, because VRHMDs have a fixed shape, they compress the subject's face after wearing, causing non-emotional deformations and introducing a large amount of noise data, which greatly affects the accuracy of facial expression recognition. Although some studies have adopted feature compensation and accuracy restoration measures for occluded areas, when key areas are occluded or the occlusion is large, and the visible areas are greatly affected by noise, existing solutions struggle to capture deep and subtle emotional changes (such as dynamic gradient changes in expressions). This will result in the reconstructed facial expressions failing to achieve high-precision temporal alignment with the physiological response data reflecting emotions obtained at high sampling rates.

[0011] 3. There is an invalid association between facial expressions / emotional states and physiological representations: Without establishing an effective correlation or mapping between facial expressions and electrophysiological signals (EEG, heart rate, EMG, etc.) and eye movement signals that reflect physiological emotions and cognitive states, it is impossible to use the emotional state of facial expression recognition to label the physiological data of emotions. Summary of the Invention

[0012] The purpose of this invention is to provide a method and system for acquiring and reconstructing facial expression information and annotating emotional and physiological responses under VR headset occlusion conditions, which solves the "compatibility conflict" between "explicit interaction" emotional state feedback and immersive experience in the prior art; the low accuracy of facial reconstruction and expression recognition under partial occlusion conditions; and the invalid correlation between expression information and physiological representation.

[0013] To achieve the above objectives, the technical solution adopted by the present invention is as follows: A method for acquiring and reconstructing facial expression information under VR headset occlusion conditions includes the following steps: Step 1: Obtain multi-source facial expression data: By setting up multimodal sensors on VR headsets, we can acquire full facial expression data under conditions where the VR headset partially obscures the face, thus obtaining multi-source facial expression data, including eye data in physically obscured areas and electromyographic signals and contact pressure data captured by the contact interface, as well as visual dynamic feature data of unobscured areas. Step 2: Reconstruct the user's full facial expression state based on multi-source facial expression data: The visual dynamic feature data of the unobstructed area, the eye data within the physical occlusion boundary, and the electromyographic signals and contact pressure data captured by the contact interface are fused to construct the user's full facial expression state.

[0014] Furthermore, the specific steps of step 1 are as follows: Step 1.1: Using the visual sensor set on the outside of the VR headset, capture the dynamic feature data of the lower half of the face that is not covered by the VR headset, that is, obtain the visual dynamic feature data of the uncovered area. Step 1.2: Use the visual sensor built into the outer periphery of the VR headset lens to obtain local state data of the eye and the surrounding area, that is, obtain eye data, including gaze direction, blink frequency, eyelid opening and closing state and pupil changes. The visual sensor on the outer periphery of the VR headset lens refers to the camera in the eye tracking module built into the VR headset. Step 1.3: Using the electromyography (EMG) sensor and flexible pressure sensor on the VR headset gasket, acquire non-visual facial muscle activity data at the HMD gasket contact interface, including EMG signals and contact pressure data captured at the contact interface.

[0015] Furthermore, in step 1.1, the visual sensor is an external camera, and the layout includes a single-camera structure protruding from the front of the head-mounted display and a symmetrical dual-camera arrangement structure below the head-mounted display; The front-mounted single-camera structure of the head-mounted display uses a protruding or folding bracket to fix an external camera to the front of the head-mounted display, so that the lens faces the wearer's face; The headset features a symmetrical dual-camera setup on the left and right sides of the lower part of the headset.

[0016] Furthermore, in step 1.3, the electromyography sensor and flexible pressure sensor on the VR helmet gasket include at least an embedded portion that is embedded within the VR helmet gasket and connected by an internal flexible connection structure; The embedded part includes an internal flexible connection structure embedded in the VR headset gasket, located on the user's forehead. The internal flexible connection structure is equipped with at least five sets of sensors: one set at the center of the forehead, at least two sets on each side of the center of the forehead corresponding to the two ends of the user's brow bones, one set of sensors on each side of the internal flexible connection structure corresponding to the area below each eye, one set of sensors on each side of the outer corner of each eye, one set of sensors on each side of the user's nose bridge, and one electromyography (EMG) sensor on each side of the internal flexible connection structure. Each set of sensors includes a flexible pressure sensor and an EMG sensor located on both sides of the flexible pressure sensor. It may also include peripheral components located outside the VR headset gasket; The peripheral device includes an external flexible connection structure connected to the internal flexible connection structure. At least one set of sensors is provided on the external flexible connection structure located above the forehead and below the hairline, and at least one set of sensors is provided on the user's cheek. The flexible pressure sensor captures local contact pressure fluctuations caused by facial muscle activity in real time to obtain contact pressure data. Electromyography (EMG) sensors monitor the electromyographic signals of facial muscles in the occluded area, directly reflecting the contraction and relaxation state of the muscles.

[0017] Furthermore, the specific steps of step 2 are as follows: Step 2.1: Construct a set of facial expressions based on basic emotions. , Indicates the first There are 10 basic emotions and their corresponding characteristics. The basic emotions include happiness, sadness, anger, fear, disgust, surprise, and contempt. The characteristics corresponding to happiness include upturned corners of the mouth, slightly curved eyes, crow's feet at the corners of the eyes, raised cheeks, and natural contraction of the muscles around the eyes. The characteristics corresponding to sadness include raised inner eyebrows, downturned corners of the mouth, and dim eyes. The characteristics corresponding to anger include lowered and gathered eyebrows, wide eyes, tightly closed or open lips, flared nostrils, flushed face, and tense muscles. The characteristics corresponding to fear include raised and gathered eyebrows, wide eyes, and horizontally stretched lips. The characteristics corresponding to disgust include wrinkled nose, upturned upper lip, and downturned corners of the mouth. The characteristics corresponding to surprise include raised eyebrows, wide eyes, and downturned jaw. The characteristic corresponding to contempt is an upturned corner of the mouth on one side. Step 2.2: Based on facial expression sets In a planar environment where the entire face is visible, facial expression images of subjects are collected to construct a planar facial expression data dataset. ; In a VR environment, where the face is occluded, multi-source facial expression data of the same subject is collected using multimodal sensors mounted on a VR headset to construct a facial expression information dataset for the VR environment. ; Step 2.3: Use a deep neural network to build and The mapping relationship is expressed by the formula: in, Represents a deep neural network. Represents a dataset of planar facial expression information Location dimensional vector space, Represents facial expression information dataset Location Dimensional vector space; Step 2.4: Eye data from the multi-source facial expression data obtained in Step 1. and visual dynamic feature data Using convolutional neural networks or feature point extraction algorithms to... Eye data at any time and visual dynamic feature data The mapping to the corresponding feature vectors, using the formula of a convolutional neural network, is as follows: in, express The eye feature vector at time step, express Facial feature vector at time step, Represents a convolutional neural network. Indicates the location of the eye feature vector dimensional vector space, This indicates the location of the facial feature vector. Dimensional vector space; Based on the non-visual facial muscle activity data from the multi-source facial expression data obtained in step 1 For each raw signal sequence from a flexible pressure sensor or electromyography sensor, within a sliding time window Extract time-domain or frequency-domain features to form sub-feature vectors, and concatenate the sub-feature vectors of all channels to obtain... The eigenvector corresponding to time step 1 is given by the formula: in, Represents the feature extraction function. express The feature vector of electromyographic signals and contact pressure data captured at the interface in real time, with dimensions... It depends on the number of channels multiplied by the number of features extracted from each channel. This indicates the number of electromyography (EMG) sensors, specifically the number of channels in the EMG sensors. This indicates the number of flexible pressure sensors, specifically the number of channels in the flexible pressure sensor. express Constantly contacting the interface to capture electromyographic signals, express Constantly monitor stress data. Indicates the first Sub-feature vectors of each electromyography sensor, Indicates the first Sub-feature vectors of a flexible pressure sensor Indicates transpose; Step 2.5: Align the results obtained in Step 2.4 on a unified time axis, and then construct a matrix of multi-source observation vectors. The formula is: in, Indicates transpose; Step 2.6, The user's full facial expression status at any given moment is defined as follows: ; Step 2.7: Matrix pairing based on multi-source observation fusion mapping function and multi-source observation vectors The formula for estimating the user's full facial expression at any given time is: in, The estimated value is... Real-time facial expression status of all users. The matrix representing the weight matrix corresponding to the aligned multi-source observation vectors. and The settings are dynamically adjusted based on how much of the subject's face is obscured by the VR headset; when the obscured area is large... Increased weight Decrease the weight, and vice versa. The weight decreases. Weight increases, The value is the total area of ​​the washer. Area of ​​the effective facial expression region of the subject The ratio value, Values , and ; Step 2.8: Based on the mapping relationship obtained in Step 2.3, the estimated... User's full facial expression status at any time Mapped to a planar environment, the reconstructed image is obtained. User's full facial expression status at any time The formula is: .

[0018] A method for labeling emotional physiological responses includes the following steps: Step S1: Generate physiological emotion or cognitive state labels based on the reconstructed user's full facial expression state: Inputting the user's full facial expressions into a pre-trained emotion recognition model or cognitive state recognition model, the system determines the subject's real-time physiological emotion or cognitive state and generates physiological emotion or cognitive state labels; or After the experiment, the sequence of users' full facial expression states will be provided to third-party professional evaluators to generate physiological emotion or cognitive state labels. Step S2: Label physiological response data using physiological emotion or cognitive state tags: By sharing a unified system clock, the online collected physiological response data and the reconstructed user facial expression state are given a time stamp that is consistent with the time. The physiological emotion or cognitive state label is automatically mapped to the physiological response data with the corresponding time stamp, so as to realize the labeling of physiological emotion or cognitive state. The physiological response data includes EEG, GSR and ECG.

[0019] A facial expression acquisition and reconstruction system under VR headset occlusion conditions includes: VR head-mounted display devices: By setting up multimodal sensors on VR headsets, full facial expression data is obtained under conditions where the VR headset partially obscures the face, thus obtaining multi-source facial expression data, including eye data in physically obscured areas and electromyographic signals and contact pressure data captured by the contact interface, as well as visual dynamic feature data of unobscured areas. Facial expression reconstruction module: It fuses the visual dynamic feature data of the unobstructed area, the eye data within the physical occlusion boundary, and the electromyographic signals and contact pressure data captured by the contact interface to construct the user's full facial expression state.

[0020] Furthermore, the VR head-mounted display device includes a VR helmet, a visual sensor disposed on the outside of the VR helmet for capturing dynamic feature data of the lower half of the face not obscured by the VR helmet, a visual sensor built into the periphery of the VR helmet lens to acquire local state data of the eyes and the area around the eyes, and an electromyography sensor and a flexible pressure sensor embedded in the VR helmet gasket to acquire non-visual facial muscle activity data of the HMD gasket contact interface. The visual sensors located on the outside of the VR headset are external cameras, and the layout includes a single camera structure protruding from the front of the headset and a symmetrical dual-camera arrangement below the headset. The front-mounted single-camera structure of the head-mounted display uses a protruding or folding bracket to fix an external camera to the front of the head-mounted display, so that the lens faces the wearer's face; The dual-camera symmetrical arrangement structure at the bottom of the headset features two external cameras symmetrically installed on the left and right sides of the lower part of the headset. The electromyography sensor and flexible pressure sensor on the VR headset gasket include at least the embedded portion that is connected within the VR headset gasket by an internal flexible connection structure. The embedded part includes an internal flexible connection structure embedded in the VR headset gasket, located on the user's forehead. The internal flexible connection structure is equipped with at least five sets of sensors: one set at the center of the forehead, at least two sets on each side of the center of the forehead corresponding to the two ends of the user's brow bones, one set of sensors on each side of the internal flexible connection structure corresponding to the area below each eye, one set of sensors on each side of the outer corner of each eye, one set of sensors on each side of the user's nose bridge, and one electromyography (EMG) sensor on each side of the internal flexible connection structure. Each set of sensors includes a flexible pressure sensor and an EMG sensor located on both sides of the flexible pressure sensor. It may also include peripheral components located outside the VR headset gasket; The peripheral device includes an external flexible connection structure connected to the internal flexible connection structure. At least one set of sensors is provided on the external flexible connection structure located above the forehead and below the hairline, and at least one set of sensors is provided on the user's cheek. The flexible pressure sensor captures local contact pressure fluctuations caused by facial muscle activity in real time to obtain contact pressure data. Electromyography (EMG) sensors monitor the electromyographic signals of facial muscles in the occluded area, directly reflecting the contraction and relaxation state of the muscles.

[0021] Furthermore, the specific implementation steps of the facial expression reconstruction module are as follows: Step 2.1: Construct a set of facial expressions based on basic emotions. , Indicates the first There are 10 basic emotions and their corresponding characteristics. The basic emotions include happiness, sadness, anger, fear, disgust, surprise, and contempt. The characteristics corresponding to happiness include upturned corners of the mouth, slightly curved eyes, crow's feet at the corners of the eyes, raised cheeks, and natural contraction of the muscles around the eyes. The characteristics corresponding to sadness include raised inner eyebrows, downturned corners of the mouth, and dim eyes. The characteristics corresponding to anger include lowered and gathered eyebrows, wide eyes, tightly closed or open lips, flared nostrils, flushed face, and tense muscles. The characteristics corresponding to fear include raised and gathered eyebrows, wide eyes, and horizontally stretched lips. The characteristics corresponding to disgust include wrinkled nose, upturned upper lip, and downturned corners of the mouth. The characteristics corresponding to surprise include raised eyebrows, wide eyes, and downturned jaw. The characteristic corresponding to contempt is an upturned corner of the mouth on one side. Step 2.2: Based on facial expression sets In a planar environment where the entire face is visible, facial expression images of subjects are collected to construct a planar facial expression data dataset. ; In a VR environment, where the face is occluded, multi-source facial expression data of the same subject is collected using multimodal sensors mounted on a VR headset to construct a facial expression information dataset for the VR environment. ; Step 2.3: Use a deep neural network to build and The mapping relationship is expressed by the formula: in, Represents a deep neural network. Represents a dataset of planar facial expression information Location dimensional vector space, Represents facial expression information dataset Location Dimensional vector space; Step 2.4: Eye data from the multi-source facial expression data obtained in Step 1. and visual dynamic feature data Using convolutional neural networks or feature point extraction algorithms to... Eye data at any time and visual dynamic feature data The mapping to the corresponding feature vectors, using the formula of a convolutional neural network, is as follows: in, express The eye feature vector at time step, express Facial feature vector at time step, Represents a convolutional neural network. Indicates the location of the eye feature vector dimensional vector space, This indicates the location of the facial feature vector. Dimensional vector space; Based on the non-visual facial muscle activity data from the multi-source facial expression data obtained in step 1 For each raw signal sequence from a flexible pressure sensor or electromyography sensor, within a sliding time window Extract time-domain or frequency-domain features to form sub-feature vectors, and concatenate the sub-feature vectors of all channels to obtain... The eigenvector corresponding to time step 1 is given by the formula: in, Represents the feature extraction function. express The feature vector of electromyographic signals and contact pressure data captured at the interface in real time, with dimensions... It depends on the number of channels multiplied by the number of features extracted from each channel. This indicates the number of electromyography (EMG) sensors, specifically the number of channels in the EMG sensors. This indicates the number of flexible pressure sensors, specifically the number of channels in the flexible pressure sensor. express Constantly contacting the interface to capture electromyographic signals, express Constantly monitor stress data. Indicates the first Sub-feature vectors of each electromyography sensor, Indicates the first Sub-feature vectors of a flexible pressure sensor Indicates transpose; Step 2.5: Align the results obtained in Step 2.4 on a unified time axis, and then construct a matrix of multi-source observation vectors. The formula is: in, Indicates transpose; Step 2.6, The user's full facial expression status at any given moment is defined as follows: ; Step 2.7: Matrix pairing based on multi-source observation fusion mapping function and multi-source observation vectors The formula for estimating the user's full facial expression at any given time is: in, The estimated value is... Real-time facial expression status of all users. The matrix representing the weight matrix corresponding to the aligned multi-source observation vectors. and The settings are dynamically adjusted based on how much of the subject's face is obscured by the VR headset; when the obscured area is large... Increased weight Decrease the weight, and vice versa. The weight decreases. Weight increases, The value is the total area of ​​the washer. Area of ​​the effective facial expression region of the subject The ratio value, Values , and ; Step 2.8: Based on the mapping relationship obtained in Step 2.3, the estimated... User's full facial expression status at any time Mapped to a planar environment, the reconstructed image is obtained. User's full facial expression status at any time The formula is: .

[0022] An emotional physiological response annotation system, comprising: Label recognition module: Based on the reconstructed full-face facial expression state inputted into a pre-trained emotion recognition model or cognitive state recognition model, it determines the subject's real-time physiological emotion or cognitive state and generates physiological emotion or cognitive state labels; or After the experiment, the sequence of users' full facial expression states will be provided to third-party professional evaluators to generate physiological emotion or cognitive state labels. Labeling module: By sharing a unified system clock, the online collected physiological response data and the reconstructed user facial expression state are assigned a time stamp that is consistent with the time. The physiological emotion or cognitive state label is automatically mapped to the physiological response data with the corresponding time stamp, so as to realize the labeling of physiological emotion or cognitive state labels. The physiological response data includes EEG, GSR and ECG.

[0023] Compared with the prior art, the advantages of the present invention are as follows: This invention addresses the critical issue of obtaining stable and reliable annotation of emotion-related physiological signals under virtual reality headset occlusion and immersive interaction conditions. It systematically reconstructs the process from three levels: perception path, user full-face expression state reconstruction, and annotation mechanism. A new emotion annotation technology system with full-face expression state as the core carrier is constructed. By organically combining multi-source facial expression perception for occluded scenarios, user full-face expression state reconstruction, and expression-based emotion annotation methods, this invention achieves the engineering implementation of the emotion annotation process without relying on explicit interaction or interfering with the immersive experience. It realizes a means to implicitly and non-intrusively obtain feedback from subjects on induced emotions or cognitive feelings, enabling high-frequency, high-quality, and non-intrusive annotation while maintaining "deep immersion" in watching induced content. Specifically, this is manifested in: I. Reconstruct the facial expression acquisition path under head-mounted display occlusion conditions to address the current issue of incomplete facial expression information acquisition: To address the issue of structural occlusion of the upper face and eye area caused by virtual reality head-mounted displays, and the failure of traditional expression acquisition methods, this invention constructs a multi-source expression perception path for occluded scenarios. By introducing eye-tracking information, visual information of the lower face, and tactile perception information of the contact pad area, it achieves continuous acquisition of complete facial expression states. This method no longer relies on the strong coupling assumption between the unoccluded and occluded areas, making it feasible to acquire facial expressions stably and continuously in real immersive VR use. It fundamentally solves the problem of incomplete facial expression information under occlusion conditions. Second, by using multi-source perception-driven full-face expression state reconstruction (i.e., achieving facial expression reconstruction), the effective correlation between facial expression and emotional state (physiological emotional or cognitive state data) and physiological response data is improved in occluded scenarios: To address the issue of insufficient stability in existing facial expression reconstruction methods under occlusion conditions, this invention uses eye-tracking information, electromyography information, and contact pressure data as direct or semi-direct observation channels for the facial expression state of the occluded area. These are then fused with visual information from the lower half of the face to achieve joint reconstruction of the entire facial expression state. This reconstruction method effectively reduces the uncertainty and systematic bias introduced by occlusion, enabling the reconstructed full-face expression state to maintain higher stability and consistency under different interaction conditions. This provides a reliable explicit representation basis for subsequent emotion analysis, facilitating the effective correlation between facial expression and physiological response data through subsequent high-precision temporal alignment. Third, by replacing interactive annotation with automatic annotation, an emotion annotation mechanism is constructed that does not interfere with the immersive experience, thus overcoming the "compatibility conflict" between existing emotion state feedback methods and immersive experiences. To address the issues of explicit emotion labeling methods based on floating interfaces or pop-up controls in virtual reality environments easily interfering with immersive experiences and lacking operational stability, this invention uses the reconstructed full-face expression state as the emotional outward carrier. It decouples the labeling process of physiological emotions or cognitive states from the subject's explicit interactive behavior, allowing it to run in parallel with the virtual reality experience as a system background perception mechanism. This approach avoids frequent switching between emotional experience and rational evaluation, reduces attention shifts and interactive interference, and makes the emotion labeling process more natural and stable, thus helping to maintain the authenticity of the emotion induction process.

[0024] Fourth, the VR headset in this invention has a large field of view, and the external visual sensors can maintain the same motion state as the headset, which can effectively reduce the blind spots in the field of view generated by traditional fixed-position cameras when the head moves. It can ensure that the unobstructed facial area can output effective data throughout the data acquisition process. The visual sensors on the periphery of the VR headset lens can also acquire eye data, and the setting of electromyography sensors and flexible pressure sensors can capture facial muscle activity in real time. This is a mechanism-level monitoring that can capture weaker and more subtle facial expression changes. At the same time, these sensors can extend to the boundary between the obstructed and unobstructed areas, which can establish a mapping relationship between facial muscle movement and visual expression changes, and ensure the acquisition of facial expression data without any blind spots. Attached Figure Description

[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as a limitation on the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a schematic diagram of the process of the present invention; Figure 2 This is a schematic diagram of the front-mounted single camera structure of the VR headset in this invention. In the figure shown in ③, the square structure is a flexible pressure sensor and the circular structure is an electromyography sensor. Figure 3 This is a schematic diagram of the symmetrical arrangement of the two cameras below the VR headset in this invention; Figure 4 This is a schematic diagram of data collection in the bonding washer area of ​​the present invention; Detailed Implementation

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] To address the issues of interference with user experience (including manipulation and emotion induction) caused by explicit interactions (such as visual scales) in immersive environments, the lack of facial expression information due to HMD facial occlusion, and the lack of correlation between reconstructed facial expression information and physiological information, this invention proposes a method for facial expression reconstruction under VR headset occlusion conditions, and a method for labeling physiological emotion or cognitive state data based on this reconstruction.

[0029] High-precision facial expression reconstruction is achieved using multi-source data (visual, electromyographic, tactile, etc.). Based on the physiological condition or cognitive state label of the reconstructed facial expression (i.e., the user's full-face expression state), (2) the physiological emotion or cognitive state label based on the user's full-face expression state is used for labeling physiological response data, replacing the "display interaction" label. This achieves high-reliability, continuous automatic or third-party assisted labeling of physiological emotion or cognitive state and physiological response data.

[0030] A method for acquiring and reconstructing facial expression information under VR headset occlusion conditions includes the following steps: This solution addresses the challenge of capturing full-face expressions due to occlusion caused by HMD devices by constructing a multi-source facial expression data acquisition system. Instead of relying on a single visual channel, this solution utilizes a visual-pressure / tactile-electromyographic (EMG) multi-source expression information detection approach. Cameras (including but not limited to eye-tracking cameras) are placed on the outside and inside of the VR headset, and EMG / flexible pressure sensors are embedded in the headset's gasket. Multi-source facial expression data is acquired using visual, EMG, and tactile detection technologies to reconstruct full-face expressions and identify emotions or cognitive responses based on these expressions. Specifically, by setting up multimodal sensors on the VR headset, full-face expression data is acquired even under partial occlusion conditions, resulting in multi-source facial expression data, including eye data from physically occluded areas, EMG signals and pressure data captured from the contact interface, and visual dynamic feature data from unoccluded areas. The specific steps are as follows: Step 1.1, Extraction of visual expression data of unobstructed face: Using the visual sensor set on the outside of the VR headset, capture the dynamic feature data of the lower half of the face that is not obstructed by the VR headset, that is, obtain the visual dynamic feature data of the unobstructed area, including mouth shape, lip shape changes, jaw contour and cheek muscle deformation features, which are used to establish the basic framework for facial and expression reconstruction. The visual sensor is an external camera. Depending on the specific interaction requirements and structural design, the layout can include a single camera structure protruding from the front of the headset and a symmetrical dual-camera arrangement below the headset. The front-mounted single-camera structure of the headset uses a protruding or folding bracket to fix an external camera to the front of the headset, so that the lens faces the wearer's face. This structure has a large field of view and can stably capture macroscopic motion features of the lower half of the face, such as... Figure 1As shown; The dual-camera symmetrical arrangement below the headset features two external cameras symmetrically mounted on the left and right sides of the lower part of the headset. This layout effectively avoids the blind spots that may exist with a single camera, improving the geometric accuracy and robustness of image acquisition. Figure 2 As shown.

[0031] Step 1.2: Acquire eye data within the physical occlusion boundary (area invisible outside the HMD but visible inside the HMD): Using the visual sensors built into the periphery of the VR headset lens, acquire local state data of the eyes and surrounding area, i.e., obtain eye data, including gaze direction, blink frequency, eyelid opening and closing state, and pupil changes. This compensates for the lack of facial expression information of the eyes and surrounding area in the visible facial expression data. The visual sensors on the periphery of the VR headset lens refer to the cameras in the eye-tracking module built into the VR headset. In conjunction with the eye-tracking module built into the HMD, real-time acquisition of the subject's gaze direction, blink frequency, eyelid opening and closing state, and pupil changes, etc., is used to capture subtle emotional fluctuations generated by the subject during immersive interaction. Because the HMD device provides long-term stable physical occlusion of the subject's forehead and periorbital area, traditional visual pathways fail in this area. The eye-tracking module bypasses the hardware shell's obstruction, enabling non-contact direct observation of the eyes and surrounding area, compensating for upper facial expression information that external cameras cannot perceive.

[0032] Step 1.3: Acquire non-visual facial data of the HMD washer contact interface: Using electromyography (EMG) sensors and flexible pressure sensors embedded in the VR headset washer, acquire non-visual facial muscle activity data of the HMD washer contact interface, including EMG signals and contact pressure data captured at the contact interface.

[0033] The electromyography sensor and flexible pressure sensor on the VR headset gasket include at least the embedded portion that is connected within the VR headset gasket by an internal flexible connection structure. The embedded part includes an internal flexible connection structure embedded in the VR headset gasket, located on the user's forehead. The internal flexible connection structure is equipped with at least five sets of sensors: one set at the center of the forehead, at least two sets on each side of the center of the forehead corresponding to the two ends of the user's brow bones, one set of sensors on each side of the internal flexible connection structure corresponding to the area below each eye, one set of sensors on each side of the outer corner of each eye, one set of sensors on each side of the user's nose bridge, and one electromyography (EMG) sensor on each side of the internal flexible connection structure. Each set of sensors includes a flexible pressure sensor and an EMG sensor located on both sides of the flexible pressure sensor. It may also include peripheral components located outside the VR headset gasket; The peripheral device includes an external flexible connection structure connected to the internal flexible connection structure. At least one set of sensors is provided on the external flexible connection structure located above the forehead and below the hairline, and at least one set of sensors is provided on the user's cheek. The flexible pressure sensor captures local contact pressure fluctuations caused by facial muscle activity in real time to obtain contact pressure data. Electromyography (EMG) sensors monitor the electromyographic signals of facial muscles in the obscured area, directly reflecting the contraction and relaxation states of the muscles. Flexible pressure sensors and EMG sensors can be installed individually or in groups as needed. The materials used for the inner and outer flexible connection structures are not limited; for example, rubber strips can be used, with the appropriate rubber strips positioned according to requirements.

[0034] This study utilizes electromyography (EMG) sensors and flexible pressure sensors to directly acquire the EMG or deformation data of facial muscles reflecting facial expressions. The correlation and mapping relationship between EMG / pressure data caused by facial expression changes in this area and facial deformation is established. Figure 3 As shown, for core areas such as the forehead, eyebrows, and around the eyes that are obscured by the head-mounted display (HMD) pad, this invention achieves indirect observation of the facial areas in contact with the HMD by integrating non-visual sensors within the pad: Flexible pressure sensors: By integrating a flexible pressure sensor array, local contact pressure fluctuations caused by facial muscle activities (such as frowning and squinting) are captured in real time. Electromyography (EMG) sensors: These sensors monitor the electrical activity signals of facial muscles in the obscured areas, directly reflecting the contraction and relaxation states of the muscles. This subsystem effectively supplements deep facial expression cues that cannot be directly perceived by the visual channel, completing the information missing due to structural occlusion.

[0035] To eliminate the uncertainty caused by physical occlusion, a multi-source facial expression data reconstruction of the user's full facial expression state is constructed: based on the visual dynamic features of the aforementioned unoccluded areas, eye data within the physical occlusion boundaries, and electromyographic signals and pressure data captured by the contact interface, a multi-source facial data-driven reconstruction of the user's full facial expression state is built. This not only eliminates the "facial blind spot" in the VR environment but also realizes the usability and credibility of facial expressions as a carrier of emotional / cognitive states in immersive interactive scenarios, thus laying a solid data foundation for the subsequent automatic / third-party annotation of high-reliability physiological response data. Specifically, the visual dynamic feature data of the unoccluded areas, eye data within the physical occlusion boundaries, and electromyographic signals and contact pressure data captured by the contact interface are fused to construct the user's full facial expression state. The specific steps are as follows: Step 2.1: Construct a set of facial expressions based on basic emotions. , Indicates the first There are four basic emotions and their corresponding characteristics. Basic emotions include happiness, sadness, anger, fear, disgust, surprise, and contempt. Characteristics of happiness include upturned corners of the mouth, slightly curved eyes, crow's feet at the corners of the eyes, lifted cheeks, and natural contraction of the muscles around the eyes. A genuine smile is usually accompanied by this natural contraction of the muscles around the eyes. Characteristics of sadness include raised inner eyebrows, downturned corners of the mouth, and dim eyes; it is often associated with feelings of loss and frustration. Characteristics of anger include lowered and furrowed eyebrows, wide eyes, tightly closed or open lips, flared nostrils, facial redness, and muscle tension. Characteristics of fear include raised and furrowed eyebrows, wide eyes, and horizontally stretched lips, possibly with an open mouth; similar to surprise, but fear is usually accompanied by leaning back or avoiding movement. Characteristics of disgust include a wrinkled nose, upturned upper lip, and downturned corners of the mouth, possibly accompanied by a pout or turning the head; it is often directed at unpleasant smells, tastes, or things. Characteristics of surprise include sharply raised eyebrows, wide eyes, and a drooping jaw; it usually lasts for a short time.

[0036] It can be positive or negative, depending on the following emotions (such as surprise or fear). Characteristics associated with contempt include one corner of the mouth turning up, which is often associated with arrogance and belittling others. Step 2.2: Based on facial expression sets In a planar environment where the entire face is visible, facial expression images of subjects are collected to construct a planar facial expression data dataset. ; In a VR environment, where the face is occluded, multi-source facial expression data of the same subject is collected using multimodal sensors mounted on a VR headset to construct a facial expression information dataset for the VR environment. ; Step 2.3: Use a deep neural network to build and The mapping relationship is expressed by the formula: in, Represents a deep neural network. Represents a dataset of planar facial expression information Location dimensional vector space, Represents facial expression information dataset Location Dimensional vector space; Step 2.4: Eye data from the multi-source facial expression data obtained in Step 1. and visual dynamic feature data Using convolutional neural networks or feature point extraction algorithms to... Eye data at any time and visual dynamic feature data The mapping to the corresponding feature vectors, using the formula of a convolutional neural network, is as follows: in, express The eye feature vector at time step, express Facial feature vector at time step, Represents a convolutional neural network. Indicates the location of the eye feature vector dimensional vector space, This indicates the location of the facial feature vector. Dimensional vector space; Based on the non-visual facial muscle activity data from the multi-source facial expression data obtained in step 1 For each raw signal sequence from a flexible pressure sensor or electromyography sensor, within a sliding time window Extract time-domain or frequency-domain features to form sub-feature vectors, and concatenate the sub-feature vectors of all channels to obtain... The eigenvector corresponding to time step 1 is given by the formula: in, Represents the feature extraction function. express The feature vector of electromyographic signals and contact pressure data captured at the interface in real time, with dimensions... It depends on the number of channels multiplied by the number of features extracted from each channel. This indicates the number of electromyography (EMG) sensors, specifically the number of channels in the EMG sensors. This indicates the number of flexible pressure sensors, specifically the number of channels in the flexible pressure sensor. express Constantly contacting the interface to capture electromyographic signals, express Constantly monitor stress data. Indicates the first Sub-feature vectors of each electromyography sensor, Indicates the first Sub-feature vectors of a flexible pressure sensor Indicates transpose; Step 2.5: Align the results obtained in Step 2.4 on a unified time axis, and then construct a matrix of multi-source observation vectors. The formula is: in, Indicates transpose; Step 2.6, The user's full facial expression status at any given moment is defined as follows: ; Step 2.7: Matrix pairing based on multi-source observation fusion mapping function and multi-source observation vectors The formula for estimating the user's full facial expression at any given time is: in, The estimated value is... Real-time facial expression status of all users. The matrix representing the weight matrix corresponding to the aligned multi-source observation vectors. and The settings are dynamically adjusted based on how much of the subject's face is obscured by the VR headset; when the obscured area is large... Increased weight Decrease the weight, and vice versa. The weight decreases. Weight increases, The value is the total area of ​​the washer. Area of ​​the effective facial expression region of the subject The ratio value, Values , and ; Step 2.8: Based on the mapping relationship obtained in Step 2.3, the estimated... User's full facial expression status at any time Mapped to a planar environment, the reconstructed image is obtained. User's full facial expression status at any time The formula is: .

[0037] Unlike traditional methods (which rely solely on the non-occluded area to make strong coupling assumptions about the occluded area), the technical breakthrough of this invention lies in treating eye movement features, electromyography, and pressure information as observation channels for the occluded area. By complementing multimodal features, it avoids the uncertainty caused by "inferring the upper face from the lower face", significantly improving the stability and reliability of expression reconstruction in dynamic and complex interactive scenarios.

[0038] To address the difficulties in aligning emotion labels and emotion data in a temporal sequence in traditional solutions, and the lack of correlation between emotional information and physiological response data after facial expression reconstruction, this solution proposes to transform the reconstructed facial expressions (i.e., the user's full-face expression state) into labels with clear emotional meanings, and to establish a temporal mapping between them and physiological response data (induced physiological response state and physiological characteristics).

[0039] A method for labeling emotional physiological responses includes the following steps: Step S1: Generate physiological emotion or cognitive state labels based on the reconstructed user's full facial expression state: Inputting the user's full facial expressions into a pre-trained emotion recognition model or cognitive state recognition model, the system determines the subject's real-time physiological emotion or cognitive state and generates physiological emotion or cognitive state labels; or After the experiment, the sequence of users' full facial expression states will be provided to third-party professional evaluators to generate physiological emotion or cognitive state labels. Step S2: Label physiological response data using physiological emotion or cognitive state tags: By sharing a unified system clock, the online collected physiological response data and the reconstructed user facial expression state are given a time stamp that is consistent with the time. The physiological emotion or cognitive state label is automatically mapped to the physiological response data with the corresponding time stamp, so as to realize the labeling of physiological emotion or cognitive state. The physiological response data includes EEG, GSR and ECG.

[0040] Specifically: First, the reconstructed full-face facial expression state of the user (including facial movement intensity or structured parameters) is input into a pre-trained emotion recognition model or cognitive state recognition model. The model automatically classifies or distinguishes the dimensions of the expression features and outputs discrete emotion categories or continuous dimension values ​​(such as Valence / Arousal). The result obtained by the emotion recognition model ("expression emotion") serves as the emotion label for the physiological response data. By sharing a unified system clock, the physiological response data (EEG, GSR, ECG, etc.) collected online and the user's full facial expression state reconstructed in each frame are given a time-consistent (time-aligned) timestamp, establishing a precise correlation (correspondence) relationship.

[0041] Based on the timestamp alignment results, the system automatically assigns the emotion tags or cognitive state tags generated in real time by the emotion recognition model or cognitive state recognition model to the synchronously collected physiological signal data segments, establishing a mapping of "user's full facial expression state - emotion tag or cognitive state tag - physiological response data".

[0042] It provides high-frequency, synchronous, and physiologically relevant objective benchmark labels for physiological response data. While subjects engage in immersive interaction, the system acquires various physiological indicators online, such as electroencephalogram (EEG), conductance of skin (GSR), electrocardiogram (ECG), or pulse, through the physiological signal acquisition module.

[0043] A unified temporal reference is established at the system's underlying layer between the physiological response data sampling sequence and the user's full-face expression state reconstruction results. The annotation process runs in parallel in the background, eliminating the need to interrupt the subject's emotional perception during critical stages of emotion induction, thus ensuring the authenticity and continuity of the induction process. Models and third-party label recognition are used to completely resolve the interference of data annotation on the immersive experience in virtual reality environments. This invention transforms the annotation process into an automated background annotation process. This process no longer relies on the subject's active assessment but instead provides a high-reliability benchmark reference for physiological signals through the full-face expression reconstruction results. Specific annotation methods are as follows... Figure 4 As shown.

Claims

1. A method for acquiring and reconstructing facial expression information under VR headset occlusion conditions, characterized in that, Includes the following steps: Step 1: Obtain multi-source facial expression data: By setting up multimodal sensors on VR headsets, we can acquire full facial expression data under conditions where the VR headset partially obscures the face, thus obtaining multi-source facial expression data, including eye data in physically obscured areas and electromyographic signals and contact pressure data captured by the contact interface, as well as visual dynamic feature data of unobscured areas. Step 2: Reconstruct the user's full facial expression state based on multi-source facial expression data: The visual dynamic feature data of the unobstructed area, the eye data within the physical occlusion boundary, and the electromyographic signals and contact pressure data captured by the contact interface are fused to construct the user's full facial expression state.

2. The method for acquiring and reconstructing facial expression information under VR headset occlusion conditions according to claim 1, characterized in that, The specific steps of step 1 are as follows: Step 1.1: Using the visual sensor set on the outside of the VR headset, capture the dynamic feature data of the lower half of the face that is not covered by the VR headset, that is, obtain the visual dynamic feature data of the uncovered area. Step 1.2: Use the visual sensor built into the outer periphery of the VR headset lens to obtain local state data of the eye and the surrounding area, that is, obtain eye data, including gaze direction, blink frequency, eyelid opening and closing state and pupil changes. The visual sensor on the outer periphery of the VR headset lens refers to the camera in the eye tracking module built into the VR headset. Step 1.3: Using the electromyography (EMG) sensor and flexible pressure sensor on the VR headset gasket, acquire non-visual facial muscle activity data at the HMD gasket contact interface, including EMG signals and contact pressure data captured at the contact interface.

3. The method for acquiring and reconstructing facial expression information under VR headset occlusion conditions according to claim 2, characterized in that: In step 1.1, the visual sensor is an external camera, and the layout includes a single-camera structure protruding from the front of the head-mounted display and a symmetrical dual-camera arrangement structure below the head-mounted display. The front-mounted single-camera structure of the head-mounted display uses a protruding or folding bracket to fix an external camera to the front of the head-mounted display, so that the lens faces the wearer's face; The headset features a symmetrical dual-camera setup on the left and right sides of the lower part of the headset.

4. The method for acquiring and reconstructing facial expression information under VR headset occlusion conditions according to claim 2, characterized in that: In step 1.3, the electromyography sensor and flexible pressure sensor on the VR helmet gasket include at least an embedded part that is embedded in the VR helmet gasket and connected by an internal flexible connection structure. The embedded part includes an internal flexible connection structure embedded in the VR headset gasket, located on the user's forehead. The internal flexible connection structure is equipped with at least five sets of sensors: one set at the center of the forehead, at least two sets on each side of the center of the forehead corresponding to the two ends of the user's brow bones, one set of sensors on each side of the internal flexible connection structure corresponding to the area below each eye, one set of sensors on each side of the outer corner of each eye, one set of sensors on each side of the user's nose bridge, and one electromyography (EMG) sensor on each side of the internal flexible connection structure. Each set of sensors includes a flexible pressure sensor and an EMG sensor located on both sides of the flexible pressure sensor. It may also include peripheral components located outside the VR headset gasket; The peripheral device includes an external flexible connection structure connected to the internal flexible connection structure. At least one set of sensors is provided on the external flexible connection structure located above the forehead and below the hairline, and at least one set of sensors is provided on the user's cheek. The flexible pressure sensor captures local contact pressure fluctuations caused by facial muscle activity in real time to obtain contact pressure data. Electromyography (EMG) sensors monitor the electromyographic signals of facial muscles in the occluded area, directly reflecting the contraction and relaxation state of the muscles.

5. The method for acquiring and reconstructing facial expression information under VR headset occlusion conditions according to claim 4, characterized in that, The specific steps of step 2 are as follows: Step 2.1: Construct a set of facial expressions based on basic emotions. , Indicates the first There are 10 basic emotions and their corresponding characteristics. The basic emotions include happiness, sadness, anger, fear, disgust, surprise, and contempt. The characteristics corresponding to happiness include upturned corners of the mouth, slightly curved eyes, crow's feet at the corners of the eyes, raised cheeks, and natural contraction of the muscles around the eyes. The characteristics corresponding to sadness include raised inner eyebrows, downturned corners of the mouth, and dim eyes. The characteristics corresponding to anger include lowered and gathered eyebrows, wide eyes, tightly closed or open lips, flared nostrils, flushed face, and tense muscles. The characteristics corresponding to fear include raised and gathered eyebrows, wide eyes, and horizontally stretched lips. The characteristics corresponding to disgust include wrinkled nose, upturned upper lip, and downturned corners of the mouth. The characteristics corresponding to surprise include raised eyebrows, wide eyes, and downturned jaw. The characteristic corresponding to contempt is an upturned corner of the mouth on one side. Step 2.2: Based on facial expression sets In a planar environment where the entire face is visible, facial expression images of subjects are collected to construct a planar facial expression data dataset. ; In a VR environment, where the face is occluded, multi-source facial expression data of the same subject is collected using multimodal sensors mounted on a VR headset to construct a facial expression information dataset for the VR environment. ; Step 2.3: Use a deep neural network to build and The mapping relationship is expressed by the formula: in, Represents a deep neural network. Represents a dataset of planar facial expression information Location dimensional vector space, Represents facial expression information dataset Location Dimensional vector space; Step 2.4: Eye data from the multi-source facial expression data obtained in Step 1. and visual dynamic feature data Using convolutional neural networks or feature point extraction algorithms to... Eye data at any time and visual dynamic feature data The mapping to the corresponding feature vectors, using the formula of a convolutional neural network, is as follows: in, express The eye feature vector at time step, express Facial feature vector at time step, Represents a convolutional neural network. Indicates the location of the eye feature vector dimensional vector space, This indicates the location of the facial feature vector. Dimensional vector space; Based on the non-visual facial muscle activity data from the multi-source facial expression data obtained in step 1 For each raw signal sequence from a flexible pressure sensor or electromyography sensor, within a sliding time window Extract time-domain or frequency-domain features to form sub-feature vectors, and concatenate the sub-feature vectors of all channels to obtain... The eigenvector corresponding to time step 1 is given by the formula: in, Represents the feature extraction function. express The feature vector of electromyographic signals and contact pressure data captured at the interface in real time, with dimensions... It depends on the number of channels multiplied by the number of features extracted from each channel. This indicates the number of electromyography (EMG) sensors, specifically the number of channels in the EMG sensors. This indicates the number of flexible pressure sensors, specifically the number of channels in the flexible pressure sensor. express Constantly contacting the interface to capture electromyographic signals, express Constantly monitor stress data. Indicates the first Sub-feature vectors of each electromyography sensor, Indicates the first Sub-feature vectors of a flexible pressure sensor Indicates transpose; Step 2.5: Align the results obtained in Step 2.4 on a unified time axis, and then construct a matrix of multi-source observation vectors. The formula is: in, Indicates transpose; Step 2.6, The user's full facial expression status at any given moment is defined as follows: ; Step 2.7: Matrix pairing based on multi-source observation fusion mapping function and multi-source observation vectors The formula for estimating the user's full facial expression at any given time is: in, The estimated value is... Real-time facial expression status of all users. The matrix representing the weight matrix corresponding to the aligned multi-source observation vectors. and The settings are dynamically adjusted based on how much of the subject's face is obscured by the VR headset; when the obscured area is large... Increased weight Decrease the weight, and vice versa. The weight decreases. Weight increases, The value is the total area of ​​the washer. Area of ​​the effective facial expression region of the subject The proportion value, Values , and ; Step 2.8: Based on the mapping relationship obtained in Step 2.3, the estimated... User's full facial expression status at any time Mapped to a planar environment, the reconstructed image is obtained. User's full facial expression status at any time The formula is: 。 6. A method for labeling emotional physiological responses, characterized in that, Includes the following steps: Step S1: Generate physiological emotion or cognitive state labels based on the reconstructed user facial expression state obtained from any one of claims 1-5: The user's full facial expression state is input into a pre-trained emotion recognition model or cognitive state recognition model to determine the subject's real-time physiological emotion or cognitive state and generate physiological emotion or cognitive state labels. or After the experiment, the sequence of users' full facial expression states will be provided to third-party professional evaluators to generate physiological emotion or cognitive state labels. Step S2: Label physiological response data using physiological emotion or cognitive state tags: By sharing a unified system clock, the online collected physiological response data and the reconstructed user facial expression state are given a time stamp that is consistent with the time. The physiological emotion or cognitive state label is automatically mapped to the physiological response data with the corresponding time stamp, so as to realize the labeling of physiological emotion or cognitive state. The physiological response data includes EEG, GSR and ECG.

7. A system for acquiring and reconstructing facial expression information under VR headset occlusion conditions, characterized in that, include: VR head-mounted display devices: By setting up multimodal sensors on VR headsets, full facial expression data is obtained under conditions where the VR headset partially obscures the face, thus obtaining multi-source facial expression data, including eye data in physically obscured areas and electromyographic signals and contact pressure data captured by the contact interface, as well as visual dynamic feature data of unobscured areas. Facial expression reconstruction module: It fuses the visual dynamic feature data of the unobstructed area, the eye data within the physical occlusion boundary, and the electromyographic signals and contact pressure data captured by the contact interface to construct the user's full facial expression state.

8. The facial expression information acquisition and reconstruction system under VR headset occlusion conditions according to claim 7, characterized in that, The VR head-mounted display device includes a VR helmet, a visual sensor disposed on the outside of the VR helmet to capture dynamic feature data of the lower half of the face not obscured by the VR helmet, a visual sensor built into the outer periphery of the VR helmet lens to acquire local state data of the eyes and the surrounding area, and an electromyography sensor and a flexible pressure sensor embedded in the VR helmet gasket to acquire non-visual facial muscle activity data of the HMD gasket contact interface. The visual sensors located on the outside of the VR headset are external cameras, and the layout includes a single camera structure protruding from the front of the headset and a symmetrical dual-camera arrangement below the headset. The front-mounted single-camera structure of the head-mounted display uses a protruding or folding bracket to fix an external camera to the front of the head-mounted display, so that the lens faces the wearer's face; The dual-camera symmetrical arrangement structure at the bottom of the headset features two external cameras symmetrically installed on the left and right sides of the lower part of the headset. The electromyography sensor and flexible pressure sensor on the VR headset gasket include at least the embedded portion that is connected within the VR headset gasket by an internal flexible connection structure. The embedded part includes an internal flexible connection structure embedded in the VR headset gasket, located on the user's forehead. The internal flexible connection structure is equipped with at least five sets of sensors: one set at the center of the forehead, at least two sets on each side of the center of the forehead corresponding to the two ends of the user's brow bones, one set of sensors on each side of the internal flexible connection structure corresponding to the area below each eye, one set of sensors on each side of the outer corner of each eye, one set of sensors on each side of the user's nose bridge, and one electromyography (EMG) sensor on each side of the internal flexible connection structure. Each set of sensors includes a flexible pressure sensor and an EMG sensor located on both sides of the flexible pressure sensor. It may also include peripheral components located outside the VR headset gasket; The peripheral device includes an external flexible connection structure connected to the internal flexible connection structure. At least one set of sensors is provided on the external flexible connection structure located above the forehead and below the hairline, and at least one set of sensors is provided on the user's cheek. The flexible pressure sensor captures local contact pressure fluctuations caused by facial muscle activity in real time to obtain contact pressure data. Electromyography (EMG) sensors monitor the electromyographic signals of facial muscles in the occluded area, directly reflecting the contraction and relaxation state of the muscles.

9. A system for acquiring and reconstructing facial expression information under VR headset occlusion conditions according to claim 8, characterized in that, The specific implementation steps of the facial expression reconstruction module are as follows: Step 2.1: Construct a set of facial expressions based on basic emotions. , Indicates the first There are 10 basic emotions and their corresponding characteristics. The basic emotions include happiness, sadness, anger, fear, disgust, surprise, and contempt. The characteristics corresponding to happiness include upturned corners of the mouth, slightly curved eyes, crow's feet at the corners of the eyes, raised cheeks, and natural contraction of the muscles around the eyes. The characteristics corresponding to sadness include raised inner eyebrows, downturned corners of the mouth, and dim eyes. The characteristics corresponding to anger include lowered and gathered eyebrows, wide eyes, tightly closed or open lips, flared nostrils, flushed face, and tense muscles. The characteristics corresponding to fear include raised and gathered eyebrows, wide eyes, and horizontally stretched lips. The characteristics corresponding to disgust include wrinkled nose, upturned upper lip, and downturned corners of the mouth. The characteristics corresponding to surprise include raised eyebrows, wide eyes, and downturned jaw. The characteristic corresponding to contempt is an upturned corner of the mouth on one side. Step 2.2: Based on facial expression sets In a planar environment where the entire face is visible, facial expression images of subjects are collected to construct a planar facial expression data dataset. ; In a VR environment, where the face is occluded, multi-source facial expression data of the same subject is collected using multimodal sensors mounted on a VR headset to construct a facial expression information dataset for the VR environment. ; Step 2.3: Use a deep neural network to build and The mapping relationship is expressed by the formula: in, Represents a deep neural network. Represents a dataset of planar facial expression information Location dimensional vector space, Represents facial expression information dataset Location Dimensional vector space; Step 2.4: Eye data from the multi-source facial expression data obtained in Step 1. and visual dynamic feature data Using convolutional neural networks or feature point extraction algorithms to... Eye data at any time and visual dynamic feature data The mapping to the corresponding feature vectors, using the formula of a convolutional neural network, is as follows: in, express The eye feature vector at time step, express Facial feature vector at time step, Represents a convolutional neural network. Indicates the location of the eye feature vector dimensional vector space, This indicates the location of the facial feature vector. Dimensional vector space; Based on the non-visual facial muscle activity data from the multi-source facial expression data obtained in step 1 For each raw signal sequence from a flexible pressure sensor or electromyography sensor, within a sliding time window Extract time-domain or frequency-domain features to form sub-feature vectors, and concatenate the sub-feature vectors of all channels to obtain... The eigenvector corresponding to time step 1 is given by the formula: in, Represents the feature extraction function. express The feature vector of electromyographic signals and contact pressure data captured at the interface in real time, with dimensions... It depends on the number of channels multiplied by the number of features extracted from each channel. This indicates the number of electromyography (EMG) sensors, specifically the number of channels in the EMG sensors. This indicates the number of flexible pressure sensors, specifically the number of channels in the flexible pressure sensor. express Constantly contacting the interface to capture electromyographic signals, express Constantly monitor stress data. Indicates the first Sub-feature vectors of each electromyography sensor, Indicates the first Sub-feature vectors of a flexible pressure sensor Indicates transpose; Step 2.5: Align the results obtained in Step 2.4 on a unified time axis, and then construct a matrix of multi-source observation vectors. The formula is: in, Indicates transpose; Step 2.6, The user's full facial expression status at any given moment is defined as follows: ; Step 2.7: Matrix pairing based on multi-source observation fusion mapping function and multi-source observation vectors The formula for estimating the user's full facial expression at any given time is: in, The estimated value is... Real-time facial expression status of all users. The matrix representing the weight matrix corresponding to the aligned multi-source observation vectors. and The settings are dynamically adjusted based on how much of the subject's face is obscured by the VR headset; when the obscured area is large... Increased weight Decrease the weight, and vice versa. The weight decreases. Weight increases, The value is the total area of ​​the washer. Area of ​​the effective facial expression region of the subject The proportion value, Values , and ; Step 2.8: Based on the mapping relationship obtained in Step 2.3, the estimated... User's full facial expression status at any time Mapped to a planar environment, the reconstructed image is obtained. User's full facial expression status at any time The formula is: 。 10. An emotional physiological response annotation system, characterized in that, include: The label recognition module: Based on the reconstructed full-face expression state of the user obtained by any one of claims 7-9, the module inputs the reconstructed expression state of the user into the pre-trained emotion recognition model or cognitive state recognition model, determines the real-time physiological emotion or cognitive state of the subject, and generates physiological emotion or cognitive state labels. or After the experiment, the sequence of users' full facial expression states will be provided to third-party professional evaluators to generate physiological emotion or cognitive state labels. Labeling module: By sharing a unified system clock, the online collected physiological response data and the reconstructed user facial expression state are assigned a time stamp that is consistent with the time. The physiological emotion or cognitive state label is automatically mapped to the physiological response data with the corresponding time stamp, so as to realize the labeling of physiological emotion or cognitive state labels. The physiological response data includes EEG, GSR and ECG.