Emotional response evaluation method, device and equipment for VR environment
By dynamically selecting preset modes in a VR environment and combining them with multimodal data evaluation, the problem of difficulty and accuracy in controlling emotion labeling in VR environment is solved, achieving efficient and reliable emotion response evaluation and ensuring the authenticity and traceability of the data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF ELECTRONICS SCI & TECH OF CHINA
- Filing Date
- 2026-01-16
- Publication Date
- 2026-04-21
AI Technical Summary
In VR environments, it is difficult for subjects to label emotions by capturing their gaze or manipulating external input devices after wearing head-mounted displays, resulting in a high error rate in the labeling results. Furthermore, existing technologies cannot effectively improve the accuracy of emotion computing models.
By dynamically selecting preset modes in a VR environment, and combining eye-tracking data, interactive controller data, voice audio data, and physiological signal data, a comprehensive confidence score is calculated, and a structured data package is output to achieve emotional response assessment.
It maximizes the protection of the purity of the induction environment, reduces the difficulty of the annotation process, significantly improves the accuracy and confidence of the data, ensures the authenticity and reliability of the emotion data, and supports efficient collection and management.
Smart Images

Figure CN121901649A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and more specifically to a method, apparatus, and device for assessing emotional responses in a VR environment. Background Technology
[0002] Virtual reality (VR) technology, with its unique immersion, interactivity, and imagination, has become an important platform for research in emotion recognition, cognitive science, and human-computer interaction. Compared to traditional flat-panel display devices, VR head-mounted displays (HMDs) can construct high-dimensional immersive experimental scenarios through multi-sensory channels and panoramic displays, effectively shielding external interference and giving subjects a strong sense of presence, thereby inducing more realistic, continuous, and profound emotional experiences. Furthermore, VR environments possess high ecological validity and controllability, providing repeatable, low-cost, and safe experimental conditions, offering an ideal new environment for quantifying emotion recognition.
[0003] In quantitative emotion recognition research, emotion fragment sampling is gradually shifting from discrete, static emotion labeling to continuous, dynamic emotion labeling that more closely resembles real-world experiences. This labeling method requires real-time tracking across multiple continuously switching and logically related emotion scenarios throughout the entire time domain. Simultaneously, building high-precision emotion (cognitive) recognition models heavily relies on high-quality emotion-induced response data and high-quality real-time labeling.
[0004] Compared to emotion induction in a two-dimensional planar environment (hereinafter referred to as "planar environment"), real-time emotion induction in a three-dimensional VR environment (hereinafter referred to as "VR environment") can improve the efficiency of collecting emotion or cognitive response data (hereinafter referred to as emotion data), and is a key technical means to improve the accuracy of emotion computing models.
[0005] However, because VR headsets (HMDs) completely block the view of the external physical environment, it is inconvenient for subjects to use eye movements to capture, locate, or control external physical input devices such as keyboards and mice to make annotations after wearing the device, resulting in a high error rate in the annotation results. Summary of the Invention
[0006] The purpose of this invention is to provide a method, apparatus, and device for assessing emotional responses in a VR environment, thereby solving the problems in the prior art.
[0007] This invention is achieved through the following technical solution:
[0008] In a first aspect, embodiments of the present invention provide a method for assessing emotional responses in a VR environment, comprising:
[0009] In a virtual reality (VR) environment, based on the current induced scene and the subject's state, a preset mode is dynamically selected to obtain the subject's emotional response label and the initial confidence level corresponding to the emotional response label. The emotional response label is a numerical value or vector representing the emotional state. The preset mode includes any of the following modes: First mode, by rendering an annotation interface that does not affect the continuous presentation and interaction of the current induced scene content, and collecting the subject's eye-tracking data and interactive handle data to complete the annotation input, the emotional response label is obtained; Second mode, by mapping the captured and parsed natural language speech of the subject to emotional response labels; Third mode, by using an emotion recognition model pre-trained in a planar environment and the immersion parameters of the current induced scene to predict and generate emotional response labels.
[0010] For the data acquisition event that generates the emotional response label, multimodal data of the subject associated with the data acquisition event is collected simultaneously. The multimodal data includes eye movement trajectory data and interactive handle operation trajectory data corresponding to the first mode, voice and audio data corresponding to the second mode, real-time behavioral features corresponding to the third mode, and physiological signal data corresponding to each mode.
[0011] Based on the multimodal data, calculate the process quality score for each mode and the physiological verification score for all modes during the assessment process;
[0012] The overall confidence score is calculated based on the initial confidence level, process quality score, and physiological verification score corresponding to each mode;
[0013] The emotion response label, the overall confidence score, and the preset pattern for generating the emotion response label are associated with a timestamp and output as a structured data packet.
[0014] Preferably, the dynamic selection of the preset mode includes:
[0015] Monitor the subject's real-time input behavior;
[0016] Based on the real-time input behavior, the currently active annotation mode is identified, wherein the first mode is executed when an action confirmation event is detected by ray tapping with the controller or eye tracking; the second mode is executed when a voice input event is detected; and the third mode is executed when no action confirmation event or voice input event is detected.
[0017] Preferably, when the first mode is executed:
[0018] The dynamic selection of preset modes to obtain the subject's emotional response label and the initial confidence level corresponding to the emotional response label includes: receiving the input signal generated by the subject through virtual ray or eye movement confirmation via a semi-transparent non-modal annotation interface to generate an emotional response label, and simultaneously recording behavioral trajectory data to calculate the initial confidence level;
[0019] And / or, the calculation of the process quality score of the subject corresponding to each mode during the assessment process includes: calculating eye movement focus based on the eye movement trajectory data, wherein the eye movement focus is calculated based on the fixation density of the fixation point falling into the marked area, fixation stability, and a penalty factor positively correlated with saccadic speed, wherein the penalty factor is used to reduce the value of the eye movement focus;
[0020] Based on the interactive handle operation trajectory data, the interactive stability is calculated, which is obtained by calculating the reciprocal of the hand tremor amplitude and the smoothness of the sliding path.
[0021] The process quality score is calculated based on the eye-tracking focus and the interaction stability.
[0022] Preferably, when the second mode is executed:
[0023] The process of dynamically selecting a preset mode to obtain the subject's emotional response label and the initial confidence level corresponding to the emotional response label includes: processing the captured natural language speech to obtain the emotional response label, and extracting acoustic features to calculate the semantic and acoustic consistency as the initial confidence level.
[0024] And / or, the calculation of the process quality score of the subject corresponding to each mode during the evaluation process includes: processing the captured speech audio data to obtain speech recognition content and extracting acoustic features;
[0025] The clarity of expression is calculated based on the speech rate and pause ratio of the speech recognition content;
[0026] The similarity between the semantic vector corresponding to the speech recognition content and the emotion tendency vector obtained from the acoustic feature analysis is calculated as semantic consistency.
[0027] The process quality score is calculated based on the clarity of expression and the semantic consistency.
[0028] Preferably, when the third mode is executed:
[0029] The dynamic selection of preset modes to obtain the subject's emotional response labels and the initial confidence level corresponding to the emotional response labels includes: obtaining a pre-trained emotion recognition model, wherein the emotion recognition model is obtained by joint training using data induced by homologous emotional stimulus materials in a planar environment and a VR environment through a domain adaptation method;
[0030] The emotional stimulus material of the current induced scenario is input into the emotion recognition model to obtain the basic emotion prediction value and the initial confidence level.
[0031] Based on the immersion parameters of the current induced scene, the basic emotion prediction value is enhanced and corrected to generate emotion response tags;
[0032] And / or, the calculation of the process quality score of the subject in each mode during the evaluation process includes: obtaining the pre-collected homologous behavioral baseline features of the subject in a planar environment, calculating the deviation of the real-time behavioral features in the VR environment from the baseline features, and calculating the cross-domain attention behavior reliability score as the process quality score based on the deviation.
[0033] Preferably, the physiological verification score is obtained in the following manner:
[0034] The physiological arousal verification index is obtained by weighting and summing the skin conductance amplitude, heart rate variability and the reciprocal of electromyography signal in the physiological signal data with preset weighting coefficients.
[0035] Calculate the matching error between the subjective arousal level represented by the emotion response label and the physiological arousal verification index;
[0036] The overall confidence score is calibrated based on the matching error, wherein the lower the matching error, the higher the overall confidence score.
[0037] Preferably, in the first mode, the annotation interface that renders without affecting the continuous presentation and interaction of the current VR scene content includes:
[0038] Based on the real-time monitoring of the task load of the induced scenario and the emotional arousal level of the subjects, the transparency and rendering scale of the annotation interface are dynamically adjusted.
[0039] In the pre-defined key nodes of the induced scenario, the labeling interface is automatically suspended or weakened.
[0040] Preferably, in the first mode, rendering the annotation interface further includes the following positioning modes:
[0041] A fixation-associated positioning mode is used to dynamically attach the labeled interface to a subcentral region outside the subject's focal point of vision based on the eye movement data.
[0042] The user-defined visual field anchor point positioning mode is used to make the annotation interface move with the subject's head rotation based on the comfortable visual field range calibrated during the initialization phase and real-time head posture data.
[0043] Secondly, embodiments of the present invention provide an emotion response assessment device for a VR environment, comprising:
[0044] An emotion response tag acquisition module is used to dynamically select a preset mode in a virtual reality (VR) environment to acquire the subject's emotion response tag and the initial confidence level corresponding to the emotion response tag, based on the current induced scene and the subject's state. The emotion response tag is a numerical value or vector representing the emotional state. The preset mode includes any of the following modes: First mode, by rendering an annotation interface that does not affect the continuous presentation and interaction of the current induced scene content, and collecting the subject's eye-tracking data and interactive handle data to complete the annotation input, the emotion response tag is obtained; Second mode, by mapping the captured and parsed natural language speech of the subject to emotion response tags; Third mode, by using an emotion recognition model pre-trained in a planar environment and the immersion parameters of the current induced scene to predict and generate emotion response tags.
[0045] The confidence data acquisition module is used to simultaneously collect multimodal data of the subject associated with the data acquisition event that generates the emotional response label. The multimodal data includes eye movement trajectory data and interactive handle operation trajectory data corresponding to the first mode, voice audio data corresponding to the second mode, real-time behavioral features corresponding to the third mode, and physiological signal data corresponding to each mode.
[0046] The process scoring module is used to calculate the process quality score corresponding to each mode and the physiological verification score corresponding to all modes for the subject during the assessment process based on the multimodal data.
[0047] The overall confidence module is used to calculate the overall confidence score based on the initial confidence score, process quality score, and physiological verification score corresponding to each mode;
[0048] The output module is used to associate the emotion response label, the comprehensive confidence score, and the preset pattern for generating the emotion response label with a timestamp, and output them as a structured data packet.
[0049] Thirdly, embodiments of the present invention provide an electronic device, including: at least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method of the first aspect described above.
[0050] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0051] 1. It maximizes the protection of the purity of the induction environment, thus solving the problem of low data confidence.
[0052] Non-modal semi-transparent rendering and task-aware scheduling were implemented, resolving the interruptions caused by "interface occlusion" and "forced pop-ups disrupting immersion" in traditional annotation schemes. The executing entity can automatically become invisible or delay triggering when the subject is at a climax of the plot or a key moment in combat. This ensures that the collected emotions are the subject's genuine reactions to the inducing material, rather than agitation or distraction caused by annotation interference, thus guaranteeing data fidelity from the source.
[0053] 2. Significantly reduced the difficulty of the annotation process, achieving efficient data collection.
[0054] An automatic prediction path based on cross-domain transfer learning was introduced, supplemented by a hands-free implicit voice path, solving the problems of "blind operation difficulty" and "device dependence" caused by physical line-of-sight obstruction in VR environments. Subjects do not need to interrupt the experiment to find physical keyboards or gamepad buttons. This "prediction-based, interaction-assisted" mode significantly reduces the frequency of manual annotation, making continuous emotion tracking in long-duration, highly immersive VR experiments economical and feasible.
[0055] 3. A refined quality assessment mechanism was established, which solved the problems of low data accuracy and confidence.
[0056] By collecting paralinguistic data of interactive behaviors (hand tremors, gaze deviation) and combining this with the Physiological Arousal Validation Index (PVI) for cross-validation, the system solves the common problems in 6DoF spatial interaction, such as "slip-on labeling" (low accuracy) and "casual completion / social expectation bias" (low confidence). Each output label is no longer an isolated numerical value, but structured data with a [0,1] confidence score. The system can automatically identify which responses are "decisive and genuine" from the participants and which are "hesitant or perfunctory" noise.
[0057] Furthermore, at the data management level, this invention also achieves hierarchical data management, enhancing the scientific research value of the dataset. It employs an adaptive fusion algorithm to output a "four-tuple data package" containing labels, confidence levels, source information, and timestamps. This avoids the information loss caused by "directly discarding low-quality data" and the "untraceable data" problems found in existing technologies. It fully preserves all interaction traces, allowing researchers to use confidence labels to hierarchically filter data (e.g., extracting only data with a confidence level > 0.8 for training), and to use low-confidence data to trace back the psychological conflicts of subjects (e.g., why did they subjectively say they weren't afraid, yet their physiological responses fluctuated drastically), providing comprehensive and traceable evidence for research on complex emotional interactions. Attached Figure Description
[0058] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the drawings:
[0059] Figure 1 A flowchart illustrating the emotional response assessment method for VR environments provided by this invention;
[0060] Figure 2 An example diagram of the emotion annotation method based on voice prompts provided by this invention;
[0061] Figure 3 This is a flowchart illustrating a gaze-point association localization method provided by the present invention.
[0062] Figure 4 This is a flowchart illustrating a user-defined viewpoint anchor point positioning method provided by the present invention.
[0063] Figure 5 Example diagram of the pop-up control form provided by the present invention;
[0064] Figure 6 A schematic diagram of the structure of the emotion response assessment device for a VR environment provided by the present invention;
[0065] Figure 7 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0066] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.
[0067] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0068] It should be noted that all actions involving the acquisition of signals, information, or data in this invention are carried out in compliance with the relevant data protection laws and regulations of the locality and with authorization from the owner of the relevant device.
[0069] Example 1
[0070] Please see Figure 1 This invention provides a method for assessing emotional responses in a VR environment, comprising:
[0071] S1. In a virtual reality (VR) environment, based on the current induced scene and the subject's state, a preset mode is dynamically selected to obtain the subject's emotional response label and the initial confidence level corresponding to the emotional response label. The emotional response label is a numerical value or vector representing the emotional state. The preset mode includes any of the following modes: First mode, by rendering an annotation interface that does not affect the continuous presentation and interaction of the current induced scene content, and collecting the subject's eye-tracking data and interactive handle data to complete the annotation input, the emotional response label is obtained; Second mode, by mapping the captured and parsed natural language speech of the subject to emotional response labels; Third mode, by using an emotion recognition model pre-trained in a planar environment and the immersion parameters of the current induced scene to predict and generate emotional response labels.
[0072] Specifically, the system first performs real-time analysis and evaluation of the induced scene content presented to the subject (such as tense game levels or calm natural scenery videos) and the subject's real-time state (such as whether they are engaged in intense interaction or quietly observing). Based on this evaluation, the system dynamically selects the most suitable one from three predefined data acquisition modes. The first mode is suitable for scenarios where the user can perform manual operations. The system renders a semi-transparent and non-modal annotation interface that does not interrupt or obscure the main scene content. The user performs operations such as selection and sliding on the interface by eye-tracking or emitting virtual rays from the controller. The system directly generates a numerical value or vector representing the current emotional state (such as pleasure or arousal) based on these inputs, i.e., an emotional response label. The second mode is suitable for highly immersive scenarios where the user's hands are occupied or it is inconvenient to operate the interface. The system captures the user's natural language speech through a microphone, and through speech recognition and emotional semantic analysis, maps statements such as "I am nervous" or "This is so boring" to corresponding emotional dimension values, generating an emotional response label. The third mode is suitable for scenarios where the user is completely immersed and has no active feedback. The system calls upon an emotion recognition model trained in a traditional planar display environment. This model can extract features and infer emotions from the audiovisual content of the current VR scene. Simultaneously, the system introduces a parameter representing the level of immersion in the current VR environment (such as visual field angle and degrees of freedom of interaction) to weight or correct the model's original output, thereby generating a predicted emotion response label. While generating the label, each mode produces a preliminary confidence assessment value, i.e., an initial confidence level, based on its internal logic. For example, in the first mode, this value might be based on the decisiveness of the user's action (such as clicking without hesitation); in the second mode, it might be based on the clarity and certainty of the speech; and in the third mode, it directly comes from the predicted probability output by the model. The core of this step lies in achieving intelligent routing of the data acquisition path and preliminary quality control.
[0073] S2. For the data acquisition event that generates the emotion response label, simultaneously collect multimodal data of the subject associated with the data acquisition event. The multimodal data includes eye movement trajectory data and interactive handle operation trajectory data corresponding to the first mode, voice audio data corresponding to the second mode, real-time behavioral features corresponding to the third mode, and physiological signal data corresponding to each mode.
[0074] Specifically, this step aims to capture an objective, multi-dimensional chain of evidence for subsequent refined quality assessment. The system will define an independent data acquisition event for each emotion tag generation and ensure that raw data streams related to the user's state and operation process are recorded synchronously within the time window of the event. The selection of these data is closely related to the currently activated acquisition mode and is highly targeted. When the system is running in the first mode, it mainly collects eye movement trajectory data (such as fixation sequence and saccade speed) and interactive handle operation trajectory data (such as displacement and button timing) during the user's operation of the labeling interface to analyze the user's visual focus and hand stability during operation. When the system is running in the second mode, it mainly collects the user's raw speech audio data to analyze their speech characteristics. When the system is running in the third mode, since there is no active interaction, it mainly collects data that reflects the user's unconscious behavioral state, such as real-time eye movements and head posture. In addition, regardless of the mode, the system can simultaneously collect physiological signal data such as skin conductance and heart rate as objective indicators reflecting the physiological arousal of emotions. The key technology in this step lies in the strict synchronization and alignment of data, ensuring that each frame of behavioral or physiological data can accurately correspond to the moment when the emotion label is generated, thus providing a foundation for building a reliable causal relationship.
[0075] In order to identify annotation errors at the "process" level, this step dynamically extracts in real time features reflecting focus and physiological consistency during annotation:
[0076] Extracting physiological objective characteristics from any selected mode: Real-time acquisition of the temporal response amplitude of skin conductance signals via an interactive handle to identify whether the subject "truly experiences emotional fluctuations." The calculation formula is as follows:
[0077] (Physiological objective characteristics):
[0078] ;
[0079] in, This is the real-time amplitude. and These represent the mean and standard deviation for the baseline phase. This indicator is used to identify whether subjects exhibit genuine physiological arousal. The presence of "genuine emotional arousal" is verified by the standardized score of the temporal amplitude of skin electrical activity.
[0080] When the first mode is selected, additional features are extracted: eye-tracking attention features: through the eye-tracking module of VRHMD, features including gaze point distribution entropy, ROI gaze duration ratio, and gaze space variance (measuring whether the gaze is drifting) are calculated in real time; interaction stability features: through real-time monitoring of the interactive controller by the executing subject, ray hover jitter, trajectory smoothness and button timing smoothness are calculated to identify whether the subject's operation is decisive and whether there is mechanical tremor.
[0081] (Characteristics of eye movement attention):
[0082] ;
[0083] in, The gaze point distribution entropy reflects the regularity of the gaze coverage.
[0084] Percentage of time spent looking at areas of interest.
[0085] : Spatial variance of fixation point; the larger the value, the more erratic the gaze.
[0086] (Interaction stability characteristics):
[0087] ;
[0088] in, It is based on the second or third derivative of the ray's end coordinates and is used to describe the "cleanliness" of the virtual ray's trajectory. It uses a high-pass filter to extract the high-frequency components in the ray pointing signal, which is used to capture high-frequency tremors of the hand at the microscopic level. It is the time interval from when the ray falls into the target area to when the wrench button is pressed, used to measure the time rhythm between the two actions of "aiming" and "clicking".
[0089] When the second mode is selected, additional features are extracted: speech para-language features: speech rate stability and speech recognition uncertainty are calculated in real time through speech recognition and recording;
[0090] (Phonetic paralinguistic features):
[0091] ;
[0092] in The paralinguistic score reflects the stability of the subject's psychological state during expression. It's the speaking speed. This is a speech rate stability function; if the speech rate suddenly increases or there are many pauses (discontinuous speech), this value will decrease. This usually indicates that the subject is in a state of tension, anxiety, or mental confusion. It is to identify uncertainty. This is the language confidence function. It is based on the confidence score output by the speech recognition algorithm (ASR). If the subject speaks unclearly, weakly, or with a lot of filler words such as "uh" or "that," the score for this item will decrease.
[0093] S3. Based on the multimodal data, calculate the process quality score for each mode and the physiological verification score for all modes for the subject during the assessment process;
[0094] Specifically, this step is the core of data quality assessment. It generates two key quality scores by calculating the raw multimodal data collected by S2. The process quality score aims to quantify the quality of user behavior during annotation or in a specific state. In the first mode, this score is calculated by analyzing the concentration and stability of eye movements and the smoothness and micro-tremor amplitude of handgrip manipulation, reflecting the user's focus and precision. In the second mode, this score is calculated by analyzing the stability of speech rate, clarity, and consistency between semantic content and acoustic features (such as tone of voice), reflecting the seriousness and authenticity of the user's verbal reports. In the third mode, this score is calculated by comparing the user's current real-time behavioral characteristics (such as saccade patterns) with their baseline behavioral characteristics in a calm state, calculating the deviation to assess whether their attention is in a normal state of immersion rather than a state of distraction. The physiological verification score serves as a universal objective verification indicator across modes. It is mainly based on synchronously collected physiological signal data (such as skin conductance and heart rate) to calculate a comprehensive physiological arousal level index, and then calculates the matching error between this index and the subjective emotional arousal level included in the emotional response label. The smaller the matching error, the closer the user's subjectively reported emotional intensity matches their objective physiological response, resulting in a higher physiological verification score. Conversely, a larger matching error indicates a potential discrepancy between the user's words and their true feelings or a distorted report, leading to a lower physiological verification score. This step transforms the user's internal state, which is difficult to observe directly, into a calculable and comparable quantitative score.
[0095] S4. Calculate the overall confidence score based on the initial confidence score, process quality score, and physiological verification score corresponding to each mode;
[0096] Specifically, this step involves the final integration and adjudication of the credibility of emotion response labels. The system integrates the initial confidence score generated in S1, the process quality score calculated in S3, and the physiological verification score according to a predetermined fusion rule (such as weighted averaging), ultimately outputting a comprehensive confidence score between 0 and 1. This score is not a simple average but reflects a trade-off between different quality dimensions. For example, a label might originate from a seemingly decisive action (high initial confidence), but process quality analysis reveals severe hand tremors during the action (low process quality score), and physiological verification shows no corresponding arousal (low physiological verification score). In such cases, the final comprehensive confidence score will be significantly lowered. Conversely, if the scores in all three dimensions are high, a high confidence score will be assigned. This mechanism ensures that each emotion response label is associated with a clear, multi-source evidence-based quality "identity card," the value of which directly reflects the usability and reliability level of the data in subsequent model training or scientific research.
[0097] The formula for calculating the overall confidence score is as follows:
[0098] ;
[0099] Q represents the initial confidence level, which measures the credibility of different annotation patterns, such as the consistency of speech and semantics or the decisiveness of 6DoF clicks.
[0100] RS stands for Process Quality Score, which measures "focus" during annotation. Examples include whether eye movements are focused and whether there is mechanical tremor in the hands.
[0101] The physiological verification score measures whether emotions are normally induced.
[0102] , , The sum of the three is 1, corresponding to the weights, and can be adjusted according to individual differences.
[0103] S5. Associate the emotion response tag, the comprehensive confidence score, and the preset pattern for generating the emotion response tag with the timestamp, and output them as a structured data packet.
[0104] Specifically, the electronic device encapsulates the emotion response label generated by S1, the comprehensive confidence score calculated by S3, the acquisition pattern identifier corresponding to the label, and the precise timestamp, outputting a structured data packet. This data packet is typically organized in the form of key-value pairs or specific field sequences to ensure the complete association and unified management of the four core elements: emotion label, its quality assessment score, source context, and time of occurrence. This structured output format facilitates data storage, transmission, and subsequent analysis. Researchers or upper-level applications can classify and filter data based on the confidence score, or combine the timestamp to accurately align and analyze emotion changes with VR scene content, thus providing a standardized and traceable data foundation for building highly robust emotion computing models or conducting in-depth psychological research.
[0105] Structured data packets ,in, Represents the final emotional response label (e.g., pleasure score); It is the comprehensive confidence score calculated through the above matching logic, ranging from 0 to 1, used to characterize the reliability of the data; Labeling the data using which preset mode was used (predictive, visual, or voice) is crucial for analyzing user behavior habits; This is to mark the timestamp of the event, ensuring accurate alignment between the emotional state and the VR scene.
[0106] In some implementations, the dynamic selection of the preset mode includes:
[0107] Monitor the subject's real-time input behavior;
[0108] Based on the real-time input behavior, the currently active annotation mode is identified, wherein the first mode is executed when an action confirmation event is detected by ray tapping with the controller or eye tracking; the second mode is executed when a voice input event is detected; and the third mode is executed when no action confirmation event or voice input event is detected.
[0109] Specifically, this is achieved by continuously monitoring the subject's real-time input behavior flow in the virtual reality environment. The electronic device analyzes input signals from multiple channels, including the head-mounted display, controllers, and microphone, and determines which emotion tag acquisition mechanism should be activated based on predefined behavior pattern recognition logic. Specifically, when the electronic device parses a controller ray click signal that matches preset characteristics from the input stream, or identifies an eye-tracking confirmation signal output by the eye-tracking module that indicates a clear selection interaction with the virtual interface controls, it determines that an action confirmation event has occurred and activates the first mode accordingly. When the electronic device recognizes the subject's voice through speech activity detection technology, and the speech signal, after preliminary analysis, contains semantic content related to the emotion description, it determines that a speech input event has occurred and activates the second mode accordingly. If the electronic device detects neither a valid action confirmation event nor a valid speech input event within a continuous time window, it infers that the subject is in a state of deep immersion without active interaction, and the electronic device automatically activates the third mode. This mode-switching mechanism, based on real-time input behavior monitoring and event judgment, enables electronic devices to adaptively respond to different user interaction states. Without interrupting or interfering with the main immersive experience, it allocates appropriate data collection methods for different situations, thereby achieving automation and contextual adaptation of the emotional data acquisition path.
[0110] In some implementations, when the first mode is dynamically selected, calculating the comprehensive confidence score for quantifying the credibility of the emotion response label based on the multimodal data includes:
[0111] When the first mode is executed:
[0112] The dynamic selection of preset modes to obtain the subject's emotional response label and the initial confidence level corresponding to the emotional response label includes: receiving the input signal generated by the subject through virtual ray or eye movement confirmation via a semi-transparent non-modal annotation interface to generate an emotional response label, and simultaneously recording behavioral trajectory data to calculate the initial confidence level;
[0113] And / or, the calculation of the process quality score of the subject corresponding to each mode during the assessment process includes: calculating eye movement focus based on the eye movement trajectory data, wherein the eye movement focus is calculated based on the fixation density of the fixation point falling into the marked area, fixation stability, and a penalty factor positively correlated with saccadic speed, wherein the penalty factor is used to reduce the value of the eye movement focus;
[0114] Based on the interactive handle operation trajectory data, the interactive stability is calculated, which is obtained by calculating the reciprocal of the hand tremor amplitude and the smoothness of the sliding path.
[0115] The process quality score is calculated based on the eye-tracking focus and the interaction stability.
[0116] Specifically, when the system executes the first mode, its workflow comprises two core and logically related parts: the joint acquisition of emotion tags and initial confidence levels, and an in-depth evaluation of the quality of the annotation process. In the first part, the system acquires data by presenting a semi-transparent, non-modal annotation interface. A key feature of this interface is that it does not interrupt the user's viewing and interaction with the main VR scene. Users interact with this interface in two main ways: first, by using a controller to fire a virtual ray for pointing and clicking; and second, by directly focusing on a specific area of the interface and coordinating with eye-tracking confirmation (such as blinking or pausing). The system receives these input signals and maps them to specific emotion dimension values, thereby generating emotion response tags. Simultaneously, the system records the complete behavioral trajectory data that generated this tag input, such as the movement path of the virtual ray, the hovering duration before and after the click, and the eye-tracking jitter before confirmation. Based on this trajectory data, the system calculates an initial confidence level using predefined algorithms (such as analyzing abrupt changes in the path and calculating the deviation between the operation duration and the expected duration). This initial confidence level reflects the smoothness and decisiveness of the operation.
[0117] In the second part, the system initiates a refined quality audit of the annotation process, the core of which is calculating a comprehensive process quality score. This calculation relies on two types of high-precision data collected simultaneously. First, eye-tracking focus is calculated based on eye-tracking trajectory data. This indicator is a composite quantity, its calculation integrating the spatiotemporal distribution density of the gaze point falling into the annotation interaction area, the stability variance of the gaze point's dwell time on the target control, and a penalty factor positively correlated with the speed of high-speed saccades. The purpose of introducing the penalty factor is to reduce the eye-tracking focus value accordingly when the system detects that the user's gaze has rapidly and non-exploratoryly jumped or drifted near the annotation area before or during the selection process, thereby quantifying potential attentional distraction or hasty operation. Second, interaction stability is calculated based on the interactive controller's operation trajectory data. This indicator is characterized by quantifying the reciprocal of the amplitude of micro-tremors in the hand and the smoothness of the operation path relative to an ideal straight line. The smaller the amplitude of the micro-tremor, the more refined the muscle control; the closer the operation path is to a straight line and the smoother it is, the clearer the user's control intention and the more decisive the action. A higher interaction stability score indicates more precise and stable physical actions by the user. Ultimately, the system uses a predefined fusion function to combine the calculated eye-tracking focus and interaction stability to generate a single process quality score. This score directly and quantitatively reflects the overall quality and level of engagement of the user's "hand-eye coordination" when annotating emotions using a visual-motor approach, providing a core basis for determining whether an emotion label stems from a focused and stable operational process.
[0118] To address the challenges of high spatial freedom but lack of physical support and susceptibility to motion interference in VR environments, particularly the inherent limitations of 6DoF (six degrees of freedom) for operation accuracy, this method implements a "gaze-driven, gesture-coordinated" annotation execution mechanism. This approach takes the user's real-time gaze coordinates, controller spatial pose, and physical button status as input; it employs multi-layout rendering, viewpoint anchor tracking, adaptive non-modal transparency adjustment, and kinematic behavior quality verification as processing steps; ultimately outputting highly accurate emotion annotation labels.
[0119] This mode obtains tag data process
[0120] The execution entity uses virtual rays as the basic annotation entry point. Users click and rate the emotion by clicking on a semi-transparent, non-modal floating window (such as a ring or polar coordinate layout) using virtual rays, or they can instantly confirm the emotion by attaching the interface to the sub-center area of the eye's focus using eye-tracking data, thus obtaining an emotion label. .
[0121] This mode obtains confidence data process
[0122] By synchronously collecting "behavioral paralinguistic" features (such as trajectory smoothness, hovering duration, hand tremors, and operational stability) in the background, the feature set is defined as follows:
[0123]
[0124] s: trajectory smoothness; d: hovering duration; t: hand tremor frequency / amplitude; σ: operational stability. The executor can detect kinematic noise during the interaction process in real time. The noise is affected by fatigue and lack of physical feedback, and can be expressed as:
[0125]
[0126] , The weighting coefficients are: a larger t and a smaller σ, the higher the noise N. Simultaneously, objective behavioral quality verification is introduced, using a logistic regression function to identify whether the subject made a "decisive choice" or a "hesitant, accidental touch," with the output value... Used for confidence assessment in subsequent step two:
[0127]
[0128] The closer to 1, the more decisive the choice, that is, the smoother the selection process, the shorter the hover, and the lower the noise. Conversely, it is considered to be a hesitant and accidental touch. k1, k2, and k3 are constants.
[0129] Process quality score in this model ;
[0130] Among them, weight and It can be adaptively adjusted through supervised learning or experimental data, usually Prioritizing the objectivity of eye-tracking signals, this formula ensures a unified computational framework for both planar environments (such as overlaid UIs) and VR environments (such as 3D holographic interfaces).
[0131] Eye movement components (Complete calculation of five original visual attention metrics):
[0132]
[0133] in, : Fixation point location and distribution score, calculates the proportion of fixation points falling into the region of interest on the annotation interface, and supports Gaussian distribution modeling to capture fixation density. The fixation duration percentage is based on the ratio of fixed fixation duration to total annotation duration, and a time window is introduced for smoothing to filter out transient noise. Fixation stability is quantified by dividing the spatial variance of the fixation point by the reciprocal of the target area size, thus representing the spatiotemporal consistency of focus. Normalized gaze offset, measures the average Euclidean distance of the gaze point relative to the center of the pop-up or labeled target, and the maximum distance. It can be calibrated according to screen / view size. : Sagging speed penalty item, using an exponential decay function to suppress abnormally high-speed sagging (which may be caused by emotional excitement or interference). For adjustable sensitivity coefficients, For environment-specific thresholds. For feature normalization and nonlinear activation functions (such as ReLU or Tanh). The weight vectors are learnable and support end-to-end training to optimize modal contributions. The function is a Sigmoid, ensuring an output range of [0,1]. This component verifies the focus of attention in real time using multi-source eye-tracking data (such as head-mounted display built-in modules, desktop eye trackers, or camera iris localization), avoiding subjective bias and ensuring label timing alignment.
[0134] Interface interaction components :
[0135]
[0136] in, Gesture stability is calculated based on the 3D position variance of key points on both hands, quantifying the clarity of the operation intention. Normalize the amplitude of hand tremors and use Fourier transform (FFT) to extract high-frequency components (>8Hz) to represent the effects of fatigue or tension. : Hand speed penalty term, an exponential function to suppress abnormal deviations from the optimal speed range, μ is the adjustment coefficient. This is within an acceptable speed range. : Smoothness of the sliding selection path, the ratio of the actual path length to the ideal straight line length, a value close to 1 indicates efficient operation. Laser pointer position and path offset, the reciprocal of the average path deviation distance, supports ray tracing data. : Continuity of button operation on the controller, evaluate the rationality of button timing and frequency abnormalities, and avoid accidental touches. The percentage of effective dwell time on the interface, relative to the expected duration, reflects cognitive load. Penalty for interface switching and adjustment frequency, quantifying operational burden. Dynamic content observation behavior score, calculated based on the continuity of user response to UI changes. For the normalization function of multi-source interaction features, This is a weighted vector. This component supports multiple input methods (such as mouse, touch, gamepad raycasting, and bare-hand interaction), effectively reducing the cost of attention switching.
[0137] RS consists of user interface interaction components With eye movement components The weighted fusion is constructed, with a range of [0,1]. Among them, It reflects the time users spend on the annotation interface, the smoothness of path selection, and the frequency of interface adjustments, thus characterizing the smoothness of operation. By comprehensively judging whether the user's gaze is truly focused on the marked target by considering gaze stability, gaze deviation, and saccade speed, the system can effectively identify erroneous operations caused by ambient lighting, head posture, or distraction.
[0138] In some implementations, when the second mode is executed:
[0139] The process of dynamically selecting a preset mode to obtain the subject's emotional response label and the initial confidence level corresponding to the emotional response label includes: processing the captured natural language speech to obtain the emotional response label, and extracting acoustic features to calculate the semantic and acoustic consistency as the initial confidence level.
[0140] And / or, the calculation of the process quality score of the subject corresponding to each mode during the evaluation process includes: processing the captured speech audio data to obtain speech recognition content and extracting acoustic features;
[0141] The clarity of expression is calculated based on the speech rate and pause ratio of the speech recognition content;
[0142] The similarity between the semantic vector corresponding to the speech recognition content and the emotion tendency vector obtained from the acoustic feature analysis is calculated as semantic consistency.
[0143] The process quality score is calculated based on the clarity of expression and the semantic consistency.
[0144] Specifically, when the system executes the second mode, its process also includes two logical stages: data acquisition and process quality assessment, but the core technology shifts to speech processing. In the first stage, namely the joint acquisition of labels and initial confidence, the system captures the subject's natural language speech stream through a microphone array or integrated microphone. This speech stream is first converted into text-based speech recognition content by an automatic speech recognition engine. Subsequently, the system uses a pre-built emotion dictionary mapping model or semantic understanding module to parse and map the text content (such as "I feel excited" or "This makes me anxious") into continuous emotion dimension coordinate values, thereby generating structured emotion response labels. At the same time, the system extracts a set of acoustic features in parallel from the original speech signal. These features include, but are not limited to, fundamental frequency, speech rate, energy spectrum, and formant structure, which together encode the paralinguistic information of the speech. The system calculates the degree of consistency between the semantic emotion tendency derived from the text content and the acoustic emotion tendency derived from the acoustic features (for example, if a user says "calm" but the tone is sharp and abrupt, the consistency is low), and uses this consistency measure as the initial confidence of this speech labeling. This initial confidence level directly reflects the superficial consistency between what the user "says" and "how they say it" in terms of emotional expression.
[0145] In the second stage, the in-depth process quality assessment, the system performs more refined analysis of the same audio data to calculate a process quality score. This process begins with a dual analysis of the audio data: on the one hand, deepening speech recognition to obtain more accurate text content; on the other hand, extracting more comprehensive acoustic features. The calculation of the process quality score mainly relies on two dimensions of indicators. The first dimension is speech clarity, which is calculated by analyzing the statistical characteristics of speech rate reflected in the speech recognition content (such as average speech rate and variance of speech rate variation) and the proportion of unnecessary pauses (such as "uh" and "ah") in the sentence. A stable speech rate and reasonable pauses usually correspond to high speech clarity, while abnormal fluctuations in speech rate or a large number of hesitant pauses will cause the index to drop, reflecting that the user may be in a state of tension, thinking, or inattentiveness. The second dimension is semantic consistency, which is a deeper verification indicator. The system transforms the speech recognition content into a semantic vector representing its literal meaning through a natural language processing model. At the same time, the system inputs the extracted acoustic features into a trained acoustic-emotion mapping model to generate an emotion tendency vector representing its implicit emotional color. Subsequently, the system calculates the similarity (e.g., cosine similarity) between the two vectors in a high-dimensional space, serving as a quantitative value for semantic consistency. If the literal meaning of a user's utterance highly matches the emotional tone and timbre conveyed, semantic consistency is high; conversely, if there is misunderstanding or insincerity, consistency significantly decreases. Finally, the system integrates these two indicators—clarity and semantic consistency—using a pre-defined algorithm (e.g., weighted combination) to output a comprehensive process quality score. This score not only assesses the fluency and comprehensibility of the user's expression but, more importantly, deeply evaluates the inherent consistency and potential authenticity of their verbal emotion reports from the perspective of content and form matching, providing crucial quality filtering criteria for speech-based emotion data.
[0146] For scenarios requiring high immersion or where participants' hands are occupied, this mode provides a contactless voice annotation path. To maintain the annotation pathway in extreme environments with limited field of vision or extremely high attention requirements (no spare energy or pop-up interaction), this mode guides participants to verbally describe their current emotional characteristics in natural language by triggering intermittent voice prompts or non-modal pop-ups, thereby utilizing user-generated emotional annotation words to achieve real-time data completion.
[0147] This mode obtains tag data process
[0148] The system automatically captures the subject's speech data using speech activity detection technology. The acquired speech signals are processed by an emotion semantic understanding unit, which performs keyword extraction and intent parsing logic to identify core emotion descriptors (such as "nervous," "excited," and "suppressed"). .
[0149] At the same time, a multi-level mapping matrix from "natural language" to "emotion space" is established: first, the parsed core emotion words are used as input vectors and entered into a pre-constructed "emotion dictionary set";
[0150]
[0151] : Target emotion space vector, representing valence, arousal, and dominance, which are the labels obtained in this model.
[0152] : Input discrete emotion word One-hot encoding or word vector.
[0153] : A pre-built emotion dictionary mapping matrix used to perform linear or non-linear transformations from discrete space to continuous space.
[0154] This mode obtains confidence data process
[0155] By using a semantic mapping function, discrete emotional words are mapped to continuous valence-arousal-dominance (VAD) coordinate values, enabling real-time data quantification without the need for a visual interface. Simultaneously, acoustic emotional features of the subject's speech (such as pitch variation, speech rate, and intonation fluctuations) are extracted.
[0156]
[0157] : Fundamental frequency (pitch variation). Speech rate. : intonation fluctuations (amplitude changes).
[0158] Finally, the cosine similarity between the semantic vector and the acoustic derivation is calculated. This is used for the voice master control reliability assessment in step two.
[0159]
[0160] Based on acoustic features through functions The predicted emotional space projection.
[0161] Consistency matching coefficient, range of values .
[0162] A value close to 1 indicates that the acoustic features are consistent with the content, meaning that the speech content truly reflects the current emotion. Figure 2 This invention provides an example of an emotion labeling method based on voice prompts.
[0163] When the second mode is selected, the process quality score is obtained by conducting a voice master reliability assessment. This score is used to quantify the quality of the subject's natural language expression and is intended to identify whether the subject is inattentive or in a perfunctory state.
[0164] ;
[0165] RS consists of speech recognition components With gesture recognition components The system is composed of components ranging from [0,1]. The speech component integrates confidence, clarity, and duration matching to resist environmental noise interference; the gesture component filters out false triggers caused by tension, fatigue, or unintentional movements through stability, speed appropriateness, and jitter suppression. This mechanism allows speech to undertake the main control and gestures to handle fine-tuning, with the two mutually verifying each other, significantly reducing cognitive load.
[0166] Among them, speech recognition components :
[0167]
[0168] The duration of speech pauses: excessively long pauses usually indicate that the user is hesitant, thinking, or distracted. The threshold can be dynamically set based on the average speech rate of the task; the closer the normalized value is to 1, the more reasonable the pause. Changes in speech rate, particularly drastic fluctuations, reflect a user's emotional excitement, tension, or external interference. The exponential decay mechanism minimizes the impact of minor changes, while drastic changes result in significant point deductions. The sensitivity coefficient is denoted as .
[0169] Pitch / volume fluctuations The amplitude of the fundamental frequency variation. This represents the amplitude of sound pressure level variation. Abnormal increases are usually associated with emotional arousal or tension, and the Euclidean distance is used as the comprehensive penalty after normalization. The voice jitter penalty item integrates the normalized indicators of local jitter, rapid jitter and flicker. A high value indicates that the vocal organs are affected by emotions or fatigue. Prosodic stability score is calculated by comparing the current speech rhythm and stress distribution with the user's personal template in a calm state using cosine similarity. The closer the value is to 1, the clearer and more focused the annotation intent is.
[0170] gesture components (Fully integrates all original hand / pointing interaction metrics):
[0171]
[0172] Gesture stability is calculated based on the standard deviation of the 3D positions of key points on both hands. The smaller the fluctuation, the clearer the operation intention. Normalized hand tremor amplitude, Fourier transform is used to extract high-frequency energy >8 Hz, reflecting fatigue, tension or physiological tremor; Hand speed rationality penalty item, exponential function to inhibit overly fast (impulsive) or overly slow (hesitant) movements. This is the adjustment coefficient; : Smoothness of the sliding selection path, the ratio of the actual trajectory length to the theoretical straight line length. The closer the value is to 1, the more efficient and deliberate the operation. : Laser pointer (or ray) position offset penalty, the normalized distance of the average deviation from the target center, the smaller the distance, the higher the score; : Analyze the continuity and regularity of button operation on the controller, assess the frequency and interval of button sequences to prevent accidental touches or panicked operation; Overall hand trajectory deviation and variance, quantifying whether the path deviates from the expected target trajectory; Repetitive or rework-related penalties: The more times the same operation is repeated in the same area, the less confident or focused the user is.
[0173] In some implementations, when the third mode is executed:
[0174] The dynamic selection of preset modes to obtain the subject's emotional response labels and the initial confidence level corresponding to the emotional response labels includes: obtaining a pre-trained emotion recognition model, wherein the emotion recognition model is obtained by joint training using data induced by homologous emotional stimulus materials in a planar environment and a VR environment through a domain adaptation method;
[0175] The emotional stimulus material of the current induced scenario is input into the emotion recognition model to obtain the basic emotion prediction value and the initial confidence level.
[0176] Based on the immersion parameters of the current induced scene, the basic emotion prediction value is enhanced and corrected to generate emotion response tags;
[0177] And / or, the calculation of the process quality score of the subject in each mode during the evaluation process includes: obtaining the pre-collected homologous behavioral baseline features of the subject in a planar environment, calculating the deviation of the real-time behavioral features in the VR environment from the baseline features, and calculating the cross-domain attention behavior reliability score as the process quality score based on the deviation.
[0178] Specifically, during the execution of the third mode, the electronic device invokes a pre-trained emotion recognition model. This model is trained using a domain-adaptive approach, with training data derived from a specific set of homologous emotional stimuli. Each piece of material in this set was collected in two environments: one in a 2D display environment with sensory enhancement, and the other in the target virtual reality environment. Through this method, the model is forced to learn, during training, to extract common emotional representation features related to the stimulus content itself from data presented in two different media, while implicitly modeling the inter-domain mapping relationship from the 2D environment to the VR environment. In the application phase, the electronic device inputs the emotional stimuli of the current inducing scene—typically referring to the audiovisual content presented to the user or its abstract representation after feature extraction—into this pre-trained model. The model infers based on its learned cross-domain features and outputs a basic emotion prediction value, reflecting the model's estimate of the possible induced emotions under baseline conditions without considering environmental differences. Subsequently, the electronic device introduces a parameter characterizing the immersion level of the current VR scene, which can be quantified based on the scene's visual field angle, interaction degrees of freedom, spatial audio intensity, or a preset experimental level. Electronic devices adjust the baseline emotion prediction value using this immersion parameter based on predefined correction rules or functions. Common corrections include non-linear amplification of the emotional arousal dimension or shifting of the emotional valence dimension. The result after this enhancement and correction process is output as the emotion response label for that mode. This process enables the automatic generation of emotion state estimates that consider both the essence of the stimulus content and the potential amplification effect of the high-immersion environment on emotional intensity without active interaction from the subject, making continuous, non-invasive emotion monitoring possible.
[0179] Due to the inconvenience of labeling and manipulation in VR environments, this model proposes an "automatic labeling" method based on transfer learning, which is a label acquisition method that does not require participant interaction. By constructing a full-chain transfer learning framework of "planar induction (source domain) - VR induction (target domain)," it achieves automated label acquisition without human intervention. This automatically and instantly acquires emotional response labels without interfering with the induction process, solving the problems of high cost and data distortion caused by the difficulty of manipulation in deep induction environments. This framework ensures high similarity between predicted labels and deep VR induction states through the following three levels of deep correction and collaboration:
[0180] The process of obtaining tag data in this mode
[0181] 1. Physical Layer: Source Domain Environment Enhancement Based on Sensory Alignment
[0182] The physical layer, as the foundation of transfer learning, is a prerequisite for subsequent feature and parameter layers. Its purpose is to reduce the "inter-domain gap" through physical means during the planar data acquisition stage, and is defined as follows:
[0183] ;
[0184] : The physical domain gap.
[0185] Sensory stimulus intensity function.
[0186] These refer to flat panel display and VR display environments, respectively.
[0187] The set of environmental enhancement parameters (including curved screen curvature, sound field immersion, ambient light shielding rate, etc.) can be adjusted to create a highly immersive planar environment by covering the visual field with a large curved screen, constructing a dark room environment, and using noise-canceling headphones and spatial audio. This reduces the differences between domains and enables the smooth transfer of prior knowledge at the physical level (such as wakefulness benchmark).
[0188] 2. Feature Layer: Implicit Domain Adaptation Based on Homologous Stimuli
[0189] After aligning the physical environment, the executing agent eliminates environmental noise using a controlled variable method, implicitly internalizing the immersion difference. That is, the executing agent does not input any labels or parameters about the "environment" into the model, but instead allows the model to learn autonomously through a training mechanism. Representative personality trait groups are selected, and data is collected in both the aligned 2D and VR environments using identical, homologous standard emotional stimulus materials (such as a high-definition video or the same image).
[0190] Feature extraction consistency constraints:
[0191] ;
[0192] : Emotionally stimulating materials of the same source (such as video and image content).
[0193] Feature mapping encoder.
[0194] Cross-domain common emotional characteristics, which are not sensitive to the environment and are only highly correlated with the content of the stimulus.
[0195] Implicit loss function:
[0196] ;
[0197] : Metric function (such as MMD distance or KL divergence). By minimizing this loss, the model is forced to automatically ignore the interference caused by differences in immersion.
[0198] This "homogeneous dual-environment acquisition" forces the model to ignore environmental interference and forcibly extract common emotional features that are highly related to the stimulus content. Through joint training across mixed domains, the executing agent transforms the induced depth differences brought about by immersion (such as mild sadness in a flat scene to deep grief in a VR scene) into implicit label enhancement weights for the model.
[0199] 3. Parameter Layer: Explicit Tag Mapping and Depth Correction Based on Immersion Parameter Enhancement
[0200] Building upon the physical and feature alignment, the execution entity introduces explicit mathematical compensation and minimizes it. Training emotion classification / regression networks Environmental differences are quantified using explicit mathematical parameters to achieve the final depth label prediction. The execution agent quantifies "immersion" into an independent enhanced input vector.
[0201] Define immersion parameters This involves constructing a composite input of "stimulus features + immersion coefficient." Through modeling methods such as deep neural networks, the model can... The changes deeply modify and enhance the initial emotional characteristics.
[0202] ;
[0203] The corrected VR deep emotion label is the label obtained in this mode.
[0204] Basic emotion classification / regression network.
[0205] Immersion parameter, used to linearly or non-linearly map the “light” sensory experience of a 2D scene to the “depth” induced result of a VR scene.
[0206] Using the above method, even in the absence of real-time manual annotation during the depth-inducing process, the executing entity can automatically generate high-fidelity emotion tags with VR depth features based on the immersion coefficient of the current scene.
[0207] The process of obtaining confidence data in this mode
[0208] This mode uses automatic label acquisition, therefore its confidence level largely depends on the classification model. Therefore, the model's classification accuracy is directly used as confidence data. It participates in the calculations of subsequent steps.
[0209] When the system executes the third mode, since this mode does not rely on the user's active interaction, its process quality score calculation logic focuses on assessing whether the user's natural behavioral patterns in the absence of instructions are within the expected attention range. The core method of this assessment is to conduct a personalized behavioral baseline comparison across environments (planar and VR). Specifically, the system first acquires the subject's pre-collected behavioral baseline characteristics in a traditional planar display environment. These characteristics are collected and calculated by devices such as eye trackers when the subject watches emotionally stimulating material (such as a video) that is of the same origin as the current VR induced scene. They typically include the subject's habitual average gaze duration, typical saccade speed patterns, and the entropy value of gaze distribution, which together constitute a baseline feature vector characterizing the user's personal "normal" attention pattern.
[0210] During real-time evaluation in a VR environment, the system simultaneously collects the user's real-time behavioral characteristics in the current induced scenario, primarily including eye-tracking data (such as fixation point and saccades) and head posture data. The system then calculates the deviation between these real-time characteristics and the aforementioned personalized baseline characteristics. This deviation calculation is not a simple difference but takes into account the inherent differences brought about by environmental transitions. For example, because the VR environment provides a panoramic view, it naturally induces more exploratory saccades and head rotations. Therefore, before comparison, the system adaptively maps the planar baseline characteristics using a predefined "environmental gain coefficient" to generate an "expected behavioral reference range in the VR environment." The calculated deviation reflects the degree of abnormality of the user's real-time behavior compared to their personal habits and the characteristics of the current environment.
[0211] Finally, the system calculates a cross-domain attention behavior reliability score based on this deviation and uses it as the process quality score in the third mode. Its calculation logic follows a non-linear threshold judgment: if the deviation between the real-time behavioral characteristics and the mapped baseline characteristics is small and within the preset "focused behavior tolerance range," the user is determined to be in a state of deep immersion and stable attention, and is given a high process quality score; conversely, if the deviation significantly exceeds the tolerance range (e.g., a large number of aimless, high-speed random glances, or prolonged head stillness), the user is determined to be in a state of inattention, distraction, or severe interference, and the process quality score will be significantly reduced. This mechanism, through personalized cross-domain calibration, effectively distinguishes between "natural exploratory behavior caused by environmental characteristics" and "abnormal behavior caused by attention loss," thus providing crucial behavioral state quality evidence for the emotion response label in the automatic prediction mode.
[0212] When the third mode is selected, the process quality is divided into cross-domain attention behavior reliability assessment. In order to prevent misjudgment caused by the difference in user behavior patterns in the planar (limited field of view) and VR (panoramic field of view) environments in step one (such as the panoramic characteristics of VRHMD will induce subjects to generate more exploratory saccades, with more intense head movements and a naturally higher saccade rate. If a unified standard is directly applied, it will be misjudged as "distraction"), the executing entity establishes a "planar-VR attention mapping benchmark".
[0213] During the pre-induction or baseline acquisition phase, planar data is used to learn the user's "normal attention pattern" as a personalized reference set, quantifying the user's "habitual" attention feature vector in a planar environment. :
[0214]
[0215] Average fixation duration in a planar environment. Average scanning speed in a planar environment. The view distribution entropy in a planar environment represents the habitual search breadth. Since the panoramic nature of VR naturally increases the scan rate, we introduce a mapping operator. Transform the planar datum into a reference datum in the VR environment. :
[0216]
[0217] Cross-domain mapping function. : Ambient gain vector (e.g., calibration coefficient for saccade velocity due to panoramic field of view) ). The Hadamard product (element-wise multiplication) represents the differential scaling of features across different dimensions.
[0218] The formal induction and data labeling phases are achieved through calculation. Real-time VR behavioral characteristics of subjects at any given moment Deviation from reference standard If the user is in a "high arousal zone" where they are more focused than usual, the reliability weight is increased; if "unconscious saccades" occur, it is marked as low reliability.
[0219]
[0220] Real-time eye and head movement characteristics of subjects in a VR environment.
[0221] Deviation. If this value is within a reasonable threshold, it indicates that the current behavior conforms to personal habits; if it significantly exceeds the range of environmental gain, it is judged as abnormal. The reliability weight of the data is determined based on the deviation and the current emotional arousal range. Perform nonlinear correction:
[0222]
[0223] The tolerance limit for focused behavior.
[0224] Weighted rewards for determining when a subject is in an "immersive / focused" state.
[0225] : Penalty coefficient for "unconscious saccades" or "gaze drift".
[0226] In some embodiments, the method further includes:
[0227] The physiological arousal verification index is obtained by weighting and summing the skin conductance amplitude, heart rate variability and the reciprocal of electromyography signal in the physiological signal data with preset weighting coefficients.
[0228] Calculate the matching error between the subjective arousal level represented by the emotion response label and the physiological arousal verification index;
[0229] The overall confidence score is calibrated based on the matching error, wherein the lower the matching error, the higher the overall confidence score.
[0230] Specifically, the electronic device calculates a physiological arousal verification index based on synchronously acquired physiological signal data, used to objectively calibrate the confidence level of emotional response labels. This index is calculated by integrating the instantaneous amplitude of the skin conductance response signal, a specific measure of heart rate variability, and the reciprocal of the electromyography (EMG) signal amplitude. These three physiological parameters reflect sympathetic nerve excitability, the dynamic balance of the autonomic nervous system, and muscle tension, respectively. The electronic device uses a set of preset weighting coefficients to weight and sum these three parameters, thus fusing them into a comprehensive physiological arousal verification index. Subsequently, the electronic device parses the subjective arousal component represented by the currently processed emotional response label. This component can be an arousal value in an emotion dimension model such as valence-arousal-dominance. The electronic device calculates the matching error between this subjective arousal value and the aforementioned physiological arousal verification index. This error can be expressed as an absolute difference, a relative difference, or a normalized distance metric. Finally, the electronic device calibrates the initially determined overall confidence score based on the calculated matching error. The calibration logic follows the principle that matching error is negatively correlated with confidence score; that is, the lower the matching error, the more closely the user's subjectively reported emotional intensity matches their physiological response state, resulting in a higher calibrated overall confidence score. Conversely, a higher matching error leads to a lower overall confidence score. This calibration step introduces an objective physiological reference independent of subjective reports and interaction behavior, effectively identifying and correcting inconsistencies between subjective and objective emotional reports caused by social expectation bias, arbitrary responses, or insufficient introspection, thereby improving the robustness of the final confidence assessment results.
[0231] Physiological arousal verification index As an "objective lie detector," it does not directly generate labels, but verifies whether the user's subjective feedback matches the body's reaction. It eliminates truth value distortion caused by social expectation bias (the feeling of filling in "correct" rather than "true") or arbitrary filling by the subjects through arousal measurement.
[0232] ;
[0233] Physiological arousal index: the higher the index, the stronger the physiological arousal.
[0234] : The current level of skin conductance.
[0235] : The mean and standard deviation of subjects under baseline conditions.
[0236] Calculate the arousal tags subjectively filled in by the subjects. In relation to actual physiological performance The residuals between them.
[0237] ;
[0238] Consistency matching error.
[0239] : Subjective arousal level (e.g., 1-9 points) as labeled by the subjects.
[0240] : The normalization function for subjective labels (mapping them to 0-1 space).
[0241] Normalization of physiological indices.
[0242] Based on the magnitude of the error, the system automatically determines the authenticity of the annotation, identifying either "social expectation bias" or "arbitrary filling":
[0243]
[0244] : The final confidence coefficient of this annotation.
[0245] Truth distortion threshold.
[0246] Case A (High Match): When the error When the score is low, the confidence level increases linearly with the matching degree.
[0247] Case B (Truth Distortion): When the error exceeds the threshold (e.g., verbally saying "very excited" but the PVI shows "calm"), it is judged as social expectation bias or arbitrary filling, and the confidence level is forcibly reduced to an extremely low value. (Close to 0).
[0248] In some implementations, in the first mode, the step of rendering a labeled interface that does not affect the continuous presentation and interaction of the current VR scene content includes:
[0249] Based on the real-time monitoring of the task load of the induced scenario and the emotional arousal level of the subjects, the transparency and rendering scale of the annotation interface are dynamically adjusted.
[0250] In the pre-defined key nodes of the induced scenario, the labeling interface is automatically suspended or weakened.
[0251] Specifically, the rendering of the annotation interface, which does not affect the continuous presentation and interaction of the current VR scene content, is achieved through a context-aware interface scheduling mechanism. The electronic device monitors two dynamic parameters in real time: first, the cognitive and perceptual task load imposed by the current induced scene itself, which can be estimated based on the number of objects in the scene, the complexity of movement, and the urgency of the interaction goal; second, the subject's real-time emotional arousal level, which can be indirectly inferred through synchronously analyzed physiological signals or prior behavioral data. Based on the real-time status of these two parameters, the electronic device dynamically adjusts the visual presentation attributes of the annotation interface, mainly including the overall transparency of the interface and the rendering size ratio in three-dimensional space. When the task load or emotional arousal level exceeds a preset threshold, the electronic device automatically increases the interface transparency or reduces its display size to reduce visual occlusion and cognitive interference with the main scene content. Furthermore, based on a pre-structured analysis of the induced content (such as narrative lines, game levels, and experimental procedures), the electronic device defines a series of key nodes, such as plot turning points, high-difficulty task execution segments, or sudden emotional stimuli. When an electronic device detects that the current scene has entered such a preset critical node, it will automatically trigger the interface suspension logic, temporarily suspending the display and interactive response of the annotation interface, or reducing its visual salience to an extremely low level. This series of mechanisms ensures that the annotation interface is only presented with high availability when the user's cognitive resources are relatively abundant and the emotional induction of the main scene is not at a critical climax. Thus, while providing annotation functionality, it maximizes the immersiveness of the virtual reality experience and the continuity of the emotional induction process, avoiding the problem of interruption or distortion of emotional response due to improper interface pop-ups.
[0252] In some implementations, in the first mode, rendering the annotation interface further includes the following positioning modes:
[0253] A fixation-associated positioning mode is used to dynamically attach the labeled interface to a subcentral region outside the subject's focal point of vision based on the eye movement data.
[0254] The user-defined visual field anchor point positioning mode is used to make the annotation interface move with the subject's head rotation based on the comfortable visual field range calibrated during the initialization phase and real-time head posture data.
[0255] Specifically, in the first mode, the spatial positioning of the annotation interface is achieved through two selectable intelligent positioning modes, aiming to solve the problems of interface drift, occlusion, or difficulty in reaching in three-dimensional space. In the gaze-associated positioning mode, the electronic device continuously receives data streams from the eye-tracking module and calculates the focal point of the subject's gaze in the virtual space in real time. Instead of directly placing the annotation interface on this focal point, the electronic device dynamically attaches it to a ring-shaped or fan-shaped area around the focal point according to predefined offset rules; this area is called the secondary central region. This region is located at the edge of the user's central visual field but can be perceived without significant gaze movement, thus achieving a visible, following interface position while avoiding direct occlusion of the core observation target. In the user-defined visual field anchor point positioning mode, the electronic device guides the user through a calibration process during the experiment or session initialization phase. During this process, the electronic device records a visually comfortable and fatigue-free visual field range determined after multiple scans of the user's head in a naturally relaxed state, serving as the comfortable visual field. In subsequent use, the electronic device calculates the ideal follow-up position of the annotation interface in the virtual world based on real-time orientation data provided by the head posture sensor and the spatial anchor point of this comfortable visual field determined during the initialization phase. This allows the annotation interface to move stably and synchronously with the user's head rotation, always remaining within the user's easily accessible field of vision. This avoids the problem of the interface moving out of the user's view and requiring active searching when the user freely explores the environment. These two positioning modes together enhance the usability and stability of two-dimensional interface interaction in a dynamic, free six-DOF VR environment, reducing the additional cognitive load and operational interruptions caused by the user searching for or locating the interface.
[0256] All interactive controls utilize a semi-transparent, non-modal floating window format. "Non-modal" means that when the interface pops up, it doesn't forcibly lock the focus of the executing entity or interrupt the current logic, ensuring that the subject can annotate without completely obscuring the induced content. Compared to "scene switching" that occupies induced space or "abrupt pop-ups" that forcibly interrupt the induced flow, this semi-transparent floating window method effectively maintains a sense of presence. The executing entity has a built-in scheduler that automatically and dynamically adjusts the transparency and rendering scale of controls based on real-time monitoring of task load and emotional arousal level; it automatically suspends or weakens UI salience at critical moments in battle or plot, achieving contextual awareness on a time scale, reducing emotional interference caused by pop-ups, and ensuring the integrity of the emotional experience.
[0257] The execution entity supports dynamic switching of various pop-up control forms (such as circular, polar coordinate layout, and upper / lower / left / right quadrant partition layout) and display positions to adapt to different users' spatial attention habits.
[0258] In terms of visual feedback, the executor maps the sliding operation into a continuous visual animation: for example, using color space gradients to represent shifts in emotional valence, and using the contraction or expansion of fluid shapes to represent changes in arousal intensity. This real-time feedback from the "visual-motor" loop compensates for the lack of physical touch, enhances the user's perceptual precision, further improves the accuracy of annotation results, and at the same time lowers the barrier to entry and reduces the interference caused by annotation.
[0259] Finally, the implementing entity provides two intelligent positioning modes: Mode A (Gaze-Associated Positioning Mode) combines eye-tracking data to attach the interface to the secondary central area of the gaze focus (the central area is the main area of the guiding material, the center of gravity of the material; the secondary central area is visible but not interfering, and is displayed as a less important part of the guiding material), achieving "no search required, instant visibility". The core of this mode is to center the interface. Positioned as the focal point Nearby, while avoiding the center of gravity of the main inducement area. .
[0260] Positioning formula:
[0261] ;
[0262] : The subject's current real-time gaze focus coordinates.
[0263] : An offset vector relative to the view direction, used to push the interface towards the edge of the field of view.
[0264] : The geometric centroid of the guiding material (main mission area).
[0265] Mode B (User-defined field of view anchor point mode) locks in a low-interference field of view through "comfortable field of view calibration" during the initialization phase, and uses head posture calculation to achieve motion tracking, effectively solving the problems of interface drift and interaction path interruption caused by scene object movement or head rotation, thus improving annotation efficiency. This mode involves initialization calibration and real-time motion tracking calculation.
[0266] During comfort field of view calibration, record the relative position anchor points of the interface in the head coordinate system. :
[0267] ;
[0268] : Initialize the head pose transformation matrix (4×4 matrix) at the instant.
[0269] The interface world coordinates when the user adjusts to a comfortable position.
[0270] During the interaction, based on the current head posture Calculate the world coordinates of the interface in real time to eliminate drift:
[0271] ;
[0272] : The real-time location of the time display.
[0273] The current head rotation and displacement matrix calculated by sensors (IMU / optical tracking). Figure 3 The diagram shown is a flowchart of a gaze-point association localization method provided by the present invention. Figure 4 The diagram shown is a flowchart of a user-defined view anchor point positioning method provided by the present invention. Figure 5 The image shown is an example of the pop-up control form provided by the present invention. This example is a ring-shaped control.
[0274] In this mode, the executing entity perceives the user's state (eye contact, hand micro-operations) and transforms the originally rigid "form filling task" into an intelligent, dynamic, automatically calibrated, and non-interfering closed-loop interactive system. This reduces interference with the induction process and increases the accuracy and confidence of the labeled data.
[0275] Example 2
[0276] Please see Figure 6 This invention provides an emotion response assessment device for a VR environment, comprising:
[0277] The emotion response label acquisition module 601 is used to dynamically select a preset mode in a virtual reality (VR) environment to acquire the subject's emotion response label and the initial confidence level corresponding to the emotion response label, based on the current induced scene and the subject's state. The emotion response label is a numerical value or vector representing the emotion state. The preset mode includes any of the following modes: First mode, by rendering an annotation interface that does not affect the continuous presentation and interaction of the current induced scene content, and collecting the subject's eye-tracking data and interactive handle data to complete the annotation input, the emotion response label is obtained; Second mode, by mapping the captured and parsed natural language speech of the subject to emotion response labels; Third mode, by using an emotion recognition model pre-trained in a planar environment and the immersion parameters of the current induced scene to predict and generate emotion response labels.
[0278] The confidence data acquisition module 602 is used to simultaneously collect multimodal data of the subject associated with the data acquisition event that generates the emotion response label. The multimodal data includes eye movement trajectory data and interactive handle operation trajectory data corresponding to the first mode, voice audio data corresponding to the second mode, real-time behavioral features corresponding to the third mode, and physiological signal data corresponding to each mode.
[0279] The process scoring module 603 is used to calculate the process quality score corresponding to each mode and the physiological verification score corresponding to all modes of the subject during the assessment process based on the multimodal data.
[0280] The comprehensive confidence module 604 is used to calculate the comprehensive confidence score based on the initial confidence score, process quality score, and physiological verification score corresponding to each mode;
[0281] The output module 605 is used to associate the emotion response label, the comprehensive confidence score, and the preset pattern for generating the emotion response label with a timestamp, and output them as a structured data packet.
[0282] It should be noted that each module and unit in the VR environment emotion response assessment device in this embodiment corresponds one-to-one with each step in the VR environment emotion response assessment method in the aforementioned embodiment. Therefore, the specific implementation of this embodiment can refer to the implementation of the aforementioned VR environment emotion response assessment method, and will not be repeated here.
[0283] Example 3
[0284] Please see Figure 7 This embodiment provides an electronic device, including at least one processor 701 and a memory 702. Optionally, the device further includes a communication component 703. The processor 701, memory 702, and communication component 703 are connected via a bus 704.
[0285] In a specific implementation, at least one processor 701 executes computer execution instructions stored in memory 702, causing at least one processor 701 to perform the above-described method.
[0286] The specific implementation process of processor 701 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0287] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0288] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0289] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0290] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0291] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0292] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0293] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0294] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0295] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0296] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0297] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0298] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0299] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for assessing emotional responses in a VR environment, characterized in that, include: In a virtual reality (VR) environment, based on the current induced scene and the subject's state, a preset mode is dynamically selected to obtain the subject's emotional response label and the initial confidence level corresponding to the emotional response label. The emotional response label is a numerical value or vector representing the emotional state. The preset mode includes any of the following modes: First mode, by rendering an annotation interface that does not affect the continuous presentation and interaction of the current induced scene content, and collecting the subject's eye movement data and interactive handle data to complete the annotation input, the emotional response label is obtained. The second mode maps the captured and parsed natural language speech of the subjects into emotion response labels; The third mode uses a pre-trained emotion recognition model in a planar environment and immersion parameters of the current induced scene to predict and generate emotion response labels. For the data acquisition event that generates the emotional response label, multimodal data of the subject associated with the data acquisition event is collected simultaneously. The multimodal data includes eye movement trajectory data and interactive handle operation trajectory data corresponding to the first mode, voice and audio data corresponding to the second mode, real-time behavioral features corresponding to the third mode, and physiological signal data corresponding to each mode. Based on the multimodal data, calculate the process quality score for each mode and the physiological verification score for all modes during the assessment process; The overall confidence score is calculated based on the initial confidence level, process quality score, and physiological verification score corresponding to each mode; The emotion response label, the overall confidence score, and the preset pattern for generating the emotion response label are associated with a timestamp and output as a structured data packet.
2. The method according to claim 1, characterized in that, The dynamic selection preset mode includes: Monitor the subject's real-time input behavior; Based on the real-time input behavior, the currently active annotation mode is identified, wherein the first mode is executed when an action confirmation event is detected by ray tapping with the controller or eye tracking; the second mode is executed when a voice input event is detected; and the third mode is executed when no action confirmation event or voice input event is detected.
3. The method according to claim 1, characterized in that, When the first mode is executed: The dynamic selection of preset modes to obtain the subject's emotional response label and the initial confidence level corresponding to the emotional response label includes: receiving the input signal generated by the subject through virtual ray or eye movement confirmation via a semi-transparent non-modal annotation interface to generate an emotional response label, and simultaneously recording behavioral trajectory data to calculate the initial confidence level; And / or, the calculation of the process quality score of the subject corresponding to each mode during the assessment process includes: calculating eye movement focus based on the eye movement trajectory data, wherein the eye movement focus is calculated based on the fixation density of the fixation point falling into the marked area, fixation stability, and a penalty factor positively correlated with saccadic speed, wherein the penalty factor is used to reduce the value of the eye movement focus; Based on the interactive handle operation trajectory data, the interactive stability is calculated, which is obtained by calculating the reciprocal of the hand tremor amplitude and the smoothness of the sliding path. The process quality score is calculated based on the eye-tracking focus and the interaction stability.
4. The method according to claim 1, characterized in that, When the second mode is executed: The process of dynamically selecting a preset mode to obtain the subject's emotional response label and the initial confidence level corresponding to the emotional response label includes: processing the captured natural language speech to obtain the emotional response label, and extracting acoustic features to calculate the semantic and acoustic consistency as the initial confidence level. And / or, the calculation of the process quality score of the subject corresponding to each mode during the evaluation process includes: processing the captured speech audio data to obtain speech recognition content and extracting acoustic features; The clarity of expression is calculated based on the speech rate and pause ratio of the speech recognition content; The similarity between the semantic vector corresponding to the speech recognition content and the emotion tendency vector obtained from the acoustic feature analysis is calculated as semantic consistency. The process quality score is calculated based on the clarity of expression and the semantic consistency.
5. The method according to claim 1, characterized in that, When the third mode is executed: The dynamic selection of preset modes to obtain the subject's emotional response labels and the initial confidence level corresponding to the emotional response labels includes: obtaining a pre-trained emotion recognition model, wherein the emotion recognition model is obtained by joint training using data induced by homologous emotional stimulus materials in a planar environment and a VR environment through a domain adaptation method; The emotional stimulus material of the current induced scenario is input into the emotion recognition model to obtain the basic emotion prediction value and the initial confidence level. Based on the immersion parameters of the current induced scene, the basic emotion prediction value is enhanced and corrected to generate emotion response tags; And / or, the calculation of the process quality score of the subject in each mode during the evaluation process includes: obtaining the pre-collected homologous behavioral baseline features of the subject in a planar environment, calculating the deviation of the real-time behavioral features in the VR environment from the baseline features, and calculating the cross-domain attention behavior reliability score as the process quality score based on the deviation.
6. The method according to claim 1, characterized in that, The physiological verification score was obtained in the following way: The physiological arousal verification index is obtained by weighting and summing the skin conductance amplitude, heart rate variability and the reciprocal of electromyography signal in the physiological signal data with preset weighting coefficients. Calculate the matching error between the subjective arousal level represented by the emotion response label and the physiological arousal verification index; The overall confidence score is calibrated based on the matching error, wherein the lower the matching error, the higher the overall confidence score.
7. The method according to claim 1, characterized in that, In the first mode, the rendering of the annotation interface that does not affect the continuous presentation and interaction of the current VR scene content includes: Based on the real-time monitoring of the task load of the induced scenario and the emotional arousal level of the subjects, the transparency and rendering scale of the annotation interface are dynamically adjusted. In the pre-defined key nodes of the induced scenario, the labeling interface is automatically suspended or weakened.
8. The method according to claim 7, characterized in that, In the first mode, rendering the annotation interface further includes the following positioning modes: A fixation-associated positioning mode is used to dynamically attach the labeled interface to a subcentral region outside the subject's focal point of vision based on the eye movement data. The user-defined visual field anchor point positioning mode is used to make the annotation interface move with the subject's head rotation based on the comfortable visual field range calibrated during the initialization phase and real-time head posture data.
9. A VR environment emotion response assessment device, characterized in that, include: The emotion response tag acquisition module is used to dynamically select a preset mode in a virtual reality (VR) environment to acquire the subject's emotion response tag and the initial confidence level corresponding to the emotion response tag, based on the current induction scene and the subject's state. The emotion response tag is a numerical value or vector representing the emotion state. The preset mode includes any of the following modes: First mode, by rendering a labeling interface that does not affect the continuous presentation and interaction of the current induction scene content, and collecting the subject's eye movement data and interactive handle data to complete the labeling input, the emotion response tag is obtained. The second mode maps the captured and parsed natural language speech of the subjects into emotion response labels; The third mode uses a pre-trained emotion recognition model in a planar environment and immersion parameters of the current induced scene to predict and generate emotion response labels. The confidence data acquisition module is used to simultaneously collect multimodal data of the subject associated with the data acquisition event that generates the emotional response label. The multimodal data includes eye movement trajectory data and interactive handle operation trajectory data corresponding to the first mode, voice audio data corresponding to the second mode, real-time behavioral features corresponding to the third mode, and physiological signal data corresponding to each mode. The process scoring module is used to calculate the process quality score corresponding to each mode and the physiological verification score corresponding to all modes for the subject during the assessment process based on the multimodal data. The overall confidence module is used to calculate the overall confidence score based on the initial confidence score, process quality score, and physiological verification score corresponding to each mode; The output module is used to associate the emotion response label, the comprehensive confidence score, and the preset pattern for generating the emotion response label with a timestamp, and output them as a structured data packet.
10. An electronic device, characterized in that, include: At least one processor, at least one memory, and computer program instructions stored in the memory, which, when executed by the processor, implement the method as described in any one of claims 1-8.