Multi-modal interactive fusion virtual reality emotion computing system

Through the virtual reality emotion computing system with multimodal interaction fusion, the distributed multimodal perception device and adaptive emotion computing model are used to solve the problem of inaccurate emotion recognition in the existing virtual reality emotion computing system, and realize real-time and accurate user emotion interaction.

CN120595940APending Publication Date: 2025-09-05GUILIN UNIV OF AEROSPACE TECH
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510688440.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-09-05

AI Technical Summary

Technical Problem

Existing virtual reality affective computing systems do not have good emotional recognition capabilities when in use, resulting in the inability to interact in real time and inaccurate interactions.

Method used

The virtual reality emotion computing system adopts multimodal interactive fusion, including distributed multimodal perception devices, user adaptation interaction modules, interactive feedback modules, emotion extraction modules, multimodal fusion modules and adaptive emotion computing models. It monitors user emotions through multiple perception methods and accurately identifies and responds to them.

Benefits of technology

It realizes real-time and accurate recognition and interaction of user emotions, and improves the interactive effect of virtual reality emotion computing system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120595940A_ABST
    Figure CN120595940A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal interactive fusion virtual reality emotion computing system, which comprises a multi-modal interactive fusion virtual reality emotion computing device, a multi-modal interactive fusion virtual reality emotion computing device and a multi-modal interactive fusion virtual reality emotion computing device, the virtual reality emotion computing device based on multi-modal interaction fusion further comprises a distributed multi-modal sensing device, a user adaptation interaction module, an interaction feedback module, an emotion extraction module, a multi-modal fusion module and a self-adaptation emotion computing model. The multi-modal fusion module is used for carrying out full-dimensional monitoring and intelligent decision making on a complex scene, the user adaptive interaction module can realize accurate identification and response to user requirements, and the multi-modal fusion module is used for making up for information limitation of a single modal, so that effective interaction and accurate identification can be realized on the whole.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of virtual reality and human-computer interaction, and specifically to a virtual reality emotion computing system integrating multimodal interaction. Background Art

[0002] Virtual reality and human-computer interaction are mainly based on computer technology, network technology, sociology, behavioral science and other disciplines to study the integrated interaction between people and content generated by computers and the Internet and content generated by computer and network reality, to realize a social space for the integration of virtual reality, and then based on this space, study the future development trend of the Internet and the social structure of our lives.

[0003] Existing virtual reality affective computing systems do not have good emotional recognition capabilities when in use, resulting in the inability to interact in real time within the virtual reality affective computing system and the inability to accurately recognize emotions during interaction.

[0004] Therefore, it is necessary to provide a multimodal interactive fusion virtual reality emotion computing system to solve the above technical problems. Summary of the Invention

[0005] The purpose of the present invention is to provide a multimodal interactive fusion virtual reality emotion computing system to solve the problem that the existing virtual reality emotion computing system proposed in the above background technology does not have good emotion recognition when in use, thereby failing to interact in real time in the virtual reality emotion computing system and failing to accurately recognize during interaction.

[0006] In order to solve the above technical problems, the present invention provides the following technical solutions: a multimodal interactive fusion virtual reality emotion computing system, comprising:

[0007] A multimodal interactive fusion virtual reality emotion computing device, the multimodal interactive fusion virtual reality emotion computing device also includes a distributed multimodal perception device, a user adaptation interaction module, an interactive feedback module, an emotion extraction module, a multimodal fusion module and an adaptive emotion computing model;

[0008] Distributed multimodal sensing devices for full-dimensional monitoring and intelligent decision-making in complex scenarios;

[0009] User adaptation interaction module, used to accurately identify and respond to user needs;

[0010] Interactive feedback module, used to interact with users;

[0011] Emotion extraction module, used to extract user's emotions;

[0012] Multimodal fusion module, used to make up for the information limitations of a single modality;

[0013] Adaptive sentiment computing model, used to improve the accuracy and scenario adaptability of sentiment analysis.

[0014] Furthermore, the distributed multimodal sensing device includes an array infrared camera with a VR headset, a distributed inertial measurement unit and a photoplethysmography sensor. The array infrared camera with the VR headset can capture the user's facial expressions, the distributed inertial measurement unit can capture the user's hand angle and torso speed using VR gloves and somatosensory clothing, and the photoplethysmography sensor can monitor the user's heart rate and skin conductivity using a VR handle.

[0015] Furthermore, the emotion extraction module includes a voice data acquisition module, a facial expression image comparison module and a body movement extraction module. The voice data acquisition module can collect the voice uttered by the user, the facial expression image comparison module can compare the collected facial expression images, and the body movement extraction module can extract body movement data.

[0016] Furthermore, the interactive feedback module includes a user emotion regulation module, a user emotion pulse module and a dialogue content and body language response module. The user emotion regulation module can emit corresponding voice to regulate user emotions, the user emotion pulse module can emit corresponding pulse frequency, and the dialogue content and body language response module can perform corresponding response control according to the dialogue content and body language.

[0017] Furthermore, the adaptive emotion computing model includes a pre-trained BERT model, an online incremental learning unit, an emotion response generator, a user adaptation optimizer and a cross-modal fusion network. The emotion response generator is used to adjust virtual emotions, the user adaptation optimizer is used to adjust weights and emotion classifications, and the cross-modal fusion network is used for the fusion of cross-modal networks.

[0018] Furthermore, the pre-trained BERT model is used to analyze the semantic sentiment tendency in user voice data.

[0019] Furthermore, the online incremental learning unit can automatically update the model classification boundary according to the user feedback label.

[0020] Furthermore, the user adaptation interaction module includes user feedback and experience evaluation, user data analysis and modeling, and interaction strategy generation and optimization. The user feedback and experience evaluation can receive user feedback and experience evaluation data, the user data analysis and modeling can perform corresponding analysis and modeling based on user data, and the interaction strategy generation and optimization can generate and optimize interaction strategies.

[0021] Furthermore, the multimodal fusion module includes a dynamic feature weighting layer based on a gating mechanism, a cross-modal attention sub-module and a fusion adversarial generative network.

[0022] Furthermore, the dynamic feature weighting layer based on the gating mechanism can adjust the contribution weight of each modal feature according to the current environmental scene type, the cross-modal attention sub-module can calculate the spatiotemporal correlation matrix of speech features and facial expression features, and the fused adversarial generative network can resolve conflicts and enhance consistency of inconsistent modal features.

[0023] Compared with the prior art, the beneficial effects achieved by the present invention are:

[0024] 1. Distributed multimodal sensing devices are used for full-dimensional monitoring and intelligent decision-making of complex scenarios. The user adaptation and interaction module can achieve accurate identification and response to user needs. The multimodal fusion module is used to compensate for the information limitations of a single modality, thereby achieving effective interaction and accurate identification.

[0025] 2. Through the structure of the emotion extraction module, user data can be effectively extracted, which facilitates the comparison of later emotion data and can be effectively reflected in the later emotion data interaction;

[0026] 3. The user emotion regulation module can emit corresponding voice to regulate user emotions, and can regulate the corresponding voice emotions according to the user's facial information. The user emotion pulse module can emit corresponding pulse frequencies, so that the corresponding pulse frequencies can be adjusted according to the collected user information, thereby effectively interacting with users. The dialogue content and body language response module can perform corresponding response control according to the dialogue content and body language, and can effectively operate the corresponding interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0028] Figure 1 It is a control schematic diagram of the present invention;

[0029] Figure 2 Schematic diagram of the structure of the distributed multimodal sensing device of the present invention;

[0030] Figure 3 It is a structural diagram of the emotion extraction module of the present invention;

[0031] Figure 4 It is a schematic diagram of the structure of the interactive feedback module of the present invention;

[0032] Figure 5 It is a schematic diagram of the structure of the adaptive emotion computing model of the present invention;

[0033] Figure 6 It is a schematic diagram of the structure of the user adaptation interaction module of the present invention;

[0034] Figure 7 It is a schematic diagram of the structure of the multimodal fusion module of the present invention.

[0035] Figure 1: 1. Virtual reality emotion computing device with multimodal interaction fusion; 2. Distributed multimodal perception device; 21. Array infrared camera with VR headset; 22. Distributed inertial measurement unit; 23. Photoplethysmography sensor; 3. User adaptation interaction module; 31. User feedback and experience evaluation; 32. User data analysis and modeling; 33. Interaction strategy generation and optimization; 4. Interaction feedback module; 41. User emotion regulation module; 42. User emotion pulse module; 43. Conversation content and body language Response module; 5. Emotion extraction module; 51. Voice data acquisition module; 52. Facial expression image comparison module; 53. Body movement extraction module; 6. Multimodal fusion module; 61. Dynamic feature weighting layer based on gating mechanism; 62. Cross-modal attention sub-module; 63. Fusion adversarial generative network; 7. Adaptive emotion computing model; 71. Pre-trained BERT model; 72. Online incremental learning unit; 73. Emotion response generator; 74. User adaptation optimizer; 75. Cross-modal fusion network. DETAILED DESCRIPTION

[0036] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0037] Example 1

[0038] See also Figure 1-Figure 5 The present invention provides a technical solution: a multimodal interactive fusion virtual reality emotion computing system, comprising a multimodal interactive fusion virtual reality emotion computing device 1; the multimodal interactive fusion virtual reality emotion computing device 1 further comprises a distributed multimodal perception device 2, a user adaptation interaction module 3, an interactive feedback module 4, an emotion extraction module 5, a multimodal fusion module 6 and an adaptive emotion computing model 7;

[0039] Distributed multimodal sensing device 2, used for full-dimensional monitoring and intelligent decision-making of complex scenes, and real-time acquisition of users' physiological signals, voice data, facial expression images, and body movement data in a virtual reality environment;

[0040] User Adaptation Interaction Module 3 dynamically adjusts the lighting, sound effects, virtual character behavior, and plot branches in the virtual reality scene based on the output of the affective computing model, and provides the user with the vibration intensity and pattern of the tactile feedback device;

[0041] Interactive feedback module 4 is used to interact with the user and dynamically adjust the lighting, sound effects, virtual character behavior and plot branches in the virtual reality scene based on the output results of the adaptive emotion computing model 7, and provide the user with the vibration intensity and pattern of the tactile feedback device;

[0042] Emotion extraction module 5, used to extract heart rate variability features from physiological signals, voiceprint spectrum features from speech data, micro-expression action unit features from facial expression images, and motion trajectory features from body movement data;

[0043] Multimodal fusion module 6, used to compensate for the information limitations of a single modality, adopts a deep neural network model based on the attention mechanism to dynamically assign weights and integrate cross-modal associations of the physiological, speech, facial expression, and action features to generate a comprehensive emotional state vector of the user;

[0044] The adaptive emotion computing model 7 is used to improve the accuracy and scenario adaptability of emotion analysis. It is built based on the transfer learning framework, inputs the comprehensive emotion state vector, combines the user's historical behavior data and environmental context information, and outputs the user's emotion category and intensity quantification value.

[0045] The distributed multimodal sensing device 2 includes an array infrared camera 21 with a VR head-mounted display, a distributed inertial measurement unit 22, and a photoplethysmography sensor 23;

[0046] The array infrared camera 21 with the VR headset can capture the user's facial expressions, so that the user's facial data information can be formed. The distributed inertial measurement unit 22 can capture the user's hand angle and torso speed using VR gloves and somatosensory clothing, and can effectively collect data information on the user's hand angle and torso speed. The photoplethysmography sensor 23 can monitor the user's heart rate and skin conductivity using the VR handle, and can effectively collect data information on the user's heart rate and skin conductivity.

[0047] The interactive feedback module 4 includes a user emotion regulation module 41, a user emotion pulse module 42, and a conversation content and body language response module 43;

[0048] The user emotion regulation module 41 can emit corresponding voice to regulate the user's emotions, and can regulate the emotions of the corresponding voice according to the user's facial information. The user emotion pulse module 42 can emit corresponding pulse frequency, so that the corresponding pulse frequency can be adjusted according to the collected user information, so as to effectively interact with users. The dialogue content and body language response module 43 can perform corresponding response control according to the dialogue content and body language, and can effectively operate the corresponding interaction.

[0049] Example 2

[0050] like Figure 1 、 Figure 3 and Figure 5 , the emotion extraction module 5 includes a voice data acquisition module 51, a facial expression image comparison module 52 and a body movement extraction module 53;

[0051] The voice data acquisition module 51 can collect the voice uttered by the user, thereby collecting and storing voice information data; the facial expression image comparison module 52 can compare the collected facial expression images, thereby collecting and storing facial expression image data; the limb movement extraction module 53 can extract limb movement data, thereby collecting and storing limb movement data.

[0052] The adaptive emotion computing model 7 includes a pre-trained BERT model 71, an online incremental learning unit 72, an emotion response generator 73, a user adaptation optimizer 74, and a cross-modal fusion network 75;

[0053] The emotional response generator 73 is used to regulate virtual emotions and dynamically adjust the emotional expression strategy of the virtual character based on the quantized value of emotional intensity, including emotional rhythm modeling of speech synthesis and continuous interpolation generation of three-dimensional facial expressions;

[0054] The user adaptation optimizer 74 is used to adjust the weights and sentiment classification, using a double-loop optimization architecture to synchronously update feature encoding weights and sentiment classification boundaries, and using a reinforcement learning reward mechanism to balance recognition accuracy and computing resource consumption;

[0055] The cross-modal fusion network 75 is used to fuse cross-modal networks, deploy a multi-scale adaptive attention mechanism, dynamically establish a spatiotemporal correlation matrix of speech, expression, and physiological parameters, and apply the GAN-Fusion algorithm to achieve reconstruction and optimization of conflicting modal features;

[0056] The pre-trained BERT model 71 can be used to analyze the semantic sentiment in user voice data, which facilitates the comparison of voice sentiment data later;

[0057] The online incremental learning unit 72 can automatically update the model classification boundary according to the user feedback label, and can effectively perform feedback and data comparison of emotional data.

[0058] Example 3

[0059] like Figure 1 、 Figure 4 and Figure 5 , the user adaptation interaction module 3 includes user feedback and experience evaluation 31, user data analysis and modeling 32, and interaction strategy generation and optimization 33;

[0060] User feedback and experience evaluation 31 can receive user feedback and experience evaluation data, so that subsequent optimization processing can be performed based on the feedback and experience evaluation data. User data analysis and modeling 32 can perform corresponding analysis and modeling based on user data, so that corresponding modeling can be formed. Interaction strategy generation and optimization 33 can generate and optimize interaction strategies, so that the generation, change and optimization of interaction strategies can be performed based on the collected data.

[0061] The multimodal fusion module 6 includes a dynamic feature weighting layer 61 based on a gating mechanism, a cross-modal attention submodule 62, and a fusion adversarial generative network 63;

[0062] The dynamic feature weighting layer 61 based on the gating mechanism can adjust the contribution weight of each modal feature according to the current environmental scene type, so as to make corresponding priority adjustments based on the contribution weight. The cross-modal attention submodule 62 can calculate the spatiotemporal correlation matrix of speech features and facial expression features, so that they can be effectively linked and associated. The fusion adversarial generative network 63 can resolve conflicts and enhance consistency of inconsistent modal features, so that consistency can be effectively enhanced.

[0063] Working principle: The multimodal interactive fusion virtual reality emotion computing device 1 is used for full-dimensional monitoring and intelligent decision-making of complex scenes through a distributed multimodal perception device 2, a user adaptation interaction module 3, which is used to achieve accurate identification and response to user needs, an interactive feedback module 4, which is used to interact with the user, an emotion extraction module 5, which is used to extract the user's emotions, a multimodal fusion module 6, which is used to make up for the information limitations of a single modality, and an adaptive emotion computing model 7, which is used to improve the accuracy of emotion analysis and scene adaptability.

[0064] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0065] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A multimodal interactive fusion virtual reality affective computing system, characterized by: include: A multimodal interactive fusion virtual reality emotion computing device (1), wherein the multimodal interactive fusion virtual reality emotion computing device (1) further comprises a distributed multimodal perception device (2), a user adaptation interaction module (3), an interaction feedback module (4), an emotion extraction module (5), a multimodal fusion module (6), and an adaptive emotion computing model (7); Distributed multimodal sensing device (2), used for full-dimensional monitoring and intelligent decision-making of complex scenes; User adaptation interaction module (3), used to accurately identify and respond to user needs; Interactive feedback module (4), used for interacting with users; Emotion extraction module (5), used to extract the user's emotional emotions; Multimodal fusion module (6), used to make up for the information limitations of a single modality; Adaptive sentiment computing model (7) is used to improve the accuracy and scenario adaptability of sentiment analysis.

2. The multimodal interactive fusion virtual reality affective computing system according to claim 1, characterized in that: The distributed multimodal sensing device (2) comprises an array infrared camera (21) with a VR head display, a distributed inertial measurement unit (22) and a photoplethysmography sensor (23). The array infrared camera (21) with the VR head display can capture the user's facial expression, the distributed inertial measurement unit (22) can capture the user's hand angle and trunk speed using VR gloves and body sensing clothing, and the photoplethysmography sensor (23) can monitor the user's heart rate and skin conductivity using the VR handle.

3. The multimodal interactive fusion virtual reality affective computing system according to claim 1, characterized in that: The emotion extraction module (5) includes a voice data acquisition module (51), a facial expression image comparison module (52) and a body movement extraction module (53). The voice data acquisition module (51) can acquire the voice uttered by the user, the facial expression image comparison module (52) can compare the acquired facial expression images, and the body movement extraction module (53) can extract body movement data.

4. The multimodal interactive fusion virtual reality affective computing system according to claim 1, characterized in that: The interactive feedback module (4) includes a user emotion regulation module (41), a user emotion pulse module (42) and a dialogue content and body language response module (43). The user emotion regulation module (41) can emit corresponding voice to regulate the user's emotion, the user emotion pulse module (42) can emit corresponding pulse frequency, and the dialogue content and body language response module (43) can perform corresponding response control according to the dialogue content and body language.

5. The multimodal interactive fusion virtual reality affective computing system according to claim 1, characterized in that: The adaptive emotion computing model (7) includes a pre-trained BERT model (71), an online incremental learning unit (72), an emotion response generator (73), a user adaptation optimizer (74) and a cross-modal fusion network (75), wherein the emotion response generator (73) is used to adjust virtual emotions, the user adaptation optimizer (74) is used to adjust weights and emotion classifications, and the cross-modal fusion network (75) is used to fuse cross-modal networks.

6. The multimodal interactive fusion virtual reality affective computing system according to claim 5, characterized in that: The pre-trained BERT model (71) is used to analyze the semantic sentiment tendency in user voice data.

7. The multimodal interactive fusion virtual reality affective computing system according to claim 5, characterized in that: The online incremental learning unit (72) can automatically update the model classification boundary according to the user feedback label.

8. The multimodal interactive fusion virtual reality affective computing system according to claim 1, characterized in that: The user adaptation interaction module (3) includes user feedback and experience evaluation (31), user data analysis and modeling (32) and interaction strategy generation and optimization (33). The user feedback and experience evaluation (31) can receive user feedback and experience evaluation data, the user data analysis and modeling (32) can perform corresponding analysis and modeling based on user data, and the interaction strategy generation and optimization (33) can generate and optimize interaction strategies.

9. The multimodal interactive fusion virtual reality affective computing system according to claim 1, characterized in that: The multimodal fusion module (6) includes a dynamic feature weighting layer (61) based on a gating mechanism, a cross-modal attention submodule (62) and a fusion adversarial generation network (63).

10. The multimodal interactive fusion virtual reality affective computing system according to claim 9, characterized in that: The dynamic feature weighting layer (61) based on the gating mechanism can adjust the contribution weight of each modal feature according to the current environmental scene type, the cross-modal attention submodule (62) can calculate the spatiotemporal correlation matrix of speech features and facial expression features, and the fusion adversarial generative network (63) can resolve conflicts and enhance consistency of inconsistent modal features.

Citation Information

Cited By

  • Intelligent pickup and speech recognition system based on multi-modal fusion

    CN120954408A

  • Intelligent sound pickup and speech recognition system based on multimodal fusion

    CN120954408B

  • Emotion recognition interaction method and device based on AI intelligent analysis and electronic equipment

    CN121919804A

  • Multi-user interaction response strategy automatic switching method fusing behaviors and emotions

    CN121934719A