VR game implementation method based on emotion analysis and related equipment

By employing multimodal data acquisition and fusion techniques, combined with LSTM neural networks and weighted voting mechanisms, the limitations of single-modal emotion analysis in existing VR games have been addressed, resulting in more accurate emotion analysis and a more personalized gaming experience.

CN121668660APending Publication Date: 2026-03-17XIAN WANXIANG ELECTRONICS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511452288.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-11
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing VR game emotion analysis solutions mostly rely on a single modality, which has limitations and low adaptability, resulting in inaccurate emotion analysis results and a poor gaming experience.

Method used

Employing multimodal data acquisition and fusion technology, visual, speech, and physiological data are collected through VR glasses and smart bracelets. Sentiment analysis is performed using LSTM neural networks and weighted voting mechanisms, and response strategies are generated in conjunction with game scenarios to provide more accurate game feedback.

Benefits of technology

It improves the accuracy of emotion analysis and the gaming experience, enabling it to better adapt to different environments and scenarios, and providing personalized game feedback and dynamic adjustment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121668660A_ABST
    Figure CN121668660A_ABST
Patent Text Reader

Abstract

The invention provides a VR game implementation method based on emotion analysis and related equipment. The method comprises the following steps: acquiring multi-modal data of a game player; wherein the multi-modal data comprises visual data, voice data and physiological data; feature values corresponding to the multi-modal data are extracted; and inputting the characteristic values corresponding to the multi-modal data into a preset training model for data analysis, and performing game related processing according to an analysis result. According to the VR game implementation method based on emotion analysis disclosed by the invention, more accurate feedback can be provided for game players, so that better VR game experience can be provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of emotion analysis technology, specifically to a VR game implementation method and related equipment based on emotion analysis. Background Technology

[0002] In the VR gaming field, existing sentiment analysis solutions mostly rely on a single modality, which has certain limitations and low adaptability. Furthermore, the sentiment analysis methods are relatively simplistic, leading to inaccurate results and ultimately a poor gaming experience when combined with sentiment analysis. The following will illustrate existing game sentiment analysis methods with specific examples:

[0003] Case 1: Relying solely on facial recognition for emotion analysis has limitations: insufficient lighting or facial occlusion can easily cause significant errors, and it cannot recognize emotions with closed eyes.

[0004] Case 2: Relying solely on heart rate monitoring has the limitation that it cannot distinguish between heart rate increases caused by exercise (such as jumping) and those caused by genuine emotions.

[0005] Case 3: Low adaptability scenarios. Its limitations are: the accuracy of voice analysis drops significantly in noisy environments, usually >35%, and single behavioral data is easily misjudged. For example, turning the head in a game may be misjudged as irritability.

[0006] Case 4: Limited by the single emotion label: most solutions only support basic emotion classification, such as joy / anger, and lack intensity quantification. Summary of the Invention

[0007] The purpose of this disclosure is to overcome the shortcomings of the prior art and provide a VR game implementation method and related equipment based on emotion analysis. This VR game implementation method based on emotion analysis can provide game players with more accurate feedback, thereby providing a better VR game experience.

[0008] According to a first aspect of the present disclosure, a VR game implementation method based on emotion analysis is provided, applied to a server, comprising the following steps:

[0009] Acquire multimodal data of game players; wherein, the multimodal data includes visual data, voice data, and physiological data;

[0010] Extract the feature values ​​corresponding to each of the multimodal data;

[0011] The feature values ​​corresponding to each of the multimodal data are input into a preset training model for data analysis, and game-related processing is performed based on the analysis results.

[0012] In one embodiment, the method further includes:

[0013] The feature values ​​corresponding to each of the multimodal data are merged into a single vector, and the single vector is input into an LSTM neural network for training to obtain the preset training model.

[0014] In one embodiment, the step of inputting the feature values ​​corresponding to each of the multimodal data into a preset training model for data analysis, and performing game-related processing based on the analysis results includes:

[0015] The feature values ​​corresponding to each of the multimodal data are input into the preset training model for analysis, and the analysis results are weighted and voted on. The final emotion of the game player is determined based on the voting results.

[0016] In one embodiment, the method further includes:

[0017] Add an emotion tag to the player based on their final emotion.

[0018] In one embodiment, the method further includes:

[0019] Detect the player's current game scene;

[0020] Generate a game response strategy based on the emotion tags and the player's current game scenario;

[0021] The game title is awarded to the player according to the game response strategy.

[0022] In one embodiment, before extracting the feature values ​​corresponding to each of the multimodal data, the method further includes:

[0023] The visual and voice data are anonymized, and the physiological data is encrypted.

[0024] In one embodiment, before extracting the feature values ​​corresponding to each of the multimodal data, the method further includes:

[0025] The visual data, voice data, and physiological data are time-stamped and synchronized.

[0026] The visual data, speech data, and physiological data are subjected to noise reduction processing.

[0027] In one embodiment, extracting the feature values ​​corresponding to each of the multimodal data includes:

[0028] For the visual data, the pupil diameter change rate and the standard deviation of the blink interval are calculated, and the pupil diameter change rate and the standard deviation of the blink interval are marked as feature values ​​of the visual data;

[0029] For the speech data, Mel frequency cepstral coefficients and fundamental frequency profiles are extracted, and the Mel frequency cepstral coefficients and fundamental frequency profiles are marked as feature values ​​of the speech data;

[0030] For the physiological data, the autonomic nervous balance index of heart rate variability and the peak value of skin conductance response are obtained, and the autonomic nervous balance index of heart rate variability and the peak value of skin conductance response are marked as the feature values ​​of the physiological data.

[0031] According to a second aspect of the present disclosure, a server is provided, the server comprising: an acquisition module, an extraction module, and a processing module; wherein...

[0032] The acquisition module is used to acquire multimodal data of game players; wherein, the multimodal data includes visual data, voice data, and physiological data;

[0033] The extraction module is used to extract the feature values ​​corresponding to each of the multimodal data;

[0034] The processing module is used to input the feature values ​​corresponding to each of the multimodal data into a preset training model for data analysis, and to perform game-related processing based on the analysis results.

[0035] In one embodiment, the server further includes: a merging module; wherein,

[0036] The merging module is used to merge the feature values ​​corresponding to each of the multimodal data into a single vector, and input the single vector into an LSTM neural network for training to obtain the preset training model.

[0037] In one embodiment, the processing module is specifically used for,

[0038] The feature values ​​corresponding to each of the multimodal data are input into the preset training model for analysis, and the analysis results are weighted and voted on. The final emotion of the game player is determined based on the voting results.

[0039] In one embodiment, the server further includes: an adding module, a detection module, a generating module, and an issuing module; wherein,

[0040] The adding module is used to add emotion tags to the game player based on the game player's final emotion;

[0041] The detection module is used to detect the game scene of the game player;

[0042] The generation module is used to generate a game response strategy based on the emotion tag and the game scene;

[0043] The issuing module is used to issue game titles to the game players according to the response strategy.

[0044] According to a third aspect of the present disclosure, a computer device is provided, comprising: a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of any of the above methods.

[0045] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the steps of the method as described in any of the above.

[0046] This disclosure provides a VR game implementation method based on emotion analysis. The method acquires multimodal data (visual, speech, and physiological data) of the game player through VR glasses and a smart bracelet, and transmits this data to a data processing layer. This data processing layer processes the data from different modalities and outputs structured feature vectors, which are then passed to a data analysis layer. The data analysis layer performs early fusion and late fusion. Early fusion is used to capture cross-modal correlations, while late fusion is used to prevent single-modal failures from affecting the overall system. Finally, the application layer processes the data and adjusts the displayed content to provide a better gaming experience for the game player. Attached Figure Description

[0047] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0048] Figure 1 This is a schematic diagram of the architecture of a VR game emotion analysis system provided in an embodiment of this disclosure.

[0049] Figure 2 A flowchart illustrating a VR game implementation method based on emotion analysis provided in this disclosure embodiment.

[0050] Figure 3 This is an architecture diagram of a server provided in an embodiment of the present disclosure.

[0051] Figure 4 This is an architecture diagram of a server provided in an embodiment of the present disclosure.

[0052] Figure 5 This is an architectural diagram of a computer device provided in an embodiment of the present disclosure. Detailed Implementation

[0053] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.

[0054] Figure 1 This is a schematic diagram of the architecture of a VR game emotion analysis system provided in an embodiment of this disclosure. Figure 1 As shown, the system comprises a data acquisition layer, a data processing layer, a data analysis layer, an application layer, and a privacy and security layer. The composition and functions of each layer are as follows:

[0055] The data acquisition layer includes devices for data acquisition, such as VR glasses and smart bracelets.

[0056] in,

[0057] VR glasses include:

[0058] Eye-tracking module: Collects data including pupil diameter, fixation point coordinates, fixation point dwell time, and blink frequency. This data can detect focus (e.g., fixation point dwell time > 15 seconds) and fear (pupil dilation > 20%).

[0059] Facial camera: Collects data, including key facial points such as the height of the corners of the mouth and the distance between the eyebrows. This data can identify the player's micro-expressions, such as an upturned mouth indicating pleasure and downturned eyebrows indicating confusion.

[0060] Microphone: Collects data, including speech spectrum, fundamental frequency, and speech rate, which can detect excited screams (fundamental frequency greater than 250Hz) and detect profanity keywords (FFT feature matching).

[0061] IMU sensor: Collects data, including head rotation speed and acceleration, which can determine the risk of jolts (such as rapid head shaking - frustration) and motion sickness (angular velocity >180° / s for 5 seconds).

[0062] Smart bracelets include:

[0063] Optical heart rate sensor: Collects data, including heart rate and HRV, which can be used for stress detection (e.g., HRV < 50 ms - high stress) and motion interference filtering (in conjunction with an accelerometer).

[0064] GSR Electrode: Data Acquisition: Includes GSR data, which is used for fear peak detection.

[0065] Triaxial accelerometer: Collects data, including the amplitude of hand movements and vibration frequency. This data is used to determine the intensity of the operation (shooting frequency > 3 times / second - excitement) and abnormal tremors (3-5Hz micro-tremors when anxious).

[0066] The data processing layer is used to process the data acquired by the data acquisition layer. Its processing steps include:

[0067] Signal alignment processing: Data acquisition timestamp synchronization, and alignment of VR glasses and wristband data via Bluetooth 5.2 CTS (Clock Transmission Service).

[0068] Noise reduction processing: Motion artifact elimination: Adaptive filters are used to remove jitter noise from the heart rate data of the wristband; Speech denoising processing: Human voice and environmental noise are separated by AI noise reduction algorithms (such as RNNoise noise suppression algorithm).

[0069] Feature extraction:

[0070] Eye characteristics: Calculate the rate of change of pupil diameter (ΔD / Δt) and the standard deviation of blink interval.

[0071] Speech features: Extract MFCC (Mel frequency cepstral coefficients) and fundamental frequency profile.

[0072] Physiological signals: RMSSD (autonomic balance index) of HRV (heart rate variability) and peak GSR count.

[0073] Output data format: includes structured feature vectors, such as pupil diameter: 4.2mm, blink rate: 12 times / min, heart rate: 85bpm, GSR peak: 3 times / second.

[0074] Data Analysis Layer: Used to analyze the data processed by the Data Processing Layer. The analysis process is as follows:

[0075] Multimodal fusion strategies include early fusion and late fusion; among which,

[0076] Early fusion (data level): Features from various modalities are merged into a single vector and input into an LSTM neural network. The advantage of this approach is that it can capture cross-modal correlations; for example, a sharp voice combined with an increased heart rate can be interpreted as a gamer's emotion being anger.

[0077] Late-stage fusion (decision level): Each modality is trained independently (e.g., CNN processes facial expression information, SVM analyzes heart rate data), and the analysis results from different modalities are weighted and voted on to make the final emotion judgment. The advantage of late-stage decision-making is that it can maintain a certain degree of robustness, that is, a failure in a single modality will not have an excessive impact on the overall model.

[0078] Emotion classification model:

[0079] Emotional tags: happiness, sadness, anger, fear, surprise, excitement, frustration, boredom, etc.

[0080] Using transfer learning: pre-trained on the AffectNet dataset and then fine-tuned in conjunction with VR scenes.

[0081] Application layer: Used to dynamically process game performance based on the obtained emotion labels and different multimodal data, such as the dynamic game adjustment mechanism in Table 1 below.

[0082] Table 1 Dynamic Game Adjustment Mechanism:

[0083]

[0084] Player profile system:

[0085] Emotional endurance: Calculates the number of times fear peaks per unit of time, unlocking the "Courage Achievement".

[0086] Operation style tags: Divide into "Power" and "Precision" types based on hand acceleration data and change the damage calculation coefficient.

[0087] Dynamic achievement system:

[0088] Complete the precision task while your heart rate is >100 bpm to earn the achievement "Master of Calm".

[0089] Detecting laughter when a game character dies earns the achievement "Humorous Response" and triggers a resurrection easter egg, among other things.

[0090] Privacy and Security Layer: Used for secure processing of game player data. The processing procedure is as follows:

[0091] Data flow control: Sensitive data (such as raw facial videos, voice, etc.) is processed locally only, without being uploaded to the cloud or used for edge computing. Data transmitted to external parties is anonymized feature vectors (such as "emotional intensity value 72" instead of using the specific heart rate).

[0092] Hardware-level protection: The wristband uses ARM TrustZone to encrypt biometric data, and VR devices support the physical camera cover being closed.

[0093] Figure 2 This is a schematic diagram illustrating the workflow of a VR game emotion analysis system provided in an embodiment of this disclosure. Figure 2 As shown, it includes the following steps:

[0094] Step 201: Obtain multimodal data of game players; wherein, the multimodal data includes visual data, voice data, and physiological data;

[0095] In this step, visual data is collected via an eye-tracking module, a facial camera, and an IMU sensor. The eye-tracking module collects data such as pupil diameter, fixation point coordinates, fixation duration, and blink frequency. This data can detect focus (e.g., fixation duration > 15 seconds) and fear (pupil dilation > 20%). The facial camera collects key facial points, such as mouth corner height and eyebrow spacing. This data can identify micro-expressions, such as an upturned mouth indicating pleasure and downturned eyebrows indicating confusion. The IMU sensor collects head rotation speed and acceleration. This data can determine the severity of actions (e.g., rapid head shaking – frustration) and the risk of motion sickness – angular velocity > 180° / s for 5 seconds. Voice data is collected via a microphone, including speech spectrum, fundamental frequency, and speech rate. This data can detect excited screams (fundamental frequency > 250Hz) and profanity keywords (FFT feature matching). Physiological data is collected via a smart bracelet equipped with an optical heart rate sensor, GSR electrodes, and a triaxial accelerometer. The optical heart rate sensor collects heart rate and HRV, which can be used for stress detection (e.g., HRV < 50ms – high stress) and motion interference filtering (in conjunction with the accelerometer). The GSR electrodes collect GSR data, which is used for fear peak detection. The triaxial accelerometer collects hand movement amplitude and vibration frequency, which is used to determine the intensity of the action (shooting frequency > 3 times / second – excitement) and abnormal tremors (3-5Hz micro-tremors during anxiety).

[0096] Step 202: Extract the feature values ​​corresponding to each of the multimodal data;

[0097] In one embodiment, extracting the feature values ​​corresponding to each of the multimodal data includes:

[0098] For the visual data, the pupil diameter change rate and the standard deviation of the blink interval are calculated, and the pupil diameter change rate and the standard deviation of the blink interval are marked as feature values ​​of the visual data;

[0099] In this step, eye features are extracted by calculating the pupil diameter change rate (ΔD / Δt) and the standard deviation of blink interval.

[0100] For the speech data, Mel frequency cepstral coefficients and fundamental frequency profiles are extracted, and the Mel frequency cepstral coefficients and fundamental frequency profiles are marked as feature values ​​of the speech data;

[0101] In this step, speech features are extracted by extracting MFCC (Mel frequency cepstral coefficients) and fundamental frequency contour.

[0102] For the physiological data, the autonomic nervous balance index of heart rate variability and the peak value of skin conductance response are obtained, and the autonomic nervous balance index of heart rate variability and the peak value of skin conductance response are marked as the feature values ​​of the physiological data.

[0103] In this step, physiological signals are extracted by obtaining the RMSSD (autonomic balance index) of HRV (heart rate variability) and the peak GSR count.

[0104] Step 203: Input the feature values ​​corresponding to each of the multimodal data into a preset training model for data analysis, and perform game-related processing based on the analysis results.

[0105] In one embodiment, before inputting the feature values ​​corresponding to each of the multimodal data into a preset training model for data analysis, the method further includes:

[0106] The feature values ​​corresponding to each of the multimodal data are merged into a single vector, and the single vector is input into an LSTM neural network for training to obtain the preset training model.

[0107] In this embodiment, early fusion (data level) is performed, specifically by merging the features of each modality into a single vector and inputting it into an LSTM neural network. The advantage of this approach is that it can capture cross-modal associations; for example, a sharp voice combined with an increased heart rate can be interpreted as a gamer's emotion being anger.

[0108] In one embodiment, the step of inputting the feature values ​​corresponding to each of the multimodal data into a preset training model for data analysis, and performing game-related processing based on the analysis results includes:

[0109] The feature values ​​corresponding to each of the multimodal data are input into the corresponding preset training model for analysis, and the analysis results are weighted and voted on. The final emotion of the game player is determined based on the voting results.

[0110] In this embodiment, late-stage fusion (decision level) is performed. Specifically, this involves independently training models for each modality (e.g., CNN for processing facial expression information, SVM for analyzing heart rate data), and then weighting and voting the analysis results from different modalities to make the final emotion judgment. The advantage of late-stage decision-making is that it can maintain a certain degree of robustness, meaning that a failure in a single modality will not have an excessive impact on the overall model.

[0111] Optionally, the method further includes:

[0112] Add an emotion tag to the player based on their final emotion.

[0113] In this embodiment, emotion labels include happiness, sadness, anger, fear, surprise, excitement, frustration, boredom, etc.

[0114] Optionally, the method further includes:

[0115] Detect the player's current game scene;

[0116] Generate a game response strategy based on the emotion tags and the player's current game scenario;

[0117] The game title is awarded to the player according to the game response strategy.

[0118] In this embodiment, the number of peak fear times per unit time is calculated to unlock the "Courage Achievement" title.

[0119] Operation style tags: Divide into "Power" and "Precision" types based on hand acceleration data and change the damage calculation coefficient.

[0120] Complete the precision task while your heart rate is >100 bpm to earn the achievement "Master of Calm".

[0121] Detecting laughter when a game character dies earns the achievement "Humorous Response" and triggers a resurrection easter egg, among other things.

[0122] Optionally, before extracting the feature values ​​corresponding to each of the multimodal data, the method further includes:

[0123] The visual and voice data are anonymized, and the physiological data is encrypted.

[0124] In this embodiment, sensitive data (such as original facial videos, voice recordings, etc.) are processed locally only, without being uploaded to the cloud or used for edge computing. Data transmitted to external parties is an anonymized feature vector (such as "emotional intensity value 72" instead of using a specific heart rate).

[0125] Hardware-level protection: The wristband uses ARM TrustZone to encrypt biometric data, and VR devices support the physical camera cover being closed.

[0126] Optionally, before extracting the feature values ​​corresponding to each of the multimodal data, the method further includes:

[0127] The visual data, voice data, and physiological data are time-stamped and synchronized.

[0128] In this step, the data timestamps are synchronized, and the VR glasses and wristband data are aligned via Bluetooth 5.2's CTS (Clock Transmission Service).

[0129] The visual data, speech data, and physiological data are subjected to noise reduction processing.

[0130] In this step, an adaptive filter is used to remove jitter noise from the heart rate data of the wristband in order to eliminate motion artifacts; and an AI noise reduction algorithm (such as RNNoise noise suppression algorithm) is used to separate human voice and environmental noise for speech denoising processing.

[0131] This disclosure provides a VR game implementation method based on emotion analysis. The method acquires multimodal data (visual, speech, and physiological data) of the game player through VR glasses and a smart bracelet, and transmits this data to a data processing layer. This data processing layer processes the data from different modalities and outputs structured feature vectors, which are then passed to a data analysis layer. The data analysis layer performs early fusion and late fusion. Early fusion is used to capture cross-modal correlations, while late fusion is used to prevent single-modal failures from affecting the overall system. Finally, the application layer processes the data and adjusts the displayed content to provide a better gaming experience for the game player.

[0132] Figure 3 This is an architecture diagram of a server provided as an embodiment of this disclosure. (For example...) Figure 3 As shown, the server includes: an acquisition module 301, an extraction module 302, and a processing module 303; wherein, the acquisition module 301 is used to acquire multimodal data of game players; wherein, the multimodal data includes visual data, voice data, and physiological data; the extraction module 302 is used to extract the feature values ​​corresponding to each of the multimodal data; the processing module 303 is used to input the feature values ​​corresponding to each of the multimodal data into a preset training model for data analysis, and perform game-related processing based on the analysis results.

[0133] Figure 4 This is an architecture diagram of a server provided as an embodiment of this disclosure. (For example...) Figure 4 As shown, the server includes: an acquisition module 401, an extraction module 402, a processing module 403, an addition module 404, a detection module 405, a generation module 406, and an awarding module 407; wherein, the addition module 404 is used to add an emotion tag to the game player based on the game player's final emotion; the detection module 405 is used to detect the game player's game scene; the generation module 406 is used to generate a game response strategy based on the emotion tag and the game scene; and the awarding module 407 is used to award a game title to the game player based on the response strategy.

[0134] In one embodiment, a computer device is provided, the internal structure of which can be as follows: Figure 5As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores data. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements the above-described method for implementing a VR game based on emotion analysis. It includes: memory and a processor; the memory stores the computer program; and the processor executes the computer program to implement any step in the above-described method for implementing a VR game based on emotion analysis.

[0135] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, can perform any of the steps in the above-described method for implementing a VR game based on emotion analysis.

[0136] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0137] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0138] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0139] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0140] Although preferred embodiments of this application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of this application.

[0141] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A VR game implementation method based on emotion analysis, characterized in that, Applied to a server, the method comprises: Obtaining multi-modal data of a game player; wherein the multi-modal data comprises visual data, voice data and physiological data; Extracting respective feature values corresponding to the multi-modal data; Inputting the respective feature values corresponding to the multi-modal data into a preset training model for data analysis, and performing game-related processing according to the analysis result.

2. The method of claim 1, wherein, Before the step of inputting the respective feature values corresponding to the multi-modal data into the preset training model for data analysis, the method further comprises: Merging the respective feature values corresponding to the multi-modal data into a single vector, and inputting the single vector into an LSTM neural network for training to obtain the preset training model.

3. The method of claim 2, wherein, The step of inputting the respective feature values corresponding to the multi-modal data into the preset training model for data analysis, and performing game-related processing according to the analysis result comprises: Inputting the respective feature values corresponding to the multi-modal data into the preset training model for analysis respectively, and performing weighted voting on the respective analysis results, and determining the final emotion of the game player according to the voting result.

4. The method of claim 3, wherein, The method further comprises: Adding an emotional label to the game player according to the final emotion of the game player.

5. The method of claim 4, wherein, The method further comprises: Detecting the current game scene of the game player; Generating a game response strategy according to the emotional label and the current game scene of the game player; Awarding a game title to the game player according to the game response strategy.

6. The method of claim 1, wherein, Before the step of extracting the respective feature values corresponding to the multi-modal data, the method further comprises: Performing desensitization processing on the visual data and voice data, and performing encryption processing on the physiological data.

7. The method of claim 1, wherein, Before the step of extracting the respective feature values corresponding to the multi-modal data, the method further comprises: Performing timestamp synchronization processing on the visual data, voice data and physiological data; Performing denoising processing on the visual data, voice data and physiological data.

8. The method of claim 1, wherein, The step of extracting the respective feature values corresponding to the multi-modal data comprises: For the visual data, calculating the pupil diameter change rate and the blink interval standard deviation, and marking the pupil diameter change rate and the blink interval standard deviation as the feature values of the visual data; For the voice data, extracting the mel-frequency cepstrum coefficient and the fundamental frequency contour, and marking the mel-frequency cepstrum coefficient and the fundamental frequency contour as the feature values of the voice data; For the physiological data, obtaining the autonomic nervous balance index of heart rate variability and the skin galvanic response peak, and marking the autonomic nervous balance index of heart rate variability and the skin galvanic response peak as the feature values of the physiological data.

9. A server, characterized by The server comprises an obtaining module, an extracting module and a processing module; wherein, The obtaining module is configured to obtain multi-modal data of a game player; wherein the multi-modal data comprises visual data, voice data and physiological data; The extracting module is configured to extract respective feature values corresponding to the multi-modal data; The processing module is configured to input the respective feature values corresponding to the multi-modal data into a preset training model for data analysis, and perform game-related processing according to the analysis result.

10. A computer device comprising: A memory and a processor, the memory storing a computer program, characterized in that the processor, when executing the computer program, implements the steps of the method according to any one of claims 1 to 8.

11. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by a processor, implements the steps of the method according to any one of claims 1 to 8.