Ai digital human interaction method and system based on emotion recognition
By acquiring multi-channel electrophysiological signals from the user's prefrontal cortex, combining them with emotion label datasets and parallel processing channels, we generate cross-modal emotional semantic features, driving the AI digital human to achieve synchronized multimodal feedback with the user's emotional state. This solves the problem of the single form of emotional feedback in existing technologies and improves the emotional authenticity and naturalness of interactions in virtual social scenarios.
Patent Information
- Application Number
- CN202510539794.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2045-04-27
AI Technical Summary
In existing technologies, it is difficult for AI digital humans to accurately capture subtle emotional changes in users in virtual social scenarios, resulting in a single form of emotional feedback and an inability to adapt to the dynamic emotional interaction needs in complex social scenarios.
By acquiring multi-channel electrophysiological signals from the user's prefrontal cortex and combining them with the emotion label dataset to generate a multi-dimensional emotion vector, the parallel processing channels are used to jointly analyze the temporal dynamic characteristics and scalp spatial topological relationships, generate cross-modal emotional semantic features, and drive the facial micro-expressions, body movements and speech rhythm parameters of the AI digital human to achieve synchronous evolution with the user's emotional state.
It achieves accurate capture of user emotions and multimodal feedback, improves the emotional authenticity and naturalness of interactions in virtual social scenarios, and ensures that digital human feedback meets the contextual requirements of social scenarios.
Smart Images

Figure CN120447735B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of emotion recognition technology, and in particular to an AI digital human interaction method and system based on emotion recognition. Background Art
[0002] In immersive virtual social scenarios, users expect AI digital humans to perceive and accurately reflect their emotional states in real time, enabling natural emotional interactions. This requires the system to possess high-precision emotion recognition capabilities, capable of simultaneously analyzing the user's physiological signals and behavioral characteristics, and generating anthropomorphic feedback that aligns with the social context, thereby enhancing the realism and immersion of the interaction.
[0003] One current approach uses emotion analysis based on a single modality, such as facial expression recognition. This involves capturing the user's facial features with a camera, combining them with a pre-trained classification model to determine the emotion category, and then prompting the digital human to provide simple feedback. This approach improves recognition speed by optimizing the neural network structure and introduces an attention mechanism to enhance the ability to capture key areas of expression.
[0004] This solution relies on visual information, making it susceptible to interference from factors like lighting and occlusion, and struggles to capture subtle emotional changes (such as fleeting micro-expressions). Furthermore, single-modal analysis lacks comprehensive assessment of the multidimensional nature of emotion, resulting in a single form of digital human feedback that is unable to adapt to the dynamic emotional interactions required in complex social scenarios. Summary of the Invention
[0005] The present application provides an AI digital human interaction method and system based on emotion recognition, which is used to solve the problems of poor emotional resonance authenticity and poor scene adaptability in human-computer interaction in the existing technology.
[0006] In a first aspect, the present application provides an AI digital human interaction method based on emotion recognition, comprising:
[0007] Acquire multi-channel electrophysiological signals from the user's prefrontal cortex in an immersive virtual social environment.
[0008] Pattern matching the multi-channel electrophysiological signal with a preset emotion tag dataset to generate a multi-dimensional emotion vector reflecting the user's emotional state;
[0009] Performing a joint analysis of the temporal dynamic characteristics and scalp spatial topological relationship of the multi-dimensional emotion vector through parallel processing channels to generate cross-modal emotion semantic features;
[0010] Generate an emotional resonance interaction strategy adapted to the current immersive virtual social scene based on the association mapping between the cross-modal emotional semantic features and the pre-stored virtual character behavior library;
[0011] Based on the emotional resonance interaction strategy, the facial micro-expression, body movement and voice prosody parameters of the AI digital person are driven to generate a dynamic feedback loop that evolves synchronously with the user's emotional state.
[0012] Optionally, the joint analysis of the time dynamic characteristics and the scalp spatial topological relationship of the multi-dimensional emotion vector through the parallel processing channel generates a cross-modal emotional semantic feature, including:
[0013] The multi-dimensional emotion vector is decomposed into a time component sequence for representing time dynamic characteristics and a spatial component set for representing scalp spatial topological relationship;
[0014] In the parallel processing channel, the dynamic correlation parameters corresponding to the time component sequence and the spatial correlation parameters corresponding to the spatial component set are generated;
[0015] The dynamic correlation parameters and the spatial correlation parameters are nonlinearly coupled to generate spatiotemporal joint analysis parameters;
[0016] Based on the spatiotemporal joint analysis parameters, signal segments that simultaneously satisfy time fluctuation consistency and spatial distribution correlation are screened from the multi-dimensional emotion vector;
[0017] The signal segments are spatiotemporally reorganized to generate a cross-modal emotional semantic feature.
[0018] Optionally, the nonlinear coupling of the dynamic correlation parameters and the spatial correlation parameters to generate spatiotemporal joint analysis parameters includes:
[0019] According to the steepness of the signal amplitude change trend reflected by the dynamic correlation parameters, the spatial correlation parameters are determined, including electrode position weights and topological weight distribution proportions;
[0020] Based on the topological weight distribution proportions, the signal amplitude change trend and the electrode position weights are superimposed layer by layer to generate a multi-level parameter set;
[0021] The level parameters in the multi-level parameter set that satisfy the preset weight amplitude condition are dynamically intercepted to generate a parameter subset;
[0022] According to the overlapping coverage range of the time fluctuation sensitive area and the spatial distribution dense area in the parameter subset, the interaction intensity coefficient of the dynamic correlation parameters and the spatial correlation parameters is calculated;
[0023] Based on the interaction intensity coefficient, the level parameters in the parameter subset are weighted and aggregated to generate spatiotemporal joint analysis parameters.
[0024] Optionally, based on the topological weight distribution ratio, the signal amplitude change trend and the electrode position weight are superimposed step by step to generate a multi-level parameter set, including:
[0025] Dividing the signal amplitude change trend into a plurality of continuous levels according to the numerical interval of the topological weight distribution ratio;
[0026] In each level, the trend intensity factor corresponding to the signal amplitude change trend and the electrode position weight are scaled respectively according to the scaling coefficient corresponding to the level;
[0027] The scaled trend intensity factor is superimposed with the scaled electrode position weight to generate independent superposition parameter units corresponding to each level;
[0028] The superposition parameter units are combined in a preset hierarchical order to generate a multi-level parameter set.
[0029] Optionally, generating, in the parallel processing channel, the dynamic correlation parameters corresponding to the time component sequence and the spatial correlation parameters corresponding to the spatial component set includes:
[0030] In the parallel processing channel, the following process is executed: the signal amplitude change trend in the time component sequence is divided into a rising phase, a stable phase and a falling phase, and the duration of each phase and the transition steepness between adjacent phases are recorded;
[0031] generating a dynamic correlation parameter for reflecting the steepness of the signal amplitude change trend based on the duration and the conversion steepness;
[0032] determining a relative distance relationship between the electrodes according to actual distribution positions of the multi-channel electrodes in the spatial component set in the prefrontal cortex;
[0033] According to the preset scalp area sensitivity distribution map, differentiated topological weights are assigned to electrode pairs with different distance intervals;
[0034] Based on the relative distance relationship and the differentiated topological weight, a spatial association parameter reflecting spatial distribution characteristics is generated.
[0035] Optionally, generating an emotional resonance interaction strategy adapted to the current immersive virtual social scene based on the association mapping between the cross-modal emotional semantic features and a pre-stored virtual character behavior library includes:
[0036] Extracting emotion intensity parameters and emotion type parameters from the cross-modal emotion semantic features;
[0037] In the virtual role behavior library, a set of action amplitude parameters corresponding to the emotion intensity parameter and a set of expression parameters corresponding to the emotion type parameter are matched;
[0038] According to the current situation type of the immersive virtual social scene, a subset of action amplitude parameters that meet the spatial constraints of the immersive virtual social scene is screened from the set of action amplitude parameters;
[0039] The action amplitude parameters in the subset of action amplitude parameters and the expression parameters in the set of expression parameters are combined in different forms to generate a set of interaction strategies;
[0040] According to the interaction relationship coefficient between the emotion intensity parameter and the emotion type parameter, the strategies in the set of interaction strategies are prioritized;
[0041] From the set of interaction strategies prioritized, the optimal strategy with the highest priority is selected as the emotional resonance interaction strategy.
[0042] Optionally, based on the emotional resonance interaction strategy, the facial micro-expression, body movement and voice prosody parameters of the AI digital person are driven to generate a dynamic feedback loop that evolves synchronously with the user's emotional state, including:
[0043] The emotional resonance interaction strategy is parsed into a set of facial control instructions, a set of body control instructions and a set of voice control instructions;
[0044] According to the expression intensity value and expression duration parameter in the set of facial control instructions, a motion trajectory sequence of the corresponding facial muscle group is generated;
[0045] Based on the joint angle change rate and action amplitude value in the set of body control instructions, the real-time motion path of each body joint is calculated;
[0046] According to the pitch modulation coefficient and speech speed adjustment factor in the set of voice control instructions, a corresponding fundamental frequency curve and syllable duration sequence are generated;
[0047] The motion trajectory sequence, the real-time motion path, the fundamental frequency curve and the syllable duration sequence are time-synchronized to generate a multi-modal output control signal of the AI digital person;
[0048] According to the user's real-time updated cross-modal emotional semantic features, the multi-modal output control signal is dynamically corrected to form a dynamic feedback loop.
[0049] In a second aspect, the present application provides an AI digital person interaction system based on emotion recognition, comprising:
[0050] An acquisition module, used to obtain multi-channel electrophysiological signals from the user's prefrontal cortex area in an immersive virtual social scene;
[0051] a pattern matching module, configured to perform pattern matching on the multi-channel electrophysiological signal and a preset emotion tag dataset to generate a multi-dimensional emotion vector reflecting the user's emotional state;
[0052] A first generation module is configured to perform a joint analysis of the temporal dynamic characteristics and the scalp spatial topological relationship of the multi-dimensional emotion vector through a parallel processing channel to generate a cross-modal emotion semantic feature;
[0053] A second generation module is used to generate an emotional resonance interaction strategy adapted to the current immersive virtual social scene based on the association mapping between the cross-modal emotional semantic features and the pre-stored virtual character behavior library;
[0054] The third generation module is used to drive the facial micro-expressions, body movements and voice rhythm parameters of the AI digital human based on the emotional resonance interaction strategy to generate a dynamic feedback loop that evolves synchronously with the user's emotional state.
[0055] In a third aspect, the present application provides a computing device comprising a processor and a memory, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute an AI digital human interaction method based on emotion recognition as described in any one of the first aspects.
[0056] In a fourth aspect, the present application provides a computer storage medium having computer program instructions stored thereon, which, when executed by a processor, implement an AI digital human interaction method based on emotion recognition as described in any one of the first aspects.
[0057] In the present application, an AI digital human interaction method based on emotion recognition is provided, which includes: obtaining multi-channel electrophysiological signals from the user's prefrontal cortex area in an immersive virtual social scene; pattern matching the multi-channel electrophysiological signals with a preset emotion label data set to generate a multidimensional emotion vector for reflecting the user's emotional state; jointly analyzing the temporal dynamic characteristics and scalp spatial topological relationships of the multidimensional emotion vector through parallel processing channels to generate cross-modal emotion semantic features; generating an emotion resonance interaction strategy adapted to the current immersive virtual social scene based on the association mapping between the cross-modal emotion semantic features and a pre-stored virtual character behavior library; and based on the emotion resonance interaction strategy, driving the facial micro-expressions, body movements and speech rhythm parameters of the AI digital human to generate a dynamic feedback loop that evolves synchronously with the user's emotional state.
[0058] The technical solution provided by this application has the following beneficial effects:
[0059] This application directly captures the physiological characteristics of emotions through prefrontal cortex signal collection, overcoming the lag of traditional external behavior monitoring; combines the emotion label data set to complete the precise mapping of signals and emotions, and establishes a quantifiable emotion representation system; through parallel processing, it simultaneously mines the temporal evolution laws and spatial distribution patterns of emotion signals to improve the robustness of recognition of complex emotional states; based on the adaptive mapping of the virtual character behavior library, it ensures that the digital human feedback meets the contextual requirements of social scenarios; realizes multimodal collaborative output of face, body, and voice, forming a natural interactive experience that is linked to the user's emotions in real time.
[0060] Furthermore, the present application also decomposes the multidimensional emotion vector into a time component sequence and a spatial component set, generates dynamic correlation parameters and spatial correlation parameters in parallel, extracts spatiotemporal joint analysis parameters through nonlinear coupling, screens signal segments with spatiotemporal consistency and reorganizes them into cross-modal emotional semantic features.
[0061] Moreover, this technology breaks through the limitations of single-dimensional analysis, improves the completeness and accuracy of emotion feature extraction, and enables digital humans to more delicately capture the spatiotemporal correlation characteristics of user emotions.
[0062] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0064] Figure 1 A flowchart of an AI digital human interaction method based on emotion recognition provided in an embodiment of the present application;
[0065] Figure 2 A schematic diagram of the structure of an AI digital human interaction system based on emotion recognition provided in an embodiment of the present application;
[0066] Figure 3 A schematic diagram of the structure of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0067] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0068] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to being different types.
[0069] Researchers have found that AI digital human interactions in existing virtual social scenarios primarily rely on single-modal emotion recognition, making it difficult to accurately capture subtle changes in users' emotions and achieve natural emotional feedback. Based on this, the present application provides an AI digital human interaction method based on emotion recognition. This method can acquire real-time user prefrontal cortex activity data through multi-channel electrophysiological signals, construct a multidimensional emotional representation based on an emotion tag dataset, and utilize parallel processing channels to synchronously analyze spatiotemporal features to generate cross-modal emotional semantics. Ultimately, this method drives the digital human to achieve multimodal dynamic feedback that is highly synchronized with the user's emotional state. The technical solution of this application is applicable to virtual social scenarios that require high immersion and emotional authenticity.
[0070] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0071] Figure 1 A flowchart of an AI digital human interaction method based on emotion recognition provided in an embodiment of the present application is shown in FIG. Figure 1 As shown, the method includes:
[0072] Step 101: In an immersive virtual social scene, obtain multi-channel electrophysiological signals from the user's prefrontal cortex area.
[0073] In this step, the prefrontal cortex represents the functional area of the brain responsible for emotion regulation, and its electrical activity can directly reflect the user's emotional state. Multi-channel electrophysiological signals refer to bioelectrical signals from the user's prefrontal cortex region collected via multiple electrodes, containing information on potential changes at different spatial locations. Each electrode corresponds to a signal acquisition channel.
[0074] In an embodiment of the present application, researchers integrated a multi-channel electrode array into a virtual social head-mounted device. When the user wears the device to enter a virtual scene, the electrode array continuously collects potential fluctuation signals in the prefrontal cortex area at a fixed sampling frequency. After removing the power frequency interference through a signal amplifier, the signals of each channel are time-aligned and finally a synchronized multi-channel electrophysiological signal data stream is output.
[0075] For example, in a virtual social platform, users wear a VR headset with integrated 16-channel electrodes, which are arranged in the forehead area (FP1, FP2, etc.) according to the international 10-20 system. When the user communicates with the virtual digital human, the system collects signals from each channel at a sampling rate of 256 times per second, filters them with a 50Hz notch filter, and outputs 16 synchronized EEG signals for subsequent processing.
[0076] Step 102: performing pattern matching on the multi-channel electrophysiological signal and a preset emotion tag data set to generate a multi-dimensional emotion vector for reflecting the user's emotional state.
[0077] In this step, the emotion labeling dataset contains a collection of standard EEG features for different emotional states. Each emotion corresponds to a specific EEG waveform pattern. The multidimensional emotion vector represents the mathematical representation obtained by mapping the original signal into the emotion feature space, with each dimension representing the intensity of an emotion feature.
[0078] In an embodiment of the present application, the collected multi-channel signal is divided into segments according to the time window, the time-frequency features (such as power spectral density) of each segment are extracted, and the similarity is calculated with the standard features in the emotion label data set. The emotion category and intensity corresponding to the current signal are determined by nearest neighbor classification, and finally a feature vector containing multiple emotion dimensions (such as pleasure and excitement) is generated.
[0079] For example, the system divides a 10-second EEG signal into 20 0.5-second segments, calculates the power ratio of theta waves (4-7Hz) and beta waves (13-30Hz) for each segment, and matches it with the reference pattern of the "happy" state in the database. When more than 8 segments are matched successfully, an emotion vector [pleasure 0.8, excitement 0.6] is generated, where the value is determined by the ratio of the number of matching segments to the total number of segments.
[0080] Step 103: performing a joint analysis of the temporal dynamic characteristics and the scalp spatial topological relationship of the multi-dimensional emotion vector through parallel processing channels to generate cross-modal emotion semantic features.
[0081] In this step, the parallel processing channel refers to a dual-path computing architecture that simultaneously and independently processes temporal dynamics (signal timing fluctuations) and scalp spatial topology (electrode position distribution). Its purpose is to separate the temporal and spatial feature extraction paths, avoiding the mutual interference of temporal and spatial features in traditional serial processing, thereby ensuring the real-time and accuracy of emotion recognition in virtual social scenarios. The prefrontal cortex is the core functional area of the brain for emotion processing. Its electrical activity is transmitted to the scalp surface through the skull, and signals are collected by electrode arrays arranged in corresponding areas of the scalp. Therefore, the spatial topology of the scalp directly reflects the spatial distribution characteristics of physiological signals in the prefrontal cortex. Temporal dynamics represent the pattern of signal amplitude changes over time, such as periodic fluctuations or sudden changes. Spatial topology represents the spatial distribution pattern of signal strength at different electrode locations. Cross-modal emotional semantic features represent high-level emotion representations that integrate temporal and spatial features.
[0082] In this embodiment, a multidimensional emotion vector is fed into a parallel processing architecture. One path uses a sliding window to analyze the temporal variation of the signal amplitude, extracting features such as the period of fluctuation. Another path calculates the gradient distribution of signal intensity based on the spatial coordinates of the electrodes. These two features are weightedly fused using an attention mechanism, filtering for feature segments with high spatiotemporal consistency and recombining them into the final emotional semantic features.
[0083] For example, for the vector [pleasure 0.8, excitement 0.6], the time channel detects a pleasure fluctuation with a period of 0.3 seconds, and the spatial channel finds that the signal intensity of the FP1 electrode is higher than that of FP2. After fusion, the feature "high intensity left forehead pleasure" is generated, where "high intensity" is determined by the spatial gradient value exceeding the threshold.
[0084] Step 104: Generate an emotional resonance interaction strategy adapted to the current immersive virtual social scene based on the association mapping between the cross-modal emotional semantic features and the pre-stored virtual character behavior library.
[0085] In this step, the virtual character behavior library contains a set of parameterized descriptions of the digital human's various emotional expression actions. The emotional resonance interaction strategy represents the multimodal behavior plan adopted by the digital human in response to the user's emotions.
[0086] In an embodiment of the present application, the behavior library is queried based on the emotional semantic features, and the basic emotion type (such as "joy") is first matched. Then, the appropriate action amplitude range is filtered according to the feature intensity. Combined with the spatial constraints of the current virtual scene (such as the size of the virtual room), a strategy set containing expression, action, and voice parameters is generated. Finally, the optimal strategy is selected based on the feature degree.
[0087] For example, for the "high-intensity left forehead happiness" feature, the system selects the "waving and smiling" strategy from the library, selects a large wave (amplitude parameter 0.7) combined with a toothy smile (expression parameter 0.8) based on the happiness level of 0.8, and adjusts the voice tone to increase by 15% to form the final interaction strategy.
[0088] Step 105: Based on the emotional resonance interaction strategy, drive the facial micro-expressions, body movements and voice rhythm parameters of the AI digital human to generate a dynamic feedback loop that evolves synchronously with the user's emotional state.
[0089] In this step, facial micro-expression parameters are quantitative indicators used to control the facial muscle movements of the AI digital human. These include parameters for subtle changes in facial expression, such as the degree of eye contraction, the angle of the mouth corners' upward movement, and the extent of the frown. These parameters are achieved by adjusting the blending weights of the virtual facial model. They can accurately reproduce the instantaneous micro-expression characteristics of real human emotional expressions (such as a brief smile within 0.5 seconds), making the digital human's expressions more emotionally realistic. Body movement parameters are physical quantities that describe the trajectory of the digital human's joints. They include dimensions such as joint rotation angle, movement amplitude, and movement speed. These parameters control the naturalness and expressiveness of the digital human's body movements (such as the amplitude of a wave or the frequency of a nod) by driving the virtual skeletal animation system, ensuring that the intensity of the movement matches the intensity of the recognized user's emotions and enhancing the expressiveness of interactive body language. Speech prosody parameters are a set of acoustic features that regulate the digital human's speech output. These include dimensions such as fundamental frequency (pitch), speech rate (syllable duration), and volume (amplitude). These parameters are adjusted in real time by the speech synthesis engine, ensuring that the digital human's voice, in terms of intonation and tempo, matches its current emotional state (for example, a 15% rise in pitch when happy), achieving a natural flow of emotional expression. This dynamic feedback loop means that the digital human's output behavior influences the user's emotions, thus forming a closed-loop interactive system.
[0090] In this embodiment, the parameters in the strategy are converted into digital human control instructions. Facial parameters drive weight changes, limb parameters control skeletal animation, and voice parameters adjust the synthesized speech characteristics. All three are kept in strict time synchronization. Simultaneously, the user's newly generated electrophysiological signals are monitored in real time, and the subsequent feedback intensity is dynamically adjusted.
[0091] For example, when the digital human performs a large wave and a toothy smile, the system detects that the signal strength of the user's FP1 electrode continues to rise by 10%, and then automatically enhances the subsequent feedback, raising the smile level from 0.8 to 0.9, forming a positive feedback loop of emotional reinforcement.
[0092] This method precisely captures the user's physiological signals, establishes a fine-grained emotional representation system, and utilizes spatiotemporal feature fusion technology to improve recognition accuracy. Ultimately, it achieves real-time synchronization and dynamic adaptation of the digital human's feedback to the user's emotional state, enhancing the naturalness and emotional realism of virtual social interactions. The system can autonomously perceive emotional changes and respond in an anthropomorphic manner, creating a continuous emotional resonance experience.
[0093] To address the problem of insufficient spatiotemporal feature fusion in existing emotion recognition methods in virtual social scenarios, in some embodiments, step 103: performing a joint analysis of temporal dynamic characteristics and scalp spatial topological relationships of the multi-dimensional emotion vector through parallel processing channels to generate cross-modal emotion semantic features includes:
[0094] Step 201: Decompose the multidimensional emotion vector into a time component sequence for characterizing temporal dynamic characteristics and a set of spatial components for characterizing scalp spatial topological relationships.
[0095] In step 201, the time component sequence refers to a waveform feature sequence extracted from the multidimensional emotion vector that reflects the temporal changes of the signal, including time-varying characteristics such as signal amplitude and frequency. The spatial component set refers to a feature set extracted from the multidimensional emotion vector that reflects the distribution of signal strength at different electrode locations, including the relative strength relationship between electrodes and spatial gradient information.
[0096] In an embodiment of the present application, a time window segmentation technique is used to divide the multidimensional emotion vector into continuous segments of fixed duration, and the time domain statistical features of each segment are extracted to form a time component sequence; at the same time, based on the physical coordinate position of the electrodes, the spatial differential features of the signal strengths of adjacent electrodes are calculated to construct a set of spatial components.
[0097] Step 202: In a parallel processing channel, generate dynamic correlation parameters corresponding to the time component sequence and spatial correlation parameters corresponding to the spatial component set.
[0098] In step 202, dynamic correlation parameters are a set of parameters that quantify the fluctuation characteristics of the time component series, including indicators such as signal change rate and cycle length. Spatial correlation parameters are a set of parameters that characterize the distribution law of the spatial component series.
[0099] In an embodiment of the present application, differential operations and extreme value detection are performed on the time component sequence to extract the key turning point features of the signal fluctuation and generate dynamic correlation parameters; density clustering and gradient calculation are performed on the spatial component set to identify the spatial aggregation area features of the signal strength and generate spatial correlation parameters.
[0100] Step 203: nonlinearly couple the dynamic correlation parameters with the spatial correlation parameters to generate spatiotemporal joint analysis parameters.
[0101] In step 203, the spatiotemporal joint analysis parameter is a composite parameter obtained by fusing spatiotemporal features, reflecting the coordinated variation of the emotion signal in the time domain and the space domain.
[0102] In an embodiment of the present application, a feature fusion method based on attention weight is adopted to dynamically adjust the fusion weight of the spatial correlation parameter according to the signal change rate in the dynamic correlation parameter, and the two types of parameters are weighted and spliced, and then a spatiotemporal joint analysis parameter is generated through nonlinear transformation.
[0103] Step 204: Based on the spatiotemporal joint analysis parameters, filter out signal segments that satisfy both temporal fluctuation consistency and spatial distribution correlation from the multidimensional emotion vector.
[0104] In step 204, temporal fluctuation consistency refers to signal segments having similar fluctuation patterns and periodic characteristics in the time domain. Temporal fluctuation consistency is manifested by the frequency of periodic changes in the signal amplitude in the time component sequence being stable within a preset range (e.g., 0.2-0.5 Hz), and the similarity of signal fluctuation trends within adjacent time windows exceeding a set threshold (e.g., a Pearson correlation coefficient ≥ 0.8), ensuring the continuity of emotional states in the temporal dimension. Spatial distribution correlation refers to the stable intensity distribution pattern of signal segments in the spatial domain. Spatial distribution correlation is manifested by the fact that the difference in signal intensity at at least three adjacent electrode locations in the spatial component set does not exceed a preset tolerance (e.g., a standard deviation ≤ 15%), and the interaction strength coefficient between electrode clusters reaches a minimum correlation threshold (e.g., ≥ 0.7), reflecting the spatial coordination of electrical activity in the prefrontal cortex. Signal segments are spatiotemporal data units extracted from a multidimensional emotion vector that have complete emotional representation significance, including the electrophysiological signal waveform within a specific time window and its distribution characteristics in scalp space. Each segment consists of three elements: (1) the original signal waveform collected continuously in the time dimension; (2) the intensity distribution diagram of the multi-electrode signal in the spatial dimension; and (3) spatiotemporal correlation metadata (such as the segment start time, duration, and main active electrode area). These segments must simultaneously meet the dual conditions of stable temporal fluctuation patterns (such as the continuous presentation of 0.3-second periodic fluctuations of pleasure) and spatial distribution characteristics (such as the left forehead electrode signal intensity being continuously higher than the right forehead by more than 30%). They are the smallest identifiable units of emotional state in the spatiotemporal dimension.
[0105] In an embodiment of the present application, the similarity score between each signal segment and the spatiotemporal joint analysis parameter is calculated, and segments with a time domain similarity higher than a threshold and a spatial domain distribution pattern matching higher than a threshold are selected as valid segments.
[0106] Step 205: Perform spatiotemporal reorganization on the signal segments to generate cross-modal emotional semantic features.
[0107] In step 205, spatiotemporal reorganization refers to the process of reintegrating the screened valid segments into a unified feature expression according to spatiotemporal coordinates.
[0108] In an embodiment of the present application, after the temporal features and spatial features of the effective segments are interpolated and aligned respectively, cross-modal emotional semantic features with temporal and spatial consistency are generated through feature splicing.
[0109] Here's a specific example:
[0110] In a specific embodiment of a virtual social platform, when processing the generated emotion vector [pleasure 0.8, excitement 0.6], the system first decomposes the 16-channel EEG signal into a temporal component sequence and a spatial component set. The temporal component sequence uses a sliding window Fourier transform to extract the theta / beta power ratio curve for each 0.5-second segment, identifying pleasure fluctuations with a 0.3-second period (the period is calculated by the inverse of the power spectrum peak frequency). The spatial component set calculates the signal intensity difference between adjacent electrode pairs, such as FP1 and FP2, and finds that the average intensity of the FP1 electrode reaches 45 μV (15 μV higher than that of FP2, exceeding the preset difference threshold of 10 μV). In parallel processing, the temporal channel generates dynamic correlation parameters based on the fluctuation period of 0.3 seconds and the amplitude change rate of 0.5 μV / ms; the spatial channel generates spatial correlation parameters based on the 3 cm electrode spacing and the intensity gradient of 5 μV / cm. Through nonlinear coupling (formula: spatiotemporal parameter = time parameter × 0.6 + space parameter × 0.4, with weight coefficients calibrated experimentally), the spatiotemporal joint analysis parameter value was 0.72 (above the effective threshold of 0.7). The system selected three 0.5-second signal segments that met the criteria (temporal fluctuation similarity > 0.8 and spatial distribution matching > 0.75). After spatiotemporal reorganization, a cross-modal signature of "high-intensity left frontal pleasure" was generated. The "high intensity" was determined when the FP1 electrode maintained a constant intensity of 40 μV or above in all three segments (the baseline for a standard pleasure state is 30 μV).
[0111] In the embodiment of the present application, accurate modeling of user emotions is achieved through the parallel extraction and deep fusion of spatiotemporal features, enabling digital humans to capture subtle spatiotemporal changes in emotions, thereby improving the accuracy of emotion recognition and the timeliness of feedback in virtual social interactions.
[0112] In order to improve the accuracy and reliability of the spatiotemporal fusion of emotional features, in some embodiments, step 203: performing nonlinear coupling of the dynamic correlation parameters and the spatial correlation parameters to generate spatiotemporal joint analysis parameters includes:
[0113] Step 301: Determine the spatial correlation parameter according to the steepness of the signal amplitude change trend reflected by the dynamic correlation parameter, where the spatial correlation parameter includes the electrode position weight and the topological weight distribution ratio.
[0114] In step 301, the steepness of the signal amplitude change trend refers to the rising or falling slope of the EEG signal waveform per unit time, reflecting the intensity of emotional changes. The electrode position weight is a preset quantitative value based on the emotional sensitivity of the scalp region, reflecting the different contributions of different brain regions to emotion recognition. This weight is set based on the functional zoning of the prefrontal cortex. For example, the FP1 electrode corresponds to the pleasure processing area with a preset weight of 0.7, and the FP2 electrode corresponds to the inhibitory control area with a preset weight of 0.5. The topological weight distribution ratio is a dynamically adjusted parameter of the inter-electrode correlation strength, determined by the base electrode weight and the distance attenuation coefficient. The calculation formula is: target electrode weight = base weight × (1-0.1 × spacing cm), ensuring that electrodes within a 3cm spacing maintain more than 90% weight transfer. When the spacing exceeds 5cm, only 50% weight is retained.
[0115] In an embodiment of the present application, the steepness is quantified by calculating the absolute value of the first-order derivative of the signal. When the slope is detected to be greater than 0.4 μV / ms, the preset weights of adjacent electrodes (spacing <4 cm) at the corresponding time point are increased by 30%, generating a spatial correlation parameter containing dynamically adjusted weights.
[0116] Step 302: Based on the topological weight distribution ratio, the signal amplitude change trend and the electrode position weight are superimposed layer by layer to generate a multi-level parameter set.
[0117] In step 302, the multi-level parameter set refers to a hierarchical feature set divided by weight value intervals, and each level contains a combination of spatiotemporal features within a specific weight range.
[0118] In an embodiment of the present application, the weight values are divided into 5 levels (0-0.2, 0.2-0.4, etc.) at intervals of 0.2, and within each level, the signal amplitude change rate is multiplied by the adjusted electrode weight to generate a hierarchical parameter unit containing the amplitude-weight product feature.
[0119] Step 303: Dynamically intercept the hierarchical parameters that meet the preset weight increase conditions in the multi-level parameter set to generate a parameter subset.
[0120] In step 303, the preset weight increase condition refers to the constraint that the weight increase between adjacent layers must meet a minimum threshold, which is used to filter the changing feature layers. The parameter subset is a set of high-value feature units selected from the multi-level parameter set, including layer parameters with weight increases (e.g., increases > 0.15). Each subset unit is composed of three elements: a time stamp (the onset of the emotional fluctuation), a spatial stamp (the set of active electrode IDs), and a fusion parameter value (the product of amplitude and weight).
[0121] In an embodiment of the present application, a threshold value of 0.15 is set for the weight increase between levels. When it is detected that the weight of a level increases by more than this value compared with the previous level, the parameter units of the level are intercepted and stored in a parameter subset to ensure that the data of the key period when the emotional characteristics suddenly change is retained.
[0122] Step 304: Calculate the interaction strength coefficient between the dynamic correlation parameter and the spatial correlation parameter according to the overlapping coverage of the time fluctuation sensitive area and the spatial distribution dense area in the parameter subset.
[0123] In step 304, the temporal fluctuation-sensitive region refers to a continuous time segment where the rate of change of the signal amplitude in the dynamic correlation parameter exceeds the emotion recognition threshold (e.g., slope ≥ 0.3 units / second). This is identified by detecting the extreme points of the first-order derivative of the signal and selecting time windows where the absolute value of the derivative is continuously greater than a preset threshold (e.g., > 0.25). The spatially densely distributed region refers to a continuous scalp region where the electrode position weights in the spatial correlation parameter exceed the spatial clustering threshold (e.g., weight ≥ 0.6) and the distance between adjacent electrodes is less than 5 cm. This is identified by calculating the spatial density gradient of the electrode weight values and selecting the closed region enclosed by the sudden drop in the gradient change rate (e.g., the second-order derivative zero point). The overlapping coverage refers to the intersection of the temporal fluctuation-sensitive region (e.g., a period of theta wave enhancement lasting 0.3 seconds) and the spatially densely distributed region (e.g., the FP1-F3-F7 electrode cluster) in the time and space dimensions. This range is determined by comparing the overlap between temporal and spatial markers. When the signal strength of an electrode exceeds a spatial threshold during a sensitive period, it is included in the overlap range. Ultimately, the coverage range is quantified by the ratio of the number of overlapping electrodes to the total number of electrodes. The interaction strength coefficient is a metric that quantifies the degree of spatiotemporal coordination, reflecting the activity of specific brain regions during critical periods of emotional fluctuation. For example, when the duration of the detected temporal fluctuation-sensitive region (T = 2 seconds) and the number of electrodes covering the spatially densely distributed region (N = 4) satisfy T × N ≥ 5, the interaction strength coefficient K = 0.7 × (mean weight of overlapping electrodes) + 0.3 × (mean rate of change of signal amplitude), where the mean weight is the arithmetic mean of the topological weight distribution ratio of the overlapping electrodes (e.g., 0.72), and the mean rate of change of signal amplitude is the average of the absolute values of the slopes within the sensitive region (e.g., 0.4 units / second). Thus, K = 0.7 × 0.72 + 0.3 × 0.4 = 0.624. The coefficient calculation reflects the coupling strength of spatiotemporal features in emotion recognition, and the threshold is set (such as K ≥ 0.6).
[0124] In an embodiment of the present application, the product of the average weight of the electrodes in the overlapping area and the signal change rate is calculated as the coefficient base value, and then multiplied by the overlapping area ratio (number of overlapping electrodes / total number of electrodes). When more than three electrodes in a certain area simultaneously meet the weight > 0.6 and the change rate > 0.3 μV / ms, it is determined that effective interaction occurs.
[0125] Step 305: Based on the interaction strength coefficient, perform weighted aggregation on the hierarchical parameters in the parameter subset to generate spatiotemporal joint analysis parameters.
[0126] In step 305, weighted aggregation refers to the process of weighting the importance of feature parameters according to the interaction strength and then fusing them.
[0127] In an embodiment of the present application, the feature units of each level in the parameter subset are weighted and summed according to their interaction intensity coefficients, and the feature weights of the coefficients > 0.7 are doubled to generate the final spatiotemporal joint analysis parameters. For example, first, the numerical range of the interaction intensity coefficient is divided into three intervals: high, medium, and low. The parameter subset belonging to the high interval is assigned a first weight value, the parameter subset belonging to the medium interval is assigned a second weight value, and the parameter subset belonging to the low interval is assigned a third weight value, wherein the first weight value is greater than the second weight value, and the second weight value is greater than the third weight value; then the weighted parameter subsets are arranged in the time order of their corresponding time fluctuation sensitive areas, and the parameter subsets with the same spatial distribution dense areas in adjacent time windows are merged to form a merged parameter subset; finally, the merged parameter subsets are sorted according to the size of their weight values, and a preset number of parameter subsets with the highest weight values are selected for arithmetic average calculation to generate the spatiotemporal joint analysis parameters.
[0128] Here's a specific example:
[0129] In an embodiment of a virtual social platform, when the system detects that the user's FP1 electrode signal rises from 40μV to 55μV within 0.3 seconds (steepness = 15μV / 0.3s = 50μV / s), the FP1 electrode position weight is increased from 0.7 to 0.85 (increase amplitude = 0.15 × steepness coefficient 1.2), and a topological weight of 0.8 is assigned to adjacent F3 electrodes at a spacing of 3cm (calculation formula: 0.85×0.9=0.765, rounded to 0.8). The signal change trend and weight are superimposed in five levels at 0.1 intervals. The 0.7-0.8 level is truncated as a parameter subset because the weight increase is 0.1 (originally 0.6-0.7 levels) and the signal change rate exceeds 40μV / s. In this subset, the signal strength of the FP1-F3 electrode clusters exceeded 45 μV within a 0.3-second period (spatially dense), completely overlapping the temporal fluctuation window (coverage = 2 / 16 = 0.125), and the interaction strength coefficient = (0.8 × 50) × 0.125 = 5. The three intercepted hierarchical parameters (with coefficients of 5, 4.2, and 3.8, respectively) were weighted and aggregated (weights 0.5, 0.3, and 0.2), resulting in a spatiotemporal joint analysis parameter value of 5 × 0.5 + 4.2 × 0.3 + 3.8 × 0.2 = 4.52, which was normalized to 0.82.
[0130] In the embodiment of the present application, through hierarchical weight adjustment and dynamic interception mechanism, accurate capture of key emotional features is achieved. Combined with the quantification of spatiotemporal interaction intensity, the generated joint parameters can reflect both the instantaneous characteristics of emotional mutations and the spatial patterns of coordinated activities of brain regions, thereby improving the accuracy and naturalness of the digital human's emotional feedback.
[0131] In order to achieve a more refined hierarchical representation of emotional features, in some embodiments, step 302: based on the topological weight distribution ratio, the signal amplitude change trend and the electrode position weight are gradually superimposed to generate a multi-level parameter set, including:
[0132] Step 401: Divide the signal amplitude variation trend into a plurality of continuous levels according to the numerical interval of the topology weight distribution ratio.
[0133] In step 401, a numerical interval is a continuous numerical range that divides the topological weight distribution ratio into equal widths or equal frequencies, with each interval corresponding to a feature level. A continuous level is a level division method in which the boundaries of the numerical intervals are connected and there are no gaps between them, ensuring a smooth transition to the adjacent level when the weight changes.
[0134] In an embodiment of the present application, the weight range of 0-1 is divided into 10 levels with an interval of 0.1 (0-0.1, 0.1-0.2, ..., 0.9-1.0). When the weight value of an electrode is detected to be 0.63, it is classified into the 0.6-0.7 level, and the boundary transition coefficient of the adjacent level is recorded at the same time (such as 0.07 from the next level).
[0135] Step 402: In each level, the trend intensity factor corresponding to the signal amplitude change trend and the electrode position weight are scaled respectively according to the scaling coefficient corresponding to the level.
[0136] In step 402, the scaling factor is a gain parameter dynamically adjusted based on the level's position, used to amplify or reduce the importance of features within the level. The trend strength factor is a derivative parameter that quantifies the severity of signal amplitude changes and is calculated as the ratio of the absolute value of the signal slope to a reference slope.
[0137] In the embodiment of the present application, the scaling factor of the 0.6-0.7 level is set to 1.2, the original trend strength factor 1.5 (slope 45 μV / s ÷ benchmark 30 μV / s) of the FP1 electrode is scaled to 1.8, and its weight 0.63 is scaled to 0.756 (0.63×1.2).
[0138] Step 403: Superimpose the scaled trend intensity factor and the scaled electrode position weight to generate a superimposed parameter unit corresponding to each level independently.
[0139] In step 403, the superposition parameter unit is the smallest unit for completing feature fusion in a single level, and includes the scaled spatiotemporal feature values and their metadata (level label, timestamp, etc.).
[0140] In the embodiment of the present application, the scaled trend intensity factor 1.8 is multiplied by the weight 0.756 to generate the superposition parameter value 1.3608 of the level, and is packaged with the level mark (0.6-0.7) and the time window mark (t1-t2) as a complete parameter unit.
[0141] Step 404: combining the superposition parameter units according to the preset level order to generate a multi-level parameter set.
[0142] In step 404, the level order refers to the order arranged from small to large according to the proportion of the topological weight distribution value interval, which is derived from the division result of the "continuous level". The relationship is: the continuous level is the division method (the value is uninterrupted), and the level order is the arrangement method (arranged according to the value size).
[0143] In the embodiment of the present application, the generated 10 level parameter units are arranged in ascending order of weight interval (0-0.1 level unit arranged first,..., 0.9-1.0 level unit arranged last), to construct a parameter set with clear hierarchical relationship.
[0144] The following is a specific example:
[0145] In the virtual social platform embodiment, the system detects that the user FP1 electrode signal rises from 40 μV to 55 μV (slope 50 μV / s) within 0.3 seconds, sets the FP1 electrode weight to 0.85, and sets the adjacent F3 electrode weight to 0.8. The 0.5-1.0 weight range is divided into 5 levels (0.5-0.6, 0.6-0.7, 0.7-0.8, 0.8-0.9, 0.9-1.0), and the scaling coefficient corresponding to the 0.7-0.8 level is 1.3 (coefficient formula: 1+0.1×level order number). In this level, the trend intensity factor 1.67 (50 μV / s ÷ 30 μV / s reference) is scaled to 2.17, and the electrode weight 0.8 is scaled to 1.04, and the product of the two is the superposition parameter value of this level, which is 2.26. Similarly, the parameter values of other levels are calculated (0.6-0.7 level 1.85, 0.8-0.9 level 2.58), and the parameter set [1.85, 2.26, 2.58] is generated by arranging them in level order, and the parameter value of the 0.7-0.8 level is 2.26=(50÷30×1.3)×(0.8×1.3).
[0146] In the embodiment of the present application, through the fine classification of the weight interval and the self-adaptive scaling of the features in the level, the smooth quantization representation of the emotion intensity is realized, so that the feedback behavior of the digital person can accurately match the subtle gradient of the user's emotion change, and the natural degree and expressiveness of emotion transmission in virtual social interaction are improved.
[0147] In order to more accurately extract the spatiotemporal features of the emotion signal, in some embodiments, step 202: generating, in the parallel processing channel, the dynamic correlation parameters corresponding to the time component sequence and the spatial correlation parameters corresponding to the spatial component set, includes:
[0148] Step 501: In the parallel processing channel, the following process is executed: the signal amplitude change trend in the time component sequence is divided into a rising phase, a stable phase and a falling phase, and the duration of each phase and the transition steepness between adjacent phases are recorded.
[0149] In step 501, the rising phase refers to the period during which the signal amplitude continuously increases. The plateau phase refers to the period during which the amplitude fluctuation is less than a threshold. The falling phase refers to the period during which the amplitude continuously decreases. The transition steepness is the absolute value of the slope change between adjacent phases, reflecting the severity of the emotional state transition. Phase refers to the temporal state division of the signal amplitude change in the time component sequence (rising / plateau / falling), and level refers to the parameter processing hierarchy divided by the numerical interval of the topological weight. Phases are used to construct the temporal dimension characteristics of the dynamic correlation parameters, and levels are used to organize the multi-level coupling relationship of the spatial correlation parameters. Adjacent phases refer to the combination of directly connected phases among the rising phase, plateau phase, and falling phase (e.g., an rising phase and the following plateau phase constitute a group of adjacent phases). The duration of each phase refers to the continuous length of time the signal amplitude remains in the rising, plateau, or falling states, reflecting the persistence characteristics of the emotional state. Specifically, it includes: the time from the signal starting to increase to reaching the peak in the rising phase; the duration during the plateau phase during which the amplitude fluctuation does not exceed a preset threshold (e.g., ±5 μV); and the time from the peak to the baseline in the falling phase. The sum of these three constitutes a complete emotional fluctuation cycle. This parameter is used to quantify the temporal characteristics of emotional reactions. For example, short-term anger is characterized by a short rise and a long fall, while sustained happiness shows a long plateau.
[0150] In an embodiment of the present application, a sliding window is used to detect the sign change of the first-order derivative of the signal. When the derivative is continuously positive and exceeds the threshold, it is marked as a rising stage. When the absolute value of the derivative is less than the threshold, it is marked as a stable stage. The stage dividing point is determined by combining the extreme point detection, and finally the duration of each stage and the average slope of the transition interval are output as the conversion steepness.
[0151] Step 502: Based on the duration and the transition steepness, a dynamic correlation parameter is generated to reflect the steepness of the signal amplitude change trend.
[0152] In step 502, the dynamic correlation parameter is a set of parameters that comprehensively characterizes the signal change pattern, including the duration proportion of each stage and the weighted value of the transition steepness. All stages (rising, steady, and falling) are combined to generate a unified dynamic correlation parameter, which integrates the weighted combination of the duration and transition steepness of each stage, rather than a separate parameter corresponding to each stage.
[0153] In the embodiment of the present application, the proportion of the rising stage duration in the total duration is taken as a basic parameter, and multiplied by the geometric mean of the conversion steepness to generate the final dynamic correlation parameter, which ensures the consideration of both emotional duration and state conversion speed.
[0154] Step 503: According to the actual distribution position of the multi-channel electrodes in the prefrontal cortex in the set of spatial components, the relative distance relationship between each electrode is determined.
[0155] In step 503, the multi-channel electrode refers to a plurality of independent signal acquisition points arranged in the corresponding region of the prefrontal cortex of the scalp (such as Fp1 / Fp2 electrodes in a 16-lead EEG cap), and each electrode corresponds to a physical signal acquisition channel. The actual distribution position of the prefrontal cortex refers to the specific spatial arrangement of the multi-channel electrode on the scalp surface corresponding to the prefrontal functional area of the brain, which is positioned according to the international 10-20 system standard. For example, the FP1 electrode is located in the corresponding region of the left prefrontal dorsal lateral cortex, and the F3 electrode is located in the left prefrontal dorsal medial region, with a distance of 3-4 cm. This distribution directly reflects the spatial topological relationship of different brain function subareas (such as the pleasure processing area and the cognitive control area), providing an anatomical basis for spatial feature extraction for emotion recognition. The relative distance relationship refers to the quantitative representation of the three-dimensional spatial distance between electrodes.
[0156] In the embodiment of the present application, according to the three-dimensional coordinates provided by the international electroencephalogram electrode positioning system, the Euclidean distance between each pair of electrodes is calculated, and a weighted distance relationship matrix is generated in combination with a preset functional area connection strength coefficient.
[0157] Step 504: According to the preset scalp region sensitivity distribution map, different topological weights are assigned to electrode pairs with different distance intervals.
[0158] In step 504, the scalp regional sensitivity distribution map is a preset brain function partition weight map. For example, the weight of the left frontal pleasure area is set to 0.8, and the weight of the right frontal inhibition area is set to 0.6; the differentiated topological weight = basic weight × distance attenuation factor (1-0.05 × distance in cm). Relative distance refers to the actual physical distance between any two electrodes. Different distance intervals are grouped according to preset intervals (such as 0-2cm, 2-4cm, etc.) for the relative distance of all electrode pairs, which are used to differentiate the topological weights. An "electrode pair" (a combination of two adjacent electrodes) corresponds to a differentiated topological weight, which is determined by the relative distance between the two electrodes and the sensitivity of the scalp area in which they are located, rather than by independently assigning weights to each electrode. For example, in a virtual social scene, a 10-electrode array is used to cover the frontal area (Fp1 / Fp2 / F3 / F4 / F7 / F8 / Fz / FC1 / FC2 / AFz), forming 15 electrode pairs (e.g., Fp1-F3 spacing 4cm weight 0.5, F3-Fz spacing 5cm weight 0.4). The total number of electrodes is set according to the EEG acquisition standard (e.g., 16 frontal electrodes in the 10-20 system), and the electrode pairs are composed of adjacent electrodes with a spacing less than a preset threshold (e.g., 6cm) (e.g., Fp1-F3, Fp2-F4, etc.). The differentiated topological weight is the initial weight value at the electrode pair level, and the topological weight distribution ratio is the final weight used for parameter coupling after normalization; the former is the calculation basis of the latter, and the latter is the standardized output of the former.
[0159] In the embodiment of the present application, a Gaussian attenuation model is used to calculate the weight. With the high-sensitivity area electrode as the center, the weight decays exponentially with increasing distance. At the same time, when connecting across functional areas, the regional synergy coefficient is additionally multiplied.
[0160] Step 505: Generate spatial correlation parameters reflecting spatial distribution characteristics based on the relative distance relationship and the differentiated topological weights.
[0161] In an embodiment of the present application, cluster analysis is performed on the weighted distance matrix to identify high-density electrode clusters with higher signal strength than the surrounding areas, and the weighted signal strength mean of each cluster is calculated as the regional activity, and finally packaged into spatial correlation parameters.
[0162] Here's a specific example:
[0163] In the virtual social platform embodiment, the system detects that the user FP1 electrode signal presents a typical emotional fluctuation within 0.5 seconds: the signal rises from 35 μV to 50 μV in the first 0.2 seconds (rising phase, slope 0.75 μV / ms), then maintains at 48-50 μV for 0.2 seconds (plateau phase), and finally drops to 42 μV in the last 0.1 seconds (falling phase, slope 0.8 μV / ms). The rising-plateau transition steepness is 0.75 μV / ms, the plateau-falling transition steepness is 0.8 μV / ms, and the dynamic correlation parameter is calculated as (0.2 s x 0.75 + 0.1 s x 0.8) x 0.8 = 0.232 (the time weighting coefficient 0.8 comes from experimental calibration). In spatial processing, the topological weight of the FP1 electrode (pleasure area weight 0.8) and the F3 electrode (distance 3.2 cm) = 0.8 x (1-0.05 x 3.2) = 0.672, rounded to 0.7, combined with the FP1 signal intensity 45 μV to generate a spatial correlation parameter 0.7 x 45 = 31.5. The two are coupled by the formula (0.232 x 0.6 + 31.5 x 0.02 = 0.139 + 0.63 = 0.769, with the spatial parameter coefficient 0.02 as the normalization factor) to obtain the space-time parameter 0.769.
[0164] In the embodiments of the present application, through fine division of signal change stages and quantitative modeling of electrode spatial relationships, accurate extraction of emotional space-time features is realized, enabling the digital person to accurately identify the intensity changes and brain activation patterns of user emotions, and improving the accuracy and adaptability of virtual social interaction.
[0165] To more accurately generate interaction strategies that conform to the virtual social scene, in some embodiments, step 104: generating an emotional resonance interaction strategy that adapts to the current immersive virtual social scene based on the association mapping between the cross-modal emotional semantic features and the pre-stored virtual character behavior library, includes:
[0166] Step 601: extracting an emotional intensity parameter and an emotional type parameter from the cross-modal emotional semantic features.
[0167] In step 601, the emotional intensity parameter is a numerical value that quantifies the degree of emotional activation, ranging from 0 to 1, calculated from the ratio of the signal amplitude in the cross-modal features to the reference value. The emotional type parameter is a classification identifier, containing basic emotional categories such as joy, anger, etc., obtained by matching the feature with an emotional label data set.
[0168] In the embodiments of the present application, the intensity parameter 0.8 is extracted from the "high intensity left forehead joy" feature (FP1 electrode 45 μV / reference value 55 μV), and the type parameter is "joy" (matching the closest label pattern in the database).
[0169] Step 602: In the virtual character behavior library, match the action amplitude parameter set corresponding to the emotion intensity parameter and the expression parameter set corresponding to the emotion type parameter.
[0170] In step 602, the action amplitude parameter set contains digital human limb action parameters of different intensity levels, such as hand waving amplitude 0.3-0.9. Different action amplitude parameters in the set represent different action amplitude levels. The expression parameter set contains facial action unit combinations corresponding to each basic emotion, such as joy corresponding to mouth corner up 0.2-0.8.
[0171] In the embodiments of the present application, the action amplitude range [0.7, 0.9] corresponding to intensity 0.8 and the expression parameter range [smile 0.6, tooth smile 0.8, laugh 1.0] corresponding to "joy" type are queried in the behavior library.
[0172] Step 603: According to the current situation type of the immersive virtual social scene, filter a subset of action amplitude parameters that meet the spatial constraints of the immersive virtual social scene from the action amplitude parameter set.
[0173] In step 603, the current situation type refers to the interactive environment category detected in real time in the immersive virtual social scene, and the judgment basis is the real-time spatial position relationship and social behavior pattern of the user and the virtual character in the scene. The spatial constraint is the maximum action range allowed by the scene, which is determined by the virtual collision body size. The action amplitude parameter subset is the effective parameter range filtered from the complete action amplitude parameter set that meets the current virtual scene space limit. The subset is generated by comparing the action physical amplitude (such as hand waving radius) and the virtual environment collision body size (such as room wall distance), for example, in a 3m x 3m room, exclude actions with a larger than 0.8 amplitude that will cause the limb to penetrate the wall, and retain a safe amplitude range of 0.7-0.8. The boundary values of the subset are dynamically calculated and determined by the ratio of the scene space size to the length of the digital human model limb.
[0174] In the embodiments of the present application, in a 3m x 3m virtual room, actions with an amplitude ≤0.8 are filtered (to avoid penetrating the wall), and the subset [0.7, 0.8] is extracted from [0.7, 0.9].
[0175] Step 604: Combine and arrange the action amplitude parameters in the action amplitude parameter subset and the expression parameters in the expression parameter set in different forms to generate an interaction strategy set.
[0176] In step 604, the action amplitude parameter is a numerical value (0-1) quantifying the range of motion of the digital human limb, and 0.5 represents moderate amplitude. The expression parameter is a mixed weight value (0-1) controlling the movement of the digital human facial muscles, and 0.8 represents high intensity expression. The combination of different forms refers to the full permutation combination of action parameters and expression parameters, including the following two forms: intensity matching type: combining actions with similar amplitudes and expressions (such as amplitude 0.8 + expression 0.8), suitable for scenarios with high consistency requirements for emotional expression; compensation enhancement type: using higher amplitude actions to match weaker expressions (such as amplitude 0.8 + expression 0.6), suitable for long-distance interaction scenarios that require to strengthen the feedback of the limbs. The interaction strategy set is the Cartesian product of all compliant parameter combinations, and each combination constitutes a complete strategy. Each strategy in the interaction strategy set is a specific permutation combination of the action amplitude parameter and the expression parameter, and the permutation determines the matching relationship between the parameters (such as high intensity anger parameters can only be combined with specific angry expressions). By enumerating all valid parameter combinations to form the strategy set, it is ensured to cover the possible emotion-behavior mapping relationship in the virtual social scene. When the action amplitude parameter subset contains [large hand waving (amplitude value 0.8), slight nodding (amplitude value 0.3)], and the expression parameter set contains [angry frown (expression ID-A), happy smile (expression ID-B)]: 4 valid combination permutations are generated: strategy 1: large hand waving (0.8) + angry frown (A), strategy 2: large hand waving (0.8) + happy smile (B), strategy 3: slight nodding (0.3) + angry frown (A), strategy 4: slight nodding (0.3) + happy smile (B), according to the emotion type parameter (such as anger) to exclude logically conflicting combinations (strategy 2, strategy 4), and finally the interaction strategy set contains strategy 1 and strategy 3.
[0177] In the embodiments of the present application, the amplitudes [0.7, 0.8] are combined with the expressions [0.6, 0.8, 1.0], and 1.0 (out of limits) is excluded, to generate the strategy set {(0.7, 0.6), (0.7, 0.8), (0.8, 0.6), (0.8, 0.8)}.
[0178] Step 605: according to the interaction relationship coefficient between the emotion intensity parameter and the emotion type parameter, the strategies in the interaction strategy set are prioritized.
[0179] In step 605, the interaction relationship coefficient is the product of the matching degree of emotion intensity and type, reflecting the adaptability of the strategy.
[0180] In the embodiments of the present application, the coefficients of each strategy are calculated: amplitude 0.8 x expression 0.8 = 0.64 (optimal), 0.7 x 0.8 = 0.56, 0.8 x 0.6 = 0.48, 0.7 x 0.6 = 0.42, and the strategies are prioritized in descending order.
[0181] Step 606: From the priority-sorted interaction strategy set, select the optimal strategy with the highest priority as the emotional resonance interaction strategy.
[0182] In step 606, the optimal strategy is the compliance strategy with the highest coefficient, which directly determines the form of digital human feedback.
[0183] In the embodiment of the present application, the (0.8, 0.8) strategy is selected, that is, a 0.8 amplitude waving is combined with a 0.8 toothy smile.
[0184] Here's a specific example:
[0185] In an embodiment of a virtual social platform, the system extracts the emotion intensity parameter 0.8 (FP1 electrode average intensity 45 μV / happiness baseline 55 μV) and the emotion type parameter "happiness" from the "high-intensity left forehead happiness" feature. The system then queries the virtual character behavior library to obtain: 1) the range of motion amplitudes corresponding to the intensity of 0.8 (0.7, 0.9) (calculated as: base amplitude 0.5 + 0.3 × intensity parameter); and 2) the set of expression parameters corresponding to the "happiness" type (smile 0.6, grin 0.8, laugh 1.0). Based on the spatial constraints of the current 3m x 3m virtual room (maximum allowable waving amplitude 0.8), the motion amplitude subset [0.7, 0.8] was selected. The subset parameters were combined with the expression parameters to generate four strategies: (0.7, 0.6), (0.7, 0.8), (0.8, 0.6), and (0.8, 0.8). The interaction coefficients (amplitude × expression matching) for each strategy were calculated as 0.42, 0.56, 0.48, and 0.64, respectively. The optimal strategy (0.8, 0.8) is selected to drive the digital human to perform a 0.8-amplitude wave (shoulder joint rotation 55 degrees = basic 30 degrees + 0.8 × 30 degrees) and a 0.8-amplitude grin (mouth corners raised 0.8 = basic 0.5 + 0.3 × intensity parameter), and the voice tone is improved by 20% (basic frequency × (1 + 0.25 × intensity parameter)).
[0186] In the embodiment of the present application, through the two-dimensional analysis of emotional parameters and the screening of scene adaptive strategies, the digital human feedback is accurately matched with the user's emotional state and the constraints of the virtual environment, thereby improving the personalization and scene adaptability of virtual social interaction.
[0187] In order to more accurately achieve collaborative control of the digital human's multimodal feedback, in some embodiments, step 105: based on the emotional resonance interaction strategy, driving the AI digital human's facial micro-expressions, body movements, and speech prosody parameters to generate a dynamic feedback loop that evolves synchronously with the user's emotional state, includes:
[0188] Step 701: Parse the emotional resonance interaction strategy into a facial control instruction set, a body control instruction set, and a voice control instruction set.
[0189] In step 701, the facial control instruction set includes an expression intensity value (0-1) and a duration parameter (seconds). The limb control instruction set includes a joint angle change rate (degrees / second) and a movement amplitude value (0-1). The voice control instruction set includes a pitch modulation coefficient (±20%) and a speech rate adjustment factor (0.8-1.2).
[0190] In the embodiment of the present application, the strategy (0.8 waving, 0.8 grinning) is parsed into: facial instructions {expression intensity 0.8, duration 1.2 seconds}, body instructions {shoulder joint change rate 90 degrees / second, amplitude 0.8}, voice instructions {pitch +15%, speaking speed 1.1 times}.
[0191] Step 702: Generate a motion trajectory sequence corresponding to the facial muscle group according to the expression intensity value and expression duration parameter in the facial control instruction set.
[0192] In step 702, the expression intensity value is a numerical value (0-1) that quantifies the expressiveness of facial expressions, with 0.8 representing a high-intensity expression. This value is derived from a nonlinear mapping of the emotion intensity parameter (intensity parameter 0.8×1.25-0.2, with the coefficient calibrated experimentally). The expression duration parameter is the number of seconds that the expression lasts, determined by the emotional fluctuation period (e.g., a 0.3-second fluctuation of joy corresponds to a 1.2-second duration, formula: period×4). Together, these two parameters constitute the spatiotemporal control benchmark for facial movements. The motion trajectory sequence is a time-varying parameter curve of facial muscle movement, generated by keyframe interpolation. Facial muscle groups refer to the collection of virtual muscles in the digital human face model that simulate human expression muscles, including: the frontalis muscle group (controls eyebrow raising); the orbicularis oculi muscle group (controls eye closure / squinting); the zygomaticus major muscle group (controls the upward movement of the corners of the mouth); and the orbicularis oris muscle group (controls lip movement). Each muscle group corresponds to an independent weight channel, and expression control is achieved through weight blending.
[0193] In the embodiment of the present application, the target position of the corner of the mouth raised is determined to be 8mm (reference value 5mm×1.6) based on the expression intensity of 0.8, and a smooth transition from the static state to the target position is performed according to an S-shaped curve within 1.2 seconds, generating a trajectory sequence containing 20 key frames.
[0194] Step 703: Calculate the real-time motion path of each limb joint based on the joint angle change rate and motion amplitude value in the limb control instruction set.
[0195] In step 703, the rate of change of the joint angle is the limb movement speed (degrees / second), which is linearly calculated from the motion amplitude value (amplitude 0.8 to 90 degrees / second = base 60 degrees / second + 0.8×30). The motion amplitude value is the normalized range of motion (0-1), which is directly taken from the amplitude parameter in the interaction strategy, and 0.8 represents 80% of the maximum range of motion (such as a waving amplitude of 0.8 = shoulder joint rotation 60 degrees / maximum 75 degrees). Limb joints specifically refer to the movable connection parts in the AI digital human model used to achieve limb movement, including but not limited to the following anatomical corresponding parts: upper limb joint group: shoulder joint (connecting the torso and upper arm), elbow joint (connecting the upper arm and forearm), wrist joint (connecting the forearm and palm); lower limb joint group: hip joint (connecting the torso and thigh), knee joint (connecting the thigh and calf), ankle joint (connecting the calf and foot). Spinal joint group: cervical joint (head movement), lumbar joint (trunk flexion). A real-time motion path is a continuous sequence of poses in the joint space that satisfies the kinematic constraints.
[0196] In an embodiment of the present application, the shoulder joint is used as the origin, the hand end point coordinates (x=0.5m, y=0.3m) are calculated according to an amplitude of 0.8, the joint angle is solved by inverse kinematics, and the joint angle data is generated every 0.01 seconds at a speed of 90 degrees / second.
[0197] Step 704: Generate a corresponding fundamental frequency curve and syllable duration sequence according to the pitch modulation coefficient and speech rate adjustment factor in the voice control instruction set.
[0198] In step 704, the pitch modulation coefficient is the percentage of the fundamental frequency change (e.g., +15%), which is determined by the emotion intensity parameter × 0.2 (0.8 × 0.2 = 16%, rounded up to 15%). The speech rate adjustment factor is a scaling factor (1.1 times) of the syllable duration, determined by the emotion type (happy type speeds up by 10%, angry type slows down by 20%), and then weighted with the intensity parameter (0.8 × 10% = 8%, rounded up to 1.1). Together, they shape the emotional speech characteristics. The fundamental frequency curve is the envelope of the speech frequency over time, and the syllable duration sequence is the adjustment scheme for the pronunciation time of each phoneme.
[0199] In an embodiment of the present application, the fundamental frequency of the basic speech "hello" is increased by 15% from 220Hz to 253Hz, the duration of the syllable "ni" is adjusted from 0.3 seconds to 0.33 seconds (1.1 times), and the duration of "hao" is adjusted from 0.4 seconds to 0.36 seconds (0.9 times), generating a pleasant intonation with ups and downs.
[0200] Step 705: Synchronize the motion trajectory sequence, the real-time motion path, the fundamental frequency curve and the syllable duration sequence to generate a multimodal output control signal of the AI digital human.
[0201] In step 705, the multi-modal output control signal is a set of coordinated control instructions under a unified timeline.
[0202] In the embodiment of the present application, the mouth corner movement trajectory (peak at 0.6 seconds), hand waving path (highest point at 0.5 seconds), and speech fundamental frequency peak (0.55 seconds) are aligned to the same timeline, ensuring that the smile, hand wave, and tone climax appear synchronously.
[0203] Step 706: dynamically correct the multi-modal output control signal according to the user's real-time updated cross-modal emotional semantic features to form a dynamic feedback loop.
[0204] In step 706, dynamic correction is the process of adjusting the control signal according to real-time emotional changes.
[0205] In the embodiment of the present application, when the user's FP1 signal strength is detected to rise from 45 μV to 50 μV (intensity parameter 0.9), the expression intensity is increased from 0.8 to 0.9 (adjusting the trajectory endpoint to 9 mm), the hand waving amplitude is increased to 0.85 (endpoint coordinates x = 0.53 m), and the pitch is increased to +18%, forming a gradual reinforcement feedback.
[0206] The following is a specific example:
[0207] In the virtual social platform embodiment, the system parses the "wave and smile" strategy as: facial instructions {expression intensity 0.8 (0.7 base value + 0.1 x 0.8), duration 1.2 seconds (0.3 seconds x 4 emotion period)}, body instructions {shoulder joint change rate 84 degrees / second (60 + 0.8 x 30), amplitude 0.7}, and voice instructions {pitch +14% (0.8 x 0.2 x 90 reference frequency), speech rate 1.08 times (base 1.0 + 0.8 x 0.1)}. Generate mouth corner movement trajectory (0-8mm up, S-curve transition 1.2 seconds) and hand waving path (shoulder joint 0-52 degrees rotation, elbow joint 0-36 degrees flexion). Speech synthesis raises the fundamental frequency of "hello" from 220 Hz to 251 Hz (220 x 1.14), and adjusts the syllable duration to "you" 0.32 seconds (original 0.3 x 1.08), "good" 0.43 seconds (original 0.4 x 1.08). When the FP1 signal rises from 45 μV to 50 μV (intensity parameter 0.85), dynamically adjust the expression intensity to 0.85 (8.5mm mouth corner displacement), hand waving amplitude 0.75 (shoulder joint 56 degrees), pitch +15% (253Hz), forming a gradually reinforced feedback loop.
[0208] In the embodiment of the present application, through the refined analysis and dynamic collaborative control of multimodal instructions, millisecond-level synchronization of the digital human's feedback behavior and the user's emotional fluctuations is achieved, so that the virtual social interaction presents a real and natural emotional resonance effect, enhancing the immersion and emotional authenticity of the interactive experience.
[0209] Figure 2 A schematic diagram of the structure of an AI digital human interaction system based on emotion recognition provided in an embodiment of the present application is shown in FIG. Figure 2 As shown, the system includes:
[0210] The acquisition module 21 is used to acquire multi-channel electrophysiological signals from the user's prefrontal cortex area in an immersive virtual social scene.
[0211] The pattern matching module 22 is used to perform pattern matching on the multi-channel electrophysiological signal and a preset emotion tag data set to generate a multi-dimensional emotion vector for reflecting the user's emotional state.
[0212] The first generating module 23 is configured to perform a joint analysis of the temporal dynamic characteristics and the scalp spatial topological relationship of the multi-dimensional emotion vector through parallel processing channels to generate cross-modal emotion semantic features.
[0213] The second generating module 24 is used to generate an emotional resonance interaction strategy adapted to the current immersive virtual social scene based on the association mapping between the cross-modal emotional semantic features and the pre-stored virtual character behavior library.
[0214] The third generation module 25 is used to drive the facial micro-expressions, body movements and voice rhythm parameters of the AI digital human based on the emotional resonance interaction strategy to generate a dynamic feedback loop that evolves synchronously with the user's emotional state.
[0215] Figure 2 The AI digital human interaction system based on emotion recognition can perform Figure 1 The implementation principles and technical effects of the AI digital human interaction method based on emotion recognition described in the illustrated embodiment are not further elaborated. The specific manner in which each module and unit performs operations in the AI digital human interaction system based on emotion recognition in the above embodiment has been described in detail in the embodiment of the method and will not be elaborated on here.
[0216] In one possible design, Figure 2 An AI digital human interaction system based on emotion recognition in the embodiment shown can be implemented as a computing device, such as Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;
[0217] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are called and executed by the processing component 32.
[0218] The processing component 32 described above Figure 1 An AI digital human interaction method based on emotion recognition of the embodiment.
[0219] The processing component 32 can include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component can also be one or more Application-Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Process Device (DSPD), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components, for executing the above method.
[0220] The storage component 31 is configured to store various types of data to support the operation of the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or their combination, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0221] Of course, the computing device can also include other components, such as an input / output interface, a display component, a communication component, etc.
[0222] The input / output interface provides an interface between the processing component and the peripheral interface module, which can be an output device, an input device, etc.
[0223] The communication component is configured to facilitate, among other things, wired or wireless communications between the computing device and other devices.
[0224] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. In this case, the computing device can refer to a cloud server, and the above-mentioned processing components, storage components, etc. can be basic server resources rented or purchased from the cloud computing platform.
[0225] The present application also provides a computer storage medium storing a computer program, wherein the computer program can achieve the above-mentioned Figure 1 The illustrated embodiment is an AI digital human interaction method based on emotion recognition.
[0226] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0227] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.
[0228] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0229] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. An AI digital human interaction method based on emotion recognition, characterized in that: include: Acquire multi-channel electrophysiological signals from the user's prefrontal cortex in an immersive virtual social environment. Pattern matching the multi-channel electrophysiological signal with a preset emotion tag dataset to generate a multi-dimensional emotion vector reflecting the user's emotional state; Performing a joint analysis of the temporal dynamic characteristics and scalp spatial topological relationship of the multi-dimensional emotion vector through parallel processing channels to generate cross-modal emotion semantic features; Generate an emotional resonance interaction strategy adapted to the current immersive virtual social scene based on the association mapping between the cross-modal emotional semantic features and the pre-stored virtual character behavior library; Based on the emotional resonance interaction strategy, the facial micro-expressions, body movements, and speech prosody parameters of the AI digital human are driven to generate a dynamic feedback loop that evolves synchronously with the user's emotional state; The method of performing a joint analysis of the temporal dynamic characteristics and the scalp spatial topological relationship of the multi-dimensional emotion vector through parallel processing channels to generate cross-modal emotion semantic features includes: Decomposing the multidimensional emotion vector into a time component sequence for representing temporal dynamic characteristics and a set of spatial components for representing scalp spatial topological relationships; In a parallel processing channel, generating dynamic correlation parameters corresponding to the time component sequence and spatial correlation parameters corresponding to the spatial component set; Nonlinearly coupling the dynamic correlation parameters with the spatial correlation parameters to generate spatiotemporal joint analytical parameters; Based on the spatiotemporal joint analysis parameters, signal segments that simultaneously satisfy temporal fluctuation consistency and spatial distribution correlation are screened out from the multidimensional emotion vector; Performing spatiotemporal reorganization on the signal segments to generate cross-modal emotional semantic features; The nonlinearly coupling the dynamic correlation parameter and the spatial correlation parameter to generate a spatiotemporal joint analysis parameter includes: Determining the spatial correlation parameter according to the steepness of the signal amplitude change trend reflected by the dynamic correlation parameter, wherein the spatial correlation parameter includes an electrode position weight and a topological weight distribution ratio; Based on the topological weight distribution ratio, the signal amplitude change trend and the electrode position weight are superimposed layer by layer to generate a multi-level parameter set; Dynamically intercepting the hierarchical parameters that meet the preset weight increase conditions in the multi-level parameter set to generate a parameter subset; Calculating the interaction intensity coefficient between the dynamic correlation parameter and the spatial correlation parameter according to the overlapping coverage of the time fluctuation sensitive area and the spatial distribution dense area in the parameter subset; Based on the interaction strength coefficient, weighted aggregation is performed on the hierarchical parameters in the parameter subset to generate spatiotemporal joint analysis parameters.
2. The method according to claim 1, characterized in that The signal amplitude change trend and the electrode position weight are superimposed step by step based on the topological weight distribution ratio to generate a multi-level parameter set, including: Dividing the signal amplitude change trend into a plurality of continuous levels according to the numerical interval of the topological weight distribution ratio; In each level, the trend intensity factor corresponding to the signal amplitude change trend and the electrode position weight are scaled respectively according to the scaling coefficient corresponding to the level; The scaled trend intensity factor is superimposed with the scaled electrode position weight to generate independent superposition parameter units corresponding to each level; The superposition parameter units are combined in a preset hierarchical order to generate a multi-level parameter set.
3. The method according to claim 1, characterized in that Generating, in the parallel processing channel, the dynamic correlation parameters corresponding to the time component sequence and the spatial correlation parameters corresponding to the spatial component set includes: In the parallel processing channel, the following process is executed: the signal amplitude change trend in the time component sequence is divided into a rising phase, a stable phase and a falling phase, and the duration of each phase and the transition steepness between adjacent phases are recorded; generating a dynamic correlation parameter for reflecting the steepness of the signal amplitude change trend based on the duration and the conversion steepness; determining a relative distance relationship between the electrodes according to actual distribution positions of the multi-channel electrodes in the spatial component set in the prefrontal cortex; According to the preset scalp area sensitivity distribution map, differentiated topological weights are assigned to electrode pairs with different distance intervals; Based on the relative distance relationship and the differentiated topological weight, a spatial association parameter reflecting spatial distribution characteristics is generated.
4. The method according to claim 1, wherein The generating of an emotional resonance interaction strategy adapted to the current immersive virtual social scene based on the association mapping between the cross-modal emotional semantic features and the pre-stored virtual character behavior library includes: Extracting emotion intensity parameters and emotion type parameters from the cross-modal emotion semantic features; In the virtual character behavior library, matching the action amplitude parameter set corresponding to the emotion intensity parameter and the expression parameter set corresponding to the emotion type parameter; According to the current situation type of the immersive virtual social scene, a subset of motion amplitude parameters that meets the spatial constraints of the immersive virtual social scene is selected from the motion amplitude parameter set; Combining and arranging the motion amplitude parameters in the motion amplitude parameter subset with the expression parameters in the expression parameter set in different forms to generate an interaction strategy set; Prioritizing the strategies in the interaction strategy set according to the interaction relationship coefficient between the emotion intensity parameter and the emotion type parameter; From the priority-sorted interaction strategy set, the optimal strategy with the highest priority is selected as the emotional resonance interaction strategy.
5. The method according to claim 1, wherein The emotional resonance interaction strategy drives the facial micro-expressions, body movements, and speech prosody parameters of the AI digital human to generate a dynamic feedback loop that evolves synchronously with the user's emotional state, including: Parsing the emotional resonance interaction strategy into a facial control instruction set, a body control instruction set, and a voice control instruction set; generating a motion trajectory sequence of corresponding facial muscle groups according to the expression intensity value and expression duration parameter in the facial control instruction set; Calculating the real-time motion path of each limb joint based on the joint angle change rate and motion amplitude value in the limb control instruction set; Generate a corresponding fundamental frequency curve and syllable duration sequence according to the pitch modulation coefficient and speech rate adjustment factor in the voice control instruction set; Performing time-series synchronization on the motion trajectory sequence, the real-time motion path, the fundamental frequency curve, and the syllable duration sequence to generate a multimodal output control signal for the AI digital human; The multimodal output control signal is dynamically modified according to the cross-modal emotional semantic features updated by the user in real time to form a dynamic feedback loop.
6. An AI digital human interaction system based on emotion recognition, characterized in that: include: An acquisition module, used to obtain multi-channel electrophysiological signals from the user's prefrontal cortex area in an immersive virtual social scene; a pattern matching module, configured to perform pattern matching on the multi-channel electrophysiological signal and a preset emotion tag dataset to generate a multi-dimensional emotion vector reflecting the user's emotional state; A first generation module is configured to perform a joint analysis of the temporal dynamic characteristics and the scalp spatial topological relationship of the multi-dimensional emotion vector through a parallel processing channel to generate a cross-modal emotion semantic feature; A second generation module is used to generate an emotional resonance interaction strategy adapted to the current immersive virtual social scene based on the association mapping between the cross-modal emotional semantic features and the pre-stored virtual character behavior library; The third generation module is used to drive the facial micro-expressions, body movements and voice prosody parameters of the AI digital human based on the emotional resonance interaction strategy to generate a dynamic feedback loop that evolves synchronously with the user's emotional state; The method of performing a joint analysis of the temporal dynamic characteristics and the scalp spatial topological relationship of the multi-dimensional emotion vector through parallel processing channels to generate cross-modal emotion semantic features includes: Decomposing the multidimensional emotion vector into a time component sequence for representing temporal dynamic characteristics and a set of spatial components for representing scalp spatial topological relationships; In a parallel processing channel, generating dynamic correlation parameters corresponding to the time component sequence and spatial correlation parameters corresponding to the spatial component set; Nonlinearly coupling the dynamic correlation parameters with the spatial correlation parameters to generate spatiotemporal joint analytical parameters; Based on the spatiotemporal joint analysis parameters, signal segments that simultaneously satisfy temporal fluctuation consistency and spatial distribution correlation are screened out from the multidimensional emotion vector; Performing spatiotemporal reorganization on the signal segments to generate cross-modal emotional semantic features; The nonlinearly coupling the dynamic correlation parameter and the spatial correlation parameter to generate a spatiotemporal joint analysis parameter includes: Determining the spatial correlation parameter according to the steepness of the signal amplitude change trend reflected by the dynamic correlation parameter, wherein the spatial correlation parameter includes an electrode position weight and a topological weight distribution ratio; Based on the topological weight distribution ratio, the signal amplitude change trend and the electrode position weight are superimposed layer by layer to generate a multi-level parameter set; Dynamically intercepting the hierarchical parameters that meet the preset weight increase conditions in the multi-level parameter set to generate a parameter subset; Calculating the interaction intensity coefficient between the dynamic correlation parameter and the spatial correlation parameter according to the overlapping coverage of the time fluctuation sensitive area and the spatial distribution dense area in the parameter subset; Based on the interaction strength coefficient, weighted aggregation is performed on the hierarchical parameters in the parameter subset to generate spatiotemporal joint analysis parameters.
7. A computing device, characterized in that It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement an AI digital human interaction method based on emotion recognition as described in any one of claims 1 to 5.
8. A computer storage medium, characterized in that A computer program is stored, and when the computer program is executed by a computer, the AI digital human interaction method based on emotion recognition as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Virtual human emotion interaction system and method based on five-dimensional emotion model
CN115543089A
Self-supervised emotion recognition method based on electroencephalogram signals in virtual reality scene
CN119293556A