Multi-modal interaction method and system of digital human intelligent agent

By building a multimodal interactive system of digital human intelligent bodies and utilizing the fusion analysis of facial images and voice signals to generate emotional holographic feature maps and voice-emotion linkage mapping spectra, we achieve a deep understanding of user emotions and personalized response, solving the integration problems of existing systems and improving the efficiency and naturalness of interaction.

CN120653118AInactive Publication Date: 2025-09-16GUANGDONG HUITONG INFORMATION TECH CO LTD
View PDF 0 Cites 9 Cited by

Patent Information

Application Number
CN202510817901.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-18
Publication Date
2025-09-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing multimodal interaction systems have difficulty effectively integrating image and voice signals and lack adaptive capabilities, resulting in low interaction efficiency and large understanding errors. They are unable to cope with complex and changing interaction scenarios and lack deep semantic understanding and emotional perception capabilities.

Method used

By acquiring the user's real-time facial images and voice signals, performing micro-expression recognition and deep emotion analysis, building a holographic feature map of user emotions, performing voice-emotion correlation analysis, and combining eye gaze tracking to generate interaction depth intention signals, dynamically adjusting interaction strategies and voice styles, building a personalized interaction adaptation model, and performing multimodal interaction feedback optimization.

Benefits of technology

It achieves synchronous processing of visual and auditory information, quickly responds to emotional changes, improves the naturalness and accuracy of interaction, adapts to interaction needs in various situations, and enhances the user experience quality and the system's adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120653118A_ABST
    Figure CN120653118A_ABST
Patent Text Reader

Abstract

The invention relates to the field of multi-modal interaction analysis, in particular to a multi-modal interaction method and system of a digital human agent. The method comprises the following steps: acquiring a real-time face image and a voice signal input stream of an interactive user based on an intelligent agent; performing real-time micro-expression recognition and deep emotion analysis based on the real-time facial image to obtain real-time emotion features of the user; performing time sequence evolution analysis on the real-time emotion characteristics of the user, performing holographic user emotion deep mining, and constructing a user emotion holographic characteristic spectrum; carrying out adaptive acoustic gain processing on the voice signal input stream, and carrying out voice-emotion association analysis based on the user emotion holographic characteristic spectrum to generate a voice-emotion linkage mapping spectrum; and carrying out eyeball fixation point migration tracking based on the user emotion holographic feature map and the real-time face image, and generating a user interaction depth intention signal. Through the real-time deep semantic understanding and emotion perception ability, the intelligent agent interaction intelligence and response accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimodal interaction analysis, and in particular to a multimodal interaction method and system for a digital human intelligent body. Background Art

[0002] With the rapid development of artificial intelligence (AI), digital human agents, as a key vehicle for human-computer interaction, are gaining widespread application in numerous fields, including education, healthcare, finance, government affairs, and virtual customer service. Driven in particular by breakthroughs in natural language processing, computer vision, and speech recognition, digital human agents have evolved from initial text-based dialogue systems to advanced interactive entities with multimodal capabilities, including visual recognition, speech perception, and emotion understanding. Digital humans demonstrate unprecedented potential for improving user experience and enhancing service intelligence. Traditional human-computer interaction methods typically rely on single-modal input, such as keyboard input, speech recognition, or gesture recognition. However, in practical applications, these single-modal approaches often fail to fully perceive user intent, leading to low interaction efficiency, large misunderstandings, and unnatural responses. As the core of a new generation of interactive systems, digital human agents must possess both image and speech recognition capabilities, acquiring user information through both visual and auditory channels to achieve a more natural, intelligent, and humanized interactive experience.

[0003] Current image recognition-based technologies are capable of performing functions such as facial recognition, expression recognition, and gaze tracking, while speech recognition technology has also made significant progress in speech transcription, semantic understanding, and sentiment analysis. However, integrating these two types of information into a unified interaction framework to support multimodal understanding and response for digital human agents still faces numerous challenges. Firstly, the temporal synchronization and semantic fusion of image and speech signals are complex, requiring an effective multimodal fusion mechanism for unified modeling. Secondly, current multimodal interaction systems often lack dynamic adaptability, making it difficult to adjust interaction strategies in real time based on user behavior, which affects the accuracy and naturalness of the agent's responses. Furthermore, existing multimodal interaction methods mostly rely on predefined rules or static models, lacking deep semantic understanding and emotion perception, making them incapable of coping with complex and ever-changing interaction scenarios. Furthermore, the robustness and intelligence of these systems also need to be improved in the face of real-world challenges such as multi-user interaction, noise interference, and incomplete input. Therefore, developing a multimodal interaction method for digital human agents with adaptive learning capabilities that can integrate image and speech recognition information has become a key direction in the development of human-computer interaction technology. Summary of the Invention

[0004] In order to solve the above technical problems, the present invention proposes a multimodal interaction method and system for a digital human intelligent body to solve at least one of the above technical problems.

[0005] To achieve the above-mentioned object, the present invention provides a multimodal interaction method of a digital human agent, comprising the following steps: Step S1: obtaining a real-time facial image and voice signal input stream of an interactive user based on an intelligent agent; performing real-time micro-expression recognition and deep emotion analysis based on the real-time facial image to obtain the user's real-time emotional characteristics; Step S2: Analyze the temporal evolution of the user's real-time emotional characteristics, conduct in-depth mining of holographic user emotions, and construct a holographic feature map of user emotions; Step S3: Adaptively perform acoustic gain processing on the voice signal input stream, and perform voice-emotion correlation analysis based on the user's emotional holographic feature map to generate a voice-emotion linkage mapping spectrum; Step S4: tracking eye gaze migration based on the user's emotional holographic feature map and real-time facial image, and performing user attention state evolution to generate a user interaction depth intention signal; Step S5: Based on the voice-emotion linkage mapping spectrum and the user's interaction depth intention signal, interactive response strategy decisions and interactive voice style dynamic adjustments are made to build a user personalized interaction adaptation model; Step S6: Perform real-time user interaction according to the user personalized interaction adaptation model, perform multimodal interaction feedback optimization, and build a multimodal interaction optimization model.

[0006] In this specification, a multimodal interaction system of a digital human agent is provided, which is used to execute the multimodal interaction method of a digital human agent as described above, including: A micro-expression recognition module is used to obtain real-time facial images and voice signal input streams of interactive users based on the intelligent agent; perform real-time micro-expression recognition and deep emotion analysis based on the real-time facial images to obtain real-time emotional characteristics of the user; The emotion deep mining module is used to analyze the time series evolution of users' real-time emotion characteristics, conduct holographic user emotion deep mining, and construct a holographic feature map of user emotions; The speech analysis module is used to perform adaptive acoustic gain processing on the speech signal input stream, and perform speech-emotion correlation analysis based on the user's emotional holographic feature map to generate a speech-emotion linkage mapping spectrum; The gaze tracking module is used to track the migration of eye gaze points based on the user's emotional holographic feature map and real-time facial images, and to monitor the evolution of the user's attention state to generate a signal of the user's interaction depth intention; The interaction adaptation module is used to make interactive response strategy decisions and dynamically adjust the interactive voice style based on the voice-emotion linkage mapping spectrum and the user's interaction depth intention signal, thereby building a user-personalized interaction adaptation model; The interaction feedback optimization module is used to perform real-time user interaction based on the user's personalized interaction adaptation model, and to perform multimodal interaction feedback optimization to build a multimodal interaction optimization model.

[0007] The beneficial effects of the present invention include: simultaneous acquisition of facial images and voice signals enables synchronous processing of visual and auditory information. At the visual level, micro-expression recognition technology is integrated to effectively capture subtle emotional responses, such as brief frowns and mouth twitches. At the audio level, features such as voice intonation, sound pressure, and frequency changes are analyzed to assist in emotion judgment, addressing the problem of single-modality misjudgment. The system can integrate the user's current facial expression and voice intonation in real time, rapidly responding to emotional changes. This system is suitable for dialogue systems with short conversation turns and frequent emotional changes. By analyzing the evolution of a user's emotions across multiple moments using time series modeling (such as LSTM and Transformer models), the system not only identifies the current state but also infers its evolving trends, such as the "calm → confused → anxious" evolution chain, assisting the system in determining the next interaction strategy. Low-level emotional signals (such as facial muscle changes and voice intonation fluctuations) are abstracted into high-level emotional states (such as trust, satisfaction, and dislike), enabling the system to understand the semantics of emotions. The emotional hologram is a multidimensional feature vector space that records typical emotional response patterns of users in different situations. This accumulates individual characteristic data for long-term interactions, enhancing the system's "memory" and continuous adaptability. Adaptive acoustic gain processing compensates for factors such as ambient noise and user volume differences in real time, ensuring speech clarity while avoiding distortion of emotional characteristics (such as rapidity, low pitch, and trembling) in the speech. Acoustic features (such as fundamental frequency F0, formants, and velocity) are correlated with the user's emotional state in the speech segment to generate a "speech-emotion linkage mapping spectrum," which can be used to predict the user's current potential emotional intent in real time. High-precision eye tracking technology analyzes user gaze trajectory, gaze duration, and skipping frequency to determine whether the user is currently paying attention, distracted, or interested in specific interface elements. Combining gaze behavior with emotional state, the system can infer whether the user wishes to continue the current topic, wishes to change content, or is dissatisfied with the system's response, thereby enabling deep intention inference. Not only does it tailor responses to user emotions, but it also dynamically adjusts speech speed, pitch, and tone within the voice output, creating more emotionally relevant responses (e.g., responding in a gentler tone when a user is feeling down). This personalized response style enhances user familiarity and trust in the digital human, effectively improving the quality of the interactive experience and avoiding the mechanical, jarring feel of traditional dialogue systems. The system records user feedback (e.g., satisfied facial expressions, affirmative voice tone) during each interaction, and combines this with emotional profiles and intention signals to continuously fine-tune the interaction model. By deeply learning multimodal feedback signals (facial, voice, and behavioral), the model gradually learns each user's communication preferences, rhythm, and emotional responses, enhancing user retention. This multimodal optimization model can adapt to interaction needs in a variety of contexts, from customer service to virtual assistants to educational companionship, forming the foundational capabilities of a widely applicable intelligent interaction platform. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 A schematic flow chart of the steps of a multimodal interaction method of a digital human agent according to the present invention; Figure 2 Detailed implementation flow chart of step S1; Figure 3 Detailed implementation flow chart of step S2; Figure 4 Schematic diagram of the detailed implementation steps of step S3. DETAILED DESCRIPTION

[0009] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0010] This application provides a multimodal interaction method and system for a digital human agent. The execution entities of this multimodal interaction method and system include, but are not limited to, the following: mechanical equipment, data processing platforms, cloud server nodes, network upload devices, etc., which can be considered as general computing nodes of this application. The data processing platform includes, but is not limited to, at least one of an audio and image management system, an information management system, and a cloud data management system.

[0011] See also Figures 1 to 4 The present invention provides a multimodal interaction method for a digital human agent, the multimodal interaction method for a digital human agent comprising the following steps: Step S1: obtaining a real-time facial image and voice signal input stream of an interactive user based on an intelligent agent; performing real-time micro-expression recognition and deep emotion analysis based on the real-time facial image to obtain the user's real-time emotional characteristics; Step S2: Analyze the temporal evolution of the user's real-time emotional characteristics, conduct in-depth mining of holographic user emotions, and construct a holographic feature map of user emotions; Step S3: Adaptively perform acoustic gain processing on the voice signal input stream, and perform voice-emotion correlation analysis based on the user's emotional holographic feature map to generate a voice-emotion linkage mapping spectrum; Step S4: tracking eye gaze migration based on the user's emotional holographic feature map and real-time facial image, and performing user attention state evolution to generate a user interaction depth intention signal; Step S5: Based on the voice-emotion linkage mapping spectrum and the user's interaction depth intention signal, interactive response strategy decisions and interactive voice style dynamic adjustments are made to build a user personalized interaction adaptation model; Step S6: Perform real-time user interaction according to the user personalized interaction adaptation model, perform multimodal interaction feedback optimization, and build a multimodal interaction optimization model.

[0012] In the embodiment of the present invention, see Figure 1 , is a flowchart of the steps of a multimodal interaction method of a digital human agent of the present invention. In this example, the steps of the method include: Step S1: obtaining a real-time facial image and voice signal input stream of an interactive user based on an intelligent agent; performing real-time micro-expression recognition and deep emotion analysis based on the real-time facial image to obtain the user's real-time emotional characteristics; In this embodiment, a high-resolution camera is used to capture the user's facial image in real time. The camera should be placed in a suitable position to ensure that the user's face is fully captured and to minimize shadows caused by direct sunlight. Furthermore, the system should have good light adaptability to cope with varying ambient lighting conditions. The image processing module captures and transmits the user's facial image in real time for subsequent analysis. Simultaneously, a high-quality microphone is used to capture the user's voice signal. These signals are converted into digital signals and transmitted to the processing system in real time. During this process, the microphone's sensitivity must be moderate to clearly capture the user's voice while filtering out background noise. To this end, a noise suppression algorithm can be used to enhance the clarity of the voice signal. Once the real-time facial image and voice signal are acquired, the next step is to perform real-time micro-expression recognition and deep emotion analysis. The goal of this process is to extract the user's emotional information by analyzing facial expression changes and voice features. Experimental parameters are set, for example, the micro-expression recognition time window is set to 200 milliseconds to capture rapidly changing facial features. When implementing micro-expression recognition, the real-time facial image must first be processed using a facial landmark detection algorithm (such as Dlib or OpenCV) to identify key facial landmarks. The movement of these feature points is used to analyze changes in facial expressions. For example, the frequency of eye blinking and the upward or downward movement of the corners of the mouth can indicate the user's emotional state. Using machine learning models (such as support vector machines or convolutional neural networks), the extracted features are classified to identify the user's micro-expressions. This is then followed by in-depth emotion analysis. By matching the micro-expression recognition results with known emotion models, the user's micro-expressions can be translated into specific emotional states. Experimental parameters are set, such as classifying emotions into basic categories such as happiness, sadness, anger, and surprise. Using an emotion recognition algorithm, the extracted micro-expression features are compared with a library of emotion features to obtain the user's real-time emotional signature.

[0013] Step S2: Analyze the temporal evolution of the user's real-time emotional characteristics, conduct in-depth mining of holographic user emotions, and construct a holographic feature map of user emotions; In this embodiment, the emotional feature data extracted from users during real-time interactions is time-series processed. Each emotional feature (such as happiness, sadness, anger, etc.) is marked on the timeline with its time of occurrence and intensity. To this end, a sliding window technique can be used to update and record the emotional features every 1 second, forming a continuous emotional time series. This process provides basic data for subsequent emotional analysis. The temporal evolution of emotions is analyzed. The Dynamic Time Warping (DTW) algorithm can be used to analyze the changing patterns between different emotional states and identify emotional fluctuation trends and transition moments. Specifically, by calculating the similarity of each emotional feature in the time series, the user's emotional fluctuations and sustained states can be identified. This helps understand the emotional ups and downs of the user during the interaction. Once the emotional temporal evolution analysis is completed, the next step is to conduct holographic user emotion deep mining. The goal of this process is to construct a holographic feature map of the user's emotions through more in-depth analysis, which comprehensively displays the user's emotional state. Experimental parameters are set, for example, setting the deep mining dimensions to five: emotion intensity, emotion duration, emotion fluctuation amplitude, emotion transition frequency, and emotion category. When implementing holographic emotion deep mining, it is first necessary to compare the user's emotional feature data with the existing emotional feature database to identify the unique pattern of the user's emotions. Cluster analysis methods can be used to compare the user's emotional features with the emotional features of other users to identify the user's emotional reactions in specific situations. In this way, a multi-dimensional emotional feature map can be constructed to show the richness and complexity of user emotions. The constructed user emotional holographic feature map is visualized for subsequent analysis and application. Graphical tools can be used to display the user's emotional feature map in the form of 3D graphics or heat maps to make emotional changes more intuitive. This map will provide the intelligent agent with rich emotional background information, helping it to adjust its response strategy in subsequent interactions and improve the user experience.

[0014] Step S3: Adaptively perform acoustic gain processing on the voice signal input stream, and perform voice-emotion correlation analysis based on the user's emotional holographic feature map to generate a voice-emotion linkage mapping spectrum; In this embodiment, the user's voice signals are captured in real time. These signals are input through a microphone and converted into digital signals. These voice signals are analyzed using digital signal processing (DSP) technology. Fast Fourier transform (FFT) methods can be used to perform spectral analysis on the signals to identify the main frequency components and noise levels. Based on this, an adaptive filtering algorithm, such as a Kalman filter, is applied to adjust the acoustic gain in real time. By analyzing the current background noise level, the gain value is dynamically adjusted to ensure that the user's voice remains clearly audible amidst the background noise. This process significantly improves the quality of the voice signal and provides a better data foundation for subsequent voice-emotion correlation analysis. Once the acoustic gain processing is completed, the next step is to conduct voice-emotion correlation analysis based on the user's emotional holographic feature map. This process aims to identify the relationship between the user's voice features and their emotional state. Experimental parameters are set, for example, the emotion categories to be analyzed are set to the four basic emotions: happiness, sadness, anger, and surprise. When conducting voice-emotion correlation analysis, features are first extracted from the processed voice signal. These features can include pitch, speaking rate, intensity, and prosody. Mel-frequency cepstral coefficients (MFCCs) can be used to extract voice features for subsequent analysis. After feature extraction, these data are compared and analyzed with the user's emotional holographic feature map. Machine learning algorithms, such as random forests or neural networks, are used to perform correlation analysis on the extracted voice features and emotional features. By training the model, the correlation between voice features and specific emotional states is identified. Specifically, the extracted pitch changes can be matched with the emotional state to generate a mapping relationship between emotions and voice features. A voice-emotion linkage mapping spectrum is generated. Visualization tools can be used to display the correlation between voice features and emotional states in the form of a graph. For example, the distribution of emotional categories and changes in voice features are plotted in a two-dimensional coordinate system, so that the manifestation of different emotions in voice features is clear at a glance. This mapping spectrum will provide important information for the intelligent agent to help it better understand and respond to user emotions during interaction.

[0015] Step S4: tracking eye gaze migration based on the user's emotional holographic feature map and real-time facial image, and performing user attention state evolution to generate a user interaction depth intention signal; In this embodiment, a high-precision eye tracking device (such as an eye tracker) is used to capture real-time eye movement data. This device illuminates the user's eyes with an infrared light source, capturing minute eye movements and converting these movements into gaze point coordinate data. To ensure data accuracy, the system requires preliminary calibration to adjust the tracking algorithm based on the user's eye characteristics. Next, the real-time eye gaze data is combined with the user's emotional holographic feature map for analysis. Experimental parameters are set, such as a 5-second time window for gaze analysis, to capture changes in user attention in specific situations. During this process, heat mapping technology can be used to visualize the user's gaze on the interface, clearly indicating the content area the user is focusing on. When implementing gaze shift tracking, the user's gaze shifts are first recorded, including the start and end positions of the gaze, and the migration speed. This data is used to analyze the user's attention evolution. When the user's attention is focused on a specific area, the system should update and record their emotional state in real time to identify signals of the user's interaction depth intention. Once gaze shift tracking is complete, the next step is to analyze the user's attention state evolution. The goal of this process is to assess the user's attention focus and changing trends during the interaction process. Set experimental parameters, such as dividing the attention state into three states: highly concentrated, moderately concentrated, and dispersed. When implementing the attention state evolution analysis, the time series analysis method can be used to calculate the proportion of the user's gaze time in different time periods. By analyzing the user's gaze time on specific content, the user's interest in the content can be assessed. Specifically, if the user stays on a certain content for a long time, it can be inferred that the user is highly concerned about the content. Based on the results of eye gaze migration tracking and attention state evolution, a user interaction depth willingness signal is generated. This signal will quantify the user's interaction willingness and help the intelligent agent understand the user's real needs in the interaction. A comprehensive index can be set, such as combining the degree of attention concentration with the emotional state to form a depth willingness signal from 0 to 100.

[0016] Step S5: Based on the voice-emotion linkage mapping spectrum and the user's interaction depth intention signal, interactive response strategy decisions and interactive voice style dynamic adjustments are made to build a user personalized interaction adaptation model; In this embodiment, the previously generated speech-emotion linkage map is combined with the user's interaction depth intention signal. This combination provides the agent with a comprehensive view of the user's emotions and intentions. By analyzing these two pieces of data, the agent can identify the user's emotional state (such as happiness, sadness, anger, etc.) and interaction depth intention (such as high engagement, moderate engagement, or low engagement). Next, based on this information, the agent needs to develop an interaction response strategy. Experimental parameters are set, such as setting the strategy decision options to include "positive response," "neutral response," and "soothing response." During this process, the agent uses a decision tree or rule-based system to select the appropriate strategy. For example, when the user displays high engagement and positive emotions, the agent may choose to respond positively, providing more information and interaction; whereas, when the user displays low emotions or shows a reluctance to engage, the agent may choose a soothing response, reducing the intensity of the interaction. Once a specific interaction response strategy is developed, the next step is to dynamically adjust the interaction voice style. This process aims to better match the agent's speech characteristics (such as intonation, speech rate, and emotional tone) to the user's emotional state. Set experimental parameters, such as adjusting the speech rate between 120 and 180 words per minute to accommodate the user's attention span. When implementing dynamic voice style adjustment, the agent will select different speech styles based on the user's emotional state. For example, for a user expressing happiness, the agent can respond with a higher pitch and faster speech rate to convey positive emotions; for a user expressing sadness or anxiety, a lower pitch and slower speech rate should be used to create a soothing atmosphere. Build a personalized user interaction adaptation model. This step aims to integrate the user's emotional characteristics, interaction willingness, and corresponding response strategies to form a dynamically adaptable model. Set experimental parameters, such as setting the model update frequency to after each interaction, to promptly reflect user changes. When building a personalized model, first collect historical user interaction data, including emotional changes, interaction feedback, and response effectiveness. Using machine learning algorithms (such as cluster analysis or reinforcement learning), the model will continuously optimize its parameters to adapt to the user's individual needs. This personalized interaction adaptation model provides the agent with real-time decision-making information, enabling more accurate and personalized responses in multimodal interactions.

[0017] Step S6: Perform real-time user interaction according to the user personalized interaction adaptation model, perform multimodal interaction feedback optimization, and build a multimodal interaction optimization model.

[0018] In this embodiment, the intelligent agent captures the user's facial expressions, voice signals, and other interaction data in real time to determine the user's current emotional state and willingness to interact. Based on a personalized interaction adaptation model, the intelligent agent analyzes this input data and generates appropriate responses. For example, when the user displays high emotional engagement, the intelligent agent may proactively provide more information or engage in in-depth discussion. Conversely, when the user appears fatigued or depressed, the intelligent agent may choose to simplify the interaction and provide a concise and gentle response. Based on this, the intelligent agent combines the acquired multimodal data (such as facial expressions, voice emotions, and eye gaze) to optimize real-time interaction feedback. Experimental parameters are set, such as assessing the user's emotional state every 5 seconds during the interaction to ensure the accuracy of real-time responses. By analyzing user feedback, the intelligent agent identifies effective elements and deficiencies in the interaction. Using machine learning algorithms (such as reinforcement learning or online learning), the intelligent agent continuously optimizes its interaction strategy. Specifically, the system records user feedback from each interaction and uses this data to adjust future responses. For example, if a user responds favorably to a particular voice style, the intelligent agent will prioritize it for subsequent interactions. In this way, the intelligent agent can gradually adapt to the user's preferences and improve the personalization of the interaction. Build a multimodal interaction optimization model. Integrate all optimization strategies and feedback into a comprehensive model for application in future interactions. Set experimental parameters, such as setting the model update frequency to after each interaction, so that user changes are reflected in a timely manner. When implementing the construction of a multimodal interaction optimization model, it is first necessary to integrate the user's historical interaction data, including emotional changes, response effects, and user preferences. Through cluster analysis, users are divided into different groups, and the model is trained according to group characteristics. The generated multimodal interaction optimization model will provide real-time decision support for the intelligent agent, enabling it to achieve more accurate responses in multimodal interactions.

[0019] In this embodiment, refer to Figure 2 , is a flowchart of the detailed implementation steps of step S1. In this embodiment, the detailed implementation steps of step S1 include: Acquire real-time facial images and voice signal input streams of interactive users based on the intelligent agent; performing contrast enhancement on the real-time facial image and adaptively optimizing the image resolution to obtain a resolution-optimized facial image; Perform image noise recognition on the resolution-optimized facial image and perform high-frequency filtering to suppress it, thereby obtaining a noise-suppressed facial image. Dynamically smoothing the noise-suppressed facial image to obtain a smoothed and optimized facial image; Perform visual recognition of facial feature points on the smoothed and optimized image and mark the user's facial feature points; Performing real-time micro-expression analysis based on the user's facial feature points to extract the user's real-time micro-expression; Perform deep emotional analysis on users' real-time micro-expressions to obtain their real-time emotional characteristics.

[0020] In this embodiment, a high-frame-rate camera (30fps or 60fps is recommended) and a high-sampling-rate microphone (44.1kHz) are used to synchronously capture video and audio data. While the camera is capturing images, the face detection module in OpenCV or MediaPipe is used to locate the face area in real time and extract the ROI (Region of Interest). Simultaneously, raw PCM data is retrieved from the microphone via an audio callback interface (e.g., based on PyAudio or WebRTC audio streaming). Video and audio frames are marked with a unified timestamp to ensure timing synchronization. All data streams are buffered to ensure uninterrupted input and provide frame-level data to subsequent modules. This stage also includes preliminary frame quality screening, such as image blur detection (using the Laplacian transform) and audio signal-to-noise ratio assessment, to eliminate substandard frames. Image enhancement processing is performed on the captured facial images. First, the CLAHE (Contrast Limited Adaptive Histogram Equalization) algorithm is used to enhance contrast in the luminance channel (Y channel) of the RGB image. Specific parameters are set to clipLimit = 2.0 and tileGridSize = (8,8). After enhancement, the RGB image is resynthesized. Image resolution is then adjusted based on system resources: when GPU utilization exceeds 80%, the image is automatically downsampled to 640×480; when resources are idle, the image is restored to 1280×720 or higher. This adjustment is dynamically triggered by calculating the difference between the current processing time and the frame interval. Resolution adjustment is performed using bilinear interpolation (OpenCV's resize function) to ensure that the scaled image is free of aliasing or noticeable distortion. After converting the image to grayscale, it is transformed into the frequency domain using a Fourier transform (FFT). High-frequency components are concentrated, identifying areas of image noise. The power spectral density of the frequency domain image is calculated, and a threshold (e.g., frequency amplitude > 150) is set to identify noise bands. These areas are then subjected to high-frequency filtering, using a bilateral filter (cv2.bilateralFilter) with parameters d=9, sigmaColor=75, and sigmaSpace=75 for denoising. Alternatively, you can use a wavelet transform to decompose the image into different frequency bands and then selectively perform soft-thresholding on the high-frequency sub-bands to suppress noise. After processing, the image is converted back to the spatial domain for the next step.

[0021] Temporal smoothing of consecutive frames is achieved to prevent feature jitter caused by inter-frame variations. Five frames, including the current frame and the two frames preceding and following it, are processed using a pixel-level sliding average. The sliding window size is 5, and the update interval is every frame. An exponential moving average (EMA) can also be used, with the formula: I_t = α*I_t + (1 - α)*I_{t-1}, where α is set to 0.6. Smoothing is performed separately on the grayscale image or the three RGB channels. To further reduce motion blur, optical flow estimation (Farneback method) is combined with pixel-level registration of the preceding and following frames, followed by weighted fusion to maintain motion consistency. The final output frame is used as a single stabilized image and input to the feature point recognition module. Keypoints are identified using a deep learning facial feature point detection model, such as the Dlib 68-point model or the MediaPipe Face Mesh model. Images are preprocessed to 256×256 size, and pixel values ​​are normalized to [0, 1] before being fed into the model for inference. The model output is a two-dimensional coordinate array of keypoints (e.g., 68 points, each with x and y coordinates). Each coordinate point corresponds to a specific facial region, such as the brow, eye corners, or mouth corners. The output is filtered and stabilized using non-maximum suppression (NMS). Euclidean distance is then used to calculate the magnitude of change in the keypoints and build a facial region motion model. The system caches the keypoint positions for the past five frames and calculates their inter-frame differences for micro-expression analysis. Using the identified facial feature points, dynamic changes between symmetrical points are calculated, such as the distance between the eyebrows, the height of the palpebral fissures, and the angle of the mouth corner's upward movement. Key regions associated with FACS-defined AUs (Action Units) are selected, and feature change modeling is performed for each AU. For example, AU1 (inner eyebrow lift) is represented by the change in the vertical distance between the upper edge of the eyebrow and the center of the eyebrows, while AU12 (mouth corner lift) is represented by the angle between the mouth corner and the center of the mouth. The change (Δd) is calculated and a threshold (e.g., Δd > 2px) is set to determine action activation. Multiple AUs in each frame are evaluated and combined to output a micro-expression label. This process is processed in frames and combines the changing trends of several past frames (such as LSTM unit input) to analyze the continuity and integrity of micro-expressions.

[0022] The extracted AU activation patterns are synchronously fused with the speech signal and fed into a multimodal emotion recognition model. Speech signal features, such as pitch (F0), volume, and MFCC (Mel-Frequency Cepstral Coefficients), totaling 13 dimensions, are extracted. A high-dimensional feature representation is then extracted using an RNN or CNN. The facial AU state vector (e.g., AU1=1, AU4=0, AU12=1) is concatenated with the speech features and fed into a Transformer architecture for cross-modal feature alignment and encoding. The context vector output by the Transformer is fed into a softmax classifier, which outputs an emotion category label (e.g., happy, sad, angry, surprised). Valence and arousal scores (ranging from -1 to 1) are also calculated to form a two-dimensional emotion feature representation. The resulting emotion output is then used by the digital human for interactive strategies such as expression control and voice intonation adaptation.

[0023] In this embodiment, refer to Figure 3 , is a flowchart of the detailed implementation steps of step S2. In this embodiment, the detailed implementation steps of step S2 include: Conduct time series evolution analysis on users' real-time emotional features and perform time gradient recursive analysis to generate a user emotional fluctuation trajectory map; Based on the user's emotional fluctuation trajectory, the user's real-time micro-expressions are verified and calculated for emotional consistency, generating a confidence level for the user's emotional authenticity. Calculate the facial region temperature of the smoothed and optimized facial image, and perform regional temperature change distribution to obtain a user's facial temperature change distribution map; Based on the user's facial temperature change distribution map, a physiological response synchronization analysis is performed to obtain the user's physiological response synchronization characteristics; Based on the confidence level of user emotion authenticity and the synchronous characteristics of user physiological reactions, deep holographic user emotion mining is carried out to construct a holographic feature map of user emotions.

[0024] In this embodiment, the emotional features extracted in the previous step (such as Valence and Arousal values ​​or discrete emotion classification results) are combined into a time series, with each emotional data point corresponding to a timestamp. These data are stored chronologically in a buffer window with a window size of T = 60 seconds and a time step of Δt of 1 second. Trajectories are updated once per second. A sliding window mechanism is used to retain the emotional state of the most recent 60 seconds. The emotional time series is input into the time gradient recursive analysis module, where first-order derivatives are used to analyze rising or falling emotional trends. Second-order derivative analysis is then performed to determine the acceleration of emotional change. For example, if Valence (V(t) - V(t-1) > θ1), an upward trend is observed. If the continuous rise lasts for more than τ seconds, it is marked as an emotional fluctuation segment. Polynomial regression (such as a third-order polynomial fit) is used to model the time series and generate a continuous trajectory curve. When plotting the fluctuation trajectory, the x-axis represents time and the y-axis represents Valence or emotion category value. If multiple emotions are involved (the Valence-Arousal dimension), a two-dimensional emotional trajectory heat map or dynamic vector graph is plotted. This step outputs a continuous emotion fluctuation curve or heat map for subsequent consistency verification. At each time t, the AU combination output by the micro-expression recognition module is mapped to a micro-expression emotion label (e.g., AU6+AU12 corresponds to happiness) and compared with the emotion classification result of the emotion analysis module at that time (e.g., Valence>0.5). If the micro-expression matches the emotion label, it is recorded as consistent (1), otherwise it is inconsistent (0). Within the time window T, the consistent matching rate P_match = ∑ consistent frames / ∑ total frames is calculated. At the same time, the consistency confidence scoring model is constructed by combining emotional stability (fluctuation derivative close to 0) and micro-expression activity (AU activation frequency). The confidence score S_conf = w1×P_match + w2×Stability + w3×AU_density, where w1, w2, and w3 are weights (e.g., 0.5, 0.3, and 0.2). The Sigmoid function is used to normalize S_conf to the interval [0,1], and the emotional credibility of each time slice is marked. If the confidence level is > 0.8, it is marked as "high authenticity", otherwise it is marked as "suspicious" or "controlled emotion". The output is a timeline graph of emotions with confidence labels.

[0025] A thermal imaging module (such as an infrared camera or computer vision estimation algorithm) is introduced to analyze the facial thermal radiation pattern. If real infrared data is unavailable, the RGB image is used to simulate the thermal intensity distribution using a skin color region modeling algorithm, and a relative temperature index is calculated based on the skin redness. The specific method involves extracting the skin region from the RGB facial image, converting it to YCbCr or HSV color space, and estimating the temperature based on the red and H channels. The example formula is: Temp(x,y) ≈ α × R(x,y) + β × H(x,y), where α and β are linear fitting coefficients. The face is divided into multiple regions (forehead, nose tip, eye corners, cheeks, lips, etc.), and the average temperature value within each region is calculated. For each region, ΔT(t) = T(t) - T(t-1) is set, and the temperature change per second is calculated. To plot the facial temperature distribution, the temperature of each region is mapped to a colormap to generate a facial heat map. The image update frequency is synchronized with the video frame rate (30 fps), and the temperature change trajectory is displayed in an animated manner. The output data structure includes a time series of regional temperature vectors and dynamic heatmap frames. The temperature change values ​​of the facial region for each frame are collected and aligned with the emotion label and AU state at the same time point. The temperature response delay ΔT_delay, triggered by the emotion change, is calculated: the time difference between the onset of the emotion and the start of the temperature change. ΔT_delay is statistically analyzed across multiple time periods to determine the temporal consistency of the physiological response. Features of facial skin blood flow redistribution (such as a drop in nose tip temperature and a rise in cheek temperature) are also extracted and labeled as physiological response events. A facial physiological event vector E(t) = [event1, event2, ...] is constructed, with each event being a Boolean value (whether it occurred or not). The synchronization between the emotion fluctuation pattern and the temperature event sequence is compared, and the dynamic time warping (DTW) distance is calculated as a synchronization match score. The synchronization score S_sync = exp(-DTW_distance); a value closer to 1 indicates higher synchronization. The final output is a list of physiological synchronization indicators for the user, such as a synchronization score, a response delay graph, and key response event markers (such as "anger → sharp drop in nose tip temperature"), which are used for high-dimensional emotion modeling. The following multi-source data is integrated: 1) Valence-Arousal emotion trajectories; 2) micro-expression emotion label sequences; 3) emotion confidence scores S_conf; 4) regional temperature change sequences T(t); and 5) physiological synchronization scores S_sync. This data is normalized and fed into a graph neural network (such as GAT or GCN) to model the user's emotional node graph. Each time node represents a snapshot of the user's emotional state, including an emotion label, an AU combination, temperature features, confidence scores, and synchronization features. The graph structure is connected by temporal dependencies and similarity edges between nodes, with edge weights representing the magnitude of emotional evolution between nodes. The graph neural network is trained to classify and cluster patterns in the graph, extracting typical user emotional patterns (e.g., "expressionally active but physiologically sluggish").The output is a multi-layered emotional graph, encompassing time, modality (visual, thermal, physiological), and consistency. This feature graph provides input for digital humans, such as emotion prediction and behavior recommendations.

[0026] In this embodiment, refer to Figure 4 , is a flowchart of the detailed implementation steps of step S3. In this embodiment, the detailed implementation steps of step S3 include: Perform environmental noise recognition on the voice signal input stream to obtain the environmental noise frequency band; Calculating the noise frequency of the ambient noise frequency band and performing dynamic noise filtering to obtain a noise-filtered speech signal; Identify user voiceprint, pitch, rhythm and pause characteristics based on noise-filtered speech signals; Performing adaptive acoustic gain processing based on the user's voiceprint, pitch, rhythm, and pause characteristics to generate an adaptively optimized voice signal; Based on the user's emotional holographic feature map, the adaptively optimized voice signal is subjected to speech-emotion correlation analysis to generate a speech-emotion linkage mapping spectrum.

[0027] In this embodiment, the original voice signal collected from the microphone is single-channel PCM data, with a sampling rate of typically 44.1kHz or 48kHz, and is framed (e.g., 25ms per frame, 10ms frame shift). A fast Fourier transform (FFT) is performed on each frame to extract its spectral information. The multi-frame spectrum is then averaged to obtain the spectral envelope over a period of time. The spectral subtraction method is used to estimate the power spectral density (PSD) of the ambient noise, and the spectral distribution of the silent background is collected in the speech-inactive segment (detected by VAD) as the noise baseline. The spectrum is divided into several bands according to frequency (e.g., one band every 100Hz), and the energy level of each band is counted. A noise energy threshold θ is set (which can be dynamically adjusted to 30% of the average energy), and frequency bands with high energy but no speech activity are identified and marked as noise bands. For example, if the energy of the 300–600Hz and 4000–5000Hz frequency bands is significantly higher than that of other bands in a silent frame, they are recorded as a noise frequency band set F_noise = {300–600, 4000–5000}Hz for subsequent filtering. Based on the identified noise frequency bands F_noise, the original speech is filtered in the frequency domain. A notch filter is constructed to block these frequency bands. The filter parameters are set based on F_noise. For example, a notch filter H(f), f∈F_noise, is designed using an IIR (second-order notch filter) or FIR implementation. Each audio frame is subjected to an FFT, multiplied by the filter response H(f), and then restored to the time domain using an iFFT. In a dynamic noise environment, the F_noise set must be updated every T seconds (e.g., 2 seconds). An update mechanism is implemented: if the average spectral energy change exceeds a δ threshold (e.g., 3dB), the noise spectrum is re-estimated and the filter is adjusted. Furthermore, Wiener filtering is incorporated for enhancement. In each frame, the signal-to-noise ratio (SNR) is estimated based on the speech and noise power spectra. A Wiener filter gain function (G(f) = SNR / (SNR+1)) is constructed. This gain is applied frequency-by-frequency to adjust the spectrum, followed by iFFT reconstruction. The final output is a low-noise speech signal (S_filtered(t)) for acoustic feature extraction. The filtered speech signal is framed and processed, and acoustic features are extracted. Voiceprint recognition uses MFCC (Mel-Frequency Cepstral Coefficient) features. MFCC coefficients are typically extracted in 13-20 dimensions, with a frame length of 25ms and a frame shift of 10ms. Delta and Delta-Delta coefficients are combined to form a 39-dimensional feature vector. This is fed into a trained voiceprint recognition model (such as an i-vector or x-vector system, or the DNN-based ECAPA-TDNN model) to identify the user and output a voiceprint vector. Pitch feature extraction includes fundamental frequency (F0) detection and formant frequency calculation. F0 tracking is performed using Praat or PyWorld, and pitch variation curves are recorded.Rhythm is extracted from the time intervals between syllables. Short-term energy envelopes are used to detect the onset and end of speech activity, and the duration of consecutive syllables and pauses is counted. Pause features include the duration, location, and density of silence segments. These features are encoded into a temporal structure: pitch sequence F0(t), rhythm template R(t), and pause interval P(t). All feature data are normalized for subsequent acoustic adjustment.

[0028] Gain adjustment is performed based on the user's acoustic characteristics to improve speech clarity and expressiveness. First, speaking style parameters corresponding to the voiceprint, such as speaking rate and articulation clarity range, are analyzed to generate a user feature profile. The pitch curve F0(t) is smoothed using a Savitzky-Golay filter (window size 11, cubic polynomial) to prevent sudden changes. Its fundamental frequency range is also fine-tuned based on the target emotion curve (e.g., increasing the average F0 by 10Hz for happiness). Rhythm and pause features are used to control speech rate adjustment. The syllable intervals in R(t) are scaled to ensure that the overall speaking rate conforms to a standard range (e.g., 150-180 words per minute). If speech rate unevenness is detected, the phoneme positions on the timeline are adjusted and silent segments are interpolated, lengthened, or compressed to achieve rhythmic balance. The energy envelope is compressed using a dynamic range compression algorithm: if the speech energy exceeds the upper limit of -5dB, the gain is reduced; if it falls below -25dB, the gain is increased to maintain consistent signal loudness. After processing, the audio frames are resynthesized and the optimized speech S_adapt(t) is output for emotion mapping. S_adapt(t) is feature-encoded and matched with the holographic emotion map. High-dimensional features are extracted from the optimized speech, including the F0 curve, formant sequence, energy envelope, rhythm, and pause vectors. These features are then matched against similar speech emotion templates in the map using cosine similarity to identify the emotion label corresponding to the current speech state (e.g., anger - high energy + fast rhythm + high fundamental frequency). A speech feature vector V_speech(t) is constructed and matched to the emotion node V_emotion in the holographic map. The similarity sim(V_speech, V_emotion) is calculated as the mapping weight. For each frame or segment of speech, a mapping spectrum matrix M(t, e) is generated, where e represents the emotion label category. This spectrum matrix is ​​used to depict the linkage between speech input and emotional state. Using t-SNE or UMAP algorithms, the linkage spectrum is reduced in dimension and visualized into a two-dimensional graph, forming a speech-emotion linkage path trajectory diagram, demonstrating the real-time correlation between emotion evolution and speech features. The final output can be used to drive real-time feedback mechanisms such as digital human voice pitch change and voice expression enhancement.

[0029] In this embodiment, the specific steps of performing speech-emotion correlation analysis on the adaptively optimized speech signal to generate a speech-emotion linkage mapping spectrum are as follows: Perform multi-level speech analysis on the adaptively optimized speech signal to extract semantic features, pragmatic features, and extralinguistic implicit features; Mining user contextual intentions based on semantic and pragmatic features to obtain the strength of user interaction intentions; Perform fuzzy intention analysis on extralinguistic implicit features to generate user interaction fuzzy intentions; Construct a user intention hierarchy tree based on the user interaction intention strength and user interaction fuzzy intention; Perform dynamic probability deduction on the user intention hierarchy tree to identify the user's explicit and implicit communication goals; Based on the user's explicit and implicit communication goals, voice-emotion correlation analysis is performed to generate a voice-emotion linkage mapping spectrum.

[0030] In this embodiment, the adaptively optimized speech signal S_adapt(t) undergoes ASR (Automatic Speech Recognition) processing. We recommend using a streaming speech recognition model based on the Transformer or Conformer architecture (such as Whisper, wav2vec 2.0, or U-S2) to transcribe the speech stream into text T(t). The resulting text undergoes word segmentation, part-of-speech tagging, and named entity recognition (NER) to establish its basic semantic structure. Semantic feature extraction includes subject-verb-object structure identification (SVO extraction), keyword extraction (TextRank or TF-IDF), and dependency tree construction (using SpaCy or Stanza). Further annotation is performed on intentional vocabulary, such as auxiliary verbs like "want," "can," and "must," which express strong intent. Pragmatic features include tone analysis (interrogative, imperative, and exclamatory), pronoun binding (resolving pronouns like "he" and "this"), and register identification (e.g., formal / informal). A rule-based BERT-based fine-tuned classifier is used to categorize pragmatic contexts. Implicit extralinguistic features are captured through joint context-speech-emotion modeling. Non-semantic aspects of speech, such as speech rate variations, sudden pauses, and non-verbal sounds (sighs, laughter, and silence), are fused with previously extracted emotional features to extract a sequence of extralinguistic behavioral features for subsequent fuzzy intent analysis. The semantic and pragmatic features of the text T(t) are input into the intent recognition module. Using a multi-label intent recognition model (such as the BERT-BiLSTM-CRF architecture), each sentence is classified at the sentence level (e.g., requesting information, expressing attitude, issuing a command, confirming, etc.), outputting the intent type and confidence level.

[0031] The intent strength assessment is calculated by combining the following: Keyword reinforcement coefficient: For example, weighting of high-intensity words such as "must", "now", and "immediately"; Intonation driving factors: sudden rise in pitch and amplification of energy are signs of enhancement; Pragmatic modality factor: imperative and wishful sentences strengthen intention judgment; Contextual continuity: If the semantics of multiple sentences continuously tend towards the same goal, the intensity of intention is enhanced.

[0032] Set a normalized intent strength score S_intent ∈ [0, 1], such as: S_intent = w1×P_sem+ w2×P_prag+w3×audio_intensity. A threshold above 0.6 indicates strong intent, 0.3-0.6 indicates medium intent, and below 0.3 indicates weak intent.

[0033] The system uses a fuzzy logic system to model the extralinguistic feature stream, including non-verbal event sequences (such as silence, laughter, and inhalation), speech rate anomaly detection, and long pause recognition. The system uses rules such as: IF (long pause > 2s) AND (pitch drops) THEN vague intent = "hesitation" IF (large fluctuations in speech speed) AND (low speech energy) THEN Fuzzy Intent = "emotional repression with rejection" The above rules are constructed into a fuzzy inference network. The input feature vector F_extra is used to output a set of fuzzy intent labels I_fuzzy = {hesitation, avoidance, suggestion, transfer, etc.}, each with a fuzzy membership μ ∈ [0,1]. The resulting fuzzy intent is expressed as a vector: ["hesitation": 0.8, "reservation": 0.6, "avoidance": 0.3], which is subsequently integrated with the explicit intent to form a hierarchical structure. The intent hierarchy Tree_intent is constructed, with the explicit semantic intent label I_sem and the fuzzy extralinguistic intent label I_fuzzy as nodes. The root node is "total interaction intent." The first layer contains the explicit intent type (e.g., "request," "confirmation," "expression"), and the second layer contains the corresponding fuzzy intent supplementary description (e.g., "request-hesitation," "expression-implicit").

[0034] Structured generation method: Explicit intent labels are sorted by confidence S_intent, and the ones with higher confidence become the main branches first; Fuzzy intent is linked based on the semantic correlation between membership degree μ and explicit intent; Each leaf node records the source timestamp, signal source (speech / pragmatic / extralingual) and weight.

[0035] The tree structure is represented using a graph database (such as Neo4j). Each node's attributes include metadata such as intent type, strength, voice support, and fuzzy feature support, enabling graph traversal, querying, and updating. This tree provides the foundational data structure for subsequent deduction and analysis. Dynamic reasoning is performed based on Tree_intent. The state probability of each node N is denoted by P(N), initially assigned by the intent recognition module (e.g., confidence level and fuzzy membership). Transition probabilities between nodes are calculated using a Bayesian network or Markov chain model, and path probabilities are derived. The dynamic update rules are as follows: Each time new input is obtained, P(N) of the current node is updated. If multiple fuzzy intents simultaneously point to an explicit intent, the explicit intent is weighted and strengthened. If explicit intents appear continuously while the fuzzy intent changes, a trend of intent transition may develop.

[0036] Finally, two types of targets are output: Explicit communication target: P>0.7 and it is an explicit node; Implicit communication goal: inferred by multiple fuzzy nodes with P cumulative > 0.5, but not clearly expressed.

[0037] For example: "I can consider it" + pause + "but..." → explicit is "expression-uncertainty", implicit is "possible rejection". Map the user's communication goals (explicit, implicit) to the existing emotional hologram and speech feature vector space. Construct a mapping matrix M(e, i), where e is the emotion label and i is the intent label (such as request-eager, rejection-guilt). Extract the emotional expression features V_emotion_speech (F0 curve, speaking speed, volume, tone) from the optimized speech signal. Perform cosine similarity matching on V_emotion_speech and the emotional expression feature template V_emotion_intent corresponding to the intent label, and calculate the matching degree S_map = cos(V1, V2). If S_map>0.7, it is marked as a strong linkage. Finally, a speech-emotion-intent is formed. Figure 3 Dimensional mapping spectrum: The triplet of emotion, voice features, and intention state for each time period is stored in the spectrogram to generate a linkage sequence trajectory. This spectrogram supports the use of linkage behavior generation modules such as digital human emotional voice response and dynamic semantic feedback.

[0038] In this embodiment, step S4 includes the following steps: Based on the adaptive optimization of voice signals, the multimodal interaction consistency evaluation of the user's emotional holographic feature map is performed to obtain the multimodal interaction consistency evaluation index; Perform eye gaze point recognition on the smoothed and optimized facial image and perform dynamic migration tracking to obtain the user's gaze point migration trajectory; Pupil constriction frequency and rate analysis is performed based on the smoothed and optimized facial image to obtain the user's pupil constriction pattern; The user's attention state evolves based on the user's pupil contraction pattern and the user's gaze point migration trajectory, generating the user's real-time attention state; Based on the multimodal interaction consistency evaluation index and the user's real-time attention status, the user's interaction willingness intensity and interaction cognition are evaluated to generate a user interaction depth willingness signal.

[0039] In this embodiment, this step performs consistency analysis on the optimized speech signal and other modalities in the atlas (image, temperature, physiology, expression, etc.) to determine whether the speech emotion matches the facial emotion, etc. The emotional features (F0 change rate, energy fluctuation, rhythm dynamics) in the adaptive speech signal are extracted and encoded as V_speech_emotion. At the same time, the facial emotion vector V_face_emotion and the physiological emotion vector V_physio_emotion of the current time period in the atlas are extracted. Construct a modal matching network and calculate the pairing similarity of multiple modal vectors: S_speech_face = cos(V_speech_emotion, V_face_emotion), S_speech_physio = cos(V_speech_emotion, V_physio_emotion). These similarities are summarized and input into the modal consistency scoring function: ΔE indicates whether the fluctuation trends of the emotion curves are consistent. The Cross-Modal Emotion Matching Index (CEMI), which ranges from 0 to 1, characterizes the degree of multimodal emotional consistency. This index serves as a key input for subsequent user intention assessment.

[0040] High-precision eye tracking is performed using the eye regions in facial images. Eye landmarks (6-8 in total) are extracted using Dlib or MediaPipe. Eye rotation direction is then estimated using the Pupil Center-Corneal Reflection (PCCR) algorithm. For environments without infrared support, deep learning-based models (such as RT-GENE or GazeML) can be used for 2D / 3D gaze point estimation. Each frame outputs a gaze point coordinate, P_t = (x_t, y_t), calibrated in image space or screen coordinates. The gaze point migration distance d = ||P_t - P_{t-1}|| is calculated between consecutive frames, and a gaze point migration sequence Track(t) is constructed. The trajectory curves can be visualized as a trajectory graph to identify saccades, gazes, and jumps. The gaze point migration trajectory data is used in subsequent attention analysis and plays a fundamental role in identifying behavioral characteristics such as "saccades," "extended fixations," and "avoidant fixations." The output includes a trajectory sequence, behavioral segment annotations, and fixation frequency. The pupil region is extracted from facial images, and the pupil edge is accurately identified using red and blue channel separation or convolutional neural network detection methods (such as ElSe or PupilNet). The pupil diameter D_t is extracted from each frame, forming a time series D(t). This series is used for frequency analysis and rate calculation. After removing eye movement artifacts using a bandpass filter (e.g., 0.1–4 Hz), D(t) is subjected to a fast Fourier transform (FFT) to obtain the dominant frequency peak f_max, which corresponds to the pupil constriction frequency (e.g., 0.6 Hz represents ~0.6 pupil constrictions per second). The constriction rate is also calculated as ΔD / Δt, the rate of change between pupil constriction and dilation. Pupil patterns are clustered and classified (e.g., using the dominant frequency, average diameter, and maximum constriction amplitude) using K-means to identify typical patterns: Type A (stable focus), Type B (rapid fluctuation), and Type C (tense constriction). This classification serves as one of the key input features for attention assessment. Pupil pattern features (e.g., constriction frequency and amplitude) are integrated with gaze trajectory behavioral features to construct an attention state assessment model. The feature set includes: 1) fixation time T_fixation; 2) jump frequency F_jump; 3) pupil constriction category; 4) eye movement velocity V_move; and 5) gaze-stimulus alignment (used in HCI applications to identify whether the user is fixating on key areas). These features are dynamically modeled using a multivariate time series model (such as LSTM+Attention), outputting the attention state A(t) ∈ {focused, distracted, saccade, interrupted, fatigued} in real time. A confidence score (e.g., based on Softmax probability) is set to assess the stability of this state judgment. The generated attention state sequence is used to perceive the user's level of engagement. A continuous "focused" output indicates high attention; frequent fluctuations indicate high attention fluctuations. This state is input into the comprehensive evaluation module as a key factor in interaction willingness.

[0041] The CEMI value and attention state distribution are used as input factors to construct an interaction intention assessment model. The following formula is constructed: Interaction_Strength = α × CEMI + β × Attention_Index + γ × Intent_Strength. Attention_Index is calculated from the proportion of "focused" time in the state sequence, gaze consistency, and pupil stability, and Intent_Strength is the strength score of the intent recognition module described above. α, β, and γ are weighting coefficients, recommended settings are (0.4, 0.4, 0.2) or can be adjusted based on specific experiments. The final output is the user's deep engagement intention signal (Engagement Score), a continuous value ranging from [0 to 1], which can be graded as high engagement (>0.75), medium engagement (0.5-0.75), and low engagement (<0.5). This signal serves as the basis for adjusting the agent's behavior, such as whether to continue guiding, switch topics, or speed up the pace.

[0042] In this embodiment, the specific steps of step S5 are: According to the user's emotional holographic feature map, the agent's interactive response strategy is decided to obtain the agent's interactive emotion modulation parameters; Dynamically adjust the interactive voice style based on the voice-emotion linkage mapping spectrum and the user's interaction depth intention signal to generate a voice intelligent style control matrix; Conduct multi-scenario interaction simulations on the agent interaction emotion modulation parameters and voice intelligence style control matrix to construct multi-scenario interaction simulation data; Perform user personalized interaction adaptation based on multi-scenario interaction simulation data and build a user personalized interaction adaptation model.

[0043] This embodiment relies on holographic modeling of the emotional characteristics expressed by users in historical and current multimodal interactions, known as a "user emotional holographic feature map." This map integrates multi-dimensional information such as voice intonation, speaking rate, facial expressions, eye movements, body language, and text semantics, and constructs a user emotional state distribution vector in a high-dimensional embedding format. For example, a user's pattern of "anxiety, high speaking rate, and strong facial expressions" expressed in interactions over different time periods would be classified as a high-turbulence, high-alertness emotional cluster. During the policy decision-making process, the system compares the user's current emotional state, collected in real time, with this map and calculates the similarity distribution of the current state across historical emotion clusters. Common methods include cosine similarity and dynamic time warping (DTW). Based on the cluster with the highest similarity, the system calls a preset interaction response strategy library to extract the corresponding interaction modulation strategy, which includes parameters such as intonation intensity, speaking rate variation, content depth, and emotional overtones (e.g., more empathy or more rationality). Through a precise mapping mechanism between voice style and emotional state, the digital human's voice response style can be dynamically adjusted, thereby enhancing the intelligent agent's emotional perception and expression capabilities. First, the system constructs a "voice-emotion linkage mapping spectrum" derived by training on a large number of voice samples and their corresponding emotion labels. A common approach uses a BiLSTM+Attention architecture to extract nonlinear mappings between acoustic features such as pitch, tone, rhythm, energy, and timbre, and emotional dimensions. For example, "low frequency + short sounds + low energy" can be mapped to the "sadness" emotional distribution. Furthermore, the system receives real-time input from the user's "interaction depth intention signal." This signal is analyzed by a semantic recognition module (such as the intent recognition BERT model) to determine the user's current intentional depth needs and determine whether they prefer "shallow inquiry" or "deep conversation." For example, the strength of the intention signal increases with continuous questioning or the use of more abstract expressions. These two inputs jointly drive the construction of the "voice intelligence style control matrix." The matrix dimensions include speech rate (slow-normal-fast), intonation emotion (neutral-positive-negative), timbre baseline frequency adjustment, pause length, semantic density adjustment, and emotional color rendering intensity. Each dimension is quantified as an adjustable parameter range. Different emotion categories and interaction depth levels will correspond to different matrix weight configurations, achieving adaptive dynamic switching of voice styles. In order to verify the adaptability and robustness of the modulation strategy, this step introduces a multi-scenario interaction simulation system to simulate the aforementioned parameter combinations in different user scenarios. The simulation environment is driven by a pre-built virtual user emotional state library, which includes typical interaction scenarios such as "high-pressure workplace communication", "child companionship education", and "elderly psychological counseling". Each scenario has a clear user emotional state distribution, intention structure, and interaction expectations.In the simulation, the system applied varying proportions of "emotion modulation parameters + voice style matrix" as input to different virtual user interaction flows. Each interaction round recorded simulated user feedback scores, including five metrics ranging from 0 to 1: "understanding," "emotional comfort," "interaction naturalness," and "emotional resonance index." To ensure simulation breadth and coverage, the experiment set up 30 different scenarios and 150 parameter combinations, with an average interaction duration of 90 seconds per combination. Each scenario simulation generated a complete multimodal interaction record, including text logs, voice output samples, emotional feedback records, and rating data. These were aggregated to form a "multi-scenario interaction simulation dataset." Based on the differences in emotional responses and voice preferences among different users, a personalized, multi-dimensional interaction style adaptation model was constructed to enhance the customized interaction capabilities of the digital human system. The input was the large-scale scenario interaction dataset obtained in step S3. The system first classified the data by user feature labels, such as age group, common emotional patterns, and voice style preferences (e.g., preference for rational or emotional tone). The modeling method used is a fusion deep learning structure with a Transformer encoder as the backbone. The input is a triplet of situational features + user labels + interaction feedback indicators, and the output is the optimal voice style and emotion modulation strategy combination. The model is trained with the interaction satisfaction score as the supervisory signal, and a user memory module is introduced to record the "successful strategies" in historical interactions to form a user emotion adaptation memory pool. After the model is trained, it can automatically classify new users into specific user group labels based on their current interaction performance, and recommend the optimal interaction parameter combination in real time. For example, when a user shows anxiety for the first time, the system can directly call the historical optimal parameters in the "30-40 years old + highly sensitive + logical demand type" user group to quickly deploy voice style and emotion regulation strategies.

[0044] In this embodiment, the specific steps of step S6 are: Conduct real-time user interaction based on the user's personalized interaction adaptation model and collect instant user feedback information; Conduct defect analysis of the entire interaction process based on real-time user feedback to identify defects in the interaction process; Evaluate the interaction effect of user instant feedback information and analyze the interaction atmosphere to obtain the interaction effect evaluation value and real-time interaction atmosphere; Based on the defects of the interaction links, the interaction effect evaluation value and the real-time interaction atmosphere, the agent's interaction emotions are adjusted, and the voice expression style is dynamically reconfigured to obtain the agent's global interaction optimization parameters; According to the global interaction optimization parameters of the intelligent agent, the user's personalized interaction adaptation model is optimized through multimodal interaction feedback to construct a multimodal interaction optimization model.

[0045] In this embodiment, online interaction is carried out in the actual operating environment based on the personalized interaction adaptation model (P-Adapter). Real-time interaction covers multiple channels such as voice output, facial expression rendering, and action generation, with sampling frequencies of 24 kHz for voice output, 30 fps for expression rendering, and 60 fps for action frames. Each interaction is based on the user intention recognition module (Transformer encoder, with input of audio + semantic embedding, a total of 1024 dimensions) to determine the interaction response strategy, and then the P-Adapter generates a comprehensive output with specific intonation, expression, and language style. During the interaction execution process, the system synchronously collects instant feedback information from the user, including: unstructured voice feedback (transcribed and recognized by the Whisper-large model); Facial micro-expression changes (using OpenFace to detect 17 types of AU variations); Physiological state fluctuations (such as heart rate variability (HRV) and skin conductance (GSR), analyzed at a sampling frequency of 128 Hz); Interaction behavior characteristics (e.g., interruption frequency, response latency, gaze direction tracking); Subjective scoring / micro-expressions (when the system detects ambiguous points in the interaction, it will proactively guide micro-interaction sampling feedback such as "Does this meet your expectations?").

[0046] Identify possible problem links and process defects in the interaction process. The interaction process can be divided into five key nodes: intent recognition → response planning → language generation → multimodal rendering → execution feedback. We construct an interaction process dependency graph based on graph neural network (GNN), where nodes represent five interaction links, edges represent their up-and-down dependencies, and edge weights represent the confidence level of the interaction information transmission. The user's immediate feedback vector U_f is fed into the trained GraphSAGE model (3 layers, 128 nodes per layer), and the defect probability vector P_defect ∈ is output for each node. , representing the potential risk of failure in each link. At the same time, in conjunction with the comparative evaluation of conversation logs and system-generated content, a multi-target anomaly detection method (Isolation Forest + One-Class SVM) was used. The anomaly detection indicators included: sentiment deviation ΔV>0.35; Semantic mismatch rate > 15%; The proportion of redundant sentences in the response is >30%; User response interruption number ≥ 2; The user's negative emotion triggers (such as frowning, mouth corners drooping, and other AU expressions) are high frequency.

[0047] Finally, an interaction defect location report is formed, which clearly points out the problematic links and their specific manifestations, such as "the voice rhythm is too different from the current user's speaking speed" and "semantic differences are not clarified in time", etc., providing direct targets for subsequent emotional regulation and style reconfiguration.

[0048] After completing defect location, we need to further evaluate the overall effect of the current interaction and analyze the interaction atmosphere. We have established a two-channel evaluation model consisting of the following two parts: The Interaction Effectiveness Score (IES) includes: Input: instant feedback vector U_f + P-Adapter output vector; Model: Multilayer Perceptron (MLP, 512-256-64-1) + Sigmoid output; The output is IES ∈ [0,1], where IES > 0.75 indicates high user satisfaction; Training data: Based on real user A / B test records, supervised regression is performed using 1200 hours of labeled conversation data.

[0049] Real-time Interaction Climate Classification (ICC) analysis includes: Model structure: BERT-base (Chinese multimodal extension) + BiLSTM combination; Input information includes: user semantic content, voice emotion lines, and non-verbal behavior; Output categories: 6 types of atmospheres including relaxation, anxiety, doubt, excitement, indifference, and trust; Each atmosphere classification is equipped with a confidence output (Softmax); The model accuracy was evaluated on the IEMOCAP + self-built multimodal atmosphere dataset and reached 86.2%.

[0050] The results of this phase are: interaction effect evaluation value IES and atmosphere label C_climate, which provide directional guidance for subsequent emotional regulation and voice style reconstruction. By integrating defect diagnosis, interaction effect, and atmosphere perception results, a unified interaction adjustment strategy is generated, which is reflected in the global interaction optimization parameters: ; This parameter is used to control the emotional style, voice rhythm, and content strategy of the digital human in the next round of interaction.

[0051] The input vector is: P_defect + IES + C_climate (encoded as one-hot); Build a policy refinement network (Policy Refinement Net) with the TransformerEncoder structure (2 layers, hidden 256, 4 heads); The network learns to map out G_opt, including: Expression emotion weight adjustment ΔE ∈ ; Multimodal output weights (voice ratio, action ratio, etc.); TTS speech rate / intonation adjustment coefficient; Response proactiveness level; The model is trained in a simulation environment using the Proximal Policy Optimization (PPO) algorithm, aiming to maximize the long-term average IES and the user's chat continuation rate. The policy is updated every 500 rounds, with a KL distance constraint of ε=0.2. The final output, G_opt, achieves significant improvements in real-world interactions—in a 200-hour real-world interaction test, the sentiment mismatch rate decreased by 24.1% and the system's question hit rate increased by 17.8%. The model is trained in a simulation environment using the Proximal Policy Optimization (PPO) algorithm, aiming to maximize the long-term average IES and the user's chat continuation rate. The policy is updated every 500 rounds, with a KL distance constraint of ε=0.2. The final output, G_opt, achieves significant improvements in real-world interactions—in a 200-hour real-world interaction test, the sentiment mismatch rate decreased by 24.1% and the system's question hit rate increased by 17.8%.

[0052] In this embodiment, a multimodal interaction system of a digital human agent is provided, which is used to execute the multimodal interaction method of a digital human agent as described above, including: A micro-expression recognition module is used to obtain real-time facial images and voice signal input streams of interactive users based on the intelligent agent; perform real-time micro-expression recognition and deep emotion analysis based on the real-time facial images to obtain real-time emotional characteristics of the user; The emotion deep mining module is used to analyze the time series evolution of users' real-time emotion characteristics, conduct holographic user emotion deep mining, and construct a holographic feature map of user emotions; The speech analysis module is used to perform adaptive acoustic gain processing on the speech signal input stream, and perform speech-emotion correlation analysis based on the user's emotional holographic feature map to generate a speech-emotion linkage mapping spectrum; The gaze tracking module is used to track the migration of eye gaze points based on the user's emotional holographic feature map and real-time facial images, and to monitor the evolution of the user's attention state to generate a signal of the user's interaction depth intention; The interaction adaptation module is used to make interactive response strategy decisions and dynamically adjust the interactive voice style based on the voice-emotion linkage mapping spectrum and the user's interaction depth intention signal, thereby building a user-personalized interaction adaptation model; The interaction feedback optimization module is used to perform real-time user interaction based on the user's personalized interaction adaptation model, and to perform multimodal interaction feedback optimization to build a multimodal interaction optimization model.

[0053] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.

[0054] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.

Claims

1. A multimodal interaction method for a digital human agent, characterized in that: The following steps are involved: Step S1: obtaining a real-time facial image and voice signal input stream of an interactive user based on an agent; Performing real-time micro-expression recognition and deep emotion analysis based on the real-time facial image to obtain the user's real-time emotional characteristics; Step S2: Analyze the temporal evolution of the user's real-time emotional characteristics, conduct in-depth mining of holographic user emotions, and construct a holographic feature map of user emotions; Step S3: Adaptively perform acoustic gain processing on the voice signal input stream, and perform voice-emotion correlation analysis based on the user's emotional holographic feature map to generate a voice-emotion linkage mapping spectrum; Step S4: tracking eye gaze migration based on the user's emotional holographic feature map and real-time facial image, and performing user attention state evolution to generate a user interaction depth intention signal; Step S5: Based on the voice-emotion linkage mapping spectrum and the user's interaction depth intention signal, interactive response strategy decisions and interactive voice style dynamic adjustments are made to build a user personalized interaction adaptation model; Step S6: Perform real-time user interaction according to the user personalized interaction adaptation model, perform multimodal interaction feedback optimization, and build a multimodal interaction optimization model.

2. The multimodal interaction method of digital human agent according to claim 1, characterized in that: The specific steps of step S1 are: Acquire real-time facial images and voice signal input streams of interactive users based on the intelligent agent; performing contrast enhancement on the real-time facial image and adaptively optimizing the image resolution to obtain a resolution-optimized facial image; Perform image noise recognition on the resolution-optimized facial image and perform high-frequency filtering to suppress it, thus obtaining a noise-suppressed facial image; = Dynamically smoothing the noise-suppressed facial image to obtain a smoothed and optimized facial image; Perform visual recognition of facial feature points on the smoothed and optimized image and mark the user's facial feature points; Performing real-time micro-expression analysis based on the user's facial feature points to extract the user's real-time micro-expression; Perform deep emotional analysis on users' real-time micro-expressions to obtain their real-time emotional characteristics.

3. The multimodal interaction method of digital human agent according to claim 1, characterized in that: The specific steps of step S2 are: Conduct time series evolution analysis on users' real-time emotional features and perform time gradient recursive analysis to generate a user emotional fluctuation trajectory map; Based on the user's emotional fluctuation trajectory, the user's real-time micro-expressions are verified and calculated for emotional consistency, generating a confidence level for the user's emotional authenticity. Calculate the facial region temperature of the smoothed and optimized facial image, and perform regional temperature change distribution to obtain a user's facial temperature change distribution map; Based on the user's facial temperature change distribution map, a physiological response synchronization analysis is performed to obtain the user's physiological response synchronization characteristics; Based on the confidence level of user emotion authenticity and the synchronous characteristics of user physiological reactions, deep holographic user emotion mining is carried out to construct a holographic feature map of user emotions.

4. The multimodal interaction method of digital human agent according to claim 1, characterized in that: The specific steps of step S3 are: Perform environmental noise recognition on the voice signal input stream to obtain the environmental noise frequency band; Calculating the noise frequency of the ambient noise frequency band and performing dynamic noise filtering to obtain a noise-filtered speech signal; Identify user voiceprint, pitch, rhythm and pause characteristics based on noise-filtered speech signals; Performing adaptive acoustic gain processing based on the user's voiceprint, pitch, rhythm, and pause characteristics to generate an adaptively optimized voice signal; Based on the user's emotional holographic feature map, the adaptively optimized voice signal is subjected to speech-emotion correlation analysis to generate a speech-emotion linkage mapping spectrum.

5. The multimodal interaction method of digital human agent according to claim 4, characterized in that: The specific steps of performing speech-emotion correlation analysis on the adaptively optimized speech signal to generate a speech-emotion linkage mapping spectrum are as follows: Perform multi-level speech analysis on the adaptively optimized speech signal to extract semantic features, pragmatic features, and extralinguistic implicit features; Mining user contextual intentions based on semantic and pragmatic features to obtain the strength of user interaction intentions; Perform fuzzy intention analysis on extralinguistic implicit features to generate user interaction fuzzy intentions; Construct a user intention hierarchy tree based on the user interaction intention strength and user interaction fuzzy intention; Perform dynamic probability deduction on the user intention hierarchy tree to identify the user's explicit and implicit communication goals; Based on the user's explicit and implicit communication goals, voice-emotion correlation analysis is performed to generate a voice-emotion linkage mapping spectrum.

6. The multimodal interaction method of digital human agent according to claim 1, characterized in that: The specific steps of step S4 are: Based on the adaptive optimization of voice signals, the multimodal interaction consistency evaluation of the user's emotional holographic feature map is performed to obtain the multimodal interaction consistency evaluation index; Perform eye gaze point recognition on the smoothed and optimized facial image and perform dynamic migration tracking to obtain the user's gaze point migration trajectory; Pupil constriction frequency and rate analysis is performed based on the smoothed and optimized facial image to obtain the user's pupil constriction pattern; The user's attention state evolves based on the user's pupil contraction pattern and the user's gaze point migration trajectory, generating the user's real-time attention state; Based on the multimodal interaction consistency evaluation index and the user's real-time attention status, the user's interaction willingness intensity and interaction cognition are evaluated to generate a user interaction depth willingness signal.

7. The multimodal interaction method of digital human agent according to claim 1, characterized in that: The specific steps of step S5 are: According to the user's emotional holographic feature map, the agent's interactive response strategy is decided to obtain the agent's interactive emotion modulation parameters; Dynamically adjust the interactive voice style based on the voice-emotion linkage mapping spectrum and the user's interaction depth intention signal to generate a voice intelligent style control matrix; Conduct multi-scenario interaction simulations on the agent interaction emotion modulation parameters and voice intelligence style control matrix to construct multi-scenario interaction simulation data; Perform user personalized interaction adaptation based on multi-scenario interaction simulation data and build a user personalized interaction adaptation model.

8. The multimodal interaction method of digital human agent according to claim 1, characterized in that: The specific steps of step S6 are: Conduct real-time user interaction based on the user's personalized interaction adaptation model and collect instant user feedback information; Conduct defect analysis of the entire interaction process based on real-time user feedback to identify defects in the interaction process; Evaluate the interaction effect of user instant feedback information and analyze the interaction atmosphere to obtain the interaction effect evaluation value and real-time interaction atmosphere; Based on the defects of the interaction links, the interaction effect evaluation value and the real-time interaction atmosphere, the agent's interaction emotions are adjusted, and the voice expression style is dynamically reconfigured to obtain the agent's global interaction optimization parameters; According to the global interaction optimization parameters of the intelligent agent, the user's personalized interaction adaptation model is optimized through multimodal interaction feedback to construct a multimodal interaction optimization model.

9. A multimodal interactive system of a digital human agent, characterized in that: For executing the method according to claim 1, comprising: A micro-expression recognition module is used to obtain real-time facial images and voice signal input streams of interactive users based on the intelligent agent; perform real-time micro-expression recognition and deep emotion analysis based on the real-time facial images to obtain real-time emotional characteristics of the user; The emotion deep mining module is used to analyze the time series evolution of users' real-time emotion characteristics, conduct holographic user emotion deep mining, and construct a holographic feature map of user emotions; The speech analysis module is used to perform adaptive acoustic gain processing on the speech signal input stream, and perform speech-emotion correlation analysis based on the user's emotional holographic feature map to generate a speech-emotion linkage mapping spectrum; The gaze tracking module is used to track the migration of eye gaze points based on the user's emotional holographic feature map and real-time facial images, and to monitor the evolution of the user's attention state to generate a signal of the user's interaction depth intention; The interaction adaptation module is used to make interactive response strategy decisions and dynamically adjust the interactive voice style based on the voice-emotion linkage mapping spectrum and the user's interaction depth intention signal, thereby building a user-personalized interaction adaptation model; The interaction feedback optimization module is used to perform real-time user interaction based on the user's personalized interaction adaptation model, and to perform multimodal interaction feedback optimization to build a multimodal interaction optimization model.

Citation Information

Cited By

  • Intelligent agent digital image interaction generation method based on multi-modal perception

    CN121187453A

  • Sales customer service data classification and arrangement method and system based on AI analysis

    CN121327605A

  • Personalized image content generation and optimization method and system fusing generative AI

    CN121353442A

  • Digital human live broadcast voice interaction system fused with emotion calculation

    CN121393436A

  • Digital human live voice interaction system fused with affective computing

    CN121393436B