A Multimodal Interactive Digital Human Training Method and System Based on Artificial Intelligence
Patent Information
- Application Number
- CN202511832377.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-08
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-12-08
AI Technical Summary
[0005]本申请公开了一种基于人工智能的多模态交互式数字人训练方法及系统,旨在解决现有智能数字人系统在复杂公共环境中,处理不同类型信息在时间和空间上对应不准、含义表达方式不统一以及不断变化的杂音干扰方面的局限性,导致系统难以准确理解用户意图,并给出恰当回复的问题
[0021] This application discloses a multimodal interactive digital human training method and system based on artificial intelligence. By acquiring and analyzing the user's multimodal interaction information (including voice, vision, and environmental information), it can comprehensively perceive changes in the user's state and environment. Based on this, the method determines the priority score of each user interaction request according to the analysis results of the multimodal interaction information, and dynamically adjusts the allocation of computing resources and the order of communication responses according to this score. This mechanism can effectively solve the problems faced by existing intelligent digital human systems in complex public places such as transportation hubs, including multi-user concurrency, information timing disorder, environmental noise interference, and information association failure. By prioritizing the processing of user interaction requests with high priority scores, the method of this application can ensure that even in noisy and changing public environments, the system can still accurately understand the user's intent and provide timely and appropriate responses, significantly improving the interaction efficiency, accuracy, and user satisfaction of the intelligent digital human system, and overcoming the shortcomings of existing technologies such as inaccurate judgment of user intent and irrelevant responses.
Smart Images

Figure CN121682271B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of intelligent human-computer interaction technology, and more specifically, to a multimodal interactive digital human training method and system based on artificial intelligence. Background Technology
[0002] Intelligent digital human systems can interact with people through various information sources, including voice, text, images, and gestures. In quiet, single-user environments, the system can accurately align various types of information, understand user intentions, and respond appropriately. However, its performance significantly declines when deployed in high-traffic public places such as transportation hubs. In such environments, the system must simultaneously handle information input from multiple users and locations, leading to temporal and spatial misalignments in data collection, transmission, and processing, making it difficult to correctly associate the voice and actions of the same user. Furthermore, the complex background noise in public places, including broadcasts, conversations, and equipment noise, severely interferes with the target speech, drastically reducing the effectiveness of traditional noise reduction and speech separation techniques and causing a sharp deterioration in speech quality.
[0003] The information fusion module relies on clear and synchronized multimodal data. However, in reality, speech is masked by noise, images are blurred due to occlusion or lighting, and information arrives asynchronously, causing the system to fail to effectively integrate information and resulting in fragmented understanding of user intent. The system learns based on this contaminated and misaligned data, further confusing different user intents with noise characteristics, impairing its judgment mechanism, and often resulting in illogical and irrelevant responses. This severely impacts user experience and exposes the fundamental limitations of existing systems in processing multimodal information in complex environments.
[0004] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention
[0005] This application discloses a multimodal interactive digital human training method and system based on artificial intelligence, aiming to solve the limitations of existing intelligent digital human systems in complex public environments, such as inaccurate temporal and spatial correspondence of different types of information, inconsistent expression of meaning, and constant noise interference, which makes it difficult for the system to accurately understand the user's intention and give an appropriate response.
[0006] The technical solution of this application is as follows:
[0007] In a first aspect, this application discloses a multimodal interactive digital human training method based on artificial intelligence, the method comprising:
[0008] Acquire multimodal interaction information for each user and perform multimodal interaction information analysis; the multimodal interaction information includes at least voice information, visual information, and environmental information;
[0009] Based on the analysis results of multimodal interaction information, the priority score of the interaction request corresponding to each user's multimodal interaction information is determined;
[0010] Based on this priority score, the allocation of computing resources and the order of communication responses are dynamically adjusted to prioritize the processing of user interaction requests with higher priority scores.
[0011] Furthermore, the multimodal interaction information analysis includes: extracting speech features from the speech information to obtain speech features; the speech features include speech rate, pitch, volume, keywords, and emotional intensity; extracting visual features from the visual information to obtain visual features; the visual features include facial expressions, body movements, and eye movements; and integrating the environmental information to obtain environmental event information.
[0012] Based on this, according to the analysis results of multimodal interaction information, the priority score of the interaction request corresponding to each user's multimodal interaction information is determined, including: determining the urgency score of the interaction request based on the emotional intensity and keywords in the voice features; determining the salience score of the interaction request based on the relevance of the body movements in the visual features and the content pointed to by the body movements to the environmental event information; determining the contextual relevance score of the interaction request based on the degree of matching between the semantic content of the interaction request and the environmental event information; determining the basic need score of the interaction request based on the preset classification of the interaction request; and weighting the urgency score, the salience score, the contextual relevance score, and the basic need score to obtain the priority score.
[0013] Furthermore, the dynamic adjustment of computing resource allocation and communication response order based on the priority score includes: allocating computing resources to the interaction request based on the priority score; the computing resources include processor core time, memory buffer, and network bandwidth; shortening the processing response time for intent determination for interaction requests with high priority scores; placing interaction requests with low priority scores into a waiting queue for delayed processing; and interrupting the generation of low-priority response messages through the response output scheduler, and instead prioritizing the generation and output of response messages for interaction requests with high priority scores.
[0014] In some preferred embodiments, determining the priority score of the interaction request corresponding to each user's multimodal interaction information based on the analysis results of multimodal interaction information further includes: calculating the intent confidence score of the interaction request corresponding to the multimodal interaction information based on the analysis results of multimodal interaction information, and detecting the consistency between different modal interaction information; when the intent confidence score is lower than a preset threshold or the consistency is conflicting, initiating a diagnostic process and re-evaluating the user intent; and determining the urgency score, the salience score, the contextual relevance score, and the basic need score based on the user intent determined by the re-evaluation, and performing a weighted calculation to obtain the priority score.
[0015] As a technological improvement, based on the analysis results of multimodal interaction information, the intent confidence score of the interaction request corresponding to the multimodal interaction information is calculated, including: obtaining the speech intent recognition confidence score corresponding to the speech feature, the visual intent recognition confidence score corresponding to the visual feature, and the contextual association confidence score corresponding to the environmental event information; and weighting and fusing the speech intent recognition confidence score, the visual intent recognition confidence score, and the contextual association confidence score to calculate a comprehensive intent confidence score.
[0016] To improve the solution, the consistency between different modal interaction information is detected, including: comparing the consistency between the speech feature and the visual feature; initiating the diagnostic process and reassessing the user intent, including: assessing whether the average fundamental frequency and speech rate in the speech feature are consistently higher than the average level of the population when the facial expressions and body movements in the visual feature remain calm; if so, generating a targeted counter-question; and reassessing the user intent based on the user's response to the targeted counter-question.
[0017] As a further improvement, the assessment of whether the average fundamental frequency and speech rate in the speech features are consistently higher than the average level of the population while the facial expressions and body movements in the visual features remain calm includes: collecting human voices in the environment surrounding the intelligent digital human; calculating the average fundamental frequency and average speech rate of all detected human voices in the current time period based on the collected human voices to obtain a dynamic population speech feature baseline; comparing the average fundamental frequency and speech rate in the speech features with the dynamic population speech feature baseline; and determining the presence of physiological characteristics when the average fundamental frequency and speech rate in the speech features are consistently higher than the dynamic population speech feature baseline and the facial expressions and body movements in the visual features remain calm.
[0018] As a system extension, the assessment of whether the average fundamental frequency and speech rate in the speech features are consistently higher than the average level of the population, while the facial expressions and body movements in the visual features remain calm, also includes: dividing the facial and body regions in the video stream to obtain expression-sensitive regions and motion-sensitive regions; dynamically tracking pixel-level brightness, color, and texture changes in the expression-sensitive regions to obtain the amount of dynamic expression changes; performing detailed analysis of key point displacement and velocity in the motion-sensitive regions to obtain the amount of dynamic motion changes; generating a visual fluctuation marker when the amount of dynamic expression changes exceeds a preset expression threshold or the amount of dynamic motion changes exceeds a preset motion threshold; and determining the presence of physiological characteristics when the average fundamental frequency and speech rate in the speech features are consistently higher than the baseline of the dynamic population speech features, the facial expressions and body movements in the user's visual features remain calm, and no visual fluctuation marker is generated.
[0019] Secondly, this application also discloses an artificial intelligence-based multimodal interactive digital human training system, which includes: an acquisition and analysis module for acquiring multimodal interaction information of each user and performing multimodal interaction information analysis; the multimodal interaction information includes at least voice information, visual information, and environmental information; a determination module for determining the priority score of the interaction request corresponding to the multimodal interaction information of each user based on the multimodal interaction information analysis results; and a processing module for dynamically adjusting the allocation of computing resources and the order of communication responses based on the priority score, so as to prioritize the processing of user interaction requests with high priority scores.
[0020] Beneficial effects
[0021] This application discloses a multimodal interactive digital human training method and system based on artificial intelligence. By acquiring and analyzing the user's multimodal interaction information (including voice, vision, and environmental information), it can comprehensively perceive changes in the user's state and environment. Based on this, the method determines the priority score of each user interaction request according to the analysis results of the multimodal interaction information, and dynamically adjusts the allocation of computing resources and the order of communication responses according to this score. This mechanism can effectively solve the problems faced by existing intelligent digital human systems in complex public places such as transportation hubs, including multi-user concurrency, information timing disorder, environmental noise interference, and information association failure. By prioritizing the processing of user interaction requests with high priority scores, the method of this application can ensure that even in noisy and changing public environments, the system can still accurately understand the user's intent and provide timely and appropriate responses, significantly improving the interaction efficiency, accuracy, and user satisfaction of the intelligent digital human system, and overcoming the shortcomings of existing technologies such as inaccurate judgment of user intent and irrelevant responses. Attached Figure Description
[0022] Figure 1This is a flowchart illustrating the steps of the multimodal interactive digital human training method based on artificial intelligence disclosed in an embodiment of the present invention;
[0023] Figure 2 This is a schematic diagram of the structure of a multimodal interactive digital human training system based on artificial intelligence disclosed in an embodiment of the present invention. Detailed Implementation
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which these embodiments belong; the terminology used herein and in the specification of the application is for the purpose of describing particular embodiments only and is not intended to limit these embodiments; the terms "comprising" and "having," and any variations thereof, in the specification of these embodiments and the foregoing drawings, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification of these embodiments and the foregoing drawings are used to distinguish different objects, not to describe a particular order.
[0025] The implementation details of the technical solution in this embodiment are described in detail below:
[0026] Traditional intelligent digital human systems face numerous challenges when dealing with complex public environments with multiple users and high noise levels. These challenges include disordered timing of multimodal information, severe information pollution, and inaccurate judgment of user intent. These problems lead to the system's inability to accurately understand user intent, resulting in responses that are often irrelevant and severely impacting user experience.
[0027] In response, this application proposes a multimodal interactive digital human training method based on artificial intelligence, such as... Figure 1 As shown, the method includes:
[0028] S101, acquire multimodal interaction information for each user and perform multimodal interaction information analysis; the multimodal interaction information includes at least voice information, visual information, and environmental information;
[0029] S102, Based on the analysis results of multimodal interaction information, determine the priority score of the interaction request corresponding to the multimodal interaction information of each user;
[0030] S103, Based on the priority score, dynamically adjust the allocation of computing resources and the order of communication responses to prioritize the processing of user interaction requests with high priority scores.
[0031] This application proposes an AI-based multimodal interactive digital human training method. By acquiring and analyzing the multimodal interaction information of each user, it achieves a more comprehensive understanding of user input. Based on the analysis results, the method determines the priority score of each user's multimodal interaction requests, thereby intelligently prioritizing user requests in complex scenarios. Furthermore, based on the priority scores, it dynamically adjusts the allocation of computing resources and the order of communication responses to prioritize high-priority user interaction requests. This effectively solves the resource allocation and response efficiency problems in multi-user concurrent interactions, significantly improving the interactive performance and user satisfaction of the intelligent digital human in complex environments.
[0032] To better understand the technical solutions proposed in this application, some key terms are explained first. "Multimodal interaction information" refers to various types of information generated when a user interacts with the intelligent digital human system, such as voice information, visual information, and environmental information. Voice information can include the user's spoken expression, tone, and speaking speed; visual information can include the user's facial expressions, body movements, and eye movements; environmental information can refer to background sounds, lighting conditions, and surrounding objects during the interaction. "Multimodal interaction information analysis" refers to the process of comprehensively processing and understanding the above-mentioned multiple modalities of information, aiming to extract useful features from different dimensions to comprehensively grasp the user's intentions and context. "Priority scoring of interaction requests" is the quantitative evaluation of each user's interaction request based on the results of multimodal interaction information analysis, reflecting its urgency, importance, or relevance. "Computational resource allocation" refers to the dynamic allocation of computing resources within the intelligent digital human system, such as processor core time, memory buffers, and network bandwidth, to different user interaction requests according to their priority scores. "Communication response order" refers to the order in which the intelligent digital human system responds to different user interaction requests; higher-priority requests will receive faster responses.
[0033] The method proposed in this application first requires acquiring multimodal interaction information for each user. This can be achieved in various ways. For example, an intelligent digital human system can be equipped with multiple microphone arrays to capture the user's voice information. These microphones can be distributed in different locations to achieve sound source localization and separation. In another implementation, the system can utilize directional microphones or beamforming technology to focus on the voice of a specific user to reduce interference from environmental noise. For acquiring visual information, the system can integrate high-definition cameras to capture the user's facial expressions, body movements, and eye movements. These cameras can use wide-angle lenses to cover a larger interaction area or PTZ (Pan-Tilt-Zoom) cameras to track specific users. In addition, depth sensors or infrared sensors can be used to acquire the user's three-dimensional pose information. Environmental information can be acquired through various sensors such as environmental microphones, light sensors, and temperature and humidity sensors to perceive background sound, lighting conditions, and surrounding environmental events during the interaction. For example, environmental microphones can continuously monitor background noise levels, and light sensors can record changes in ambient brightness.
[0034] After acquiring multimodal interaction information, it is necessary to perform multimodal interaction information analysis on this information. For speech information, speech recognition technology can be used to convert it into text, and further speech features such as speech rate, pitch, volume, keywords, and emotional intensity can be extracted. These features can be analyzed using acoustic and language models. For example, speech rate can be obtained by calculating the number of syllables or words per unit time; pitch can be obtained through fundamental frequency analysis; volume can be measured through sound pressure level; keywords can be matched using a pre-set dictionary or topic model; and emotional intensity can be judged based on the prosody and spectral features of speech using emotion recognition algorithms. For visual information, computer vision technology can be used to extract visual features such as facial expressions, body movements, and eye movements. Facial expressions can be analyzed using facial keypoint detection and expression recognition models; body movements can be identified using skeleton extraction and pose estimation techniques; and eye movements can be obtained using eye-tracking technology. For environmental information, it can be integrated to obtain environmental event information. For example, by analyzing the sound collected by environmental microphones, environmental events such as background music, broadcast announcements, and crowd noise can be identified; and changes in ambient light can be determined using data from light sensors.
[0035] After completing the multimodal interaction information analysis, the priority score of the interaction request corresponding to each user's multimodal interaction information needs to be determined based on the analysis results. This can be done through various methods. For example, the urgency score of the interaction request can be determined based on the emotional intensity and keywords in the voice features. When the emotional intensity is high or contains keywords such as "urgent" or "help," the urgency score will increase accordingly. Simultaneously, the salience score of the interaction request can be determined based on the body language in the visual features and the relevance of the content pointed to by the body language to the environmental event information. For example, when a user points to a display screen that is playing an emergency notification, the salience score will increase. Furthermore, the contextual relevance score of the interaction request can be determined based on the degree of matching between the semantic content of the interaction request and the environmental event information. For example, when a user asks "When does the next train depart?" and the environmental information shows that they are currently at a train station, the contextual relevance score will be high. The basic need score of the interaction request can also be determined based on the preset classification of the interaction request. For example, querying flight information and seeking emergency medical assistance can be preset as high basic needs. Finally, the urgency score, salience score, contextual relevance score, and basic need score are weighted and calculated to obtain the final priority score. The weighting coefficients can be adjusted according to the needs of the actual application scenario.
[0036] Finally, based on priority scores, the allocation of computing resources and the order of communication responses are dynamically adjusted to prioritize user interaction requests with high priority scores. For example, for high-priority interaction requests, the system can allocate more processor core time, larger memory buffers, and higher network bandwidth to ensure they can be processed quickly. Simultaneously, the processing response time for intent determination can be shortened, allowing them to proceed to the subsequent response generation stage more quickly. For low-priority interaction requests, they can be placed in a waiting queue for delayed processing to avoid consuming critical resources. Furthermore, the response output scheduler can interrupt the generation of low-priority response messages and prioritize the generation and output of response messages for high-priority interaction requests. For example, when the intelligent digital human is responding to a low-priority request about weather inquiries, if it suddenly receives a high-priority request for emergency medical assistance, the system will immediately interrupt the weather inquiry response and prioritize processing and responding to the medical assistance request.
[0037] The proposed AI-based multimodal interactive digital human training method aims to address the challenges of multimodal information processing faced by existing intelligent digital human systems in complex public environments. Traditional intelligent digital human systems, in high-traffic public places such as transportation hubs, struggle to accurately link a user's voice with their corresponding body language due to issues such as simultaneous multi-user communication, environmental noise interference, and disordered information processing sequences. This leads to fundamental difficulties in understanding user intentions. Furthermore, the system introduces incorrect patterns and connections during learning, resulting in a decreased ability to judge user intentions. Ultimately, the digital human's responses are often irrelevant, severely impacting the user experience.
[0038] The core innovation of this application lies in the introduction of a priority scoring mechanism, which dynamically adjusts the allocation of computing resources and the order of communication responses. Specifically, this method acquires multimodal interaction information for each user, including voice, visual, and environmental information, and performs in-depth analysis on this information. For example, by extracting voice features from the voice information, it obtains voice features such as speech rate, pitch, volume, keywords, and emotional intensity; by extracting visual features from the visual information, it obtains visual features such as facial expressions, body movements, and eye movements; and by integrating environmental information, it obtains environmental event information. These detailed analysis results provide comprehensive data support for subsequent priority scoring.
[0039] Compared to existing technologies, the advantage of this application lies in that it no longer simply processes user requests according to the order of receipt or fixed rules. Instead, it assigns a priority score to each interaction request based on the analysis results of multimodal interaction information, comprehensively considering multiple dimensions such as the urgency, salience, contextual relevance, and basic needs of the interaction request. For example, when a user exhibits highly anxious emotions (judged by the intensity of emotional voice), while their body language points towards an emergency exit (judged by visual features), and environmental information indicates an emergency broadcast (judged by environmental event information), then that user's interaction request will be assigned a very high priority.
[0040] After determining the priority score, this application further dynamically adjusts the allocation of computing resources and the order of communication responses based on the score. This means that high-priority requests will receive more processor core time, memory buffers, and network bandwidth, thereby ensuring that they can be processed and responded to quickly. For example, for high-priority requests, the system will shorten the processing response time for intent judgment and prioritize the generation and output of response messages, and may even interrupt ongoing low-priority responses. This dynamic adjustment mechanism effectively solves the resource contention problem in multi-user concurrent interaction, avoids the delayed processing of important requests, and significantly improves the response efficiency and user interaction experience of the intelligent digital human system in complex public environments. In this way, this application can more accurately understand user intent and provide more timely and appropriate responses, thereby overcoming the limitations of existing technologies in handling complex multimodal interaction scenarios and improving the practicality and user satisfaction of the intelligent digital human system.
[0041] In some embodiments described above in this application, a method is proposed to acquire multimodal interaction information for each user and perform multimodal interaction information analysis. Specifically, the aforementioned multimodal interaction information analysis includes: extracting speech features from the speech information to obtain speech features; the speech features include speech rate, pitch, volume, keywords, and emotional intensity; extracting visual features from the visual information to obtain visual features; the visual features include facial expressions, body movements, and eye movements; and integrating the environmental information to obtain environmental event information.
[0042] Specifically, speech feature extraction refers to identifying and quantifying the acoustic and linguistic attributes of a user's speech input using acoustic models and natural language processing techniques. Speech rate reflects the user's rhythm of expression, pitch and volume reveal their emotional state or emphasis, keywords directly point to the user's intent or focus, and emotional intensity is quantified through acoustic features (such as fundamental frequency and formants) and sentiment lexicon analysis to determine the strength of the user's emotions. The extraction of these speech features aims to provide detailed linguistic and emotional dimension data for subsequent priority assessment of interactive requests.
[0043] Furthermore, visual feature extraction refers to the analysis of video streams or image information during user interaction using computer vision technology. Facial expressions can be identified through facial landmark detection and expression recognition algorithms to recognize emotional states such as joy, anger, and surprise; body movements are analyzed using skeleton tracking and behavior recognition technologies to understand nonverbal information such as user posture and gestures, including pointing gestures or the openness of body posture; eye tracking technology is used to determine the user's focus of attention, reflecting their attention distribution and points of interest. The extraction of these visual features aims to capture user interaction signals at the nonverbal level, providing important supplementary information for understanding user intentions.
[0044] Furthermore, environmental information integration refers to the collection and processing of contextual data from sensors, external systems, or pre-defined scenarios. Environmental event information can include the current time, location, weather, surrounding crowd activity, and device status, which helps to construct the specific context in which the user is situated. By integrating environmental information, intelligent digital humans can be provided with more comprehensive background knowledge, thereby more accurately understanding the deeper meaning and potential needs of user interaction requests.
[0045] This application's solution refines multimodal interaction information into quantifiable voice features, visual features, and environmental event information, enabling intelligent digital humans to comprehensively perceive and understand users' interaction intentions and contexts from multiple dimensions. Traditionally, single-modal analysis may lead to missing information or misjudgments; for example, voice alone may not accurately determine whether a user's rapid speech is due to physiological reasons or emotional excitement. By extracting speech rate, pitch, volume, keywords, and emotional intensity from voice information, a deeper analysis of the user's language expression and emotional state can be achieved. Simultaneously, by extracting facial expressions, body movements, and eye movements from visual information, nonverbal signals can be captured, compensating for potential deficiencies in voice information. Furthermore, by integrating environmental information and acquiring environmental event information, rich contextual background can be provided for user interactions. This multi-dimensional, fine-grained feature extraction and information integration lays a solid data foundation for accurately prioritizing user interaction requests, thereby improving the depth and accuracy of the intelligent digital human's understanding of user intentions.
[0046] Through the aforementioned technical solutions, intelligent digital humans can obtain richer and more refined multimodal user interaction data. This meticulous feature extraction and information integration significantly improves the accuracy and comprehensiveness of user intent recognition, avoiding misjudgments caused by insufficient information from a single modality. Specifically, by quantifying voice features, the user's emotions and level of urgency can be captured more accurately; by analyzing visual features, the user's nonverbal expressions and focus of attention can be understood more intuitively; and by integrating environmental information, the user's situation can be grasped more comprehensively. This provides more reliable and insightful data support for subsequent priority scoring, enabling intelligent digital humans to respond to user needs more intelligently and human-like, thus improving the quality and efficiency of the overall interactive experience.
[0047] Specifically, in the aforementioned AI-based multimodal interactive digital human training method, determining the priority score of the interaction request corresponding to each user's multimodal interaction information based on the multimodal interaction information analysis results includes:
[0048] The urgency score of the interaction request is determined based on the emotional intensity and keywords in the voice features.
[0049] The salience score of the interaction request is determined based on the body movements in the visual features and the correlation between the content indicated by the body movements and the environmental event information.
[0050] The contextual relevance score of the interaction request is determined based on the degree of matching between the semantic content of the interaction request and the environmental event information.
[0051] Based on the preset classification of the interaction request, determine the basic requirement score of the interaction request;
[0052] The priority score is obtained by weighting the urgency score, the salience score, the contextual relevance score, and the basic need score.
[0053] Specifically, the urgency score is determined by analyzing the emotional intensity and keywords in the speech features. Emotional intensity can be understood as the degree to which the speech conveys feelings of urgency, anxiety, or excitement, for example, through dramatic changes in tone, increased speech rate, or increased volume. Keywords refer to words in the speech information that clearly indicate urgency, such as "help," "urgent," or "hurry up." Through comprehensive analysis of these features, the immediate processing needs of interactive requests can be quantified.
[0054] The salience score is determined by evaluating body language in visual features and the relevance of the content the body language points to to environmental event information. Body language can include user gestures, pointing behaviors, or body postures, which may indicate the user's focus. When these body language actions point to a specific object or area in the environment, and that object or area is associated with currently detected environmental event information, such as a user pointing to a smoking appliance, the salience of the interaction request is considered high.
[0055] In practical applications, the contextual relevance score is determined by measuring the degree of matching between the semantic content of an interaction request and environmental event information. For example, if a user says "It's too hot here," and the environmental information shows an abnormally high indoor temperature, it indicates that the request is highly relevant to the current context. The higher this degree of matching, the higher the contextual relevance score, meaning that the request has a higher processing priority in the current environment.
[0056] Furthermore, the basic requirement score is determined based on preset categories of interaction requests. Certain types of requests, such as those involving security, health, or system-critical operations, may be preset to have a higher basic priority, regardless of their urgency, prominence, or contextual relevance. These preset categories can be configured according to application scenarios and business needs.
[0057] Finally, the priority score is obtained by weighting the urgency score, the salience score, the contextual relevance score, and the basic need score. This weighted calculation allows different types of scores to be assigned different levels of importance based on actual needs. For example, in emergency situations, the urgency score may be given higher weight, while in everyday interactions, the contextual relevance score may be more important.
[0058] This application's solution overcomes the limitations of single-modality or simple rule-based priority determination by comprehensively considering multi-dimensional information. By using emotional intensity and keywords from voice features to determine urgency scores, it captures the user's immediate emotional state and clear distress signals. By analyzing body language in visual features and its correlation with environmental event information to determine salience scores, the system can understand the user's focus and intentions in the physical space. By evaluating the matching degree between the semantic content of the interaction request and environmental event information to determine contextual relevance scores, it ensures that the digital human's perception and understanding of the current environment can effectively guide interaction processing. By introducing basic need scores, a safety net is provided for requests of critical or pre-defined importance. The weighted calculation of these scores allows the final priority score to flexibly reflect the complexity and diversity of user needs in different scenarios, thereby achieving more intelligent and human-centered response scheduling.
[0059] Through the aforementioned technical solutions, intelligent digital humans can perform more refined and multi-dimensional priority assessments of user interaction requests. This comprehensive scoring mechanism enables digital humans not only to recognize the surface-level language content of users but also to deeply understand their underlying emotions, behavioral intentions, and the environmental context, thereby avoiding misjudgments or response delays caused by incomplete information. As a result, digital humans can more accurately identify requests that truly require priority processing, significantly improving the efficiency, accuracy, and user satisfaction of interactions, especially effectively enhancing their interaction management capabilities in multi-user, high-concurrency, or complex scenarios.
[0060] Furthermore, this application proposes dynamically adjusting the allocation of computing resources and the order of communication responses based on the aforementioned priority scores, including: allocating computing resources to the interaction requests according to the priority scores; the computing resources include processor core time, memory buffers, and network bandwidth; shortening the processing response time for intent determination for interaction requests with high priority scores; placing interaction requests with low priority scores into a waiting queue for delayed processing; and interrupting the generation of low-priority response messages through a response output scheduler, and instead prioritizing the generation and output of response messages for interaction requests with high priority scores.
[0061] Specifically, allocating computing resources based on priority scores means that when an interaction request is assigned a higher priority score, the system allocates more processor core time, a larger memory buffer, and wider network bandwidth to it. For example, processor core time can be understood as the length of the time slice or the number of cores that the CPU can use to process the request; the memory buffer can be set up as a dedicated or extended area for storing data and model parameters related to the request; and network bandwidth can be prioritized to ensure that the data transmission of the request is not delayed.
[0062] Furthermore, for high-priority scoring interaction requests, shortening their intent judgment processing response time can be achieved in several ways. For example, relevant intent recognition models or knowledge bases can be preloaded for high-priority requests to reduce model loading time; or, a simplified but sufficiently accurate intent judgment algorithm can be used, sacrificing a small amount of accuracy for a faster response speed. Conversely, for low-priority scoring interaction requests, they can be placed in a waiting queue for delayed processing. This means that these requests will only be processed when system resources are relatively idle, or will only be activated after a preset waiting time threshold has been reached.
[0063] Furthermore, by interrupting the generation of low-priority response messages by the response output scheduler and prioritizing the generation and output of response messages for high-priority scoring interaction requests, the timely delivery of critical information is ensured. The response output scheduler can be configured to monitor all ongoing response generation tasks and their corresponding request priorities in real time. Once a higher-priority request requiring a response is detected, the scheduler will immediately pause or terminate the generation process of the current low-priority response and quickly initiate the generation and output process of the high-priority response.
[0064] This application's solution ensures that the intelligent digital human prioritizes user interaction requests with higher priority ratings by dynamically and intelligently allocating computing resources and flexibly adjusting the order of communication responses. High-priority requests receive more ample computing resource guarantees, and their intent-based processing time is effectively shortened, allowing for faster understanding and response. Simultaneously, the delayed processing of low-priority requests and the interruption mechanism for low-priority response messages prevent system resources from being occupied by non-urgent tasks for extended periods, enabling the system to concentrate limited resources on the most critical and urgent user needs.
[0065] Through the aforementioned technical solutions, the intelligent digital human system can significantly improve its real-time response capabilities and resource utilization efficiency in multi-user, high-concurrency scenarios. This solution ensures a smooth and satisfactory user experience, especially when handling urgent or important interactions, providing more timely and accurate feedback, thereby effectively avoiding user dissatisfaction or potential risks caused by improper resource allocation or response delays. This dynamic adjustment mechanism enables the intelligent digital human to more intelligently adapt to complex and ever-changing interactive environments, demonstrating higher interaction efficiency and a higher level of intelligence.
[0066] In some preferred embodiments, suppose the intelligent digital human is having a casual chat with one user (low-priority interaction). Suddenly, another user expresses an urgent need for help via voice and visual information (high-priority interaction). For example, the voice contains keywords such as "help" or "fire," and is delivered in a high-pitched, rapid tone, while the user visually displays an anxious facial expression and exaggerated body movements. According to the solution of this application, the system will immediately identify the high-priority score of the emergency help request. The response output scheduler will immediately interrupt the response message being generated for the casual chat user and place the casual chat request into a waiting queue. Simultaneously, the system will allocate more processor core time, memory buffer, and network bandwidth to the emergency help request, performing intent determination as quickly as possible and prioritizing the generation and output of response messages tailored to the emergency, such as providing emergency contact information or reassuring guidance, thereby ensuring that the emergency is handled promptly and effectively.
[0067] Furthermore, this application proposes the above-mentioned method of determining the priority score of interaction requests corresponding to each user's multimodal interaction information based on the analysis results of multimodal interaction information, and also includes:
[0068] Based on the analysis results of multimodal interaction information, the intent confidence score of the interaction request corresponding to the multimodal interaction information is calculated, and the consistency between different modal interaction information is detected.
[0069] When the confidence score of the intent is lower than a preset threshold or when there is a conflict in the consistency, the diagnostic process is initiated and the user intent is reassessed.
[0070] Based on the user intent determined through reassessment, the urgency score, the salience score, the contextual relevance score, and the basic need score are determined and weighted to obtain the priority score.
[0071] Specifically, the intent confidence score refers to the degree of certainty or reliability with which an intelligent digital human system identifies the intent of a user's interaction request. This score reflects the uncertainty the system faces in understanding the user's intent. For example, the intent confidence score may be low when the user's expression is ambiguous, contains multiple meanings, or does not conform to the system's known patterns. Detecting consistency between different modalities of interaction information involves comparing and verifying the user's intent pointed to by features extracted from speech, visual, and environmental information. For example, if the speech content expresses a positive emotion, but visual information (such as facial expressions or body language) shows negativity or confusion, there may be inconsistency. When the intent confidence score is below a preset threshold, indicating significant uncertainty in the system's understanding of the current user's intent, or when there are conflicts between different modalities, indicating contradictory information received by the system, the system will initiate a diagnostic process. This diagnostic process aims to clarify the user's intent through further analysis or interaction, such as guiding the user to provide more specific information by generating targeted counter-questions. After the diagnostic process is completed, the urgency score, salience score, contextual relevance score, and basic need score will be re-determined based on the user intent determined by the reassessment, and a weighted calculation will be performed to obtain a more accurate priority score.
[0072] This application's solution effectively addresses the problem of inaccurate user intent recognition in complex interaction scenarios by introducing intent confidence scoring and modal consistency detection mechanisms. Specifically, after the intelligent digital human receives multimodal interaction information from a user, it not only performs preliminary feature extraction and priority score calculation but also calculates its confidence in the user's intent in parallel and cross-validates whether there are contradictions between different modal information. The introduction of intent confidence scoring allows the system to quantify the certainty of its understanding of the user's intent, avoiding blind responses under conditions of high uncertainty. Modal consistency detection ensures that the system's understanding of the user's intent is comprehensive and conflict-free. For example, when the user's voice expression and body language are inconsistent, the system can promptly detect this inconsistency. Once low intent confidence or conflicts between modalities are detected, the system proactively initiates a diagnostic process, such as by asking questions or requesting clarification from the user, thereby obtaining more accurate and clearer user intent information. Thus, based on the re-evaluated user intent, the scores for each item can be calculated more accurately, and the final priority score can be obtained, ensuring the rationality and effectiveness of subsequent allocation of computing resources and the order of communication responses.
[0073] In some preferred embodiments, suppose a user, in a calm tone, says to the intelligent digital human, "I'm not feeling well, could you help me find a nearby pharmacy?" At this point, the intelligent digital human, through voice information analysis, initially identifies the user's intention to "find a pharmacy," but the confidence score of this intention may be moderate. Simultaneously, through visual information analysis, the system detects that the user's facial expression suggests slight pain, and their body movements appear slightly stiff, which is somewhat inconsistent with the calm tone. In this situation, because the confidence score of the intention does not reach a preset threshold and there is an inconsistency between modalities, the system will initiate a diagnostic process. The intelligent digital human may generate a targeted reverse question, such as, "Are you feeling unwell? Should I call a doctor or emergency contact for you?" The user might reply, "Yes, I'm a little dizzy, but right now I just want to know where the pharmacy is." Through this clarification, the intelligent digital human reassesses and confirms that the user's primary intention is "finding a pharmacy," but simultaneously identifies the secondary information that the user is feeling unwell. Based on the reassessed user intent, the system can more accurately determine the urgency score, salience score, contextual relevance score, and basic need score, and perform weighted calculations to obtain a more reasonable priority score. It can then prioritize providing users with information on nearby pharmacies and may also advise users to pay attention to their health.
[0074] Specifically, the above-mentioned calculation of the intent confidence score of the interaction request corresponding to the multimodal interaction information based on the analysis results of the multimodal interaction information includes the following steps:
[0075] Obtain the speech intent recognition confidence level corresponding to the speech features, the visual intent recognition confidence level corresponding to the visual features, and the context association confidence level corresponding to environmental event information;
[0076] The confidence scores for voice intent recognition, visual intent recognition, and contextual association are weighted and fused to calculate a comprehensive intent confidence score.
[0077] Voice intent recognition confidence refers to the reliability or accuracy of the intent recognition result obtained through in-depth analysis of voice information, such as using speech recognition technology to convert speech into text, and further using a natural language processing (NLP) model to identify intent within the text content. This confidence reflects the certainty with which the intelligent digital human understands the user's voice commands or questions. Visual intent recognition confidence refers to the reliability of judging the user's potential intent by analyzing visual information, such as recognizing facial expressions, body movements, and eye movements, and combining this with a pre-set behavioral pattern library or machine learning model. For example, when a user's facial expression shows confusion or their body points to a specific object, the system will assess the confidence of their visual intent. Contextual relevance confidence refers to assessing the relevance and reliability of the user's interaction request to the current context by integrating environmental information, such as recognizing the current scene, time, location, or surrounding events. For example, in a smart home environment, when a user issues a "open" command in the kitchen, the system will assess the confidence of the command's relevance to kitchen appliances.
[0078] Furthermore, the weighted fusion of the voice intent recognition confidence, the visual intent recognition confidence, and the context-related confidence refers to assigning different weights to each confidence level based on the importance of different modal information to user intent recognition in a specific context. These weighted confidence values are then summed or fused using other algorithms to obtain a single numerical value that comprehensively reflects the certainty of the user's intent. For example, in scenarios where voice commands are dominant, the voice intent recognition confidence can be given a higher weight; while in scenarios where visual interaction is more important, the visual intent recognition confidence may have a higher weight. In this way, the overall confidence of the user's intent can be assessed more comprehensively and accurately.
[0079] This application's solution overcomes the limitations or uncertainties that may exist with single-modal information by separately acquiring intent recognition confidence scores from three different modalities: speech, vision, and environmental information, and then weighted and fusing them. For example, when speech recognition confidence is low due to environmental noise, the high confidence scores of visual or environmental information can compensate for this deficiency, thereby improving the overall accuracy of intent judgment. This multimodal fusion mechanism enables intelligent digital humans to cross-validate user intent from multiple dimensions, thus forming a more robust judgment of user intent and calculating a more reliable intent confidence score in complex and ever-changing interactive environments. This provides a solid data foundation for subsequent intent confidence assessment and potential diagnostic processes.
[0080] In some embodiments described above, this application proposes calculating the intent confidence score of an interaction request based on multimodal interaction information analysis results and detecting the consistency between different modal interaction information. However, in practical applications, users may sometimes exhibit inconsistencies between modalities. For example, their facial expressions and body language may remain calm, but their speech features may show a high speech rate or tone. This may suggest that the user has underlying anxiety, urgency, or physiological characteristics. Simple modal consistency detection may not be able to effectively identify such deep-seated, non-obvious intent conflicts, leading to misjudgment of the user's true intent and affecting the accuracy of priority scoring and subsequent resource allocation and response efficiency.
[0081] To address this, this application further proposes detecting the consistency between different modal interaction information, including: comparing the consistency between the speech features and the visual features; initiating the diagnostic process and reassessing the user's intent, including: assessing whether the average fundamental frequency and speech rate in the speech features are consistently higher than the average level of the population when the facial expressions and body movements in the visual features remain calm; if so, generating a targeted counter-question; and reassessing the user's intent based on the user's response to the targeted counter-question.
[0082] Specifically, comparing the consistency between the speech features and the visual features involves analyzing speech features extracted from speech information (such as speech rate, pitch, volume, keywords, and emotional intensity) and visual features extracted from visual information (such as facial expressions, body movements, and eye movements) to determine whether the information conveyed by different modalities is mutually supportive or conflicting. For example, if the speech expresses an urgent request, but the user appears relaxed visually, this is a potential inconsistency. Specifically, assessing whether the average fundamental frequency and speech rate in the speech features are consistently higher than the population average when the facial expressions and body movements in the visual features remain calm aims to identify a specific modal conflict pattern. A consistently calm facial expression and body movements can be understood as a period of time during which the user's facial muscle activity and body posture changes are relatively small, without showing obvious excitement, tension, or large movements. Simultaneously, a consistently higher average fundamental frequency and speech rate in the speech features means that the user's voice frequency and speaking speed are significantly higher than the average of the general population or their own historical average over a period of time. This combination may indicate that the user is experiencing some degree of physiological or psychological tension or urgency, but is attempting to conceal it through external behavior. When a specific modal inconsistency is detected, the system generates targeted counter-questions. Targeted counter-questions involve posing guiding or clarifying questions to the user based on the detected modal conflict, directly probing the user's true intentions or potential needs. For example, questions might include, "You sound a bit rushed, is there some emergency?" or "You seem calm, but you're speaking a bit fast, is there anything I should pay special attention to?" The aim is to eliminate the uncertainty caused by the modal conflict through direct user feedback. Based on the user's response to the targeted counter-questions, the system reassesses the user's intent, thereby gaining a more accurate understanding of that intent.
[0083] This application's solution effectively addresses the potential misjudgment of user intent in basic solutions by introducing a mechanism for recognizing specific modal conflict patterns. Specifically, when the intelligent digital human observes that a user's facial expressions and body movements remain consistently calm, but their average fundamental frequency and speech rate in their speech characteristics are consistently higher than the average level of the population, the system can identify this as a potential signal of physiological or psychological tension, rather than a simple modal consistency or inconsistency. It is precisely this meticulous recognition that enables the system to proactively generate targeted counter-questions, directly seeking clarification from the user. Through the user's response, the system obtains more direct and accurate intent information, thus avoiding the bias that may result from judging solely based on surface modal information and ensuring the accuracy of subsequent intent assessment.
[0084] In some preferred embodiments, a specific example is given below. Suppose a user is communicating with an intelligent digital human. Their facial expressions and body movements are identified as consistently calm in the video stream, without any obvious changes in expression or large movements. However, by analyzing the user's voice information, the system detects that their average fundamental frequency and speech rate are consistently higher than the system's preset average level for the general population. At this point, the system recognizes the inconsistency between this visual calmness and the abnormal speech, and determines that the user may be experiencing potential urgency or tension. To accurately assess the user's intent, the intelligent digital human immediately generates and outputs a targeted reverse question to the user, such as: "Hello, I noticed that your speech rate is a bit fast. Is there anything urgent that you need my assistance with?" Upon receiving this question, the user might reply: "Yes, I need to check my flight information immediately because my plane is about to take off." Based on the user's explicit response, the system can reassess and accurately identify the urgency of the user's current intent, thereby adjusting the priority score of the interaction request to high priority and immediately allocating computing resources for processing to ensure that the user's needs are responded to in a timely manner.
[0085] Furthermore, this application proposes a more precise evaluation method, in which the average fundamental frequency and speech rate in the speech features are consistently higher than the average level of the population when the facial expressions and body movements in the visual features remain calm, including:
[0086] Collect human voices in the environment surrounding the intelligent digital human, and calculate the average fundamental frequency and average speech rate of all detected human voices in the current time period in real time based on the collected human voices to obtain a dynamic crowd speech feature baseline;
[0087] The average fundamental frequency and speech rate in the speech features are compared with the dynamic crowd speech feature baseline; and when the average fundamental frequency and speech rate in the speech features are consistently higher than the dynamic crowd speech feature baseline, and the facial expressions and body movements in the visual features remain calm, it is determined that there is a physiological characteristic influence.
[0088] Specifically, in this embodiment, "collecting human voices in the environment surrounding the intelligent digital human" refers to continuously and uninterruptedly acquiring all detectable human voice signals within the physical space where the intelligent digital human is located, using the microphone array or other audio acquisition devices equipped with it. The purpose is to provide a real-time, localized data source for constructing a dynamic speech feature baseline.
[0089] The phrase "calculating the average fundamental frequency and average speaking rate of all detected human voices in real time within the current time period based on the collected human voices to obtain a dynamic crowd speech feature baseline" can be understood as follows: the system processes the real-time collected human voices, extracts speech features such as fundamental frequency and speaking rate, and statistically averages these features within the current time window to establish a dynamic reference standard reflecting the current environment and the speech habits of the crowd. This dynamic crowd speech feature baseline is adaptive and can be updated in real time according to changes in the environment or the interacting crowd, aiming to provide a more context-relevant comparison benchmark.
[0090] In practical applications, "comparing the average fundamental frequency and speech rate in the speech features with the dynamic crowd speech feature baseline" means quantitatively comparing the average fundamental frequency and speech rate exhibited by the current user during the interaction with the dynamic crowd speech feature baseline calculated in real time. This comparison aims to determine whether the user's speech features significantly deviate from the "normal" level in the current environment.
[0091] Therefore, the statement "when the average fundamental frequency and speech rate in the speech features are consistently higher than the baseline of the dynamic crowd speech features, and the facial expressions and body movements in the visual features remain calm, it is determined that there is a physiological influence" means that if a user's speech exhibits a consistently high fundamental frequency and high speech rate, while their facial expressions and body movements remain calm, this inconsistency between multimodal information is considered a strong indication of potential physiological influences. The purpose of this judgment is to more accurately identify potential tension, anxiety, or other non-emotional physiological states in the user, thereby triggering subsequent diagnostic procedures.
[0092] This application's solution effectively addresses the misjudgment problem that can arise from traditional fixed baselines by introducing a dynamic crowd speech feature baseline. Specifically, by real-time acquisition of human voices in the environment surrounding the intelligent digital human and calculating a dynamic baseline, the system can establish a more adaptive and accurate reference standard based on the current acoustic environment and the speech habits of the crowd. When a user's speech features (average fundamental frequency and speech rate) are compared with this dynamic baseline, it can more accurately determine whether their speech is truly "consistently above the average level of the crowd," rather than simply due to environmental noise or individual differences. This dynamically adjusted comparison mechanism makes the detection of physiological characteristics more robust and reliable, avoids misdiagnosis caused by baseline mismatch, and thus improves the accuracy of the intelligent digital human's understanding of the user's state.
[0093] Through the above technical solution, this application can significantly improve the accuracy and robustness of intelligent digital humans in detecting the impact of user physiological characteristics. Compared with using a fixed average level of the population as a baseline, the introduction of a dynamic population voice feature baseline allows the system to adaptively adjust the evaluation criteria, effectively avoiding the risk of misjudgment caused by environmental changes, background noise, or individual differences in vocal habits. As a result, intelligent digital humans can more accurately identify the potential physiological characteristics in multimodal interaction information, thereby generating more targeted reverse questions in the subsequent diagnostic process, ultimately achieving a more accurate assessment of user intent and improving the intelligence level and user experience of intelligent digital human interaction.
[0094] In some preferred embodiments, a specific example is given below. Suppose an intelligent digital human is interacting with multiple users in an open-plan office environment. In this environment, due to background conversations and occasional ambient noise, the overall speech frequency and rate may be slightly higher than in a quiet private space. If a fixed average level of a group of people, preset based on a quiet environment, is used as a baseline, then many users' normal conversations may be misinterpreted as "consistently above the average level of the group," thus incorrectly triggering the diagnostic process.
[0095] However, according to the scheme of this application, the intelligent digital human continuously collects human voices in the open office environment and calculates in real time the average fundamental frequency and average speech rate of all detected human voices in the current time period, thereby establishing a dynamic baseline of human voice characteristics that reflects the characteristics of the current office environment. When a user interacts with the intelligent digital human, even if their voice characteristics are slightly higher than a general static average, as long as they do not consistently exceed the currently dynamically calculated baseline of the office environment, and their facial expressions and body movements remain calm, the system will not incorrectly determine that there is a physiological influence. Conversely, if the user maintains visual calmness while their voice characteristics (such as average fundamental frequency and speech rate) consistently and significantly exceed the current dynamic human voice characteristic baseline, the system can accurately identify that this may be a physiological influence and initiate corresponding diagnostic processes, such as generating targeted reverse questions to further confirm the user's state. This dynamic adaptability enables the intelligent digital human to understand the user's true intentions and state in complex environments more intelligently and accurately.
[0096] In some embodiments of this application, when the average fundamental frequency and speech rate in the speech features are consistently higher than the baseline of dynamic crowd speech features, and facial expressions and body movements in the visual features remain consistently calm, it is determined that there is an influence of physiological characteristics. However, in practical applications, simply judging "consistent calmness" of facial expressions and body movements may have certain ambiguity or inaccuracy. For example, users may exhibit slight, non-emotional changes in their face or body due to non-physiological reasons (such as high concentration, slight discomfort, etc.), but these changes are not sufficient to be simply classified as "uncalm," but may affect the accurate judgment of the influence of physiological characteristics. If the above problems are not addressed, it may lead to misjudgment of user intentions, thereby affecting the accuracy of intelligent digital human responses and user experience. In this regard, this application further proposes a more refined visual feature analysis method to more accurately assess the calmness state of user visual features, thereby improving the accuracy of judging the influence of physiological characteristics.
[0097] The above assessment of whether the average fundamental frequency and speech rate in the speech features are consistently higher than the average level of the population when the facial expressions and body movements in the visual features remain calm further includes: dividing the facial and body regions in the video stream to obtain expression-sensitive regions and motion-sensitive regions; dynamically tracking pixel-level brightness, color, and texture changes in the expression-sensitive regions to obtain the amount of dynamic change in facial expressions; performing fine analysis of key point displacement and velocity in the motion-sensitive regions to obtain the amount of dynamic change in motion; generating a visual fluctuation marker when the amount of dynamic change in facial expressions exceeds a preset expression threshold or the amount of dynamic change in motion exceeds a preset motion threshold; and determining the presence of physiological characteristics when the average fundamental frequency and speech rate in the speech features are consistently higher than the baseline of the dynamic population speech features, and the user's facial expressions and body movements in the visual features remain calm, and no visual fluctuation marker is generated.
[0098] Specifically, region segmentation of the face and limbs in a video stream involves using image processing and computer vision techniques to identify and select specific areas of the face (e.g., eyes, eyebrows, mouth) as expression-sensitive areas and major body parts (e.g., arms, hands, head, torso) as motion-sensitive areas. This region segmentation aims to focus on visual information most relevant to changes in emotion and physiological state. Dynamic tracking of pixel-level brightness, color, and texture changes in these expression-sensitive areas involves analyzing changes in pixel values over time to quantify subtle dynamics of facial expressions. For example, the difference in pixel values between consecutive frames, the average rate of change in brightness of local areas, or the fluctuation amplitude of texture features (e.g., Gabor filter response) can be calculated to obtain the amount of dynamic change in facial expressions. This amount of change reflects the activity level of facial expressions. Simultaneously, fine analysis of keypoint displacement and velocity in the motion-sensitive areas involves using skeleton tracking or keypoint detection algorithms to identify and track the spatial position changes and movement speed of limb keypoints (e.g., joints, hand feature points). By calculating the displacement distance, velocity vector, or acceleration of these key points across consecutive frames, the dynamic changes in movement can be obtained. These changes reflect the activity level of the body movements. In practical applications, the preset facial expression threshold and preset movement threshold are set based on extensive user behavior data and expert experience to distinguish between normal, non-emotional minor fluctuations and significant changes that truly indicate emotion or a non-calm state. When the dynamic changes in facial expression exceed the preset facial expression threshold, or the dynamic changes in movement exceed the preset movement threshold, it indicates the presence of non-calm visual activity, and a visual fluctuation marker is generated. Therefore, only when the average fundamental frequency and speech rate in the speech features are consistently higher than the baseline of the dynamic crowd speech features, and the user's facial expressions and body movements in the visual features remain calm, and no visual fluctuation marker is generated, is it finally determined that there is an influence of physiological characteristics. This means that only when the speech features show an abnormal increase, and visual analysis confirms that there are no significant facial or movement fluctuations, is it considered to be caused by physiological characteristics.
[0099] This application's solution addresses the potential ambiguity in traditional methods' "continuous calmness" assessment by introducing refined dynamic analysis of visual features. Specifically, by dynamically tracking pixel-level brightness, color, and texture changes in expression-sensitive and motion-sensitive areas, as well as performing refined analysis of keypoint displacement and velocity, it can capture even subtle, non-emotional dynamic changes on the user's face and limbs. When these dynamic changes exceed a preset threshold, a visual fluctuation marker is generated, clearly indicating a visually non-calm state. Only when voice features exhibit a consistently higher-than-average level, and after this refined analysis confirms no significant visual fluctuations (i.e., no visual fluctuation marker is generated), can the illusion of visual "calmness" caused by non-physiological factors such as concentration or slight discomfort be more reliably ruled out, thus more accurately attributing the abnormality of voice features to physiological characteristics.
[0100] The above technical solution significantly improves the accuracy and robustness of judging the impact of user physiological characteristics. This solution avoids misjudgments caused by inaccurate assessment of a visually "calm" state, ensuring that abnormalities in voice features are only attributed to physiological characteristics after truly excluding other non-physiological interference factors. This allows intelligent digital humans to more accurately understand the user's true state, avoiding incorrectly initiating intent diagnosis processes or generating inappropriate responses when the user's voice is abnormal due to physiological reasons, thereby enhancing the intelligence level of intelligent digital human interaction and the user experience.
[0101] In some preferred embodiments, a specific example is given below. Suppose an intelligent digital human is communicating with a user. During a certain time period, the system detects that the user's average fundamental frequency and speech rate are consistently higher than the dynamic crowd speech characteristic baseline. At this point, the system initiates a detailed evaluation of the user's visual characteristics. First, using facial recognition and skeleton tracking technology, the user's facial area in the video stream is divided into expression-sensitive areas, and active areas such as the arms and hands are divided into motion-sensitive areas. Next, the system continuously tracks minute changes in the brightness, color, and texture of pixels within the expression-sensitive areas between consecutive frames and calculates the amount of dynamic change in expression. Simultaneously, a detailed analysis of the displacement and velocity of key points (such as the elbow and wrist joints) within the motion-sensitive areas is performed to calculate the amount of dynamic change in motion. For example, if the user is thinking, their facial muscles may twitch slightly, or their hands may unconsciously sway slightly; these minute changes are calculated as expression dynamic changes or motion dynamic changes. The system then compares these dynamic changes with preset expression thresholds and motion thresholds. If none of these dynamic changes exceed their respective thresholds, it indicates that the user is indeed in a highly calm visual state and no visual fluctuation markers are generated. In this case, combined with the anomalies in voice features, the system can more confidently determine that the user is currently affected by physiological characteristics, such as tension, fatigue, or other physiological factors that may cause voice changes, rather than emotional fluctuations or intentional expression.
[0102] Furthermore, specific embodiments of this application also disclose a multimodal interactive digital human training system based on artificial intelligence, such as... Figure 2 As shown, the system includes:
[0103] The acquisition and analysis module 201 is used to acquire multimodal interaction information for each user and perform multimodal interaction information analysis; the multimodal interaction information includes at least voice information, visual information, and environmental information;
[0104] The determination module 202 is used to determine the priority score of the interaction request corresponding to the multimodal interaction information of each user based on the analysis results of the multimodal interaction information.
[0105] The processing module 203 is used to dynamically adjust the allocation of computing resources and the order of communication responses based on the priority scores, so as to prioritize the processing of user interaction requests with high priority scores.
[0106] Based on this, this application proposes an AI-based multimodal interactive digital human training system. Through modular design, it achieves effective acquisition, analysis, priority evaluation, and dynamic resource adjustment of user multimodal interaction information, thereby enabling a more accurate understanding of user intent and providing timely and appropriate responses. This significantly improves the interactive performance and user satisfaction of the intelligent digital human in complex environments. The system acquires and analyzes the multimodal interaction information of each user through an acquisition and analysis module, enabling a more comprehensive understanding of user input. Furthermore, a determination module determines the priority score of the interaction requests corresponding to each user's multimodal interaction information based on the analysis results, thus intelligently prioritizing user requests in complex scenarios. Finally, a processing module dynamically adjusts the allocation of computing resources and the order of communication responses based on the priority scores, prioritizing user interaction requests with high priority scores, effectively solving the problems of resource allocation and response efficiency in multi-user concurrent interaction.
[0107] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A multimodal interactive digital human training method based on artificial intelligence, characterized in that, The method includes: Acquire multimodal interaction information for each user and perform multimodal interaction information analysis; the multimodal interaction information includes at least voice information, visual information, and environmental information; Based on the analysis results of multimodal interaction information, the priority score of the interaction request corresponding to each user's multimodal interaction information is determined; Based on the priority scores, the allocation of computing resources and the order of communication responses are dynamically adjusted to prioritize the processing of user interaction requests with high priority scores. The process of performing multimodal interaction information analysis includes: Speech features are extracted from the speech information to obtain speech features; the speech features include speech rate, pitch, volume, keywords, and emotional intensity; Visual features are extracted from the visual information to obtain visual features; the visual features include facial expressions, body movements, and eye movements. The environmental information is integrated to obtain environmental event information; The step of determining the priority score of the interaction request corresponding to each user's multimodal interaction information based on the analysis results of multimodal interaction information includes: The urgency score of the interaction request is determined based on the emotional intensity and keywords in the voice features. The salience score of the interaction request is determined based on the body movements in the visual features and the correlation between the content indicated by the body movements and the environmental event information. The contextual relevance score of the interaction request is determined based on the degree of matching between the semantic content of the interaction request and the environmental event information. Based on the preset classification of the interaction request, determine the basic requirement score of the interaction request; The priority score is obtained by weighting the urgency score, the salience score, the contextual relevance score, and the basic need score.
2. The multimodal interactive digital human training method based on artificial intelligence according to claim 1, characterized in that, The step of dynamically adjusting the allocation of computing resources and the order of communication responses based on the priority score includes: Based on the priority score, computing resources are allocated to the interaction request; the computing resources include processor core time, memory buffer, and network bandwidth. For high-priority scoring interaction requests, shorten the processing response time for intent determination; for low-priority scoring interaction requests, place them in a waiting queue for delayed processing; and, through the reply output scheduler, interrupt the generation of low-priority reply messages and instead prioritize the generation and output of reply messages for high-priority scoring interaction requests.
3. The multimodal interactive digital human training method based on artificial intelligence according to claim 1, characterized in that, The step of determining the priority score of the interaction request corresponding to each user's multimodal interaction information based on the analysis results of multimodal interaction information also includes: Based on the analysis results of multimodal interaction information, the intent confidence score of the interaction request corresponding to the multimodal interaction information is calculated, and the consistency between different modal interaction information is detected. When the confidence score of the intent is lower than a preset threshold or when there is a conflict in the consistency, the diagnostic process is initiated and the user intent is reassessed. Based on the user intent determined through reassessment, the urgency score, the salience score, the contextual relevance score, and the basic need score are determined and weighted to obtain the priority score.
4. The multimodal interactive digital human training method based on artificial intelligence according to claim 3, characterized in that, The step of calculating the intent confidence score of the interaction request corresponding to the multimodal interaction information based on the analysis results of the multimodal interaction information includes: Obtain the speech intent recognition confidence level corresponding to the speech features, the visual intent recognition confidence level corresponding to the visual features, and the context association confidence level corresponding to environmental event information; The confidence scores for voice intent recognition, visual intent recognition, and contextual association are weighted and fused to calculate a comprehensive intent confidence score.
5. The multimodal interactive digital human training method based on artificial intelligence according to claim 3, characterized in that, The detection of consistency between different modal interaction information includes: comparing the consistency between the speech features and the visual features; The initiation of the diagnostic process and reassessment of user intent includes: When the facial expressions and body movements in the visual features remain calm, assess whether the average fundamental frequency and speech rate in the speech features are consistently higher than the average level of the population. If so, a targeted counter-question is generated; based on the user's response to the targeted counter-question, the user's intent is reassessed.
6. The multimodal interactive digital human training method based on artificial intelligence according to claim 5, characterized in that, The assessment of whether the average fundamental frequency and speech rate in the speech features are consistently higher than the average level of the population when the facial expressions and body movements in the visual features remain calm includes: Collect human voices in the environment surrounding the intelligent digital human, and calculate the average fundamental frequency and average speech rate of all detected human voices in the current time period in real time based on the collected human voices to obtain a dynamic crowd speech feature baseline; The average fundamental frequency and speech rate in the speech features are compared with the dynamic crowd speech feature baseline; and when the average fundamental frequency and speech rate in the speech features are consistently higher than the dynamic crowd speech feature baseline, and the facial expressions and body movements in the visual features remain calm, it is determined that there is a physiological characteristic influence.
7. The multimodal interactive digital human training method based on artificial intelligence according to claim 6, characterized in that, The assessment of whether the average fundamental frequency and speech rate in the speech features are consistently higher than the average level of the population when the facial expressions and body movements in the visual features remain calm also includes: The facial and limb regions in the video stream are divided into expression-sensitive regions and motion-sensitive regions. The expression-sensitive region is dynamically tracked for pixel-level brightness, color, and texture changes to obtain the amount of dynamic expression changes; the motion-sensitive region is finely analyzed for key point displacement and velocity to obtain the amount of dynamic motion changes. When the amount of dynamic change in facial expression exceeds a preset facial expression threshold, or the amount of dynamic change in action exceeds a preset action threshold, a visual fluctuation marker is generated; When the average fundamental frequency and speech rate in the speech features are consistently higher than the baseline of the dynamic crowd speech features, and the user's facial expressions and body movements in the visual features remain calm and no visual fluctuation markers are generated, it is determined that there is an influence of physiological characteristics.
8. A multimodal interactive digital human training system based on artificial intelligence, characterized in that, The system includes: The acquisition and analysis module is used to acquire multimodal interaction information for each user and perform multimodal interaction information analysis; the multimodal interaction information includes at least voice information, visual information, and environmental information; The process of performing multimodal interaction information analysis includes: Speech features are extracted from the speech information to obtain speech features; the speech features include speech rate, pitch, volume, keywords, and emotional intensity; Visual features are extracted from the visual information to obtain visual features; the visual features include facial expressions, body movements, and eye movements. The environmental information is integrated to obtain environmental event information; The determination module is used to determine the priority score of the interaction request corresponding to each user's multimodal interaction information based on the analysis results of multimodal interaction information; The step of determining the priority score of the interaction request corresponding to each user's multimodal interaction information based on the analysis results of multimodal interaction information includes: The urgency score of the interaction request is determined based on the emotional intensity and keywords in the voice features. The salience score of the interaction request is determined based on the body movements in the visual features and the correlation between the content indicated by the body movements and the environmental event information. The contextual relevance score of the interaction request is determined based on the degree of matching between the semantic content of the interaction request and the environmental event information. Based on the preset classification of the interaction request, determine the basic requirement score of the interaction request; The priority score is obtained by weighting the urgency score, the salience score, the contextual relevance score, and the basic need score. The processing module is used to dynamically adjust the allocation of computing resources and the order of communication responses based on the priority scores, so as to prioritize the processing of user interaction requests with high priority scores.
Citation Information
Patent Citations
Multi-person participation-based man-machine interaction method and apparatus
CN107831903A
Blind person auxiliary interaction method and system based on multi-modal large model and storage medium
CN119537909A
Intelligent doll-oriented multi-modal data processing task scheduling optimization method and system
CN120256146A