An AI-based virtual digital human interaction system
By analyzing voice, text, facial and body data, dynamically adjusting the performance of virtual digital people, the problem that existing systems are difficult to accurately judge user emotions and attention under multi-modal input, and improve the naturalness of interaction and teaching effect.
Patent Information
- Application Number
- CN202510408504.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2045-04-02
AI Technical Summary
When facing complex multimodal inputs, the existing virtual digital human interaction system is difficult to accurately judge the user's emotional changes and attention state, which affects the teaching effect.
The interactive data is analyzed through the feature extraction module, combined with voice, text, facial and body data, the voice emotion type, text emotion type, facial emotion type and attention state are calculated, and the optimization and adjustment module dynamically adjusts the voice, text and visual performance of virtual digital people to improve emotional resonance and interaction nature.
It realizes accurate judgment of user emotions and attention, and improves the naturalness of interaction between virtual digital people and users and the teaching effect.
Smart Images

Figure CN119902625B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to an AI-based virtual digital human interaction system. Background Art
[0002] Currently, virtual digital human interaction systems are widely used in the field of education. They mainly construct virtual character images based on technologies such as computer graphics, natural language processing, and deep learning to achieve intelligent interaction. Among them, the generation of virtual digital humans usually relies on 3D modeling, motion capture, and speech synthesis technologies. The appearance of the character is constructed through modeling tools, and the rendering engine is used to enhance the realism. In terms of interaction, natural language processing technology enables virtual digital humans to understand and respond to the text or speech input by users. Emotion recognition algorithms further enhance their interactive experience, enabling them to adjust intonation, expressions, and even actions according to the emotions of users.
[0003] Although existing virtual digital humans can recognize text or voice commands, there are still limitations in the understanding of real-time, multi-modal inputs. For example, in an educational scenario, the speech, tone, facial expressions, and body movements of students can usually convey complex emotions and needs. Existing systems often only rely on voice or text inputs for response. This single-modal processing method makes it difficult for virtual digital humans to make accurate judgments in the face of complex interaction scenarios, and they may not be able to accurately perceive the attention state or emotional changes of students, affecting the teaching effect. Summary of the Invention
[0004] The purpose of the present invention is to provide an AI-based virtual digital human interaction system, aiming to solve the problems mentioned in the background art.
[0005] To solve the above technical problems, the technical solution of the present invention is as follows:
[0006] An AI-based virtual digital human interaction system, the system includes:
[0007] A feature extraction module, used to extract features from the preprocessed interaction data set to obtain a user feature data set;
[0008] A speech analysis module, used to calculate the audio change rate and speech rate stability according to the user feature data set, and analyze the speech emotion type of the user to obtain speech emotion data;
[0009] A text analysis module, used to calculate the syntactic complexity and count the number of sentiment words according to the user feature data set, and analyze the text emotion type of the user to obtain text emotion data;
[0010] A facial analysis module, used to calculate the facial expression symmetry degree according to the user feature data set, and analyze the facial emotion type of the user to obtain facial emotion data;
[0011] The limb analysis module is used to calculate body stability, identify the visual fixation area, analyze the user's attention state, and obtain attention state data according to the user feature dataset.
[0012] The optimization and adjustment module is used to adjust the pitch and speech rate of the virtual digital human according to the speech emotion data; optimize the sentence complexity and emotional words in the virtual digital human's text according to the text emotion data; adjust the changes in the eyes, eyebrows, and corners of the mouth of the virtual digital human according to the facial emotion data; adjust the information display order and visual focus guidance according to the attention state data.
[0013] Further, the feature extraction module includes:
[0014] The speech feature extraction unit is used to calculate the fundamental frequency curve of the speech data, obtain the speech fundamental frequency feature, perform frequency domain conversion on it, obtain the speech time-frequency feature, and calculate the energy value of the speech data according to the speech time-frequency feature to obtain the speech energy feature according to the interaction dataset.
[0015] The text feature extraction unit is used to perform structural analysis on the text data, extract the basic structural information of the text data, obtain the text structure feature, and extract the emotional words in the text data according to it to obtain the text emotional word feature according to the interaction dataset.
[0016] The facial feature extraction unit is used to extract the position information of the eyes, eyebrows, corners of the mouth, and facial contour, identify the key points of the facial data, obtain the facial key point feature, and analyze the displacement of the key points between adjacent time frames to obtain the expression change feature according to the interaction dataset.
[0017] The limb feature extraction unit is used to extract the position information of the head, arms, torso, and legs, obtain the limb joint point feature, and draw the trajectory curve of the user's actions according to it, analyze the user's movement pattern, and obtain the limb action trajectory feature according to the interaction dataset.
[0018] Further, the speech analysis module includes:
[0019] The audio change calculation unit is used to analyze the change trend of the speech fundamental frequency, calculate the audio change rate of the speech fundamental frequency, and obtain the speech fundamental frequency change data according to the speech fundamental frequency feature.
[0020] The time series unit is used to calculate the energy value of the speech data within a fixed time window according to the speech energy feature to obtain the speech energy time series data.
[0021] The energy fluctuation calculation unit is used to calculate the energy change rate between consecutive time windows according to the speech energy time series data to obtain the speech energy fluctuation data.
[0022] A speech rate stability analysis unit, which is used to calculate the standard deviation of the energy fluctuation of speech data according to the speech energy fluctuation data, analyze the degree of fluctuation of the energy value over time, and obtain the speech rate stability data;
[0023] A speech emotion judgment unit, which is used to judge the type of the user's speech emotion according to the speech fundamental frequency change data and the speech rate stability data, and obtain the speech emotion data.
[0024] Further, the speech rate stability analysis unit includes:
[0025] A standard deviation calculation unit, which is used to calculate the mean value of the energy fluctuation of speech data according to the speech energy fluctuation data, and calculate the standard deviation of the speech data according to the mean value of the energy fluctuation, so as to obtain the speech energy standard deviation data;
[0026] A normalization processing unit, which is used to perform normalization processing on the speech energy standard deviation data to obtain the normalized speech rate stability data;
[0027] A speech rate evaluation unit, which is used to judge the speech rate stability of the user according to the normalized speech rate stability data. When the normalized speech rate stability data tends to 1, the speech rate tends to be stable; when the normalized speech rate stability data tends to 0, the speech rate tends to be unstable.
[0028] Further, the text analysis module includes:
[0029] A text syntax analysis unit, which is used to perform syntax analysis on the text data according to the text structure characteristics, calculate the syntax complexity, and obtain the text syntax data;
[0030] A text sentiment word analysis unit, which is used to count the number of sentiment words in the text data according to the text sentiment word characteristics, and calculate the sentiment word frequency, so as to obtain the text sentiment word data;
[0031] A text emotion bias calculation unit, which is used to analyze the emotion tendency of the text data according to the text syntax data and the text sentiment word data, calculate the text emotion bias, and obtain the text emotion bias data;
[0032] A text emotion judgment unit, which is used to judge the type of the user's text emotion according to the text emotion bias data, and obtain the text emotion data.
[0033] Further, the text emotion bias calculation unit includes:
[0034] An emotion score calculation unit, which is used to calculate the emotion score of the sentiment word according to the text sentiment word data and the preset emotion dictionary;
[0035] A sentiment word weight calculation unit, which is used to determine the text weight of the sentiment word in the text data according to the emotion score;
[0036] The sentence semantic consistency calculation unit is used to determine the sentence weight of the sentiment word in the sentence according to the text syntactic data, and calculate the sentence semantic consistency of the sentiment word in the sentence;
[0037] The text emotion deviation calculation formula unit is used to calculate the text emotion deviation according to the sentence semantic consistency. The calculation formula of the text emotion deviation is:
[0038] ,
[0039] wherein, is the text emotion deviation, is the text weight of the th sentiment word, is the emotion score of the th sentiment word, is the total number of sentiment words, is the index of the sentiment word, is the total number of words in the text data, is the number of sentences containing the subject in the text data, is the total number of sentences in the text data, is the proportion of sentiment words in the text data, is the syntactic complexity, is the sentence semantic consistency of the th sentiment word in the sentence, and are coefficients.
[0040] Furthermore, the facial analysis module includes:
[0041] The facial key point position unit is used to determine the relative positions of the eyes, eyebrows, and corners of the mouth according to the facial key point features, and obtain the facial key point position data;
[0042] The facial expression change unit is used to calculate the key point offset between adjacent time frames according to the facial key point position data and the expression change features, analyze the facial expression change trend, and obtain the facial expression dynamic data;
[0043] The facial expression symmetry calculation unit is used to calculate the facial left-right symmetry index according to the facial expression dynamic data, analyze the stability of the facial expression, and obtain the facial expression symmetry data;
[0044] The facial emotion judgment unit is used to judge the facial emotion type of the user according to the facial symmetry data, and obtain the facial emotion data.
[0045] Furthermore, the limb analysis module includes:
[0046] A body stability calculation unit, configured to analyze the user's movement trajectory according to the limb movement trajectory characteristics, calculate the movement frequency per unit time, and obtain body stability data;
[0047] A visual attention analysis unit, configured to identify the user's visual fixation area according to the head position and eye direction, and obtain visual attention data;
[0048] An attention state judgment unit, configured to judge the user's attention state according to the body stability data and the visual attention data, and obtain attention state data.
[0049] Further, the optimization and adjustment module includes:
[0050] A voice fundamental frequency adjustment unit, configured to increase the upper limit of the voice fundamental frequency of the virtual digital human according to the voice emotion data, making the voice data more high-pitched;
[0051] A speech rate and pitch adjustment unit, configured to dynamically optimize the speech rate and pitch of the virtual digital human according to the voice emotion data, making the voice data more fluent;
[0052] A syntactic complexity adjustment unit, configured to optimize the syntactic complexity in the text of the virtual digital human according to the text emotion data, making it conform to the user's comprehension ability and emotional state;
[0053] An emotional word frequency adjustment unit, configured to adjust the usage frequency of emotional words in the text of the virtual digital human according to the text emotion data, enhancing the user's emotional resonance.
[0054] Further, the optimization and adjustment module further includes:
[0055] A blink frequency adjustment unit, configured to optimize the blink frequency of the virtual digital human according to the facial emotion data, making the virtual digital human show a focused state;
[0056] A mouth corner and eyebrow adjustment unit, configured to optimize the mouth corner curvature and eyebrow angle of the virtual digital human according to the facial emotion data, making it conform to the user's emotional state;
[0057] An information presentation order adjustment unit, configured to adjust the information presentation order of the virtual digital human according to the attention state data, making it conform to the user's attention concentration state;
[0058] A visual guidance adjustment unit, configured to adjust the line of sight direction and head movement of the virtual digital human according to the attention state data, making it conform to the user's visual focus.
[0059] The above solution of the present invention has at least the following beneficial effects:
[0060] The present invention analyzes the user's speech emotion type by calculating the audio change rate and speech rate stability, thereby obtaining speech emotion data. It can accurately judge the user's current speech emotion state through a quantitative calculation method, not only paying attention to the text information of the speech content, but also being able to combine the prosodic features of the speech for emotion judgment. By analyzing by combining multiple speech features, it can provide a more accurate judgment than a single speech emotion recognition method, improve the emotional resonance between the virtual digital human and the user, and the analysis results can be used as input data for the optimization and adjustment module to make the interaction more natural and smooth.
[0061] The present invention analyzes the user's text emotion type by calculating the syntactic complexity and the number of emotion words, obtaining text emotion data, and introducing more complex text analysis methods, such as syntactic structure analysis and emotion word frequency statistics, to improve the accuracy of text emotion recognition. When the text input by the user contains more complex sentence patterns and more negative emotion words, the system can judge that the user may be in an anxious or uneasy state; while a text using more short sentences and containing positive emotion words may represent a happy or excited emotion. It can not only improve the accuracy of text emotion analysis, but also combine the results of the speech analysis module to further enhance the reliability of multi-modal emotion judgment.
[0062] The present invention analyzes the user's facial emotion type by calculating the facial expression symmetry, obtaining facial emotion data, increasing the dynamic tracking of facial key points, and being able to more accurately capture the changes in the user's emotions. By analyzing the displacement of key points between adjacent time frames, it is possible to detect expression changes such as smiling, frowning, and surprise, thereby inferring the user's emotion state, and introducing facial symmetry analysis to evaluate the stability of the expression. This refined expression analysis method can significantly improve the virtual digital human's ability to perceive the user's emotions, and thus make interaction feedback more in line with the user's psychological state.
[0063] The present invention analyzes the user's attention state by calculating body stability and identifying the visual fixation area, obtaining attention state data. By tracking the movement trajectories of key joints such as the user's head, arms, and legs, calculating limb stability, and combining information on head orientation and visual fixation area, the degree of user concentration is inferred. When the user's head moves frequently and the visual fixation time is short, it may mean inattention; when the user's line of sight is fixed on the screen for a long time, it may represent high concentration, enabling the virtual digital human to accurately judge the user's attention state, and thus provide more personalized content push in scenarios such as teaching or training, improving the user's learning efficiency and interaction experience.
[0064] The present invention dynamically adjusts the speech, text, expressions, and visual guidance methods of a virtual digital human through speech emotion data, text emotion data, facial emotion data, and attention state data, and can adjust the interaction performance of the virtual digital human according to the user's real-time state, making it more personalized and emotionally resonant. In terms of speech, it can dynamically adjust the speech rate and pitch of the virtual digital human to conform to the user's emotional state; in terms of text, it can optimize the sentence complexity to make the text expression more natural; in terms of visual performance, it can adjust the eye contact, mouth curvature, etc. of the virtual digital human to more realistically simulate human emotional responses. Compared with the simple interaction method based on preset rules in the prior art, this function can significantly improve the naturalness of the interaction between the virtual digital human and the user. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 is a flowchart of an AI-based virtual digital human interaction system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0066] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0067] As Figure 1 shown, an embodiment of the present invention provides an AI-based virtual digital human interaction system, and the system includes:
[0068] A feature extraction module, configured to extract features from the preprocessed interaction data set to obtain a user feature data set;
[0069] A speech analysis module, configured to calculate the audio change rate and speech rate stability according to the user feature data set, and analyze the speech emotion type of the user to obtain speech emotion data;
[0070] A text analysis module, configured to calculate the syntactic complexity and count the number of emotion words according to the user feature data set, and analyze the text emotion type of the user to obtain text emotion data;
[0071] A facial analysis module, configured to calculate the facial expression symmetry according to the user feature data set, and analyze the facial emotion type of the user to obtain facial emotion data;
[0072] A limb analysis module, configured to calculate the body stability and identify the visual fixation area according to the user feature data set, and analyze the attention state of the user to obtain attention state data;
[0073] An optimization and adjustment module, which is used to adjust the tone and speech rate of the virtual digital human according to the speech emotion data; optimize the sentence complexity and emotional words in the virtual digital human's text according to the text emotion data; adjust the changes of the eyes, eyebrows and corners of the mouth of the virtual digital human according to the facial emotion data; adjust the information display order and visual focus guidance according to the attention state data.
[0074] In the embodiment of the present invention, a feature extraction module is used to extract features from the preprocessed interaction data set to obtain a user feature data set, ensuring that the feature data used by subsequent modules has high-quality input; a speech analysis module is used to calculate the audio change rate and speech rate stability according to the user feature data set, analyze the speech emotion type of the user, and obtain speech emotion data, through a data calculation method, rather than simply relying on keyword matching, thereby improving the accuracy of speech emotion recognition; a text analysis module is used to calculate the syntactic complexity and count the number of emotional words according to the user feature data set, analyze the text emotion type of the user, and obtain text emotion data, which can reduce misjudgment caused by individual words and improve the accuracy of emotion recognition; a facial analysis module is used to calculate the facial expression symmetry according to the user feature data set, analyze the facial emotion type of the user, and obtain facial emotion data. By dynamically tracking facial key points, the system can capture the instantaneous changes of the user's emotions and improve the timeliness of expression recognition; a limb analysis module is used to calculate the body stability and identify the visual fixation area according to the user feature data set, analyze the attention state of the user, and obtain attention state data. By combining the body stability and the visual fixation area, the attention state of the user can be more accurately inferred; through the linkage optimization of multi-modal data, the virtual digital human can be more anthropomorphic, enhancing the user's immersion and interaction experience.
[0075] Among them, the interaction data set includes:
[0076] The user's speech data, text data, facial data and limb movement data;
[0077] The speech data captures the user's speech in real time through a microphone or an in-ear microphone of a headset and stores it as an audio file in mp3 format;
[0078] The text data collects the user's input content through a text input interface;
[0079] The facial data collects the user's facial image data in real time through a high-definition camera or an infrared camera;
[0080] The limb movement data collects the user's limb movement data through a depth camera sensor;
[0081] Data denoising: For speech, image, and limb movement data, a filtering algorithm is used to remove background noise and improve data quality;
[0082] Multi-device synchronization: Adopt the timestamp mechanism to keep the voice, text, facial, and body movement data in time synchronization, reducing data mismatch problems.
[0083] In a preferred embodiment of the present invention, the feature extraction module includes:
[0084] A voice feature extraction unit, which is used to calculate the fundamental frequency curve of voice data according to the interaction data set, obtain the voice fundamental frequency feature, perform frequency domain conversion on it to obtain the voice time-frequency feature, and calculate the energy value of the voice data according to the voice time-frequency feature to obtain the voice energy feature;
[0085] A text feature extraction unit, which is used to perform structural analysis on the text data according to the interaction data set, extract the basic structure information of the text data to obtain the text structure feature, and extract the sentiment words in the text data according to it to obtain the text sentiment word feature;
[0086] A facial feature extraction unit, which is used to extract the position information of the eyes, eyebrows, corners of the mouth, and facial contour according to the interaction data set, identify the key points of the facial data to obtain the facial key point feature, and obtain the expression change feature by analyzing the displacement of the key points between adjacent time frames;
[0087] A body feature extraction unit, which is used to extract the position information of the head, arms, torso, and legs according to the interaction data set to obtain the body joint point feature, and draw the trajectory curve of the user's movement according to it to analyze the user's movement pattern to obtain the body movement trajectory feature.
[0088] In the embodiments of the present invention, a voice feature extraction unit is configured to calculate the fundamental frequency curve of voice data according to an interaction data set, obtain the voice fundamental frequency feature, perform frequency domain conversion on it to obtain the voice time-frequency feature, calculate the energy value of the voice data according to the voice time-frequency feature to obtain the voice energy feature, so that the system can recognize the voice features corresponding to different emotions and improve the accuracy of voice emotion recognition; a text feature extraction unit is configured to perform structural analysis on text data according to the interaction data set, extract the basic structure information of the text data to obtain the text structure feature, and extract the emotional words in the text data according to it to obtain the text emotional word feature, which can help the system recognize the emotional tendency of the user; a facial feature extraction unit is configured to extract the position information of the eyes, eyebrows, corners of the mouth, and facial contour according to the interaction data set, identify the key points of the facial data to obtain the facial key point feature, and obtain the expression change feature by analyzing the displacement of the key points between adjacent time frames, which can accurately capture the facial emotions of the user; a limb feature extraction unit is configured to extract the position information of the head, arms, torso, and legs according to the interaction data set to obtain the limb joint point feature, and draw the trajectory curve of the user's action according to it, analyze the user's movement pattern to obtain the limb action trajectory feature, which can detect the user's attention.
[0089] Among them, the voice feature extraction unit is configured to calculate the fundamental frequency curve of voice data according to an interaction data set, obtain the voice fundamental frequency feature, perform frequency domain conversion on it to obtain the voice time-frequency feature, calculate the energy value of the voice data according to the voice time-frequency feature to obtain the voice energy feature, and specifically includes:
[0090] Fundamental frequency curve calculation: First, preprocess the voice data in the interaction data set, including noise reduction, normalization, and time-frequency conversion; perform frequency domain analysis on the voice signal through fast Fourier transform to extract the fundamental frequency of the voice data; calculate the fundamental frequency curve of the voice data changing with time to obtain the pitch fluctuation situation.
[0091] Time-frequency feature extraction: Use Mel cepstral coefficients for feature extraction to reduce noise interference and enhance the emotion expression features of voice data; use short-time Fourier transform to analyze the time-frequency features of the voice signal and extract the key information in the voice spectrum; normalize the extracted time-frequency features to ensure the stability of the feature values under different input conditions.
[0092] Voice energy calculation: Calculate the energy value of the voice signal in each frame and construct a curve of the voice energy changing with time; use a sliding window calculation method to extract the energy fluctuation feature of the voice signal to reflect the loudness change of the voice; combine the voice time-frequency feature and the fundamental frequency curve to form a complete voice feature data set.
[0093] Among them, the text feature extraction unit is used to perform structural analysis on the text data according to the interaction dataset, extract the basic structural information of the text data to obtain text structure features, and extract sentiment words in the text data to obtain text sentiment word features, specifically including:
[0094] Text structure analysis: Parse the text data, perform word segmentation and part-of-speech tagging to extract sentence structure information; identify the subject-predicate-object structure and analyze the basic grammatical features of the text; calculate the syntactic complexity of the text, including the use of subordinate clauses, compound sentences, and rhetorical devices.
[0095] Text sentiment word extraction: Perform sentiment analysis on the text data using a preset sentiment dictionary; count the number of sentiment words and their distribution positions in the text; combine the syntactic structure information to calculate the influence weight of the sentiment words.
[0096] Among them, the facial feature extraction unit is used to extract the position information of the eyes, eyebrows, corners of the mouth, and facial contour according to the interaction dataset, identify the key points of the facial data to obtain facial key point features, and obtain the expression change features by analyzing the displacement of the key points between adjacent time frames, specifically including;
[0097] Facial key point extraction: Detect the facial key points of the user through a deep learning algorithm; identify the position information of the eyes, eyebrows, corners of the mouth, and facial contour to construct a facial feature point dataset of the user; calculate the relative positions of the facial key points to ensure the spatial consistency of the feature points.
[0098] Expression change feature calculation: Calculate the displacement of the key points between adjacent time frames to analyze the trend of facial expression changes; count the expression change rates of different time frames to distinguish static expressions and dynamic expressions; calculate the symmetry of the eyes and corners of the mouth to determine whether there are asymmetrical expressions, such as an asymmetrical smile.
[0099] Among them, the limb feature extraction unit is used to extract the position information of the head, arms, torso, and legs according to the interaction dataset to obtain limb joint point features, and draw the trajectory curve of the user's actions based on this to analyze the user's movement pattern to obtain limb action trajectory features, specifically including:
[0100] Limb joint point feature extraction: Use human pose recognition technology to extract the position information of the head, arms, legs, and torso; construct a joint point dataset of the user and calculate the relative positions of each joint point; perform multi-frame time series analysis to identify the movement pattern.
[0101] User movement trajectory analysis: Calculate the joint point movement trajectory within a unit time to extract limb action features; identify the user's movement pattern, such as whether they move frequently or maintain a stable posture; analyze the head orientation to judge the user's visual focus.
[0102] In a preferred embodiment of the present invention, the speech analysis module includes:
[0103] An audio change calculation unit, configured to analyze the change trend of the speech fundamental frequency according to the speech fundamental frequency feature, calculate the audio change rate of the speech fundamental frequency, and obtain speech fundamental frequency change data;
[0104] A time series unit, configured to calculate the energy value of speech data within a fixed time window according to the speech energy feature, and obtain speech energy time series data;
[0105] An energy fluctuation calculation unit, configured to calculate the energy change rate between consecutive time windows according to the speech energy time series data, and obtain speech energy fluctuation data;
[0106] A speech rate stability analysis unit, configured to calculate the standard deviation of the energy fluctuation of speech data according to the speech energy fluctuation data, analyze the degree of fluctuation of the energy value over time, and obtain speech rate stability data;
[0107] A speech emotion judgment unit, configured to judge the type of the user's speech emotion according to the speech fundamental frequency change data and the speech rate stability data, and obtain speech emotion data.
[0108] In the embodiment of the present invention, the audio change calculation unit, configured to analyze the change trend of the speech fundamental frequency according to the speech fundamental frequency feature, calculate the audio change rate of the speech fundamental frequency, and obtain speech fundamental frequency change data, improves the accuracy of speech feature analysis; the time series unit, configured to calculate the energy value of speech data within a fixed time window according to the speech energy feature, and obtain speech energy time series data, provides the dynamic characteristics of the speech signal changing over time; the energy fluctuation calculation unit, configured to calculate the energy change rate between consecutive time windows according to the speech energy time series data, and obtain speech energy fluctuation data, can accurately capture the change trend of the speech energy over time; the speech rate stability analysis unit, configured to calculate the standard deviation of the energy fluctuation of speech data according to the speech energy fluctuation data, analyze the degree of fluctuation of the energy value over time, and obtain speech rate stability data, and the quantitative speech rate stability analysis enables the system to more scientifically evaluate whether the user's speech rate is stable; the speech emotion judgment unit, configured to judge the type of the user's speech emotion according to the speech fundamental frequency change data and the speech rate stability data, and obtain speech emotion data, can automatically learn the emotion patterns of different speech features and improve the recognition accuracy.
[0109] Among them, the speech emotion judgment unit, configured to judge the type of the user's speech emotion according to the speech fundamental frequency change data and the speech rate stability data, and obtain speech emotion data, specifically includes:
[0110] The voice fundamental frequency change data and speech rate stability data are combined to form a multi-dimensional feature vector for emotion classification. The fundamental frequency change data mainly reflects the pitch fluctuations of the user's voice. For example, an angry emotion is usually characterized by a high fundamental frequency change rate, while a sad emotion has a smaller fundamental frequency change and tends to be at a lower frequency. The speech rate stability data is used to analyze the fluency and rhythm of the user's speech. For example, tense or anxious speech usually has a higher speech rate fluctuation, while calm or confident speech has a higher speech rate stability. To improve the classification accuracy, the principal component analysis method is used to reduce feature redundancy, and the feature standardization method is used to make the dimensions of different feature data consistent, ensuring that the subsequent model calculations are not affected by different numerical scales and do not affect the judgment results.
[0111] The feature data is classified to determine the user's emotion type. First, the feature vector is input into a support vector machine for preliminary classification. For example, according to the fundamental frequency change rate and speech rate stability, emotions are classified into categories such as happy, angry, sad, surprised, etc.; then, a long short-term memory network is used for time series analysis to judge the change trend of emotions, such as whether the emotion is in a gradually strengthening or weakening state. For the situation of multiple emotions intersecting, a soft classification method can be introduced to calculate the confidence of each emotion category and select the category with the highest confidence as the final judgment result.
[0112] Combined with context information and historical data, the emotion recognition result is further optimized to improve the robustness of the system. The time-weighted smoothing method is used to ensure that the emotion judgment within a short period does not show drastic fluctuations, thereby enhancing the coherence of emotion recognition. For example, if a user's speech shows a happy emotion within a short period and suddenly changes to a sad emotion in the next second, the system will consider the current environment and combine the front and back speech features to adjust the final emotion judgment result. To reduce noise interference, the system introduces a confidence threshold mechanism, and only when the confidence of the emotion recognition result is higher than the set threshold, the emotion category is output to avoid misjudgment.
[0113] In a preferred embodiment of the present invention, the speech rate stability analysis unit includes:
[0114] A standard deviation calculation unit for calculating the mean value of the energy fluctuation of the speech data according to the speech energy fluctuation data and calculating the standard deviation of the speech data based on the mean value of the energy fluctuation to obtain the speech energy standard deviation data;
[0115] A normalization processing unit for normalizing the speech energy standard deviation data to obtain the normalized speech rate stability data;
[0116] A speech rate evaluation unit for judging the speech rate stability of the user according to the normalized speech rate stability data. When the normalized speech rate stability data tends to 1, the speech rate tends to be stable; when the normalized speech rate stability data tends to 0, the speech rate tends to be unstable.
[0117] In an embodiment of the present invention, a standard deviation calculation unit is configured to calculate the mean value of the energy fluctuation of speech data according to the speech energy fluctuation data, and calculate the standard deviation of the speech data according to the mean value of the energy fluctuation, so as to obtain the speech energy standard deviation data. By using the standard deviation calculation method to quantify the fluctuation of the speech energy, it can reflect the emotional state of the speech; a normalization processing unit is configured to perform normalization processing on the speech energy standard deviation data to obtain the normalized speech rate stability data, which can effectively reduce the influence caused by individual differences and make the speech energy standard deviation data of different users comparable;
[0118] A speech rate evaluation unit is configured to judge the speech rate stability of a user according to the normalized speech rate stability data. When the normalized speech rate stability data tends to 1, the speech rate tends to be stable; when the normalized speech rate stability data tends to 0, the speech rate tends to be unstable. By using a quantitative method to accurately evaluate the speech rate stability of the user, it provides an important basis for subsequent speech adjustment.
[0119] In a preferred embodiment of the present invention, the text analysis module includes:
[0120] A text syntax analysis unit is configured to perform syntax analysis on text data according to the text structure characteristics, calculate the syntax complexity, and obtain the text syntax data;
[0121] A text emotion word analysis unit is configured to count the number of emotion words in the text data according to the text emotion word characteristics and calculate the emotion word frequency to obtain the text emotion word data;
[0122] A text emotion bias degree calculation unit is configured to analyze the emotion tendency of the text data according to the text syntax data and the text emotion word data, calculate the text emotion bias degree, and obtain the text emotion bias data;
[0123] A text emotion judgment unit is configured to judge the text emotion type of the user according to the text emotion bias data to obtain the text emotion data.
[0124] In the embodiments of the present invention, a text syntactic analysis unit is configured to perform syntactic analysis on text data according to text structure features, calculate syntactic complexity, obtain text syntactic data, quantify the structural complexity of the text, so as to judge the user's language expression ability and their emotional state; a text sentiment word analysis unit is configured to count the number of sentiment words in the text data according to text sentiment word features and calculate sentiment word frequency, obtain text sentiment word data, quantify the emotional tendency in the user's text, and analyze the user's true emotional state in combination with syntactic complexity; a text emotion deviation degree calculation unit is configured to analyze the emotional tendency of the text data according to the text syntactic data and the text sentiment word data, calculate the text emotion deviation degree, obtain text emotion deviation data, which can quantify the overall emotional trend of the text, enabling the virtual digital human to adjust the interaction strategy more accurately; a text emotion judgment unit is configured to judge the text emotion type of the user according to the text emotion deviation data, obtain text emotion data, realize the precise classification of the user's text emotion, and provide input data for the subsequent optimization and adjustment module.
[0125] Among them, the text syntactic analysis unit is configured to perform syntactic analysis on text data according to text structure features, calculate syntactic complexity, obtain text syntactic data, and specifically includes:
[0126] First, preprocess the input text, including word segmentation, stop word removal, and part-of-speech tagging, to ensure the accuracy of subsequent syntactic analysis. Then, use natural language processing techniques, such as a dependency-based syntactic analyzer or context-free grammar parsing techniques, to decompose the sentence structure, identify the subject-predicate-object structure, modification relationships, and clause nesting levels.
[0127] After completing the basic syntactic parsing, quantify the complexity of the text structure. First, traverse each sentence in the text, calculate its word count, and count the proportion of complex sentences containing multiple clauses. If a sentence contains multiple subordinate relationships, such as "because...so..." or "although...but...", then the sentence is regarded as a complex sentence, and the syntactic complexity weight is increased according to its nesting level. Secondly, analyze the distribution of different parts of speech, such as verbs, nouns, and adjectives, in the text to judge the syntactic richness at the lexical level. For example, if a sentence contains more modifying components, such as adjectives and adverbs, it indicates that the sentence structure is more complex.
[0128] In addition, a syntactic tree parsing algorithm is adopted to decompose the sentence structure to calculate the nesting depth. The core of syntactic tree parsing lies in constructing the hierarchical structure in the sentence, dividing it into multiple clauses, and calculating the number of levels. For example, the sentence "I know he likes reading books" can be parsed into a two-layer structure, where "he likes reading books" is a nested clause, which has a higher syntactic complexity compared to a simple sentence. During the analysis process, the logical relationships between sentences are also considered, and inter-sentence connectives such as "therefore" and "however" are identified to determine whether the text contains complex reasoning logic.
[0129] Among them, the text emotion judgment unit is used to judge the text emotion type of the user according to the text emotion bias data to obtain the text emotion data, specifically including:
[0130] First, data normalization processing is performed to ensure that the emotion data is within a reasonable numerical range. Since multiple variables are involved in the calculation process of text emotion bias, the data distributions of different texts may vary greatly. Therefore, the maximum-minimum normalization method is used to map the bias values to the interval [0, 1]. After processing, the emotion biases of all texts are within the standardized range, eliminating the bias caused by different numerical ranges and making the subsequent classification judgment more accurate.
[0131] After completing the normalization processing, the text emotion judgment unit uses a classification model to identify the text emotion type. A classification method based on decision trees is adopted. Through a pre-trained emotion data set, the mapping relationship between different text emotion bias data and actual emotion categories is learned. After inputting the normalized bias data, the classification model calculates the probability distribution of the text belonging to different emotion categories, such as positive, neutral, and negative, and selects the category with the highest probability as the final text emotion category. For example, if the calculated probability distribution is (0.8, 0.15, 0.05), the system will judge the text emotion as a positive emotion.
[0132] In a preferred embodiment of the present invention, the text emotion bias calculation unit includes:
[0133] An emotion score calculation unit for calculating the emotion score of the sentiment word according to the text sentiment word data and a preset emotion dictionary;
[0134] A sentiment word weight calculation unit for determining the text weight of the sentiment word in the text data according to the emotion score;
[0135] A sentence meaning consistency calculation unit for determining the sentence weight of the sentiment word in the sentence according to the text syntactic data and calculating the sentence meaning consistency of the sentiment word in the sentence;
[0136] A text emotion bias calculation formula unit is used to calculate the text emotion bias according to the sentence meaning consistency. The calculation formula of the text emotion bias is as follows:
[0137] ,
[0138] where, is the text emotion bias, is the text weight of the th sentiment word, is the emotion score of the th sentiment word, is the total number of sentiment words, is the index of the sentiment word, is the total number of words in the text data, is the number of sentences containing the subject in the text data, is the total number of sentences in the text data, is the proportion of sentiment words in the text data, is the syntactic complexity, is the th sentence meaning consistency of the sentiment word in the sentence, and are coefficients.
[0139] In the embodiments of the present invention, the refined processing of text emotions. By combining the weights and emotion scores of sentiment words, the system can identify the core emotional expressions in the text, rather than relying solely on simple word statistics. For example, in the sentence "This performance is very good", the sentiment word "good" may be a neutral emotion when viewed alone, but when combined with the adverb "very", its emotional intensity is significantly enhanced. When calculating, this formula can accurately identify this enhanced relationship through weight correction and avoid misjudgment. In addition, the introduction of sentence meaning consistency ensures the rationality of sentiment words in the context. If a sentiment word is inconsistent with the overall meaning of the sentence, its influence weight will be reduced, thereby improving the stability of the calculation.
[0140] Among them, the sentiment word weight calculation unit is used to determine the weight of the sentiment word in the text data according to the emotion score, specifically including:
[0141] First, preprocess the input text, including word segmentation, stop word removal, part-of-speech tagging, and dependency syntactic analysis, to obtain the basic structural information of the sentence. In the word segmentation stage, the text is split into independent words or phrases, laying the foundation for subsequent matching analysis. Stop word removal can reduce the interference of invalid words on sentiment analysis and improve computational efficiency. The role of part-of-speech tagging is to identify important components such as adjectives, verbs, and adverbs that have a greater impact on emotions, providing a basis for the screening of sentiment words. Dependency syntactic analysis is used to determine the subject-predicate-object relationship in the sentence, clarify the object modified by the sentiment word, and thus affect its weight calculation. After completing the text preprocessing, by means of emotion dictionary matching, screen out the sentiment words in the text and assign a basic emotion score to each sentiment word. This score usually comes from the emotion dictionary, where positive words correspond to positive values and negative words correspond to negative values. For example, in the sentence "The weather today is very good and I feel very happy.", the matched sentiment words "good" and "happy" correspond to emotion scores of 0.8 and 0.9 respectively. These scores represent their basic intensity in emotion expression.
[0142] To reasonably allocate the influence weight of sentiment words in the text, it is necessary to calculate their contribution ratio according to their emotion scores. Specifically, the weight of a single sentiment word is determined by the ratio of its emotion score to the total score of all sentiment words in the text. Suppose the total sum of the emotion scores of all sentiment words in the text is S. The weight of a single sentiment word can be calculated as the ratio of the basic intensity of the sentiment word to the total score. In the above example, the score of the sentiment word "good" is 0.8, the score of "happy" is 0.9, and the total score is 1.7. Then the calculated weight of "good" is 0.47 and the weight of "happy" is 0.53. Through this calculation method, it is ensured that sentiment words with higher emotion intensity have a greater impact on the text emotion, and at the same time, the misleading of low-frequency sentiment words on the overall emotion judgment is avoided.
[0143] Among them, the sentence meaning consistency calculation unit is used to determine the weight of the sentiment word in the sentence according to the text syntactic data and calculate the sentence meaning consistency of the sentiment word in the sentence. Specifically, it includes:
[0144] When calculating the emotional bias of text, only considering the weights and scores of sentiment words still has limitations. Because the role of sentiment words in a sentence is affected by syntactic structures. For example, the distance between a sentiment word and the subject or verb, whether it is modified by an emphasizing adverb, etc. will all affect the intensity of the emotion it actually expresses. Therefore, it is necessary to adjust the weights of sentiment words in a sentence according to the syntactic data of the text. First, determine the relevance between sentiment words and the core sentence components through dependency analysis. The closer a sentiment word is to the core components, the greater its impact on the overall sentence meaning. For example, in the two sentences "I like this book very much" and "This book is liked by many people", "like" is closely connected to the subject "I" in the first sentence and plays a leading role in emotion expression, while in the second sentence, it is a general description and has a weaker reflection on individual emotions. Based on this characteristic, a dependency distance factor can be defined , and its value can be set as , so as to ensure that the higher the relevance between a sentiment word and the core sentence components, the greater its weight
[0145] In addition to dependency relationships, the intensity of emotion expression is also affected by modifiers. For example, adverbs such as "very" and "extremely" will enhance emotion expression, while words such as "a little" and "possibly" will weaken the emotion intensity. In order to accurately reflect this influence, an emotion intensification factor needs to be introduced, and its value is dynamically adjusted according to the modification situation of the sentiment word. For example, in the two sentences "I am very happy" and "I am a little happy", "very" can enhance the intensity of "happy", so it is set as , "a little" weakens the emotion expression and is set as . Finally, the weight of a sentiment word in a sentence can be calculated as , where is the weight of the th sentiment word in the sentence is the weight of the th sentiment word in the text data, and is equivalent to in the formula for calculating the text emotional bias, and the two have the same meaning and value is the emotion intensification factor of the th sentiment word in the sentence. Taking the previous example for calculation, assume that the dependency distance between "good" and the core component is 2, and its dependency distance factor is 0.5, while "happy" directly modifies the subject "I" with a dependency distance of 1 and a dependency distance factor of 1. Since "happy" is modified by "extremely", the emotion intensification factor is set as 1.2, and "good" has no modification, so the emotion intensification factor is set as 1. Finally, it is calculated that the sentence weight of "good" is 0.235, and the final weight of "happy" is 0.636. It can be seen that the emotional influence of "happy" is stronger than that of "good"
[0146] Sentence consistency measures the degree of match between sentiment words and the main structure of the sentence, avoiding individual sentiment words from causing deviations in the overall sentiment judgment. For example, in the two sentences "This meal is really unpalatable" and "This meal is not bad, but the environment is poor", the former's "unpalatable" is the core evaluation word of the sentence, which is highly matched with the subject, while the latter's "poor" is a negative word, but it mainly modifies "environment" rather than "this meal", so it cannot be directly used as a decisive factor in the overall sentiment of the text. To achieve this judgment, the word vector calculation method can be used to quantify the degree of consistency by calculating the cosine similarity between the sentiment word and the overall sentence vector. The closer the value is to 1, the higher the consistency is, and the closer it is to 0, the lower the consistency is. For example, assuming that the cosine similarity between the vectors of "good" and "weather" is 0.85, and the similarity between the vectors of "happy" and "me" is 0.95, the final weight of "good" is calculated to be 0.2, and the final weight of "happy" is 0.6, thereby further optimizing the accuracy of emotion recognition.
[0147] In a preferred embodiment of the present invention, the facial analysis module includes:
[0148] The facial key point location unit is used to determine the relative positions of eyes, eyebrows, and mouth corners according to the facial key point features, and obtain the facial key point location data;
[0149] A facial expression change unit is used to calculate the key point offset between adjacent time frames based on the facial key point position data and expression change characteristics, analyze the facial expression change trend, and obtain facial expression dynamic data;
[0150] A facial expression symmetry calculation unit is used to calculate the left-right symmetry index of the face according to the dynamic data of facial expression, analyze the stability of the facial expression, and obtain the facial expression symmetry data;
[0151] The facial emotion judgment unit is used to judge the user's facial emotion type according to the facial symmetry data to obtain facial emotion data.
[0152] In the embodiment of the present invention, the facial key point position unit is used to determine the relative positions of the eyes, eyebrows, and the corners of the mouth based on the facial key point features, and obtain the facial key point position data. By extracting the facial key points, the system can accurately capture the user's expression details and provide data support for subsequent expression analysis; the facial expression change unit is used to calculate the key point offset between adjacent time frames according to the facial key point position data and the expression change features, analyze the facial expression change trend, and obtain the facial expression dynamic data, which can effectively capture the dynamic changes of facial emotions; the facial expression symmetry calculation unit is used to calculate the left-right symmetry index of the face according to the facial expression dynamic data, analyze the stability of the facial expression, and obtain the facial expression symmetry data, which helps the system distinguish real emotions from disguised emotions; the facial emotion judgment unit is used to judge the type of the user's facial emotion according to the facial symmetry data, and obtain the facial emotion data, which can achieve more accurate emotion classification and improve the virtual digital human's perception ability of the user's emotions.
[0153] Among them, the facial expression symmetry calculation unit is used to calculate the left-right symmetry index of the face according to the facial expression dynamic data, analyze the stability of the facial expression, and obtain the facial expression symmetry data, specifically including:
[0154] Based on the facial key point position data and the expression dynamic data, calculate the symmetry index of the left and right faces, and analyze the expression stability of the user;
[0155] The calculation method is as follows:
[0156] The symmetry of the opening and closing of the left and right eyes = the height-width ratio of the left eye / the height-width ratio of the right eye;
[0157] The position deviation of the left and right corners of the mouth = |the vertical coordinate of the left corner of the mouth - the vertical coordinate of the right corner of the mouth|;
[0158] The difference in the inclination of the eyebrows = |the angle of the left eyebrow - the angle of the right eyebrow|;
[0159] By comparing the positions of the feature points on the left and right faces, calculate the facial symmetry score. The higher the score, the more uniform and natural the expression is.
[0160] Among them, the facial emotion judgment unit is used to judge the type of the user's facial emotion according to the facial symmetry data, and obtain the facial emotion data, specifically including:
[0161] The system uses the data from the facial symmetry analysis unit to make emotional judgments. Facial symmetry calculations usually include parameters such as the symmetry of left and right eye opening and closing, the position deviation of the left and right corners of the mouth, and the difference in eyebrow inclination. The symmetry score is calculated by mathematical methods. For example, a symmetry index threshold is set. When the position deviation of the left and right corners of the mouth exceeds a certain critical value, it may indicate that the user's emotions tend to be doubtful or dissatisfied. Similarly, if one side of the eyebrow is obviously raised while the other side remains unchanged, it may indicate that the user is showing doubt or thinking. In addition to static symmetry analysis, the system also combines the dynamic changes of the user's face to judge the authenticity of the emotions. For example, in a real smile, the eye muscles usually produce a natural contraction, and if the corners of the mouth are raised but the eye muscles do not change synchronously, it may indicate that the smile is deliberate rather than spontaneous.
[0162] Subsequently, the system will input all the calculated facial feature data into the pre-trained emotion classification model to automatically identify the final facial emotion type. The classification model usually adopts machine learning or deep learning methods, such as support vector machines or emotion recognition models based on convolutional neural networks. These models are usually trained on large-scale facial emotion datasets to ensure high-accuracy emotion classification capabilities. The model will combine multiple facial feature data of the user to classify emotions. For example, when the user shows raised eyebrows, moderately raised corners of the mouth, and relaxed eyes, the system will judge that he is in a "happy" or "relaxed" state, and when the corners of the mouth droop, the brows are furrowed, and the eyes are dull, the system will judge that he is in a "sad" or "tired" state.
[0163] In a preferred embodiment of the present invention, the limb analysis module includes:
[0164] A body stability calculation unit, used to analyze the user's motion trajectory according to the limb motion trajectory characteristics, calculate the motion frequency per unit time, and obtain body stability data;
[0165] A visual attention analysis unit, used to identify the user's visual focus area based on the head position and eye direction, and obtain visual attention data;
[0166] The attention state judgment unit is used to judge the user's attention state according to the body stability data and the visual attention data to obtain attention state data.
[0167] In an embodiment of the present invention, a body stability calculation unit is configured to analyze the user's movement trajectory according to the limb movement trajectory characteristics, calculate the movement frequency per unit time, obtain body stability data, accurately capture the user's limb movement through non-contact sensing technology, and then through trajectory feature analysis, it can identify whether the user has redundant actions to help judge the user's concentration; a visual attention analysis unit is configured to identify the user's visual fixation area according to the head position and eye direction, obtain visual attention data, and by counting the visual fixation duration, the system can evaluate whether the user is interested in the screen content; an attention state judgment unit is configured to judge the user's attention state according to the body stability data and the visual attention data, obtain attention state data, and combining multi-modal information fusion can more accurately evaluate the user's attention state than a single-modal method.
[0168] Among them, the body stability calculation unit is configured to analyze the user's movement trajectory according to the limb movement trajectory characteristics, calculate the movement frequency per unit time, obtain body stability data, and specifically includes:
[0169] First, extract the time series coordinates of key joint points from the user's limb movement data, and calculate the movement trajectory of the joint points through a trajectory tracking algorithm. The system uses low-pass filtering to smooth the data to reduce noise interference and ensure the continuity of the trajectory. Subsequently, calculate the rate of change of the movement amplitude per unit time, that is, by measuring the displacement amplitude of the joint points between adjacent time frames to judge the stability of the limbs. For example, in a standing state, if the trajectory changes of the torso and legs are small, it indicates high body stability; on the contrary, if frequent small fluctuations are detected, it indicates that the user may be in an unstable state. The system further analyzes the overall stability trend of the user within a specific time window, such as using Fourier transform or time-frequency analysis methods to extract periodic oscillation characteristics to identify whether there are regular fluctuations or redundant actions. Finally, the system calculates the user's body stability index based on statistical analysis methods and standardizes it so that it can be horizontally compared with other user data to provide a quantitative reference for further evaluation of the attention state.
[0170] Among them, the calculation formula for body stability is:
[0171] ,
[0172] Among them, is the body stability, is the length of the time window, is the index of time, is the limb movement amplitude at the moment, is the mean value of the limb movement amplitude within the time window, is the standard deviation of the limb movement amplitude within the time window, is the energy of the dominant oscillation frequency, is the sum of the energies of all frequencies, is the joint velocity at time , is the joint point position at time is the joint point position at time is the difference between the time and the time is the average value of the joint velocity within the time window, is the coefficient.
[0173] Among them, the visual attention analysis unit is used to identify the visual fixation area of the user based on the head position and eye direction, and obtain visual attention data, specifically including:
[0174] By capturing the user's head direction and eye movement data, the system determines the visual fixation area in real time. First, the system uses facial key point detection technology to extract the orientation angle of the head and the pupil center coordinates of the eyes, constructs the user's line-of-sight vector, and combines with a depth camera or an infrared eye tracking device. The system measures the user's gaze direction and infers the user's main visual attention area according to the position of the fixation point projected on the screen or in the environment. For example, if the user's line of sight is fixed in the central area of the screen for a long time, the system determines that the user is in a focused state; if the user's line of sight frequently deviates or quickly jumps between multiple areas, it may indicate that the user's attention is scattered. In addition, the system also combines the blink frequency and the micro-eye movement pattern to further optimize the recognition accuracy of the visual attention area. For example, in a learning or reading scenario, not blinking for a long time may represent high concentration, while frequent blinking or wandering of the line of sight may indicate fatigue or distraction. Finally, the visual attention data will be provided as input to the attention state judgment unit to comprehensively evaluate the user's concentration.
[0175] Among them, the attention state judgment unit is used to judge the attention state of the user based on the body stability data and the visual attention data, and obtain attention state data, specifically including:
[0176] First, establish a multimodal fusion algorithm to perform weighted calculations on body stability data and visual attention data, and combine time series analysis methods to extract the changing trend of the user's attention. For example, if the user has a high body stability index and the visual attention data shows that their line of sight stays in the key area, the system will determine that the user is in a focused state; conversely, if it is detected that the user's body shakes significantly and the line of sight is unstable, it may be determined that the user's attention is distracted. The system further uses a machine learning model to train an attention classifier to adapt to the behavior patterns of different users and improve the accuracy of prediction. In addition, the analysis weights can be dynamically adjusted. For example, in a quiet reading environment, the visual attention weight is higher, while in a sports training scenario, the body stability weight is higher. Finally, output the user's attention state data to provide a reference for optimizing the interaction method of the virtual digital human.
[0177] In a preferred embodiment of the present invention, the optimization and adjustment module includes:
[0178] A voice fundamental frequency adjustment unit for increasing the upper limit of the voice fundamental frequency of the virtual digital human according to the voice emotion data to make the voice data more high-pitched;
[0179] A speech rate and pitch adjustment unit for dynamically optimizing the speech rate and pitch of the virtual digital human according to the voice emotion data to make the voice data more fluent;
[0180] A syntactic complexity adjustment unit for optimizing the syntactic complexity in the text of the virtual digital human according to the text emotion data to make it conform to the user's comprehension ability and emotional state;
[0181] An emotional word frequency adjustment unit for adjusting the usage frequency of emotional words in the text of the virtual digital human according to the text emotion data to enhance the user's emotional resonance.
[0182] In an embodiment of the present invention, a voice fundamental frequency adjustment unit is configured to increase the upper limit of the voice fundamental frequency of the virtual digital human according to voice emotion data, making the voice data more high-pitched. By adjusting the voice fundamental frequency of the virtual digital human, its voice performance becomes more humanized, improving the user's interaction experience; a speech rate and pitch adjustment unit is configured to dynamically optimize the speech rate and pitch of the virtual digital human according to voice emotion data, making the voice data more fluent, and making the speech rate and pitch of the virtual digital human conform to the user's communication habits, avoiding the impact of too fast or too slow speech on the interaction experience; a syntactic complexity adjustment unit is configured to optimize the syntactic complexity in the virtual digital human's text according to text emotion data, making it conform to the user's comprehension ability and emotional state. By adaptively adjusting the syntactic complexity, the virtual digital human's expression becomes more in line with the user's comprehension ability; an emotional word frequency adjustment unit is configured to adjust the usage frequency of emotional words in the virtual digital human's text according to text emotion data, enhancing the user's emotional resonance. By adjusting the emotional word frequency, the virtual digital human's dialogue becomes more in line with the user's current emotional needs.
[0183] Among them, the voice fundamental frequency adjustment unit is configured to increase the upper limit of the voice fundamental frequency of the virtual digital human according to voice emotion data, making the voice data more high-pitched, and specifically includes:
[0184] According to the voice fundamental frequency change data, calculate the fundamental frequency change rate, and in combination with the speech rate stability data, determine the user's voice emotion category;
[0185] When the voice emotion category is "excited", increase the upper limit of the voice fundamental frequency to make the voice data more high-pitched;
[0186] When the voice emotion category is "low", lower the lower limit of the voice fundamental frequency to make the voice data more steady;
[0187] When the voice emotion category is "calm", keep the fundamental frequency within the normal range of the conversation to avoid excessive voice fluctuations.
[0188] Among them, the speech rate and pitch adjustment unit is configured to dynamically optimize the speech rate and pitch of the virtual digital human according to voice emotion data, making the voice data more fluent, and specifically includes:
[0189] According to the standard deviation of the voice energy fluctuation, analyze the voice energy stability to judge whether the user's speech rate is stable;
[0190] When the user's speech rate stability data tends to 0, it means the speech rate is unstable, then lower the speech rate to make the voice expression more fluent;
[0191] When the user's speech rate stability data tends to 1, it means the speech rate is stable, then keep the normal speech rate and adjust the pitch in combination with the voice fundamental frequency to make the voice more expressive;
[0192] When the speech emotion category is "excited", increase the speech rate and enhance the amplitude of pitch fluctuations to make the speech more vivid;
[0193] When the speech emotion category is "steady", reduce the amplitude of pitch fluctuations to make the speech sound more peaceful.
[0194] Among them, the syntactic complexity adjustment unit is used to optimize the syntactic complexity in the virtual digital human text according to the text emotion data to make it conform to the user's comprehension ability and emotional state, specifically including:
[0195] Calculate the sentence nesting depth and the number of subordinate clauses according to the text syntactic complexity data to determine the syntactic complexity of the text;
[0196] When the user's text emotion category is "doubtful", simplify the text structure, reduce the nested sentences, and improve the readability of the sentences;
[0197] When the user's text emotion category is "formal", increase the sentence structure complexity and introduce more modifying elements to make the expression more logical and professional;
[0198] When the user's text emotion category is "neutral", keep the sentence structure stable to make the text expression clear and natural.
[0199] Among them, the emotional word frequency adjustment unit is used to adjust the usage frequency of emotional words in the virtual digital human text according to the text emotion data to enhance the user's emotional resonance, specifically including:
[0200] Calculate the proportion of positive and negative emotional words in the text according to the text emotional word data to determine the current text emotion deviation degree;
[0201] When the text emotion deviation degree tends to be negative, such as sadness or anger, reduce the use of negative emotional words and use neutral or positive modifiers to ease;
[0202] When the text emotion deviation degree tends to be positive, such as excitement or encouragement, increase the use of positive emotional words and enhance the emotional expressiveness;
[0203] When the text emotion deviation degree is close to 0, that is, neutral, keep the use of emotional words at a moderate level to make the text expression more objective.
[0204] In a preferred embodiment of the present invention, the optimization and adjustment module further includes:
[0205] The blink frequency adjustment unit is used to optimize the blink frequency of the virtual digital human according to the facial emotion data to make the virtual digital human show a focused state;
[0206] The mouth corner and eyebrow adjustment unit is used to optimize the curvature of the virtual digital person's mouth corner and the eyebrow angle according to the facial emotion data to make it consistent with the user's emotional state;
[0207] An information presentation sequence adjustment unit, used to adjust the information presentation sequence of the virtual digital human according to the attention state data so as to make it consistent with the user's attention concentration state;
[0208] The visual guidance adjustment unit is used to adjust the line of sight and head movement of the virtual digital person according to the attention state data to make it consistent with the user's visual focus.
[0209] In an embodiment of the present invention, a blinking frequency adjustment unit is used to optimize the blinking frequency of a virtual digital person according to facial emotion data, so that the virtual digital person shows a state of concentration, and the blinking action of the virtual digital person is more in line with the natural reaction of humans, thereby enhancing the realism and intimacy of the interaction, avoiding dull or unnatural eye expressions, and improving the user's immersion in the virtual image; a mouth corner and eyebrow adjustment unit is used to optimize the curvature of the mouth corners and the eyebrow angle of the virtual digital person according to facial emotion data, so that it conforms to the emotional state of the user, so that the virtual digital person can synchronize or adaptively adjust the expression according to the facial emotion data of the user, thereby improving the authenticity of its emotional expression and the naturalness of the interaction, enabling the user to more intuitively feel the emotional feedback of the virtual digital person and increase the emotional resonance of the interaction.
[0210] The information presentation order adjustment unit is used to adjust the information presentation order of the virtual digital human according to the attention state data to make it consistent with the user's concentration state, ensuring that the information presentation order can match the user's attention state, thereby improving the efficiency of information absorption, avoiding understanding deviations caused by distraction, and making information transmission more efficient; the visual guidance adjustment unit is used to adjust the virtual digital human's line of sight and head movement according to the attention state data to make it consistent with the user's visual focus. Through visual guidance adjustment, the virtual digital human can actively attract the user's attention, improve the interactive experience, make it easier for users to obtain key information, and enhance the system's human-computer interaction intelligence and naturalness.
[0211] The blinking frequency adjustment unit is used to optimize the blinking frequency of the virtual digital person according to the facial emotion data so that the virtual digital person can show a focused state, which specifically includes:
[0212] According to the facial key point position data, the user's average blinking time interval is calculated, and the user's current blinking frequency is extracted;
[0213] When the blinking frequency is higher than normal and the facial emotion category is "nervous", the blinking frequency of the virtual digital person is reduced to make it appear focused or anxious;
[0214] When the blink frequency is lower than the normal value and the facial emotion category is "relaxed", increase the blink frequency of the virtual digital human to make it show a natural and relaxed state;
[0215] When the user's facial emotion category is "normal", keep the blink frequency within the average range to make the expression of the virtual digital human more realistic.
[0216] Among them, the mouth-corner and eyebrow adjustment unit is used to optimize the mouth-corner curvature and eyebrow angle of the virtual digital human according to the facial emotion data to make them conform to the user's emotion state, specifically including:
[0217] Calculate the relative positions of the mouth-corner and eyebrows based on the facial expression dynamic data, and analyze the trend of expression changes;
[0218] When the user's facial emotion category is "happy", increase the mouth-corner curvature, enhance the smiling amplitude, and at the same time raise the eyebrow angle to make the expression more bright;
[0219] When the user's facial emotion category is "sad", decrease the mouth-corner curvature to make the mouth-corner droop slightly, and at the same time decrease the eyebrow angle to make the virtual digital human show a melancholy expression;
[0220] When the user's facial emotion category is "surprised", adjust the upward angle of the eyebrows and appropriately raise the mouth-corner to make the expression more variable.
[0221] Among them, the information presentation order adjustment unit is used to adjust the information presentation order of the virtual digital human according to the attention state data to make it conform to the user's attention concentration state, specifically including:
[0222] Analyze the user's attention concentration area based on the user's visual fixation time and content attention degree, and calculate the current attention level;
[0223] When the user's attention is highly concentrated, present the core information first and reduce redundant content to make the interaction more efficient;
[0224] When the user's attention is dispersed, adjust the information presentation rhythm, increase the information repetition times, and use visual guidance to attract the user's attention;
[0225] When the user's attention is at a medium level, adopt a balanced information display method to enable the user to adapt to the information rhythm.
[0226] Among them, the visual guidance adjustment unit is used to adjust the line-of-sight direction and head movement of the virtual digital human according to the attention state data to make them conform to the user's visual focus, specifically including:
[0227] Calculate the user's current visual fixation area based on the visual attention data, and determine the user's line-of-sight focus;
[0228] When the user's visual focus shifts, adjust the line of sight direction of the virtual digital human to align it with the user's line of sight to enhance the interactive perception;
[0229] When the user's attention is low, adjust the small head movements of the virtual digital human, such as nodding or tilting, to attract the user's attention;
[0230] When the user's attention is concentrated, keep the line of sight of the virtual digital human stable and reduce unnecessary head movements to make the interaction more natural.
[0231] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art of the present technology, without departing from the principle described in the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An AI-based virtual digital human interaction system, characterized in that, The system includes: A feature extraction module, which is used to extract features from the preprocessed interaction dataset to obtain a user feature dataset; A speech analysis module, which is used to calculate the audio change rate and speech rate stability according to the user feature dataset, analyze the speech emotion type of the user, and obtain speech emotion data; A text analysis module, which is used to calculate the syntactic complexity and count the number of sentiment words according to the user feature dataset, calculate the text emotion bias degree according to it, analyze the text emotion type of the user, and obtain text emotion data; A facial analysis module, which is used to calculate the facial expression symmetry according to the user feature dataset, analyze the facial emotion type of the user, and obtain facial emotion data; A limb analysis module, which is used to calculate the body stability and identify the visually fixated area according to the user feature dataset, analyze the attention state of the user, and obtain attention state data; An optimization and adjustment module, which is used to adjust the pitch and speech rate of the virtual digital human according to the speech emotion data; optimize the syntactic complexity and sentiment words in the text of the virtual digital human according to the text emotion data; adjust the changes in the eyes, eyebrows and mouth corners of the virtual digital human according to the facial emotion data; adjust the information display order and visual focus guidance according to the attention state data; The calculation formula of the text emotion bias degree is: , Among them, is the text emotion bias degree, is the text weight of the th sentiment word, is the emotion score of the th sentiment word, is the total number of sentiment words, is the index of the sentiment word, is the total number of words in the text data, is the number of sentences containing the subject in the text data, is the total number of sentences in the text data, is the proportion of sentiment words in the text data, is the syntactic complexity, is the semantic consistency of the th sentiment word in the sentence, and are coefficients; The sentence semantic consistency is obtained by multiplying the cosine similarity between the sentiment word and the overall vector of the sentence by the weight of the sentiment word in the sentence. The weight of the sentiment word in the sentence can be calculated as , where is the weight of the -th sentiment word in the sentence, is the weight of the -th sentiment word in the text data, and is equivalent to in the formula for the text emotion bias degree. The two have the same meaning and value. is the emotion intensification factor of the -th sentiment word in the sentence, is the dependency distance factor, .
2. The AI-based virtual digital human interaction system according to claim 1, wherein The feature extraction module includes: A speech feature extraction unit, which is used to calculate the fundamental frequency curve of the speech data according to the interaction dataset, obtain the speech fundamental frequency feature, perform frequency domain conversion on it, obtain the speech time-frequency feature, and calculate the energy value of the speech data according to the speech time-frequency feature to obtain the speech energy feature; A text feature extraction unit, which is used to perform structural analysis on the text data according to the interaction dataset, extract the basic structural information of the text data to obtain the text structure feature, and extract the sentiment words in the text data according to it to obtain the text sentiment word feature; A facial feature extraction unit, which is used to extract the position information of the eyes, eyebrows, mouth corners and facial contour according to the interaction dataset, identify the key points of the facial data to obtain the facial key point feature, and obtain the expression change feature by analyzing the displacement of the key points between adjacent time frames; A limb feature extraction unit, which is used to extract the position information of the head, arms, torso and legs according to the interaction dataset to obtain the limb joint point feature, and draw the trajectory curve of the user's action according to it to analyze the user's movement pattern to obtain the limb action trajectory feature.
3. An AI-based virtual digital human interaction system according to claim 2, wherein, The speech analysis module includes: An audio change calculation unit, which is used to analyze the change trend of the speech fundamental frequency according to the speech fundamental frequency feature, calculate the audio change rate of the speech fundamental frequency, and obtain the speech fundamental frequency change data; A time series unit, which is used to calculate the energy value of the speech data within a fixed time window according to the speech energy feature to obtain the speech energy time series data; An energy fluctuation calculation unit, which is used to calculate the energy change rate between consecutive time windows according to the speech energy time series data to obtain the speech energy fluctuation data; A speech rate stability analysis unit, which is used to calculate the standard deviation of the energy fluctuation of the speech data according to the speech energy fluctuation data, analyze the degree of energy value fluctuation over time, and obtain the speech rate stability data; A voice emotion judgment unit, which is used to judge the type of the user's voice emotion according to the voice fundamental frequency change data and the speech rate stability data, and obtain voice emotion data.
4. An AI-based virtual digital human interaction system according to claim 3, characterized in that The speech rate stability analysis unit includes: A standard deviation calculation unit, which is used to calculate the average value of the energy fluctuation of the voice data according to the voice energy fluctuation data, and calculate the standard deviation of the voice data according to the average value of the energy fluctuation, so as to obtain voice energy standard deviation data; A normalization processing unit, which is used to perform normalization processing on the voice energy standard deviation data to obtain normalized speech rate stability data; A speech rate evaluation unit, which is used to judge the speech rate stability of the user according to the normalized speech rate stability data. When the normalized speech rate stability data tends to 1, the speech rate tends to be stable; when the normalized speech rate stability data tends to 0, the speech rate tends to be unstable.
5. An AI-based virtual digital human interaction system according to claim 4, characterized in that, The text analysis module includes: A text syntax analysis unit, which is used to perform syntax analysis on the text data according to the text structure characteristics, calculate the syntax complexity, and obtain text syntax data; A text emotion word analysis unit, which is used to count the number of emotion words in the text data according to the text emotion word characteristics, and calculate the emotion word frequency, so as to obtain text emotion word data; An emotion score calculation unit, which is used to calculate the emotion score of the emotion word according to the text emotion word data and a preset emotion dictionary; An emotion word weight calculation unit, which is used to determine the text weight of the emotion word in the text data according to the emotion score; A sentence meaning consistency calculation unit, which is used to determine the sentence weight of the emotion word in the sentence according to the text syntax data, and calculate the sentence meaning consistency of the emotion word in the sentence according to it; A text emotion bias degree calculation formula unit, which is used to calculate the text emotion bias degree according to the sentence meaning consistency, so as to obtain text emotion bias data; A text emotion judgment unit, which is used to judge the type of the user's text emotion according to the text emotion bias data, and obtain text emotion data.
6. The AI-based virtual digital human interaction system according to claim 5, wherein, The facial analysis module includes: A facial key point position unit, which is used to determine the relative positions of the eyes, eyebrows, and mouth corners according to the facial key point characteristics, and obtain facial key point position data; A facial expression change unit, which is used to calculate the key point offset between adjacent time frames according to the facial key point position data and the expression change characteristics, analyze the facial expression change trend, and obtain facial expression dynamic data; A facial expression symmetry calculation unit, which is used to calculate the left-right symmetry index of the face according to the facial expression dynamic data, analyze the stability degree of the facial expression, and obtain facial expression symmetry data; A facial emotion judgment unit, which is used to judge the type of the user's facial emotion according to the facial symmetry data, and obtain facial emotion data.
7. An AI-based virtual digital human interaction system according to claim 6, characterized in that, The limb analysis module includes: A body stability calculation unit, which is used to analyze the movement trajectory of the user according to the limb movement trajectory characteristics, calculate the movement frequency per unit time, and obtain body stability data; A visual attention analysis unit, which is used to identify the visual fixation area of the user according to the head position and eye direction, and obtain visual attention data; An attention state judgment unit, which is used to judge the attention state of the user according to the body stability data and the visual attention data, and obtain attention state data.
8. An AI-based virtual digital human interaction system according to claim 7, characterized in that, The optimization and adjustment module includes: A voice fundamental frequency adjustment unit, which is used to increase the upper limit of the voice fundamental frequency of the virtual digital human according to the voice emotion data, making the voice data more high-pitched; A speech rate and pitch adjustment unit, which is used to dynamically optimize the speech rate and pitch of the virtual digital human according to the voice emotion data, making the voice data more fluent; A syntactic complexity adjustment unit, which is used to optimize the syntactic complexity in the text of the virtual digital human according to the text emotion data, making it conform to the user's comprehension ability and emotional state; An emotional word frequency adjustment unit, which is used to adjust the usage frequency of emotional words in the text of the virtual digital human according to the text emotion data, enhancing the user's emotional resonance.
9. An AI-based virtual digital human interaction system according to claim 8, characterized in that, The optimization and adjustment module further includes: A blink frequency adjustment unit, which is used to optimize the blink frequency of the virtual digital human according to the facial emotion data, making the virtual digital human show a focused state; A mouth corner and eyebrow adjustment unit, which is used to optimize the curvature of the mouth corners and the angles of the eyebrows of the virtual digital human according to the facial emotion data, making them conform to the user's emotional state; An information presentation order adjustment unit, which is used to adjust the information presentation order of the virtual digital human according to the attention state data, making it conform to the user's attention concentration state; A visual guidance adjustment unit, which is used to adjust the line of sight direction and head movement of the virtual digital human according to the attention state data, making them conform to the user's visual focus.
Citation Information
Patent Citations
Virtual digital human interaction method and system
CN118426593A
Digital human control method and device based on multiple modes and electronic equipment
CN119441403A