Multi-modal digital human interaction method and system based on large language model

The multi-modal digital human interaction method addresses high workload and emotional understanding gaps by analyzing video, audio, and text data to generate emotionally responsive interactions, enhancing user experience.

CN120317879AActive Publication Date: 2025-07-15SHENYANG HANHUA SOFTWARE CO LTD

Patent Information

Application Number
CN202510804391.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-15
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

Traditional manual customer service has high intensity and low efficiency, and digital human interaction lacks active perception and dynamic understanding of the user's emotional state, resulting in mechanized interaction process and poor user experience.

Method used

A multimodal digital human interaction method based on a large language model is adopted to analyze video, audio and text data, and combine emotional dialogue generation models to generate replies that match the user's emotional state.

Benefits of technology

It improves user interaction experience, accurately perceives user emotions through multimodal data fusion, generates emotional replies that are more in line with user needs, and reduces deviations in user intentions and emotional understanding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120317879A_ABST
    Figure CN120317879A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large language models, in particular to a multi-modal digital human interaction method and system based on a large language model, and the method comprises the steps: obtaining video data, audio signals and text data of a single conversation when a user interacts with a digital human; extracting an envelope line of the audio signal, and determining an emotion evaluation value and a semantic contrast ratio of a single dialogue; identifying emotions contained in each frame of image in the video data to form an emotion sequence of a single dialogue; obtaining a semantic deviation value of a single dialogue; obtaining a semantic guide value of a single dialogue; and generating a corresponding emotional dialogue in combination with the emotional dialogue generation model. According to the method and the device, the deviation between the understanding of the intention expressed by the user and the perception of the emotion can be reduced, and the emotion state of the user can be more accurately judged, so that the generated emotion reply content is matched with the emotion state of the user, and the user interaction experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of large language models, and in particular to a multimodal digital human interaction method and system based on a large language model. Background Art

[0002] Due to the problems of large number of customer consultations and high repetitiveness of consultation content in the service mode of manual customer service, the workload of traditional manual customer service is relatively high, which is prone to low service efficiency and unstable service quality. Therefore, the introduction of digital human technology has alleviated this problem, but in the traditional interaction scenarios between people and digital humans, a single-mode passive response mode is usually adopted, lacking the ability to actively perceive and dynamically understand the user's emotional state, resulting in a mechanized interaction process and a stiff user experience.

[0003] Due to the ambiguity of text characters in the natural language processing process and the multidimensional expression of human emotions, such as tone and speaking speed, this unimodal sentiment analysis method will have deviations in the understanding of human intentions and the perception of emotions when faced with the ambiguity of natural language and the complexity of human emotional expression. It cannot accurately identify human emotions and intentions, resulting in a mismatch between the emotional response content generated by the digital human and the user's emotional state, making the interaction process lack emotional resonance and the user interaction experience poor. Summary of the invention

[0004] In order to solve the above technical problems, a multimodal digital human interaction method and system based on a large language model are provided to solve the existing problems.

[0005] The solution to the technical problem of this application is to provide a multimodal digital human interaction method and system based on a large language model, including the following steps: In a first aspect, an embodiment of the present application provides a multimodal digital human interaction method based on a large language model, the method comprising the following steps: Obtain video data, audio signals, and text data of a single conversation between a user and a digital human; Extracting the envelope of the audio signal, analyzing the extreme range and average level of the interval time between adjacent peaks on the envelope, and combining the discrete conditions of all peak values to determine the emotion evaluation value of a single conversation; Identify polysemous words, reduplication words and antonyms contained in the text data; calculate the semantic contrast of a single conversation according to the proportion of the number of polysemous words, reduplication words and antonyms in the text data; Identify the emotions contained in each frame of the video data to form an emotion sequence of a single conversation; analyze the discreteness of the elements in the emotion sequence, and combine the semantic contrast with the emotion evaluation value to obtain the semantic deviation of the single conversation; Obtain the semantic guidance value of a single conversation by combining the deviation of the maximum element within the emotion sequence of a single conversation and the difference in the emotion sequences between two adjacent conversations with the semantic deviation amount. Generate a corresponding emotional conversation based on the semantic guidance value and the text data, in combination with an emotional dialogue generation model.

[0006] Preferably, an envelope extraction algorithm is used to extract the envelope line of the audio signal.

[0007] Preferably, the determination of the emotion evaluation value of a single conversation includes: Obtain all the wave peaks on the envelope line. Calculate the time interval between the moments corresponding to two adjacent wave peaks on the envelope line; calculate the ratio of the range to the mean of all the time intervals between two adjacent wave peaks on the envelope line, denoted as the relative ratio. Calculate the degree of dispersion of the peak values of all the wave peaks on the envelope line, denoted as the first degree of dispersion. The emotion evaluation value is the normalized result of the sum of the relative ratio and the first degree of dispersion.

[0008] Preferably, the identification of polysemous words, reduplicated words, and antonyms contained in the text data includes: identifying the polysemous words, reduplicated words, and antonyms contained in the text data through a polysemous word library and an antonym pair library using regular expressions.

[0009] Preferably, the calculation of the semantic contrast of a single conversation includes: taking the ratio between the sum of the total number of characters of all polysemous words, the total number of characters of all reduplicated words, and the total number of characters of all antonyms in the text data and the total number of words in the text data as the semantic contrast of a single conversation.

[0010] Preferably, the further acquisition process of the emotion sequence is: perform frame-by-frame processing on the video data to obtain each frame of image, identify the emotion label corresponding to each frame of image through a neural network algorithm, and form an emotion sequence with the emotion labels of all the frames of image corresponding to a single conversation.

[0011] Preferably, the obtaining of the semantic deviation amount of a single conversation includes: Calculate the degree of dispersion of the emotion sequence, denoted as the second degree of dispersion. Calculate the cumulative sum of the semantic contrast and the emotion evaluation value, and take the normalized result of the product of the cumulative sum and the second degree of dispersion as the semantic deviation amount of a single conversation.

[0012] Preferably, the obtaining of the semantic guidance value of a single conversation includes: Calculate the difference between the mode of the emotion sequence of a single conversation and the mode of the emotion sequence of the previous conversation, which is denoted as the emotion difference amount; Denote the difference between the maximum value in the emotion sequence of a single conversation and a preset value as the emotion deviation; Calculate the sum value of the emotion deviation and the emotion difference amount, and take the normalized value of the product of the sum value and the semantic deviation amount as the semantic guidance value of a single conversation.

[0013] Preferably, the generating the corresponding emotional conversation includes: During the backpropagation process of the emotional conversation generation model, introduce a random variable when updating the weights, so that the formula for weight update is: , where is the updated weight, is the weight before update, is the learning rate, is the gradient calculated by the loss function, is the semantic guidance value, is the introduced random variable, where obeys the normal distribution; Use the text data and the semantic guidance value as the input of the emotional conversation generation model, and introduce a random variable to update the weights based on the semantic guidance value during the backpropagation process, and generate the corresponding emotional conversation in combination with the updated weights.

[0014] In a second aspect, an embodiment of the present application further provides a multimodal digital human interaction system based on a large language model, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of the multimodal digital human interaction method based on the large language model described in any one of the above.

[0015] The present application has at least the following beneficial effects: This application obtains multiple modal data, including video, text, and audio data, to perceive the user's emotions from multiple dimensions, thereby being able to more accurately grasp the user's emotions; by analyzing the time interval between adjacent wave peaks on the envelope line in the audio signal, the emotional evaluation value of a single conversation is calculated, and its beneficial effect is that it takes into account the eagerness of the user's speech rate and tone during the expression process to reflect the user's emotional fluctuations; secondly, by calculating the semantic contrast of a single conversation based on the polysemous words, antonyms, and reduplicated words in the text data, its beneficial effect is that it takes into account the complex context in the user's expression sentences and the situation of irony in the context, further reflecting the user's emotional state; by analyzing the changes in the user's facial expressions in the video data, combining the semantic contrast and the emotional evaluation value, the semantic deviation amount of a single conversation is obtained, and its beneficial effect is that it takes into account the user's facial expressions during expression, thereby perceiving the user's emotional state. By fusing the results of multiple modal data for emotion perception, the user's emotions and the ambiguity of semantic expressions are described to reflect the possible deviations in the understanding of the user's semantics. By analyzing the change difference in the user's facial expressions in two adjacent conversations, the semantic guidance value of a single conversation is obtained, and its beneficial effect is that it analyzes the situation where the user's emotions change due to dissatisfaction with the robot's expression content during the previous and subsequent conversations to reflect the situation where the reply content generated by the robot should match the user's emotions at this time; based on the semantic guidance value and the text data, combined with the emotional dialogue generation model, the corresponding emotional dialogue is generated, and its beneficial effect is that by adding a random variable for updating the weight based on the semantic guidance value during the model backpropagation process, the weight has a certain degree of randomness during the update process, so as to explore a larger solution space and improve the emotional dialogue of the model under the complex text requirements of the user. Therefore, through the emotional analysis of multiple modalities, the deviation in the understanding of the user's expressed intention and emotion perception can be reduced, and the user's emotional state can be more accurately judged, so that the generated emotional reply content of the model matches the user's emotional state and improves the user interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The following further details the multi-modal digital human interaction method based on the large language model of the present application with reference to the drawings.

[0017] Figure 1 It is a flowchart of the steps of the multi-modal digital human interaction method based on the large language model provided by the embodiment of the present application; Figure 2 It is a flowchart of the steps of the method for obtaining the semantic guidance value provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] To make the objectives, technical solutions, and advantages of this application more clear and understandable, the multi-modal digital human interaction method and system based on large language models proposed in this application will be further described in detail below in combination with the accompanying drawings and implementation examples. It should be understood that the specific implementation examples described herein are only used to explain this application and are not used to limit this application.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.

[0020] Please refer to Figure 1 , which shows the step flow chart of the multi-modal digital human interaction method based on large language models provided by an embodiment of this application. The method includes the following steps: Step 1, obtain the video data, audio signal, and text data of a single conversation when a user interacts with a digital human.

[0021] In the process of information exchange between humans and digital humans, information interaction is mainly achieved through emotional conversations. The emotional conversation task can be divided into two subtasks: dialogue emotion perception and emotional conversation generation. The former focuses on the classification of emotions in conversations, and the latter focuses on generating emotional responses. The emotional conversation generation technology includes two methods: pure text conversation generation and multi-modal conversation generation. Among them, pure text conversation generation generates responses through the analysis and processing of text data. However, when perceiving emotions in text information, semantic ambiguity may occur due to the polysemous nature of the text, resulting in inaccurate semantic perception of the text information and affecting the generation quality of emotional conversations. Multi-modal conversation generation comprehensively processes multi-modal data such as audio, video, and text, and then performs emotion perception on the conversation to improve the conversation quality between humans and digital humans.

[0022] Based on the above analysis, deploy the digital human on the robot. Through the camera and microphone array on the robot, collect the video data and audio signal during the interaction between a single user and the robot. After filtering out environmental noise from the audio signal through a filtering algorithm, convert the audio signal into text data; regard the process from when the user asks a question to when the robot makes a response as a single conversation. Thus, obtain the video data, audio signal, and text data of the single conversation; In this embodiment, the Kalman filtering algorithm is used to filter the audio signal. Among them, the Kalman filtering algorithm and the process of converting audio to text are well-known technologies and will not be elaborated here.

[0023] So far, the video data, audio signal, and text data of the single conversation are obtained.

[0024] Step 2: Extract the envelope of the audio signal, analyze the extreme range and average level of the time intervals between adjacent wave peaks on the envelope, and combine the dispersion of all wave peak values to determine the emotional evaluation value of a single conversation.

[0025] During the human-computer interaction process, the user's intention can be recognized through the user's audio signal and text data. The audio signal can reflect the user's emotional level, while the text data can more intuitively express the user's specific intention. Among them, the amplitude and speech rate of the user's voice in the audio signal will generate corresponding electrical signals, and this electrical signal can express the user's emotional state. After the user expresses the corresponding intention in a relatively gentle tone, if the robot's answer is not sufficient to solve the user's problem, the user will have an emotional expression at this time, which may lead to an increase in the speech rate, cadence between each audio, large fluctuations in the audio symbols, and at the same time, there may be large fluctuations in the audio amplitude. Therefore, by analyzing the changes in the signal amplitude in the audio signal and calculating the emotional evaluation value, the user's emotion can be evaluated. Specifically: Adopt an envelope extraction algorithm to extract the envelope of the audio signal; In this embodiment, the Hilbert transform algorithm is used to extract the envelope. Among them, the Hilbert transform algorithm is a well-known technology and will not be elaborated here. As other implementation manners, implementers can use other methods of existing technologies, such as the cepstrum method, etc. This embodiment does not make special restrictions on this.

[0026] Obtain all the wave peaks on the envelope; In this embodiment, the AMPD (Automatic multiscale-based peak detection) algorithm is used to obtain the wave peaks. Among them, the AMPD algorithm is a well-known technology and will not be elaborated here.

[0027] Calculate the time interval between the moments corresponding to adjacent wave peaks on the envelope; Calculate the ratio of the range to the mean value of the time intervals of all adjacent wave peaks on the envelope, denoted as the relative ratio; Calculate the dispersion degree of all wave peak values on the envelope, denoted as the first dispersion degree; In this embodiment, the dispersion degree is measured by calculating the variance of all wave peak values on the envelope. As other implementation manners, implementers can use other methods of existing technologies, such as the standard deviation, etc. This embodiment does not make special restrictions on this.

[0028] Take the normalized result of the sum of the relative ratio and the first dispersion degree as the emotional evaluation value of a single conversation; In this embodiment, the sigmoid function is used for normalization processing, wherein the sigmoid function is a well-known technology and will not be described in detail here. As other implementation methods, the implementer can adopt other methods of the prior art, such as the softmax function, etc. This embodiment does not impose any special restrictions on this.

[0029] It should be noted that the time interval reflects the time interval between adjacent syllables in the user's language expression process. The larger the range, the uneven the speaking speed of the user, and the ups and downs between different syllable segments. The smaller the range, the relatively stable time interval between adjacent peaks, and the user's speaking speed is relatively stable during the entire conversation. The smaller the mean, the faster the speaking speed of the user, and the more urgent the expression. The larger the first discreteness, the more drastic the change in the strength of the voice signal, and the larger the obtained emotion evaluation value, reflecting that the user's emotional fluctuations are more significant.

[0030] At this point, the emotion evaluation value of a single conversation is obtained.

[0031] Step 3, identifying polysemous words, reduplication words and antonyms contained in the text data; calculating the semantic contrast of a single conversation according to the proportion of the number of polysemous words, reduplication words and antonyms in the text data.

[0032] Secondly, the intention expressed by the user needs to be reflected through text. Among them, text is composed of words or characters, including emotional vocabulary and semantic information, and is an important carrier of emotional expression. Due to the multidimensionality of human emotions and the complex characteristics of text, the same words may have different meanings in different contexts. When human emotions are expressed in text, they often contain sarcasm and polysemy. Therefore, when users express sarcasm or ridicule in their words, they usually express it with reduplicated words or sentences containing antonyms, such as "right, right, right", "You are so smart that you messed up such a simple thing", etc., expressing the opposite meaning of the word through statements. At the same time, polysemous words appear in the text data of a single conversation to express ideas that are opposite to the meaning of the statement.

[0033] Based on the above analysis, the semantic contrast is calculated by analyzing the characteristics of word meanings in text data, specifically: Based on the polysemous word library and the antonym pair library, regular expressions are used to identify polysemous words, reduplications, and antonyms contained in the text data; It should be noted that the polysemous word library and the antonym pair library can be constructed by collecting and organizing a large number of polysemous words and antonym pairs, or the publicly available polysemous word dictionary and antonym pair dictionary can be used as the polysemous word library and the antonym pair library, and regular expressions are used to find antonyms, polysemous words and reduplicated words in the text data. Among them, regular expressions are well-known technologies and will not be elaborated here.

[0034] The ratio between the sum of the total number of characters of all polysemous words, the total number of characters of all reduplicated words and the total number of characters of all antonyms in the text data and the total number of words in the text data is used as the semantic contrast ratio of a single conversation. It should be noted that since Chinese characters pay more attention to context, when the user expresses ironic or sarcastic meanings in a declarative way to express dissatisfaction with the robot's answer, there will often be more ironic text features at this time, such as the appearance of certain reduplicated words to emphasize the attitude, the appearance of more antonyms to reflect irony, and the appearance of more synonyms to construct a complex context, etc. Therefore, the proportion of these three types of words in the text data is relatively large, and the resulting semantic contrast ratio is relatively high, indicating that the text data in this conversation contains more complex semantics and the user expresses more ironic meanings, thereby reflecting the user's emotional state.

[0035] Thus, the semantic contrast ratio of a single conversation is obtained.

[0036] Step 4: Identify the emotions contained in each frame of the video data to form an emotion sequence of a single conversation; analyze the dispersion of the elements in the emotion sequence, and combine the semantic contrast ratio and the emotion evaluation value to obtain the semantic deviation amount of a single conversation.

[0037] During a single conversation, when the user expresses their intention, there will be emotional information in the facial expression. The emotional state of the user is perceived by recognizing the user's facial expression in the video data. Therefore, it is necessary to recognize the user's expression in the video data. Specifically: Perform frame-by-frame processing on the video data to obtain each frame of image, recognize each frame of image through a neural network algorithm, obtain the emotion label corresponding to each frame of image, and form an emotion sequence with the emotion labels of all frames of images corresponding to a single conversation. In this embodiment, the ResNet18 algorithm is used to recognize the emotion of each frame of image, and the emotion label of each frame of image is obtained. Among them, the ResNet18 algorithm is trained through the publicly available dataset FER2013. Among them, 7 emotions are labeled in the dataset, including anger, disgust, sadness, fear, neutral, surprise and happiness. Each emotion corresponds to an emotion label, specifically: anger corresponds to emotion label 1, disgust corresponds to emotion label 2, sadness corresponds to emotion label 3, fear corresponds to emotion label 4, neutral corresponds to emotion label 5, surprise corresponds to emotion label 6, and happiness corresponds to emotion label 7. Therefore, the emotion of each frame of image is recognized by the trained ResNet18 algorithm, and the emotion label of each frame of image is obtained; among them, the ResNet18 algorithm is a well-known technology and will not be elaborated here.

[0038] Secondly, calculate the semantic deviation amount based on the user's facial expression, the semantic situation of the text, and the emotional state in the audio, specifically: Calculate the dispersion degree of the emotion sequence, denoted as the second dispersion degree; In this embodiment, the dispersion degree is measured by calculating the standard deviation of the emotion sequence. As other implementation manners, implementers can adopt other methods in the prior art, such as variance, coefficient of variation, etc. This embodiment does not make special restrictions on this.

[0039] Calculate the cumulative sum of the semantic contrast and the emotion evaluation value, and take the normalized result of the product of the cumulative sum and the second dispersion degree as the semantic deviation amount of a single conversation; In this embodiment, the sigmoid function is used for normalization processing. Among them, the sigmoid function is a well-known technology and will not be elaborated here. As other implementation manners, implementers can adopt other methods in the prior art, such as softmax function, tanh function, etc. This embodiment does not make special restrictions on this.

[0040] It should be noted that when the content expressed by the robot is not sufficient to solve the user's actual problem, or when the user is relatively eager during the questioning process, it will affect the user's language expression, which may cause changes in the user's emotions. Due to the multi-dimensionality of emotions and the complexity of Chinese characters in the multi-modal data, the expressed word meaning may be contrary to the true intention. Correspondingly, the fluctuation of the emotion label of each frame of image in the video data is relatively large, and the obtained second dispersion degree is larger; the speech rate in the audio data is relatively fast and fluctuates greatly, and the obtained emotion evaluation value is larger. The semantic complexity in the text data is relatively high, and the obtained semantic contrast is larger. Therefore, the obtained semantic deviation amount is larger, indicating that the user's emotion is intense and the semantic expression is implicit, and the deviation of the robot's understanding of the user's semantics may be larger.

[0041] Step 5: Based on the deviation of the maximum element within the emotion sequence of a single conversation and the difference in the emotion sequences between two adjacent conversations, and in combination with the semantic deviation amount, obtain the semantic guidance value of the single conversation; based on the semantic guidance value and the text data, and in combination with the emotional dialogue generation model, generate the corresponding emotional dialogue.

[0042] Furthermore, affected by the previous and subsequent conversations, it may be that the expression method of the robot in a certain conversation cannot satisfy the user, or the answer content is not good enough, resulting in a change in the user's emotion. Therefore, based on the emotional change of the user's facial expression in two adjacent conversations, and in combination with the semantic deviation amount, determine the semantic guidance value. Specifically: Calculate the difference between the mode of the emotion sequence of a single conversation and the mode of the emotion sequence of the previous conversation, and denote it as the emotion difference amount. In this embodiment, calculate the absolute value of the difference between the mode of the emotion sequence of a single conversation and the mode of the emotion sequence of the previous conversation, and denote it as the emotion difference amount.

[0043] It should be noted that if the user and the robot are having a conversation for the first time, the emotion difference amount is assigned a value of 0.

[0044] Denote the difference between the maximum value in the emotion sequence of a single conversation and a preset value as the emotion deviation. In this embodiment, since the emotion label corresponding to the happy emotion in the facial expression is 7, the preset value is set to 7. As other implementation manners, the implementer can set it according to the actual situation; therefore, denote the absolute value of the difference between the maximum value in the emotion sequence of a single conversation and the preset value as the emotion deviation.

[0045] Calculate the sum value of the emotion deviation and the emotion difference amount, and take the normalized value of the product of the sum value and the semantic deviation amount as the semantic guidance value of the single conversation. In this embodiment, the calculation formula for the semantic guidance value of a single conversation is: Wherein, is the semantic guidance value of a single conversation, is the semantic deviation amount of a single conversation, is the emotion difference amount of a single conversation, is the emotion deviation of a single conversation, is a preset value greater than 0 to avoid the result of being 0. In this embodiment, the preset value greater than 0 is set to 1. As other implementation manners, the implementer can set it according to the actual situation, is a normalization function. In this embodiment, the sigmoid function is used for normalization processing. The sigmoid function is a well-known technology and will not be elaborated here. As other implementation manners, implementers can adopt other methods of the prior art. For example, the softmax function, the tanh function, etc. This embodiment does not make special restrictions on this.

[0046] It should be noted that during the two consecutive conversations, the greater the emotional fluctuation of the user, that is, the greater the emotional difference, and the greater the deviation between the emotional label corresponding to the happy label, that is, the greater the emotional deviation. At this time, the content expressed by the robot cannot be recognized by the customer, and it may affect the logical analysis of the user's semantics by the robot with a more complex expression, and it is more likely to have the problem of semantic ambiguity. The greater the obtained semantic guidance value, the more it indicates that the robot should generate a conversation with an appeasing emotion. Among them, the step flow chart of the method for obtaining the semantic guidance value provided by the embodiment of the present application is as Figure 2 shown.

[0047] Train the emotional dialogue generation model. In this embodiment, train the Variational Autoencoders (VAE) model. Thus, the training process is as follows: Collect multi-modal data of multiple users having multiple conversations with the robot. The multi-modal data includes video data, audio data, and text data to form an experimental data set; calculate the semantic guidance value of each conversation in the experimental data set; In this embodiment, collect the multi-modal data of 2000 users during conversations to form an experimental data set. As other implementation manners, implementers can set it by themselves according to the actual situation.

[0048] In the traditional VAE model during the backpropagation process, the weight update formula is , and in this embodiment, during the backpropagation process of the VAE model, when updating the weights, a random variable is introduced. Based on the semantic guidance value, the weight update formula is: , where is the updated weight, is the weight before update, is the learning rate, is the gradient calculated through the loss function, is the semantic guidance value, is the introduced random variable, where obeys the normal distribution.

[0049] In this embodiment, the learning rate is set to 0.01; where obeys the normal distribution with a mean of 0 and a standard deviation of the gradient. As other implementation manners, implementers can set it by themselves according to the actual situation.

[0050] Thus, the VAE model is trained using all the text data in the experimental dataset and the corresponding semantic guidance values. The text data and semantic guidance value of a single conversation are used as the input to the trained VAE model. During the backpropagation process, based on the semantic guidance value, a random variable is introduced to update the weights. Combining the updated weights, a corresponding emotional conversation is then generated. In this embodiment, the VAE model is a well-known technology and will not be elaborated here.

[0051] It should be noted that when training the VAE model, multimodal data of multiple conversations between multiple users and the robot is collected. The multimodal data includes video data, audio data, and text data, which constitute an experimental dataset. The semantic guidance value of each conversation in the experimental dataset is calculated. Thus, in the backpropagation process of the VAE model, a random variable for updating the weights based on the semantic guidance value is added, making the weights have a certain degree of randomness during the update process. This enables exploration of a larger solution space, improves the emotional conversation of the model under complex text requirements of users, and enhances the quality of emotional conversations. Therefore, the VAE model is trained using the experimental dataset, and the trained VAE model is deployed on the robot. The robot collects multimodal data in real-time during each conversation with the user, and then generates corresponding emotional conversations according to the customer's needs and emotional conditions, improving the accuracy of user intention recognition and enhancing the service quality and service efficiency for users.

[0052] Based on the same inventive concept as the above method, an embodiment of the present application further provides a multimodal digital human interaction system based on a large language model, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, it implements the steps of any one of the above methods of the multimodal digital human interaction method based on a large language model.

[0053] It should be understood that although Figure 1 the steps in the flowchart are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 1 at least a part of the steps in

[0054] The technical features of the above embodiments can be combined arbitrarily. For the sake of concise description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0055] The above-described embodiments only represent several implementation manners of the present application, and their descriptions are relatively specific and detailed. However, it should not be construed as a limitation to the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made. Therefore, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application all belong to the protection scope of the technical solution of the present application.

Claims

1. A multimodal digital human interaction method based on a large language model, characterized in that The method includes the following steps: Obtain the video data, audio signal, and text data of a single conversation when a user interacts with a digital human; Extract the envelope of the audio signal, analyze the extreme range and average level of the time interval between adjacent wave peaks on the envelope, and combine the discrete situation of all wave peak values to determine the emotion evaluation value of a single conversation; Identify the polysemous words, reduplicated words, and antonyms contained in the text data; calculate the semantic contrast of a single conversation based on the proportion of the number of characters of the polysemous words, reduplicated words, and antonyms in the text data; Identify the emotion contained in each frame of the video data to form an emotion sequence of a single conversation; analyze the discrete situation of the elements in the emotion sequence, and combine the semantic contrast and the emotion evaluation value to obtain the semantic deviation amount of a single conversation; Obtain the semantic guidance value of a single conversation through the deviation situation of the maximum element in the emotion sequence of a single conversation and the difference situation of the emotion sequences between two adjacent conversations, combined with the semantic deviation amount; Generate a corresponding emotional conversation based on the semantic guidance value and the text data, combined with an emotional dialogue generation model; The obtaining of the semantic deviation amount of a single conversation includes: Calculate the degree of dispersion of the emotion sequence, denoted as the second dispersion degree; Calculate the cumulative sum of the semantic contrast and the emotion evaluation value, and use the normalized result of the product of the cumulative sum and the second dispersion degree as the semantic deviation amount of a single conversation; The obtaining of the semantic guidance value of a single conversation includes: Calculate the difference between the mode of the emotion sequence of a single conversation and the mode of the emotion sequence of the previous conversation, denoted as the emotion difference amount; Denote the difference between the maximum value in the emotion sequence of a single conversation and a preset value as the emotion deviation; Calculate the sum value of the emotion deviation and the emotion difference amount, and use the normalized value of the product of the sum value and the semantic deviation amount as the semantic guidance value of a single conversation.

2. The multimodal digital human interaction method based on a large language model according to claim 1, wherein Adopt an envelope extraction algorithm to extract the envelope of the audio signal.

3. The multimodal digital human interaction method based on a large language model according to claim 1, wherein, The determination of the emotion evaluation value of a single conversation includes: Obtain all the wave peaks on the envelope; Calculate the time interval between the moments corresponding to two adjacent wave peaks on the envelope; calculate the ratio of the range and the mean of all the time intervals between two adjacent wave peaks on the envelope, denoted as the relative ratio; Calculate the degree of dispersion of all the wave peak values on the envelope, denoted as the first dispersion degree; The emotion evaluation value is the normalized result of the sum of the relative ratio and the first dispersion degree.

4. The multimodal digital human interaction method based on a large language model according to claim 1, wherein The identification of the polysemous words, reduplicated words, and antonyms contained in the text data includes: identifying the polysemous words, reduplicated words, and antonyms contained in the text data through a polysemous word library and an antonym pair library using regular expressions.

5. The multimodal digital human interaction method based on a large language model according to claim 1, characterized in that The calculation of the semantic contrast of a single conversation includes: using the ratio between the sum of the total number of characters of all polysemous words, all reduplicated words, and all antonyms in the text data and the total number of all words in the text data as the semantic contrast of a single conversation.

6. The multimodal digital human interaction method based on a large language model according to claim 1, wherein The further process of obtaining the emotion sequence is as follows: The video data is frame-divided to obtain each frame of image, and the emotion label corresponding to each frame of image is identified through a neural network algorithm. The emotion labels of all frames of images corresponding to a single conversation are combined to form an emotion sequence.

7. The multimodal digital human interaction method based on a large language model according to claim 1, characterized in that The generation of the corresponding emotional conversation includes: During the backpropagation process of the emotional dialogue generation model, a random variable is introduced when updating the weights, so that the formula for weight update is: , where is the updated weight, is the weight before update, is the learning rate, is the gradient calculated through the loss function, is the semantic guidance value, is the introduced random variable, where follows a normal distribution; The text data and the semantic guidance value are used as the inputs of the emotional conversation generation model. During the backpropagation process, based on the semantic guidance value, a random variable is introduced to update the weights. Combining the updated weights, the corresponding emotional conversation is generated.

8. A multimodal digital human interaction system based on a large language model, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the multi-modal digital human interaction method based on the large language model according to any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-round dialogue semantic comprehension subsystem based on multi-modal emotion identification system

    CN108877801A

  • Intelligent real-time emotion evaluation method for social media and online text data based on multi-modal knowledge graph

    CN119202270A

  • Emotion change judgment method, device, equipment and medium based on semantic recognition

    CN119740583A

  • User emotion recognition method based on AI and voice data

    CN120148561A

  • Sentence generation method, electronic device and storage medium

    WO2023108994A1

Cited By

  • Evaluation method and device for generative model content

    CN122154668A

  • A method and apparatus for evaluating generative model content

    CN122154668B