Digital human emotion processing method and system based on AI model

By using an AI model-based approach, the type of emotion of the digital human is determined according to the application context, and the replica object is selected and switched. This solves the problem of insufficient accuracy of digital human emotion processing in different situations, and achieves efficient use of training data and reliable processing results.

CN120895060BActive Publication Date: 2026-03-20HANGZHOU BEIMING AURORA TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2026-03-20

AI Technical Summary

Technical Problem

In existing technologies, the accuracy of digital human emotion processing replication is insufficient due to the different needs of different application scenarios, and it is difficult to determine the replication target according to the needs of the scenario.

Method used

By using an AI model-based approach, we can determine the emotion types of digital humans in different application scenarios. By combining training data and emotion recognition results from voice data, we can select replicas that meet the requirements of similarity and data volume, and switch to other replicas for training in abnormal situations.

Benefits of technology

It improves the accuracy and reliability of digital human emotion processing, avoids the waste of training data, ensures that in abnormal situations, it can switch to a suitable replica object for training in a timely manner, and guarantees the reliability of training processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120895060B_ABST
    Figure CN120895060B_ABST
Patent Text Reader

Abstract

The application provides a digital human emotion processing method and system based on an AI model, and belongs to the technical field of digital humans, and specifically comprises the following steps: determining the problem emotion type of a target emotion type based on voice data in a training data combination; determining a screening replica object in a target replica object based on the training data of the problem emotion type of the target replica object; determining a replica object in the screening replica object based on the similarity of the emotion recognition results of the voice data of the screening replica object and other screening replica objects in the training data combination; and when the training processing result of the replica object is abnormal, determining the switching target of the replica object based on the training processing result of the replica object and the similarity with other screening replica objects, thereby improving the reliability and accuracy of the emotion processing of the digital human.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of digital people, and particularly relates to a digital person emotion processing method and system based on an AI model. BACKGROUND

[0002] In the prior art, the construction of a digital person can greatly reduce the pressure of manual customer service, and how to realize the replication processing of the digital person becomes a technical problem to be solved.

[0003] In order to realize the replication processing of the digital person, in the invention patent application CN202510013307.3 "Digital person image replication method and system based on NeRF", the expanded digital person image data is used to train a data reconstruction model, the features of the digital person image replication result output by the autoencoder are extracted, and the feature vectors after the feature extraction are input into a classifier model for digital person image data screening, so that the accuracy of the replication processing is improved.

[0004] During the replication processing of the digital person, the differences in application occasions will lead to differences in the degree of emotional processing demand, and therefore how to determine the replication object of the digital person according to the requirements of the application occasion to ensure the replication processing accuracy of the emotional processing demand with a higher degree becomes a technical problem to be solved.

[0005] In view of the above technical problems, the application provides a digital person emotion processing method and system based on an AI model. SUMMARY

[0006] To achieve the object of the application, the application adopts the following technical solutions:

[0007] Specifically, the application provides a digital person emotion processing method based on an AI model, which specifically includes:

[0008] S1, based on the application occasion of a digital person, determining the use data of the emotion type of the digital person in various application scenarios, and determining a target emotion type in the emotion type based on the use data;

[0009] S2, dividing the training data of the target emotion type into training data combinations according to the similarity of the face images corresponding to the training data, determining the problem emotion type of the target emotion type based on the voice data between the training data combinations, and determining a screening replication object in the target replication object according to the training data of the problem emotion type.

[0010] S3, determining a replication object in the screening replication object by using the similarity of the emotion recognition results of the voice data under the training data combination between the screening replication object and other screening replication objects.

[0011] S4 When the training processing result of the replica object is abnormal, determining the switching target of the replica object according to the training processing result of the replica object.

[0012] The application has the following beneficial effects:

[0013] The replica object is determined by screening the similarity of the emotion recognition results of the voice data of the replica object and other screened replica objects under the training data combination, thereby ensuring that the similarity of the voice data of the replica object under different target emotion types meets the requirements and the data quantity meets the requirements, and further considering the similarity with other screened replica objects, so that when the training processing result is abnormal, the other screened replica object can be switched to in time and effectively, and the reliability of the digital human training processing is ensured.

[0014] According to the training processing result of the replica object and the similarity with other screened replica objects, the switching target of the replica object is determined, which not only ensures the reliability of the data quantity of the training data of the switching target under the target emotion type of the replica object when the training processing result of the replica object is abnormal, but also ensures the similarity between the switching target and the training data of the replica object, thereby avoiding the waste of the original training data of the replica object and ensuring the reliability of the training processing.

[0015] The further technical solution is that the application scenarios are divided according to the business types of the digital human processing under the application occasions.

[0016] The further technical solution is that the emotion types include anger, confusion, anger, contempt, fear, joy, pain, sadness, expectation, anxiety, excitement and surprise.

[0017] The further technical solution is that the use data is determined according to the emotion recognition results of the artificial customer service in various application scenarios.

[0018] The further technical solution is that the method for determining the target emotion type in the emotion type is:

[0019] The use times of the emotion types of the artificial customer service in various application scenarios are determined by the use data of the emotion types in various application scenarios.

[0020] The matching value of the emotion types in various application scenarios is determined by the ratio of the use times of the emotion types of the artificial customer service in various application scenarios to the processing times of the artificial customer service in the application scenarios.

[0021] Determine whether the emotion type is a target emotion type according to an average value of matching values of emotion types in various application scenarios.

[0022] A further technical solution is that when the average value of matching values of emotion types in various application scenarios is greater than a preset matching value threshold, it is determined that the emotion type is a target emotion type.

[0023] A further technical solution is that the switching target of the cloned object is another filtered cloned object with the largest average value of data amounts of training data of different abnormal emotion types.

[0024] A further technical solution is that when switching to the switching target of the cloned object, the switching target of the cloned object is subjected to augmentation processing on voice data of different target emotion types based on voice data of the cloned object, and the digital person is subjected to training processing based on the augmented voice data.

[0025] In a second aspect, the present application provides a computer system, comprising a memory and a processor connected in communication, and a computer program stored on the memory and capable of running on the processor, wherein the processor executes the computer program to perform the above-mentioned AI model-based digital person emotion processing method.

[0026] Other features and advantages will be set forth in the following description of the application, and in part will be apparent from the description and the drawings, or can be learned by practice of the application as claimed in the claims.

[0027] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the following preferred embodiments are specifically described with reference to the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0028] The above-mentioned and other features and advantages of the present application will become more apparent by describing in detail example embodiments thereof with reference to the attached drawings.

[0029] Figure 1 is a flowchart of an AI model-based digital person emotion processing method;

[0030] Figure 2 is a flowchart of a method for determining a target emotion type in emotion types;

[0031] Figure 3 is a flowchart of a method for determining a problem emotion type of a target emotion type;

[0032] Figure 4 is a flowchart of a method for determining a filtered cloned object in a target cloned object;

[0033] Figure 5It is a framework diagram of a computer system. DETAILED DESCRIPTION

[0034] In order for those skilled in the technical field to better understand the technical solutions in the specification, the technical solutions in the specification will be clearly and completely described below in combination with the drawings in the specification. Obviously, the described embodiments are only part of the embodiments of the specification, not all. Based on the embodiments of the specification, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the specification.

[0035] In the present application, according to the variation of the voice data of the selected replication object in the target emotion type, the replication object with higher emotion stability of the voice data is selected, and the final replication object is selected based on the emotion stability of the voice data, which improves the accuracy of the replication processing.

[0036] Embodiment 1

[0037] As shown in Figure 1 The present application provides an AI model-based digital human emotion processing method, which specifically includes:

[0038] S1 determines the use data of the emotion type of the digital human in various application scenarios based on the application scenarios of the digital human, and determines the target emotion type in the emotion type based on the use data;

[0039] S2 divides the training data of the target emotion type into training data combinations based on the similarity of the face images corresponding to the training data, determines the problem emotion type of the target emotion type based on the voice data between the training data combinations, and determines the screening replication object in the target replication object based on the training data of the problem emotion type.

[0040] S3 determines the replication object in the screening replication object based on the similarity of the emotion recognition results of the voice data under the training data combination between the screening replication object and other screening replication objects.

[0041] S4 determines the switching target of the replication object based on the training processing result of the replication object when the training processing result of the replication object is abnormal.

[0042] Further, the application scenarios are divided according to the business types processed by the digital human in the application scenarios, as shown in Table 1, the target emotion types under different business types in the bank system.

[0043] Table 1 Target emotion types under different business types

[0044]

[0045] Specifically, the emotion types include anger, confusion, anger, contempt, fear, joy, pain, sadness, expectation, anxiety, excitement and surprise.

[0046] It should be noted that the use data is determined according to the emotion recognition result of the artificial customer service in various application scenarios.

[0047] Further, the voice features include acoustic features, speech intonation features, speech rhythm features, and speech spectrum features.

[0048] Specifically, as shown in Figure 2 The method for determining the target emotion type in the emotion types is:

[0049] The use data of the emotion types in various application scenarios is used to determine the use times of the emotion types of the artificial customer service in various application scenarios.

[0050] The use times of the emotion types of the artificial customer service in various application scenarios and the processing times of the artificial customer service in the application scenarios are used to determine whether the emotion types are target emotion types.

[0051] Further, when the proportion of the use times of the emotion types to the processing times of the artificial customer service in the application scenarios meets the required application scenarios, in one possible embodiment, greater than 0.4, it is determined that the emotion types are target emotion types.

[0052] In another possible embodiment, the method for determining the target emotion type in the emotion types is:

[0053] The use data of the emotion types in various application scenarios is used to determine the use times of the emotion types of the artificial customer service in various application scenarios.

[0054] The use times of the emotion types of the artificial customer service in various application scenarios are used to determine the total use times of the emotion types in the application scenarios.

[0055] According to the total use times of the emotion types in the application scenarios, it is determined whether the emotion types are target emotion types.

[0056] Further, when the total use times of the emotion types in the application scenarios are greater than a preset use time threshold, it is determined that the emotion types are target emotion types.

[0057] It should be noted that when there is no use data, the target emotion type is determined according to the preset emotion type corresponding to the application scenario.

[0058] Specifically, the similarity of the face image corresponding to the training data is determined according to the deviation of the image features of the face image.

[0059] Specifically, the training data of the target emotion type is divided into a training data combination, specifically including:

[0060] The training data corresponding to the face image with the deviation of the image features in the preset image deviation range is divided into the same training data combination, and specifically, the face images with the deviation rate of the image features between each other less than 0.05 are divided into the same training data combination. Specifically, the determination is made by using the Euclidean distance function.

[0061] It should be noted that, as shown in Figure 3 The method for determining the problem emotion type of the target emotion type is:

[0062] Based on the voice data of the training data combination, the voice features of each voice data in the training data combination are determined;

[0063] Based on the deviation of the voice features of each voice data, the voice data with the deviation of the voice features less than the preset voice feature deviation is divided into the same voice data combination, and the matching data combination of the training data combination is determined according to the voice data combination with the largest data amount of voice data in the training data combination.

[0064] According to the deviation of the voice features between the matching data combination of the training data combination, it is determined whether the target emotion type is a problem emotion type.

[0065] Further, the voice features of the matching data combination are determined according to the average value of the voice features of each voice data in the matching data combination. It can be understood that the voice features are MFCC features of the voice data, and specifically, the average value is determined after the MFCC features of the voice data are used.

[0066] It can be understood that, according to the deviation of the voice features between the matching data combination of the training data combination, it is determined whether the target emotion type is a problem emotion type, specifically including:

[0067] According to the deviation of the voice features between the matching data combination of the training data combination, the deviation of the voice features between the training data combinations is determined.

[0068] According to the average value of the deviation of the voice features between the training data combinations, it is determined whether the target emotion type is a problem emotion type.

[0069] Further, when the average value of the absolute value of the bias amount of the voice feature between the training data combinations is greater than a preset bias amount threshold, it is determined that the target emotion type is a problem emotion type. In a possible embodiment, if the ratio of the average value of the absolute value of the bias amount of the voice feature between the training data combinations to the average value of the feature amount of the voice feature of each training data combination is greater than 0.15 or more, it is determined that the target emotion type is a problem emotion type.

[0070] Specifically, as shown in Figure 4 the method for determining the screening replication object in the target replication object is:

[0071] determining the data amount of the training data of the target replication object in the problem emotion type based on the training data of the target replication object in the problem emotion type;

[0072] determining the data amount matching value of the training data of different problem emotion types according to the data amount of the training data in the problem emotion type;

[0073] determining whether the target replication object is a screening replication object according to the minimum value of the data amount matching value of the training data of different problem emotion types.

[0074] Further, when the minimum value of the data amount matching value of the training data of the target replication object in the problem emotion type is less than a preset matching value threshold, in a possible embodiment, less than 0.3, it is determined that the target replication object does not belong to the screening replication object. The preset data amount is determined according to the frequency of reply processing of the customer service personnel, wherein the higher the frequency of reply processing of the customer service personnel, the greater the preset data amount, and in a possible embodiment, the value thereof is 3000.

[0075] In another possible embodiment, the method for determining the screening replication object in the target replication object is:

[0076] determining the data amount of the training data of the target replication object in the problem emotion type based on the training data of the target replication object in the problem emotion type;

[0077] determining the average value of the data amount of the training data of different problem emotion types according to the data amount of the training data in the problem emotion type, and taking it as a data amount average value;

[0078] determining whether the target replication object is a screening replication object based on the data amount average value.

[0079] Further, when the data amount average value is less than a preset data amount threshold, it is determined that the target replication object does not belong to the screening replication object.

[0080] Optionally, the method for determining the copycat object in the copycat objects comprises:

[0081] Based on the voice data of the training data combination, the voice features of each voice data in the training data combination are determined, and based on the deviation amount of the voice features of each voice data, the voice data with a voice feature deviation amount less than a preset voice feature deviation amount is divided into the same voice data combination, and the matching data combination of the training data combination is determined according to the voice data combination with the largest data amount in the training data combination;

[0082] Based on the similarity of the emotion recognition results of the voice data of the copycat object and other copycat objects under the training data combination, the deviation amount of the voice features of the matching data combination of the training data combination is determined, and the voice feature deviation value between the copycat object and other copycat objects is determined according to the average value of the voice feature deviation amount between the copycat object and other copycat objects;

[0083] Based on the voice feature deviation value between the copycat object and other copycat objects, it is determined whether the copycat object is a copycat object.

[0084] Table 2 is the distribution of the training data amount of a certain object under different emotion types

[0085]

[0086] Further, the copycat object is the copycat object with the minimum average value of the voice feature deviation value between the copycat object and other copycat objects.

[0087] Further, it is determined that the training processing result of the copycat object is abnormal, which specifically comprises:

[0088] When the voice feature under any target emotion type does not match the image feature of the face, it is determined that the training processing result of the copycat object is abnormal.

[0089] It should be noted that the correspondence between the voice feature and the image feature of the face is determined according to the preset correspondence between the voice feature and the image feature of the copycat object.

[0090] Specifically, the method for determining the switching target of the copycat object comprises:

[0091] Based on the training data processing result of the copycat object, the target emotion type of the copycat object with an abnormal training processing result is determined, and the target emotion type is taken as an abnormal emotion type;

[0092] determine a deviation amount of the voice feature of the matching data combination of the training data combination according to the similarity of the emotion recognition result of the voice data of the other screening cloned object and the cloned object under the training data combination, and determine the average value of the voice feature deviation amount according to the average value of the voice feature deviation amount of the other screening cloned object and the cloned object under the training data combination;

[0093] According to the data amount of the training data of different abnormal emotion types and the average value of the voice feature deviation amount, determine whether the other screening cloned object is the switching target of the cloned object.

[0094] Further, according to the data amount of the training data of different abnormal emotion types and the average value of the voice feature deviation amount, determine whether the other screening cloned object is the switching target of the cloned object, specifically including:

[0095] When the average value of the voice feature deviation amount of the other screening cloned object is greater than the preset voice deviation amount threshold, it is determined that the other screening cloned object does not belong to the switching target of the cloned object.

[0096] When the average value of the voice feature deviation amount of the other screening cloned object is not greater than the preset voice deviation amount threshold, it is determined whether the other screening cloned object is the switching target of the cloned object according to the data amount of the training data of different abnormal emotion types.

[0097] Table 3 is a data comparison table of different screening cloned objects and cloned objects

[0098]

[0099] In this case, the screening cloned object is the object D with the largest data amount whose average value of the voice feature deviation amount is less than 0.4.

[0100] It can be understood that the switching target of the cloned object is the other screening cloned object with the largest average value of the data amount of the training data of different abnormal emotion types.

[0101] Specifically, when switching to the switching target of the cloned object, the augmented processing of the voice data of the switching target of the cloned object under different target emotion types is performed based on the voice data of the cloned object, and the training processing of the digital human is performed based on the augmented voice data.

[0102] Embodiment 2

[0103] In a second aspect, as Figure 5As shown, the present application provides a computer system, comprising a memory and a processor connected in communication, and a computer program stored on the memory and capable of running on the processor, wherein the processor executes the computer program to perform the AI model-based digital human emotion processing method described above.

[0104] In one specific embodiment thereof:

[0105] The target emotion type is an emotion type used in various application scenarios.

[0106] When, in the target emotion type, there is a training data combination whose deviation amount of the speech features of the speech data combined with other training data is greater than 0.2, it is determined that the target emotion type is a problem emotion type.

[0107] The screening duplication object is a target duplication object that does not exist in the problem emotion type or the number of problem emotion types is less than 3.

[0108] The screening duplication object with the minimum average of the deviation amount of the speech features of the speech data under the training data combination with other screening duplication objects is taken as the duplication object.

[0109] The other screening duplication object with the minimum deviation amount of the speech features of the speech data of the duplication object is taken as the switching target of the duplication object.

[0110] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each embodiment focuses on the difference from other embodiments. In particular, for the device, apparatus, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiment.

[0111] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous.

[0112] The above description is merely one or more embodiments of this specification and is not intended to limit this specification. Various modifications and variations can be made to the one or more embodiments of this specification by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of one or more embodiments of this specification should be included within the scope of the claims of this specification.

Claims

1. A digital human emotion processing method based on an AI model, characterized in that, Specifically, it includes: Based on the application scenarios of digital humans, the usage data of the emotional types of the digital humans in various application scenarios is determined, and the target emotional type is determined based on the usage data. Based on the similarity of the facial images corresponding to the training data, the training data of the target emotion type is divided into training data combinations. Based on the voice data between the training data combinations, the problem emotion type of the target emotion type is determined. Based on the training data of the target replica object in the problem emotion type, the selected replica objects among the target replica objects are determined. By using the similarity between the selected replica objects and other selected replica objects in the emotion recognition results of the speech data under the training data combination, the replica objects among the selected replica objects are determined. When the training result of the replicated object is abnormal, the target for switching the replicated object is determined based on the training result of the replicated object. The similarity of the facial images corresponding to the training data is determined based on the deviation of the image features of the facial images; The training data for the target emotion type is divided into training data sets, specifically including: Training data for facial images whose image feature deviations fall within a preset range are grouped into the same training data set. Facial images with a deviation rate of less than 0.05 in their image features are grouped into the same training data set, specifically determined by the Euclidean distance function; The method for determining the problem emotion type of the target emotion type is as follows: Based on the speech data of the training data combination, determine the speech features of each speech data in the training data combination; Based on the deviation of speech features of each speech data, speech data with a deviation of less than the preset speech feature deviation are grouped into the same speech data combination. The matching data combination of the training data combination is determined according to the speech data combination with the largest amount of speech data in the training data combination. Based on the deviation between the target emotion type and the speech features of the matching data combination of the training data combination, it is determined whether the target emotion type is a problem emotion type.

2. The digital human emotion processing method based on an AI model as described in claim 1, characterized in that, The application scenarios are categorized based on the business types of digital human processing within those application contexts.

3. The digital human emotion processing method based on an AI model as described in claim 1, characterized in that, The emotion types include anger, confusion, rage, contempt, fear, joy, pain, sadness, anticipation, anxiety, excitement, and surprise.

4. The digital human emotion processing method based on an AI model as described in claim 1, characterized in that, The usage data is determined based on the emotion recognition results of human customer service representatives in various application scenarios.

5. The digital human emotion processing method based on an AI model as described in claim 1, characterized in that, The method for determining the target emotion type in the emotion types is as follows: Based on the usage data of emotion types in various application scenarios, determine the frequency of use of emotion types by human customer service in various application scenarios; The matching value of the emotion type in each application scenario is determined by the ratio of the number of times the emotion type of human customer service is used to the number of times human customer service is processed in the application scenario. Based on the average value of the emotion type matching in various application scenarios, it is determined whether the emotion type is the target emotion type.

6. The digital human emotion processing method based on an AI model as described in claim 5, characterized in that, When the average value of the matching values ​​of the emotion type in various application scenarios is greater than the preset matching value threshold, the emotion type is determined to be the target emotion type.

7. The digital human emotion processing method based on an AI model as described in claim 1, characterized in that, The method for determining the switching target of the replicated object is as follows: Based on the training data processing results of the replicated object, it is determined that the training processing results of the replicated object contain an abnormal target emotion type, and this is identified as an abnormal emotion type. Based on the similarity of the emotion recognition results of the speech data of the replicated object and other selected replicated objects under the training data combination, the deviation of the speech features in the matching data combination of the training data combination is determined. Based on the average value of the deviation of the speech features between the other selected replicated objects and the replicated object under the training data combination, the average value of the speech feature deviation is determined. Based on the amount of training data and the average deviation of speech features for different abnormal emotion types, it is determined whether the other selected replica objects are the switching targets for replica objects.

8. The digital human emotion processing method based on an AI model as described in claim 7, characterized in that, Based on the amount of training data and the average deviation of speech features for different abnormal emotion types, it is determined whether the other selected replica objects are switching targets for replica objects, specifically including: When the average voice feature deviation of the other selected replica objects is greater than the preset voice deviation threshold, it is determined that the other selected replica objects are not the switching targets of the replica objects. When the average deviation of the speech features of the other selected replica objects is not greater than the preset speech deviation threshold, the other selected replica objects are determined as the switching target of the replica object based on the amount of training data in different abnormal emotion types.

9. The digital human emotion processing method based on an AI model as described in claim 8, characterized in that, The target for switching replicas is to select the replicas that have the highest average data volume among the training data of different abnormal emotion types.

10. A computer system, comprising: A memory and processor connected in communication, and a computer program stored in the memory and capable of running on the processor, characterized in that, when the processor runs the computer program, it executes a digital human emotion processing method based on an AI model as described in any one of claims 1-9.

Citation Information

Patent Citations

  • A digital human image reproduction method and system based on NeRF

    CN119399603B

  • AI digital human interaction method, device and system based on emotion recognition

    CN116560513A

  • AI digital human interaction system based on emotion recognition

    CN119473003A