Digital human emotion processing method and system based on AI model
By using an AI model-based approach to screen and switch digital human replicas, the accuracy of emotional processing in different application scenarios of digital humans was solved, ensuring the reliability and efficiency of training data.
Patent Information
- Application Number
- CN202511083738.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-04
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-08-04
AI Technical Summary
In existing technologies, the replication of digital humans is difficult to meet the different needs of emotional processing in different application scenarios, resulting in insufficient replication accuracy.
By using an AI model-based approach, the emotional types of digital humans in different application scenarios are determined. By combining training data and emotion recognition results from voice data, suitable replicas are selected, and when training results are abnormal, other replicas are switched to ensure the reliability and accuracy of the training data.
It achieves accuracy and reliability in digital human emotion processing in different application scenarios, avoids the waste of training data, and improves the efficiency of replication processing.
Smart Images

Figure CN120895060A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of digital people, and particularly relates to a digital person emotion processing method and system based on an AI model. BACKGROUND
[0002] In the prior art, the construction of a digital person can greatly reduce the pressure of manual customer service, and how to realize the replication processing of the digital person becomes a technical problem to be solved.
[0003] In order to realize the replication processing of the digital person, in the invention patent application CN202510013307.3 "Digital person image replication method and system based on NeRF", the expanded digital person image data is used to train a data reconstruction model, the features of the digital person image replication result output by the autoencoder are extracted, and the feature vectors after feature extraction are input into a classifier model for digital person image data screening, thereby improving the accuracy of the replication processing.
[0004] During the replication processing of the digital person, the differences in application occasions will lead to differences in the degree of emotional processing needs, and therefore how to determine the replication object of the digital person according to the needs of the application occasion to ensure the replication processing accuracy of the emotional processing needs with a higher degree becomes a technical problem to be solved.
[0005] In view of the above technical problems, the application provides a digital person emotion processing method and system based on an AI model. SUMMARY
[0006] To achieve the purpose of the application, the application adopts the following technical solutions: Specifically, the application provides a digital person emotion processing method based on an AI model, which specifically includes: S1, based on the application occasion of a digital person, determining the usage data of the emotion types of the digital person in various application scenarios, and determining a target emotion type in the emotion types based on the usage data; S2, dividing the training data of the target emotion type into training data combinations according to the similarity of the face images corresponding to the training data, determining a problem emotion type of the target emotion type based on the voice data between the training data combinations, and determining a screening replication object in the target replication object according to the training data of the problem emotion type; S3, determining a replication object in the screening replication object by using the similarity of the emotion recognition results of the voice data under the training data combinations between the screening replication object and other screening replication objects; S4 When the training processing result of the replica object is abnormal, determining a switching target of the replica object according to the training processing result of the replica object.
[0007] The application has the advantages of: The replica object is determined by screening the emotion recognition results of the voice data of the replica object and other screened replica objects under the training data combination, thereby ensuring that the similarity of the voice data of the replica object under different target emotion types meets the requirements and the data quantity meets the requirements, and further considering the similarity with other screened replica objects, so that when the training processing result is abnormal, the other screened replica object can be switched to in time and effectively, and the reliability of the digital human training processing is ensured.
[0008] According to the training processing result of the replica object and the similarity with other screened replica objects, the switching target of the replica object is determined, which not only ensures the reliability of the data quantity of the training data of the switching target under the target emotion type of the replica object when the training processing result of the replica object is abnormal, but also ensures the similarity between the switching target and the training data of the replica object, thereby avoiding the waste of the original training data of the replica object and ensuring the reliability of the training processing.
[0009] The further technical solution is that the application scenarios are divided according to the business types of the digital human processing under the application occasions.
[0010] The further technical solution is that the emotion types include anger, confusion, anger, contempt, fear, joy, pain, sadness, expectation, anxiety, excitement and surprise.
[0011] The further technical solution is that the use data is determined according to the emotion recognition results of the artificial customer service in various application scenarios.
[0012] The further technical solution is that the method for determining the target emotion type in the emotion types is: The use times of the emotion types of the artificial customer service in various application scenarios are determined by the use data of the emotion types in various application scenarios. The matching values of the emotion types in various application scenarios are determined by the ratio of the use times of the emotion types of the artificial customer service in various application scenarios to the processing times of the artificial customer service in the application scenarios. Whether the emotion type is a target emotion type is determined according to the average value of the matching values of the emotion types in various application scenarios.
[0013] Further, when the average of the matching values of the emotion types in various application scenarios is greater than a preset matching value threshold, the emotion type is determined as a target emotion type.
[0014] Further, the switching target of the cloned object is another filtered cloned object with the largest average of data amounts of training data of different abnormal emotion types.
[0015] Further, when switching to the switching target of the cloned object, the switching target of the cloned object is subjected to augmentation processing of voice data of different target emotion types based on voice data of the cloned object, and the digital person is subjected to training processing based on the augmented voice data.
[0016] In a second aspect, the present application provides a computer system, comprising a memory and a processor connected in communication, and a computer program stored on the memory and capable of running on the processor, wherein the processor executes the computer program to perform the AI model-based digital person emotion processing method described above.
[0017] Other features and advantages will be set forth in the accompanying description of the application, taken in connection with the accompanying drawings and a preferred embodiment. Such features and advantages of the application can also be obtained by carrying out the application in accordance with the teachings of the application as set forth in the claims.
[0018] So that the manner in which the above recited features and advantages of the present application can be understood in detail, a more particular description of the application, briefly summarized above, can be had by reference to the drawings, which are illustrated in the accompanying drawings, which form a part of the detailed description. BRIEF DESCRIPTION OF DRAWINGS
[0019] The above and other features and advantages of the present application will become more apparent by describing in detail exemplary embodiments thereof with reference to the attached drawings in which:
[0020] Figure 1 is a flowchart of an AI model-based digital person emotion processing method; Figure 2 is a flowchart of a method for determining a target emotion type in emotion types; Figure 3 is a flowchart of a method for determining a problem emotion type of a target emotion type; Figure 4 is a flowchart of a method for determining a filtered cloned object in a target cloned object; Figure 5 is a framework diagram of a computer system. DETAILED DESCRIPTION
[0021] In order for those skilled in the art to better understand the technical solutions in the specification, the technical solutions in the specification will be clearly and completely described below in combination with the drawings in the specification. Obviously, the described embodiments are only part of the embodiments of the specification, not all. Based on the embodiments of the specification, all other embodiments obtained by those skilled in the art without creative labor should belong to the protection scope of the specification.
[0022] In the present application, according to the variation of the voice data of the selected replication object in the target emotion type, the replication object with higher emotion stability of the voice data is selected, and the final replication object is selected based on the emotion stability of the voice data, which improves the accuracy of the replication processing.
[0023] Embodiment 1 As Figure 1 shown, the present application provides a digital human emotion processing method based on an AI model, which specifically includes: S1 determines the use data of the emotion type of the digital human in various application scenarios based on the application scenarios of the digital human, and determines the target emotion type in the emotion type based on the use data; S2 divides the training data of the target emotion type into training data combinations according to the similarity of the face images corresponding to the training data, determines the problem emotion type of the target emotion type based on the voice data between the training data combinations, and determines the screening replication object in the target replication object according to the training data of the problem emotion type. S3 determines the replication object in the screening replication object by using the similarity of the emotion recognition results of the voice data under the training data combination between the screening replication object and other screening replication objects; S4 determines the switching target of the replication object according to the training processing result of the replication object when the training processing result of the replication object is abnormal.
[0024] Further, the application scenarios are divided according to the business types processed by the digital human in the application scenarios, as shown in Table 1, the target emotion types under different business types in the bank system.
[0025] Table 1 Target emotion types under different business types
[0026] Specifically, the emotion type includes anger, confusion, anger, contempt, fear, joy, pain, sadness, expectation, anxiety, excitement and surprise.
[0027] It should be noted that the use data is determined according to the emotion recognition result of the artificial customer service in various application scenarios.
[0028] Further, the voice features include acoustic features, voice tone features, voice rhythm features, and voice spectrum features.
[0029] Specifically, as shown in the method for determining a target emotion type in the emotion types is: Figure 2 determining, with the use data of the emotion types in various application scenarios, the use times of the emotion types of the artificial customer service in various application scenarios; determining, with the use times of the emotion types of the artificial customer service in various application scenarios and the handling times of the artificial customer service in the application scenarios, whether the emotion type is a target emotion type.
[0030] Further, when there is an application scenario in which the proportion of the use times of the emotion type to the handling times of the artificial customer service in the application scenario meets the requirement, in one possible embodiment, greater than 0.4, it is determined that the emotion type is a target emotion type.
[0031] In another possible embodiment, the method for determining a target emotion type in the emotion types is: determining, with the use data of the emotion types in various application scenarios, the use times of the emotion types of the artificial customer service in various application scenarios; determining, with the use times of the emotion types of the artificial customer service in various application scenarios, the total use times of the emotion types in the application scenarios; determining, according to the total use times of the emotion types in the application scenarios, whether the emotion type is a target emotion type.
[0032] Further, when the total use times of the emotion types in the application scenarios are greater than a preset use time threshold, it is determined that the emotion type is a target emotion type.
[0033] It should be noted that when there is no use data, the target emotion type is determined according to the preset emotion type corresponding to the application scenario.
[0034] Specifically, the similarity of the face image corresponding to the training data is determined according to the deviation amount of the image features of the face image.
[0035] Specifically, the training data of the target emotion type is divided into a training data combination, which specifically includes: The face images corresponding to the deviation amount of the image features within the preset image deviation amount interval are divided into the same training data combination. Specifically, the face images with a deviation rate of image features less than 0.05 are divided into the same training data combination. Specifically, the determination is made by using the Euclidean distance function.
[0036] It should be noted that, as shown in Figure 3 The method for determining the problem emotion type of the target emotion type is as follows: The speech features of each speech data in the training data combination are determined based on the speech data of the training data combination. Based on the deviation amount of the speech features of each speech data, the speech data with a speech feature deviation amount less than a preset speech feature deviation amount are divided into the same speech data combination. The matching data combination of the training data combination is determined according to the speech data combination with the largest amount of speech data in the training data combination. The deviation of the target emotion type between the speech features of the matching data combination of the training data combination is determined, and it is determined whether the target emotion type is a problem emotion type.
[0037] Further, the speech features of the matching data combination are determined according to the average value of the speech features of each speech data in the matching data combination. It can be understood that the speech features are MFCC features of the speech data, and the determination is made after the average value of the MFCC features of the speech data is obtained.
[0038] It can be understood that the deviation of the target emotion type between the speech features of the matching data combination of the training data combination is determined, and it is determined whether the target emotion type is a problem emotion type. Specifically, the determination includes: The deviation amount of the speech features between the training data combinations is determined according to the deviation of the target emotion type between the speech features of the matching data combination of the training data combination. The average value of the deviation amount of the speech features between the training data combinations is determined, and it is determined whether the target emotion type is a problem emotion type.
[0039] Further, when the average value of the absolute value of the deviation amount of the speech features between the training data combinations is greater than a preset deviation amount threshold, it is determined that the target emotion type is a problem emotion type. In a possible embodiment, if the ratio of the average value of the absolute value of the deviation amount of the speech features between the training data combinations to the average value of the feature amount of the speech features of each training data combination is greater than 0.15 or more, it is determined that the target emotion type is a problem emotion type.
[0040] Specifically, as shown in Figure 4As shown, the method for determining the screening replica object in the target replica object is: Based on the training data of the target replica object in the question emotion type, determine the data amount of the training data of the target replica object in the question emotion type; According to the ratio of the data amount of the training data in the question emotion type and the preset data amount, determine the data amount matching value of the training data of different question emotion types; According to the minimum value of the data amount matching value of the training data of different question emotion types, determine whether the target replica object is a screening replica object.
[0041] Further, when the minimum value of the data amount matching value of the training data of the target replica object in the question emotion type is less than the preset matching value threshold, in one possible embodiment, less than 0.3, then it is determined that the target replica object does not belong to the screening replica object, and the preset data amount is determined according to the frequency of reply processing of the customer service personnel, wherein the higher the frequency of reply processing of the customer service personnel, the larger the preset data amount, in one possible embodiment, the value is 3000.
[0042] In another possible embodiment, the method for determining the screening replica object in the target replica object is: Based on the training data of the target replica object in the question emotion type, determine the data amount of the training data of the target replica object in the question emotion type; According to the data amount of the training data in the question emotion type, determine the average value of the data amount of the training data of different question emotion types, and take it as the data amount average value; Based on the data amount average value, determine whether the target replica object is a screening replica object.
[0043] Further, when the data amount average value is less than the preset data amount threshold, it is determined that the target replica object does not belong to the screening replica object.
[0044] Optionally, the method for determining the replica object in the screening replica object is: Based on the voice data of the training data combination, determine the voice feature of each voice data in the training data combination, based on the deviation amount of the voice feature of each voice data, divide the voice data with the voice feature deviation amount less than the preset voice feature deviation amount to the same voice data combination, and according to the voice data combination with the largest data amount of the voice data in the training data combination, determine the matching data combination of the training data combination; determine the deviation of the voice feature of the matching data combination of the training data combination according to the similarity of the emotion recognition result of the voice data of the screening duplication object and other screening duplication objects under the training data combination, and determine the voice feature deviation value between the screening duplication object and other screening duplication objects according to the average value of the voice feature deviation value between the screening duplication object and other screening duplication objects under the training data combination; determine whether the screening duplication object is a duplication object based on the voice feature deviation value between the screening duplication object and other screening duplication objects.
[0045] Table 2 is the distribution of the amount of training data of a certain object under different emotion types
[0046] Further, the duplication object is the screening duplication object with the minimum average value of the voice feature deviation value between the screening duplication object and other screening duplication objects.
[0047] Further, the determination of the abnormality of the training processing result of the duplication object specifically includes: When the voice feature under any target emotion type does not match the image feature of the face, it is determined that the training processing result of the duplication object is abnormal.
[0048] It should be noted that the correspondence between the voice feature and the image feature of the face is determined according to the preset correspondence between the voice feature and the image feature of the duplication object.
[0049] Specifically, the method for determining the switching target of the duplication object is: determine the target emotion type of the training processing result of the duplication object, and take it as an abnormal emotion type; determine the deviation of the voice feature of the matching data combination of the training data combination according to the similarity of the emotion recognition result of the voice data of the screening duplication object and other screening duplication objects under the training data combination, and determine the voice feature deviation value between the screening duplication object and other screening duplication objects according to the average value of the voice feature deviation value between the screening duplication object and other screening duplication objects under the training data combination; determine whether the other screening duplication object is the switching target of the duplication object according to the data amount of the training data under different abnormal emotion types and the average value of the voice feature deviation value.
[0050] Further, according to the data amount of the training data under different abnormal emotion types and the average value of the voice feature deviation value, it is determined whether the other screening duplication object is the switching target of the duplication object, specifically including: When the average voice feature deviation of the other selected replica objects is greater than the preset voice deviation threshold, it is determined that the other selected replica objects are not the switching targets of the replica objects. When the average deviation of the speech features of the other selected replica objects is not greater than the preset speech deviation threshold, the other selected replica objects are determined as the switching target of the replica object based on the amount of training data in different abnormal emotion types.
[0051] Table 3 is a comparison table of data for different selection and replication targets.
[0052] In this case, the object to be replicated will be the object D with the largest amount of data whose average speech feature deviation is less than 0.4.
[0053] It is understandable that the target for switching replicas is to select the replicas that have the highest average amount of training data for different abnormal emotion types.
[0054] Specifically, when switching to the target of the replica object, the voice data of the replica object is used as the basis to amplify the voice data of the target of the replica object in different target emotion types, and the digital human is trained based on the amplified voice data.
[0055] Example 2 Secondly, such as Figure 5 As shown, the present invention provides a computer system comprising: a memory and a processor connected in communication, and a computer program stored in the memory and capable of running on the processor, wherein the processor executes the aforementioned digital human emotion processing method based on an AI model when running the computer program.
[0056] In one specific embodiment: The target emotion type is an emotion type that is used in various application scenarios.
[0057] If, within the target emotion type, there exists a training data combination whose speech features deviate from other training data combinations by a factor greater than 0.2, then the target emotion type is determined to be a problem emotion type.
[0058] The target replicas are those that do not have any problematic emotion types or have fewer than 3 problematic emotion types.
[0059] The candidate that minimizes the average deviation of speech features from other selected candidates in the speech data under the training data combination is selected as the candidate to be copied.
[0060] The other screening duplication object that minimizes the amount of deviation of the voice feature of the voice data of the duplication object is selected as a switching target of the duplication object.
[0061] Each of the embodiments in the specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other. Each of the embodiments focuses on the difference from other embodiments. In particular, for the device, apparatus, and non-transitory computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the part of the method embodiments.
[0062] The above describes specific embodiments of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different than the order in which they are recited, and still achieve desirable results. In addition, the processes depicted in the figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In some implementations, multitasking and parallel processing can be advantageous or necessary.
[0063] The above only describes one or more embodiments of the specification, and is not intended to limit the specification. One or more embodiments of the specification can have various modifications and changes for those skilled in the art. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of one or more embodiments of the specification shall be included in the scope of the claims of the specification.
Claims
1. A digital human emotion processing method based on an AI model, characterized in that, Specifically, it includes: Based on the application scenarios of digital humans, the usage data of the emotional types of the digital humans in various application scenarios is determined, and the target emotional type is determined based on the usage data. Based on the similarity of the facial images corresponding to the training data, the training data of the target emotion type is divided into training data combinations. Based on the voice data between the training data combinations, the problem emotion type of the target emotion type is determined. Based on the training data of the target replica object in the problem emotion type, the selected replica objects among the target replica objects are determined. By using the similarity between the selected replica objects and other selected replica objects in the emotion recognition results of the speech data under the training data combination, the replica objects among the selected replica objects are determined. When the training result of the replicated object is abnormal, the switching target of the replicated object is determined based on the training result of the replicated object.
2. The digital human emotion processing method based on an AI model as described in claim 1, characterized in that, The application scenarios are categorized based on the business types of digital human processing within those application contexts.
3. The digital human emotion processing method based on an AI model as described in claim 1, characterized in that, The emotion types include anger, confusion, rage, contempt, fear, joy, pain, sadness, anticipation, anxiety, excitement, and surprise.
4. The digital human emotion processing method based on an AI model as described in claim 1, characterized in that, The usage data is determined based on the emotion recognition results of human customer service representatives in various application scenarios.
5. The digital human emotion processing method based on an AI model as described in claim 1, characterized in that, The method for determining the target emotion type in the emotion types is as follows: Based on the usage data of emotion types in various application scenarios, determine the frequency of use of emotion types by human customer service in various application scenarios; The matching value of the emotion type in each application scenario is determined by the ratio of the number of times the emotion type of human customer service is used to the number of times human customer service is processed in the application scenario. Based on the average value of the emotion type matching in various application scenarios, it is determined whether the emotion type is the target emotion type.
6. The digital human emotion processing method based on an AI model as described in claim 5, characterized in that, When the average value of the matching values of the emotion type in various application scenarios is greater than the preset matching value threshold, the emotion type is determined to be the target emotion type.
7. The digital human emotion processing method based on an AI model as described in claim 1, characterized in that, The method for determining the switching target of the replicated object is as follows: Based on the training data processing results of the replicated object, it is determined that the training processing results of the replicated object contain an abnormal target emotion type, and this is identified as an abnormal emotion type. Based on the similarity of the emotion recognition results of the speech data of the replicated object and other selected replicated objects under the training data combination, the deviation of the speech features in the matching data combination of the training data combination is determined. Based on the average value of the deviation of the speech features between the other selected replicated objects and the replicated object under the training data combination, the average value of the speech feature deviation is determined. Based on the amount of training data and the average deviation of speech features in different abnormal emotion types, it is determined whether the other selected replica objects are the switching targets for replica objects.
8. The digital human emotion processing method based on an AI model as described in claim 7, characterized in that, Based on the amount of training data and the average deviation of speech features for different abnormal emotion types, it is determined whether the other selected replica objects are switching targets for replica objects, specifically including: When the average voice feature deviation of the other selected replica objects is greater than the preset voice deviation threshold, it is determined that the other selected replica objects are not the switching targets of the replica objects. When the average deviation of the speech features of the other selected replica objects is not greater than the preset speech deviation threshold, the other selected replica objects are determined as the switching target of the replica object based on the amount of training data in different abnormal emotion types.
9. The digital human emotion processing method based on an AI model as described in claim 8, characterized in that, The target for switching replicas is to select the replicas that have the highest average data volume among the training data of different abnormal emotion types.
10. A computer system, comprising: A memory and processor connected in communication, and a computer program stored in the memory and capable of running on the processor, characterized in that, when the processor runs the computer program, it executes a digital human emotion processing method based on an AI model as described in any one of claims 1-9.
Citation Information
Patent Citations
A digital human image reproduction method and system based on NeRF
CN119399603B
AI digital human interaction method, device and system based on emotion recognition
CN116560513A
AI digital human interaction method and system based on emotion recognition
CN118519538A
Digital human image copying method and system based on NeRF
CN119399603A
AI digital human interaction system based on emotion recognition
CN119473003A