LLM-based three-dimensional digital human expression generation method and system
By creating an expression library and a large language model, combined with expression animation and keyframe technology, the accuracy and naturalness issues of three-dimensional digital human expression generation in complex contexts were solved, and high-precision and vivid expression generation was achieved.
Patent Information
- Application Number
- CN202510686395.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-09
AI Technical Summary
Existing methods for generating expressions for three-dimensional digital humans lack expressiveness when faced with complex and changing contexts, are unable to accurately generate corresponding expressions, and lack specific analysis of each sentence spoken by the three-dimensional digital human.
Create an expression library and a large language model, train the large language model to identify emotional tags, match expression animations and play them, adopt the LLM architecture of hierarchical feature fusion and composite loss function, combine the expression library with keyframe animation technology, and realize emotion analysis and expression generation.
It improves the accuracy, naturalness and vividness of the expression generation of three-dimensional digital humans, can accurately match the emotional needs in complex contexts, avoids abrupt expression switching, and improves the smoothness and naturalness of animation.
Smart Images

Figure CN120612409A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for generating three-dimensional digital human expressions based on LLM. Background Art
[0002] Against the backdrop of the profound convergence of computer graphics and artificial intelligence, 3D digital humans, a cutting-edge form of digital avatar technology, have gradually moved from the laboratory to industrial applications. Their core technology system integrates multiple technical fields, including 3D modeling, motion capture, speech synthesis, and natural language processing. Current mainstream 3D digital human systems construct human topology through high-precision 3D scanning and modeling, combine skeletal rigging and skinning weighting techniques to simulate basic movements, and employ deep learning algorithms to optimize facial features, resulting in a digital human's appearance that can achieve millimeter-level accuracy.
[0003] To make 3D digital humans more vivid and lifelike, it's necessary to generate expressions for them. Traditionally, this has relied on predefined expression libraries or keypoint-based animation techniques. While these techniques can realistically reproduce specific emotions or expressions, they often lack expressiveness in complex and changing contexts. Furthermore, they lack the ability to analyze each sentence spoken by the 3D digital human to accurately generate expressions that convey the correct emotion.
[0004] Therefore, how to provide a method and system for generating 3D digital human expressions based on LLM to improve the accuracy, naturalness and vividness of 3D digital human expression generation has become a technical problem that needs to be solved urgently. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a method and system for generating three-dimensional digital human expressions based on LLM, so as to improve the accuracy, naturalness and vividness of three-dimensional digital human expression generation.
[0006] In a first aspect, the present invention provides a method for generating expressions of a three-dimensional digital human based on LLM, comprising the following steps:
[0007] Step S1, creating an expression library, creating a large number of expression animations and storing them in the expression library;
[0008] Step S2: creating a large language model for sentiment analysis and setting a loss function for the large language model;
[0009] Step S3: obtaining a large amount of historical interaction texts of the three-dimensional digital human, pre-processing and annotating each of the historical interaction texts to construct a data set, training a large language model using the data set and a loss function, and deploying the trained large language model;
[0010] Step S4: obtaining the real-time interactive text of the three-dimensional digital human, pre-processing the real-time interactive text, and inputting it into the deployed large language model to obtain recognized emotion labels;
[0011] Step S5: Match the corresponding facial expression animation from the facial expression library using the emotion tag, and load the matched facial expression animation onto the face of the three-dimensional digital human for playback.
[0012] Furthermore, in step S1, the expression themes of the expression animation include at least laughter, happiness, sweetness, excitement, surprise, pleasure, enthusiasm, complacency, satisfaction, showing off, requesting, encouraging, cheering, comforting, dependence, trust, relaxation, sympathy, pity, distress, guilt, touching, respecting, cute, pitiful, asking, praying, pleading, acting coquettishly, enjoying, longing, appreciating, approving, admiring, worshipping, loving, envying, persisting, strong, firm, resolute, witty, clever, wise, naughty, playful, calm, composed, surprised, startled, confused, hesitant, Confusion, ignorance, helplessness, silence, seriousness, solemnity, caution, thinking, contemplation, need, desire, shyness, shame, anger, joy, dissatisfaction, slight anger, sadness, grief, sorrow, anger, rage, shock and anger, disbelief, regret, disbelief, questioning, denial, pain, suffering, torment, helplessness, bitterness, tension, anxiety, worry, concern, depression, melancholy, fear, fear, fear, sorrow, rejection, grievance, reluctance, indifference, disdain, alienation, indifference, impatience, restlessness, irritability, worry, distress, smile and greet, comfort and teasing;
[0013] The expression animation is provided with key frames including at least expression start, expression display and expression end.
[0014] Furthermore, the step S2 is specifically as follows:
[0015] Creating a large language model for sentiment analysis based on the interactive text encoding layer, the dialogue structure perception layer, the feature fusion layer, and the sentiment prediction layer, and setting a loss function for the large language model;
[0016] The interactive text encoding layer is used to extract text sequence features, dialogue structure features, and context-dependent features from the input interactive text, splice the text sequence features, dialogue structure features, and context-dependent features to obtain spliced features, and input the spliced features into the dialogue structure perception layer;
[0017] The dialogue structure perception layer is used to model the dialogue roles, temporal relationships, and context dependencies of the spliced features to obtain dialogue perception features, and input the dialogue perception features into the feature fusion layer;
[0018] The feature fusion layer is used to perform multi-dimensional fusion of the dialogue perception features and the pre-trained language model knowledge to obtain the emotion discrimination features, and input the emotion discrimination features into the emotion prediction layer;
[0019] The emotion prediction layer is used to map the emotion discrimination features into a probability distribution of emotion labels, and output an emotion prediction result carrying the emotion label based on the probability distribution;
[0020] The formula of the loss function is:
[0021] L_total=λ1*L_ce+λ2*L_struct+λ3*L_align+λ4*L_reg;
[0022] Among them, L_total represents the loss value of the loss function; L_ce represents the cross entropy loss; L_struct represents the structural consistency loss; L_align represents the knowledge alignment loss; L_reg represents the regularization term; λ1, λ2, λ3, and λ4 all represent weight coefficients.
[0023] Furthermore, the step S3 is specifically as follows:
[0024] Acquire a large amount of historical interaction texts of the three-dimensional digital human, perform preprocessing on each of the historical interaction texts, including at least text cleaning, text standardization, and word segmentation, annotate each of the preprocessed historical interaction texts with emotional tags, and then construct a data set;
[0025] Dividing the data set into a training set, a validation set, and a test set based on a preset ratio, and training the large language model using the training set until the loss value of the loss function is less than a preset loss threshold;
[0026] Verifying the trained large language model using the validation set, testing the verified large language model using the test set, and deploying the tested large language model using containerization technology;
[0027] The emotional tags include at least 0-admiration, 1-amusement, 2-anger, 3-annyance, 4-approval, 5-caring, 6-confusion, 7-curiosity, 8-desire, 9-disappointment, 10-disapproval, 11-disgust, 12-embarrassment, 13-excitingment, 14-fear, 15-gratitude, 16-grief, 17-joy, 18-love, 19-nervousness, 20-optimism, 21-pride, 22-realization, 23-rel ief, 24-remorse, 25-sadness, 26-surprise, and 27-neutral; each of the emotional tags corresponds to at least one facial expression animation.
[0028] Furthermore, the step S5 further includes:
[0029] The real-time interactive text is converted into interactive voice through EdgeTTS, and the interactive voice is played synchronously when the three-dimensional digital human plays the expression animation.
[0030] In a second aspect, the present invention provides a 3D digital human expression generation system based on LLM, comprising the following modules:
[0031] An expression animation creation module is used to create an expression library, create a large number of expression animations and store them in the expression library;
[0032] A large language model creation module, used to create a large language model for sentiment analysis and set a loss function for the large language model;
[0033] A large language model training module is used to obtain a large amount of historical interaction texts of 3D digital humans, pre-process and annotate each of the historical interaction texts to construct a data set, train the large language model using the data set and loss function, and deploy the trained large language model;
[0034] The emotion tag recognition module is used to obtain the real-time interactive text of the 3D digital human, pre-process the real-time interactive text, and then input it into the deployed large language model to obtain the recognized emotion tag;
[0035] The expression animation playback module is used to match the corresponding expression animation from the expression library through the emotion tag, and load the matched expression animation into the face of the three-dimensional digital human for playback.
[0036] Furthermore, in the expression animation creation module, the expression themes of the expression animation include at least laughter, happiness, sweetness, excitement, surprise, pleasure, enthusiasm, complacency, satisfaction, showing off, requesting, encouraging, cheering, comforting, dependence, trust, relaxation, sympathy, pity, distress, guilt, touching, respecting, acting cute, pitiful, asking, praying, pleading, acting coquettish, enjoying, longing, appreciating, approving, admiring, worshipping, loving, envying, persisting, being strong, firm, resolute, witty, clever, wise, naughty, playful, calm, composed, surprised, startled, confused, hesitant, etc. Hesitation, confusion, ignorance, helplessness, silence, seriousness, solemnity, caution, thinking, contemplation, need, desire, shyness, shame, shameful anger, shameful joy, dissatisfaction, slight anger, sadness, grief, sorrow, anger, rage, shock and anger, disbelief, regret, disbelief, doubt, denial, pain, suffering, torment, helplessness, bitterness, tension, anxiety, worry, concern, melancholy, depression, fear, fear, terror, sorrow, rejection, grievance, reluctance, indifference, disdain, alienation, indifference, impatience, restlessness, irritability, worry, distress, annoyance, distress, smiling hello, comforting and teasing;
[0037] The expression animation is provided with key frames including at least expression start, expression display and expression end.
[0038] Furthermore, the large language model creation module is specifically used to:
[0039] Creating a large language model for sentiment analysis based on the interactive text encoding layer, the dialogue structure perception layer, the feature fusion layer, and the sentiment prediction layer, and setting a loss function for the large language model;
[0040] The interactive text encoding layer is used to extract text sequence features, dialogue structure features, and context-dependent features from the input interactive text, splice the text sequence features, dialogue structure features, and context-dependent features to obtain spliced features, and input the spliced features into the dialogue structure perception layer;
[0041] The dialogue structure perception layer is used to model the dialogue roles, temporal relationships, and context dependencies of the spliced features to obtain dialogue perception features, and input the dialogue perception features into the feature fusion layer;
[0042] The feature fusion layer is used to perform multi-dimensional fusion of the dialogue perception features and the pre-trained language model knowledge to obtain the emotion discrimination features, and input the emotion discrimination features into the emotion prediction layer;
[0043] The emotion prediction layer is used to map the emotion discrimination features into a probability distribution of emotion labels, and output an emotion prediction result carrying the emotion label based on the probability distribution;
[0044] The formula of the loss function is:
[0045] L_total=λ1*L_ce+λ2*L_struct+λ3*L_align+λ4*L_reg;
[0046] Among them, L_total represents the loss value of the loss function; L_ce represents the cross entropy loss; L_struct represents the structural consistency loss; L_align represents the knowledge alignment loss; L_reg represents the regularization term; λ1, λ2, λ3, and λ4 all represent weight coefficients.
[0047] Furthermore, the large language model training module is specifically used to:
[0048] Acquire a large amount of historical interaction texts of the three-dimensional digital human, perform preprocessing on each of the historical interaction texts, including at least text cleaning, text standardization, and word segmentation, annotate each of the preprocessed historical interaction texts with emotional tags, and then construct a data set;
[0049] Dividing the data set into a training set, a validation set, and a test set based on a preset ratio, and training the large language model using the training set until the loss value of the loss function is less than a preset loss threshold;
[0050] Verifying the trained large language model using the validation set, testing the verified large language model using the test set, and deploying the tested large language model using containerization technology;
[0051] The emotional labels include at least 0-admiration, 1-amusement, 2-anger, 3-annyance, 4-approval, 5-caring, 6-confusion, 7-curiosity, 8-desire, 9-disappointment, 10-disapproval, 11-disgust, 12-embarrassment, 13-excitingment, 14-fear, 15-gratitude, 16-grief, 17-joy, 18-love, 19-nervousness, 20-optimism, 21-pride, 22-realization, 23-relief, 24-remorse, 25-sadness, 26-surprise, and 27-neutral; each of the emotional labels corresponds to at least one facial expression animation.
[0052] Furthermore, the expression animation playback module is also used to:
[0053] The real-time interactive text is converted into interactive voice through EdgeTTS, and the interactive voice is played synchronously when the three-dimensional digital human plays the expression animation.
[0054] The advantages of the present invention are:
[0055] 1. Create a large number of facial expression animations and store them in an expression library, then create a large language model for sentiment analysis and set the loss function of the large language model; obtain a large amount of historical interaction texts of 3D digital humans, pre-process and annotate each historical interaction text to build a data set, train the large language model through the data set and loss function, and deploy the trained large language model; then obtain the real-time interaction text of the 3D digital human, pre-process the real-time interaction text and input it into the deployed large language model to obtain the recognized emotion label, match the corresponding facial expression animation from the expression library through the emotion label, and load the matched facial expression animation onto the face of the 3D digital human for playback; that is, the emotional label recognition of the real-time interaction text through the pre-trained large language model can effectively cope with complex and changing contexts, and can perform specific analysis on each sentence of the real-time interaction text to obtain the corresponding emotion label, and then match the expression animation from the expression library that stores rich facial expression animation based on the emotion label for playback, thereby greatly improving the accuracy, naturalness and vividness of the 3D digital human's expression generation.
[0056] 2. By defining facial expression animations that encompass a wide range of emotional themes (such as laughter, anger, dependence, and pity), a broad spectrum of emotions from positive to negative is covered, far exceeding the single classification of traditional facial expression libraries (such as only joy, anger, sadness, and happiness), and can accurately match emotional needs in complex contexts. By setting facial expression animations to include a start→display→end keyframe sequence, the physiological dynamics of human expression are simulated (such as the gradual rise and fall of the corners of the mouth when smiling), avoiding abrupt expression switches and improving the smoothness and naturalness of animations.
[0057] 3. By setting up a four-layer model architecture of the large language model (encoding layer → perception layer → fusion layer → prediction layer): the interactive text encoding layer extracts text sequence, dialogue structure, and context-dependent features; the dialogue structure perception layer models role relationships and temporal logic; the feature fusion layer integrates pre-trained knowledge to enhance generalization capabilities; this enables the large language model to capture the implicit semantics and long-term dependencies in the dialogue, thereby improving the accuracy of sentiment label prediction.
[0058] 4. By setting the loss function to integrate cross entropy loss (L_ce), structural consistency loss (L_struct), knowledge alignment loss (L_align) and regularization term (L_reg), and dynamically balancing different optimization objectives through weight coefficients (λ1-λ4), it effectively prevents overfitting and improves the model's adaptability to complex dialogue scenarios.
[0059] 5. By cleaning, standardizing, segmenting, and labeling historical interaction texts, we construct a high-quality dataset covering 28 categories of sentiment labels (such as admiration, grief, neutral, etc.) to ensure the diversity of model training data and consistency of labeling.
[0060] 6. By using containerization technology to deploy trained LLM, it supports rapid elastic scaling, adapts to high-concurrency real-time interaction scenarios (such as virtual customer service, game NPCs), and facilitates subsequent model updates and maintenance.
[0061] 7. By integrating the Large Language Model (LLM) with 3D animation technology, the intelligent and high-precision generation of 3D digital human expressions is achieved. Its core advantages lie in: the first LLM architecture with hierarchical feature fusion (interactive text encoding → dialogue structure perception → multi-dimensional knowledge fusion → emotion prediction), combined with a refined expression library covering 84 emotional themes and dynamic optimization of key frames, which solves the problems of rough emotion classification and abrupt expression switching in traditional methods; at the same time, the robustness of the model is improved through a composite loss function (cross entropy + structural consistency + knowledge alignment), and low-latency, multi-modal real-time interaction is achieved by using containerized deployment and voice-expression synchronization technology (EdgeTTS), which combines engineering implementation efficiency and cross-domain generalization capabilities (virtual customer service, education, games, etc.), significantly reducing development costs while improving user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0063] Figure 1 This is a flow chart of a method for generating expressions of a three-dimensional digital human based on LLM according to the present invention.
[0064] Figure 2 It is a structural diagram of a three-dimensional digital human expression generation system based on LLM of the present invention. DETAILED DESCRIPTION
[0065] The technical solution in the embodiments of the present application has the following overall idea: by using a pre-trained large language model to identify the emotional labels of real-time interactive text, it can effectively deal with complex and changing contexts, and can perform specific analysis on each sentence of the real-time interactive text to obtain the corresponding emotional label, and then match the emotional animation from the expression library that stores rich emotional animations based on the emotional label for playback, so as to improve the accuracy, naturalness and vividness of the expression generation of three-dimensional digital humans.
[0066] Please refer to Figures 1 to 2 As shown, a preferred embodiment of the present invention's method for generating three-dimensional digital human expressions based on LLM includes the following steps:
[0067] Step S1, creating an expression library, creating a large number of expression animations and storing them in the expression library;
[0068] Step S2: creating a large language model for sentiment analysis and setting a loss function for the large language model;
[0069] Step S3: obtaining a large amount of historical interaction texts of the three-dimensional digital human, pre-processing and annotating each of the historical interaction texts to construct a data set, training a large language model using the data set and a loss function, and deploying the trained large language model;
[0070] Step S4: obtaining the real-time interactive text of the three-dimensional digital human, pre-processing the real-time interactive text, and inputting it into the deployed large language model to obtain recognized emotion labels;
[0071] Step S5: Match the corresponding facial expression animation from the facial expression library using the emotion tag, and load the matched facial expression animation onto the face of the three-dimensional digital human for playback.
[0072] In step S1, the expression themes of the expression animation include at least laughter, happiness, sweetness, excitement, surprise, pleasure, enthusiasm, complacency, satisfaction, showing off, requesting, encouraging, cheering, comforting, dependence, trust, relaxation, sympathy, pity, distress, guilt, touching, respecting, cute, pitiful, asking, praying, begging, acting coquettishly, enjoying, longing, appreciating, approving, admiring, worshipping, loving, envying, persisting, strong, firm, resolute, witty, clever, wise, naughty, playful, calm, composed, surprised, startled, confused, hesitant, bewildered, Ignorance, helplessness, silence, seriousness, solemnity, caution, thinking, contemplation, need, desire, shyness, shame, shame-anger, shame-joy, dissatisfaction, slight anger, sadness, grief, sorrow, anger, rage, shock and anger, disbelief, regret, disbelief, questioning, denial, pain, suffering, torment, helplessness, bitterness, tension, anxiety, worry, concern, depression, melancholy, fear, fear, terror, sorrow, rejection, grievance, reluctance, indifference, disdain, alienation, indifference, impatience, restlessness, irritability, worry, distress, annoyance, distress, smiling hello, comforting and teasing;
[0073] The expression animation is provided with key frames including at least expression start, expression display and expression end.
[0074] By defining facial expression animations that include a large number of emotional themes (such as laughter, anger, dependence, pity, etc.), covering a wide range of emotions from positive to negative, far exceeding the single classification of traditional facial expression libraries (such as only joy, anger, sadness, and happiness), it can accurately match emotional needs in complex contexts; by setting facial expression animations to include a start→display→end key frame sequence, imitating the physiological dynamic process of human expression (such as the gradual rise and fall of the corners of the mouth when smiling), it avoids abrupt expression switches and improves the smoothness and naturalness of the animation.
[0075] The step S2 is specifically as follows:
[0076] Creating a large language model for sentiment analysis based on the interactive text encoding layer, the dialogue structure perception layer, the feature fusion layer, and the sentiment prediction layer, and setting a loss function for the large language model;
[0077] The interactive text encoding layer is used to extract text sequence features, dialogue structure features, and context-dependent features from the input interactive text, splice the text sequence features, dialogue structure features, and context-dependent features to obtain spliced features, and input the spliced features into the dialogue structure perception layer;
[0078] The dialogue structure perception layer is used to model the dialogue roles, temporal relationships, and context dependencies of the spliced features to obtain dialogue perception features, and input the dialogue perception features into the feature fusion layer;
[0079] The feature fusion layer is used to perform multi-dimensional fusion of the dialogue perception features and the pre-trained language model knowledge to obtain the emotion discrimination features, and input the emotion discrimination features into the emotion prediction layer;
[0080] The emotion prediction layer is used to map the emotion discrimination features into a probability distribution of emotion labels, and output an emotion prediction result carrying the emotion label based on the probability distribution;
[0081] The formula of the loss function is:
[0082] L_total=λ1*L_ce+λ2*L_struct+λ3*L_align+λ4*L_reg;
[0083] Among them, L_total represents the loss value of the loss function; L_ce represents the cross entropy loss; L_struct represents the structural consistency loss; L_align represents the knowledge alignment loss; L_reg represents the regularization term; λ1, λ2, λ3, and λ4 all represent weight coefficients.
[0084] By setting up a four-layer model architecture of the large language model (encoding layer → perception layer → fusion layer → prediction layer): the interactive text encoding layer extracts text sequence, dialogue structure, and context-dependent features; the dialogue structure perception layer models role relationships and temporal logic; the feature fusion layer integrates pre-trained knowledge to enhance generalization ability; this enables the large language model to capture the implicit semantics and long-term dependencies in the dialogue, thereby improving the accuracy of sentiment label prediction.
[0085] By setting the loss function to integrate cross entropy loss (L_ce), structural consistency loss (L_struct), knowledge alignment loss (L_align) and regularization term (L_reg), and dynamically balancing different optimization objectives through weight coefficients (λ1-λ4), overfitting is effectively prevented and the model's adaptability to complex dialogue scenarios is improved.
[0086] The step S3 is specifically as follows:
[0087] Acquire a large amount of historical interaction texts of the three-dimensional digital human, perform preprocessing on each of the historical interaction texts, including at least text cleaning, text standardization, and word segmentation, annotate each of the preprocessed historical interaction texts with emotional tags, and then construct a data set;
[0088] By cleaning, standardizing, segmenting, and labeling historical interaction texts, we construct a high-quality dataset covering 28 categories of sentiment labels (such as admiration, grief, neutral, etc.), ensuring the diversity of model training data and consistency of labeling.
[0089] Dividing the data set into a training set, a validation set, and a test set based on a preset ratio, and training the large language model using the training set until the loss value of the loss function is less than a preset loss threshold;
[0090] Verifying the trained large language model using the validation set, testing the verified large language model using the test set, and deploying the tested large language model using containerization technology;
[0091] By using containerization technology to deploy trained LLMs, it supports rapid elastic scaling, adapts to high-concurrency real-time interaction scenarios (such as virtual customer service and game NPCs), and facilitates subsequent model updates and maintenance.
[0092] The emotional tags include at least 0-admiration, 1-amusement, 2-anger, 3-annyance, 4-approval, 5-caring, 6-confusion, 7-curiosity, 8-desire, 9-disappointment, 10-disapproval, 11-disgust, 12-embarrassment, 13-excitingment, 14-fear, 15-gratitude, 16-grief, 17-joy, 18-love, 19-nervousness, 20-optimism, 21-pride, 22-realization, 23-rel ief, 24-remorse, 25-sadness, 26-surprise, and 27-neutral; each of the emotional tags corresponds to at least one facial expression animation.
[0093] The step S5 further includes:
[0094] The real-time interactive text is converted into interactive voice through EdgeTTS, and the interactive voice is played synchronously when the three-dimensional digital human plays the expression animation.
[0095] By integrating the Large Language Model (LLM) with 3D animation technology, the intelligent and high-precision generation of 3D digital human expressions is achieved. Its core advantages lie in: the first LLM architecture with hierarchical feature fusion (interactive text encoding → dialogue structure perception → multi-dimensional knowledge fusion → emotion prediction), combined with a refined expression library covering 84 emotional themes and dynamic optimization of key frames, which solves the problems of rough emotion classification and abrupt expression switching in traditional methods; at the same time, the robustness of the model is improved through a composite loss function (cross entropy + structural consistency + knowledge alignment), and containerized deployment and voice-expression synchronization technology (EdgeTTS) are used to achieve low-latency, multi-modal real-time interaction, combining engineering implementation efficiency and cross-domain generalization capabilities (virtual customer service, education, games, etc.), significantly reducing development costs while improving user experience.
[0096] A preferred embodiment of the present invention's LLM-based three-dimensional digital human expression generation system includes the following modules:
[0097] An expression animation creation module is used to create an expression library, create a large number of expression animations and store them in the expression library;
[0098] A large language model creation module, used to create a large language model for sentiment analysis and set a loss function for the large language model;
[0099] A large language model training module is used to obtain a large amount of historical interaction texts of 3D digital humans, pre-process and annotate each of the historical interaction texts to construct a data set, train the large language model using the data set and loss function, and deploy the trained large language model;
[0100] The emotion tag recognition module is used to obtain the real-time interactive text of the 3D digital human, pre-process the real-time interactive text, and then input it into the deployed large language model to obtain the recognized emotion tag;
[0101] The expression animation playback module is used to match the corresponding expression animation from the expression library through the emotion tag, and load the matched expression animation into the face of the three-dimensional digital human for playback.
[0102] In the expression animation creation module, the expression themes of the expression animation include at least laughter, happiness, sweetness, excitement, surprise, pleasure, enthusiasm, complacency, satisfaction, showing off, requesting, encouraging, cheering, comforting, dependence, trust, relaxation, sympathy, pity, distress, guilt, touching, respecting, cute, pitiful, asking, praying, pleading, acting like a spoiled child, enjoying, longing, appreciating, approving, admiring, worshipping, loving, envying, persisting, strong, firm, resolute, witty, clever, wise, naughty, playful, calm, composed, surprised, startled, confused, hesitant, trapped, etc. Confusion, ignorance, helplessness, silence, seriousness, solemnity, caution, thinking, meditation, need, desire, shyness, shame, shame-anger, shame-joy, dissatisfaction, slight anger, sadness, grief, sorrow, anger, rage, shock and anger, disbelief, regret, disbelief, questioning, denial, pain, suffering, torment, helplessness, bitterness, tension, anxiety, worry, concern, depression, melancholy, fear, fear, fear, sorrow, rejection, grievance, reluctance, indifference, disdain, alienation, indifference, impatience, restlessness, irritability, worry, distress, smile and say hello, comfort and teasing;
[0103] The expression animation is provided with key frames including at least expression start, expression display and expression end.
[0104] By defining facial expression animations that include a large number of emotional themes (such as laughter, anger, dependence, pity, etc.), covering a wide range of emotions from positive to negative, far exceeding the single classification of traditional facial expression libraries (such as only joy, anger, sadness, and happiness), it can accurately match emotional needs in complex contexts; by setting facial expression animations to include a start→display→end key frame sequence, imitating the physiological dynamic process of human expression (such as the gradual rise and fall of the corners of the mouth when smiling), it avoids abrupt expression switches and improves the smoothness and naturalness of the animation.
[0105] The large language model creation module is specifically used to:
[0106] Creating a large language model for sentiment analysis based on the interactive text encoding layer, the dialogue structure perception layer, the feature fusion layer, and the sentiment prediction layer, and setting a loss function for the large language model;
[0107] The interactive text encoding layer is used to extract text sequence features, dialogue structure features, and context-dependent features from the input interactive text, splice the text sequence features, dialogue structure features, and context-dependent features to obtain spliced features, and input the spliced features into the dialogue structure perception layer;
[0108] The dialogue structure perception layer is used to model the dialogue roles, temporal relationships, and context dependencies of the spliced features to obtain dialogue perception features, and input the dialogue perception features into the feature fusion layer;
[0109] The feature fusion layer is used to perform multi-dimensional fusion of the dialogue perception features and the pre-trained language model knowledge to obtain the emotion discrimination features, and input the emotion discrimination features into the emotion prediction layer;
[0110] The emotion prediction layer is used to map the emotion discrimination features into a probability distribution of emotion labels, and output an emotion prediction result carrying the emotion label based on the probability distribution;
[0111] The formula of the loss function is:
[0112] L_total=λ1*L_ce+λ2*L_struct+λ3*L_align+λ4*L_reg;
[0113] Among them, L_total represents the loss value of the loss function; L_ce represents the cross entropy loss; L_struct represents the structural consistency loss; L_align represents the knowledge alignment loss; L_reg represents the regularization term; λ1, λ2, λ3, and λ4 all represent weight coefficients.
[0114] By setting up a four-layer model architecture of the large language model (encoding layer → perception layer → fusion layer → prediction layer): the interactive text encoding layer extracts text sequence, dialogue structure, and context-dependent features; the dialogue structure perception layer models role relationships and temporal logic; the feature fusion layer integrates pre-trained knowledge to enhance generalization ability; this enables the large language model to capture the implicit semantics and long-term dependencies in the dialogue, thereby improving the accuracy of sentiment label prediction.
[0115] By setting the loss function to integrate cross entropy loss (L_ce), structural consistency loss (L_struct), knowledge alignment loss (L_align) and regularization term (L_reg), and dynamically balancing different optimization objectives through weight coefficients (λ1-λ4), overfitting is effectively prevented and the model's adaptability to complex dialogue scenarios is improved.
[0116] The large language model training module is specifically used for:
[0117] Acquire a large amount of historical interaction texts of the three-dimensional digital human, perform preprocessing on each of the historical interaction texts, including at least text cleaning, text standardization, and word segmentation, annotate each of the preprocessed historical interaction texts with emotional tags, and then construct a data set;
[0118] By cleaning, standardizing, segmenting, and labeling historical interaction texts, we construct a high-quality dataset covering 28 categories of sentiment labels (such as admiration, grief, neutral, etc.), ensuring the diversity of model training data and consistency of labeling.
[0119] Dividing the data set into a training set, a validation set, and a test set based on a preset ratio, and training the large language model using the training set until the loss value of the loss function is less than a preset loss threshold;
[0120] Verifying the trained large language model using the validation set, testing the verified large language model using the test set, and deploying the tested large language model using containerization technology;
[0121] By using containerization technology to deploy trained LLMs, it supports rapid elastic scaling, adapts to high-concurrency real-time interaction scenarios (such as virtual customer service and game NPCs), and facilitates subsequent model updates and maintenance.
[0122] The emotional labels include at least 0-admiration, 1-amusement, 2-anger, 3-annyance, 4-approval, 5-caring, 6-confusion, 7-curiosity, 8-desire, 9-disappointment, 10-disapproval, 11-disgust, 12-embarrassment, 13-excitingment, 14-fear, 15-gratitude, 16-grief, 17-joy, 18-love, 19-nervousness, 20-optimism, 21-pride, 22-realization, 23-relief, 24-remorse, 25-sadness, 26-surprise, and 27-neutral; each of the emotional labels corresponds to at least one facial expression animation.
[0123] The expression animation playback module is also used for:
[0124] The real-time interactive text is converted into interactive voice through EdgeTTS, and the interactive voice is played synchronously when the three-dimensional digital human plays the expression animation.
[0125] By integrating the Large Language Model (LLM) with 3D animation technology, the intelligent and high-precision generation of 3D digital human expressions is achieved. Its core advantages lie in: the first LLM architecture with hierarchical feature fusion (interactive text encoding → dialogue structure perception → multi-dimensional knowledge fusion → emotion prediction), combined with a refined expression library covering 84 emotional themes and dynamic optimization of key frames, which solves the problems of rough emotion classification and abrupt expression switching in traditional methods; at the same time, the robustness of the model is improved through a composite loss function (cross entropy + structural consistency + knowledge alignment), and containerized deployment and voice-expression synchronization technology (EdgeTTS) are used to achieve low-latency, multi-modal real-time interaction, combining engineering implementation efficiency and cross-domain generalization capabilities (virtual customer service, education, games, etc.), significantly reducing development costs while improving user experience.
[0126] In summary, the advantages of the present invention are:
[0127] 1. Create a large number of facial expression animations and store them in an expression library, then create a large language model for sentiment analysis and set the loss function of the large language model; obtain a large amount of historical interaction texts of 3D digital humans, pre-process and annotate each historical interaction text to build a data set, train the large language model through the data set and loss function, and deploy the trained large language model; then obtain the real-time interaction text of the 3D digital human, pre-process the real-time interaction text and input it into the deployed large language model to obtain the recognized emotion label, match the corresponding facial expression animation from the expression library through the emotion label, and load the matched facial expression animation onto the face of the 3D digital human for playback; that is, the emotional label recognition of the real-time interaction text through the pre-trained large language model can effectively cope with complex and changing contexts, and can perform specific analysis on each sentence of the real-time interaction text to obtain the corresponding emotion label, and then match the expression animation from the expression library that stores rich facial expression animation based on the emotion label for playback, thereby greatly improving the accuracy, naturalness and vividness of the 3D digital human's expression generation.
[0128] 2. By defining facial expression animations that encompass a wide range of emotional themes (such as laughter, anger, dependence, and pity), a broad spectrum of emotions from positive to negative is covered, far exceeding the single classification of traditional facial expression libraries (such as only joy, anger, sadness, and happiness), and can accurately match emotional needs in complex contexts. By setting facial expression animations to include a start→display→end keyframe sequence, the physiological dynamics of human expression are simulated (such as the gradual rise and fall of the corners of the mouth when smiling), avoiding abrupt expression switches and improving the smoothness and naturalness of animations.
[0129] 3. By setting up a four-layer model architecture of the large language model (encoding layer → perception layer → fusion layer → prediction layer): the interactive text encoding layer extracts text sequence, dialogue structure, and context-dependent features; the dialogue structure perception layer models role relationships and temporal logic; the feature fusion layer integrates pre-trained knowledge to enhance generalization capabilities; this enables the large language model to capture the implicit semantics and long-term dependencies in the dialogue, thereby improving the accuracy of sentiment label prediction.
[0130] 4. By setting the loss function to integrate cross entropy loss (L_ce), structural consistency loss (L_struct), knowledge alignment loss (L_align) and regularization term (L_reg), and dynamically balancing different optimization objectives through weight coefficients (λ1-λ4), it effectively prevents overfitting and improves the model's adaptability to complex dialogue scenarios.
[0131] 5. By cleaning, standardizing, segmenting, and labeling historical interaction texts, we construct a high-quality dataset covering 28 categories of sentiment labels (such as admiration, grief, neutral, etc.) to ensure the diversity of model training data and consistency of labeling.
[0132] 6. By using containerization technology to deploy trained LLM, it supports rapid elastic scaling, adapts to high-concurrency real-time interaction scenarios (such as virtual customer service, game NPCs), and facilitates subsequent model updates and maintenance.
[0133] 7. By integrating the Large Language Model (LLM) with 3D animation technology, the intelligent and high-precision generation of 3D digital human expressions is achieved. Its core advantages lie in: the first LLM architecture with hierarchical feature fusion (interactive text encoding → dialogue structure perception → multi-dimensional knowledge fusion → emotion prediction), combined with a refined expression library covering 84 emotional themes and dynamic optimization of key frames, which solves the problems of rough emotion classification and abrupt expression switching in traditional methods; at the same time, the robustness of the model is improved through a composite loss function (cross entropy + structural consistency + knowledge alignment), and low-latency, multi-modal real-time interaction is achieved by using containerized deployment and voice-expression synchronization technology (EdgeTTS), which combines engineering implementation efficiency and cross-domain generalization capabilities (virtual customer service, education, games, etc.), significantly reducing development costs while improving user experience.
[0134] Although the specific embodiments of the present invention are described above, those skilled in the art should understand that the specific embodiments described are merely illustrative and are not intended to limit the scope of the present invention. Equivalent modifications and changes made by those skilled in the art in accordance with the spirit of the present invention should be included within the scope of protection of the claims of the present invention.
Claims
1. A method for generating expressions of a 3D digital human based on LLM, characterized by: The steps include: Step S1, creating an expression library, creating a large number of expression animations and storing them in the expression library; Step S2: creating a large language model for sentiment analysis and setting a loss function for the large language model; Step S3: obtaining a large amount of historical interaction texts of the three-dimensional digital human, pre-processing and annotating each of the historical interaction texts to construct a data set, training a large language model using the data set and a loss function, and deploying the trained large language model; Step S4: obtaining the real-time interactive text of the three-dimensional digital human, pre-processing the real-time interactive text, and inputting it into the deployed large language model to obtain recognized emotion labels; Step S5: Match the corresponding facial expression animation from the facial expression library using the emotion tag, and load the matched facial expression animation onto the face of the three-dimensional digital human for playback.
2. The method for generating expressions of a three-dimensional digital human based on LLM according to claim 1, wherein: In step S1, the expression themes of the expression animation include at least laughter, happiness, sweetness, excitement, surprise, pleasure, enthusiasm, complacency, satisfaction, showing off, requesting, encouraging, cheering, comforting, dependence, trust, relaxation, sympathy, pity, distress, guilt, touching, respecting, cute, pitiful, asking, praying, begging, acting coquettishly, enjoying, longing, appreciating, approving, admiring, worshipping, loving, envying, persisting, strong, firm, resolute, witty, clever, wise, naughty, playful, calm, composed, surprised, startled, confused, hesitant, bewildered, Ignorance, helplessness, silence, seriousness, solemnity, caution, thinking, contemplation, need, desire, shyness, shame, shame-anger, shame-joy, dissatisfaction, slight anger, sadness, grief, sorrow, anger, rage, shock and anger, disbelief, regret, disbelief, questioning, denial, pain, suffering, torment, helplessness, bitterness, tension, anxiety, worry, concern, depression, melancholy, fear, fear, terror, sorrow, rejection, grievance, reluctance, indifference, disdain, alienation, indifference, impatience, restlessness, irritability, worry, distress, annoyance, distress, smiling hello, comforting and teasing; The expression animation is provided with key frames including at least expression start, expression display and expression end.
3. The method for generating expressions of a three-dimensional digital human based on LLM according to claim 1, wherein: The step S2 is specifically as follows: Creating a large language model for sentiment analysis based on the interactive text encoding layer, the dialogue structure perception layer, the feature fusion layer, and the sentiment prediction layer, and setting a loss function for the large language model; The interactive text encoding layer is used to extract text sequence features, dialogue structure features, and context-dependent features from the input interactive text, splice the text sequence features, dialogue structure features, and context-dependent features to obtain spliced features, and input the spliced features into the dialogue structure perception layer; The dialogue structure perception layer is used to model the dialogue roles, temporal relationships, and context dependencies of the spliced features to obtain dialogue perception features, and input the dialogue perception features into the feature fusion layer; The feature fusion layer is used to perform multi-dimensional fusion of the dialogue perception features and the pre-trained language model knowledge to obtain the emotion discrimination features, and input the emotion discrimination features into the emotion prediction layer; The emotion prediction layer is used to map the emotion discrimination features into a probability distribution of emotion labels, and output an emotion prediction result carrying the emotion label based on the probability distribution; The formula of the loss function is: L_total=λ1*L_ce+λ2*L_struct+λ3*L_align+λ4*L_reg; Among them, L_total represents the loss value of the loss function; L_ce represents the cross entropy loss; L_struct represents the structural consistency loss; L_align represents the knowledge alignment loss; L_reg represents the regularization term; λ1, λ2, λ3, and λ4 all represent weight coefficients.
4. The method for generating expressions of a 3D digital human based on LLM according to claim 1, wherein: The step S3 is specifically as follows: Acquire a large amount of historical interaction texts of the three-dimensional digital human, perform preprocessing on each of the historical interaction texts, including at least text cleaning, text standardization, and word segmentation, annotate each of the preprocessed historical interaction texts with emotional tags, and then construct a data set; Dividing the data set into a training set, a validation set, and a test set based on a preset ratio, and training the large language model using the training set until the loss value of the loss function is less than a preset loss threshold; Verifying the trained large language model using the validation set, testing the verified large language model using the test set, and deploying the tested large language model using containerization technology; The emotional labels include at least 0-admiration, 1-amusement, 2-anger, 3-annyance, 4-approval, 5-caring, 6-confusion, 7-curiosity, 8-desire, 9-disappointment, 10-disapproval, 11-disgust, 12-embarrassment, 13-excitingment, 14-fear, 15-gratitude, 16-grief, 17-joy, 18-love, 19-nervousness, 20-optimism, 21-pride, 22-realization, 23-relief, 24-remorse, 25-sadness, 26-surprise, and 27-neutral; each of the emotional labels corresponds to at least one facial expression animation.
5. The method for generating expressions of a 3D digital human based on LLM according to claim 1, wherein: The step S5 further includes: The real-time interactive text is converted into interactive voice through EdgeTTS, and the interactive voice is played synchronously when the three-dimensional digital human plays the expression animation.
6. A 3D digital human expression generation system based on LLM, characterized by: Includes the following modules: An expression animation creation module is used to create an expression library, create a large number of expression animations and store them in the expression library; A large language model creation module, used to create a large language model for sentiment analysis and set a loss function for the large language model; A large language model training module is used to obtain a large amount of historical interaction texts of 3D digital humans, pre-process and annotate each of the historical interaction texts to construct a data set, train the large language model using the data set and loss function, and deploy the trained large language model; The emotion tag recognition module is used to obtain the real-time interactive text of the 3D digital human, pre-process the real-time interactive text, and then input it into the deployed large language model to obtain the recognized emotion tag; The expression animation playback module is used to match the corresponding expression animation from the expression library through the emotion tag, and load the matched expression animation into the face of the three-dimensional digital human for playback.
7. The LLM-based 3D digital human expression generation system according to claim 6, characterized in that: In the expression animation creation module, the expression themes of the expression animation include at least laughter, happiness, sweetness, excitement, surprise, pleasure, enthusiasm, complacency, satisfaction, showing off, requesting, encouraging, cheering, comforting, dependence, trust, relaxation, sympathy, pity, distress, guilt, touching, respecting, cute, pitiful, asking, praying, pleading, acting like a spoiled child, enjoying, longing, appreciating, approving, admiring, worshipping, loving, envying, persisting, strong, firm, resolute, witty, clever, wise, naughty, playful, calm, composed, surprised, startled, confused, hesitant, trapped, etc. Confusion, ignorance, helplessness, silence, seriousness, solemnity, caution, thinking, meditation, need, desire, shyness, shame, shame-anger, shame-joy, dissatisfaction, slight anger, sadness, grief, sorrow, anger, rage, shock and anger, disbelief, regret, disbelief, questioning, denial, pain, suffering, torment, helplessness, bitterness, tension, anxiety, worry, concern, depression, melancholy, fear, fear, fear, sorrow, rejection, grievance, reluctance, indifference, disdain, alienation, indifference, impatience, restlessness, irritability, worry, distress, smile and say hello, comfort and teasing; The expression animation is provided with key frames including at least expression start, expression display and expression end.
8. The LLM-based 3D digital human expression generation system according to claim 6, characterized in that: The large language model creation module is specifically used to: Creating a large language model for sentiment analysis based on the interactive text encoding layer, the dialogue structure perception layer, the feature fusion layer, and the sentiment prediction layer, and setting a loss function for the large language model; The interactive text encoding layer is used to extract text sequence features, dialogue structure features, and context-dependent features from the input interactive text, splice the text sequence features, dialogue structure features, and context-dependent features to obtain spliced features, and input the spliced features into the dialogue structure perception layer; The dialogue structure perception layer is used to model the dialogue roles, temporal relationships, and context dependencies of the spliced features to obtain dialogue perception features, and input the dialogue perception features into the feature fusion layer; The feature fusion layer is used to perform multi-dimensional fusion of the dialogue perception features and the pre-trained language model knowledge to obtain the emotion discrimination features, and input the emotion discrimination features into the emotion prediction layer; The emotion prediction layer is used to map the emotion discrimination features into a probability distribution of emotion labels, and output an emotion prediction result carrying the emotion label based on the probability distribution; The formula of the loss function is: L_total=λ1*L_ce+λ2*L_struct+λ3*L_align+λ4*L_reg; Among them, L_total represents the loss value of the loss function; L_ce represents the cross entropy loss; L_struct represents the structural consistency loss; L_align represents the knowledge alignment loss; L_reg represents the regularization term; λ1, λ2, λ3, and λ4 all represent weight coefficients.
9. The LLM-based 3D digital human expression generation system according to claim 6, characterized in that: The large language model training module is specifically used for: Acquire a large amount of historical interaction texts of the three-dimensional digital human, perform preprocessing on each of the historical interaction texts, including at least text cleaning, text standardization, and word segmentation, annotate each of the preprocessed historical interaction texts with emotional tags, and then construct a data set; Dividing the data set into a training set, a validation set, and a test set based on a preset ratio, and training the large language model using the training set until the loss value of the loss function is less than a preset loss threshold; Verifying the trained large language model using the validation set, testing the verified large language model using the test set, and deploying the tested large language model using containerization technology; The emotional labels include at least 0-admiration, 1-amusement, 2-anger, 3-annyance, 4-approval, 5-caring, 6-confusion, 7-curiosity, 8-desire, 9-disappointment, 10-disapproval, 11-disgust, 12-embarrassment, 13-excitingment, 14-fear, 15-gratitude, 16-grief, 17-joy, 18-love, 19-nervousness, 20-optimism, 21-pride, 22-realization, 23-relief, 24-remorse, 25-sadness, 26-surprise, and 27-neutral; each of the emotional labels corresponds to at least one facial expression animation.
10. The LLM-based 3D digital human expression generation system according to claim 6, characterized in that: The expression animation playback module is also used for: The real-time interactive text is converted into interactive voice through EdgeTTS, and the interactive voice is played synchronously when the three-dimensional digital human plays the expression animation.