Digital human face driving method and device, electronic equipment, storage medium and program product
By obtaining the lip shape and emotion parameters of the phonemes to calculate the control weights and drive the digital mouth shape, the problem of lip movement being disturbed by emotions in the existing technology is solved, and full emotional expression and improved naturalness of lip shape are achieved.
Patent Information
- Application Number
- CN202510828525.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-19
- Publication Date
- 2025-10-03
AI Technical Summary
Existing 3D digital human facial driving technology has limitations in emotional expression. Lip movements are easily disturbed by emotional curves, resulting in unnatural and insufficient emotional expression.
By obtaining the phoneme sequence corresponding to the text to be broadcasted by the target digital human, the lip shape parameters and emotional parameters of the phonemes are obtained, the control weight of the phonemes is calculated, and the lip shape of the target digital human is driven according to the control weight. The emotional information is integrated to generate an emotional and stylized lip shape driving effect.
Without affecting the accuracy of lip fitting, it achieves full expression of emotions and improves the naturalness of lip movement, generating more vivid and natural lip animation.
Smart Images

Figure CN120747313A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a digital human face driving method, a digital human face driving device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] In recent years, with the rapid development of computer graphics, artificial intelligence, and virtual reality, 3D (three-dimensional) digital human technology has been widely applied in various fields. From intelligent customer service to exhibition hall explanations, from virtual idols to live streaming sales, 3D digital humans, with their realistic appearance and natural interaction, have brought new user experiences. In conversational scenarios, digital humans need to generate corresponding facial expressions in real time based on the input text content.
[0003] Currently, research on 3D digital human facial driving technology primarily focuses on improving the synchronization and accuracy between lip movements and text content. While this research has made significant progress in improving the naturalness of lip movements, it still has limitations in terms of emotional expression. Some existing solutions attempt to separate lip movements from emotions, achieving emotional expression by simply overlaying emotion curves with lip movement curves. However, this approach often results in lip movements being disrupted by the emotion curves, resulting in unnatural lip movements and insufficient emotional expression. Summary of the Invention
[0004] The present disclosure provides a digital human face driving technology solution.
[0005] According to one aspect of the present disclosure, a digital human face driving method is provided, comprising:
[0006] Obtain the phoneme sequence corresponding to the text to be broadcasted by the target digital human;
[0007] For any phoneme in the phoneme sequence, obtaining the lip shape parameter and emotion parameter corresponding to the phoneme;
[0008] Determining a control weight corresponding to the phoneme according to the lip shape parameter and the emotion parameter;
[0009] The lip shape of the target digital human is driven according to the control weight corresponding to the phoneme.
[0010] In one possible implementation,
[0011] The lip shape parameters include the pronunciation start time and pronunciation end time of the phoneme;
[0012] Determining the control weight corresponding to the phoneme according to the lip shape parameter and the emotion parameter includes: for any target time between the pronunciation start time and the pronunciation end time, determining the control weight corresponding to the phoneme at the target time according to the lip shape parameter and the emotion parameter;
[0013] Driving the lip shape of the target digital human according to the control weight corresponding to the phoneme includes: driving the lip shape of the target digital human according to the control weight corresponding to the phoneme at the target moment.
[0014] In a possible implementation, at the target moment, driving the lip shape of the target digital human according to the control weight corresponding to the phoneme includes:
[0015] At the target moment, the lip shape of the target digital human is driven according to the control weights corresponding to the various phonemes whose pronunciation time periods cover the target moment.
[0016] In a possible implementation, the lip shape parameters further include a key frame time corresponding to the phoneme, wherein the key frame time represents a time when the articulatory organ reaches a target pronunciation position corresponding to the phoneme;
[0017] The step of determining, for any target moment between the pronunciation start moment and the pronunciation end moment, the control weight corresponding to the phoneme at the target moment according to the lip shape parameter and the emotion parameter, includes:
[0018] For any target moment between the pronunciation start moment and the pronunciation end moment, calculating the time interval between the key frame moment and the target moment;
[0019] Determine a control weight corresponding to the phoneme at the target moment according to the lip shape parameter, the emotion parameter, and the time interval.
[0020] In a possible implementation, determining the control weight corresponding to the phoneme at the target moment according to the lip shape parameter, the emotion parameter, and the time interval includes:
[0021] In response to the target time being later than or equal to the key frame time, determining a control weight corresponding to the phoneme at the target time according to the lip shape parameter, the emotion parameter, the time interval, and a speed parameter of a backward decay of a pronunciation action curve corresponding to the phoneme;
[0022] or,
[0023] In response to the target moment being earlier than the key frame moment, the control weight corresponding to the phoneme at the target moment is determined according to the lip shape parameter, the emotion parameter, the time interval, and the speed parameter of the forward attenuation of the pronunciation action curve corresponding to the phoneme.
[0024] In a possible implementation, the lip shape parameter includes a phoneme influence weight corresponding to the phoneme, and the control weight corresponding to the phoneme is positively correlated with the phoneme influence weight.
[0025] In a possible implementation, the emotion parameter includes an emotion type influence weight corresponding to the phoneme, and the control weight corresponding to the phoneme is positively correlated with the emotion type influence weight.
[0026] In a possible implementation, obtaining the emotion parameter corresponding to the phoneme includes:
[0027] determining the sentiment type of the sentence to which the phoneme belongs;
[0028] The emotion type influence weight corresponding to the emotion type of the sentence is determined as the emotion type influence weight corresponding to the phoneme.
[0029] In a possible implementation, the emotion parameter includes the emotion intensity corresponding to the phoneme, and the control weight corresponding to the phoneme is positively correlated with the emotion intensity.
[0030] In a possible implementation, obtaining the emotion parameter corresponding to the phoneme includes:
[0031] Obtaining a keyword detection result of the sentence to which the phoneme belongs;
[0032] Determine the emotion intensity corresponding to the phoneme according to the keyword detection result.
[0033] In a possible implementation, the emotion parameter includes a speech rate influence factor corresponding to the phoneme, wherein the speech rate influence factor is used to indicate the degree to which the emotion affects the speed at which the pronunciation organ reaches the target pronunciation position corresponding to the phoneme.
[0034] In a possible implementation, obtaining the emotion parameter corresponding to the phoneme includes:
[0035] determining the sentiment type of the sentence to which the phoneme belongs;
[0036] The speaking rate influence factor corresponding to the emotion type of the sentence is determined as the speaking rate influence factor corresponding to the phoneme.
[0037] In a possible implementation, the method further includes:
[0038] Inputting the text to be broadcast into a pre-trained language model, and outputting expression labels corresponding to at least part of the phonemes in the text to be broadcast through the language model;
[0039] According to the expression labels corresponding to at least part of the phonemes in the text to be broadcasted, the expression of the target digital person when broadcasting the text to be broadcasted is driven.
[0040] In a possible implementation, before inputting the to-be-broadcasted text into a pre-trained language model, the method further includes:
[0041] Inputting a training text into the language model, and outputting, through the language model, expression labels corresponding to at least some phonemes in the training text, and the meanings corresponding to the expression labels in the training text;
[0042] The language model is trained according to expression labels corresponding to at least some of the phonemes in the training text and annotation information corresponding to the training text.
[0043] According to one aspect of the present disclosure, a digital human face driving device is provided, comprising:
[0044] An acquisition module, used to obtain the phoneme sequence corresponding to the text to be broadcasted by the target digital human;
[0045] An acquisition module, configured to obtain, for any phoneme in the phoneme sequence, a lip shape parameter and an emotion parameter corresponding to the phoneme;
[0046] A determination module, configured to determine a control weight corresponding to the phoneme according to the lip shape parameter and the emotion parameter;
[0047] The first driving module is used to drive the lip shape of the target digital human according to the control weight corresponding to the phoneme.
[0048] In one possible implementation,
[0049] The lip shape parameters include the pronunciation start time and pronunciation end time of the phoneme;
[0050] The determination module is configured to: determine, for any target moment between the pronunciation start moment and the pronunciation end moment, the control weight corresponding to the phoneme at the target moment according to the lip shape parameter and the emotion parameter;
[0051] The first driving module is used to drive the lip shape of the target digital human at the target moment according to the control weight corresponding to the phoneme.
[0052] In a possible implementation, the first driving module is configured to:
[0053] At the target moment, the lip shape of the target digital human is driven according to the control weights corresponding to the various phonemes whose pronunciation time periods cover the target moment.
[0054] In a possible implementation, the lip shape parameters further include a key frame time corresponding to the phoneme, wherein the key frame time represents a time when the articulatory organ reaches a target pronunciation position corresponding to the phoneme;
[0055] The determining module is used for:
[0056] For any target moment between the pronunciation start moment and the pronunciation end moment, calculating the time interval between the key frame moment and the target moment;
[0057] Determine a control weight corresponding to the phoneme at the target moment according to the lip shape parameter, the emotion parameter, and the time interval.
[0058] In a possible implementation, the determining module is configured to:
[0059] In response to the target time being later than or equal to the key frame time, determining a control weight corresponding to the phoneme at the target time according to the lip shape parameter, the emotion parameter, the time interval, and a speed parameter of a backward decay of a pronunciation action curve corresponding to the phoneme;
[0060] or,
[0061] In response to the target moment being earlier than the key frame moment, the control weight corresponding to the phoneme at the target moment is determined according to the lip shape parameter, the emotion parameter, the time interval, and the speed parameter of the forward attenuation of the pronunciation action curve corresponding to the phoneme.
[0062] In a possible implementation, the lip shape parameter includes a phoneme influence weight corresponding to the phoneme, and the control weight corresponding to the phoneme is positively correlated with the phoneme influence weight.
[0063] In a possible implementation, the emotion parameter includes an emotion type influence weight corresponding to the phoneme, and the control weight corresponding to the phoneme is positively correlated with the emotion type influence weight.
[0064] In a possible implementation, the obtaining module is configured to:
[0065] determining the sentiment type of the sentence to which the phoneme belongs;
[0066] The emotion type influence weight corresponding to the emotion type of the sentence is determined as the emotion type influence weight corresponding to the phoneme.
[0067] In a possible implementation, the emotion parameter includes the emotion intensity corresponding to the phoneme, and the control weight corresponding to the phoneme is positively correlated with the emotion intensity.
[0068] In a possible implementation, the obtaining module is configured to:
[0069] Obtaining a keyword detection result of the sentence to which the phoneme belongs;
[0070] Determine the emotion intensity corresponding to the phoneme according to the keyword detection result.
[0071] In a possible implementation, the emotion parameter includes a speech rate influence factor corresponding to the phoneme, wherein the speech rate influence factor is used to indicate the degree to which the emotion affects the speed at which the pronunciation organ reaches the target pronunciation position corresponding to the phoneme.
[0072] In a possible implementation, the obtaining module is configured to:
[0073] determining the sentiment type of the sentence to which the phoneme belongs;
[0074] The speaking rate influence factor corresponding to the emotion type of the sentence is determined as the speaking rate influence factor corresponding to the phoneme.
[0075] In a possible implementation, the apparatus further includes:
[0076] A first prediction module is configured to input the text to be broadcast into a pre-trained language model, and output expression labels corresponding to at least part of the phonemes in the text to be broadcast through the language model;
[0077] The second driving module is used to drive the expression of the target digital person when reading the text to be read according to the expression tags corresponding to at least part of the phonemes in the text to be read.
[0078] In a possible implementation, the apparatus further includes a training module, wherein the training module is configured to:
[0079] Inputting a training text into the language model, and outputting, through the language model, expression labels corresponding to at least some phonemes in the training text, and the meanings corresponding to the expression labels in the training text;
[0080] The language model is trained according to expression labels corresponding to at least some of the phonemes in the training text and annotation information corresponding to the training text.
[0081] According to one aspect of the present disclosure, an electronic device is provided, comprising: one or more processors; a memory for storing executable instructions; wherein the one or more processors are configured to call the executable instructions stored in the memory to execute the above method.
[0082] According to one aspect of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the above method is implemented.
[0083] According to one aspect of the present disclosure, a computer program product is provided, including a computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes the above method.
[0084] In the embodiment of the present disclosure, by obtaining the phoneme sequence corresponding to the text to be broadcasted by the target digital human, for any phoneme in the phoneme sequence, the lip shape parameters and emotion parameters corresponding to the phoneme are obtained, and the control weight corresponding to the phoneme is determined according to the lip shape parameters and the emotion parameters. According to the control weight corresponding to the phoneme, the lip shape of the target digital human is driven. In this way, in the process of calculating the lip shape movement curve of the digital human, the emotional information is aligned and integrated to generate an emotional and stylized lip shape driving effect, so that the emotion can be fully expressed without affecting the accuracy of the lip shape fitting.
[0085] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure.
[0086] Further features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments with reference to the attached drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0087] The accompanying drawings herein are incorporated into and constitute a part of the specification. These drawings illustrate embodiments consistent with the present disclosure and, together with the specification, are used to explain the technical solutions of the present disclosure.
[0088] Figure 1 A flowchart of a digital human face driving method provided by an embodiment of the present disclosure is shown.
[0089] Figure 2 A curve showing the changes in the degree of jaw opening and closing when speech is produced with different emotions.
[0090] Figure 3 A schematic diagram showing a curve of changes in the degree of mandibular opening and closing obtained by using the digital human face driving method provided by an embodiment of the present disclosure.
[0091] Figure 4A schematic diagram illustrating an application scenario of the digital human face driving method provided by an embodiment of the present disclosure.
[0092] Figure 5 A block diagram of a digital human face driving device provided by an embodiment of the present disclosure is shown.
[0093] Figure 6 A block diagram of an electronic device 1900 provided by an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0094] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numerals in the accompanying drawings represent elements with the same or similar functions. Although various aspects of the embodiments are shown in the accompanying drawings, the drawings are not necessarily drawn to scale unless otherwise indicated.
[0095] The word “exemplary” is used exclusively herein to mean “serving as an example, example, or illustration.” Any embodiment described herein as “exemplary” is not necessarily to be construed as preferred or advantageous over other embodiments.
[0096] The term "and / or" herein simply describes an association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent the existence of A alone, the simultaneous existence of A and B, and the existence of B alone. Furthermore, the term "at least one" herein refers to any combination of at least two of any one or more of a plurality of items. For example, "at least one of A, B, and C" can represent any one or more elements selected from the set consisting of A, B, and C.
[0097] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.
[0098] In the related art, lip movement is separated from emotion, and the emotion curve is combined with the lip movement curve in a simple superposition manner, resulting in the lip shape being affected and the emotion expression being insufficient.
[0099] The inventors of the present application discovered that emotions and lip movements are not completely decoupled. For example, when expressing "surprise", the amplitude of lip movement often increases, while when expressing "fear", the amplitude decreases.
[0100] In order to solve technical problems similar to those described above, the embodiments of the present disclosure provide a digital human facial driving method, which obtains a phoneme sequence corresponding to the text to be broadcasted by the target digital human, obtains the lip shape parameters and emotion parameters corresponding to any phoneme in the phoneme sequence, determines the control weight corresponding to the phoneme based on the lip shape parameters and the emotion parameters, and drives the lip shape of the target digital human based on the control weight corresponding to the phoneme. In this way, in the process of calculating the lip shape motion curve of the digital human, the emotional information is aligned and integrated to generate an emotional and stylized lip shape driving effect, which allows the emotions to be fully expressed without affecting the accuracy of the lip shape fitting.
[0101] The digital human face driving method provided by the embodiment of the present disclosure is described in detail below with reference to the accompanying drawings.
[0102] Figure 1 A flow chart of a digital human face driving method provided by an embodiment of the present disclosure is shown. In one possible implementation, the execution subject of the digital human face driving method may be a digital human face driving device. For example, the digital human face driving method may be executed by a terminal device or a server or other electronic device. The terminal device may be a user equipment (UE), a mobile device, a user terminal, a terminal, a cellular phone, a cordless phone, a personal digital assistant (PDA), a handheld device, a computing device, a vehicle-mounted device or a wearable device, etc. In some possible implementations, the digital human face driving method may be implemented by a processor calling computer-readable instructions stored in a memory. As Figure 1 As shown, the digital human face driving method includes steps S11 to S14.
[0103] In step S11, a phoneme sequence corresponding to the text to be broadcasted by the target digital human is obtained.
[0104] In step S12, for any phoneme in the phoneme sequence, the lip shape parameter and emotion parameter corresponding to the phoneme are obtained.
[0105] In step S13, the control weight corresponding to the phoneme is determined according to the lip shape parameter and the emotion parameter.
[0106] In step S14, the lip shape of the target digital human is driven according to the control weight corresponding to the phoneme.
[0107] In the embodiment of the present disclosure, the target digital human can be any digital human to be facially driven. The target digital human can be a three-dimensional digital human or a two-dimensional digital human, without limitation.
[0108] The to-be-announced text can represent the text content to be announced or read aloud by the target digital human. For example, in the application scenario of intelligent customer service, the text content of the to-be-announced text can include answers to user questions; in the application scenario of exhibition hall explanations, the text content of the to-be-announced text can include the script of the exhibit introduction; in the application scenario of virtual idols, the text content of the to-be-announced text can include the virtual idol's lines; in the application scenario of live streaming sales, the text content of the to-be-announced text can include the product description of the live streaming sales; and so on.
[0109] Phonemes are the smallest units of speech, similar to letters in written language. In linguistics, phonemes are used to describe the sounds or speech in a language. Each language has its own unique phoneme system, and these phonemes combine to form words, phrases, and sentences. In the disclosed embodiment, after obtaining the text to be broadcasted by the target digital human, the text to be broadcasted can be converted into a phoneme sequence.
[0110] In the disclosed embodiment, for any phoneme in the phoneme sequence, the corresponding lip shape parameters and emotion parameters are obtained. The lip shape parameters corresponding to the phoneme can reflect the characteristics of the mouth shape when the phoneme is pronounced. The emotion parameters can reflect the changes in lip movement under different emotional states. The use of emotion parameters allows lip movement to be more vivid and natural, not just mechanical movement based on the phoneme.
[0111] In the disclosed embodiment, the control weight corresponding to the phoneme can be determined based on the lip shape parameters and emotion parameters corresponding to the phoneme. By adjusting the control weights corresponding to different phonemes, a more delicate and natural lip shape animation can be achieved.
[0112] In one possible implementation, the lip shape parameters include the pronunciation start time and pronunciation end time of the phoneme; determining the control weight corresponding to the phoneme based on the lip shape parameters and the emotion parameters includes: for any target moment between the pronunciation start time and the pronunciation end time, determining the control weight corresponding to the phoneme at the target moment based on the lip shape parameters and the emotion parameters; driving the lip shape of the target digital human based on the control weight corresponding to the phoneme includes: driving the lip shape of the target digital human at the target moment based on the control weight corresponding to the phoneme.
[0113] In this implementation, the pronunciation start time and pronunciation end time of the phoneme define the time range of the phoneme pronunciation. In some application scenarios, the pronunciation start time of the phoneme can also be called the start frame time of the phoneme, and the pronunciation end time of the phoneme can also be called the end frame time of the phoneme, which is not limited here. In one example, the pronunciation start time of the phoneme can be ts Indicates that the pronunciation end time of the phoneme can be expressed as t e express.
[0114] For each target moment between the start and end of pronunciation of the phoneme, the control weight corresponding to the phoneme at that target moment can be determined based on the mouth shape parameters and the emotion parameters. The control weight corresponding to the phoneme changes over time, thereby more carefully reflecting the dynamic changes in the mouth shape during the pronunciation process.
[0115] At the target moment, the target digital human's mouth shape can be driven according to the calculated control weight corresponding to the phoneme. That is, at each pronunciation point, the target digital human's mouth shape can be adjusted in real time based on the current phoneme and control weight, achieving synchronization between mouth shape and speech.
[0116] This implementation considers the start and end times of phoneme pronunciation and combines them with emotional parameters to determine control weights, ensuring that the target digital human's lip movements closely match the pronunciation process of the speech. Furthermore, because the control weights vary over time, the lip animation can more smoothly and delicately display subtle changes in pronunciation, enhancing the animation's realism and naturalness.
[0117] In a possible implementation, at the target moment, according to the control weight corresponding to the phoneme, driving the lip shape of the target digital human includes: at the target moment, according to the control weight corresponding to each phoneme whose pronunciation time period covers the target moment, driving the lip shape of the target digital human.
[0118] In this implementation, at any target moment, the pronunciation time periods of multiple phonemes may cover that target moment, and each phoneme has a corresponding control weight at that target moment. At the target moment, the target digital human's mouth shape is driven not only by the control weight corresponding to a single phoneme, but also by the control weights of all phonemes whose pronunciation time periods cover that target moment. In other words, the target digital human's mouth shape is influenced by multiple phonemes, and each phoneme contributes to the final mouth shape according to its control weight.
[0119] This implementation enables more accurate simulation of the complex changes in human mouth shape during pronunciation. In human speech, changes in mouth shape are often not caused by a single phoneme alone, but rather by the combined effects of multiple phonemes. By comprehensively considering the control weights of multiple phonemes to drive the target digital human's mouth shape, the changes in human mouth shape during pronunciation can be more realistically restored, making the target digital human's lip animation more natural and lifelike. This not only improves the visual quality of the target digital human's broadcast or voice animation, but also helps enhance the audience's immersion and interactive experience.
[0120] In one possible implementation, the ControlRig system can be used to control and drive the facial movements of the target digital human. In one example, 36 ControlRig elements related to the mouth and 3 ControlRig elements related to the chin can be selected to create viseme assets corresponding to the target digital human. At the target moment, the control weights corresponding to each phoneme covering the pronunciation time period at the target moment can be used to drive the target digital human's lip shape based on the created assets.
[0121] In a possible implementation, the lip shape parameters also include a key frame moment corresponding to the phoneme, wherein the key frame moment represents the moment when the pronunciation organ reaches the target pronunciation position corresponding to the phoneme; for any target moment between the pronunciation start moment and the pronunciation end moment, determining the control weight corresponding to the phoneme at the target moment according to the lip shape parameters and the emotion parameters, including: for any target moment between the pronunciation start moment and the pronunciation end moment, calculating the time interval between the key frame moment and the target moment; determining the control weight corresponding to the phoneme at the target moment according to the lip shape parameters, the emotion parameters and the time interval.
[0122] During speech production, the vocal organs (including the lips, tongue, and throat) perform a series of movements to form specific phonemes. These phonemes are the basic units of speech, such as vowels and consonants. Each phoneme has a specific pronunciation position and method. The vocal organs must move to that position accurately and vibrate or emit airflow in the correct manner to produce the correct phoneme. The key frame moment corresponding to a phoneme can represent the moment when the vocal organs reach the target pronunciation position required by the phoneme during the pronunciation process.
[0123] In one example, the key frame time can be t k express.
[0124] In this implementation, by introducing keyframe moments (the moments when the vocal organs reach the target pronunciation position corresponding to the phoneme) and considering the impact of the time interval between the keyframe moment and the target moment on the control weight, the accuracy and naturalness of the target digital human's lip animation can be further improved. The introduction of keyframe moments enables the lip animation to more accurately capture the position changes of the vocal organs at specific moments, thereby more realistically restoring the pronunciation process. At the same time, considering the impact of the time interval between the keyframe moment and the target moment on the control weight, the lip animation can transition more smoothly between different pronunciation stages, reducing abruptness and unnaturalness.
[0125] In one possible implementation, the determining the control weight corresponding to the phoneme at the target moment based on the lip shape parameters, the emotional parameters and the time interval includes: in response to the target moment being later than or equal to the key frame moment, determining the control weight corresponding to the phoneme at the target moment based on the lip shape parameters, the emotional parameters, the time interval and the speed parameter of the pronunciation action curve corresponding to the phoneme attenuating backward; or, in response to the target moment being earlier than the key frame moment, determining the control weight corresponding to the phoneme at the target moment based on the lip shape parameters, the emotional parameters, the time interval and the speed parameter of the pronunciation action curve corresponding to the phoneme attenuating forward.
[0126] In this implementation, the pronunciation action curve corresponding to the phoneme can describe the movement trajectory and speed change of the pronunciation organ corresponding to the phoneme during the pronunciation process.
[0127] In this implementation, the speed parameter of the pronunciation action curve corresponding to the phoneme attenuating backward can represent the speed of the pronunciation action curve corresponding to the phoneme attenuating backward, and the speed parameter of the pronunciation action curve corresponding to the phoneme attenuating forward can represent the speed of the pronunciation action curve corresponding to the phoneme attenuating forward. In one example, the speed parameter of the pronunciation action curve corresponding to the phoneme attenuating backward can be expressed as θ bw Indicates that the speed parameter of the forward decay of the pronunciation action curve corresponding to the phoneme can be θ fw express.
[0128] As an example of this implementation, when the target moment is later than or equal to the key frame moment, the control weight corresponding to the phoneme at the target moment can be determined based on the mouth shape parameter, the emotion parameter, the time interval, and the speed parameter of the pronunciation action curve corresponding to the phoneme attenuating backward. In this example, if the target moment is later than or equal to the key frame moment, it can be determined that the pronunciation organ corresponding to the phoneme has reached the target pronunciation position. Therefore, when the target moment is later than or equal to the key frame moment, the speed parameter of the pronunciation action curve corresponding to the phoneme attenuating backward is taken into consideration.
[0129] As an example of this implementation, when the target moment is earlier than the key frame moment, the control weight corresponding to the phoneme at the target moment can be determined based on the lip shape parameter, the emotion parameter, the time interval, and the speed parameter of the forward decay of the pronunciation action curve corresponding to the phoneme. In this example, if the target moment is earlier than the key frame moment, it can be determined that the pronunciation organ corresponding to the phoneme has not yet reached the target pronunciation position. Therefore, when the target moment is earlier than the key frame moment, the speed parameter of the forward decay of the pronunciation action curve corresponding to the phoneme is taken into consideration.
[0130] This implementation introduces a decay rate parameter into the pronunciation action curve, which more accurately reflects the dynamic characteristics of mouth shape changes during pronunciation. Whether the articulators are approaching or moving away from the target position, the corresponding decay rate is factored into the control weights, allowing for more precise control of the lip animation. This not only enhances the realism and naturalness of the animation, but also makes the digital human's lip shape changes smoother and more coherent.
[0131] In another possible implementation, the determining the control weight corresponding to the phoneme at the target moment based on the lip shape parameters, the emotional parameters and the time interval includes: in response to the target moment being later than the key frame moment, determining the control weight corresponding to the phoneme at the target moment based on the lip shape parameters, the emotional parameters, the time interval and the speed parameter of the pronunciation action curve corresponding to the phoneme attenuating backward; or, in response to the target moment being earlier than or equal to the key frame moment, determining the control weight corresponding to the phoneme at the target moment based on the lip shape parameters, the emotional parameters, the time interval and the speed parameter of the pronunciation action curve corresponding to the phoneme attenuating forward.
[0132] In a possible implementation, the lip shape parameter includes a phoneme influence weight corresponding to the phoneme, and the control weight corresponding to the phoneme is positively correlated with the phoneme influence weight.
[0133] The phoneme influence weight corresponding to any phoneme may represent the degree of influence of the phoneme on the lip shape, and the phoneme influence weight corresponding to any phoneme may be a fixed value. In one example, the phoneme influence weight may be represented by α.
[0134] In this implementation, the control weight is positively correlated with the phoneme influence weight, that is, the larger the phoneme influence weight is, the larger the corresponding control weight is.
[0135] This implementation introduces phoneme influence weights and establishes a positive correlation with the control weights, enabling more precise control of the target digital human's lip movements during the broadcast. Highlighting important phonemes not only helps viewers better understand the broadcast content but also enhances the animation's visual impact and appeal. This also makes the target digital human's lip movements more closely aligned with real human pronunciation habits, further enhancing the animation's realism and naturalness.
[0136] In a possible implementation, the emotion parameter includes an emotion type influence weight corresponding to the phoneme, and the control weight corresponding to the phoneme is positively correlated with the emotion type influence weight.
[0137] In this implementation, the emotion type influence weight corresponding to the phoneme can represent the degree of influence of the emotion type contained in the phoneme on the lip animation. The emotion type may be joy, anger, sadness, etc., and each emotion type will have a specific impact on the pronunciation of the phoneme and the change of the lip shape.
[0138] In one example, the sentiment type influence weight can be α emo express.
[0139] This implementation introduces emotion type influence weights and establishes a positive correlation with the control weights, enabling more accurate simulation of the changes in human lip movements when expressing different emotions. This not only enhances the realism and naturalness of lip-sync animation, but also makes the target digital human's expressions and emotions richer and more vivid.
[0140] In a possible implementation, obtaining the emotion parameter corresponding to the phoneme includes: determining the emotion type of the sentence to which the phoneme belongs; and determining the emotion type influence weight corresponding to the emotion type of the sentence as the emotion type influence weight corresponding to the phoneme.
[0141] For example, emotion types can be amazement, anger, joy, fear, sadness, disgust, neutral, etc., and the emotion type influence weight corresponding to each emotion type can be a preset fixed value. In one example, the emotion type influence weight corresponding to each emotion type can be determined by analyzing a large amount of emotional facial video data.
[0142] In this implementation, the emotion type of each sentence in the text to be broadcast can be determined separately, and the emotion type influence weight corresponding to the emotion type of the sentence can be directly assigned to each phoneme in the sentence as the emotion type influence weight corresponding to these phonemes. Since each phoneme in a sentence is part of the emotional expression of the sentence, it can be considered that they jointly carry the emotional information of the sentence. By directly applying the emotion type influence weight corresponding to the emotion type of the sentence to the phonemes, it can be ensured that the emotional expression of the entire sentence is consistently and accurately reflected in the lip animation.
[0143] This implementation directly applies the emotion-type influence weights corresponding to the sentence's emotion type to the phonemes, avoiding the tedious process of analyzing the emotion type phoneme by phoneme and improving processing efficiency. Furthermore, because the phonemes in a sentence collectively carry the sentence's emotional information, this implementation ensures that the emotional expression of the entire sentence is coherently and uniformly reflected in the lip-sync animation, enhancing the animation's emotional expressiveness and the audience's sense of immersion.
[0144] In a possible implementation, the emotion parameter includes the emotion intensity corresponding to the phoneme, and the control weight corresponding to the phoneme is positively correlated with the emotion intensity.
[0145] In this implementation, emotional intensity is a quantitative indicator that represents the strength of the emotion conveyed by a phoneme. Different phonemes may express emotions with varying intensity, and this difference directly impacts the presentation of the lip-sync animation.
[0146] In this implementation, the control weight is positively correlated with the emotional intensity. That is, the greater the emotional intensity of a phoneme, the greater its corresponding control weight. This way, in the lip-sync animation, phonemes with high emotional intensity will receive more attention, and their lip-sync changes will be more obvious and prominent.
[0147] In one example, sentiment intensity can be expressed as β emo express.
[0148] This implementation introduces emotional intensity as an emotional parameter and establishes a positive correlation with the control weight, allowing the lip-sync animation to more accurately reflect the emotional intensity of the phonemes. Differences in emotional intensity often convey richer emotional information and enhance the audience's perceptual experience. By highlighting phonemes with high emotional intensity, not only does this enhance the visual impact of the lip-sync animation, but it also makes the emotional expression of the 3D digital human more delicate and realistic.
[0149] In a possible implementation, obtaining the emotion parameter corresponding to the phoneme includes: obtaining a keyword detection result of the sentence to which the phoneme belongs; and determining the emotion intensity corresponding to the phoneme according to the keyword detection result.
[0150] In this implementation, to determine the emotional intensity corresponding to a phoneme, sentence-level keyword detection can be performed first. Keyword detection is a natural language processing technique that identifies keywords or phrases related to emotional expression within a sentence. These keywords or phrases often reflect the emotional orientation or intensity of the sentence. After obtaining the keyword detection results, the emotional intensity corresponding to the phoneme can be determined based on the keyword detection results.
[0151] This implementation utilizes keyword detection technology to determine the emotional intensity corresponding to phonemes, enabling automated acquisition of emotional parameters and improving processing efficiency and accuracy. Furthermore, because keyword detection directly reflects the emotional information of a sentence, the emotional parameters obtained in this way are more closely aligned with actual emotional expression, helping to enhance the emotional expressiveness of lip-sync animations.
[0152] In a possible implementation, the emotion parameter includes a speech rate influence factor corresponding to the phoneme, wherein the speech rate influence factor is used to indicate the degree to which the emotion affects the speed at which the pronunciation organ reaches the target pronunciation position corresponding to the phoneme.
[0153] In this implementation, the speech rate influence factor indicates how quickly the vocal organs reach the target pronunciation position corresponding to a phoneme under the influence of emotion. In other words, when emotions change, people's speech rate often adjusts, and this adjustment reflects the impact of emotion on the pronunciation process. The introduction of the speech rate influence factor enables the generation of the target digital human's lip animation to more accurately reflect the changes in speech rate under the influence of emotion. By analyzing the text content and emotion type of the text to be broadcast, the speech rate influence factor corresponding to each phoneme can be determined.
[0154] In one example, the speech rate influence factor can be expressed as τ emo express.
[0155] This implementation, by introducing speech rate factors corresponding to phonemes, allows for more precise control over the generation of the target digital human's lip animation, particularly in terms of changes in speech rate under the influence of emotion. This allows the target digital human's lip animation to not only accurately reflect the pronunciation process but also vividly demonstrate the impact of emotion on pronunciation. This helps enhance the realism and naturalness of the target digital human's announcements or speech animations, and strengthens the audience's sense of immersion and emotional resonance.
[0156] In a possible implementation, obtaining the emotion parameter corresponding to the phoneme includes: determining the emotion type of the sentence to which the phoneme belongs; and determining the speech rate influence factor corresponding to the emotion type of the sentence as the speech rate influence factor corresponding to the phoneme.
[0157] In this implementation, in order to obtain the speech rate influence factor corresponding to the phoneme, the emotion type of the sentence to which the phoneme belongs can be determined first. The emotion type can be determined by analyzing the content, context and vocabulary used in the sentence. For example, the sentence may express different emotions such as joy, anger, sadness, etc. After determining the emotion type of the sentence, the speech rate influence factor corresponding to this emotion type can be directly assigned to each phoneme in the sentence. The speech rate influence factor is a quantitative indicator that reflects how fast the pronunciation organs reach the target pronunciation position corresponding to the phoneme under a specific emotion. Different emotion types often lead to different changes in speech rate. For example, the speech rate may increase when angry, and the speech rate may slow down when sad.
[0158] In this implementation, the speech rate influencing factor corresponding to each emotion type may be a preset fixed value.
[0159] In one example, a large amount of emotional facial video data can be analyzed to obtain speech rate influencing factors corresponding to various emotion types.
[0160] This implementation simplifies the acquisition of emotional parameters by directly using sentence-level emotion types to determine the speech rate influencing factors corresponding to phonemes, while ensuring the consistency and accuracy of emotional expression. This allows the target digital human's lip animation to more realistically reflect the changes in human speech rate under different emotional states, enhancing the animation's emotional expressiveness and the audience's sense of immersion.
[0161] In one example, Formula 1 can be used to determine the control weight corresponding to a certain phoneme at the target time t:
[0162]
[0163] Where τ = t k -t,t k It can represent the key frame time corresponding to the phoneme;
[0164] α may represent the phoneme influence weight corresponding to the phoneme;
[0165] α emo The influence weight of the emotion type corresponding to the phoneme can be represented;
[0166] θ bw A speed parameter that can represent the backward decay of the pronunciation action curve corresponding to the phoneme;
[0167] θ fw A speed parameter that can represent the forward decay of the pronunciation action curve corresponding to the phoneme;
[0168] τ emo The speech rate influence factor corresponding to the phoneme may be represented;
[0169] C can be a constant, for example, 1 or 2;
[0170] β emo It can represent the emotional intensity corresponding to the phoneme.
[0171] In one possible implementation, the method further includes: inputting the text to be broadcast into a pre-trained language model, and outputting expression labels corresponding to at least some of the phonemes in the text to be broadcast through the language model; and driving the expression of the target digital person when broadcasting the text to be broadcast based on the expression labels corresponding to at least some of the phonemes in the text to be broadcast.
[0172] As an example of this implementation, the language model is a large language model (LLM).
[0173] As an example of this implementation, the expression corresponding to the expression tag may be a micro-expression.
[0174] In this implementation, the expression tags may include expression tags corresponding to facial parts such as eyebrows, eyes, cheeks, and nose.
[0175] For example, the expression tags corresponding to eyebrows may include the inner corners of eyebrows rising, the outer corners of eyebrows rising, the eyebrows moving horizontally, the eyebrows moving vertically, etc. The expression tags corresponding to eyes may include tightening eyelids, widening eyes, squinting eyes, eyes turning to one side, pupil dilation, increased blinking frequency, etc. The expression tags corresponding to cheeks may include bulging cheeks, sunken cheeks, rosy cheeks, shaking cheeks, deepening smile, etc. The expression tags corresponding to noses may include dilated nostrils, upturned nose, wrinkled nose, twitching nose, heavier nasal breathing, etc.
[0176] For example, raised inner corners of eyebrows can be used to express sadness, doubt, uncertainty, prayer, etc.; widened eyes can be used to express surprise, fear, or excitement; puffed cheeks can express anger or dissatisfaction; flared nostrils can be used to express anger, rage, or surprise, and so on.
[0177] In this implementation, a series of expression tags can be predefined, and different expression tags can represent different facial expressions or emotions. Each expression tag can correspond to a pre-made expression animation clip, which can be used to express a specific emotion or state. According to the time corresponding to the phoneme (for example, the pronunciation start time or key frame time corresponding to the phoneme, etc.), the insertion time of the expression animation clip corresponding to the expression tag in the text to be broadcast can be determined. According to the insertion time and the expression animation clip, an expression animation curve can be generated, wherein the expression animation curve can describe the change process of the expression animation clip from the beginning to the end. The expression animation curve can ensure the smooth transition and natural performance of the expression animation clip, so that the changes in micro-expressions are coordinated with the character's voice and movements. The expression animation curve can be used to drive the facial animation of the target digital human, that is, the facial model of the target digital human can move and deform according to the instructions of the expression animation curve, thereby presenting the facial expression corresponding to the expression tag.
[0178] Related technologies for 3D digital human facial animation generally lack consideration of micro-expression movements in the upper face. Micro-expression movements, such as eyebrow movement and eye movement, are closely related to emotional expression and context. However, most existing research ignores the importance of these micro-expression movements, resulting in a lack of vividness and expressiveness in digital human facial animation.
[0179] In this implementation, the text to be broadcast is input into a pre-trained language model, and the expression labels corresponding to at least some of the phonemes in the text to be broadcast are output through the language model. Based on the expression labels corresponding to at least some of the phonemes in the text to be broadcast, the expression of the target digital human when broadcasting the text to be broadcast is driven. Thus, the micro-expressions of the target digital human are generated through the pre-trained language model, which can make the facial organs (such as eyebrows, eyes, cheeks, nose, etc.) of the target digital human move naturally following the text content of the text to be broadcast, thereby improving the expressiveness and vividness of the digital human's facial animation.
[0180] In one possible implementation, before inputting the text to be broadcast into a pre-trained language model, the method further includes: inputting a training text into the language model, outputting, through the language model, expression labels corresponding to at least some of the phonemes in the training text, and the corresponding meanings of the expression labels in the training text; and training the language model based on the expression labels corresponding to at least some of the phonemes in the training text and the annotation information corresponding to the training text.
[0181] In this implementation, training text is fed into a language model, which then outputs emoji labels corresponding to at least some of the phonemes in the training text, along with the meanings of these emoji labels in the training text. These meanings are primarily used as a reference during manual calibration to determine whether the emoji labels output by the language model match the emotional content of the training text, but are not directly used in the language model training process.
[0182] For example, the output of the language model can be "Once upon a time, there was a little boy who {inner corners of eyebrows rise} (meaning: as if recalling the beautiful village) lived in a beautiful village. His parents were farmers, and they had to work hard in the fields every day {eyebrows move horizontally} (meaning: expressing their hard work and hardship). However, the little boy was not satisfied with such a life. He had a dream, {eyebrows move vertically} (meaning: as if thinking about his dream), which was to go to the big city. So, he {outer corners of eyebrows rise} (meaning: expressing his determination and courage) He decided to leave his hometown to pursue his dream. He encountered various difficulties in the city, but he never gave up. Finally, he succeeded and became a successful businessman. He returned to his hometown and used his wealth to help his family and fellow villagers. He told them that as long as they had a dream, it would come true. This is his story, the story of an ordinary person's struggle and success.
[0183] In this implementation, a language model can be trained based on the expression labels corresponding to phonemes in a training text and the corresponding annotation information (e.g., manually annotated correct expression labels). By comparing the expression labels output by the language model with the annotation information, the value of the language model's loss function can be calculated. Algorithms such as backpropagation can then be used to adjust the language model's parameters, thereby optimizing the language model's performance.
[0184] In one example, the training text may be obtained by collecting or generating, etc. For example, another large language model may be used to generate the training text.
[0185] In this implementation, the trained language model can better understand the emotional content of text and accurately predict the expression labels corresponding to phonemes. This helps improve the emotional expression of the target digital human, making it more closely resemble the emotions expressed by real people. Furthermore, by providing the meaning of the expression labels as a reference for manual calibration, the language model's transparency and interpretability are increased, making calibration and adjustment more convenient and efficient.
[0186] In one possible implementation, the language model input may also include the target digital human's character setting information. This character setting information may include personality, age, identity, and other information, without limitation. By incorporating this character setting information, micro-expressions that better match the character's characteristics can be generated, making the digital human's expressions more vivid, natural, and personalized.
[0187] Figure 2 The figure shows a curve showing the change of jaw opening and closing degree when the speech is produced with different emotions. Figure 2 In the example, the horizontal axis is the time axis, and the unit can be frames; the vertical axis can represent the degree of jaw opening and closing. Figure 2 It can be seen that emotions and lip movements are not linearly superimposed. By adopting the digital human face driving method provided by the embodiment of the present disclosure, it is possible to align emotions with phoneme lip movements while ensuring that the lip pronunciation movements conform to the actual pronunciation process.
[0188] Figure 3 A schematic diagram showing a curve showing the change in the degree of mandibular opening and closing obtained by using the digital human face driving method provided by the embodiment of the present disclosure. Figure 3 In the figure, the horizontal axis is the time axis, and the unit can be frame; the vertical axis can represent the degree of mandibular opening and closing.
[0189] The digital human face driving method provided by the disclosed embodiments can be implemented within existing systems or frameworks without requiring the addition of or reliance on other complex algorithmic modules. In other words, the digital human face driving method provided by the disclosed embodiments can be efficiently integrated into existing technology systems, leveraging existing resources and functions to achieve facial animation generation without requiring complex system reconstruction or introducing additional computational burdens.
[0190] The digital human face driving method provided in the embodiments of the present disclosure can be applied to technical fields such as 3D digital human face animation, Lipsync, three-dimensional lip animation, virtual humans, text driving, large language models, computer vision, etc., and is not limited here.
[0191] The following describes the digital human face driving method provided by the embodiment of the present disclosure through a specific application scenario. Figure 4 A schematic diagram illustrating an application scenario of the digital human face driving method provided by an embodiment of the present disclosure is shown. In this application scenario, the digital human face driving method can be applied to a conversation scenario.
[0192] In this application scenario, a pre-trained large language model (LLM) can be used to respond to conversations and generate emoticon labels. For example, the input (Q) of the large language model is "Tell a story," and the output (A) is "Once upon a time, there was a little boy who lived in a beautiful village. His parents were farmers, and they worked hard in the fields every day. However, the little boy was not satisfied with such a life. He had a dream: to visit a big city. So, he decided to leave his hometown to pursue his dream. He encountered various difficulties in the city, but he never gave up. In the end, he succeeded and became a successful businessman. He returned to his hometown and used his wealth to help his family and villagers. He told them that as long as you have a dream, it is possible to achieve it. This is his story, a story of the struggle and success of an ordinary person."
[0193] In this application scenario, the text output by the large language model can serve as the target digital human's intended broadcast text. The emotion recognition module can perform emotion recognition on each sentence in the intended broadcast text to determine the corresponding emotion type. The intended broadcast text can then be converted into a phoneme sequence.
[0194] For each phoneme in the phoneme sequence, the control weight corresponding to each phoneme at each moment can be determined using Formula 1 above. At any moment, the control weight corresponding to each phoneme covering the pronunciation time period can be used to drive the target digital human's mouth shape, and the micro-expressions of the 3D digital human can be controlled based on the expression labels output by the large language model.
[0195] It is understood that the above-mentioned various method embodiments mentioned in this disclosure can be combined with each other to form combined embodiments without violating the principle logic. Due to space limitations, this disclosure will not go into details. It is understood by those skilled in the art that in the above-mentioned methods of specific implementation, the specific execution order of each step should be determined by its function and possible internal logic.
[0196] In addition, the present disclosure also provides a digital human face driving device, an electronic device, a computer-readable storage medium, and a computer program product, all of which can be used to implement any digital human face driving method provided by the present disclosure. The corresponding technical solutions and technical effects can be found in the corresponding records in the method section and will not be repeated here.
[0197] Figure 5 FIG. 1 is a block diagram of a digital human face driving device provided by an embodiment of the present disclosure. Figure 5 As shown, the digital human face driving device includes:
[0198] An acquisition module 51 is used to acquire a phoneme sequence corresponding to the text to be broadcasted by the target digital human;
[0199] An obtaining module 52 is configured to obtain, for any phoneme in the phoneme sequence, a lip shape parameter and an emotion parameter corresponding to the phoneme;
[0200] A determination module 53, configured to determine a control weight corresponding to the phoneme according to the lip shape parameter and the emotion parameter;
[0201] The first driving module 54 is configured to drive the lip shape of the target digital human according to the control weight corresponding to the phoneme.
[0202] In one possible implementation,
[0203] The lip shape parameters include the pronunciation start time and pronunciation end time of the phoneme;
[0204] The determining module 53 is configured to: for any target moment between the pronunciation start moment and the pronunciation end moment, determine the control weight corresponding to the phoneme at the target moment according to the lip shape parameter and the emotion parameter;
[0205] The first driving module 54 is used to drive the lip shape of the target digital human at the target moment according to the control weight corresponding to the phoneme.
[0206] In a possible implementation, the first driving module 54 is configured to:
[0207] At the target moment, the lip shape of the target digital human is driven according to the control weights corresponding to the various phonemes whose pronunciation time periods cover the target moment.
[0208] In a possible implementation, the lip shape parameters further include a key frame time corresponding to the phoneme, wherein the key frame time represents a time when the articulatory organ reaches a target pronunciation position corresponding to the phoneme;
[0209] The determining module 53 is used to:
[0210] For any target moment between the pronunciation start moment and the pronunciation end moment, calculating the time interval between the key frame moment and the target moment;
[0211] Determine a control weight corresponding to the phoneme at the target moment according to the lip shape parameter, the emotion parameter, and the time interval.
[0212] In a possible implementation, the determining module 53 is configured to:
[0213] In response to the target time being later than or equal to the key frame time, determining a control weight corresponding to the phoneme at the target time according to the lip shape parameter, the emotion parameter, the time interval, and a speed parameter of a backward decay of a pronunciation action curve corresponding to the phoneme;
[0214] or,
[0215] In response to the target moment being earlier than the key frame moment, the control weight corresponding to the phoneme at the target moment is determined according to the lip shape parameter, the emotion parameter, the time interval, and the speed parameter of the forward attenuation of the pronunciation action curve corresponding to the phoneme.
[0216] In a possible implementation, the lip shape parameter includes a phoneme influence weight corresponding to the phoneme, and the control weight corresponding to the phoneme is positively correlated with the phoneme influence weight.
[0217] In a possible implementation, the emotion parameter includes an emotion type influence weight corresponding to the phoneme, and the control weight corresponding to the phoneme is positively correlated with the emotion type influence weight.
[0218] In a possible implementation, the obtaining module 52 is configured to:
[0219] determining the sentiment type of the sentence to which the phoneme belongs;
[0220] The emotion type influence weight corresponding to the emotion type of the sentence is determined as the emotion type influence weight corresponding to the phoneme.
[0221] In a possible implementation, the emotion parameter includes the emotion intensity corresponding to the phoneme, and the control weight corresponding to the phoneme is positively correlated with the emotion intensity.
[0222] In a possible implementation, the obtaining module 52 is configured to:
[0223] Obtaining a keyword detection result of the sentence to which the phoneme belongs;
[0224] Determine the emotion intensity corresponding to the phoneme according to the keyword detection result.
[0225] In a possible implementation, the emotion parameter includes a speech rate influence factor corresponding to the phoneme, wherein the speech rate influence factor is used to indicate the degree to which the emotion affects the speed at which the pronunciation organ reaches the target pronunciation position corresponding to the phoneme.
[0226] In a possible implementation, the obtaining module 52 is configured to:
[0227] determining the sentiment type of the sentence to which the phoneme belongs;
[0228] The speaking rate influence factor corresponding to the emotion type of the sentence is determined as the speaking rate influence factor corresponding to the phoneme.
[0229] In a possible implementation, the apparatus further includes:
[0230] A first prediction module is configured to input the text to be broadcast into a pre-trained language model, and output expression labels corresponding to at least part of the phonemes in the text to be broadcast through the language model;
[0231] The second driving module is used to drive the expression of the target digital person when reading the text to be read according to the expression tags corresponding to at least part of the phonemes in the text to be read.
[0232] In a possible implementation, the apparatus further includes a training module, wherein the training module is configured to:
[0233] Inputting a training text into the language model, and outputting, through the language model, expression labels corresponding to at least some phonemes in the training text, and the meanings corresponding to the expression labels in the training text;
[0234] The language model is trained according to expression labels corresponding to at least some of the phonemes in the training text and annotation information corresponding to the training text.
[0235] In some embodiments, the functions or modules included in the device provided by the embodiments of the present disclosure can be used to execute the method described in the above method embodiments. Its specific implementation and technical effects can refer to the description of the above method embodiments. For the sake of brevity, they will not be repeated here.
[0236] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement the above method. The computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium.
[0237] The embodiment of the present disclosure further provides a computer program, comprising a computer-readable code. When the computer-readable code is executed in an electronic device, a processor in the electronic device executes the above method.
[0238] An embodiment of the present disclosure further provides a computer program product, including computer-readable code, or a non-volatile computer-readable storage medium carrying the computer-readable code. When the computer-readable code runs in an electronic device, a processor in the electronic device executes the above method.
[0239] An embodiment of the present disclosure also provides an electronic device, comprising: one or more processors; a memory for storing executable instructions; wherein the one or more processors are configured to call the executable instructions stored in the memory to execute the above method.
[0240] The electronic device may be provided as a terminal, a server, or other forms of devices.
[0241] Figure 6 FIG1 shows a block diagram of an electronic device 1900 provided by an embodiment of the present disclosure. For example, the electronic device 1900 can be provided as a server or a terminal. Figure 6 The electronic device 1900 includes a processing component 1922, which further includes one or more processors, and a memory resource represented by a memory 1932 for storing instructions executable by the processing component 1922, such as an application. The application stored in the memory 1932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above-described method.
[0242] The electronic device 1900 may further include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 may operate based on an operating system stored in the memory 1932, such as a Microsoft Server operating system (Windows Server 2003). TM ), a graphical user interface operating system launched by Apple (MacOS X TM ), a multi-user, multi-process computer operating system (Unix TM ), a free and open source Unix-like operating system (Linux TM ), an open-source Unix-like operating system (FreeBSD TM ) or similar.
[0243] In an exemplary embodiment, a non-volatile computer-readable storage medium is also provided, such as a memory 1932 including computer program instructions that can be executed by the processing component 1922 of the electronic device 1900 to perform the above method.
[0244] The present disclosure may be a system, method and / or computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.
[0245] A computer-readable storage medium can be a tangible device that can hold and store instructions for use by an instruction execution device. A computer-readable storage medium can be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination thereof. As used herein, a computer-readable storage medium is not to be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through an electrical wire.
[0246] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, fiber optic transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions to be stored in the computer-readable storage medium in each computing / processing device.
[0247] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, and conventional procedural programming languages such as "C" language or similar programming languages. Computer-readable program instructions may be executed entirely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., utilizing an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of the computer-readable program instructions. The electronic circuit may execute the computer-readable program instructions, thereby realizing various aspects of the present disclosure.
[0248] Various aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0249] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device is generated that implements the functions / actions specified in one or more blocks in the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium, where these instructions cause the computer, programmable data processing device, and / or other device to operate in a specific manner. Thus, the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing various aspects of the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0250] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more blocks in the flowchart and / or block diagram.
[0251] The flow charts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flow chart or block diagram can represent a part of a module, program segment or instruction, and the part of the module, program segment or instruction contains one or more executable instructions for realizing the prescribed logical function. In some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the prescribed function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0252] The computer program product may be implemented in hardware, software, or a combination thereof. In one embodiment, the computer program product is implemented as a computer storage medium. In another embodiment, the computer program product is implemented as a software product, such as a software development kit (SDK).
[0253] The above description of the various embodiments tends to emphasize the differences between the various embodiments. The same or similar aspects can be referenced with each other and will not be repeated herein for the sake of brevity.
[0254] If the technical solutions of the embodiments of the present disclosure involve personal information, the products applying the technical solutions of the embodiments of the present disclosure have clearly informed the personal information processing rules and obtained the individual's voluntary consent before processing the personal information. If the technical solutions of the embodiments of the present disclosure involve sensitive personal information, the products applying the technical solutions of the embodiments of the present disclosure have obtained the individual's separate consent before processing the sensitive personal information, and at the same time meet the "explicit consent" requirement. For example, on personal information collection devices such as cameras, clear and prominent signs are set to inform that the personal information collection scope has been entered and personal information will be collected. If the individual voluntarily enters the collection scope, it is deemed that they agree to the collection of their personal information; or on the personal information processing device, when the personal information processing rules are notified by obvious signs / information, the individual's authorization is obtained through pop-up information or by asking the individual to upload their personal information. The personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the type of personal information processed.
[0255] While various embodiments of the present disclosure have been described above, the above descriptions are illustrative, non-exhaustive, and not intended to be limiting of the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or improvements to existing technologies, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A digital human face driving method, characterized in that: include: Obtain the phoneme sequence corresponding to the text to be broadcasted by the target digital human; For any phoneme in the phoneme sequence, obtaining the lip shape parameter and emotion parameter corresponding to the phoneme; Determining a control weight corresponding to the phoneme according to the lip shape parameter and the emotion parameter; The lip shape of the target digital human is driven according to the control weight corresponding to the phoneme.
2. The method according to claim 1, characterized in that The lip shape parameters include the pronunciation start time and pronunciation end time of the phoneme; Determining the control weight corresponding to the phoneme according to the lip shape parameter and the emotion parameter includes: for any target time between the pronunciation start time and the pronunciation end time, determining the control weight corresponding to the phoneme at the target time according to the lip shape parameter and the emotion parameter; Driving the lip shape of the target digital human according to the control weight corresponding to the phoneme includes: driving the lip shape of the target digital human according to the control weight corresponding to the phoneme at the target moment.
3. The method according to claim 2, characterized in that At the target moment, driving the lip shape of the target digital human according to the control weight corresponding to the phoneme includes: At the target moment, the lip shape of the target digital human is driven according to the control weights corresponding to the various phonemes whose pronunciation time periods cover the target moment.
4. The method according to claim 2, characterized in that The lip shape parameters also include a key frame time corresponding to the phoneme, wherein the key frame time represents the time when the pronunciation organ reaches the target pronunciation position corresponding to the phoneme; The step of determining, for any target moment between the pronunciation start moment and the pronunciation end moment, the control weight corresponding to the phoneme at the target moment according to the lip shape parameter and the emotion parameter, includes: For any target moment between the pronunciation start moment and the pronunciation end moment, calculating the time interval between the key frame moment and the target moment; Determine a control weight corresponding to the phoneme at the target moment according to the lip shape parameter, the emotion parameter, and the time interval.
5. The method according to claim 4, characterized in that The determining, based on the lip shape parameter, the emotion parameter, and the time interval, a control weight corresponding to the phoneme at the target moment includes: In response to the target time being later than or equal to the key frame time, determining a control weight corresponding to the phoneme at the target time according to the lip shape parameter, the emotion parameter, the time interval, and a speed parameter of a backward decay of a pronunciation action curve corresponding to the phoneme; or, In response to the target moment being earlier than the key frame moment, the control weight corresponding to the phoneme at the target moment is determined according to the lip shape parameter, the emotion parameter, the time interval, and the speed parameter of the forward attenuation of the pronunciation action curve corresponding to the phoneme.
6. The method according to any one of claims 1 to 5, characterized in that The lip shape parameter includes a phoneme influence weight corresponding to the phoneme, and the control weight corresponding to the phoneme is positively correlated with the phoneme influence weight.
7. The method according to any one of claims 1 to 5, characterized in that The emotion parameter includes an emotion type influence weight corresponding to the phoneme, and the control weight corresponding to the phoneme is positively correlated with the emotion type influence weight.
8. The method according to claim 7, characterized in that Obtaining the emotion parameter corresponding to the phoneme includes: determining the sentiment type of the sentence to which the phoneme belongs; The emotion type influence weight corresponding to the emotion type of the sentence is determined as the emotion type influence weight corresponding to the phoneme.
9. The method according to any one of claims 1 to 5, characterized in that The emotion parameter includes the emotion intensity corresponding to the phoneme, and the control weight corresponding to the phoneme is positively correlated with the emotion intensity.
10. The method according to claim 9, characterized in that Obtaining the emotion parameter corresponding to the phoneme includes: Obtaining a keyword detection result of the sentence to which the phoneme belongs; Determine the emotion intensity corresponding to the phoneme according to the keyword detection result.
11. The method according to any one of claims 1 to 5, characterized in that The emotion parameter includes a speech rate influence factor corresponding to the phoneme, wherein the speech rate influence factor is used to indicate the degree to which emotion affects the speed at which the pronunciation organ reaches the target pronunciation position corresponding to the phoneme.
12. The method according to claim 11, characterized in that Obtaining the emotion parameter corresponding to the phoneme includes: determining the sentiment type of the sentence to which the phoneme belongs; The speaking rate influence factor corresponding to the emotion type of the sentence is determined as the speaking rate influence factor corresponding to the phoneme.
13. The method according to any one of claims 1 to 5, characterized in that The method further comprises: Inputting the text to be broadcast into a pre-trained language model, and outputting expression labels corresponding to at least part of the phonemes in the text to be broadcast through the language model; According to the expression labels corresponding to at least part of the phonemes in the text to be broadcasted, the expression of the target digital person when broadcasting the text to be broadcasted is driven.
14. The method according to claim 13, characterized in that Before inputting the to-be-broadcasted text into the pre-trained language model, the method further includes: Inputting a training text into the language model, and outputting, through the language model, expression labels corresponding to at least some phonemes in the training text, and the meanings corresponding to the expression labels in the training text; The language model is trained according to expression labels corresponding to at least some of the phonemes in the training text and annotation information corresponding to the training text.
15. A digital human face driving device, characterized in that: include: An acquisition module, used to obtain the phoneme sequence corresponding to the text to be broadcasted by the target digital human; An acquisition module, configured to obtain, for any phoneme in the phoneme sequence, a lip shape parameter and an emotion parameter corresponding to the phoneme; A determination module, configured to determine a control weight corresponding to the phoneme according to the lip shape parameter and the emotion parameter; The first driving module is used to drive the lip shape of the target digital human according to the control weight corresponding to the phoneme.
16. An electronic device, characterized in that: include: one or more processors; a memory for storing executable instructions; The one or more processors are configured to call the executable instructions stored in the memory to execute the method according to any one of claims 1 to 14.
17. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 14 is implemented.
18. A computer program product comprising computer-readable code, or a non-volatile computer-readable storage medium carrying computer-readable code, characterized in that: When the computer-readable code is executed in an electronic device, a processor in the electronic device executes the method according to any one of claims 1 to 14.
Citation Information
Patent Citations
Expression animation generation method and device, equipment, storage medium and program product
CN117557696A
Mouth shape driving method, device and equipment of digital human and storage medium
CN117935807A
Digital human face driving method, device, equipment, medium and program product
CN118505864A
A digital human expression driving method, system and program product
CN119760088A
Method for animation synthesis, electronic device and storage medium
US20220375456A1
Cited By
Virtual digital population broadcast optimization method and device
CN121564153A