Digital human expression driving method and device, electronic equipment and storage medium
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING YINGZHI TECH CO LTD
- Filing Date
- 2026-04-20
- Publication Date
- 2026-08-04
AI Technical Summary
[0004]为此,本申请的第一个目的在于提出一种数字人的表情驱动方法,以实现解决数字人表情僵硬的问题,使得数字人在进行播报的过程中的表情具有情绪的自然起伏
[0014] The digital human expression-driving method, apparatus, electronic device, and storage medium provided in this application acquire the text to be read by the digital human and perform sentiment analysis on the text to obtain sentiment tags and sentiment intensity information. During the digital human's reading of the text, its facial expressions are synchronously driven based on the sentiment tags and sentiment intensity information. Therefore, by introducing time-related sentiment intensity information, this solution can solve the problem of stiff digital human expressions, enabling the digital human's expressions to have natural emotional fluctuations during the reading process, thus improving the realism and emotional credibility of the digital human.
Smart Images

Figure CN122510409A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of intelligent cockpit, in-vehicle intelligent agent, digital human interaction, scene arrangement, etc., and in particular to a digital human expression driving method, device, electronic device and storage medium. Background Technology
[0002] Driving a digital human usually involves performing sentiment analysis on a piece of text to obtain a global sentiment label, and then driving the digital human to maintain a "happy" expression until the entire passage is read. This results in a very unnatural presentation of the digital human and fails to simulate the dynamic characteristics of natural emotional fluctuations when a real person speaks. Summary of the Invention
[0003] The purpose of this application is to at least partially solve one of the technical problems in the related art.
[0004] Therefore, the first objective of this application is to propose a digital human expression-driven method to solve the problem of stiff digital human expressions, so that the digital human's expressions during broadcasting have natural emotional fluctuations.
[0005] The second objective of this application is to propose a digital human facial expression driving device.
[0006] The third objective of this application is to propose an electronic device.
[0007] The fourth objective of this application is to provide a computer-readable storage medium.
[0008] The fifth objective of this application is to provide a computer program product.
[0009] To achieve the above objectives, a first aspect of this application proposes a digital human expression-driven method, comprising: acquiring a text to be read by the digital human; performing sentiment analysis on the text to acquire sentiment tags and sentiment intensity information of the text, wherein the sentiment intensity information includes the sentiment intensity at each reading moment covered by the sentiment tags; and synchronously driving the digital human's expressions based on the sentiment tags and the sentiment intensity information during the process of the digital human reading the text.
[0010] To achieve the above objectives, a second aspect of this application provides a digital human expression-driving device, comprising: an acquisition module for acquiring a text to be read by the digital human; a sentiment analysis module for performing sentiment analysis on the text to acquire sentiment tags and sentiment intensity information of the text, wherein the sentiment intensity information includes the sentiment intensity at each reading moment covered by the sentiment tags; and an expression-driving module for synchronously driving the digital human's expressions based on the sentiment tags and the sentiment intensity information during the reading of the text by the digital human.
[0011] To achieve the above objectives, a third aspect of this application provides an electronic device, comprising: a processor; and a memory communicatively connected to the processor; the memory storing computer-executable instructions; and the processor executing the computer-executable instructions stored in the memory to enable the processor to execute the digital human expression-driven method described in the first aspect of the application.
[0012] To achieve the above objectives, a fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, the computer instructions being used to cause the computer to execute the digital human expression-driven method described in the above aspect of the embodiment.
[0013] To achieve the above objectives, a fifth aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the digital human expression-driving method described in the above aspect of the embodiment.
[0014] The digital human expression-driving method, apparatus, electronic device, and storage medium provided in this application acquire the text to be read by the digital human and perform sentiment analysis on the text to obtain sentiment tags and sentiment intensity information. During the digital human's reading of the text, its facial expressions are synchronously driven based on the sentiment tags and sentiment intensity information. Therefore, by introducing time-related sentiment intensity information, this solution can solve the problem of stiff digital human expressions, enabling the digital human's expressions to have natural emotional fluctuations during the reading process, thus improving the realism and emotional credibility of the digital human.
[0015] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0016] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0017] Figure 1A flowchart illustrating a digital human expression-driven method provided in an embodiment of this application; Figure 2 A flowchart illustrating another digital human expression-driven method provided in an embodiment of this application; Figure 3 A flowchart illustrating another digital human expression-driven method provided in an embodiment of this application; Figure 4 This is a schematic diagram of the emotional intensity time curve provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of a digital human expression driving device provided in an embodiment of this application. Detailed Implementation
[0018] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0019] The following description, with reference to the accompanying drawings, describes a digital human expression-driving method and apparatus according to embodiments of this application.
[0020] Figure 1 This is a flowchart illustrating a digital human expression-driven method provided in an embodiment of this application, as shown below. Figure 1 As shown, the digital human expression-driving method of this application embodiment includes, but is not limited to, the following steps: S101, Obtain the text to be read by the digital human.
[0021] It should be noted that the execution subject of the digital human expression-driven method provided in this application embodiment is an electronic device, which can be a terminal device. Optionally, the terminal device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be personal computers (PCs), televisions, etc. This application embodiment does not impose specific limitations.
[0022] It should be noted that the digital human in this application embodiment can be applied to scenarios requiring user responses, such as intelligent cockpits in vehicles, customer service scenarios, and consultation scenarios. In other words, the digital human can respond to user questions or requests, and the text it broadcasts can be the response text corresponding to the question or request.
[0023] In some embodiments, by receiving user input information and combining it with a large language model to parse and respond to the input information, a response text corresponding to the input information is generated as the broadcast text that the digital human needs to broadcast.
[0024] It should be noted that the digital human in this embodiment can also be applied to other scenarios where user response is not required, such as information display and educational scenarios. In other words, if the digital human can broadcast received information, the text it broadcasts can be the user's input.
[0025] In some embodiments, user input is received and used as the text to be read by the digital human.
[0026] In this embodiment of the application, no specific limitations are made on the application scenarios of digital humans.
[0027] S102, Perform sentiment analysis on the broadcast text to obtain the sentiment tags and sentiment intensity information of the broadcast text. The sentiment intensity information includes the sentiment intensity of each broadcast moment covered by the sentiment tags.
[0028] In some embodiments, Natural Language Processing (NLP) techniques can be used to perform sentiment analysis on the broadcast text to obtain sentiment tags and sentiment intensity information. Sentiment tags are used to describe, classify, or label human emotional states; for example, sentiment tags could be happiness, surprise, sadness, etc.
[0029] The emotional intensity information includes the emotional intensity of each broadcast time covered by the emotional tag. This can be obtained by determining the emotional tag of the broadcast text from the broadcast text and generating the emotional intensity and broadcast time corresponding to the emotional tag.
[0030] In some embodiments, the emotional intensity information can be an emotional intensity time curve related to the broadcast time and emotional intensity, which can intuitively display the emotional intensity corresponding to the emotional tag at different times based on the emotional intensity time curve.
[0031] In some embodiments, since sentiment tags are used to describe, classify, or label human emotional states, the broadcast text can be sentiment-classified to determine the sentiment tag of the broadcast text based on the classification result. Optionally, a sentiment tag for each sentiment can be output by extracting features from the broadcast text and performing sentiment classification based on the features.
[0032] For example, sentiment classification models can be used to classify sentiment based on features. Another approach is to obtain predefined classification rules and then perform sentiment classification based on those rules and the defined features.
[0033] In some embodiments, the emotional intensity of the broadcast text can be determined based on the words included in the broadcast text. The emotional intensity of each word can be evaluated and combined with the broadcast time of the word to generate emotional intensity covering all broadcast times, as emotional intensity information.
[0034] In some embodiments, the text to be broadcast can be converted into corresponding speech data, thereby determining the time when each word appears in the speech data as the broadcast time.
[0035] S103, during the process of the digital human broadcasting the text, the digital human's facial expressions are synchronously driven based on emotion tags and emotion intensity information.
[0036] In some embodiments, the facial expressions that the digital human needs to use during the reading of text, as well as the changes in those expressions, can be determined and synchronized to enable the digital human to have natural emotional fluctuations during the reading of text, thereby enhancing the effect of the digital human's reading.
[0037] In some embodiments, the facial expressions used by the digital human during the reading of text can be determined based on emotion tags. For example, if the emotion tag is "happy," expression 1 is used; if the emotion tag is "sad," expression 2 is used.
[0038] In other words, a correspondence between different emotion tags and expressions can be established in advance. Once the emotion tag of the text to be broadcast is determined, the expression that the digital human needs to use when broadcasting the text can be determined by querying the correspondence.
[0039] In some embodiments, the changes in facial expressions of a digital human during the reading of text can be determined based on emotion intensity information. The greater the emotion intensity, the more exaggerated the expression; the smaller the emotion intensity, the more neutral the expression.
[0040] In some embodiments, since the emotional intensity information includes the emotional intensity of each broadcast moment covered by the emotional label, the emotional intensity corresponding to each broadcast moment may be the same or different, and the expression corresponding to each broadcast moment may change.
[0041] For example, if the emotion label is "happy" and the expression 1 is selected, with the emotion intensity corresponding to expression 1 being intensity a, and if it is determined that during the broadcast of the text, the emotion intensity information indicates that the emotional mildness at broadcast time 1 is intensity a, the emotional mildness at broadcast time 2 is intensity b, and the emotional mildness at broadcast time 3 is intensity c, where intensity a is less than intensity b is less than intensity c, then expression 1 is most exaggerated at broadcast time 3 and most calm at broadcast time 1.
[0042] In some embodiments, based on emotional intensity information, changes in the digital human's facial expressions during the reading of text can be obtained, and facial expression-driven data can be generated based on these changes to drive the digital human's facial expressions. For example, the digital human can be driven to display expression 1, and expression 1 can be displayed in order from calm to exaggerated during the reading process.
[0043] Optionally, facial expression-driven data can be used to render the digital human, enabling synchronized facial expression control. During the process, facial expressions and spoken audio can be aligned based on the broadcast time to improve the accuracy and effectiveness of the digital human's broadcast.
[0044] For example, facial expression data can be input into the digital human's driving engine to drive the digital human's facial expressions to synchronize.
[0045] In some embodiments, if the emotional tag corresponding to the broadcast text includes at least one, then during the process of the digital human broadcasting the text, the digital human can present different expressions and the changes in each expression, so that the digital human's broadcast has emotional fluctuations and the broadcast effect is more natural.
[0046] The digital human expression-driven method provided in this application obtains the text to be read by the digital human and performs sentiment analysis on the text to obtain its sentiment tags and sentiment intensity information. During the digital human's reading of the text, its facial expressions are synchronously driven based on the sentiment tags and sentiment intensity information. Therefore, by introducing time-related sentiment intensity information, this solution can solve the problem of stiff digital human expressions, enabling the digital human's expressions to have natural emotional fluctuations during the reading process, thus improving the realism and emotional credibility of the digital human.
[0047] Figure 2 This is a flowchart illustrating another digital human expression-driven method provided in an embodiment of this application, as shown below. Figure 2 As shown, the digital human expression-driving method of this application embodiment includes, but is not limited to, the following steps: S201, Obtain the text to be read by the digital human.
[0048] In the embodiments of this application, step S201 can be implemented in any of the ways described in the embodiments of this application. This is not limited here and will not be described in detail.
[0049] S202, perform sentiment analysis on the broadcast text to obtain the sentiment tags, sentiment trigger words and sentiment peak words of the broadcast text.
[0050] It should be noted that sentiment tags refer to the sentiment classification of the broadcast text; sentiment trigger words refer to words that directly express emotions, that is, which words give the broadcast text emotional color; sentiment peak words refer to words with the highest sentiment intensity, that is, words that best reflect the intensity of emotions, such as "This dish is amazing!" which is a sentiment peak word.
[0051] In some embodiments, semantic understanding and feature extraction are performed on the broadcast text to obtain the emotional information carried in the broadcast text, and emotional tags are extracted from the broadcast text based on the emotional information to obtain the emotional tags of the broadcast text.
[0052] Optionally, sentiment classification can be performed based on sentiment information, and the classification results can be used as sentiment labels. For example, by extracting features from the broadcast text and using the extracted features as sentiment information, sentiment classification can be performed based on the sentiment information to output a sentiment label for each sentiment.
[0053] In some embodiments, trigger words can be extracted from the broadcast text to obtain the emotional trigger words. For example, the attention mechanism of a large model can be used to locate words that can express emotions as emotional trigger words.
[0054] Optionally, a first word library containing emotional trigger words can be pre-set, and different words can be extracted from the broadcast text. If the extracted words are the same as the words in the first word library, the words can be identified as emotional trigger words.
[0055] Optionally, if none of the words in the broadcast text match the words in the first word database, the similarity between each word and the words in the first word database can be calculated, thereby filtering out emotion trigger words from the broadcast text based on the similarity.
[0056] In some embodiments, a second word library containing sentiment peak words can be predefined, and different words can be extracted from the broadcast text. If the extracted words are the same as the words in the second word library, the words can be identified as sentiment peak words.
[0057] Optionally, if none of the words in the broadcast text match the words in the second word database, the similarity between each word and the words in the second word database can be calculated, thereby filtering out the sentiment peak words from the broadcast text based on the similarity.
[0058] S203, based on the emotional trigger words and emotional peak words, determine the emotional intensity of each broadcast moment covered by the emotional tag.
[0059] In some embodiments, after obtaining the emotional trigger words and emotional peak words, the emotional intensity of each broadcast moment covered by the emotional tag can be determined based on the emotional trigger words and emotional peak words, that is, the emotional intensity information can be determined.
[0060] In some embodiments, text-to-speech (TTS) can be used to convert the broadcast text into the speech that the data person needs to broadcast. Then, based on the speech that the data person needs to broadcast, the broadcast time of each word in the broadcast text can be determined.
[0061] In other words, by performing sentiment analysis on the broadcast text, the broadcast time of each word in the broadcast text on the TTS speech timeline is obtained.
[0062] In some embodiments, the emotional intensity of each word can be determined based on the emotional trigger word and the emotional peak word, thereby determining the emotional intensity of each broadcast time covered by the emotional tag based on the emotional intensity of each word and the broadcast time.
[0063] It should be noted that the emotional intensity of the emotional trigger word is greater than that of the emotional peak word, which in turn is greater than that of other words. At the time of broadcast, words closer to the emotional trigger word or emotional peak word have a greater emotional intensity.
[0064] In some embodiments, the emotional intensity of a word at each broadcast moment in the broadcast text can be determined by setting a baseline emotional intensity in the words and combining the emotional intensity of emotional trigger words and emotional peak words with the broadcast time.
[0065] For example, the first and last words in the broadcast text are set as the baseline sentiment intensity. The closer the broadcast time of a word is to the broadcast time of the sentiment trigger word and the sentiment peak word, the greater its corresponding sentiment intensity. If the sentiment intensity of the sentiment trigger word is intensity a, and the sentiment intensity of the sentiment peak word is intensity b, then intensity b is greater than intensity a.
[0066] If the words in the broadcast text are ordered by broadcast time as follows: first word, word 1, word 2, emotional trigger word, word 3, emotional peak word, word 4, and last word, then the emotional intensity of the first word and the last word is intensity A.
[0067] In other words, the emotional intensity of the above words, ordered by broadcast time, is as follows: Intensity A, Intensity C, Intensity D, Intensity a, Intensity e, Intensity b, Intensity f, and Intensity A again. Intensity A is less than Intensity c, less than Intensity d, less than Intensity a, less than Intensity b; Intensity f is less than Intensity b; and Intensity f is greater than Intensity A. Intensity f can be less than, equal to, or greater than Intensity e.
[0068] In some embodiments, the broadcast time and emotional intensity of the first and last words in the broadcast text are obtained, wherein the emotional intensity of the first and last words can be set as the baseline emotional intensity.
[0069] Furthermore, the broadcast time and emotional intensity of the emotional trigger words and emotional peak words are obtained respectively. Based on the broadcast time and emotional intensity of the first word and the last word, as well as the broadcast time and emotional intensity of the emotional trigger words and emotional peak words, the emotional intensity of each broadcast time covered by the emotional tag is determined.
[0070] Among them, the emotional intensity of the emotional peak word is greater than that of the emotional trigger word, and the emotional intensity of the emotional trigger word is greater than that of the first word and the last word.
[0071] In some embodiments, the emotional intensity of each word at each broadcast time covered by the emotional tag is determined based on the broadcast time and emotional intensity of the first and last words, as well as the broadcast time and emotional intensity of the emotional trigger word and the emotional peak word. Thus, emotional intensity information, that is, the emotional intensity of each broadcast time covered by the emotional tag, is determined based on the broadcast time and emotional intensity.
[0072] In this embodiment, the emotional intensity information can be an emotional intensity time curve, which can be generated based on the emotional intensity of each word at each broadcast time. The horizontal axis of the emotional intensity time curve represents the broadcast time, and the vertical axis represents the emotional intensity.
[0073] S204, during the process of the digital human broadcasting the text, the digital human's facial expressions are synchronously driven based on emotion tags and emotion intensity information.
[0074] In the embodiments of this application, step S204 can be implemented in any of the ways described in the various embodiments of this application. This is not limited here and will not be described in detail.
[0075] The digital human expression-driven method provided in this application involves performing sentiment analysis on the broadcast text to obtain its sentiment tags, sentiment trigger words, and sentiment peak words. Based on the sentiment trigger words and sentiment peak words, the emotional intensity of each broadcast moment covered by the sentiment tags is determined, i.e., the emotional intensity information. In determining the emotional intensity information, the broadcast moment is introduced, and the emotional intensity at each broadcast moment allows the emotional intensity to change over time, resulting in natural emotional fluctuations in the digital human's expressions during the broadcast, thus enhancing the realism of the digital human.
[0076] In the above embodiments, the emotion intensity information is an emotion intensity time curve. The generation process of the emotion intensity time curve can be discussed in conjunction with... Figure 3 To understand further, Figure 3 This is a flowchart illustrating another digital human expression-driven method provided in an embodiment of this application, as shown below. Figure 3 As shown, the digital human expression-driving method of this application embodiment includes, but is not limited to, the following steps: S301, within the first time interval between the broadcast time of the first word and the broadcast time of the sentiment peak word, starting from the sentiment intensity corresponding to the first word, the intensity increases to the sentiment intensity corresponding to the sentiment peak word, so as to generate a sentiment intensity time curve segment within the first time interval.
[0077] S302, within the second time interval between the broadcast time of the sentiment peak word and the broadcast time of the last word, starting from the sentiment intensity corresponding to the sentiment peak word, decays to the sentiment intensity corresponding to the last word, so as to generate a sentiment intensity time curve segment within the second time interval.
[0078] In some embodiments, an emotional intensity time curve segment can be generated with the broadcast time as the horizontal axis and the emotional intensity as the vertical axis. Optionally, the emotional intensity time curve can be generated by determining the coordinate points corresponding to the emotional intensity at each broadcast time and connecting each coordinate point.
[0079] In some embodiments, to make the generated emotional intensity time curve more accurate, when generating the emotional intensity time curve, a first time interval of emotional intensity enhancement and a second time interval of emotional intensity decay can be divided, and emotional intensity time curve segments in the first time interval and the second time interval can be generated respectively, so as to determine the emotional intensity time curve based on the emotional intensity time curve segments.
[0080] In some embodiments, time intervals can be divided based on the broadcast times of different words. The broadcast time of the first word marks the beginning of the first time interval, and the broadcast time of the last word marks the end of the second time interval. Since the emotional peak word corresponds to the highest emotional intensity, the broadcast time of the emotional peak word can be used as the end of the first time interval and the beginning of the second time interval.
[0081] In other words, the first time interval is from the time the first word is broadcast to the time the sentiment peak word is broadcast; the second time interval is from the time the sentiment peak word is broadcast to the time the last word is broadcast.
[0082] Alternatively, a smoothing algorithm can be used to smooth the curve connecting the coordinate points to obtain the emotional intensity time curve.
[0083] In some embodiments, starting from the emotional intensity corresponding to the first word, the emotional intensity at each broadcast moment can be smoothly increased according to a Bézier curve or linear interpolation to generate an emotional intensity time curve segment within a first time interval.
[0084] In some embodiments, starting from the emotional intensity corresponding to the emotional peak word, the emotional intensity can be smoothly decayed according to a Bézier curve or a Gaussian distribution function to generate an emotional intensity time curve segment within a second time interval.
[0085] In some embodiments, when smoothing the attenuation of emotional intensity according to a Bezier curve or a Gaussian distribution function, the attenuation slope can be determined and used to smooth the attenuation.
[0086] For example, if a sentence has ended, the decay rate drops rapidly; if a sentence has not ended, the decay rate drops slowly. In other words, the decay rate of a sentence that has ended is greater than the decay rate of a sentence that has not ended.
[0087] In some embodiments, the decay slope can be determined based on the remaining text quantity and syntactic structure. The decay slope corresponding to a Bézier curve or Gaussian distribution function is determined by determining the remaining text quantity from the word following the sentiment peak to the last word, and based on the remaining text quantity and the syntactic structure of the broadcast text.
[0088] In some embodiments, the remaining text quantity can be the number of remaining words, and the syntactic structure of the broadcast text can be determined based on the grammatical features of the broadcast text. For example, grammatical features include transition words, "only this," and end markers.
[0089] Optionally, a decay adjustment coefficient can be assigned to each grammatical feature, and the decay slope can be calculated based on the decay adjustment coefficient and the amount of remaining text.
[0090] Figure 4 This is a schematic diagram of the emotional intensity time curve provided in an embodiment of this application. Figure 4 It can be seen that the emotional intensity gradually increases before the broadcast of the emotional peak word, and gradually decreases after the broadcast of the emotional peak word.
[0091] In some embodiments, after obtaining the emotion intensity time curve, the digital human's expressions can be synchronized based on the emotion intensity time curve. Expression-driven data can be generated based on emotion tags and the emotion intensity time curve, and the digital human's expressions can be synchronized using the expression-driven data.
[0092] In some embodiments, the expression that the digital human needs to use when broadcasting text can be determined first based on emotion tags, and expression-driven data for that expression can be generated based on the emotion intensity time curve.
[0093] In some embodiments, the emotional intensity corresponding to a word can be determined from the emotional intensity time curve based on the broadcast time of each word in the broadcast text. That is, based on the broadcast time of each word in the broadcast text, the emotional intensity corresponding to that broadcast time is found from the emotional intensity time curve, thereby determining the expression-driven data corresponding to the word based on the emotional intensity of the word.
[0094] Optionally, the facial driving unit of the digital human can be determined based on the emotional intensity at each broadcast moment, thereby determining the expression driving data corresponding to each word based on the facial driving unit.
[0095] For example, each emotion intensity corresponds to a different facial driving unit. Therefore, after determining the emotion intensity of each word in the broadcast text, the facial driving unit for each word can be determined based on this correspondence. Facial driving units include driving units for different emotion labels; for example, if the emotion label is "happy," facial driving units include upturned corners of the mouth and outward-facing corners of the mouth.
[0096] Furthermore, during the process of the digital human broadcasting text, the facial expression-driven data corresponding to the broadcast words can be used simultaneously to drive the digital human's facial expressions. For example, facial expression-driven data can be used to render the digital human's facial expressions, thus driving the digital human's facial expressions.
[0097] In the digital human expression driving method provided in this application embodiment, by drawing the emotional intensity time curve segments of the first time interval and the second time interval respectively, and generating an emotional intensity curve based on the emotional intensity time curve segments, the expression of the digital human is synchronously driven based on the emotional intensity curve, so that the expression of the digital human during the broadcast has natural emotional fluctuations.
[0098] Corresponding to the digital human expression driving methods proposed in the above embodiments, an embodiment of this application also proposes a digital human expression driving device. Since the digital human expression driving device proposed in this application corresponds to the digital human expression driving methods proposed in the above embodiments, the implementation methods of the above digital human expression driving methods are also applicable to the digital human expression driving device proposed in this application, and will not be described in detail in the following embodiments.
[0099] Figure 5 This is a schematic diagram of the structure of a digital human expression driving device provided in an embodiment of this application.
[0100] like Figure 5 As shown, the digital human's expression-driving device 500 includes: The acquisition module 501 is used to acquire the text to be read by the digital human. The sentiment analysis module 502 is used to perform sentiment analysis on the broadcast text and obtain the sentiment tags and sentiment intensity information of the broadcast text. The sentiment intensity information includes the sentiment intensity of each broadcast moment covered by the sentiment tags. The expression driving module 503 is used to synchronously drive the expression of the digital human based on the emotion tag and the emotion intensity information during the process of the digital human broadcasting the broadcast text.
[0101] In one possible implementation of this application embodiment, the sentiment analysis module 502 is further configured to: perform sentiment analysis on the broadcast text to obtain the sentiment tags, sentiment trigger words and sentiment peak words of the broadcast text; and determine the sentiment intensity of each broadcast moment covered by the sentiment tags based on the sentiment trigger words and sentiment peak words.
[0102] In one possible implementation of this application embodiment, the sentiment analysis module 502 is further configured to: obtain the broadcast time and sentiment intensity of the first word and the last word in the broadcast text; obtain the broadcast time and sentiment intensity of the sentiment trigger word and the sentiment peak word; determine the sentiment intensity of each broadcast time covered by the sentiment tag based on the broadcast time and sentiment intensity of the first word and the last word, as well as the broadcast time and sentiment intensity of the sentiment trigger word and the sentiment peak word; wherein the sentiment intensity of the sentiment peak word is greater than the sentiment intensity of the sentiment trigger word, and the sentiment intensity of the sentiment trigger word is greater than the sentiment intensity of the first word and the last word.
[0103] In one possible implementation of this application embodiment, the emotion intensity information is an emotion intensity time curve. The emotion analysis module 502 is further configured to: in a first time interval between the broadcast time of the first word and the broadcast time of the emotion peak word, increase the emotion intensity from the emotion intensity corresponding to the first word to the emotion intensity corresponding to the emotion peak word to generate an emotion intensity time curve segment in the first time interval; and in a second time interval between the broadcast time of the emotion peak word and the broadcast time of the last word, decrease the emotion intensity from the emotion intensity corresponding to the emotion peak word to the emotion intensity corresponding to the last word to generate an emotion intensity time curve segment in the second time interval.
[0104] In one possible implementation of this application embodiment, the sentiment analysis module 502 is further configured to: starting from the sentiment intensity corresponding to the sentiment peak word, smoothly decay the sentiment intensity according to a Bezier curve or a Gaussian distribution function to generate a sentiment intensity time curve segment within a second time interval.
[0105] In one possible implementation of this application embodiment, the sentiment analysis module 502 is further configured to: determine the number of remaining texts from the next word after the sentiment peak word to the last word; and determine the decay slope corresponding to the Bezier curve or Gaussian distribution function based on the number of remaining texts and the syntactic structure of the broadcast text.
[0106] In one possible implementation of this application embodiment, the sentiment analysis module 502 is further configured to: perform sentiment analysis on the broadcast text and obtain the broadcast time of each word in the broadcast text on the speech time axis of the text-to-speech (TTS).
[0107] In one possible implementation of this application embodiment, the emotion intensity information is an emotion intensity time curve, and the expression driving module 503 is further configured to: determine the expression that the digital human needs to use when broadcasting the text based on the emotion tag; determine the emotion intensity corresponding to each word from the emotion intensity time curve based on the broadcast time of each word in the broadcast text; determine the expression driving data corresponding to the word based on the emotion intensity of the word; and simultaneously drive the expression of the digital human using the expression driving data corresponding to the broadcast words during the process of the digital human broadcasting the text.
[0108] The digital human expression-driving device provided in this application acquires the text to be read by the digital human and performs sentiment analysis on the text to obtain its sentiment tags and sentiment intensity information. During the digital human's reading of the text, its facial expressions are synchronously driven based on the sentiment tags and sentiment intensity information. Therefore, by introducing time-related sentiment intensity information, this solution can solve the problem of stiff digital human expressions, enabling the digital human's expressions to have natural emotional fluctuations during the reading process, thus improving the realism and emotional credibility of the digital human.
[0109] It should be noted that the foregoing explanation of the digital human expression driving method embodiment also applies to the digital human expression driving device of this embodiment, and will not be repeated here.
[0110] To implement the above embodiments, this application also proposes an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments.
[0111] To implement the above embodiments, this application also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the foregoing embodiments.
[0112] To implement the above embodiments, this application also proposes a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the foregoing embodiments.
[0113] The collection, storage, use, processing, transmission, provision, and application of user personal information involved in this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0114] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.
[0115] This application is intended to provide an implementation scheme for users to selectively prevent the use or access to their personal information data. Specifically, this application is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information is de-identified to protect user privacy.
[0116] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0117] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0118] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0119] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0120] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any of the following techniques known in the art, or a combination thereof: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0121] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0122] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0123] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A digital human expression driving method, characterized by, The method further includes: Obtain the text to be read by the digital human; Sentiment analysis is performed on the broadcast text to obtain the sentiment tags and sentiment intensity information of the broadcast text, wherein the sentiment intensity information includes the sentiment intensity of each broadcast moment covered by the sentiment tags; During the process of the digital human broadcasting the text, the digital human's facial expressions are synchronously driven based on the emotion tags and the emotion intensity information.
2. The method of claim 1, wherein, The step of performing sentiment analysis on the broadcast text to obtain the sentiment tag and sentiment intensity information of the broadcast text includes: Sentiment analysis is performed on the broadcast text to obtain the sentiment tags, sentiment trigger words, and sentiment peak words of the broadcast text; Based on the emotional trigger words and emotional peak words, the emotional intensity of each broadcast moment covered by the emotional tag is determined.
3. The method of claim 2, wherein, The step of determining the emotional intensity of each broadcast moment covered by the emotional tag based on the emotional trigger words and emotional peak words includes: Obtain the broadcast time and emotional intensity of the first and last words in the broadcast text; Obtain the broadcast time and emotional intensity of the emotional trigger words and the emotional peak words respectively; Based on the broadcast time and emotional intensity of the first and last words, as well as the broadcast time and emotional intensity of the emotional trigger word and the emotional peak word, the emotional intensity of each broadcast time covered by the emotional tag is determined; wherein, the emotional intensity of the emotional peak word is greater than the emotional intensity of the emotional trigger word, and the emotional intensity of the emotional trigger word is greater than the emotional intensity of the first and last words.
4. The method of claim 3, wherein, The emotional intensity information is an emotional intensity time curve. The determination of the emotional intensity of each broadcast time covered by the emotional tag, based on the broadcast time and emotional intensity of the first and last words, as well as the broadcast time and emotional intensity of the emotional trigger word and the emotional peak word, includes: Within a first time interval between the broadcast time of the first word and the broadcast time of the emotional peak word, the emotional intensity is increased from the emotional intensity corresponding to the first word to the emotional intensity corresponding to the emotional peak word to generate an emotional intensity time curve segment within the first time interval; Within a second time interval between the broadcast time of the emotional peak word and the broadcast time of the last word, the emotional intensity is reduced from the emotional intensity corresponding to the emotional peak word to the emotional intensity corresponding to the last word to generate an emotional intensity time curve segment within the second time interval.
5. The method of claim 4, wherein, The step of generating an emotional intensity time curve segment within the second time interval, starting from the emotional intensity corresponding to the emotional peak word and decreasing to the emotional intensity corresponding to the last word, includes: Starting from the emotional intensity corresponding to the emotional peak word, the emotional intensity is smoothly decayed according to a Bezier curve or Gaussian distribution function to generate an emotional intensity time curve segment within the second time interval.
6. The method of claim 5, wherein, The method further includes: Determine the amount of text remaining from the next word after the sentiment peak word to the last word; Based on the remaining text quantity and the syntactic structure of the broadcast text, determine the decay slope corresponding to the Bézier curve or Gaussian distribution function.
7. The method according to any one of claims 1 to 6, characterized in that, The method further includes: Sentiment analysis is performed on the broadcast text to obtain the broadcast time of each word in the broadcast text on the speech timeline of the text-to-speech (TTS) process.
8. The method according to claim 7, characterized in that, The emotion intensity information is an emotion intensity time curve. During the process of the digital human broadcasting the text, the digital human's facial expressions are synchronously driven based on the emotion tag and the emotion intensity information, including: Based on the emotion tags, determine the facial expressions that the digital human needs to use when broadcasting the text; Based on the broadcast time of each word in the broadcast text, the emotional intensity corresponding to the word is determined from the emotional intensity time curve; Based on the emotional intensity of the words, determine the expression-driven data corresponding to the words; During the process of the digital human broadcasting the text, the facial expression driving data corresponding to the broadcast words is used simultaneously to drive the digital human's facial expressions.
9. A digital human facial expression driving device, characterized in that, The device further includes: The acquisition module is used to acquire the text that the digital human needs to read aloud; The sentiment analysis module is used to perform sentiment analysis on the broadcast text, and obtain the sentiment tags and sentiment intensity information of the broadcast text. The sentiment intensity information includes the sentiment intensity of each broadcast moment covered by the sentiment tags. An expression-driven module is used to synchronously drive the digital human's expressions based on the emotion tags and the emotion intensity information during the process of the digital human broadcasting the text.
10. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-8.
12. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1-8.