Methods, devices, and electronic devices for controlling digital human facial expressions based on interjections
By extracting and correcting the interjections in the digital human's broadcast text, target facial expression-driven data is generated, solving the problem of lack of facial expression variation during digital human broadcasts and achieving more dynamic and realistic digital human interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING YINGZHI TECH CO LTD
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-31
AI Technical Summary
The lack of facial expressions when digital humans pronounce interjections results in a stiff and unresponsive performance.
By extracting interjections from the broadcast text, the candidate expression-driven data is corrected using these interjections to generate target expression-driven data. This data is then used to synchronously drive the expressions during the digital human's broadcast, thereby capturing and combining micro-expressions.
It enhances the dynamism and realism of the digital human's broadcasts, and strengthens the interactive experience between the digital human and the user.
Smart Images

Figure CN122491246A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of intelligent cockpit, in-vehicle intelligent agent, digital human interaction, scene arrangement, etc., and in particular to a digital human expression control method, device and electronic device based on interjections. Background Technology
[0002] In everyday conversation, interjections carry rich emotional information, and different subtle facial expressions appear when interjections are used. Currently, the existence of interjections is completely ignored when driving digital humans, and they are treated the same as ordinary text without producing any special facial expressions. As a result, when digital humans use interjections, their expressions remain unchanged, making them appear stiff and lacking in dynamism. Summary of the Invention
[0003] The purpose of this application is to at least partially solve one of the technical problems in the related art.
[0004] Therefore, the first objective of this application is to propose a digital human facial expression control method based on interjections, so as to capture the micro-expressions carried in the broadcast text during the digital human's broadcasting process, and combine the micro-expressions to drive the digital human, thereby improving the agility and realism of the digital human's broadcasting.
[0005] The second objective of this application is to propose a digital human facial expression control device based on interjections.
[0006] The third objective of this application is to propose an electronic device.
[0007] The fourth objective of this application is to provide a computer-readable storage medium.
[0008] The fifth objective of this application is to provide a computer program product.
[0009] To achieve the above objectives, a first aspect of this application proposes a digital human facial expression control method based on interjections, comprising: extracting interjections from the broadcast text of the digital human to obtain interjections in the broadcast text; modifying candidate facial expression driving data of the broadcast text based on the interjections to obtain target facial expression driving data of the broadcast text; and synchronously driving the facial expressions of the digital human based on the target facial expression driving data during the broadcast of the broadcast text by the digital human.
[0010] To achieve the above objectives, a second aspect of this application proposes a digital human facial expression control device based on interjections, comprising: an extraction module for extracting interjections from the broadcast text of the digital human to obtain interjections in the broadcast text; a correction module for correcting candidate facial expression driving data of the broadcast text based on the interjections to obtain target facial expression driving data of the broadcast text; and a driving module for synchronously driving the facial expressions of the digital human based on the target facial expression driving data during the broadcast of the broadcast text by the digital human.
[0011] To achieve the above objectives, a third aspect of this application provides an electronic device, comprising: a processor; and a memory communicatively connected to the processor; the memory storing computer-executable instructions; and the processor executing the computer-executable instructions stored in the memory to enable the processor to execute the digital human facial expression control method based on interjections described in the first aspect of the application.
[0012] To achieve the above objectives, a fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, the computer instructions being used to cause the computer to execute the digital human facial expression control method based on interjections described in the above aspect of the embodiment.
[0013] To achieve the above objectives, a fifth aspect of this application provides a computer program product, including a computer program that, when executed by a processor, implements the digital human facial expression control method based on interjections described in the above aspect of the embodiment.
[0014] The digital human facial expression control method, apparatus, and electronic device provided in this application extract interjections from the digital human's broadcast text to obtain the interjections in the broadcast text. Based on these interjections, candidate facial expression driving data for the broadcast text is corrected to obtain target facial expression driving data for the broadcast text. This allows for synchronized facial expression control of the digital human based on the target facial expression driving data during the broadcast process. Therefore, this solution can capture micro-expressions carried in the broadcast text and use these micro-expressions to drive the digital human, thereby improving the liveliness and realism of the digital human's broadcast.
[0015] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0016] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:
[0017] Figure 1 A flowchart illustrating a digital human facial expression control method based on interjections provided in this application embodiment; Figure 2 A flowchart illustrating another digital human facial expression control method based on interjections provided in this application embodiment; Figure 3 A flowchart illustrating another digital human facial expression control method based on interjections provided in this application embodiment; Figure 4 This is a schematic diagram of the final emotional intensity time curve of the broadcast text provided in the embodiments of this application; Figure 5 This is a schematic diagram of the structure of a digital human facial expression control device based on interjections, provided in an embodiment of this application. Detailed Implementation
[0018] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0019] The following describes a digital human facial expression control method and apparatus based on interjections according to embodiments of this application, with reference to the accompanying drawings.
[0020] Figure 1 This is a flowchart illustrating a digital human facial expression control method based on interjections provided in an embodiment of this application, as shown below. Figure 1 As shown, the digital human facial expression control method based on interjections in this application includes, but is not limited to, the following steps: S101, extract the modal particles from the digital human's broadcast text to obtain the modal particles in the broadcast text.
[0021] In some embodiments, the voice words can be extracted from the broadcast text that the digital human needs to broadcast, thereby obtaining the voice words in the broadcast text.
[0022] It should be noted that the digital human in this application embodiment can be applied to scenarios requiring user responses, such as intelligent cockpits in vehicles, customer service scenarios, and consultation scenarios. In other words, the digital human can respond to user questions or requests, and the text it broadcasts can be the response text corresponding to the question or request.
[0023] In some embodiments, by receiving user input information and combining it with a large language model to parse and respond to the input information, a response text corresponding to the input information is generated as the broadcast text that the digital human needs to broadcast.
[0024] It should be noted that the digital human in the embodiments of the present application can also be applied to other scenarios where there is no need to respond to the user, such as information display scenarios, education and popular science scenarios, etc. That is to say, the digital human can broadcast the received information, and the broadcast text that the digital human needs to broadcast can be the input information of the user.
[0025] In some embodiments, by receiving the input information of the user and using the input information as the broadcast text that the digital human needs to broadcast.
[0026] In the embodiments of the present application, the application scenarios of the digital human are not specifically limited.
[0027] In some embodiments, the modal particles in the broadcast text can be identified and located to extract the modal particles from the broadcast text. Optionally, a modal particle library containing different modal particles can be established in advance, and the words in the broadcast text can be extracted. By comparing the extracted words with the modal particles in the modal particle library, the modal particles can be determined from the extracted words.
[0028] In some embodiments, the similarity between the words and the modal particles can be calculated, and the words with a similarity greater than the set similarity threshold are used as the modal particles extracted from the broadcast text.
[0029] For example, the modal particle library includes modal particles such as "wow", "alas", "oh", "ah", "um", "ya", "ei", "ha", "hei", etc.
[0030] If words 1 and words 2 are extracted from the broadcast text, the similarity between word 1 and each of the above modal particles is calculated, and the similarity between word 2 and each of the above modal particles is calculated. When the similarity between word 1 and "wow" is greater than the similarity threshold, it can be determined that word 1 is the modal particle "wow".
[0031] S102. Based on the modal particles, the candidate expression driving data of the broadcast text is corrected to obtain the target expression driving data of the broadcast text.
[0032] In some embodiments, the candidate expression driving data of the broadcast text can be determined based on the emotional intensity information of the broadcast text. Natural Language Processing (NLP) technology can be used to perform emotional analysis on the broadcast text to obtain the emotional intensity information of the broadcast text.
[0033] Among them, the emotional intensity information includes the emotional intensity of each broadcast moment covered by the emotional label. The emotional label of the broadcast text can be determined from the broadcast text, and the emotional intensity and broadcast moment corresponding to the emotional label can be generated, so as to obtain the emotional intensity information.
[0034] In some embodiments, emotion intensity information can be used to obtain changes in facial expressions of the digital human during the reading of text, and candidate expression-driven data can be generated based on these changes. In some embodiments, the emotional intensity corresponding to a word can be determined from the emotional intensity information based on the broadcast time of each word in the broadcast text. That is, based on the broadcast time of each word in the broadcast text, the emotional intensity corresponding to that broadcast time is searched from the emotional intensity information, thereby determining the expression-driven data corresponding to the word based on the emotional intensity of the word, and determining the candidate expression-driven data of the broadcast text based on the expression-driven data corresponding to each word.
[0035] It should be noted that since interjections carry different emotional information, each interjection corresponds to a different micro-expression. Therefore, the micro-expression driving data of the interjection can be used as a basis to drive the data, and the candidate expression driving data of the broadcast text can be corrected based on the micro-expression driving data to obtain the target expression driving data of the broadcast text.
[0036] In some embodiments, since the candidate expression-driven data contains expression-driven data for different words, if there is a word that is a modal particle, the micro-expression-driven data of the modal particle can be used to replace the expression-driven data, thereby correcting the candidate expression-driven data and obtaining the target expression-driven data for the broadcast text.
[0037] Continuing with the example above, word 1 and word 2 are extracted from the broadcast text, where word 1 is an interjection. The candidate expression-driven data of the broadcast text includes the expression-driven data of word 1 and word 2. Therefore, the micro-expression-driven data of the interjection can be used to replace the expression-driven data of word 1 to obtain the target expression-driven data, which includes the micro-expression-driven data and the expression-driven data of word 2.
[0038] In some embodiments, micro-expression-driven data for interjections can be generated based on the micro-expression features corresponding to the interjections.
[0039] By defining a corresponding micro-expression feature combination for each interjection, and generating corresponding micro-expression-driven data through these feature combinations, the micro-expression feature combination comprises multiple micro-expression features.
[0040] In other words, a single interjection can correspond to multiple micro-expression features. By weighted fusion of each micro-expression feature, we can obtain the micro-expression-driven data corresponding to the interjection.
[0041] S103, during the process of the digital human broadcasting the text, the digital human's facial expressions are synchronously driven based on the target facial expression driving data.
[0042] In some embodiments, target facial expression driving data can be used to render the digital human to achieve synchronized facial expression control. During the driving process, facial expressions and spoken audio can be aligned based on the broadcast time to improve the accuracy and effectiveness of the digital human's broadcast.
[0043] For example, target facial expression data can be input into the digital human's driving engine to drive the digital human's facial expressions to synchronize.
[0044] The digital human facial expression control method based on interjections provided in this application extracts interjections from the digital human's broadcast text, obtains the interjections in the broadcast text, and corrects the candidate facial expression driving data of the broadcast text based on the interjections to obtain the target facial expression driving data of the broadcast text. Thus, during the digital human's broadcast of the text, its facial expressions can be synchronously driven based on the target facial expression driving data. Therefore, this solution can capture micro-expressions carried in the broadcast text during the digital human's broadcasting process and drive the digital human based on these micro-expressions, thereby improving the agility and realism of the digital human's broadcast.
[0045] Figure 2 This is a flowchart illustrating another digital human facial expression control method based on interjections provided in an embodiment of this application, as shown below. Figure 2 As shown, the digital human facial expression control method based on interjections in this application includes, but is not limited to, the following steps: S201, extract the modal particles from the digital human's broadcast text to obtain the modal particles in the broadcast text.
[0046] In the embodiments of this application, step S201 can be implemented in any of the ways described in the embodiments of this application. This is not limited here and will not be described in detail.
[0047] S202, Determine the micro-expression-driven data of interjections.
[0048] In some embodiments, micro-expression feature combinations corresponding to each interjection can be established, and micro-expression driven data of the interjection can be generated based on each micro-expression feature in the micro-expression feature combination.
[0049] For example, the micro-expression features corresponding to the interjection "wow" include: raising eyebrows, widening eyes, and slightly opening the mouth. The micro-expression features corresponding to the interjection "sigh" include: slightly furrowed brows, downturned corners of the mouth, and slight sighing movements. The micro-expression features corresponding to the interjection "oh" include: slight nodding, raising one eyebrow, etc. The micro-expression feature combinations corresponding to the modal particle "ah" include micro-expression features such as an open mouth and widened eyes; The micro-expression feature combinations corresponding to the modal particle "um" include micro-expression features such as a slight nod and slightly closed lips.
[0050] By performing weighted fusion on the micro-expression features corresponding to the modal particle, micro-expression driving data for the modal particle can be generated. That is, different weight values are set for the micro-expression features corresponding to the modal particle, and weighted fusion is performed based on the weight values.
[0051] S203. Based on the micro-expression driving data of the modal particle, correct the candidate expression driving data to obtain the target expression driving data for the broadcast text.
[0052] In some embodiments, the micro-expression driving data of the modal particle can be added to the candidate expression driving data to correct the candidate expression driving data, thereby obtaining the corrected expression driving data as the target expression driving data for the broadcast text.
[0053] In some embodiments, based on the broadcast time of the modal particle and based on the broadcast time, add the micro-expression driving data of the modal particle to the candidate expression driving data.
[0054] In some embodiments, Text-to-Speech (TTS) can be used to convert the broadcast text into the voice that the digital human needs to broadcast. Then, based on the voice that the digital human needs to broadcast, determine the broadcast time when the modal particle in the broadcast text needs to be broadcast.
[0055] In some embodiments, after determining the broadcast time of the modal particle, based on the micro-expression driving data, under the constraint of the broadcast time of the modal particle, correct the candidate expression driving data to obtain the target expression driving data for the broadcast text.
[0056] That is, based on the broadcast time of the modal particle, determine the position where the micro-expression driving data in the candidate expression driving data needs to be added, and add the micro-expression driving data to this position, so that the candidate expression driving data can include the micro-expression driving data, thereby obtaining the target expression driving data.
[0057] In some embodiments, under the constraint of the broadcast time of the modal particle, that is, perform time alignment on the micro-expression driving data and the candidate expression driving data, to avoid the micro-expression driving data being at the wrong time, resulting in inconsistent emotions corresponding to the micro-expression and the expression of the digital human.
[0058] S204. During the process of the digital human broadcasting the broadcast text, synchronously drive the expression of the digital human based on the target expression driving data.
[0059] In the embodiments of this application, step S204 can be implemented in any of the ways described in the various embodiments of this application. This is not limited here and will not be described in detail.
[0060] In the above embodiments, the process of determining candidate expression-driven data can be combined with... Figure 3 To understand further, Figure 3 This is a flowchart illustrating another digital human facial expression control method based on interjections provided in an embodiment of this application, as shown below. Figure 3 As shown, the digital human facial expression control method based on interjections in this application includes, but is not limited to, the following steps: S301, determine the emotional intensity information of the broadcast text, which includes the emotional label of the broadcast text and the emotional intensity of each broadcast moment covered by the emotional label.
[0061] In some embodiments, sentiment analysis can be performed on the broadcast text to obtain the sentiment tags, sentiment trigger words, and sentiment peak words of the broadcast text, and the sentiment intensity of each broadcast moment covered by the sentiment tags can be determined based on the sentiment trigger words and sentiment peak words.
[0062] Furthermore, based on the emotional intensity of each broadcast moment covered by the emotional tags, an emotional intensity time curve of the broadcast text can be generated. In other words, the emotional intensity information includes at least the emotional intensity time curve.
[0063] It should be noted that sentiment tags refer to the sentiment classification of the broadcast text; sentiment trigger words refer to words that directly express emotions, that is, which words give the broadcast text emotional color; sentiment peak words refer to words with the highest sentiment intensity, that is, words that best reflect the intensity of emotions, such as "This dish is amazing!" which is a sentiment peak word.
[0064] In some embodiments, semantic understanding and feature extraction are performed on the broadcast text to obtain the emotional information carried in the broadcast text, and emotional tags are extracted from the broadcast text based on the emotional information to obtain the emotional tags of the broadcast text.
[0065] Optionally, a first word library containing emotional trigger words can be pre-set, and different words can be extracted from the broadcast text. If the extracted words are the same as the words in the first word library, the words can be identified as emotional trigger words.
[0066] Optionally, if none of the words in the broadcast text match the words in the first word database, the similarity between each word and the words in the first word database can be calculated, thereby filtering out emotion trigger words from the broadcast text based on the similarity.
[0067] In some embodiments, a second word library containing sentiment peak words can be predefined, and different words can be extracted from the broadcast text. If the extracted words are the same as the words in the second word library, the words can be identified as sentiment peak words.
[0068] Optionally, if none of the words in the broadcast text match the words in the second word database, the similarity between each word and the words in the second word database can be calculated, thereby filtering out the sentiment peak words from the broadcast text based on the similarity.
[0069] In some embodiments, after obtaining the emotional trigger words and emotional peak words, the emotional intensity of each broadcast moment covered by the emotional tag can be determined based on the emotional trigger words and emotional peak words, that is, the emotional intensity information can be determined.
[0070] In this embodiment, the emotional intensity information can be an emotional intensity time curve, which can be generated based on the emotional intensity of each word at each broadcast time. The horizontal axis of the emotional intensity time curve represents the broadcast time, and the vertical axis represents the emotional intensity.
[0071] S302, Based on the emotional intensity information, determine the candidate facial expression driving data for the broadcast text.
[0072] In some embodiments, the emotion intensity information is an emotion intensity time curve. The emotion intensity corresponding to a word can be determined from the emotion intensity time curve based on the broadcast time of each word in the broadcast text. That is, based on the broadcast time of each word in the broadcast text, the emotion intensity corresponding to that broadcast time is found from the emotion intensity time curve, thereby determining the expression-driven data corresponding to the word based on the emotion intensity of the word.
[0073] Optionally, the facial driving unit of the digital human can be determined based on the emotional intensity at each broadcast moment, thereby determining the expression driving data corresponding to each word based on the facial driving unit, and fusing the expression driving data corresponding to each word according to the broadcast moment to obtain candidate expression driving data for the broadcast text.
[0074] The digital human facial expression control method based on interjections provided in this application determines the emotional intensity information of the broadcast text and, based on the emotional intensity information, determines candidate facial expression driving data for the broadcast text. Introducing emotional intensity information during the determination of candidate facial expression driving data allows the emotional intensity to change over time, resulting in natural emotional fluctuations in the digital human's facial expressions during broadcasting, thus enhancing the realism of the digital human.
[0075] Based on any of the above embodiments, after generating the emotional intensity time curve of the broadcast text, an intensity pulse curve corresponding to each interjection can also be generated, and the intensity pulse curve is superimposed with the emotional intensity time curve to obtain the final emotional intensity time curve of the broadcast text, thereby realizing the fusion of the emotional intensity of micro-expressions and digital human expressions.
[0076] Figure 4 This is a schematic diagram of the final emotional intensity time curve of the broadcast text provided in the embodiments of this application.
[0077] In some embodiments, by generating an intensity pulse curve corresponding to each interjection, the emotional information of the interjection can be captured, thereby enabling the digital human to have corresponding micro-expression reactions when broadcasting interjections, thus improving the agility of the digital human's broadcast.
[0078] In some embodiments, during the sentiment analysis of the broadcast text, the broadcast time of the interjection can be determined, thereby determining the intensity pulse curve of the interjection based on the interjection and the corresponding broadcast time.
[0079] Optionally, the intensity pulse curve of the modal particle can be generated based on the emotional intensity of the modal particle and the corresponding broadcast time.
[0080] Among them, the intensity pulse curve of the modal particle can be dynamically adjusted according to the pronunciation duration of the modal particle. In other words, when generating the intensity pulse curve of the modal particle, the duration of the intensity pulse curve of the modal particle can be predetermined, and the intensity pulse curve of the modal particle can be generated according to the emotional intensity of the modal particle and the corresponding broadcast time.
[0081] In some embodiments, an intensity pulse curve can be superimposed on the emotion intensity time curve based on the broadcast time of the interjection to obtain the final emotion intensity time curve of the broadcast text. Optionally, different weights can be assigned to the intensity pulse curve and the emotion intensity time curve during the superposition process for weighted fusion.
[0082] The sum of the weights of the intensity pulse curve and the emotion intensity time curve should be less than or equal to 1, so as to avoid the expression of the superimposed emotion intensity time curve being too exaggerated.
[0083] In some embodiments, during the process of generating the intensity pulse curve of the modal particle, the reference emotional intensity of the modal particle at the time of its broadcast can also be determined, and the intensity pulse curve of the modal particle can be determined based on the broadcast time of the modal particle and the reference emotional intensity, wherein the period of the intensity pulse curve is less than or equal to a set duration.
[0084] In some embodiments, the emotional information carried by the interjection can be extracted, the emotional intensity level corresponding to the interjection can be determined, and the emotional information can be quantified according to the emotional intensity level and the corresponding broadcast time to obtain the reference emotional intensity of the interjection at the broadcast time.
[0085] For example, the emotional intensity level can be corrected based on the broadcast time of the interjection, and the corrected emotional intensity level can be used to quantify the emotional information to obtain the reference emotional intensity of the interjection at its broadcast time.
[0086] In some embodiments, the duration of the intensity pulse curve can be determined based on the pronunciation duration of different interjections, thereby generating the intensity pulse curve of the interjection within the duration.
[0087] Optionally, the duration of the intensity pulse curve can be determined by measuring the pronunciation duration of the interjection and based on the pronunciation duration of the interjection and the broadcast time. For example, the duration of the intensity pulse curve can be defined as the set time range before the broadcast time and the set time range after the broadcast time.
[0088] Furthermore, an intensity pulse curve can be generated over a duration based on a reference emotional intensity. In other words, the reference emotional intensity refers to the maximum emotional intensity corresponding to the modal particle, thus allowing the intensity to rise from any emotional intensity to the reference emotional intensity and fall back to any emotional intensity over the duration, resulting in an intensity pulse curve.
[0089] In some embodiments, any emotional intensity can be a default emotional intensity. By determining the start and end times of the duration, and between the start and broadcast times of the duration, the emotional intensity increases from the default emotional intensity to a reference emotional intensity at a first rate, and then between the broadcast time and the end time of the duration, the emotional intensity decreases from the reference emotional intensity back to the default emotional intensity at a first rate, thus obtaining an intensity pulse curve.
[0090] The first rate can be a set value.
[0091] In some embodiments, when superimposing the emotional intensity time curve and the intensity pulse curve, the superposition can be based on the emotional intensity, that is, the emotional intensity of the intensity pulse curve can be superimposed on the emotional intensity of the emotional intensity time curve to obtain the final emotional intensity time curve.
[0092] It should be noted that the superposition period for superimposing the emotional intensity time curve can be determined in advance, so that the emotional intensity can be superimposed within the superposition period.
[0093] In some embodiments, the superposition period of the emotional intensity time curve and the intensity pulse curve can be determined based on the duration of the intensity pulse curve. For example, the broadcast time of the interjection can be determined from the emotional intensity time curve, and the superposition period of the emotional intensity time curve and the intensity pulse curve can be determined based on the broadcast time and the duration of the intensity pulse curve.
[0094] Furthermore, by weighted fusion of the emotional intensity of the intensity pulse curve and the emotional intensity time curve at the same broadcast time within the superimposed period, the target emotional intensity at the broadcast time within the superimposed period is obtained. Optionally, a first weight is set for the emotional intensity of the intensity pulse curve, a second weight is set for the emotional intensity of the emotional intensity time curve, and the first and second weights are used for weighted fusion.
[0095] The sum of the first weight and the second weight is less than or equal to 1.
[0096] In some embodiments, after obtaining the target emotional intensity at the broadcast time within the superimposed time period, the final emotional intensity time curve can be obtained based on the target emotional intensity at the broadcast time within the superimposed time period and the target emotional intensity at the broadcast time outside the superimposed time period.
[0097] In other words, the target emotional intensity at the broadcast time within the superimposed time period can be spliced with the target emotional intensity at the broadcast time within the non-superimposed time period to obtain the final emotional intensity time curve.
[0098] It should be noted that after determining the superposition period of the emotional intensity time curve and the intensity pulse curve, the target expression driving data can be determined based on the micro-expression driving data and candidate expression driving data within the superposition period.
[0099] In some embodiments, candidate expression-driven data for the same broadcast moment can be weighted and fused based on micro-expression-driven data for each broadcast moment within the superimposed time period to obtain target expression-driven data for the same broadcast moment.
[0100] Corresponding to the digital human expression control methods based on interjections proposed in the above embodiments, an embodiment of this application also proposes a digital human expression control device based on interjections. Since the digital human expression control device based on interjections proposed in this application corresponds to the digital human expression control methods based on interjections proposed in the above embodiments, the implementation methods of the above-mentioned digital human expression control methods based on interjections are also applicable to the digital human expression control device based on interjections proposed in this application, and will not be described in detail in the following embodiments.
[0101] Figure 5 This is a schematic diagram of the structure of a digital human facial expression control device based on interjections, provided in an embodiment of this application.
[0102] like Figure 5 As shown, the digital human facial expression control device 500 based on interjections includes: Extraction module 501 is used to extract interjections from the broadcast text of the digital human to obtain the interjections in the broadcast text; The correction module 502 is used to correct the candidate expression-driven data of the broadcast text based on the interjection, so as to obtain the target expression-driven data of the broadcast text. The driving module 503 is used to synchronously drive the digital human's facial expressions based on the target facial expression driving data during the process of the digital human broadcasting the text.
[0103] In one possible implementation of this application, the correction module 502 is further configured to: determine the micro-expression driving data of the tone word; and based on the micro-expression driving data of the tone word, correct the candidate expression driving data to obtain the target expression driving data of the broadcast text.
[0104] In one possible implementation of this application, the correction module 502 is further configured to: determine the broadcast time of the interjection; and, based on the micro-expression driving data, correct the candidate expression driving data under the constraint of the broadcast time of the interjection to obtain the target expression driving data of the broadcast text.
[0105] In one possible implementation of this application embodiment, the correction module 502 is further configured to: determine the emotional intensity information of the broadcast text, the emotional intensity information including the emotional tag of the broadcast text and the emotional intensity of each broadcast moment covered by the emotional tag; and determine candidate expression-driven data of the broadcast text based on the emotional intensity information.
[0106] In one possible implementation of this application embodiment, the correction module 502 is further configured to: perform sentiment analysis on the broadcast text to obtain sentiment tags, sentiment trigger words, and sentiment peak words of the broadcast text; determine the sentiment intensity of each broadcast moment covered by the sentiment tags based on the sentiment trigger words and sentiment peak words; and generate a sentiment intensity time curve of the broadcast text based on the sentiment intensity of each broadcast moment covered by the sentiment tags, wherein the sentiment intensity information includes at least the sentiment intensity time curve.
[0107] In one possible implementation of this application embodiment, the correction module 502 is further configured to: determine the broadcast time of the interjection during the sentiment analysis of the broadcast text; determine the intensity pulse curve of the interjection based on the interjection and the corresponding broadcast time; and superimpose the intensity pulse curve on the sentiment intensity time curve based on the broadcast time of the interjection to obtain the final sentiment intensity time curve of the broadcast text.
[0108] In one possible implementation of this application embodiment, the correction module 502 is further configured to: determine the reference emotional intensity of the modal word at the time of its broadcast; and determine the intensity pulse curve of the modal word based on the broadcast time of the modal word and the reference emotional intensity, wherein the period of the intensity pulse curve is less than or equal to a set duration.
[0109] In one possible implementation of this application embodiment, the correction module 502 is further configured to: determine the pronunciation duration of the interjection, and determine the duration of the intensity pulse curve based on the pronunciation duration of the interjection and the broadcast time; and generate the intensity pulse curve within the duration based on the reference emotional intensity.
[0110] In one possible implementation of this application, the correction module 502 is further configured to: between the start time of the duration and the broadcast time, increase the emotional intensity from the default emotional intensity to the reference emotional intensity at a first rate; and between the broadcast time and the end time of the duration, decrease the emotional intensity from the reference emotional intensity to the default emotional intensity at the first rate.
[0111] In one possible implementation of this application, the correction module 502 is further configured to: determine the superposition period of the emotional intensity time curve and the intensity pulse curve based on the duration of the intensity pulse curve; perform weighted fusion on the emotional intensities of the intensity pulse curve and the emotional intensity time curve at the same broadcast time within the superposition period to obtain the target emotional intensity at the broadcast time within the superposition period; and obtain the final emotional intensity time curve based on the target emotional intensity at the broadcast time within the superposition period and the target emotional intensity at the broadcast time outside the superposition period.
[0112] In one possible implementation of this application embodiment, the correction module 502 is further configured to: perform weighted fusion of candidate expression-driven data for the same broadcast time based on the micro-expression-driven data for each broadcast time within the superimposed time period, so as to obtain the target expression-driven data for the same broadcast time.
[0113] It should be noted that the foregoing explanation of the embodiment of the digital human expression control method based on interjections also applies to the digital human expression control device based on interjections in this embodiment, and will not be repeated here.
[0114] The digital human facial expression control device based on interjections provided in this application extracts interjections from the digital human's broadcast text, obtains the interjections in the broadcast text, and corrects the candidate facial expression driving data of the broadcast text based on the interjections to obtain the target facial expression driving data of the broadcast text. Thus, during the digital human's broadcast of the text, its facial expressions can be synchronously driven based on the target facial expression driving data. Therefore, this solution can capture micro-expressions carried in the broadcast text during the digital human's broadcasting process and drive the digital human based on these micro-expressions, thereby improving the agility and realism of the digital human's broadcast.
[0115] To implement the above embodiments, this application also proposes an electronic device, including: a processor and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments.
[0116] To implement the above embodiments, this application also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the foregoing embodiments.
[0117] To implement the above embodiments, this application also proposes a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the foregoing embodiments.
[0118] The collection, storage, use, processing, transmission, provision, and application of user personal information involved in this application all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0119] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.
[0120] This application is intended to provide an implementation scheme for users to selectively prevent the use or access to their personal information data. Specifically, this application is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information is de-identified to protect user privacy.
[0121] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0122] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0123] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.
[0124] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0125] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any of the following techniques known in the art, or a combination thereof: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0126] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.
[0127] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0128] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.
Claims
1. A digital human facial expression control method based on interjections, characterized in that, The method includes: Extract interjections from the broadcast text of the digital human to obtain the interjections in the broadcast text; Based on the interjections, the candidate facial expression driving data of the broadcast text is corrected to obtain the target facial expression driving data of the broadcast text. During the process of the digital human broadcasting the text, the digital human's facial expressions are synchronously driven based on the target facial expression driving data.
2. The method according to claim 1, characterized in that, The step of correcting the candidate facial expression-driven data of the broadcast text based on the interjections to obtain the target facial expression-driven data of the broadcast text includes: Determine the micro-expression driving data of the interjections; Based on the micro-expression-driven data of the interjections, the candidate expression-driven data is corrected to obtain the target expression-driven data of the broadcast text.
3. The method according to claim 2, characterized in that, The micro-expression-driven data based on the interjections is used to correct the candidate expression-driven data to obtain the target expression-driven data for the broadcast text, including: Determine the broadcast time of the aforementioned modal particle; Based on the micro-expression driving data, and under the constraint of the broadcast time of the interjection, the candidate expression driving data is modified to obtain the target expression driving data of the broadcast text.
4. The method according to claim 2, characterized in that, Before correcting the candidate expression-driven data based on the micro-expression-driven data of the interjection to obtain the target expression-driven data of the broadcast text, the method further includes: Determine the emotional intensity information of the broadcast text, wherein the emotional intensity information includes the emotional tag of the broadcast text and the emotional intensity of each broadcast moment covered by the emotional tag; Based on the emotional intensity information, candidate facial expression-driven data for the broadcast text is determined.
5. The method according to claim 4, characterized in that, The determination of the emotional intensity information of the broadcast text includes: Sentiment analysis is performed on the broadcast text to obtain the sentiment tags, sentiment trigger words, and sentiment peak words of the broadcast text; Based on the emotional trigger words and emotional peak words, determine the emotional intensity of each broadcast moment covered by the emotional tag; Based on the emotional intensity of each broadcast time covered by the emotional tags, an emotional intensity time curve of the broadcast text is generated, wherein the emotional intensity information includes at least the emotional intensity time curve.
6. The method according to claim 5, characterized in that, After generating the emotional intensity time curve of the broadcast text, the method further includes: During the sentiment analysis of the broadcast text, the broadcast time of the interjection is determined; Based on the modal particle and its corresponding broadcast time, determine the intensity pulse curve of the modal particle; Based on the broadcast time of the interjection, the intensity pulse curve is superimposed on the emotional intensity time curve to obtain the final emotional intensity time curve of the broadcast text.
7. The method according to claim 6, characterized in that, The step of determining the intensity pulse curve of the modal particle based on the modal particle and its corresponding broadcast time includes: Determine the reference emotional intensity of the modal particle at the time of its broadcast; Based on the broadcast time of the interjection and the reference emotional intensity, the intensity pulse curve of the interjection is determined, and the period of the intensity pulse curve is less than or equal to a set duration.
8. The method according to claim 7, characterized in that, The step of determining the intensity pulse curve of the modal particle based on the broadcast time of the modal particle and the reference emotional intensity includes: The duration of the pronunciation of the interjection is determined, and the duration of the intensity pulse curve is determined based on the duration of the pronunciation of the interjection and the broadcast time. Based on the reference emotional intensity, the intensity pulse curve is generated within the duration.
9. The method according to claim 8, characterized in that, The step of generating the intensity pulse curve within the duration based on the reference emotional intensity includes: Between the start time of the duration and the broadcast time, the emotional intensity increases from the default emotional intensity to the reference emotional intensity at a first rate; Between the broadcast time and the end time of the duration, the emotional intensity decreases from the reference emotional intensity to the default emotional intensity at the first rate.
10. The method according to any one of claims 6-9, characterized in that, The broadcast time based on the interjection, by superimposing the intensity pulse curve on the emotion intensity time curve, yields the final emotion intensity time curve of the broadcast text, including: Based on the duration of the intensity pulse curve, the superposition period of the emotion intensity time curve and the intensity pulse curve is determined; The emotional intensity of the intensity pulse curve and the emotional intensity time curve at the same broadcast time within the superimposed time period is weighted and fused to obtain the target emotional intensity at the broadcast time within the superimposed time period; The final emotional intensity time curve is obtained based on the target emotional intensity of the broadcast time within the superimposed time period and the target emotional intensity of the broadcast time outside the superimposed time period.
11. The method according to claim 10, characterized in that, The method further includes: Based on the micro-expression driving data of each broadcast moment within the superimposed time period, the candidate expression driving data of the same broadcast moment are weighted and fused to obtain the target expression driving data of the same broadcast moment.
12. A digital human facial expression control device based on interjections, characterized in that, The device includes: The extraction module is used to extract interjections from the broadcast text of the digital human to obtain the interjections in the broadcast text; The correction module is used to correct the candidate expression-driven data of the broadcast text based on the interjection, so as to obtain the target expression-driven data of the broadcast text. The driving module is used to synchronously drive the digital human's facial expressions based on the target facial expression driving data during the process of the digital human broadcasting the text.
13. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-11.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-11.
15. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1-11.