A method, apparatus, device, storage medium, and program product for determining lip shape.

By dividing the lip-syncing method of 3D digital humans into three time periods and utilizing phoneme-lip mapping tables and interpolation algorithms, the problems of unnatural lip shapes and poor real-time performance in existing technologies are solved, achieving more natural, stable and efficient lip-syncing.

CN118692484BActive Publication Date: 2025-10-31ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410940948.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-12
Publication Date
2025-10-31
Estimated Expiration
2044-07-12

AI Technical Summary

Technical Problem

Existing lip-syncing methods for 3D digital humans suffer from unnatural and unrealistic behavior, especially in noisy audio data and digital human models with different driving standards, where lip-syncing parameters are unstable and have poor real-time performance.

Method used

The digital human lip-sync method is divided into three time periods. The key frame lip shape is determined using a phoneme-lip shape mapping table, and the transition is performed in each phoneme's broadcast time period using an interpolation algorithm. The first time period is used to transition with the lip shape of the previous phoneme, the second time period maintains the key frame lip shape, and the third time period is used to transition with the lip shape of the next phoneme.

Benefits of technology

It improves the naturalness and realism of lip shapes, reduces costs, enhances the stability and real-time performance of lip-shape driving, and adapts to digital human models with different driving standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118692484B_ABST
    Figure CN118692484B_ABST
Patent Text Reader

Abstract

This specification provides a lip-shape determination method. It acquires text data from a digital human and determines the playback time segment for each phoneme in the text data. The playback time segment for each phoneme is divided into three segments. The middle segment can be determined by consulting a phoneme-lip shape mapping table to identify the keyframe lip shape to be maintained. The other two segments can be determined by interpolation between the keyframe lip shape of the phoneme and the keyframe lip shapes before and after it. By dividing the playback time segment for each phoneme into three parts, only the middle segment is used to maintain the lip shape of the current phoneme, while the other segments are used for transitions between the lip shape of the preceding or following phoneme. This results in more realistic and natural lip shape changes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to one or more embodiments in the field of digital human technology, and more particularly to a lip shape determination method, apparatus, device, storage medium, and program product. Background Technology

[0002] A 3D digital human is a virtual 3D image generated through digital technology. With the development of the concept of the "metaverse," 3D digital humans are finding increasingly wider applications.

[0003] 3D digital humans can be driven by computers, which can make their limbs and faces move, making them more realistic and natural. Facial driving of 3D digital humans includes lip-syncing, which requires driving the 3D digital human's lips to correspond with the audio being spoken when the 3D digital human is giving a speech.

[0004] In existing technologies, the lip-driving methods used to drive 3D digital humans still produce unnatural and unrealistic lip shapes. Summary of the Invention

[0005] In view of the above, this specification provides a lip shape determination method, apparatus, device, storage medium, and program product through one or more embodiments.

[0006] According to a first aspect of one or more embodiments of this specification, a lip shape determination method is provided, the method comprising:

[0007] Acquire the text data to be broadcast by the digital human, and determine the broadcast time period of the phoneme corresponding to each text unit in the text data; the text unit is a character and / or a word;

[0008] For each phoneme in the text data, the broadcast time period of that phoneme is divided into a first time period, a second time period including keyframes, and a third time period;

[0009] Based on a preset phoneme-lip shape mapping table, determine the keyframe lip shape corresponding to each phoneme of the text data;

[0010] For each phoneme in the text data, the lip shape for the first time period is determined based on the interpolation of the lip shape of the previous keyframe and the keyframe lip shape of the phoneme. The lip shape for the second time period is set to maintain the keyframe lip shape. The lip shape for the third time period is determined based on the interpolation of the keyframe lip shape of the phoneme and the next keyframe lip shape.

[0011] According to a second aspect of one or more embodiments of this specification, a lip shape determining device is provided, the device comprising:

[0012] The time period determination module is used to acquire the text data to be broadcast by the digital human and determine the broadcast time period of the phoneme corresponding to each text unit in the text data; the text unit is a character and / or a word;

[0013] The time period segmentation module is used to divide the broadcast time period of each phoneme in the text data into a first time period, a second time period including keyframes, and a third time period.

[0014] The keyframe lip shape determination module is used to determine the keyframe lip shape corresponding to each phoneme of the text data according to a preset phoneme-lip shape mapping table.

[0015] The lip shape determination module is used to determine the lip shape of a first time period for each phoneme of the text data based on the interpolation of the lip shape of the previous keyframe and the keyframe lip shape of the phoneme, set the lip shape of the second time period to maintain the keyframe lip shape, and determine the lip shape of a third time period based on the interpolation of the keyframe lip shape of the phoneme and the next keyframe lip shape.

[0016] According to a third aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer instructions thereon, which, when executed by a processor, implement the lip shape determination method as described in the first aspect of the embodiments of this specification.

[0017] According to a fourth aspect of the embodiments of this specification, a computer device is provided, the computer device comprising:

[0018] processor;

[0019] Memory used to store processor-executable instructions;

[0020] The processor implements the lip shape determination method as described in the first aspect of the embodiments of this specification by running executable instructions.

[0021] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including computer program instructions that, when executed by a processor, implement the lip shape determination method as described in the first aspect of the embodiments of this specification.

[0022] This specification provides a lip-shape determination method. It acquires text data from a digital human and determines the playback time segment for each phoneme in the text data. The playback time segment for each phoneme is divided into three segments. The middle segment can be determined by consulting a phoneme-lip shape mapping table to identify the keyframe lip shape to be maintained. The other two segments can be determined by interpolation between the keyframe lip shape of the phoneme and the keyframe lip shapes before and after it. By dividing the playback time segment for each phoneme into three parts, only the middle segment is used to maintain the lip shape of the current phoneme, while the other segments are used for transitions between the lip shape of the preceding or following phoneme. This results in more realistic and natural lip shape changes.

[0023] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description

[0024] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this specification and, together with the description, serve to explain the principles of this specification.

[0025] Figure 1 This is a schematic diagram of a voice-driven method in a related art, as shown in this specification.

[0026] Figure 2 This is a flowchart illustrating a lip shape determination method according to an exemplary embodiment of this specification.

[0027] Figure 3 This is a schematic diagram illustrating the effects of the two interpolation methods shown in this specification.

[0028] Figure 4 This is a schematic diagram illustrating a lip shape determination method according to a specific embodiment of this specification.

[0029] Figure 5 This manual presents a comparison diagram of the effects of phoneme demonstration lip-shape and lip-shape driving parameters on a digital human, based on a specific embodiment.

[0030] Figure 6 This is a schematic diagram illustrating a three-segment interpolation according to a specific embodiment of this specification.

[0031] Figure 7 This is a block diagram illustrating a lip shape determining device according to an exemplary embodiment of this specification.

[0032] Figure 8 This is a hardware structure diagram of a computer device illustrated in this specification according to an exemplary embodiment. Detailed Implementation

[0033] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.

[0034] It should be noted that the steps of the corresponding methods are not necessarily performed in the order shown and described in this specification in other embodiments. In some other embodiments, the methods may include more or fewer steps than described in this specification. Furthermore, a single step described in this specification may be broken down into multiple steps in other embodiments; and multiple steps described in this specification may be combined into a single step in other embodiments.

[0035] A 3D digital human (hereinafter referred to as a digital human) is a 3D model that exists in the digital world and is driven by computer programs. The applications of digital humans are becoming increasingly widespread; for example, they can be used as customer service representatives or tour guides. Digital human actuation is a method of controlling the movement of a digital human through computer programs. A digital human has multiple key points on its body; by controlling the displacement and rotation of these key points, the body and face of the digital human can be altered, thus making the digital human move.

[0036] Digital human actuation includes facial actuation and limb actuation. In digital human applications, users' perception of the digital human's face is more pronounced; therefore, the facial actuation effect is crucial for digital human applications. Facial actuation for digital humans includes lip-syncing actuation and expression actuation. In related technologies, expression actuation is generally achieved by playing preset facial animations. Lip-syncing actuation is generally based on audio or text data, using automated techniques to generate lip-syncing parameters that are semantically consistent with the audio or text data. These lip-syncing parameters then drive the digital human to produce lip movements consistent with the audio or text data, making the digital human more lifelike and natural.

[0037] In related technologies, lip-syncing methods for digital humans include speech-driven and phoneme-driven methods.

[0038] Speech-driven lip-sync is a method that predicts lip-sync parameters based on input audio data, typically implemented using artificial neural networks. For example... Figure 1 As shown, Figure 1A schematic diagram of a speech-driven lip-sync method is shown. First, audio or video data is input into a predefined audio feature extractor (i.e., encoder) to extract audio features from the audio or video data. These audio features can be an audio representation matrix. Then, a trained decoder is used to input the audio features into the decoder to obtain the lip-sync driving parameters for each frame. Figure 1 The dashed boxes in the diagram represent data, and the solid boxes represent artificial neural network models.

[0039] The encoder and decoder described above can be artificial neural network models trained using audio-lip alignment training data. Specifically, for the decoder, as... Figure 1 As shown, it can use an autoregressive method to output lip-shape driving parameters based on the input audio features. That is, it takes the lip-shape driving parameters of the first N-1 frames as input and uses a decoder to predict the lip-shape driving parameters of the Nth frame. The lip-shape driving parameters can be driving data in vector form, such as a 52-dimensional vector, where each value in the vector can represent the magnitude or direction of the displacement of a key point of the digital human's lips relative to its initial position.

[0040] Although the above-mentioned speech-based end-to-end lip-syncing method can conveniently generate lip-syncing parameters, this method has the following problems:

[0041] First, speech-based lip-syncing methods utilize artificial neural networks, and training these networks requires a large amount of high-quality audio-lip alignment training data. Most of this training data is collected in laboratory environments, which is costly. Furthermore, while the audio quality of the training data is relatively high due to laboratory collection, real-world applications may use noisy audio data. The characteristics of this noisy audio data differ from those of laboratory-collected data. Models trained on laboratory-collected audio data may exhibit unstable performance when processing this different type of out-of-domain data, potentially leading to semantically incompatible lip-syncing parameters output by the decoder when given noisy audio data as input.

[0042] Secondly, in practical applications, the types and sources of the digital human models that need to be driven vary, leading to differences in the driving standards used for different digital humans. Under different driving standards, the format of the lip-sync driving parameters may differ; for example, some standards use 52-dimensional lip-sync driving parameters, while others use 108-dimensional parameters. The keypoints selected also differ under different driving standards. Therefore, for digital humans with different driving standards, it is necessary to collect training data under multiple driving standards to train the encoder and encoder, resulting in significant model training overhead and high costs.

[0043] Third, the aforementioned decoder models generally employ autoregressive prediction, which is relatively slow and cannot guarantee the real-time generation of lip-shaped driving parameters. To improve the model's inference efficiency, specialized modifications to accelerate inference are needed; however, ensuring high real-time performance is prohibitively expensive.

[0044] Besides voice-driven methods, phoneme-based lip-syncing can also be used to drive digital humans, overcoming the aforementioned problems of voice-driven methods. Phoneme-driven methods divide the text into multiple phonemes and drive the digital human's lips according to preset lip-syncing parameters for each phoneme.

[0045] Compared to speech-driven methods, phoneme-based lip-syncing methods maintain consistent lip-syncing parameters for the same phoneme at different times, resulting in more stable driving effects and stronger interpretability. Furthermore, since this method drives the digital human based on pre-defined lip-syncing parameters for each phoneme, it is more efficient than autoregressive prediction models, ensuring real-time processing. Moreover, for different driving standards, it eliminates the need to collect training data under multiple standards; only the lip-syncing parameters for each phoneme under different standards need to be obtained. These parameters can be derived from the lip-syncing parameters of a single standard through simple processing, resulting in lower costs.

[0046] However, in phoneme-based lip-syncing methods, the lip movements of digital humans still lack naturalness and realism. The main reason for this is the lack of natural transitions between phonemes during existing digital human broadcasting, which easily leads to abrupt changes in lip shape.

[0047] To address the issue of unnatural and unrealistic lip movements in digital humans, this specification provides a lip movement determination method. This method acquires text data from a digital human and determines the broadcast time segment for each phoneme within the text data. Each phoneme's broadcast time segment is divided into three segments. The middle segment of these three segments can be determined by consulting a phoneme-lip movement mapping table to identify the keyframe lip movements to be maintained. The other two segments can be determined using interpolation between the keyframe lip movements of the phoneme and the keyframe lip movements preceding and following the phoneme.

[0048] In this way, the broadcast time of each phoneme is divided into three parts. Only the middle time is used to maintain the lip shape of the current phoneme, and the other time is used to transition between the lip shape of the previous phoneme or the lip shape of the next phoneme. This makes the lip shape changes more realistic and natural.

[0049] Compared to interpolation methods in related technologies, the method in this specification takes into account the problem of lip shape abrupt changes that are ignored in related technologies, and solves this problem by interpolating a portion of the broadcast time period of each phoneme, making the lip shape more realistic and natural.

[0050] Moreover, compared to lip-driven speech, phoneme-driven methods can also solve the problems of poor stability, poor real-time performance, and high cost associated with speech-driven methods.

[0051] The following section will provide a detailed explanation of one method for determining lip shape as described in this specification. Figure 2 As shown, Figure 2 This is a flowchart illustrating a lip shape determination method according to an exemplary embodiment of this specification, including:

[0052] Step 201: Obtain the text data to be broadcast by the digital human, and determine the broadcast time period of the phoneme corresponding to each text unit in the text data.

[0053] The text units are characters and / or words.

[0054] In step 201, in order to finally obtain the lip shape of the digital human, it is necessary to first obtain the text data that the digital human needs to broadcast, and then break the text data into each phoneme and determine the duration of each phoneme.

[0055] The text data can be a sentence or a word. The length of the text data can be determined by the processor's processing power; for example, a powerful processor can generate lip movements for longer sentences in a shorter time, allowing for the selection of longer text data. The length of the text data can also be preset, such as setting it to input one sentence at a time. Alternatively, the length of the text data can be determined based on sentence segmentation, such as using punctuation marks to determine the length of text data to be processed in a single operation.

[0056] A phoneme is a concept in linguistics, referring to the smallest, indivisible, semantically distinct unit of speech in a language. A phoneme is an abstract category of speech sounds and does not directly correspond to any specific sound; rather, it refers to the functional attribute of speech sounds that can distinguish word meaning. In Chinese, phonemes can include initials, finals, etc. In English, phonemes can include individual phonetic symbols.

[0057] A text unit is a sentence that cannot be further divided and consists of at least one phoneme. For example, in Chinese, a text unit can be a single character; in English, a text unit can be a word; and in Japanese, a text unit can be a single hiragana or katakana character (here, a hiragana or katakana character can be considered as a single character).

[0058] As for the method of obtaining text data, it can be to directly obtain text data marked with the start and end timestamps of each text unit, or text data marked with the start and end timestamps of each phoneme. The start and end timestamps can be marked manually or by other automated methods.

[0059] Another method for obtaining text data is to first obtain the audio data that the digital human needs to broadcast, then perform speech recognition processing on the audio data to obtain the text data, as well as the start and end timestamps of each text unit in the audio data broadcast, then split the text unit into multiple phonemes, and determine the broadcast time period of each phoneme based on the start and end timestamps of each text unit.

[0060] In other words, step 201 includes: performing speech recognition on the target audio to obtain the text data and the start and end timestamps corresponding to each text unit in the text data; determining the phonemes corresponding to each text unit in the text data; and dividing the time period corresponding to the start and end timestamps of each text unit into broadcast time periods for multiple phonemes corresponding to that text unit according to a preset division rule.

[0061] Determining the phoneme corresponding to each text unit in the text data can be achieved by pre-storing a phoneme table, which records the phoneme corresponding to each text unit, and then looking up the phoneme table to determine the phoneme corresponding to each text unit. Alternatively, the phoneme corresponding to each text unit can be determined by first converting the text unit into pinyin or phonetic symbols, and then splitting the pinyin or phonetic symbols into multiple phonemes according to the splitting rules of pinyin or phonetic symbols. The above examples are not intended to limit this specification; any method for determining the phoneme corresponding to a text unit in related technologies can be used as the method in this specification.

[0062] Furthermore, the preset division rule can be an average division, that is, dividing the time period corresponding to the start and end timestamps of a text unit equally among the phonemes of that text unit, thereby determining the broadcast time period for each phoneme. Alternatively, the division rule can be determined based on pronunciation patterns. For example, when the text data is in Chinese, since the pronunciation of vowels is generally longer than that of consonants during reading, a certain pre-set division ratio can be used based on these pronunciation patterns. For instance, 40% of the time period corresponding to the start and end timestamps of a text unit could be allocated to consonants, and 60% to vowels.

[0063] The above two division methods are not intended to limit this instruction manual. This instruction manual does not limit the preset division rules for the broadcast time periods of phonemes.

[0064] Step 203: For each phoneme of the text data, the broadcast time period of the phoneme is divided into a first time period, a second time period including keyframes, and a third time period.

[0065] After determining the broadcast time period for each phoneme in step 201, the broadcast time period for each phoneme needs to be further divided into a time period for maintaining the keyframe lip shape of the phoneme (i.e., the second time period) and a time period for smooth transition (i.e., the first time period and the third time period). The first time period is used to connect with the previous keyframe lip shape, and the third time period is used to connect with the next keyframe lip shape.

[0066] By dividing the broadcast time of each phoneme into multiple segments and transitioning the lip shape between the first and third segments, abrupt changes in lip shape can be avoided, making the lip shape of the digital human more realistic and natural. Furthermore, the second segment can be the middle segment of the broadcast time; maintaining the keyframe lip shape during the second segment also ensures a stronger semantic correspondence between the lip shape and the broadcast text content.

[0067] The system divides the speech into three time periods. This can be achieved by dividing the broadcast time of each phoneme into three equal segments, numbered sequentially as the first, second, and third time periods. Alternatively, in some cases, the digital human is customized. The division ratio of the three time periods can be determined based on the lip-shape changes of the customized digital human. For example, if the customized digital human is based on a specific person whose lip movements are slower and do not maintain a particular lip shape for a long time, a preset division ratio can be used to determine the first, second, and third time periods. The second time period can account for less than one-third of this ratio, ensuring a closer correspondence between the customized digital human and the specific person. Of course, other methods can also be used to divide the three time periods; the above examples are not intended to limit the scope of this instruction manual.

[0068] The specific implementation of dividing the three time periods can be to divide the broadcast frame corresponding to each phoneme into three segments, with each time period including several frames.

[0069] Step 205: Determine the keyframe lip shape corresponding to each phoneme of the text data according to the preset phoneme-lip shape mapping table.

[0070] Step 205 determines the keyframe lip shape that needs to be maintained in the second time period.

[0071] Regarding the execution order of steps 203 and 205, steps 203 can be executed first, followed by steps 205, or steps 205 can be executed first, followed by steps 203, or steps 203 and 205 can be executed in parallel. This manual does not limit the execution order of steps 203 and 205.

[0072] In this context, a keyframe refers to a specific frame defined on the timeline. Between keyframes, interpolation and other methods can be used to determine the corresponding lip-sync driving parameters. The keyframe lip shape is the lip shape on the keyframe. The keyframe lip shape can be directly obtained using methods such as table lookup.

[0073] A phoneme-lip shape mapping table is a database that stores the lip shape corresponding to each phoneme. The lip shape defined in this specification can be a lip shape driving parameter, which is the driving data used to drive the movement of key points on the digital human's lips. Alternatively, the lip shape defined in this specification can also be other types of data, such as data describing the position coordinates of key points on the lips in a three-dimensional coordinate system. After obtaining other types of lip shape data, the lip shape data can be mapped to lip shape driving parameters through methods such as table lookup, thereby driving changes in the lip shape.

[0074] For Chinese and English, Chinese and English phonemes can include 16 whole-syllable recognition, 23 initials, 34 finals, and 39 phonetic symbols. When the text data includes Chinese and English, the phoneme-lip shape mapping table can include the lip shape data corresponding to the above 16 whole-syllable recognition, 23 initials, 34 finals, and 39 phonetic symbols.

[0075] The method for obtaining the phoneme-lip mapping table will be explained next. Digital humans are generally divided into custom digital humans and ordinary digital humans. Custom digital humans are those customized for a specific image and can be driven by a computer, such as digital humans customized for a celebrity or a cartoon character. Ordinary digital humans are those not customized for a specific image and use a general driving standard. It should be noted that custom digital humans may also use a general driving standard.

[0076] For customized digital humans, in order to make the digital human correspond more closely to the specific image, customized digital humans are generally pre-made with a phoneme-lip shape mapping table that includes basic phonemes. In this case, the lip shape corresponding to other phonemes can be obtained by combining the lip shape lines of multiple phonemes according to the phoneme-lip shape mapping table that includes basic phonemes.

[0077] Specifically, if the digital human is a digital human model customized for a specific image, this digital human model generally corresponds to lip-shape driving parameters of each phoneme in a pre-customized first phoneme set. In one embodiment, the range of the preset phoneme set corresponding to the phoneme-lip shape mapping table required in step 205 is larger than the customized first phoneme set. In this case, the phoneme-lip shape mapping table is obtained in the following way: the lip-shape driving parameters corresponding to the intersection phonemes of the first phoneme set and the preset phoneme set are set as lip shape data in the phoneme-lip shape mapping table; for target phonemes in the preset phoneme set that do not belong to the intersection phonemes, they are decomposed into a combination of multiple phonemes in the first phoneme set, and the combination of the lip-shape driving parameters of each of the multiple phonemes is set as the lip shape data of the target phoneme in the phoneme-lip shape mapping table.

[0078] For example, in Chinese, the first phoneme set can include all initials, vowels, and finals. It may exclude compound finals such as an, ai, and ing. For ai, it can be obtained by linearly combining the lip shapes of the finals a and i. Of course, the first phoneme set can also include other phonemes; this specification does not limit the specific form of the first phoneme set.

[0079] In some cases, customized digital humans may also have a pre-made phoneme-lip mapping table that covers all phonemes. In this case, the pre-made phoneme-lip mapping table can be directly used as the phoneme-lip mapping table in step 205.

[0080] For typical general-purpose digital human models, there is usually no pre-built phoneme-lip shape mapping table. Therefore, the phoneme-lip shape mapping table can be obtained through motion capture. In other words, the phoneme-lip shape mapping table is obtained by acquiring the lip shape driving parameters corresponding to each phoneme in a preset phoneme set through motion capture, thereby generating the phoneme-lip shape mapping table.

[0081] For digital humans with different driving standards, the motion capture tool corresponding to each driving standard can be used to obtain the phoneme-lip shape mapping table for each driving standard. For example, for a digital human model using the ARKit52 driving standard, the Livelink motion capture software corresponding to the ARKit52 driving standard can be used to capture the lip shapes of the motion capture actor when demonstrating different phonemes, thereby obtaining the phoneme-lip shape mapping table.

[0082] Furthermore, for digital humans with different driving standards, the lip shapes captured by motion capture under one driving standard can be extended to obtain phoneme-lip shape mapping tables under other driving standards. For example, if the lip shape is a lip shape driving parameter, and there is already a 52-dimensional phoneme-lip shape mapping table, then the lip shape driving parameters of this phoneme-lip shape mapping table can be upgraded or down-dimensioned to obtain phoneme-lip shape mapping tables of other dimensions under other driving standards.

[0083] For various digital humans, the cost of obtaining the phoneme-lip shape mapping table is relatively low. Compared with the speech-driven methods in related technologies, which require obtaining audio-lip shape alignment training data under each driving standard, it is lower in cost and easier to extend to multiple driving standards.

[0084] Step 207: For each phoneme of the text data, determine the lip shape for the first time period based on the interpolation of the lip shape of the previous keyframe and the keyframe lip shape of the phoneme, set the lip shape for the second time period to maintain the keyframe lip shape, and determine the lip shape for the third time period based on the interpolation of the keyframe lip shape of the phoneme and the next keyframe lip shape.

[0085] After obtaining the keyframe lip shape corresponding to each phoneme through the phoneme-lip shape mapping table, and determining the first time period, second time period, and third time period corresponding to each phoneme, the lip shape can be set to transition with the previous keyframe lip shape in the first time period, the keyframe lip shape corresponding to the current phoneme can be maintained in the second time period, and the lip shape can be set to transition with the next keyframe lip shape in the third time period.

[0086] There are two cases for the lip shape in the previous keyframe and the lip shape in the next keyframe: one is the lip shape corresponding to the keyframe of the phoneme adjacent to the phoneme, and the other is a closed lip shape. The implementation of step 207 in these two cases will be explained below.

[0087] In one scenario, the speech in the broadcast audio is continuous, meaning there is no silence between adjacent phonemes, or the silence is very short. In this case, the lip shape of the previous keyframe is the lip shape of the phoneme preceding that phoneme in the text data, and the lip shape of the next keyframe is the lip shape of the phoneme following that phoneme in the text data.

[0088] For example, if the text data is "virtual digital human", the text data can be broken down into several phonemes such as "x", "u", "n", "i", "sh", "u", "zi", "r" and "en". For the phoneme "n", the previous keyframe lip shape can be the keyframe lip shape of "u", and the next keyframe lip shape can be the keyframe lip shape of "i".

[0089] Since the first time interval of each phoneme is used to smoothly transition from the previous keyframe lip shape to the current phoneme's keyframe lip shape, and the third time interval is used to smoothly transition from the current phoneme's keyframe lip shape to the next keyframe lip shape, then when the previous keyframe lip shape is the keyframe lip shape corresponding to the previous phoneme in the text data, and the next keyframe lip shape is the keyframe lip shape corresponding to the next phoneme in the text data, the third time interval of the previous phoneme and the first time interval of the current phoneme are combined to transition from the previous phoneme's keyframe lip shape to the current phoneme's keyframe lip shape, and the third time interval of the current phoneme and the first time interval of the next phoneme are combined to transition from the current phoneme's keyframe lip shape to the next phoneme's keyframe lip shape.

[0090] Accordingly, step 207, determining the lip shape of the first time period, specifically includes: determining the interpolation period between the start time of the third time period of the preceding phoneme and the start time of the second time period of the phoneme; obtaining the keyframe lip shape of the preceding phoneme and the interpolated lip shape data of the keyframe lip shape of the phoneme obtained according to the target interpolation algorithm in the interpolation period, including the lip shape of the first time period.

[0091] Specifically, in one example, since the third time segment of the previous phoneme and the first time segment of the current phoneme are merged into an interpolation segment, after obtaining the interpolated lip shape data of the interpolation segment, the interpolated lip shape data is divided according to the duration ratio of the third time segment of the previous phoneme and the first time segment of the current phoneme, and used as the lip shape data of the third time segment of the previous phoneme and the lip shape data of the first time segment of the current phoneme.

[0092] Similarly, step 207 determines the lip shape of the third time period, which may specifically include determining the interpolation period between the end time of the second time period of the phoneme and the end time of the first time period of the next phoneme; obtaining the interpolated lip shape data of the keyframe lip shape of the phoneme and the keyframe lip shape of the next phoneme in the interpolation period according to the target interpolation algorithm, including the lip shape of the third time period.

[0093] In another scenario, during the process of a digital human broadcasting text data, there may still be pauses between some text units. During these pauses, the digital human remains silent. In this case, if the digital human is still emitting interpolated lip-shape data from the keyframe lip shape of the previous phoneme to the keyframe lip shape of the next phoneme, the range of lip-shape changes will be small, which is obviously unnatural.

[0094] Therefore, when the interval between two adjacent phonemes in the text data is relatively long, a keyframe corresponding to a closed lip shape can be inserted. In other words, for adjacent first and second phonemes in the text data, if the time interval between the end of the broadcast time period of the first phoneme and the start of the broadcast time period of the second phoneme is greater than a preset time threshold, a keyframe corresponding to a closed lip shape is inserted between the end and start times.

[0095] The keyframes corresponding to the closed lip shape can be inserted midway between the end and start times, making the closed lip shape more natural. The inserted keyframes can be one frame or multiple frames, or the number of keyframes can be determined based on the time interval between the end and start times. If the time interval is greater than a preset time threshold, the number of inserted closed lip keyframes is positively correlated with the time interval. Of course, the above examples are not intended to limit this specification.

[0096] This avoids the problem of visually small changes in lip shape due to overly smooth lip shape changes.

[0097] Besides the aforementioned situations where closed lip shapes may exist, when a digital human pronounces the first or last sound of text data, the lip shape in the previous keyframe or the lip shape in the next keyframe may also be a closed lip shape.

[0098] If the lip shape in the previous keyframe or the lip shape in the next keyframe is a closed lip shape, the lip shape in the first time period or the third time period can be obtained according to the following method.

[0099] When the lip shape in the previous keyframe is a closed lip shape, determining the lip shape in the first time period may include: determining the interpolation period between the keyframe time corresponding to the closed lip shape and the start time of the second time period of the phoneme; obtaining the interpolated lip shape data of the closed lip shape and the keyframe lip shape of the phoneme obtained according to the target interpolation algorithm in the interpolation period, including the lip shape of the first time period.

[0100] The method for determining the lip shape of the third time segment when the next keyframe lip shape is a closed lip shape is similar to the method for determining the lip shape of the first time segment, and will not be repeated here.

[0101] The above describes the implementation of step 207 when the previous keyframe lip shape is either the keyframe lip shape corresponding to the previous phoneme or a closed lip shape. Furthermore, for any phoneme, its previous keyframe lip shape and next keyframe lip shape can be any combination of the phoneme's keyframe lip shape and closed lip shape. For example, the previous keyframe lip shape is the keyframe lip shape of the previous phoneme, and the next keyframe lip shape is the keyframe lip shape of the next phoneme; or both the previous and next keyframe lip shapes are closed lip shapes; or the previous keyframe lip shape is the keyframe lip shape of the previous phoneme, and the next keyframe lip shape is a closed lip shape; or the previous keyframe lip shape is a closed lip shape, and the next keyframe lip shape is the keyframe lip shape of the next phoneme.

[0102] Furthermore, the target interpolation algorithm used in step 207 can be a second-order interpolation algorithm and / or a Gaussian smoothing algorithm with coefficients less than a preset coefficient threshold (hereinafter referred to as a Gaussian smoothing algorithm with smaller coefficients). In other words, the above interpolation algorithm can be a second-order interpolation algorithm plus a Gaussian smoothing algorithm with smaller coefficients, or a second-order interpolation algorithm plus other smoothing algorithms, or other interpolation algorithms plus a Gaussian smoothing algorithm with smaller coefficients. The above processing ensures that the interpolation result does not affect the overall opening and closing amplitude of the lip shape, making the overall visual effect of the lip shape more natural. Especially when the lip shape in the previous keyframe or the lip shape in the next keyframe is a closed lip shape, these two interpolation algorithms will not erase the closed lip shape.

[0103] like Figure 3 As shown, Figure 3 In the two figures, the horizontal axis represents time, and the vertical axis represents the overall opening and closing degree of the lips. Figure 3 (a) is the lip shape after linear interpolation and Gaussian smoothing with large coefficients. Figure 3 (b) shows the lip shape after processing with second-order interpolation and a Gaussian smoothing algorithm with small coefficients. From Figure 3 As can be seen, compared with conventional linear interpolation and Gaussian smoothing algorithm with large coefficients, the lip shape changes more significantly after processing with second-order interpolation and Gaussian smoothing algorithm with smaller coefficients. It does not erase the closed lip shape in the middle, ensuring that the overall opening and closing range of the lip shape is not too small, making the lip shape change more natural.

[0104] The lip shapes obtained in step 207 for the first and third time periods can be generated in real time during processing.

[0105] In addition, to improve processing efficiency, the lip shapes in the first and third time periods mentioned above can also be obtained from pre-cached lip shape interpolation data.

[0106] Specifically, the target buffer stores lip shape interpolation data of keyframe lip shape combinations corresponding to various phoneme combinations in multiple interpolation time periods; the lip shape of the first time period and the lip shape of the third time period are read from the target buffer according to the corresponding keyframe lip shape combination and interpolation time period.

[0107] Since the interpolation algorithm has a larger overall time cost compared to other steps in this method, the real-time performance of the processing can be further improved by pre-storing the lip-shaped interpolation data in the target cache.

[0108] Specifically, regarding the data in the target buffer, since the broadcast time of each phoneme at normal speaking speed is generally between 20ms and 50ms, the lip-interpolation data of each phoneme combination interpolation time period of 0-20 frames can be cached, which can handle most situations.

[0109] Furthermore, to reduce the space occupied by the target cache, since some phonemes are generally not adjacent (e.g., two initials in Chinese are usually not adjacent), the frequency of occurrence of each phoneme combination can be obtained in advance using artificial neural networks or other methods. The lip-interpolation data of phoneme combinations with a frequency greater than a certain threshold are cached in the target cache across multiple interpolation periods, while phoneme combinations with too low a frequency are not cached. This improves real-time performance while further reducing storage space usage.

[0110] After performing step 207, the following can also be done: After determining the lip shape for each phoneme's broadcast time period, the resulting lip shape sequence can be smoothed. Smoothing can be done using a first-order Gaussian filter, or other smoothing algorithms can be employed.

[0111] Smoothing can make the overall lip shape changes more natural and realistic, improving the display effect of digital humans.

[0112] Compared to speech-driven methods in related technologies, the above method has several advantages: First, it only requires creating a phoneme-lip shape mapping table, eliminating the need for acquiring large amounts of high-quality audio-lip alignment data, thus reducing costs. Second, compared to encoder-decoder models trained on laboratory data, the speech recognition results based on text data are less affected by noise, resulting in more stable and interpretable output lip shapes. Third, for digital human models with different driving standards, it eliminates the need to train separate encoder-decoder models for each standard. Instead, it uses a phoneme-lip shape mapping table based on a specific driving standard, employing cost-effective methods such as dimensionality increase or decrease to obtain the mapping tables for different standards, further reducing costs. Fourth, addressing the issue of poor real-time performance caused by the high time consumption of autoregressive inference in speech-driven methods, determining lip shapes by looking up the phoneme-lip shape mapping table is more efficient overall. Furthermore, the pre-cached lip shape interpolation data in the target buffer further improves processing efficiency, achieving real-time driving performance.

[0113] Furthermore, compared to phoneme-driven methods in related technologies, the method in this specification divides each phoneme's broadcast time period into three segments and performs interpolation processing on different segments, avoiding abrupt changes in lip shape and making the lip shape more natural and realistic.

[0114] The following will describe a lip shape determination method shown in this specification with reference to a specific embodiment. In this embodiment, the lip shape is a lip shape driving parameter, and the text data is Chinese.

[0115] like Figure 4 As shown, this embodiment mainly includes the following steps:

[0116] First, the input text data and the start and end timestamps of each character in the text data are converted into initial and final phonemes and their corresponding start and end timestamps.

[0117] like Figure 4 As shown, this transformation can begin by converting each character into pinyin, and then further breaking down the pinyin into initials and finals. Here, the text data is "virtual digital person," with the overall start and end timestamps of the virtual digital person ranging from 0 to 0.97 seconds. Based on the pinyin library, the pinyin for "virtual digital person" can be determined as "xu," "ni," "shu," "zi," and "ren." Furthermore, these pinyin are broken down according to initials, finals, and whole-syllable recognition into several phonemes: "x," "u," "n," "i," "sh," "u," "zi," "r," and "en."

[0118] Second, the lip-driving parameters corresponding to each phoneme can be found through a predefined phoneme-lip mapping table. The lip-driving parameters corresponding to each phoneme are the keyframe lip shapes mentioned above.

[0119] The phoneme-lip mapping table stores the most suitable lip-shape driving parameters for each phoneme. For example... Figure 4 As shown, the lip-driving parameters of x and u can be determined by looking up the phoneme-lip mapping table.

[0120] The methods for obtaining phoneme-lip mapping tables mainly fall into two common scenarios. One is a custom digital human with a predefined phoneme-lip controller during modeling, and the other is a general digital human that adopts a common driving standard (taking ARKit52 standard as an example) without a specially preset phoneme-lip controller.

[0121] For custom 3D digital humans with predefined controllers, the predefined controllers usually cover all initials, vowels, and some English phonetic symbols. Therefore, the driving parameters of most phonemes can be obtained directly. For compound vowels such as an / ai / ing and some phonetic symbols, the lip-shape driving parameters of these phonemes can be obtained by combining them based on known phonetic controllability through expert experience.

[0122] For digital humans using a universal facial motion capture standard, since there are no preset controllers for specific phonemes, it is necessary to use motion capture tools corresponding to that standard (such as Livelink motion capture software corresponding to the ARKit52 standard) to collect facial motion capture data from professional motion capture actors demonstrating different phonemes, thereby obtaining the lip movement parameters for each phoneme. For example... Figure 5 As shown, Figure 5 The image on the left shows the lip shape of a motion capture actor during a demonstration. The effect of using lip-driving parameters determined based on this lip shape to drive a digital human is shown below. Figure 5 As shown in the diagram on the right.

[0123] Third, the lip-shaped interpolation data is filled using a three-segment interpolation algorithm.

[0124] The phoneme-lip mapping table only stores the lip movement driving parameters for one frame for each phoneme. However, in practical applications, if the lip opening only lasts for one frame, the amplitude of the lip movement will appear very small to the user due to the short duration.

[0125] To address the aforementioned issues, a three-segment interpolation method is employed to obtain the interpolation sequence between every two lip-shape keyframes. Specifically, the broadcast time period for each phoneme is divided into three equal segments: the first segment is used for the transition from the lip shape of the previous keyframe to the lip shape of the current phoneme keyframe; the second segment is used to maintain the lip shape of the current phoneme keyframe; and the third segment is used for the transition from the lip shape of the current phoneme keyframe to the lip shape of the next keyframe.

[0126] like Figure 4As shown, x can be used as a keyframe, u as a keyframe, and interpolation transitions can be performed between the two keyframes. It should be noted that, for ease of explanation, Figure 4 The example only uses the case where the second segment of x and u occupies only one frame, and the third segment of x and the first segment of u share one frame. However, in actual applications, the keyframe lip shape of x and u may not be maintained for more than one frame, and the interpolation transition between the two keyframe lip shapes may also occupy more than one frame.

[0127] For specific three-segment interpolation methods, such as Figure 6 As shown, for phoneme x, the first time period is the transition frame sequence from the closed lip shape (the previous keyframe lip shape) to the x keyframe lip shape, the second time period is the keyframe lip shape preservation sequence of phoneme x, and the third time period is the transition sequence from the keyframe lip shape of phoneme x to the keyframe lip shape of phoneme u. In addition, since the first time period of phoneme u is also the transition sequence from the keyframe lip shape of phoneme x to the keyframe lip shape of phoneme u, the first time period of phoneme u and the third time period of phoneme x together serve as the transition sequence from the keyframe lip shape of phoneme x to the keyframe lip shape of phoneme u.

[0128] Furthermore, considering the transition between phonemes, if the interval between the broadcasting time of two phonemes is too long, the lip shape will appear too smooth, resulting in a smaller visual change in lip shape. For example, Figure 6 The first segment of phoneme u and the third segment of phoneme x together serve as the transition sequence from the lip shape of the phoneme x keyframe to the lip shape of the phoneme u keyframe. If there is a long time interval between the third segment of x and the first segment of u, the transition sequence from the lip shape of the phoneme x keyframe to the lip shape of the phoneme u keyframe will be too smooth and unnatural. Therefore, for cases where the time interval between phoneme broadcast time segments is greater than a preset threshold of 100ms, a closed lip shape keyframe will be inserted in the middle.

[0129] Furthermore, to further improve processing efficiency, the lip-shape interpolation sequence can be pre-cached. Specific caching methods and interpolation algorithms are detailed above and will not be repeated here.

[0130] Fourth, a first-order Gaussian filter is used to smooth the entire lip shape sequence, further enhancing the visual naturalness and smoothness of the lip shape.

[0131] Figure 4 The three vectors obtained at the end are the lip-shaped driving parameters of the interpolation frames x, u, and xu in the above processing.

[0132] Corresponding to the embodiments of the foregoing methods, this specification also provides embodiments of the apparatus and the terminal to which it is applied.

[0133] like Figure 7 As shown, Figure 7 This is a block diagram illustrating a lip shape determining device according to an exemplary embodiment of this specification, the device comprising:

[0134] The time period determination module 710 is used to acquire the text data to be broadcast by the digital human and determine the broadcast time period of the phoneme corresponding to each text unit in the text data; the text unit is a character and / or a word;

[0135] The time period segmentation module 720 is used to divide the broadcast time period of each phoneme in the text data into a first time period, a second time period including keyframes, and a third time period.

[0136] The keyframe lip shape determination module 730 is used to determine the keyframe lip shape corresponding to each phoneme of the text data according to a preset phoneme-lip shape mapping table.

[0137] The lip shape determination module 740 is used to determine the lip shape of a first time period for each phoneme of the text data based on the interpolation of the lip shape of the previous keyframe and the keyframe lip shape of the phoneme, set the lip shape of the second time period to maintain the keyframe lip shape, and determine the lip shape of a third time period based on the interpolation of the keyframe lip shape of the phoneme and the next keyframe lip shape.

[0138] In an optional embodiment, the time period determination module 710 is specifically used for: performing speech recognition on the target audio to obtain the text data and the start and end timestamps corresponding to each text unit in the text data; determining the phonemes corresponding to each text unit in the text data; and dividing the time period corresponding to the start and end timestamps of each text unit into broadcast time periods for multiple phonemes corresponding to the text unit according to a preset division rule.

[0139] In an optional embodiment, the digital human is a digital human model customized for a specific image. This digital human model corresponds to lip-shape driving parameters of each phoneme in a pre-customized first phoneme set. The range of the preset phoneme set corresponding to the phoneme-lip shape mapping table is larger than that of the first phoneme set. The above device further includes a first mapping table acquisition module (not shown in the figure), used to: set the lip-shape driving parameters corresponding to the intersection phonemes of the first phoneme set and the preset phoneme set as lip-shape data in the phoneme-lip shape mapping table; for a target phoneme in the preset phoneme set that does not belong to the intersection phonemes, decompose it into a combination of multiple phonemes in the first phoneme set, and set the combination of the lip-shape driving parameters of each of the multiple phonemes as the lip-shape data of the target phoneme in the phoneme-lip shape mapping table.

[0140] In an optional embodiment, the above-mentioned device further includes a second mapping table acquisition module (not shown in the figure), which is used to: acquire the lip-shape driving parameters corresponding to each phoneme in the preset phoneme set by motion capture, thereby generating the phoneme-lip-shape mapping table.

[0141] In an optional embodiment, the previous keyframe lip shape is the keyframe lip shape corresponding to the previous phoneme in the text data, and the next keyframe lip shape is the keyframe lip shape corresponding to the next phoneme in the text data.

[0142] In an optional embodiment, when the lip shape of the previous keyframe is the lip shape of the keyframe corresponding to the previous phoneme in the text data, the lip shape determination module 740 is specifically used to: determine the interpolation period between the start time of the third time period of the previous phoneme and the start time of the second time period of the phoneme; and obtain the interpolated lip shape data of the keyframe lip shape corresponding to the previous phoneme and the keyframe lip shape of the phoneme in the interpolation period, obtained according to the target interpolation algorithm, including the lip shape of the first time period.

[0143] In an optional embodiment, the above-mentioned device further includes a keyframe insertion module (not shown in the figure), which is used to: for adjacent first and second phonemes in the text data, if the time interval between the end time of the broadcast time period of the first phoneme and the start time of the broadcast time period of the second phoneme is greater than a preset time threshold, insert a keyframe corresponding to the closed lip shape between the end time and the start time.

[0144] In an optional embodiment, the previous keyframe lip shape or the next keyframe lip shape is a closed lip shape.

[0145] In an optional embodiment, the previous keyframe lip shape is a closed lip shape; the lip shape determination module 740 is specifically used to: determine the interpolation period between the keyframe time corresponding to the closed lip shape and the start time of the second time period of the phoneme; and obtain the interpolated lip shape data of the closed lip shape and the keyframe lip shape of the phoneme obtained according to the target interpolation algorithm in the interpolation period, including the lip shape of the first time period.

[0146] In one optional embodiment, the target interpolation algorithm is a second-order interpolation algorithm and / or a Gaussian smoothing algorithm with coefficients less than a preset coefficient threshold.

[0147] In one optional embodiment, the target buffer stores lip-shape interpolation data of keyframe lip-shape combinations corresponding to various phoneme combinations in multiple interpolation time periods; the lip shape of the first time period and the lip shape of the third time period are read from the target buffer according to the corresponding keyframe lip-shape combination and the interpolation time period.

[0148] In an optional embodiment, the above-described apparatus further includes a smoothing module (not shown in the figure) for smoothing the resulting lip shape sequence after determining the lip shape for each phoneme's broadcast time period.

[0149] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0150] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0151] like Figure 8 As shown, Figure 8 A hardware structure diagram of a computer device containing the lip-shape determining device of an embodiment is shown. This device may include: a processor 1010, a memory 1020 for storing processor-executable instructions, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are internally connected to each other via the bus 1050.

[0152] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits. The processor executes the executable instructions to implement the aforementioned lip-shape determination method.

[0153] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant executable instructions are stored in the memory 1020.

[0154] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.

[0155] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0156] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.

[0157] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.

[0158] This specification also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the lip shape determination method described above.

[0159] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0160] This specification also provides a computer program product, including computer program instructions that, when executed by a processor, implement the lip shape determination method described above.

[0161] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0162] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0163] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.

Claims

1. A method for determining lip shape, the method comprising: Acquire the text data to be broadcast by the digital human, and determine the broadcast time period of the phoneme corresponding to each text unit in the text data; The text units are characters and / or words; For each phoneme in the text data, the broadcast time period of that phoneme is divided into a first time period, a second time period including keyframes, and a third time period; Based on a preset phoneme-lip shape mapping table, determine the keyframe lip shape corresponding to each phoneme of the text data; For each phoneme in the text data, the lip shape for the first time period is determined based on the interpolation of the lip shape of the previous keyframe and the keyframe lip shape of the phoneme. The lip shape for the second time period is set to maintain the keyframe lip shape. The lip shape for the third time period is determined based on the interpolation of the keyframe lip shape of the phoneme and the next keyframe lip shape.

2. The method according to claim 1, wherein, The process of acquiring the text data to be broadcast by the digital human and determining the broadcast time period for the phoneme corresponding to each character in the text data includes: Speech recognition is performed on the target audio to obtain the text data, and the start and end timestamps corresponding to each text unit in the text data; Determine the phoneme corresponding to each text unit in the text data; According to the preset division rules, the time period corresponding to the start and end timestamps of each text unit is divided into the broadcast time periods of multiple phonemes corresponding to that text unit.

3. The method according to claim 1, wherein, The digital human corresponds to a digital human model customized for a specific image. This digital human model corresponds to lip-shape driving parameters for each phoneme in a pre-customized first phoneme set. The range of the preset phoneme set corresponding to the phoneme-lip shape mapping table is larger than the first phoneme set. The phoneme-lip shape mapping table is obtained in the following way: The lip-shape driving parameters corresponding to the intersection phonemes of the first phoneme set and the preset phoneme set are set as the lip-shape data in the phoneme-lip-shape mapping table. For a target phoneme in the preset phoneme set that does not belong to the intersection phonemes, it is decomposed into a combination of multiple phonemes in the first phoneme set, and the combination of the lip-shape driving parameters of each of the multiple phonemes is set as the lip-shape data of the target phoneme in the phoneme-lip-shape mapping table.

4. The method according to claim 1, wherein, The phoneme-lip shape mapping table is obtained in the following way: The lip-shape driving parameters corresponding to each phoneme in the preset phoneme set are obtained by motion capture, thereby generating the phoneme-lip shape mapping table.

5. The method according to claim 1, wherein the previous keyframe lip shape is the keyframe lip shape corresponding to the previous phoneme in the text data, and the next keyframe lip shape is the keyframe lip shape corresponding to the next phoneme in the text data.

6. The method according to claim 5, wherein, Determining the lip shape in the first time period includes: Determine the interpolation period between the start time of the third time period of the preceding phoneme and the start time of the second time period of the phoneme; Obtain the keyframe lip shape corresponding to the previous phoneme obtained according to the target interpolation algorithm and the interpolated lip shape data of the keyframe lip shape of the phoneme in the interpolation period, including the lip shape of the first time period.

7. The method according to claim 1, further comprising: For adjacent first and second phonemes in the text data, if the time interval between the end time of the broadcast time period of the first phoneme and the start time of the broadcast time period of the second phoneme is greater than a preset time threshold, a keyframe corresponding to the closed lip shape is inserted between the end time and the start time.

8. The method according to claim 7, wherein, The lip shape in the previous keyframe or the lip shape in the next keyframe is a closed lip shape.

9. The method according to claim 8, wherein the lip shape of the previous keyframe is a closed lip shape; determining the lip shape of the first time period includes: Determine the interpolation period between the keyframe moment corresponding to the closed lip shape and the start moment of the second time segment of the phoneme; Obtain the interpolated lip shape data of the closed lip shape and the keyframe lip shape of the phoneme obtained according to the target interpolation algorithm during the interpolation period, including the lip shape of the first time period.

10. The method according to claim 6 or 9, wherein, The target interpolation algorithm is a second-order interpolation algorithm and / or a Gaussian smoothing algorithm with coefficients less than a preset coefficient threshold.

11. The method according to claim 1, wherein, The target buffer stores lip shape interpolation data of keyframe lip shape combinations corresponding to various phoneme combinations in multiple interpolation time periods; the lip shape of the first time period and the lip shape of the third time period are read from the target buffer according to the corresponding keyframe lip shape combination and interpolation time period.

12. The method according to claim 1, further comprising: After determining the lip shape for each phoneme's broadcast time period, the resulting lip shape sequence is smoothed.

13. A lip shape determining device, the device comprising: The time period determination module is used to acquire the text data to be broadcast by the digital human and determine the broadcast time period of the phoneme corresponding to each text unit in the text data. The text units are characters and / or words; The time period segmentation module is used to divide the broadcast time period of each phoneme in the text data into a first time period, a second time period including keyframes, and a third time period. The keyframe lip shape determination module is used to determine the keyframe lip shape corresponding to each phoneme of the text data according to a preset phoneme-lip shape mapping table. The lip shape determination module is used to determine the lip shape of a first time period for each phoneme of the text data based on the interpolation of the lip shape of the previous keyframe and the keyframe lip shape of the phoneme, set the lip shape of the second time period to maintain the keyframe lip shape, and determine the lip shape of a third time period based on the interpolation of the keyframe lip shape of the phoneme and the next keyframe lip shape.

14. A computer device, comprising: processor; Memory used to store processor-executable instructions; The processor implements the lip shape determination method as described in any one of claims 1-12 by running the executable instructions.

15. A computer-readable storage medium storing computer instructions that, when executed by a processor, implement the lip shape determination method as described in any one of claims 1-12.

16. A computer program product comprising computer program instructions that, when executed by a processor, implement the lip shape determination method as described in any one of claims 1-12.

Citation Information

Patent Citations

  • Lip language combination method and device, electronic device and storage medium

    CN108831463A

  • Image processing method and device and electronic equipment

    CN111277912A