Model acquisition method, lip coefficient generation method, device, equipment and medium
By training the mapping layer and style encoding layer of the lip coefficient generation model and utilizing both unstyled and styled speech samples, the problem of recording large amounts of speech data in the existing technology is solved, and the target style lip coefficients are efficiently generated, thereby improving the efficiency of animation production.
Patent Information
- Application Number
- CN202211288333.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-20
- Publication Date
- 2025-09-26
- Estimated Expiration
- 2042-10-20
AI Technical Summary
In the prior art, in order to obtain lip-shape coefficients with a speaking style, a large amount of voice data needs to be recorded and trained, resulting in high work costs and low animation production efficiency.
The method of using long-term styleless speech samples to train the mapping layer and short-term styled speech samples to train the style coding vector is used to generate the target lip coefficient generation model, reducing work costs.
The target style lip-sync coefficient can be generated through a small amount of voice collection, which improves the efficiency of animation production and reduces work costs.
Smart Images

Figure CN115938352B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to a technology for synthesizing virtual image speech animation, and specifically to a method, device, equipment and medium for obtaining a lip coefficient generation model. The present application also relates to a method, device, equipment and medium for generating a lip coefficient. Background Art
[0002] In the field of animation synthesis, achieving complete avatar animation requires synthesizing speech into lip-sync coefficients. This involves taking a speech input and synthesizing the corresponding lip-sync coefficients to ensure that the avatar's lip movements match the speech content. However, given that different avatars require different speaking styles, determining the lip-sync coefficients requires ensuring they correspond to the avatar's speaking style.
[0003] In the existing technology, in order to obtain lip-shape coefficients with speaking styles, for different speaking styles, it is necessary to record at least 30 minutes of speech data that retains the speaking style as training samples to train the deep neural network. This method requires a lot of work costs to obtain speech motion capture data.
[0004] Therefore, how to reduce the work cost in the process of obtaining the lip shape coefficients that retain the speaking style and improve the production efficiency of animation has become a technical problem that those skilled in the art urgently need to solve. Summary of the Invention
[0005] The present application provides a method, device, electronic device, and storage medium for obtaining a lip-shape coefficient generation model. Long-term, styleless speech samples are used to train the mapping layer in a target lip-shape coefficient generation model, and short-term, speech-styled speech samples are used to train the style encoding vector in the target lip-shape coefficient generation model. This reduces the labor cost in obtaining lip-shape coefficients with a speaking style and improves animation production efficiency.
[0006] A first aspect of an embodiment of the present application provides a method for obtaining a lip shape coefficient generation model, comprising:
[0007] Obtain the first voice of the first preset duration and the first lip shape coefficient corresponding to the first voice, the first voice For unstyled voice.
[0008] A first phoneme sequence corresponding to the first speech is determined, and the first phoneme sequence and a first lip shape coefficient are determined as a first training sample.
[0009] A second speech of a second preset duration and a second lip shape coefficient corresponding to the second speech are obtained, where the second speech has a target style. The second preset duration is shorter than the first preset duration.
[0010] A second phoneme sequence corresponding to the second speech is determined, and the second phoneme sequence and the second lip shape coefficient are determined as a second training sample.
[0011] The target lip shape coefficient generation model is trained using the first training sample and the second training sample, wherein the target lip shape coefficient generation model is used to generate target lip shape coefficients with a target style corresponding to the target speech.
[0012] The first training samples are used to train the mapping layer in the target lip-shape coefficient generation model. The mapping layer is used to determine the mapping relationship between each phoneme in the first phoneme sequence and each lip-shape coefficient in the first lip-shape coefficient. The second training samples are used to train the style encoding layer in the target lip-shape coefficient generation model. The style encoding layer is used to extract the style features corresponding to the second phoneme sequence based on the second lip-shape coefficient and the mapping relationship.
[0013] A second aspect of the present application provides a method for generating a lip coefficient, including:
[0014] A target phoneme sequence of the target speech is obtained, where the target phoneme sequence is used to represent semantic information of the target speech.
[0015] The target phoneme sequence is input into the target lip-shape coefficient generation model to obtain the target lip-shape coefficient output by the target lip-shape coefficient generation model. The target lip-shape coefficient is used to drive the mouth movements of the target virtual character.
[0016] The target lip shape coefficient generation model is generated according to the method for obtaining the lip shape coefficient generation model provided in the first aspect.
[0017] A third aspect of the embodiments of the present application provides a device for obtaining a lip shape coefficient generation model, comprising:
[0018] The acquisition unit is configured to acquire a first voice of a first preset duration and a first lip shape coefficient corresponding to the first voice, wherein the first voice is a styleless voice.
[0019] The determining unit is configured to determine a first phoneme sequence corresponding to the first speech, and determine the first phoneme sequence and the first lip shape coefficient as a first training sample.
[0020] The acquiring unit is further configured to acquire a second speech of a second preset duration and a second lip shape coefficient corresponding to the second speech, wherein the second speech is a speech of a target style and the second preset duration is shorter than the first preset duration.
[0021] The determining unit is further configured to determine a second phoneme sequence corresponding to the second speech, and determine the second phoneme sequence and the second lip shape coefficient as a second training sample.
[0022] The training unit is configured to train a target lip-shape coefficient generation model using the first training sample and the second training sample, wherein the target lip-shape coefficient generation model is configured to generate target lip-shape coefficients having a target style corresponding to the target speech.
[0023] The first training sample is used to train the mapping layer in the target lip-shape coefficient generation model. The mapping layer is used to determine the mapping relationship between each phoneme in the first phoneme sequence and each lip-shape coefficient in the first lip-shape coefficient. The second training sample is used to train the style encoding layer in the target lip-shape coefficient generation model. The style encoding layer is used to extract style features corresponding to the second phoneme sequence based on the second lip-shape coefficient and the mapping relationship.
[0024] A fourth aspect of the present application provides a device for generating a lip coefficient, including:
[0025] The acquisition unit is used to acquire a target phoneme sequence of the target speech, where the target phoneme sequence is used to represent semantic information of the target speech.
[0026] The processing unit is used to input the target phoneme sequence into the target lip shape coefficient generation model to obtain the target lip shape coefficient output by the target lip shape coefficient generation model, and the target lip shape coefficient is used to drive the mouth movement change of the target virtual character.
[0027] The target lip-shape coefficient generation model is obtained according to the device provided in the third aspect.
[0028] A fifth aspect of the embodiments of the present application provides an electronic device, including:
[0029] processor;
[0030] The memory is used to store a method program, and when the program is read and executed by the processor, any one of the above methods is executed.
[0031] The present application also provides a computer storage medium, which stores a computer program, and when the program is executed, any of the above methods is implemented.
[0032] Compared with the prior art, this application has the following advantages:
[0033] The method for obtaining the lip coefficient generation model provided by the present application is to obtain the lip coefficient generation model by using the first language of the first preset time length.The first phoneme sequence of the first speech provides the semantic information required for synthesizing lip-shape coefficients, while the lip-shape coefficients of the second speech with a first target style provide the style information required for synthesizing the lip-shape coefficients. Next, based on the first phoneme sequence and first lip-shape coefficients of the first speech, a mapping relationship between different phonemes and different lip-shape coefficients is determined. Then, based on the second phoneme sequence and second lip-shape coefficients of the second speech with a second preset duration, the lip-shape coefficients corresponding to each phoneme in the target style after style feature encoding are obtained. This method first uses the first phoneme sequence and first lip-shape coefficients of the unstyled first speech to ensure the original mapping relationship between semantic phonemes and lip-shape coefficients. Therefore, if lip-shape coefficients for a target virtual character with a style are to be obtained, only a small number of speech samples of the target virtual character need to be collected and the corresponding style encoding vectors can be obtained from these speech samples. In this way, the style encoding vectors style-encode each phoneme, and then, based on the original mapping relationship, the target lip-shape coefficients corresponding to each style-encoded phoneme can be obtained. In this method, the original mapping relationship between semantic phonemes and lip-shape coefficients is fundamental; only a single unstyled speech sample is required to obtain the original mapping relationship. If you need to obtain different lip-shape coefficients for different characters, you only need to obtain a short-term speech sample of a certain character to train different style encoding vectors. This greatly reduces the work cost of obtaining lip-shape coefficients with speaking styles and improves the efficiency of animation production. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A schematic flow chart of a method for obtaining a lip coefficient generation model provided in an embodiment of the present application;
[0035] Figure 2 A schematic diagram of the structure of a lip coefficient generation model provided in an embodiment of the present application;
[0036] Figure 3 A flow chart of a method for generating lip coefficients provided in an embodiment of the present application;
[0037] Figure 4 A schematic diagram of the structure of a device for obtaining a lip coefficient generation model provided in an embodiment of the present application;
[0038] Figure 5 A schematic structural diagram of a device for generating lip coefficients provided in an embodiment of the present application;
[0039] Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0040] The present application provides a method, device, electronic device, and storage medium for obtaining a lip-shape coefficient generation model. Long-term, styleless speech samples are used to train the mapping layer in a target lip-shape coefficient generation model, and short-term, speech-styled speech samples are used to train the style encoding vector in the target lip-shape coefficient generation model. This reduces the labor cost in obtaining lip-shape coefficients with a speaking style and improves animation production efficiency.
[0041] The following description sets forth many specific details to facilitate a full understanding of the present application. However, the present application can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present application. Therefore, the present application is not limited to the specific embodiments disclosed below.
[0042] The terms used in this application are for the purpose of describing specific embodiments only and are not intended to limit this application. Descriptive terms such as "a," "a," "a first," and "a second," etc., used in this application and the appended claims, are not intended to limit quantity or sequence, but are used to distinguish information of the same type from one another.
[0043] In the field of animation synthesis, achieving complete avatar animation requires synthesizing speech into lip-sync coefficients. This involves taking a speech input and synthesizing the corresponding lip-sync coefficients to ensure that the avatar's lip movements match the speech content. However, given that different avatars require different speaking styles, determining the lip-sync coefficients requires ensuring they correspond to the avatar's speaking style.
[0044] In the existing technology, in order to obtain lip-shape coefficients with speaking styles, for different speaking styles, it is necessary to record at least 30 minutes of speech data that retains the speaking style as training samples. Then, a large number of training samples are required to train the deep neural network. This method requires a lot of work costs to obtain speech motion capture data.
[0045] In order to solve the above problems, the present application provides a method for obtaining a lip-shape coefficient generation model to reduce the work cost in the process of obtaining the lip-shape coefficients that retain the speaking style and improve the production efficiency of animation.
[0046] The core of the method for obtaining the lip coefficient generation model provided in this application lies in:
[0047] During the sample collection process, since the style information of the lip shape can be obtained based on the lip shape coefficient, and when the style information is obtained based on the lip shape coefficient, only a small amount of speech collection is required.
[0048] Therefore, in the process of training the lip shape coefficient generation model, the present application uses the first phoneme sequence and the first lip shape coefficient of the first speech with a first preset duration, and the second lip shape coefficient of the second speech with a first target style as training samples to train the lip shape coefficient generation model.
[0049] The semantic information required for synthesizing lip shape coefficients is provided by the first phoneme sequence of the first speech of a first preset duration, and the style information required for synthesizing lip shape coefficients is provided by the lip shape coefficients of the second speech having a first target style. At the same time, based on the first phoneme sequence and the first lip shape coefficients of the first speech, the mapping relationship between different phonemes and different lip shape coefficients is determined. This method can not only ensure the mapping relationship between different semantic information through the first phoneme sequence and the first lip shape coefficients of the first speech, but also perform a small amount of speech collection according to the requirements of the first target style. This method reduces the work cost in the process of obtaining lip shape coefficients that retain the speaking style and improves the efficiency of animation production.
[0050] The embodiment of the present application first provides a method for obtaining a lip coefficient generation model. The implementing entity of this method can be various types of user terminals such as laptops, tablets, desktop computers, set-top boxes, mobile devices (for example, mobile phones, personal digital assistants, dedicated messaging devices, game consoles), or a combination of any two or more of these data processing devices, or a server.
[0051] Please refer to Figure 1 , Figure 1 This is a flow chart of a method for obtaining a lip shape coefficient generation model provided in one embodiment of the present application. The method includes the following steps S101 to S105:
[0052] Step S101: Obtain a first speech of a first preset duration and a first lip-shape coefficient corresponding to the first speech.
[0053] In an optional implementation manner of the present application, the first voice of the first preset duration may be a segment of voice data without a language style.
[0054] The first lip shape coefficient can be understood as data used to drive the change in the shape of the virtual character's mouth when the virtual character plays the first voice. In a specific application, the first lip shape coefficient can be obtained by collecting the lip shape change data of a specified virtual character.
[0055] Step S102: Determine a first phoneme sequence corresponding to the first speech.
[0056] Phonemes are the smallest speech units divided according to the natural properties of speech. From the perspective of human physiology, phonemes are generated based on pronunciation actions. For example, [ma] contains two pronunciation actions, [m] and [a], which are two different factors. In the process of specific application, all the contents in a language can be obtained by combining different factors. For example, among the phonemes involved in Chinese, the entire set of phonemes is {'ao','l','r','j','h','z','iao','t','m','uo','ei','er','o','v','s','ng','d','uai','w','ai','x','i','f','n','a','iu','u','ui','zh','ia','ch','ou','g','ie','k','sil','e','y','q','sh','ua','b','p','c','ue'}. Furthermore, the phoneme sequence of the first speech can be understood as being used to represent the content of speech information in the speech, that is, the first phoneme sequence is used to represent the semantic information of the first speech, which can be obtained by eliminating the speech timbre information and other content in the first speech.
[0057] In an optional implementation manner of the present application, the phoneme sequence of the first speech may be obtained by following steps 1 to 3:
[0058] Step S1, obtaining duration information of each phoneme in the first speech according to the pronunciation order of each phoneme in the first speech;
[0059] The duration information of each phoneme in the first voice can be understood as the duration information of each phoneme in the first voice when the designated virtual character plays the first voice. For example, the duration information of each phoneme can be expressed as: {'sil': 0ms~20ms;
[0060] 'w': 20ms~50ms; ...}, where 'sil' and 'w' represent different phoneme categories, 0ms~20ms means that the virtual character is reading the phoneme 'sil' during the time period 0ms~20ms at the beginning of the speech, and 20ms~50ms further means that the virtual character is reading the phoneme 'w' during the time period 20ms~50ms.
[0061] Step S2, determining the lip-sync frame rate information of the target virtual character;
[0062] In the field of imaging, frame rate refers to the number of frames transmitted per second. Furthermore, the target virtual character can be understood as the target of the stylized lip-sync coefficients generated by the lip-sync coefficient generation model. In practical applications, the target virtual character is also a virtual character. The lip-sync frame rate information of the target virtual character can be understood as the frame rate of the animation when the target virtual character plays a speech in the form of an animation.
[0063] Step S3: aligning the phonemes in the first speech according to the target virtual character's lip-syncing frame rate information and the duration information of the phonemes, and using the aligned phonemes in the first speech as the first phoneme sequence.
[0064] The aligning of the phonemes in the first speech according to the lip frame rate information of the target virtual character and the duration information of each phoneme refers to determining the number of frames of each phoneme in the animation frame according to the frame rate of the target virtual character when playing the voice animation and the pronunciation time of each phoneme.
[0065] For example, assuming that the frame rate of the target virtual character when playing the voice animation is 100fps, then the first voice with the phoneme time information {'sil': 0ms~20ms; 'w': 20ms~50ms; ...} is aligned, and the first phoneme sequence obtained can be further represented as {'sil', 'sil', 'w',
[0066] 'w', 'w'...}, that is, the phoneme 'sil' lasts for 20ms, and when the target virtual character plays the voice animation, the phoneme occupies 2 frames of the voice animation; the phoneme 'w' lasts for 30ms, and when the target virtual character plays the voice animation, the phoneme occupies 3 frames of the voice animation.
[0067] Step S103: Obtain a second speech of a second preset duration and a second lip-shape coefficient corresponding to the second speech.
[0068] The second speech is a speech with a target style, and the time of collecting the second speech is much shorter than that of collecting the first speech.
[0069] In the embodiment of the present application, the purpose is to train and obtain a lip-shape coefficient generation model that can synthesize lip-shape coefficients with style. Therefore, in the stage of collecting training samples, it is essential to collect style information of speech.
[0070] In an embodiment of the present application, the second voice is a carrier of the target style. For example, assuming that the target style represents the lively personality of the virtual character, the second voice can be a piece of voice data of another character with the same personality as the virtual character; for example, assuming that the target style represents the dialect characteristics of the virtual character, the second voice can be a piece of voice data of another character with the same dialect characteristics.
[0071] In an optional embodiment of the present application, the speech with the target style can also be obtained by relevant staff synthesizing the speech of the target virtual character. Similar to the above step S101, the target virtual character also refers to the object of action of the stylized lip coefficient generated based on the lip coefficient generation model.
[0072] Furthermore, the second lip-shape coefficient is primarily used to determine the stylistic characteristics of the target style. These stylistic characteristics are obtained during the training of the initial lip-shape coefficient generation model. To obtain these stylistic characteristics, only one to five minutes of speech data is required.
[0073] Step S104: Determine a second phoneme sequence corresponding to the second speech.
[0074] Similarly, this step can refer to the process of determining whether the first speech corresponds to the first phoneme sequence in the above step S102, and will not be described in detail here.
[0075] Step S105 : taking the first phoneme sequence and the first lip shape coefficient as a first training sample, taking the second phoneme sequence and the second lip shape coefficient as a second training sample, and using the first training sample and the second training sample to train an initial target lip shape coefficient generation model.
[0076] The first training samples are used to train the mapping layer in the target lip-shape coefficient generation model. This mapping layer is used to determine the mapping relationship between each phoneme in the first phoneme sequence and each lip-shape coefficient in the first lip-shape coefficient. The second training samples are used to train the style encoding layer in the target lip-shape coefficient generation model. The style encoding layer is used to extract the style features corresponding to the second phoneme sequence based on the second lip-shape coefficient and the mapping relationship.
[0077] In the specific application process, this application uses machine learning (ML) to train the initial lip coefficient generation model. Machine learning (a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines) is specifically used to study how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their own new abilities. Machine learning generally includes artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning and other technologies. Machine learning is a branch of artificial intelligence (AI) technology.
[0078] Before the initial lip shape coefficient generation model is trained using the first phoneme sequence, the first lip shape coefficient, and the second lip shape coefficient, since the first phoneme sequence exists in the form of phoneme classification features (for example, the phoneme sequence can be {'sil', 'sil', 'w', 'w',
[0079] 'w'...}, where 'sil' and 'w' represent different factor categories, respectively), but in each phoneme sequence, different phonemes are not continuous, but discrete and disordered. Therefore, in an optional embodiment of the present application, the first phoneme sequence needs to be further preprocessed.
[0080] Specifically, the process of preprocessing the first phoneme sequence is: performing feature encoding on the first phoneme sequence to obtain a phoneme feature vector corresponding to the first phoneme sequence.
[0081] In an optional embodiment of the present application, the phoneme feature vector can be obtained in the form of one-hot encoding. One-hot encoding mainly uses an N-bit register to encode N states. In the embodiment of the present application, by The one-hot encoding form of processing the first phoneme sequence is to The discrete factor categories in each phoneme sequence are converted into a vector form, so as to learn the mapping relationship between different phonemes and different mouth shape coefficients in the process of training the initial mouth shape coefficient generation model.
[0082] Specifically, the first lip shape coefficient is used to determine the correspondence between different phonemes and different lip shape coefficients in combination with the first phoneme sequence; and the second lip shape coefficient is used to determine the style features of the target style of the second speech.
[0083] Furthermore, in order to facilitate understanding of the method for obtaining the lip-shape coefficient generation model provided in this application, the training process of the lip-shape coefficient generation model is specifically introduced below in combination with the specific structure of the lip-shape coefficient generation model.
[0084] Please refer to Figure 2 , Figure 2 A schematic diagram of the structure of a lip coefficient generation model provided in an embodiment of the present application.
[0085] The lip coefficient generation model includes: a semantic coding network 201 , a style coding layer 202 , a feature transmission module 203 , and a mapping layer 204 .
[0086] The lip shape coefficient generation model is mainly used to output lip shape coefficients with style characteristics corresponding to the input voice information. The lip shape coefficients are specifically used to drive the change of the lip shape state of the virtual character so that the lip shape state of the virtual character when playing the voice can match the voice style characteristics of the virtual character.
[0087] In practical applications, the changes in the avatar's lip shape are primarily controlled by the semantic information of the avatar's voice and the avatar's style information. Therefore, when training the initial lip shape coefficient generation model, it is necessary to complete the entire training process based on training sample data that contains both semantic information and avatar style information.
[0088] In the embodiment of the present application, the first phoneme sequence is mainly used to provide semantic information required for training the initial lip shape coefficient generation model, and the second lip shape coefficient is used to provide style information required for training the initial lip shape coefficient generation model.
[0089] Specifically, the semantic information in the first phoneme sequence can be extracted through the semantic coding network 201. The semantic coding network 201 can be understood as a convolutional neural network. In an optional embodiment of the present application, the semantic coding network 201 includes a 6-layer 1-dimensional convolutional neural network, and the kernel size of each network is 3 and the step size is 1.
[0090] Furthermore, after the semantic encoding network 201 obtains the semantic information in the first phoneme sequence, the semantic information and the first lip shape coefficient are input into the mapping layer 204 so that the mapping layer can learn the mapping relationship between different lip shape coefficients and each semantic information.
[0091] Furthermore, the style encoding layer 202 is used to extract the style features of the target style from the second lip coefficient and store the style features of the target style. Specifically, the style encoding layer 202 mainly consists of a style encoding vector for storing the style features of the target style. The style encoding vector can be represented by T = {μ, σ}, where μ represents the translation coefficient of the target style and σ represents the scaling coefficient of the target style.
[0092] In a specific application process, the style coding layer 202 can train the ability of the style coding layer 202 to extract style features while training the initial lip coefficient generation model. Specifically, the training of the style feature learning module requires obtaining the second phoneme sequence of the second speech with the target style, which is the same as the above-mentioned first phoneme sequence. The second phoneme sequence is used to represent the semantic information of the second speech.
[0093] Furthermore, the second phoneme sequence can be obtained in a manner similar to steps S1 to S3 above, that is, first, according to the pronunciation order of each phoneme in the second speech, the duration information of each phoneme in the second speech is obtained; secondly, the lip frame rate information of the target virtual character is determined; finally, according to the lip frame rate information of the target virtual character and the duration information of each phoneme, the phonemes in the second speech are aligned, and the phonemes in the second speech after the alignment are used as the second phoneme sequence.
[0094] Afterwards, the second phoneme sequence is feature-encoded to obtain a phoneme feature vector corresponding to the second phoneme sequence; finally, the phoneme feature vector is input into the semantic encoding network 201, and the semantic features in the second phoneme sequence are extracted through the semantic encoding network 201 (for the sake of ease of description, the semantic features are represented by the symbol F below).
[0095] After obtaining the semantic features in the second phoneme sequence, the semantic features are input into the style coding layer 202, and the style coding vector T = {μ, σ} is used by the style coding layer 202 to perform AdaIN operation on the feature F to obtain the stylized target semantic features. Wherein, 256 represents the dimension of the stylized target semantic feature, and n represents the length of the stylized target semantic feature; finally, the stylized target semantic feature is input into the subsequent mapping layer.
[0096] The operation of AdaIN is expressed by the following formula (1):
[0097]
[0098] Wherein, m(F) represents the mean of the semantic feature, and v(F) represents the standard deviation of the semantic feature.
[0099] Furthermore, after different semantic information and target style information are learned based on the first phoneme sequence and the second lip shape coefficient, a mapping relationship between different semantics and different lip shape coefficients needs to be learned.
[0100] In the embodiment of the present application, the mapping relationship between the different semantics and the different lip coefficients is obtained by the mapping layer 204 .
[0101] Furthermore, in order to determine the degree of training of the lip coefficient generation network, the loss function of the lip coefficient generation model can be constructed by the following formula (2):
[0102]
[0103] Wherein, B represents the lip shape coefficient of the sample speech input into the lip shape coefficient generation model; represents the lip coefficients of the sample speech with the target style after being processed by the lip coefficient generation model, where B∈R m×n , m represents the characteristic dimension of the first lip-shape coefficient, and n represents the fragmentation degree of the lip-shape coefficient.
[0104] Based on the above description of the method for obtaining the lip shape coefficient generation model, it can be seen that in the above training samples, the second lip shape coefficient is used to determine the style characteristics of the first target style, and the first phoneme sequence and the first lip shape coefficient are used to provide semantic information and the mapping relationship between different semantic information and different lip shape coefficients.
[0105] Furthermore, for a lip shape coefficient generation model that generates lip shape coefficients of different styles, it is only necessary to re-obtain speech with style characteristics when collecting samples.
[0106] As can be seen from the above description, to obtain lip shape coefficients of different styles, there is no need to train the mapping layer. Instead, it is sufficient to collect short bursts of speech of different styles based on the first speech and train the style coding vectors in the style coding layer. Therefore, in an optional embodiment of the present application, the method for obtaining the lip shape coefficient generation model further includes the following steps S106 and S108:
[0107] Step S106: Obtain a third speech with another target style and a third lip shape coefficient corresponding to the third speech.
[0108] Step S107: Obtain a third phoneme sequence corresponding to the third speech.
[0109] Step S108, after the mapping layer of the lip coefficient generation model is trained using the first phoneme sequence and the first lip coefficient, the style encoding layer of the lip coefficient generation model is trained using the third phoneme sequence and the third lip coefficient to obtain a target lip coefficient generation model for generating a target speech with another target style.
[0110] The third phoneme sequence is input into a lip-shape coefficient generation model for generating lip-shape coefficients having another target style, and a third lip-shape coefficient having the other target style is output by the lip-shape coefficient generation model. The third speech having the other target style is similar to the second speech mentioned in the above step, except that the style of the third speech is different from the style of the second speech.
[0111] That is to say, in the method for obtaining the lip shape coefficient generation model provided in the embodiment of the present application, for the lip shape coefficients of different style types, it is only necessary to collect the speech corresponding to the style in a targeted manner to obtain the lip shape coefficients of the speech to obtain the corresponding style characteristics, without repeatedly using a large amount of speech information to determine the mapping relationship between various semantic information and the lip shape coefficients.
[0112] In summary, the method for obtaining the lip shape coefficient generation model provided in the embodiment of the present application uses the first phoneme sequence and first lip shape coefficient of the first speech with a preset duration, and the second lip shape coefficient of the second speech with a target style as training samples to train the lip shape coefficient generation model.
[0113] The semantic information required for synthesizing lip-shape coefficients is provided by the first phoneme sequence of a first speech of a first preset duration, and the stylistic information required for synthesizing lip-shape coefficients is provided by the lip-shape coefficients of a second speech of a target style. Simultaneously, based on the first phoneme sequence and the first lip-shape coefficients of the first speech, a mapping relationship between different phonemes and different lip-shape coefficients is determined. This method not only ensures the mapping relationship between different semantic information using the first phoneme sequence and the first lip-shape coefficients of the first speech, but also allows for a small amount of speech to be collected for the target style as required by the speech style. This reduces the labor cost of obtaining lip-shape coefficients that preserve the speaking style and improves animation production efficiency.
[0114] This application also provides a method for generating lip coefficients, please refer to Figure 3 , Figure 3 This is a flow chart of a method for generating lip coefficients provided in another embodiment of the present application. Since this method embodiment is basically similar to the above Mouth-shape coefficient generation model The method is described briefly, and the relevant parts are referred to in this application. The relevant introduction of the method for obtaining the lip shape coefficient generation model is sufficient and will not be repeated here.
[0115] The method includes the following steps S301 and S302:
[0116] Step S301: Obtain a target phoneme sequence of a target speech.
[0117] The target phoneme sequence is used to represent the semantic information of the target speech.
[0118] Step S302: input the target phoneme sequence into a target lip-shape coefficient generation model for generating lip-shape coefficients with a target style, and obtain target lip-shape coefficients with a target style corresponding to the target speech.
[0119] In an optional implementation manner of the present application, the target phoneme sequence of the target speech may be obtained in the following manner:
[0120] According to the pronunciation order of each phoneme in the target speech, the duration information of each phoneme in the target speech is obtained;
[0121] Determining the lip-sync frame rate information of the target virtual character corresponding to the target voice;
[0122] According to the lip frame rate information of the target virtual character and the duration information of each phoneme, each phoneme in the target speech is processed, and each phoneme in the processed target speech is used as a target phoneme sequence.
[0123] In an optional embodiment of the present application, the method further includes:
[0124] The target phoneme sequence is feature-encoded to obtain a phoneme feature vector corresponding to the target phoneme sequence.
[0125] Since the target lip-shape coefficient generation model is a two-branch neural network, when inputting the target lip-shape coefficients into the target lip-shape coefficient generation model, the speaking style of the target virtual character corresponding to the target speech must first be determined. If the target virtual character has no speaking style requirements, the target phoneme sequence corresponding to the target speech is directly input into the mapping layer of the target lip-shape coefficient generation model through the first branch. If the target virtual character has a speaking style, the target speech is input into the target lip-shape coefficient generation model with that style and first input into the style encoding layer of the target lip-shape coefficient generation model through the second branch. The output of the intermediate style encoding layer is then input into the mapping layer of the target lip-shape coefficient generation model to obtain the final target lip-shape coefficients.
[0126] The present application also provides a device for obtaining a lip coefficient generation model. Figure 4 , Figure 4 This is a schematic diagram of the structure of a device for obtaining a lip-shape coefficient generation model provided in an embodiment of the present application. Because this device embodiment is substantially similar to the aforementioned method embodiment for obtaining a lip-shape coefficient generation model, the description is relatively brief. For relevant details, please refer to the description of the aforementioned method embodiment. The following description of this device embodiment is merely illustrative.
[0127] The device for obtaining the lip coefficient generation model includes:
[0128] The acquisition unit 401 is configured to acquire a first speech of a first preset duration and a first lip shape coefficient corresponding to the first speech. The first speech is a styleless speech.
[0129] The determining unit 402 is configured to determine a first phoneme sequence corresponding to the first speech, and determine the first phoneme sequence and the first lip shape coefficient as a first training sample.
[0130] The acquisition unit 401 is further configured to acquire a second speech of a second preset duration and a second lip shape coefficient corresponding to the second speech, wherein the second speech is a speech of a target style and the second preset duration is shorter than the first preset duration.
[0131] The determining unit 402 is further configured to determine a second phoneme sequence corresponding to the second speech, and determine the second phoneme sequence and the second lip shape coefficient as a second training sample.
[0132] The training unit 403 is configured to train a target lip-shape coefficient generation model using the first training sample and the second training sample, wherein the target lip-shape coefficient generation model is configured to generate target lip-shape coefficients having a target style corresponding to the target speech.
[0133] The first training samples are used to train the mapping layer in the target lip-shape coefficient generation model. The mapping layer is used to determine the mapping relationship between each phoneme in the first phoneme sequence and each lip-shape coefficient in the first lip-shape coefficient. The second training samples are used to train the style encoding layer in the target lip-shape coefficient generation model. The style encoding layer is used to extract the style features corresponding to the second phoneme sequence based on the second lip-shape coefficient and the mapping relationship.
[0134] In an optional embodiment, the training unit 403 is specifically configured to directly input the first phoneme sequence in the first training sample into the mapping layer to obtain a first output result corresponding to the first phoneme sequence. A first loss value is determined based on the first output result and the first lip shape coefficient. The network layer coefficients of the mapping layer are updated based on the first loss value.
[0135] In an optional embodiment, the training unit 403 is specifically configured to, after the mapping layer is trained using the first training sample, input the second phoneme sequence in the second training sample into the style encoding layer to obtain an intermediate output result corresponding to the second phoneme sequence. Input the intermediate output result into the mapping layer to obtain a second output result corresponding to the second phoneme sequence. Determine a second loss value based on the second output result and the second lip shape coefficient. Update the network layer coefficients of the style encoding layer based on the second loss value.
[0136] In an optional embodiment, the style encoding layer includes a trainable style encoding vector.
[0137] The training unit 403 is specifically configured to obtain a second phoneme feature vector corresponding to the second phoneme sequence. The semantic features corresponding to the second phoneme sequence are extracted based on the phoneme feature vector. The semantic features are encoded using the style encoding vector to obtain stylized target semantic features. The stylized target semantic features are input into the mapping layer.
[0138] In an optional embodiment, the first phoneme sequence or the second phoneme sequence is obtained by:
[0139] Obtain the pronunciation order of each phoneme and the pronunciation duration information of each phoneme in the first voice or the second voice.
[0140] The lip-sync frame rate information of the target avatar is determined, and the target avatar corresponds to the target style.
[0141] The phonemes in the first speech or the second speech are aligned according to the lip-sync frame rate information of the target virtual character and the duration information of the phonemes in the first speech or the second speech.
[0142] The first phoneme sequence or the second phoneme sequence is determined according to the arrangement order of the phonemes in the first speech or the second speech after the alignment process.
[0143] In an optional embodiment, the target lip coefficient generation model further includes a semantic encoding network.
[0144] The acquisition unit 401 is further configured to perform feature encoding on the first phoneme sequence using a semantic coding network to obtain a first phoneme feature vector corresponding to the first phoneme sequence, and to perform feature encoding on the second phoneme sequence using the semantic coding network to obtain a second phoneme feature vector corresponding to the second phoneme sequence.
[0145] The training unit 403 is specifically configured to directly input the first phoneme feature vector into the mapping layer.
[0146] In an optional embodiment, the acquiring unit 401 is specifically configured to cut the first phoneme sequence into a plurality of first phoneme sequence segments according to a preset phoneme sequence length, perform feature encoding on each of the plurality of first phoneme sequence segments, and obtain a first phoneme feature vector corresponding to each first phoneme sequence segment.
[0147] In an optional embodiment, the acquiring unit 401 is specifically configured to truncate the second phoneme sequence into at least one second phoneme sequence segment according to a preset phoneme sequence length, perform feature encoding on each of the at least one second phoneme sequence segments, and obtain a second phoneme feature vector corresponding to each second phoneme sequence segment.
[0148] It should be noted that the information interaction and execution process between the modules / units in the device for obtaining the lip coefficient generation model are the same as those in the present application. Figures 1 to 3 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0149] The present application also provides a device for generating lip coefficients. Figure 5 , Figure 5 This is a schematic diagram of the structure of a device for generating lip coefficients provided in an embodiment of the present application. Since this device embodiment is substantially similar to the aforementioned method embodiment for generating lip coefficients, the description is relatively simple. For relevant details, please refer to the description of the aforementioned method embodiment. The following description of this device embodiment is merely illustrative.
[0150] The device for generating the lip coefficient includes:
[0151] The acquisition unit 501 is configured to acquire a target phoneme sequence of a target speech, where the target phoneme sequence is used to represent semantic information of the target speech.
[0152] The processing unit 502 is used to input the target phoneme sequence into the target lip-shape coefficient generation model to obtain the target lip-shape coefficient output by the target lip-shape coefficient generation model, and the target lip-shape coefficient is used to drive the mouth movement changes of the target virtual character.
[0153] Among them, the target lip coefficient generation model is composed of Figure 4 The embodiment shown provides a device for obtaining a lip-shape coefficient generation model.
[0154] In an optional implementation, the device for generating lip coefficients further includes a determination unit 503 .
[0155] The determining unit 503 is configured to determine the style requirement of the target virtual character.
[0156] The processing unit 502 is specifically configured to directly input the target phoneme sequence into the mapping layer of the target lip-shape coefficient generation model if the style requirement is no style.
[0157] If the style requirement is the target style, the target phoneme sequence is input into the style encoding layer of the target lip coefficient generation model, and then the intermediate output result of the style encoding layer is input into the mapping layer.
[0158] It should be noted that the information interaction and execution process between the modules / units in the lip coefficient generation device are the same as those in the present application. Figures 1 to 3 The corresponding method embodiments are based on the same concept. For specific contents, please refer to the description in the method embodiments shown above in this application, which will not be repeated here.
[0159] The present application also provides an electronic device. Figure 6 , Figure 6 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application.
[0160] The electronic device includes:
[0161] At least one processor 601, at least one memory 603, at least one communication interface 602, and at least one communication bus 604;
[0162] Optionally, the communication interface 602 may be an interface of a communication module, such as an interface of a GSM module;
[0163] The processor 601 may be a CPU, or an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement the embodiments of the present application;
[0164] The memory 603 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.
[0165] Among them, the memory 603 is a program for storing the method, and when the program is read and executed by the processor, the method provided by the method embodiment of the present application is executed.
[0166] Another embodiment of the present application also provides a computer storage medium, which stores a computer program. When the program is executed, it implements the method provided in the above method embodiment.
[0167] Although the present application is disclosed as above with the preferred embodiments, it is not intended to limit the present application. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the scope of protection of the present application shall be based on the scope defined by the claims of the present application.
[0168] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0169] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0170] 1. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include non-transitory media such as modulated data signals and carrier waves.
[0171] 2. Those skilled in the art will appreciate that the embodiments of the present application may be provided as systems or electronic devices. Therefore, the present application may take the form of a fully hardware embodiment, a fully software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
Claims
1. A method for obtaining a lip coefficient generation model, characterized in that: include: Obtaining a first speech of a first preset duration and a first lip shape coefficient corresponding to the first speech; The first voice is a styleless voice; Determining a first phoneme sequence corresponding to the first speech, and determining the first phoneme sequence and the first lip shape coefficient as a first training sample; Obtaining a second speech of a second preset duration and a second lip shape coefficient corresponding to the second speech; The second speech is a speech with a target style; The second preset duration is shorter than the first preset duration; Determining a second phoneme sequence corresponding to the second speech, and determining the second phoneme sequence and the second lip shape coefficient as a second training sample; Using the first training sample and the second training sample to train a target lip shape coefficient generation model; wherein the target lip shape coefficient generation model is used to generate target lip shape coefficients having the target style corresponding to the target speech; Among them, the first training sample is used to train the mapping layer in the target lip shape coefficient generation model; the mapping layer is used to determine the mapping relationship between each phoneme in the first phoneme sequence and each lip shape coefficient in the first lip shape coefficient; the second training sample is used to train the style encoding layer in the target lip shape coefficient generation model; the style encoding layer is used to extract the style features corresponding to the second phoneme sequence based on the second lip shape coefficient and the mapping relationship.
2. The method according to claim 1, characterized in that The step of using the first training sample and the second training sample to train a target lip coefficient generation model includes: Inputting the first phoneme sequence in the first training sample directly into the mapping layer to obtain a first output result corresponding to the first phoneme sequence; determining a first loss value according to the first output result and the first lip coefficient; The network layer coefficients of the mapping layer are updated according to the first loss value.
3. The method according to claim 2, characterized in that The step of using the first training sample and the second training sample to train a target lip coefficient generation model includes: After the mapping layer is trained using the first training sample, the second phoneme sequence in the second training sample is input into the style encoding layer to obtain an intermediate output result corresponding to the second phoneme sequence; Inputting the intermediate output result into the mapping layer to obtain a second output result corresponding to the second phoneme sequence; determining a second loss value according to the second output result and the second lip coefficient; The network layer coefficients of the style encoding layer are updated according to the second loss value.
4. The method according to claim 3, characterized in that The style coding layer includes a trainable style coding vector; and inputting the second phoneme sequence in the second training sample into the style coding layer to obtain an intermediate output result corresponding to the second phoneme sequence includes: Obtaining a second phoneme feature vector corresponding to the second phoneme sequence; Extracting semantic features corresponding to the second phoneme sequence according to the phoneme feature vector; Encoding the semantic feature using the style encoding vector to obtain a stylized target semantic feature; Inputting the intermediate output result into the mapping layer includes: The stylized target semantic features are input into the mapping layer.
5. The method according to claim 1, wherein The first phoneme sequence or the diphoneme sequence is obtained by: Obtaining the pronunciation order of each phoneme in the first speech or the second speech and the pronunciation duration information of each phoneme; Determining lip-sync frame rate information of a target virtual character; the target virtual character corresponds to the target style; performing alignment processing on each phoneme in the first speech or the second speech according to the lip-sync frame rate information of the target virtual character and the duration information of each phoneme in the first speech or the second speech; The first phoneme sequence or the second phoneme sequence is determined according to the arrangement order of the phonemes in the first speech or the second speech after the alignment process.
6. The method according to any one of claims 2 to 4, characterized in that The first phoneme sequence or the diphoneme sequence is obtained by: Obtaining the pronunciation order of each phoneme in the first speech or the second speech and the pronunciation duration information of each phoneme; Determining lip-sync frame rate information of a target virtual character; the target virtual character corresponds to the target style; performing alignment processing on each phoneme in the first speech or the second speech according to the lip-sync frame rate information of the target virtual character and the duration information of each phoneme in the first speech or the second speech; The first phoneme sequence or the second phoneme sequence is determined according to the arrangement order of the phonemes in the first speech or the second speech after the alignment process.
7. The method according to claim 6, characterized in that The target lip coefficient generation model also includes a semantic coding network; the method also includes: Performing feature encoding on the first phoneme sequence using the semantic coding network to obtain a first phoneme feature vector corresponding to the first phoneme sequence; and performing feature encoding on the second phoneme sequence using the semantic coding network to obtain a second phoneme feature vector corresponding to the second phoneme sequence; The step of directly inputting the first phoneme sequence in the first training sample into the mapping layer comprises: The first phoneme feature vector is directly input into the mapping layer.
8. The method according to claim 7, characterized in that The step of performing feature encoding on the first phoneme sequence by using the semantic coding network to obtain a first phoneme feature vector corresponding to the first phoneme sequence includes: According to a preset phoneme sequence length, cutting the first phoneme sequence into a plurality of first phoneme sequence segments; Feature encoding is performed on each of the plurality of first phoneme sequence segments to obtain the first phoneme feature vector corresponding to each first phoneme sequence segment.
9. The method according to claim 8, characterized in that The performing feature encoding on the second phoneme sequence by using the semantic coding network to obtain the second phoneme feature vector corresponding to the second phoneme sequence includes: According to the preset phoneme sequence length, cutting the second phoneme sequence into at least one second phoneme sequence segment; Feature encoding is performed on the at least one second phoneme sequence segment to obtain the second phoneme feature vector corresponding to each second phoneme sequence segment.
10. A method for generating lip coefficients, characterized in that: The generation method comprises: Obtaining a target phoneme sequence of a target speech, wherein the target phoneme sequence is used to represent semantic information of the target speech; Inputting the target phoneme sequence into a target lip-shape coefficient generation model to obtain target lip-shape coefficients output by the target lip-shape coefficient generation model; the target lip-shape coefficients are used to drive changes in the mouth movements of the target virtual character; The target lip-shape coefficient generation model is generated according to the method for obtaining the lip-shape coefficient generation model according to any one of claims 1 to 9.
11. The method according to claim 10, characterized in that The method further comprises: Determining the style requirements of the target virtual character; If the style requirement is no style, the target phoneme sequence is directly input into the mapping layer of the target lip coefficient generation model; If the style requirement is the target style, the target phoneme sequence is input into the style encoding layer of the target lip coefficient generation model, and then the intermediate output result of the style encoding layer is input into the mapping layer.
12. A device for obtaining a lip-shape coefficient generation model, characterized in that: include: An acquiring unit, configured to acquire a first speech of a first preset duration and a first lip shape coefficient corresponding to the first speech; The first voice is a styleless voice; a determining unit, configured to determine a first phoneme sequence corresponding to the first speech, and determine the first phoneme sequence and the first lip shape coefficient as a first training sample; The acquisition unit is further configured to acquire a second speech of a second preset duration and a second lip shape coefficient corresponding to the second speech; the second speech is a speech having a target style; and the second preset duration is shorter than the first preset duration; The determining unit is further configured to determine a second phoneme sequence corresponding to the second speech, and determine the second phoneme sequence and the second lip shape coefficient as a second training sample; A training unit, configured to train a target lip shape coefficient generation model using the first training sample and the second training sample; wherein the target lip shape coefficient generation model is configured to generate target lip shape coefficients having the target style corresponding to the target speech; Among them, the first training sample is used to train the mapping layer in the target lip shape coefficient generation model; the mapping layer is used to determine the mapping relationship between each phoneme in the first phoneme sequence and each lip shape coefficient in the first lip shape coefficient; the second training sample is used to train the style encoding layer in the target lip shape coefficient generation model; the style encoding layer is used to extract the style features corresponding to the second phoneme sequence based on the second lip shape coefficient and the mapping relationship.
13. A device for generating lip coefficients, characterized in that: include: an acquiring unit, configured to acquire a target phoneme sequence of a target speech, wherein the target phoneme sequence is used to represent semantic information of the target speech; a processing unit, configured to input the target phoneme sequence into a target lip-shape coefficient generation model to obtain target lip-shape coefficients output by the target lip-shape coefficient generation model; the target lip-shape coefficients are used to drive changes in the mouth movements of the target virtual character; Wherein, the target lip coefficient generation model is obtained according to the device provided in claim 12.
14. An electronic device, characterized in that: include: processor; A memory for storing a method program, wherein when the program is read and executed by the processor, the method according to any one of claims 1 to 11 is executed.
15. A computer storage medium, characterized in that The computer storage medium stores a computer program, and when the program is executed, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Voice-based mouth shape animation synthesis device and method and readable storage medium
CN108763190A
Speech synthesis method, model training method, equipment and storage medium
CN114283783A