Method for generating a facial image based on mouth shape, method and device for training a model
The method and device enhance the accuracy of facial image generation by using speaking rate and semantic features to adapt mouth shapes in digital humans, addressing challenges in matching mouth shapes with audio data and improving image quality.
Patent Information
- Application Number
- JP2024099687
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-08-17
- Filing Date
- 2024-06-20
- Publication Date
- 2025-08-27
- Estimated Expiration
- 2044-06-20
AI Technical Summary
Existing digital human applications face challenges in accurately matching mouth shapes in facial images with audio data, particularly when dealing with changes in speaking rate, leading to issues such as character omission and multiple reading.
A method and device for generating facial images based on mouth shapes, utilizing speaking rate and semantic features to process facial images, and a model training process that includes feature extraction and determination of speaking rate and semantic features to accurately match mouth shapes with audio data.
Improves the accuracy and efficiency of generating facial images by accurately adapting mouth shapes to audio data, reducing issues like character omission and multiple reading, especially at varying speaking rates.
Smart Images

Figure 0007730402000001 
Figure 0007730402000002 
Figure 0007730402000003
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to the field of cloud computing and digital humans in the field of artificial intelligence, and in particular to a method for generating a facial image based on mouth shape, a method and a device for training a model. [Background technology]
[0002] With the rapid development of artificial intelligence technology, digital human applications have become a mainstream research topic. The digital human's face can be changed by voice, for example, the facial expression and mouth shape in the digital human's facial image can be changed by voice.
[0003] One of the key technologies in digital human applications is the technology of driving facial mouth shapes with audio, and how to accurately match the mouth shapes in facial images with audio data is a technical challenge that needs to be solved as soon as possible. Summary of the Invention [Problem to be solved by the invention]
[0004] The present disclosure provides a method for generating a facial image based on mouth shape, a method and a device for training a model. [Means for solving the problem]
[0005] According to a first aspect of the present disclosure, there is provided a method for generating a face image based on a mouth shape, the method for generating a face image based on a mouth shape comprising: acquiring recognition target audio data and a preset face image; determining audio features of the recognition target audio data, the audio features including speaking rate features and semantic features; and processing the predetermined face image based on the speech rate feature and the semantic feature to generate a face image having a mouth shape.
[0006] According to a second aspect of the present disclosure, there is provided a method for training a face and mouth typing model, the method comprising: acquiring training target image data and a predetermined facial image, wherein the training target image data includes training target audio data and a training target facial image, and the training target facial image has a mouth shape corresponding to the training target audio data; determining audio features of the training subject audio data, the audio features including speaking rate features and semantic features; training an initial face and mouth shape determination model based on the speech rate feature, the semantic feature, and the predetermined face image to obtain a face image with a mouth shape; and determining that a training-completed face and mouth shape determination model is obtained if the face image having the mouth shape matches the training target face image.
[0007] According to a third aspect of the present disclosure, there is provided an apparatus for generating a face image based on a mouth shape, the apparatus for generating a face image based on a mouth shape comprising: a data acquisition unit for acquiring audio data to be recognized and a predetermined face image; a feature determination unit used for determining audio features of the recognition target audio data, the audio features including speech rate features and semantic features; and an image generating unit for processing the predetermined face image based on the speech rate feature and the semantic feature to generate a face image having a mouth shape.
[0008] According to a fourth aspect of the present disclosure, there is provided an apparatus for training a face and mouth typing model, the apparatus for training the face and mouth typing model comprising: an image acquisition unit used to acquire training target image data and a predetermined face image, the training target image data including training target audio data and a training target face image, and the training target face image has a mouth shape corresponding to the training target audio data; a feature extraction unit used for determining audio features of the training target audio data, the audio features including speaking rate features and semantic features; a model training unit, which is used to train an initial face and mouth shape determination model according to the speech rate feature, the semantic feature and the predetermined face image, to obtain a face image with a mouth shape; and a model obtaining unit, which is used for determining that a face image with a mouth shape matches the training target face image, and obtains a face mouth shape determination model that has been trained.
[0009] According to a fifth aspect of the present disclosure, there is provided an electronic device, the electronic device comprising: at least one processor; and a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, cause the at least one processor to perform the methods described in the first and second aspects of the present disclosure.
[0010] According to a sixth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium having computer instructions stored thereon, ,Ko The computer instructions are for causing the computer to carry out the methods of the first and second aspects of the present disclosure.
[0011] According to a seventh aspect of the present disclosure, a computer program M Provide ,beforeWhen the computer program is executed by a processor, the methods according to the first and second aspects of the present disclosure are realized.
[0012] The technology of the present disclosure improves the accuracy of generating a facial image based on mouth shape.
[0013] It should be understood that the contents set forth herein are not intended to determine key or critical features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily apparent from the following specification. [Brief explanation of the drawings]
[0014] The drawings are used for better understanding of the present invention and are not intended to limit the present disclosure. [Figure 1] 1 is a flowchart of a method for generating a face image based on a mouth shape provided by an embodiment of the present disclosure. [Figure 2] 1 is a flowchart of a method for generating a face image based on a mouth shape provided by an embodiment of the present disclosure. [Figure 3] 1 is a flowchart of a method for generating a face image based on a mouth shape provided by an embodiment of the present disclosure. [Figure 4] 1 is a flowchart of a method for training a face and mouth type determination model provided by an embodiment of the present disclosure. [Figure 5] 1 is a flowchart of a method for training a face and mouth type determination model provided by an embodiment of the present disclosure. [Figure 6] FIG. 1 is a configuration diagram of an apparatus for generating a facial image based on a mouth shape provided by an embodiment of the present disclosure. [Figure 7] FIG. 1 is a configuration diagram of an apparatus for generating a facial image based on a mouth shape provided by an embodiment of the present disclosure. [Figure 8] FIG. 1 is a configuration diagram of an apparatus for training a face and mouth type determination model provided by an embodiment of the present disclosure. [Figure 9]FIG. 1 is a block diagram of an electronic device for implementing a method for generating a facial image based on mouth shape and a method for training a model according to an embodiment of the present disclosure. [Figure 10] FIG. 1 is a block diagram of an electronic device for implementing a method for generating a facial image based on mouth shape and a method for training a model according to an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0015]
[0023] Exemplary embodiments of the present disclosure will now be described with reference to the drawings. Various details of the embodiments of the present disclosure are included for ease of understanding, but they should be considered merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, in the following description, for the sake of clarity and conciseness, descriptions of well-known functions and structures will be omitted.
[0016] In current digital human applications, one of the core technologies is to drive the facial mouth shape with audio, that is, to change the mouth shape in the facial image through audio data, and then adapt the mouth shape in the facial image to the audio data. Therefore, how to drive the facial mouth shape more realistically and accurately is a technical issue that needs to be solved as soon as possible.
[0017] Related Methods for generating facial images based on mouth shapes have difficulty dealing with changes in speaking rate, and the speaking rate of audio data has a significant impact on mouth shapes. When the same sentence is spoken at different speaking rates, the corresponding mouth shapes may be completely different. When the speaking rate is slow, the mouth shapes for each character can be perfectly aligned with the pronunciation. However, when the speaking rate increases, the mouth shapes in the facial image do not accelerate at an equal rate, and it is possible that one mouth shape cannot be completed in time before the next character needs to be pronounced. This results in changes in the mouth shapes of many characters, leading to various phenomena such as "character omission" and "multiple reading," and many mouth shapes being lost, merged, or simplified, which affects the accuracy of facial image generation.
[0018] The present disclosure provides a method for generating a facial image based on mouth shape, a method and a device for training a model, which are applied in the fields of cloud computing and digital humans in the field of artificial intelligence, to improve the accuracy of generating a facial image with mouth shape.
[0019] It should be noted that the model in this example is not targeted to any specific user and does not reflect any personal information of any specific user. Furthermore, the facial images in this example are from a public dataset.
[0020] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision and disclosure of relevant users' personal information shall all comply with the provisions of relevant laws and regulations and shall not violate public order and morals.
[0021] To help readers better understand the implementation principles of the present disclosure, the embodiments will be further broken down with reference to the following FIGS. 1 to 10. FIG.
[0022] 1 is a flowchart of a method for generating a facial image based on a mouth shape provided by an embodiment of the present disclosure, which can be performed by an apparatus for generating a facial image based on a mouth shape. As shown in FIG. 1, the method includes the following steps:
[0023] S101: Acquire audio data to be recognized and a preset face image.
[0024] For example, the face of the digital human can be designed in advance, for example, the shape of the digital human's face, eyes, nose, mouth, etc., can be designed, and a preset facial image can be generated. The digital human can change the shape of its mouth based on the preset facial image, for example, in the preset facial image, the digital human's mouth is closed, and the shape of the digital human's mouth can change as audio data is transmitted.
[0025] The audio data to be recognized is prepared in advance, and in the facial image of the digital human, the mouth shape needs to change according to the audio data to be recognized. Pre-set audio data to be recognized and a pre-set facial image are acquired. The audio data to be recognized is an audio stream, and the pre-set facial image may be a two-dimensional or three-dimensional image.
[0026] S102: determining audio features of the recognition target audio data, where the audio features include speaking rate features and semantic features;
[0027] For example, after obtaining the audio data to be recognized, feature extraction is performed on the audio data to be recognized to obtain audio features of the audio data to be recognized. The audio features may include speaking rate features, semantic features, etc. The speaking rate features may be used to represent the rate of change of phonemes in the audio data to be recognized. For example, the speaking rate features may be represented as the number of phonemes output per second, i.e., the number of phonemes in the audio data to be recognized and the duration of the audio data to be recognized may be determined. The number of phonemes is determined by the duration of the audio data. In this embodiment, the speech rate feature may be determined as an average speech rate feature of the audio data to be recognized, or speech rate features corresponding to different phonemes of the audio data to be recognized may be determined.
[0028] The semantic features can be used to represent the meaning expressed by the phonemes in the audio data to be recognized. The audio data to be recognized may include multiple phonemes, and semantic features for each phoneme can be determined for the audio data to be recognized. That is, the audio data to be recognized can be segmented into phonemes to obtain each phoneme in the audio data to be recognized, and semantic recognition can be performed on the phonemes to determine the semantic features. For example, semantic recognition can be performed using a preset semantic recognition model, which can be a neural network model. Associations between phonemes and meanings can be preset, and the semantic features of each phoneme in the audio data to be recognized can be searched for as semantic features of the audio data to be recognized based on the preset associations.
[0029] S103: Processing is performed on a preset face image based on the speech rate feature and the semantic feature to generate a face image having a mouth shape.
[0030] For example, after obtaining the speech rate feature and the semantic feature, a predetermined facial image is processed based on the speech rate feature and the semantic feature, and a change in the mouth shape in the predetermined facial image is controlled to obtain a facial image with a mouth shape. For example, if the sound emitted by the audio data to be recognized is "a," the mouth shape in the facial image will be that of "a." In this embodiment, the mouth shape in the facial image is determined based on the semantic feature and the speech rate feature, and multiple facial images corresponding to the audio data to be recognized can be obtained. A facial video of the audio data to be recognized can also be determined based on the multiple facial images.
[0031] The relationship between mouth shapes and speech rate features and the relationship between mouth shapes and semantic features may be set in advance, or the relationship between mouth shapes, speech rate features, and semantic features may be set in advance. Based on the previously set relationship, a mouth shape corresponding to the speech rate features and semantic features is determined, and a facial image having the mouth shape is generated. A neural network model for determining the mouth shape may be trained in advance, and the speech rate features and semantic features may be input into this neural network model as input data, and a facial image having the mouth shape may be output.
[0032] In this embodiment, the method further includes, when it is determined that the numerical value represented by the speech rate feature of the recognition target audio data is smaller than the predetermined speech rate threshold, processing the predetermined facial image based on the semantic feature to generate a facial image having a mouth shape.
[0033] Specifically, when speaking slowly, the mouth shapes of each character can be perfectly aligned with the pronunciation, but when speaking quickly, it may be necessary to pronounce the next character before completing one mouth shape, resulting in many missing, merging, or simplification of mouth shapes.
[0034] The speech rate threshold is preset, and after the speech rate feature is obtained, the value represented by the speech rate feature can be compared with the preset speech rate threshold. If it is determined that the numerical value represented by the speech rate feature of the recognition target audio data is equal to or greater than the preset speech rate threshold, it indicates that the speech rate is fast, and a predetermined facial image can be processed based on the speech rate feature and semantic feature to generate a facial image with a mouth shape.
[0035] If it is determined that the numerical value represented by the speech rate feature of the recognition target audio data is smaller than a predetermined speech rate threshold, it is determined that the speech rate of the recognition target audio data is slow, and a predetermined facial image can be processed using only the semantic features to generate a facial image with a mouth shape. For example, by using only the semantic features as input data for a predetermined neural network model and performing processing such as convolution of the semantic features, the amount of calculation required to process the facial image can be reduced.
[0036] The beneficial effect of such a setting is that when the speech rate of the audio data to be recognized is slow, accurate mouth shapes can be obtained based on semantic features alone, reducing the amount of calculation and improving the efficiency of generating facial images.
[0037] In an embodiment of the present disclosure, audio data to be recognized is acquired, and speech rate features and semantic features are determined from the audio data to be recognized. The speech rate features and semantic features are combined to perform processing on a preset facial image. Here, the preset facial image is an initial image that serves as a basis when the mouth shape changes and can represent the facial appearance. Based on the speech rate features and semantic features, facial images with different mouth shapes are generated, and the mouth shapes of the facial images are matched with the audio data to be recognized. The problem of missing characters and continuous reading in the mouth shapes of the facial images when the speaking rate is fast is solved. Accurate driving of the mouth shapes in the facial images is realized, improving the accuracy of determining the facial image.
[0038] FIG. 2 is a flowchart of a method for generating a face image based on mouth shape provided by an embodiment of the present disclosure, which is an alternative embodiment based on the above embodiment.
[0039] In this embodiment, determining audio features of the audio data to be recognized can be subdivided as follows: determining speaking rate features of the audio data to be recognized based on a predetermined first feature extraction model, which is used to extract the speaking rate features from the audio data to be recognized; and determining semantic features of the audio data to be recognized based on a predetermined second feature extraction model, which is used to extract the semantic features from the audio data to be recognized.
[0040] As shown in FIG. 2, the method includes the following steps:
[0041] S201: Acquire audio data to be recognized and a preset face image.
[0042] Illustratively, this step can refer to the above step S101, and will not be further described.
[0043] S202: determining a speaking rate feature of the recognition target audio data based on a preset first feature extraction model, where the first feature extraction model is used to extract the speaking rate feature from the recognition target audio data;
[0044] For example, a first feature extraction model, which may be a predetermined neural network model for extracting speaking rate features from audio data to be recognized, is preset. The audio data to be recognized is input to the first feature extraction model and processed to obtain speaking rate features of the audio data to be recognized. For example, the first feature extraction model may include network layers such as a convolutional layer and a pooling layer, and convolution processing and feature extraction are performed on the audio data to be recognized to obtain speaking rate features of the audio data to be recognized. In this embodiment, the network structure of the first feature extraction model is not particularly limited.
[0045] In this embodiment, determining a speaking rate feature of the recognition-target audio data based on a predetermined first feature extraction model includes: inputting the recognition-target audio data into the predetermined first feature extraction model to extract features and obtain speech posterior probability features of the recognition-target audio data, where the speech posterior probability features represent information of phoneme categories of the recognition-target audio data; and determining a speaking rate feature of the recognition-target audio data based on the speech posterior probability features of the recognition-target audio data.
[0046] Specifically, the first feature extraction model may be an ASR (Automatic Speech Recognition) model, which may include multiple network layers, such as a convolutional layer, a pooling layer, and a fully connected layer. The target audio data is input to a predetermined ASR model to perform feature extraction, for example, using a convolutional layer, to obtain PPG (Phonetic Posterioram) features of the target audio data. The PPG features are a matrix of time versus category that can represent the posterior probability of each phonetic category for each specific time frame of an utterance. The PPG features can be represented using a two-dimensional coordinate axis image, representing phoneme category information of the target audio data, with the horizontal axis representing time and the vertical axis representing phoneme category.
[0047] After obtaining the PPG features, calculations can be performed on the PPG features based on a preset speaking rate determination algorithm, and the PPG features can be converted into speaking rate features of the audio data to be recognized. The phoneme change rate can be calculated and used as the speaking rate magnitude to achieve explicit modeling of the speaking rate features. In this embodiment, the preset speaking rate determination algorithm is not particularly limited.
[0048] The beneficial effect of this configuration is that the target audio data is input to the automatic speech recognition model for processing, PPG features of the target audio data are obtained, and further calculations are performed on the PPG features to obtain speaking rate features, which realizes explicit modeling of speaking rate, thereby introducing speaking rate features and significantly improving the accuracy and authenticity of audio-driven mouth movements when speaking rate changes.
[0049] In this embodiment, determining a speaking rate feature of the recognition-target audio data based on the speech posterior probability feature of the recognition-target audio data includes: performing a fast Fourier transform on the speech posterior probability feature to obtain a frequency domain signal feature, where the frequency domain signal feature represents information of the phoneme category of the recognition-target audio data; dividing the frequency domain signal feature into frequency domain signal features of at least two frequency bands based on a predetermined frequency band size; and performing an integration process on the frequency domain signal features of the at least two frequency bands to obtain a speaking rate feature of the recognition-target audio data.
[0050] Specifically, PPG features are time-domain signals. After obtaining the PPG features of the target audio data, the PPG features can be subjected to a fast Fourier transform (FFT). That is, the PPG features are transformed into the frequency domain using FFT (Fast Fourier Transform), and frequency-domain signal features corresponding to the PPG features are obtained. These frequency-domain signal features may be represented as phoneme category information of the target audio data.
[0051] The frequency domain signal features are integrated for each frequency band to calculate the desired frequency as the speaking rate size, i.e., to obtain the speaking rate feature of the audio data to be recognized. When calculating the speaking rate feature, the frequency band size can be set in advance, and the frequency domain signal features are divided based on the set frequency band size to obtain frequency domain signal features for multiple frequency band sizes. An integration process is performed on the frequency domain signal features for each frequency band size one by one, and the integration result can be used as an embodiment of the phoneme change rate in the audio data to be recognized, i.e., the speaking rate feature.
[0052] The beneficial effect of such a setting is that by performing FFT processing and integral calculation, the PPG features can be converted into a specific speaking rate magnitude, realizing the determination of speaking rate features, thereby improving the accuracy of generating face images.
[0053] S203: determining semantic features of the recognition target audio data based on a preset second feature extraction model, where the second feature extraction model is used to extract the semantic features from the recognition target audio data;
[0054] For example, the second feature extraction model may be a pre-trained neural network model, such as a pre-set semantic recognition model, which includes a feature extraction network and extracts semantic features from the recognition target audio data based on the pre-set second feature extraction model, thereby obtaining the semantic features of the recognition target audio data.
[0055] Through the first feature extraction model and the second feature extraction model, the speech rate feature and the semantic feature can be quickly obtained, and the speech rate feature and the semantic feature can be extracted separately, which improves the efficiency of feature extraction and further improves the efficiency of facial image generation.
[0056] In this embodiment, determining semantic features of the recognition-target audio data based on the preset second feature extraction model includes inputting the recognition-target audio data into the preset second feature extraction model to extract features, and outputting the semantic features of the recognition-target audio data.
[0057] Specifically, the second feature extraction model may be a semantic recognition model, which may include a network layer such as a multi-layer convolutional layer to form a feature extraction network. The audio data to be recognized is input to a pre-defined semantic recognition model for processing, and feature extraction is performed, for example, using a convolutional layer, to obtain semantic features of the audio data to be recognized. The audio data to be recognized may be streaming data, and the extracted semantic features may be streaming features. In this embodiment, the model structure of the semantic recognition model is not particularly limited.
[0058] The beneficial effect of such a setting is to automatically extract semantic features from the input audio stream data, improve the efficiency and accuracy of determining semantic features, and further improve the efficiency and accuracy of generating facial images.
[0059] S204: Processing the preset face image based on the speech rate feature and semantic feature to generate a face image with a mouth shape.
[0060] Illustratively, this step can refer to step S103 above, and will not be further described.
[0061] In an embodiment of the present disclosure, audio data to be recognized is acquired, and speech rate features and semantic features are determined from the audio data to be recognized. The speech rate features and semantic features are combined and processed on a preset facial image. Here, the preset facial image is an initial image that serves as a basis when the mouth shape changes and can represent the facial appearance. Based on the speech rate features and semantic features, facial images with different mouth shapes are generated, and the mouth shapes of the facial image are matched with the audio data to be recognized. The problem of missing characters and continuous reading in the mouth shape of the facial image when the speaking rate is fast is solved. Accurate driving of the mouth shape in the facial image is realized, improving the accuracy of determining the facial image.
[0062] FIG. 3 is a flowchart of a method for generating a face image based on mouth shape provided by an embodiment of the present disclosure, which is an alternative embodiment based on the above embodiment.
[0063] In this embodiment, processing a predetermined facial image based on speech rate features and semantic features to generate a facial image with a mouth shape can be subdivided as follows: the speech rate features and semantic features are input into a predetermined face and mouth shape determination model for processing, and a facial image with a mouth shape is generated based on the processed result and the predetermined facial image.
[0064] As shown in FIG. 3, the method includes the following steps:
[0065] S301: Acquire audio data to be recognized and a preset face image.
[0066] Illustratively, this step can refer to the above step S101, and will not be further described.
[0067] S302: determining audio features of the recognition target audio data, where the audio features include speaking rate features and semantic features;
[0068] Illustratively, this step can refer to step S102 above, and will not be further described.
[0069] S303: The speech rate feature and the semantic feature are inputted into a preset face and mouth shape determination model for processing, and a face image having a mouth shape is generated based on the processed result and the preset face image.
[0070] For example, a face and mouth pattern determination model, which is a neural network model that can be used to output a facial image with a mouth shape, is constructed and trained in advance. Speech rate features and semantic features are input as input data into a predetermined face and mouth pattern determination model for processing. After processing, the face and mouth pattern determination model changes the mouth shape of the predetermined facial image based on the processing result, thereby obtaining a facial image with a mouth shape. For example, the processing result determined by the face and mouth pattern determination model based on the speech rate features and semantic features may be mouth shape size and shape information. The predetermined facial image can be rendered based on the determined mouth shape size and shape information to generate a facial image including the mouth shape. Using the face and mouth pattern determination model allows for quick generation of a facial image, and by combining speech rate features and semantic features, the problem of reduced effectiveness of audio-driven face and mouth patterns due to changes in speech rate can be avoided, improving the efficiency and accuracy of facial image generation.
[0071] In this embodiment, inputting the speech rate features and semantic features into a predetermined face and mouth type determination model for processing, and generating a facial image with a mouth shape based on the processed result and the predetermined facial image includes: performing a combination process on the speech rate features and semantic features based on the predetermined face and mouth type determination model to obtain a combination feature of the recognition target audio data, where the combination feature represents the speech rate feature and the semantic feature; performing feature extraction on the combination feature based on a convolution layer in the predetermined face and mouth type determination model to obtain facial driving parameters, where the facial driving parameters are used to represent parameters required to drive changes in the mouth shape in the facial image; and performing image rendering on the predetermined facial image based on the facial driving parameters to generate a facial image with a mouth shape.
[0072] Specifically, the speech rate feature and the semantic feature are input into a preset face and mouth type determination model. Based on the face and mouth type determination model, a combining process can be performed on the speech rate feature and the semantic feature, for example, a matrix represented by the speech rate feature and a matrix represented by the semantic feature can be combined. The combined data is determined as a combined feature of the recognition target audio data. That is, the combined feature can represent the speech rate feature and the semantic feature.
[0073] The face and mouth type determination model is configured with a network layer such as a convolutional layer. When the combined features pass through the convolutional layer of the face and mouth type determination model, feature extraction is performed on the combined features based on the convolutional layer, and facial drive parameters are calculated. The facial drive parameters are parameters required to drive changes in the mouth shape in the facial image. For example, the facial drive parameters may be position information and size information of a target frame including the mouth shape in the facial image. After obtaining the facial drive parameters, image rendering is performed on the preset facial image, and the mouth shape in the preset facial image is changed from its original closed shape to a shape corresponding to the facial drive parameters, thereby obtaining a facial image with a mouth shape. For some of the audio data to be recognized, multiple facial images with different mouth shapes can be generated.
[0074] The beneficial effect of such a setting is that the speaking rate features and semantic features are combined, and the parameters required to drive the face and mouth shapes are obtained through the driving network of the face and mouth shape determination model, so that the mouth shapes in the generated face image are adapted to the audio data to be recognized, thereby reducing the influence of the speaking rate on the mouth shapes in the face image, and improving the efficiency and accuracy of generating face images.
[0075] In this embodiment, the face driving parameters are weight parameters of a mixed deformation, and performing image rendering on a predetermined face image based on the face driving parameters to generate a face image having a mouth shape includes determining face 3D mesh data corresponding to the predetermined face image based on the weight parameters of the mixed deformation, where the face 3D mesh data is data representing a 3D mesh model of the face surface in the face image, and performing image rendering on the predetermined face image based on the face 3D mesh data to generate a face image having a mouth shape.
[0076] Specifically, the facial driving parameters may be blend shape weights, and the blend shape weights are obtained by a driving network in a face-mouth shape determination model. Based on the blend shape weight parameters, a preset rendering engine can generate a facial image with a mouth shape based on a preset facial image. For example, the preset rendering engine may be an Unreal rendering engine.
[0077] When performing image rendering, 3D facial mesh data can be first determined based on the blend shape weights. The 3D facial mesh data can be data for representing a 3D mesh model of the facial surface in the facial image. The 3D facial mesh can be determined based on the blend shape weights and the blend shape base. Here, the blend shape base is related to the binding of the human image and is a fixed, invariant, predetermined parameter. After obtaining the 3D facial mesh data, image rendering is performed on the facial image to obtain a facial image with a mouth shape.
[0078] The beneficial effect of this setup is that we first obtain a 3D facial mesh based on the blend shape weights, and then obtain a facial image based on the 3D facial mesh, which realizes accurate generation of facial images and is convenient for users to experience the digital human.
[0079] In an embodiment of the present disclosure, audio data to be recognized is acquired, and speech rate features and semantic features are determined from the audio data to be recognized. The speech rate features and semantic features are combined and processed on a preset facial image. Here, the preset facial image is an initial image that serves as a basis when the mouth shape changes and can represent the facial appearance. Based on the speech rate features and semantic features, facial images with different mouth shapes are generated, and the mouth shapes of the facial image are matched with the audio data to be recognized. The problem of missing characters and continuous reading in the mouth shape of the facial image when the speaking rate is fast is solved. Accurate driving of the mouth shape in the facial image is realized, improving the accuracy of determining the facial image.
[0080] FIG. 4 is a flowchart of a method for training a face and mouth typing model provided by an embodiment of the present disclosure, which includes: Model As shown in Figure 4, this method includes the following steps:
[0081] S401: obtaining training target image data and a preset face image, the training target image data including training target audio data and a training target face image, and the training target face image having a mouth shape corresponding to the training target audio data;
[0082] For example, a deep learning-based face and mouth type determination model can be used to determine a face image with a mouth shape. The face and mouth type determination model can implement the method for generating a face image described in any of the above embodiments, and the face and mouth type determination model needs to be pre-trained before use. Pre-collected training target image data and pre-set face images are obtained. The training target image data may include training target audio data and training target face images, where the training target audio data is an audio stream for training the model, and the training target face images have mouth shapes that match the training target audio data.
[0083] The preset facial image is a facial image of a digital human with a pre-designed mouth, and the preset facial image may also include facial features such as eyes and nose. The digital human's facial shape, eyes, nose, mouth, etc. can be designed to generate a preset facial image. The digital human can change its mouth shape based on the preset facial image. For example, in the preset facial image, the digital human's mouth is closed, and the digital human's mouth shape can change as audio data is transmitted. The difference between the training target facial image and the preset facial image is that the mouth shape has changed.
[0084] In this embodiment, obtaining training target image data includes obtaining training target audio data, performing a 3D reconstruction process of a face image based on the training target audio data to obtain 3D face mesh data corresponding to the training target audio data, and obtaining a training target face image based on the 3D face mesh data corresponding to the training target audio data.
[0085] Specifically, a pre-collected training set is obtained, which may be training target audio data. Based on the training target audio data, a training target face image is generated. The training target face image has a mouth shape, and the mouth shape in the training target face image matches the training target audio data.
[0086] A 3D reconstruction process for a facial image can be performed based on the training target audio data. For example, a 3D reconstruction process for each frame of a facial image is performed based on each phoneme of the training target audio data. In this embodiment, the procedure for the 3D reconstruction process is not particularly limited. By determining 3D facial mesh data for each frame, a 3D facial mesh for multiple frames corresponding to the training target audio data can be obtained. A training target facial image for multiple frames is obtained based on the 3D facial mesh corresponding to the training target audio data.
[0087] The beneficial effect of such a setting is that predetermining a facial image corresponding to the training audio data makes it easier to train the face-mouth type determination model, thereby improving the training efficiency and accuracy of the face-mouth type determination model.
[0088] S402, determining audio features of the training audio data, where the audio features include speaking rate features and semantic features.
[0089] For example, after obtaining training audio data, feature extraction is performed on the training audio data to obtain audio features of the training audio data. The audio features may include speaking rate features and semantic features. The speaking rate features can be used to represent the rate of change of phonemes in the training audio data. For example, the speaking rate features can be expressed as the number of phonemes output per second, i.e., the number of phonemes in the training audio data and the duration of the training audio data can be determined. The number of phonemes is calculated based on the duration of the training audio data. In this embodiment, the speaking rate feature of the training target audio data may be determined as an average speaking rate feature of the training target audio data, or speaking rate features corresponding to different phonemes of the training target audio data may be determined.
[0090] The semantic features can be used to represent the meaning expressed by the training audio data. The training audio data may include multiple phonemes, and semantic features for each phoneme can be determined for the training audio data. That is, the training audio data can be segmented into phonemes to obtain each phoneme in the training audio data, and semantic recognition can be performed on the phonemes to determine the semantic features. For example, semantic recognition can be performed using a preset semantic recognition model, which can be a neural network model. Alternatively, association relationships between phonemes and meanings can be preset, and the semantic features of all phonemes in the training audio data can be retrieved based on the preset association relationships as the semantic features of the training audio data.
[0091] S403: Based on the speech rate feature, semantic feature and the preset face image, an initial face and mouth shape determination model is trained to obtain a face image with a mouth shape.
[0092] For example, the speech rate and semantic features of the training audio data are input to the training face and mouth shape determination model for repeated training, and a face image with a mouth shape is generated based on the processed result and a preset face image at each iteration.
[0093] A training target face and mouth type determination model is constructed in advance, and speaking rate features and semantic features are input as input data into the training target face and mouth type determination model for processing. After processing, the face and mouth type determination model changes the mouth shape of a predetermined facial image based on the processing result, thereby obtaining a facial image with a different mouth shape. For example, the processing result determined by the face and mouth type determination model based on the speaking rate features and semantic features is mouth size and shape information, and the predetermined facial image can be rendered based on the determined mouth size and shape information to generate a facial image including that mouth shape. The training target audio data includes multiple phonemes, and a facial image with a mouth shape corresponding to each phoneme can be generated.
[0094] S404: If the face image having the mouth shape matches the training target face image, it is determined that the training is completed and a face and mouth shape determination model is obtained.
[0095] For example, after obtaining a facial image with a mouth shape output from the model, the facial image with a mouth shape corresponding to the phoneme is compared with the training target facial image corresponding to the phoneme. If they match, it is determined that the training of the face and mouth type determination model is complete. If they do not match, it is determined that the face and mouth type determination model needs to be trained. Then, the semantic features and speaking rate features of the training target audio data are input into the face and mouth type determination model, and training is performed based on a preset back propagation algorithm until the output facial image with a mouth shape matches the corresponding training target facial image.
[0096] In addition, a similarity threshold may be preset, and the similarity threshold may be used to determine whether the training of the face and mouth type determination model is complete. After obtaining a face image with a mouth shape, the similarity between the face image with the mouth shape and the corresponding training target face image is determined. If the determined similarity is equal to or greater than the preset similarity threshold, it is determined that the training of the face and mouth type determination model is complete. If the similarity is less than the preset similarity threshold, it is determined that the training of the face and mouth type determination model is not complete.
[0097] In an embodiment of the present disclosure, training target audio data and training target facial images are obtained, and speaking rate features and semantic features are determined from the training target audio data. The speaking rate features and semantic features are combined to train a training target face and mouth pattern determination model. Based on the speaking rate features and semantic features, facial images with different mouth shapes are generated, and the mouth shapes in the output facial images are matched with the training target audio data by training. The model learns the effects of different speaking rates on mouth shapes, which greatly improves the accuracy and authenticity of audio-driven mouth shapes when speaking rate changes, and is convenient for improving the accuracy of facial image determination when the face and mouth pattern determination model is subsequently used.
[0098] FIG. 5 is a flowchart of a method for training a face and mouth type determination model provided by an embodiment of the present disclosure, which is an alternative embodiment based on the above embodiment.
[0099] In this embodiment, determining audio features of the training-target audio data can be subdivided as follows: determining speaking rate features of the training-target audio data based on a predetermined first feature extraction model, which is used to extract the speaking rate features from the training-target audio data; determining semantic features of the training-target audio data based on a predetermined second feature extraction model, which is used to extract the semantic features from the training-target audio data.
[0100] As shown in FIG. 5, the method includes the following steps:
[0101] S501: obtaining training target image data and a preset face image, the training target image data including training target audio data and a training target face image, and the training target face image having a mouth shape corresponding to the training target audio data;
[0102] Illustratively, this step can refer to step S401 above, and will not be further described.
[0103] S502, determining speaking rate features of the training target audio data based on a preset first feature extraction model, where the first feature extraction model is used to extract the speaking rate features from the training target audio data.
[0104] For example, a first feature extraction model is preset, which may be a neural network model predetermined to extract speaking rate features from training target audio data. The training target audio data is input to the first feature extraction model and processed to obtain speaking rate features of the training target audio data. For example, the first feature extraction model may include network layers such as a convolutional layer and a pooling layer, and convolution processing and feature extraction are performed on the training target audio data to obtain speaking rate features of the training target audio data. In this embodiment, the network structure of the first feature extraction model is not particularly limited.
[0105] In this embodiment, determining a speaking rate feature of the training-target audio data based on a predetermined first feature extraction model includes: inputting the training-target audio data into the predetermined first feature extraction model to perform feature extraction, and obtaining a speech posterior probability feature of the training-target audio data, where the speech posterior probability feature represents information of the phoneme category of the training-target audio data; and determining a speaking rate feature of the training-target audio data based on the speech posterior probability feature of the training-target audio data.
[0106] Specifically, the first feature extraction model may be an ASR model, which may include multiple network layers, such as a convolutional layer, a pooling layer, and a fully connected layer. The training audio data is input to a predetermined ASR model to perform feature extraction, for example, by a convolutional layer, to obtain PPG features of the training audio data. The PPG features are a matrix of time versus category that can represent the posterior probability of each phonetic category for each specific time frame of an utterance. The PPG features can be represented using a two-dimensional coordinate axis image, where the horizontal axis represents time and the vertical axis represents phoneme category information of the training audio data.
[0107] After obtaining the PPG features, calculations can be performed on the PPG features based on a preset speaking rate determination algorithm, and the PPG features can be converted into speaking rate features of the training audio data. The phoneme change rate can be calculated and used as the speaking rate magnitude to achieve explicit modeling of the speaking rate features. In this embodiment, the preset speaking rate determination algorithm is not particularly limited.
[0108] The beneficial effect of this configuration is that the training audio data is input to the automatic speech recognition model and processed to obtain the PPG features of the training audio data, and then further calculations are performed on the PPG features to obtain the speaking rate features, which realizes explicit modeling of speaking rate, thereby introducing the speaking rate features and significantly improving the accuracy and authenticity of the audio-driven mouth movements when the speaking rate changes.
[0109] In this embodiment, determining the speaking rate feature of the training target audio data based on the speech posterior probability feature of the training target audio data includes: performing a fast Fourier transform on the speech posterior probability feature to obtain a frequency domain signal feature, where the frequency domain signal feature represents information of the phoneme category of the training target audio data; dividing the frequency domain signal feature into frequency domain signal features of at least two frequency bands based on a predetermined frequency band size; and performing an integration process on the frequency domain signal features of the at least two frequency bands to obtain the speaking rate feature of the training target audio data.
[0110] Specifically, PPG features are time-domain signals. After obtaining the PPG features of the training audio data, the PPG features can be subjected to a fast Fourier transform (FFT) process. The PPG features are then transformed into the frequency domain using FFT to obtain frequency-domain signal features corresponding to the PPG features. These frequency-domain signal features can then be used to represent phoneme category information for the training audio data.
[0111] The frequency domain signal features are integrated for each frequency band to calculate the desired frequency as the speaking rate size, i.e., obtain the speaking rate feature of the training audio data. When calculating the speaking rate feature, the frequency band size can be set in advance, and the frequency domain signal features are divided based on the set frequency band size to obtain frequency domain signal features of multiple frequency band sizes. An integration process is performed on the frequency domain signal features of each frequency band size one by one, and the integration result can be used as an embodiment of the phoneme change rate in the training audio data, i.e., the speaking rate feature.
[0112] The beneficial effect of this setting is that by performing FFT processing and integration calculation, the PPG features can be converted into specific speaking rate magnitudes, realizing the determination of speaking rate features, thereby improving the training accuracy of the face and mouth type determination model.
[0113] S503, determining semantic features of the training target audio data based on a preset second feature extraction model, where the second feature extraction model is used to extract the semantic features from the training target audio data.
[0114] For example, the second feature extraction model may be a pre-trained neural network model, for example, a pre-set semantic recognition model, and includes a feature extraction network, which extracts semantic features from the training audio data based on the feature extraction network in the second feature extraction model, thereby obtaining the semantic features of the training audio data.
[0115] Through the first feature extraction model and the second feature extraction model, the speech rate feature and the semantic feature can be quickly obtained, the speech rate feature and the semantic feature can be extracted separately, the efficiency of the feature extraction can be improved, and the training efficiency of the face and mouth type determination model can be further improved.
[0116] In this embodiment, determining semantic features of the training-target audio data based on the preset second feature extraction model includes: inputting the training-target audio data into the preset second feature extraction model to extract features, and outputting the semantic features of the training-target audio data.
[0117] Specifically, the second feature extraction model may be a semantic recognition model, which may include a network layer such as a multi-layer convolutional layer to form a feature extraction network. The training target audio data may be input to a pre-configured semantic recognition model for processing. For example, feature extraction may be performed using a convolutional layer to obtain semantic features of the training target audio data. The training target audio data may be streaming data, and the extracted semantic features may be streaming features. In this embodiment, the model structure of the semantic recognition model is not particularly limited.
[0118] The beneficial effect of such a setting is to automatically extract semantic features from the input audio stream data, improve the efficiency and accuracy of semantic feature determination, and further improve the efficiency and accuracy of training the face-mouth pattern determination model.
[0119] S504: Based on the speech rate feature, the semantic feature and the preset face image, an initial face and mouth shape determination model is trained to obtain a face image with a mouth shape.
[0120] For example, the speech rate feature and the semantic feature are input to a training target face and mouth type determination model for training, which processes the semantic feature and the speech rate feature and generates a face image with a mouth shape based on the processing result and a preset face image.
[0121] In this embodiment, training an initial face and mouth type determination model based on a speech rate feature, a semantic feature, and a predetermined facial image to obtain a facial image with a mouth shape includes: performing a combination process on the speech rate feature and the semantic feature based on the initial face and mouth type determination model to obtain a combination feature of the training target audio data, where the combination feature represents the speech rate feature and the semantic feature; performing feature extraction on the combination feature based on a convolutional layer in the initial face and mouth type determination model to obtain facial driving parameters, where the facial driving parameters are used to represent parameters required to drive changes in the mouth shape in the facial image; and performing image rendering on the predetermined facial image based on the facial driving parameters to obtain a facial image with a mouth shape.
[0122] Specifically, the speech rate feature and the semantic feature are input into a face / mouth type determination model to be trained. Based on the face / mouth type determination model, a combining process can be performed on the speech rate feature and the semantic feature, for example, a matrix represented by the speech rate feature and a matrix represented by the semantic feature can be combined. The combined data is determined as a combined feature of the training audio data. That is, the combined feature can represent the speech rate feature and the semantic feature.
[0123] The face and mouth type determination model is configured with a network layer such as a convolutional layer. When the combined features pass through the convolutional layer of the face and mouth type determination model, feature extraction is performed on the combined features based on the convolutional layer, and facial driving parameters can be calculated and obtained. The facial driving parameters are parameters required to drive changes in the mouth shape in the facial image. For example, the facial driving parameters may be position information and size information of a target frame including the mouth shape in the facial image. After obtaining the facial driving parameters, image rendering is performed on the preset facial image, and the mouth shape in the preset facial image is changed from its original closed shape to a shape corresponding to the facial driving parameters, thereby obtaining a facial image with a mouth shape.
[0124] The beneficial effect of such a setting is that by combining the speaking rate features and semantic features and obtaining and training the parameters required to drive the face and mouth shapes through the driving network of the face and mouth shape determination model, the mouth shapes in the generated face image can be adapted to the training target audio data, the influence of the speaking rate on the mouth shapes in the face image can be reduced, and the training accuracy of the face and mouth shape determination model can be improved.
[0125] In this embodiment, the face driving parameters are weight parameters of a mixed deformation, and performing image rendering on a predetermined face image based on the face driving parameters to obtain a face image having a mouth shape includes determining face 3D mesh data corresponding to the predetermined face image based on the weight parameters of the mixed deformation, where the face 3D mesh data is data representing a 3D mesh model of the face surface in the face image, and performing image rendering on the predetermined face image based on the face 3D mesh data to generate a face image having a mouth shape.
[0126] Specifically, the facial driving parameters may be blend shape weights, and the blend shape weights are obtained by a driving network in a face and mouth shape determination model. Based on the blend shape weight parameters, a preset rendering engine can generate a facial image with a mouth shape based on a preset facial image. For example, the preset rendering engine may be an Unreal rendering engine.
[0127] When performing image rendering, 3D facial mesh data can be first determined based on the blend shape weights. The 3D facial mesh data can be data for representing a 3D mesh model of the facial surface in the facial image. The 3D facial mesh can be determined based on the blend shape weights and the blend shape base. Here, the blend shape base is related to the binding of the human image and is a fixed, invariant, predetermined parameter. After obtaining the 3D facial mesh data, image rendering is performed on the facial image to obtain a facial image with a mouth shape.
[0128] The beneficial effect of such a setting is to first obtain a 3D facial mesh based on the blend shape weights, and then obtain a facial image based on the 3D facial mesh, thereby realizing accurate generation of facial images and improving the training accuracy of the face and mouth type determination model.
[0129] S505: if the face image having the mouth shape matches the training target face image, it is determined that the training is completed and a face and mouth shape determination model is obtained.
[0130] Illustratively, this step can refer to step S404 above, and will not be further described.
[0131] In an embodiment of the present disclosure, training target audio data and training target facial images are obtained, and speaking rate features and semantic features are determined from the training target audio data. The speaking rate features and semantic features are combined to train a training target face and mouth pattern determination model. Facial images with different mouth shapes are generated based on the speaking rate features and semantic features, and the mouth shapes in the output facial images are matched with the training target audio data through training. The model learns the effects of different speaking rates on mouth shapes, which greatly improves the accuracy and authenticity of audio-driven mouth shapes when speaking rate changes, and is convenient for improving the accuracy of facial image determination when the face and mouth pattern determination model is subsequently used.
[0132] 6 is a block diagram of an apparatus for generating a facial image based on a mouth shape provided by an embodiment of the present disclosure. For ease of explanation, only parts related to the embodiment of the present disclosure are shown. Referring to FIG. 6, the apparatus 600 for generating a facial image based on a mouth shape includes a data acquisition unit 601, a feature determination unit 602, and an image generation unit 603.
[0133] The data acquisition unit 601 is used to acquire the target audio data and the preset face image; The feature determining unit 602 is used to determine audio features of the recognition target audio data, the audio features including speaking rate features and semantic features; The image generating unit 603 processes the predetermined face image according to the speech rate feature and the semantic feature, and generates a face image with a mouth shape.
[0134] FIG. 7 is a structural diagram of an apparatus for generating a facial image based on mouth shape provided by an embodiment of the present disclosure. As shown in FIG. 7, the apparatus 700 for generating a facial image based on mouth shape includes a data acquisition unit 701, a feature determination unit 702, and an image generation unit 703, where the feature determination unit 702 includes a first determination module 7021 and a second determination module 7022.
[0135] The first determination module 7021 is used to determine the speech rate feature of the recognition target audio data based on a preset first feature extraction model. , th The feature extraction model 1 is used to extract speech rate features from the target audio data, The second determination module 7022 is used to determine semantic features of the recognition target audio data based on a preset second feature extraction model, and the second feature extraction model is used to extract semantic features from the recognition target audio data.
[0136] In one example, the first determination module 7021: a feature extraction sub-module for inputting the recognition-target audio data into a predetermined first feature extraction model to perform feature extraction and obtain phonetic posterior probability features of the recognition-target audio data, the phonetic posterior probability features representing phoneme category information of the recognition-target audio data; a feature determining sub-module for determining a speaking rate feature of the recognition target audio data based on the speech posterior probability feature of the recognition target audio data.
[0137] In one example, the feature determination sub-module specifically: performing a fast Fourier transform process on the speech posterior probability features to obtain frequency domain signal features, the frequency domain signal features representing phoneme category information of the recognition target audio data; Dividing the frequency domain signal features into frequency domain signal features of at least two frequency bands based on a predetermined frequency band size; and performing an integration process on the frequency domain signal features of the at least two frequency bands to obtain the speech rate feature of the recognition target audio data.
[0138] In one example, the second determination module 7022 specifically: The recognition target audio data is input to a preset second feature extraction model to extract features, and semantic features of the recognition target audio data are obtained as output.
[0139] In one example, the image generation unit 703 and an image generation module used to input the speech rate features and the semantic features into a predetermined face and mouth shape determination model for processing, and generate a face image having a mouth shape based on the processed result and the predetermined face image.
[0140] In one example, the image generation module: a feature combining sub-module for performing combining processing on the speech rate feature and the semantic feature based on the predetermined face and mouth type determination model to obtain combined features of the recognition target audio data, the combined features representing the speech rate feature and the semantic feature; a parameter determination sub-module, which is used to perform feature extraction on the combined features based on a convolution layer in the preset face and mouth shape determination model, to obtain face driving parameters, the face driving parameters being used to represent parameters required to drive changes in mouth shape in a face image; and an image rendering sub-module, which is used to perform image rendering on the preset face image according to the face driving parameters to generate a face image with a mouth shape.
[0141] In one example, the face driving parameters are weight parameters of the blending deformation, and the image rendering sub-module specifically includes: determining three-dimensional facial mesh data corresponding to the predetermined facial image based on weight parameters of the mixed transformation, the three-dimensional facial mesh data being data representing a three-dimensional mesh model of a facial surface in the facial image; and performing image rendering on the preset face image based on the face 3D mesh data to generate a face image having a mouth shape.
[0142] One example further includes: and a semantic processing unit, which is used to process the predetermined facial image based on the semantic features to generate a facial image with a mouth shape when it is determined that the numerical value represented by the speech rate feature of the recognition target audio data is smaller than a predetermined speech rate threshold.
[0143] 8 is a block diagram of an apparatus for training a face and mouth type determination model provided by an embodiment of the present disclosure. For ease of explanation, only parts relevant to the embodiment of the present disclosure are shown. Referring to FIG. 8, the apparatus 800 for training a face and mouth type determination model includes an image acquisition unit 801, a feature extraction unit 802, a model training unit 803, and a model acquisition unit 804.
[0144] The image acquisition unit 801 is used to acquire training target image data and a predetermined face image, the training target image data including training target audio data and a training target face image, and the training target face image has a mouth shape corresponding to the training target audio data; The feature extraction unit 802 is used to determine audio features of the training audio data, the audio features including speaking rate features and semantic features; The model training unit 803 is used to train an initial face and mouth shape determination model according to the speech rate feature, the semantic feature and the predetermined face image, to obtain a face image with a mouth shape; The model obtaining unit 804 is used to determine that if the face image with mouth shape matches the training target face image, a trained face mouth shape determination model is obtained.
[0145] In one example, the feature extraction unit 802 a first extraction module used to determine speaking rate features of the training target audio data based on a predetermined first feature extraction model, wherein the first feature extraction model is used to extract speaking rate features from the training target audio data; a second extraction module used to determine semantic features of the training target audio data based on a predetermined second feature extraction model, where the second feature extraction model is used to extract the semantic features from the training target audio data.
[0146] In one example, the first extraction module comprises: a probability determination sub-module, which is used to input the training target audio data into a predetermined first feature extraction model to perform feature extraction and obtain phonetic posterior probability features of the training target audio data, wherein the phonetic posterior probability features represent information of phoneme categories of the training target audio data; a speaking rate determining sub-module for determining speaking rate features of the training target audio data based on the speech posterior probability features of the training target audio data.
[0147] In one example, the speaking rate determination sub-module: performing a fast Fourier transform on the speech posterior probability features to obtain frequency domain signal features, the frequency domain signal features representing phoneme category information of the training audio data; Dividing the frequency domain signal features into frequency domain signal features of at least two frequency bands based on a predetermined frequency band size; and performing an integration process on the frequency domain signal features of the at least two frequency bands to obtain the speaking rate feature of the training audio data.
[0148] In one example, the second extraction module specifically: The training target audio data is input to a preset second feature extraction model to extract features, and semantic features of the training target audio data are obtained as an output.
[0149] In one example, the model training unit 803 a feature combining module, which is used to perform combining processing on the speech rate features and the semantic features based on the initial face and mouth type determination model to obtain combined features of the training target audio data, wherein the combined features represent speech rate features and semantic features; a parameter determination module used to perform feature extraction on the combined features based on a convolutional layer in the initial face and mouth shape determination model to obtain face driving parameters, the face driving parameters being used to represent parameters required to drive changes in mouth shape in a face image; and an image rendering module for performing image rendering on the preset face image based on the face driving parameters to obtain a face image having a mouth shape.
[0150] In one example, the face driving parameters are blending deformation weight parameters, and the image rendering module: a data determination sub-module for determining 3D facial mesh data corresponding to the predetermined facial image based on the weight parameters of the mixed transformation, the 3D facial mesh data being data representing a 3D mesh model of a facial surface in the facial image; and an image rendering sub-module, which is used to perform image rendering on the preset face image based on the face 3D mesh data, to generate a face image with a mouth shape.
[0151] In one example, the image acquisition unit 801 a data acquisition module used for acquiring the training target audio data; a 3D reconstruction module used to perform 3D reconstruction processing of a face image based on the training target audio data, and obtain 3D face mesh data corresponding to the training target audio data; an image acquisition module used to obtain the training target face image based on the facial 3D mesh data corresponding to the training target audio data.
[0152] FIG. 9 is a block diagram of an electronic device provided by an embodiment of the present disclosure. As shown in FIG. 9, the electronic device 900 includes at least one processor 902 and a memory 901 communicatively connected to the at least one processor 902, wherein the memory stores instructions executable by the at least one processor 902, and the instructions, when executed by the at least one processor 902, cause the at least one processor 902 to perform the method for generating a facial image based on a mouth shape and the method for training a model of the present disclosure.
[0153] The electronic device 900 further includes a receiver 903 and a transmitter 904. The receiver 903 is used to receive commands and data transmitted from other devices, and the transmitter 904 is used to transmit commands and data to external devices.
[0154] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program. M provide.
[0155] According to an embodiment of the present disclosure, the present disclosure further provides a computer program M Provides computer programs M is , stored in a readable storage medium R, At least one processor of the electronic device can read a computer program from a readable storage medium, and the computer program is executed by the at least one processor to cause the electronic device to implement the technical solutions provided by any of the above embodiments.
[0156] 10 is a schematic block diagram of an exemplary electronic device 1000 that can be used to implement embodiments of the present disclosure. The electronic device may be a laptop computer, a desktop computer, a work desk, a personal digital assistant, or any other device. assistant The term "electronic device" is intended to represent various forms of digital computers, such as computers, servers, blade servers, mainframes, and other suitable computers. Electronic devices may also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown herein are merely exemplary and are not intended to limit the practice of the present disclosure as described and / or claimed herein.
[0157] As shown in FIG. 10, the device 1000 includes a computing unit 1001, which has a read-only memory ( Read Only Memory, A computer program stored in the random access memory (ROM) 1002 or the storage unit 1008 Random Access Memory, The computing unit 1001, the ROM 1002, and the RAM 1003 can execute various appropriate operations and processes based on a computer program loaded into the RAM 1003. The RAM 1003 may also store various programs and data necessary for the operation of the device 1000. The computing unit 1001, the ROM 1002, and the RAM 1003 are interconnected by a bus 1004. Input / Output, An I / O interface 1005 is also connected to the bus 1004 .
[0158] A number of components in the device 1000 are connected to an I / O interface 1005, including an input unit 1006 such as a keyboard, a mouse, etc., an output unit 1007 such as various types of displays, speakers, etc., a storage unit 1008 such as a magnetic disk, an optical disk, etc., and a communication unit 1009 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 enables the device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0159] The computing unit 1001 may be a variety of general-purpose and / or special-purpose processing components having processing and computing capabilities. Some examples of the computing unit 1001 include a central processing unit ( Central Processing Unit, CPU), graphics processing unit ( Graphics Processing Unit, GPU, various dedicated artificial intelligence ( Artificial Intelligence, AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors ( Digital Signal Processor,The computing unit 1001 may include, but is not limited to, a processor, a DSP, and any suitable processor, controller, microcontroller, etc. The computing unit 1001 performs the methods and processes described above, such as the method for generating a facial image based on mouth shape and the method for training a model. For example, in some embodiments, the method for generating a facial image based on mouth shape and the method for training a model are embodied as computer software programs and tangibly stored in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer programs may be loaded and / or installed into the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer programs are loaded into the RAM 1003 and executed by the computing unit 1001, they may cause one or more steps of the method for generating a facial image based on mouth shape and the method for training a model described above to be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to perform the method for generating a facial image based on mouth shape and the method for training a model in any other suitable manner (e.g., by firmware).
[0160] Various embodiments of the systems and techniques described herein may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (e.g., Field-Programmable Gate Array, FPGA), dedicated integrated circuits ( Application Specific Integrated Circuit, ASIC), dedicated standard products ( Application Specific Standard Product, ASSP), System-on-Chip System ( System On Chip, SOC), Complex Programmable Logic Device ( Complex Programmable Logic Device,The present invention may be implemented in a variety of computer-implemented systems, including a CPLD, computer hardware, firmware, software, and / or combinations thereof. In these various embodiments, the present invention may be embodied in one or more computer programs that can be executed and / or interpreted by a programmable system that includes at least one programmable processor, which may be a special-purpose or general-purpose programmable processor, and that can receive data and instructions from and transmit data and instructions to a storage system, at least one input device, and at least one output device.
[0161] Program codes for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus, so that when the program code is executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are performed. The program code can be executed entirely on the device, partially on the device, as a separate software package, partially on the device and partially on a remote device, or entirely on a remote device or server.
[0162] In the context of this disclosure, a machine-readable medium may be a tangible medium, and may contain or store a program for use by or in combination with an instruction execution system, device, or apparatus. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination of the above. More specific examples of machine-readable storage media include one or more wire-based electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory ( Erasable Programmable Read Only Memory, EPROM ) or flash memory, optical fiber, portable compact disc read-only memory ( Compact Disc Read Only Memory, The content may include a CD-ROM, optical storage, magnetic storage, or any suitable combination of the above.
[0163] To provide for user interaction, the systems and techniques described herein can be implemented in a computer, which may include a display device (e.g., ,shadow polar ray tube (Cathode Ray Tube, CRT )also Liquid LCD display (Liquid Crystal Display, LCD ) monitor), and a keyboard and pointing device (e.g., a mouse or trackball) through which a user can provide input to the computer. Other types of devices may also be used to provide interaction with the user. For example, feedback provided to the user may be any form of sensing feedback (e.g., visual feedback, auditory feedback, or haptic feedback), and may receive input from the user in any form (including acoustic input, voice input, or haptic input).
[0164] The systems and techniques described herein may be implemented in a computing system including a back-end component (e.g., a data server), or a computing system including a middleware component (e.g., an application server), or a computing system including a front-end component (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system including any combination of such back-end, middleware, or front-end components. The components of the system may be connected to each other by any form or medium of digital data communication (e.g., a communications network). An example of a communications network is a local area network ( Local Area Networks, LAN), Wide Area Network ( Wide Area Networks, WAN) and the Internet.
[0165] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. The relationship between the client and the server is created by computer programs running on corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in a cloud computing service system, thereby resolving the drawbacks of traditional physical hosts and VPS services (also referred to as "Virtual Private Server" or "VPS"), such as high management difficulty and poor service scalability. The server may be a server in a distributed system or a server combined with blockchain.
[0166] As can be understood, the order of steps can be changed, added, or deleted using the various forms of flow shown above. For example, the steps described in the present disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present application can be achieved, and are not limited to the present specification.
[0167] The above specific embodiments do not limit the scope of protection of the present application. It should be understood that those skilled in the art can make various modifications, combinations, sub-combinations, and substitutions based on design requirements and other factors. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present disclosure should be included in the scope of protection of the present disclosure.
Claims
1. 1. A method for generating a facial image based on mouth shape, comprising: acquiring recognition target audio data and a preset face image; determining audio features of the recognition target audio data, the audio features including speaking rate features and semantic features; and processing the predetermined face image based on the speech rate feature and the semantic feature to generate a face image having a mouth shape, Determining audio features of the recognition target audio data includes: determining a speaking rate feature of the recognition-target audio data based on a predetermined first feature extraction model, the first feature extraction model being used to extract the speaking rate feature from the recognition-target audio data; determining semantic features of the recognition-target audio data based on a predetermined second feature extraction model, wherein the second feature extraction model is used to extract semantic features from the recognition-target audio data; A method for generating facial images based on mouth shapes.
2. determining a speech rate feature of the recognition target audio data based on a predetermined first feature extraction model, inputting the recognition-target audio data into a predetermined first feature extraction model to extract features, thereby obtaining phonetic posterior probability features of the recognition-target audio data, the phonetic posterior probability features representing phoneme category information of the recognition-target audio data; determining a speech rate feature of the recognition target audio data based on a speech posterior probability feature of the recognition target audio data; The method of claim 1.
3. determining a speech rate feature of the recognition target audio data based on a speech posterior probability feature of the recognition target audio data, performing a fast Fourier transform process on the speech posterior probability features to obtain frequency domain signal features, the frequency domain signal features representing phoneme category information of the recognition target audio data; dividing the frequency domain signal features into frequency domain signal features of at least two frequency bands based on a predetermined frequency band size; and performing an integration process on the frequency domain signal features of the at least two frequency bands to obtain a speech rate feature of the recognition target audio data. The method of claim 2.
4. determining semantic features of the recognition target audio data based on a second predetermined feature extraction model, inputting the recognition target audio data into a predetermined second feature extraction model to extract features, and outputting semantic features of the recognition target audio data. The method of claim 1.
5. Processing the predetermined face image based on the speech rate feature and the semantic feature to generate a face image having a mouth shape, inputting the speech rate feature and the semantic feature into a predetermined face and mouth shape determination model to process the feature, and generating a face image having a mouth shape based on the processed result and the predetermined face image. The method of claim 1.
6. The speech rate feature and the semantic feature are inputted into a predetermined face and mouth shape determination model for processing, and a face image having a mouth shape is generated based on the processed result and the predetermined face image, performing a combining process on the speech rate feature and the semantic feature based on the predetermined face and mouth type determination model to obtain a combined feature of the recognition target audio data, the combined feature representing the speech rate feature and the semantic feature; Performing feature extraction on the combined features based on a convolutional layer in the preset face and mouth shape determination model to obtain face driving parameters, the face driving parameters being used to represent parameters required to drive changes in mouth shape in a face image; performing image rendering on the preset face image based on the face driving parameters to generate a face image having a mouth shape; The method of claim 5.
7. The face driving parameters are weight parameters of a mixed deformation, and image rendering is performed on the predetermined face image based on the face driving parameters to generate a face image having a mouth shape, determining three-dimensional face mesh data corresponding to the predetermined face image based on weight parameters of the mixed transformation, the three-dimensional face mesh data being data representing a three-dimensional mesh model of a face surface in the face image; and performing image rendering on the preset face image based on the face three-dimensional mesh data to generate a face image having a mouth shape. The method of claim 6.
8. and when it is determined that the numerical value represented by the speech rate feature of the recognition target audio data is smaller than a predetermined speech rate threshold, processing the predetermined face image based on the semantic feature to generate a face image having a mouth shape. The method of claim 1.
9. 1. A method for training a face and mouth typing model, comprising: acquiring training target image data and a predetermined facial image, wherein the training target image data includes training target audio data and a training target facial image, and the training target facial image has a mouth shape corresponding to the training target audio data; determining audio features of the training subject audio data, the audio features including speaking rate features and semantic features; training an initial face and mouth shape determination model based on the speech rate feature, the semantic feature, and the predetermined face image to obtain a face image with a mouth shape; and determining that a training-completed face and mouth shape determination model is obtained when the face image having the mouth shape matches the training target face image. A method for training a face-mouth typing model.
10. Determining audio features of the training subject audio data includes: determining a speaking rate feature of the training target audio data based on a predetermined first feature extraction model, wherein the first feature extraction model is used to extract the speaking rate feature from the training target audio data; and determining semantic features of the training target audio data based on a second predetermined feature extraction model, wherein the second feature extraction model is used to extract semantic features from the training target audio data; 10. The method of claim 9.
11. determining a speaking rate feature of the training target audio data based on a predetermined first feature extraction model, inputting the training target audio data into a predetermined first feature extraction model to perform feature extraction, and obtaining phonetic posterior probability features of the training target audio data, the phonetic posterior probability features representing phoneme category information of the training target audio data; determining a speaking rate feature of the training target audio data based on speech posterior probability features of the training target audio data; The method of claim 10.
12. determining a speaking rate feature of the training target audio data based on speech posterior probability features of the training target audio data, performing a fast Fourier transform on the speech posterior probability features to obtain frequency domain signal features, the frequency domain signal features representing phoneme category information of the training audio data; dividing the frequency domain signal features into frequency domain signal features of at least two frequency bands based on a predetermined frequency band size; and performing an integration process on the frequency domain signal features of the at least two frequency bands to obtain a speaking rate feature of the training audio data. The method of claim 11.
13. determining semantic features of the training target audio data based on a predetermined second feature extraction model, inputting the training target audio data into a predetermined second feature extraction model to extract features, and outputting semantic features of the training target audio data; The method of claim 10.
14. training an initial face and mouth shape determination model based on the speech rate feature, the semantic feature, and the predetermined face image to obtain a face image with a mouth shape; performing a combining process on the speech rate feature and the semantic feature based on the initial face and mouth type determination model to obtain a combined feature of the training target audio data, wherein the combined feature represents a speech rate feature and a semantic feature; Based on a convolutional layer in the initial face and mouth shape determination model, perform feature extraction on the combined features to obtain face driving parameters, which are used to represent parameters required to drive changes in mouth shape in a face image; and performing image rendering on the preset face image based on the face driving parameters to obtain a face image having a mouth shape; 10. The method of claim 9.
15. The face driving parameters are weight parameters of a mixed deformation, and image rendering is performed on the predetermined face image based on the face driving parameters to obtain a face image having a mouth shape, determining three-dimensional face mesh data corresponding to the predetermined face image based on weight parameters of the mixed transformation, the three-dimensional face mesh data being data representing a three-dimensional mesh model of a face surface in the face image; and performing image rendering on the preset face image based on the face three-dimensional mesh data to generate a face image having a mouth shape.
15. The method of claim 14.
16. Obtaining training subject image data includes: obtaining the training target audio data; performing a three-dimensional reconstruction process of a facial image based on the training target audio data to obtain three-dimensional facial mesh data corresponding to the training target audio data; obtaining the training target face image based on three-dimensional facial mesh data corresponding to the training target audio data; 10. The method of claim 9.
17. An apparatus for generating a facial image based on a mouth shape, comprising: a data acquisition unit for acquiring audio data to be recognized and a predetermined face image; a feature determination unit used for determining audio features of the recognition target audio data, the audio features including speech rate features and semantic features; an image generating unit used to process the predetermined facial image based on the speech rate feature and the semantic feature to generate a facial image having a mouth shape; The feature determining unit includes a first determining module and a second determining module; the first determination module is used to determine a speaking rate feature of the recognition target audio data based on a predetermined first feature extraction model, and the first feature extraction model is used to extract the speaking rate feature from the recognition target audio data; the second determination module is used to determine semantic features of the recognition target audio data based on a predetermined second feature extraction model, and the second feature extraction model is used to extract semantic features from the recognition target audio data; A device for generating facial images based on mouth shapes.
18. 1. An apparatus for training a face-mouth type determination model, comprising: an image acquisition unit used to acquire training target image data and a predetermined face image, the training target image data including training target audio data and a training target face image, and the training target face image has a mouth shape corresponding to the training target audio data; a feature extraction unit used for determining audio features of the training target audio data, the audio features including speaking rate features and semantic features; a model training unit, which is used to train an initial face and mouth shape determination model according to the speech rate feature, the semantic feature and the predetermined face image, to obtain a face image with a mouth shape; a model acquisition unit, which is used for determining that a training-completed face and mouth shape determination model is obtained when the face image having the mouth shape matches the training target face image; A device for training face-mouth typing models.
19. An electronic device, at least one processor; and a memory communicatively coupled to the at least one processor; An electronic device in which instructions executable by the at least one processor are stored in the memory, and the instructions are executed by the at least one processor to cause the at least one processor to perform the method described in any one of claims 1 to 8 or claims 9 to 16.
20. 17. A non-transitory computer readable storage medium storing computer instructions, the computer instructions causing a computer to perform a method according to any one of claims 1 to 8 or claims 9 to 16.
21. A computer program, which, when executed by a processor, causes the steps of the method according to any one of claims 1 to 8 or 9 to 16 to be implemented.
Citation Information
Patent Citations
Live broadcast method, device and system of virtual anchor
CN116095357A