Method, apparatus and device for generating dynamic image based on audio, and storage medium

By generating dynamic images based on single images and audio, and using a generative network model to extract features and adjust the network, the problems of video acquisition and data cleaning in digital human production are solved, resulting in reduced costs and shorter production cycles.

CN118279453BActive Publication Date: 2025-11-28JIAXING SILICON INTELLIGENT TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410303751.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-08
Publication Date
2025-11-28
Estimated Expiration
2044-01-08

AI Technical Summary

Technical Problem

The current process of creating digital humans requires a lot of video capture and data cleaning, resulting in high costs and long cycles.

Method used

By generating dynamic images based on single images and audio, a trained generative network model is used to extract target head movement features and expression coefficient features. The network model is then adjusted to generate target dynamic images, avoiding video capture and data cleaning.

Benefits of technology

It reduced the cost of creating digital humans, shortened the production cycle, and reduced the amount of data required.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118279453B_ABST
    Figure CN118279453B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method and device for generating a dynamic image based on audio, an apparatus, and a storage medium, relating to the field of natural human-computer interaction. The method comprises: first obtaining a reference image and a reference audio input by a user; then, based on the reference image and a trained generation network model, determining a target head action feature and a target expression coefficient feature, and adjusting the trained generation network model based on the target head action feature and the target expression coefficient feature to obtain a target generation network model; finally, based on the reference audio, the reference image, and the target generation network model, processing a to-be-processed image to obtain a target dynamic image; wherein the to-be-processed image is the same as an image object in the reference image; in this way, a corresponding digital person can be obtained based on a single picture of a target person; in this way, video acquisition work and data cleaning work are not required, the production cost of the digital person can be reduced, and the production cycle of the digital person is shortened.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application, the original application's application number is 202410022841.6, the original application's original date is January 8, 2024, and the original application's entire content is incorporated by reference in this application. TECHNICAL FIELD

[0002] The application relates to the field of natural human-computer interaction, and in particular to a method and device for generating dynamic images based on audio, equipment and storage medium. BACKGROUND

[0003] Digital human (Digital Human / Meta Human) is a digitalized character image close to human image created by using digital technology. At present, the production process of digital human is as follows: video data of a target person speaking is collected; then, a deep learning network (such as a GAN network model) is used to learn the corresponding relationship between the voice and the lip shape of the target person in the video data, so as to obtain a trained network model; finally, a new audio is input into the trained network model, so that the trained network model generates the lip shape animation corresponding to the audio, thereby completing the production of the digital human.

[0004] However, the above-mentioned method of producing digital human needs a lot of video collection work and data cleaning work; that is, when a user wants to generate a corresponding digital human, a lot of video data of the user speaking needs to be obtained; and in order to ensure the effect of the obtained digital human, there is a certain requirement for the quality of the video data of the user speaking; thus, when generating the digital human corresponding to the user, it is relatively troublesome, the cost is too high, and the period is relatively long. SUMMARY

[0005] The application provides a method and device for generating dynamic images based on audio, equipment and storage medium, which can obtain a dynamic image of a target person based on a single picture of the target person, thereby obtaining a digital human; thus, video collection work and data cleaning work are not needed, the production cost of the digital human can be reduced, and the production period of the digital human is shortened.

[0006] In a first aspect, the application provides a method for generating dynamic images based on audio, comprising:

[0007] obtaining a reference image and a reference audio input by a user;

[0008] determining a target head action feature and a target expression coefficient feature based on the reference image and a trained generation network model;

[0009] adjusting the trained generation network model based on the target head action feature and the target expression coefficient feature to obtain a target generation network model;

[0010] Based on reference audio, reference image, and target generation network model, the image to be processed is processed to obtain a target dynamic image; the target dynamic image represents the dynamic image of the target person in the image to be processed changing facial expressions based on the reference audio; the image to be processed and the image objects in the reference image are the same.

[0011] In the above technical solution, based on the reference image and the trained generative network model, the target head action features and target expression coefficient features are determined, including:

[0012] Based on the reference image, reference data is obtained;

[0013] The trained generative network model extracts target head movement features and target facial expression coefficient features from the reference data.

[0014] In the above technical solution, the target head action features and target expression coefficient features are extracted from the reference data through a trained generative network model, including:

[0015] The trained generative network model determines multiple frames of facial images of the target person from the reference data.

[0016] The trained generative network model extracts features from the facial images of the target person in each frame to obtain the target head action features and the target expression coefficient features.

[0017] In the above technical solution, the target generation network model includes an affine subnetwork and a driving subnetwork;

[0018] Based on reference audio, reference image, and target generation network model, the image to be processed is processed to obtain a dynamic image of the target, including:

[0019] The image to be processed is processed by an affine sub-network to obtain a feature map to be processed, and a deformation feature map is obtained by the affine sub-network based on the reference audio, reference image and feature map to be processed.

[0020] By driving the sub-network based on the deformation feature map, the image to be processed is processed to obtain the target dynamic image.

[0021] In the above technical solution, the affine subnetwork includes a speech processing layer, a feature extraction layer, a feature fusion layer, and a feature affine layer;

[0022] The image to be processed is processed using an affine subnetwork to obtain a feature map to be processed. Then, based on the reference audio, reference image, and the feature map to be processed, a deformation feature map is obtained using the affine subnetwork, including:

[0023] The target Mel-spectral coefficient features corresponding to the reference audio are determined through the speech processing layer;

[0024] The reference image is feature-extracted through a feature extraction layer to obtain a reference feature map;

[0025] The to-be-processed image is feature-extracted through a feature extraction layer to obtain a to-be-processed feature map;

[0026] The reference feature map and the to-be-processed feature map are feature-stacked and aligned through a feature fusion layer to obtain a fusion feature map;

[0027] An affine coefficient is determined based on the fusion feature map and the target mel-frequency cepstrum coefficient feature through a feature affine layer, and a spatial deformation of the reference feature map is performed through affine transformation based on the affine coefficient to obtain a deformation feature map.

[0028] In the technical solution, the driving sub-network includes a feature processing layer, a feature synchronization layer, and an image driving layer;

[0029] The to-be-processed image is processed based on the deformation feature map through the driving sub-network to obtain a target dynamic image, including:

[0030] An initial driving feature is obtained based on the target mel-frequency cepstrum coefficient feature through the feature processing layer;

[0031] The to-be-processed feature map is driven and processed based on the initial driving feature through the image driving layer to obtain an initial feature map;

[0032] The deformation feature map and the initial feature map are stacked to determine a feature synchronization parameter between the deformation feature map and the initial feature map through the feature synchronization layer;

[0033] The initial driving feature is adjusted based on the feature synchronization parameter through the feature processing layer to obtain a target driving feature;

[0034] The to-be-processed image is driven and processed based on the target driving feature through the image driving layer to obtain the target dynamic image.

[0035] In the technical solution, the method further includes:

[0036] A sample video is obtained; wherein a video object in the sample video is different from an image object in the to-be-processed image;

[0037] The sample video is processed through the to-be-trained generation network model to extract sample audio data and sample image data;

[0038] The sample audio data and the sample image data are processed based on the to-be-trained generation network model to obtain a prediction training result;

[0039] The prediction training result is taken as an initial training output of the to-be-trained generative network model, the sample image data is taken as supervision information, and the to-be-trained generative network model is iteratively trained to obtain the trained generative network model.

[0040] In the technical solution, the sample audio data and the sample image data are processed based on the to-be-trained generative network model to obtain the prediction training result, including:

[0041] The reference mel-frequency cepstral coefficient feature is extracted from the sample audio data based on the to-be-trained generative network model.

[0042] The reference head action, the reference expression coefficient feature and the reference face feature are extracted from the sample image data based on the to-be-trained generative network model.

[0043] The prediction training result is obtained based on the reference mel-frequency cepstral coefficient feature, the reference head action, the reference expression coefficient feature and the reference face feature through the to-be-trained generative network model.

[0044] In the technical solution, the prediction training result is taken as an initial training output of the to-be-trained generative network model, the sample image data is taken as supervision information, and the to-be-trained generative network model is iteratively trained to obtain the trained generative network model, including:

[0045] The loss value is determined according to the prediction training result and the sample image data.

[0046] The to-be-trained generative network model is iteratively updated according to the loss value to obtain the trained generative network model.

[0047] In the second aspect of the application, an apparatus for generating a dynamic image based on audio is provided, including:

[0048] The acquisition module is configured to acquire a reference image and reference audio input by a user.

[0049] The processing module is configured to determine target head action features and target expression coefficient features based on the reference image and the trained generative network model.

[0050] The adjustment module is configured to adjust the trained generative network model based on the target head action features and the target expression coefficient features to obtain a target generative network model.

[0051] The processing module is further configured to process a to-be-processed image based on the reference audio, the reference image and the target generative network model to obtain a target dynamic image; the target dynamic image represents a dynamic image of a target person changing facial expressions based on the reference audio in the to-be-processed image; the to-be-processed image and an image object in the reference image are the same.

[0052] In a third aspect, the application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method of any of the above embodiments when executing the program.

[0053] In a fourth aspect, the application provides a non-transitory computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the method of any of the above embodiments.

[0054] The embodiments of the application provide a method, device and equipment for generating a dynamic image based on audio and a storage medium, wherein the method for generating a dynamic image based on audio comprises the following steps: obtaining a reference image and reference audio input by a user; determining target head action features and target expression coefficient features based on the reference image and a trained generation network model; adjusting the trained generation network model based on the target head action features and the target expression coefficient features to obtain a target generation network model; processing a to-be-processed image based on the reference audio, the reference image and the target generation network model to obtain a target dynamic image; the target dynamic image represents a dynamic image of a target person changing facial expressions based on the reference audio in the to-be-processed image; the image object in the to-be-processed image is the same as that in the reference image; in this way, a digital person (a dynamic image of the target person changing facial expressions based on the reference audio) can be obtained based on a single picture (the reference image) of the target person; in this way, video acquisition work and data cleaning work are not required, the production cost of the digital person can be reduced, and the production cycle of the digital person is shortened. BRIEF DESCRIPTION OF DRAWINGS

[0055] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the application, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0056] Figure 1 A flowchart of a method for generating a dynamic image based on audio provided by the embodiments of the application is shown in the figure.

[0057] Figure 2 A flowchart of another method for generating a dynamic image based on audio provided by the embodiments of the application is shown in the figure.

[0058] Figure 3 A flowchart of still another method for generating a dynamic image based on audio provided by the embodiments of the application is shown in the figure.

[0059] Figure 4 A structure diagram of a target generation network model provided by the embodiments of the application is shown in the figure.

[0060] Figure 5A A flowchart of another method for generating a dynamic image based on audio provided by an embodiment of the present application is shown in FIG. 6;

[0061] Figure 5B A flowchart of another method for generating a dynamic image based on audio provided by an embodiment of the present application is shown in FIG. 6;

[0062] Figure 5C A flowchart of another method for generating a dynamic image based on audio provided by an embodiment of the present application is shown in FIG. 6;

[0063] Figure 6 A flowchart of another method for generating a dynamic image based on audio provided by an embodiment of the present application is shown in FIG. 6;

[0064] Figure 7 A flowchart of another method for generating a dynamic image based on audio provided by an embodiment of the present application is shown in FIG. 6;

[0065] Figure 8 A flowchart of another method for generating a dynamic image based on audio provided by an embodiment of the present application is shown in FIG. 6;

[0066] Figure 9 A flowchart of another method for generating a dynamic image based on audio provided by an embodiment of the present application is shown in FIG. 6;

[0067] Figure 10 A flowchart of a method for training a generation network model provided by an embodiment of the present application is shown in FIG. 7;

[0068] Figure 11 A structural diagram of a device for generating a dynamic image based on audio provided by an embodiment of the present application is shown in FIG. 8;

[0069] Figure 12 A structural diagram of an electronic device provided by an embodiment of the present application is shown in FIG. 9. DETAILED DESCRIPTION

[0070] In order to make the purpose, technical solutions and advantages of the present application clearer, the technical solutions in the present application will be described clearly and completely below with reference to the drawings in the present application.

[0071] Digital Human / Meta Human is a digital character image close to human image created by using digital technology. With the popularization of the concept of metaverse, digital human enters the public life. At present, the production process of digital human is as follows: video data of a target person speaking is collected; then, a deep learning network (such as a GAN network model) is used to learn the corresponding relationship between the voice and the lip shape of the target person in the video data, so as to obtain a trained network model; finally, a new audio is input into the trained network model, so that the trained network model generates the lip shape animation corresponding to the audio, thereby completing the production of the digital human.

[0072] However, the above-mentioned digital human production method requires a large amount of video collection work and data cleaning work; that is, when a user wants to generate a corresponding digital human, a large amount of video data of the user speaking needs to be obtained, and in order to ensure the effect of the obtained digital human, there is also a certain requirement for the quality of the video data of the user speaking. And because the video data needs to be learned frame by frame, the data volume is large, and the hardware configuration requirement of the device is also high. Therefore, when generating a digital human corresponding to the user, it is troublesome, the cost is too high, and the period is too long.

[0073] In order to solve the above-mentioned technical problems, the present application provides a method for generating dynamic images based on audio, which can obtain a digital human (dynamic image of a target person changing facial expressions based on reference audio) based on a single picture (reference image) of the target person. In this way, video collection work and data cleaning work are not required, the production cost of the digital human can be reduced, and the production period of the digital human is shortened.

[0074] Referring to Figure 1 The embodiment of the present application provides a method for generating dynamic images based on audio, which includes S101-S104.

[0075] S101, obtaining a reference image and a reference audio input by a user.

[0076] In some embodiments, the reference image input by the user is a single photo of the target person, that is, the number of reference images is one. The reference image shows the front information of the head of the target person, that is, the reference image can completely expose the face of the person or completely expose the mouth area of the person. The reference image can be an image downloaded from the Internet, or an image of the person taken by the user through a mobile terminal with a camera function, or a virtual digital person image, an animation character image, etc. The embodiment of the present application does not limit the acquisition method of the reference image and the type of the person in the reference image. The following embodiments will be exemplarily described taking the acquisition method of the reference image as an example.

[0077] In some embodiments, the reference audio input by the user can be an audio downloaded from the Internet or an audio recorded by the user. The embodiments of the present application do not limit the acquisition method of the reference audio. In the following embodiments, the acquisition method of the reference audio is taken as an example of downloading from the Internet.

[0078] In some embodiments, the obtained reference image can be a pre-processed reference image or an unprocessed reference image, and the obtained reference audio can be a pre-processed reference audio or an unprocessed reference audio. If the reference image and / or the reference audio are unprocessed, the reference image and / or the reference audio are pre-processed after S101. The pre-processing method of the reference image can be cropping, noise reduction, etc., and the pre-processing method of the reference audio can be noise reduction, audio enhancement, editing, etc. The embodiments of the present application do not limit the pre-processing time and the pre-processing method of the reference image and / or the reference audio. In the following embodiments, the reference image and the reference audio input by the user in S101 are taken as an example of being pre-processed.

[0079] S102, determining the target head action feature and the target expression coefficient feature based on the reference image and the trained generative network model.

[0080] In some embodiments, the reference image is input into the trained network model. The trained generative network model processes the reference image to determine the target head action feature and the target expression coefficient feature from the reference image. The trained generative network model includes a plurality of generators. The trained generative network model generates a plurality of predicted images based on the reference image through each generator. Then, the target head action feature and the target expression coefficient feature are learned based on the difference between each predicted image and the reference image.

[0081] It should be noted that when there is only one person in the reference image, the trained generative network model can directly process the reference image to determine the target head action feature and the target expression coefficient feature. When there are multiple persons in the reference image, the trained generative network model can identify the number of persons in the reference image and select a target person whose face is fully exposed.

[0082] Referring to Figure 2 In some embodiments, S102 can include S1021-S1022.

[0083] S1021, obtaining reference data based on the reference image.

[0084] In some embodiments, the preset time length of video is generated based on the reference image, and the video is processed to obtain the reference data. The preset time length of video generated based on the reference image can be a silent video, i.e., the audio in the video is blank audio.

[0085] For example, a 1-minute video is generated based on the reference image, and the 1-minute video is processed in units of frames. Each frame of image is a reference image.

[0086] In some embodiments, the reference data can be obtained based on the reference image by the trained generation network model, or the reference data can be obtained based on the reference image by the electronic device provided with the trained generation network model before the reference data is input into the trained generation network model.

[0087] The generation method of the reference data in the embodiments of the present application is not limited. In the embodiments of the present application, the reference data obtained based on the reference image by the trained generation network model will be exemplarily described.

[0088] S1022, extracting target head action features and target expression coefficient features from the reference data by the trained generation network model.

[0089] In some embodiments, the reference data is input into the trained generation network model to determine the target human face part from the reference data by the trained generation network model; then, the target head action features and the target expression coefficient features are extracted from the target human face part.

[0090] For example, the target head action features can represent the orientation of the head of the target person, and can also represent the action of the head of the target person. The target expression coefficient features can represent at least one mouth action of the target person, and can also represent the action of the remaining organs of the target person other than the mouth, such as the eye. The content of the target head action features and the content of the target expression coefficient features are not limited in the embodiments of the present application. In the embodiments of the present application, the target head action features include the head action information of the target person, and the target expression coefficient features include the mouth action information of the target person will be exemplarily described.

[0091] Referring to Figure 3 In some embodiments, S1022 can include S401-S402.

[0092] S401, determining the face images of the target person in multiple frames from the reference data by the trained generation network model.

[0093] In some embodiments, after the reference data is input into the trained generative network model, the trained generative network model learns based on the reference data to determine the facial images of the multiple frames of the target person from the reference data.

[0094] S402, feature extraction is performed on the facial images of each frame of the target person based on the trained generative network model to obtain target head action features and target expression coefficient features.

[0095] In some embodiments, the corresponding features are extracted from the facial images of each frame of the target person by the trained generative network model, i.e., feature extraction is performed on the facial images of each frame of the target person; to obtain target head action features and target expression coefficient features.

[0096] S103, adjusting the trained generative network model based on the target head action features and the target expression coefficient features to obtain a target generative network model.

[0097] In some embodiments, the parameters of the trained generative network model are adjusted based on the target head action features and the target expression coefficient features, so that the trained generative network model is more suitable for the target person. After adjusting the trained network model, a target generative network model is obtained.

[0098] S104, processing the to-be-processed image based on the reference audio, the reference image and the target generative network model to obtain a target dynamic image.

[0099] The target dynamic image represents a dynamic image of the target person changing facial expressions based on the reference audio in the to-be-processed image; the image object in the to-be-processed image is the same as that in the reference image.

[0100] In some embodiments, the reference audio and the reference image are input into the target generative network model, and the target generative network processes the to-be-processed image based on the reference audio and the reference image to obtain a target dynamic image. In the target dynamic image, the target region changes according to the reference audio, so that the facial expression of the target person in the to-be-processed image can change according to the reference audio.

[0101] For example, the target region includes at least one of a mouth region, an eye region, a nose region, an ear region or a brow region. The target region is not limited in the embodiments of the present application. The target region includes a mouth region, an eye region and a brow region in the embodiments of the present application.

[0102] In some embodiments, the reference audio includes a text voice.

[0103] Exemplarily, if the reference voice is "Welcome to experience", the target dynamic image obtained by the target generation network based on the text voice "Welcome to experience" has a facial expression change of the target person corresponding to the text voice "Welcome to experience".

[0104] Since the target generation network model is obtained by adjusting the trained generation network model based on the target head action feature and the target expression coefficient feature, the model structure of the target generation network model is the same as that of the trained generation network model. As shown in Figure 4 In some embodiments, the target generation network model includes an affine subnetwork and a driving subnetwork. The affine subnetwork includes a feature voice processing layer, a feature extraction layer, a feature fusion layer, and a feature affine layer; the driving subnetwork includes a feature processing layer, a feature synchronization layer, and an image driving layer.

[0105] Based on Figure 4 , as shown in Figure 5A In some embodiments, S104 can include S1041-S1042.

[0106] S1041, processes the to-be-processed image through the affine subnetwork to obtain a to-be-processed feature map, and obtains a deformation feature map based on the reference audio, the reference image, and the to-be-processed feature map through the affine subnetwork.

[0107] In some embodiments, the affine subnetwork includes a voice processing layer, a feature extraction layer, a feature fusion layer, and a feature affine layer.

[0108] Referring to Figure 5B , in some embodiments, S1041 can include S10411-S10415.

[0109] S10411, determines the target mel-frequency cepstrum coefficient feature corresponding to the reference audio through the voice processing layer.

[0110] In some embodiments, the voice processing layer converts the energy spectrum of the reference audio in the frequency domain into the energy distribution in the mel-frequency scale to determine the target mel-frequency cepstrum coefficient feature corresponding to the reference audio from the reference audio.

[0111] S10412, extracts features of the reference image through the feature extraction layer to obtain a reference feature map.

[0112] In some embodiments, the feature extraction layer extracts features of the reference image to obtain feature vectors of each pixel in the reference image, thereby obtaining the reference feature map corresponding to the reference image.

[0113] S10413, extracts features of the to-be-processed image through the feature extraction layer to obtain a to-be-processed feature map.

[0114] In some embodiments, the feature extraction layer performs feature extraction on the to-be-processed image to obtain a feature vector of each pixel in the to-be-processed image, thereby obtaining a to-be-processed feature map corresponding to the to-be-processed image.

[0115] In S10414, the feature fusion layer performs feature stacking alignment processing on the reference feature map and the to-be-processed feature map to obtain a fused feature map.

[0116] For example, the feature fusion layer stacks the reference feature map and the to-be-processed feature map along the feature channel, and then inputs the stacked reference feature map and to-be-processed feature map into an alignment encoder to obtain the fused feature map.

[0117] In S10415, the feature affine layer determines an affine coefficient based on the fused feature map and the target mel-frequency cepstrum coefficient feature, and performs spatial deformation of the reference feature map by affine transformation based on the affine coefficient to obtain a deformed feature map.

[0118] For example, after the affine coefficient is determined, the feature affine layer performs spatial deformation of each feature channel in the reference feature map by affine transformation based on the affine coefficient to obtain the deformed feature map.

[0119] In S1042, the driving sub-network processes the to-be-processed image based on the deformed feature map to obtain a target dynamic image.

[0120] In some embodiments, the driving sub-network includes a feature processing layer, a feature synchronization layer, and an image driving layer.

[0121] Referring to Figure 5C In some embodiments, S1042 can include S10421-S10425.

[0122] In S10421, the feature processing layer obtains an initial driving feature based on the target mel-frequency cepstrum coefficient feature.

[0123] In S10422, the image driving layer performs driving processing on the to-be-processed feature map based on the initial driving feature to obtain an initial feature map.

[0124] In S10423, the feature synchronization layer performs stacking processing on the deformed feature map and the initial feature map to determine a feature synchronization parameter between the deformed feature map and the initial feature map.

[0125] In S10424, the feature processing layer adjusts the initial driving feature based on the feature synchronization parameter to obtain a target driving feature.

[0126] In S10425, the image driving layer performs driving processing on the to-be-processed image based on the target driving feature to obtain a target dynamic image.

[0127] Exemplarily, the target driving feature characterizes a target region in the to-be-processed image. In this way, the target region in the to-be-processed image is driven based on the target driving feature by the image driving layer to obtain a target dynamic image. For example, the mouth region in the target region is driven so that the mouth of the target person in the to-be-processed image changes from a closed state to an "O" shape; the eye region in the target region is driven so that the eyes of the target person in the to-be-processed image change from an open state to a closed state; and the eyebrow region in the target region is driven so that the eyebrows of the target person in the to-be-processed image are raised. It should be noted that, in the case where the target region includes multiple regions, the multiple target regions in the to-be-processed image can be driven in sequence by the image driving layer, for example, the mouth region is driven first, then the eye region is driven, and finally the eyebrow region is driven; or the multiple target regions can be driven simultaneously. The present embodiment does not limit this.

[0128] In some embodiments, the target mel-frequency cepstrum coefficient feature corresponding to the reference audio can be obtained by feature extraction on the reference audio through the target generation network model, or can be obtained by preprocessing the reference audio, or can be obtained by processing the reference audio through other models in the electronic device provided with the target generation network model. The present embodiment does not limit the manner of determining the target mel-frequency cepstrum coefficient feature corresponding to the reference audio. The present embodiment takes the target mel-frequency cepstrum coefficient feature corresponding to the reference audio obtained by processing the reference audio through other models in the electronic device provided with the target generation network model as an example for illustrative description.

[0129] It can be understood that the method for generating a dynamic image based on audio proposed in the present application can learn the correspondence between the head movement, expression, and audio lip shape of others through the trained generation network model, and can be migrated to a single to-be-processed image (photo). In this way, a digital person with head and expression changes can be made using a single to-be-processed image. This can reduce the production cost of the digital person without the need for video acquisition and data cleaning work. Moreover, because the to-be-processed image is one, i.e., one frame, the amount of data to be processed can be reduced, thereby shortening the production cycle of the digital person.

[0130] Reference is made to Figure 6 In some embodiments, the method for generating a dynamic image based on audio provided in the present application further includes S601-S604.

[0131] S601, acquiring a sample video.

[0132] In this way, the video object in the sample video is different from the image object in the to-be-processed image.

[0133] In some embodiments, the sample video shows the frontal information of the head of the person (any third person), and the sample video includes the number of audio to be extracted; that is, the sample video shows the facial information of the person when speaking. In this way, through the sample video, the facial expression of the person during speaking can be determined.

[0134] The sample video can be a video downloaded from a public channel (such as the Internet), a video of the person taken by the user through a mobile terminal with a camera function, or a virtual digital person image, an animation person image, and the like. The embodiments of the present application do not limit the acquisition method of the sample video and the type of the person in the sample video. In the following embodiments, the sample video downloaded from the network will be taken as an example for illustrative description.

[0135] In some embodiments, one sample video can be acquired, or multiple sample videos can be acquired to ensure the training effect of the subsequent generated network model to be trained. The embodiments of the present application do not limit the number of acquired sample videos. In the following embodiments, multiple sample videos will be taken as an example for illustrative description.

[0136] S602, processing the sample video through the generated network model to be trained to extract sample audio data and sample image data.

[0137] In some embodiments, the sample video is input into the generated network model to be trained as training data, and the sample video is processed through the generated network model to be trained to extract sample audio data and sample image data from the sample video. The sample image data includes at least one sample image. The sample image shows the frontal information of the head of the person.

[0138] In some embodiments, the facial expressions of the person in each sample image can be the same or different. The embodiments of the present application do not limit the number of sample images in the sample image data and the content of the sample images. In the following embodiments, it will be taken as an example for illustrative description that the sample image data includes multiple sample images, and the facial expressions of the person in each sample image are different.

[0139] S603, processing the sample audio data and the sample image data based on the generated network model to be trained to obtain a predicted training result.

[0140] In some embodiments, the sample audio data and the sample image data are processed through the generated network model to be trained to obtain a predicted training result.

[0141] Referring to Figure 7 In some embodiments, S603 can include S6031-S6033.

[0142] S6031, extract reference mel-frequency cepstral coefficient features from the sample audio data based on the to-be-trained generative network model.

[0143] In some embodiments, the sample audio data is feature-extracted by the to-be-trained generative network model to extract the reference mel-frequency cepstral coefficient features from the sample audio data. Wherein, the manner of extracting the reference mel-frequency cepstral coefficient features from the sample audio data by the to-be-trained generative network model can be the same as the manner of determining the target mel-frequency cepstral coefficient features from the reference audio in S1041.

[0144] S6032, extract reference head motion, reference expression coefficient features and reference face features from the sample image data based on the to-be-trained generative network model.

[0145] In some embodiments, the sample image data is feature-extracted by the to-be-trained generative network model to extract the reference head motion, reference expression coefficient features and reference face features from the sample image data.

[0146] S6033, obtain a predicted training result based on the reference mel-frequency cepstral coefficient features, the reference head motion, the reference expression coefficient features and the reference face features by the to-be-trained generative network model.

[0147] In some embodiments, the reference mel-frequency cepstral coefficient features, the reference head motion, the reference expression coefficient features and the reference face features are input into the to-be-trained generative network model, and a predicted training result is obtained based on the reference mel-frequency cepstral coefficient features, the reference head motion, the reference expression coefficient features and the reference face features by the to-be-trained generative network model.

[0148] Wherein, the structure of the to-be-trained generative network model is the same as that of the trained generative network model, which will not be repeated here.

[0149] S604, take the predicted training result as the initial training output of the to-be-trained generative network model, take the sample image data as the supervision information, and iteratively train the to-be-trained generative network model to obtain the trained generative network model.

[0150] Referring to Figure 8 In some embodiments, S604 can include:

[0151] S6041, determine a loss value according to the predicted training result and the sample image data.

[0152] In some embodiments, the loss value between the predicted training result and the sample image data is determined based on a preset loss function. Wherein, the preset loss function can be a cross-entropy loss function.

[0153] The embodiments of the present application do not limit the manner of determining the loss value of the predicted training result and the sample image data. The following embodiments will be exemplarily described by taking the manner of determining the loss value of the predicted training result and the sample image data through a preset loss function.

[0154] S6042, according to the loss value, iteratively updating the to-be-trained generative network model to obtain the trained generative network model.

[0155] In some embodiments, according to the loss value, the to-be-trained generative network model is iteratively updated until the loss value no longer increases, or the loss value is lower than a preset loss threshold, to obtain the trained generative network model.

[0156] As shown in Figure 9 In some embodiments, the present application provides another method for generating a dynamic image based on audio, which comprises: first obtaining reference video data (reference data) and speaking audio (reference audio). Wherein, the reference video data and the speaking audio can be obtained simultaneously, or the reference video data and the speaking audio can be obtained separately, and the present application does not limit the order of obtaining the reference video data and the speaking audio. The following embodiments will be exemplarily described by taking the reference video data and the speaking audio obtained separately as an example.

[0157] After obtaining the reference video data and the speaking audio separately, the reference video data is processed to obtain a target face part. Then, feature extraction is performed on the target face part to obtain head pose (target head action feature) and expression coefficient (target expression coefficient feature). Feature extraction is also performed on the speaking audio to obtain target MFCC (Mel-Frequency Cepstral Coefficient) feature. The present application does not limit the order of feature extraction of the speaking audio and the target face part. The present application will be exemplarily described by taking the feature extraction of the speaking audio and the target face part simultaneously as an example.

[0158] After obtaining the expression coefficient, the head pose and the target MFCC feature, the expression coefficient, the head pose, the target MFCC feature and a single photo (to-be-processed image) are input into the trained generator network (target generative network model) to generate a dynamic photo (target dynamic image) based on the expression coefficient, the head pose, the target MFCC feature and the single photo through the trained generator network. Wherein, the trained generator network is obtained based on the trained generative network model, that is, the head pose and the expression coefficient are input into the trained generative network model to fine-tune the model parameters in the trained generative network model, thereby obtaining the post-trained generator network.

[0159] AsFigure 10 As shown in some embodiments, the present application also provides a method for training a generative network model, comprising: first obtaining sample video data; then processing the video sample data to obtain standard digital audio data (wav) and sample face part (sample image data). After that, extracting features from the wav (also referred to as sample audio data) to obtain sample MFCC features; processing the sample face part to obtain reference head pose (reference head action), reference expression coefficient features and reference face frame (reference face features). The order of feature extraction of the wav and the sample face part is not limited in the embodiments of the present application. The embodiments of the present application will be exemplarily described by taking the example of simultaneously extracting features from the wav and the sample face part.

[0160] After obtaining the reference head pose, the reference expression coefficient features, the reference face frame and the sample MFCC features, the reference head pose, the reference expression coefficient features, the reference face frame and the sample MFCC features are input into the generator network (the generative network model to be trained) to generate predicted face data (predicted training result) through the generator network. Finally, based on the loss value between the predicted face data and the sample face part, the trained generative network model is obtained, thereby completing the training of the generative network model.

[0161] Corresponding to the foregoing embodiments of the method for generating dynamic images based on audio, the present application also provides embodiments of a device for generating dynamic images based on audio.

[0162] Reference Figure 11 The embodiments of the present application provide a device for generating dynamic images based on audio, comprising:

[0163] The acquisition module 1101 is configured to acquire reference images and reference audio input by a user.

[0164] The processing module 1102 is configured to determine target head action features and target expression coefficient features based on the reference images and the trained generative network model.

[0165] The adjustment module 1103 is configured to adjust the trained generative network model based on the target head action features and the target expression coefficient features to obtain a target generative network model.

[0166] The processing module 1102 is further configured to process a to-be-processed image based on the reference audio, the reference images and the target generative network model to obtain a target dynamic image; the target dynamic image represents a dynamic image of a target person changing facial expressions based on the reference audio in the to-be-processed image; the to-be-processed image and the image object in the reference images are the same.

[0167] In some embodiments, the processing module 1102 is further configured to obtain reference data based on the reference image; and extract target head motion features and target expression coefficient features from the reference data by using the trained generative network model.

[0168] In some embodiments, the processing module 1102 is further configured to determine the face images of the target person in multiple frames from the reference data by using the trained generative network model; and extract features based on the face images of the target person in each frame by using the trained generative network model to obtain the target head motion features and the target expression coefficient features.

[0169] In some embodiments, the target generative network model comprises an affine sub-network and a driving sub-network.

[0170] The processing module 1102 is further configured to process the to-be-processed image by using the affine sub-network to obtain a to-be-processed feature map, and obtain a deformation feature map based on the reference audio, the reference image and the to-be-processed feature map by using the affine sub-network; and process the to-be-processed image based on the deformation feature map by using the driving sub-network to obtain the target dynamic image.

[0171] In some embodiments, the affine sub-network comprises a speech processing layer, a feature extraction layer, a feature fusion layer and a feature affine layer.

[0172] The processing module 1102 is further configured to determine target mel-frequency cepstral coefficient features corresponding to the reference audio by using the speech processing layer; extract features of the reference image by using the feature extraction layer to obtain a reference feature map; extract features of the to-be-processed image by using the feature extraction layer to obtain a to-be-processed feature map; perform feature stacking and alignment processing on the reference feature map and the to-be-processed feature map by using the feature fusion layer to obtain a fusion feature map; determine affine coefficients based on the fusion feature map and the target mel-frequency cepstral coefficient features by using the feature affine layer, and perform spatial deformation of affine transformation on the reference feature map based on the affine coefficients by using the feature affine layer to obtain the deformation feature map.

[0173] In some embodiments, the driving sub-network comprises a feature processing layer, a feature synchronization layer and an image driving layer.

[0174] The processing module 1102 is further configured to obtain initial driving features based on the target mel-frequency cepstrum coefficient features through the feature processing layer; perform driving processing on the to-be-processed feature map based on the initial driving features through the image driving layer to obtain an initial feature map; perform stacking processing on the morphological feature map and the initial feature map through the feature synchronization layer to determine feature synchronization parameters between the morphological feature map and the initial feature map; adjust the initial driving features based on the feature synchronization parameters through the feature processing layer to obtain target driving features; and perform driving processing on the to-be-processed image based on the target driving features through the image driving layer to obtain the target dynamic image.

[0175] In some embodiments, the acquisition module 1101 is further configured to acquire a sample video; wherein a video object in the sample video is different from an image object in the to-be-processed image.

[0176] The processing module 1102 is further configured to process the sample video through the to-be-trained generation network model to extract sample audio data and sample image data; further configured to process the sample audio data and the sample image data based on the to-be-trained generation network model to obtain a predicted training result; further configured to take the predicted training result as an initial training output of the to-be-trained generation network model, take the sample image data as supervision information, and iteratively train the to-be-trained generation network model to obtain the trained generation network model.

[0177] In some embodiments, the processing module 1102 is further configured to extract reference mel-frequency cepstrum coefficient features from the sample audio data based on the to-be-trained generation network model; further configured to extract reference head motion, reference expression coefficient features, and reference face features from the sample image data based on the to-be-trained generation network model; and further configured to obtain the predicted training result based on the reference mel-frequency cepstrum coefficient features, the reference head motion, the reference expression coefficient features, and the reference face features through the to-be-trained generation network model.

[0178] In some embodiments, the processing module 1102 is further configured to determine a loss value according to the predicted training result and the sample image data; and further configured to iteratively update the to-be-trained generation network model according to the loss value to obtain the trained generation network model.

[0179] As Figure 12As shown, the electronic device provided by the embodiment of the present application can include a processor 1210, a communications interface 1220, a memory 1230 and a communications bus 1240, wherein the processor 1210, the communications interface 1220 and the memory 1230 complete mutual communication through the communications bus 1240. The processor 1210 can invoke the logical instructions in the memory 1230 to execute the above-mentioned methods.

[0180] In addition, the logical instructions in the memory 1230 described above can be implemented in the form of a software functional unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, includes several instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the switch device mechanical state monitoring method described in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk and various program codes that can be stored in the medium.

[0181] In another aspect, the present application also provides a non-transitory computer readable storage medium having a computer program stored thereon, which is executed by a processor to implement the above-mentioned methods.

[0182] The device embodiments described above are only schematic, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e. they can be located in one place, or distributed on a plurality of network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement it without creative labor.

[0183] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0184] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for generating dynamic images based on audio, characterized in that, include: Obtain the user-input reference audio, reference image, and image to be processed; The image to be processed is the same as the image object in the reference image; The image to be processed is processed by the affine sub-network in the target generation network model to obtain the feature map to be processed, and the deformation feature map is obtained by the affine sub-network based on the reference audio, the reference image and the feature map to be processed; The affine subnetwork is used to determine the target Mel-frequency cepstral coefficient features corresponding to the reference audio, extract features from the reference image to obtain a reference feature map, and perform an affine transformation on the reference feature map to obtain the deformation feature map; The target dynamic image is obtained by processing the image to be processed based on the deformation feature map through the driving sub-network in the target generation network model; The driving sub-network is used to drive the image to be processed to obtain the target dynamic image.

2. The method according to claim 1, characterized in that, The affine subnetwork includes a speech processing layer, a feature extraction layer, a feature fusion layer, and a feature affine layer; The process of processing the image to be processed through an affine sub-network in the target generation network model to obtain a feature map to be processed, and obtaining a deformation feature map based on the reference audio, the reference image, and the feature map to be processed through the affine sub-network, includes: The target Mel-spectral coefficient features corresponding to the reference audio are determined through the speech processing layer; The reference image is used to extract features through the feature extraction layer to obtain a reference feature map; The feature extraction layer extracts features from the image to be processed to obtain a feature map to be processed. The reference feature map and the feature map to be processed are subjected to feature stacking and alignment processing through the feature fusion layer to obtain a fused feature map; The affine layer determines the affine coefficients based on the fused feature map and the target Mel-Cepstral coefficients, and performs spatial deformation of the reference feature map by affine transformation based on the affine coefficients to obtain the deformed feature map.

3. The method according to claim 2, characterized in that, The driving sub-network includes a feature processing layer, a feature synchronization layer, and an image driving layer; The step of processing the image to be processed based on the deformation feature map through the driving sub-network in the target generation network model to obtain the target dynamic image includes: The initial driving features are obtained through the feature processing layer based on the target Mel-frequency cepstral coefficient features; The image driving layer performs driving processing on the feature map to be processed based on the initial driving features to obtain the initial feature map. The deformation feature map and the initial feature map are stacked through the feature synchronization layer to determine the feature synchronization parameters between the deformation feature map and the initial feature map; The initial driving features are adjusted by the feature processing layer based on the feature synchronization parameters to obtain the target driving features; The image driving layer performs driving processing on the image to be processed based on the target driving features to obtain the target dynamic image.

4. The method according to claim 1, characterized in that, The method further includes: Based on the reference image and the trained generative network model, the target head action features and target expression coefficient features are determined; the trained generative network model is used to generate multiple prediction images based on the input reference image, and based on the differences between each prediction image and the reference image, the target head action features and target expression coefficient features are determined. The trained generative network model is adjusted based on the target head action features and the target facial expression coefficient features to obtain the target generative network model.

5. The method according to claim 4, characterized in that, The determination of target head action features and target expression coefficient features based on the reference image and the trained generative network model includes: A video of a preset duration is generated based on the reference image to obtain reference data; The trained generative network model extracts the target head movement features and the target facial expression coefficient features from the reference data.

6. The method according to claim 5, characterized in that, The step of extracting the target head action features and the target facial expression coefficient features from the reference data using the trained generative network model includes: The trained generative network model determines multiple frames of facial images of the target person from the reference data. The trained generative network model extracts features from the facial images of the target person in each frame to obtain the target head movement features and the target expression coefficient features.

7. The method according to any one of claims 4-6, characterized in that, The method further includes: Obtain a sample video; wherein the video object in the sample video is different from the image object in the image to be processed; The sample video is processed by the generative network model to be trained to extract sample audio data and sample image data. The sample audio data and sample image data are processed based on the generative network model to be trained to obtain the prediction training result; The predicted training results are used as the initial training output of the generative network model to be trained, and the sample image data is used as supervision information. The generative network model to be trained is iteratively trained to obtain the trained generative network model.

8. The method according to claim 7, characterized in that, The process of the sample audio data and sample image data based on the generative network model to be trained to obtain the prediction training result includes: Based on the generative network model to be trained, reference Mel-frequency cepstral coefficient features are extracted from the sample audio data; Based on the generative network model to be trained, reference head movements, reference facial expression coefficients, and reference facial features are extracted from the sample image data. The predictive training result is obtained by using the generative network model to be trained based on the reference Mel-Cepstral coefficient features, the reference head motion, the reference facial expression coefficient features, and the reference face features.

9. The method according to claim 7, characterized in that, The step of using the predicted training result as the initial training output of the generative network model to be trained, and the sample image data as supervision information, to iteratively train the generative network model to be trained to obtain the trained generative network model includes: The loss value is determined based on the prediction training results and the sample image data; Based on the loss value, the generator network model to be trained is iteratively updated to obtain the trained generator network model.

10. An apparatus for generating dynamic images based on audio, characterized in that, include: The acquisition module is used to acquire the reference audio, reference image, and image to be processed input by the user; The image to be processed is the same as the image object in the reference image; The processing module is used to process the image to be processed through the affine sub-network in the target generation network model to obtain the feature map to be processed, and to obtain the deformation feature map based on the reference audio, the reference image and the feature map to be processed through the affine sub-network. The affine subnetwork is used to determine the target Mel-frequency cepstral coefficient features corresponding to the reference audio, extract features from the reference image to obtain a reference feature map, and perform an affine transformation on the reference feature map to obtain the deformation feature map; The adjustment module is used to process the image to be processed based on the deformation feature map through the driving sub-network in the target generation network model to obtain the target dynamic image; The driving sub-network is used to drive the image to be processed to obtain the target dynamic image.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method as described in any one of claims 1-9.

12. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1-9.

Citation Information

Patent Citations

  • Image recognition method and device, computer equipment and computer readable storage medium

    CN110210571A

  • Face image animation method and system based on action and voice features

    CN114445529A