Image generation method and related device

By extracting spatial features from original facial image frames and temporal features from audio-driven information, and combining them with an image generation model for facial reconstruction, the problem of time-consuming and labor-intensive generation of images of different facial postures in existing technologies is solved, and efficient and accurate facial image generation is achieved.

CN115131849BActive Publication Date: 2025-09-16TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210477320.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-04
Publication Date
2025-09-16
Estimated Expiration
2042-05-04

AI Technical Summary

Technical Problem

In the existing technology, generating facial images under different facial postures is time-consuming and labor-intensive, and the image generation efficiency is low.

Method used

By obtaining the original facial image frame and audio driving information of the target object, spatial feature extraction and temporal feature extraction are performed, and facial reconstruction processing is performed in combination with the image generation model to generate the target facial image frame.

Benefits of technology

The generation efficiency and accuracy of target facial image frames are improved, facial posture details are captured, and efficient facial image adjustment is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115131849B_ABST
    Figure CN115131849B_ABST
Patent Text Reader

Abstract

This application discloses an image generation method and related equipment. The relevant embodiments can be applied to various scenarios such as cloud technology, artificial intelligence, smart transportation, and assisted driving. The method can obtain the original facial image frame of the target object and the audio driving information of the target facial image frame to be generated; perform spatial feature extraction on the original facial image frame to obtain the original facial spatial features; perform temporal feature extraction on the audio driving information to obtain the facial local posture features; and perform facial reconstruction processing on the target object based on the original facial spatial features and the facial local posture features to generate the target facial image frame. This application can capture the facial posture details of the target object by extracting features from the audio driving information, and then perform facial adjustments on the original facial image frame based on the captured information, thereby obtaining the target facial image frame corresponding to the audio driving information. This is conducive to improving the generation efficiency and accuracy of the target facial image frame.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to an image generation method and related equipment. Background Art

[0002] With the development of computer technology, image processing technology has been applied to more and more fields. For example, image processing technology can include image generation, specifically facial image generation, which can be applied to fields such as animation production.

[0003] In the current related technologies, if you want to obtain facial images of the same target object in different facial postures, modelers and animators are required to draw the facial images in each facial posture separately. This image generation method is relatively time-consuming and labor-intensive, and the image generation efficiency is low. Summary of the Invention

[0004] Embodiments of the present application provide an image generation method and related equipment. The related equipment may include an image generation device, an electronic device, a computer-readable storage medium, and a computer program product, which can improve the generation efficiency and accuracy of target facial image frames.

[0005] The present invention provides an image generation method, including:

[0006] Obtaining audio driving information corresponding to an original facial image frame of a target object and a target facial image frame to be generated;

[0007] Performing spatial feature extraction on the original facial image frame to obtain original facial spatial features corresponding to the original facial image frame;

[0008] Performing temporal feature extraction on the audio driving information to obtain local facial posture features corresponding to the target facial image frame;

[0009] Based on the original facial spatial features and the local facial posture features, facial reconstruction processing is performed on the target object to generate the target facial image frame.

[0010] Accordingly, an embodiment of the present application provides an image generating device, including:

[0011] an acquisition unit, configured to acquire audio driving information corresponding to an original facial image frame of a target object and a target facial image frame to be generated;

[0012] a first extraction unit, configured to extract spatial features from the original facial image frame to obtain original facial spatial features corresponding to the original facial image frame;

[0013] A second extraction unit is configured to extract temporal features from the audio driving information to obtain local facial posture features corresponding to the target facial image frame;

[0014] The reconstruction unit is configured to perform facial reconstruction processing on the target object based on the original facial spatial features and the facial local posture features to generate the target facial image frame.

[0015] Optionally, in some embodiments of the present application, the second extraction unit may include an extraction subunit, a processing subunit, and a first fusion subunit, as follows:

[0016] The extraction subunit is used to extract features from each audio frame in the audio driving information to obtain audio semantic feature information of each audio frame;

[0017] a processing subunit, configured to process the audio semantic feature information of each audio frame based on the audio semantic feature information of the preceding and following audio frames of each audio frame;

[0018] The first fusion subunit is used to fuse the processed audio semantic feature information of each audio frame to obtain the facial local posture feature corresponding to the target facial image frame.

[0019] Optionally, in some embodiments of the present application, the reconstruction unit may include a second fusion subunit, a reconstruction subunit, and a generation subunit, as follows:

[0020] The second fusion subunit is configured to fuse the original facial spatial features with the facial local posture features to obtain fused facial spatial features;

[0021] A reconstruction subunit, configured to perform facial reconstruction processing on the target object based on the fused facial spatial features to obtain a reference facial image frame corresponding to the target object;

[0022] A generating subunit is configured to generate the target facial image frame based on the original facial image frame, the fused facial spatial features, and the reference facial image frame.

[0023] Optionally, in some embodiments of the present application, the reconstruction subunit can be specifically used to perform facial reconstruction processing on the target object based on the fused facial spatial features to obtain a reconstructed three-dimensional facial image corresponding to the target object; and perform rendering mapping processing on the reconstructed three-dimensional facial image to obtain a reference facial image frame corresponding to the target object.

[0024] Optionally, in some embodiments of the present application, the generating subunit can be specifically used to perform multi-scale feature extraction on the original facial image frame to obtain original facial feature maps at multiple scales corresponding to the original facial image frame; perform multi-scale feature extraction on the reference facial image frame to obtain reference facial feature maps at multiple scales corresponding to the reference facial image frame; perform encoding mapping processing on the fused facial spatial features to obtain latent feature information corresponding to the fused facial spatial features; and fuse the original facial feature maps at multiple scales, the reference facial feature maps at multiple scales, and the latent feature information to obtain the target facial image frame.

[0025] Optionally, in some embodiments of the present application, the step of “fusing the original facial feature maps at the multiple scales, the reference facial feature maps at the multiple scales, and the latent feature information to obtain the target facial image frame” may include:

[0026] fusing the latent feature information, the original facial feature map at a target scale, and the reference facial feature map at the target scale to obtain a fused facial feature map corresponding to the target scale, where the target scale is a scale selected from the multiple scales;

[0027] The fused facial feature map corresponding to the target scale, the original facial feature map at the adjacent scale, and the reference facial feature map at the adjacent scale are fused to obtain the target facial image frame.

[0028] Optionally, in some embodiments of the present application, the step of “fusing the corresponding fused facial feature map at the target scale, the original facial feature map at the adjacent scale, and the reference facial feature map at the adjacent scale to obtain the target facial image frame” may include:

[0029] Based on the latent feature information, performing style modulation processing on the fused facial feature map corresponding to the target scale to obtain a modulated style feature;

[0030] The modulated style features, the original facial feature map at adjacent scales, and the reference facial feature map at adjacent scales are fused to obtain the target facial image frame.

[0031] Optionally, in some embodiments of the present application, the first extraction unit may be specifically configured to perform spatial feature extraction on the original facial image frame using an image generation model to obtain original facial spatial features corresponding to the original facial image frame;

[0032] The second extraction unit may be specifically configured to perform temporal feature extraction on the audio driving information using the image generation model to obtain local facial posture features corresponding to the target facial image frame;

[0033] The reconstruction unit can be specifically configured to perform facial reconstruction processing on the target object based on the original facial spatial features and the facial local posture features through the image generation model to generate the target facial image frame.

[0034] Optionally, in some embodiments of the present application, the image generation device may further include a training unit, and the training unit may be used to train the image generation model;

[0035] The training unit can be specifically used to obtain training data, which includes original facial image frame samples of the sample object, target-driven facial image frame samples, and audio-driven information samples corresponding to the target-driven facial image frame samples; through a preset image generation model, spatial feature extraction is performed on the original facial image frame samples to obtain original facial spatial features corresponding to the original facial image frame samples; temporal feature extraction is performed on the audio-driven information samples to obtain facial local posture features corresponding to the target-driven facial image frame samples; based on the original facial spatial features and the facial local posture features, facial reconstruction processing is performed on the sample object to obtain a predicted-driven facial image frame; based on the target-driven facial image frame samples and the predicted-driven facial image frame, the parameters of the preset image generation model are adjusted to obtain a trained image generation model.

[0036] Optionally, in some embodiments of the present application, the step of “adjusting parameters of a preset image generation model based on the target driving facial image frame sample and the predicted driving facial image frame to obtain a trained image generation model” may include:

[0037] Performing spatial feature extraction on the target-driven facial image frame sample to obtain target facial spatial features corresponding to the target-driven facial image frame sample;

[0038] Determining first loss information based on the facial local posture features corresponding to the target driven facial image frame sample and the target facial spatial features;

[0039] determining second loss information based on the target driving facial image frame sample and the predicted driving facial image frame;

[0040] According to the first loss information and the second loss information, the parameters of the preset image generation model are adjusted to obtain a trained image generation model.

[0041] Optionally, in some embodiments of the present application, the step of “determining second loss information based on the target driving facial image frame sample and the predicted driving facial image frame” may include:

[0042] respectively predicting the probability that the target driving facial image frame sample and the predicted driving facial image frame belong to the real driving facial image frame, and determining adversarial loss information of a preset image generation model based on the probabilities;

[0043] Determining reconstruction loss information of a preset image generation model based on a similarity between the target driving facial image frame sample and the predicted driving facial image frame;

[0044] Performing identity recognition on the target driving facial image frame sample and the predicted driving facial image frame respectively, and determining identity loss information of a preset image generation model based on the identity recognition results;

[0045] Second loss information is determined according to the adversarial loss information, the reconstruction loss information, and the identity loss information.

[0046] An electronic device provided in an embodiment of the present application includes a processor and a memory, wherein the memory stores a plurality of instructions, and the processor loads the instructions to execute the steps in the image generation method provided in the embodiment of the present application.

[0047] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps in the image generation method provided in the embodiment of the present application are implemented.

[0048] In addition, an embodiment of the present application also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implements the steps in the image generation method provided in the embodiment of the present application.

[0049] The embodiments of the present application provide an image generation method and related equipment, which can obtain the original facial image frame of the target object and the audio driving information corresponding to the target facial image frame to be generated; perform spatial feature extraction on the original facial image frame to obtain the original facial spatial features corresponding to the original facial image frame; perform temporal feature extraction on the audio driving information to obtain the facial local posture features corresponding to the target facial image frame; based on the original facial spatial features and the facial local posture features, perform facial reconstruction processing on the target object to generate the target facial image frame. The present application can capture the facial posture details of part of the target object by performing feature extraction on the audio driving information, and then perform facial adjustments on the original facial image frame based on the captured information, thereby obtaining the target facial image frame corresponding to the audio driving information, which is conducive to improving the generation efficiency and accuracy of the target facial image frame. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0051] Figure 1a Schematic diagram of a scene of the image generation method provided in an embodiment of the present application;

[0052] Figure 1b is a flowchart of the image generation method provided in an embodiment of the present application;

[0053] Figure 1c is an illustration of the image generation method provided in an embodiment of the present application;

[0054] Figure 1d is another illustration of the image generation method provided in an embodiment of the present application;

[0055] Figure 1e This is a model structure diagram of the image generation method provided in the embodiment of the present application;

[0056] Figure 1f is another model structure diagram of the image generation method provided in an embodiment of the present application;

[0057] Figure 1g is another model structure diagram of the image generation method provided in an embodiment of the present application;

[0058] Figure 2 is another flow chart of the image generation method provided in an embodiment of the present application;

[0059] Figure 3is a structural diagram of an image generating device provided in an embodiment of the present application;

[0060] Figure 4 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0061] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.

[0062] The present invention provides an image generation method and related devices, which may include an image generation device, an electronic device, a computer-readable storage medium, and a computer program product. The image generation device may be integrated into an electronic device, which may be a terminal or a server.

[0063] It is understandable that the image generation method of this embodiment can be executed on a terminal, on a server, or jointly by a terminal and a server. The above examples should not be construed as limiting the present application.

[0064] like Figure 1a As shown, the image generation method is performed jointly by a terminal and a server as an example. The image generation system provided in the embodiment of the present application includes a terminal 10 and a server 11, etc. The terminal 10 and the server 11 are connected via a network, such as a wired or wireless network, wherein the image generation device can be integrated into the server.

[0065] Among them, the server 11 can be used to: obtain the original facial image frame of the target object and the audio driving information corresponding to the target facial image frame to be generated; perform spatial feature extraction on the original facial image frame to obtain the original facial spatial features corresponding to the original facial image frame; perform temporal feature extraction on the audio driving information to obtain the facial local posture features corresponding to the target facial image frame; based on the original facial spatial features and the facial local posture features, perform facial reconstruction processing on the target object to generate the target facial image frame, and send the target facial image frame to the terminal 10. Among them, the server 11 can be a single server, or a server cluster or cloud server composed of multiple servers. In the image generation method or device disclosed in this application, multiple servers can be composed of a blockchain, and the server is a node on the blockchain.

[0066] The terminal 10 can be configured to receive the target facial image frame sent by the server 11. The terminal 10 can include a mobile phone, a smart TV, a tablet computer, a laptop computer, a personal computer (PC), an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, or an aircraft. The terminal 10 can also be configured with a client, which can be an application client or a browser client.

[0067] The step of generating the target facial image frame by the server 11 may also be performed by the terminal 10 .

[0068] The image generation method provided in the embodiments of the present application relates to computer vision technology and speech technology in the field of artificial intelligence.

[0069] Artificial Intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive field of computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. AI software technologies primarily include computer vision, speech processing, natural language processing, machine learning / deep learning, autonomous driving, and smart transportation.

[0070] Computer vision (CV) is the science of making machines "see." Specifically, it refers to machine vision, which uses cameras and computers to replace the human eye to identify and measure targets, and then further processes images to make them more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems that can obtain information from images or multidimensional data. Computer vision technology typically includes image processing, image recognition, image semantic understanding, image retrieval, optical character recognition (OCR), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous positioning and mapping, autonomous driving, and smart transportation. It also includes common biometric recognition technologies such as facial recognition and fingerprint recognition.

[0071] Key speech technologies include automatic speech recognition, speech synthesis, and voiceprint recognition. They enable computers to hear, see, speak, and feel, and are the future direction of human-computer interaction. Speech technology is one of the most promising methods of human-computer interaction.

[0072] It should be noted that the order of description of the following embodiments is not intended to limit the preferred order of the embodiments.

[0073] This embodiment will be described from the perspective of an image generating device. The image generating device may be integrated into an electronic device, which may be a server, a terminal or other device.

[0074] It is understandable that in the specific implementation of this application, related data such as user information is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0075] This embodiment can be applied to various scenarios such as cloud technology, artificial intelligence, smart transportation, and assisted driving.

[0076] like Figure 1b As shown, the specific process of the image generation method can be as follows:

[0077] 101. Obtain audio driving information corresponding to an original facial image frame of a target object and a target facial image frame to be generated.

[0078] The target object may be an object whose facial posture is to be adjusted, and the original facial image frame may be an image containing the target object's face. The target facial image frame may specifically be a facial image corresponding to the facial posture adjusted in the original facial image frame based on the audio driving information. The facial posture herein may specifically refer to facial expression, such as mouth shape, eye expression, and other facial information of the object, which is not limited in this embodiment.

[0079] The audio driving information is audio information used to adjust the facial posture of the original facial image frame. Specifically, it can be used to replace the facial posture of the target subject in the original facial image frame with the facial posture corresponding to the target subject speaking, thereby obtaining a target facial image frame. The audio information corresponding to the target subject speaking serves as the audio driving information. The audio length corresponding to the audio driving information can be 1 second or 2 seconds, which is not limited in this embodiment.

[0080] In this embodiment, the target object's lip shape and other changes contained in the audio driving information can be used to determine the target object's facial posture changes; in addition, the target object's speech content and volume contained in the audio driving information can be used to judge the target object's emotional changes, and then determine the target object's facial posture changes. Therefore, the facial posture information of the target object when speaking can be obtained by extracting the audio semantic feature information from the audio driving information, and then a target facial image corresponding to the audio driving information can be generated.

[0081] In a specific scenario, this embodiment can obtain multiple segments of audio driving information of the target object, and for each segment of audio driving information, generate a target facial image frame corresponding to each segment of audio driving information based on each segment of audio driving information and the original facial image frame, and then splice the target facial image frames corresponding to each audio driving information to generate a target facial video clip corresponding to the target object. The target facial video clip includes the facial posture change process of the target object when speaking, and the audio information corresponding to the target object's speech is each segment of audio driving information. It should be noted that the protagonist in the target facial video clip is still the object face in the original facial image frame, and the expression (especially the mouth shape) of the target object in the generated target facial video clip corresponds to each segment of audio driving information.

[0082] In one embodiment, the present application can also be applied in the scenario of video repair. For example, if a speech video about the target object is damaged and some video frames in the speech video are lost, the image generation method provided by the present application can be used to utilize other video frames and corresponding audio information in the speech video to generate and repair the lost video frames. The audio information used for repair can specifically be an audio clip of 1 second before and after the lost video frame in the speech video. The audio clip is also the audio driving information in the above embodiment.

[0083] Specifically, if Figure 1c As shown in FIG, taking the target object as a human face as an example, a target facial image frame generated based on the original facial image frame and audio driving information is shown. The original facial image frame can be regarded as the original facial image to be driven.

[0084] 102. Perform spatial feature extraction on the original facial image frame to obtain original facial spatial features corresponding to the original facial image frame.

[0085] The original facial spatial features may specifically include 3D (3-dimensional) facial coefficients corresponding to the original facial image frame, such as identity information, lighting, texture, expression, pose, gaze, etc. The original facial image frame can be reconstructed based on these facial coefficients. Figure 1d shown.

[0086] The spatial feature extraction of the original facial image frame may specifically include performing convolution processing and pooling processing on the original facial image frame, which is not limited in this embodiment.

[0087] In this embodiment, spatial features of the original facial image frame can be extracted through an image feature extraction network, wherein the image feature extraction network can specifically be a neural network model, and the neural network can be a visual geometry group network (VGGNet, Visual Geometry Group Network), a residual network (ResNet, Residual Network) and a densely connected convolutional network (DenseNet, Dense Convolutional Network), etc. However, it should be understood that the neural network of this embodiment is not limited to the types listed above.

[0088] Specifically, the image feature extraction network is pre-trained, and the three-dimensional facial coefficients corresponding to the facial image can be predicted through the image feature extraction network.

[0089] In specific scenarios, ResNet50 or other network structures can be used to extract the original facial spatial features corresponding to the original facial image frame. The feature extraction process can be expressed by the following formula (1):

[0090]

[0091] Among them, coeff is the three-dimensional facial coefficient, that is, the original facial spatial feature corresponding to the original facial image frame, Represents the ResNet50 network, I face Represents the original facial image frame.

[0092] After extracting the initial raw facial spatial features, it is possible to filter out features associated with the facial posture of the target object from the raw facial spatial features. For example, identity information, light and shadow, texture, expression, posture, eye contact and other three-dimensional facial coefficients can be extracted from the raw facial spatial features and used as the final raw facial spatial features.

[0093] 103. Perform temporal feature extraction on the audio driving information to obtain local facial posture features corresponding to the target facial image frame.

[0094] Optionally, in this embodiment, the step of “extracting temporal features from the audio driving information to obtain local facial posture features corresponding to the target facial image frame” may include:

[0095] Performing feature extraction on each audio frame in the audio driving information to obtain audio semantic feature information of each audio frame;

[0096] Processing the audio semantic feature information of each audio frame based on the audio semantic feature information of the preceding and following audio frames of each audio frame;

[0097] The processed audio semantic feature information of each audio frame is fused to obtain the facial local posture feature corresponding to the target facial image frame.

[0098] Among them, the step of "extracting features from each audio frame in the audio driving information to obtain audio semantic feature information of each audio frame" may include: performing convolution operations and pooling operations on each audio frame in the audio driving information through a neural network to obtain audio semantic feature information of each audio frame.

[0099] The audio semantic feature information of each audio frame can be processed using a memory network model based on the audio semantic feature information of the preceding and following audio frames. The memory network model can be a long short-term memory (LSTM) network or a two-layer gated recurrent unit (GRU) network.

[0100] LSTMs, through their three-gate structure (input gate, forget gate, and output gate), selectively forget some historical data, add some current input data, and ultimately integrate it into the current state to generate the output state. LSTMs are well-suited for extracting semantic features from time series data and are often used to extract semantic features from contextual information in natural language processing tasks. GRUs, a type of recurrent neural network, like LSTMs, were designed to address issues such as long-term memory and gradients in backpropagation.

[0101] Optionally, in this embodiment, the step of “fusing the processed audio semantic feature information of each audio frame to obtain the local facial posture feature corresponding to the target facial image frame” may include:

[0102] A weighted transformation is performed on the audio semantic feature information of each processed audio frame to obtain a facial local posture feature corresponding to the target facial image frame.

[0103] It is understandable that other fusion methods can also be used to fuse the audio semantic feature information of each processed audio frame. For example, a splicing processing method can be used to splice the audio semantic feature information of each processed audio frame to obtain the local facial posture features corresponding to the target facial image frame.

[0104] It should be noted that in some embodiments, instead of using audio driving information, text driving information corresponding to the target facial image frame can be used to obtain the facial local posture features corresponding to the target facial image frame. In a specific scenario, for example, a speech video of the target object is damaged and some video frames in the speech video are lost. The subtitle information corresponding to the lost video frame can be used as text driving information, and the other non-lost video frames in the speech video can be used as the original facial image frames and the text driving information to generate and repair the lost video frame. Specifically, feature extraction can be performed on the text driving information to obtain the facial local posture features corresponding to the target facial image frame to be generated; spatial feature extraction can be performed on the original facial image frame to obtain the original facial spatial features corresponding to the original facial image frame, and facial reconstruction processing can be performed based on the original facial spatial features and the facial local posture features to generate the lost target facial image frame.

[0105] 104. Perform facial reconstruction processing on the target object based on the original facial spatial features and the facial local posture features to generate the target facial image frame.

[0106] Optionally, in this embodiment, the step of “performing facial reconstruction processing on the target object based on the original facial spatial features and the facial local posture features to generate the target facial image frame” may include:

[0107] Fusing the original facial spatial features with the facial local posture features to obtain fused facial spatial features;

[0108] Based on the fused facial spatial features, performing facial reconstruction processing on the target object to obtain a reference facial image frame corresponding to the target object;

[0109] The target facial image frame is generated based on the original facial image frame, the fused facial spatial features and the reference facial image frame.

[0110] Among them, the local facial posture feature contains part of the facial posture information in the target facial image frame to be generated. For example, the local facial posture feature can include relevant three-dimensional facial coefficients such as expression, pose, and gaze; while the original facial spatial feature contains the facial posture information of the original facial image frame.

[0111] There are many ways to fuse the original facial spatial features and the facial local posture features. For example, the fusion method can be splicing processing or weighted fusion, etc., which is not limited in this embodiment.

[0112] Optionally, in this embodiment, the step of “performing facial reconstruction processing on the target object based on the fused facial spatial features to obtain a reference facial image frame corresponding to the target object” may include:

[0113] Based on the fused facial spatial features, performing facial reconstruction processing on the target object to obtain a reconstructed three-dimensional facial image corresponding to the target object;

[0114] The reconstructed three-dimensional facial image is subjected to rendering mapping processing to obtain a reference facial image frame corresponding to the target object.

[0115] Specifically, a facial reconstruction process can be performed on the target object based on the fused facial spatial features using a 3DMM (3D Morphable Model) model to obtain a reconstructed three-dimensional facial image. This facial reconstruction process is also known as 3D reconstruction. 3D reconstruction can represent the input two-dimensional facial image with a 3D mesh (three-dimensional mesh model). The 3D mesh can contain vertex coordinates and colors of a three-dimensional mesh structure. This embodiment can also project the 3D mesh (i.e., the reconstructed three-dimensional facial image) onto a two-dimensional plane by rendering, such as Figure 1d shown.

[0116] Among them, the texture and light of the reconstructed three-dimensional facial image can come from the original facial image frame, and the posture and expression of the reconstructed three-dimensional facial image can be driven by audio information. The reconstructed three-dimensional facial image is rendered and mapped, and the three-dimensional image can be projected onto a two-dimensional plane to obtain a reference facial image frame corresponding to the target object.

[0117] Specifically, the fused facial spatial features may include geometric features and texture features of the target object, and a reconstructed three-dimensional facial image may be constructed based on the geometric features and texture features. The geometric features may be understood as the coordinate information of the key points of the 3D mesh structure of the target object, and the texture features may be understood as the features indicating the texture information of the target object. There are many ways to perform facial reconstruction based on the fused facial spatial features and obtain the reconstructed three-dimensional facial image corresponding to the target object. For example, the position information of at least one facial key point may be extracted from the fused facial spatial features, the position information of the facial key point may be converted into geometric features, and the texture features of the target object may be extracted from the fused facial spatial features. Specifically, the method may be as shown in formula (2):

[0118]

[0119] Among them, Coeff cat represents the facial spatial features after fusion, represents the 3DMM model, S represents the geometric features, and T represents the texture features.

[0120] After converting the geometric features and texture features, a three-dimensional object model of the target object can be constructed, that is, the reconstructed three-dimensional facial image in the above embodiment, and the three-dimensional object model can be projected onto a two-dimensional plane to obtain a reference facial image frame. There are many ways to obtain a reference facial image frame. For example, the three-dimensional model parameters of the target object can be determined based on the geometric features and texture features, and the three-dimensional object model of the target object can be constructed based on the three-dimensional model parameters. The three-dimensional object model can then be projected onto a two-dimensional plane to obtain a reference facial image frame, as shown in formula (3):

[0121]

[0122] Among them, I rendered is the reference facial image frame, S represents the geometric features, T represents the texture features, Represents a function that transforms a 3D model into a 2D image.

[0123] Optionally, in this embodiment, the step of “generating the target facial image frame based on the original facial image frame, the fused facial spatial features, and the reference facial image frame” may include:

[0124] Performing multi-scale feature extraction on the original facial image frame to obtain original facial feature maps at multiple scales corresponding to the original facial image frame;

[0125] Performing multi-scale feature extraction on the reference facial image frame to obtain reference facial feature maps at multiple scales corresponding to the reference facial image frame;

[0126] Performing encoding mapping processing on the fused facial spatial features to obtain latent feature information corresponding to the fused facial spatial features;

[0127] The original facial feature maps at the multiple scales, the reference facial feature maps at the multiple scales, and the latent feature information are fused to obtain the target facial image frame.

[0128] Multi-scale feature extraction can be used to obtain spatial features of the original facial image frame or the reference facial image frame at each preset resolution. The image scales of the original facial feature maps corresponding to different resolutions are different. Similarly, the image scales of the reference facial feature maps corresponding to different resolutions are different. In this embodiment, multi-scale extraction can be used to obtain original spatial features of the original facial image frame and reference spatial features of the reference facial image frame. The original spatial features include original facial feature maps at multiple scales, and the reference spatial features include reference facial feature maps at multiple scales. Therefore, the original spatial features of the original facial image frame and the reference spatial features of the reference facial image frame are both multi-layer spatial features.

[0129] There are multiple ways to extract original facial feature maps at multiple scales of the original facial image frame and reference facial feature maps at multiple scales of the reference facial image frame, as follows:

[0130] For example, the encoding network (Enc Block) of the trained image generation model can be used to spatially encode the original facial image frame and the reference facial image frame at each preset resolution, thereby obtaining the original facial feature map and the reference facial feature map at each resolution.

[0131] Among them, the encoding network (Enc Block) can include multiple sub-encoding networks, each sub-encoding network corresponds to a preset resolution. The sub-encodings can be arranged in order from small to large according to the size of the resolution, thereby obtaining an encoding network. When the original facial image frame and the reference facial image frame are input into the encoding network for network encoding, each sub-encoding network can output a spatial feature corresponding to the preset resolution. The encoding network for the original facial image frame and the reference facial image frame can be the same or different encoding networks, but different encoding networks share network parameters. The structure of the encoding sub-network can be various, for example, it can be composed of a simple one-layer convolutional network, or it can be other encoding network structures. The preset resolution can be set according to the actual application, for example, the resolution can be from 4*4 to 512*512.

[0132] Among them, the latent feature information is specifically the intermediate feature w obtained by encoding and mapping the fused facial spatial features. Different elements of the intermediate feature w control different visual features, thereby reducing the correlation between features (decoupling and feature separation). The encoding and mapping process can be used to extract the deep-level relationships hidden under the surface features from the fused facial spatial features, decouple these relationships, and thus obtain the latent features (latent code). There are many ways to use the trained image generation model to map the fused facial spatial features to latent feature information. For example, the mapping network of the trained image generation model can be used ( ) maps the fused facial spatial features into latent feature information (w).

[0133] Optionally, in this embodiment, the step of “fusing the original facial feature maps at the multiple scales, the reference facial feature maps at the multiple scales, and the latent feature information to obtain the target facial image frame” may include:

[0134] fusing the latent feature information, the original facial feature map at a target scale, and the reference facial feature map at the target scale to obtain a fused facial feature map corresponding to the target scale, where the target scale is a scale selected from the multiple scales;

[0135] The fused facial feature map corresponding to the target scale, the original facial feature map at the adjacent scale, and the reference facial feature map at the adjacent scale are fused to obtain the target facial image frame.

[0136] Optionally, in this embodiment, the step of “fusing the corresponding fused facial feature map at the target scale, the original facial feature map at the adjacent scale, and the reference facial feature map at the adjacent scale to obtain the target facial image frame” may include:

[0137] Based on the latent feature information, performing style modulation processing on the fused facial feature map corresponding to the target scale to obtain a modulated style feature;

[0138] The modulated style features, the original facial feature map at adjacent scales, and the reference facial feature map at adjacent scales are fused to obtain the target facial image frame.

[0139] The adjacent scale can be a scale larger than the target scale among the multiple scales. Specifically, if the multiple scales include 4*4, 8*8, 16*16, 32*32, and 64*64, if the target scale is 16*16, the adjacent scale can be 32*32; if the target scale is 4*4, the adjacent scale can be 8*8.

[0140] Specifically, this embodiment can adjust the preset basic style features based on the implicit feature information to obtain the modulated style features. The preset basic style features can be understood as the style features in the constant tensor (Const) pre-set during the image driving process. The so-called style features can be understood as the feature information used to generate an image of a specific style.

[0141] There are many ways to perform style modulation processing, such as adjusting the size of the basic style features to obtain the initial style features, modulating the latent feature information to obtain the convolution weights corresponding to the initial style features, and adjusting the initial style features based on the convolution weights to obtain the modulated style features.

[0142] The convolution weights can be understood as the weight information used when convolving the initial style features. There are many ways to modulate the latent feature information. For example, the base convolution weights can be obtained and adjusted based on the latent feature information to obtain the convolution features corresponding to the initial style features. Adjusting the convolution weights based on latent feature information can be achieved primarily using the Mod and Demod modules in the decoding network of StyleGAN v2 (a style transfer model).

[0143] After modulating the latent feature information, the initial style features can be adjusted based on the convolution weights obtained after the modulation process. This adjustment can be done in a variety of ways. For example, the target style convolution network corresponding to the resolution of the base facial image can be selected from the style convolution network (StyleConv) of the trained image generation model. Based on the convolution weights, the initial style features are adjusted to obtain the modulated style features. The base facial image at the initial resolution is generated based on the preset base style features.

[0144] The step of “fusing the modulated style features, the original facial feature map at adjacent scales, and the reference facial feature map at adjacent scales to obtain the target facial image frame” may include:

[0145] fusing the modulated style feature, the original facial feature map at an adjacent scale, and the reference facial feature map at the adjacent scale to obtain a fused facial feature map at the adjacent scale;

[0146] The facial feature maps at adjacent scales are fused with the basic facial image to generate a target facial image frame.

[0147] Among them, the fused facial feature map can also be regarded as a fused style feature.

[0148] The step of "generating a target facial image frame using the fused facial feature maps at adjacent scales and the base facial image" may include: using the fused facial feature maps at adjacent scales as the fused facial feature maps at a new target scale, returning to execute the step of performing style modulation processing on the corresponding fused facial feature maps at the target scale based on the latent feature information to obtain modulated style features, until the scale of the obtained target facial image frame meets a preset scale condition.

[0149] The preset scale condition may specifically enable the scale of the target facial image frame to be the largest scale among the multiple scales.

[0150] Specifically, in this embodiment, there are many ways to filter out the target original spatial features (i.e., the original facial feature map at the target scale) from the original spatial features based on the preset resolution, and to filter out the target reference spatial features (i.e., the reference facial feature map at the target scale) from the reference spatial features. For example, the original spatial features and the reference spatial features can be sorted based on the preset resolution, and based on the sorting information, the original spatial features with the smallest resolution can be filtered out from the original spatial features as the target original spatial features, and the original spatial features with the smallest resolution can be filtered out from the reference spatial features as the target original spatial features. After filtering out the target original spatial features and the target reference spatial features, the target original spatial features can be deleted from the original spatial features and deleted from the reference spatial features. In this way, the spatial features with the smallest resolution can be filtered out from the original spatial features and the reference spatial features each time, thereby obtaining the target original spatial features and the target reference spatial features.

[0151] After the target original space features and the target reference space features are screened out, the modulated style features, the target original space features and the target reference space features can be fused. There are many ways to fuse them. For example, the target original space features, the target reference space features and the modulated style features can be directly spliced ​​to obtain the fused style features at the current resolution. Specifically, it can be shown as formula (4):

[0152]

[0153] in, is the fused style feature, which can be the style feature corresponding to the next preset resolution of the basic style feature. As the basic style features, is the original spatial feature of the target at the preset resolution, is the target reference space feature at the preset resolution, Concat means concatenating or splicing the features, and StyleConv is the style convolutional network.

[0154] After obtaining the fused style features at the current resolution, the target facial image frame at the target resolution can be generated based on the fused style features and the basic facial image. There are many ways to generate the target facial image frame at the target resolution. For example, the current facial image can be generated based on the fused style features, and the current facial image and the basic facial image can be fused to obtain a fused facial image at the current resolution. The fused style features can be used as the preset basic style features, and the fused facial image can be used as the basic facial image. The step of adjusting the preset basic style features based on the latent feature information is returned to execute until the current resolution is the target resolution, and the target facial image frame is obtained.

[0155] Among them, in the process of generating the target facial image frame of the target resolution, it can be found that the current facial image at different resolutions is superimposed with the basic facial image in sequence, and in the superposition process, the resolution is increased in sequence, so that a high-definition target facial image frame can be output.

[0156] Optionally, a basic optical flow field at the initial resolution can be generated based on the basic style features, thereby outputting a target optical flow field at the target resolution. The basic optical flow field can be understood as a visual field indicating the movement of key points in the facial image at the initial resolution. There are multiple ways to output the target optical flow field. For example, a basic optical flow field at the initial style resolution can be generated based on the basic style features. Based on the basic optical flow field, the modulated style features, the original spatial features, and the reference spatial features are fused to obtain the target optical flow field at the target resolution.

[0157] Among them, there are many ways to fuse the modulated style features, original space features, and reference space features. For example, based on the preset resolution, the target original space features can be screened out from the original space features, and the target reference space features can be screened out from the reference space features. The modulated style features and the target style features can be fused to obtain the fused style features at the current resolution. Based on the fused style features and the basic optical flow field, the target optical flow field at the target resolution is generated.

[0158] Among them, there are many ways to generate a target optical flow field at the target resolution based on the fused style features and the basic optical flow field. For example, the current optical flow field can be generated based on the fused style features, and the current optical flow field and the basic optical flow field can be fused to obtain the fused optical flow field at the current resolution. The fused style features are used as the preset basic style features, and the fused optical flow field is used as the basic optical flow field. The step of adjusting the preset basic style features based on the latent feature information is returned to execute until the current resolution is the target resolution to obtain the target optical flow field.

[0159] Among them, the target facial image and the target optical flow field can be generated simultaneously based on the preset basic style features, or the target facial image or the target optical flow field can be generated separately based on the preset style features. Taking the simultaneous generation of the target object and the target optical flow field as an example, the preset basic style features, the basic facial image and the basic optical flow field can be processed by the decoding network of the trained image generation model. Each decoding network can include a decoding sub-network corresponding to each preset resolution. The resolution can be increased from 4*4 to 512*512. The network structure of the decoding sub-network can be as follows: Figure 1e As shown, the fused style features output by the previous decoding sub-network are received ( ), fusion facial image (I i ) and fused optical flow field (f i ), the latent feature information w is modulated to obtain The corresponding convolution weight, and based on the convolution weight, Perform convolution processing to obtain the modulated style features ( ), based on the resolution corresponding to the decoding sub-network, the target original spatial features corresponding to the decoding sub-network are selected from the original spatial features ( ), and select the target reference space features corresponding to the decoding sub-network from the reference space features ( ), The spatial resolution is the same as and By connecting them in series, we can obtain the fused style features output by the decoding sub-network ( ), then, based on Generate the current facial image and the current optical flow field, and combine the current facial image and the I output of the previous decoding sub-network i By fusing, we can get the fused facial image (I i+1 ), the fused optical flow field (f i ) are fused to obtain the fused optical flow field (f i+1 ), and then output it to the decoding sub-network of the next layer until the decoding sub-network corresponding to the target resolution outputs the fused facial image and the fused optical flow field. The fused facial image at the target resolution (specifically, resolution 512) can be used as the target facial image frame, and the fused optical flow field at the target resolution can be used as the target optical flow field.

[0160] Optionally, in this embodiment, the step of “extracting spatial features from the original facial image frame to obtain original facial spatial features corresponding to the original facial image frame” may include:

[0161] Extracting spatial features of the original facial image frame using an image generation model to obtain original facial spatial features corresponding to the original facial image frame;

[0162] The step of “extracting temporal features from the audio driving information to obtain local facial posture features corresponding to the target facial image frame” may include:

[0163] Performing temporal feature extraction on the audio driving information through the image generation model to obtain local facial posture features corresponding to the target facial image frame;

[0164] The step of “performing facial reconstruction processing on the target object based on the original facial spatial features and the facial local posture features to generate the target facial image frame” may include:

[0165] The target object is subjected to facial reconstruction processing by the image generation model based on the original facial spatial features and the facial local posture features to generate the target facial image frame.

[0166] Among them, the image generation model can be a visual geometry group network (VGGNet, Visual Geometry Group Network), a residual network (ResNet, Residual Network) and a densely connected convolutional network (DenseNet, Dense Convolutional Network), etc., but it should be understood that the image generation model of this embodiment is not limited to the types listed above.

[0167] It should be noted that the image generation model can be trained by multiple sets of training data. The image generation model can be trained by other equipment and then provided to the image generation device, or it can be trained by the image generation device itself.

[0168] If the image generation device performs training by itself, the following steps may be further included before the step of "extracting spatial features from the original facial image frame using the image generation model to obtain original facial spatial features corresponding to the original facial image frame":

[0169] Acquire training data, the training data including original facial image frame samples of a sample object, target-driven facial image frame samples, and audio driving information samples corresponding to the target-driven facial image frame samples;

[0170] Extracting spatial features of the original facial image frame samples using a preset image generation model to obtain original facial spatial features corresponding to the original facial image frame samples;

[0171] Performing temporal feature extraction on the audio driving information sample to obtain facial local posture features corresponding to the target driving facial image frame sample;

[0172] Based on the original facial spatial features and the local facial posture features, performing facial reconstruction processing on the sample object to obtain a predicted driving facial image frame;

[0173] Based on the target driving facial image frame sample and the predicted driving facial image frame, the parameters of the preset image generation model are adjusted to obtain a trained image generation model.

[0174] The target driven facial image frame sample may be regarded as label information, and specifically may be the expected driven facial image frame corresponding to the audio driven information sample.

[0175] There are many ways to obtain the original facial image frame samples of the sample object, the target-driven facial image frame samples, and the audio driving information samples corresponding to the target-driven facial image frame samples, which are not limited in this embodiment.

[0176] For example, any two video frames containing the object's face can be extracted from a speech video about the sample object, one of which is used as the original facial image frame sample, the remaining frame is used as the target-driven facial image frame sample, and the audio information corresponding to the target-driven facial image frame sample 1 second before and after in the speech video is used as the audio-driven information sample.

[0177] Optionally, in this embodiment, the step of “adjusting parameters of a preset image generation model based on the target driving facial image frame sample and the predicted driving facial image frame to obtain a trained image generation model” may include:

[0178] Performing spatial feature extraction on the target-driven facial image frame sample to obtain target facial spatial features corresponding to the target-driven facial image frame sample;

[0179] Determining first loss information based on the facial local posture features corresponding to the target driven facial image frame sample and the target facial spatial features;

[0180] determining second loss information based on the target driving facial image frame sample and the predicted driving facial image frame;

[0181] According to the first loss information and the second loss information, the parameters of the preset image generation model are adjusted to obtain a trained image generation model.

[0182] The spatial feature extraction of the target-driven facial image frame samples may include performing convolution processing and pooling processing on the target-driven facial image frame samples, which is not limited in this embodiment. In this embodiment, the spatial feature extraction of the target-driven facial image frame samples may be performed using a trained image feature extraction network.

[0183] The extracted target facial spatial features corresponding to the target-driven facial image frame samples may specifically include three-dimensional (3D) facial coefficients corresponding to the target-driven facial image frame samples, such as identity information, lighting, texture, expression, pose, gaze, etc.

[0184] In this embodiment, the target facial spatial features corresponding to the target-driven facial image frame samples can be used as supervisory signals for the local facial posture features extracted from the audio-driven information samples. Specifically, this embodiment can calculate the vector distance between the local facial posture features corresponding to the target-driven facial image frame samples and the target facial spatial features, and determine first loss information based on this vector distance. A greater vector distance indicates a greater loss value corresponding to the first loss information, and conversely, a smaller vector distance indicates a smaller loss value corresponding to the first loss information.

[0185] The step of “adjusting parameters of a preset image generation model according to the first loss information and the second loss information to obtain a trained image generation model” may include:

[0186] fusing the first loss information and the second loss information to obtain total loss information;

[0187] According to the total loss information, the parameters of the preset image generation model are adjusted to obtain a trained image generation model.

[0188] There are many ways to fuse the first loss information and the second loss information, which are not limited in this embodiment. For example, the fusion method may be weighted fusion.

[0189] The training process of the preset image generation model first calculates the total loss information. Then, the backpropagation algorithm is used to adjust the parameters of the preset image generation model. The parameters of the image generation model are optimized based on the total loss information, so that the loss value corresponding to the total loss information is less than the preset loss value, thereby obtaining a trained image generation model. The preset loss value can be set according to actual conditions. For example, if the accuracy of the image generation model is required to be higher, the preset loss value should be smaller.

[0190] Optionally, in this embodiment, the step of “determining second loss information based on the target driving facial image frame sample and the predicted driving facial image frame” may include:

[0191] respectively predicting the probability that the target driving facial image frame sample and the predicted driving facial image frame belong to the real driving facial image frame, and determining adversarial loss information of a preset image generation model based on the probabilities;

[0192] Determining reconstruction loss information of a preset image generation model based on a similarity between the target driving facial image frame sample and the predicted driving facial image frame;

[0193] Performing identity recognition on the target driving facial image frame sample and the predicted driving facial image frame respectively, and determining identity loss information of a preset image generation model based on the identity recognition results;

[0194] Second loss information is determined according to the adversarial loss information, the reconstruction loss information, and the identity loss information.

[0195] In this embodiment, the step of “respectively predicting the probabilities that the target driving facial image frame sample and the predicted driving facial image frame belong to the real driving facial image frame, and determining the adversarial loss information of the preset image generation model based on the probabilities” may include:

[0196] Predicting first probability information that the target driving facial image frame sample belongs to a real driving facial image frame through a preset discriminant model;

[0197] Predicting, by means of the preset discriminant model, second probability information that the predicted driving facial image frame belongs to the real driving facial image frame;

[0198] Based on the first probability information and the second probability information, adversarial loss information of a preset image generation model is determined.

[0199] The preset discriminant model is the discriminator D. During training, the target driving facial image frame samples are real images, and the predicted driving facial image frames are the results generated by the preset image generation model. The discriminator needs to judge the generated results as false and the real images as true. The preset image generation model can be regarded as the entire driving network G. During training and learning, the images generated by the driving network G need to be able to deceive the discriminator D. In other words, the probability that the discriminator D judges the predicted driving facial image frames generated by the driving network G as real driving facial image frames is 1 as much as possible.

[0200] The discriminator takes as input a real image or the output of a generative model. Its goal is to distinguish the generative model output from real images as closely as possible. The generative model, on the other hand, aims to deceive the discriminator as much as possible. The generative model and the discriminator compete with each other, continuously adjusting parameters to produce a trained generative model.

[0201] In this embodiment, the step of “determining reconstruction loss information of a preset image generation model based on the similarity between the target driving facial image frame sample and the predicted driving facial image frame” may include:

[0202] Performing feature extraction on the target-driven facial image frame sample to obtain first feature information corresponding to the target-driven facial image frame sample;

[0203] performing feature extraction on the predicted driven facial image frame to obtain second feature information corresponding to the predicted driven facial image frame;

[0204] Reconstruction loss information of a preset image generation model is determined according to the similarity between the first feature information and the second feature information.

[0205] Among them, the vector distance between the feature vector corresponding to the first feature information and the feature vector corresponding to the second feature information can be calculated, and the similarity between the first feature information and the second feature information can be determined based on the vector distance. The larger the vector distance, the lower the similarity, and the greater the loss value corresponding to the reconstruction loss information; conversely, the smaller the vector distance, the higher the similarity, and the smaller the loss value corresponding to the reconstruction loss information.

[0206] In this embodiment, the step of “performing identity recognition on the target driving facial image frame sample and the predicted driving facial image frame respectively, and determining identity loss information of the preset image generation model based on the identity recognition results” may include:

[0207] Performing identity recognition on the target-driven facial image frame sample to obtain a first identity recognition result;

[0208] Performing identity recognition on the predicted driven facial image frame to obtain a second identity recognition result;

[0209] The first identity recognition result and the second identity recognition result are compared to obtain identity loss information of a preset image generation model.

[0210] If the first identity recognition result and the second identity recognition result are the same, the identity loss information of the preset image generation model is 0.

[0211] The step of “determining second loss information based on the adversarial loss information, the reconstruction loss information, and the identity loss information” may include:

[0212] The adversarial loss information, the reconstruction loss information, and the identity loss information are fused to obtain second loss information.

[0213] There are many ways to fuse the adversarial loss information, the reconstruction loss information, and the identity loss information, which are not limited in this embodiment. For example, the fusion method may be weighted fusion.

[0214] Specifically, during the training process of the preset image generation model G, the original facial image frame sample of the sample object can be recorded as I source , the target driven facial image frame sample can be recorded as I drive , the audio driving information sample corresponding to the target driving facial image frame sample can be recorded as V drive , and the predicted driving facial image frame generated by the preset image generation model can be recorded as G(I source , V drive ), the preset discriminant model is recorded as D, then the adversarial loss information in the above embodiment can be expressed as the following formula (5):

[0215]

[0216] Among them, D(I drive ) represents the first probability information of the discriminator predicting that the target driving facial image frame sample belongs to the real driving facial image frame in the above embodiment, and D(G(I source , V drive )) represents the second probability information of the discriminator predicting that the predicted driving facial image frame belongs to the real driving facial image frame in the above embodiment, L GAN Represents the adversarial loss information of the preset image generation model.

[0217] The reconstruction loss information can be expressed by formula (6):

[0218]

[0219] Among them, the reconstruction loss information L rec The L1 loss function and LPIPS (perceptual loss) loss function are used in the reconstruction of loss information to make the image generated by the trained image generation model consistent with the real driving image (i.e., the target driving facial image frame sample).

[0220] The identity loss information can be expressed as follows:

[0221]

[0222] Among them, since the identity information of the target facial image frame generated in the trained image generation model should be consistent with the identity information of the driving image (that is, the target driving facial image frame corresponding to the audio driving information), the identity loss information L is set during the training process. ID .

[0223] in, represents the identity feature extraction network.

[0224] The first loss information can be expressed by the following formula (8):

[0225]

[0226] Among them, the first loss information adopts the L2 loss function. The local facial posture features predicted by the trained image generation model through the audio driving information should be the same as the target facial spatial features of the driving image (that is, the target driving facial image frame corresponding to the audio driving information). Therefore, the first loss information is set during the training process.

[0227] Among them, exp represents expression features, pose represents posture features, and gaze represents eye features.

[0228] In some embodiments, the total loss information L may be expressed as follows:

[0229] L=L GAN +L rec +L 3d +L ID (9)

[0230] In a specific scenario, the overall training process of the preset image generation model can be as follows: Figure 1f As shown, specifically, any two video frames containing the face of the subject can be extracted from a speech video about the sample subject, and one of the frames is used as the original facial image frame sample (I s ), the remaining frame is used as the target driving facial image frame sample (I d ), and the audio information corresponding to the target driving facial image frame sample before and after 1 second in the speech video is used as the audio driving information sample, wherein it should be noted that the target driving facial image frame sample (I d ) can be used as GroudTruth (GT). In machine learning, GT can represent the classification accuracy of the training set of supervised learning.

[0231] Among them, in the specific training process of the preset image generation model, the image feature extraction network in the preset image generation model can be used to perform spatial feature extraction on the original facial image frame sample to obtain the original facial spatial feature corresponding to the original facial image frame sample, and the original facial spatial feature can include facial three-dimensional coefficients such as identity, light and shadow, and texture; then the audio semantic extraction network in the preset image generation model can be used to perform temporal feature extraction on the audio-driven information sample to obtain the facial local posture feature corresponding to the target-driven facial image frame sample, and the facial local posture feature can include partial facial posture information such as posture, eyes, and expression, and then the original facial spatial feature and the facial local posture feature can be fused to obtain the fused facial spatial feature, and then the reconstruction network can be used to perform facial reconstruction processing on the sample object based on the fused facial spatial feature to obtain the reference facial image frame of the sample object.

[0232] After obtaining the reference facial image frame, the original facial image frame sample, the reference facial image frame and the fused facial spatial features can be fused through the generation network in the preset image generation model to obtain the predicted driving facial image frame. Specifically, the generation network is mainly based on Stylegan v2 and includes the encoding network (Enc Block), the mapping network ( ) and decoding network. The encoding network can perform multi-scale feature extraction on the original facial image frame samples and the reference facial image frame respectively, and obtain multi-layer spatial features of the original facial image frame samples and the reference facial image frame. The lowest resolution of the encoding network output is 4*4 and the highest resolution is 512*512, which can be matched one-to-one with the output features of the decoding module. The encoding module can be composed of a simple layer of convolutional network. The mapping network ( ) The fused facial spatial features can be mapped to latent features w in the latent space. The spatial features corresponding to each preset resolution obtained by spatial encoding are decoded through a decoding network. During the decoding process, basic style features, basic facial images, and basic optical flow fields need to be generated based on a constant tensor (const). The basic facial image and the current facial image are superimposed according to the resolution from small to large, thereby obtaining a predicted driving facial image frame at the target resolution. Based on the target driving facial image frame samples and the predicted driving facial image frame, the total loss information is calculated, and the parameters of the preset image generation model are adjusted based on the total loss information to obtain a trained image generation model.

[0233] When training a preset image generation model, the generation network (mapping network and decoding network) and discriminator in the preset image generation model can be pre-trained, while the encoding network needs to be trained from scratch. Therefore, during training, the learning rates of the three are different, and the ratio of the learning rates can be set according to actual applications. For example, the learning rate ratio of the encoding network, generation network and discriminator can be 100:10:1 or other ratios.

[0234] Among them, the trained image generation model can output the target facial image frame and the target optical flow field at the same time, or can output the target facial image frame or the target optical flow field separately.

[0235] Taking the trained image generation model outputting the target facial image frame as an example, the process of outputting the target facial image frame at the target resolution through the trained image generation model can be as follows: Figure 1g Specifically, the image feature extraction network in the image generation model can be used to perform spatial feature extraction on the original facial image frame to obtain the original facial spatial features corresponding to the original facial image frame. The original facial spatial features can include facial three-dimensional coefficients such as identity, light and shadow, and texture; then the audio semantic extraction network in the image generation model can be used to perform temporal feature extraction on the audio driving information to obtain the facial local posture features corresponding to the target facial image frame. The facial local posture features can include partial facial posture information such as posture, eye contact, and expression. Then, the original facial spatial features and the facial local posture features can be fused to obtain the fused facial spatial features. Then, the reconstruction network can be used to perform facial reconstruction processing on the target object based on the fused facial spatial features to obtain the reference facial image frame of the sample object.

[0236] After obtaining the reference facial image frame, the original facial image frame, the reference facial image frame and the fused facial spatial features can be fused through the generative network in the image generation model to obtain the target facial image frame. Specifically, the generative network is mainly based on Stylegan v2 and includes the encoding network (Enc Block), the mapping network ( ) and decoding network. The encoding network can extract multi-scale features of the original facial image frame and the reference facial image frame respectively, and obtain multi-layer spatial features of the original facial image frame and the reference facial image frame. The lowest resolution of the encoding network output is 4*4 and the highest resolution is 512*512, which can be matched one-to-one with the output features of the decoding module. The mapping network ( ) The fused facial spatial features can be mapped to latent features w in the latent space, and the spatial features corresponding to each preset resolution encoded in the space are decoded through the decoding network. During the decoding process, basic style features, basic facial images and basic optical flow fields need to be generated according to the constant tensor (const), and the basic facial image and the current facial image are superimposed according to the resolution from small to large, so as to obtain the target facial image frame at the target resolution. The texture of the target facial image frame is consistent with the texture of the original facial image frame, and the posture and mouth shape of the target facial image frame conform to the audio driving information.

[0237] As can be seen from the above, this embodiment can obtain the original facial image frame of the target object and the audio driving information corresponding to the target facial image frame to be generated; perform spatial feature extraction on the original facial image frame to obtain the original facial spatial features corresponding to the original facial image frame; perform temporal feature extraction on the audio driving information to obtain the facial local posture features corresponding to the target facial image frame; based on the original facial spatial features and the facial local posture features, perform facial reconstruction processing on the target object to generate the target facial image frame. The present application can capture the facial posture details of part of the target object by performing feature extraction on the audio driving information, and then perform facial adjustments on the original facial image frame based on the captured information, thereby obtaining the target facial image frame corresponding to the audio driving information, which is conducive to improving the generation efficiency and accuracy of the target facial image frame.

[0238] According to the method described in the previous embodiment, the image generating device will be further described in detail below by taking the example of being specifically integrated into a server.

[0239] The present application embodiment provides an image generation method, such as Figure 2 As shown, the specific process of the image generation method can be as follows:

[0240] 201. The server obtains audio driving information corresponding to an original facial image frame of a target object and a target facial image frame to be generated.

[0241] The target object may be an object whose facial posture is to be adjusted, and the original facial image frame may be an image containing the target object's face. The target facial image frame may specifically be a facial image corresponding to the facial posture adjusted in the original facial image frame based on the audio driving information. The facial posture herein may specifically refer to facial expression, such as mouth shape, eye expression, and other facial information of the object, which is not limited in this embodiment.

[0242] The audio driving information is audio information used to adjust the facial posture of the original facial image frame. Specifically, it can be used to replace the facial posture of the target subject in the original facial image frame with the facial posture corresponding to the target subject's speech, thereby obtaining a target facial image frame. The audio information corresponding to the target subject's speech is the audio driving information. In this embodiment, the target subject's facial posture changes can be determined by using information such as changes in lip shape when the target subject speaks, which is included in the audio driving information. In addition, the target subject's speech content and volume, which are included in the audio driving information, can be used to judge the target subject's emotional changes and further determine the target subject's facial posture changes. Therefore, the facial posture information of the target subject when speaking can be obtained by extracting audio semantic feature information from the audio driving information, thereby generating a target facial image corresponding to the audio driving information.

[0243] 202. The server extracts spatial features from the original facial image frame to obtain original facial spatial features corresponding to the original facial image frame.

[0244] The original facial spatial features may specifically include three-dimensional (3D) facial coefficients corresponding to the original facial image frame, such as identity information, lighting, texture, expression, pose, gaze, etc.

[0245] 203. The server performs temporal feature extraction on the audio driving information to obtain local facial posture features corresponding to the target facial image frame.

[0246] Optionally, in this embodiment, the step of “extracting temporal features from the audio driving information to obtain local facial posture features corresponding to the target facial image frame” may include:

[0247] Performing feature extraction on each audio frame in the audio driving information to obtain audio semantic feature information of each audio frame;

[0248] Processing the audio semantic feature information of each audio frame based on the audio semantic feature information of the preceding and following audio frames of each audio frame;

[0249] The processed audio semantic feature information of each audio frame is fused to obtain the facial local posture feature corresponding to the target facial image frame.

[0250] It should be noted that in some embodiments, instead of using audio driving information, text driving information corresponding to the target facial image frame can be used to obtain the facial local posture features corresponding to the target facial image frame. In a specific scenario, for example, a speech video of the target object is damaged and some video frames in the speech video are lost. The subtitle information corresponding to the lost video frame can be used as text driving information, and the other non-lost video frames in the speech video can be used as the original facial image frames and the text driving information to generate and repair the lost video frame. Specifically, feature extraction can be performed on the text driving information to obtain the facial local posture features corresponding to the target facial image frame to be generated; spatial feature extraction can be performed on the original facial image frame to obtain the original facial spatial features corresponding to the original facial image frame, and facial reconstruction processing can be performed based on the original facial spatial features and the facial local posture features to generate the lost target facial image frame.

[0251] 204. The server fuses the original facial spatial features and the facial local posture features to obtain fused facial spatial features.

[0252] Among them, the local facial posture feature contains part of the facial posture information in the target facial image frame to be generated. For example, the local facial posture feature can include relevant three-dimensional facial coefficients such as expression, pose, and gaze; while the original facial spatial feature contains the facial posture information of the original facial image frame.

[0253] There are many ways to fuse the original facial spatial features and the facial local posture features. For example, the fusion method can be splicing processing or weighted fusion, etc., which is not limited in this embodiment.

[0254] 205. The server performs facial reconstruction processing on the target object based on the fused facial spatial features to obtain a reference facial image frame corresponding to the target object.

[0255] Optionally, in this embodiment, the step of “performing facial reconstruction processing on the target object based on the fused facial spatial features to obtain a reference facial image frame corresponding to the target object” may include:

[0256] Based on the fused facial spatial features, performing facial reconstruction processing on the target object to obtain a reconstructed three-dimensional facial image corresponding to the target object;

[0257] The reconstructed three-dimensional facial image is subjected to rendering mapping processing to obtain a reference facial image frame corresponding to the target object.

[0258] Among them, the texture and light of the reconstructed three-dimensional facial image can come from the original facial image frame, and the posture and expression of the reconstructed three-dimensional facial image can be driven by audio information. The reconstructed three-dimensional facial image is rendered and mapped, and the three-dimensional image can be projected onto a two-dimensional plane to obtain a reference facial image frame corresponding to the target object.

[0259] 206. The server generates the target facial image frame based on the original facial image frame, the fused facial spatial features, and the reference facial image frame.

[0260] Optionally, in this embodiment, the step of “generating the target facial image frame based on the original facial image frame, the fused facial spatial features, and the reference facial image frame” may include:

[0261] Performing multi-scale feature extraction on the original facial image frame to obtain original facial feature maps at multiple scales corresponding to the original facial image frame;

[0262] Performing multi-scale feature extraction on the reference facial image frame to obtain reference facial feature maps at multiple scales corresponding to the reference facial image frame;

[0263] Performing encoding mapping processing on the fused facial spatial features to obtain latent feature information corresponding to the fused facial spatial features;

[0264] The original facial feature maps at the multiple scales, the reference facial feature maps at the multiple scales, and the latent feature information are fused to obtain the target facial image frame.

[0265] Multi-scale feature extraction can be used to obtain spatial features of the original facial image frame or the reference facial image frame at each preset resolution. The image scales of the original facial feature maps corresponding to different resolutions are different. Similarly, the image scales of the reference facial feature maps corresponding to different resolutions are different. In this embodiment, multi-scale extraction can be used to obtain original spatial features of the original facial image frame and reference spatial features of the reference facial image frame. The original spatial features include original facial feature maps at multiple scales, and the reference spatial features include reference facial feature maps at multiple scales. Therefore, the original spatial features of the original facial image frame and the reference spatial features of the reference facial image frame are both multi-layer spatial features.

[0266] Optionally, in this embodiment, the step of “fusing the original facial feature maps at the multiple scales, the reference facial feature maps at the multiple scales, and the latent feature information to obtain the target facial image frame” may include:

[0267] fusing the latent feature information, the original facial feature map at a target scale, and the reference facial feature map at the target scale to obtain a fused facial feature map corresponding to the target scale, where the target scale is a scale selected from the multiple scales;

[0268] The fused facial feature map corresponding to the target scale, the original facial feature map at the adjacent scale, and the reference facial feature map at the adjacent scale are fused to obtain the target facial image frame.

[0269] Optionally, in this embodiment, the step of “fusing the corresponding fused facial feature map at the target scale, the original facial feature map at the adjacent scale, and the reference facial feature map at the adjacent scale to obtain the target facial image frame” may include:

[0270] Based on the latent feature information, performing style modulation processing on the fused facial feature map corresponding to the target scale to obtain a modulated style feature;

[0271] The modulated style features, the original facial feature map at adjacent scales, and the reference facial feature map at adjacent scales are fused to obtain the target facial image frame.

[0272] The adjacent scale can be a scale larger than the target scale among the multiple scales. Specifically, if the multiple scales include 4*4, 8*8, 16*16, 32*32, and 64*64, if the target scale is 16*16, the adjacent scale can be 32*32; if the target scale is 4*4, the adjacent scale can be 8*8.

[0273] The step of “fusing the modulated style features, the original facial feature map at adjacent scales, and the reference facial feature map at adjacent scales to obtain the target facial image frame” may include:

[0274] fusing the modulated style feature, the original facial feature map at an adjacent scale, and the reference facial feature map at the adjacent scale to obtain a fused facial feature map at the adjacent scale;

[0275] The facial feature maps at adjacent scales are fused with the basic facial image to generate a target facial image frame.

[0276] Optionally, in this embodiment, the step of “extracting spatial features from the original facial image frame to obtain original facial spatial features corresponding to the original facial image frame” may include:

[0277] Extracting spatial features of the original facial image frame using an image generation model to obtain original facial spatial features corresponding to the original facial image frame;

[0278] The step of “extracting temporal features from the audio driving information to obtain local facial posture features corresponding to the target facial image frame” may include:

[0279] Performing temporal feature extraction on the audio driving information through the image generation model to obtain local facial posture features corresponding to the target facial image frame;

[0280] The step of “performing facial reconstruction processing on the target object based on the original facial spatial features and the facial local posture features to generate the target facial image frame” may include:

[0281] The target object is subjected to facial reconstruction processing by the image generation model based on the original facial spatial features and the facial local posture features to generate the target facial image frame.

[0282] It should be noted that the image generation model can be trained by multiple sets of training data. The image generation model can be trained by other equipment and then provided to the image generation device, or it can be trained by the image generation device itself.

[0283] If the image generation device performs training by itself, the following steps may be further included before the step of "extracting spatial features from the original facial image frame using the image generation model to obtain original facial spatial features corresponding to the original facial image frame":

[0284] Acquire training data, the training data including original facial image frame samples of a sample object, target-driven facial image frame samples, and audio driving information samples corresponding to the target-driven facial image frame samples;

[0285] Extracting spatial features of the original facial image frame samples using a preset image generation model to obtain original facial spatial features corresponding to the original facial image frame samples;

[0286] Performing temporal feature extraction on the audio driving information sample to obtain facial local posture features corresponding to the target driving facial image frame sample;

[0287] Based on the original facial spatial features and the local facial posture features, performing facial reconstruction processing on the sample object to obtain a predicted driving facial image frame;

[0288] Based on the target driving facial image frame sample and the predicted driving facial image frame, the parameters of the preset image generation model are adjusted to obtain a trained image generation model.

[0289] The target driven facial image frame sample may be regarded as label information, and specifically may be the expected driven facial image frame corresponding to the audio driven information sample.

[0290] There are many ways to obtain the original facial image frame samples of the sample object, the target-driven facial image frame samples, and the audio driving information samples corresponding to the target-driven facial image frame samples, which are not limited in this embodiment.

[0291] For example, any two video frames containing the object's face can be extracted from a speech video about the sample object, one of which is used as the original facial image frame sample, the remaining frame is used as the target-driven facial image frame sample, and the audio information corresponding to the target-driven facial image frame sample 1 second before and after in the speech video is used as the audio-driven information sample.

[0292] Optionally, in this embodiment, the step of “adjusting parameters of a preset image generation model based on the target driving facial image frame sample and the predicted driving facial image frame to obtain a trained image generation model” may include:

[0293] Performing spatial feature extraction on the target-driven facial image frame sample to obtain target facial spatial features corresponding to the target-driven facial image frame sample;

[0294] Determining first loss information based on the facial local posture features corresponding to the target driven facial image frame sample and the target facial spatial features;

[0295] determining second loss information based on the target driving facial image frame sample and the predicted driving facial image frame;

[0296] According to the first loss information and the second loss information, the parameters of the preset image generation model are adjusted to obtain a trained image generation model.

[0297] The step of “adjusting parameters of a preset image generation model according to the first loss information and the second loss information to obtain a trained image generation model” may include:

[0298] fusing the first loss information and the second loss information to obtain total loss information;

[0299] According to the total loss information, the parameters of the preset image generation model are adjusted to obtain a trained image generation model.

[0300] There are many ways to fuse the first loss information and the second loss information, which are not limited in this embodiment. For example, the fusion method may be weighted fusion.

[0301] The training process of the preset image generation model first calculates the total loss information. Then, the backpropagation algorithm is used to adjust the parameters of the preset image generation model. The parameters of the image generation model are optimized based on the total loss information, so that the loss value corresponding to the total loss information is less than the preset loss value, thereby obtaining a trained image generation model. The preset loss value can be set according to actual conditions. For example, if the accuracy of the image generation model is required to be higher, the preset loss value should be smaller.

[0302] Optionally, in this embodiment, the step of “determining second loss information based on the target driving facial image frame sample and the predicted driving facial image frame” may include:

[0303] respectively predicting the probability that the target driving facial image frame sample and the predicted driving facial image frame belong to the real driving facial image frame, and determining adversarial loss information of a preset image generation model based on the probabilities;

[0304] Determining reconstruction loss information of a preset image generation model based on a similarity between the target driving facial image frame sample and the predicted driving facial image frame;

[0305] Performing identity recognition on the target driving facial image frame sample and the predicted driving facial image frame respectively, and determining identity loss information of a preset image generation model based on the identity recognition results;

[0306] Second loss information is determined according to the adversarial loss information, the reconstruction loss information, and the identity loss information.

[0307] As can be seen from the above, this embodiment can obtain the original facial image frame of the target object and the audio driving information corresponding to the target facial image frame to be generated through the server; perform spatial feature extraction on the original facial image frame to obtain the original facial spatial features corresponding to the original facial image frame; perform temporal feature extraction on the audio driving information to obtain the facial local posture features corresponding to the target facial image frame; fuse the original facial spatial features and the facial local posture features to obtain fused facial spatial features; based on the fused facial spatial features, perform facial reconstruction processing on the target object to obtain the reference facial image frame corresponding to the target object; generate the target facial image frame based on the original facial image frame, the fused facial spatial features and the reference facial image frame. This application can capture the facial posture details of the target object by performing feature extraction on the audio driving information, and then perform facial adjustments on the original facial image frame based on the captured information, thereby obtaining the target facial image frame corresponding to the audio driving information, which is conducive to improving the generation efficiency and accuracy of the target facial image frame.

[0308] In order to better implement the above method, the embodiment of the present application also provides an image generating device, such as Figure 3 As shown, the image generation device may include an acquisition unit 301, a first extraction unit 302, a second extraction unit 303, and a reconstruction unit 304, as follows:

[0309] (1) Acquisition unit 301;

[0310] The acquisition unit is used to acquire the original facial image frame of the target object and the audio driving information corresponding to the target facial image frame to be generated.

[0311] (2) a first extraction unit 302;

[0312] The first extraction unit is configured to extract spatial features from the original facial image frame to obtain original facial spatial features corresponding to the original facial image frame.

[0313] (3) second extraction unit 303;

[0314] The second extraction unit is used to perform time series feature extraction on the audio driving information to obtain local facial posture features corresponding to the target facial image frame.

[0315] Optionally, in some embodiments of the present application, the second extraction unit may include an extraction subunit, a processing subunit, and a first fusion subunit, as follows:

[0316] The extraction subunit is used to extract features from each audio frame in the audio driving information to obtain audio semantic feature information of each audio frame;

[0317] a processing subunit, configured to process the audio semantic feature information of each audio frame based on the audio semantic feature information of the preceding and following audio frames of each audio frame;

[0318] The first fusion subunit is used to fuse the processed audio semantic feature information of each audio frame to obtain the facial local posture feature corresponding to the target facial image frame.

[0319] (4) reconstruction unit 304;

[0320] The reconstruction unit is configured to perform facial reconstruction processing on the target object based on the original facial spatial features and the facial local posture features to generate the target facial image frame.

[0321] Optionally, in some embodiments of the present application, the reconstruction unit may include a second fusion subunit, a reconstruction subunit, and a generation subunit, as follows:

[0322] The second fusion subunit is configured to fuse the original facial spatial features with the facial local posture features to obtain fused facial spatial features;

[0323] A reconstruction subunit, configured to perform facial reconstruction processing on the target object based on the fused facial spatial features to obtain a reference facial image frame corresponding to the target object;

[0324] A generating subunit is configured to generate the target facial image frame based on the original facial image frame, the fused facial spatial features, and the reference facial image frame.

[0325] Optionally, in some embodiments of the present application, the reconstruction subunit can be specifically used to perform facial reconstruction processing on the target object based on the fused facial spatial features to obtain a reconstructed three-dimensional facial image corresponding to the target object; and perform rendering mapping processing on the reconstructed three-dimensional facial image to obtain a reference facial image frame corresponding to the target object.

[0326] Optionally, in some embodiments of the present application, the generating subunit can be specifically used to perform multi-scale feature extraction on the original facial image frame to obtain original facial feature maps at multiple scales corresponding to the original facial image frame; perform multi-scale feature extraction on the reference facial image frame to obtain reference facial feature maps at multiple scales corresponding to the reference facial image frame; perform encoding mapping processing on the fused facial spatial features to obtain latent feature information corresponding to the fused facial spatial features; and fuse the original facial feature maps at multiple scales, the reference facial feature maps at multiple scales, and the latent feature information to obtain the target facial image frame.

[0327] Optionally, in some embodiments of the present application, the step of “fusing the original facial feature maps at the multiple scales, the reference facial feature maps at the multiple scales, and the latent feature information to obtain the target facial image frame” may include:

[0328] fusing the latent feature information, the original facial feature map at a target scale, and the reference facial feature map at the target scale to obtain a fused facial feature map corresponding to the target scale, where the target scale is a scale selected from the multiple scales;

[0329] The fused facial feature map corresponding to the target scale, the original facial feature map at the adjacent scale, and the reference facial feature map at the adjacent scale are fused to obtain the target facial image frame.

[0330] Optionally, in some embodiments of the present application, the step of “fusing the corresponding fused facial feature map at the target scale, the original facial feature map at the adjacent scale, and the reference facial feature map at the adjacent scale to obtain the target facial image frame” may include:

[0331] Based on the latent feature information, performing style modulation processing on the fused facial feature map corresponding to the target scale to obtain a modulated style feature;

[0332] The modulated style features, the original facial feature map at adjacent scales, and the reference facial feature map at adjacent scales are fused to obtain the target facial image frame.

[0333] Optionally, in some embodiments of the present application, the first extraction unit may be specifically configured to perform spatial feature extraction on the original facial image frame using an image generation model to obtain original facial spatial features corresponding to the original facial image frame;

[0334] The second extraction unit may be specifically configured to perform temporal feature extraction on the audio driving information using the image generation model to obtain local facial posture features corresponding to the target facial image frame;

[0335] The reconstruction unit can be specifically configured to perform facial reconstruction processing on the target object based on the original facial spatial features and the facial local posture features through the image generation model to generate the target facial image frame.

[0336] Optionally, in some embodiments of the present application, the image generation device may further include a training unit, and the training unit may be used to train the image generation model;

[0337] The training unit can be specifically used to obtain training data, which includes original facial image frame samples of the sample object, target-driven facial image frame samples, and audio-driven information samples corresponding to the target-driven facial image frame samples; through a preset image generation model, spatial feature extraction is performed on the original facial image frame samples to obtain original facial spatial features corresponding to the original facial image frame samples; temporal feature extraction is performed on the audio-driven information samples to obtain facial local posture features corresponding to the target-driven facial image frame samples; based on the original facial spatial features and the facial local posture features, facial reconstruction processing is performed on the sample object to obtain a predicted-driven facial image frame; based on the target-driven facial image frame samples and the predicted-driven facial image frame, the parameters of the preset image generation model are adjusted to obtain a trained image generation model.

[0338] Optionally, in some embodiments of the present application, the step of “adjusting parameters of a preset image generation model based on the target driving facial image frame sample and the predicted driving facial image frame to obtain a trained image generation model” may include:

[0339] Performing spatial feature extraction on the target-driven facial image frame sample to obtain target facial spatial features corresponding to the target-driven facial image frame sample;

[0340] Determining first loss information based on the facial local posture features corresponding to the target driven facial image frame sample and the target facial spatial features;

[0341] determining second loss information based on the target driving facial image frame sample and the predicted driving facial image frame;

[0342] According to the first loss information and the second loss information, the parameters of the preset image generation model are adjusted to obtain a trained image generation model.

[0343] Optionally, in some embodiments of the present application, the step of “determining second loss information based on the target driving facial image frame sample and the predicted driving facial image frame” may include:

[0344] respectively predicting the probability that the target driving facial image frame sample and the predicted driving facial image frame belong to the real driving facial image frame, and determining adversarial loss information of a preset image generation model based on the probabilities;

[0345] Determining reconstruction loss information of a preset image generation model based on a similarity between the target driving facial image frame sample and the predicted driving facial image frame;

[0346] Performing identity recognition on the target driving facial image frame sample and the predicted driving facial image frame respectively, and determining identity loss information of a preset image generation model based on the identity recognition results;

[0347] Second loss information is determined according to the adversarial loss information, the reconstruction loss information, and the identity loss information.

[0348] As can be seen from the above, this embodiment can obtain the original facial image frame of the target object and the audio driving information corresponding to the target facial image frame to be generated through the acquisition unit 301; the first extraction unit 302 performs spatial feature extraction on the original facial image frame to obtain the original facial spatial features corresponding to the original facial image frame; the second extraction unit 303 performs temporal feature extraction on the audio driving information to obtain the facial local posture features corresponding to the target facial image frame; the reconstruction unit 304 performs facial reconstruction processing on the target object based on the original facial spatial features and the facial local posture features to generate the target facial image frame. The present application can capture the facial posture details of the target object by performing feature extraction on the audio driving information, and then perform facial adjustments on the original facial image frame based on the captured information, thereby obtaining the target facial image frame corresponding to the audio driving information, which is conducive to improving the generation efficiency and accuracy of the target facial image frame.

[0349] The present application also provides an electronic device, such as Figure 4 , which shows a schematic diagram of the structure of an electronic device involved in an embodiment of the present application. The electronic device may be a terminal or a server, etc. Specifically:

[0350] The electronic device may include one or more processing core processors 401, one or more computer-readable storage media memories 402, a power supply 403, an input unit 404 and other components. Those skilled in the art will understand that Figure 4 The electronic device structure shown in the figure does not constitute a limitation of the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange components differently.

[0351] Processor 401 is the control center of the electronic device. It connects all parts of the electronic device using various interfaces and circuits. It performs various functions of the electronic device and processes data by running or executing software programs and / or modules stored in memory 402 and accessing data stored in memory 402. Optionally, processor 401 may include one or more processing cores. Preferably, processor 401 may integrate an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 401.

[0352] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.

[0353] The electronic device also includes a power supply 403 for supplying power to various components. Preferably, the power supply 403 can be logically connected to the processor 401 via a power management system, thereby enabling the power management system to manage charging, discharging, and power consumption. The power supply 403 can also include one or more DC or AC power supplies, a recharging system, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components.

[0354] The electronic device may further include an input unit 404, which may be configured to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0355] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 will run the application programs stored in the memory 402 to implement various functions as follows:

[0356] An original facial image frame of a target object and audio driving information corresponding to a target facial image frame to be generated are obtained; spatial feature extraction is performed on the original facial image frame to obtain original facial spatial features corresponding to the original facial image frame; temporal feature extraction is performed on the audio driving information to obtain local facial posture features corresponding to the target facial image frame; facial reconstruction processing is performed on the target object based on the original facial spatial features and the local facial posture features to generate the target facial image frame.

[0357] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0358] As can be seen from the above, this embodiment can obtain the original facial image frame of the target object and the audio driving information corresponding to the target facial image frame to be generated; perform spatial feature extraction on the original facial image frame to obtain the original facial spatial features corresponding to the original facial image frame; perform temporal feature extraction on the audio driving information to obtain the facial local posture features corresponding to the target facial image frame; based on the original facial spatial features and the facial local posture features, perform facial reconstruction processing on the target object to generate the target facial image frame. The present application can capture the facial posture details of part of the target object by performing feature extraction on the audio driving information, and then perform facial adjustments on the original facial image frame based on the captured information, thereby obtaining the target facial image frame corresponding to the audio driving information, which is conducive to improving the generation efficiency and accuracy of the target facial image frame.

[0359] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0360] To this end, an embodiment of the present application provides a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute the steps of any of the image generation methods provided in the embodiments of the present application. For example, the instructions can execute the following steps:

[0361] An original facial image frame of a target object and audio driving information corresponding to a target facial image frame to be generated are obtained; spatial feature extraction is performed on the original facial image frame to obtain original facial spatial features corresponding to the original facial image frame; temporal feature extraction is performed on the audio driving information to obtain local facial posture features corresponding to the target facial image frame; facial reconstruction processing is performed on the target object based on the original facial spatial features and the local facial posture features to generate the target facial image frame.

[0362] The specific implementation of the above operations can be found in the previous embodiments and will not be repeated here.

[0363] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0364] Since the instructions stored in the computer-readable storage medium can execute the steps in any image generation method provided in the embodiments of the present application, the beneficial effects that can be achieved by any image generation method provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0365] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various optional implementations of the aforementioned image generation aspects.

[0366] The above is a detailed introduction to an image generation method and related equipment provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, based on the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. An image generation method, characterized in that: include: Obtaining audio driving information corresponding to an original facial image frame of a target object and a target facial image frame to be generated; Performing spatial feature extraction on the original facial image frame to obtain original facial spatial features corresponding to the original facial image frame; Performing temporal feature extraction on the audio driving information to obtain local facial posture features corresponding to the target facial image frame; Fusing the original facial spatial features with the facial local posture features to obtain fused facial spatial features; Based on the fused facial spatial features, performing facial reconstruction processing on the target object to obtain a reference facial image frame corresponding to the target object; Performing multi-scale feature extraction on the original facial image frame to obtain original facial feature maps at multiple scales corresponding to the original facial image frame; Performing multi-scale feature extraction on the reference facial image frame to obtain reference facial feature maps at multiple scales corresponding to the reference facial image frame; Performing encoding mapping processing on the fused facial spatial features to obtain latent feature information corresponding to the fused facial spatial features; The method comprises fusing the original facial feature maps at the multiple scales, the reference facial feature maps at the multiple scales, and the latent feature information to obtain the target facial image frame, including: fusing the latent feature information, the original facial feature map at the target scale, and the reference facial feature map at the target scale to obtain a fused facial feature map corresponding to the target scale, where the target scale is a scale selected from the multiple scales; performing style modulation processing on the fused facial feature map corresponding to the target scale based on the latent feature information to obtain a modulated style feature; and fusing the modulated style feature, the original facial feature map at an adjacent scale, and the reference facial feature map at the adjacent scale to obtain the target facial image frame.

2. The method according to claim 1, characterized in that The extracting time series features of the audio driving information to obtain local facial posture features corresponding to the target facial image frame includes: Performing feature extraction on each audio frame in the audio driving information to obtain audio semantic feature information of each audio frame; Processing the audio semantic feature information of each audio frame based on the audio semantic feature information of the preceding and following audio frames of each audio frame; The processed audio semantic feature information of each audio frame is fused to obtain the facial local posture feature corresponding to the target facial image frame.

3. The method according to claim 1, characterized in that The step of performing facial reconstruction on the target object based on the fused facial spatial features to obtain a reference facial image frame corresponding to the target object includes: Based on the fused facial spatial features, performing facial reconstruction processing on the target object to obtain a reconstructed three-dimensional facial image corresponding to the target object; The reconstructed three-dimensional facial image is subjected to rendering mapping processing to obtain a reference facial image frame corresponding to the target object.

4. The method according to claim 1, wherein The extracting spatial features of the original facial image frame to obtain original facial spatial features corresponding to the original facial image frame includes: Extracting spatial features of the original facial image frame using an image generation model to obtain original facial spatial features corresponding to the original facial image frame; The extracting time series features of the audio driving information to obtain local facial posture features corresponding to the target facial image frame includes: Performing temporal feature extraction on the audio driving information through the image generation model to obtain local facial posture features corresponding to the target facial image frame; The step of performing facial reconstruction on the target object based on the original facial spatial features and the facial local posture features to generate the target facial image frame includes: The target object is subjected to facial reconstruction processing by the image generation model based on the original facial spatial features and the facial local posture features to generate the target facial image frame.

5. The method according to claim 4, characterized in that Before extracting spatial features from the original facial image frame using the image generation model to obtain original facial spatial features corresponding to the original facial image frame, the method further includes: Acquire training data, the training data including original facial image frame samples of a sample object, target-driven facial image frame samples, and audio driving information samples corresponding to the target-driven facial image frame samples; Extracting spatial features of the original facial image frame samples using a preset image generation model to obtain original facial spatial features corresponding to the original facial image frame samples; Performing temporal feature extraction on the audio driving information sample to obtain facial local posture features corresponding to the target driving facial image frame sample; Based on the original facial spatial features and the local facial posture features, performing facial reconstruction processing on the sample object to obtain a predicted driving facial image frame; Based on the target driving facial image frame sample and the predicted driving facial image frame, the parameters of the preset image generation model are adjusted to obtain a trained image generation model.

6. The method according to claim 5, characterized in that The method further comprises adjusting parameters of a preset image generation model based on the target driving facial image frame sample and the predicted driving facial image frame to obtain a trained image generation model, including: Performing spatial feature extraction on the target-driven facial image frame sample to obtain target facial spatial features corresponding to the target-driven facial image frame sample; Determining first loss information based on the facial local posture features corresponding to the target driven facial image frame sample and the target facial spatial features; determining second loss information based on the target driving facial image frame sample and the predicted driving facial image frame; According to the first loss information and the second loss information, the parameters of the preset image generation model are adjusted to obtain a trained image generation model.

7. The method according to claim 6, characterized in that The determining of second loss information based on the target driving facial image frame sample and the predicted driving facial image frame includes: respectively predicting the probability that the target driving facial image frame sample and the predicted driving facial image frame belong to the real driving facial image frame, and determining adversarial loss information of a preset image generation model based on the probabilities; Determining reconstruction loss information of a preset image generation model based on a similarity between the target driving facial image frame sample and the predicted driving facial image frame; Performing identity recognition on the target driving facial image frame sample and the predicted driving facial image frame respectively, and determining identity loss information of a preset image generation model based on the identity recognition results; Second loss information is determined according to the adversarial loss information, the reconstruction loss information, and the identity loss information.

8. An image generating device, characterized in that: include: an acquisition unit, configured to acquire audio driving information corresponding to an original facial image frame of a target object and a target facial image frame to be generated; a first extraction unit, configured to extract spatial features from the original facial image frame to obtain original facial spatial features corresponding to the original facial image frame; A second extraction unit is configured to extract temporal features from the audio driving information to obtain local facial posture features corresponding to the target facial image frame; A reconstruction unit, configured to perform facial reconstruction processing on the target object based on the original facial spatial features and the facial local posture features to generate the target facial image frame; A second fusion unit is used to fuse the original facial spatial features and the facial local posture features to obtain a fused facial spatial feature; a reconstruction unit, configured to perform facial reconstruction processing on the target object based on the fused facial spatial features to obtain a reference facial image frame corresponding to the target object; a third extraction unit, configured to perform multi-scale feature extraction on the original facial image frame to obtain original facial feature maps at multiple scales corresponding to the original facial image frame; a fourth extraction unit, configured to perform multi-scale feature extraction on the reference facial image frame to obtain reference facial feature maps at multiple scales corresponding to the reference facial image frame; a coding mapping unit, configured to perform coding mapping processing on the fused facial spatial features to obtain latent feature information corresponding to the fused facial spatial features; A third fusion unit is configured to fuse the original facial feature maps at the multiple scales, the reference facial feature maps at the multiple scales, and the latent feature information to obtain the target facial image frame, including: fusing the latent feature information, the original facial feature map at the target scale, and the reference facial feature map at the target scale to obtain a fused facial feature map corresponding to the target scale, where the target scale is a scale selected from the multiple scales; performing style modulation processing on the fused facial feature map corresponding to the target scale based on the latent feature information to obtain a modulated style feature; and fusing the modulated style feature, the original facial feature map at an adjacent scale, and the reference facial feature map at the adjacent scale to obtain the target facial image frame.

9. An electronic device, characterized in that: The method comprises a memory and a processor; the memory stores an application program, and the processor is used to run the application program in the memory to perform the operation in the image generation method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the steps in the image generation method according to any one of claims 1 to 7.

11. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed by a processor, the steps of the image generation method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Video generation method and device, server and storage medium

    CN113901894A