Image generation method and related device
By performing feature encoding and audio-driven information processing under basic expression states, facial images of the target object under different expression states are generated, solving the problem of low efficiency in facial image generation in existing technologies and achieving efficient and accurate image generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- TENCENT TECHNOLOGY (SHENZHEN) CO LTD
- Filing Date
- 2024-11-28
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, acquiring facial images of the same target object under different facial poses is time-consuming and labor-intensive, resulting in low image generation efficiency.
By acquiring the base facial image frame of the target object in the basic expression state, and using audio-driven information for feature extraction and adjustment, the target facial image frame of the target object in the expression state corresponding to the audio-driven information is generated, including feature encoding, feature extraction and expression adjustment processing.
It significantly reduces computational load, improves the efficiency and accuracy of target facial image frame generation, and simplifies the facial image generation process.
Smart Images

Figure CN122115701A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, specifically to an image generation method and related equipment. Background Technology
[0002] With the development of computer technology, image processing technology has been applied to more and more fields. For example, image processing technology can include image generation, specifically facial image generation, which can be applied to fields such as animation production.
[0003] In current related technologies, if it is necessary to obtain facial images of the same target object under different facial poses, such as facial images of the target object under different facial poses during speech, the modeler and animator need to draw the facial images under each facial pose separately. This image generation method is relatively time-consuming and labor-intensive, and the image generation efficiency is low. Summary of the Invention
[0004] This application provides an image generation method and related equipment. The related equipment may include an image generation device, an electronic device, a computer-readable storage medium, and a computer program product, which can significantly reduce the amount of computation and improve the generation efficiency and accuracy of target facial image frames.
[0005] This application provides an image generation method, including:
[0006] Obtain the base facial image frame of the target object in its basic expression state, and obtain the audio driving information used for image generation;
[0007] The basic facial image frame is subjected to feature encoding processing to obtain a basic facial encoded feature map;
[0008] The audio driving information is subjected to feature extraction processing to obtain the audio features corresponding to the audio driving information;
[0009] Based on the audio features, the basic facial coding feature map is processed to adjust the facial expression, so as to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0010] Accordingly, embodiments of this application provide an image generation apparatus, including:
[0011] The acquisition unit is used to acquire the base facial image frame of the target object in the basic expression state, and to acquire the audio driving information for image generation;
[0012] An encoding unit is used to perform feature encoding processing on the basic facial image frame to obtain a basic facial encoded feature map;
[0013] The feature extraction unit is used to perform feature extraction processing on the audio driving information to obtain the audio features corresponding to the audio driving information.
[0014] The generation unit is used to perform expression adjustment processing on the basic facial coding feature map based on the audio features, so as to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0015] Optionally, in some embodiments of this application, the acquisition unit may include an acquisition subunit, an expression coefficient determination subunit, a pose feature extraction subunit, and a reconstruction subunit, as follows:
[0016] The acquisition subunit is used to acquire the initial facial image frame of the target object in its initial expression state;
[0017] The expression coefficient determination subunit is used to determine the preset basic expression coefficient corresponding to the basic expression state;
[0018] The pose feature extraction subunit is used to extract facial pose features from the initial facial image frame to obtain the initial facial coefficients corresponding to the initial facial image frame.
[0019] The reconstruction subunit is used to perform facial reconstruction processing on the target object based on the preset basic expression coefficients and the initial facial coefficients to obtain the basic facial image frame of the target object in the basic expression state.
[0020] Optionally, in some embodiments of this application, the initial facial coefficients include initial expression coefficients and initial facial morphology coefficients;
[0021] The reconstruction subunit can be specifically used to replace the initial expression coefficients in the initial facial coefficients based on the preset basic expression coefficients to obtain target facial coefficients, wherein the target facial coefficients include the preset basic expression coefficients and the initial facial morphology coefficients; and to perform facial reconstruction processing on the target object according to the target facial coefficients to obtain the basic facial image frame of the target object in the basic expression state.
[0022] Optionally, in some embodiments of this application, the step "performing facial reconstruction processing on the target object based on the target facial coefficients to obtain a basic facial image frame of the target object in a basic expression state" may include:
[0023] Based on the target facial coefficients, the target object is subjected to facial reconstruction processing to obtain the reconstructed three-dimensional facial image of the target object in the basic expression state;
[0024] The reconstructed 3D facial image is rendered and mapped to obtain the base facial image frame of the target object in its basic expression state.
[0025] Optionally, in some embodiments of this application, the encoding unit may include a downsampling subunit and an attention processing subunit, as follows:
[0026] The downsampling subunit is used to perform multiple downsampling processes on the basic facial image frame to obtain a facial feature map;
[0027] The attention processing subunit is used to perform attention processing on the facial feature map to obtain the basic facial coding feature map.
[0028] Optionally, in some embodiments of this application, the generation unit may include a feature fusion subunit and a generation subunit, as follows:
[0029] The feature fusion subunit is used to perform feature fusion processing on the audio features and the basic facial coding feature map to obtain initial fused feature information;
[0030] A generation subunit is used to inject the audio features into the initial fused feature information for feature interaction processing, so as to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0031] Optionally, in some embodiments of this application, the generation subunit may specifically be used to perform convolution processing on the initial fusion feature information to obtain convolution result information; and to perform fusion processing on the convolution result information and the audio features to obtain target fusion feature information; to perform convolution processing on the target fusion feature information to obtain target convolution result information; to perform fusion processing on the target convolution result information and the basic facial coding feature map to obtain a processed facial coding feature map; and to perform feature decoding processing on the processed facial coding feature map according to the audio features to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0032] Optionally, in some embodiments of this application, the encoding unit may be used to perform feature encoding processing on the basic facial image frame through an image generation model to obtain a basic facial encoded feature map.
[0033] Optionally, in some embodiments of this application, the image generation apparatus may further include a training unit for training the image generation model; specifically, the training unit may include a training data acquisition subunit, a feature encoding subunit, a feature extraction subunit, an expression adjustment subunit, and a parameter adjustment subunit, as follows:
[0034] The training data acquisition subunit is used to acquire training data, which includes basic facial image frame samples of the sample object, target-driven facial image frame samples, and audio driving information samples corresponding to the target-driven facial image frame samples.
[0035] The feature encoding subunit is used to perform feature encoding processing on the basic facial image frame samples through an image generation model to obtain the basic facial encoding feature map corresponding to the basic facial image frame samples.
[0036] The feature extraction subunit is used to perform feature extraction processing on the audio driving information sample to obtain sample audio features;
[0037] The expression adjustment subunit is used to perform expression adjustment processing on the basic facial coding feature map based on the sample audio features, so as to generate a predicted driven facial image frame of the sample object in the corresponding expression state of the audio driving information sample.
[0038] The parameter adjustment subunit is used to adjust the parameters of the image generation model based on the target-driven facial image frame samples and the prediction-driven facial image frames to obtain the trained image generation model.
[0039] Optionally, in some embodiments of this application, the parameter adjustment subunit may specifically be used to calculate image loss information between the target-driven facial image frame sample and the prediction-driven facial image frame; extract target facial feature information corresponding to the target-driven facial image frame sample and prediction facial feature information corresponding to the prediction-driven facial image frame; calculate feature loss information based on the target facial feature information and the prediction facial feature information; perform realism discrimination processing on the target-driven facial image frame sample and the prediction-driven facial image frame respectively, so as to determine the generative adversarial loss information of the image generation model based on the discrimination result; and adjust the parameters of the image generation model based on the image loss information, the feature loss information and the generative adversarial loss information to obtain the trained image generation model.
[0040] Optionally, in some embodiments of this application, the image generation apparatus may further include a training data construction unit, as follows:
[0041] The training data construction unit is used to acquire a sample generation model, at least one audio-driven information sample, and a basic facial image frame sample of the sample object; extract facial pose features from the basic facial image frame sample using the sample generation model to obtain basic facial coefficients corresponding to the basic facial image frame sample; extract temporal features from the audio-driven information sample to obtain target expression coefficients corresponding to the expression state of the audio-driven information sample; perform facial reconstruction processing on the sample object based on the basic facial coefficients and the target expression coefficients to obtain target driven facial image frame samples of the sample object in the expression state corresponding to the audio-driven information sample; and construct training data based on the target driven facial image frame samples, the audio-driven information sample, and the basic facial image frame samples.
[0042] Optionally, in some embodiments of this application, the image generation apparatus may further include an audio sample generation unit, as follows:
[0043] The audio sample generation unit is used to acquire preset facial expression guidance text information; and to perform audio conversion processing on the preset facial expression guidance text information in multiple languages to obtain audio driving information samples in the multiple languages.
[0044] Optionally, in some embodiments of this application, the step "performing facial reconstruction processing on the sample object based on the basic facial coefficients and the target expression coefficients to obtain a target driven facial image frame sample of the sample object in the expression state corresponding to the audio-driven information sample" may include:
[0045] Based on the target expression coefficient, the expression coefficients in the basic facial coefficients are replaced to obtain the target facial coefficients;
[0046] The sample object is reconstructed based on the target facial coefficients to obtain a target driven facial image frame sample of the sample object in the expression state corresponding to the audio driven information sample.
[0047] In this application embodiment, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.
[0048] An electronic device provided in this application includes a processor and a memory. The memory stores multiple instructions, and the processor loads the instructions to execute the steps in the image generation method provided in this application.
[0049] This application also provides a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the steps in the image generation method provided in this application.
[0050] Furthermore, embodiments of this application also provide a computer program product, including a computer program or instructions, which, when executed by a processor, implement the steps in the image generation method provided in embodiments of this application.
[0051] This application provides an image generation method and related equipment, which can acquire a basic facial image frame of a target object in a basic expression state, and acquire audio driving information for image generation; perform feature encoding processing on the basic facial image frame to obtain a basic facial encoding feature map; perform feature extraction processing on the audio driving information to obtain audio features corresponding to the audio driving information; and perform expression adjustment processing on the basic facial encoding feature map based on the audio features to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0052] This application can utilize some facial pose details contained in audio features to perform facial expression adjustment processing on basic facial image frames, thereby obtaining target facial image frames corresponding to audio-driven information. This is beneficial to improving the generation efficiency and accuracy of target facial image frames. Moreover, this application generates images based on basic facial image frames under basic expression states. Compared with directly adjusting complex expression states, this expression adjustment under basic expression states can significantly reduce the amount of computation and further improve image generation efficiency. Attached Figure Description
[0053] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0054] Figure 1a This is a schematic diagram of a scene illustrating the image generation method provided in an embodiment of this application;
[0055] Figure 1b This is a flowchart of the image generation method provided in the embodiments of this application;
[0056] Figure 1c These are illustrative diagrams illustrating the image generation method provided in the embodiments of this application;
[0057] Figure 1d This is another illustrative diagram of the image generation method provided in the embodiments of this application;
[0058] Figure 1e This is another illustrative diagram of the image generation method provided in the embodiments of this application;
[0059] Figure 1f This is a model architecture diagram of the image generation method provided in the embodiments of this application;
[0060] Figure 2 This is another flowchart of the image generation method provided in the embodiments of this application;
[0061] Figure 3 This is a schematic diagram of the structure of the image generation apparatus provided in the embodiments of this application;
[0062] Figure 4 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation
[0063] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0064] This application provides an image generation method and related equipment. The related equipment may include an image generation apparatus, an electronic device, a computer-readable storage medium, and a computer program product. Specifically, the image generation apparatus may be integrated into an electronic device, which may be a terminal or a server, etc.
[0065] It is understood that the image generation method of this embodiment can be executed on a terminal, on a server, or jointly by a terminal and a server. The above examples should not be construed as limiting this application.
[0066] like Figure 1a As shown, taking the joint execution of an image generation method by a terminal and a server as an example, the image generation system provided in this application embodiment includes a terminal 10 and a server 11, etc.; the terminal 10 and the server 11 are connected via a network, such as a wired or wireless network, etc., wherein the image generation device can be integrated into the server.
[0067] Server 11 can be used to: acquire a basic facial image frame of a target object in a basic expression state, and acquire audio driving information for image generation; perform feature encoding processing on the basic facial image frame to obtain a basic facial encoding feature map; perform feature extraction processing on the audio driving information to obtain audio features corresponding to the audio driving information; and perform expression adjustment processing on the basic facial encoding feature map based on the audio features to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information. Server 11 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.
[0068] Terminal 10 can be used to: send a basic facial image frame of the target object in a basic expression state, and audio driving information for image generation to server 11, so that server 11 can generate a target facial image frame of the target object in the expression state corresponding to the audio driving information; terminal 10 can also receive the target facial image frame sent by server 11. Terminal 10 may include a mobile phone, vehicle terminal, aircraft, tablet computer, laptop computer, or personal computer (PC), etc. A client can also be set on terminal 10, which can be an application client or a browser client, etc.
[0069] The image generation and other steps in the aforementioned server 11 can also be performed by the terminal 10.
[0070] The image generation method provided in this application relates to machine learning and computer vision technologies in the field of artificial intelligence.
[0071] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.
[0072] This embodiment will be described from the perspective of an image generating device, which can be integrated into an electronic device, such as a server or a terminal.
[0073] It is understood that in the specific embodiments of this application, data such as user information are involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0074] like Figure 1b As shown, the specific process of this image generation method can be as follows:
[0075] 101. Obtain the base facial image frame of the target object in its basic expression state, and obtain the audio driving information used for image generation.
[0076] The target object can be an object whose facial posture needs to be adjusted. Facial posture specifically refers to facial expressions, such as mouth shape and eye contact; this embodiment does not impose such limitations.
[0077] The basic facial image frame is an image containing the face of the target object, specifically the facial image to be driven. Specifically, the basic facial image frame is a facial image of the target object in a basic expression state. The basic expression state can be the facial posture corresponding to not speaking, such as a facial posture with the mouth closed. The basic facial image frame is also the image frame corresponding to a silent face, which specifically refers to a face in a state where the mouth is not open.
[0078] The audio driving information is audio used to adjust the facial pose in the base facial image frame. Specifically, it can be used to replace the facial pose of the target object in the base facial image frame with the facial pose corresponding to the target object speaking, thereby obtaining the target facial image frame. The audio information corresponding to the target object speaking is the audio driving information. The audio length corresponding to the audio driving information can be 1 second or 2 seconds; this embodiment does not impose any limitation on this.
[0079] In this embodiment, the facial posture changes of the target object can be determined by utilizing the changes in lip movements and other information contained in the audio driving information.
[0080] In some embodiments, the target's emotions can be determined based on the content and volume of the target's speech contained in the audio driving information, and the facial posture changes of the target, such as changes in eye expression, can be determined in combination with the determination result.
[0081] In a specific scenario, this embodiment can acquire multiple audio driving information segments of the target object, and for each audio driving information segment, generate a target facial image frame corresponding to the target expression state based on each audio driving information segment and the basic facial image frame. Then, the target facial image frames corresponding to the target expression states of each audio driving information segment are spliced together to generate a target facial video segment corresponding to the target object. The target facial video segment contains the facial posture change process of the target object when speaking (expressing the content in these audio driving information segments). The audio information corresponding to the target object speaking is the audio driving information segment. It should be noted that the main subject in the target facial video segment is still the object's face in the basic facial image frame, and the expression (especially the mouth shape) of the target object in the generated target facial video segment corresponds to each audio driving information segment.
[0082] In one embodiment, this application can also be applied to video repair scenarios. For example, if a speech video about a target object is damaged and some video frames in the speech video are lost, the lost video frames can be repaired by using the image generation method provided by this application, along with other video frames in the speech video and the audio information corresponding to the lost video frames. Specifically, the audio information used for repair can be an audio segment one second before and after the lost video frame in the speech video, which is also the audio driving information in the above embodiment.
[0083] The image generation method provided in this application can generate images from basic facial image frames under basic expression states. Compared with directly adjusting complex expression states, adjusting expressions under basic expression states can significantly reduce the amount of computation and improve image generation efficiency.
[0084] Optionally, in this embodiment, the step of "obtaining the basic facial image frame of the target object in a basic expression state" may include:
[0085] Obtain the initial facial image frame of the target object in its initial expression state;
[0086] Determine the preset basic expression coefficients corresponding to the basic expression states;
[0087] Facial pose features are extracted from the initial facial image frame to obtain the initial facial coefficients corresponding to the initial facial image frame;
[0088] Based on the preset basic expression coefficients and the initial facial coefficients, the target object is subjected to facial reconstruction processing to obtain the basic facial image frame of the target object in the basic expression state.
[0089] In some embodiments, the facial image frames of the target object initially acquired may not be in a basic expression state. In this case, it is necessary to adjust the expression state of the initial facial image frames to the basic expression state to reduce subsequent computational costs. If the facial image frames initially acquired are in a basic expression state, no adjustment is required.
[0090] The initial facial expression state can be the facial posture corresponding to the target object when speaking, such as a facial posture with the mouth open.
[0091] Among them, the preset basic expression coefficient corresponding to the basic expression state can be a preset value, which can be determined based on expert experience. For example, all coefficients in the preset basic expression coefficient can be 0.
[0092] The step "extracting facial pose features from the initial facial image frame to obtain the initial facial coefficients corresponding to the initial facial image frame" can specifically involve extracting spatial features from the initial facial image frame to obtain its facial spatial features. These facial spatial features can specifically include three-dimensional (3D) facial coefficients corresponding to the initial facial image frame, such as identity information, lighting, texture, expression, and pose. Expression information can include eye contact, lip shape, and eyebrow shape. Based on these facial coefficients, the face of the initial facial image frame can be reconstructed. Because 3D has good decoupling properties for facial information, the shape of the face and its expression can be decoupled.
[0093] Specifically, spatial feature extraction of the initial facial image frame can involve convolution and pooling processes, etc., and this embodiment is not limited to this. In this embodiment, spatial features of the initial facial image frame can be extracted using an image feature extraction network. This image feature extraction network can be a neural network model, such as a Visual Geometry Group Network (VGGNet), a Residual Network (ResNet), or a Dense Convolutional Network (DenseNet), etc. However, it should be understood that the neural network in this embodiment is not limited to the types listed above. Specifically, the image feature extraction network is pre-trained and can predict the three-dimensional facial coefficients corresponding to the facial image.
[0094] Optionally, in this embodiment, the initial facial coefficients include initial expression coefficients and initial facial morphology coefficients;
[0095] The step "based on the preset basic expression coefficients and the initial facial coefficients, perform facial reconstruction processing on the target object to obtain a basic facial image frame of the target object in a basic expression state" may include:
[0096] Based on the preset basic expression coefficients, the initial expression coefficients in the initial facial coefficients are replaced to obtain the target facial coefficients, which include the preset basic expression coefficients and the initial facial morphology coefficients.
[0097] The target object is reconstructed based on the target facial coefficients to obtain a basic facial image frame of the target object in a basic expression state.
[0098] Specifically, the initial facial coefficients may include identity information, lighting, texture, posture, eye expression, mouth shape, eyebrow shape, etc. Among them, eye expression, mouth shape, and eyebrow shape can belong to the initial expression coefficients. The initial facial morphology coefficients can be other coefficients in the initial facial coefficients besides the initial expression coefficients, such as identity information, lighting, texture, and posture.
[0099] In this embodiment, after obtaining the initial facial coefficients, coefficients related to the expression of the target object can be selected from the initial facial coefficients, namely initial expression coefficients. For example, initial expression coefficients such as eye expression, mouth shape, and eyebrow shape can be extracted from the initial facial coefficients. Then, these initial expression coefficients are replaced with preset basic expression coefficients to obtain the replaced target facial coefficients.
[0100] Optionally, in this embodiment, the step "performing facial reconstruction processing on the target object based on the target facial coefficients to obtain a basic facial image frame of the target object in a basic expression state" may include:
[0101] Based on the target facial coefficients, the target object is subjected to facial reconstruction processing to obtain the reconstructed three-dimensional facial image of the target object in the basic expression state;
[0102] The reconstructed 3D facial image is rendered and mapped to obtain the base facial image frame of the target object in its basic expression state.
[0103] Specifically, facial reconstruction can be performed on the target object based on the target facial coefficients using a three-dimensional facial driving model (such as 3DMM, 3D Morphable Model). This facial reconstruction process is also known as 3D reconstruction. 3D reconstruction can represent the input two-dimensional facial image using a 3D mesh (three-dimensional mesh model). The 3D mesh can contain the vertex coordinates and colors of the three-dimensional mesh structure.
[0104] The texture and lighting of the reconstructed 3D facial image can be derived from the initial facial image frame, and the pose and expression of the reconstructed 3D facial image can be derived from the preset basic expression coefficients. By performing rendering and mapping processing on the reconstructed 3D facial image, the 3D image can be projected onto a 2D plane to obtain the basic facial image frame of the target object in the basic expression state.
[0105] The target facial coefficients can include the geometric and texture features of the target object. Based on these features, a reconstructed 3D facial image can be constructed. Geometric features can be understood as the coordinate information of key points in the 3D mesh structure of the target object, while texture features can be understood as features indicating the texture information of the target object. For geometric features, the positional information of at least one facial key point can be extracted from the target facial coefficients, and this positional information can be converted into geometric features.
[0106] Specifically, after converting the geometric and texture features, the three-dimensional model parameters of the target object are determined based on the geometric and texture features. Based on the three-dimensional model parameters, the three-dimensional object model of the target object can be constructed, which is the reconstructed three-dimensional facial image in the above embodiment. The three-dimensional object model is then projected onto a two-dimensional plane to obtain the basic facial image frame of the target object in its basic expression state.
[0107] 102. Perform feature encoding processing on the basic facial image frame to obtain the basic facial encoded feature map.
[0108] Optionally, in this embodiment, the step "performing feature encoding processing on the basic facial image frame to obtain a basic facial encoding feature map" may include:
[0109] The basic facial image frame is downsampled multiple times to obtain a facial feature map;
[0110] Attention processing is applied to the facial feature map to obtain the basic facial coding feature map.
[0111] Among these methods, attention processing can capture key information in images and improve the accuracy of feature extraction.
[0112] In one specific embodiment, the basic facial image frame can be first convolved, then downsampled for the convolved basic facial image frame, and then convolved again, followed by a second downsampling process to obtain a facial feature map. After obtaining the facial feature map, it is convolved for the first time to obtain a convolved facial feature map, then attention is applied to the convolved facial feature map, followed by a second convolving and attention process to obtain the basic facial encoding feature map.
[0113] 103. Perform feature extraction processing on the audio driving information to obtain the audio features corresponding to the audio driving information.
[0114] Specifically, the audio driving information can be an audio driving sequence, which may include at least one audio frame.
[0115] The step "perform feature extraction processing on the audio driving information to obtain the audio features corresponding to the audio driving information" can specifically include: vectorizing each audio frame in the audio driving information to obtain an initial feature vector corresponding to each audio frame, and then performing dimension mapping on the initial feature vectors of each audio frame to obtain the audio feature vectors corresponding to each audio frame. Dimension mapping can transform the vector dimension of the initial feature vectors, thus ensuring that the vector dimension of the audio feature vectors of each audio frame is consistent.
[0116] In the process of vectorizing each audio frame in the audio driving information, position embedding can be performed on each audio frame to obtain the positional encoding information corresponding to each audio frame, and feature embedding can be performed on each audio frame itself to obtain the audio feature information corresponding to each audio frame. Based on the positional encoding information and audio feature information of each audio frame in the audio driving information, the initial feature vector of each audio frame in the audio driving information is determined. Specifically, the positional encoding information and audio feature information can be concatenated to obtain the initial feature vector. The positional encoding information corresponding to the audio driving information can include the frame number of each audio frame in the audio driving information. Adding positional encoding information during the vectorization process can fuse the temporal features of the audio driving information into the initial feature vector.
[0117] 104. Based on the audio features, perform expression adjustment processing on the basic facial coding feature map to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0118] Specifically, the target facial image frame can be a facial image corresponding to the facial pose (specifically, expression) adjusted from the base facial image frame based on audio-driven information.
[0119] The audio features can include information such as lip movements when the target speaks, as well as the content and volume of the speech. Based on the content and volume, the target's emotional changes can be determined. By adjusting facial expressions based on these audio features, the lip movements in the resulting facial image frames can match the audio-driven information, and the facial expressions in the frames can correspond to the audio-driven information.
[0120] Optionally, in this embodiment, the step "based on the audio features, performing expression adjustment processing on the basic facial coding feature map to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information" may include:
[0121] The audio features and the basic facial coding feature map are subjected to feature fusion processing to obtain initial fused feature information;
[0122] The audio features are injected into the initial fused feature information for feature interaction processing to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0123] There are various ways to fuse audio features and basic facial coding feature maps, and this embodiment does not limit this approach. Specifically, the feature fusion method can be concatenation or weighted operations, etc.
[0124] In one specific embodiment, the step "performing feature fusion processing on the audio features and the basic facial coding feature map to obtain initial fused feature information" may include:
[0125] The audio features are linearly processed to obtain the processed audio features;
[0126] The basic facial coding feature map is normalized to obtain a normalized basic facial coding feature map.
[0127] The processed audio features and the normalized basic facial coding feature map are fused together to obtain initial fused feature information.
[0128] In a specific scenario, the audio features and the basic facial coding feature map can be fused using AdaIN (Adaptive Instance Normalization), as shown in Equation (1):
[0129] AdaIN(x,e)=(1+FC1(e))*Norm(x)+FC2(e) (1)
[0130] Where 'e' represents the audio feature, FC i For a linear layer, Norm(.) represents the normalization function, which can be BatchNorm, LayerNorm, InstanceNorm, GroupNorm, etc. When the input x represents the basic facial encoding feature map, AdaIN(x,e) represents the initial fused feature information obtained after feature fusion processing. Through AdaIN processing, the content of the basic facial encoding feature map and the style of the audio features can be effectively combined.
[0131] Optionally, in this embodiment, the step "injecting the audio features into the initial fused feature information for feature interaction processing to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information" may include:
[0132] The initial fusion feature information is subjected to convolution processing to obtain convolution result information; and the convolution result information is fused with the audio features to obtain target fusion feature information.
[0133] The target fusion feature information is subjected to convolution processing to obtain the target convolution result information;
[0134] The target convolution result information is fused with the basic facial coding feature map to obtain the processed facial coding feature map;
[0135] Based on the audio features, the processed facial encoding feature map is subjected to feature decoding to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0136] There are various ways to fuse convolution result information and audio features, and this embodiment does not limit this. The fusion method can be concatenation or weighted operation, etc. Specifically, convolution result information and audio features can also be fused using AdaIN. Fusion using AdaIN can include:
[0137] The audio features are linearly processed to obtain the processed audio features;
[0138] The convolution result information is normalized to obtain normalized convolution result information;
[0139] The processed audio features and the normalized convolution result information are fused together to obtain the target fused feature information.
[0140] Specifically, you can refer to the above formula (1). When the input x represents the convolution result information, AdaIN(x,e) is the target fusion feature information.
[0141] In this embodiment, the target convolution result information and the basic facial coding feature map are fused. Specifically, the target convolution result information and the basic facial coding feature map are weighted and fused to obtain the processed facial coding feature map.
[0142] The step "based on the audio features, performing feature decoding on the processed facial encoding feature map to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information" may include:
[0143] The processed facial encoding feature map is subjected to feature processing to obtain the target facial feature map;
[0144] The target facial feature map and the audio feature are fused together to obtain the first fused feature information;
[0145] The first fused feature information is subjected to convolution processing to obtain the first convolution result information; and the first convolution result information is fused with the audio feature to obtain the second fused feature information.
[0146] The second fused feature information is convolved to obtain the second convolution result information;
[0147] The second convolution result information is fused with the target facial feature map to obtain the processed target facial feature map;
[0148] The processed target facial feature map is used as a new processed facial coding feature map. The step of performing feature processing on the processed facial coding feature map to obtain the target facial feature map is returned to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0149] The feature processing of the processed facial coding feature map can be either attention processing or upsampling processing; this embodiment does not impose any restrictions on this.
[0150] Upsampling can restore an image to a higher resolution, specifically by restoring the target facial image frame to the same size as the base facial image frame. Upsampling essentially involves image magnification and image interpolation. Interpolation methods can include nearest neighbor interpolation, bilinear interpolation, and cubic convolution interpolation, among others.
[0151] There are various ways to fuse the target facial feature map and audio features, and this embodiment does not limit this. For example, the feature fusion can be a splicing process or a weighted fusion. In some embodiments, the AdaIN method can be used to fuse the target facial feature map and audio features. Referring to the above formula (1), when the input x represents the target facial feature map, AdaIN(x,e) is the first fused feature information.
[0152] There are various ways to fuse the first convolution result information with the audio features, and this embodiment does not limit this. For example, the feature fusion can be a concatenation process or a weighted fusion. In some embodiments, the AdaIN method can be used to fuse the first convolution result information with the audio features. Referring to the above formula (1), when the input x represents the first convolution result information, AdaIN(x,e) is the second fused feature information.
[0153] Optionally, in this embodiment, the step "performing feature encoding processing on the basic facial image frame to obtain a basic facial encoding feature map" may include:
[0154] The basic facial image frame is processed by feature encoding using an image generation model to obtain a basic facial encoded feature map.
[0155] The image generation model can be a neural network model, such as a UNet (U-shaped network) model or a Transformer model, etc. However, it should be understood that the image generation model in this embodiment is not limited to the types listed above.
[0156] It should be noted that the image generation model can be trained from multiple sets of training data. Specifically, the image generation model can be trained by other devices and then provided to the image generation device, or it can be trained by the image generation device itself.
[0157] If the image generation device performs the training itself, then before the step "using the image generation model to perform feature encoding processing on the basic facial image frame to obtain the basic facial encoding feature map", the following steps are also included:
[0158] Acquire training data, which includes basic facial image frame samples of the sample object, target-driven facial image frame samples, and audio driving information samples corresponding to the target-driven facial image frame samples.
[0159] The basic facial image frame samples are processed by feature encoding using an image generation model to obtain the basic facial encoding feature map corresponding to the basic facial image frame samples.
[0160] The audio driving information samples are subjected to feature extraction processing to obtain sample audio features;
[0161] Based on the sample audio features, the basic facial coding feature map is processed for expression adjustment to generate a predicted driven facial image frame of the sample object in the corresponding expression state of the audio driving information sample.
[0162] Based on the target-driven facial image frame samples and the prediction-driven facial image frames, the parameters of the image generation model are adjusted to obtain the trained image generation model.
[0163] Among them, the target-driven facial image frame sample can be regarded as label information, specifically the expected-driven facial image frame corresponding to the audio-driven information sample.
[0164] There are multiple ways to obtain the basic facial image frame sample, the target-driven facial image frame sample, and the audio-driven information sample corresponding to the target-driven facial image frame sample of the sample object, and this embodiment does not limit these methods.
[0165] For example, a speaker's video can be used to train the model. Specifically, a video frame containing the subject's face but not speaking can be extracted from a speech video about the subject as a base facial image frame sample, and a video frame containing the subject's face and speaking can be extracted from the speech video as a target-driven facial image frame sample. The audio information corresponding to the target-driven facial image frame sample one second before and after in the speech video can be used as audio-driven information samples.
[0166] Optionally, in this embodiment, the step "adjusting the parameters of the image generation model based on the target-driven facial image frame samples and the predicted-driven facial image frames to obtain the trained image generation model" may include:
[0167] Calculate the image loss information between the target-driven face image frame sample and the predicted-driven face image frame;
[0168] Extract the target facial feature information corresponding to the target-driven facial image frame sample and the predicted facial feature information corresponding to the predicted facial image frame, respectively; and calculate feature loss information based on the target facial feature information and the predicted facial feature information.
[0169] The target-driven facial image frame samples and the predicted-driven facial image frames are subjected to realism discrimination processing respectively, so as to determine the generative adversarial loss information of the image generation model based on the discrimination results;
[0170] Based on the image loss information, the feature loss information, and the generative adversarial loss information, the parameters of the image generation model are adjusted to obtain the trained image generation model.
[0171] Specifically, the image loss information refers to the image reconstruction loss. In some embodiments, the image reconstruction loss of the image generation model can be determined based on the similarity between the target-driven facial image frame samples and the prediction-driven facial image frames. This similarity can be cosine similarity or histogram similarity measurement; this embodiment does not impose any restrictions on this.
[0172] The feature loss information refers to the loss value at the feature level. Specifically, spatial features are extracted from both the target-driven facial image frame samples and the predicted-driven facial image frames to obtain the first facial spatial features of the target-driven facial image frame samples (i.e., the aforementioned target facial feature information) and the second facial spatial features of the predicted-driven facial image frames (i.e., the aforementioned predicted facial feature information). The similarity between the first and second facial spatial features is then calculated to determine the feature loss information. A higher similarity results in a lower feature loss, and vice versa. This spatial feature extraction can be performed on both the target-driven and predicted-driven facial image frame samples using a pre-trained image feature extraction network.
[0173] Among them, for generating adversarial loss information (denoted as l) adv It can predict the probability that the target-driven facial image frame sample and the driving facial image frame belong to the real driving facial image frame, and determine the generative adversarial loss information of the image generation model based on the probability.
[0174] In this embodiment, the step of "predicting the probability that the target-driven facial image frame sample and the probability that the driven facial image frame belong to the real driven facial image frame respectively, and determining the generative adversarial loss information of the image generation model based on the probability" may include:
[0175] By using a pre-defined discrimination model, the first probability information of whether the target-driven facial image frame sample belongs to the real-driven facial image frame is predicted.
[0176] The preset discrimination model is used to predict the second probability information of whether the predicted driving face image frame belongs to the real driving face image frame.
[0177] Based on the first probability information and the second probability information, the generative adversarial loss information of the image generation model is determined.
[0178] Specifically, the preset discriminant model can be a discriminator D. During training, the target-driven facial image frame samples are real images, and the predicted-driven facial image frames are the generated results of the image generation model. The discriminator needs to determine whether the generated result is fake and whether the real image is real. The image generation model can be regarded as a whole generative network G. During training, the images generated by the generative network G need to be able to fool the discriminator D. That is, the probability that the discriminator D judges the predicted-driven facial image frame generated by the generative network G to be a real-driven facial image frame should be 1.
[0179] The discriminator takes as input either a real image or the output of an image generation model, and its goal is to distinguish the output of the image generation model from the real image as much as possible. The image generation model, on the other hand, aims to deceive the discriminator as much as possible. The image generation model and the discriminator work against each other, constantly adjusting their parameters to obtain a well-trained image generation model.
[0180] In this embodiment, the step "adjusting the parameters of the image generation model according to the image loss information, the feature loss information, and the generative adversarial loss information to obtain the trained image generation model" may include:
[0181] The image loss information, the feature loss information, and the generative adversarial loss information are fused together to obtain the total loss information.
[0182] Based on the total loss information, the parameters of the image generation model are adjusted to obtain the trained image generation model.
[0183] Specifically, the fusion processing of image loss information, feature loss information, and generative adversarial loss information can be weighted fusion, etc.
[0184] The training process for the image generation model involves first calculating the total loss information, then using the backpropagation algorithm to adjust the parameters of the image generation model. The parameters are optimized based on the total loss information, ensuring that the loss value corresponding to the total loss information is less than a preset loss value, thus obtaining a trained image generation model. This preset loss value can be set according to specific requirements; for example, a smaller preset loss value indicates a higher requirement for the accuracy of the image generation model.
[0185] In one specific embodiment, the prediction-driven facial image frames generated by the image generation model during training are denoted as x. r Its label—target-driven facial image frame sample—is denoted as x. h The calculation of the above image loss information, feature loss information, and generative adversarial loss information can be shown in equations (2), (3), and (4), respectively:
[0186] l pix =||x r -x h || 1 (2)
[0187] l per =||φ(x) r )-φ(x h )|| 1 (3)
[0188]
[0189] Where, x r For the face output by the network, x h To train real faces (labels), This indicates that for all data samples x r The average value. pix Represents image loss information, l per Representing feature loss information, l adv This represents the adversarial loss information. φ is the network used to extract the depth features of the image, such as a VGG (Visual Geometry Group) network. D is the discriminator network. The discriminator network structure can be varied, such as a VGG network structure or a PatchGAN structure. The total loss information of the image generation model can be shown in equation (5):
[0190] l = l pix +l per +l adv (5)
[0191] Regarding the acquisition of training data for the image generation model, the image generation method of this application can also construct training data using a pre-trained sample generation model.
[0192] Optionally, in this embodiment, the image generation method may include:
[0193] Acquire a sample generation model, at least one audio-driven information sample, and a base facial image frame sample of the sample object;
[0194] The sample generation model is used to extract facial pose features from the basic facial image frame samples to obtain the basic facial coefficients corresponding to the basic facial image frame samples.
[0195] Temporal features are extracted from the audio-driven information samples to obtain the target expression coefficients of the corresponding facial expression states of the audio-driven information samples.
[0196] Based on the basic facial coefficients and the target expression coefficients, the sample object is subjected to facial reconstruction processing to obtain the target driven facial image frame sample of the sample object in the expression state corresponding to the audio driven information sample;
[0197] Training data is constructed based on the target-driven facial image frame samples, the audio-driven information samples, and the base facial image frame samples.
[0198] Specifically, the sample generation model is a 3D face-driven model, which can be used as a tool for generating training data. It generates training data by constructing audio and corresponding driving results, resulting in paired training data such as (silent face, driving audio i) → (driving face i), where the silent face is the base facial image frame sample. Multiple sets of paired training data can be obtained based on different driving audios. Driving audio can be obtained by collecting other spoken audio or by generating audio from text.
[0199] This involves extracting facial pose features from basic facial image frame samples. Specifically, this can involve extracting spatial features from the basic facial image frame samples to obtain their facial spatial features. These facial spatial features can include three-dimensional (3D) facial coefficients corresponding to the basic facial image frame samples, such as identity information, lighting, texture, expression, and pose. Based on these facial coefficients, the face of the basic facial image frame samples can be reconstructed.
[0200] Specifically, the extraction of spatial features from basic facial image frame samples can involve performing convolution and pooling processes on the basic facial image frame samples, and this embodiment does not impose any limitations on this.
[0201] Specifically, the process for generating training data for this 3D facial driving model can be as follows:
[0202] 1. For the target's driving audio, the audio is converted into audio features, which are then input into the network module that converts audio features into 3D expression coefficients to obtain the target's 3D expression coefficients;
[0203] 2. For the original input face, the 3D coefficients of the face will be calculated. These coefficients include: shape coefficient, expression coefficient, rotation coefficient, camera coefficient, etc. Among them, the rotation coefficient can specifically refer to the rotation angle of the head, and the camera coefficient refers to the distance relationship between the head and the world camera.
[0204] 3. Replace the original 3D expression coefficients with the target's 3D expression coefficients, and obtain the target's 3D face rendering image through 3D rendering;
[0205] 4. Render the target's 3D face using a deep network to obtain a realistic, high-definition driven face.
[0206] Optionally, in this embodiment, the image generation method may include:
[0207] Retrieve preset emoji guide text information;
[0208] The preset emoticon guidance text information is subjected to audio conversion processing in multiple languages to obtain audio driving information samples in the multiple languages.
[0209] In this embodiment, in order to make the constructed audio have richer lip-syncing, text-generated audio can be used. The text can cover words with multiple pronunciations and various languages, such as Chinese and English, so that the training data is more fully distributed and the learned model has better generalization ability.
[0210] Optionally, in this embodiment, the step "performing facial reconstruction processing on the sample object based on the basic facial coefficients and the target expression coefficients to obtain a target driven facial image frame sample of the sample object in the expression state corresponding to the audio driven information sample" may include:
[0211] Based on the target expression coefficient, the expression coefficients in the basic facial coefficients are replaced to obtain the target facial coefficients;
[0212] The sample object is reconstructed based on the target facial coefficients to obtain a target driven facial image frame sample of the sample object in the expression state corresponding to the audio driven information sample.
[0213] The basic facial coefficients can include basic expression coefficients and basic facial morphology coefficients. Specifically, the basic facial coefficients can include identity information, lighting, texture, posture, eye expression, mouth shape, eyebrow shape, etc. Among them, eye expression, mouth shape, and eyebrow shape can belong to the basic expression coefficients, while the basic facial morphology coefficients can be other coefficients in the basic facial coefficients besides the basic expression coefficients. For example, identity information, lighting, texture, and posture can belong to the basic facial morphology coefficients.
[0214] In this embodiment, after obtaining the basic facial coefficients, coefficients related to the expression of the sample object can be selected from the basic facial coefficients, namely, basic expression coefficients. For example, basic expression coefficients such as eye expression, mouth shape, and eyebrow shape can be extracted from the basic facial coefficients. Then, the target expression coefficients corresponding to the audio-driven information sample are used to replace these basic expression coefficients, thereby obtaining the replaced target facial coefficients. The replaced target facial coefficients include target expression coefficients and basic facial morphology coefficients.
[0215] Optionally, in this embodiment, the step "performing facial reconstruction processing on the sample object based on the target facial coefficients to obtain a target driven facial image frame sample of the sample object in the expression state corresponding to the audio driven information sample" may include:
[0216] The sample object is reconstructed based on the target facial coefficients to obtain the reconstructed three-dimensional facial image of the sample object corresponding to the expression state of the audio-driven information sample.
[0217] The reconstructed 3D facial image is rendered and mapped to obtain the target driven facial image frame sample of the sample object in the expression state corresponding to the audio driven information sample.
[0218] Among them, facial reconstruction of sample objects through a three-dimensional facial driving model, also known as 3D reconstruction, can represent the input two-dimensional facial image with a 3D mesh (three-dimensional mesh model). The 3D mesh can contain the vertex coordinates and colors of the three-dimensional mesh structure.
[0219] The texture and lighting of the reconstructed 3D facial image can be derived from the base facial image frame sample, and the pose and expression of the reconstructed 3D facial image can be derived from the audio-driven information sample. By performing rendering and mapping processing on the reconstructed 3D facial image, the 3D image can be projected onto a 2D plane to obtain the target driven facial image frame sample of the sample object in the corresponding expression state of the audio-driven information sample.
[0220] The target facial coefficients can include the geometric and texture features of the sample object. Based on these features, a reconstructed 3D facial image can be constructed. Geometric features can be understood as the coordinate information of key points in the 3D mesh structure of the sample object, while texture features can be understood as features indicating the texture information of the sample object. For geometric features, the positional information of at least one facial key point can be extracted from the target facial coefficients, and this positional information can be converted into geometric features.
[0221] Specifically, after converting the geometric and texture features, the three-dimensional model parameters of the sample object are determined based on the geometric and texture features. Based on the three-dimensional model parameters, the three-dimensional object model of the sample object can be constructed, which is the reconstructed three-dimensional facial image in the above embodiment. The three-dimensional object model is then projected onto a two-dimensional plane to obtain the target driven facial image frame sample of the sample object in the expression state corresponding to the audio driven information sample.
[0222] Optionally, in this embodiment, the sample generation model includes a mapping module;
[0223] The step "extracting temporal features from the audio-driven information samples to obtain the target expression coefficients corresponding to the facial expression states of the audio-driven information samples" may include:
[0224] The mapping module extracts temporal features from the audio driving information samples to obtain the target expression coefficients corresponding to the facial expression states of the audio driving information samples.
[0225] Before performing temporal feature extraction on the audio-driven information sample through the mapping module to obtain the target expression coefficient corresponding to the expression state of the audio-driven information sample, the method further includes:
[0226] Acquire training sample data, which includes sample audio information and expression coefficient label information corresponding to the sample audio information;
[0227] The temporal features of the sample audio information are extracted using the mapping module to obtain audio feature information.
[0228] Based on the audio feature information, predict the actual facial expression coefficient corresponding to the sample audio information;
[0229] Based on the actual facial expression coefficients and the facial expression coefficient label information, the parameters of the mapping module are adjusted to obtain the trained mapping module.
[0230] In training the sample generation model, this embodiment can first learn the mapping from different audio to different facial expression coefficients, so that for any audio input, a corresponding facial expression coefficient can be obtained. Then, it learns the mapping from the 3D rendered image to the 2D image (real face image), which can also be learned separately.
[0231] The training process for the mapping module involves first calculating the loss value between the actual facial expression coefficients and the label information of those coefficients. Then, the parameters of the mapping module are adjusted using the backpropagation algorithm. Based on the loss value, the parameters are optimized to ensure that the loss value is less than a preset loss value, resulting in a trained mapping module. This preset loss value can be set according to actual conditions.
[0232] Based on the image generation method provided in this application, changes in the face within any facial image frame can be flexibly driven according to given audio to generate a target facial image frame whose lip movements and other features are consistent with the audio content. This solution can realize voice-driven portraiture, reflecting arbitrary audio onto any human face.
[0233] This application provides a lightweight lip-sync driven model design and training scheme. The image generation model can modify the lip movements of the original person in the silent face based on the silent face and audio provided by the user, so that the modified lip movements match the audio content. The image generation model is directly driven by the silent face, without the need for 3D face reconstruction, which greatly improves the model inference efficiency. This significantly increases the concurrency on cloud GPU (Graphics Processing Unit) servers, specifically, the concurrency can be increased to more than 20. It can also be deployed on CPU (Central Processing Unit) servers, and can also be deployed on mobile devices to perform real-time calculations on the mobile device.
[0234] The 3D face-driving model of this application requires face reconstruction during image generation. It can be used as a tool to generate training data for image generation models, generating corresponding silent faces as input, and then using predefined audio to generate the face image of the target as a label. This allows training a model that can directly obtain the corresponding driven face by only needing to input a silent face and audio, which greatly improves the inference efficiency of the model.
[0235] In some embodiments, during the online application of the image generation model, if the user provides a non-silent face, a 3D face-driving model can be used to generate a corresponding silent face based on the non-silent face as input to the image generation model.
[0236] like Figure 1c The diagram illustrates the lightweight lip-sync driving framework of this application. The lightweight lip-sync driving framework only requires a silent face and the audio to be driven as input, which are then fed into a deep network (i.e., the image generation model described above) to obtain the driven image, eliminating the need for 3D dependencies. Here, a silent face refers to a face without mouth slurring. Specifically, directly changing a speaking expression to another speaking expression is very difficult for the model. Therefore, the original face in a non-silent state can be converted into a silent state first. The model only needs to map from the silent state to the corresponding speaking state, thus significantly reducing the model's learning cost and further shrinking the model's parameter size, achieving a lightweight standard. During the training of the image generation model, the model parameters can be adjusted based on the loss between the generated driven face and the desired target face.
[0237] like Figure 1d The diagram illustrates the silent face generation process. Using a 3D facial driving model, the 3D facial coefficients of the original face are obtained. Then, the 3D expression coefficients of the original face are determined from these coefficients. These 3D expression coefficients are then converted into silent 3D expression coefficients (i.e., the preset basic expression coefficients in the above embodiment), thus obtaining the target facial coefficients. These target facial coefficients include the silent expression coefficients and other original 3D facial coefficients besides the 3D expression coefficients. Based on these target facial coefficients, 3D rendering is performed to obtain a 3D rendered face image (i.e., the reconstructed 3D facial image). Then, deep network rendering is performed on the reconstructed 3D facial image to obtain a silent face (i.e., the basic facial image frame in the basic expression state). In this way, the facial expression of each frame in the training video can also be modified to a silent expression, thus turning the training video into a silent video, which serves as the input video for the lightweight model (image generation model).
[0238] For training image generation models, after obtaining the silent input video, it is also necessary to construct training label data, i.e., lip-sync data corresponding to the audio. Firstly, the original video itself is a label. Using the audio of the original video as training audio and the lip-sync data of the original video as the desired target face label, we can construct training data from (silent face, original audio) to (original face). However, with only this one set of data, the network may not fully perceive the relationship between audio and lip-sync when learning the mapping. It might predict the lip-sync based on certain features of the silent face, such as the mouth being open when looking up. Learning the lip-sync through such dependencies requires more training data.
[0239] like Figure 1e The diagram shown illustrates the training data generation process, illustrating how to obtain more training data. Specifically, training data is generated by constructing audio and corresponding driving results using a trained 3D facial driving model, resulting in paired training data such as (silent face, driving audio i) → (driving face i). Multiple sets of paired training data can be obtained based on different driving audios. Driving audio can be obtained by collecting other speech audio or by generating audio from text.
[0240] like Figure 1f The diagram illustrates the image generation model, which includes multiple residual convolutional blocks (ResBlock), a downsampling module, an attention-aware module (AttnBlock), an instance normalization-convolutional block (AdaIN-ResBlock), and an upsampling module. The downsampling module can be replaced by a convolution with a stride of 2, and the attention-aware module is specifically a convolutional self-attention module.
[0241] The audio features can be injected via AdaIN-ResBlock. Specifically, the audio is first converted into features by an MLP (Multilayer Perceptron) and then injected into the features of the ResBlock via AdaIN. The operation of AdaIN is described in equation (1) above.
[0242] The calculation method of AdaIN-ResBlock is shown in equation (6):
[0243] y=Conv2(AdaIN2(Conv1(AdaIN1(x,e)),e))+x (6)
[0244] Where x represents the input feature, e represents the injected audio feature, and Conv i y represents a 2D convolutional layer and y represents the module's output. The AdaIN-ResBlock module contains two AdaIN operations and two convolutional computations.
[0245] It should be noted that the network structure of an image generation model does not necessarily have to be like... Figure 1f The design is the same; either the basic structure of UNet or the Transformer structure can be used. This embodiment does not impose any restrictions on either.
[0246] In addition, the method of injecting audio features is not limited to AdaIN. CrossAttention can also be used. Alternatively, audio features can be copied from one-dimensional features to two-dimensional features according to the image size, concatenated with image features, and then fused through convolution.
[0247] This application presents a lightweight lip-sync driven model design that can directly output the corresponding driven face by inputting target audio and a silent face. This application allows for training with different network parameters for different devices, enabling the lip-sync driven algorithm to run on various devices, including GPU servers, CPU servers, and even mobile devices. For example, the resolution of the input image frame on a mobile device can be set to 160x160, while on the server, it can be set to 320x320 for better image clarity. Furthermore, network parameters can be adjusted, such as 8M / 4M / 2M, depending on the device's computing power. Generally, higher parameters result in better image clarity.
[0248] As can be seen from the above, this embodiment can obtain a basic facial image frame of the target object in a basic expression state, and obtain audio driving information for image generation; perform feature encoding processing on the basic facial image frame to obtain a basic facial encoding feature map; perform feature extraction processing on the audio driving information to obtain the audio features corresponding to the audio driving information; and perform expression adjustment processing on the basic facial encoding feature map based on the audio features to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0249] This application can utilize some facial pose details contained in audio features to perform facial expression adjustment processing on basic facial image frames, thereby obtaining target facial image frames corresponding to audio-driven information. This is beneficial to improving the generation efficiency and accuracy of target facial image frames. Moreover, this application generates images based on basic facial image frames under basic expression states. Compared with directly adjusting complex expression states, this expression adjustment under basic expression states can significantly reduce the amount of computation and further improve image generation efficiency.
[0250] Based on the method described in the preceding embodiments, the following will provide a more detailed explanation by taking the specific integration of the image generation device into a server as an example.
[0251] This application provides an image generation method, such as... Figure 2 As shown, the specific process of this image generation method can be as follows:
[0252] 201. The server obtains the initial facial image frame of the target object in its initial expression state and obtains the audio driving information used for image generation.
[0253] The target object can be an object whose facial posture needs to be adjusted. Facial posture specifically refers to facial expressions, such as mouth shape and eye contact; this embodiment does not impose such limitations.
[0254] The initial facial expression state can be the facial posture corresponding to the target object when speaking, such as a facial posture with the mouth open.
[0255] Specifically, the facial image frames of the target object initially acquired may not be in a basic expression state. In this case, the expression state of the initial facial image frames needs to be adjusted to the basic expression state to reduce subsequent computational costs. If the facial image frames initially acquired are in a basic expression state, no adjustment is required.
[0256] 202. The server determines the preset basic expression coefficient corresponding to the basic expression state.
[0257] Specifically, the basic facial expression state can be the facial posture corresponding to not speaking, such as a facial posture with the mouth closed.
[0258] Among them, the preset basic expression coefficient corresponding to the basic expression state can be a preset value, which can be determined based on expert experience. For example, all coefficients in the preset basic expression coefficient can be 0.
[0259] 203. The server extracts facial pose features from the initial facial image frame to obtain the initial facial coefficients corresponding to the initial facial image frame.
[0260] The step "extracting facial pose features from the initial facial image frame to obtain the initial facial coefficients corresponding to the initial facial image frame" can specifically involve extracting spatial features from the initial facial image frame to obtain its facial spatial features. These facial spatial features can specifically include three-dimensional (3D) facial coefficients corresponding to the initial facial image frame, such as identity information, lighting, texture, expression, and pose. Expression information can include eye contact, lip shape, and eyebrow shape. Based on these facial coefficients, the face of the initial facial image frame can be reconstructed. Because 3D has good decoupling properties for facial information, the shape of the face and its expression can be decoupled.
[0261] Specifically, the spatial feature extraction of the initial facial image frame can be performed by convolution and pooling processes on the initial facial image frame, and this embodiment does not impose any restrictions on this.
[0262] 204. The server performs facial reconstruction processing on the target object based on the preset basic expression coefficients and the initial facial coefficients to obtain the basic facial image frame of the target object in the basic expression state.
[0263] Optionally, in this embodiment, the initial facial coefficients include initial expression coefficients and initial facial morphology coefficients;
[0264] The step "based on the preset basic expression coefficients and the initial facial coefficients, perform facial reconstruction processing on the target object to obtain a basic facial image frame of the target object in a basic expression state" may include:
[0265] Based on the preset basic expression coefficients, the initial expression coefficients in the initial facial coefficients are replaced to obtain the target facial coefficients, which include the preset basic expression coefficients and the initial facial morphology coefficients.
[0266] The target object is reconstructed based on the target facial coefficients to obtain a basic facial image frame of the target object in a basic expression state.
[0267] Optionally, in this embodiment, the step "performing facial reconstruction processing on the target object based on the target facial coefficients to obtain a basic facial image frame of the target object in a basic expression state" may include:
[0268] Based on the target facial coefficients, the target object is subjected to facial reconstruction processing to obtain the reconstructed three-dimensional facial image of the target object in the basic expression state;
[0269] The reconstructed 3D facial image is rendered and mapped to obtain the base facial image frame of the target object in its basic expression state.
[0270] Specifically, a 3D facial reconstruction process can be performed on the target object based on the target facial coefficients using a 3D facial driving model to obtain a reconstructed 3D facial image. This facial reconstruction process is also known as 3D reconstruction. 3D reconstruction can represent the input 2D facial image using a 3D mesh (3D mesh model). The 3D mesh can contain the vertex coordinates and colors of the 3D mesh structure.
[0271] 205. The server performs feature encoding processing on the basic facial image frame to obtain a basic facial encoded feature map.
[0272] Optionally, in this embodiment, the step "performing feature encoding processing on the basic facial image frame to obtain a basic facial encoding feature map" may include:
[0273] The basic facial image frame is downsampled multiple times to obtain a facial feature map;
[0274] Attention processing is applied to the facial feature map to obtain the basic facial coding feature map.
[0275] Among these methods, attention processing can capture key information in images and improve the accuracy of feature extraction.
[0276] 206. The server performs feature extraction processing on the audio driving information to obtain the audio features corresponding to the audio driving information.
[0277] 207. Based on the audio features, the server performs expression adjustment processing on the basic facial coding feature map to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0278] Specifically, the target facial image frame can be a facial image corresponding to the facial pose (specifically, expression) adjusted from the base facial image frame based on audio-driven information.
[0279] The audio features can include information such as lip movements when the target speaks, as well as the content and volume of the speech. Based on the content and volume, the target's emotional changes can be determined. By adjusting facial expressions based on these audio features, the lip movements in the resulting facial image frames can match the audio-driven information, and the facial expressions in the frames can correspond to the audio-driven information.
[0280] Optionally, in this embodiment, the step "based on the audio features, performing expression adjustment processing on the basic facial coding feature map to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information" may include:
[0281] The audio features and the basic facial coding feature map are subjected to feature fusion processing to obtain initial fused feature information;
[0282] The audio features are injected into the initial fused feature information for feature interaction processing to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0283] There are various ways to fuse audio features and basic facial coding feature maps, and this embodiment does not limit this approach. Specifically, the feature fusion method can be concatenation or weighted operations, etc.
[0284] In one specific embodiment, the step "performing feature fusion processing on the audio features and the basic facial coding feature map to obtain initial fused feature information" may include:
[0285] The audio features are linearly processed to obtain the processed audio features;
[0286] The basic facial coding feature map is normalized to obtain a normalized basic facial coding feature map.
[0287] The processed audio features and the normalized basic facial coding feature map are fused together to obtain initial fused feature information.
[0288] Optionally, in this embodiment, the step "injecting the audio features into the initial fused feature information for feature interaction processing to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information" may include:
[0289] The initial fusion feature information is subjected to convolution processing to obtain convolution result information; and the convolution result information is fused with the audio features to obtain target fusion feature information.
[0290] The target fusion feature information is subjected to convolution processing to obtain the target convolution result information;
[0291] The target convolution result information is fused with the basic facial coding feature map to obtain the processed facial coding feature map;
[0292] Based on the audio features, the processed facial encoding feature map is subjected to feature decoding to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0293] As can be seen from the above, this embodiment can obtain the initial facial image frame of the target object in the initial expression state through the server, and obtain the audio driving information for image generation; determine the preset basic expression coefficients corresponding to the basic expression state; extract facial pose features from the initial facial image frame to obtain the initial facial coefficients corresponding to the initial facial image frame; perform facial reconstruction processing on the target object according to the preset basic expression coefficients and the initial facial coefficients to obtain the basic facial image frame of the target object in the basic expression state; perform feature encoding processing on the basic facial image frame to obtain the basic facial encoding feature map; perform feature extraction processing on the audio driving information to obtain the audio features corresponding to the audio driving information; and perform expression adjustment processing on the basic facial encoding feature map based on the audio features to generate the target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0294] This application can utilize some facial pose details contained in audio features to perform facial expression adjustment processing on basic facial image frames, thereby obtaining target facial image frames corresponding to audio-driven information. This is beneficial to improving the generation efficiency and accuracy of target facial image frames. Moreover, this application generates images based on basic facial image frames under basic expression states. Compared with directly adjusting complex expression states, this expression adjustment under basic expression states can significantly reduce the amount of computation and further improve image generation efficiency.
[0295] To better implement the above methods, embodiments of this application also provide an image generation apparatus, such as... Figure 3 As shown, the image generation device may include an acquisition unit 301, an encoding unit 302, a feature extraction unit 303, and a generation unit 304, as follows:
[0296] (1) Obtain unit 301;
[0297] The acquisition unit is used to acquire the base facial image frame of the target object in the basic expression state, and to acquire the audio driving information for image generation.
[0298] Optionally, in some embodiments of this application, the acquisition unit may include an acquisition subunit, an expression coefficient determination subunit, a pose feature extraction subunit, and a reconstruction subunit, as follows:
[0299] The acquisition subunit is used to acquire the initial facial image frame of the target object in its initial expression state;
[0300] The expression coefficient determination subunit is used to determine the preset basic expression coefficient corresponding to the basic expression state;
[0301] The pose feature extraction subunit is used to extract facial pose features from the initial facial image frame to obtain the initial facial coefficients corresponding to the initial facial image frame.
[0302] The reconstruction subunit is used to perform facial reconstruction processing on the target object based on the preset basic expression coefficients and the initial facial coefficients to obtain the basic facial image frame of the target object in the basic expression state.
[0303] Optionally, in some embodiments of this application, the initial facial coefficients include initial expression coefficients and initial facial morphology coefficients;
[0304] The reconstruction subunit can be specifically used to replace the initial expression coefficients in the initial facial coefficients based on the preset basic expression coefficients to obtain target facial coefficients, wherein the target facial coefficients include the preset basic expression coefficients and the initial facial morphology coefficients; and to perform facial reconstruction processing on the target object according to the target facial coefficients to obtain the basic facial image frame of the target object in the basic expression state.
[0305] Optionally, in some embodiments of this application, the step "performing facial reconstruction processing on the target object based on the target facial coefficients to obtain a basic facial image frame of the target object in a basic expression state" may include:
[0306] Based on the target facial coefficients, the target object is subjected to facial reconstruction processing to obtain the reconstructed three-dimensional facial image of the target object in the basic expression state;
[0307] The reconstructed 3D facial image is rendered and mapped to obtain the base facial image frame of the target object in its basic expression state.
[0308] (2) Encoding unit 302;
[0309] The encoding unit is used to perform feature encoding processing on the basic facial image frame to obtain a basic facial encoded feature map.
[0310] Optionally, in some embodiments of this application, the encoding unit may include a downsampling subunit and an attention processing subunit, as follows:
[0311] The downsampling subunit is used to perform multiple downsampling processes on the basic facial image frame to obtain a facial feature map;
[0312] The attention processing subunit is used to perform attention processing on the facial feature map to obtain the basic facial coding feature map.
[0313] (3) Feature extraction unit 303;
[0314] The feature extraction unit is used to perform feature extraction processing on the audio driving information to obtain the audio features corresponding to the audio driving information.
[0315] (4) Generation unit 304;
[0316] The generation unit is used to perform expression adjustment processing on the basic facial coding feature map based on the audio features, so as to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0317] Optionally, in some embodiments of this application, the generation unit may include a feature fusion subunit and a generation subunit, as follows:
[0318] The feature fusion subunit is used to perform feature fusion processing on the audio features and the basic facial coding feature map to obtain initial fused feature information;
[0319] A generation subunit is used to inject the audio features into the initial fused feature information for feature interaction processing, so as to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0320] Optionally, in some embodiments of this application, the generation subunit may specifically be used to perform convolution processing on the initial fusion feature information to obtain convolution result information; and to perform fusion processing on the convolution result information and the audio features to obtain target fusion feature information; to perform convolution processing on the target fusion feature information to obtain target convolution result information; to perform fusion processing on the target convolution result information and the basic facial coding feature map to obtain a processed facial coding feature map; and to perform feature decoding processing on the processed facial coding feature map according to the audio features to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0321] Optionally, in some embodiments of this application, the encoding unit may be used to perform feature encoding processing on the basic facial image frame through an image generation model to obtain a basic facial encoded feature map.
[0322] Optionally, in some embodiments of this application, the image generation apparatus may further include a training unit for training the image generation model; specifically, the training unit may include a training data acquisition subunit, a feature encoding subunit, a feature extraction subunit, an expression adjustment subunit, and a parameter adjustment subunit, as follows:
[0323] The training data acquisition subunit is used to acquire training data, which includes basic facial image frame samples of the sample object, target-driven facial image frame samples, and audio driving information samples corresponding to the target-driven facial image frame samples.
[0324] The feature encoding subunit is used to perform feature encoding processing on the basic facial image frame samples through an image generation model to obtain the basic facial encoding feature map corresponding to the basic facial image frame samples.
[0325] The feature extraction subunit is used to perform feature extraction processing on the audio driving information sample to obtain sample audio features;
[0326] The expression adjustment subunit is used to perform expression adjustment processing on the basic facial coding feature map based on the sample audio features, so as to generate a predicted driven facial image frame of the sample object in the corresponding expression state of the audio driving information sample.
[0327] The parameter adjustment subunit is used to adjust the parameters of the image generation model based on the target-driven facial image frame samples and the prediction-driven facial image frames to obtain the trained image generation model.
[0328] Optionally, in some embodiments of this application, the parameter adjustment subunit may specifically be used to calculate image loss information between the target-driven facial image frame sample and the prediction-driven facial image frame; extract target facial feature information corresponding to the target-driven facial image frame sample and prediction facial feature information corresponding to the prediction-driven facial image frame; calculate feature loss information based on the target facial feature information and the prediction facial feature information; perform realism discrimination processing on the target-driven facial image frame sample and the prediction-driven facial image frame respectively, so as to determine the generative adversarial loss information of the image generation model based on the discrimination result; and adjust the parameters of the image generation model based on the image loss information, the feature loss information and the generative adversarial loss information to obtain the trained image generation model.
[0329] Optionally, in some embodiments of this application, the image generation apparatus may further include a training data construction unit, as follows:
[0330] The training data construction unit is used to acquire a sample generation model, at least one audio-driven information sample, and a basic facial image frame sample of the sample object; extract facial pose features from the basic facial image frame sample using the sample generation model to obtain basic facial coefficients corresponding to the basic facial image frame sample; extract temporal features from the audio-driven information sample to obtain target expression coefficients corresponding to the expression state of the audio-driven information sample; perform facial reconstruction processing on the sample object based on the basic facial coefficients and the target expression coefficients to obtain target driven facial image frame samples of the sample object in the expression state corresponding to the audio-driven information sample; and construct training data based on the target driven facial image frame samples, the audio-driven information sample, and the basic facial image frame samples.
[0331] Optionally, in some embodiments of this application, the image generation apparatus may further include an audio sample generation unit, as follows:
[0332] The audio sample generation unit is used to acquire preset facial expression guidance text information; and to perform audio conversion processing on the preset facial expression guidance text information in multiple languages to obtain audio driving information samples in the multiple languages.
[0333] Optionally, in some embodiments of this application, the step "performing facial reconstruction processing on the sample object based on the basic facial coefficients and the target expression coefficients to obtain a target driven facial image frame sample of the sample object in the expression state corresponding to the audio-driven information sample" may include:
[0334] Based on the target expression coefficient, the expression coefficients in the basic facial coefficients are replaced to obtain the target facial coefficients;
[0335] The sample object is reconstructed based on the target facial coefficients to obtain a target driven facial image frame sample of the sample object in the expression state corresponding to the audio driven information sample.
[0336] As can be seen from the above, in this embodiment, the acquisition unit 301 can acquire a basic facial image frame of the target object in a basic expression state and acquire audio driving information for image generation; the encoding unit 302 performs feature encoding processing on the basic facial image frame to obtain a basic facial encoding feature map; the feature extraction unit 303 performs feature extraction processing on the audio driving information to obtain the audio features corresponding to the audio driving information; and the generation unit 304 performs expression adjustment processing on the basic facial encoding feature map based on the audio features to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0337] This application can utilize some facial pose details contained in audio features to perform facial expression adjustment processing on basic facial image frames, thereby obtaining target facial image frames corresponding to audio-driven information. This is beneficial to improving the generation efficiency and accuracy of target facial image frames. Moreover, this application generates images based on basic facial image frames under basic expression states. Compared with directly adjusting complex expression states, this expression adjustment under basic expression states can significantly reduce the amount of computation and further improve image generation efficiency.
[0338] This application also provides an electronic device, such as... Figure 4 The diagram shows a structural schematic of an electronic device involved in an embodiment of this application. This electronic device can be a terminal or a server, specifically:
[0339] The electronic device may include components such as a processor 401 with one or more processing cores, a memory 402 with one or more computer-readable storage media, a power supply 403, and an input unit 404. Those skilled in the art will understand that... Figure 4 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0340] The processor 401 is the control center of the electronic device, connecting various parts of the device via various interfaces and lines. It executes software programs and / or modules stored in the memory 402, and calls data stored in the memory 402, to perform various functions and process data. Optionally, the processor 401 may include one or more processing cores; preferably, the processor 401 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 401.
[0341] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 402 may also include a memory controller to provide the processor 401 with access to the memory 402.
[0342] The electronic device also includes a power supply 403 that supplies power to the various components. Preferably, the power supply 403 can be logically connected to the processor 401 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 403 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0343] The electronic device may also include an input unit 404, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0344] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 401 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 402 according to the following instructions, and the processor 401 runs the applications stored in the memory 402 to realize various functions, as follows:
[0345] A basic facial image frame of the target object in a basic expression state is obtained, and audio driving information for image generation is obtained; the basic facial image frame is subjected to feature encoding processing to obtain a basic facial encoding feature map; the audio driving information is subjected to feature extraction processing to obtain audio features corresponding to the audio driving information; based on the audio features, the basic facial encoding feature map is subjected to expression adjustment processing to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0346] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0347] As can be seen from the above, this embodiment can obtain a basic facial image frame of the target object in a basic expression state, and obtain audio driving information for image generation; perform feature encoding processing on the basic facial image frame to obtain a basic facial encoding feature map; perform feature extraction processing on the audio driving information to obtain the audio features corresponding to the audio driving information; and perform expression adjustment processing on the basic facial encoding feature map based on the audio features to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0348] This application can utilize some facial pose details contained in audio features to perform facial expression adjustment processing on basic facial image frames, thereby obtaining target facial image frames corresponding to audio-driven information. This is beneficial to improving the generation efficiency and accuracy of target facial image frames. Moreover, this application generates images based on basic facial image frames under basic expression states. Compared with directly adjusting complex expression states, this expression adjustment under basic expression states can significantly reduce the amount of computation and further improve image generation efficiency.
[0349] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0350] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the image generation methods provided in embodiments of this application. For example, the instructions can execute the following steps:
[0351] A basic facial image frame of the target object in a basic expression state is obtained, and audio driving information for image generation is obtained; the basic facial image frame is subjected to feature encoding processing to obtain a basic facial encoding feature map; the audio driving information is subjected to feature extraction processing to obtain audio features corresponding to the audio driving information; based on the audio features, the basic facial encoding feature map is subjected to expression adjustment processing to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
[0352] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0353] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0354] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the image generation methods provided in the embodiments of this application, the beneficial effects that any of the image generation methods provided in the embodiments of this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0355] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations of the image generation aspect described above.
[0356] The above provides a detailed description of an image generation method and related equipment provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An image generation method, characterized in that, include: Obtain the base facial image frame of the target object in its basic expression state, and obtain the audio driving information used for image generation; The basic facial image frame is subjected to feature encoding processing to obtain a basic facial encoded feature map; The audio driving information is subjected to feature extraction processing to obtain the audio features corresponding to the audio driving information; Based on the audio features, the basic facial coding feature map is processed to adjust the facial expression, so as to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
2. The method according to claim 1, characterized in that, The acquisition of the base facial image frame of the target object in its basic expression state includes: Obtain the initial facial image frame of the target object in its initial expression state; Determine the preset basic expression coefficients corresponding to the basic expression states; Facial pose features are extracted from the initial facial image frame to obtain the initial facial coefficients corresponding to the initial facial image frame; Based on the preset basic expression coefficients and the initial facial coefficients, the target object is subjected to facial reconstruction processing to obtain the basic facial image frame of the target object in the basic expression state.
3. The method according to claim 2, characterized in that, The initial facial coefficients include initial expression coefficients and initial facial morphology coefficients; The step of performing facial reconstruction processing on the target object based on the preset basic expression coefficients and the initial facial coefficients to obtain a basic facial image frame of the target object in a basic expression state includes: Based on the preset basic expression coefficients, the initial expression coefficients in the initial facial coefficients are replaced to obtain the target facial coefficients, which include the preset basic expression coefficients and the initial facial morphology coefficients. The target object is reconstructed based on the target facial coefficients to obtain a basic facial image frame of the target object in a basic expression state.
4. The method according to claim 3, characterized in that, The step of performing facial reconstruction processing on the target object based on the target facial coefficients to obtain a basic facial image frame of the target object in a basic expression state includes: Based on the target facial coefficients, the target object is subjected to facial reconstruction processing to obtain the reconstructed three-dimensional facial image of the target object in the basic expression state; The reconstructed 3D facial image is rendered and mapped to obtain the base facial image frame of the target object in its basic expression state.
5. The method according to claim 1, characterized in that, The step of performing feature encoding processing on the basic facial image frame to obtain a basic facial encoded feature map includes: The basic facial image frame is downsampled multiple times to obtain a facial feature map; Attention processing is applied to the facial feature map to obtain the basic facial coding feature map.
6. The method according to claim 1, characterized in that, The step of performing expression adjustment processing on the basic facial coding feature map based on the audio features to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information includes: The audio features and the basic facial coding feature map are subjected to feature fusion processing to obtain initial fused feature information; The audio features are injected into the initial fused feature information for feature interaction processing to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
7. The method according to claim 6, characterized in that, The step of injecting the audio features into the initial fused feature information for feature interaction processing to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information includes: The initial fusion feature information is subjected to convolution processing to obtain convolution result information; and the convolution result information is fused with the audio features to obtain target fusion feature information. The target fusion feature information is subjected to convolution processing to obtain the target convolution result information; The target convolution result information is fused with the basic facial coding feature map to obtain the processed facial coding feature map; Based on the audio features, the processed facial encoding feature map is subjected to feature decoding to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
8. The method according to claim 1, characterized in that, The step of performing feature encoding processing on the basic facial image frame to obtain a basic facial encoded feature map includes: The basic facial image frame is processed by feature encoding using an image generation model to obtain a basic facial encoded feature map. Before performing feature encoding processing on the base facial image frame using an image generation model to obtain the base facial encoded feature map, the method further includes: Acquire training data, which includes basic facial image frame samples of the sample object, target-driven facial image frame samples, and audio driving information samples corresponding to the target-driven facial image frame samples. The basic facial image frame samples are processed by feature encoding using an image generation model to obtain the basic facial encoding feature map corresponding to the basic facial image frame samples. The audio driving information samples are subjected to feature extraction processing to obtain sample audio features; Based on the sample audio features, the basic facial coding feature map is processed for expression adjustment to generate a predicted driven facial image frame of the sample object in the corresponding expression state of the audio driving information sample. Based on the target-driven facial image frame samples and the prediction-driven facial image frames, the parameters of the image generation model are adjusted to obtain the trained image generation model.
9. The method according to claim 8, characterized in that, The step of adjusting the parameters of the image generation model based on the target-driven facial image frame samples and the predicted-driven facial image frames to obtain the trained image generation model includes: Calculate the image loss information between the target-driven face image frame sample and the predicted-driven face image frame; Extract the target facial feature information corresponding to the target-driven facial image frame sample and the predicted facial feature information corresponding to the predicted facial image frame, respectively; and calculate feature loss information based on the target facial feature information and the predicted facial feature information. The target-driven facial image frame samples and the predicted-driven facial image frames are subjected to realism discrimination processing respectively, so as to determine the generative adversarial loss information of the image generation model based on the discrimination results; Based on the image loss information, the feature loss information, and the generative adversarial loss information, the parameters of the image generation model are adjusted to obtain the trained image generation model.
10. The method according to claim 8, characterized in that, The method further includes: Acquire a sample generation model, at least one audio-driven information sample, and a base facial image frame sample of the sample object; The sample generation model is used to extract facial pose features from the basic facial image frame samples to obtain the basic facial coefficients corresponding to the basic facial image frame samples. Temporal features are extracted from the audio-driven information samples to obtain the target expression coefficients of the corresponding facial expression states of the audio-driven information samples. Based on the basic facial coefficients and the target expression coefficients, the sample object is subjected to facial reconstruction processing to obtain the target driven facial image frame sample of the sample object in the expression state corresponding to the audio driven information sample; Training data is constructed based on the target-driven facial image frame samples, the audio-driven information samples, and the base facial image frame samples.
11. The method according to claim 10, characterized in that, The method further includes: Retrieve preset emoji guide text information; The preset emoticon guidance text information is subjected to audio conversion processing in multiple languages to obtain audio driving information samples in the multiple languages.
12. The method according to claim 10, characterized in that, The step of performing facial reconstruction processing on the sample object based on the basic facial coefficients and the target expression coefficients to obtain a target driven facial image frame sample of the sample object in the expression state corresponding to the audio driven information sample includes: Based on the target expression coefficient, the expression coefficients in the basic facial coefficients are replaced to obtain the target facial coefficients; The sample object is reconstructed based on the target facial coefficients to obtain a target driven facial image frame sample of the sample object in the expression state corresponding to the audio driven information sample.
13. An image generation apparatus, characterized in that, include: The acquisition unit is used to acquire the base facial image frame of the target object in the basic expression state, and to acquire the audio driving information for image generation; An encoding unit is used to perform feature encoding processing on the basic facial image frame to obtain a basic facial encoded feature map; The feature extraction unit is used to perform feature extraction processing on the audio driving information to obtain the audio features corresponding to the audio driving information. The generation unit is used to perform expression adjustment processing on the basic facial coding feature map based on the audio features, so as to generate a target facial image frame of the target object in the expression state corresponding to the audio driving information.
14. An electronic device, characterized in that, It includes a memory and a processor; the memory stores an application program, and the processor runs the application program within the memory to perform the operations in the image generation method according to any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the image generation method according to any one of claims 1 to 12.
16. A computer program product comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by a processor, they implement the steps of the image generation method according to any one of claims 1 to 12.