Image processing method, device and storage medium
By training facial expression and visual prediction models and utilizing video data from multiple reference speakers and target speakers, the audio and motion generalization capabilities of digital human generation technology have been improved. This solves the problem of digital humans generated in existing technologies not conforming to audio content, and achieves efficient and accurate digital human generation.
Patent Information
- Application Number
- CN202510983251.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-07-17
AI Technical Summary
Existing digital human generation technologies suffer from insufficient audio and motion generalization capabilities, resulting in digital human expressions and head movements that do not match the input audio content, and also lead to slow training speed and low efficiency.
By training facial expression prediction and visual prediction models, and using video data from multiple reference speakers to improve the generalization ability of the models, personalized training is then performed using video data from the target speaker to generate a digital human image synchronized with the target audio.
It achieves a high degree of accuracy in matching digital human images with audio content, can adapt to different languages and head movements, simplifies and speeds up the training process, and generates realistic visual effects.
Smart Images

Figure CN120495479B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to an image processing method, device, and storage medium. Background Art
[0002] With advances in artificial intelligence, digital human generation technology has made significant progress. Based on user-provided facial and voice information, this technology can generate a digital human image that matches the facial information, and its dynamic information, such as expressions, movements, and lip movements, also matches the voice information. However, related digital human generation technologies suffer from insufficient generalization capabilities.
[0003] For example, the relevant technology can only generate the lip shape and expression of a digital human based on a certain type of audio, which indicates that it has the defect of insufficient audio generalization ability. Or, the digital human generated by the relevant technology cannot flexibly set head movements, otherwise it may produce problems such as artifacts or color collapse, which indicates that it has the defect of insufficient movement generalization ability. Summary of the Invention
[0004] In order to solve at least one of the above technical problems, the present disclosure provides an image processing method, device, and storage medium. The technical solutions of the present disclosure are as follows:
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided an image processing method, including:
[0006] Acquire a plurality of first videos corresponding to a plurality of reference speakers, wherein the first videos include a first audio recording the voice of the corresponding reference speaker and a first frame sequence representing a digital image of the reference speaker synchronized with the first audio;
[0007] Training an expression prediction model based on the multiple first videos, wherein the expression prediction model is used to predict a correspondence between sound and expression;
[0008] Based on the multiple first videos and the expression prediction model, training a visual prediction model, wherein the visual prediction model is used to predict the correspondence between expressions and facial texture features;
[0009] Obtaining a second video and an image processing model, and training the image processing model based on the second video; the second video includes a second audio recording a target speaker's voice and a second image sequence representing a digital image of the target speaker synchronized with the second audio; and the image processing model includes the expression prediction model and the visual prediction model.
[0010] Get the target audio;
[0011] By inputting the target audio into the image processing model, a target picture sequence is obtained, where the target picture sequence represents a digital image of a target speaker synchronized with the target audio.
[0012] In an exemplary embodiment, the expression prediction model includes an audio feature extractor, an individual feature extractor, and an expression generator, and training the expression prediction model based on the multiple first videos includes:
[0013] For each of the first videos, inputting the corresponding first audio into the audio feature extractor to obtain a first audio feature; extracting a reference picture from the corresponding first picture sequence and inputting the reference picture into the individual feature extractor to obtain a reference speaker individual feature; inputting the first audio feature and the reference speaker individual feature into the expression generator to obtain a first predicted expression sequence;
[0014] Extracting the reference speaker's expression in each frame from each of the first frame sequences to obtain a reference expression sequence;
[0015] Based on the difference between each of the first predicted expression sequences and the corresponding reference expression sequence, the parameters of the audio feature extractor, the individual feature extractor and the expression generator are adjusted.
[0016] In an exemplary embodiment, the training of the visual prediction model based on the plurality of first videos and the expression prediction model includes:
[0017] For each of the first videos, extract a reference picture from the corresponding first picture sequence; input the reference picture into the visual prediction model to obtain a first head texture; input the corresponding first audio and the reference picture into the expression prediction model to obtain a second predicted expression sequence; and generate a reference predicted image based on the first head texture and the expression corresponding to the reference picture in the second predicted expression sequence;
[0018] Based on the difference between the reference prediction image and the foreground of the reference picture, the parameters of the visual prediction model are adjusted.
[0019] In an exemplary embodiment, the training the image processing model based on the second video includes:
[0020] Inputting the second audio into the expression prediction model to obtain a third predicted expression sequence;
[0021] Extracting a target picture from the second picture sequence; inputting the target picture into the visual prediction model to obtain a second head texture; generating a target prediction image based on the second head texture and the expression corresponding to the target picture in the third predicted expression sequence;
[0022] The image processing model is trained based on the third predicted expression sequence, the second picture sequence, the target predicted image and the target picture.
[0023] In an exemplary embodiment, the training of the image processing model based on the third predicted expression sequence, the second picture sequence, the target predicted image, and the target picture includes:
[0024] extracting the target speaker's expression in each frame of the second frame sequence to obtain a target expression sequence;
[0025] Adjusting parameters of the expression prediction model based on a difference between the third predicted expression sequence and the target expression sequence;
[0026] Based on the difference between the target predicted image and the foreground of the target picture, the second head texture is adjusted, or based on the difference between the target predicted image and the foreground of the target picture, the parameters of the visual prediction model are adjusted.
[0027] In an exemplary embodiment, the image processing model includes a texture harmony model, and training the image processing model based on the third predicted expression sequence, the second picture sequence, the target predicted image, and the target picture includes:
[0028] In a case where a target predicted expression meeting a preset requirement exists in the third predicted expression sequence, inputting the target predicted expression and the second head texture into a texture blending model to obtain a third head texture;
[0029] generating a blended image based on the third head texture and the target predicted expression;
[0030] Extracting the original picture corresponding to the target predicted expression from the second picture sequence;
[0031] Based on the difference between the blended image and the original picture, the parameters of the texture blending model are adjusted.
[0032] In an exemplary embodiment, the image processing model includes an individual background generator, and training the image processing model based on the third predicted expression sequence, the second picture sequence, the target predicted image, and the target picture includes:
[0033] For each non-target predicted expression in the third predicted expression sequence, generating a corresponding first predicted image based on the non-target predicted expression and the second head texture, and inputting the first predicted image into the individual background generator to obtain a corresponding first visual image;
[0034] For each target predicted expression in the third predicted expression sequence, generating a corresponding second predicted image based on the target predicted expression and the corresponding third head texture, and inputting the second predicted image into the individual background generator to obtain a corresponding second visual image;
[0035] obtaining a visual image sequence based on the first visual image and the second visual image;
[0036] Based on the difference between the visual image sequence and the second picture sequence, the parameters of the individual background generator are adjusted.
[0037] In an exemplary embodiment, the expression prediction model includes an audio feature extractor, an individual feature extractor, and an expression generator, and inputting the second audio into the expression prediction model to obtain a third predicted expression sequence includes:
[0038] Inputting the second audio into the audio feature extractor to obtain a second audio feature; inputting the target picture into the individual feature extractor to obtain an individual feature of the target speaker; inputting the second audio feature and the individual feature of the target speaker into the expression generator to obtain the third predicted expression sequence;
[0039] The adjusting the parameters of the expression prediction model based on the difference between the third predicted expression sequence and the target expression sequence includes:
[0040] Based on the difference between the third predicted expression sequence and the target expression sequence, the parameters of the individual feature extractor and the expression generator are adjusted while the parameters of the audio feature extractor are frozen.
[0041] In an exemplary embodiment, the plurality of first videos include audio information in multiple languages or motion information including multiple head movements; and the method uses FLAME Mesh for expression representation.
[0042] In an exemplary embodiment, the step of inputting the target audio into the image processing model to obtain a target picture sequence includes:
[0043] The expression prediction model obtains an expression sequence based on the target audio;
[0044] For each non-target expression in the expression sequence that does not meet the preset requirements, generating a corresponding first foreground image based on the non-target expression and the second head texture, and inputting the first foreground image into the individual background generator to obtain a corresponding first image;
[0045] For a target expression in the expression sequence that meets the preset requirements, input the target expression and the second head texture into the texture harmony model to obtain a target harmonious texture; generate a corresponding second foreground image based on the target expression and the target harmonious texture, and input the second foreground image into the individual background generator to obtain a corresponding second image;
[0046] The target picture sequence is obtained based on the first image and the second image.
[0047] According to a second aspect of an embodiment of the present disclosure, there is provided an image processing apparatus, including:
[0048] The generalization training module is configured to perform:
[0049] Acquire a plurality of first videos corresponding to a plurality of reference speakers, wherein the first videos include a first audio recording the voice of the corresponding reference speaker and a first frame sequence representing a digital image of the reference speaker synchronized with the first audio;
[0050] Training an expression prediction model based on the multiple first videos, wherein the expression prediction model is used to predict a correspondence between sound and expression;
[0051] The personalized capability training module is configured to perform:
[0052] Based on the multiple first videos and the expression prediction model, training a visual prediction model, wherein the visual prediction model is used to predict the correspondence between expressions and facial texture features;
[0053] Obtaining a second video and an image processing model, and training the image processing model based on the second video; the second video includes a second audio recording a target speaker's voice and a second image sequence representing a digital image of the target speaker synchronized with the second audio; and the image processing model includes the expression prediction model and the visual prediction model.
[0054] The image generation module is configured to perform:
[0055] Get the target audio;
[0056] By inputting the target audio into the image processing model, a target picture sequence is obtained, where the target picture sequence represents a digital image of a target speaker synchronized with the target audio.
[0057] In an exemplary embodiment, the expression prediction model includes an audio feature extractor, an individual feature extractor, and an expression generator, and the generalization ability training module is configured to perform:
[0058] For each of the first videos, inputting the corresponding first audio into the audio feature extractor to obtain a first audio feature; extracting a reference picture from the corresponding first picture sequence and inputting the reference picture into the individual feature extractor to obtain a reference speaker individual feature; inputting the first audio feature and the reference speaker individual feature into the expression generator to obtain a first predicted expression sequence;
[0059] Extracting the reference speaker's expression in each frame from each of the first frame sequences to obtain a reference expression sequence;
[0060] Based on the difference between each of the first predicted expression sequences and the corresponding reference expression sequence, the parameters of the audio feature extractor, the individual feature extractor and the expression generator are adjusted.
[0061] In an exemplary embodiment, the generalization ability training module is configured to execute:
[0062] For each of the first videos, extract a reference picture from the corresponding first picture sequence; input the reference picture into the visual prediction model to obtain a first head texture; input the corresponding first audio and the reference picture into the expression prediction model to obtain a second predicted expression sequence; and generate a reference predicted image based on the first head texture and the expression corresponding to the reference picture in the second predicted expression sequence;
[0063] Based on the difference between the reference prediction image and the foreground of the reference picture, the parameters of the visual prediction model are adjusted.
[0064] In an exemplary embodiment, the personalized ability training module is configured to execute:
[0065] Inputting the second audio into the expression prediction model to obtain a third predicted expression sequence;
[0066] Extracting a target picture from the second picture sequence; inputting the target picture into the visual prediction model to obtain a second head texture; generating a target prediction image based on the second head texture and the expression corresponding to the target picture in the third predicted expression sequence;
[0067] The image processing model is trained based on the third predicted expression sequence, the second picture sequence, the target predicted image and the target picture.
[0068] In an exemplary embodiment, the personalized ability training module is configured to execute:
[0069] extracting the target speaker's expression in each frame of the second frame sequence to obtain a target expression sequence;
[0070] Adjusting parameters of the expression prediction model based on a difference between the third predicted expression sequence and the target expression sequence;
[0071] Based on the difference between the target predicted image and the foreground of the target picture, the second head texture is adjusted, or based on the difference between the target predicted image and the foreground of the target picture, the parameters of the visual prediction model are adjusted.
[0072] In an exemplary embodiment, the image processing model includes a texture harmony model, and the personalized ability training module is configured to execute:
[0073] In a case where a target predicted expression meeting a preset requirement exists in the third predicted expression sequence, inputting the target predicted expression and the second head texture into a texture blending model to obtain a third head texture;
[0074] generating a blended image based on the third head texture and the target predicted expression;
[0075] Extracting the original picture corresponding to the target predicted expression from the second picture sequence;
[0076] Based on the difference between the blended image and the original picture, the parameters of the texture blending model are adjusted.
[0077] In an exemplary embodiment, the image processing model includes an individual background generator, and the personalized capability training module is configured to perform:
[0078] For each non-target predicted expression in the third predicted expression sequence, generating a corresponding first predicted image based on the non-target predicted expression and the second head texture, and inputting the first predicted image into the individual background generator to obtain a corresponding first visual image;
[0079] For each target predicted expression in the third predicted expression sequence, generating a corresponding second predicted image based on the target predicted expression and the corresponding third head texture, and inputting the second predicted image into the individual background generator to obtain a corresponding second visual image;
[0080] obtaining a visual image sequence based on the first visual image and the second visual image;
[0081] Based on the difference between the visual image sequence and the second picture sequence, the parameters of the individual background generator are adjusted.
[0082] In an exemplary embodiment, the expression prediction model includes an audio feature extractor, an individual feature extractor, and an expression generator, and the personalized ability training module is configured to perform:
[0083] Inputting the second audio into the audio feature extractor to obtain a second audio feature; inputting the target picture into the individual feature extractor to obtain an individual feature of the target speaker; inputting the second audio feature and the individual feature of the target speaker into the expression generator to obtain the third predicted expression sequence;
[0084] The adjusting the parameters of the expression prediction model based on the difference between the third predicted expression sequence and the target expression sequence includes:
[0085] Based on the difference between the third predicted expression sequence and the target expression sequence, the parameters of the individual feature extractor and the expression generator are adjusted while the parameters of the audio feature extractor are frozen.
[0086] In an exemplary embodiment, the plurality of first videos include audio information in multiple languages or motion information including multiple head movements; and the method uses FLAME Mesh for expression representation.
[0087] In an exemplary embodiment, the image generation module is configured to execute:
[0088] The expression prediction model obtains an expression sequence based on the target audio;
[0089] For each non-target expression in the expression sequence that does not meet the preset requirements, generating a corresponding first foreground image based on the non-target expression and the second head texture, and inputting the first foreground image into the individual background generator to obtain a corresponding first image;
[0090] For a target expression in the expression sequence that meets the preset requirements, input the target expression and the second head texture into the texture harmony model to obtain a target harmonious texture; generate a corresponding second foreground image based on the target expression and the target harmonious texture, and input the second foreground image into the individual background generator to obtain a corresponding second image;
[0091] The target picture sequence is obtained based on the first image and the second image.
[0092] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute instructions to implement any one of the methods in the above-mentioned first aspect.
[0093] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute any one of the methods in the first aspect of the embodiment of the present disclosure.
[0094] According to a fifth aspect of an embodiment of the present disclosure, a computer program product is provided, which includes a computer program, the computer program is stored in a readable storage medium, and at least one processor of a computer device reads and executes the computer program from the readable storage medium, so that the computer device performs any one of the methods in the first aspect of the embodiment of the present disclosure.
[0095] The technical solutions provided by the embodiments of the present disclosure bring at least the following beneficial effects:
[0096] This disclosure proposes an image processing method. In the first phase of training, the method improves the generalization capabilities of both the expression prediction model and the visual prediction model by collecting a large number of first videos. Building on this foundation, the method then proceeds to the second phase of training. In the second phase, the image processing model is trained using a second video recorded by a target speaker to improve the model's ability to generate a digital human for the target speaker. With the foundation laid by the first phase, the second phase's training data no longer needs to consider generalization. Therefore, the target speaker does not need to record videos in various languages or perform various head movements; instead, a simple recording of the second video is sufficient. This reduces the difficulty of recording the second video for the target speaker, speeds up the training process, and reduces the difficulty of the second phase. Furthermore, the recorded image processing model is capable of generating a highly accurate digital human for the target speaker, and the digital human can adapt to any language and head movement of the target speaker. Next, target audio is obtained; the target audio is input into the image processing model to obtain a target image sequence, which represents a digital image of the target speaker synchronized with the target audio. The target image sequence presents the target speaker's image, i.e., a digital human of the target speaker. The digital human's lip movements and expressions match the target audio content. Playing the target image sequence and the target audio simultaneously creates a realistic visual effect of the digital human voice broadcasting the target audio.
[0097] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0098] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the description are used to explain the principles of the present disclosure, and do not constitute an improper limitation of the present disclosure.
[0099] Figure 1 is a schematic diagram of an implementation environment of an image processing method according to an exemplary embodiment;
[0100] Figure 2 is a flowchart of an image processing method according to an exemplary embodiment;
[0101] Figure 3 is a schematic diagram of a first-stage expression prediction model training method according to an exemplary embodiment;
[0102] Figure 4 is a structural diagram of an image processing model according to an exemplary embodiment;
[0103] Figure 5 is a schematic diagram of a first-stage visual prediction model training method according to an exemplary embodiment;
[0104] Figure 6 is a schematic diagram of a second-stage training method according to an exemplary embodiment;
[0105] Figure 7 is a schematic diagram of an image processing process according to an exemplary embodiment;
[0106] Figure 8 is a block diagram of an image processing apparatus according to an exemplary embodiment;
[0107] Figure 9 A structural block diagram of a computer device according to an exemplary embodiment is shown. Figure 1 ;
[0108] Figure 10 A structural block diagram of a computer device according to an exemplary embodiment is shown. Figure 2 . DETAILED DESCRIPTION
[0109] In order to enable ordinary persons in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0110] It should be noted that the terms "first," "second," and the like in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar first objects and are not necessarily used to describe a specific order or precedence. It should be understood that the numbers used in this manner are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. Instead, they are merely examples of devices and methods consistent with certain aspects of the present disclosure as detailed in the appended claims.
[0111] Before introducing the method embodiments provided in the present application, a brief introduction is first given to the relevant terms or nouns that may be involved in the method embodiments of the present application to facilitate understanding by those skilled in the art in the field of the present application.
[0112] 3DGS: 3D Gaussian Splatting, also known as 3D Gaussian Splatting, is a 3D rendering technology proposed in 2023 that enables high-definition, real-time, and highly realistic 3D scene visualization. 3D Gaussian Splatting models and renders the 3D Gaussian distribution within a scene, enabling efficient and high-quality 3D scene visualization. 3DGS technology delivers lifelike 3D images, providing users with a more immersive visual experience.
[0113] FLAME (Faces Learned with an Articulated Model and Expressions) is a 3D deformable face model. This model collects a large amount of 3D facial scan data and uses statistical methods to derive the principal components of the face. This allows for control of facial shape, expression, blinking, and posture.
[0114] FLAME Mesh: Mesh, which means grid in Chinese, is a three-dimensional mesh structure composed of vertices, edges, and faces in computer graphics. FLAME Mesh is a three-dimensional mesh structure built on the FLAME model, which can accurately represent the basic gestures of a human face's three-dimensional expressions.
[0115] Before describing the embodiments of the present disclosure in detail, the relevant technical background related to the embodiments of the present disclosure is introduced to facilitate understanding by those skilled in the art in the art of this application.
[0116] When training image processing models used to generate digital humans in related technologies, generalization processing is not required, resulting in insufficient generalization capabilities. For example, if user A records a video of them speaking English and trains an image processing model based on this video, the image processing model can only generate a digital human for user A speaking English. If audio of someone speaking Chinese is input into the image processing model, the facial expressions and lip movements of the digital human output by the image processing model will not match the Chinese audio, and the model will not be able to well represent the natural state of the digital human corresponding to the Chinese content. This indicates that the audio generalization capabilities of the image processing model are insufficient. Moreover, the digital human output by the image processing model is difficult to control to perform head movements, otherwise it may crash and other problems. This is because the image processing model's movement generalization capabilities are also insufficient. In addition, the training speed of such image processing models that have not been generalized is also very slow and the training efficiency is also low.
[0117] In view of this, in order to obtain an image processing model with strong generalization ability and enable the image processing model to quickly, efficiently and accurately generate a digital human image for a specific speaker and an image sequence based on the digital human image, the present disclosure provides an image processing method.
[0118] All data or information involved in this disclosure are authorized by the user or fully authorized by all parties.
[0119] Figure 1 FIG. 1 is a schematic diagram showing an implementation environment of an image processing method according to an exemplary embodiment. Assume that the electronic device is provided as a terminal. Figure 1 , the implementation environment specifically includes: a terminal 101 and a server 102.
[0120] The terminal 101 may be at least one of a smartphone, a smartwatch, a desktop computer, a laptop computer, and a portable computer. An application program that utilizes the image processing method of the present disclosure may be installed and run on the terminal 101, and a user may log in to the application program through the terminal 101 to obtain services related to the image processing method of the present disclosure. The terminal 101 may generally refer to one of a plurality of terminals. This embodiment only uses the terminal 101 as an example. Those skilled in the art will appreciate that the number of the above-mentioned terminals may be greater or lesser. For example, the above-mentioned terminals may be only a few, or the above-mentioned terminals may be dozens, hundreds, or even more. The embodiments of this disclosure do not limit the number of terminals and the types of devices.
[0121] Server 102 can be at least one of a single server, multiple servers, a cloud computing platform, and a virtualization center. Server 102 can be connected to terminal 101 and other terminals via a wireless or wired network. Server 102 can implement the image processing method disclosed herein and transmit the image processing results to terminal 101, which can then use or display the image processing results. Of course, server 102 can also include other functional servers to provide more comprehensive and diverse services.
[0122] Figure 2 is a flowchart of an image processing method according to an exemplary embodiment. Figure 2 As shown, the above method includes:
[0123] In step S210 , a plurality of first videos corresponding to a plurality of reference speakers are obtained, wherein the first videos include a first audio recording the voice of the corresponding reference speaker and a first frame sequence representing a digital image of the reference speaker synchronized with the first audio.
[0124] The purpose of using multiple first videos of multiple reference speakers in this disclosure is to improve the generalization ability of the image processing model. Therefore, there is no limit on the number of first videos, as long as these first videos can effectively help the model learn the characteristics of different reference speakers. To further optimize the training effect of the model, first videos can be collected from different scenes, different shooting angles, different lighting conditions, etc. to enrich the diversity of training data. At the same time, when obtaining the first video, the synchronization accuracy of the first audio and first picture sequence must be ensured to avoid the model's learning of features being affected by the lack of synchronization between the audio and the picture.
[0125] In order to improve the model's generalization ability for audio, so that it can effectively generate corresponding expressions and lip shapes for audio spoken in different languages, audio in multiple languages can be collected and corresponding first videos can be obtained, that is, the multiple first videos include audio information in multiple languages. These audio information can cover speech in different languages, different accents, different speaking speeds, and different emotional expressions, so that the model can more accurately learn and understand the correspondence between audio and expressions and lip shapes in different languages.
[0126] To improve the model's ability to generalize movements, you can collect videos of various head movements as the first videos. These movements can include tilting the head at different angles, turning the head, nodding, and so on. That is, the multiple first videos can include information about various head movements. This allows the model to learn a rich variety of movement information, preventing problems such as corruption when generating digital humans with head movements, and enhancing the model's generalization ability in movement generation.
[0127] The present disclosure does not limit the source of the multiple first videos; they can be derived from an open source dataset that includes many user-recorded videos of people speaking in various languages, with synchronized audio and video, and users making various head movements during recording. Alternatively, the multiple first videos can be collected through recording, for example, by arranging multiple users as reference speakers and recording them according to preset text in multiple languages, with different voice emotions, speech speed requirements, and rich head movement instructions, to obtain high-quality and diverse first videos.
[0128] In step S220, an expression prediction model is trained based on the multiple first videos, where the expression prediction model is used to predict the correspondence between sound and expression.
[0129] The present disclosure does not limit the structure of the expression prediction model and its training method. In an exemplary embodiment, the expression prediction model may include an audio feature extractor, an individual feature extractor, and an expression generator. The audio feature extractor is used to extract information that can represent audio features from the input audio data, such as pitch, timbre, audio rhythm, etc., to provide an audio basis for subsequent expression generation. The individual feature extractor can extract features related to the reference speaker individual, such as its facial features. The expression generator will combine the information extracted by the audio feature extractor and the individual feature extractor to predict the corresponding expression. All expressions in the present disclosure can be represented by FLAME Mesh. FLAME Mesh can constrain expression information to the area around the head, so that the model can focus more on learning head information, and focus on predicting the correspondence between sound and expression in the head area in step S220.
[0130] Please refer to Figure 3 , which is a schematic diagram of a first-stage expression prediction model training method according to an exemplary embodiment of the present disclosure. The expression prediction model training based on the plurality of first videos includes:
[0131] In step S310, for each of the first videos, the corresponding first audio is input into the audio feature extractor to obtain the first audio feature; a reference picture is extracted from the corresponding first picture sequence and input into the individual feature extractor to obtain the reference speaker individual feature; the first audio feature and the reference speaker individual feature are input into the expression generator to obtain a first predicted expression sequence.
[0132] The present disclosure does not limit the reference picture, for example, it can be the first picture in the first picture sequence, or the last picture. The present disclosure does not limit the structures of the audio feature extractor, individual feature extractor, and expression generator, as long as they can meet their corresponding functional requirements.
[0133] Please refer to Figure 4, which shows a schematic diagram of the structure of the image processing model in an exemplary embodiment of the present disclosure. The audio feature extractor may include sequentially connected Wav2Vec and linear layers. Wav2Vec is an unsupervised speech representation learning model that can learn rich speech features from raw audio signals. In the audio feature extractor of the present disclosure, Wav2Vec can convert the input audio signal into a feature representation with semantic information. The features obtained after Wav2Vec processing are further processed by the linear layer to obtain the first audio feature. The individual feature extractor may include an embedded feature extraction layer, and the expression generator may be a diffusion model with a cross-distribution of self-attention layers and cross-attention layers. The diffusion model with a cross-distribution of self-attention layers and cross-attention layers as an expression generator can use the first audio feature and the individual features of the reference speaker to accurately and efficiently generate a first predicted expression sequence. The self-attention layer allows the model to focus on the internal relationships of the input features during the expression generation process and capture long-range dependency information, while the cross-attention layer can establish connections between different features, fusing audio features and individual features, so that the generated expressions are more consistent with the individual speaker and the lip expressions corresponding to the audio.
[0134] In step S320, the reference speaker expression of each frame is extracted from each of the first frame sequences to obtain a reference expression sequence; based on the differences between each of the first predicted expression sequences and the corresponding reference expression sequences, the parameters of the audio feature extractor, the individual feature extractor and the expression generator are adjusted.
[0135] The reference expression sequence is extracted from each picture in the first picture sequence, which is a FLAME Mesh sequence extracted based on these pictures, and is used as a true value for adjusting the parameters of the audio feature extractor, the individual feature extractor and the expression generator. The present disclosure does not limit the method for adjusting the loss of the parameters of the audio feature extractor, the individual feature extractor and the expression generator based on the calculation of the difference between the first predicted expression sequence and the corresponding reference expression sequence. For example, the mean square error loss function or the cross entropy loss function can be used. By minimizing these loss functions, the parameters of the audio feature extractor, the individual feature extractor and the expression generator can be continuously optimized by the gradient descent method, so that the audio feature extractor, the individual feature extractor and the expression generator have a wide range of audio to expression generalization capabilities.
[0136] In step S230, a visual prediction model is trained based on the multiple first videos and the expression prediction model, where the visual prediction model is used to predict the correspondence between expressions and facial texture features.
[0137] The present disclosure does not limit the structure of the visual prediction model and its training method. In an exemplary embodiment, reference Figure 5 , which shows a schematic diagram of the first stage visual prediction model training method of the present disclosure. The visual prediction model is trained based on the multiple first videos and the expression prediction model, including:
[0138] S510. For each of the first videos, extract a reference picture from the corresponding first picture sequence; input the reference picture into the visual prediction model to obtain a first head texture; input the corresponding first audio and the reference picture into the expression prediction model to obtain a second predicted expression sequence; and generate a reference predicted image based on the first head texture and the expression corresponding to the reference picture in the second predicted expression sequence.
[0139] The present disclosure does not limit the reference picture. For example, it can be the first picture in the first picture sequence, or the last picture. Step S510 can be implemented on the basis of the first stage of the expression prediction model implemented in step S220. After the expression prediction model has been adjusted in the aforementioned step S320, it has a relatively strong audio generalization capability. In this case, the first audio and the reference picture are input into the expression prediction model again. The process of obtaining the second predicted expression sequence is based on the same inventive concept as the process of obtaining the first predicted expression sequence in the aforementioned step S310, and no further details are given here. The second predicted expression sequence is also composed of a series of FLAME Mesh, and the expression corresponding to the reference picture in the second predicted expression sequence is the FLAME Mesh corresponding to the reference picture. For example, if the reference picture is the first picture in the first picture sequence, the expression corresponding to the reference picture in the second predicted expression sequence is the first FLAME Mesh of the first predicted expression sequence.
[0140] The visual prediction model can output a first head texture based on the reference image. This first head texture can be a two-dimensional texture. This disclosure does not limit the structure of the visual prediction model. It is an "individual Gaussian generator" used to generate two-dimensional texture features corresponding to the reference speaker, as long as the functional purpose is achieved.
[0141] The present disclosure can achieve the process of generating a reference prediction image based on the expression corresponding to the first head texture and the reference picture in the second predicted expression sequence by Gaussian splashing. Specifically, the two-dimensional first head texture is first mapped to a three-dimensional space based on the UV unfolding principle to construct a three-dimensional skin. UV unfolding is an important concept in computer graphics, which is equivalent to peeling the surface of a three-dimensional object and then laying this layer of skin flat on a two-dimensional plane, except that the reverse operation is performed in the present disclosure. On the basis of obtaining the three-dimensional skin, the three-dimensional skin is bound to the FLAME Mesh determined based on the reference picture, and 3D Gaussian splashing is performed to render a three-dimensional reference prediction image.
[0142] S520. Adjust the parameters of the visual prediction model based on the difference between the reference prediction image and the foreground of the reference picture.
[0143] By extracting the foreground of the reference picture, the corresponding reference speaker's head image can be obtained. The difference between the head image and the three-dimensional reference prediction image can be quantified as a corresponding loss, which can be used to adjust the parameters of the visual prediction model by gradient descent. The present disclosure does not limit the calculation method for the loss used to quantify the difference between the reference prediction image and the foreground of the reference picture. Exemplarily, the mean square error loss function or the cross entropy loss function can be used for calculation. Based on steps S510-S520, Gaussian splashing and binding with the FLAME Mesh output by the expression prediction model can be used to focus on fitting the head visual information, thereby improving the model's action generalization ability and visual generalization ability based on the trained expression prediction model. The FLAME Mesh output by the expression prediction model is used as an intermediate feature used in the training of the visual prediction model to improve the accuracy of texture prediction.
[0144] In step S240, a second video and an image processing model are obtained, and the image processing model is trained based on the second video; the second video includes a second audio recording the voice of the target speaker and a second picture sequence representing a digital image of the target speaker synchronized with the second audio, and the image processing model includes the expression prediction model and the visual prediction model.
[0145] Steps S210-S230 constitute the first phase of training in this disclosure. By collecting a large number of first videos, the generalization capabilities of both the expression prediction model and the visual prediction model are improved. Building on this foundation, the second phase of training can begin. In the second phase, the image processing model is trained using the second video recorded by the target speaker to improve its ability to generate a digital human for the target speaker. With the foundation laid in the first phase, the second phase's training data no longer needs to consider generalization. Therefore, the target speaker does not need to record videos in various languages or perform various head movements; instead, a simple recording of the second video is sufficient. This reduces the difficulty for the target speaker to record the second video, speeds up the second phase's training, and reduces its difficulty. Furthermore, the recorded image processing model can generate a highly accurate digital human for the target speaker, and this digital human can adapt to any language spoken and any head movements made by the target speaker.
[0146] In step S250, a target audio is acquired; and a target picture sequence is obtained by inputting the target audio into the image processing model. The target picture sequence represents a digital image of a target speaker synchronized with the target audio.
[0147] The target audio can be generated based on the target text. This disclosure does not limit the target text; it can be any content that the digital human is expected to express, such as daily conversations, professional explanations, storytelling, etc. This disclosure does not limit the method for obtaining the target audio; it can be the target audio obtained by speech synthesis of the target text, or it can be the target audio obtained by any speaker reciting the target text. The target image sequence presents the image of the target speaker, that is, the digital human of the target speaker, and the lip shape and expression of the digital human are consistent with the content of the target audio. By playing the target image sequence and the target audio synchronously, a realistic visual effect of the target speaker's own voice broadcasting the target audio can be achieved.
[0148] The present disclosure does not limit the application scenarios of the image processing method. For example, in various live broadcast scenarios such as e-commerce, news, and entertainment, the video recorded by a real anchor can be used as the second video, and the present disclosure can be used to train the image processing model corresponding to the real anchor. Then, it is only necessary to input any target audio into the image processing model to obtain the corresponding target picture sequence. When the target picture sequence is played, it can reflect the broadcast effect of the real anchor, thereby reducing the workload of the real anchor. For another example, the present disclosure can be used to easily replace the role played by actor 1 in the movie with actor 2, and there is no sense of disobedience. For another example, you can also use videos of deceased relatives and friends to reconstruct lifelike digital people and enhance the user's emotional experience.
[0149] This disclosure does not limit the method for training the image processing model based on the second video. Figure 6, which is a schematic diagram of the second stage training method of the present disclosure. The training of the image processing model based on the second video includes:
[0150] S610. Input the second audio into the expression prediction model to obtain a third predicted expression sequence.
[0151] This process is consistent with the aforementioned inventive concept of obtaining the second predicted expression sequence and the first predicted expression sequence, and will not be elaborated on. In an exemplary embodiment, the expression prediction model includes an audio feature extractor, an individual feature extractor, and an expression generator. The inputting of the second audio into the expression prediction model to obtain the third predicted expression sequence includes: inputting the second audio into the audio feature extractor to obtain the second audio feature; inputting the target picture into the individual feature extractor to obtain the individual feature of the target speaker; and inputting the second audio feature and the individual feature of the target speaker into the expression generator to obtain the third predicted expression sequence. The target picture has been described above and will not be elaborated on.
[0152] S620. Extract a target picture from the second picture sequence; input the target picture into the visual prediction model to obtain a second head texture; and generate a target prediction image based on the second head texture and the expression corresponding to the target picture in the third predicted expression sequence.
[0153] The meanings and acquisition methods of the target picture, the second head texture, and the target predicted image are consistent with the inventive concepts of the meanings and acquisition methods of the aforementioned reference picture, the first head texture, and the reference predicted image, and are not elaborated on.
[0154] S630. Train the image processing model based on the third predicted expression sequence, the second picture sequence, the target predicted image and the target picture.
[0155] Based on the third predicted expression sequence, the second picture sequence, the target predicted image and the target picture, the image processing is trained to obtain an image processing model. This process does not require the second video to provide information for training generalization capabilities, but only requires it to provide specific information related to the target speaker. Therefore, the training process is simple and the training speed is fast, thereby realizing the training of an image generation model with both specific digital human generation capabilities and generalization capabilities.
[0156] In an exemplary embodiment, please refer to Figure 7 , which shows a schematic diagram of the image processing process of the present disclosure. The training of the image processing model based on the third predicted expression sequence, the second picture sequence, the target predicted image and the target picture includes:
[0157] S710. Extract the target speaker's expression in each frame of the second frame sequence to obtain a target expression sequence.
[0158] This process is consistent with the aforementioned inventive concept of obtaining a reference expression sequence and will not be elaborated on.
[0159] S720. Adjust the parameters of the expression prediction model based on the difference between the third predicted expression sequence and the target expression sequence.
[0160] In an exemplary embodiment, adjusting the parameters of the expression prediction model based on the difference between the third predicted expression sequence and the target expression sequence includes: adjusting the parameters of the individual feature extractor and the expression generator while freezing the parameters of the audio feature extractor based on the difference between the third predicted expression sequence and the target expression sequence. Figure 4 By freezing the parameters of the audio feature extractor, the audio generalization ability learned by the expression prediction model in the first training stage is not affected to the greatest extent. Only the parameters of the individual feature extractor and the expression generator are adjusted, so that the expression prediction model can further learn the target speaker-specific audio information and expression information in the second training stage. The expression prediction model trained in the second training stage has the ability to generate accurate expressions regardless of the target speaker's speech.
[0161] S730. Adjust the second head texture based on the difference between the target predicted image and the foreground of the target picture, or adjust the parameters of the visual prediction model based on the difference between the target predicted image and the foreground of the target picture.
[0162] Both the method for adjusting the expression prediction model parameters in step S720 and the method for adjusting the second head texture or visual prediction model parameters in step S730 employ gradient descent. This disclosure does not limit the method for quantifying the differences in steps S720 and S730, as they are based on the same inventive concept as previously described and are omitted for clarity. By adjusting the second head texture or visual prediction model parameters, the visual prediction model can further learn target speaker-specific visual and textural information during the second training phase. This allows the visual prediction model trained in the second training phase to generate accurate textures regardless of the target speaker's speech.
[0163] In an exemplary embodiment, the image processing model includes a texture blending model, see Figure 4The texture blending model is connected to the visual prediction model and is used to blend the corresponding texture when the target speaker makes a large-scale expression, so that the blended texture can be accurately covered on the corresponding FLAME Mesh, avoiding texture collapse in the case of large-scale expressions. Accordingly, the image processing model is trained based on the third predicted expression sequence, the second picture sequence, the target predicted image, and the target picture, including:
[0164] S740. When there is a target predicted expression that meets preset requirements in the third predicted expression sequence, input the target predicted expression and the second head texture into a texture blending model to obtain a third head texture.
[0165] The present disclosure does not limit the preset requirements. For example, the expression amplitude can be set. If the expression amplitude exceeds a certain threshold, it is considered to meet the preset requirements. There is no limitation on the quantification method of the expression amplitude. For example, the displacement change of the facial key points can be used to quantify the expression amplitude. The target predicted expression refers to the expression with a large expression amplitude. For this type of expression, it is also necessary to determine the corresponding second head texture and render it through Gaussian splashing. However, this type of expression is prone to rendering collapse, so the second head texture needs to be specially harmonized based on this type of expression. First, the target predicted expression and the second head texture are input into the texture harmonization model to obtain the third head texture. The third head texture is based on the second head texture and is a harmonized texture obtained after specific adjustment based on the target predicted expression.
[0166] S750. Generate a blended image based on the third head texture and the target predicted expression.
[0167] This process is the same as the aforementioned inventive concept of predicting the avatar based on the second head texture rendering target, and will not be elaborated on.
[0168] S760. Extract the original picture corresponding to the target predicted expression from the second picture sequence; and adjust the parameters of the texture blending model based on the difference between the blended image and the original picture.
[0169] By training the texture blending model, we can perform targeted blending on large-scale expressions based on the second head texture generated by the visual prediction model to obtain a texture that is suitable for large-scale expressions. Using the blended texture for Gaussian rendering on large-scale expressions can avoid texture cracks or texture collapse caused by large-scale expressions.
[0170] In an exemplary embodiment, the image processing model further includes an individual background generator, see Figure 4The individual background generator is connected to the expression prediction model, the visual prediction model, and the texture harmony model, and is used to render an image with a background, regardless of whether the target speaker makes a large or small expression. The background image includes not only the digital human head generated based on texture and expression, but also the digital human torso and the overall background of the digital human. Accordingly, the image processing model is trained based on the third predicted expression sequence, the second picture sequence, the target predicted image, and the target picture, including:
[0171] S770. For each non-target predicted expression in the third predicted expression sequence, generate a corresponding first predicted image based on the non-target predicted expression and the second head texture, input the first predicted image into the individual background generator, and obtain a corresponding first visual image.
[0172] The present disclosure does not limit the structure of the background generator. For example, it can be a lightweight UNet that seamlessly fuses the torso of the digital human and the background based on the first predicted image including only the head to obtain the first visual image.
[0173] S780. For each target predicted expression in the third predicted expression sequence, generate a corresponding second predicted image based on the target predicted expression and the corresponding third head texture, input the second predicted image into the individual background generator, and obtain a corresponding second visual image.
[0174] This process is based on the same inventive concept as S770 and will not be elaborated on here.
[0175] S790. Obtain a visual image sequence based on the first visual image and the second visual image; and adjust parameters of the individual background generator based on differences between the visual image sequence and the second image sequence.
[0176] The method for quantifying the difference between the image sequences and the image differences can be based on the same inventive concept and will not be elaborated on here. By training the individual background generator, the generated digital human can have not only a head, but also a torso and background, making it more realistic.
[0177] For this disclosure Figure 4The image processing model shown, after training, can have both the ability to generalize to any head movement and any language, and the ability to accurately fit the personalized style of the target speaker. For any target audio, the target audio can be input into the image processing model to obtain a target picture sequence. Specifically, the target picture sequence is obtained by inputting the target audio into the image processing model, including: the expression prediction model obtains an expression sequence based on the target audio; for each non-target expression in the expression sequence that does not meet the preset requirements, a corresponding first foreground image is generated based on the non-target expression and the second head texture, and the first foreground image is input into the individual background generator to obtain a corresponding first image; for the target expression in the expression sequence that meets the preset requirements, the target expression and the second head texture are input into the texture harmony model to obtain a target harmonious texture; based on the target expression and the target harmonious texture, a corresponding second foreground image is generated, and the second foreground image is input into the individual background generator to obtain a corresponding second image; based on the first image and the second image, the target picture sequence is obtained.
[0178] This disclosure uses FLAME Mesh for expression control, and FLAME Mesh itself can support editing of head movements. Because the image processing model has the ability to generalize head movements, head movements can be edited during the generation of the aforementioned target image sequence without any distortion. During the target image sequence generation process, large-scale expressions are also harmonized to ensure that no distortion occurs under any expression. Clearly, this disclosure can generate highly realistic digital humans. The target image sequence composed of this digital human image, combined with the target audio, is sufficiently realistic to achieve an effect close to that of a live broadcast.
[0179] In summary, the image processing method disclosed herein enables personalized training for any target speaker (second-stage training) based on generalized training (first-stage training). This generates a target image sequence of the target speaker for any target audio, while maintaining audio and visual generalization capabilities. This enables high-fidelity, strong generalization, low-cost, and real-time digital human generation. Through two-stage training, the method learns global prior knowledge of digital faces and character-specific information, respectively, improving training efficiency. In the first stage, audio generalization is enhanced, making it applicable to all audio types (Chinese, English, singing, etc.), expanding the range of applicable scenarios. Furthermore, generalization for large head movements is enhanced, allowing digital humans to perform large head movements such as lowering, raising, and turning their heads, rather than simply looking straight ahead, increasing their appeal. In the second stage, portions of the prior network are frozen to reduce errors and improve training accuracy. By introducing a harmonic model, details are optimized for large expressions, and a background generator further enhances background generation capabilities, improving realism.
[0180] Figure 8 FIG. 1 is a block diagram of an image processing apparatus according to an exemplary embodiment. Figure 8 , the device comprises:
[0181] The generalization ability training module 810 is configured to execute:
[0182] Acquire a plurality of first videos corresponding to a plurality of reference speakers, wherein the first videos include a first audio recording the voice of the corresponding reference speaker and a first frame sequence representing a digital image of the reference speaker synchronized with the first audio;
[0183] Training an expression prediction model based on the multiple first videos, wherein the expression prediction model is used to predict a correspondence between sound and expression;
[0184] The personalized capability training module 820 is configured to execute:
[0185] Based on the multiple first videos and the expression prediction model, training a visual prediction model, wherein the visual prediction model is used to predict the correspondence between expressions and facial texture features;
[0186] Obtaining a second video and an image processing model, and training the image processing model based on the second video; the second video includes a second audio recording a target speaker's voice and a second image sequence representing a digital image of the target speaker synchronized with the second audio; and the image processing model includes the expression prediction model and the visual prediction model.
[0187] The image generation module 830 is configured to perform:
[0188] Get the target audio;
[0189] By inputting the target audio into the image processing model, a target picture sequence is obtained, where the target picture sequence represents a digital image of a target speaker synchronized with the target audio.
[0190] In an exemplary embodiment, the expression prediction model includes an audio feature extractor, an individual feature extractor, and an expression generator, and the generalization ability training module 810 is configured to perform:
[0191] For each of the first videos, inputting the corresponding first audio into the audio feature extractor to obtain a first audio feature; extracting a reference picture from the corresponding first picture sequence and inputting the reference picture into the individual feature extractor to obtain a reference speaker individual feature; inputting the first audio feature and the reference speaker individual feature into the expression generator to obtain a first predicted expression sequence;
[0192] Extracting the reference speaker's expression in each frame from each of the first frame sequences to obtain a reference expression sequence;
[0193] Based on the difference between each of the first predicted expression sequences and the corresponding reference expression sequence, the parameters of the audio feature extractor, the individual feature extractor and the expression generator are adjusted.
[0194] In an exemplary embodiment, the generalization ability training module 810 is configured to execute:
[0195] For each of the first videos, extract a reference picture from the corresponding first picture sequence; input the reference picture into the visual prediction model to obtain a first head texture; input the corresponding first audio and the reference picture into the expression prediction model to obtain a second predicted expression sequence; and generate a reference predicted image based on the first head texture and the expression corresponding to the reference picture in the second predicted expression sequence;
[0196] Based on the difference between the reference prediction image and the foreground of the reference picture, the parameters of the visual prediction model are adjusted.
[0197] In an exemplary embodiment, the personalized ability training module 820 is configured to execute:
[0198] Inputting the second audio into the expression prediction model to obtain a third predicted expression sequence;
[0199] Extracting a target picture from the second picture sequence; inputting the target picture into the visual prediction model to obtain a second head texture; generating a target prediction image based on the second head texture and the expression corresponding to the target picture in the third predicted expression sequence;
[0200] The image processing model is trained based on the third predicted expression sequence, the second picture sequence, the target predicted image and the target picture.
[0201] In an exemplary embodiment, the personalized ability training module 820 is configured to execute:
[0202] extracting the target speaker's expression in each frame of the second frame sequence to obtain a target expression sequence;
[0203] Adjusting parameters of the expression prediction model based on a difference between the third predicted expression sequence and the target expression sequence;
[0204] Based on the difference between the target predicted image and the foreground of the target picture, the second head texture is adjusted, or based on the difference between the target predicted image and the foreground of the target picture, the parameters of the visual prediction model are adjusted.
[0205] In an exemplary embodiment, the image processing model includes a texture harmony model, and the personalized ability training module 820 is configured to perform:
[0206] In a case where a target predicted expression meeting a preset requirement exists in the third predicted expression sequence, inputting the target predicted expression and the second head texture into a texture blending model to obtain a third head texture;
[0207] generating a blended image based on the third head texture and the target predicted expression;
[0208] Extracting the original picture corresponding to the target predicted expression from the second picture sequence;
[0209] Based on the difference between the blended image and the original picture, the parameters of the texture blending model are adjusted.
[0210] In an exemplary embodiment, the image processing model includes an individual background generator, and the personalized capability training module 820 is configured to perform:
[0211] For each non-target predicted expression in the third predicted expression sequence, generating a corresponding first predicted image based on the non-target predicted expression and the second head texture, and inputting the first predicted image into the individual background generator to obtain a corresponding first visual image;
[0212] For each target predicted expression in the third predicted expression sequence, generating a corresponding second predicted image based on the target predicted expression and the corresponding third head texture, and inputting the second predicted image into the individual background generator to obtain a corresponding second visual image;
[0213] obtaining a visual image sequence based on the first visual image and the second visual image;
[0214] Based on the difference between the visual image sequence and the second picture sequence, the parameters of the individual background generator are adjusted.
[0215] In an exemplary embodiment, the expression prediction model includes an audio feature extractor, an individual feature extractor, and an expression generator, and the personalized ability training module 820 is configured to perform:
[0216] Inputting the second audio into the audio feature extractor to obtain a second audio feature; inputting the target picture into the individual feature extractor to obtain an individual feature of the target speaker; inputting the second audio feature and the individual feature of the target speaker into the expression generator to obtain the third predicted expression sequence;
[0217] The adjusting the parameters of the expression prediction model based on the difference between the third predicted expression sequence and the target expression sequence includes:
[0218] Based on the difference between the third predicted expression sequence and the target expression sequence, the parameters of the individual feature extractor and the expression generator are adjusted while the parameters of the audio feature extractor are frozen.
[0219] In an exemplary embodiment, the plurality of first videos include audio information in multiple languages or motion information including multiple head movements; and the method uses FLAME Mesh for expression representation.
[0220] In an exemplary embodiment, the image generation module 830 is configured to execute:
[0221] The expression prediction model obtains an expression sequence based on the target audio;
[0222] For each non-target expression in the expression sequence that does not meet the preset requirements, generating a corresponding first foreground image based on the non-target expression and the second head texture, and inputting the first foreground image into the individual background generator to obtain a corresponding first image;
[0223] For a target expression in the expression sequence that meets the preset requirements, input the target expression and the second head texture into the texture harmony model to obtain a target harmonious texture; generate a corresponding second foreground image based on the target expression and the target harmonious texture, and input the second foreground image into the individual background generator to obtain a corresponding second image;
[0224] The target picture sequence is obtained based on the first image and the second image.
[0225] Regarding the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments of the relevant methods and will not be elaborated on here.
[0226] Please refer to Figure 9 , which shows a structural block diagram of a computer device provided by an exemplary embodiment of the present disclosure Figure 1 The computer device may be a terminal. The computer device is used to implement the image processing method provided in the above embodiment. Specifically:
[0227] Typically, the computer device 900 includes a processor 901 and a memory 902 .
[0228] Processor 901 may include one or more processing cores, such as a quad-core processor or an octa-core processor. Processor 901 may be implemented in hardware using at least one of a DSP (Digital Signal Processing), an FPGA (Field Programmable Gate Array), and a PLA (Programmable Logic Array). Processor 901 may also include a main processor and a coprocessor. The main processor is used to process data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor used to process data in the standby state. In one exemplary embodiment, processor 901 may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing content displayed on the display screen. In one exemplary embodiment, processor 901 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0229] The memory 902 may include one or more computer-readable storage media, which may be non-transitory. The memory 902 may also include high-speed random access memory and non-volatile memory, such as one or more magnetic disk storage devices or flash memory storage devices. In an exemplary embodiment, the non-transitory computer-readable storage medium in the memory 902 is used to store at least one instruction, at least one program, code set, or instruction set, and is configured to be executed by one or more processors to implement the above-mentioned image processing method.
[0230] In an exemplary embodiment, computer device 900 may optionally include a peripheral device interface 903 and at least one peripheral device. Processor 901, memory 902, and peripheral device interface 903 may be connected via a bus or signal lines. Each peripheral device may be connected to peripheral device interface 903 via a bus, signal lines, or circuit boards. Specifically, the peripheral device includes at least one of a radio frequency circuit 904, a touchscreen display 905, a camera assembly 906, an audio circuit 907, a positioning assembly 908, and a power supply 909.
[0231] Those skilled in the art will understand that Figure 9 The structure shown in the figure does not constitute a limitation on the computer device 900, and the computer device 900 may include more or fewer components than shown in the figure, or combine some components, or adopt a different component arrangement.
[0232] Please refer to Figure 10 It shows a structural block diagram of a computer device provided by another exemplary embodiment of the present disclosure. Figure 2 The computer device may be a server for executing the above-mentioned image processing method. Specifically:
[0233] Computer device 1000 includes a central processing unit (CPU) 1001, a system memory 1004 including random access memory (RAM) 1002 and read-only memory (ROM) 1003, and a system bus 1005 connecting system memory 1004 and CPU 1001. Computer device 1000 also includes a basic input / output system (I / O) 1006 that facilitates information transfer between various components within the computer, and a mass storage device 1007 for storing an operating system 1013, application programs 1014, and other program modules 1011.
[0234] The basic input / output system 1006 includes a display 1008 for displaying information and an input device 1009, such as a mouse and keyboard, for user input. Both the display 1008 and the input device 1009 are connected to the central processing unit 1001 via an input / output controller 1100 connected to the system bus 1005. The basic input / output system 1006 may also include an input / output controller 1100 for receiving and processing input from a variety of other devices, such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1100 also provides output to a display screen, printer, or other types of output devices.
[0235] Mass storage device 1007 is connected to central processing unit 1001 via a mass storage controller (not shown) connected to system bus 1005. Mass storage device 1007 and its associated computer-readable media provide non-volatile storage for computer device 1000. In other words, mass storage device 1007 may include computer-readable media (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.
[0236] Without loss of generality, computer-readable media may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic cassettes, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that computer storage media is not limited to the aforementioned types. The aforementioned system memory 1004 and mass storage device 1007 may be collectively referred to as memory.
[0237] According to various embodiments of the present disclosure, the computer device 1000 may also be connected to a remote computer on a network such as the Internet for operation. That is, the computer device 1000 may be connected to the network 1012 via the network interface unit 1011 connected to the system bus 1005. Alternatively, the network interface unit 1011 may be used to connect to other types of networks or remote computer systems (not shown).
[0238] The memory further includes a computer program, which is stored in the memory and configured to be executed by one or more processors to implement the image processing method.
[0239] In an exemplary embodiment, a computer-readable storage medium is also provided, in which at least one instruction, at least one program, a code set or an instruction set is stored. When the at least one instruction, the at least one program, the code set or the instruction set is executed by a processor, the image processing method is implemented.
[0240] Optionally, the computer-readable storage medium may include: ROM (Read Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or an optical disk. The random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0241] In an exemplary embodiment, a computer-readable storage medium including program code is also provided, such as a memory including the program code. The program code can be executed by a processor to perform the image processing method. Alternatively, the computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, an optical data storage device, or the like.
[0242] In an exemplary embodiment, a computer program product is further provided, including a computer program, which implements the above-mentioned image processing method when executed by a processor.
[0243] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0244] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. An image processing method, characterized in that: The method comprises: Acquire a plurality of first videos corresponding to a plurality of reference speakers, wherein the first videos include a first audio recording the voice of the corresponding reference speaker and a first frame sequence representing a digital image of the reference speaker synchronized with the first audio; Training an expression prediction model based on the multiple first videos, wherein the expression prediction model is used to predict a correspondence between sound and expression; Based on the multiple first videos and the expression prediction model, training a visual prediction model, wherein the visual prediction model is used to predict the correspondence between expressions and facial texture features; Obtaining a second video and an image processing model, wherein the second video includes a second audio recording a target speaker's voice and a second image sequence representing a digital image of the target speaker synchronized with the second audio, and the image processing model includes the expression prediction model and the visual prediction model; Training the image processing model based on the second video to obtain an image processing model corresponding to the target speaker; Get the target audio obtained by any speaker reciting the target text; By inputting the target audio into an image processing model corresponding to the target speaker, a target picture sequence is obtained, wherein the target picture sequence represents a digital image of the target speaker synchronized with the target audio.
2. The image processing method according to claim 1, wherein: The expression prediction model includes an audio feature extractor, an identity feature extractor, and an expression generator, and the expression prediction model is trained based on the multiple first videos, including: For each of the first videos, inputting the corresponding first audio into the audio feature extractor to obtain a first audio feature; extracting a reference picture from the corresponding first picture sequence and inputting the reference picture into the identity feature extractor to obtain a reference speaker identity feature; inputting the first audio feature and the reference speaker identity feature into the expression generator to obtain a first predicted expression sequence; Extracting the reference speaker's expression in each frame from each of the first frame sequences to obtain a reference expression sequence; Based on the difference between each of the first predicted expression sequences and the corresponding reference expression sequence, the parameters of the audio feature extractor, the identity feature extractor, and the expression generator are adjusted.
3. The image processing method according to claim 1 or 2, characterized in that: The step of training a visual prediction model based on the plurality of first videos and the expression prediction model comprises: For each of the first videos, extract a reference picture from the corresponding first picture sequence; input the reference picture into the visual prediction model to obtain a first head texture; input the corresponding first audio and the reference picture into the expression prediction model to obtain a second predicted expression sequence; and generate a reference predicted image based on the first head texture and the expression corresponding to the reference picture in the second predicted expression sequence; Based on the difference between the reference prediction image and the foreground of the reference picture, the parameters of the visual prediction model are adjusted.
4. The image processing method according to claim 3, wherein: The training of the image processing model based on the second video includes: Inputting the second audio into the expression prediction model to obtain a third predicted expression sequence; Extracting a target picture from the second picture sequence; inputting the target picture into the visual prediction model to obtain a second head texture; generating a target prediction image based on the second head texture and the expression corresponding to the target picture in the third predicted expression sequence; The image processing model is trained based on the third predicted expression sequence, the second picture sequence, the target predicted image and the target picture.
5. The image processing method according to claim 4, characterized in that The training of the image processing model based on the third predicted expression sequence, the second picture sequence, the target predicted image, and the target picture includes: extracting the target speaker's expression in each frame of the second frame sequence to obtain a target expression sequence; Adjusting parameters of the expression prediction model based on a difference between the third predicted expression sequence and the target expression sequence; Based on the difference between the target predicted image and the foreground of the target picture, the second head texture is adjusted, or based on the difference between the target predicted image and the foreground of the target picture, the parameters of the visual prediction model are adjusted.
6. The image processing method according to claim 4, wherein: The image processing model includes a texture harmony model, and training the image processing model based on the third predicted expression sequence, the second picture sequence, the target predicted image, and the target picture includes: In a case where a target predicted expression meeting a preset requirement exists in the third predicted expression sequence, inputting the target predicted expression and the second head texture into a texture blending model to obtain a third head texture; generating a blended image based on the third head texture and the target predicted expression; Extracting the original picture corresponding to the target predicted expression from the second picture sequence; Based on the difference between the blended image and the original picture, the parameters of the texture blending model are adjusted.
7. The image processing method according to claim 6, characterized in that: The image processing model includes an identity background generator, and the training of the image processing model based on the third predicted expression sequence, the second picture sequence, the target predicted image, and the target picture includes: For each non-target predicted expression in the third predicted expression sequence, generating a corresponding first predicted image based on the non-target predicted expression and the second head texture, and inputting the first predicted image into the identity background generator to obtain a corresponding first visual image; For each target predicted expression in the third predicted expression sequence, generating a corresponding second predicted image based on the target predicted expression and the corresponding third head texture, and inputting the second predicted image into the identity background generator to obtain a corresponding second visual image; obtaining a visual image sequence based on the first visual image and the second visual image; Based on the difference between the visual image sequence and the second picture sequence, the parameters of the identity background generator are adjusted.
8. The image processing method according to claim 5, wherein: The expression prediction model includes an audio feature extractor, an identity feature extractor, and an expression generator. Inputting the second audio into the expression prediction model to obtain a third predicted expression sequence includes: Inputting the second audio into the audio feature extractor to obtain a second audio feature; inputting the target picture into the identity feature extractor to obtain a target speaker identity feature; inputting the second audio feature and the target speaker identity feature into the expression generator to obtain the third predicted expression sequence; The adjusting the parameters of the expression prediction model based on the difference between the third predicted expression sequence and the target expression sequence includes: Based on the difference between the third predicted expression sequence and the target expression sequence, the parameters of the identity feature extractor and the expression generator are adjusted while the parameters of the audio feature extractor are frozen.
9. The image processing method according to claim 1, wherein The multiple first videos include audio information in multiple languages or motion information including multiple head movements; and the method uses FLAME Mesh for expression representation.
10. The image processing method according to claim 7, wherein: The step of inputting the target audio into an image processing model corresponding to the target speaker to obtain a target picture sequence includes: The expression prediction model obtains an expression sequence based on the target audio; For each non-target expression in the expression sequence that does not meet the preset requirements, generating a corresponding first foreground image based on the non-target expression and the second head texture, and inputting the first foreground image into the identity background generator to obtain a corresponding first image; For a target expression in the expression sequence that meets the preset requirements, input the target expression and the second head texture into the texture harmony model to obtain a target harmonious texture; generate a corresponding second foreground image based on the target expression and the target harmonious texture, and input the second foreground image into the identity background generator to obtain a corresponding second image; The target picture sequence is obtained based on the first image and the second image.
11. An image processing device, characterized in that: The device comprises: The generalization training module is configured to perform: Acquire a plurality of first videos corresponding to a plurality of reference speakers, wherein the first videos include a first audio recording the voice of the corresponding reference speaker and a first frame sequence representing a digital image of the reference speaker synchronized with the first audio; Training an expression prediction model based on the multiple first videos, wherein the expression prediction model is used to predict a correspondence between sound and expression; The personalized capability training module is configured to perform: Based on the multiple first videos and the expression prediction model, training a visual prediction model, wherein the visual prediction model is used to predict the correspondence between expressions and facial texture features; Obtaining a second video and an image processing model, wherein the second video includes a second audio recording a target speaker's voice and a second image sequence representing a digital image of the target speaker synchronized with the second audio, and the image processing model includes the expression prediction model and the visual prediction model; Training the image processing model based on the second video to obtain an image processing model corresponding to the target speaker; The image generation module is configured to perform: Get the target audio obtained by any speaker reciting the target text; By inputting the target audio into an image processing model corresponding to the target speaker, a target picture sequence is obtained, wherein the target picture sequence represents a digital image of the target speaker synchronized with the target audio.
12. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the image processing method according to any one of claims 1 to 10.
13. A computer-readable storage medium, characterized in that When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the image processing method according to any one of claims 1 to 10.
14. A computer program product, characterized in that The computer program product includes a computer program, which is stored in a readable storage medium. At least one processor of a computer device reads and executes the computer program from the readable storage medium, so that the computer device performs the image processing method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Digital human generation method and device, equipment and medium
CN113886641A
Digital human image generation method and device based on voice driving and storage medium
CN116758189A