Lip Shape Generation Method, Device, Equipment and Medium
By extracting sub-videos from the speaking video sample for audio and video separation and training the lip shape generation model, the problem that lip shape generation is limited by a specific speaker in the prior art is solved, and high-quality lip shapes are generated without restraint on any character, meeting the video creation needs of multiple users.
Patent Information
- Application Number
- CN202210439599.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2042-04-25
AI Technical Summary
The prior art is unable to generate lip shapes unrestrainedly on multiple users, resulting in video creation being limited by training of specific speakers and unable to adapt to the needs of multiple users.
By extracting sub-videos from different time periods from the speaking video sample for audio and video separation, the lip image and audio sequence are obtained, the lip shape generation model is trained, and the lip shape is generated on any character using the trained model.
It realizes the unrestrained generation of lip shapes on any character, meets the video creation needs of multiple users, and improves video quality and generation efficiency.
Smart Images

Figure CN114820891B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of video processing, and in particular, to a lip shape generation method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] With the rapid growth of short video content consumption, quickly creating video content has become a typical requirement, and the quality of the video largely depends on the generated lip shapes of the speakers.
[0003] In order to quickly construct high-quality videos, in related technologies, through deep learning, the mapping from language expression to lip landmarks is learned from a long-time speaking video of a single speaker. Since it is only trained with a specific speaker, new identities or voices cannot be synthesized.
[0004] However, in practical applications, video creation is expected to serve multiple users. Therefore, there is a need to provide a method that is independent of the speaker and can freely generate lip shapes on any person. Summary of the Invention
[0005] The main objective of the embodiments of the present application is to propose a lip shape generation method, apparatus, electronic device, and computer-readable storage medium that can freely generate lip shapes on any person.
[0006] To achieve the above objective, a first aspect of the embodiments of the present application proposes a lip shape generation method, and the method includes:
[0007] Obtain a speaking video sample containing the lip shapes of a speaker;
[0008] Extract a first sub-video and a second sub-video from the speaking video sample, where the first sub-video and the second sub-video are sub-videos of different time periods in the speaking video sample;
[0009] Separate the audio and video of the first sub-video to obtain a first lip shape image sequence and a first speaking audio sequence;
[0010] Separate the audio and video of the second sub-video to obtain a second lip shape image sequence;
[0011] Use the first speaking audio sequence and the second lip shape image sequence as the inputs of the lip shape generation model, and use the first lip shape image sequence as the result expected to be output by the lip shape generation model, and train the lip shape generation model to obtain a trained lip shape generation model;
[0012] Obtain an initial lip shape image sequence and a target speaking audio sequence, where the initial lip shape image sequence contains the lip shapes of a speaker;
[0013] Input the initial lip image sequence and the target speech audio sequence into the trained lip generation model to obtain a target lip image sequence in which the speaker's lip shape matches the target speech audio sequence.
[0014] According to the lip generation method provided by some embodiments of the present invention, training the lip generation model by using the first speech audio sequence and the second lip image sequence as inputs of the lip generation model, and using the first lip image sequence as the result expected to be output by the lip generation model, to obtain a trained lip generation model, includes:
[0015] Input the first speech audio sequence and the second lip image sequence into the lip generation model to obtain a predicted lip image sequence;
[0016] Based on the predicted lip image sequence and the first lip image sequence, determine the discrimination result of the first lip image sequence;
[0017] When the discrimination result meets the preset training end condition, end the training to obtain a trained lip generation model;
[0018] When the discrimination result does not meet the preset training end condition, update the parameters of the lip generation model according to the discrimination result, and continue to train the lip generation model until the discrimination result meets the preset training end condition.
[0019] According to the lip generation method provided by some embodiments of the present invention, the discrimination result includes the image authenticity probability;
[0020] The determining the discrimination result of the first lip image sequence based on the predicted lip image sequence and the first lip image sequence includes:
[0021] Obtain a preset image quality discriminator;
[0022] Input the predicted lip image sequence and the first lip image sequence into the image quality discriminator to obtain the image authenticity probability of the predicted lip image sequence relative to the first lip image sequence through the image quality discriminator.
[0023] According to the lip generation method provided by some embodiments of the present invention, the discrimination result includes the optical flow feature difference value;
[0024] The determining the discrimination result of the first lip image sequence based on the predicted lip image sequence and the first lip image sequence includes:
[0025] Obtain the first optical flow features of two adjacent frames in the predicted lip image sequence and the second optical flow features of two adjacent frames in the first lip image sequence;
[0026] Determine the optical flow feature difference value between the predicted lip image sequence and the first lip image sequence according to the first optical flow feature and the second optical flow feature.
[0027] According to the lip shape generation method provided by some embodiments of the present invention, the discrimination result includes the audio-visual synchronization rate;
[0028] The method further includes:
[0029] Obtain a trained SyncNet model;
[0030] Input the first speech audio sequence and the predicted lip image sequence into the SyncNet model to obtain the audio-visual synchronization rate of the speaker's lip shape in the predicted lip image sequence relative to the first speech audio sequence through the SyncNet model.
[0031] According to the lip shape generation method provided by some embodiments of the present invention, the lip shape generation model includes an audio encoder, an image encoder, and an image decoder, wherein,
[0032] The audio encoder is used to encode the input first speech audio sequence to obtain an audio representation vector, and the audio representation vector contains audio feature information;
[0033] The image encoder is used to encode the input second lip image sequence to obtain an image representation vector, and the image representation vector contains lip shape feature information;
[0034] The image decoder is used to decode the concatenated vector of the audio representation vector and the image representation vector to generate the predicted lip image sequence.
[0035] According to the lip shape generation method provided by some embodiments of the present invention, the method further includes:
[0036] Perform audio-visual merging on the target speech audio sequence and the target lip image sequence to obtain a target video.
[0037] To achieve the above object, a second aspect of the embodiments of the present application proposes a lip shape generation device, and the device includes:
[0038] A video sample acquisition module, configured to acquire a speech video sample including a speaker's lip shape;
[0039] A sub-video extraction module, configured to extract a first sub-video and a second sub-video from the speech video sample, where the first sub-video and the second sub-video are sub-videos of different time periods in the speech video sample;
[0040] A first audio-video separation module, configured to perform audio-video separation on the first sub-video to obtain a first lip image sequence and a first speech audio sequence;
[0041] A second audio-video separation module, configured to perform audio-video separation on the second sub-video to obtain a second lip image sequence;
[0042] A model training module, configured to use the first speech audio sequence and the second lip image sequence as inputs of a lip generation model, and use the first lip image sequence as the result expected to be output by the lip generation model, and train the lip generation model to obtain a trained lip generation model;
[0043] A sequence acquisition module, configured to acquire an initial lip image sequence and a target speech audio sequence, where the initial lip image sequence includes the lip shape of the speaker;
[0044] A lip generation module, configured to input the initial lip image sequence and the target speech audio sequence into the trained lip generation model to obtain a target lip image sequence in which the lip shape of the speaker matches the target speech audio sequence.
[0045] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, where the electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the method described in the first aspect above is implemented.
[0046] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, where the storage medium is a computer-readable storage medium for computer-readable storage, and the storage medium stores one or more computer programs, and the one or more computer programs can be executed by one or more processors to implement the method described in the first aspect above.
[0047] The present application provides a lip shape generation method, apparatus, electronic device, and computer-readable storage medium. The lip shape generation method obtains a speaking video sample containing the lip shape of a speaker, extracts a first sub-video and a second sub-video from different time periods in the speaking video sample, then separates the audio and video of the first sub-video and the second sub-video to obtain a first lip image sequence, a first speaking audio sequence corresponding to the first lip image sequence, and a second lip image sequence respectively. After that, the first speaking audio sequence and the second lip image sequence are used as inputs to a lip shape generation model, and the first lip image sequence is used as the result expected to be output by the lip shape generation model to train the lip shape generation model, obtaining a trained lip shape generation model. Then, an initial lip image sequence and a target speaking audio sequence are obtained. The initial lip image sequence contains the lip shape of the speaker, and the initial lip image sequence and the target speaking audio sequence are input into the trained lip shape generation model to obtain a target lip image sequence in which the lip shape of the speaker matches the target speaking audio sequence. By extracting the first speaking audio sequence and the second lip image sequence from different time periods in the same speaking video sample and using them as inputs to the lip shape generation model, and using the first lip image sequence in which the lip shape of the speaker matches the first speaking audio sequence as the result expected to be output by the lip shape generation model to train the lip shape generation model, the trained lip shape generation model is independent of the speaker and can perform lip shape generation unrestrictedly on any person. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is a schematic flowchart of a lip shape generation method provided by an embodiment of the present application;
[0049] Figure 2 is Figure 1 a schematic diagram of sub-steps of step S150 in;
[0050] Figure 3 is Figure 2 a schematic diagram of sub-steps of step S220 in;
[0051] Figure 4 is Figure 2 a schematic diagram of sub-steps of step S220 in;
[0052] Figure 5 is a schematic flowchart of a lip shape generation method provided by another embodiment of the present application;
[0053] Figure 6 is a schematic flowchart of a lip shape generation method provided by another embodiment of the present application;
[0054] Figure 7 is a schematic structural diagram of a lip shape generation apparatus provided by an embodiment of the present application;
[0055] Figure 8 This is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0056] In order to make the objectives, technical solutions and advantages of the present application more clearly understood, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0057] It should be noted that unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0058] First, the nouns involved in the present application are analyzed:
[0059] Optical Flow: The movement of the image corresponding in two consecutive frames of an image sequence caused by the movement of the target object or the camera is called optical flow. It is a 2D vector field that can be used to show the movement of a point from the first frame of the image to the second frame. By using the change of pixels in the time domain in the image sequence and the correlation between adjacent frames, the corresponding relationship between the previous frame and the current frame can be found, so as to calculate the movement information of the object between adjacent frames.
[0060] With the rapid growth of short video content consumption, quickly creating video content has become a typical requirement, and the quality of the video largely depends on the generated lip shapes of the speakers.
[0061] In order to quickly construct high-quality videos, in related technologies, through deep learning, the mapping from language expression to lip landmarks is learned from the long-term speaking videos of a single speaker. Since it is only trained on a specific speaker, new identities or voices cannot be synthesized.
[0062] However, in practical applications, video creation is expected to serve multiple users. Therefore, a method that is independent of the speaker and can freely generate lip shapes on any person needs to be provided.
[0063] Based on this, the embodiments of the present application provide a lip shape generation method, apparatus, electronic device and computer-readable storage medium, which can freely generate lip shapes on any person.
[0064] A lip shape generation method, apparatus, electronic device and computer-readable storage medium provided by the embodiments of the present application will be specifically described through the following embodiments. First, the lip shape generation method in the embodiments of the present application is described.
[0065] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0066] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0067] The lip shape generation method provided by the embodiments of this application can be applied to terminals, can also be applied to the server side, or can be software running on the terminal or the server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, etc.; the server side can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the lip shape generation method, etc., but is not limited to the above forms.
[0068] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment, where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0069] Please refer to Figure 1 , Figure 1 which shows a schematic flowchart of a lip shape generation method provided by the embodiments of this application. As Figure 1As shown, the lip shape generation method includes but is not limited to steps S110 to S170:
[0070] Step S110, obtaining a speaking video sample containing the lip shape of the speaker.
[0071] Step S120, extracting a first sub-video and a second sub-video from the speaking video sample, where the first sub-video and the second sub-video are sub-videos of different time periods in the speaking video sample.
[0072] Step S130, performing audio-visual separation on the first sub-video to obtain a first lip shape image sequence and a first speaking audio sequence.
[0073] Step S140, performing audio-visual separation on the second sub-video to obtain a second lip shape image sequence.
[0074] Step S150, using the first speaking audio sequence and the second lip shape image sequence as the inputs of the lip shape generation model, and using the first lip shape image sequence as the result expected to be output by the lip shape generation model, training the lip shape generation model to obtain a trained lip shape generation model.
[0075] Step S160, obtaining an initial lip shape image sequence and a target speaking audio sequence, where the initial lip shape image sequence contains the lip shape of the speaker.
[0076] Step S170, inputting the initial lip shape image sequence and the target speaking audio sequence into the trained lip shape generation model to obtain a target lip shape image sequence in which the lip shape of the speaker matches the target speaking audio sequence.
[0077] It can be understood that two sub-videos of different time periods, namely a first sub-video and a second sub-video, are randomly extracted from a speaking video sample containing the lip shape of the speaker, and the first sub-video and the second sub-video have the same time length. Audio-visual separation is performed on the first sub-video and the second sub-video, and the separated image video is divided into multiple image frames at a preset frame interval to obtain a lip shape image sequence and a speaking audio sequence corresponding to the lip shape image sequence.
[0078] Exemplarily, performing audio-visual separation on the first sub-video to obtain a first lip shape image sequence I r ,I r+1 ,…,I r+n and a first speaking audio sequence A r ,A r+1 ,…,A r+n , performing audio-visual separation on the second sub-video to obtain a second lip shape image sequence I t ,I t+1 ,…,I t+n, take the first lip image sequence I r , I r+1 , …, I r+n as the reference image sequence, and take the second lip image sequence I t , I t+1 , …, I t+n as the image sequence to be processed.
[0079] In some embodiments, as Figure 2 shown, the step S150: taking the first speech audio sequence and the second lip image sequence as the inputs of the lip generation model, taking the first lip image sequence as the result expected to be output by the lip generation model, and training the lip generation model to obtain a trained lip generation model, including but not limited to steps S210 to S240:
[0080] Step S210, inputting the first speech audio sequence and the second lip image sequence into the lip generation model to obtain a predicted predicted lip image sequence.
[0081] Step S220, based on the predicted lip image sequence and the first lip image sequence, determining the discrimination result of the first lip image sequence.
[0082] Step S230, when the discrimination result meets the preset training end condition, end the training to obtain a trained lip generation model.
[0083] Step S240, when the discrimination result does not meet the preset training end condition, update the parameters of the lip generation model according to the discrimination result, and continue to train the lip generation model until the discrimination result meets the preset training end condition.
[0084] In some embodiments, the lip generation model includes an audio encoder, an image encoder, and an image decoder, where
[0085] the audio encoder is used to encode the input first speech audio sequence to obtain an audio representation vector, and the audio representation vector contains audio feature information;
[0086] the image encoder is used to encode the input second lip image sequence to obtain an image representation vector, and the image representation vector contains lip feature information;
[0087] the image decoder is used to decode the concatenated vector of the audio representation vector and the image representation vector to generate the predicted lip image sequence.
[0088] Exemplarily, please refer to Figure 6 , Figure 6The flowchart in a lip shape generation method provided by another embodiment of the present application is shown. As Figure 6 shown, in step S210, the first speech audio sequence A r , A r+1 , …, A r+n , the second lip shape image sequence I t , I t+1 , …, I t+n are input into the lip shape generation model to obtain the predicted lip shape image sequence, that is, through the audio encoder E ψ to encode the first speech audio sequence A r , A r+1 , …, A r+n to obtain the corresponding audio representation vector R1, and through the image encoder to encode the first lip shape image sequence I t , I t+1 , …, I t+n to obtain the corresponding image representation vector R2, and then through the image decoder D ω to splice the audio representation vector R1 and the image representation vector R2 to obtain the spliced vector of the audio representation vector R1 and the image representation vector R2, and then decode the spliced vector to generate the predicted lip shape image sequence
[0089] It can be understood that in step S220, based on the predicted lip shape image sequence and the first lip shape image sequence, the discrimination result of the first lip shape image sequence is determined, and the parameters of the lip shape generation model are updated according to the discrimination result, that is, the lip shape generation model is optimized through the discrimination result. Among them, the discrimination result can be parameter information such as the image quality, audio synchronization, and image sequence smoothness of the predicted lip shape image sequence. By setting the training end condition corresponding to the discrimination result, when the discrimination result does not meet the corresponding training end condition, the parameters of the lip shape generation model are updated according to the discrimination result, and when the discrimination result meets the corresponding training end condition, the training ends to obtain the trained lip shape generation model.
[0090] In some embodiments, the discrimination result includes the image authenticity probability. Please refer to Figure 3 , Figure 3 which shows the schematic diagram of the sub-steps of step S220 in a lip shape generation method provided by an embodiment of the present application. As Figure 3 shown, the lip shape generation method includes but is not limited to steps S310 to S320.
[0091] Step S310, obtain a preset image quality discriminator.
[0092] Step S320: Input the predicted lip image sequence and the first lip image sequence into the image quality discriminator, so as to obtain the image authenticity probability of the predicted lip image sequence relative to the first lip image sequence through the image quality discriminator.
[0093] It can be understood that by obtaining a preset image quality discriminator to perform image authenticity discrimination on the predicted lip image sequence, inputting the predicted lip image sequence and the real image (i.e., the first lip image sequence) into the image quality discriminator, so as to output the image authenticity probability value of the predicted lip image sequence relative to the first lip image sequence through the image quality discriminator. Then, according to the image authenticity probability value, during the training process, update the parameters of the lip generation model to guide the lip generation model to generate a predicted lip image sequence with better image quality until the image authenticity probability obtained through the image quality discriminator meets the preset training end condition.
[0094] Exemplarily, as Figure 6 shown, input the predicted lip image sequence and the first lip image sequence I t ,I t+1 ,…,I t+n into the preset image quality discriminator D δ , so as to obtain the image authenticity probability of the predicted lip image sequence relative to the real image (the first lip image sequence) through the image quality discriminator D δ .
[0095] Exemplarily, the preset training end condition is 45%, that is, when the image authenticity probability obtained through the image quality discriminator D δ is less than 45%, obtain the corresponding Loss loss function according to the discrimination result and constrain the lip generation model, so that the lip generation model generates a more realistic lip image sequence; and when the image authenticity probability is equal to or greater than 45%, end the training to obtain the trained lip generation model.
[0096] It should be noted that the training end condition corresponding to the image authenticity probability can be adjusted according to the actual application situation, and no specific limitation is made here.
[0097] In some embodiments, the discrimination result includes an optical flow feature difference value. Please refer to Figure 4 , Figure 4 which shows a schematic diagram of sub-steps of step S220 in a lip generation method provided by an embodiment of the present application. As Figure 4 shown, this lip generation method includes but is not limited to steps S410 to S420.
[0098] Step S410: Obtain the first optical flow features of two adjacent frames in the predicted lip shape image sequence and the second optical flow features of two adjacent frames in the first lip shape image sequence.
[0099] Step S420: Determine the optical flow feature difference value between the predicted lip shape image sequence and the first lip shape image sequence according to the first optical flow feature and the second optical flow feature.
[0100] It can be understood that the optical flow features of two adjacent frames in the predicted lip shape image sequence and the first lip shape image sequence are obtained respectively, that is, the first optical flow feature and the second optical flow feature. The optical flow feature difference value between the predicted lip shape image sequence and the first lip shape image sequence is determined according to the first optical flow feature and the second optical flow feature. Then, according to the optical flow feature difference value, during the training process, the parameters of the lip shape generation model are updated to guide the lip shape generation model to generate a smoother and more fluent lip shape image sequence until the optical flow feature difference value of the predicted lip shape image sequence relative to the first lip shape image sequence meets the preset training end condition.
[0101] Exemplarily, as Figure 6 shown, input the predicted lip shape image sequence and the first lip shape image sequence I t , I t+1 , …, I t+n into the optical flow feature extractor D to obtain the optical flow feature difference value between two adjacent frames of the predicted lip shape image sequence and the real image (the first lip shape image sequence) through the optical flow feature extractor D.
[0102] In some embodiments, based on the optical flow feature difference between the predicted lip shape image sequence and the first lip shape image sequence, the L2 Loss mean square error function (MEAN Square Error, MSE) is used to constrain the lip shape generation model, so that the lip shape generation model generates a smooth and fluent lip shape image sequence.
[0103] It should be noted that the training end condition corresponding to the optical flow feature difference value can be adjusted according to the actual application situation, and no specific limitation is made here.
[0104] In some embodiments, the discrimination result includes the audio-video synchronization rate. Please refer to Figure 5 , Figure 5 which shows the schematic flow chart of a lip shape generation method provided by an embodiment of the present application. As Figure 5 shown, the lip shape generation method includes but is not limited to steps S510 to S520.
[0105] Step S510: Obtain the trained SyncNet model.
[0106] Step S520: Input the first speech audio sequence and the predicted lip image sequence into the SyncNet model to obtain the audio-visual synchronization rate of the speaker's lip shape in the predicted lip image sequence relative to the first speech audio sequence through the SyncNet model.
[0107] It can be understood that by pre-learning the synchronization rates of relevant datasets, a trained SyncNet model is obtained. Then, the first speech audio sequence and the predicted lip image sequence are input into the trained SyncNet model to obtain the audio-visual synchronization rate of the speaker's lip shape in the predicted lip image sequence relative to the first speech audio sequence through the SyncNet model.
[0108] Exemplarily, as Figure 6 shown, input the first speech audio A r , A r+1 , …, A r+n and the predicted lip image sequence into the audio-visual synchronizer S to obtain the audio-visual synchronization rate of the speaker's lip shape in the predicted lip image sequence relative to the first speech audio through the audio-visual synchronizer S.
[0109] Exemplarily, the preset training end condition is 95%. That is, when the audio-visual synchronization rate obtained through the audio-visual synchronizer S is less than 95%, obtain the corresponding Loss loss function according to the discriminant structure and constrain the lip generation model to make the lip generation model generate a lip image sequence with a higher audio-visual synchronization rate; when the audio-visual synchronization rate is equal to or greater than 95%, the training ends to obtain a trained lip generation model.
[0110] It should be noted that the training end condition corresponding to the audio-visual synchronization rate can be adjusted according to the actual application situation, and no specific limitation is made here.
[0111] In some embodiments, the lip generation method further includes:
[0112] Perform audio-visual merging on the target speech audio sequence and the target lip image sequence to obtain a target video.
[0113] It can be understood that through the trained lip generation model, a target lip image sequence in which the speaker's lip shape matches the target speech audio sequence is obtained. Then, performing audio-visual merging on the target speech audio sequence and the target lip image sequence can quickly construct a high-quality video with corresponding lips and synchronized audio to meet the demand for quickly creating video content brought by the rapid growth of short video content consumption.
[0114] In some embodiments, the discriminant results corresponding to the predicted lip image sequence in the lip generation method include the probability of image authenticity and the optical flow feature difference value.
[0115] It is understandable that based on the predicted lip image sequence and the first lip image sequence, the image authenticity probability of the predicted lip image sequence relative to the first lip image sequence and the optical flow feature difference value between the predicted lip image sequence and the first lip image sequence are obtained. During the training process of the lip generation model, the parameters of the lip generation model are updated based on the image authenticity probability and the optical flow feature difference value, so that the lip generation model generates a lip image sequence with higher image quality and smoothness, until the image authenticity probability and the optical flow feature difference value respectively meet the corresponding training end conditions, and a trained lip generation model is obtained.
[0116] Exemplarily, as Figure 6 shown, the predicted lip image sequence and the first lip image sequence I t , I t+1 , …, I t+n are input into a preset image quality discriminator D δ and an optical flow feature extractor D respectively, to obtain the image authenticity probability of the predicted lip image sequence relative to the real image (the first lip image sequence) and the optical flow feature difference value between two adjacent frames of the predicted lip image sequence and the real image (the first lip image sequence).
[0117] In some embodiments, the discrimination result corresponding to the predicted lip image sequence in the lip generation method includes an optical flow feature difference value and an audio-video synchronization rate.
[0118] It is understandable that the optical flow feature difference value between the predicted lip image sequence and the first lip image sequence is obtained based on the predicted lip image sequence and the first lip image sequence. At the same time, the audio-video synchronization rate between the speaker's lip shape in the predicted lip image sequence and the first speech audio sequence is obtained based on the first speech audio sequence and the predicted lip image sequence. During the training process of the lip generation model, the parameters of the lip generation model are updated based on the optical flow feature difference value and the audio-video synchronization rate, so that the lip generation model generates a lip image sequence with higher smoothness and audio-video synchronization rate, until the optical flow feature difference value and the audio-video synchronization rate respectively meet the corresponding training end conditions, and a trained lip generation model is obtained.
[0119] Exemplarily, as Figure 6 shown, the predicted lip image sequence and the first lip image sequence I t , I t+1 , …, I t+n are input into the optical flow feature extractor D, and the predicted lip image sequence and the first speech audio sequence A r , A r+1,…,A r+n Input into the audio - video synchronizer S, and respectively obtain the optical flow feature difference values between two adjacent frames of the predicted lip - shape image sequence and the real image (the first lip - shape image sequence), and the audio - video synchronization rate between the speaker's lip - shape in the predicted lip - shape image sequence and the first speech audio sequence.
[0120] In some embodiments, the discrimination results corresponding to the predicted lip - shape image sequence in the lip - shape generation method include the image authenticity probability, the optical flow feature difference value, and the audio - video synchronization rate.
[0121] It can be understood that, according to the predicted lip - shape image sequence and the first lip - shape image sequence, the image authenticity probability of the predicted lip - shape image sequence relative to the first lip - shape image sequence and the optical flow feature difference value between the predicted lip - shape image sequence and the first lip - shape image sequence are obtained. At the same time, according to the first speech audio sequence and the predicted lip - shape image sequence, the audio - video synchronization rate between the speaker's lip - shape in the predicted lip - shape image sequence and the first speech audio sequence is obtained. During the training process of the lip - shape generation model, the parameters of the lip - shape generation model are updated based on the image authenticity probability, the optical flow feature difference value, and the audio - video synchronization rate, so that the lip - shape generation model generates a lip - shape image sequence with higher image quality, video smoothness, and audio - video synchronization rate, until the image authenticity probability, the optical flow feature difference value, and the audio - video synchronization rate respectively meet the corresponding training end conditions, and a trained lip - shape generation model is obtained.
[0122] Exemplarily, as Figure 6 shown, input the predicted lip - shape image sequence and the first lip - shape image sequence I r ,I r+1 ,…,I r+n into the preset image quality discriminator D δ and the optical flow feature extractor D, and input the predicted lip - shape image sequence and the first speech audio sequence A r ,A r+1 ,…,A r+n into the audio - video synchronizer S, and respectively obtain the image authenticity probability of the predicted lip - shape image sequence relative to the real image (the first lip - shape image sequence), the optical flow feature difference values between two adjacent frames of the predicted lip - shape image sequence and the real image (the first lip - shape image sequence), and the audio - video synchronization rate between the speaker's lip - shape in the predicted lip - shape image sequence and the first speech audio sequence.
[0123] The following describes the lip - shape generation method provided by the embodiments of the present application through a specific embodiment:
[0124] See Figure 6, the method extracts the first sub-video and the second sub-video of different time periods from the speech video sample containing the lip shape of the speaker, and separates the audio and video of the first sub-video and the second sub-video to obtain the first lip shape image sequence I r ,I r+1 ,…,I r+n and the first speech audio sequence A r ,A r+1 ,…,A r+n , and the second lip shape image sequence I t ,I t+1 ,…,I t+n , input the first speech audio sequence A r ,A r+1 ,…,A r+n and the second lip shape image sequence I t ,I t+1 ,…,I t+n into the lip shape generation model, and respectively obtain the audio representation vector R1 corresponding to the first speech audio sequence and the image representation vector R2 corresponding to the second lip shape image sequence through the audio encoder E ψ and the image encoder , and use the image decoder D ω to decode the concatenated vector of the audio representation vector R1 and the image representation vector R2 to obtain the predicted lip shape image sequence
[0125] Based on the predicted lip shape image sequence, the first speech audio sequence and the first lip shape image sequence, through the image quality discriminator D δ , the optical flow feature extractor D and the audio-video synchronizer S, obtain the discriminant results in three aspects of image quality, video smoothness and audio-video synchronization rate, and use this to constrain the lip shape generation model to make it generate a lip shape image sequence with higher image quality, video smoothness and audio-video synchronization rate until the discriminant results respectively meet the corresponding training end conditions, end the training, and obtain the trained lip shape generation model. Finally, use the trained lip shape generation model for lip shape generation.
[0126] The present application provides a lip shape generation method. The lip shape generation method includes obtaining a speaking video sample containing the lip shape of a speaker, extracting a first sub-video and a second sub-video from the speaking video sample at different time periods, then performing audio-visual separation on the first sub-video and the second sub-video to respectively obtain a first lip image sequence, a first speaking audio sequence corresponding to the first lip image sequence, and a second lip image sequence. After that, taking the first speaking audio sequence and the second lip image sequence as inputs of a lip shape generation model, and taking the first lip image sequence as the result expected to be output by the lip shape generation model, training the lip shape generation model to obtain a trained lip shape generation model. Then, obtaining an initial lip image sequence and a target speaking audio sequence, where the initial lip image sequence contains the lip shape of the speaker, and inputting the initial lip image sequence and the target speaking audio sequence into the trained lip shape generation model to obtain a target lip image sequence in which the lip shape of the speaker matches the target speaking audio sequence. By extracting the first speaking audio sequence and the second lip image sequence at different time periods from the same speaking video sample and using them as inputs of the lip shape generation model, and taking the first lip image sequence in which the lip shape of the speaker matches the first speaking audio sequence as the result expected to be output by the lip shape generation model, the lip shape generation model is trained, so that the trained lip shape generation model is independent of the speaker and can freely generate lip shapes on any person.
[0127] Please refer to Figure 7 , an embodiment of the present application further provides a lip shape generation device 100, and the lip shape generation device 100 includes:
[0128] A video sample acquisition module 110, configured to acquire a speaking video sample containing the lip shape of a speaker;
[0129] A sub-video extraction module 120, configured to extract a first sub-video and a second sub-video from the speaking video sample, where the first sub-video and the second sub-video are sub-videos at different time periods of the speaking video sample;
[0130] A first audio-visual separation module 130, configured to perform audio-visual separation on the first sub-video to obtain a first lip image sequence and a first speaking audio sequence;
[0131] A second audio-visual separation module 140, configured to perform audio-visual separation on the second sub-video to obtain a second lip image sequence;
[0132] A model training module 150, configured to take the first speaking audio sequence and the second lip image sequence as inputs of a lip shape generation model, and take the first lip image sequence as the result expected to be output by the lip shape generation model, and train the lip shape generation model to obtain a trained lip shape generation model;
[0133] A sequence acquisition module 160 is configured to acquire an initial lip image sequence and a target speech audio sequence, where the initial lip image sequence includes the lips of a speaker.
[0134] A lip shape generation module 170 is configured to input the initial lip image sequence and the target speech audio sequence into the trained lip shape generation model to obtain a target lip image sequence in which the lips of the speaker match the target speech audio sequence.
[0135] In some embodiments, the lip shape generation device 100 further includes:
[0136] A video generation module is configured to perform audio-visual merging on the target speech audio sequence and the target lip image sequence to obtain a target video.
[0137] This application proposes a lip shape generation device. The lip shape generation device acquires a speaking video sample including the lips of a speaker through a video sample acquisition module, and uses a first audio-visual separation module and a second audio-visual separation module to extract a first sub-video and a second sub-video in different time periods from the speaking video sample. Then, the first sub-video and the second sub-video are subjected to audio-visual separation to respectively obtain a first lip image sequence, a first speech audio sequence corresponding to the first lip image sequence, and a second lip image sequence. After that, a model training module uses the first speech audio sequence and the second lip image sequence as inputs to the lip shape generation model, and uses the first lip image sequence as the result expected to be output by the lip shape generation model to train the lip shape generation model to obtain a trained lip shape generation model. Then, a sequence acquisition module acquires an initial lip image sequence and a target speech audio sequence, where the initial lip image sequence includes the lips of a speaker, and a lip shape generation module inputs the initial lip image sequence and the target speech audio sequence into the trained lip shape generation model to obtain a target lip image sequence in which the lips of the speaker match the target speech audio sequence. By extracting a first speech audio sequence and a second lip image sequence in different time periods from the same speaking video sample and using them as inputs to the lip shape generation model, and using the first lip image sequence in which the lips of the speaker match the first speech audio sequence as the result expected to be output by the lip shape generation model to train the lip shape generation model, the trained lip shape generation model is independent of the speaker and can perform lip shape generation freely on any person.
[0138] It should be noted that for the information interaction, execution process, etc. between the modules of the above device, since they are based on the same concept as the method embodiment of this application, their specific functions and the technical effects brought are specifically described in the method embodiment part and will not be elaborated here.
[0139] Please refer to Figure 8 , Figure 8The hardware structure of an electronic device provided by an embodiment of the present application is shown. The electronic device includes:
[0140] A processor 210, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant computer programs to implement the technical solutions provided by the embodiments of the present application;
[0141] A memory 220, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 220 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 220 and are called by the processor 210 to execute the lip shape generation method of the embodiments of the present application;
[0142] An input / output interface 230, which is used to implement information input and output;
[0143] A communication interface 240, which is used to implement communication interaction between this device and other devices. It can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); and a bus 250, which transmits information between each component of the device (such as the processor 210, the memory 220, the input / output interface 230, and the communication interface 240);
[0144] Among them, the processor 210, the memory 220, the input / output interface 230, and the communication interface 240 achieve communication connections with each other inside the device through the bus 250.
[0145] An embodiment of the present application also provides a storage medium. The storage medium is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more computer programs, and the one or more computer programs can be executed by one or more processors to implement the above lip shape generation method.
[0146] As a computer-readable storage medium, the memory can be used to store software programs and computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0147] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0148] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0149] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0150] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0151] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0152] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or similar expressions refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0153] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or can be integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0154] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0155] In addition, in each embodiment of this application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0156] If the assembled unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of each embodiment of this application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0157] The preferred embodiments of the embodiments of this application have been described above with reference to the accompanying drawings, which does not limit the scope of rights of the embodiments of this application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of this application shall be within the scope of rights of the embodiments of this application.
Claims
1. A lip shape generation method, characterized in that The method includes: Obtaining a speaking video sample containing the speaker's lip shape; Extracting a first sub-video and a second sub-video from the speaking video sample, where the first sub-video and the second sub-video are sub-videos of different time periods in the speaking video sample; Separating audio and video of the first sub-video to obtain a first lip image sequence and a first speaking audio sequence; Separating audio and video of the second sub-video to obtain a second lip image sequence; Inputting the first speaking audio sequence and the second lip image sequence into a lip generation model to obtain a predicted lip image sequence; Determining a discrimination result of the first lip image sequence based on the predicted lip image sequence and the first lip image sequence; When the discrimination result meets a preset training end condition, ending the training to obtain a trained lip generation model; Obtaining an initial lip image sequence and a target speaking audio sequence, where the initial lip image sequence contains the speaker's lip shape; Inputting the initial lip image sequence and the target speaking audio sequence into the trained lip generation model to obtain a target lip image sequence in which the speaker's lip shape matches the target speaking audio sequence; Wherein, the discrimination result includes an optical flow feature difference value, and determining the discrimination result of the first lip image sequence based on the predicted lip image sequence and the first lip image sequence includes: Obtaining a first optical flow feature of two adjacent frames in the predicted lip image sequence and a second optical flow feature of two adjacent frames in the first lip image sequence; Determining an optical flow feature difference value between the predicted lip image sequence and the first lip image sequence according to the first optical flow feature and the second optical flow feature.
2. The lip shape generation method according to claim 1, wherein The method further includes: When the discrimination result does not meet the preset training end condition, updating the parameters of the lip generation model according to the discrimination result, and continuing to train the lip generation model until the discrimination result meets the preset training end condition.
3. The lip shape generation method according to claim 2, wherein The discrimination result includes an image authenticity probability; Determining the discrimination result of the first lip image sequence based on the predicted lip image sequence and the first lip image sequence includes: Obtaining a preset image quality discriminator; Inputting the predicted lip image sequence and the first lip image sequence into the image quality discriminator to obtain an image authenticity probability of the predicted lip image sequence relative to the first lip image sequence through the image quality discriminator.
4. The lip shape generation method according to claim 2, wherein The discrimination result includes an audio-video synchronization rate; The method further includes: Obtaining a trained SyncNet model; Inputting the first speaking audio sequence and the predicted lip image sequence into the SyncNet model to obtain an audio-video synchronization rate of the speaker's lip shape in the predicted lip image sequence relative to the first speaking audio sequence through the SyncNet model.
5. The lip shape generation method according to claim 2, characterized in that, The lip generation model includes an audio encoder, an image encoder, and an image decoder, where The audio encoder is used to encode the input first speech audio sequence to obtain an audio representation vector, and the audio representation vector contains audio feature information; The image encoder is used to encode the input second lip image sequence to obtain an image representation vector, and the image representation vector contains lip feature information; The image decoder is used to decode the concatenated vector of the audio representation vector and the image representation vector to generate the predicted lip image sequence.
6. The lip shape generation method according to claim 1, characterized in that The method further includes: Performing audio-visual merging on the target speech audio sequence and the target lip image sequence to obtain a target video.
7. A lip shape generating device, characterized in that, The device includes: A video sample acquisition module, configured to acquire a speech video sample including the lips of a speaker; A sub-video extraction module, configured to extract a first sub-video and a second sub-video from the speech video sample, where the first sub-video and the second sub-video are sub-videos of different time periods in the speech video sample; A first audio-visual separation module, configured to perform audio-visual separation on the first sub-video to obtain a first lip image sequence and a first speech audio sequence; A second audio-visual separation module, configured to perform audio-visual separation on the second sub-video to obtain a second lip image sequence; A model training module, configured to input the first speech audio sequence and the second lip image sequence into a lip generation model to obtain a predicted lip image sequence; determining a discrimination result of the first lip image sequence based on the predicted lip image sequence and the first lip image sequence; when the discrimination result meets a preset training end condition, ending the training to obtain a trained lip generation model; A sequence acquisition module, configured to acquire an initial lip image sequence and a target speech audio sequence, where the initial lip image sequence includes the lips of a speaker; A lip generation module, configured to input the initial lip image sequence and the target speech audio sequence into the trained lip generation model to obtain a target lip image sequence in which the lips of the speaker match the target speech audio sequence; Wherein, the discrimination result includes an optical flow feature difference value, and determining the discrimination result of the first lip image sequence based on the predicted lip image sequence and the first lip image sequence includes: Obtaining a first optical flow feature of two adjacent frames in the predicted lip image sequence and a second optical flow feature of two adjacent frames in the first lip image sequence; Determining an optical flow feature difference value between the predicted lip image sequence and the first lip image sequence according to the first optical flow feature and the second optical flow feature.
8. An electronic device, characterized in that, Includes: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program, and the computer program is executed by the at least one processor so that the at least one processor can execute the lip generation method according to any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the lip generation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method, system and device for converting voice into lip shape and storage medium
CN111370020A
AI anchor video generation method and device, electronic equipment and storage medium
CN113256765A