Video synthesis method, device, storage medium and electronic device

By extracting phoneme sequences and converting them into target lip-type parameter sequences, the existing video synthesis methods are solved, and the problem of time-consuming, complex and low realism is achieved, achieving more efficient and realistic video synthesis effects.

CN116193052BActive Publication Date: 2025-05-27NETEASE (HANGZHOU) NETWORK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310183219.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-21
Publication Date
2025-05-27
Estimated Expiration
2043-02-21

AI Technical Summary

Technical Problem

The existing video synthesis methods are time-consuming and labor-intensive, the process is complex, and the authenticity of the synthesised video is poor, which is prone to misleading problems and affecting the user's visual experience.

Method used

By obtaining the pending text and converting it into an audio file, the phoneme sequence is extracted using the pre-trained phoneme prediction model, and converting it into the target lip parameter sequence based on the correspondence between the phoneme and lip parameter of the target anchor, and finally processing the lip region in the original video of the target anchor to generate a synthetic video.

Benefits of technology

It improves the efficiency of video synthesis and the authenticity of the display effect of the synthetic video, making the synthetic video more in line with the lip characteristics of the target anchor when speaking, and improves the user's visual experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116193052B_ABST
    Figure CN116193052B_ABST
Patent Text Reader

Abstract

The present disclosure provides a video synthesis method, apparatus, storage medium, and electronic device, which relate to the field of computer technologies. The video synthesis method includes: obtaining a text to be processed and converting the text to be processed into an audio file; inputting the audio file into a pre-trained phoneme prediction model to extract phonemes, so as to obtain a phoneme sequence corresponding to the audio file; converting the phoneme sequence into a target mouth shape parameter sequence corresponding to a target anchor according to the correspondence between the phonemes and mouth shape parameters of the target anchor; and processing the mouth shape region in the original video of the target anchor according to the target mouth shape parameter sequence to obtain a synthesized video about the target anchor. The present disclosure improves the efficiency of video synthesis and the authenticity of the display effect of the synthesized video.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] With the development of artificial intelligence technology, synthetic anchors can solve the problem of time-consuming and laborious traditional audio-visual production, and thus are widely used in news or short video production.

[0003] In related technologies, the provided video synthesis method usually first converts text into audio text, determines the lip shape at each moment according to the audio text to obtain a lip shape sequence, and finally generates lip shape images according to the lip shape sequence and combines them into the final video.

[0004] However, the video synthesis process provided in related technologies usually takes a lot of time, the production process is complex, and the authenticity of the synthesized video is poor, and problems such as continuity errors are likely to occur, affecting the visual experience of users. Summary of the Invention

[0005] The present disclosure provides a video synthesis method, a video synthesis device, a computer-readable storage medium and an electronic device, thereby at least to some extent improving the efficiency of video synthesis and the authenticity of the display effect of the synthesized video.

[0006] According to a first aspect of the present disclosure, there is provided a video synthesis method, including:

[0007] Obtain a text to be processed, and convert the text to be processed into an audio file;

[0008] Input the audio file into a pre-trained phoneme prediction model for phoneme extraction to obtain a phoneme sequence corresponding to the audio file, where the phoneme sequence includes the phonemes included in the audio file and the time point information of the phonemes in the audio file;

[0009] Convert the phoneme sequence into a target lip shape parameter sequence corresponding to the target anchor according to the correspondence between the phonemes and lip shape parameters of the target anchor;

[0010] Process the lip shape area in the original video of the target anchor according to the target lip shape parameter sequence to obtain a synthesized video of the target anchor.

[0011] According to a second aspect of the present disclosure, there is provided a video synthesis device, including:

[0012] An acquisition module, configured to obtain a text to be processed and convert the text to be processed into an audio file;

[0013] An extraction module, configured to input the audio file into a pre-trained phoneme prediction model for phoneme extraction to obtain a phoneme sequence corresponding to the audio file, where the phoneme sequence includes the phonemes included in the audio file and the time point information of the phonemes in the audio file;

[0014] A conversion module, configured to convert the phoneme sequence into a target lip parameter sequence corresponding to the target anchor according to the correspondence between the phonemes and lip parameters of the target anchor;

[0015] A processing module, configured to process the lip region in the original video of the target anchor according to the target lip parameter sequence to obtain a synthesized video of the target anchor.

[0016] According to a third aspect of the present disclosure, there is provided a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the method of the first aspect is implemented.

[0017] According to a fourth aspect of the present disclosure, there is provided an electronic device, including:

[0018] A processor; and

[0019] A memory for storing executable instructions of the processor;

[0020] Wherein, the processor is configured to execute the method of the first aspect by executing the executable instructions.

[0021] The technical solution of the present disclosure has the following beneficial effects:

[0022] On the one hand, the video synthesis method, device, storage medium and electronic device provided by the embodiments of the present disclosure do not need to rely on the correspondence between text and phonemes. Through a pre-trained phoneme prediction model, a phoneme sequence can be directly obtained, which can improve the acquisition efficiency of the phoneme sequence; on the other hand, since the phoneme prediction model determines the time point information of phonemes in the audio file, rather than the duration information of each phoneme of different anchors' oral broadcasts, the phoneme prediction model can be widely applied to different anchors, and the generalization ability of the phoneme prediction model is strong; on the other hand, according to the correspondence between the phonemes and lip parameters of the target anchor, a lip parameter sequence corresponding to the target anchor can be determined to obtain a synthesized video of the target anchor, making the synthesized video more conform to the lip characteristics of the target anchor when speaking and improving the visual effect of the synthesized video.

[0023] It should be understood that the above general description and subsequent detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure. Obviously, the accompanying drawings in the following description are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings based on these drawings without creative efforts.

[0025] Figure 1 Showing a schematic architecture diagram of a video synthesis system in this exemplary embodiment;

[0026] Figure 2 Showing a flowchart of a video synthesis method in this exemplary embodiment;

[0027] Figure 3 Showing the process of another video synthesis method in this exemplary embodiment;

[0028] Figure 4 Showing a block diagram of a video synthesis device in this exemplary embodiment;

[0029] Figure 5 Showing a schematic structural diagram of an electronic device in this exemplary embodiment. Detailed Embodiments

[0030] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, the exemplary embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be more complete and comprehensive, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The features, structures, or characteristics described can be combined in any suitable manner in one or more embodiments. In the following description, numerous specific details are provided to give a thorough understanding of the embodiments of the present disclosure. However, those skilled in the art will realize that one or more of the specific details can be omitted in practicing the technical solutions of the present disclosure, or other methods, components, devices, steps, etc. can be adopted. In other cases, well-known technical solutions are not shown or described in detail to avoid obscuring the various aspects of the present disclosure.

[0031] In addition, the accompanying drawings are only schematic illustrations of the present disclosure and are not necessarily drawn to scale. The same reference numerals in the figures denote the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0032] The flowcharts shown in the accompanying drawings are only illustrative and do not necessarily include all steps. For example, some steps can be further decomposed, while some steps can be combined or partially combined. Therefore, the actual execution order may change according to the actual situation.

[0033] In the related art, a solution for generating a synthetic video using the original video of the anchor is provided. When determining the lip sequence according to the audio text, it usually relies on the correspondence between phonemes and characters; it is necessary to first extract the text information in the audio text and rely on the correspondence between phonemes and characters to determine the phoneme sequence contained in the audio text; among them, the correspondence between phonemes and text also includes the duration of the phonemes associated with the characters when the anchor broadcasts the text, which is applicable to a specific anchor, resulting in a complex processing process for the audio-to-lip conversion and poor generalization ability.

[0034] At the same time, in the process of generating lip images according to the lips, three methods can be provided in the related art. The first is to obtain a three-dimensional face model through three-dimensional reconstruction, and then drive the lips of the three-dimensional face model to change through an animation driving method, and render the lips to obtain lip images to form a synthetic video; however, this method usually requires expensive three-dimensional model acquisition equipment, with a complex production process, long cycle, and the authenticity of the final rendering result being limited by the equipment accuracy and the production process level of the staff.

[0035] The second is to pre-establish a lip dataset, and then retrieve the lip image most similar to the lip shape in the lip dataset according to the lip shape, and use the retrieved lip image to fuse the lip area with the remaining facial areas to obtain a rendering result. However, this method requires a large amount of human and time costs to establish the lip dataset, and in the retrieval process, once there is no most similar lip image in the lip dataset, the video synthesis will fail.

[0036] The third uses a deep learning method to train a conversion network for lip shape and lip image, so as to obtain a real image after inputting the lip shape into the conversion network; however, this process lacks the guidance of three-dimensional face information and cannot show the subtle changes in the lips when the anchor is speaking, resulting in a poor authenticity of the generated image display effect, being easily exposed, and affecting the user's viewing experience.

[0037] In view of the above problems, an exemplary embodiment of the present disclosure provides a video synthesis method, which is applicable to video synthesis services. The application scenarios of this video synthesis method include, but are not limited to: during the process of synthesizing a news video of a target anchor, obtaining a text to be processed and converting the text to be processed into an audio file; inputting the audio file into a pre-trained phoneme prediction model for phoneme extraction to obtain a phoneme sequence corresponding to the audio file; converting the phoneme sequence into a target lip parameter sequence according to the correspondence between the phonemes and lip parameters of the target anchor; and processing the lip region in the original video of the target anchor according to the target lip parameter sequence to obtain a synthesized video of the target anchor. Among them, the phoneme sequence includes the phonemes contained in the audio file and the time point information of the phonemes in the audio file; since there is no need to rely on the correspondence between text and phonemes, the pre-trained phoneme prediction model can directly obtain the phoneme sequence, which can improve the acquisition efficiency of the phoneme sequence; since the phoneme prediction model determines the time point information of the phonemes in the audio file, rather than the duration information of each phoneme spoken by different anchors, the phoneme prediction model can be widely applied to different anchors, and the generalization ability of the phoneme prediction model is strong; further, the target lip parameter sequence corresponding to the target anchor can be determined according to the correspondence between the phonemes and lip parameters of the target anchor to obtain the synthesized video of the target anchor, making the synthesized video more conform to the lip characteristics of the target anchor when speaking and improving the visual effect of the synthesized video.

[0038] To implement the above service processing method, an exemplary embodiment of the present disclosure provides a video synthesis system. Figure 1 The schematic architecture diagram of the service processing system is shown. As Figure 1 shown, the service processing system 100 may include a server 110 and a terminal device 120. Among them, the server 110 may be a background server deployed by a video synthesis service provider. The terminal device 120 may be a terminal device used by a user who needs to synthesize a video. More specifically, the terminal device may be a smart phone, a personal computer, a tablet computer, etc. The server 110 and the terminal device 120 may be connected through a network to implement video synthesis.

[0039] It should be noted that the video synthesis solution provided in the embodiments of the present disclosure may be applied to a terminal device or a server; among them, when the video synthesis solution provided in the embodiments of the present disclosure is applied to the server, the terminal device 120 may send a video synthesis request to the server 110. The video synthesis request may include the text to be processed and the original video of the target anchor. After receiving the video synthesis request, the server 110 may parse the video synthesis request to obtain the text to be processed and the original video of the target anchor, and obtain the synthesized video of the target anchor according to the solution provided in the embodiments of the present disclosure, and send the synthesized video to the terminal device 120.

[0040] Figure 2 It is a flowchart of a video synthesis method shown according to an exemplary embodiment. In the embodiments of the present disclosure, taking the application of the video synthesis method in a terminal device as an example, the video synthesis method is described. As Figure 2 shown, it includes the following steps:

[0041] Step S201, obtain the text to be processed and convert the text to be processed into an audio file;

[0042] Step S202, input the audio file into a pre-trained phoneme prediction model for phoneme extraction to obtain a phoneme sequence corresponding to the audio file;

[0043] Among them, the phoneme sequence includes the phonemes contained in the audio file and the time point information of the phonemes in the audio file;

[0044] Step S203, according to the correspondence between the phonemes and lip parameter of the target anchor, convert the phoneme sequence into a target lip parameter sequence corresponding to the target anchor;

[0045] Step S204, process the lip region in the original video of the target anchor according to the target lip parameter sequence to obtain a synthesized video of the target anchor.

[0046] In summary, for the video synthesis method provided by the embodiments of the present disclosure, on the one hand, without relying on the correspondence between text and phonemes, through a pre-trained phoneme prediction model, a phoneme sequence can be directly obtained, which can improve the acquisition efficiency of the phoneme sequence; on the other hand, since the phoneme prediction model determines the time point information of the phonemes in the audio file, rather than the duration information of each phoneme pronounced by different anchors, the phoneme prediction model can be widely applied to different anchors, and the generalization ability of the phoneme prediction model is strong; on the other hand, according to the correspondence between the phonemes and lip parameter of the target anchor, a lip parameter sequence corresponding to the target anchor can be determined to obtain a synthesized video of the target anchor, making the synthesized video more conform to the lip characteristics of the target anchor when speaking and improving the visual effect of the synthesized video.

[0047] The following will specifically describe each step in Figure 2 respectively.

[0048] In step S201, the terminal device can obtain the text to be processed and convert the text to be processed into an audio file.

[0049] In the embodiments of the present disclosure, the text to be processed is the text of the text information in the synthesized video to be obtained, such as the text information in a weather forecast video.

[0050] In an alternative embodiment, the process by which the terminal device can obtain the text to be processed and convert the text to be processed into an audio file may include: The terminal device obtains the text to be processed in response to a video synthesis instruction and converts the text to be processed into an audio file. Among them, the process of converting the text to be processed into an audio file may be implemented based on a Text To Speech (TTS) conversion system.

[0051] It should be noted that, in the embodiments of the present disclosure, the text to be processed is pre-stored in the terminal device. The video synthesis instruction is an input operation for detecting the text address of the text to be processed in the video processing software after the terminal device runs the video processing software. It can be understood that after the terminal device obtains the text address of the text to be processed, it can obtain the text to be processed according to the storage location corresponding to the text address in the terminal device.

[0052] In step S202, the terminal device may input the audio file into a pre-trained phoneme prediction model for phoneme extraction to obtain a phoneme sequence corresponding to the audio file;

[0053] In the embodiments of the present disclosure, the phoneme prediction model is used to determine the phonemes in the audio file and the time points of the phonemes in the audio file; it can be understood that the phoneme sequence includes the phonemes contained in the audio file and the time point information of the phonemes in the audio file.

[0054] It should be noted that, in the embodiments of the present disclosure, a phoneme refers to the smallest speech unit divided according to the natural attributes of speech. It is divided according to the pronunciation actions in a syllable, and one pronunciation action constitutes one phoneme. Phonemes include two types: vowels and consonants. For Chinese characters, the syllable corresponding to the Chinese character "o" is "o", and this Chinese character corresponds to one phoneme. Another example is that the syllable corresponding to the Chinese character "ni" is "ni", and this Chinese character corresponds to two phonemes. Optionally, in the embodiments of the present disclosure, the phonemes used may include ten, including "a", "o", "e", "g", "zh", "i", "q", "f", "w", and "b". Any combination of these ten phonemes can represent the pronunciation of a character.

[0055] In an alternative embodiment, the terminal device inputs the audio file into a pre-trained phoneme prediction model for phoneme extraction to obtain a phoneme sequence corresponding to the audio file; for example, the phoneme sequence obtained by the terminal device may be Among them, is the phoneme at time t in the audio file 1 in the audio file, is the phoneme at time t in the audio file 2 in the audio file, and is the phoneme at time t in the audio file n in the audio file.

[0056] In step S203, the terminal device may convert the phoneme sequence into a target lip parameter sequence corresponding to the target anchor according to the correspondence between the phonemes and lip parameters of the target anchor.

[0057] In the embodiments of the present disclosure, for each anchor, the correspondence between the phonemes and lip parameters of each anchor may be pre-constructed in advance. The lip parameters are used to characterize three-dimensional lip features. For example, the lip parameters may be 3DMM (3D Morphable Face Model) lip parameters; the target anchor is the anchor that needs to appear in the finally obtained synthesized video.

[0058] In an alternative embodiment, the process by which the terminal device converts the phoneme sequence into a target lip parameter sequence corresponding to the target anchor according to the correspondence between the phonemes and lip parameters of the target anchor may include: obtaining the correspondence between the phonemes and lip parameters of the target anchor, and in the correspondence between the phonemes and lip parameters, searching for the target lip parameters corresponding to the phonemes at each time point in the phoneme sequence, and arranging the multiple target lip parameters in the order of the time points to obtain the target lip parameter sequence corresponding to the target anchor. Among them, the target lip parameters are used to characterize the three-dimensional lip features of the target anchor; for example, the target lip parameter sequence obtained by the terminal device may be Among them, represents the lip parameters of the target anchor at time t 1 , represents the lip parameters of the target anchor at time t 2 , and represents the lip parameters of the target anchor at time t n .

[0059] In an alternative embodiment, if the terminal device may be based on the input target identity identifier of the target anchor, the process by which the terminal device obtains the correspondence between the phonemes and lip parameters of the target anchor may include: in response to the obtained target identity information of the target anchor, searching in the correspondence information library of the phonemes and lip parameters for the correspondence between the phonemes and lip parameters corresponding to the target identity information to obtain the correspondence between the phonemes and lip parameters of the target anchor.

[0060] In an alternative embodiment, the video synthesis instruction may include the target identity identifier of the target anchor. The process by which the terminal device obtains the correspondence between the phonemes and lip parameters of the target anchor may include: based on the target identity identifier obtained by parsing the video synthesis instruction, searching in the correspondence information library of the phonemes and lip parameters for the correspondence between the phonemes and lip parameters corresponding to the target identity information to obtain the correspondence between the phonemes and lip parameters of the target anchor.

[0061] In step S204, the terminal device may process the mouth region in the original video of the target anchor according to the target mouth shape parameter sequence to obtain a synthesized video of the target anchor.

[0062] In the embodiments of the present disclosure, the target mouth shape parameter sequence is used to represent the mouth shape parameters of the target anchor at each time point in the video to be synthesized. Processing the mouth region in the original video of the target anchor based on the mouth shape parameters of the target anchor can obtain a synthesized video of the target anchor.

[0063] In an alternative embodiment, as Figure 3 shown, the process by which the terminal device processes the mouth region in the original video of the target anchor according to the target mouth shape parameter sequence to obtain a synthesized video of the target anchor may include:

[0064] Step S301, for each original video frame in the original video, determine the face shape parameter and the head pose parameter in the original video frame according to the face region in the original video frame;

[0065] In the embodiments of the present disclosure, the face shape parameter is used to represent the three-dimensional face shape feature, and the head pose parameter is used to represent the three-dimensional head pose feature.

[0066] In an alternative embodiment, the face shape parameter may be a 3DMM face shape parameter, and the head pose parameter may be a 3DMM head pose parameter. Then the process by which the terminal device determines the face shape parameter and the head pose parameter in the original video frame according to the face region in the original video frame may include: using 3DMM technology to fit the face in the original image for three-dimensional reconstruction to obtain the face shape parameter and the head pose parameter of the portrait in the original image.

[0067] Step S302, obtain the target mouth shape parameter at each time point in the target mouth shape parameter sequence, and determine the target three-dimensional face model corresponding to each time point according to the target mouth shape parameter and the face shape parameter corresponding to each time point;

[0068] In the embodiments of the present disclosure, the face shape parameter corresponding to each time point is the face shape parameter in the original video frame corresponding to each time point; the target three-dimensional face model corresponding to each time point is the three-dimensional face model of the target face in the video frame to be synthesized corresponding to each time point.

[0069] In an alternative embodiment, the process by which the terminal device determines the target three-dimensional face model corresponding to each time point based on the target mouth shape parameters and face shape parameters corresponding to each time point may include: performing face reconstruction based on the target mouth shape parameters and face shape parameters corresponding to each time point to obtain the target three-dimensional face model corresponding to each time point. Optionally, the face reconstruction process may be implemented based on the 3DMM technology.

[0070] Step S303: Determine the rendering position map corresponding to each time point according to the head pose parameters and the target three-dimensional face model corresponding to each time point;

[0071] In the embodiments of the present disclosure, the rendering position map corresponding to each time point includes the position information of the mouth area in the video frame to be synthesized corresponding to the time point; optionally, the position information may be UV coordinates.

[0072] In an alternative embodiment, the process by which the terminal device determines the rendering position map corresponding to each time point based on the head pose parameters and the target three-dimensional face model corresponding to each time point may include: projecting the target three-dimensional face model corresponding to each time point into a two-dimensional image according to the head pose parameters corresponding to each time point to obtain the rendering position map corresponding to each time point. Optionally, the projection process is implemented based on a mask of a partial face area, and the partial face area includes the mouth area, so as to obtain a rendering position map including the mouth area, thereby reducing the data processing amount during the generation of the synthesized video.

[0073] Step S304: Determine the synthesized video frame corresponding to each time point according to the rendering position map and the target video frame corresponding to each time point, and obtain the synthesized video of the target anchor;

[0074] In the embodiments of the present disclosure, the target video frame corresponding to each time point is the video frame other than the mouth area in the original video frame corresponding to the time point. The target three-dimensional face model may be reconstructed based on the face shape parameters and the target mouth shape parameters in the original video frame, and the mouth area to be rendered may be determined according to the reconstructed target three-dimensional face model, so as to render the synthesized video frame; the mouth area of the face in the video frame to be synthesized may be rendered according to the three-dimensional information, improving the naturalness and authenticity of the mouth area in the obtained synthesized video frame and enhancing the visual effect when the user views the synthesized video.

[0075] In an alternative embodiment, the process by which the terminal device determines the synthesized video frame corresponding to each time point based on the rendering position map and the target video frame corresponding to each time point to obtain the synthesized video of the target anchor may include: inputting the rendering position map corresponding to each time point into a pre-trained texture image for texture sampling to obtain a rendering color map; inputting the rendering color information map into a pre-trained rendering network for lip rendering to obtain a target lip map; and inputting the target lip map and the target video frame into a pre-trained fusion network for video frame fusion to obtain the synthesized video frame corresponding to each time point. Among them, the rendering color map includes the color information and feature information of the mouth area in the video frame to be synthesized corresponding to the time point. The rendering position map can be processed based on the texture image, the rendering network, and the fusion network to obtain the synthesized video frame corresponding to each time point, without the need to establish a complex three-dimensional model, which can reduce the complexity of obtaining the synthesized video frame. At the same time, since the rendering position map is determined based on three-dimensional information, the authenticity of the mouth area in the obtained synthesized video frame can be guaranteed.

[0076] It should be noted that in the embodiments of the present disclosure, the network structures of the rendering network and the fusion network can be determined according to actual needs, and the embodiments of the present disclosure do not limit this. For example, the rendering network can be a Unet network, and the fusion network can also be a Unet network. The size of the texture image can be an image of 1024×1024×16. The 16-channel image can carry more information to improve the authenticity of the determined target lip map.

[0077] In an alternative embodiment, the texture image, the rendering network, and the fusion network are trained in the following manner, including: for each original video frame in the original video, determining the facial shape parameters, head pose parameters, and original lip parameters in the original video frame according to the facial area in the original video frame; determining the sample three-dimensional facial model associated with each original video frame according to the original lip parameters and facial shape parameters associated with each original video frame; determining the sample rendering position map associated with each original video frame according to the head pose parameters and the sample three-dimensional facial model associated with each original video frame; and iteratively training the texture image, the rendering network, and the fusion network according to each original video frame and the sample rendering position map associated with each original video frame to obtain the trained texture image, rendering network, and fusion network. Among them, the sample rendering position map includes the position information of the mouth area in the original video frame. The texture image, the rendering network, and the fusion network can be trained according to the original video frames of the target anchor, making the rendering network and the fusion network more targeted when determining the synthesized video of the target anchor, so as to obtain a synthesized video that better conforms to the facial features of the anchor.

[0078] It should be noted that the process by which the terminal device determines the facial shape parameters, head pose parameters, and original mouth shape parameters in the original video frame based on the facial region in the original video frame is similar to the above-mentioned step S301; the process of determining the sample three-dimensional facial model associated with each original video frame based on the original mouth shape parameters and facial shape parameters associated with each original video frame is similar to the above-mentioned step S302; the process of determining the sample rendering position map associated with each original video frame based on the head pose parameters and the sample three-dimensional facial model associated with each original video frame is similar to the above-mentioned step S303, and the embodiments of the present disclosure will not elaborate on this.

[0079] In an alternative embodiment, the process by which the terminal device iteratively trains the rendering network based on each original video frame and the sample rendering position map associated with each original video frame includes: inputting the sample rendering position map associated with each original video frame into the texture image to be trained for texture sampling to obtain a sample rendering color map; inputting the sample rendering color information map into the rendering network to be trained for mouth shape rendering to obtain a predicted mouth shape map associated with each original video frame; and adjusting the image parameters of the texture image to be trained and the model parameters of the rendering network to be trained based on the predicted mouth shape map and the mouth shape map in the original video frame to obtain a trained rendering network. The texture image and the rendering network can be pre-trained to improve the efficiency of determining the target mouth shape image during the video synthesis process.

[0080] In an alternative embodiment, the process by which the terminal device iteratively trains the fusion network based on each original video frame and the sample rendering position map associated with each original video frame may include: inputting the predicted mouth shape map and the sample target video frame associated with each original video frame into the fusion network to be trained for video frame fusion to obtain a sample synthesized video frame, where the sample target video frame is the video frame in the original video frame except for the mouth region; further, adjusting the model parameters of the fusion network to be trained based on the sample synthesized video frame and the original video frame to obtain a trained fusion network. The fusion network can be pre-trained to improve the efficiency of determining the target synthesized video frame during the video synthesis process.

[0081] In an alternative embodiment, before processing the mouth region in the original video of the target anchor according to the target mouth shape parameter sequence to obtain a synthesized video of the target anchor, the terminal device may further: determine the target mouth shape parameter of the target time point between adjacent time points according to the target mouth shape parameters of adjacent time points in the target mouth shape parameter sequence; add the target mouth shape parameter of the target time point to the target mouth shape parameter sequence to obtain an updated target mouth shape parameter sequence; then the process by which the terminal device processes the mouth region in the original video of the target anchor according to the target mouth shape parameter sequence to obtain a synthesized video of the target anchor may include: processing the mouth region in the original video of the target anchor according to the updated target mouth shape parameter sequence to obtain a synthesized video of the target anchor. The target mouth shape parameter sequence may be smoothed to make the transition of the mouth region in the synthesized video frames obtained more natural, further enhancing the visual effect when the user views the synthesized video.

[0082] Among them, the process by which the terminal device determines the target mouth shape parameter of the target time point between adjacent time points according to the target mouth shape parameters of adjacent time points in the target mouth shape parameter sequence may include: inputting the target mouth shape parameters of adjacent time points in the target mouth shape parameter sequence into a preset formula to obtain the target mouth shape parameter of the target time point between adjacent time points. The preset formula (1) may be:

[0083]

[0084]

[0085] Among them, e t is the target mouth shape parameter of the target time point, is the target mouth shape parameter of the next time point adjacent to the target time point, is the mouth shape parameter of the next time point adjacent to the target time point; t is the target time point, t i+1 is the next time point adjacent to the target time point, and t i is the previous time point adjacent to the target time point.

[0086] In an alternative embodiment, after obtaining the synthesized video of the target anchor, post-processing such as tone adjustment, background replacement, logo or subtitle embedding, etc. may be performed on the synthesized video based on actual requirements to meet the video publishing requirements.

[0087] This disclosure provides an embodiment of a video synthesis device. As Figure 4 shown, the video synthesis device 400 may include:

[0088] An acquisition module 401, configured to acquire the text to be processed and convert the text to be processed into an audio file;

[0089] The extraction module 402 is configured to input an audio file into a pre-trained phoneme prediction model for phoneme extraction, obtaining a phoneme sequence corresponding to the audio file, where the phoneme sequence includes the phonemes contained in the audio file and the time point information of the phonemes in the audio file;

[0090] The conversion module 403 is configured to convert the phoneme sequence into a target lip parameter sequence corresponding to the target anchor according to the correspondence between the phonemes of the target anchor and the lip parameter;

[0091] The processing module 404 is configured to process the lip region in the original video of the target anchor according to the target lip parameter sequence, obtaining a synthesized video of the target anchor.

[0092] Optionally, the processing module 404 is configured to:

[0093] For each original video frame in the original video, determine the face shape parameter and the head pose parameter in the original video frame according to the face region in the original video frame;

[0094] Obtain the target lip parameter at each time point in the target lip parameter sequence, and determine the target three-dimensional face model corresponding to each time point according to the target lip parameter and the face shape parameter corresponding to each time point;

[0095] Determine the rendering position map corresponding to each time point according to the head pose parameter and the target three-dimensional face model corresponding to each time point, where the rendering position map corresponding to each time point includes the position information of the mouth region in the video frame to be synthesized corresponding to the time point;

[0096] Determine the synthesized video frame corresponding to each time point according to the rendering position map corresponding to each time point and the target video frame, obtaining a synthesized video of the target anchor, where the target video frame corresponding to each time point is the video frame other than the mouth region in the original video frame corresponding to the time point.

[0097] Optionally, the processing module 404 is configured to:

[0098] Input the rendering position map corresponding to each time point into a pre-trained texture image for texture sampling, obtaining a rendering color map, where the rendering color map includes the color information and feature information of the mouth region in the video frame to be synthesized corresponding to the time point;

[0099] Input the rendering color information map into a pre-trained rendering network for lip rendering, obtaining a target lip map;

[0100] Input the target mouth shape map and the target video frame into the pre-trained fusion network for video frame fusion to obtain a synthetic video frame corresponding to each time point.

[0101] Optionally, as Figure 4 shown, the video synthesis device 400 further includes a model training module 405, configured to:

[0102] For each original video frame in the original video, determine the face shape parameters, head pose parameters, and original mouth shape parameters in the original video frame according to the face region in the original video frame;

[0103] Determine a sample three-dimensional face model associated with each original video frame according to the original mouth shape parameters and face shape parameters associated with each original video frame;

[0104] Determine a sample rendering position map associated with each original video frame according to the head pose parameters and the sample three-dimensional face model associated with each original video frame. The sample rendering position map includes the position information of the mouth region in the original video frame;

[0105] Iteratively train the texture image, the rendering network, and the fusion network according to each original video frame and the sample rendering position map associated with each original video frame to obtain a trained rendering network and a trained fusion network.

[0106] Optionally, the model training module 405 is configured to:

[0107] Input the sample rendering position map associated with each original video frame into the texture image to be trained for texture sampling to obtain a sample rendering color map;

[0108] Input the sample rendering color information map into the rendering network to be trained for mouth shape rendering to obtain a predicted mouth shape map associated with each original video frame;

[0109] Adjust the image parameters of the texture image to be trained and the model parameters of the rendering network to be trained according to the predicted mouth shape map and the mouth shape map in the original video frame to obtain a trained rendering network.

[0110] Optionally, the model training module 405 is configured to:

[0111] Input the predicted mouth shape map associated with each original video frame and the sample target video frame into the fusion network to be trained for video frame fusion to obtain a sample synthetic video frame. The sample target video frame is the video frame in the original video frame except for the mouth region;

[0112] Adjust the model parameters of the fusion network to be trained according to the sample synthetic video frame and the original video frame to obtain a trained fusion network.

[0113] Optionally, as Figure 4 shown, the video synthesis device 400 further includes a smoothing module 406 configured to:

[0114] Determine the target lip parameters at the target time points between adjacent time points according to the target lip parameters at adjacent time points in the target lip parameter sequence;

[0115] Add the target lip parameters at the target time points to the target lip parameter sequence to obtain an updated target lip parameter sequence;

[0116] A processing module 404 configured to:

[0117] Process the lip region in the original video of the target anchor according to the updated target lip parameter sequence to obtain a synthesized video of the target anchor.

[0118] Exemplary embodiments of the present disclosure also provide a computer-readable storage medium, which can be implemented in the form of a program product, including program code. When the program product runs on an electronic device, the program code is used to cause the electronic device to execute the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section of this specification. In one embodiment, the program product can be implemented as a portable compact disc read-only memory (CD-ROM) and includes program code, and can run on an electronic device, such as a personal computer. However, the program product of the present disclosure is not limited to this. In this document, the readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0119] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0120] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device.

[0121] The program code contained on the readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0122] The program code for performing the operations of the present disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).

[0123] Exemplary embodiments of the present disclosure also provide an electronic device, which can be a terminal device or a server. The following refers to Figure 5 for an explanation of this electronic device. It should be understood that Figure 5 the electronic device 500 shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.

[0124] As Figure 5 shown, the electronic device 500 is presented in the form of a general-purpose computing device. The components of the electronic device 500 may include but are not limited to: at least one processing unit 510, at least one storage unit 520, and a bus 530 connecting different system components (including the storage unit 520 and the processing unit 510).

[0125] Among them, the storage unit stores program code, which can be executed by the processing unit 510, so that the processing unit 510 executes the steps according to various exemplary embodiments of the present invention described in the above "Exemplary Method" section of this specification. For example, the processing unit 510 can execute as Figure 2 or Figure 3The method steps shown, etc.

[0126] The storage unit 520 may include volatile storage units, such as a random access storage unit (RAM) 521 and / or a cache storage unit 522, and may further include a read-only storage unit (ROM) 523.

[0127] The storage unit 520 may also include a program / utilities 524 having a set (at least one) of program modules 525. Such program modules 525 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0128] The bus 530 may include a data bus, an address bus, and a control bus.

[0129] The electronic device 500 may also communicate with one or more external devices 600 (such as a keyboard, a pointing device, a Bluetooth device, etc.). Such communication may be performed through an input / output (I / O) interface 540. The electronic device 500 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 550. As shown in the figure, the network adapter 550 communicates with other modules of the electronic device 500 through the bus 530. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0130] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, such a division is not mandatory. In fact, according to the exemplary embodiments of the present disclosure, the features and functions of two or more modules or units described above may be embodied in one module or unit. Conversely, the features and functions of one module or unit described above may be further divided and embodied by multiple modules or units.

[0131] Those skilled in the art can understand that various aspects of the present disclosure can be implemented as a system, a method, or a program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here. After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily think of other implementations of the present disclosure. This application aims to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include well-known common general knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and the embodiments are only regarded as exemplary, and the true scope and spirit of the present disclosure are pointed out by the claims.

[0132] It should be understood that the present disclosure is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only defined by the appended claims.

Claims

1. A video synthesis method, characterized in that, it includes: Obtain the text to be processed and convert the text to be processed into an audio file; Input the audio file into a pre-trained phoneme prediction model for phoneme extraction to obtain a phoneme sequence corresponding to the audio file, where the phoneme sequence includes the phonemes contained in the audio file and the time point information of the phonemes in the audio file; According to the correspondence between the phonemes and lip movement parameters of the target anchor, convert the phoneme sequence into a target lip movement parameter sequence corresponding to the target anchor; For each original video frame in the original video of the target anchor, determine the face shape parameter and head pose parameter in the original video frame according to the face area in the original video frame; Obtain the target lip movement parameter at each time point in the target lip movement parameter sequence, and determine the target three-dimensional face model corresponding to each time point according to the target lip movement parameter and the face shape parameter corresponding to each time point; According to the head pose parameter and the target three-dimensional face model corresponding to each time point, determine the rendering position map corresponding to each time point, and the rendering position map corresponding to each time point includes the position information of the mouth area in the video frame to be synthesized corresponding to the time point; According to the rendering position map and the target video frame corresponding to each time point, determine the synthesized video frame corresponding to each time point to obtain the synthesized video of the target anchor, and the target video frame corresponding to each time point is the video frame other than the mouth area in the original video frame corresponding to the time point.

2. The method according to claim 1, characterized in that, The determining the synthesized video frame corresponding to each time point according to the rendering position map and the target video frame corresponding to each time point includes: Input the rendering position map corresponding to each time point into a pre-trained texture image for texture sampling to obtain a rendered color map, where the rendered color map includes the color information and feature information of the mouth area in the video frame to be synthesized corresponding to the time point; Input the rendered color map into a pre-trained rendering network for lip movement rendering to obtain a target lip movement map; Input the target lip movement map and the target video frame into a pre-trained fusion network for video frame fusion to obtain the synthesized video frame corresponding to each time point.

3. The method according to claim 2, characterized in that, The texture image, the rendering network and the fusion network are trained in the following way, including: For each original video frame in the original video, determine the face shape parameter, head pose parameter and original lip movement parameter in the original video frame according to the face area in the original video frame; According to the original lip movement parameter and the face shape parameter associated with each original video frame, determine the sample three-dimensional face model associated with each original video frame; Determine a sample rendering position map associated with each original video frame according to the head pose parameters and the sample three-dimensional face model associated with each original video frame, where the sample rendering position map includes position information of the mouth region in the original video frame; Iteratively train the texture image, the rendering network, and the fusion network according to each original video frame and the sample rendering position map associated with each original video frame to obtain a trained rendering network and a fusion network.

4. The method according to claim 3, wherein, Iteratively training the texture image and the rendering network according to each original video frame and the sample rendering position map associated with each original video frame includes: Input the sample rendering position map associated with each original video frame into the texture image to be trained for texture sampling to obtain a sample rendering color map; Input the sample rendering color map into the rendering network to be trained for mouth shape rendering to obtain a predicted mouth shape map associated with each original video frame; Adjust the image parameters of the texture image to be trained and the model parameters of the rendering network to be trained according to the predicted mouth shape map and the mouth shape map in the original video frame to obtain the trained texture image and rendering network.

5. The method according to claim 4, wherein, Iteratively training the fusion network according to each original video frame and the sample rendering position map associated with each original video frame includes: Input the predicted mouth shape map and the sample target video frame associated with each original video frame into the fusion network to be trained for video frame fusion to obtain a sample synthesized video frame, where the sample target video frame is the video frame in the original video frame except for the mouth region; Adjust the model parameters of the fusion network to be trained according to the sample synthesized video frame and the original video frame to obtain the trained fusion network.

6. The method according to claim 1, wherein, Before processing the mouth region in the original video of the target anchor according to the target mouth shape parameter sequence to obtain a synthesized video of the target anchor, the method further includes: Determine the target mouth shape parameter of the target time point between adjacent time points according to the target mouth shape parameters of adjacent time points in the target mouth shape parameter sequence; Add the target mouth shape parameter of the target time point to the target mouth shape parameter sequence to obtain an updated target mouth shape parameter sequence; The processing the mouth region in the original video of the target anchor according to the target mouth shape parameter sequence to obtain a synthesized video of the target anchor includes: Process the mouth region in the original video of the target anchor according to the updated target mouth shape parameter sequence to obtain a synthesized video of the target anchor.

7. A video synthesis device, wherein, comprises: An acquisition module, configured to acquire a text to be processed and convert the text to be processed into an audio file; An extraction module, configured to input the audio file into a pre-trained phoneme prediction model for phoneme extraction, to obtain a phoneme sequence corresponding to the audio file, where the phoneme sequence includes phonemes contained in the audio file and time point information of the phonemes in the audio file; A conversion module, configured to convert the phoneme sequence into a target lip parameter sequence corresponding to the target anchor according to the correspondence between the phonemes and lip parameter of the target anchor; A processing module, configured to, for each original video frame in the original video of the target anchor, determine a face shape parameter and a head pose parameter in the original video frame according to a face region in the original video frame; obtain a target lip parameter at each time point in the target lip parameter sequence, and determine a target 3D face model corresponding to each time point according to the target lip parameter and the face shape parameter corresponding to each time point; determine a rendering position map corresponding to each time point according to the head pose parameter and the target 3D face model corresponding to each time point, where the rendering position map corresponding to each time point includes position information of a mouth region in a to-be-synthesized video frame corresponding to the time point; Determine a synthesized video frame corresponding to each time point according to the rendering position map and a target video frame corresponding to each time point, to obtain a synthesized video of the target anchor, where the target video frame corresponding to each time point is a video frame other than the mouth region in the original video frame corresponding to the time point.

8. A computer-readable storage medium, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

9. An electronic device, characterized in that, comprising: a processor; and a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the method according to any one of claims 1 to 6 by executing the executable instructions.

Citation Information

Patent Citations

  • Video synthesis method, device and equipment and storage medium

    CN111741326A

  • Audio recognition method based on acoustic model and language model

    CN114171000A