Video generation method and device, electronic equipment and storage medium

By dynamically extracting and fusing image reference frames, high-definition processing, and adjusting the lip generation parameters of the video generation model, the problem of poor video quality in the digital human video lip synchronization solution is solved, and higher synchronization accuracy and image quality clarity are achieved.

CN120186428APending Publication Date: 2025-06-20SHENZHEN YISHIHUOLALA TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510244095.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing digital human video lip synchronization solution is difficult to take into account the accuracy of lip synchronization, image fusion coordination and image quality clarity, resulting in poor video quality.

Method used

By obtaining the reference frame video and audio to be converted, the image reference frame is dynamically extracted and inputted to the video generation model, generating the lip-shaped generation image, fused to generate the initial generated frame, performing high-definition processing, detecting the lip-shaped alignment deviation, and adjusting the lip-shaped generation parameters of the video generation model to generate the target video.

Benefits of technology

It improves the visual display effect of the video, adapts to the high-definition display needs, improves the synchronization accuracy and nature of the lip shape and audio data frames, reduces the impact of the inconsistency of the lip shape on the video quality, and further enhances the realism and credibility of the video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120186428A_ABST
    Figure CN120186428A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and apparatus, an electronic device and a storage medium. The method comprises the steps of obtaining a reference frame video and a to-be-converted audio; dynamically extracting a plurality of image reference frames in the reference frame video, inputting the image reference frames into a video generation model, and reasoning to generate a plurality of lip shape generation images of the audio to be converted; fusing each lip-shaped generated image with the corresponding image reference frame to obtain a plurality of initial generated frames; performing high-definition processing on each initial generation frame to obtain a plurality of image generation frames; detecting lip shape alignment deviation of each image generation frame, determining a target reference frame, and adjusting lip shape generation parameters of a video generation model according to the target reference frame, so that the remodeled lip shape is aligned with the corresponding audio data frame, and a target video is generated; according to the method and the device, the visual display effect can be improved through high-definition processing of the initial generation frame so as to adapt to the high-definition display requirement, the lip shape is remolded after the lip shape generation parameters are reversely adjusted through the target reference frame, and the influence of lip shape discordance on the video quality is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of audio - video processing, and in particular, to a video generation method, apparatus, electronic device, and storage medium. Background Art

[0002] Currently, in the field of digital human lip - sync, most are based on deep learning technologies. For example, the voice - to - lip conversion model (Wav2Lip) analyzes the input speech features and reference - frame video to generate lip - shape movements synchronized with the speech features, thereby achieving highly realistic digital human videos.

[0003] However, in the actual application process, whether the video lip - sync of a digital human is coordinated depends on the matching accuracy between the speech features and the lip - shape movements. At the same time, the videos generated by this model require high - quality image synthesis and post - processing. In other words, the existing video lip - sync solutions for digital humans are difficult to balance in terms of lip - sync accuracy, image fusion coordination, and image quality clarity. Summary of the Invention

[0004] The present invention provides a video generation method, apparatus, electronic device, and storage medium to solve the technical problem of poor video quality in the above - mentioned video lip - sync solution.

[0005] In an embodiment of this application, a video generation method is provided, including: obtaining a reference - frame video and an audio to be converted; dynamically extracting a plurality of image reference frames from the reference - frame video and inputting them into a video generation model to infer and generate a plurality of lip - shape generation images for the audio to be converted; fusing each of the lip - shape generation images with the corresponding image reference frame to obtain a plurality of initial generation frames; performing high - definition processing on each of the initial generation frames to obtain a plurality of image generation frames; detecting the lip - shape alignment deviation of each of the image generation frames, determining a target reference frame, and adjusting the lip - shape generation parameters of the video generation model according to the target reference frame so that the reshaped lip - shape aligns with the corresponding audio data frame to generate a target video.

[0006] In an embodiment of this application, dynamically extracting a plurality of image reference frames from the reference - frame video includes: reversing the frame sequence of a plurality of image reference frames in the reference - frame video to obtain a reversed - frame video; segmenting the audio to be converted into a plurality of segments to be generated; determining the reference source of each segment to be generated according to the alternating reference between the reversed - frame video and the reference - frame video for input into the video generation model.

[0007] In one embodiment of the present application, determining the reference source of each to-be-generated segment according to the alternating reference of the inverted frame video and the reference frame video includes: using the inverted frame video as the reference source of the to-be-generated segments of odd segments, and using the reference frame video as the reference source of the to-be-generated segments of even segments; or, using the reference frame video as the reference source of the to-be-generated segments of odd segments, and using the inverted frame video as the reference source of the to-be-generated segments of even segments.

[0008] In one embodiment of the present application, detecting the lip alignment deviation of each image generation frame and determining the target reference frame includes: detecting the lip alignment deviation between each image generation frame and the corresponding audio data frame; if a lip alignment deviation is detected in an image generation frame, determining the target reference frame corresponding to the image generation frame; wherein, the target reference frame includes one of a first reference frame, a second reference frame, and a third reference frame, the first reference frame is obtained based on an image reference frame with a closed lip state, the second reference frame is generated based on the audio data frame corresponding to the image generation frame with a lip alignment deviation, and the third reference frame is obtained based on the image difference between each image reference frame and the image generation frame with the lip alignment deviation.

[0009] In one embodiment of the present application, determining the target reference frame corresponding to the image generation frame includes: if the first reference frame exists in the reference frame video, confirming the first reference frame as the target reference frame; if the first reference frame does not exist in the reference frame video, confirming the second reference frame or the third reference frame as the target reference frame.

[0010] In one embodiment of the present application, confirming the second reference frame or the third reference frame as the target reference frame includes: extracting the spectral feature value corresponding to the audio data frame for voice state confirmation; if the voice state is characterized as silent, generating the second reference frame and determining the second reference frame as the target reference frame; if the voice state is characterized as active, determining the third reference frame as the target reference frame.

[0011] In one embodiment of the present application, fusing the lip generation image with the corresponding image reference frame to obtain an initial generation frame includes: creating a first mask with the same size as the lip generation image; drawing a rectangular area in the central area of the first mask; performing a transition process on the rectangular area to obtain a target mask; fusing the lip generation image and the corresponding image reference frame in channels based on the normalized weight of the target mask to obtain an initial generation frame.

[0012] The present application provides a video generation device, including: an audio-video acquisition module for acquiring a reference frame video and an audio to be converted; a lip shape generation module for dynamically extracting a plurality of image reference frames from the reference frame video and inputting them into a video generation model to infer and generate a plurality of lip shape generation images for the audio to be converted; an image fusion module for fusing each of the lip shape generation images with the corresponding image reference frame respectively to obtain a plurality of initial generation frames; a high-definition processing module for performing high-definition processing on each of the initial generation frames to obtain a plurality of image generation frames; and a lip shape alignment module for detecting the lip shape alignment deviation of each of the image generation frames, determining a target reference frame, and adjusting the lip shape generation parameters of the video generation model according to the target reference frame so that the reshaped lip shapes are aligned with the corresponding audio data frames to generate a target video.

[0013] In an embodiment of the present application, the present application provides an electronic device, which includes: one or more processors; a storage device for storing one or more programs, and when the one or more programs are executed by the one or more processors, the electronic device realizes the video generation method as described in any one of the above.

[0014] In an embodiment of the present application, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor of a computer, the computer executes the video generation method as described in any one of the above.

[0015] The beneficial effects of the embodiments of the present invention: The present invention provides a video generation method, device, electronic device and storage medium. By performing high-definition processing on the initial generation frames, the visual display effect can be improved in the embodiments of the present invention to adapt to the high-definition display requirements and enhance the visual display effect; and, by determining the target reference frame and adjusting the lip shape generation parameters of the video generation model, the lip shapes can be reshaped in reverse to align with the audio data frames, making the generated lip shapes more accurate and natural in synchronization with the audio data frames, reducing the impact of uncoordinated lip shapes on the video quality, and further enhancing the video realism and credibility.

[0016] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing the embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application. Obviously, the accompanying drawings in the following description are only some embodiments of the present application, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts. In the drawings:

[0018] Figure 1Shows a schematic flowchart of a video generation method according to an embodiment of the present application;

[0019] Figure 2 Shows a schematic flowchart of a video generation method according to another embodiment of the present application;

[0020] Figure 3 Shows a block diagram of a video generation device according to an embodiment of the present application;

[0021] Figure 4 Shows a schematic diagram of the structure of a computer system of an electronic device suitable for implementing the embodiments of the present application. Detailed implementation manners

[0022] The following uses specific specific examples to illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0023] It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present application in a schematic manner. Therefore, only the components related to the present application are shown in the diagrams, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0024] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present application. However, it is obvious to those skilled in the art that the embodiments of the present application can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present application difficult to understand.

[0025] To solve the above technical problems, the present application provides a video generation method, device, electronic device, and storage medium. The implementation details of the technical solutions of the embodiments of the present application are elaborated in detail below.

[0026] Please refer to Figure 1 , Figure 1 Shows a schematic flowchart of a video generation method according to an embodiment of the present application. As Figure 1 shown, in an exemplary embodiment, the video generation method at least includes steps S110 to S150, which are introduced in detail as follows:

[0027] Step S110: Obtain a reference frame video and an audio to be converted.

[0028] In an embodiment of the present application, the reference frame video is a pre-prepared video reference segment, which contains a face area for reference. Among them, the reference frame video can be actively submitted after the user reads the video generation instructions. This is only an example.

[0029] In an embodiment of the present application, the audio to be converted is the audio that the user expects to convert into a target video. Among them, the audio to be converted can be obtained by the user's actual recording or generated by the text provided by the user. This is only an example.

[0030] Step S120: Dynamically extract multiple image reference frames from the reference frame video and input them into a video generation model to infer and generate multiple lip shape generation images of the audio to be converted.

[0031] In an embodiment of the present application, dynamically extracting multiple image reference frames from the reference frame video includes: reversing the frame sequence of multiple image reference frames in the reference frame video to obtain a reversed frame video; segmenting the audio to be converted into multiple segments to be generated; and determining the reference source of each segment to be generated according to the alternating reference of the reversed frame video and the reference frame video, so as to input it into the video generation model.

[0032] In an embodiment of the present application, determining the reference source of each segment to be generated according to the alternating reference of the reversed frame video and the reference frame video includes: using the reversed frame video as the reference source of the odd-numbered segments to be generated, and using the reference frame video as the reference source of the even-numbered segments to be generated; or, using the reference frame video as the reference source of the odd-numbered segments to be generated, and using the reversed frame video as the reference source of the even-numbered segments to be generated.

[0033] In an embodiment of the present application, in the audio-visual generation scenario, the image reference frame plays a key reference role. In the related art, due to the insufficient segment duration of the reference frame video, it is impossible to meet the actual requirements of generating the target video, resulting in difficulties in picture connection. The picture frequently shows breaks and freezes, seriously destroying the video coherence and greatly affecting the viewing experience.

[0034] In an embodiment of the present application, the video generation model includes a voice-to-lip conversion model (Wav2Lip).

[0035] In an embodiment of the present application, reversing the frame sequence means extracting the image reference frames in reverse to obtain a reversed frame video.

[0036] In one embodiment of the present application, for even-numbered segments of the to-be-generated segments, the picture information is accurately extracted frame by frame strictly following the original order of the reference-frame video, and the audio features of the audio data frames corresponding to the corresponding time periods are immediately associated and input into Wav2Lip in an orderly manner, providing a stable and coherent data source for Wav2Lip; for odd-numbered segments of the to-be-generated segments, the frame video is reversed and the audio features of the audio data frames are closely matched and then transported to Wav2Lip. The limited resources of the reference-frame video can be compensated by the fetching mode of alternating odd and even numbers and complementing positive and reverse orders. Wav2Lip can also continuously obtain a highly compatible reference source, ensuring smooth connection of the pictures, overcoming the picture connection problem caused by insufficient duration of the reference-frame video, improving the video quality, and enhancing the video viewing experience and user experience.

[0037] In one embodiment of the present application, the reversed frame video can also be referred to in the even-numbered segments of the to-be-generated segments, and the original order of the reference-frame video can be referred to in the odd-numbered segments of the to-be-generated segments.

[0038] In one embodiment of the present application, when inputting into the video generation model, it includes: corresponding reference sources need to be input into the video generation model for each to-be-generated segment.

[0039] Step S130, respectively fuse each lip shape generation image with the corresponding image reference frame to obtain a plurality of initial generation frames.

[0040] In one embodiment of the present application, fusing the lip shape generation image with the corresponding image reference frame to obtain the initial generation frame includes: creating a first mask with the same size as the lip shape generation image; drawing a rectangular area in the central area of the first mask; performing a transition process on the rectangular area to obtain the target mask; based on the normalized weight of the target mask, fuse the lip shape generation image and the corresponding image reference frame channel by channel to obtain the initial generation frame.

[0041] In one embodiment of the present application, during the process of fitting the lip shape generation image generated by Wav2Lip inference with the image reference frame, the problem of edge defects is particularly prominent. When directly superimposing, there are often abrupt black edges. The reason is that there are significant differences in color, texture, and edge characteristics between the lip shape generation image and the image reference frame, resulting in the inability of the two to be naturally fused. Visually, it seems like "two separate layers", seriously reducing the overall visual perception of the video, and the problem of detail loss also follows, greatly reducing the refinement of the picture.

[0042] In one embodiment of the present application, the above-mentioned image fusion technology is called edge feathering fusion, which can make the fusion edge of the lip shape generation image and the target image soft.

[0043] In one embodiment of the present application, the image reference frame fused with a lip shape generation image is used as the target image.

[0044] In one embodiment of the present application, mask creation includes: creating a first mask with all zeros and the same size as the lip generation image, drawing a rectangular area at the center of the second mask by using a rectangle drawing function and combining with a feather width set by fusing the expected and image edge characteristics, where the rectangular area is a white rectangle, and performing Gaussian blur on the rectangular area to shape an edge feathered mask, that is, obtaining the target mask.

[0045] In one embodiment of the present application, channel-by-channel fusion includes: focusing on each color channel of the lip generation image and the target image, and constructing a fusion formula according to the weight of the target mask as follows:

[0046] f2[y1:y2,x1:x2,c] = f[y1:y2,x1:x2,c] × (1.0 - mask[:,:,c] / 255.0) +

[0047] p[:,:,c] × (mask[:,:,c] / 255.0) Equation (1)

[0048] where f2[y1:y2,x1:x2,c] is the pixel value of the face region in the target image on color channel c after fusion, f[y1:y2,x1:x2,c] is the pixel value of the face region in the target image on color channel c before fusion, mask[:,:,c] is the pixel value of the target mask on color channel c, and p[:,:,c] is the pixel value of the lip generation image on color channel c

[0049] In one embodiment of the present application, color channel c can take values from 0 to 2 to represent different color channels, such as red, green, and blue.

[0050] In one embodiment of the present application, the pixel value of the target mask on color channel c is normalized from the range of [0, 255] to [0, 1] by mask[:,:,c] / 255.0 to introduce a part of the lip generation image.

[0051] In one embodiment of the present application, 1.0 - mask[:,:,c] / 255.0 is used to retain a part of the target image.

[0052] In one embodiment of the present application, the pixel values of all color channels are obtained after addition fusion to achieve a soft edge transition.

[0053] In one embodiment of the present application, the lip-generated image is fused with the corresponding image reference frame to obtain an initial generated frame, which further includes: preprocessing the lip-generated image, creating a second mask with the same size as the lip-generated image, and fusing the preprocessed lip-generated image with the target image based on the second mask and the central position to obtain the initial generated frame, where the central position is used to represent the central position information of the lip-generated image in the target image.

[0054] In one embodiment of the present application, the above image fusion technology is called Poisson fusion, which can make the lip-generated image and the target image be fused with a natural boundary transition.

[0055] In one embodiment of the present application, the preprocessing of the lip-generated image includes: blurring the lip-generated image through a mean blur function, accurately adjusting the blur parameters according to the image noise level and edge sharpness, effectively eliminating noise points and softening edge serrations; and, by virtue of an intelligent pixel traversal algorithm, accurately locking black pixels. For example, replacing the black pixel [0, 0, 0] represented in the RGB format with a preset color value, such as [235, 166, 164], to avoid the interference of black pixels on the fusion effect.

[0056] In one embodiment of the present application, the creation of the second mask includes: constructing a second mask with the same size as the lip-generated image and all values of 255, performing Gaussian blur on the first mask through a Gaussian blur function, and flexibly adjusting the parameters of the Gaussian blur based on the image edge complexity to create a suitable soft-edge mask to help the lip-generated image and the target image be seamlessly connected.

[0057] In one embodiment of the present application, the fusion process includes comprehensively considering the sizes and content layouts of the lip-generated image and the target image, accurately positioning the central position of the lip-generated image in the target image, and calling a seamless cloning function, using the preprocessed lip-generated image, the target image, the first mask after Gaussian blur, and the central position as key parameters to achieve a natural and smooth fusion effect.

[0058] In one embodiment of the present application, the lip-generated image is fused with the corresponding image reference frame to obtain an initial generated frame, which further includes: respectively cropping the face regions in the lip-generated image and the target image, and respectively performing multiple downsamplings on the cropped face image and the lip-cropped image to obtain an image pyramid, performing size matching adjustment on the face image and the lip-cropped image at each resolution in the image pyramid, performing weighted fusion on the adjusted face image and the lip-cropped image at each resolution, and performing upsampling and position restoration of the fused image to the target image to obtain the initial generated frame.

[0059] In one embodiment of the present application, the above image fusion technology is called Laplacian pyramid fusion, which can balance the fusion of the lip-shaped generated image and the target image while retaining details.

[0060] In one embodiment of the present application, the face region in the target image is accurately cropped to obtain a face image, and the lip-shaped generated image is cropped to obtain a lip-shaped cropped image. For example, the face image can be cropped according to the vertical and horizontal coordinate ranges of the face image in the target image. The pyramid downsampling function is called, and the downsampling parameters are flexibly adjusted according to the image resolution and size characteristics. The downsampling parameters include the resolution, and a hierarchical downsampled image, that is, an image pyramid, is obtained.

[0061] In one embodiment of the present application, size matching adjustment: By resetting the image size function, referring to the size of the downsampled image with the same downsampling parameters, an adaptive interpolation algorithm is selected to accurately adjust the size of the downsampled lip-shaped cropped image to ensure that it fits the corresponding downsampled face image.

[0062] In one embodiment of the present application, weighted fusion includes: According to the image content weight and visual expectation, the fusion weights of the weighted superposition function are reasonably set, such as 0.8 and 0.2, and the adjusted face image and lip-shaped cropped image in each resolution are weighted and fused to balance multi-dimensional information such as color and brightness.

[0063] In one embodiment of the present application, upsampling and position restoration include upsampling the fused image (res) through the pyramid upsampling function to restore it to the original size (for example, the restored width is represented by size[1], and the restored height is represented by size[0]), carefully controlling the upsampling parameters to ensure detail restoration, and then accurately positioning it to the corresponding position of the face region in the target image through size[1] and size[0].

[0064] In one embodiment of the present application, a variety of image fusion technologies provide comprehensive solutions for different fusion requirements and scenarios, enabling the generated lip-shaped generated image and the image reference frame to be seamlessly fused, and both the edge transition and detail performance are more natural and delicate, greatly improving the quality of video visual display.

[0065] Step S140, perform high-definition processing on each initial generated frame to obtain multiple image generated frames.

[0066] In one embodiment of the present application, in the generation process of the target video, the poor clarity of the training set materials is a major constraint, resulting in limited clarity of the lip-shaped generated image generated by model inference, directly leading to blurred image quality and hazy details in the finally synthesized target generated video, making it difficult to adapt to the high-definition display requirements, greatly reducing the visual effect, and unable to meet the current stringent demands for high-quality video content.

[0067] In one embodiment of the present application, the present application performs high-definition processing on the initial generated frame through a preset multi-layer neural network.

[0068] In one embodiment of the present application, when the initial generated frame input containing the face area has a built-in multi-layer neural network architecture, high-definition processing is performed: the low-level layer relies on the convolution kernel and reasonable step size to deeply mine the basic features of the image; the middle layer uses deconvolution operations to flexibly adjust parameters according to the image magnification requirements and detail enhancement expectations, gradually restore the image size and carve details; the whole process is assisted by residual connections to fuse the original and processed features in real time to reduce information loss. After multiple layers of processing, high-definition and detailed image generation frames are output, which significantly improves the image quality.

[0069] In one embodiment of the present application, the application of a multi-layer neural network architecture effectively improves the clarity and texture of the facial area, and can present a more realistic and clear picture effect on a high-definition display device, meeting application scenarios with high requirements for video quality.

[0070] Step S150, detect the lip shape alignment deviation of each image generation frame, determine the target reference frame, and adjust the lip shape generation parameters of the video generation model according to the target reference frame to align the reshaped lip shape with the corresponding audio data frame to generate the target video.

[0071] In one embodiment of the present application, the lip alignment deviation of each image generated frame is detected and a target reference frame is determined, including: detecting the lip alignment deviation between each image generated frame and the corresponding audio data frame; if a lip alignment deviation is detected in an image generated frame, determining the target reference frame corresponding to the image generated frame; wherein the target reference frame includes one of a first reference frame, a second reference frame and a third reference frame, the first reference frame is obtained based on an image reference frame with lips in a closed mouth state, the second reference frame is generated based on an audio data frame corresponding to an image generated frame with lip alignment deviation, and the third reference frame is obtained based on the image difference between each image reference frame and the image generated frame with lip alignment deviation.

[0072] In one embodiment of the present application, the lip shape in the reference frame video changes frequently and the image consistency is poor, which makes it difficult to accurately synchronize the generated lip shape image with the audio data frame. Due to the misaligned lip shape and strong sense of incongruity, the authenticity and credibility of the target generated video are greatly reduced, seriously affecting the audio and video interaction effect and information transmission efficiency.

[0073] In one embodiment of the present application, a high-sensitivity monitoring module is deployed during the entire process of generating the target video, integrating cutting-edge image analysis and audio processing technologies to compare the rhythm and tempo information of the current lip shape generated image with the audio data frame in real time, and accurately capture the moment of alignment deviation between the lip shape and the audio data frame, that is, the lip shape alignment deviation.

[0074] In one embodiment of the present application, determining the target reference frame corresponding to the image generation frame includes: if the first reference frame exists in the reference frame video, confirming the first reference frame as the target reference frame; if the first reference frame does not exist in the reference frame video, confirming the second reference frame or the third reference frame as the target reference frame.

[0075] In one embodiment of the present application, confirming the second reference frame or the third reference frame as the target reference frame includes: extracting the spectral feature value corresponding to the audio data frame for voice state confirmation; if the voice state is characterized as silent, generating the second reference frame and determining the second reference frame as the target reference frame; if the voice state is characterized as active, determining the third reference frame as the target reference frame.

[0076] In one embodiment of the present application, the spectral feature value includes audio features extracted from the Mel spectrum, such as Mel spectrum energy.

[0077] In one embodiment of the present application, extracting the spectral feature value corresponding to the audio data frame for voice state confirmation includes: extracting the Mel spectrum of the audio data frame and determining the spectral feature value according to the Mel spectrum; if the spectral feature value is less than the preset threshold, determining that the voice state is characterized as silent; if the spectral feature value is greater than or equal to the preset threshold, determining that the voice state is characterized as active.

[0078] In one embodiment of the present application, determining the spectral feature value according to the Mel spectrum includes: determining the current feature value of the Mel spectrum; performing exponential smoothing on the current feature value, or performing moving average on the initial feature value, to obtain the spectral feature value.

[0079] In one embodiment of the present application, performing exponential smoothing on the initial feature value includes: determining the spectral feature value based on the current feature value, the preset smoothing factor, and the smoothed feature value corresponding to the audio data frame at the previous moment.

[0080] In one embodiment of the present application, performing moving average on the initial feature value: based on the current feature value within a preset time window, selecting the initial feature values corresponding to multiple audio data frames for weighted average to obtain the spectral feature value.

[0081] In one embodiment of the present application, exponential smoothing or moving average can optimize the matching degree between the audio feature and the image generation frame, effectively avoiding lip-sync anomalies caused by audio problems, significantly improving the overall coordination of the audio and video, and creating a more realistic and harmonious audio-visual effect.

[0082] In one embodiment of the present application, the third reference frame is obtained based on the image gap between each image reference frame and the frames generated from the images with lip alignment deviation, including: respectively performing image feature scoring on the frames generated from the images with lip alignment deviation and each image reference frame to obtain the first score of the frame generated from the image and the second scores of each image reference frame, and determining the image reference frame with a small gap from the first score among the second scores as the target reference frame.

[0083] In one embodiment of the present application, the image feature scoring includes measuring the gap through key image features; if there are multiple types of gap measurements and multiple types of key image features, weighted scoring can be performed after assigning weights.

[0084] In one embodiment of the present application, the key image features include color histogram, edge detection, local binary pattern, depth features, etc. The gap measurements include mean square error, structural similarity index, perceptual hashing, feature distance, etc.

[0085] In one embodiment of the present application, closely connecting with the underlying mechanism of Wav2Lip, deeply analyzing the subtle differences in the shape, position, and motion trajectory of the lips in the target reference frame and the frame generated from the image, applying the backpropagation principle of the deep learning model, precisely fine-tuning the lip generation parameters associated within Wav2Lip, reshaping the lips to accurately fit the audio data frame, comprehensively ensuring the natural and smooth lip-sync effect in the target generated video, so as to endow the digital human video with a realistic texture and vivid vitality.

[0086] In one embodiment of the present application, the lip generation parameters include the weight parameters of the loss function, the audio feature vectors corresponding to the audio data frames, the convolutional layer parameters, etc.

[0087] In one embodiment of the present application, by performing lip alignment adjustment through the target reference frame, the problem that it is difficult to fully align the lips in the frame generated from the image with the audio data frame is effectively solved, making the generated lips more accurately and naturally synchronized with the audio, reducing the impact of uncoordinated lips on the video quality, and further improving the realism and credibility of the video.

[0088] In one embodiment of the present application, the present application can also solve the problem that there are lip shape changes when the audio data frame is in a silent state.

[0089] In one embodiment of the present application, after obtaining the target video, it further includes: replacing the original background in the frame generated from the image with a preset background, and / or converting the face area in the frame generated from the image into a cartoon face.

[0090] In one embodiment of the present application, the post-processing link can also be extended: the background removal operation is carried out through the built-in image segmentation algorithm, the background area is intelligently identified according to the complexity of the background and the color distribution characteristics of the frames generated from the images, and the algorithm parameters are fine-tuned as needed to achieve a clean background stripping; then, the custom preset background is loaded as needed, and the fusion parameters are finely adjusted to achieve the ideal background synthesis effect. On the other hand, with the help of the portrait cartoonization model, a neural network structure trained with a large amount of portrait data, the facial contour, color and texture are accurately reshaped, and the real face area is cleverly transformed into a unique style cartoon image, endowing the target video with a special artistic charm and comprehensively expanding the video application scenarios and visual expressiveness.

[0091] In one embodiment of the present application, the present application can address a series of problems in audio-visual synchronization processing in related technologies, and comprehensively reshape the audio-visual generation process from reference frame allocation, image fusion, image quality improvement, and lip synchronization, and can deliver high-quality and super-realistic audio-visual results for multiple fields such as digital human applications and video content creation, strongly promoting the technological advancement and application expansion of the industry.

[0092] In one embodiment of the present application, the present application can be applied to fields such as content creation, live broadcast, intelligent customer service, training and education, etc. Among them, in content creation, when producing video content related to digital humans, such as animated short films, virtual character demonstration videos, etc., high-quality audio-visual effects can be generated, enhancing the attractiveness and professionalism of the works. In live broadcasts, whether it is a virtual anchor live broadcast or a digital human assistant in a live e-commerce scenario, the synchronization and high quality of audio-visual during the live broadcast can be ensured, enhancing the viewing experience of the audience and improving the interactivity and credibility of the live broadcast. In intelligent customer service, when the intelligent customer service interacts with users in the form of a digital human, clear and accurate audio-visual synchronization can enable users to better understand the information conveyed by the customer service, improving the service quality and user satisfaction. In training and education, when using digital human lecturers in enterprise internal training and online education courses, the present application can make the teaching videos more vivid and realistic, helping students better understand and absorb knowledge.

[0093] In one embodiment of the present application, please refer to Figure 2 , Figure 2 shows an implementation schematic diagram of a video generation method according to an embodiment of the present application. As Figure 2As shown in the figure, in step S210, the audio to be converted and the reference frame video are input; in step S220, the audio to be converted is processed: the audio to be converted is segmented into multiple segments to be generated; in step S230, the reference frame video is processed: the frame sequence of the reference frame video is reversed, and based on the alternating reference of the reversed frame video and the reference frame video, the reference source of each segment to be generated is determined and input into the video generation model; in step S240, the video generation model is called; in step S250, the inference lip shape is generated: based on the input reference frame video and the reversed frame video sequence, multiple lip shape generation images corresponding to the audio to be converted are generated; in step S260, it is overlaid on the original video: the lip shape generation images are overlaid on the target image; in step S270, image fusion processing is performed: the overlaid lip shape generation images and the target image are fused through Poisson fusion, Laplacian pyramid fusion or edge feathering fusion to obtain the initial generated frames; in step S280, image high-definition processing is performed: the image quality of the image generated frames is improved through a preset multi-layer neural network; in step S290, the result is generated: the lip shape alignment deviation of each image generated frame is detected, the target reference frame is determined, and the lip shape generation parameters of the video generation model are adjusted according to the target reference frame so that the reshaped lip shapes are aligned with the corresponding audio data frames to generate the target video.

[0094] Please refer to Figure 3 , Figure 3 which shows a block diagram of a video generation device according to an embodiment of the present application. This embodiment does not limit the application environment applicable to this device.

[0095] As Figure 3 shown, a video generation device 300 according to an embodiment of the present application includes: an audio-video acquisition module 301, a lip shape generation module 302, an image fusion module 303, a high-definition processing module 304, and a lip shape alignment module 305.

[0096] Among them, the audio-video acquisition module 301 is used to acquire the reference frame video and the audio to be converted;

[0097] The lip shape generation module 302 is used to dynamically extract multiple image reference frames from the reference frame video and input them into the video generation model to infer and generate multiple lip shape generation images of the audio to be converted;

[0098] The image fusion module 303 is used to fuse each lip shape generation image with the corresponding image reference frame respectively to obtain multiple initial generated frames;

[0099] The high-definition processing module 304 is used to perform high-definition processing on each initial generated frame to obtain multiple image generated frames;

[0100] A lip alignment module 305 is configured to detect the lip alignment deviation of each image generation frame, determine a target reference frame, and adjust the lip generation parameters of the video generation model according to the target reference frame, so as to align the reshaped lips with the corresponding audio data frames and generate a target video.

[0101] It should be noted that the video generation device provided in the above embodiments and the video generation method provided in the above embodiments belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiments, and will not be elaborated herein. In practical applications, the video generation device provided in the above embodiments may, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure into different functional modules to complete all or part of the functions described above. This is not limited herein either.

[0102] An embodiment of the present application further provides an electronic device, including: one or more processors; a storage device for storing one or more programs, which, when executed by the one or more processors, cause the electronic device to implement the video generation methods provided in the above embodiments.

[0103] Please refer to Figure 4 , Figure 4 , which shows a schematic structural diagram of a computer system of an electronic device suitable for implementing the embodiments of the present application. It should be noted that Figure 4 the computer system 400 of the electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0104] As Figure 4 shown, the computer system 400 includes a central processing unit (CPU) 401, which can perform various appropriate actions and processes according to the program stored in a read-only memory (ROM) 402 or the program loaded from a storage section 408 into a random access memory (RAM) 403, such as executing the method in the above embodiments. In the RAM 403, various programs and data required for system operations are also stored. The CPU 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output (I / O) interface 405 is also connected to the bus 404.

[0105] The following components are connected to the I / O interface 405: an input section 406 including a keyboard, a mouse, etc.; an output section 407 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 408 including a hard disk, etc.; and a communication section 409 including a network interface card such as a LAN (Local Area Network) card, a modem, etc. The communication section 409 performs communication processing via a network such as the Internet. A drive 410 is also connected to the I / O interface 405 as required. A removable medium 411 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive 410 as required so that a computer program read therefrom is installed into the storage section 408 as required.

[0106] Specifically, according to an embodiment of the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product including a computer program carried on a computer-readable medium, the computer program including a computer program for performing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 409, and / or installed from the removable medium 411. When the computer program is executed by a central processing unit (CPU) 401, various functions defined in the system of the present application are executed.

[0107] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0108] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0109] The units involved in the embodiments described in this application can be implemented in software or in hardware, and the described units can also be provided in a processor. Among them, the names of these units do not, in some cases, limit the units themselves. Therefore, the technical solution according to the embodiments of this application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a touch terminal, or a network device, etc.) to execute the method according to the embodiments of this application.

[0110] Another aspect of this application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor of a computer, it enables the computer to execute the video generation method provided in each of the above embodiments. The computer-readable storage medium can be included in the electronic device described in the above embodiments, or can exist separately without being assembled into the electronic device.

[0111] In the above embodiments, unless otherwise specified, when using serial numbers such as "first" and "second" to describe a common object, it only represents different instances referring to the same object, rather than indicating that the object to be described must be in a given order, whether in terms of time, space, sorting, or any other way.

[0112] The above embodiments are only used to exemplarily illustrate the principles and effects of this application, rather than to limit this application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of this application. Therefore, all equivalent modifications or changes completed by those with ordinary knowledge in the technical field without departing from the spirit and technical ideas disclosed in this application should still be covered by the claims of this application.

Claims

1. A video generation method, characterized in that: The method comprises: Get the reference frame video and the audio to be converted; Dynamically extract multiple image reference frames from the reference frame video, and input them into a video generation model to infer and generate multiple lip shape generation images of the audio to be converted; fusing each of the lip shape generated images with the corresponding image reference frames to obtain a plurality of initial generated frames; High-definition processing of each of the initial generated frames to obtain a plurality of image generated frames; The lip shape alignment deviation of each image generation frame is detected, a target reference frame is determined, and the lip shape generation parameters of the video generation model are adjusted according to the target reference frame so that the reshaped lip shape is aligned with the corresponding audio data frame to generate the target video.

2. The video generation method according to claim 1, characterized in that: Dynamically extracting multiple image reference frames in the reference frame video, including: Reversing a frame sequence of a plurality of image reference frames in the reference frame video to obtain a reversed frame video; Segmenting the audio to be converted to obtain multiple segments to be generated; According to the alternating reference between the inverted frame video and the reference frame video, a reference source of each of the to-be-generated segments is determined to be input into the video generation model.

3. The video generation method according to claim 2, characterized in that: Determining a reference source of each of the to-be-generated segments according to the alternating references of the inverted frame video and the reference frame video includes: Using the inverted frame video as a reference source for odd-numbered segments to be generated, and using the reference frame video as a reference source for even-numbered segments to be generated; or, The reference frame video is used as a reference source for odd-numbered segments to be generated, and the inverted frame video is used as a reference source for even-numbered segments to be generated.

4. The video generation method according to claim 1, characterized in that: Detecting the lip alignment deviation of each of the image generation frames and determining the target reference frame includes: Detecting the lip alignment deviation between each of the image generation frames and the corresponding audio data frame; If a lip alignment deviation is detected in an image generation frame, determining a target reference frame corresponding to the image generation frame; Among them, the target reference frame includes one of a first reference frame, a second reference frame and a third reference frame, the first reference frame is obtained based on an image reference frame with lips in a closed mouth state, the second reference frame is generated based on an audio data frame corresponding to an image generation frame with lip alignment deviation, and the third reference frame is obtained based on the image difference between each of the image reference frames and the image generation frame with lip alignment deviation.

5. The video generation method according to claim 4, characterized in that: Determining a target reference frame corresponding to the image generation frame includes: If the first reference frame exists in the reference frame video, confirming the first reference frame as the target reference frame; If the first reference frame does not exist in the reference frame video, the second reference frame or the third reference frame is confirmed as the target reference frame.

6. The video generation method according to claim 5, characterized in that: Confirming the second reference frame or the third reference frame as the target reference frame includes: Extracting the frequency spectrum feature value corresponding to the audio data frame to confirm the voice status; If the speech state is characterized as silence, generating the second reference frame, and determining the second reference frame as the target reference frame; If the speech state is characterized as active, the third reference frame is determined as the target reference frame.

7. The video generation method according to any one of claims 1 to 6, characterized in that: The lip shape generated image is fused with the corresponding image reference frame to obtain an initial generated frame, including: Create a first mask of the same size as the lip shape generation image; Draw a rectangular area in the center area of ​​the first mask; Performing transition processing on the rectangular area to obtain a target mask; The lip shape generated image and the corresponding image reference frame are fused channel by channel based on the normalized weight of the target mask to obtain an initial generated frame.

8. A video generating device, characterized in that: The device comprises: An audio and video acquisition module, used to acquire reference frame video and audio to be converted; A lip shape generation module, used for dynamically extracting a plurality of image reference frames from the reference frame video, and inputting the plurality of image reference frames into a video generation model, and inferring and generating a plurality of lip shape generation images of the audio to be converted; An image fusion module, used for fusing each of the lip shape generation images with the corresponding image reference frame to obtain a plurality of initial generation frames; A high-definition processing module, used for processing each of the initial generated frames in high definition to obtain a plurality of image generation frames; The lip alignment module is used to detect the lip alignment deviation of each image generation frame, determine the target reference frame, and adjust the lip generation parameters of the video generation model according to the target reference frame so that the reshaped lip shape is aligned with the corresponding audio data frame to generate the target video.

9. An electronic device, characterized in that: The electronic device comprises: one or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enables the electronic device to implement the video generation method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the video generation method according to any one of claims 1 to 7.