Image generation methods, apparatus, devices and storage media
By performing facial mesh parameter conversion and mouth shape smoothing processing on voice and video signals, the problem of lost mouth shape expression in existing technologies has been solved, generating mouth shape driven images that combine pronunciation accuracy and emotional expressiveness, thereby enhancing the realism and interactive experience of virtual digital humans.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GUANGZHOU HUYA INFORMATION TECH CO LTD
- Filing Date
- 2026-03-12
- Publication Date
- 2026-07-31
AI Technical Summary
Existing voice-driven lip-syncing technology fails to effectively preserve the micro-expressions of the mouth in the original video, resulting in stiff and lifeless expressions in the generated digital human, lacking emotional expressiveness.
By acquiring frame data from audio and video signals, facial mesh parameters are converted to generate third-party facial mesh data. Mouth shape smoothing and conditional rendering are then performed. Finally, the rendered image of the mouth area is fused with the video frame to ensure accurate pronunciation and natural facial expressions.
It effectively preserves the natural facial expression details of the original video, generating mouth-driven images that combine accurate pronunciation with rich emotional expression, thus enhancing the realism and interactive experience of the virtual digital human.
Smart Images

Figure CN122492902A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of virtual digital human technology, and in particular to an image generation method, apparatus, device, and storage medium. Background Technology
[0002] With the rapid development of digital human technology, virtual reality, and online content creation, voice-driven lip-syncing animation technology has become a key element in achieving lifelike human-computer interaction and efficient video production. This technology aims to automatically generate precisely matched mouth movements based on input audio signals, thereby giving digital humans or virtual avatars the ability to speak vividly.
[0003] Most current voice-driven lip-sync solutions employ a frame-by-frame lip-sync inference and overlay processing method. Specifically, the model regenerates a completely new lip region image for each frame of the video based on the current audio frame, and then pastes it onto the image. Figure 1 This method directly overlays the mouth area onto the original image. Essentially, this process replaces rather than merges the original video's mouth, thus discarding the subtle, expressive features inherent in the character's mouth that carry rich emotional information. This results in a stiff, lifeless digital human with expressions lacking emotional depth and expressiveness. Viewers can clearly perceive an unnatural mechanical feel, severely impacting the immersive experience and the credibility of the content. Summary of the Invention
[0004] The main objective of this invention is to provide an image generation method, apparatus, device, and storage medium, aiming to solve the technical problem that existing voice-driven lip-syncing technology fails to effectively preserve the micro-expressions of the mouth in the original video, resulting in stiff and lifeless digital human expressions that lack emotional expressiveness.
[0005] The first aspect of the present invention provides an image generation method, the image generation method comprising: Acquire audio frames from the audio signal and video frames from the original video signal; The voice frame is converted to face grid parameters to obtain first face grid data, and the video frame is converted to face grid parameters to obtain second face grid data. If the voice frame is located in a preset transition window, then a third face grid data is generated based on the first face grid data and the second face grid data; Using the third face mesh data as a constraint, conditional rendering is performed on the preset mouth shape driving reference map to obtain a rendered image of the mouth region; The rendered image of the mouth area and the video frame are fused together to obtain the mouth shape driven image of the current frame.
[0006] Optionally, in a first implementation of the first aspect of the present invention, the image generation further includes... Speech activity detection is performed on the speech signal, and based on the speech activity detection results, silent segments and speech segments in the speech signal are marked; At the boundary between each silent segment and each speech segment of the speech signal, a transition window of a preset number of speech frames is set. The transition window includes speech frames of some silent segments and speech frames of some speech segments.
[0007] Optionally, in a second implementation of the first aspect of the present invention, generating third face mesh data based on the first face mesh data and the second face mesh data includes: If the voice frame is located in the transition window, then a mouth smoothing transition is performed based on the first face grid data and the second face grid data to obtain the third face grid data.
[0008] Optionally, in a third implementation of the first aspect of the present invention, the step of performing mouth smoothing processing based on the first face mesh data and the second face mesh data to obtain the third face mesh data includes: If the transition window corresponds to the transition from a speech segment to a silence segment, then obtain the first weight of the first face grid data and the second weight of the second face grid data; Based on the first weight and the second weight, a first weighting strategy is adopted to perform a weighted average processing on the first face grid data and the second face grid data to obtain the third face grid data; The first weighting strategy includes: referring to the processing order of each speech frame in the speech signal, the weight of the first face grid data corresponding to each speech frame decreases frame by frame; referring to the processing order of each video frame in the original video signal, the weight of the second face grid data corresponding to each video frame increases frame by frame.
[0009] Optionally, in a fourth implementation of the first aspect of the present invention, the step of performing mouth smoothing transition processing based on the first face mesh data and the second face mesh data to obtain the third face mesh data further includes: If the transition window corresponds to the transition from a silent segment to a speech segment, then the third weight of the first face grid data and the fourth weight of the second face grid data are obtained; Based on the third weight and the fourth weight, a second weighting strategy is adopted to perform a weighted average processing on the first face grid data and the third face grid data to obtain the third face grid data. The second weighting strategy includes: referring to the processing order of each speech frame in the speech signal, the weight of the first face grid data corresponding to each speech frame increases frame by frame; referring to the processing order of each video frame in the original video signal, the weight of the second face grid data corresponding to each video frame decreases frame by frame.
[0010] Optionally, in a fifth implementation of the first aspect of the present invention, after fusing the rendered image of the mouth region and the video frame to obtain the mouth shape driving image of the current frame, the method further includes: If the transition window corresponds to the transition from a speech segment to a silence segment, then the fifth weight of the video frame and the sixth weight of the lip-sync image are obtained. Based on the fifth weight and the sixth weight, the third weighting strategy is used to perform weighted averaging on the video frame and the lip-shape driving image to obtain the optimized lip-shape driving image of the current frame. The third weighting strategy includes: referring to the processing order of each video frame in the original video signal, the weight of each video frame increases frame by frame, and the weight of each lip-sync image decreases frame by frame.
[0011] Optionally, in a sixth implementation of the first aspect of the present invention, after fusing the rendered image of the mouth region and the video frame to obtain the mouth shape driving image of the current frame, the method further includes: If the transition window corresponds to the transition from a silent segment to a speech segment, then the seventh weight of the video frame and the eighth weight of the lip-sync image are obtained. Based on the seventh weight and the eighth weight, the fourth weighting strategy is used to perform weighted averaging on the video frame and the mouth shape driving image to obtain the optimized mouth shape driving image of the current frame. The fourth weighting strategy includes: referring to the processing order of each video frame in the original video signal, the weight of each video frame decreases frame by frame, and the weight of each lip-sync image increases frame by frame.
[0012] A second aspect of the present invention also provides an image generating apparatus, the image generating apparatus comprising: The acquisition module is used to acquire the audio frames of the audio signal and the video frames of the original video signal. The conversion module is used to convert the face grid parameters of the audio frame to obtain the first face grid data, and to convert the face grid parameters of the video frame to obtain the second face grid data. The lip-shape optimization module is used to generate third face grid data based on the first face grid data and the second face grid data if the voice frame is located in a preset transition window; The rendering module is used to perform conditional rendering on a preset mouth shape driving reference map using the third face mesh data as a constraint, so as to obtain a rendered image of the mouth region. The fusion module is used to fuse the rendered image of the mouth region and the video frame to obtain the mouth shape driven image of the current frame.
[0013] Optionally, in a first implementation of the second aspect of the present invention, the image generation apparatus further includes: The speech detection module is used to perform speech activity detection on the speech signal and mark the silent segments and speech segments in the speech signal according to the speech activity detection results; at the boundary between each silent segment and each speech segment of the speech signal, a transition window of a preset number of speech frames is set, and the transition window includes speech frames of some silent segments and speech frames of some speech segments.
[0014] Optionally, in a second implementation of the second aspect of the present invention, the mouth shape optimization module is specifically used for: If the voice frame is located in the transition window, then a mouth smoothing transition is performed based on the first face grid data and the second face grid data to obtain the third face grid data.
[0015] Optionally, in a third implementation of the second aspect of the present invention, the mouth shape optimization module is specifically used for: If the voice frame is located in the transition window, and the transition window corresponds to the transition from a voice segment to a silence segment, then the first weight of the first face grid data and the second weight of the second face grid data are obtained. Based on the first weight and the second weight, a first weighting strategy is adopted to perform a weighted average processing on the first face grid data and the second face grid data to obtain the third face grid data; The first weighting strategy includes: referring to the processing order of each speech frame in the speech signal, the weight of the first face grid data corresponding to each speech frame decreases frame by frame; referring to the processing order of each video frame in the original video signal, the weight of the second face grid data corresponding to each video frame increases frame by frame.
[0016] Optionally, in a fourth implementation of the second aspect of the present invention, the mouth shape optimization module is further used for: If the voice frame is located in the transition window, and the transition window corresponds to the transition from a silent segment to a voice segment, then the third weight of the first face grid data and the fourth weight of the second face grid data are obtained. Based on the third weight and the fourth weight, a second weighting strategy is adopted to perform a weighted average processing on the first face grid data and the third face grid data to obtain the third face grid data. The second weighting strategy includes: referring to the processing order of each speech frame in the speech signal, the weight of the first face grid data corresponding to each speech frame increases frame by frame; referring to the processing order of each video frame in the original video signal, the weight of the second face grid data corresponding to each video frame decreases frame by frame.
[0017] Optionally, in a fifth implementation of the second aspect of the present invention, the image generation apparatus further includes: The color optimization module is used to obtain the fifth weight of the video frame and the sixth weight of the lip-sync image if the voice frame is located in the transition window and the transition window corresponds to the transition from the voice segment to the silence segment; based on the fifth weight and the sixth weight, a third weighting strategy is used to perform weighted average processing on the video frame and the lip-sync image to obtain the optimized lip-sync image of the current frame. The third weighting strategy includes: referring to the processing order of each video frame in the original video signal, the weight of each video frame increases frame by frame, and the weight of each lip-sync image decreases frame by frame.
[0018] Optionally, in a sixth implementation of the second aspect of the present invention, the color optimization module is further configured to: If the speech frame is located in the transition window, and the transition window corresponds to the transition from a silent segment to a speech segment, then the seventh weight of the video frame and the eighth weight of the lip-sync image are obtained. Based on the seventh weight and the eighth weight, the fourth weighting strategy is used to perform weighted averaging on the video frame and the mouth shape driving image to obtain the optimized mouth shape driving image of the current frame. The fourth weighting strategy includes: referring to the processing order of each video frame in the original video signal, the weight of each video frame decreases frame by frame, and the weight of each lip-sync image increases frame by frame.
[0019] A third aspect of the present invention provides a computer device, comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor invokes the instructions in the memory to cause the computer device to perform the image generation method described above.
[0020] A fourth aspect of the present invention provides a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the image generation method described above.
[0021] The image generation method provided in this invention acquires both the audio frame of the audio signal and the video frame of the original video signal. Then, it performs face mesh parameter conversion on both the audio frame and the video frame. For audio frames located within a preset transition window, this invention does not directly use the face mesh data converted from the audio frame to generate a mouth-shaped image. Instead, it first uses the face mesh data converted from the original video frame to optimize the mouth region of the face mesh data converted from the audio frame. Then, using the optimized face mesh data as a constraint, it performs conditional rendering on a preset mouth-shaped reference image to obtain the mouth region rendering image corresponding to the current image frame to be generated. Finally, it fuses the mouth region rendering image with the original video frame to obtain the final mouth-shaped image. This invention optimizes the face mesh parameters of the audio and video frames to find the optimal balance between pronunciation accuracy and the original facial expression state, ensuring that the optimized face mesh data can match the speech while retaining key micro-expression features such as upturned corners of the mouth and tight lips, thus fundamentally solving the root cause problem of lost facial expression in mouth-shaped images. The embodiments of the present invention resolve the technical contradiction between audio-visual synchronization and facial expression preservation, effectively preserving the natural facial expression details of the original video, and ultimately outputting a mouth-driven image that combines accurate pronunciation and rich emotional expression, significantly enhancing the realistic sensory and interactive experience of the virtual digital human. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of one embodiment of the image generation method in this invention; Figure 2 This is a schematic diagram of an embodiment of the transition window in the audio signal according to the present invention; Figure 3 This is a schematic diagram of one embodiment of the image generation device in this invention; Figure 4 This is a schematic diagram of one embodiment of the computer device in this invention. Detailed Implementation
[0023] The terms "first," "second," "third," "fourth," etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" or "having" and any variations thereof are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0024] For ease of understanding, the specific process of the embodiments of the present invention is described below. Please refer to [link / reference]. Figure 1 One embodiment of the image generation method in this invention includes: 101. Acquire the audio frames of the audio signal and the video frames of the original video signal; In this embodiment, in order to output a mouth-shaped driven image that has both accurate pronunciation and rich emotional expression, the mouth-shaped driven image is not generated by the speech signal alone, but by the speech signal and the original video signal corresponding to the speech signal at the same time.
[0025] In this embodiment, one audio frame corresponds to one lip-sync image frame. Therefore, the audio signal needs to be segmented into frames to obtain audio frames. For example, the audio signal to be processed can be segmented into frames of a fixed duration (e.g., 20-40 milliseconds) to obtain the corresponding audio frame sequence. When reading the original video signal, it is also necessary to extract the image frame by frame according to the timestamp or frame number and synchronize it with the audio frames to ensure that the time of each video frame is consistent with the corresponding audio frame. For example, a buffer or timestamp alignment strategy can be set to ensure strict synchronization of audio and video and prevent misalignment of lip movements and sound.
[0026] 102. Perform face grid parameter conversion on the audio frame to obtain first face grid data, and perform face grid parameter conversion on the video frame to obtain second face grid data; In this embodiment, to address the issue of lost mouth expression information in the original video, face mesh parameter conversion is performed on both the audio and video frames simultaneously for smooth mouth transition processing. The parameter conversion model is a pre-trained model, with different models used for audio and video frames. Specifically, the speech features of the input audio frame or the image features of the video frame are mapped to high-precision face mesh data, thereby achieving the conversion of speech signals or video signals into facial parameters.
[0027] Parametric transformation models typically employ deep learning models to map speech features or image features from video frames to face mesh data. Training data can be supervised using a large amount of synchronized audio and video data to enable the model to learn the mapping relationship between speech or image features and mouth movements. The output of the parametric transformation model is a set of face mesh vertex coordinates or parameters corresponding to the input speech or video frame, used to describe the geometric shape of the mouth in the current speech or video frame (e.g., open mouth, closed mouth, lip stretching, etc.), thus obtaining the corresponding face mesh data.
[0028] 103. If the voice frame is located in a preset transition window, then a third face grid data is generated based on the first face grid data and the second face grid data; In this embodiment, to address the issue of lost mouth expression information in the original video, a mouth shape smoothing transition process is introduced. This process smooths the transition of the face mesh data generated based on the speech signal, preventing abrupt changes in mouth shape and thus preserving the mouth expression features in the original video. It should be noted that this embodiment does not perform mouth shape smoothing transition processing on all face mesh data corresponding to all speech frames. Therefore, it is necessary to pre-determine whether the currently processed speech frame meets the pre-set mouth shape optimization conditions. If the conditions are met, mouth shape optimization processing is performed; otherwise, it is not.
[0029] In an optional embodiment, it is preferable to perform lip-shape smoothing processing on the face grid data corresponding to the audio frame located between the speech segment and the silence segment in the speech signal. That is, if the currently processed speech frame is located between the speech segment and the silence segment, it is determined that the lip-shape optimization condition is met.
[0030] Before performing lip-sync smoothing, it is necessary to perform speech activity detection and transition window settings on the speech signal, as follows: (1) Perform speech activity detection on the speech signal, and mark the silent segments and speech segments in the speech signal according to the speech activity detection results; (2) At the junction of each silent segment and each speech segment of the speech signal, a transition window of a preset number of speech frames is set. The transition window includes speech frames of some silent segments and speech frames of some speech segments.
[0031] This optional embodiment is specifically used to segment an input continuous audio stream into segments of two states: speech and silence, and to mark them with timestamps or frame indices. Specifically, it analyzes the short-term characteristics of the audio signal to determine whether each audio segment contains human voice. If it contains human voice, the audio segment is marked as a speech segment; otherwise, it is marked as a silence segment.
[0032] Simultaneously, at each boundary between "mute" and "speak" or "speak" and "mute," a buffer containing both states is created for subsequent smooth lip-sync processing. Considering that switching lip-sync directly at the boundary of speech activity (e.g., instantly switching from a raw facial expression to audio-driven lip-sync) would create a visual jarring effect, this optional embodiment introduces a transition window to define a gradient region within which smooth lip-sync processing is performed.
[0033] like Figure 2 As shown, a transition window is provided at the boundary between two segments, whether transitioning from a silence segment to a speech segment or vice versa. Assuming the transition window has a length of 10 frames, these 10 frames contain not only some speech frames from the silence segment but also some speech frames from the speech segment. The proportion of silence segments to speech segments within the transition window is unlimited. For example, the 10 frames of the transition window may contain 6 speech frames from the silence segment and 4 speech frames from the speech segment.
[0034] In this embodiment, the transition window is set as the lip-shape optimization condition. That is, when executing step 103, it is necessary to first determine whether the currently processed voice frame is located in the transition window. If the voice frame is located in the transition window, it is determined that the voice frame meets the lip-shape optimization condition, and it is necessary to further perform lip-shape smooth transition processing on the face grid data corresponding to the currently processed voice frame. Otherwise, it is determined that the currently processed voice frame does not meet the lip-shape optimization condition, that is, it is not necessary to perform lip-shape smooth transition processing on the face grid data corresponding to the currently processed voice frame.
[0035] The smooth transition processing in this embodiment can employ a temporal smoothing algorithm to interpolate or filter the grid data of consecutive frames, eliminating jitter and abrupt changes. By combining the facial key points of the original video frames, the facial grid data corresponding to the audio frames is fine-tuned to ensure that the mouth shape is consistent with the original facial posture and expression. Through mouth shape smooth transition processing, a set of smooth, continuous facial grid data can be obtained, which reflects both the speech-driven mouth shape changes and maintains visual consistency with the original video.
[0036] The preferred mouth shape smooth transition process in this embodiment specifically includes the following two stages: Phase 1: Extracting the second face mesh data from the original video frames In this stage, each video frame of the original video signal needs to be fed into a pre-trained face detection model for face detection to obtain the second face mesh data. Then, based on the second face mesh data corresponding to the original video frame and the first face mesh data corresponding to the audio frame, lip shape smoothing processing is performed.
[0037] Face mesh data consists of hundreds of points that conform to the contours of a human face. It defines the shape and position of eyebrows, eyes, nose, lips, and cheeks. The original mouth shape can be determined using face mesh data. Specifically, this stage involves using a face detection model to detect faces in video frames to extract the original video's face mesh data.
[0038] Phase Two: Smooth Transition of Lip Shape This stage performs lip-sync smoothing based on the second face mesh data extracted from the original video in the previous stage and the first face mesh data corresponding to the audio frames. Specifically, it weights and fuses the lip-sync generated based on the speech signal (first face mesh data) with the natural facial expressions of the original video (second face mesh data). Lip-sync generated based on speech frames (first face mesh data) is used in speech segments, while lip-sync generated based on video frames (second face mesh data) is used in silent segments. In the transition window, the first and second face mesh data are smoothly blended according to their weights, ultimately generating a natural face mesh (third face mesh data) that retains the details of the mouth in the original video.
[0039] The smooth transition of the mouth shape in this stage is specifically divided into two situations: Scenario 1: The transition window corresponds to the transition from a speech segment to a silence segment. In this scenario, if the speech signal transitions from a speech segment to a silent segment during the transition window, the first weight of the first face grid data and the second weight of the second face grid data are obtained. Then, based on the first and second weights, a first weighting strategy is adopted to perform a weighted average processing on the first face grid data corresponding to the current speech frame and the second face grid data corresponding to the current video frame to obtain the third face grid data.
[0040] The first weighting strategy used in this scenario includes: referring to the processing order of each speech frame in the speech signal, the weight of the first face grid data corresponding to each speech frame decreases frame by frame; referring to the processing order of each video frame in the original video signal, the weight of the second face grid data corresponding to each video frame increases frame by frame.
[0041] For example, the speech signal transitions from a speech segment to a silence segment. The transition window contains 10 speech frames. The values of the first weight and the second weight range from 0.1 to 1. If the current speech frame is the first frame processed within the transition window, the weight of the first face grid data corresponding to that speech frame is the maximum value of 1, while the weight of the second face grid data corresponding to the video frame is the minimum value of 0.1. If the current speech frame is the second frame processed within the transition window, the weight of the first face grid data corresponding to that speech frame is 0.9, while the weight of the second face grid data corresponding to the video frame is 0.2. The weights corresponding to other speech frames and video frames are calculated in the same way.
[0042] In this scenario, the third face grid data is obtained by weighted averaging the first face grid data corresponding to the current audio frame and the second face grid data corresponding to the current video frame. This can be represented by the following formula: T = (M*k1 + N*k2) / (k1 + k2), where T represents the third face grid data, M represents the first face grid data, N represents the second face grid data, k1 represents the first weight of the first face grid data, and k2 represents the second weight of the second face grid data.
[0043] Scenario 2: The transition window corresponds to the transition from a silent segment to a speech segment. In this scenario, if the speech signal transitions from a silent segment to a speech segment within the transition window, the third weight of the first face grid data and the fourth weight of the second face grid data are obtained. Then, based on the third and fourth weights, a second weighting strategy is adopted to perform a weighted average of the first face grid data corresponding to the current speech frame and the second face grid data corresponding to the current video frame to obtain the third face grid data.
[0044] The second weighting strategy used in this scenario includes: referencing the processing order of each speech frame in the speech signal, the weight of the first face grid data corresponding to each speech frame increases frame by frame; and referring to the processing order of each video frame in the original video signal, the weight of the second face grid data corresponding to each video frame decreases frame by frame.
[0045] For example, the speech signal transitions from a silence segment to a speech segment. The transition window contains 10 speech frames. The values of the third and fourth weights range from 0.1 to 1. If the current speech frame is the first frame processed within the transition window, the weight of the first face grid data corresponding to that speech frame is the minimum value of 0.1, while the weight of the second face grid data corresponding to the video frame is the maximum value of 1. If the current speech frame is the second frame processed within the transition window, the weight of the first face grid data corresponding to that speech frame is 0.2, while the weight of the second face grid data corresponding to the video frame is 0.9. The weights corresponding to other speech frames and video frames follow the same pattern.
[0046] In this embodiment, if the current speech frame is not located within the preset transition window, there is no need to perform mouth shape smoothing transition processing. The corresponding mouth shape driving image can be directly generated using the face grid data corresponding to the current speech frame.
[0047] 105. Using the third face mesh data as a constraint, perform conditional rendering on the preset mouth shape driving reference map to obtain a rendered image of the mouth region. 106. Fuse the rendered image of the mouth area with the video frame to obtain the mouth shape driven image of the current frame.
[0048] In this embodiment, within the transition window, the third-person face mesh data after mouth shape smoothing is used as a constraint to conditionally render a preset mouth shape driving reference image, thereby obtaining the rendered image of the mouth region corresponding to the current image frame to be generated. The mouth shape driving reference image can be the previous frame's mouth shape driving image generated based on the speech signal, or a fixed frame's mouth shape driving image.
[0049] Conditional rendering refers to using a graphics rendering engine to render mesh data into a 2D / 3D image. Conditions include lighting, materials, and textures (such as skin and lips). By using mesh masks or UV coordinate clipping, only the mouth area is rendered, thus generating a high-precision mouth image.
[0050] In this embodiment, using the smoothed third-person face mesh data as a constraint, after obtaining the mouth region image corresponding to the current image frame to be generated through conditional rendering, the mouth region image and the current video frame are input into a neural rendering network for image fusion to obtain the final mouth-driven image. The neural rendering network can be a deep learning model used to fuse the mouth region image with the original video frame. The final mouth-driven image is generated through image fusion to ensure that the mouth is synchronized with the speech and blends naturally with the original video frame.
[0051] In addition, to further address the issue of ineffective processing of the mouth occlusion area in the original video, this embodiment introduces color smoothing processing to smoothly transition the color of the mouth area between speech segments and silent segments, preventing color jumps and improving the consistency of the rendering result with the mouth color in the original video, thereby enabling the direct and effective application of occluded video material in silent segments.
[0052] In an optional embodiment, to address the issue of ineffective processing of the mouth occlusion area in the original video, after image fusion and obtaining the mouth-driven image, it is further determined whether color optimization should be performed, specifically including: (1) Determine whether the speech frame meets the color optimization conditions; (2) If the speech frame satisfies the color optimization condition, then based on the video frame, the mouth shape driving image is subjected to color smoothing transition processing to obtain the optimized mouth shape driving image.
[0053] In this optional embodiment, color optimization and mouth optimization preferably use the same conditions, that is, it is determined whether the currently processed speech frame is located in the transition window; if the currently processed speech frame is located in the transition window, it is determined that the currently processed speech frame meets the color optimization conditions, that is, the mouth shape driven image needs to be color smooth transition processing; otherwise, it is determined that the currently processed speech frame does not meet the color optimization conditions, that is, the mouth shape driven image does not need to be color smooth transition processing.
[0054] The color smoothing transition processing in this optional embodiment is specifically divided into two cases: Scenario 1: The transition window corresponds to the transition from a speech segment to a silence segment. In this scenario, if the speech signal transitions from a speech segment to a silence segment within the transition window, the fifth weight of the acquired video frame and the sixth weight of the lip-sync image are obtained. Then, based on the fifth and sixth weights, a third weighting strategy is used to perform a weighted average on the current video frame and the lip-sync image to obtain the optimized lip-sync image.
[0055] The third weighting strategy used in this case includes: referring to the processing order of each video frame in the original video signal, the weight of each video frame increases frame by frame, and the weight of each lip-sync image decreases frame by frame.
[0056] For example, the speech signal transitions from a speech segment to a silence segment. The transition window contains 10 speech frames. The values of the fifth and sixth weights range from 0.1 to 1. If the current speech frame is the first frame processed within the transition window, the weight of the corresponding video frame is the minimum value of 0.1, while the weight of the corresponding lip-sync image is the maximum value of 1. If the current speech frame is the second frame processed within the transition window, the weight of the corresponding video frame is 0.2, while the weight of the corresponding lip-sync image is 0.9. The weights of other video frames and lip-sync images follow the same pattern.
[0057] In this case, the optimized lip-sync image is obtained by weighted averaging the video frame corresponding to the current speech frame and the lip-sync image, which can be expressed by the following formula: S = (P*h1 + Q*h2) / (h1 + h2), where S represents the optimized mouth shape driven image, P represents the video frame, Q represents the mouth shape driven image, h1 represents the fifth weight of the video frame, and h2 represents the sixth weight of the mouth shape driven image.
[0058] Scenario 2: The transition window corresponds to the transition from a silent segment to a speech segment. In this scenario, if the speech signal transitions from a silent segment to a speech segment within the transition window, the seventh weight of the acquired video frame and the eighth weight of the lip-sync image are obtained. Then, based on the seventh and eighth weights, a fourth weighting strategy is used to perform a weighted average on the current video frame and the lip-sync image to obtain the optimized lip-sync image.
[0059] The fourth weighting strategy used in this case includes: referring to the processing order of each video frame in the original video signal, the weight of each video frame decreases frame by frame, and the weight of each lip-sync image increases frame by frame.
[0060] For example, the speech signal transitions from a silence segment to a speech segment. The transition window contains 10 speech frames. The values of the seventh and eighth weights range from 0.1 to 1. If the current speech frame is the first frame processed within the transition window, the weight of the corresponding video frame is the maximum value of 1, while the weight of the corresponding lip-sync image is the minimum value of 0.1. If the current speech frame is the second frame processed within the transition window, the weight of the corresponding video frame is 0.9, while the weight of the corresponding lip-sync image is 0.2. The weights of other video frames and lip-sync images follow the same pattern.
[0061] In this embodiment, if the current audio frame is not within the preset transition window, color smoothing transition processing is unnecessary. In this optional embodiment, by performing color smoothing transition processing on the lip-sync image, it ensures that even when the user is holding an object or their face is partially obscured during live streaming, a visually coordinated and seamless composite effect can still be generated. This overcomes the limitations of lip-sync driving on user expressions and movements, allowing live stream creators to use more expressive original materials. Users can naturally use gestures, props, and other elements in live streams without worrying about technical limitations causing continuity errors, thus providing greater freedom for content creation and helping to create a more personalized and recognizable live stream style, enriching the content formats of AI live streaming.
[0062] The above describes the mouth shape-driven image generation method in the embodiments of the present invention. The following describes the image generation apparatus in the embodiments of the present invention. Please refer to [link / reference]. Figure 3 One embodiment of the image generation apparatus in this invention includes: The acquisition module 301 is used to acquire the audio frames of the audio signal and the video frames of the original video signal; The conversion module 302 is used to convert the face grid parameters of the audio frame to obtain first face grid data, and to convert the face grid parameters of the video frame to obtain second face grid data. The lip-shape optimization module 303 is used to generate third face grid data based on the first face grid data and the second face grid data if the voice frame is located in a preset transition window; The rendering module 304 is used to perform conditional rendering on a preset mouth shape driving reference map with the third face mesh data as a constraint to obtain a mouth region rendering image. The fusion module 305 is used to fuse the rendered image of the mouth region and the video frame to obtain the mouth shape driven image of the current frame.
[0063] In an optional embodiment, the mouth shape optimization module 303 is specifically used for: If the voice frame is located within the transition window, then a mouth smoothing transition is performed based on the first face grid data and the second face grid data to obtain the third face grid data.
[0064] In an optional embodiment, the mouth shape-driven image generation device further includes: The speech detection module is used to perform speech activity detection on the speech signal and mark the silent segments and speech segments in the speech signal according to the speech activity detection results; at the boundary between each silent segment and each speech segment of the speech signal, a transition window of a preset number of speech frames is set, and the transition window includes speech frames of some silent segments and speech frames of some speech segments.
[0065] In an optional embodiment, the mouth shape optimization module 303 is specifically used for: If the voice frame is located in a preset transition window, and the transition window corresponds to the transition from a voice segment to a silence segment, then the first weight of the first face grid data and the second weight of the second face grid data are obtained. Based on the first weight and the second weight, a first weighting strategy is adopted to perform a weighted average processing on the first face grid data and the second face grid data to obtain the third face grid data; The first weighting strategy includes: referring to the processing order of each speech frame in the speech signal, the weight of the first face grid data corresponding to each speech frame decreases frame by frame; referring to the processing order of each video frame in the original video signal, the weight of the second face grid data corresponding to each video frame increases frame by frame.
[0066] In an optional embodiment, the mouth shape optimization module 303 is further configured to: If the voice frame is located in a preset transition window, and the transition window corresponds to the transition from a silent segment to a voice segment, then the third weight of the first face grid data and the fourth weight of the second face grid data are obtained. Based on the third weight and the fourth weight, a second weighting strategy is adopted to perform a weighted average processing on the first face grid data and the third face grid data to obtain the third face grid data. The second weighting strategy includes: referring to the processing order of each speech frame in the speech signal, the weight of the first face grid data corresponding to each speech frame increases frame by frame; referring to the processing order of each video frame in the original video signal, the weight of the second face grid data corresponding to each video frame decreases frame by frame.
[0067] In an optional embodiment, the mouth shape-driven image generation device further includes: The color optimization module is used to obtain the fifth weight of the video frame and the sixth weight of the lip-sync image if the voice frame is located in the transition window and the transition window corresponds to the transition from the voice segment to the silence segment; based on the fifth weight and the sixth weight, a third weighting strategy is used to perform weighted average processing on the video frame and the lip-sync image to obtain the optimized lip-sync image of the current frame. The third weighting strategy includes: referring to the processing order of each video frame in the original video signal, the weight of each video frame increases frame by frame, and the weight of each lip-sync image decreases frame by frame.
[0068] In an optional embodiment, the color optimization module is further configured to: If the speech frame is located in the transition window, and the transition window corresponds to the transition from a silent segment to a speech segment, then the seventh weight of the video frame and the eighth weight of the lip-sync image are obtained. Based on the seventh weight and the eighth weight, the fourth weighting strategy is used to perform weighted averaging on the video frame and the mouth shape driving image to obtain the optimized mouth shape driving image of the current frame. The fourth weighting strategy includes: referring to the processing order of each video frame in the original video signal, the weight of each video frame decreases frame by frame, and the weight of each lip-sync image increases frame by frame.
[0069] Since the embodiments of the device part correspond to the embodiments of the above method, the description of the image generation device provided by the present invention should refer to the above method embodiments. The present invention will not be described again here, but it has the same beneficial effects as the above image generation method.
[0070] above Figure 3 The image generation device in the embodiments of the present invention will be described in detail from the perspective of modular functional entities. The computer device in the embodiments of the present invention will be described in detail from the perspective of hardware processing.
[0071] Figure 4 This is a schematic diagram of the structure of a computer device 500 provided in an embodiment of the present invention. The computer device 500 can vary significantly due to different configurations or performance characteristics. It may include one or more central processing units (CPUs) 510 (e.g., one or more processors) and a memory 520, and one or more storage media 530 (e.g., one or more mass storage devices) for storing application programs 533 or data 532. The memory 520 and storage media 530 can be temporary or persistent storage. The program stored in the storage media 530 may include one or more modules (not shown in the diagram), each module including a series of instruction operations on the computer device 500. Furthermore, the processor 510 may be configured to communicate with the storage media 530 and execute the series of instruction operations in the storage media 530 on the computer device 500.
[0072] Computer device 500 may also include one or more power supplies 540, one or more wired or wireless network interfaces 550, one or more input / output interfaces 560, and / or one or more operating systems 531, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that... Figure 4 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0073] The present invention also provides a computer device, the computer device including a memory and a processor, the memory storing computer-readable instructions, which, when executed by the processor, cause the processor to perform the steps of the image generation method in the above embodiments. The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, the computer-readable storage medium storing instructions, which, when executed on a computer, cause the computer to perform the steps of the image generation method.
[0074] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0075] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0076] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An image generation method characterized by, The image generation method includes: Acquire audio frames from the audio signal and video frames from the original video signal; The voice frame is converted to face grid parameters to obtain first face grid data, and the video frame is converted to face grid parameters to obtain second face grid data. If the voice frame is located in a preset transition window, then a third face grid data is generated based on the first face grid data and the second face grid data; Using the third face mesh data as a constraint, conditional rendering is performed on the preset mouth shape driving reference map to obtain a rendered image of the mouth region; The rendered image of the mouth area and the video frame are fused together to obtain the mouth shape driven image of the current frame.
2. The image generation method according to claim 1, characterized by, The image generation also includes: Speech activity detection is performed on the speech signal, and based on the speech activity detection results, silent segments and speech segments in the speech signal are marked; At the boundary between each silent segment and each speech segment of the speech signal, a transition window of a preset number of speech frames is set. The transition window includes speech frames of some silent segments and speech frames of some speech segments.
3. The image generation method according to claim 2, characterized by, The step of generating third face grid data based on the first face grid data and the second face grid data includes: Based on the first and second face mesh data, a smooth transition process for the mouth is performed to obtain the third face mesh data.
4. The image generation method of claim 3, wherein, The step of performing mouth smoothing processing based on the first face mesh data and the second face mesh data to obtain the third face mesh data includes: If the transition window corresponds to the transition from a speech segment to a silence segment, then obtain the first weight of the first face grid data and the second weight of the second face grid data; Based on the first weight and the second weight, a first weighting strategy is adopted to perform a weighted average processing on the first face grid data and the second face grid data to obtain the third face grid data; The first weighting strategy includes: referring to the processing order of each speech frame in the speech signal, the weight of the first face grid data corresponding to each speech frame decreases frame by frame; referring to the processing order of each video frame in the original video signal, the weight of the second face grid data corresponding to each video frame increases frame by frame.
5. The image generation method of claim 3, wherein, The step of performing mouth smoothing processing based on the first face mesh data and the second face mesh data to obtain the third face mesh data further includes: If the transition window corresponds to the transition from a silent segment to a speech segment, then the third weight of the first face grid data and the fourth weight of the second face grid data are obtained; Based on the third weight and the fourth weight, a second weighting strategy is adopted to perform a weighted average processing on the first face grid data and the third face grid data to obtain the third face grid data. The second weighting strategy includes: referring to the processing order of each speech frame in the speech signal, the weight of the first face grid data corresponding to each speech frame increases frame by frame; referring to the processing order of each video frame in the original video signal, the weight of the second face grid data corresponding to each video frame decreases frame by frame.
6. The image generation method according to any one of claims 2-5, characterized in that, After fusing the rendered image of the mouth region with the video frame to obtain the mouth shape driven image of the current frame, the method further includes: If the speech frame is located in the transition window, and the transition window corresponds to the transition from a speech segment to a silence segment, then the fifth weight of the video frame and the sixth weight of the lip-sync image are obtained. Based on the fifth weight and the sixth weight, the third weighting strategy is used to perform weighted averaging on the video frame and the lip-shape driving image to obtain the optimized lip-shape driving image of the current frame. The third weighting strategy includes: referring to the processing order of each video frame in the original video signal, the weight of each video frame increases frame by frame, and the weight of each lip-sync image decreases frame by frame.
7. The image generation method according to any one of claims 2-5, characterized by, After fusing the rendered image of the mouth region with the video frame to obtain the mouth shape driven image of the current frame, the method further includes: If the speech frame is located in the transition window, and the transition window corresponds to the transition from a silent segment to a speech segment, then the seventh weight of the video frame and the eighth weight of the lip-sync image are obtained. Based on the seventh weight and the eighth weight, the fourth weighting strategy is used to perform weighted averaging on the video frame and the mouth shape driving image to obtain the optimized mouth shape driving image of the current frame. The fourth weighting strategy includes: referring to the processing order of each video frame in the original video signal, the weight of each video frame decreases frame by frame, and the weight of each lip-sync image increases frame by frame.
8. An image generation apparatus characterized by comprising: The image generation device includes: The acquisition module is used to acquire the audio frames of the audio signal and the video frames of the original video signal. The conversion module is used to convert the face grid parameters of the audio frame to obtain the first face grid data, and to convert the face grid parameters of the video frame to obtain the second face grid data. The lip-shape optimization module is used to generate third face grid data based on the first face grid data and the second face grid data if the voice frame is located in a preset transition window; The rendering module is used to perform conditional rendering on a preset mouth shape driving reference map using the third face mesh data as a constraint, so as to obtain a rendered image of the mouth region. The fusion module is used to fuse the rendered image of the mouth region and the video frame to obtain the mouth shape driven image of the current frame.
9. A computer device, characterized in that, The computer device includes: a memory and at least one processor, the memory storing instructions; the at least one processor invokes the instructions in the memory to cause the computer device to perform the image generation method as described in any one of claims 1-7.
10. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions are executed by the processor, they implement the image generation method as described in any one of claims 1-7.