Virtual digital human lip synchronization optimization method, device, equipment and storage medium

By acquiring and determining the type of audio clips, generating and optimizing 3D facial lip shape parameter frame sequences, the problems of lip jitter and unnatural transitions in 2D digital human lip synchronization are solved, achieving a more natural and smooth lip synchronization effect.

CN118678135BActive Publication Date: 2025-09-09GUANGZHOU HUYA TECH CO LTD

Patent Information

Application Number
CN202410787161.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-18
Publication Date
2025-09-09
Estimated Expiration
2044-06-18

AI Technical Summary

Technical Problem

The existing 2D digital human real-time audio feedback technology has problems with lip synchronization, such as lip jitter and unnatural transitions, which are particularly evident in real-time interactive live broadcasts.

Method used

By obtaining the type of the target audio clip and generating a 3D face mouth shape parameter frame sequence based on a preset lip synchronization optimization strategy, the Euro filter algorithm is used for smoothing. Combined with the topological structure and semantic information of the 3D face model, lip synchronization is optimized, and the corresponding 3D face mouth shape image frame sequence is generated and rendered into the virtual digital human.

Benefits of technology

It achieves the smoothness and naturalness of the virtual digital human's mouth shape, improves the alignment effect of lip synchronization, reduces mouth shape jitter and unnatural transitions, and improves the user's interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118678135B_ABST
    Figure CN118678135B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of computer vision technology, and discloses a virtual digital human lip synchronization optimization method, device, equipment and storage medium. The virtual digital human lip synchronization optimization method includes: obtaining the target audio segment to be output by the virtual digital human at the next moment; judging whether the target audio segment belongs to the audio type to be processed; if the target audio segment belongs to the audio type to be processed, then based on a preset lip synchronization optimization strategy, generating a 3D human face mouth shape parameter frame sequence corresponding to the target audio segment; based on the 3D human face mouth shape parameter frame sequence, generating a corresponding 3D human face mouth shape image frame sequence and rendering it into the virtual digital human. The present invention can adapt to various audio types and improve the fluency and naturalness of the virtual digital human's mouth shape under different audio types.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision technology, and in particular to a virtual digital human voice lip synchronization optimization method, device, equipment and storage medium. Background Art

[0002] Against the backdrop of today's rapidly developing digital and entertainment industries, real-time interactive live streaming platforms have become a new option for user entertainment and communication. In particular, live interactive broadcasts using 2D digital humans (avatars) have gained market traction due to their unique appeal and broad application potential. However, current 2D digital human real-time audio feedback technology faces several challenges, particularly in achieving lip synchronization. Existing lip synchronization technologies still suffer from issues such as lip jitter and unnatural transitions, resulting in poor driving performance. Summary of the Invention

[0003] The main purpose of the present invention is to provide a virtual digital human lip synchronization optimization method, device, equipment and storage medium, aiming to solve the technical problem of poor driving effect of existing lip synchronization technology.

[0004] A first aspect of the present invention provides a virtual digital human voice lip synchronization optimization method, the virtual digital human voice lip synchronization optimization method comprising:

[0005] Obtain the target audio clip to be output by the virtual digital human at the next moment;

[0006] Determining whether the target audio segment belongs to an audio type to be processed, wherein the audio type to be processed includes a short audio type and a silent audio type;

[0007] If the target audio segment belongs to the audio type to be processed, generating a 3D human face mouth shape parameter frame sequence corresponding to the target audio segment based on a preset lip synchronization optimization strategy;

[0008] Based on the 3D human face mouth shape parameter frame sequence, a corresponding 3D human face mouth shape image frame sequence is generated, and the 3D human face mouth shape image frame sequence is rendered into the virtual digital human.

[0009] Optionally, in a first implementation of the first aspect of the present invention, if the target audio segment belongs to the audio type to be processed, generating a 3D facial lip shape parameter frame sequence corresponding to the target audio segment based on a preset lip synchronization optimization strategy includes:

[0010] If the target audio segment is a short-duration audio segment, inputting the target audio segment into a preset audio lip shape conversion model for processing, and outputting a first 3D human face lip shape parameter sequence;

[0011] Acquire topological structure information of a 3D face model corresponding to the virtual digital human, wherein the topological structure information includes a plurality of vertices constituting the 3D face model;

[0012] Based on the semantic information corresponding to each of the vertices, identifying each first target vertex in the first 3D face mouth shape parameter frame sequence, wherein the first target vertex includes a vertex in a mouth area;

[0013] A Euro-filter algorithm is used to smooth each of the first target vertices in the first 3D human face mouth shape parameter frame sequence to obtain a 3D human face mouth shape parameter frame sequence corresponding to the target audio segment.

[0014] Optionally, in a second implementation of the first aspect of the present invention, if the target audio segment belongs to the audio type to be processed, generating a 3D facial lip shape parameter frame sequence corresponding to the target audio segment based on a preset lip synchronization optimization strategy includes:

[0015] If the target audio segment is of the silent audio type, determining whether an audio segment preceding the target audio segment is of the silent audio type;

[0016] If the audio segment preceding the target audio segment is a silent audio type, the second 3D human face mouth shape parameter frame sequence with a preset closed mouth state is used as the 3D human face mouth shape parameter frame sequence corresponding to the target audio segment.

[0017] Optionally, in a third implementation manner of the first aspect of the present invention, after determining whether an audio segment preceding the target audio segment is of the silent audio type if the target audio segment is of the silent audio type, the method further includes:

[0018] If the audio segment preceding the target audio segment is a non-silent audio type, inputting the preset all-zero audio array into the preset audio lip shape conversion model for processing, and outputting a third 3D human face lip shape parameter frame sequence;

[0019] Acquire topological structure information of a 3D face model corresponding to the virtual digital human, wherein the topological structure information includes a plurality of vertices constituting the 3D face model;

[0020] Based on the semantic information corresponding to each of the vertices, identifying each second target vertex in the third 3D face mouth shape parameter frame sequence, wherein the second target vertex includes a vertex in the mouth area;

[0021] A Euro-filter algorithm is used to smooth each of the second target vertices in the third 3D human face mouth shape parameter frame sequence to obtain a 3D human face mouth shape parameter frame sequence corresponding to the target audio segment.

[0022] Optionally, in a fourth implementation manner of the first aspect of the present invention, the audio segment corresponding to the silent audio type has an audio duration of N seconds and corresponds to K frames of video, where N is less than 1 and K is greater than 1.

[0023] Optionally, in a fifth implementation of the first aspect of the present invention, before obtaining the target audio segment to be output by the virtual digital human at the next moment, the method further includes:

[0024] Obtain multiple non-silent audio clips with time sequence and labels as original training samples, where each audio clip corresponds to a 3D face mouth shape label;

[0025] Randomly selecting one or more audio clips from each of the original training samples and performing mute processing to obtain new training samples with time sequence and labels, wherein the audio clips after mute processing retain the corresponding 3D face and mouth shape labels before mute processing;

[0026] The new training samples with time sequence and the corresponding 3D human face mouth shape labels are input into the preset network model for training to obtain a trained audio mouth shape conversion model.

[0027] Optionally, in a sixth implementation of the first aspect of the present invention, determining whether the target audio segment belongs to the audio type to be processed includes:

[0028] Performing audio duration and audio energy detection on the target audio segment respectively;

[0029] If the audio duration corresponding to the target audio segment is less than the preset duration threshold, determining that the target audio segment belongs to the short audio type;

[0030] If the audio energy corresponding to the target audio segment is less than a preset energy threshold, it is determined that the target audio segment belongs to a silent audio type.

[0031] A second aspect of the present invention provides a virtual digital human lip synchronization optimization device, the virtual digital human lip synchronization optimization device comprising:

[0032] An acquisition module is used to obtain the target audio segment to be output by the virtual digital human at the next moment;

[0033] a determination module, configured to determine whether the target audio segment belongs to an audio type to be processed, wherein the audio type to be processed includes a short audio type and a silent audio type;

[0034] an optimization module, configured to generate a 3D human face mouth shape parameter frame sequence corresponding to the target audio segment based on a preset lip synchronization optimization strategy if the target audio segment belongs to the audio type to be processed;

[0035] A rendering module is used to generate a corresponding 3D human face mouth shape image frame sequence based on the 3D human face mouth shape parameter frame sequence, and render the 3D human face mouth shape image frame sequence into the virtual digital human.

[0036] A third aspect of the present invention provides a computer device comprising: a memory and at least one processor, wherein the memory stores instructions; the at least one processor calls the instructions in the memory so that the computer device executes the above-mentioned virtual digital human lip synchronization optimization method.

[0037] A fourth aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when executed on a computer, enable the computer to execute the above-mentioned virtual digital human lip synchronization optimization method.

[0038] The embodiments of the present invention provide a method, device, equipment and storage medium for optimizing lip synchronization of a virtual digital human. When driving a virtual digital human, the target audio segment to be output by the virtual digital human at the next moment is first obtained, and the audio type to which the target audio segment belongs is determined. Since some special audio types may cause the virtual digital human to have problems such as lip jitter and unnatural transition, lip synchronization optimization needs to be performed. First, it is determined whether the target audio segment belongs to the audio type to be processed, and the audio types to be processed include short-term audio types and silent audio types. If the target audio segment belongs to the audio type to be processed, a 3D human face lip shape parameter frame sequence corresponding to the target audio segment is generated based on a preset lip synchronization optimization strategy. Finally, based on the optimized 3D human face lip shape parameter frame sequence, a corresponding 3D human face lip shape image frame sequence is generated, and the 3D human face lip shape image frame sequence is rendered into the virtual digital human. Since the 3D human face lip shape parameter frame sequence corresponding to the audio segment to be output is optimized in advance, it can better match the audio segment, thereby ensuring lip synchronization alignment. The present invention can adapt to various audio types and improve the fluency and naturalness of the lip shape of virtual digital humans under different audio types. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 This is a flowchart of a first embodiment of a method for optimizing lip synchronization of a virtual digital human according to an embodiment of the present invention;

[0040] Figure 2 1 is a flow chart of a second embodiment of a method for optimizing lip synchronization of a virtual digital human according to an embodiment of the present invention;

[0041] Figure 3 1 is a flow chart of a third embodiment of a method for optimizing lip synchronization of a virtual digital human according to an embodiment of the present invention;

[0042] Figure 44 is a flowchart of a fourth embodiment of a method for optimizing lip synchronization of a virtual digital human according to an embodiment of the present invention;

[0043] Figure 5 1 is a flowchart of a fifth embodiment of a method for optimizing lip synchronization of a virtual digital human according to an embodiment of the present invention;

[0044] Figure 6 This is a functional module diagram of an embodiment of a device for optimizing lip synchronization of a virtual digital human according to an embodiment of the present invention;

[0045] Figure 7 FIG. 1 is a schematic diagram of an embodiment of a computer device in an embodiment of the present invention. DETAILED DESCRIPTION

[0046] The terms "first," "second," "third," "fourth," and the like (if any) in the description and claims of the present invention and in the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate so that the embodiments described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "including" or "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units that are not explicitly listed or that are inherent to these processes, methods, products, or apparatus.

[0047] For ease of understanding, the specific process of the embodiment of the present invention is described below. Figure 1 , Figure 1 This is a flow chart of a first embodiment of a method for optimizing lip synchronization of a virtual digital human voice according to an embodiment of the present invention. In this embodiment, the method for optimizing lip synchronization of a virtual digital human voice includes:

[0048] 101. Obtain the target audio segment to be output by the virtual digital human at the next moment;

[0049] The virtual digital human in this embodiment is a digitized humanoid created using digital technology that closely resembles a human. The virtual digital human can interact with people through conversations, expressions, and other interactions, synchronizing audio and lip movements, creating a near-real-life interaction. The virtual digital human in this embodiment can be applied to a variety of human-computer interaction scenarios, such as live streaming.

[0050] In this embodiment, before driving the virtual digital human, the driving signal needs to be optimized first. Specifically, the image driving signal is optimized based on the audio driving signal so that the optimized image driving signal can achieve lip synchronization alignment in the display effect.

[0051] In this embodiment, when a user interacts with a virtual digital human in a live broadcast, the virtual digital human needs to make a corresponding response. Therefore, before responding, it is necessary to first obtain the target audio clip to be output by the virtual digital human at the next moment. The corresponding content of the target audio clip is the audio content that the virtual digital human needs to respond to the user at the next moment.

[0052] In this embodiment, the target audio clip is stored in chunks, each of which consists of three parts: (1) id: the chunk identifier; (2) data size: the size of the chunk's data portion, in bytes; and (3) data: the chunk's data portion. Chunking is a method of dividing a large file into multiple small chunks for transmission, thereby improving data transmission efficiency and speed.

[0053] 102. Determine whether the target audio segment belongs to a to-be-processed audio type, where the to-be-processed audio type includes a short audio type and a silent audio type;

[0054] In this embodiment, after obtaining the target audio segment to be output by the virtual digital human at the next moment through step 101, it is necessary to first determine whether lip synchronization optimization is required, specifically, to determine whether the target audio segment belongs to the audio type to be processed. If so, lip synchronization optimization is performed, otherwise, lip synchronization optimization is not performed.

[0055] In one embodiment, in step 102, it is preferred to determine whether the target audio segment belongs to the audio type to be processed by the following method:

[0056] 1021. Perform audio duration and audio energy detection on the target audio segment respectively;

[0057] 1022. If the audio duration corresponding to the target audio segment is less than a preset duration threshold, determine that the target audio segment is a short audio type;

[0058] 1023. If the audio energy corresponding to the target audio segment is less than a preset energy threshold, determine that the target audio segment belongs to a silent audio type.

[0059] In this optional embodiment, to distinguish whether the target audio segment is short or silent, audio duration and energy detection are performed separately. Prior to detection, corresponding thresholds must be set, such as a duration threshold for short audio segments and an energy threshold for silent audio. Then, audio duration detection and audio energy detection are performed on the target audio segment.

[0060] (1) Audio duration detection: This can be achieved through audio processing libraries or tools, such as Python's pydub, librosa, or ffmpeg. The duration of the target audio can be quickly obtained through the audio processing library or tool, and then compared with the pre-set duration threshold. If the detected audio duration is less than or equal to the duration threshold (for example, 0.2 seconds), then the audio is short-duration audio, otherwise it is long-duration audio.

[0061] (2) Audio energy detection: The target audio segment is pre-divided into multiple small frames (e.g., 20-30 milliseconds per frame). For each frame, the energy (usually the sum of the squares of all sample values) is calculated using Python libraries such as librosa, pydub, and scipy. If the energy of a frame is lower than a pre-set energy threshold, the frame is considered silent. If multiple consecutive frames of audio are considered silent (e.g., for a period exceeding a certain length, such as 1 second), the entire target audio segment is considered silent.

[0062] 103. If the target audio segment belongs to the audio type to be processed, generating a 3D human face mouth shape parameter frame sequence corresponding to the target audio segment based on a preset lip synchronization optimization strategy;

[0063] In this embodiment, short-duration audio and silent audio are used as the audio types to be processed. If the target audio segment is short-duration audio or silent audio, lip synchronization optimization of the virtual digital human is required.

[0064] Lip synchronization specifically refers to the consistency of the audio content and the character's lip movements in audio-visual media. For example, in a real-time interactive live broadcast scene of a 2D virtual digital human, the audience expects the virtual image to respond and feedback audio information in real time and accurately, such as the barrage sent by the audience or the host's voice, and the lip movements are consistent with the sound to enhance the sense of reality. However, traditional technologies have a series of challenges in dealing with lip synchronization, especially when facing real-time interactive live broadcasts, the technical requirements are more stringent. This embodiment mainly optimizes lip synchronization for the problems of jitter in the transition between short-term audio and the inability of silent audio to respond in real time. Different lip synchronization optimization strategies are adopted for different audio types. For specific optimization methods, please refer to the following embodiment description.

[0065] 104. Generate a corresponding 3D human face mouth shape image frame sequence based on the 3D human face mouth shape parameter frame sequence, and render the 3D human face mouth shape image frame sequence into the virtual digital human.

[0066] In this embodiment, the 3D face mouth parameter frame sequence is composed of multiple time-sequential 3D face mouth parameter frames, and each 3D face mouth parameter frame corresponds to a frame of 3D face mouth image. The 3D face mouth parameters are used to describe the 3D face mouth image in the form of characteristic parameters used by the face mouth. The 3D face mouth parameters of each frame are input into a pre-trained image generation model to generate the 3D face mouth image of the corresponding frame. Among them, the image generation model used in this embodiment can use an existing image generation model, so it will not be described in detail.

[0067] This embodiment is based on a preset lip synchronization optimization strategy to generate a 3D human face mouth shape parameter frame sequence corresponding to the target audio clip, that is, through the optimized 3D human face mouth shape parameter frame sequence, lip alignment synchronization is achieved from the parameter data, and a 3D human face mouth shape image frame sequence generated by the 3D human face mouth shape parameter frame sequence is used as the image driving signal, thereby achieving lip synchronization alignment between the image driving signal and the audio driving signal, and then rendering it into the virtual digital human allows the user to sensorily feel the consistent synchronization between the virtual digital human's voice and mouth shape, thereby improving the fluency and naturalness of the virtual digital human's mouth shape.

[0068] Reference Figure 2 , Figure 2 This is a flow chart of the second embodiment of the virtual digital human voice lip synchronization optimization method in an embodiment of the present invention. Compared with the first embodiment of the virtual digital human voice lip synchronization optimization method, this embodiment further details the specific optimization process of short-term audio. The transition anti-shake strategy between short-term audio segments provided in this embodiment can ensure smooth and natural output lip animation while maintaining real-time interactive characteristics. In this embodiment, the virtual digital human voice lip synchronization optimization method includes:

[0069] 201. Obtain the target audio segment to be output by the virtual digital human at the next moment;

[0070] 202. Determine whether the target audio segment belongs to a to-be-processed audio type, where the to-be-processed audio type includes a short audio type and a silent audio type;

[0071] In this embodiment, the process of obtaining the target audio segment and determining the audio type is similar to the process described in steps 101 and 102, and thus will not be described in detail.

[0072] 203. If the target audio segment is a short audio segment, input the target audio segment into a preset audio lip shape conversion model for processing, and output a first 3D human face lip shape parameter sequence;

[0073] In this embodiment, the audio lip shape conversion model is a neural network model that takes an audio clip as input and outputs a sequence of 3D facial lip shape parameter frames corresponding to each audio frame sequence in the audio clip. The 3D facial lip shape parameter frames contain 3D facial lip shape feature parameter data. The 3D facial lip shape parameter frame sequence can be used to generate image drive signals for a virtual digital human. Different facial lip shape feature parameter data corresponds to different facial lip shapes.

[0074] In this embodiment, the target audio clip to be output by the virtual digital human at the next moment is input into a pre-trained audio lip shape conversion model to convert the audio into human face lip shape parameters, thereby obtaining a first 3D human face lip shape parameter frame sequence that matches the input target audio clip. The audio lip shape conversion model in this embodiment is preferably based on the FaceFormer model, trained with a partial causal multi-head self-attention mechanism and periodic position encoding.

[0075] 204. Acquire topological structure information of the 3D face model corresponding to the virtual digital human, wherein the topological structure information includes a plurality of vertices constituting the 3D face model;

[0076] A 3D face model refers to a three-dimensional model of a face created through specific technical means. Compared to a two-dimensional face image, it has an additional dimension and can more realistically restore the three-dimensional form of the face. 3D face models can be generated in a variety of ways, including image-based modeling technology and three-dimensional scanning-based technology. The topological structure of a 3D face model specifically refers to the structural distribution on the surface of the model, that is, the wiring method. The topological structure information specifically refers to the vertices, edges and faces in the model and their mutual connections. In a 3D face model, the topological structure needs to pay special attention to the wiring of key feature areas such as the eyes, mouth, and nose to ensure the smoothness of the animation and the expressiveness of details.

[0077] In this embodiment, in order to optimize the 3D face model of the mouth area, the topological structure information of the 3D face model corresponding to the virtual digital human is obtained, and then the multiple vertices constituting the 3D face model can be optimized.

[0078] 205. Identify first target vertices in the first 3D face mouth shape parameter frame sequence based on semantic information corresponding to each vertex, wherein the first target vertices include vertices in the mouth area;

[0079] In this embodiment, in 3D face modeling, vertices are one of the basic elements that make up the model and define the spatial position of the model. For example, in model-based face reconstruction, such as using the CANDIDE-3 model, the model consists of 113 vertices and 168 faces. These vertices can be adjusted to match the features of the image to be reconstructed. By modifying the position and attributes of the vertices, the shape and expression of the model can be changed.

[0080] In 3D face modeling, semantic information refers to information with practical meaning associated with geometric elements such as model vertices and faces. Semantic information includes characteristic information such as the position, size, and shape of the facial features, as well as dynamic information such as facial expression and posture. Semantic information is crucial for achieving more realistic and natural face reconstruction and animation effects. In this embodiment, if the mouth area needs to be optimized, the vertices of the mouth area (i.e., the first target vertices) with semantic information can be identified based on the semantic information corresponding to each vertex.

[0081] 206. Use a Euro-filter algorithm to smooth each of the first target vertices in the first 3D human face mouth shape parameter frame sequence to obtain a 3D human face mouth shape parameter frame sequence corresponding to the target audio segment.

[0082] In this embodiment, after identifying the first target vertices in the first 3D face mouth parameter frame sequence through the semantic information corresponding to each vertex in the 3D face model, a Euro filtering algorithm can be applied to perform smooth transition processing on each first target vertex, thereby achieving seamless 3D mouth dynamics.

[0083] The OneEuroFilter is a low-pass filter used for real-time noise filtering that adaptively adjusts its cutoff frequency based on the speed of the signal. The OneEuroFilter utilizes exponential smoothing, based on the fact that people are more sensitive to jitter at low speeds and more sensitive to lag at high speeds. This method defines an adaptive smoothing factor that adjusts according to displacement speed. The smoothing factor α is an adaptive value, not a constant, and is dynamically calculated using information about the signal's rate of change (speed). Thus, when the signal changes slowly, the filter primarily suppresses high-frequency components to reduce jitter; when the signal changes rapidly, the filter reduces latency, allowing the signal to pass more quickly.

[0084] In this embodiment, a Eurofilter is used to smooth the vertices in the mouth area of ​​the 3D face model. This balances the stability and sensitivity of the input data, reduces the jitter of mouth movements, and ensures timely response and low processing latency. To address the issue of lip jitter in the output frame sequence between each audio clip due to short audio sequences, this embodiment does not directly process the audio level, but instead performs transition processing at the level of the virtual digital human's drive signal. This is a "what you see is what you get" transition solution, significantly reducing lip jitter and providing a smoother visual experience for the viewer.

[0085] Reference Figure 3 , Figure 3 This is a flow chart of the third embodiment of the virtual digital human lip synchronization optimization method according to the present invention. Compared to the second embodiment of the virtual digital human lip synchronization optimization method, this embodiment further details the specific optimization process for silent audio. The silent audio segment timing optimization logic and transition processing strategy provided in this embodiment are designed to address the issue of silent segments in audio input.

[0086] In this embodiment, the virtual digital human lip synchronization optimization method includes:

[0087] 301. Obtain the target audio segment to be output by the virtual digital human at the next moment;

[0088] 302. Determine whether the target audio segment belongs to a to-be-processed audio type, where the to-be-processed audio type includes a short audio type and a silent audio type;

[0089] In this embodiment, the process of obtaining the target audio segment and determining the audio type is similar to the process described in steps 101 and 102, and thus will not be described in detail.

[0090] In this embodiment, silent audio refers to audio with no audio content. If the audio content is empty within a certain time period in an audio file, it is a silent audio segment.

[0091] In one embodiment, to optimize the complete link time for long silent segments, a design method is used for silent audio segments with an audio duration of N seconds and corresponding to K frames of video, where N and K are both positive numbers, N is less than 1 and K is greater than 1. The specific values ​​can be set according to actual needs. For example, N = 0.2 and K = 5.

[0092] For example, each silent audio segment is designed with N being a 0.2-second audio duration, corresponding to K being 5 video frames. This design effectively shortens audio processing time, making the entire chain from receiving user interaction content to outputting a response as efficient as possible. This avoids delayed responses while ensuring the continuity of animation effects during live interactions with virtual digital humans.

[0093] 303. If the target audio segment is of the silent audio type, determine whether an audio segment preceding the target audio segment is of the silent audio type;

[0094] In this embodiment, in order to solve the problem of lip shape jump caused by the transition between a silent audio segment (no audio input) and a non-silent audio segment (with audio input), it is necessary to take corresponding optimization processing measures according to different situations before and after the silent state.

[0095] In this embodiment, if the audio type of the target audio segment to be output by the virtual digital human at the next moment is silent audio, it is necessary to further determine the state of the audio segment before the target audio segment (silent state or non-silent state). For example, if it is detected that there is an audio signal (non-silent state) before a silent segment, it can be determined that the previous audio segment is of non-silent audio type; otherwise, it is of silent audio type.

[0096] 304. If the audio segment preceding the target audio segment is a silent audio type, use a second 3D human face mouth shape parameter frame sequence with a preset closed mouth state as the 3D human face mouth shape parameter frame sequence corresponding to the target audio segment.

[0097] In this embodiment, if the audio segment preceding the target audio segment is of the silent audio type (no sound, no need for the virtual digital human to make lip movements), and the target audio segment is also of the silent audio type (no sound, no need for the virtual digital human to make lip movements), then the virtual digital human does not need to make lip movements for two consecutive audio segments. In other words, no additional audio lip shape conversion processing is required for the virtual digital human when outputting the target audio segment. Therefore, this embodiment directly uses the second 3D human face lip shape parameter frame sequence in the closed mouth state without lip shape movements as the 3D human face lip shape parameter frame sequence corresponding to the target audio segment. This not only further optimizes processing time and avoids unnecessary calculations, ensuring the processing efficiency of the entire system when there is no audio input, but also makes the virtual digital human's lip shape movements more natural when transitioning between two adjacent audio segments.

[0098] Reference Figure 4 , Figure 4This is a flow chart of the fourth embodiment of the virtual digital human lip synchronization optimization method according to an embodiment of the present invention. Compared to the third embodiment of the virtual digital human lip synchronization optimization method, this embodiment further details another scenario of silent audio. The silent audio segment transition processing strategy provided in this embodiment is used to address the issue of silent segments in audio input. In this embodiment, the virtual digital human lip synchronization optimization method includes:

[0099] 401. Obtain the target audio segment to be output by the virtual digital human at the next moment;

[0100] 402. Determine whether the target audio segment belongs to a to-be-processed audio type, where the to-be-processed audio type includes a short audio type and a silent audio type.

[0101] 403. If the target audio segment is of the silent audio type, determine whether an audio segment preceding the target audio segment is of the silent audio type;

[0102] In this embodiment, the process of obtaining the target audio segment, determining the audio type, and determining the audio type of the previous audio segment is similar to the process described in steps 101, 102, and 303, and thus will not be repeated.

[0103] 404. If the audio segment preceding the target audio segment is a non-silent audio type, input the preset all-zero audio array into the preset audio lip shape conversion model for processing, and output a third 3D human face lip shape parameter frame sequence;

[0104] In this embodiment, if the audio clip preceding the target audio clip is of a non-silent audio type (there is sound, and the virtual digital human needs to make lip movements), and the target audio clip is of a silent audio type (no sound, and the virtual digital human does not need to make lip movements), the transition from the virtual digital human needing to make lip movements to the virtual digital human not needing to make lip movements will cause the user to perceive that the virtual digital human has an abrupt mouth-closing movement. Therefore, it is necessary to perform lip synchronization optimization on the 3D human face lip shape corresponding to the target audio clip. Specifically, the preset all-zero audio array is input into the preset audio lip shape conversion model for processing, and then the third 3D human face lip shape parameter frame sequence corresponding to the target audio clip is output.

[0105] An all-zero audio array is an array of audio data, where all elements are zero. This array indicates the absence of an audio signal. The muted audio output by a virtual human during live interaction isn't completely silent; it's simply low in amplitude, similar to silence. Directly inputting silent audio clips into the audio lip-shape conversion model for processing could result in the virtual human's mouth shaping similar to a closed mouth during transitions between audio clips. Therefore, a preset all-zero audio array is used to replace the silent audio input into the audio lip-shape conversion model, generating a 3D human face lip-shape parameter frame sequence for a completely closed mouth.

[0106] 405. Obtain topological structure information of the 3D face model corresponding to the virtual digital human, wherein the topological structure information includes a plurality of vertices constituting the 3D face model;

[0107] 406. Identify second target vertices in the third 3D face mouth shape parameter frame sequence based on semantic information corresponding to each vertex, wherein the second target vertices include vertices in the mouth area;

[0108] 407 : Use a Euro-filter algorithm to smooth each of the second target vertices in the third 3D human face mouth shape parameter frame sequence to obtain a 3D human face mouth shape parameter frame sequence corresponding to the target audio segment.

[0109] In this embodiment, a 3D human face mouth parameter frame sequence in a completely closed mouth state is obtained through step 404. However, the transition between the virtual digital human requiring lip movements and the virtual human not requiring lip movements still causes the user to perceive an abrupt lip closure of the virtual digital human. Therefore, this embodiment further performs transition processing on the third 3D human face mouth parameter frame sequence in a completely closed mouth state obtained in step 404, thereby creating a smooth transition of the mouth shape from sound to silence, reducing the abrupt lip closure movement that the audience may perceive and enhancing the sense of naturalness during viewing. The above-mentioned smooth transition processing process for the mouth shape is similar to the process described in steps 204, 205, and 206, and will not be repeated here.

[0110] The silent audio segment transition processing strategy of this embodiment not only improves the real-time response performance technically, but also visually ensures the natural and smooth mouth movements of the virtual digital human during live interaction, greatly improving the user experience and optimizing the overall performance of the real-time live broadcast system.

[0111] Reference Figure 5 , Figure 5This is a flow chart of the fifth embodiment of the virtual digital human voice lip synchronization optimization method in an embodiment of the present invention. Compared with the first to fourth embodiments of the virtual digital human voice lip synchronization optimization method, this embodiment further provides an audio data enhancement processing strategy. The audio data enhancement processing strategy provided in this embodiment is used to solve the problem of the influence of short-term audio clip input on the learning and generalization ability of the neural network model (i.e., the audio lip shape conversion model), which leads to a decrease in the robustness of the model training. In this embodiment, the virtual digital human voice lip synchronization optimization method also includes:

[0112] 501. Acquire multiple non-silent audio clips with time sequence and labels as original training samples, wherein each audio clip corresponds to a 3D face mouth shape label;

[0113] In this embodiment, the original training samples used to train the audio lip shape conversion model are multiple non-silent audio clips with time sequence and labels, and each audio clip corresponds to a 3D face mouth shape label. The 3D face mouth shape label is represented by different 3D face mouth shape parameters.

[0114] 502. Randomly select one or more audio clips from each of the original training samples and perform mute processing to obtain new training samples with time sequence and labels, wherein the audio clips after mute processing retain the corresponding 3D face mouth shape labels before mute processing;

[0115] Taking the application of virtual digital humans in live interactive broadcasts as an example, considering the various audio interruptions that may occur in real live interactive broadcasts, the training samples used in existing model training cannot well cover various situations of audio interruptions, thus reducing the accuracy of the model output. Based on this, in order to improve the generalization ability of the model, during the training process, silent segments are randomly inserted into the audio data to simulate the various audio interruptions that may occur in real live interactive broadcasts. The specific implementation method is: randomly select one or more audio segments from each original training sample for muting. This muting process is equivalent to silently "pausing" the audio signal during training, but does not change the corresponding 3D face labels, that is, these labels still represent the 3D face mouth movements that should be in a normal speaking state.

[0116] For example, if the original training samples are Audio A, Audio B, Audio C, and Audio D, then a random audio segment is silenced, resulting in the following new training samples: Silence Audio A, Audio B, Audio C, Audio D; Audio A, Silent Audio B, Audio C, Audio D; Audio A, Audio B, Silent Audio C, Audio D; Audio A, Audio B, Audio C, Silent Audio D. And so on. By combining multiple silence intervals, more new training samples can be generated, which can better cover various audio interruption scenarios, thereby improving model training effectiveness.

[0117] 503. Input the new training samples with time sequence and the corresponding 3D human face mouth shape labels into a preset network model for training to obtain a trained audio mouth shape conversion model.

[0118] In this embodiment, the above steps 501-503 belong to the training phase of the audio lip shape conversion model and need to be performed before step 101.

[0119] The audio data enhancement strategy of this embodiment enables the network to accurately predict future lip movements even during brief silences, enhancing the network's adaptability to short audio interruptions and its prediction accuracy. Through this embodiment's silence processing approach, the model is trained to accurately infer and generate coherent lip animations despite the insertion of silence, achieving smooth transitions and avoiding jitter or unnatural artifacts in the actual output caused by short audio inputs.

[0120] In this embodiment, the audio lip-conversion model fully leverages the inherent correlation between audio and 3D facial motion sequences during training, maintaining consistency between the training and testing phases. This ensures stable and natural lip-conversion output in real time, even for short audio clips in actual use. This model training strategy increases the model's adaptability to the complex audio environments of live broadcasts, significantly improving the effectiveness and stability of virtual digital human technology.

[0121] The above describes the virtual digital human voice lip synchronization optimization method according to the embodiment of the present invention. The following describes the virtual digital human voice lip synchronization optimization device according to the embodiment of the present invention. Figure 6 , Figure 6 This is a functional module diagram of an embodiment of a virtual digital human lip synchronization optimization device according to an embodiment of the present invention. In this embodiment, the virtual digital human lip synchronization optimization device includes:

[0122] An acquisition module 601 is used to acquire a target audio segment to be output by the virtual digital human at the next moment;

[0123] A determination module 602 is configured to determine whether the target audio segment belongs to an audio type to be processed, wherein the audio type to be processed includes a short audio type and a silent audio type;

[0124] an optimization module 603 configured to generate a 3D human face mouth shape parameter frame sequence corresponding to the target audio segment based on a preset lip synchronization optimization strategy if the target audio segment belongs to the audio type to be processed;

[0125] The rendering module 604 is configured to generate a corresponding 3D human face mouth shape image frame sequence based on the 3D human face mouth shape parameter frame sequence, and render the 3D human face mouth shape image frame sequence into the virtual digital human.

[0126] Optionally, in one embodiment, the optimization module 603 includes:

[0127] A short-time audio optimization unit is used to input the target audio segment into a preset audio lip shape conversion model for processing if the target audio segment belongs to the short-time audio type, and output a first 3D human face lip shape parameter sequence; obtain the topological structure information of the 3D human face model corresponding to the virtual digital human, wherein the topological structure information includes multiple vertices constituting the 3D human face model; based on the semantic information corresponding to each of the vertices, identify each first target vertex in the first 3D human face lip shape parameter frame sequence, wherein the first target vertex includes the vertex of the mouth area; use a Euro filtering algorithm to smooth each first target vertex in the first 3D human face lip shape parameter frame sequence to obtain a 3D human face lip shape parameter frame sequence corresponding to the target audio segment.

[0128] Optionally, in one embodiment, the optimization module 603 further includes:

[0129] The silent audio optimization unit is configured to determine whether the previous audio segment of the target audio segment is of the silent audio type if the target audio segment is of the silent audio type; if the previous audio segment of the target audio segment is of the silent audio type, use the second 3D human face mouth shape parameter frame sequence with a preset closed mouth state as the 3D human face mouth shape parameter frame sequence corresponding to the target audio segment.

[0130] Optionally, in one embodiment, the silent audio optimization unit is further configured to:

[0131] If the previous audio clip of the target audio clip is a non-silent audio type, the preset all-zero audio array is input into the preset audio lip shape conversion model for processing, and a third 3D human face lip shape parameter frame sequence is output; the topological structure information of the 3D human face model corresponding to the virtual digital human is obtained, wherein the topological structure information includes multiple vertices constituting the 3D human face model; based on the semantic information corresponding to each of the vertices, each second target vertex in the third 3D human face lip shape parameter frame sequence is identified, wherein the second target vertex includes the vertex of the mouth area; each second target vertex in the third 3D human face lip shape parameter frame sequence is smoothed using a Euro-filter algorithm to obtain a 3D human face lip shape parameter frame sequence corresponding to the target audio clip.

[0132] Optionally, in one embodiment, the audio segment corresponding to the silent audio type has an audio duration of N seconds and corresponds to K frames of video, where N is less than 1 and K is greater than 1.

[0133] Optionally, in one embodiment, the virtual digital human lip synchronization optimization device further includes:

[0134] A sample enhancement module is configured to obtain a plurality of non-silent audio clips with time sequence and labels as original training samples, wherein each audio clip corresponds to a 3D human face mouth shape label; randomly select one or more audio clips from each of the original training samples and perform silence processing to obtain new training samples with time sequence and labels, wherein the audio clips after silence processing retain the 3D human face mouth shape labels corresponding to the audio clips before silence processing;

[0135] The model training module is used to input the new training samples with time sequence and the corresponding 3D face mouth shape labels into the preset network model for training to obtain a trained audio mouth shape conversion model.

[0136] Optionally, in one embodiment, the determining module 602 is configured to:

[0137] Performing audio duration and audio energy detection on the target audio segment respectively;

[0138] If the audio duration corresponding to the target audio segment is less than the preset duration threshold, determining that the target audio segment belongs to the short audio type;

[0139] If the audio energy corresponding to the target audio segment is less than a preset energy threshold, it is determined that the target audio segment belongs to a silent audio type.

[0140] Since the embodiments of the device part correspond to the embodiments of the above-mentioned method, please refer to the above-mentioned method embodiments for the introduction of the pseudo-digital human voice lip synchronization optimization device provided by the present invention. The present invention will not be repeated here. It has the same beneficial effects as the above-mentioned pseudo-digital human voice lip synchronization optimization method.

[0141] above Figure 6 The virtual digital human voice lip synchronization optimization device in the embodiment of the present invention is described in detail from the perspective of modular functional entities. The computer device in the embodiment of the present invention is described in detail from the perspective of hardware processing.

[0142] Figure 77 is a schematic diagram of the structure of a computer device provided in an embodiment of the present invention. The computer device 700 may have relatively large differences due to different configurations or performances, and may include one or more processors (central processing units, CPU) 710 (for example, one or more processors) and a memory 720, and one or more storage media 730 (for example, one or more mass storage devices) storing application programs 733 or data 732. Among them, the memory 720 and the storage medium 730 can be temporary storage or permanent storage. The program stored in the storage medium 730 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations in the computer device 700. Furthermore, the processor 710 can be configured to communicate with the storage medium 730 to execute a series of instruction operations in the storage medium 730 on the computer device 700.

[0143] The computer device 700 may further include one or more power supplies 740, one or more wired or wireless network interfaces 750, one or more input and output interfaces 760, and / or one or more operating systems 731, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc. It will be appreciated by those skilled in the art that Figure 7 The illustrated computer device structure does not limit the computer device and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0144] The present invention also provides a computer device, which includes a memory and a processor. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor executes the steps of the virtual digital human voice lip synchronization optimization method in the above-mentioned embodiments.

[0145] The present invention also provides a computer-readable storage medium, which may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions, which, when executed on a computer, cause the computer to execute the steps of the virtual digital human lip synchronization optimization method.

[0146] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0147] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc., various media that can store program code.

[0148] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that the technical solutions described in the above embodiments can still be modified, or some of the technical features thereof can be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A virtual digital human lip synchronization optimization method, characterized in that: The virtual digital human voice lip synchronization optimization method comprises: Obtain the target audio clip to be output by the virtual digital human at the next moment; Determining whether the target audio segment belongs to an audio type to be processed, wherein the audio type to be processed includes a short audio type and a silent audio type; If the target audio segment is a short-duration audio segment, inputting the target audio segment into a preset audio lip shape conversion model for processing, and outputting a first 3D human face lip shape parameter sequence; Acquire topological structure information of a 3D face model corresponding to the virtual digital human, wherein the topological structure information includes a plurality of vertices constituting the 3D face model; Based on the semantic information corresponding to each of the vertices, identifying each first target vertex in the first 3D face mouth shape parameter frame sequence, wherein the first target vertex includes a vertex in the mouth area; Using a Euro-filter algorithm to smooth each of the first target vertices in the first 3D human face mouth shape parameter frame sequence to obtain a 3D human face mouth shape parameter frame sequence corresponding to the target audio clip; If the target audio segment is of the silent audio type, determining whether an audio segment preceding the target audio segment is of the silent audio type; If the audio segment preceding the target audio segment is a silent audio type, using the second 3D human face mouth shape parameter frame sequence with a preset closed mouth state as the 3D human face mouth shape parameter frame sequence corresponding to the target audio segment; Based on the 3D human face mouth shape parameter frame sequence, a corresponding 3D human face mouth shape image frame sequence is generated, and the 3D human face mouth shape image frame sequence is rendered into the virtual digital human.

2. The virtual digital human lip synchronization optimization method according to claim 1, characterized in that: After determining whether an audio segment preceding the target audio segment is of the silent audio type if the target audio segment is of the silent audio type, the method further includes: If the audio segment preceding the target audio segment is a non-silent audio type, inputting the preset all-zero audio array into the preset audio lip shape conversion model for processing, and outputting a third 3D human face lip shape parameter frame sequence; Acquire topological structure information of a 3D face model corresponding to the virtual digital human, wherein the topological structure information includes a plurality of vertices constituting the 3D face model; Based on the semantic information corresponding to each of the vertices, identifying each second target vertex in the third 3D face mouth shape parameter frame sequence, wherein the second target vertex includes a vertex in the mouth area; A Euro-filter algorithm is used to smooth each of the second target vertices in the third 3D human face mouth shape parameter frame sequence to obtain a 3D human face mouth shape parameter frame sequence corresponding to the target audio segment.

3. The virtual digital human lip synchronization optimization method according to claim 1 or 2, characterized in that: The audio segment corresponding to the silent audio type has an audio duration of N seconds and corresponds to K frames of video, where N is less than 1 and K is greater than 1.

4. The virtual digital human lip synchronization optimization method according to claim 1, characterized in that: Before obtaining the target audio segment to be output by the virtual digital human at the next moment, the method further includes: Obtain multiple non-silent audio clips with time sequence and labels as original training samples, where each audio clip corresponds to a 3D face mouth shape label; Randomly selecting one or more audio clips from each of the original training samples and performing mute processing to obtain new training samples with time sequence and labels, wherein the audio clips after mute processing retain the corresponding 3D face and mouth shape labels before mute processing; The new training samples with time sequence and the corresponding 3D human face mouth shape labels are input into the preset network model for training to obtain a trained audio mouth shape conversion model.

5. The virtual digital human lip synchronization optimization method according to claim 1, characterized in that: Determining whether the target audio segment belongs to the audio type to be processed includes: Performing audio duration and audio energy detection on the target audio segment respectively; If the audio duration corresponding to the target audio segment is less than the preset duration threshold, determining that the target audio segment belongs to the short audio type; If the audio energy corresponding to the target audio segment is less than a preset energy threshold, it is determined that the target audio segment belongs to a silent audio type.

6. A virtual digital human lip synchronization optimization device, characterized in that: The virtual digital human lip synchronization optimization device comprises: An acquisition module is used to obtain the target audio segment to be output by the virtual digital human at the next moment; a determination module, configured to determine whether the target audio segment belongs to an audio type to be processed, wherein the audio type to be processed includes a short audio type and a silent audio type; an optimization module, configured to generate a 3D human face mouth shape parameter frame sequence corresponding to the target audio segment based on a preset lip synchronization optimization strategy if the target audio segment belongs to the audio type to be processed; A rendering module, configured to generate a corresponding 3D human face mouth shape image frame sequence based on the 3D human face mouth shape parameter frame sequence, and render the 3D human face mouth shape image frame sequence into the virtual digital human; The optimization module includes: A short-term audio optimization unit is configured to input the target audio segment into a preset audio lip shape conversion model for processing if the target audio segment belongs to a short-term audio type, and output a first 3D human face lip shape parameter sequence; obtain topological structure information of the 3D human face model corresponding to the virtual digital human, wherein the topological structure information includes a plurality of vertices constituting the 3D human face model; identify each first target vertex in the first 3D human face lip shape parameter frame sequence based on semantic information corresponding to each vertex, wherein the first target vertex includes a vertex in the mouth area; and smooth each first target vertex in the first 3D human face lip shape parameter frame sequence using a Euro filtering algorithm to obtain a 3D human face lip shape parameter frame sequence corresponding to the target audio segment; The optimization module also includes: The silent audio optimization unit is configured to determine whether the previous audio segment of the target audio segment is of the silent audio type if the target audio segment is of the silent audio type; if the previous audio segment of the target audio segment is of the silent audio type, use the second 3D human face mouth shape parameter frame sequence with a preset closed mouth state as the 3D human face mouth shape parameter frame sequence corresponding to the target audio segment.

7. A computer device, characterized in that: The computer device includes: a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory to enable the computer device to execute the virtual digital human lip synchronization optimization method according to any one of claims 1 to 5.

8. A computer-readable storage medium having instructions stored thereon, characterized in that: When the instructions are executed by the processor, the virtual digital human lip synchronization optimization method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Voice-based mouth shape and / or expression simulation method and device

    CN108538308A

Cited By

  • Wind-liquid fusion heat dissipation structure for communication equipment and control method

    CN120730704A

  • Air-liquid fusion heat dissipation structure and control method for communication device

    CN120730704B