Audio-driven object speaking video generation method and device, equipment and medium

CN122601942APending Publication Date: 2026-08-18PENG CHENG LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610606226.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

然而,这种方式建模与渲染过程通常仅聚焦于头部区域,对应非头部区域(如颈部及身体部分),往往直接采用静态或简单的二维图像进行后期拼接与融合,导致在头颈连接处极易出现明显的几何断裂、颜色不协调或运动不一致的视觉伪影,进而降低了合成的对象说话视频的精确度与自然度

Benefits of technology

[0016]本申请实施例通过获取目标音频以及对应的目标对象的人脸参数集合,并将目标音频和人脸参数集合输入至预设的目标模型中,得到对应的人脸参数网格集合;通过目标模型,确定人脸参数网格集合对应的多个面片采样点,并基于每个面片采样点对应的预设语义标签,对每个面片采样点沿法线方向进行长度缩放处理,得到初始点云集合;通过目标模型,对初始点云集合进行图像渲染对齐,得到目标点云集合;从人脸参数集合中确定眼部动作参数,基于眼部动作参数生成第一网格子集合,基于目标音频生成第二网格子集合,并基于第一网格子集合与第二网格子集合进行融合,得到形变网格集合;基于形变网格集合,对目标点云集合进行位置映射更新,得到目标图像帧,并对目标音频对应的多个目标图像帧进行依序拼接,得到目标对象在目标音频驱动下的说话视频。以此,能够通过语义引导的点云空间扩展与面部运动分区域解耦驱动,实现头颈区域的几何完整性建模与口型、表情的精准同步。具体来说,通过基于语义标签对面片采样点沿法线方向的长度缩放处理,使得点云能够根据面部与非面部区域的几何特性进行差异化空间扩展,从几何表征层面保证了头部至颈部乃至躯干的连续覆盖,避免了头颈连接处因几何缺失导致的断裂与视觉伪影;而将人脸参数网格集合解耦为眼部动作参数驱动的第一网格子集合与音频驱动的第二网格子集合进行融合,可以使得眼部区域能够依据动作参数保持稳定的表情控制,口部区域则能够充分响应音频信号实现灵活的唇形变化,从而在运动驱动层面兼顾了不同面部区域与音频相关性的差异,确保了面部整体运动的协调性与唇形同步的准确性。最后,基于形变网格集合,对目标点云集合进行位置映射更新,能够将形变网格集合作为中间几何载体,将音频驱动的非刚性形变与眼部动作参数驱动的形变融合为统一的顶点位移信息,进而通过位置映射将该形变信息精确传递至目标点云集合的每个采样点,使得高斯点云能够在保持几何连续性的同时同步响应面部动态变化,从而避免了隐式表征中因缺乏显式几何约束而导致的运动失真与细节丢失。综上,本申请能够提高合成的对象说话视频的精确度与自然度。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122601942A_ABST
    Figure CN122601942A_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide an audio-driven object speaking video generation method, device, equipment and medium. The method comprises: inputting a target audio and a face parameter set of a target object into a target model to obtain a face parameter grid set; determining, by the target model, a plurality of face sheet sampling points of the face parameter grid set, performing length scaling processing on each face sheet sampling point along a normal direction to obtain an initial point cloud set; performing image rendering alignment on the initial point cloud set by the target model to obtain a target point cloud set; generating a first grid sub-set based on eye movement parameters, generating a second grid sub-set based on the target audio, and fusing the first grid sub-set and the second grid sub-set to obtain a deformation grid set; obtaining a target image frame based on the deformation grid set, and sequentially splicing a plurality of target image frames to obtain a speaking video of the target object. In this way, the accuracy and naturalness of the synthesized object speaking video can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to an audio-driven method, apparatus, device, and medium for generating audio-driven video of an object speaking. Background Technology

[0002] Audio-driven digital human synthesis technology is a core technology in fields such as virtual reality, digital entertainment, and remote interaction. Its main goal is to automatically generate digital human speaking videos with precise lip-sync, vivid and natural expressions, and three-dimensional spatial consistency based on input speech. Therefore, achieving high-quality audio-driven speaking video synthesis is of great significance for enhancing the user's immersive experience and improving the efficiency of digital content production.

[0003] In related technologies, implicit representation based on neural radiation fields is generally used to synthesize audio-driven speaker head videos. Specifically, this method maps audio features to a dynamic neural radiation field and generates speaker facial images frame by frame using volume rendering. However, this modeling and rendering process usually focuses only on the head region. For non-head regions (such as the neck and body), static or simple two-dimensional images are often directly used for post-processing stitching and fusion. This results in visual artifacts such as obvious geometric breaks, color inconsistencies, or inconsistent motion at the head-neck junction, which in turn reduces the accuracy and naturalness of the synthesized speaking video. Summary of the Invention

[0004] This application proposes an audio-driven method, apparatus, device, and medium for generating audio-driven speech videos, which can improve the accuracy and naturalness of synthesized speech videos.

[0005] To achieve the above objectives, a first aspect of this application proposes an audio-driven method for generating audio-driven object speaking videos, the method comprising: Obtain the target audio and the corresponding set of facial parameters of the target object, and input the target audio and the set of facial parameters into a preset target model to obtain the corresponding set of facial parameter meshes; Using the target model, multiple patch sampling points corresponding to the face parameter mesh set are determined, and based on the preset semantic label corresponding to each patch sampling point, the length of each patch sampling point is scaled along the normal direction to obtain an initial point cloud set. Using the target model, the initial point cloud set is image rendered and aligned to obtain the target point cloud set; Eye movement parameters are determined from the face parameter set, a first mesh subset is generated based on the eye movement parameters, a second mesh subset is generated based on the target audio, and the first mesh subset and the second mesh subset are fused to obtain a deformable mesh set; Based on the deformable mesh set, the target point cloud set is updated by position mapping to obtain target image frames, and multiple target image frames corresponding to the target audio are sequentially stitched together to obtain the speaking video of the target object driven by the target audio.

[0006] Accordingly, a second aspect of this application provides an audio-driven object speaking video generation apparatus, the apparatus comprising: The acquisition module is used to acquire the target audio and the corresponding set of facial parameters of the target object, and input the target audio and the set of facial parameters into a preset target model to obtain the corresponding set of facial parameter meshes; The determination module is used to determine multiple patch sampling points corresponding to the face parameter mesh set through the target model, and to perform length scaling processing on each patch sampling point along the normal direction based on the preset semantic label corresponding to each patch sampling point to obtain an initial point cloud set. The alignment module is used to perform image rendering alignment on the initial point cloud set using the target model to obtain the target point cloud set; The fusion module is used to determine eye movement parameters from the face parameter set, generate a first mesh subset based on the eye movement parameters, generate a second mesh subset based on the target audio, and fuse the first mesh subset and the second mesh subset to obtain a deformable mesh set. The update module is used to update the position mapping of the target point cloud set based on the deformable mesh set to obtain the target image frame, and to sequentially stitch together multiple target image frames corresponding to the target audio to obtain the speaking video of the target object driven by the target audio.

[0007] In some implementations, the determining module is further configured to: Based on the preset semantic label corresponding to each patch sampling point, determine the length parameter and scaling parameter corresponding to each patch sampling point; Obtain the unit normal vector corresponding to each patch sampling point, and perform length scaling processing on the corresponding patch sampling points along the normal direction based on the product between the length parameter, the unit normal vector and the scaling parameter to obtain facial difference sampling points; An initial point cloud set is obtained based on the facial difference sampling points corresponding to each patch sampling point.

[0008] In some embodiments, the alignment module is further configured to: Using the target model, a corresponding filling attribute is attached to each facial difference sampling point contained in the initial point cloud set, wherein the filling attribute includes a color attribute, a radius attribute, and a density attribute; Determine the preset semantic label corresponding to each facial difference sampling point contained in the initial point cloud set, and divide the initial point cloud set into a facial point cloud subset and a non-facial point cloud subset; Using a preset point cloud rasterizer, the facial point cloud subset is rendered based on the filling attributes associated with each facial difference sampling point to obtain a corresponding first rendered image, and the non-facial point cloud subset is rendered to obtain a corresponding second rendered image. The first rendered image and the second rendered image are fused together to obtain a character rendered image; Based on the rendered image of the person, the initial point cloud set is aligned to obtain the target point cloud set.

[0009] In some implementations, the target model includes an upper face deformation network and a lower face deformation network, and the fusion module is further configured to: The eye movement parameters are input into the upper face deformation network to obtain the corresponding first mesh deformation subset; The target audio is input into the lower half-face deformation network to obtain the corresponding second mesh deformation subset; The first mesh subset is obtained by summing the preset upper surface reference mesh subset with the first mesh deformation subset; The second mesh subset is obtained by summing the preset lower reference mesh subset and the second mesh deformation subset.

[0010] In some implementations, the fusion module is further configured to: A facial mesh set is obtained by fusing the first mesh subset and the second mesh subset; Based on the set of face parameters, the pose parameters are determined and input into a preset pose mixing function to obtain a set of pose displacement meshes; Linear hybrid skinning is performed based on the facial mesh set and the pose displacement mesh set to obtain a facial deformation mesh subset. A subset of non-facial deformation meshes is determined from the set of facial parameter meshes, and linear hybrid skinning is performed based on the subset of non-facial deformation meshes and the set of pose displacement meshes to obtain the subset of non-facial deformation meshes. The deformable mesh set is obtained by fusing the facial deformable mesh subset and the non-facial deformable mesh subset.

[0011] In some implementations, the update module is further configured to: Based on the deformed mesh set, the target point cloud set is updated by position mapping to obtain the deformed Gaussian point set; Obtain the Gaussian attribute parameters corresponding to the target object, and perform Gaussian transformation on the set of deformed Gaussian points based on the Gaussian attribute parameters to obtain the target set of deformed Gaussian points; The target deformed Gaussian point set is rendered to obtain an object rendering image; Obtain the object mask of the object rendering image, and adjust the preset background image based on the difference between the preset reference matrix and the object mask to obtain the target background image; The rendered image of the object is fused with the target background image to obtain the target image frame.

[0012] In some embodiments, the audio-driven object speaking video generation apparatus further includes a training module for: Obtain sample audio and the corresponding sample face parameter set of the target object, and input the sample audio and the sample face parameter set into a preset model to obtain the corresponding sample face parameter mesh set; Using the preset model, multiple sample patch sampling points corresponding to the sample face parameter grid set are determined, and based on the sample semantic label corresponding to each sample patch sampling point, the length of each sample patch sampling point is scaled along the normal direction to obtain the initial point cloud set of the sample. Using the preset model, the initial point cloud set of samples is image rendered and aligned to obtain the sample point cloud set; The sample eye movement parameters are determined from the sample face parameter set, a first sample grid subset is generated based on the sample eye movement parameters, a second sample grid subset is generated based on the sample audio, and the first sample grid subset and the second sample grid subset are fused to obtain a sample deformation grid set. Based on the sample deformation mesh set, the position mapping of the sample point cloud set is updated to obtain the predicted image frame, and multiple predicted image frames corresponding to the sample audio are sequentially stitched together to obtain the sample speaking video of the target object driven by the sample audio. Obtain multiple reference image frames corresponding to the sample audio, and construct a target loss based on the difference between each predicted image frame and the corresponding reference image frame; Based on the target loss, the parameters of the preset model are adjusted to obtain the target model.

[0013] In some implementations, the training module is further configured to: The reconstruction sub-loss is calculated based on the pixel difference between each predicted image frame and the corresponding reference image frame; Obtain the first image feature corresponding to each predicted image frame and the second image feature of the corresponding reference image frame, and construct the perceptron loss based on the difference between the first image feature and the second image feature; Obtain the scale parameters during the location mapping update process of the sample point cloud set, and construct a regularization constraint term for the scale parameters to obtain the scaling sub-loss; For each vertex deformation vector contained in the first and second sample grid subsets, a corresponding spatial weight is determined. A scale value is calculated based on each vertex deformation vector and its corresponding spatial weight. An offset loss is calculated based on the mean of multiple scale values ​​corresponding to multiple vertex deformation vectors. The first spatial weight of each first vertex deformation vector contained in the first sample grid subset is greater than the second spatial weight of each second vertex deformation vector contained in the second sample grid subset. For the difference between any two adjacent sample deformation vertices contained in the sample deformation mesh set, a smoothing sub-loss is constructed; The target loss is constructed based on the reconstruction sub-loss, the perception sub-loss, the scaling sub-loss, the offset sub-loss, and the smoothing sub-loss.

[0014] Accordingly, a third aspect of the present application provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the audio-driven object speaking video generation method of any of the embodiments of the first aspect of the present application.

[0015] Accordingly, a fourth aspect of the embodiments of this application proposes a computer-readable storage medium storing a computer program that, when executed by a processor, implements the audio-driven object speaking video generation method of any one of the embodiments of the first aspect of this application.

[0016] This application embodiment obtains the target audio and the corresponding set of facial parameters of the target object, and inputs the target audio and the set of facial parameters into a preset target model to obtain a corresponding set of facial parameter meshes. Using the target model, multiple patch sampling points corresponding to the set of facial parameter meshes are determined, and based on the preset semantic labels corresponding to each patch sampling point, the length of each patch sampling point is scaled along the normal direction to obtain an initial point cloud set. Using the target model, the initial point cloud set is image rendered and aligned to obtain a target point cloud set. Eye motion parameters are determined from the set of facial parameters, and a first mesh subset is generated based on the eye motion parameters. A second mesh subset is generated based on the target audio, and the first and second mesh subsets are fused to obtain a deformable mesh set. Based on the deformable mesh set, the target point cloud set is updated by position mapping to obtain target image frames. Multiple target image frames corresponding to the target audio are sequentially stitched together to obtain a speaking video of the target object driven by the target audio. In this way, through semantically guided point cloud spatial expansion and facial motion regional decoupling, geometric integrity modeling of the head and neck region and precise synchronization of lip movements and expressions can be achieved. Specifically, by scaling the length of face sampling points along the normal direction based on semantic tags, the point cloud can be spatially expanded differently according to the geometric characteristics of facial and non-facial regions. This ensures continuous coverage from the head to the neck and even the torso from a geometric representation perspective, avoiding breaks and visual artifacts caused by geometric gaps at the head-neck junction. Furthermore, by decoupling the face parameter mesh set into a first mesh subset driven by eye motion parameters and a second mesh subset driven by audio, the eye region can maintain stable expression control based on motion parameters, while the mouth region can fully respond to audio signals to achieve flexible lip shape changes. This takes into account the differences in the correlation between different facial regions and audio at the motion-driven level, ensuring the coordination of overall facial movement and the accuracy of lip synchronization. Finally, based on the deformable mesh set, the target point cloud set is updated by position mapping. This allows the deformable mesh set to serve as an intermediate geometric carrier, fusing audio-driven non-rigid deformation with eye motion parameter-driven deformation into unified vertex displacement information. This deformation information is then precisely transmitted to each sampling point in the target point cloud set through position mapping. This enables the Gaussian point cloud to synchronously respond to facial dynamic changes while maintaining geometric continuity, thus avoiding motion distortion and detail loss caused by the lack of explicit geometric constraints in implicit representations. In summary, this application can improve the accuracy and naturalness of synthesized spoken video. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the architecture of the audio-driven object speaking video generation system provided in the embodiments of this application; Figure 2This is a flowchart of the audio-driven object speaking video generation method provided in the embodiments of this application; Figure 3 This is an example diagram illustrating the process of generating a target point cloud set provided in an embodiment of this application; Figure 4 This is an example diagram illustrating the process of generating a target image frame provided in an embodiment of this application; Figure 5 This is a flowchart illustrating the different stages of the processing for generating a video of an object speaking, as provided in the embodiments of this application. Figure 6 This is a general flowchart of the audio-driven object speaking video generation method provided in the embodiments of this application; Figure 7 This is a schematic diagram of the functional modules of the audio-driven object speaking video generation device provided in the embodiments of this application; Figure 8 This is a schematic diagram of the hardware structure of the computer device provided in the embodiments of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0019] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0021] Audio-driven digital human synthesis technology is a core technology in fields such as virtual reality, digital entertainment, and remote interaction. Its main goal is to automatically generate digital human speaking videos with precise lip-sync, vivid and natural expressions, and three-dimensional spatial consistency based on input speech. Therefore, achieving high-quality audio-driven speaking video synthesis is of great significance for enhancing the user's immersive experience and improving the efficiency of digital content production.

[0022] In related technologies, implicit representation based on neural radiation fields is generally used to synthesize audio-driven speaker head videos. Specifically, this method maps audio features to a dynamic neural radiation field and generates speaker facial images frame by frame using volume rendering. However, this modeling and rendering process usually focuses only on the head region. For non-head regions (such as the neck and body), static or simple two-dimensional images are often directly used for post-processing stitching and fusion. This results in visual artifacts such as obvious geometric breaks, color inconsistencies, or inconsistent motion at the head-neck junction, which in turn reduces the accuracy and naturalness of the synthesized speaking video.

[0023] Based on this, embodiments of this application provide an audio-driven method, apparatus, device, and medium for generating audio-driven speech videos, which can improve the accuracy and naturalness of synthesized speech videos.

[0024] The audio-driven object speaking video generation method, apparatus, device, and medium provided in this application are specifically described through the following embodiments. First, the audio-driven object speaking video generation system in this application is described.

[0025] Please refer to Figure 1 In some embodiments, this application provides an audio-driven object speaking video generation system, including a terminal 11 and a server 12.

[0026] In some implementations, terminal 11 can be used to collect audio data input by the user and display the generated 3D digital human speaking video. For example, it can be a hardware device with audio input and image display functions, such as a smartphone, tablet, personal computer, virtual reality headset, or augmented reality glasses. Terminal 11 can collect voice signals through a built-in or external microphone and transmit the audio data to server 12 via a network, as well as receive and render the 3D digital human video stream returned by server 12, to provide the user with a real-time, interactive audio-driven digital human generation experience.

[0027] In some implementations, server 12 can be a cloud computing server, server cluster, or distributed computing platform equipped with a high-performance GPU. Server 12 can process the audio data uploaded by terminal 11 by loading and running a pre-trained semantic-guided point cloud generation module, facial motion decoupling network, 3D Gaussian attribute optimizer, and rendering pipeline, and generate corresponding 3D digital human video sequences with high synchronization and natural posture.

[0028] Furthermore, the terminal 11 and the server 12 can communicate and exchange data via the network. The terminal 11 can upload the collected raw audio stream to the server 12. After completing audio feature extraction, semantic point cloud deformation, 3D Gaussian rendering and background fusion, the server 12 will send the generated video frames or video stream back to the terminal 11. Through the above division of labor and cooperation, the two can jointly achieve low-latency, high-quality, real-time interactive object speaking video generation.

[0029] The audio-driven object speaking video generation method in this application can be illustrated by the following embodiments.

[0030] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent will be obtained first. Furthermore, the collection, use, and processing of this data will comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user will be obtained through pop-ups or redirects to confirmation pages. Only after obtaining the user's separate permission or consent will the necessary user-related data for the normal operation of the embodiments of this application be obtained.

[0031] In this embodiment, the description will focus on an audio-driven object speaking video generation device, which can be integrated into a computer device. See [link to relevant documentation]. Figure 2 , Figure 2 This is a flowchart illustrating the steps of an audio-driven object speaking video generation method provided in this application embodiment. Taking the audio-driven object speaking video generation device specifically integrated into a terminal or server as an example, the specific process when the processor on the terminal or server executes the program instructions corresponding to the audio-driven object speaking video generation method is as follows: Step 101: Obtain the target audio and the corresponding set of face parameters of the target object, and input the target audio and the set of face parameters into the preset target model to obtain the corresponding set of face parameter meshes.

[0032] In some implementations, in order to establish a deformable geometric basis for audio-driven high-fidelity digital humans, the target audio and parameterized face model parameters can be input into a preset target model to generate an initial three-dimensional mesh (i.e., a set of face parameter meshes) that contains the identity features of the target object (such as face shape) and is in a neutral pose and expression state, thereby providing accurate and structured initial geometry for subsequent semantically perceptual shape construction and rigid-flexible coupling deformation.

[0033] The target audio can be any speech signal used to drive the digital human to produce corresponding lip movements and facial expressions; for example, it can be a .wav format audio file.

[0034] The target can be a specific individual (e.g., an actor or the user) for whom a speaking video needs to be generated.

[0035] The face parameter set can be a set of standardized low-dimensional parameter vectors used to describe the facial shape, expression, and pose of a target object, obtained by analyzing single or multiple frames of the target object using a parameterized face model (such as the FLAME model). The face parameter set can be automatically estimated from any video frame corresponding to the target object by a pre-trained face tracking model (such as DECA).

[0036] The target model can be a pre-trained neural network model system used to generate high-fidelity 3D speaking object videos from audio. The target model integrates multiple processing modules, capable of mapping input audio and facial parameters into dynamic, high-quality 3D digital human representations.

[0037] The face parameter mesh set can be the corresponding 3D mesh data generated by the target model based on the input face parameter set. For example, it can be a set of vertices and faces based on the FLAME topology. The face parameter mesh set can be used to represent the static 3D face shape of the target object under the face parameter set.

[0038] Specifically, the target audio can be a speech signal of any duration, such as an audio file in .wav or .mp3 format containing the speech of the target subject. To extract effective deep features from the target audio that can be used to drive facial movements, a pre-trained audio coding model can be used to obtain the driving signal of the target audio. For example, a pre-trained HuBERT (Hidden-unit BERT) model can be used to extract features from the audio waveform of the target audio, resulting in a temporal, high-dimensional audio feature sequence. This audio feature sequence can capture semantic information such as phonemes and intonation in the audio, providing crucial information for accurately driving lip movements and facial expressions. Therefore, the target audio can be processed to obtain the corresponding audio feature sequence before being input into the target model, or a feature extraction component can be set within the target model to process the input target audio and obtain the corresponding audio feature sequence. This application actually processes the audio feature sequence of the target audio.

[0039] In some implementations, the face parameter set can be constructed based on the parameter system of a parametric face model (e.g., the FLAME model). Specifically, it can be obtained by analyzing one or more calibration videos of the target object (e.g., short videos containing neutral expressions and various poses of the object). For example, existing 3D face reconstruction and tracking tools, such as DECA (Detailed Expression Capture and Animation), can be used to estimate the various parameters of the FLAME model from video frames. These parameters collectively constitute the face parameter set, which can include at least: shape parameters, used to characterize the unique facial contours and skeletal structure of the target object; expression parameters, used to describe deformations caused by facial muscle activity, such as smiling and frowning; and pose parameters, used to define the global rotation and translation states of the head, neck, and eyeballs.

[0040] Furthermore, the shape, expression, and pose parameters contained in the face parameter set can be input into the generation function of the FLAME model within the target model. This function, based on Linear Blend Skinning (LBS) and morphological deformation principles, outputs a 3D mesh with a standard topological structure composed of vertices and triangular faces. This 3D mesh serves as the initial face parameter mesh set. In this way, a 3D mesh reference strictly consistent with the target object's identity, initial expression, and pose can be obtained.

[0041] The above methods can provide accurate surface topology and semantic partitioning basis, thus laying a precise and stable geometric foundation for subsequent steps such as fine point cloud sampling based on this mesh set and implementing rigid-flexible coupling deformation driving, ensuring the identity consistency and motion rationality of the final synthesis result.

[0042] Step 102: Using the target model, determine multiple patch sampling points corresponding to the face parameter mesh set, and based on the preset semantic label corresponding to each patch sampling point, perform length scaling processing on each patch sampling point along the normal direction to obtain the initial point cloud set.

[0043] In some implementations, in order to generate a spatially extended set of 3D point clouds that can cover and represent fine geometry beyond the original mesh, a base set of points can be obtained by randomly sampling on the triangular facets of the face parameter mesh. Based on pre-defined semantic region labels, a semantically related controllable length scaling is applied to each sampled point along its normal direction to provide a fine and complete geometric basis for subsequent high-fidelity 3D digital object generation.

[0044] Among them, the face sampling points can be a series of three-dimensional spatial points obtained by random sampling from the triangular face patches contained in the face parameter mesh set.

[0045] The preset semantic labels can be category identifiers pre-assigned to different regions of the parameterized face mesh, used to indicate the function or attribute of different regions. For example, mesh regions can be labeled as facial regions (which can be further subdivided into mouth regions, eye regions, etc.) and non-facial regions (such as neck, back of the head, limbs, etc.) to facilitate the implementation of differentiated strategies in subsequent processing, such as controlling the degree of sampling expansion of different regions.

[0046] The initial point cloud set can be a set of three-dimensional spatial points obtained by offsetting the sampling points of each facet along the normal direction of the triangle facet. The offset length of each facet sampling point is determined by its preset semantic label, thus initially forming a more complete geometric shape coverage of the target figure.

[0047] Specifically, a series of sample points can be obtained by uniformly or randomly sampling the surface of each triangular facet in the face parameter mesh set. Meanwhile, to characterize details that the mesh cannot directly express (such as hair regions of a certain thickness), after determining the preset semantic label for each sample point, length scaling processing can be performed on each sample point along the normal direction of the triangular facet based on the preset semantic label corresponding to each sample point.

[0048] For example, for any patch sampling point Let its corresponding unit normal vector be and the length scaling factor determined according to its preset semantic tags. Then, new sampling points are generated along the normal direction. The position can be obtained by the following formula: ; in, This represents a preset or learnable length parameter along the normal direction, which defines the overall range of sampling. It can be a set of parameter values ​​that the target model learns through training and is best suited to describe the geometric details of the target object. It can be used to reconstruct the geometric shape of specific parts of the target person (such as the thickness of the hair, the fullness of the cheeks, the degree of protrusion of the glasses, etc.). The value can be determined by the preset semantic label corresponding to the sampling point of the patch. For example, for non-facial areas that require large volume expansion, such as hair or the back of the head, It can be set to a positive number (such as between 0.5 and 1.5) to create a sense of volume for things like hair; for facial skin areas, It can be set to close to 0 or a very small positive value to maintain the compactness of the surface, and is only used to increase the sampling density for areas that need to be expanded inward to model the internal structures of the oral cavity (such as teeth). It can be set to a negative value to shrink the sampling points inward.

[0049] In some implementations, for the oral cavity region, facet sampling points with tooth semantic labels can be directly generated by adding new facets representing teeth inside the face parameter mesh set (i.e., in the opposite direction of the normal), and then the facet sampling points are processed along the normal as described above.

[0050] Furthermore, after scaling the length of each sampling point along the normal direction, the corresponding facial difference sampling points can be obtained. By summing up all the scaled facial difference sampling points, an initial point cloud set can be obtained.

[0051] In some implementations, an initial point cloud set can be constructed by summing the face parameter mesh set, the corresponding multiple face patch sampling points, and the corresponding multiple facial difference sampling points scaled along the normal direction. This allows for the construction of a dense and volumetric hybrid point cloud.

[0052] Using the above methods, an initial point cloud set with a wider geometric coverage can be dynamically generated based on the topology of the parametric mesh and a semantically guided differential normal extension strategy. This point cloud set not only inherits the structured semantic information of the original mesh but also possesses the potential to represent high-frequency details and complete head and neck contours, overcoming the limitations of simple mesh representation.

[0053] In some implementations, to perform refined and differentiated geometric modeling of different facial and head-neck regions, a specific length reference and semantically related scaling factor can be configured for each semantically labeled facet sampling point, and a parametric offset can be performed along the normal direction of the triangular facet containing that point. This generates new sampling points with semantically perceptive characteristics in spatial distribution, thus initially constructing an extended point cloud that can adapt to the geometric feature requirements of different regions. For example, step 102, "based on the preset semantic label corresponding to each facet sampling point, perform length scaling processing on each facet sampling point along the normal direction to obtain an initial point cloud set," may include: (102.1) Based on the preset semantic label corresponding to each patch sampling point, determine the length parameter and scaling parameter corresponding to each patch sampling point; (102.2) Obtain the unit normal vector corresponding to each patch sampling point, and perform length scaling processing on the corresponding patch sampling points along the normal direction based on the product between the length parameter, the unit normal vector and the scaling parameter to obtain the facial difference sampling points; (102.3) Based on the facial difference sampling points corresponding to each patch sampling point, an initial point cloud set is obtained.

[0054] The length parameter can be a base distance scalar value used to control the offset of sampling points in the normal direction.

[0055] The scaling parameter can be an adjustment coefficient determined based on the preset semantic label to which the sampling point belongs. For example, for areas that need to simulate the geometry of teeth inside the oral cavity, the scaling coefficient can be set to a larger value to allow for a greater outward expansion of the sampling; for areas such as the cheeks that only require slight supplementation of details, a smaller value can be set.

[0056] The unit normal vector can be the direction vector of the 3D mesh triangular facet where the sampling point is located in space, and the magnitude (length) of the vector is 1. It defines the orientation of the facet in space and is used to determine the direction of the offset of each facet sampling point.

[0057] Among them, facial difference sampling points can be new 3D coordinate points generated by adaptively offsetting the patch sampling points of the original mesh surface along its normal direction through semantically guided length scaling processing, which are used to characterize the details and volume beyond the basic mesh.

[0058] In some implementations, to overcome the inherent limitations of facial parametric mesh sets in representing high-frequency details (such as hair and mouth) and complete body structures, this application proposes a semantically guided geometric expansion method. Specifically, instead of expanding all facial regions equally, this application assigns corresponding semantic labels (such as lips, cheeks, hair, or more specific categories like nose tip and alar) to the sampling points of each triangular facet on the facial parametric mesh set, intelligently determining the extent of expansion for each sampling point. Specifically, for sampling point regions requiring rich detail or volume (such as hair), a larger degree of normal expansion is allowed; for sampling point regions requiring precise shape preservation (such as facial contours), a smaller degree or even zero expansion is permitted. Thus, while remaining faithful to the basic facial topology, point cloud data that more completely and precisely describes the 3D shape of the target object can be adaptively generated.

[0059] For example, for any patch sampling point The new sampling points generated along the normal direction The position can be obtained by the following formula: ; in, These can be optimization parameters within the target model, representing a preset or learnable length parameter along the normal direction. For example, for a sampling point located in a hair region, It could be optimized to 0.05 meters, allowing the point to be expanded outwards to simulate the actual thickness of hair; while a sampling point located on a smooth part of the cheek, It can be optimized to be close to 0.001 meters, producing only a tiny expansion to maintain the accuracy of facial contours; Represents the unit normal vector; This represents the scaling parameters determined based on the preset semantic labels of the patch sampling points.

[0060] By employing the above methods, semantic information (preset semantic labels) can be transformed into specific parameters (scaling parameters) controlling geometric generation, enabling refined and differentiated control over the sampling expansion degree of different facial regions. This allows for significant expansion of areas requiring high-frequency details or exceeding the original mesh boundaries, such as the mouth and hairline, while appropriately supplementing main areas like the cheeks. This results in a more reasonable and targeted spatial distribution of the generated initial point cloud set. Consequently, it provides a good starting point for subsequent optimization stages, offering more comprehensive geometric coverage and already containing semantic structural information, effectively improving the geometric integrity and realistic detail of the final 3D digital human model.

[0061] Step 103: Using the target model, perform image rendering alignment on the initial point cloud set to obtain the target point cloud set.

[0062] In some implementations, in order to select and optimize a set of key points that can faithfully represent the complete 3D geometry of a target object from an initial point cloud generated by semantic guidance, which may have redundant or inaccurate distribution, the initial point cloud set can be projected into a 2D image using differentiable point cloud rendering technology. This drives the iterative optimization and selection of point cloud attributes (such as position, color, and density) to obtain a geometrically consistent, detailed, concise, and efficient 3D point cloud representation.

[0063] The target point cloud set can be the set of three-dimensional spatial points that are finally retained after the initial point cloud set has undergone alignment optimization and filtering processes based on image rendering. The points in the target point cloud set can be given optimized visual attributes (such as color and radius), and their spatial distribution and density are adjusted to model the complete geometry of the head, face and neck of a specific target person with higher fidelity and efficiency.

[0064] In some implementations, a target model can be used to attach corresponding fill attributes to each facial difference sampling point in the initial point cloud set. These fill attributes include color, radius, and density attributes, which are fixed parameters within the target model specific to the target person. For each point in the initial point cloud set, the target model can associate a color attribute with it. (It can be a 3D vector representing RGB color), radius attribute (A scalar that controls the pixel range occupied by the point during rendering) and density attribute (A scalar representing the opacity or presence confidence of the point) These properties are used to define the contribution of each point to the final rendered image.

[0065] In some implementations, based on preset semantic labels (such as lips, cheeks, hair, and neck) assigned to each point, all points whose labels belong to facial regions (e.g., regions defined within the facial vertices of the FLAME model, including cheeks, mouth, eyes, nose, etc.) can be classified into a facial point cloud subset; points whose labels belong to non-facial regions (such as hair, external ears, neck, and torso extensions) can be classified into a non-facial point cloud subset. This facilitates partitioned rendering, thereby implementing differentiated rendering and subsequent deformation strategies for different semantic regions (such as movable faces and relatively static non-faces). This ensures accurate lip-syncing while maintaining the integrity and stability of the head and torso geometry, fundamentally avoiding breaks or artifacts at the head-neck junction in the synthesis result.

[0066] Furthermore, two independent point cloud rasterizers (e.g., point cloud rasterization modules implemented using the PyTorch3D library) can be used to process the facial point cloud subset and the non-facial point cloud subset separately. Specifically, the facial point cloud subset can be processed by the first rasterizer, which can read the spatial coordinates of each facial difference sampling point in the subset. The color, radius, and density attributes are used to render these points onto a two-dimensional image plane using a differentiable rasterization algorithm, generating a first rendered image containing only the facial region. and the corresponding mask Meanwhile, the second rasterizer processes the non-facial point cloud subset in the same way, generating a second rendered image that includes parts such as hair and neck. and mask .

[0067] In some implementations, since the facial point cloud subset and the non-facial point cloud subset are complementary and seamlessly connected in 3D space (based on a unified semantic partition), the complete image of the person can be obtained by superimposing the images rendered separately by the two subsets. The fusion process can be represented as: This ensures a natural transition between the two parts at the interface, ultimately resulting in a complete rendered image of the character. .

[0068] Furthermore, after obtaining the rendered image of the person, it can be compared with the original reference frame (or a standard pose frame) of the target task corresponding to the input face parameter set to verify the rendered image. For example, the structural similarity (SSIM) between the rendered image and the reference image can be calculated. If it is higher than a threshold, the geometric consistency of the current initial point cloud set is considered good, and it can be directly output as the target point cloud set for subsequent deformation driving; otherwise, a warning can be triggered or another pre-aligned point cloud version can be selected.

[0069] In some implementations, after obtaining the rendered image of the person, it can be input together with a standard reference image of the target object (e.g., a clear frontal image extracted from a neutral expression video frame) into a pre-trained lightweight geometric alignment network. The alignment network can quickly infer a sparse 3D displacement field based on pixel-level differences between the two images (e.g., contours, key feature point locations). This 3D displacement field is a vector field defined in three-dimensional space, where each vector indicates the small three-dimensional spatial offset required for the coordinates of the corresponding point in the initial point cloud set. The points can be mapped back to the semantic identity of each point in the initial point cloud set. A one-time, minor adjustment is made to the 3D coordinates of each point in the initial point cloud set to achieve a higher degree of conformity between its projected shape and the geometric contour of the standard reference image. The point cloud set after this fine-tuning becomes the final high-quality, geometrically aligned target point cloud set used for audio driving.

[0070] In some implementations, the target point cloud set can be obtained by directly adding a corresponding filling attribute to each facial difference sampling point in the initial point cloud set through the target model, and filtering out point clouds with a preset density attribute, thus obtaining the target point cloud set.

[0071] The above methods can align the target point cloud set with the real appearance of the target person, laying a precise and compact geometric and appearance foundation for subsequent steps to convert it into a high-quality, drivable 3D Gaussian representation.

[0072] In some implementations, to efficiently and accurately optimize the initial point cloud generated by semantic guidance, aligning its rendered image with the real appearance of the target person, optimizable visual and geometric attributes can be added to the point cloud. Then, the point cloud is independently rendered and fused based on semantic label partitions. The real image serves as a supervisory signal to drive the iterative updating and filtering of these attributes, thereby achieving convergence from a coarse point cloud to a high-fidelity, structurally clear 3D shape. For example, step 103 may include: (103.1) Using the target model, attach corresponding filling attributes to each facial difference sampling point contained in the initial point cloud set, wherein the filling attributes include color attributes, radius attributes and density attributes; (103.2) Determine the preset semantic label corresponding to each facial difference sampling point contained in the initial point cloud set, and divide the initial point cloud set into a facial point cloud subset and a non-facial point cloud subset; (103.3) Using a preset point cloud rasterizer, based on the filling attributes associated with each facial difference sampling point, the facial point cloud subset is rendered to obtain the corresponding first rendered image, and the non-facial point cloud subset is rendered to obtain the corresponding second rendered image. (103.4) Perform channel fusion between the first rendered image and the second rendered image to obtain the character rendered image; (103.5) The initial point cloud set is shaped and aligned based on the rendered image of the person to obtain the target point cloud set.

[0073] The fill attribute can be a set of learnable parameters attached to each point in the initial point cloud set, used to control the visual effect and geometric influence range of that point during rendering.

[0074] The color attribute can be used to represent the color information carried by each point in the initial point cloud set (e.g., RGB three-channel values). During rendering, this attribute affects the color of that point and its surrounding area in the final rendered image of the character.

[0075] The radius attribute can be used to define the range or size of influence that each point in the initial point cloud set occupies when projected onto a two-dimensional image plane. Points with larger radii will affect the pixel color of a larger area.

[0076] The density attribute can be used to represent the confidence or contribution weight of each point's existence. It can be a scalar with a value between [0,1], representing the opacity or probability of existence of the current point, and is used to control the visibility weight of that point in the final image synthesis. During the optimization process, points with a density attribute below a certain threshold can be considered redundant points and removed, thereby simplifying the point cloud in the initial point cloud set and improving the confidence of the generated target point cloud set.

[0077] Among them, the facial point cloud subset can be a subset of all points selected from the initial point cloud set and marked as belonging to facial regions (such as cheeks, mouth, and eyes).

[0078] Among them, the non-facial point cloud subset can be a subset of all points that are selected from the initial point cloud set based on the preset semantic labels associated with the points and are marked as not belonging to the facial region (such as hair, neck, etc.).

[0079] The point cloud rasterizer can be a differentiable computer graphics module that converts 3D point cloud data with attributes (such as color and radius) into a 2D image. This process allows gradients to propagate back from the image space to the point cloud attributes, thereby enabling image-supervised optimization.

[0080] The first rendered image can be a two-dimensional image generated by a point cloud rasterizer using a subset of facial point clouds and their fill attributes. This image mainly contains facial region information of the target person.

[0081] The second rendered image can be a two-dimensional image generated by a point cloud rasterizer using a subset of non-facial point clouds and their fill attributes. This image mainly contains information about the non-facial regions (such as hair and neck) of the target person.

[0082] The character rendering image can be a complete two-dimensional character image obtained by blending the first rendering image and the second rendering image on the color channels (e.g., weighted overlay or alpha blending), which combines the rendering results of the face and non-face areas.

[0083] In some implementations, the fill attribute includes at least the color attribute. radius attribute Density attribute The fill attribute of each point in the initial point cloud set can be determined based on the point's preset semantic label. During the iteration of the model, the fill attribute corresponding to points with different semantic labels can be adjusted according to the target loss.

[0084] Please refer to Figure 3 In some implementations, based on the preset semantic labels assigned to each point (such as lips, cheeks, hair, and neck), all points whose labels belong to the facial region (e.g., the region defined within the facial vertices of the FLAME model, including cheeks, mouth, eyes, nose, etc.) can be classified into a facial point cloud subset; points whose labels belong to the non-facial region (such as hair, external ears, neck, and torso extensions) can be classified into a non-facial point cloud subset. This facilitates partitioned rendering and allows for differentiated rendering and subsequent deformation strategies for different semantic regions (such as movable faces and relatively static non-faces). This ensures accurate lip-syncing while maintaining the integrity and stability of the head and torso geometry, fundamentally preventing breaks or artifacts at the head-neck connection in the synthesis result.

[0085] Furthermore, two independent point cloud rasterizers (e.g., point cloud rasterization modules implemented using the PyTorch3D library) can be used to process the facial point cloud subset and the non-facial point cloud subset separately. Specifically, the facial point cloud subset can be processed by the first rasterizer, which can read the spatial coordinates of each facial difference sampling point in the subset. The color, radius, and density attributes are used to render these points onto a two-dimensional image plane using a differentiable rasterization algorithm, generating a first rendered image containing only the facial region. and the corresponding mask Meanwhile, the second rasterizer processes the non-facial point cloud subset in the same way, generating a second rendered image that includes parts such as hair and neck. and mask .

[0086] In some implementations, since the facial point cloud subset and the non-facial point cloud subset are complementary and seamlessly connected in 3D space (based on a unified semantic partition), the complete image of the person can be obtained by superimposing the images rendered separately by the two subsets. The fusion process can be represented as: This ensures a natural transition between the two parts at the interface, ultimately resulting in a complete rendered image of the character. .

[0087] In some implementations, a person-rendered image can be used. The difference between the image and the corresponding real video frame (which can be determined based on the task parameter set) is used to construct a loss to inversely optimize the filling attributes (color, radius, density) of each point.

[0088] Specifically, the reconstruction loss between the rendered image of the person and the real image can be calculated, such as L1 or L2 norm loss; this application does not limit the form of the loss. Backpropagation is then performed based on this loss, which affects the color attribute of each point in the initial point cloud set. radius attribute and density properties Gradients are generated. Optimizers (like Adam) can update these attribute values ​​based on the gradients. In each iteration, the density attribute of each point is examined. If its value is lower than a set threshold (e.g., 0.5), it is considered that the point contributes little to the rendering and is removed from the point cloud set, thereby achieving adaptive sparsity of the point cloud and improving representation efficiency.

[0089] Furthermore, after multiple iterations (e.g., thousands of steps), the properties of the point cloud gradually converge, ensuring that the rendered image of the person is highly consistent with the corresponding real image from multiple perspectives. At this point, the optimization process stops, and the final point cloud with stable optimized properties constitutes the target point cloud set. The points in this set are not only spatially distributed (covering the entire head and neck with appropriate density), but also possess accurate color and transparency, enabling high-quality rendering of the target object. This provides a stable and realistic 3D geometric and appearance foundation for the next stage of audio-driven dynamic deformation.

[0090] By employing the above methods, the complex overall point cloud optimization problem can be decomposed into more targeted sub-optimization problems for facial (dynamic) and non-facial (static) regions. This allows for independent rendering of different regions, enabling differentiated optimization strategies for different semantic regions. Furthermore, by leveraging differentiable rendering and real-image supervision, the geometric position (through density filtering), visual appearance (color), and projection properties (radius) of the point cloud can be optimized simultaneously. This efficiently drives the initial point cloud to converge into a geometrically accurate, realistically appearing, and structurally semantically clear target point cloud set, providing high-quality input for subsequent transformation into a driveable Gaussian representation.

[0091] Step 104: Determine eye movement parameters from the face parameter set, generate a first mesh subset based on the eye movement parameters, generate a second mesh subset based on the target audio, and fuse the first mesh subset and the second mesh subset to obtain a deformable mesh set.

[0092] In some implementations, in order to achieve refined and decoupled drive control of complex facial movements to improve the accuracy of lip-sync and the naturalness of facial expressions, facial movements can be decomposed into upper face (eye) movements that are weakly correlated with audio and lower face (mouth) movements that are strongly correlated with audio. Independent deformation networks are then driven using eye motion parameters and target audio signals to predict the corresponding local mesh deformations. Finally, these local deformations are fused to generate a complete and coordinated facial deformation mesh.

[0093] The eye movement parameters can be feature parameters extracted from the facial parameter set corresponding to the target object, used to describe the muscle activity state in the eye area, such as action unit parameters related to eyelid and eyebrow movements in a facial motion coding system. Eye movement parameters can be extracted from sample video frames corresponding to the target object, or they can be parameters set by the user. The correlation between eye movement parameters and speech audio signals is weak; they are mainly used to drive natural facial expressions such as blinking and frowning.

[0094] The first mesh subset can be a result of a series of three-dimensional displacement vectors of the vertices of the upper facial region (such as around the eyes and eyebrows) in the face mesh model, calculated by a dedicated deformation prediction network (such as an upper face deformation network) based on eye movement parameters. It can be used to characterize the local facial deformation caused by eye movement.

[0095] The second mesh subset can be a series of three-dimensional displacement vectors calculated based on the target audio through another independent deformation prediction network (the lower face deformation network), which are used to characterize the deformation of the lip and chin regions synchronized with the target audio content.

[0096] The deformable mesh set can be a complete 3D mesh obtained by applying the predicted local vertex displacements of the first and second mesh subsets to the corresponding facial reference mesh regions and then fusing them. The vertex positions of the deformable mesh set have changed relative to the facial reference mesh, and it contains composite facial expressions driven by the eyes and audio respectively.

[0097] In some implementations, eye movement parameters may include at least eye movement unit signals (encoded vectors describing actions such as eyelid opening and closing, eyebrow raising, etc.). Specifically, the eye movement unit signals can be input into a pre-trained upper face deformation network of the target model, which can be based on the eye movement unit signals. Mapping is performed to predict the deformation offset of each vertex in the eye region and the related upper face, and then a first mesh deformation subset is constructed based on multiple deformation offsets. Ultimately, the first grid subset... This can be achieved by comparing the predicted offset with a subset of the baseline grid. Adding them together, we get:

[0098] in, It can represent the upper half of the face deformation network.

[0099] This enables precise and natural control of upper facial movements such as blinking and raising eyebrows, which are weakly related to the audio content but are crucial.

[0100] Furthermore, the target audio can be processed in parallel to generate a second subset of grids that drive lip movements. This process can be handled by another motion prediction subnetwork ( Specifically, the target audio (corresponding to audio feature input) can be input into the lower face deformation network. In the process, the deformation offset of each vertex of the lower half of the face is predicted, and then a second mesh deformation subset is constructed based on multiple deformation offsets. Ultimately, the second grid subset... This can be achieved by comparing the predicted offset with a subset of the baseline grid. Adding them together, we get: ; in, It can represent the lower half of the face deformation network.

[0101] This achieves accurate and synchronized phoneme-to-lip movement, ensuring the readability and authenticity of lip reading.

[0102] For example, the first and second mesh subsets generated separately can be merged to form a complete and coherent set of deformable meshes.

[0103] By employing the above methods, the originally coupled facial movements can be decoupled into independent components driven by different signal sources, and then accurately predicted using a dedicated network. This ensures that lip movements strictly follow the audio content, while allowing eye movements to occur naturally according to facial expression patterns, avoiding motion coupling anomalies that can occur with a single driving signal (such as random eye movements while speaking). This results in a deformable mesh set with accurate local motion and natural overall coordination, providing a precise geometric motion foundation for subsequently mapping detailed facial expressions onto a complete point cloud or Gaussian model to achieve high-fidelity audio-driven audio.

[0104] In some implementations, to achieve independent and accurate modeling of strongly correlated (lips) and weakly correlated (eyes) components in audio-driven lower facial motion, two proprietary networks can be designed and used to process eye motion parameters and audio signals respectively, predicting the vertex displacement increments of their corresponding facial regions. These displacement increments are then added to the corresponding neutral reference meshes to generate mesh subsets containing only local deformations, providing accurate and decoupled geometric deformation inputs for subsequent fusion and overall driving. For example, the target model may include an upper half-face deformation network and a lower half-face deformation network. Step 104, "generating a first mesh subset based on eye motion parameters and a second mesh subset based on the target audio," may include: (104.a1) Input the eye movement parameters into the upper half face deformation network to obtain the corresponding first mesh deformation subset; (104.a2) Input the target audio into the lower half-face deformation network to obtain the corresponding second mesh deformation subset; (104.a3) The first mesh subset is obtained based on the sum of the preset upper surface reference mesh subset and the first mesh deformation subset; (104.a4) The second mesh subset is obtained based on the sum of the preset lower part reference mesh subset and the second mesh deformation subset.

[0105] The upper face deformation network can be a sub-network within the target model, specifically designed to predict the required 3D deformation displacements of vertices in the upper face region (e.g., around the eyes and eyebrows) of the face mesh model based on input eye movement parameters. This network is independent of the lower face deformation network and is used to handle facial expressions that are weakly correlated with speech content.

[0106] The first mesh deformation subset can be a set of three-dimensional displacement vectors predicted for each vertex in the preset upper face reference mesh subset, output by the upper face deformation network.

[0107] Among them, the lower face deformation network can be a sub-network in the target model, which is specifically used to predict the three-dimensional deformation displacement required for the vertices of the lower face region (especially around the mouth and chin) of the face mesh model based on the input speech audio features.

[0108] The second mesh deformation subset can be a set of three-dimensional displacement vectors predicted for each vertex in the preset lower face reference mesh subset, output by the lower face deformation network. It can be used to characterize the amount of deformation change of the lips and surrounding area following the changes in audio content.

[0109] The first mesh subset can be a new set of three-dimensional vertex coordinates obtained by adding the displacement vectors in the first mesh deformation subset to the coordinates of the corresponding vertices in the preset upper face reference mesh subset one by one. It represents the complete upper face mesh state after applying deformation driven by eye movements.

[0110] The second mesh subset can be a new set of three-dimensional vertex coordinates obtained by adding the displacement vectors in the second mesh deformation subset to the coordinates of the corresponding vertices of the preset lower face reference mesh subset one by one. It represents the complete lower face mesh state after applying audio-driven deformation.

[0111] Among them, the upper face reference mesh subset can be a set of three-dimensional vertex coordinates corresponding to the eyes, eyebrows and surrounding areas extracted from the neutral expression mesh of the parametric face model (such as FLAME) based on semantic partitioning, which can be used as a stable geometric reference for the motion deformation of the upper face.

[0112] Among them, the lower reference mesh subset can be a set of three-dimensional vertex coordinates corresponding to the mouth, chin and surrounding areas extracted from the same neutral expression mesh based on semantic partitioning, which can be used as a stable geometric reference for audio-driven lower face deformation.

[0113] In some implementations, eye movement parameters may include at least eye movement unit signals (encoded vectors describing actions such as eyelid opening and closing, eyebrow raising, etc.). Specifically, the eye movement unit signals can be input into a pre-trained upper face deformation network of the target model, which can be based on the eye movement unit signals. Mapping is performed to predict the deformation offset of each vertex in the eye region and the related upper face, and then a first mesh deformation subset is constructed based on multiple deformation offsets. Ultimately, the first grid subset... This can be achieved by combining the predicted offset with a subset of the upper face reference mesh. Adding them together, we get: ; in, It can represent the upper half of the face deformation network.

[0114] This enables precise and natural control of upper facial movements such as blinking and raising eyebrows, which are weakly related to the audio content but are crucial.

[0115] Furthermore, the target audio can be processed in parallel to generate a second subset of grids that drive lip movements. This process can be handled by another motion prediction subnetwork ( Specifically, the target audio (corresponding to audio feature input) can be input into the lower face deformation network. In the process, the deformation offset of each vertex of the lower half of the face is predicted, and then a second mesh deformation subset is constructed based on multiple deformation offsets. Ultimately, the second grid subset... This can be achieved by comparing the predicted offset with the lower reference grid subset. Adding them together, we get: ; in, It can represent the lower half of the face deformation network.

[0116] This achieves accurate and synchronized phoneme-to-lip movement, ensuring the readability and authenticity of lip reading.

[0117] In some implementations, to enhance the network's ability to model subtle geometric changes around the eyes (such as eyelid folds and fine lines at the corners of the eyes) and overcome the spectral bias problem that neural networks tend to learn low-frequency signals, this application uses a subset of upper facial reference meshes. With the lower reference mesh subset Before each vertex coordinate is input into the corresponding deformation network, it undergoes position encoding. This encoding process maps each vertex coordinate from a low-dimensional coordinate to a high-dimensional feature space using a set of sine and cosine functions. For each vertex coordinate p (normalized to the range [-1, 1]), its encoding... It can be determined in the following ways: ; Here, L represents the number of frequency bands for positional encoding (e.g., it can be set to 10). This encoding not only preserves the original coordinate information (p) of each vertex but also introduces sine and cosine components of multiple frequencies, thus significantly improving the model's sensitivity to high-frequency spatial changes. The encoded high-dimensional vertex features are concatenated with the original vertex coordinates to preserve the original spatial information. Therefore, for the upper face, the network can accurately predict local geometric deformations caused by eyelid opening and closing, eyebrow raising, etc., based on eye movement signals; high-dimensional mapping of the lower face reference mesh subset enhances the fitting ability for high-frequency details such as lip contours and exposed tooth edges. After high-dimensional mapping of the lower face reference mesh subset and the lower face reference mesh subset, they are added to the first and second mesh deformation subsets output by the upper and lower face deformation networks, respectively.

[0118] By using the above method, two structurally independent and functionally specialized networks can be used to separate the complex mapping relationship between different driving signals and facial deformation. In this way, high-fidelity upper half facial expression mesh and highly synchronized lower half facial lip shape mesh are generated respectively. This lays a key foundation for the next step of seamlessly integrating these two decoupled, high-quality regional deformations into a coordinated and natural complete facial deformation mesh.

[0119] In some implementations, to address the coupling and fusion problem between audio-driven local non-rigid facial deformation and overall rigid head motion (such as rotation), and to achieve natural and stable head posture control, the decoupled facial expression mesh and the rigid displacement mesh calculated based on posture parameters can be subjected to linear hybrid skinning on the facial and non-facial regions, respectively. This organically integrates local detail deformation with global rigid body motion, generating a unified and coordinated complete head deformation mesh. For example, step 104, "fusion based on the first mesh subset and the second mesh subset to obtain a deformation mesh set," may include: (104.b1) A facial mesh set is obtained by fusing the first mesh subset and the second mesh subset; (104.b2) Determine the pose parameters based on the face parameter set, and input the pose parameters into the preset pose blending function to obtain the pose displacement mesh set; (104.b3) Linear hybrid skinning is performed based on the facial mesh set and the pose displacement mesh set to obtain a facial deformation mesh subset; (104.b4) Determine the non-facial deformation mesh subset from the face parameter mesh set, and perform linear hybrid skinning based on the non-facial deformation mesh subset and the pose displacement mesh set to obtain the non-facial deformation mesh subset; (104.b5) The deformable mesh set is obtained by fusing the facial deformable mesh subset and the non-facial deformable mesh subset.

[0120] The facial mesh set can be a complete and continuous three-dimensional facial mesh obtained by splicing or merging the first mesh subset (upper half face deformation mesh) and the second mesh subset (lower half face deformation mesh) at the shared vertex boundary. It can contain all facial expression changes driven by the eyes and audio respectively.

[0121] Among them, the pose parameters can be a set of parameters extracted from the face parameter set to describe the overall orientation of the head and the rotation state of the joints, such as the global head rotation and neck joint rotation parameters in the FLAME model, which can be used to control the rigid head posture of the digital human.

[0122] The pose blending function can be an algorithm that calculates the rigid displacement of each vertex in a 3D mesh due to the rotation and translation of bone joints, based on input pose parameters. It can be a linear blending skinning algorithm that calculates the new position of the vertex based on predefined skinning weights and joint transformation matrices.

[0123] The pose displacement mesh set can be a set of three-dimensional displacement vectors of each vertex relative to its neutral position, which are output after processing the reference mesh (or its vertices) under neutral pose by the pose mixing function. It can be used to characterize the overall rigid motion caused by changes in head and neck pose.

[0124] The facial deformation mesh subset can be the portion of the final 3D mesh belonging to the facial region obtained after inputting the facial mesh set (which already includes surface deformation) and the pose displacement mesh set (rigid displacement) into a linear hybrid skinning process. The vertex positions of the facial deformation mesh subset are simultaneously affected by both non-rigid surface deformation and rigid pose transformation.

[0125] The non-facial deformation mesh subset can be the portion of the face parameter mesh set belonging to non-facial regions (such as the back of the head and neck), which is then input into the linear hybrid skinning process along with the pose displacement mesh set, resulting in the corresponding non-facial region portion in the final 3D mesh. The vertex positions of the non-facial deformation mesh subset are mainly affected by rigid pose transformations to maintain overall consistency with facial motion.

[0126] Please refer to Figure 4 In some implementations, to achieve a natural and stable fusion between audio-driven facial expressions and user-controlled head postures, facial soft tissue deformation (flexible parts) driven purely by audio and eye signals can be predicted in a neutral posture space. Then, through the classic linear blend skinning (LBS) technique, this deformation is combined with rigid skeletal movements controlled by posture parameters, ultimately resulting in a three-dimensional head mesh that is both expressive and posture-correct. This ensures that no matter how the head turns, facial details such as smiles and mouth shapes can change correctly, fundamentally avoiding visual artifacts caused by expressions pasted in the wrong position.

[0127] Specifically, the first grid subset of the prediction can be... With the second grid subset The data is then aggregated to obtain a set of facial meshes under neutral head poses. This forms the basis for subsequent deformations.

[0128] However, No head rotation information is yet included. To combine facial expression with the target pose, a pose-parameter-driven deformation term can be introduced to calculate vertex displacements due to skeletal joint rotation via a pose blending function using the deformation prior of the parameterized face model. This process can be represented as: ;in, , , These are the target head pose, neck pose, and eye pose parameters determined from the input set of face parameters, respectively. P can represent the pose space basis of the parametric model (such as FLAME). It is a preset pose blending function that performs linear blending on the pose basis based on the input pose parameters, and outputs a shape that is similar to the pose base. A set of pose displacement meshes of the same dimension , used to describe vertex offset caused purely by skeletal motion.

[0129] Furthermore, a linear blending skinning technique can be used to fuse the neutral expression mesh with pose displacement and apply a rigid transformation. LBS is a standard technique in computer graphics for driving character animation; it can weight and blend multiple joint transformations associated with vertices based on their skinning weights, thereby achieving smooth surface deformation. Applying the aforementioned results to the facial region, the calculation is as follows: ; In the formula, This represents an LBS function. Its input can include the coordinates of the vertex to be deformed, i.e. and The coordinates of each vertex in the array; It is a joint regressor, which can be used to calculate joint positions from vertices; It includes , , The complete pose vector, including [variables]; This can be a predefined skinning weight matrix, defining the degree to which each vertex is affected by each joint. The function's output... It is a set of facial deformation meshes in the world coordinate system that combines expression and posture.

[0130] In some implementations, deformation of non-facial regions (such as the back of the head and neck) is primarily driven by rigid head movements, without the need for complex facial expressions. Therefore, a corresponding subset of non-facial deformable meshes can be directly generated from the parametric model based on identity and pose parameters, and the same LBS rigid transformation can be applied to them to maintain continuity with facial movements. ; Here, Represents a subset of non-facial deformable meshes. It can be a base mesh for non-facial regions (head regions) directly determined from a set of facial parameter meshes based on semantic segmentation. The same LBS process as for facial regions is applied to it, the specific process of which will not be elaborated here. This ensures that when the head turns, parts such as the neck and hair can move naturally with it, so as to achieve visual continuity of the head, neck and torso.

[0131] Finally, the LBS-processed facial deformation mesh subset is merged with the original facial deformation mesh subset to obtain the complete deformation mesh set V that can be used for rendering. V=[ , ]; This set contains complete 3D geometric information of the target person under specific audio-driven facial expressions and specific head poses, and is key data connecting the preceding deformation calculations and subsequent efficient rendering.

[0132] By employing the above methods and leveraging the mature graphics technology of linear hybrid skinning, predicted local non-rigid facial deformations can be stably and naturally integrated with parametrically controlled global rigid pose motion. This ensures that when the head rotates significantly, fine local deformations such as lip movements and facial expressions correctly move along with the head coordinate system, avoiding artifacts such as head-neck separation or facial misalignment. This generates a deformable mesh set with high degrees of freedom in local (facial) detail variations and flexible overall posture parameter control, providing the final and accurate geometric driving target for achieving highly realistic and posture-controllable audio-driven digital human generation.

[0133] Step 105: Based on the deformable mesh set, update the position mapping of the target point cloud set to obtain the target image frame, and sequentially stitch together multiple target image frames corresponding to the target audio to obtain the speaking video of the target object driven by the target audio.

[0134] In some implementations, in order to ultimately transform the generated set of deformable meshes into a viewable, dynamic visual output, the geometric information of the driven deformable meshes can be mapped to update the position of the target point cloud set, thereby rendering a single frame image and combining all corresponding audio frames in sequence, thus realizing a complete synthesis process from input audio to output high-fidelity, lip-synced speaker video.

[0135] The target image frame can be a static color image generated by performing position-driven and attribute processing on points in the target point cloud set based on the deformable mesh set at the current moment, and then using a 3D-to-2D differentiable rendering technique (such as a Gaussian splash renderer). The target image frame can be used to characterize the visual state of the target object's facial expression, lip shape, and head posture at a specific moment under the driving force of the corresponding audio segment.

[0136] In this context, "speaking video" can be a dynamic video sequence formed by segmenting the features of the target audio in chronological order, generating a series of consecutive target image frames in chronological order, and then splicing and encoding these image frames in chronological order (e.g., combining them into MP4, AVI, or other format files). Speaking video can showcase the entire process of a target object seemingly speaking the target audio content, featuring synchronized audio and video.

[0137] In some implementations, the deformable mesh set can be transformed into a discrete geometric representation that can be rendered in real time and then synthesized into a final image. Since 3D Gaussian Splash (3DGS) rendering technology has significant advantages in terms of high fidelity and high efficiency, this embodiment uses a position mapping function ρ to accurately transfer continuous mesh deformation to a discrete target point cloud set, thereby driving the movement of the Gaussian center point and achieving synchronous updates of geometry and appearance.

[0138] Specifically, a stable and differentiable correspondence can be established between mesh vertices and point cloud points, specifically through... To determine the location of the Gaussian center; among which, and Together they form a deformable mesh set, representing the vertex coordinates of non-facial (such as hair and neck) and facial regions after deformation. This is the core function for implementing position mapping updates. Specifically, each point P in the target point cloud set establishes a static, semantically partitioned, and spatially distance-based correspondence with one or more vertices in the deformed mesh set (e.g., through nearest neighbor matching or barycenter coordinate encoding). Based on these predefined correspondences, we can find the mesh vertices associated with P. or The new coordinates in the matrix are used to calculate the updated position of P through interpolation. All updated points The set of points, which is the deformed Gaussian point set, constitutes the center coordinates of the 3D Gaussian sphere in the current image frame. In this way, the geometric state of the point cloud set is strictly synchronized with the deformed mesh after the process.

[0139] Furthermore, after completing the position mapping, each 3D Gaussian sphere can be rendered to give it complete visual attributes. These attributes, together with the center point, constitute the final rendering unit (the set of target deformable Gaussian points). Specifically, it can be represented as follows: ; in, These represent the color, scaling, rotation, and opacity attributes of each Gaussian center point. These attributes are learned along with the geometric network during model training and are treated as fixed parameters during inference, along with the dynamically updated center points. The combination of these factors determines the distribution of the radiation field in three-dimensional space.

[0140] Furthermore, after obtaining the target deformed Gaussian point set... Then, a differentiable Gaussian splash renderer can be used to project it onto a two-dimensional image plane to generate an initial object rendering image. By merging the object rendering image with the target background image, the corresponding target image frame can be generated.

[0141] This application embodiment obtains the target audio and the corresponding set of facial parameters of the target object, and inputs the target audio and the set of facial parameters into a preset target model to obtain a corresponding set of facial parameter meshes. Using the target model, multiple patch sampling points corresponding to the set of facial parameter meshes are determined, and based on the preset semantic labels corresponding to each patch sampling point, the length of each patch sampling point is scaled along the normal direction to obtain an initial point cloud set. Using the target model, the initial point cloud set is image rendered and aligned to obtain a target point cloud set. Eye motion parameters are determined from the set of facial parameters, and a first mesh subset is generated based on the eye motion parameters. A second mesh subset is generated based on the target audio, and the first and second mesh subsets are fused to obtain a deformable mesh set. Based on the deformable mesh set, the target point cloud set is updated by position mapping to obtain target image frames. Multiple target image frames corresponding to the target audio are sequentially stitched together to obtain a speaking video of the target object driven by the target audio. In this way, through semantically guided point cloud spatial expansion and facial motion regional decoupling, geometric integrity modeling of the head and neck region and precise synchronization of lip movements and expressions can be achieved. Specifically, by scaling the length of face sampling points along the normal direction based on semantic tags, the point cloud can be spatially expanded differently according to the geometric characteristics of facial and non-facial regions. This ensures continuous coverage from the head to the neck and even the torso from a geometric representation perspective, avoiding breaks and visual artifacts caused by geometric gaps at the head-neck junction. Furthermore, by decoupling the face parameter mesh set into a first mesh subset driven by eye motion parameters and a second mesh subset driven by audio, the eye region can maintain stable expression control based on motion parameters, while the mouth region can fully respond to audio signals to achieve flexible lip shape changes. This takes into account the differences in the correlation between different facial regions and audio at the motion-driven level, ensuring the coordination of overall facial movement and the accuracy of lip synchronization. Finally, based on the deformable mesh set, the target point cloud set is updated by position mapping. This allows the deformable mesh set to serve as an intermediate geometric carrier, fusing audio-driven non-rigid deformation with eye motion parameter-driven deformation into unified vertex displacement information. This deformation information is then precisely transmitted to each sampling point in the target point cloud set through position mapping. This enables the Gaussian point cloud to synchronously respond to facial dynamic changes while maintaining geometric continuity, thus avoiding motion distortion and detail loss caused by the lack of explicit geometric constraints in implicit representations. In summary, this application can improve the accuracy and naturalness of synthesized spoken video.

[0142] In some implementations, to efficiently and faithfully convert the 3D facial deformation geometry driven by rigid-flexible coupling into a photorealistic 2D image and achieve natural and seamless integration with the background, this step generates a Gaussian representation by mapping the deformation mesh to the point cloud, generates the foreground and its mask using a differentiable Gaussian splash renderer, and then fuses it with a mask-guided background adjustment and image inpainting network to synthesize a single-frame output image with a clear foreground, harmonious background, and natural edges. Step 105, "updating the position mapping of the target point cloud set based on the deformation mesh set to obtain the target image frame," may include: (105.1) Based on the deformed mesh set, the target point cloud set is updated by position mapping to obtain the deformed Gaussian point set; (105.2) Obtain the Gaussian attribute parameters corresponding to the target object, and perform Gaussian transformation on the set of deformed Gaussian points based on the Gaussian attribute parameters to obtain the target set of deformed Gaussian points; (105.3) Render the object image by performing image rendering on the set of Gaussian points of the target deformation; (105.4) Obtain the object mask of the object rendering image, and adjust the preset background image based on the difference between the preset reference matrix and the object mask to obtain the target background image; (105.5) The object rendering image is fused with the target background image to obtain the target image frame.

[0143] The deformed Gaussian point set can be a new set of three-dimensional spatial points obtained by updating the position of each point in the target point cloud set according to the deformation relationship of the deformed mesh set (e.g., through interpolation). Each point in the deformed Gaussian point set can be used as the initial position of the three-dimensional Gaussian representation.

[0144] The Gaussian attribute parameters can be a set of parameters used to define and describe the visual appearance and geometric properties of each Gaussian sphere in the 3D Gaussian splash representation, typically including color, scaling, rotation (orientation), and opacity. These Gaussian attribute parameters can be learned through model training.

[0145] The target deformable Gaussian point set can be a complete three-dimensional Gaussian splatter representation data set formed by combining each point in the deformable Gaussian point set with the corresponding Gaussian attribute parameters (color, scaling, rotation, opacity). This set can be directly input into the Gaussian splatter renderer for image synthesis.

[0146] The object rendering image can be a color image that only contains the target object (such as a head) after the target deformed Gaussian point set is input into a differentiable Gaussian splash renderer and undergoes a three-dimensional to two-dimensional projection and rasterization process. This image does not contain the background.

[0147] The object mask can be a binary or probabilistic mask image with the same size as the rendered image, output by the renderer at the same time as the rendered object image. In this mask, the pixel area occupied by the target object has a value of 1 (or a high value), and the background area has a value of 0 (or a low value), which is used to accurately separate the foreground and background.

[0148] The reference matrix can be a matrix with all elements equal to 1 and the same size as the object mask. In the calculation, subtracting the object mask from this matrix yields the mask for the background region, which indicates the portion of the background image that should be retained.

[0149] The target background image can be a background image that is coordinated with the foreground, obtained by adjusting a preset background image (e.g., a static background extracted from the original video) according to a background mask calculated by a reference matrix and an object mask (e.g., setting the foreground region to zero or blurring), and then inputting it into an image inpainting network (such as U-Net) for edge optimization and detail restoration.

[0150] In some implementations, the deformable mesh set can be transformed into a discrete geometric representation that can be rendered in real time and then synthesized into a final image. Since 3D Gaussian Splash (3DGS) rendering technology has significant advantages in terms of high fidelity and high efficiency, this embodiment uses a position mapping function ρ to accurately transfer continuous mesh deformation to a discrete target point cloud set, thereby driving the movement of the Gaussian center point and achieving synchronous updates of geometry and appearance.

[0151] Please refer to Figure 4 Specifically, a stable and differentiable correspondence can be established between mesh vertices and point cloud points, specifically through... To determine the location of the Gaussian center; among which, and Together they form a deformable mesh set, representing the vertex coordinates of non-facial (such as hair and neck) and facial regions after deformation. This is the core function for implementing position mapping updates. Specifically, each point P in the target point cloud set establishes a static, semantically partitioned, and spatially distance-based correspondence with one or more vertices in the deformed mesh set (e.g., through nearest neighbor matching or barycenter coordinate encoding). Based on these predefined correspondences, we can find the mesh vertices associated with P. or The new coordinates in the matrix are used to calculate the updated position of P through interpolation. All updated points The set of points, which is the deformed Gaussian point set, constitutes the center coordinates of the 3D Gaussian sphere in the current image frame. In this way, the geometric state of the point cloud set is strictly synchronized with the deformed mesh after the process.

[0152] Furthermore, after completing the position mapping, each 3D Gaussian sphere can be rendered to give it complete visual attributes. These attributes, together with the center point, constitute the final rendering unit (the set of target deformable Gaussian points). Specifically, it can be represented as follows: ; in, These represent the color, scaling, rotation, and opacity attributes of each Gaussian center point. These attributes are learned along with the geometric network during model training and are treated as fixed parameters during inference, along with the dynamically updated center points. The combination of these factors determines the distribution of the radiation field in three-dimensional space.

[0153] In some implementations, the target deformable Gaussian point set can be processed by a differentiable Gaussian splash renderer. The renderer projects all 3D Gaussian spheres onto a 2D image plane based on camera parameters. Then, through tile-based rasterization and alpha blending techniques, the contribution of each Gaussian sphere to pixel color is accumulated from front to back (or using other depth-sorting strategies), ultimately synthesizing an object rendering image with correct geometry, appearance, and transparency. At the same time, the rendering process will output a result that is similar to... Object masks of the same size Each pixel value can represent the cumulative opacity (typically between 0 and 1) contributed by the foreground Gaussian point at that location, to clearly identify the foreground area occupied by the digital human.

[0154] It should be noted that this is obtained through direct rendering. Foreground objects often have sharp edges, and their lighting and color tone may not perfectly match the background image, leading to a noticeable pasting effect when simply overlaid. To address this issue, this application first preprocesses the background image by using an object mask in the object rendering image to appropriately darken or blur the boundary areas, simulating the occlusion and shadow effect of foreground objects, making the background recede more naturally. Second, a lightweight neural network is used to soften the edges, enhance details, and harmonize the colors of the rendered foreground objects, making them visually easier to blend with the processed background.

[0155] The mathematical expression for this fusion process is as follows: ; in, It is a pre-defined lightweight U-net network, which serves as a fusion and detail repair module, responsible for optimization. To enhance visual detail and edge blending, the U-net network, which is part of the target model, has its parameters determined through model training. It is a preset background image. It is an object mask. This represents an inverse mask, with a value close to 1 in the background region, close to 0 in the foreground region, and a smooth change in the transition zone. It is used to extract the regions that need to be preserved and adjusted from the background image. This indicates the target background image; This represents the final merged target image frame where the foreground and background blend seamlessly.

[0156] Using the above methods, high-performance Gaussian splash rendering technology can be employed to rapidly render the driven 3D digital human geometry into a high-detail foreground image, simultaneously obtaining an accurate object segmentation mask. This allows for intelligent adjustment and repair of the background image based on the mask, effectively avoiding edge artifacts and color mismatches that may occur when the foreground and background are directly superimposed. Furthermore, image fusion technology is used to synthesize target image frames with a prominent foreground, a natural background, and overall visual coherence, significantly enhancing the visual realism and scene immersion of the final synthesized speaking video.

[0157] In some implementations, to enable the model to master the comprehensive ability to generate high-fidelity speaking videos from audio and human parameters (including semantically guided shape reconstruction and rigid-flexible coupled motion driving), this application constructs a sample training set containing real audio-video pairs. Using the aforementioned inventive scheme as the forward propagation process, and employing the multi-dimensional differences between the generated video and the real video as supervision signals, it drives end-to-end joint optimization of the parameters of all components in the model, thereby obtaining a target model capable of generalization. For example, the target model can be trained in the following manner: (A.1) Obtain the sample audio and the corresponding sample face parameter set of the target object, and input the sample audio and sample face parameter set into the preset model to obtain the corresponding sample face parameter mesh set; (A.2) By using a preset model, multiple sample patch sampling points corresponding to the sample face parameter grid set are determined, and based on the sample semantic label corresponding to each sample patch sampling point, the length of each sample patch sampling point is scaled along the normal direction to obtain the initial point cloud set of the sample. (A.3) Using a preset model, perform image rendering and alignment on the initial point cloud set of the samples to obtain the sample point cloud set; (A.4) Determine the sample eye action parameters from the sample face parameter set, generate the first sample grid subset based on the sample eye action parameters, generate the second sample grid subset based on the sample audio, and fuse the first sample grid subset and the second sample grid subset to obtain the sample deformation grid set; (A.5) Based on the sample deformable mesh set, the position mapping of the sample point cloud set is updated to obtain the predicted image frame, and multiple predicted image frames corresponding to the sample audio are sequentially stitched together to obtain the sample speaking video of the target object driven by the sample audio. (A.6) Obtain multiple reference image frames corresponding to the sample audio, and construct the target loss based on the difference between each predicted image frame and the corresponding reference image frame; (A.7) Based on the target loss, the parameters of the preset model are adjusted to obtain the target model.

[0158] The sample audio can be a known speech recording used during the model training phase, usually synchronized with a video, as the driving signal input in the training data.

[0159] The sample face parameter set can be a set of parameterized vectors, such as shape, expression, and pose, extracted from reference image frames synchronized with the sample audio by a parameterized face model analyzer (such as DECA).

[0160] The preset model can be an initial neural network model to be trained, whose internal parameters have not yet been optimized and need to be adjusted through the target loss during the training phase.

[0161] The sample face parameter mesh set can be the initial 3D face mesh data generated by the preset model based on the input sample face parameter set.

[0162] Among them, the sample patch sampling points can be a series of three-dimensional spatial points sampled from the triangular patches of the sample face parameter mesh set during the training process.

[0163] Among them, the sample semantic labels can be semantic category identifiers pre-assigned to different regions of the sample face parameter grid set for the training phase, and their meanings are the same as the preset semantic labels in the inference phase.

[0164] The initial point cloud set can be an unoptimized 3D point cloud set obtained by scaling the sample points of the sample patch along the normal according to their sample semantic labels during the training process.

[0165] The sample point cloud set can be the final 3D point cloud set obtained after the initial sample point cloud set has undergone image rendering alignment optimization during the training phase.

[0166] Among them, the sample eye motion parameters can be eye region motion unit parameters extracted or derived from the sample face parameter set and used to drive the movement of the upper half of the face.

[0167] The first grid subset of samples can be a grid subset generated during training based on the eye action parameters of the samples, after applying upper face deformation.

[0168] The second grid subset of samples can be a grid subset generated based on sample audio during the training process, after applying lower face deformation.

[0169] The sample deformation mesh set can be the complete head deformation mesh obtained by fusing the first sample mesh subset and the second sample mesh subset during the training process and then driving it with rigid body motion.

[0170] The predicted image frame can be a single output image corresponding to a certain moment, generated by the preset model with current parameters based on sample data during the training process.

[0171] Among them, the sample speaking video can be a dynamic video sequence spliced ​​together from a series of predicted image frames generated by a preset model based on sample audio during the training process.

[0172] The reference image frame can be a real video frame recorded synchronously with the sample audio, when the target object speaks the corresponding content, and used as the ground truth image for supervising model training.

[0173] The target loss can be a comprehensive scalar function value, which is constructed by calculating the differences between the predicted image frame and the reference image frame in many aspects (such as pixel differences, feature differences, and geometric rule differences). It is used to quantify the gap between the current model output and the real situation, so as to guide the update direction of the model parameters.

[0174] In some implementations, the process of generating sample audio and corresponding sample facial parameter sets of the target object through a preset model to form a sample speaking video is the same as the process of generating target audio and corresponding speaking video through a target model as described above. The only difference is the data state and model state. For details, please refer to the above embodiments, which will not be repeated here.

[0175] Specifically, when constructing the target loss, the reconstruction loss can first be obtained by calculating the sum of the absolute differences of all pixels between each predicted image frame and its corresponding true reference image frame. This constrains the accuracy of overall color and structure. Simultaneously, the predicted image frame and the reference image frame are input into the pre-trained VGG network to extract high-level features and calculate the mean square error between their feature maps, which serves as the perceptron loss. This enhances the visual detail and naturalness of the synthesized image; furthermore, it can calculate the regularization constraints of the scale parameters generated during the prediction process. To control the scaling of Gaussian points and avoid outliers, different weights can be assigned based on the semantic labels of facial vertices (e.g., lips and non-lip regions), and the weighted sum of squares of the predicted vertex offsets can be calculated to obtain the offset loss. This encourages flexible mouth movements while suppressing unnatural shaking in areas such as the cheeks; finally, the Laplacian smoother loss can be calculated based on the sum of squared position differences between all adjacent vertex pairs in the deformed mesh. This ensures the smoothness of the facial surface geometry. The five loss terms are then linearly combined using preset weighting coefficients to obtain the target loss used for training. During training, the gradient descent algorithm can be used, based on... Backpropagation and iterative optimization are performed on all learnable parameters (such as network weights, Gaussian properties, deformation parameters, etc.) in the preset model until the model converges. The result is the target model trained to complete the audio-driven speech video generation task.

[0176] By employing the above methods, a complete end-to-end training framework consistent with the inference process can be constructed. This allows the model to simultaneously learn all subtasks, including semantically guided point cloud reconstruction, facial motion decoupling prediction, rigid-flexible coupling deformation, and high-quality rendering, under the supervision of real data. In this way, the parameters of each part of the model can be optimized collaboratively, resulting in a target model that possesses the generalization ability to robustly convert any input audio and target object parameters into high-fidelity, highly synchronous speaking videos.

[0177] In some implementations, to enable the model to learn not only pixel-level image reconstruction during training, but also high-level visual features, reasonable three-dimensional geometric deformation rules, and motion smoothness, a target loss function composed of multi-dimensional and multi-level constraints can be designed. This function simultaneously supervises model optimization from multiple perspectives, such as image appearance, geometric stability, and motion naturalness, thereby guiding the model to learn to generate speaker videos that are both realistic and conform to physical and visual laws. For example, "constructing a target loss based on the difference between each predicted image frame and the corresponding reference image frame" in (A.6) can include: (A.6.1) Calculate the reconstruction sub-loss based on the pixel difference between each predicted image frame and the corresponding reference image frame; (A.6.2) Obtain the first image feature corresponding to each predicted image frame and the second image feature of the corresponding reference image frame, and construct the perceptron loss based on the difference between the first image feature and the second image feature; (A.6.3) Obtain the scale parameters during the location mapping update process of the sample point cloud set, and construct a regularization constraint term for the scale parameters to obtain the scaling sub-loss; (A.6.4) For each vertex deformation vector contained in the first and second grid subsets of the samples, determine the corresponding spatial weight, calculate the scale value based on each vertex deformation vector and its corresponding spatial weight, and calculate the offset loss based on the mean of multiple scale values ​​corresponding to multiple vertex deformation vectors; wherein, the first spatial weight of each first vertex deformation vector contained in the first grid subset of the samples is greater than the second spatial weight of each second vertex deformation vector contained in the second grid subset of the samples; (A.6.5) Construct a smoothing sub-loss for the difference between any two adjacent sample deformation vertices contained in the sample deformation mesh set; (A.6.6) Construct the target loss based on the reconstruction sub-loss, perception sub-loss, scaling sub-loss, offset sub-loss and smoothing sub-loss.

[0178] The reconstruction sub-loss can be a loss function used to measure the difference in pixel color values ​​between the predicted image frame and the reference image frame, such as mean squared error loss. It directly constrains the generated image to be close to the real image at the pixel level, ensuring the basic structure and color accuracy of the synthesized result.

[0179] The perceptron loss can be a loss function used to measure the difference between the high-level feature maps extracted by the pre-trained deep neural network (such as the VGG network) and the reference image frame. It can be used to constrain the generated image to be perceptually similar to the real image in terms of visual content and style, thereby improving the visual realism and naturalness of the synthesized predicted image frame.

[0180] The scaling sub-loss can be a regularization constraint term applied to the scale parameter of each Gaussian sphere in the 3D Gaussian representation. It is used to prevent the Gaussian sphere from being over-scaled during the optimization process, which would lead to geometric distortion and maintain the rationality and stability of the 3D representation.

[0181] The vertex deformation vector can be a three-dimensional vector output by the deformation prediction network, representing the displacement of each vertex in the 3D mesh from its reference position. The vertex deformation vector can include a first vertex deformation vector and a second vertex deformation vector.

[0182] The spatial weight can be a scalar coefficient pre-set or learned based on the semantic region to which the facial vertex belongs (e.g., the lip region and non-lip region), used to apply differentiated penalty strength to the vertex deformation magnitude of different regions in the loss function. The spatial weight includes a first spatial weight and a second spatial weight.

[0183] Among them, the offset loss can be a geometric regularization loss that applies a weighted penalty to the magnitude of the vertex deformation vector based on spatial weights. It can maintain the stability of other facial regions while ensuring the flexibility of the mouth shape by assigning lower weights to the mouth region to allow for large deformations and higher weights to non-mouth regions to suppress unreasonable jitter.

[0184] The first vertex deformation vector can be the vertex deformation vector corresponding to the vertex belonging to the first grid subset of the sample (i.e., the upper half of the face region).

[0185] The first spatial weight can be a spatial weight applied to the deformation vector of the first vertex. Since the upper face motion is weakly correlated with audio, this weight can be set to a large penalty coefficient to constrain excessive or unnatural deformation in this region.

[0186] The second vertex deformation vector can be the vertex deformation vector corresponding to the vertex belonging to the second grid subset of the sample (i.e., the lower half of the face region).

[0187] The second spatial weight can be a spatial weight applied to the deformation vector of the second vertex. To accurately fit complex mouth shape changes, this weight can be set to a very small value to maximize the deformation degrees of freedom of the mouth region.

[0188] Among them, the smoothing sub-loss can be a loss function used to constrain the smoothness of the 3D mesh surface after deformation, such as the Laplacian smoothing sub-loss. It can suppress high-frequency noise and discontinuous deformation of the mesh surface by penalizing excessive displacement differences between adjacent vertices, thus ensuring the naturalness and smoothness of the synthesized expression.

[0189] In some implementations, to ensure that the predicted image frame remains consistent with the reference image frame of the real scene at the pixel level, a basic reconstruction sub-loss can be constructed. This loss directly measures the predicted image frame... With the corresponding reference image frame (i.e., the ground truth). The absolute difference in color value at each pixel forces the model to generate RGB outputs that are as close as possible to the actual captured image, ensuring the fundamental constraint of correct color and contour alignment in the synthesized result. For example, the reconstruction sub-loss... The calculation method can be as follows: ; Where t represents the time frame index, which is the summation of all predicted image frames in the image frame sequence. p can represent the pixel position in the image. and Let p represent the RGB color vectors at position p in the predicted and reference images of frame t, respectively. This is achieved by minimizing... Despite the loss, the model is able to learn accurate pixel-level color mapping.

[0190] Understandably, relying solely on pixel-level differences can sometimes lead to overly smooth results lacking texture detail. Therefore, this application also introduces a perceptron loss. Specifically, a pre-trained deep convolutional neural network (such as VGG-19) can be used to extract high-level semantic features of the predicted and reference image frames. By comparing the distances between the predicted and reference image frames in these feature spaces, the model can be guided to generate images that are visually more natural and richer in detail, rather than simply pixel-by-pixel matching. Perceptron Loss The calculation can be based on the Euclidean distance between feature maps, and can take the following form: ; in, L represents the feature map output of the l-th layer of the pre-trained VGG network, where L is the set of intermediate layers selected (e.g., relu1_1, relu2_1, etc.). These are weighting coefficients used to balance the contributions of different layers; minimizing Each predicted image frame is driven to be similar to the ground truth in perceptual properties such as texture, contour, and structure.

[0191] In some implementations, when rendering using 3D Gaussian splashing, the scale parameter of each Gaussian point... Controlling its spatial extent. To avoid excessive expansion (leading to blurred rendering) or excessive contraction (leading to numerical instability) of certain Gaussian points during training, a regularization constraint needs to be imposed on the scale parameter, namely the scaling sub-loss. The principle is to impose a soft upper limit on the scale of all Gaussian points, encouraging them to remain within a reasonable and compact range, thereby ensuring rendering stability and efficiency. This scaling sub-loss... A logarithmic penalty term can be used, which can take the following form: ; in, Represents the set of all 3D Gaussian points. It represents the number of points. It is the scale vector of the g-th Gaussian (usually three-dimensional). This indicates taking the natural logarithm for each component. This can be a preset maximum allowable scale threshold. The scaling sub-loss can penalize only Gaussian points whose scale logarithm exceeds the threshold. The squared term makes the penalty increase rapidly with the magnitude of the exceedance, effectively suppressing the occurrence of abnormally large scale values.

[0192] In some implementations, to accurately drive lip movements while maintaining stability in other facial regions, this application designs a semantically weighted offset loss. Specifically, for the lower face region (especially the mouth) driven by audio, large vertex displacements can be allowed to match complex lip movements, thus requiring smaller constraint weights; while for the upper face and cheeks, regions less correlated with audio, unnecessary jitter needs to be suppressed, thus requiring larger penalty weights. This differentiated constraint strategy utilizes a spatial weight mask bound to vertex semantic labels. To achieve this. For example, offset loss. The calculation formula can be: ; Where N represents the total number of facial vertices involved in the deformation. It is the deformation displacement vector predicted for the i-th vertex (i.e., the vertex deformation vector), output by the deformation network. It is the spatial weight corresponding to the i-th vertex.

[0193] For example, for vertices belonging to the second grid subset of the sample (i.e., the mouth region), their weights... It can be set to a minimum value to release degrees of freedom; for vertices belonging to the first grid subset of the sample (i.e., non-mouth facial regions), its weight... It can be set to a large value (such as 0.1) to suppress jitter. Therefore, the first spatial weight is greater than the second spatial weight.

[0194] In some implementations, to ensure the deformed 3D mesh surface remains smooth and avoids unnatural local wrinkles or high-frequency noise, a smoothing sub-loss, namely Laplacian smoothing loss, can be introduced. This constrains the relative position of each vertex in the mesh with its adjacent vertices, making the local curvature change of the deformed mesh gradual, conforming to the smooth characteristics of a real human face surface. For example, the smoothing sub-loss... Calculations can be performed based on the topological connectivity of the mesh, as follows: ; in, This represents the unique set of edges extracted from a parametric face model (such as FLAME), where each edge is represented by the index (i,j) of a pair of adjacent vertices. It represents the total number of edges. and These are the three-dimensional coordinate vectors of the i-th and j-th vertices after deformation (i.e., in the sample deformed mesh set). Calculate the square of the distance between these two adjacent vertices. Minimizing this loss means penalizing excessive stretching or compression between adjacent vertices, thereby causing the entire mesh surface to deform uniformly and smoothly.

[0195] In some implementations, the final target loss This is a weighted sum of the five sub-losses mentioned above, used to simultaneously optimize multiple objectives during training, including rendering quality, geometric plausibility, and deformation stability. By assigning appropriate weights to each sub-loss, the importance of different constraints can be balanced. During training, the total loss is minimized using the gradient descent algorithm to update all learnable parameters in the pre-defined model (including motion network weights, Gaussian properties, fusion network parameters, etc.). The loss formula is as follows: ; in, , , , , These are the non-negative weighted coefficients corresponding to the reconstruction sub-loss, perceptual sub-loss, scaling sub-loss, offsetting sub-loss, and smoothing sub-loss, respectively. These coefficients are hyperparameters and can be adjusted empirically or through a validation set before training. By jointly optimizing this target loss, the model can progressively learn the complex mapping from audio to a high-fidelity, lip-synced, and naturally expressive 3D speaking digital human, ultimately obtaining the trained target model. Using the target model, a speaking video of the target object driven by the target audio can be output based on the target audio and the target object's facial parameter set.

[0196] Through the above methods, a multi-layered, comprehensive supervision system can be constructed, encompassing pixel-to-perception and appearance-to-geometry. In this system, the reconstruction sub-loss and the perception sub-loss jointly ensure the global realism and detail quality of the synthesized image; the scaling sub-loss maintains the geometric rationality of the 3D representation; the offset sub-loss, combined with semantic weights, cleverly balances the flexibility of lip-syncing with the overall stability of the face; and the smoothing sub-loss further guarantees the temporal and spatial coherence of the motion. Furthermore, these loss terms work synergistically to guide the pre-defined model to learn a complex mapping relationship that can accurately fit audio-driven lip-syncing while maintaining reasonable facial structure and motion smoothness, ultimately training a high-performance, highly robust target model.

[0197] Please refer to Figure 5 and Figure 6 In some implementations, combined with Figure 5 and Figure 6 The overall process of the embodiments of this application will be described.

[0198] For example, the audio-driven object speaking video generation method provided in this application can be divided into two closely connected core stages: Stage 1 is a shape reconstruction stage based on semantically guided point cloud (SAPS), and Stage 2 is a digital generation stage based on a rigid-flexible morphological mechanism. The two stages work together, taking the target audio and the target object's facial parameter set (shown as "3DMM parameter estimation" in the figure) as input, and gradually generating a high-fidelity, lip-synced, and complete head and neck geometry speaker video through a preset target model.

[0199] Specifically, in Phase 1, a complete face parameter mesh set in neutral space can be reconstructed based on the input face parameter set (such as FLAME parameters). This mesh is shown as the "FLAME complete mesh" in the figure, containing 5023 vertices and 9976 faces. The vertices of this mesh are then divided into face region vertices and non-face region vertices according to preset semantic labels. The principle behind this division is that facial regions require fine expression-driven processing, while non-facial regions (such as the neck and back of the head) are more concerned with geometric integrity, thus laying a semantic foundation for subsequent differential processing.

[0200] Furthermore, a point cloud-based shape expansion method can be used to sample the mesh surface (corresponding to sampled points). Based on the predefined semantic labels (such as lips, cheeks, hair, etc.) corresponding to each sampled point, a learnable length scaling process is applied along its normal direction. This explicitly constructs an initial point cloud set covering the head and neck region while maintaining the topological structure. Thus, semantically guided differential spatial expansion allows the hair region to extend outwards to build a sense of volume, the oral cavity region to contract inwards to model the tooth structure, and the facial skin region to remain compact. This applies reasonable geometric extensions to different facial and non-facial regions, preventing breaks and visual artifacts caused by geometric deficiencies at the head-neck connection during subsequent rendering, ensuring the geometric integrity of the head and neck region during modeling.

[0201] For example, in stage two, rigid-flexible coupled motion-driven processing and high-quality rendering can be performed. Specifically, eye motion parameters (corresponding to eye-related AUs in the figure) and target audio can be input into the upper face deformation network (corresponding to the Spatial-AU Attention module in the figure) and the lower face deformation network (corresponding to the Spatial-AudioAttention module in the figure), respectively, to decouple the generation of mouth deformation strongly correlated with audio and periorbital facial deformation weakly correlated with audio. Then, a complete facial deformation mesh is obtained through weighted fusion. In this way, the weak correlation between upper face movements (such as blinking and raising eyebrows) and speech content can be fully considered, and the independent driving by eye motion parameters can ensure the natural stability of facial expressions; while the high correlation between lower face movements (such as lip shape changes) and audio can be considered, and the accurate mapping from phonemes to lip shapes can be achieved by processing with a dedicated network, thus taking into account the differences in the correlation between different facial regions and audio at the motion-driven level. At the same time, pose parameters can be extracted from the set of facial parameters, and the above-mentioned non-rigid facial deformation can be combined with the rigid head movement through standard linear hybrid skinning technology to finally form a unified and coordinated set of deformation meshes.

[0202] Furthermore, a spatial mapping function can be used to transfer the vertex displacements of the deformable mesh to the target point cloud set generated in stage one, driving its dynamic updates. Then, a 3D Gaussian sputtering (3DGS) renderer is used to fuse the background image to generate each frame. After all frames are stitched together sequentially, they are synthesized into a video of the target object speaking under the target audio. In this way, by using the deformable mesh set as an intermediate geometric carrier, the non-rigid deformation driven by audio and the deformation driven by eye motion parameters can be fused into unified vertex displacement information. The deformation information is accurately transferred to the point cloud through explicit geometric constraints, effectively avoiding motion distortion and loss of detail caused by the lack of explicit geometric constraints in implicit representation.

[0203] also, Figure 6 The end-to-end training process of the target model is further illustrated: using sample audio and the corresponding sample face parameter set of the target object, a sequence of predicted image frames is generated using the same two-stage process described above. The target loss (including pixel reconstruction loss, perceptual loss, geometric regularization loss, etc.) is calculated in multiple dimensions with reference video frames to jointly optimize the parameters of each component in the model. Specifically, the reconstruction sub-loss constrains pixel-level color accuracy, the perceptual sub-loss improves visual realism, and the offset and smoothing sub-losses ensure the rationality of deformation and surface smoothness through semantic weighting and Laplacian regularization. This allows the model to learn a robust mapping ability from audio to high-fidelity spoken video, ultimately resulting in a well-trained target model.

[0204] It should be noted that, in order to protect the privacy of the people in the images, the figures in the accompanying drawings of this patent application... Figure 3, Figure 4 and Figure 5 The figures in the image are rendered in a simplified style, but in actual operation, they are represented by real-life image frames.

[0205] Please see Figure 7 This application also provides an audio-driven object speaking video generation apparatus, which can implement the above-described audio-driven object speaking video generation method. The audio-driven object speaking video generation apparatus includes: The acquisition module 71 is used to acquire the target audio and the corresponding set of face parameters of the target object, and input the target audio and the set of face parameters into the preset target model to obtain the corresponding set of face parameter meshes; The determination module 72 is used to determine multiple patch sampling points corresponding to the face parameter mesh set through the target model, and to perform length scaling processing on each patch sampling point along the normal direction based on the preset semantic label corresponding to each patch sampling point to obtain an initial point cloud set. Alignment module 73 is used to perform image rendering alignment on the initial point cloud set using the target model to obtain the target point cloud set; The fusion module 74 is used to determine eye movement parameters from the face parameter set, generate a first grid subset based on the eye movement parameters, generate a second grid subset based on the target audio, and fuse the first grid subset and the second grid subset to obtain a deformable grid set. The update module 75 is used to update the position mapping of the target point cloud set based on the deformable mesh set to obtain the target image frame, and to sequentially stitch together multiple target image frames corresponding to the target audio to obtain the speaking video of the target object driven by the target audio.

[0206] The specific implementation of the audio-driven object speaking video generation device is basically the same as the specific embodiment of the audio-driven object speaking video generation method described above, and will not be repeated here. Subject to meeting the requirements of the embodiments of this application, the audio-driven object speaking video generation device may also be equipped with other functional modules to implement the audio-driven object speaking video generation method in the above embodiments.

[0207] This application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described audio-driven object speaking video generation method. This computer device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0208] Please see Figure 8 , Figure 8 The hardware structure of a computer device according to another embodiment is illustrated. The computer device includes: The processor 81 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 82 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 82 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 82, and the processor 81 calls and executes the audio-driven object speaking video generation method of the embodiments of this application. Input / output interface 83 is used to implement information input and output; Communication interface 84 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 85 transmits information between various components of the device (e.g., processor 81, memory 82, input / output interface 83, and communication interface 84); The processor 81, memory 82, input / output interface 83, and communication interface 84 are connected to each other within the device via bus 85.

[0209] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described audio-driven object speaking video generation method.

[0210] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0211] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0212] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0213] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0214] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0215] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0216] It should be understood that in this application, "at least one" and "several" refer to one or more, and "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0217] In the embodiments provided in this application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0218] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0219] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0220] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0221] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A method for generating audio-driven object speaking video, characterized in that, The method includes: Obtain the target audio and the corresponding set of facial parameters of the target object, and input the target audio and the set of facial parameters into a preset target model to obtain the corresponding set of facial parameter meshes; Using the target model, multiple patch sampling points corresponding to the face parameter mesh set are determined, and based on the preset semantic label corresponding to each patch sampling point, the length of each patch sampling point is scaled along the normal direction to obtain an initial point cloud set. Using the target model, the initial point cloud set is image rendered and aligned to obtain the target point cloud set; Eye movement parameters are determined from the face parameter set, a first mesh subset is generated based on the eye movement parameters, a second mesh subset is generated based on the target audio, and the first mesh subset and the second mesh subset are fused to obtain a deformable mesh set; Based on the deformable mesh set, the target point cloud set is updated by position mapping to obtain target image frames, and multiple target image frames corresponding to the target audio are sequentially stitched together to obtain the speaking video of the target object driven by the target audio.

2. The audio-driven object speaking video generation method according to claim 1, characterized in that, The initial point cloud set is obtained by scaling the length of each sampling point along the normal direction based on the preset semantic label corresponding to each sampling point, including: Based on the preset semantic label corresponding to each patch sampling point, determine the length parameter and scaling parameter corresponding to each patch sampling point; Obtain the unit normal vector corresponding to each patch sampling point, and perform length scaling processing on the corresponding patch sampling points along the normal direction based on the product between the length parameter, the unit normal vector and the scaling parameter to obtain facial difference sampling points; An initial point cloud set is obtained based on the facial difference sampling points corresponding to each patch sampling point.

3. The audio-driven object speaking video generation method according to claim 1, characterized in that, The step of performing image rendering and alignment on the initial point cloud set using the target model to obtain the target point cloud set includes: Using the target model, a corresponding filling attribute is attached to each facial difference sampling point contained in the initial point cloud set, wherein the filling attribute includes a color attribute, a radius attribute, and a density attribute; Determine the preset semantic label corresponding to each facial difference sampling point contained in the initial point cloud set, and divide the initial point cloud set into a facial point cloud subset and a non-facial point cloud subset; Using a preset point cloud rasterizer, the facial point cloud subset is rendered based on the filling attributes associated with each facial difference sampling point to obtain a corresponding first rendered image, and the non-facial point cloud subset is rendered to obtain a corresponding second rendered image. The first rendered image and the second rendered image are fused together to obtain a character rendered image; Based on the rendered image of the person, the initial point cloud set is aligned to obtain the target point cloud set.

4. The audio-driven object speaking video generation method according to claim 1, characterized in that, The target model includes an upper face deformation network and a lower face deformation network. The generation of a first mesh subset based on the eye movement parameters and a second mesh subset based on the target audio include: The eye movement parameters are input into the upper face deformation network to obtain the corresponding first mesh deformation subset; The target audio is input into the lower half-face deformation network to obtain the corresponding second mesh deformation subset; The first mesh subset is obtained by summing the preset upper surface reference mesh subset with the first mesh deformation subset; The second mesh subset is obtained by summing the preset lower reference mesh subset and the second mesh deformation subset.

5. The audio-driven object speaking video generation method according to claim 1, characterized in that, The process of fusing the first mesh subset and the second mesh subset to obtain a deformable mesh set includes: A facial mesh set is obtained by fusing the first mesh subset and the second mesh subset; Based on the set of face parameters, the pose parameters are determined and input into a preset pose mixing function to obtain a set of pose displacement meshes; Linear hybrid skinning is performed based on the facial mesh set and the pose displacement mesh set to obtain a facial deformation mesh subset. A subset of non-facial deformation meshes is determined from the set of facial parameter meshes, and linear hybrid skinning is performed based on the subset of non-facial deformation meshes and the set of pose displacement meshes to obtain the subset of non-facial deformation meshes. The deformable mesh set is obtained by fusing the facial deformable mesh subset and the non-facial deformable mesh subset.

6. The audio-driven object speaking video generation method according to claim 1, characterized in that, The step of updating the position mapping of the target point cloud set based on the deformed mesh set to obtain the target image frame includes: Based on the deformed mesh set, the target point cloud set is updated by position mapping to obtain the deformed Gaussian point set; Obtain the Gaussian attribute parameters corresponding to the target object, and perform Gaussian transformation on the set of deformed Gaussian points based on the Gaussian attribute parameters to obtain the target set of deformed Gaussian points; The target deformed Gaussian point set is rendered to obtain an object rendering image; Obtain the object mask of the object rendering image, and adjust the preset background image based on the difference between the preset reference matrix and the object mask to obtain the target background image; The rendered image of the object is fused with the target background image to obtain the target image frame.

7. The audio-driven object speaking video generation method according to claim 1, characterized in that, The target model is trained in the following way: Obtain sample audio and the corresponding sample face parameter set of the target object, and input the sample audio and the sample face parameter set into a preset model to obtain the corresponding sample face parameter mesh set; Using the preset model, multiple sample patch sampling points corresponding to the sample face parameter grid set are determined, and based on the sample semantic label corresponding to each sample patch sampling point, the length of each sample patch sampling point is scaled along the normal direction to obtain the initial point cloud set of the sample. Using the preset model, the initial point cloud set of samples is image rendered and aligned to obtain the sample point cloud set; The sample eye movement parameters are determined from the sample face parameter set, a first sample grid subset is generated based on the sample eye movement parameters, a second sample grid subset is generated based on the sample audio, and the first sample grid subset and the second sample grid subset are fused to obtain a sample deformation grid set. Based on the sample deformation mesh set, the position mapping of the sample point cloud set is updated to obtain the predicted image frame, and multiple predicted image frames corresponding to the sample audio are sequentially stitched together to obtain the sample speaking video of the target object driven by the sample audio. Obtain multiple reference image frames corresponding to the sample audio, and construct a target loss based on the difference between each predicted image frame and the corresponding reference image frame; Based on the target loss, the parameters of the preset model are adjusted to obtain the target model.

8. The audio-driven object speaking video generation method according to claim 7, characterized in that, The target loss is constructed based on the difference between each predicted image frame and the corresponding reference image frame, including: The reconstruction sub-loss is calculated based on the pixel difference between each predicted image frame and the corresponding reference image frame; Obtain the first image feature corresponding to each predicted image frame and the second image feature of the corresponding reference image frame, and construct the perceptron loss based on the difference between the first image feature and the second image feature; Obtain the scale parameters during the location mapping update process of the sample point cloud set, and construct a regularization constraint term for the scale parameters to obtain the scaling sub-loss; For each vertex deformation vector contained in the first and second sample grid subsets, a corresponding spatial weight is determined. A scale value is calculated based on each vertex deformation vector and its corresponding spatial weight. An offset loss is calculated based on the mean of multiple scale values ​​corresponding to multiple vertex deformation vectors. The first spatial weight of each first vertex deformation vector contained in the first sample grid subset is greater than the second spatial weight of each second vertex deformation vector contained in the second sample grid subset. For the difference between any two adjacent sample deformation vertices contained in the sample deformation mesh set, a smoothing sub-loss is constructed; The target loss is constructed based on the reconstruction sub-loss, the perception sub-loss, the scaling sub-loss, the offset sub-loss, and the smoothing sub-loss.

9. An audio-driven object speaking video generation device, characterized in that, The device includes: The acquisition module is used to acquire the target audio and the corresponding set of facial parameters of the target object, and input the target audio and the set of facial parameters into a preset target model to obtain the corresponding set of facial parameter meshes; The determination module is used to determine multiple patch sampling points corresponding to the face parameter mesh set through the target model, and to perform length scaling processing on each patch sampling point along the normal direction based on the preset semantic label corresponding to each patch sampling point to obtain an initial point cloud set. The alignment module is used to perform image rendering alignment on the initial point cloud set using the target model to obtain the target point cloud set; The fusion module is used to determine eye movement parameters from the face parameter set, generate a first mesh subset based on the eye movement parameters, generate a second mesh subset based on the target audio, and fuse the first mesh subset and the second mesh subset to obtain a deformable mesh set. The update module is used to update the position mapping of the target point cloud set based on the deformable mesh set to obtain the target image frame, and to sequentially stitch together multiple target image frames corresponding to the target audio to obtain the speaking video of the target object driven by the target audio.

10. A computer device, characterized in that, The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the audio-driven object speaking video generation method according to any one of claims 1 to 8.

11. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the audio-driven object speaking video generation method according to any one of claims 1 to 8.