A digital human construction method and device based on dual-flow coupling mechanism and posture adaptation
Patent Information
- Application Number
- CN202610625470.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-08
- Publication Date
- 2026-08-14
AI Technical Summary
[0006]发明目的:针对现有数字人构建技术中音频信息与人物姿态信息关联建模不足、音频与姿态同步性差以及生成动作不自然等问题,本发明提供一种基于双流耦合机制与姿态自适应的数字人构建方法及装置,通过对音频数据与视频姿态数据进行联合建模,构建语音驱动流与姿态先验流,并利用双流耦合机制实现跨模态特征交互,在姿态自适应约束条件下生成与音频内容相匹配的自适应姿态序列,进而通过姿态条件驱动的基础视频生成骨干网络生成数字人视频,用于提升数字人动作的自然性与音频姿态匹配一致性
[0044]1、本发明通过分别构建语音驱动流与姿态先验流,对音频特征与姿态特征进行独立建模,并引入双流耦合机制实现音频表达特征与姿态表达特征之间的双向特征交互,通过耦合注意力分配与动态权重调节,有效增强了音频信息与人物姿态变化之间的关联性,提升了音频驱动姿态生成的精确性,解决了现有技术中音频与姿态匹配度不足、驱动关系弱的问题,使生成的数字人在口型、头部及上半身动作上与语音内容保持更高一致性。
Smart Images

Figure CN122574176A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of computer vision and multimodal intelligent interaction technology, specifically relating to a method and device for constructing a digital human based on a dual-flow coupling mechanism and posture adaptation. Background Technology
[0002] With the rapid development of artificial intelligence and computer vision technologies, digital humans, as a virtual human-computer interaction technology product that integrates voice, vision, and interactive capabilities, have gradually gained widespread application in various fields such as entertainment, education, virtual live streaming, and intelligent customer service. The creation of digital humans typically involves the joint processing of audio information and video information such as human posture and facial expressions. By generating virtual avatars with natural movement and voice synchronization capabilities, a more immersive and realistic human-computer interaction experience can be achieved. However, existing digital human creation technologies still face many technical challenges in terms of natural movement and multimodal consistency.
[0003] Existing digital human creation technologies typically rely on large amounts of labeled audio and video data to train models. This approach not only suffers from high data acquisition and labeling costs but is also susceptible to variations in data quality, leading to inconsistent results across different scenarios. Furthermore, precise temporal alignment between audio information and human posture remains a key challenge in constructing natural-looking digital human avatars. Improper handling can result in asynchrony between voice and movement, or abrupt changes in posture.
[0004] On the other hand, many existing research methods still focus on modeling single-modal information such as audio or posture, lacking in-depth characterization of the correlation between audio and posture multimodal information. Although some methods attempt to jointly model or simply fuse audio and posture features, they mostly use feature concatenation or one-way conditional constraints, failing to fully consider the mutual influence between audio and posture information over time. In the process of generating human motion, traditional methods often rely on post-processing smoothing or rule constraints to suppress posture jitter, which to some extent limits the naturalness and flexibility of human motion, making it difficult to achieve stable and continuous dynamic performance.
[0005] Therefore, how to effectively model the dynamic relationship between audio information and human posture information during the construction of digital humans, and introduce adaptive adjustment and constraint mechanisms in the posture generation stage so that human movements can be naturally adjusted according to changes in speech content, thereby improving the overall effect of digital humans in terms of movement coherence, performance naturalness, and consistency of audio and posture matching, has become a key technical problem that urgently needs to be solved in the current field of digital human research. Summary of the Invention
[0006] Purpose of the Invention: To address the problems of insufficient modeling of the correlation between audio information and human posture information, poor synchronization between audio and posture, and unnatural generated actions in existing digital human construction technologies, this invention provides a digital human construction method and apparatus based on a dual-stream coupling mechanism and posture adaptation. By jointly modeling audio data and video posture data, a speech-driven stream and a posture prior stream are constructed. The dual-stream coupling mechanism is used to realize cross-modal feature interaction. Under posture adaptation constraints, an adaptive posture sequence matching the audio content is generated. Then, a posture-condition-driven basic video generation backbone network is used to generate digital human videos, thereby improving the naturalness of digital human actions and the consistency of audio posture matching.
[0007] Technical Solution: This invention proposes a digital human construction method based on a dual-flow coupling mechanism and posture adaptation, comprising the following steps:
[0008] Step 1: Obtain the input data of the digital human to be constructed, including audio data A1 and video pose data P1. Preprocess the input data to obtain the audio feature set F1 and the initial pose sequence P2.
[0009] Step 2: Construct a speech-driven flow S1 based on the audio feature set F1, construct a pose prior flow S2 based on the initial pose sequence P2, and perform temporal modeling to obtain the audio expression feature set E1 and the pose expression feature set E2.
[0010] Step 3: Use a dual-stream coupling mechanism to perform cross-stream feature interaction between the audio expression feature set E1 and the gesture expression feature set E2, and generate a fused feature set C1 by coupling attention allocation and dynamic weight adjustment;
[0011] Step 4: Dynamically adjust the attitude of the fused feature set C1 based on the attitude adaptation module. The attitude adaptation module generates an adaptive attitude sequence P3 through attitude constraint mapping, attitude parameter prediction and action smoothing optimization.
[0012] Step 5: Input the adaptive pose sequence P3 into the pose conditional video frame generation network to generate the key frame sequence Seq_K of the target digital human. Perform texture consistency compensation and temporal detail enhancement processing on the key frame sequence Seq_K to output the final digital human video result D1.
[0013] Furthermore, the specific method of step 1 is as follows:
[0014] Step 1.1: Obtain the original audio and video materials of the digital human to be constructed, and extract the mono audio data A1; at the same time, extract frames from the video to obtain the original video frame sequence I;
[0015] Step 1.2: Perform feature quantization on audio data A1, segment the audio data into frames using a sliding window mechanism, and extract the Mel-frequency cepstral coefficients (MFCC) and fundamental frequency F0. Concatenate these frames in chronological order to construct a high-dimensional audio feature set. ;
[0016] Step 1.3: Perform pose parameterization analysis on the video frame sequence to extract expression coefficient vectors representing subtle facial movements. Coordinates of key lip points representing the closed state of the mouth Euler angles and translation vectors representing head spatial pose and the coordinates of joints representing the skeletal structure of the upper body. And construct temporally sequenced video pose data P1 in chronological order;
[0017] Step 1.4: Perform temporal cleaning and smoothing on the video pose data P1 to obtain the initial pose sequence P2.
[0018] Furthermore, the front-end feature extraction subnetwork in the basic video generation backbone network includes a speech coding network EncoderA and a pose coding network EncoderP, specifically as follows:
[0019] Step 2.1: Based on the audio feature set F1, the speech coding network EncoderA is used to output the speech driving stream S1. The speech coding network EncoderA includes three one-dimensional convolutional layers connected in sequence to extract the prosodic changes and semantic features of the audio features in the time dimension and output the speech driving stream S1.
[0020] Step 2.2: Based on the initial pose sequence P2, the pose encoding network EncoderP outputs the pose prior flow S2. The pose encoding network EncoderP includes three fully connected layers, which map the geometric topology information of the pose to the action manifold features and output the pose prior flow S2.
[0021] Step 2.3: Based on the speech coding network EncoderA, introduce a time dependency modeling module based on multi-head self-attention mechanism to model the long-term temporal dependency of speech driving flow S1 and output audio expression feature set E1;
[0022] Step 2.4: Model the dynamic changes of the pose prior flow S2 output by the pose coding network EncoderP in the time dimension, and output the pose expression feature set E2; the time step of the pose expression feature set E2 is consistent with that of the audio expression feature set E1.
[0023] Furthermore, the specific method of the dual-flow coupling mechanism in step 3 is as follows:
[0024] Step 3.1: Perform mean-variance standardization on each feature dimension in the audio expression feature set E1 and the posture expression feature set E2 in the time dimension to obtain a cross-modal aligned feature matrix that corresponds one-to-one in the time dimension;
[0025] Step 3.2: Construct a two-stream coupled attention model based on the cross-modal feature matrix. Perform linear mapping on the audio representation feature set E1 to generate the query vector Q1, and perform linear mapping on the posture representation feature set E2 to generate the key vector K1 and value vector V1. Calculate the audio-to-posture coupled attention weight matrix W1, which represents the driving strength of audio features on posture features.
[0026] ;
[0027] Where t represents the current time step, i and j represent the time step indices corresponding to the pose features, d is the feature dimension, and T is the total number of time steps;
[0028] Step 3.3: Based on the cross-modal feature matrix, perform a linear mapping on the posture expression feature set E2 to generate a query vector Q2, and perform a linear mapping on the audio expression feature set E1 to generate a key vector K2 and a value vector V2. Then, calculate the posture-to-audio coupling attention weight matrix W2, which represents the constraint strength of the posture features on the audio features in the time dimension.
[0029] ;
[0030] Step 3.4: Based on W1, perform a weighted summation on the value vector V1 corresponding to the posture expression feature set E2 to obtain the posture weighted feature set E2′ driven by audio features; based on W2, perform a weighted summation on the value vector V2 corresponding to the audio expression feature set E1 to obtain the audio weighted feature set E1′ constrained by posture features.
[0031] Step 3.5: Concatenate the audio weighted feature set E1′ and the pose weighted feature set E2′ in the feature dimension direction to form a joint feature vector, and input the joint feature vector into a fusion mapping network composed of two fully connected layers connected in sequence to obtain the fusion feature set C1.
[0032] Furthermore, the attitude adaptation module includes an attitude constraint mapping network, an attitude parameter prediction network, and an action smoothing optimization unit connected in sequence.
[0033] A pose constraint mapping network is constructed based on the fused feature set C1 to generate the target pose driving vector and region constraint weights. The pose constraint mapping network consists of three fully connected layers connected in series. The first layer is used to perform dimensionality up-mapping of the fused features and connect them to the ReLU activation function. The second layer introduces the pose structure prior weight matrix to apply structural constraints to the feature dimensions related to the key points of the mouth, head and upper body skeleton. The third layer is used to output the target pose driving vector D(t) corresponding to each time step.
[0034] The target attitude driving vector D(t) is input into the attitude parameter prediction network. The attitude parameter prediction network adopts a shared feature layer and a multi-branch prediction head structure. The shared feature layer is used to perform a unified feature transformation on the target attitude driving vector D(t). The multi-branch prediction head includes a mouth parameter prediction head, a head attitude prediction head, and a shoulder and neck displacement prediction head. The predicted attitude parameters of each time step are combined in chronological order to form the original predicted attitude sequence P3′.
[0035] Attitude boundary constraints are applied to the original predicted attitude sequence P3′ to limit the range of values for mouth opening and closing amplitude, head pitch angle, yaw angle and displacement parameters of key joints of the upper body, and the attitude change rate is constrained by setting a preset attitude change constraint threshold for the attitude change amplitude between adjacent time steps.
[0036] For the attitude sequence that satisfies the attitude boundary constraints, motion smoothing optimization is performed through the motion smoothing optimization unit to generate a continuous and natural adaptive attitude sequence P3.
[0037] Furthermore, the specific method of step 5 is as follows:
[0038] Step 5.1: Select a clear reference frame of the target person from the original video frame sequence I, and extract the identity feature encoding of the target person based on the reference frame; input the adaptive pose sequence P3 and the identity feature encoding into the pre-trained U-Net encoder-decoder pose conditional video frame generation network to generate the original image frames of the digital human frame by frame, and combine the original image frames in chronological order to obtain the original keyframe sequence Seq_K′ of the digital human.
[0039] Step 5.2: Perform texture consistency compensation on the pixel features between adjacent frames in the original keyframe sequence Seq_K′ to obtain the texture-corrected keyframe sequence Seq_K″;
[0040] Step 5.3: Use a BasicVSR-type video enhancement network to perform temporal detail enhancement on the texture-corrected keyframe sequence Seq_K″ to obtain enhanced digital human video frames. Combine all enhanced digital human video frames in chronological order to obtain the final digital human video result D1.
[0041] Furthermore, the texture consistency compensation specifically involves: establishing a texture correction mapping relationship for local texture alignment between adjacent frames; extracting pixel features or shallow convolution features of the corresponding local regions from the current frame and its adjacent frames based on the region localization results of the face region and key joint region; obtaining the region difference quantity that characterizes the degree of local texture change by calculating the feature difference between the corresponding regions; and then generating a texture compensation quantity for the corresponding texture region of the current frame and correcting it based on the region difference quantity.
[0042] The present invention also discloses a digital human construction device based on a dual-stream coupling mechanism and posture adaptation, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded onto the processor, it is used to execute the above-mentioned digital human construction method based on a dual-stream coupling mechanism and posture adaptation.
[0043] Beneficial effects:
[0044] 1. This invention constructs a speech-driven flow and a posture prior flow separately, independently models audio features and posture features, and introduces a dual-flow coupling mechanism to realize bidirectional feature interaction between audio expression features and posture expression features. By coupling attention allocation and dynamic weight adjustment, it effectively enhances the correlation between audio information and changes in human posture, improves the accuracy of audio-driven posture generation, and solves the problems of insufficient audio-posture matching and weak driving relationship in the prior art. This enables the generated digital human to maintain a higher consistency with the speech content in terms of lip movements, head and upper body movements.
[0045] 2. This invention performs posture constraint mapping and motion smoothing optimization on the fused features through a posture adaptive module. By introducing structural constraints and temporal continuity constraints during the posture generation process, it effectively suppresses posture jitter and unnatural deformation, improves the continuity and stability of digital human movements in the temporal dimension, and makes the generated human movements more natural and smooth.
[0046] 3. This invention improves the consistency of generated video frames in terms of spatial texture and temporal continuity by using an adaptive pose sequence as input to drive the basic video generation backbone network, combined with texture consistency compensation and temporal detail enhancement processing, thereby further improving the overall quality of digital human videos. This invention provides a method for constructing digital humans with highly coordinated audio and pose, resulting in natural and stable movements, effectively enhancing the realism and practicality of digital human generation effects. Attached Figure Description
[0047] Figure 1 This is the overall flowchart of the present invention;
[0048] Figure 2 This is a diagram of the posture adaptation mechanism model architecture.
[0049] Figure 3 This is a flowchart illustrating the specific implementation of the present invention. Detailed Implementation
[0050] The present invention will be further illustrated below with reference to specific embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention. After reading the present invention, any modifications of the present invention in various equivalent forms by those skilled in the art will fall within the scope defined by the appended claims.
[0051] This invention discloses a method and apparatus for constructing a digital human based on a dual-flow coupling mechanism and posture adaptation. The method for constructing a digital human based on a dual-flow coupling mechanism and posture adaptation includes the following steps:
[0052] Step 1: Obtain the input data for the digital human to be constructed, including audio data A1 and video pose data P1. Preprocess the input data to obtain the audio feature set F1 and the initial pose sequence P2.
[0053] Step 1.1: Select a 252-second video clip of a single person giving a frontal presentation as the original audio and video material. Use the ffmpeg audio and video processing tool to extract mono audio data A1 from the original audio and video material. In this embodiment, the audio sampling rate is set to 16kHz. At the same time, use the OpenCV video processing library to extract frames from the video at 25fps to obtain approximately 6300 original video frame sequences I.
[0054] Step 1.2: Perform feature quantization on the audio data A1. Use a sliding window mechanism to divide the audio data into frames, with a window length of 25ms and a frame shift of 10ms. Utilize the Librosa audio feature extraction library to perform short-time Fourier transform and Mel filter bank mapping on each window, extracting the Mel frequency cepstral coefficients (MFCC) and the fundamental frequency F0. Concatenate these features in chronological order to construct a high-dimensional audio feature set F1. T represents the total number of audio time steps. Let be the audio feature vector at time step t.
[0055] Step 1.3: Perform pose parameterization analysis on the original video frame sequence. Use the OpenPose-based human keypoint detection tool to detect human pose in each frame of the video and extract expression coefficient vectors representing subtle facial movements. Coordinates of key lip points representing the closed state of the mouth Euler angles and translation vectors representing head spatial pose and the coordinates of joints representing the skeletal structure of the upper body. And construct temporally sequenced video pose data P1 in chronological order.
[0056] Step 1.4: Perform temporal cleaning and smoothing on the video pose data P1. By setting the pose detection confidence threshold to 0.6, abnormal jitter frames with a confidence level below 0.6 are removed. The Savitzky-Golay temporal smoothing filter algorithm is then used to smooth the head pose parameters and body skeleton parameters to eliminate abrupt pose changes between adjacent frames, resulting in a temporally aligned and continuously stable initial pose sequence P2. After this step, local jitter in the original pose sequence is suppressed, providing stable input for subsequent pose encoding and adaptive pose generation.
[0057] Step 2: Construct a speech-driven flow S1 based on the audio feature set F1, construct a pose prior flow S2 based on the initial pose sequence P2, and further perform temporal modeling to obtain the audio expression feature set E1 and the pose expression feature set E2.
[0058] Step 2.1: Input the audio feature set F1 into the speech coding network EncoderA to obtain the speech driving flow S1. The speech coding network EncoderA consists of three one-dimensional convolutional layers connected in series. Each convolutional layer has a kernel size of 3, a stride of 1, and padding of 1. The number of convolutional channels is set to 64, 128, and 256 respectively. After each convolutional layer, a batch normalization layer and a ReLU activation function are connected in sequence to encode the prosodic variations and semantic features of the audio features in the temporal dimension.
[0059] Step 2.2: Input the initial pose sequence P2 into the pose encoding network EncoderP to obtain the pose prior flow S2. The pose encoding network EncoderP consists of three fully connected layers. The input and output dimensions of the first fully connected layer are set to [512, 1024], the second fully connected layer is set to [1024, 1024], and the third fully connected layer is set to [1024, 1024]. A ReLU activation function is connected after each fully connected layer to map the geometric topology information of the pose to the action manifold features.
[0060] Step 2.3: By introducing a time-dependent modeling module based on a multi-head self-attention mechanism on the basis of the speech coding network EncoderA, the long-term temporal dependency of the speech driving flow S1 is modeled, and the audio expression feature set E1 is output to represent the changes in pronunciation intensity and potential mouth shape driving information of speech content in the time dimension.
[0061] Step 2.4: By modeling the dynamic changes of the pose prior flow S2 output by the pose coding network EncoderP in the time dimension, the pose expression feature set E2 is output. The time step of the pose expression feature set E2 is consistent with that of the audio expression feature set E1, so that the pose expression feature set E2 and the audio expression feature set E1 correspond one-to-one in the time dimension.
[0062] Step 3: Use a dual-stream coupling mechanism to perform cross-stream feature interaction between the audio expression feature set E1 and the posture expression feature set E2, and generate a fused feature set C1 by coupling attention allocation and dynamic weight adjustment.
[0063] Step 3.1: Perform mean-variance standardization on the time dimension for each feature dimension in the audio expression feature set E1 and the posture expression feature set E2 respectively. That is, the values of the same feature dimension at all time steps are normalized by removing the mean and normalizing the variance, so that the distribution of each feature dimension on the time axis satisfies zero mean and unit variance, and obtain the cross-modal aligned feature matrix that corresponds one-to-one in the time dimension.
[0064] Step 3.2: Construct a two-stream coupled attention model based on the cross-modal aligned feature matrix. Perform linear mapping on the audio representation feature set E1 to generate the query vector Q1, and perform linear mapping on the pose representation feature set E2 to generate the key vector K1 and value vector V1. Calculate the audio-to-pose coupled attention weight matrix W1 according to the following formula, which represents the driving strength of audio features on pose features:
[0065]
[0066] Where d is the feature dimension, T is the total number of time steps, t represents the current time step, and i and j represent the time step indices corresponding to the pose features.
[0067] Step 3.3: Based on the cross-modal alignment feature matrix, perform a linear mapping on the pose representation feature set E2 to generate the query vector Q2, and perform a linear mapping on the audio representation feature set E1 to generate the key vector K2 and the value vector V2. Then, calculate the pose-to-audio coupling attention weight matrix W2 according to the following formula, which is used to represent the constraint strength of the pose features on the audio features in the time dimension:
[0068]
[0069] Step 3.4: Based on the audio-to-pose coupled attention weight matrix W1, perform a weighted summation of the value vector V1 corresponding to the pose representation feature set E2 to obtain the pose weighted feature set E2′ driven by audio features; simultaneously, based on the pose-to-audio coupled attention weight matrix W2, perform a weighted summation of the value vector V2 corresponding to the audio representation feature set E1 to obtain the audio weighted feature set E1′ constrained by pose features.
[0070]
[0071]
[0072] Step 3.5: Concatenate the audio weighted feature set E1′ and the pose weighted feature set E2′ in the feature dimension direction to form a joint feature vector. Input the joint feature vector into a fusion mapping network consisting of two fully connected layers connected in series. The number of neurons in the two fully connected layers is set to 512 and 256 respectively, to obtain the final 256-dimensional fusion feature set C1, which provides joint driving input for the subsequent pose adaptation module.
[0073] Step 4: Dynamically adjust the attitude of the fused feature set C1 based on the attitude adaptation module to generate an adaptive attitude sequence P3. The attitude adaptation module includes an attitude constraint mapping network, an attitude parameter prediction network, and an action smoothing optimization unit.
[0074] Step 4.1: Construct a pose constraint mapping network based on the fused feature set C1 to generate the target pose driving vector and region constraint weights. The pose constraint mapping network consists of three fully connected layers connected in series, with the output dimensions of each layer set to 512, 256, and 128, respectively. The first layer is used to perform dimensionality upscaling mapping on the fused features and connect them to the ReLU activation function. The second layer introduces a prior weight matrix of pose structure to apply structural constraints to the feature dimensions related to the key points of the mouth, head, and upper body skeleton. The third layer is used to output the target pose driving vector D(t) corresponding to each time step.
[0075] Step 4.2: Input the target attitude driving vector D(t) into the attitude parameter prediction network. The attitude parameter prediction network adopts a shared feature layer and a multi-branch prediction head structure. The shared feature layer is used to perform a unified feature transformation on the target attitude driving vector D(t). The multi-branch prediction head includes a mouth parameter prediction head, a head attitude prediction head, and a shoulder and neck displacement prediction head. The predicted attitude parameters of each time step are combined in time order to form the original predicted attitude sequence P3′.
[0076] Step 4.3: Apply attitude boundary constraints to the original predicted attitude sequence P3′, limit the range of values for mouth opening and closing amplitude, head pitch angle, yaw angle and displacement parameters of key joints of the upper body, and constrain the attitude change rate by setting a preset attitude change constraint threshold for the attitude change amplitude between adjacent time steps to prevent unnatural deformation.
[0077] Step 4.4: Perform motion smoothing optimization on the attitude sequence that satisfies the attitude boundary constraints. Minimize the difference terms of the attitude parameters between adjacent time steps by constructing the following objective function:
[0078]
[0079] To reduce abrupt changes in pose between consecutive frames, a continuous and natural adaptive pose sequence P3 is generated.
[0080] Step 5: Input the adaptive pose sequence P3 into the pose condition video frame generation network, generate target human image frames frame by frame according to the input pose conditions, and perform texture consistency compensation and temporal detail enhancement processing on the generated keyframe sequence Seq_K′ to output the final digital human video result D1.
[0081] Step 5.1: Select a clear reference frame of the target person from the original video frame sequence I, and input the adaptive pose sequence P3 and the reference frame into the pre-trained U-Net encoder-decoder generator network (pose-conditional video frame generation network) to generate the original image frames of the digital human frame by frame; combine the original image frames in chronological order to obtain the original keyframe sequence Seq_K′ of the digital human.
[0082] Step 5.2: Perform texture consistency compensation on the pixel features between adjacent frames in the original keyframe sequence Seq_K′. Specifically, construct a texture correction mapping model for local texture alignment between adjacent frames. Based on the region localization results of the face region and key joint region, extract the pixel features or shallow convolutional features of the corresponding local regions from the current frame and its adjacent frames. By calculating the feature differences between corresponding regions, obtain the region difference quantity representing the degree of local texture change. The region difference quantity can be expressed as:
[0083]
[0084] in, This indicates the facial area or key joint area. Indicates the first Mid-frame region Extracted pixel features, Indicates the first Corresponding region in the frame Extracted pixel features, Indicates adjacent frames in the region The feature difference is measured; then, based on the regional difference, a texture compensation amount is generated for the corresponding texture region of the current frame and corrected to obtain the texture-corrected keyframe sequence Seq_K″. After this step, the local texture difference between consecutive frames is reduced, and the texture continuity of the generated video in the facial region and key joint region is improved.
[0085] Step 5.3: Perform temporal detail enhancement processing on the texture-corrected keyframe sequence Seq_K″. Specifically, the temporal detail enhancement processing involves constructing a temporal detail enhancement network, extracting high-frequency texture residual information between consecutive frames, and superimposing this high-frequency texture residual information onto the current frame image to enhance the temporal continuity of details in the mouth, eyes, and facial edge regions, outputting the final digital human video result D1. The temporal detail enhancement network is a post-processing network used for video detail enhancement.
[0086] The present invention also discloses a digital human construction device based on a dual-flow coupling mechanism and posture adaptation, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the computer program is loaded onto the processor, it is used to execute the above-mentioned digital human construction method based on a dual-flow coupling mechanism and posture adaptation.
[0087] The following experiments demonstrate the digital human construction method based on dual-stream coupling mechanism and posture adaptation of this invention. Using the same input video and audio data, we conducted comparative experiments on digital human generation effects, posture continuity, and ablation using Wav2Lip, SyncTalk, and the method of this invention, respectively. The comparison results are shown in Tables 1, 2, and 3.
[0088] Table 1. Comparison of digital human generation effects using different methods
[0089] Wav2Lip 30.883 0.0690 4.306 9.853 3.454 SyncTalk 36.098 0.0311 2.644 7.856 5.751 Method of the present invention 36.063 0.0310 2.633 7.474 5.966
[0090] Table 2 Comparison of pose continuity of different methods
[0091] Wav2Lip 5.021 3.951 4.000 SyncTalk 2.687 2.694 2.533 Method of the present invention 2.734 2.714 2.456
[0092] Table 3 Comparison of ablation experiments using the method of the present invention
[0093] Remove dual-flow coupling mechanism 19.2073 0.271168 3.8496 11.759 1.273 Remove attitude adaptation module 19.2070 0.271208 3.8856 11.762 1.239 Method of the present invention 36.1041 0.030964 2.585 7.918 5.691
[0094] The experimental results above show that the method of the present invention is superior to existing methods in terms of digital human generation quality, audio-visual synchronization consistency, and posture continuity. Meanwhile, ablation experiments show that both the dual-stream coupling mechanism and the posture adaptation module have a significant effect on improving the final digital human generation effect.
[0095] The above embodiments are only for illustrating the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent modifications or alterations made in accordance with the spirit and essence of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A method for constructing a digital human based on a dual-flow coupling mechanism and posture adaptation, characterized in that, Includes the following steps: Step 1: Obtain the input data of the digital human to be constructed, including audio data A1 and video pose data P1. Preprocess the input data to obtain the audio feature set F1 and the initial pose sequence P2. Step 2: Construct a speech-driven flow S1 based on the audio feature set F1, construct a pose prior flow S2 based on the initial pose sequence P2, and perform temporal modeling to obtain the audio expression feature set E1 and the pose expression feature set E2. Step 3: Use a dual-stream coupling mechanism to perform cross-stream feature interaction between the audio expression feature set E1 and the gesture expression feature set E2, and generate a fused feature set C1 by coupling attention allocation and dynamic weight adjustment; Step 4: Dynamically adjust the attitude of the fused feature set C1 based on the attitude adaptation module. The attitude adaptation module generates an adaptive attitude sequence P3 through attitude constraint mapping, attitude parameter prediction and action smoothing optimization. Step 5: Input the adaptive pose sequence P3 into the pose conditional video frame generation network to generate the key frame sequence Seq_K of the target digital human. Perform texture consistency compensation and temporal detail enhancement processing on the key frame sequence Seq_K to output the final digital human video result D1.
2. The digital human construction method based on dual-flow coupling mechanism and posture adaptation according to claim 1, characterized in that, The specific method for step 1 is as follows: Step 1.1: Obtain the original audio and video materials of the digital human to be constructed, and extract the mono audio data A1; at the same time, extract frames from the video to obtain the original video frame sequence I; Step 1.2: Perform feature quantization on audio data A1, segment the audio data into frames using a sliding window mechanism, and extract the Mel-frequency cepstral coefficients (MFCC) and fundamental frequency F0. Concatenate these features in chronological order to construct a high-dimensional audio feature set. ; Step 1.3: Perform pose parameterization analysis on the video frame sequence to extract expression coefficient vectors representing subtle facial movements. Coordinates of key lip points representing the closed state of the mouth Euler angles and translation vectors representing head spatial pose and the coordinates of joints representing the skeletal structure of the upper body. And construct temporally sequenced video pose data P1 in chronological order; Step 1.4: Perform temporal cleaning and smoothing on the video pose data P1 to obtain the initial pose sequence P2.
3. The digital human construction method based on dual-flow coupling mechanism and posture adaptation according to claim 1, characterized in that, The front-end feature extraction subnetwork in the basic video generation backbone network includes a speech coding network EncoderA and a pose coding network EncoderP. The specific method is as follows: Step 2.1: Based on the audio feature set F1, the speech coding network EncoderA is used to output the speech driving stream S1. The speech coding network EncoderA includes three one-dimensional convolutional layers connected in sequence to extract the prosodic changes and semantic features of the audio features in the time dimension and output the speech driving stream S1. Step 2.2: Based on the initial pose sequence P2, the pose encoding network EncoderP outputs the pose prior flow S2. The pose encoding network EncoderP includes three fully connected layers, which map the geometric topology information of the pose to the action manifold features and output the pose prior flow S2. Step 2.3: Based on the speech coding network EncoderA, introduce a time dependency modeling module based on multi-head self-attention mechanism to model the long-term temporal dependency of speech driving flow S1 and output audio expression feature set E1; Step 2.4: Model the dynamic changes of the pose prior flow S2 output by the pose coding network EncoderP in the time dimension, and output the pose expression feature set E2; the time step of the pose expression feature set E2 is consistent with that of the audio expression feature set E1.
4. The digital human construction method based on dual-flow coupling mechanism and posture adaptation according to claim 1, characterized in that, The specific method of the dual-flow coupling mechanism in step 3 is as follows: Step 3.1: Perform mean-variance standardization on each feature dimension in the audio expression feature set E1 and the posture expression feature set E2 in the time dimension to obtain a cross-modal aligned feature matrix that corresponds one-to-one in the time dimension; Step 3.2: Construct a two-stream coupled attention model based on the cross-modal feature matrix. Perform linear mapping on the audio representation feature set E1 to generate the query vector Q1, and perform linear mapping on the posture representation feature set E2 to generate the key vector K1 and value vector V1. Calculate the audio-to-posture coupled attention weight matrix W1, which represents the driving strength of audio features on posture features. ; Where t represents the current time step, i and j represent the time step indices corresponding to the pose features, d is the feature dimension, and T is the total number of time steps; Step 3.3: Based on the cross-modal feature matrix, perform a linear mapping on the posture expression feature set E2 to generate a query vector Q2, and perform a linear mapping on the audio expression feature set E1 to generate a key vector K2 and a value vector V2. Then, calculate the posture-to-audio coupling attention weight matrix W2, which represents the constraint strength of the posture features on the audio features in the time dimension. ; Step 3.4: Based on W1, perform a weighted summation on the value vector V1 corresponding to the posture expression feature set E2 to obtain the posture weighted feature set E2′ driven by audio features; based on W2, perform a weighted summation on the value vector V2 corresponding to the audio expression feature set E1 to obtain the audio weighted feature set E1′ constrained by posture features. Step 3.5: Concatenate the audio weighted feature set E1′ and the pose weighted feature set E2′ in the feature dimension direction to form a joint feature vector, and input the joint feature vector into a fusion mapping network composed of two fully connected layers connected in sequence to obtain the fusion feature set C1.
5. The digital human construction method based on dual-flow coupling mechanism and posture adaptation according to claim 1, characterized in that, The attitude adaptation module includes an attitude constraint mapping network, an attitude parameter prediction network, and a motion smoothing optimization unit connected in sequence. A pose constraint mapping network is constructed based on the fused feature set C1 to generate the target pose driving vector and region constraint weights. The pose constraint mapping network consists of three fully connected layers connected in series. The first layer is used to perform dimensionality up-mapping of the fused features and connect them to the ReLU activation function. The second layer introduces the pose structure prior weight matrix to apply structural constraints to the feature dimensions related to the key points of the mouth, head and upper body skeleton. The third layer is used to output the target pose driving vector D(t) corresponding to each time step. The target attitude driving vector D(t) is input into the attitude parameter prediction network. The attitude parameter prediction network adopts a shared feature layer and a multi-branch prediction head structure. The shared feature layer is used to perform a unified feature transformation on the target attitude driving vector D(t). The multi-branch prediction head includes a mouth parameter prediction head, a head attitude prediction head, and a shoulder and neck displacement prediction head. The predicted attitude parameters of each time step are combined in chronological order to form the original predicted attitude sequence P3′. Attitude boundary constraints are applied to the original predicted attitude sequence P3′ to limit the range of values for mouth opening and closing amplitude, head pitch angle, yaw angle and displacement parameters of key joints of the upper body, and the attitude change rate is constrained by setting a preset attitude change constraint threshold for the attitude change amplitude between adjacent time steps. For the attitude sequence that satisfies the attitude boundary constraints, motion smoothing optimization is performed through the motion smoothing optimization unit to generate a continuous and natural adaptive attitude sequence P3.
6. The digital human construction method based on dual-flow coupling mechanism and posture adaptation according to claim 1, characterized in that, The specific method for step 5 is as follows: Step 5.1: Select a clear reference frame of the target person from the original video frame sequence I, and extract the identity feature encoding of the target person based on the reference frame; input the adaptive pose sequence P3 and the identity feature encoding into the pre-trained U-Net encoder-decoder pose conditional video frame generation network to generate the original image frames of the digital human frame by frame, and combine the original image frames in chronological order to obtain the original keyframe sequence Seq_K′ of the digital human. Step 5.2: Perform texture consistency compensation on the pixel features between adjacent frames in the original keyframe sequence Seq_K′ to obtain the texture-corrected keyframe sequence Seq_K″; Step 5.3: Use a BasicVSR-type video enhancement network to perform temporal detail enhancement on the texture-corrected keyframe sequence Seq_K″ to obtain enhanced digital human video frames. Combine all enhanced digital human video frames in chronological order to obtain the final digital human video result D1.
7. The digital human construction method based on dual-flow coupling mechanism and posture adaptation according to claim 6, characterized in that, The texture consistency compensation specifically involves: establishing a texture correction mapping relationship for local texture alignment between adjacent frames; extracting pixel features or shallow convolution features of the corresponding local regions from the current frame and its adjacent frames based on the region localization results of the face region and key joint region; obtaining the region difference quantity that characterizes the degree of local texture change by calculating the feature difference between the corresponding regions; and then generating a texture compensation quantity for the corresponding texture region of the current frame and correcting it according to the region difference quantity.
8. A digital human construction device based on a dual-stream coupling mechanism and posture adaptation, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed, it causes the processor to implement the steps in the digital human construction method based on dual-stream coupling mechanism and posture adaptation as described in any one of claims 1 to 7.