Digital human facial expression synchronization method and device and medium

Through the implementation of the synchronization method of digital human facial expressions, the problem of low authenticity and humanization of digital human simulations is solved, and the synchronization conversion of audio and video features and the naturalness of digital human facial expressions is realized, which significantly enhances the expressiveness and immersion of digital humans.

CN119991887APending Publication Date: 2025-05-13山东浪潮智慧建筑科技有限公司
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510057372.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In the prior art, the analog authenticity and humanization of digital people are low, especially in the audio encoder and audio-driven digital people's lips movement links. The feature mapping relationship is not effectively aligned, resulting in the failure of synchronizing audio and video feature spaces, which affects the overall performance and coordination of digital people.

Method used

Through a digital human facial expression synchronization method, it includes dividing and preprocessing the target video, extracting audio features and video features, calculating cosine similarity, and combining binary cross entropy loss function, the conversion of audio and video features in the same parameter space is realized. Then, based on three-dimensional morphological model modeling and multi-resolution hash encoding, the density and color of the body are constructed and rendered to generate the audio-driven digital human facial expression image.

Benefits of technology

It significantly enhances the synchronization between audio and video content, ensures the precise matching of lips movement and audio content, improves the realism and nature of the facial expressions of digital people, and enhances the expressiveness and immersion of digital people in interactive scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991887A_ABST
    Figure CN119991887A_ABST
Patent Text Reader

Abstract

The invention discloses a digital human facial expression synchronization method and device and a medium. A target video is divided to obtain a corresponding video set, and the video set is preprocessed to obtain a standard video set; extracting audio features and video features based on an audio and lip motion synchronization discriminator to calculate cosine similarity, and realizing conversion of the audio features and the video features in the same parameter space by combining a binary cross entropy loss function; performing three-dimensional shape model modeling on the face according to the standard video set, obtaining a three-dimensional space rotation translation matrix, a camera observation direction and an expression coefficient of each frame of face image and a shape coefficient of a face three-dimensional model, and obtaining an average face shape and key point coordinates; and according to the key point coordinates and the camera observation direction, obtaining spatial geometric features of multi-resolution Hash coding, combining audio features to obtain conditional features, and based on the conditional features and high-frequency coding information, constructing and rendering volume density and color to generate an audio-driven digital human facial expression image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device and medium for synchronizing digital human facial expressions. Background Art

[0002] At present, digital human technology has been widely used in many industries. In the field of customer service, some companies have begun to deploy digital humans to replace traditional manual customer service roles and achieve interactive communication with customers; in the field of virtual assistants, many smartphones and smart speakers have integrated digital human assistant functions, which can assist users in completing various tasks; as for the field of education and training, some educational institutions have begun to adopt digital humans as teaching tools to provide students with exclusive customized teaching experiences.

[0003] However, the simulation realism and humanization of digital humans in the existing technology still need to be improved. In particular, there are problems with the audio encoder and the audio-driven digital human lip movement, such as the failure to effectively align the two feature mapping relationships, and the failure to synchronize the audio and video feature spaces, which affects the overall performance and coordination of the digital human. Summary of the invention

[0004] The embodiments of the present application provide a digital human facial expression synchronization method, device and medium to solve the above technical problems.

[0005] On the one hand, an embodiment of the present application provides a method for synchronizing digital human facial expressions, comprising:

[0006] Dividing the target video collected under the specified environment to obtain a corresponding video set, and preprocessing the video set to obtain a standard video set;

[0007] Based on a pre-built audio and lip motion synchronization discriminator, extract audio features and video features to calculate the cosine similarity between the audio features and the video features, and combine with a binary cross entropy loss function to achieve conversion of the audio features and the video features in the same parameter space;

[0008] Modeling a three-dimensional morphological model of the face according to the standard video set, obtaining the three-dimensional space rotation and translation matrix, camera observation direction, expression coefficient and shape coefficient of the three-dimensional face model of each frame of the face image, and obtaining the average face shape and key point coordinates;

[0009] According to the key point coordinates and the camera observation direction, the spatial geometric features of multi-resolution hash coding are obtained, and the conditional features are obtained in combination with the audio features. Based on the conditional features and high-frequency coding information, the volume density and color are constructed and rendered to generate a digital human facial expression image driven by audio.

[0010] In one implementation of the present application, the target video collected in a specified environment is divided to obtain a corresponding video set, and the video set is preprocessed to obtain a standard video set, which specifically includes:

[0011] Under a specified environment, a target video is acquired through a preset acquisition device; wherein the preset acquisition device includes a high-definition camera and a noise reduction microphone, and the target video includes a speaking video and a closed-mouth video;

[0012] Performing endpoint detection on the target video based on automatic speech endpoint detection technology to identify all spoken video clips, and cutting the target video according to the detected time points;

[0013] The cut videos are classified to obtain a speaking video set and a closed mouth video set, and for each video set, background noise separation and noise reduction processing are performed on the video set to form a cleaned target video set.

[0014] In one implementation of the present application, audio features and video features are extracted based on a pre-built audio and lip movement synchronization discriminator, specifically including:

[0015] Performing shot detection on the input video in the standard video set to identify scene switching in the video, performing face detection on the video in the standard video set to identify face areas in the video, and performing face tracking on the video in the standard video set to distinguish different speakers in a multi-person conversation scene;

[0016] Receiving a first preset number of target images, splicing the first preset number of target images in a channel dimension, retaining time series information of the target images, and performing feature encoding on the spliced ​​target images to generate a video face feature vector of a target number of dimensions;

[0017] Audio features are extracted from the input video of the standard video set through Mel-cepstral coefficients, and feature encoding is performed on a second preset number of Mel-cepstral coefficient audio features to generate an audio feature vector of a target number of dimensions.

[0018] In one implementation of the present application, the cosine similarity between the audio feature and the video feature is calculated, and combined with a binary cross entropy loss function, the conversion of the audio feature and the video feature in the same parameter space is realized, specifically including:

[0019] Calculate the cosine similarity between the video face feature vector and the audio feature vector in the target number of dimensions;

[0020] The cosine similarity calculation formula of audio and video features is as follows:

[0021]

[0022] Among them, Sim(FaceF 5 , AudF 20 ) represents the cosine similarity of audio and video features, FaceF 5 Represents the video face feature vector, AudF 20 represents the audio feature vector, FaceF 5 .AudF 20 represents the dot product of two eigenvectors, ||FaceF 5 ||2||AudF 20 ||2 represents the product of the L2 norm of two vectors;

[0023] Based on the cosine similarity and in combination with a binary cross entropy loss function, a binary cross entropy loss between the video face feature vector and the audio feature vector is calculated;

[0024] The binary cross entropy loss calculation formula for audio and video features is as follows:

[0025] Loss av-sync =-(y log Sim(FaceF 5 , AudF 20 )+(1-y)log(1-Sim(FaceF 5 , AudF 20 )))

[0026] Among them, Loss av-sync represents the binary cross entropy loss of audio and video features, y represents the real distribution of audio and video synchronization, 1-y represents the real distribution of audio and video asynchrony, Sim(FaceF 5 , AudF 20 ) represents the predicted probability distribution of audio and video synchronization, 1-Sim(FaceF 5 , AudF 20 ) represents the predicted probability distribution of audio and video asynchrony.

[0027] In one implementation of the present application, a three-dimensional morphological model of a human face is modeled according to the standard video set, and the three-dimensional space rotation and translation matrix, camera observation direction, expression coefficient and shape coefficient of the three-dimensional face model of each frame of the face image are obtained, and the average face shape and key point coordinates are obtained, specifically including:

[0028] Modeling a three-dimensional morphological model of the face according to the face image, the two-dimensional key point coordinates and the open source three-dimensional coordinates of the face in the standard data set, and determining the three-dimensional key points corresponding to the two-dimensional key points to associate the two-dimensional key points with the corresponding three-dimensional key points;

[0029] For each frame of face image, construct the corresponding three-dimensional space rotation and translation matrix and expression coefficient vector, and construct the shape coefficient vector for all face images;

[0030] According to the shape coefficients corresponding to the shape coefficient vectors, weighted summation is performed on all three-dimensional face shapes to obtain an average face shape and three-dimensional key point coordinates.

[0031] In one implementation of the present application, after obtaining the average face shape and key point coordinates, the method further includes:

[0032] In a preset focal length range, a focal length is recursively selected at a fixed interval, and at each selected focal length, all three-dimensional facial expressions are weighted summed according to the expression coefficient corresponding to the expression coefficient vector to obtain an average facial expression;

[0033] According to the average facial expression and the three-dimensional spatial rotation and translation matrix, spatially deforming the three-dimensional key point coordinates to obtain new three-dimensional key point coordinates;

[0034] Performing a two-dimensional projection according to the new three-dimensional key point coordinates and the selected focal length to obtain new two-dimensional key point coordinates, and constructing a loss constraint according to the new two-dimensional key point coordinates and the three-dimensional key point coordinates;

[0035] The losses corresponding to each selected focal length are compared to take the selected focal length corresponding to the minimum loss among all losses as the optimal focal length, and based on the optimal focal length, the three-dimensional space rotation and translation matrix, expression coefficient and shape coefficient are optimized until the model converges to obtain the optimal three-dimensional space rotation and translation matrix and camera observation direction as well as expression coefficient and shape coefficient corresponding to each frame of the face image.

[0036] In one implementation of the present application, obtaining the spatial geometric features of the multi-resolution hash code according to the key point coordinates and the camera observation direction specifically includes:

[0037] Through the three-plane decoupling method of three-dimensional space, the three-dimensional space is reduced to two-dimensional space;

[0038] Through the multi-resolution hash coding voxel grid method, the coordinates of several key points of the three-dimensional face model corresponding to each frame of the face image and the camera observation direction are frequency encoded to obtain the spatial geometric features of the multi-resolution hash coding.

[0039] In one implementation of the present application, the conditional features are obtained in combination with the audio features, and based on the conditional features and high-frequency coding information, the body density and color are constructed and rendered to generate a digital human facial expression image driven by audio, specifically including:

[0040] The spatial geometric features are input into a first multi-layer perceptron to construct a weighted mapping from audio features to blink features, and the audio features and blink features are weighted in channel dimension according to the output weighted values ​​to obtain corresponding conditional features; the conditional features and high-frequency coding information are input into a second multi-layer perceptron to construct a mapping of volume density and color, and an image mapping of the volume density and the color to a two-dimensional space is constructed through neural rendering to generate an audio-driven digital human facial expression image.

[0041] On the other hand, the embodiment of the present application further provides a digital human facial expression synchronization device, the device comprising:

[0042] at least one processor;

[0043] and, a memory communicatively coupled to the at least one processor;

[0044] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the digital human facial expression synchronization method as described above.

[0045] On the other hand, an embodiment of the present application further provides a non-volatile computer storage medium storing computer executable instructions, which, when executed, implements a digital human facial expression synchronization method as described above.

[0046] The present application provides a method, device and medium for synchronizing digital human facial expressions, which at least have the following beneficial effects:

[0047] By dividing and preprocessing the target video collected in a specified environment, a standard video set is obtained, which effectively improves the processing efficiency of video data and ensures the accuracy and consistency of subsequent analysis, helps reduce noise interference, and improves the stability of subsequent feature extraction and model building; through the audio and lip movement synchronization discriminator, audio features and video features can be accurately extracted, and by calculating cosine similarity and combining binary cross entropy loss function, efficient conversion of audio and video features in the same parameter space is achieved, which significantly enhances the synchronization of audio and video content, especially in the digital human generation scene, ensuring the accurate matching of lip movement and audio content, and improving the realism of visual and auditory experience; based on the standard video set, the three-dimensional morphological model of the face is modeled, which can obtain detailed three-dimensional information of each frame of the face image, including spatial rotation and translation moments. The array, camera observation direction, expression coefficient and shape coefficient of the three-dimensional face model not only provide a basis for the creation of personalized digital humans, but also make the facial expressions of digital humans more delicate and realistic, and can more accurately reflect the emotional changes in the original video; the spatial geometric features of multi-resolution hash coding are obtained according to the key point coordinates and camera observation direction, which not only improves the efficiency of feature extraction, but also retains key spatial information, which helps to maintain the accuracy and consistency of facial features in the subsequent digital human generation process; combining audio features and conditional features, as well as high-frequency coding information, constructing and rendering volume density and color, and finally generating audio-driven digital human facial expression images, which not only realizes real-time interaction between audio and facial expressions, but also makes the facial expressions of digital humans more vivid and natural, enhancing the expressiveness and immersion of digital humans in interactive scenes. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0049] Figure 1 A schematic diagram of a flow chart of a digital human facial expression synchronization method provided in an embodiment of the present application;

[0050] Figure 2 A schematic diagram of the internal structure of a digital human facial expression synchronization device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.

[0052] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.

[0053] Figure 1 A flowchart of a digital human facial expression synchronization method provided in an embodiment of the present application.

[0054] The analysis method involved in the embodiments of the present application can be implemented by a terminal device or a server, and the present application does not impose any special restrictions on this. For the convenience of understanding and description, the following embodiments are described in detail by taking a server as an example.

[0055] It should be noted that the server may be a single device or a system consisting of multiple devices, that is, a distributed server, and this application does not make any specific limitation on this.

[0056] like Figure 1 As shown, a digital human facial expression synchronization method provided in an embodiment of the present application includes:

[0057] 101. Divide the target video collected in the specified environment to obtain a corresponding video set, and pre-process the video set to obtain a standard video set.

[0058] The audio encoder uses a pre-trained speech recognition model to extract audio features, achieving feature mapping from audio to text. However, audio-driven digital human lip movement involves feature mapping from audio to movement. This design raises two main problems:

[0059] First, the two different feature mapping relationships, audio to text and audio to motion, are not effectively aligned in the parameter encoding space. This means that although both are derived from the same audio input, they have differences in feature representation and processing, making it difficult to unify and coordinate.

[0060] Second, the latent feature spaces of the two different modalities, audio and video, also failed to achieve alignment, which suggests that although audio and video may be related in content, they are not effectively integrated and synchronized at the feature level, affecting the overall performance and coordination.

[0061] This application constructs an audio and lip movement synchronization discriminator called ALmD, which can evaluate the similarity between audio signals and lip shapes in a common parameter space and realize the conversion of audio and video features in the same parameter space. ALmD's lip shape judgment is a time-dependent dynamic process, which compares the changes in the speaker's voice and his lip movements within a certain time range. In order to deal with this time series problem, ALmD incorporates the features of the time dimension in the feature extraction stage and merges the frames of the time series into the channel dimension.

[0062] Specifically, in one embodiment of the present application, the target video collected in a specified environment is divided to obtain a corresponding video set, and the video set is preprocessed to obtain a standard video set, which specifically includes:

[0063] Under a specified environment, a target video is acquired through a preset acquisition device; wherein the preset acquisition device includes a high-definition camera and a noise reduction microphone, and the target video includes a speaking video and a closed-mouth video;

[0064] Perform endpoint detection on the target video based on automatic speech endpoint detection technology to identify all spoken video clips and cut the target video according to the detected time points;

[0065] The cut videos are classified to obtain a speaking video set and a closed mouth video set, and for each video set, background noise separation and noise reduction are performed on the video set to form a cleaned target video set.

[0066] In one embodiment, to ensure the quality of data collection, the collection should be carried out in a closed environment with soft light, sufficient natural light and no background noise; a green screen should be used as the background; a high-definition 4K resolution camera should be used for shooting; the microphone must have a noise reduction function; the face of the person being collected should be clearly visible and unobstructed, and the imaging of the face, mouth and teeth must reach a pixel quality of 1080P or above; the yaw angle of the head must be controlled within 30°; the length of the frontal face video should account for at least 70% of the total video length; avoid wearing complex or green accessories; hair should be neat, and it is recommended to use styling products to reduce streaking; clothing should be kept clean and tidy, and avoid wearing green or green patterned, striped and rough-edged clothing.

[0067] In one embodiment, 30 employees are selected from within the company to participate in data collection to ensure the diversity of the sample. All participants must meet the above collection specifications. Each employee is required to provide a total of 23 minutes of video material, including 20 minutes of speaking video and 3 minutes of closed-mouth video. The 20-minute speaking video is further divided into two segments: one 17 minutes and one 3 minutes. The 3-minute speaking video and the 3-minute closed-mouth video must be shot continuously without interruption. In the 17-minute speaking video, it should be ensured that the speaking time is interrupted after 3 to 9 seconds, and then start speaking again after an interval of 3 to 5 seconds to facilitate subsequent video editing.

[0068] The collected videos will be carefully divided into two different datasets. The first dataset, named GeneDatasets, contains 30 videos, each of which is 17 minutes long and contains only speaking segments. The second dataset, named CusDatasets, contains 30 videos, each of which is 6 minutes long and contains both speaking and closed-mouth segments.

[0069] Then, we used the automatic voice endpoint detection (AVD) technology to perform endpoint detection on all videos in GeneDatasets to identify all the speaking video clips. According to the detected time points, we used the open source tool FFmpeg to accurately cut the original video. Subsequently, we used the open source tool UVR5 and the open source algorithm FSMN to separate the background noise and reduce the noise of the cut videos, and finally formed a cleaned dataset named PGeneDatasets.

[0070] The open source tool UVR5 and the open source algorithm FSMN are used to separate the background noise and reduce the noise of CusDatasets. After processing, the open source face detection model S3FD and the face segmentation model FaceParse are used to extract and segment the face area in the video. Then, the open source tool OpenFace is used to extract 68 key points and action units (AU) from the face in the video. After completing this series of processing steps, the preprocessed dataset will be obtained, named PCusDatasets.

[0071] 102. Based on the pre-built audio and lip motion synchronization discriminator, audio features and video features are extracted to calculate the cosine similarity between audio features and video features, and combined with the binary cross entropy loss function, the conversion of audio features and video features in the same parameter space is realized.

[0072] Specifically, in one embodiment of the present application, based on a pre-built audio and lip movement synchronization discriminator, audio features and video features are extracted, specifically including:

[0073] Perform shot detection on input videos in the standard video set to identify scene changes in the video, perform face detection on videos in the standard video set to identify face areas in the video, and perform face tracking on videos in the standard video set to distinguish different speakers in multi-person conversation scenarios;

[0074] Receiving a first preset number of target images, splicing the first preset number of target images in a channel dimension, retaining time series information of the target images, and performing feature encoding on the spliced ​​target images to generate a video face feature vector of a target number of dimensions;

[0075] Audio features are extracted from the input video of the standard video set through Mel-cepstral coefficients, and feature encoding is performed on a second preset number of Mel-cepstral coefficient audio features to generate an audio feature vector of a target number of dimensions.

[0076] In one embodiment, the input video is processed frame by frame, including video lens detection, face detection and face tracking. Through these operations, the lower half of the RGB image of the face in the video can be accurately captured. These five consecutive RGB images will be sent to the video encoding module in chronological order. The video lens detection function is used to monitor the scene switching in the video in real time, the face detection is used to identify the face area in the video, and the face tracking is used to distinguish different speakers in a multi-person conversation scene.

[0077] 5 frames of RGB images of the lower half of the face are continuously received and spliced ​​in the channel dimension to retain the time series information. These images are then sent to the video encoding module for feature encoding, and finally a 256-dimensional video face feature vector FaceF is generated. 5 The video encoding module consists of 7 serially connected sub-modules, including convolutional pooling layer, convolutional layer, fully convolutional layer and fully connected layer.

[0078] Audio information is extracted from the input video, and audio features are extracted using Mel-frequency cepstral coefficients (MFCC). The audio feature extraction process includes multiple steps such as pre-emphasis, framing, windowing, fast Fourier transform, Mel filter bank filtering, logarithmic operation, discrete cosine transform, and dynamic feature extraction, and finally generates 13-dimensional Mel-frequency cepstral coefficient audio features, which will be used as the input of the audio encoding module.

[0079] Responsible for obtaining the audio features of Mel cepstral coefficients of 20 consecutive frames and inputting them into the audio encoding module for feature encoding. Similar to the video encoding module, the audio encoding module finally outputs a 256-dimensional audio feature vector AudF 20 The audio encoding module also consists of 7 serially connected sub-modules, including convolutional layer, convolutional pooling layer, full convolutional layer and fully connected layer.

[0080] In one embodiment, the construction of the three submodules of video lens detection, face detection and face tracking adopts two methods: one is based on the existing open source model, and the other is self-coding using Python programming language combined with OpenCV, an open source framework.

[0081] The 7-layer network of the video encoding module includes the following layers in sequence: standard convolution and pooling layer, standard convolution and pooling layer, standard convolution layer, standard convolution layer, standard convolution and pooling layer, full convolution layer, and fully connected layer. The 7-layer network of the audio encoding module includes the following layers in sequence: standard convolution layer, standard convolution and pooling layer, standard convolution layer, standard convolution layer, standard convolution and pooling layer, full convolution layer, and fully connected layer.

[0082] In one embodiment of the present application, the cosine similarity between the audio features and the video features is calculated, and combined with the binary cross entropy loss function, the conversion of the audio features and the video features in the same parameter space is realized, specifically including:

[0083] Calculate the cosine similarity between the video face feature vector and the audio feature vector in the target number of dimensions;

[0084] The cosine similarity calculation formula of audio and video features is as follows:

[0085]

[0086] Among them, Sim(FaceF 5 , AudF 20 ) represents the cosine similarity of audio and video features, FaceF 5 Represents the video face feature vector, AudF 20 represents the audio feature vector, FaceF 5 .AudF 20 represents the dot product of two eigenvectors, ||FaceF 5 ||2||AudF 20 ||2 represents the product of the L2 norm of two vectors;

[0087] Based on cosine similarity and combined with the binary cross entropy loss function, the binary cross entropy loss between the video face feature vector and the audio feature vector is calculated;

[0088] The binary cross entropy loss calculation formula for audio and video features is as follows:

[0089] Loss av-sync =-(y log Sim(FaceF 5 , AudF 20 )+(1-y)log(1-Sim(FaceF 5, AudF 20 )))

[0090] Among them, Loss av-sync represents the binary cross entropy loss of audio and video features, y represents the real distribution of audio and video synchronization, 1-y represents the real distribution of audio and video asynchrony, Sim(FaceF 5 , AudF 20 ) represents the predicted probability distribution of audio and video synchronization, 1-Sim(FaceF 5 , AudF 20 ) represents the predicted probability distribution of audio and video asynchrony.

[0091] In one embodiment, the cosine similarity binary cross entropy loss is used to calculate the cosine value of the 256-dimensional video face and audio feature vectors. If the audio and video features overlap -- y = 1 -- it indicates that the audio and video are synchronized, and if they do not overlap -- y = 0 -- it indicates that the audio and video are not synchronized. The formula for calculating the cosine similarity of audio and video features is:

[0092]

[0093] Using binary cross entropy loss, the distance between synchronized samples is reduced while the interval between asynchronous samples is expanded. The binary cross entropy loss calculation formula for audio and video features is:

[0094] Loss av-sync =-(y log Sim(FaceF 5 , AudF 20 )+(1-y)log(1-Sim(FaceF 5 , AudF 20 ))).

[0095] In one embodiment, the ALmD model is pre-trained using the PGeneDatasets dataset. 40% of the audio and video data is extracted from PGeneDatasets, and the audio and video offsets are created using the open source tool FFmpeg to generate negative samples. The remaining 60% of the data is used as positive samples and merged with the 40% of negative samples to construct a dataset D for training the ALmD model. Then, the D dataset is divided into a training set, a validation set, and a test set in a ratio of 70%, 20%, and 10% to facilitate model training and evaluation. This strategy aims to optimize the generalization ability and accuracy of the model by constructing and balancing positive and negative samples.

[0096] The feedforward inference process of training, validation and test data is consistent. It includes: using the open source FFmpeg tool to split the audio and video to obtain audio data and video data. Perform video lens detection, face detection and face tracking on the video data. Through these operations, the RGB image of the lower half of the face in the video can be accurately captured. These 5 consecutive RGB images will be sent to the video encoding module in chronological order.

[0097] The audio data is processed in sequence by pre-emphasis, framing, windowing, fast Fourier transform, Mel filter bank filtering, logarithmic operation, discrete cosine transform and dynamic feature extraction to finally generate a 13-dimensional Mel cepstral coefficient audio feature.

[0098] The RGB images of the lower half of the face of 5 consecutive frames are spliced ​​in the channel dimension to retain the time series information. These images are then sent to the video encoding module for feature encoding, and finally a 256-dimensional video face feature vector FaceF is generated. 5 The video encoding module consists of 7 serially connected submodules, including standard convolution and pooling layer, standard convolution and pooling layer, standard convolution layer, standard convolution layer, standard convolution and pooling layer, full convolution layer and fully connected layer.

[0099] Get the audio features of Mel cepstral coefficients of 20 consecutive frames and input them into the audio encoding module for feature encoding. Similar to the video encoding module, the audio encoding module finally outputs a 256-dimensional audio feature vector AudF 20 The audio coding module is also composed of 7 serially connected sub-modules, including standard convolution layer, standard convolution and pooling layer, standard convolution layer, standard convolution layer, standard convolution and pooling layer, full convolution layer and fully connected layer.

[0100] Calculate the cosine value of the 256-dimensional video face and audio feature vectors. If the audio and video features overlap, it means the audio and video are synchronized. If they do not overlap, it means the audio and video are not synchronized. Substitute the cosine value into the binary cross entropy loss, construct the loss constraint, and use the Adam optimizer to perform gradient backpropagation to update the parameters until the model converges.

[0101] For the binary classification results of synchronization / asynchronous, human evaluators were used to verify the performance of the model on the test set. The audio encoding module of the trained ALmD model was used as the audio encoder of ER-NeRF.

[0102] 103. The three-dimensional morphological model of the face is built according to the standard video set, and the three-dimensional space rotation and translation matrix, camera observation direction, expression coefficient and shape coefficient of the three-dimensional face model of each frame of the face image are obtained, and the average face shape and key point coordinates are obtained.

[0103] Specifically, in one embodiment of the present application, a 3D morphological model of a face is modeled according to a standard video set, a 3D space rotation and translation matrix, a camera observation direction, an expression coefficient, and a shape coefficient of a 3D face model of each frame of a face image are obtained, and an average face shape and key point coordinates are obtained, specifically including:

[0104] According to the face images, two-dimensional key point coordinates and open source three-dimensional coordinates of the face in the standard data set, a three-dimensional morphological model of the face is built, and the three-dimensional key points corresponding to the two-dimensional key points are determined to associate the two-dimensional key points with the corresponding three-dimensional key points;

[0105] For each frame of face image, construct the corresponding three-dimensional space rotation and translation matrix and expression coefficient vector, and construct the shape coefficient vector for all face images;

[0106] According to the shape coefficients corresponding to the shape coefficient vector, all three-dimensional face shapes are weighted summed to obtain the average face shape and three-dimensional key point coordinates.

[0107] In one embodiment, 3DMM modeling is performed using face images in the PCusDatasets dataset and their corresponding 2D 68 key point coordinates and face 3D coordinates in the open source BFM2009 dataset, and 3D key points corresponding to 2D key points are selected to establish a one-to-one correspondence.

[0108] For each frame of the face image, a 3D spatial rotation and translation matrix and an expression coefficient vector are constructed, and a shape coefficient vector is constructed for all face images. The shape coefficient is used to perform a weighted summation of all 3D face shapes in BFM2009 to obtain the average face shape and the three-dimensional geometric coordinates of 68 key points.

[0109] In one embodiment of the present application, after obtaining the average face shape and key point coordinates, the method further includes:

[0110] In a preset focal length range, a focal length is recursively selected at a fixed interval, and at each selected focal length, all three-dimensional facial expressions are weighted summed according to the expression coefficient corresponding to the expression coefficient vector to obtain an average facial expression;

[0111] According to the average facial expression and the three-dimensional space rotation and translation matrix, the three-dimensional key point coordinates are spatially deformed to obtain new three-dimensional key point coordinates;

[0112] Perform a two-dimensional projection according to the new three-dimensional key point coordinates and the selected focal length to obtain new two-dimensional key point coordinates, and construct a loss constraint according to the new two-dimensional key point coordinates and the three-dimensional key point coordinates;

[0113] The losses corresponding to each selected focal length are compared, and the selected focal length corresponding to the minimum loss among all losses is taken as the optimal focal length. Based on the optimal focal length, the three-dimensional space rotation and translation matrix, expression coefficient and shape coefficient are optimized until the model converges to obtain the optimal three-dimensional space rotation and translation matrix and camera observation direction, expression coefficient and shape coefficient corresponding to each frame of the face image.

[0114] In one embodiment, the present application sets an optional focal length interval, recursively selects the focal length at fixed intervals, and iteratively performs the following steps on all face images at each focal length.

[0115] Firstly, the expression coefficients are used to perform weighted summation of all 3D facial expressions in BFM2009 to obtain the average facial expression. The average facial expression, 3D spatial rotation and translation matrices are used to perform spatial deformation operations on the geometric coordinates of the 3D 68 key points to obtain new 3D 68 key points.

[0116] Then, the new 3D 68 key points and focal length are used for 2D projection to obtain new 2D 68 key point coordinates. Based on the new 2D 68 key point coordinates and 68 key point coordinates, loss constraints are constructed to optimize the 3D space rotation and translation matrices, shape and expression coefficients.

[0117] Afterwards, compare the losses at all focal lengths, select the focal length corresponding to the minimum loss as the optimal focal length, and fix the optimal focal length. Follow the steps again to optimize the 3D space rotation and translation matrix, shape and expression coefficients until convergence. You can get the optimal 3D space rotation and translation matrix and camera observation direction, expression coefficient and shape coefficient of the 3D face model for each frame of the face image.

[0118] The 50,000 3D key points in BFM2009 are spatially deformed using the optimal 3D spatial rotation and translation matrix of each face frame, expression coefficient, and shape coefficient of the 3D face model to obtain the coordinates of the 50,000 deformed 3D key points corresponding to each face frame.

[0119] 104. According to the key point coordinates and the camera observation direction, the spatial geometric features of multi-resolution hash coding are obtained, and the conditional features are obtained in combination with the audio features. Based on the conditional features and high-frequency coding information, the volume density and color are constructed and rendered to generate the digital human facial expression image driven by audio.

[0120] Specifically, in one embodiment of the present application, the spatial geometric features of multi-resolution hash coding are obtained according to the key point coordinates and the camera observation direction, which specifically includes:

[0121] Through the three-plane decoupling method of three-dimensional space, the three-dimensional space is reduced to two-dimensional space;

[0122] Through the multi-resolution hash coding voxel grid method, the coordinates of several key points of the three-dimensional face model and the camera observation direction corresponding to each frame of face image are frequency encoded to obtain the spatial geometric features of multi-resolution hash coding.

[0123] In one embodiment, the coordinates of 50,000 deformed 3D key points and camera observation directions corresponding to each frame of the face image are frequency encoded to introduce high-frequency information. The frequency encoding adopts a multi-resolution hash coding voxel grid method. At the same time, the 3D space is reduced to 2D by a three-plane decoupling method of stereo space to minimize hash conflicts, and finally the spatial geometric features of the three-plane multi-resolution hash coding are obtained.

[0124] In one embodiment of the present application, conditional features are obtained by combining audio features, and body density and color are constructed and rendered based on the conditional features and high-frequency coding information to generate a digital human facial expression image driven by audio, specifically including:

[0125] The spatial geometric features are input into the first multi-layer perceptron to construct a weighted mapping from audio features to blink features, and the audio features and blink features are weighted in channel dimension according to the output weighted values ​​to obtain the corresponding conditional features; the conditional features and high-frequency coding information are input into the second multi-layer perceptron to construct a mapping of body density and color, and an image mapping of body density and color to two-dimensional space is constructed through neural rendering to generate an audio-driven digital human facial expression image.

[0126] In one embodiment, the audio is brought into the audio coding module of the ALmD model to obtain audio features. The three-plane multi-resolution hash coding is brought into the MLP to construct a weighted mapping to the audio and blink features, and the output weighted values ​​are used to weight the audio and blink features in the channel dimension to obtain conditional features.

[0127] The conditional features and high-frequency coding information are brought into the "second multi-layer perceptron" to construct a mapping to volume density and color. The neural radiation field is an MLP network. The first half is replaced by high-frequency coding and the second half is the "second multi-layer perceptron" network here.

[0128] Then, neural rendering is used to construct a mapping of density and color to 2D spatial images, and finally the face image driven by audio is obtained. In terms of time sequence, all driven face images and conditional audio are integrated according to FPS to obtain the final audio-driven digital human speaking video.

[0129] The above is an embodiment of the method proposed in this application. Based on the same inventive concept, this application embodiment also provides a digital human facial expression synchronization device, whose structure is as follows: Figure 2 shown.

[0130] Figure 2This is a schematic diagram of the internal structure of a digital human facial expression synchronization device provided in an embodiment of the present application. Figure 2 As shown, the device includes:

[0131] at least one processor;

[0132] and, a memory communicatively coupled to the at least one processor;

[0133] The memory stores instructions that can be executed by at least one processor, and the instructions are executed by at least one processor to enable the at least one processor to:

[0134] Divide the target video collected in the specified environment to obtain the corresponding video set, and pre-process the video set to obtain the standard video set;

[0135] Based on the pre-built audio and lip motion synchronization discriminator, audio features and video features are extracted to calculate the cosine similarity between audio features and video features, and combined with the binary cross entropy loss function, the audio features and video features are converted in the same parameter space;

[0136] The 3D morphological model of the face is built based on the standard video set, and the 3D space rotation and translation matrix, camera observation direction, expression coefficient and shape coefficient of the 3D face model of each frame of the face image are obtained, and the average face shape and key point coordinates are obtained;

[0137] According to the key point coordinates and the camera observation direction, the spatial geometric features of multi-resolution hash coding are obtained, and the conditional features are obtained by combining the audio features. Based on the conditional features and high-frequency coding information, the volume density and color are constructed and rendered to generate the audio-driven digital human facial expression image.

[0138] The present application also provides a non-volatile computer storage medium storing computer executable instructions. When the computer executable instructions are executed, they can:

[0139] Divide the target video collected in the specified environment to obtain the corresponding video set, and pre-process the video set to obtain the standard video set;

[0140] Based on the pre-built audio and lip motion synchronization discriminator, audio features and video features are extracted to calculate the cosine similarity between audio features and video features, and combined with the binary cross entropy loss function, the audio features and video features are converted in the same parameter space;

[0141] The 3D morphological model of the face is built based on the standard video set, and the 3D space rotation and translation matrix, camera observation direction, expression coefficient and shape coefficient of the 3D face model of each frame of the face image are obtained, and the average face shape and key point coordinates are obtained;

[0142] According to the key point coordinates and the camera observation direction, the spatial geometric features of multi-resolution hash coding are obtained, and the conditional features are obtained by combining the audio features. Based on the conditional features and high-frequency coding information, the volume density and color are constructed and rendered to generate the audio-driven digital human facial expression image.

[0143] Each embodiment in this application is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.

[0144] The above describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0145] The devices and media provided in the embodiments of the present application correspond one-to-one to the methods. Therefore, the devices and media also have similar beneficial technical effects as the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the devices and media will not be repeated here.

[0146] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0147] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0148] These computer program instructions may also be stored in a computer readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture including an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0149] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the computer or other programmable device. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0150] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0151] The memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0152] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.

[0153] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0154] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.

Claims

1. A digital human facial expression synchronization method, characterized in that: The method comprises: Dividing the target video collected under the specified environment to obtain a corresponding video set, and preprocessing the video set to obtain a standard video set; Based on a pre-built audio and lip motion synchronization discriminator, extract audio features and video features to calculate the cosine similarity between the audio features and the video features, and combine with a binary cross entropy loss function to achieve conversion of the audio features and the video features in the same parameter space; Modeling a three-dimensional morphological model of the face according to the standard video set, obtaining the three-dimensional space rotation and translation matrix, camera observation direction, expression coefficient and shape coefficient of the three-dimensional face model of each frame of the face image, and obtaining the average face shape and key point coordinates; According to the key point coordinates and the camera observation direction, the spatial geometric features of multi-resolution hash coding are obtained, and the conditional features are obtained in combination with the audio features. Based on the conditional features and high-frequency coding information, the volume density and color are constructed and rendered to generate a digital human facial expression image driven by audio.

2. A digital human facial expression synchronization method according to claim 1, characterized in that: The target video collected under the specified environment is divided to obtain a corresponding video set, and the video set is preprocessed to obtain a standard video set, which specifically includes: Under a specified environment, a target video is acquired through a preset acquisition device; wherein the preset acquisition device includes a high-definition camera and a noise reduction microphone, and the target video includes a speaking video and a closed-mouth video; Performing endpoint detection on the target video based on automatic speech endpoint detection technology to identify all spoken video clips, and cutting the target video according to the detected time points; The cut videos are classified to obtain a speaking video set and a closed mouth video set, and for each video set, background noise separation and noise reduction processing are performed on the video set to form a cleaned target video set.

3. A digital human facial expression synchronization method according to claim 1, characterized in that: Extract audio and video features based on the pre-built audio and lip motion synchronization discriminator, including: Performing shot detection on the input video in the standard video set to identify scene switching in the video, performing face detection on the video in the standard video set to identify face areas in the video, and performing face tracking on the video in the standard video set to distinguish different speakers in a multi-person conversation scene; Receiving a first preset number of target images, splicing the first preset number of target images in a channel dimension, retaining time series information of the target images, and performing feature encoding on the spliced ​​target images to generate a video face feature vector of a target number of dimensions; Audio features are extracted from the input video of the standard video set through Mel-cepstral coefficients, and feature encoding is performed on a second preset number of Mel-cepstral coefficient audio features to generate an audio feature vector of a target number of dimensions.

4. A digital human facial expression synchronization method according to claim 1, characterized in that: Calculating the cosine similarity between the audio feature and the video feature, and combining the binary cross entropy loss function to achieve the conversion of the audio feature and the video feature in the same parameter space, specifically including: Calculate the cosine similarity between the video face feature vector and the audio feature vector in the target number of dimensions; The cosine similarity calculation formula of audio and video features is as follows: Among them, Sim(FaceF 5 ,AudF 20 ) represents the cosine similarity of audio and video features, FaceF 5 Represents the video face feature vector, AudF 20 represents the audio feature vector, FaceF 5 ·AudF 20 represents the dot product of two eigenvectors, ‖FaceF 5 ‖2‖AudF 20 ‖2 represents the product of the L2 norm of two vectors; Based on the cosine similarity and in combination with a binary cross entropy loss function, a binary cross entropy loss between the video face feature vector and the audio feature vector is calculated; The binary cross entropy loss calculation formula for audio and video features is as follows: Loss av-sync =-(ylogSim(FaceF 5 ,AudF 20 )+(1-y)log(1-Sim(FaceF 5 ,AudF 20 ))) Among them, Loss av-sync represents the binary cross entropy loss of audio and video features, y represents the real distribution of audio and video synchronization, 1-y represents the real distribution of audio and video asynchrony, Sim(FaceF 5 ,AudF 20 ) represents the predicted probability distribution of audio and video synchronization, 1-Sim(FaceF 5 ,AudF 20 ) represents the predicted probability distribution of audio and video asynchrony.

5. A digital human facial expression synchronization method according to claim 1, characterized in that: A three-dimensional morphological model of the face is constructed according to the standard video set, and the three-dimensional space rotation and translation matrix, camera observation direction, expression coefficient and shape coefficient of the three-dimensional face model of each frame of the face image are obtained, and the average face shape and key point coordinates are obtained, specifically including: Modeling a three-dimensional morphological model of the face according to the face image, the two-dimensional key point coordinates and the open source three-dimensional coordinates of the face in the standard data set, and determining the three-dimensional key points corresponding to the two-dimensional key points to associate the two-dimensional key points with the corresponding three-dimensional key points; For each frame of face image, construct the corresponding three-dimensional space rotation and translation matrix and expression coefficient vector, and construct the shape coefficient vector for all face images; According to the shape coefficients corresponding to the shape coefficient vectors, weighted summation is performed on all three-dimensional face shapes to obtain an average face shape and three-dimensional key point coordinates.

6. A digital human facial expression synchronization method according to claim 5, characterized in that: After obtaining the average face shape and key point coordinates, the method further includes: In a preset focal length range, a focal length is recursively selected at a fixed interval, and at each selected focal length, all three-dimensional facial expressions are weighted summed according to the expression coefficient corresponding to the expression coefficient vector to obtain an average facial expression; According to the average facial expression and the three-dimensional spatial rotation and translation matrix, spatially deforming the three-dimensional key point coordinates to obtain new three-dimensional key point coordinates; Performing a two-dimensional projection according to the new three-dimensional key point coordinates and the selected focal length to obtain new two-dimensional key point coordinates, and constructing a loss constraint according to the new two-dimensional key point coordinates and the three-dimensional key point coordinates; The losses corresponding to each selected focal length are compared to take the selected focal length corresponding to the minimum loss among all losses as the optimal focal length, and based on the optimal focal length, the three-dimensional space rotation and translation matrix, expression coefficient and shape coefficient are optimized until the model converges to obtain the optimal three-dimensional space rotation and translation matrix and camera observation direction as well as expression coefficient and shape coefficient corresponding to each frame of the face image.

7. A digital human facial expression synchronization method according to claim 1, characterized in that: According to the key point coordinates and the camera observation direction, the spatial geometric features of the multi-resolution hash code are obtained, specifically including: Through the three-plane decoupling method of three-dimensional space, the three-dimensional space is reduced to two-dimensional space; Through the multi-resolution hash coding voxel grid method, the coordinates of several key points of the three-dimensional face model corresponding to each frame of the face image and the camera observation direction are frequency encoded to obtain the spatial geometric features of the multi-resolution hash coding.

8. A digital human facial expression synchronization method according to claim 1, characterized in that: Combining the audio features to obtain conditional features, constructing and rendering body density and color based on the conditional features and high-frequency coding information to generate an audio-driven digital human facial expression image, specifically including: The spatial geometric features are input into a first multi-layer perceptron to construct a weighted mapping from audio features to blink features, and the audio features and blink features are weighted in channel dimension according to the output weighted values ​​to obtain corresponding conditional features; the conditional features and high-frequency coding information are input into a second multi-layer perceptron to construct a mapping of volume density and color, and an image mapping of the volume density and the color to a two-dimensional space is constructed through neural rendering to generate an audio-driven digital human facial expression image.

9. A digital human facial expression synchronization device, characterized in that: The device comprises: at least one processor; and, a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute a digital human facial expression synchronization method as described in any one of claims 1-8.

10. A non-volatile computer storage medium storing computer executable instructions, characterized in that: When the computer executable instructions are executed, a digital human facial expression synchronization method as described in any one of claims 1 to 8 is implemented.

Citation Information

Cited By

  • Digital human-driven key point generation method and device, and medium

    CN120876883A

  • Digital human coding method and device based on region of interest and long-term reference frame

    CN121644814A