Method, device and equipment for driving digital human by audio, and storage medium

By extracting audio features and fusing 3D key points and portrait information, high-quality speaker videos are generated, which solves the shortcomings of end-to-end models in terms of sound lip synchronization and image quality, and improves the credibility and information transmission effect of digital people.

CN120220719APending Publication Date: 2025-06-27PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510253097.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, the end-to-end model of the speaker video synthesis method has problems in lip synchronization and image quality, and it is difficult to accurately match the voice and lip shape in long facial motion sequences, resulting in the lip shape and voice that do not correspond to the voice, affecting the credibility of digital people and the information transmission effect.

Method used

By extracting the audio characteristics of the target audio, input the first preset model to obtain the target face 3D key points, and fuse it with the portrait information of the target person, and input the second preset model to generate the image information of the target person saying the target audio.

Benefits of technology

This method can improve the accuracy and image quality of sound and lip synchronization while reducing excessive dependence on historical data and reducing sample period selection sensitivity. It is suitable for explaining complex content for a long time, enhancing the credibility and information transmission effect of digital people.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220719A_ABST
    Figure CN120220719A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence and the field of medical treatment and health, and discloses a method, device and equipment for driving a digital human through an audio, and a storage medium, and the method comprises the steps: extracting the audio features of a target audio; inputting the audio features into a first preset model to obtain target human face 3D key points corresponding to the target audio; fusing the target face 3D key point and the portrait information of the target person to obtain fused face information; and inputting the fused face information into a second preset model to obtain image information of the target person speaking the target audio. The invention provides a method, a device and equipment for driving a digital human through audio, and a storage medium, and solves the problems existing in a speaker video synthesis method based on an end-to-end model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and medical health, and particularly to a method, device, equipment and storage medium for driving a digital human by audio. Background Art

[0002] In the field of digital humans, the video synthesis of a speaker driven by speech is a popular research topic with great research value. Its core goal is to accurately synthesize the target face speaking video according to the input speech. In the field of medical health, this technology also has a wide range of application prospects. For example, a virtual health advisor answers medical insurance policies and explains health knowledge for customers, or a virtual claims assistant assists customers in completing the claims process, etc.

[0003] In the prior art, when accurately synthesizing the target face speaking video according to the input speech, it is usually to input the speech into the model and then directly output the target face speaking video through the model, that is, to generate the target face speaking video by using an end-to-end model. However, this method of speaker video synthesis based on the end-to-end model exposes some problems in practical applications.

[0004] On the one hand, in terms of audio-lip synchronization, the end-to-end model usually directly maps from the input speech to the output speaker video, lacking more detailed modeling and constraint mechanisms for intermediate processes such as audio-lip synchronization. When dealing with long facial motion sequences, the end-to-end model is difficult to capture the complex long-term dependencies and temporal consistency between speech and lip movements well. Taking the medical insurance policy explanation video as an example, when the virtual digital human needs to explain complex medical insurance reimbursement processes, coverage of different insurance types, etc. for a long time and coherently, due to the difficulty of the model in accurately matching speech and lip shapes in long sequences, there will be an embarrassing situation where the lip shape does not match the speech. This will not only reduce the credibility of the digital human, but may also lead to deviations in the customer's understanding of important medical insurance information, affecting the customer's accurate grasp of medical insurance policies, and further may affect their insurance participation decisions or rights enjoyment.

[0005] On the other hand, the inconsistent goals of lip movement and image quality in the training process make training difficult. In the production of health knowledge popular science videos, if you want to make the virtual digital human image vivid and clear, and at the same time ensure that its lip movement is completely consistent with the explanation voice, these two goals conflict with each other in the training process of the end-to-end model. For example, in order to improve the image quality, the model may pay too much attention to the details and beauty of the picture, but to a certain extent sacrifice the synchronization accuracy of lip movement and voice; on the contrary, if you focus on optimizing the accuracy of lip movement, it may lead to a decline in image quality, blurring, distortion and other problems. This is extremely unfavorable for the application in the field of medical insurance and health, because whether it is low-quality images or inaccurate lip synchronization, it will reduce the effectiveness of digital people in conveying information, weaken customers' acceptance of health knowledge and medical insurance services, and hinder the further promotion and application of digital people in the field of medical insurance and health. Summary of the invention

[0006] The present invention provides a method, device, equipment and storage medium for driving a digital human through audio, so as to solve the problems existing in a method for synthesizing a speaker video based on an end-to-end model.

[0007] In a first aspect, the present invention provides a method for driving a digital human using audio, comprising:

[0008] Extracting audio features of the target audio;

[0009] Input the audio features into a first preset model to obtain 3D key points of the target face corresponding to the target audio;

[0010] Fusing the target person's face 3D key points and the target person's portrait information to obtain fused face information;

[0011] The fused face information is input into a second preset model to obtain image information of the target person speaking the target audio.

[0012] In a second aspect, the present invention provides a device for driving a digital human using audio, comprising:

[0013] An extraction module, used for extracting audio features of the target audio;

[0014] A 3D key point output module, used to input the audio features into a first preset model to obtain a target face 3D key point corresponding to the target audio;

[0015] A fusion module is used to fuse the target face 3D key points and the portrait information of the target person to obtain fused face information;

[0016] The image information output module is used to input the fused face information into a second preset model to obtain the image information of the target person speaking the target audio.

[0017] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the above method for driving a digital human by audio are implemented.

[0018] In a fourth aspect, the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the above method for driving a digital human by audio are implemented.

[0019] In the solution implemented by the above method for driving a digital human by audio, firstly, over-reliance on historical data can be reduced. By comprehensively considering transaction data and influencing factors, it can adapt to market changes and avoid underestimation of risks. Secondly, the sensitivity of sample period selection can be reduced. Through a preset model, multi-dimensional data fusion and adaptive learning are realized to make the prediction more stable. Thirdly, the applicability to new products or assets lacking historical data can be improved. The influencing factors supplement information, and the model generalization ability makes up for the lack of data. Therefore, the problems existing in the historical simulation method in the prior art can be avoided.

[0020] Moreover, in the solution implemented by the above method for driving a digital human by audio, the audio features of the target audio are input into a first preset model to obtain the target face 3D key points corresponding to the target audio, so as to obtain the facial movement situation of a person speaking the target audio. Further, the fused face information obtained by fusing the target face 3D key points and the portrait information of the target person is input into a second preset model to obtain the video information of the final target person speaking the target audio, so as to further incorporate the character image finally required by the digital human on the basis of obtaining the facial movement situation of the target audio. It can be understood from the above that since the facial movement situation of the target audio and the video information for displaying the character image are output through different preset models, the problems brought by using an end-to-end model to output the speaker's video in the prior art can be avoided. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings without creative efforts based on these drawings.

[0022] Figure 1 is a flowchart of a method for driving a digital human by audio in an embodiment of the present invention;

[0023] Figure 2 is Figure 1 a flowchart of step S110 in

[0024] Figure 3 is Figure 1 a schematic diagram of a training process of the first preset model in

[0025] Figure 4 is Figure 3 a schematic diagram of a process of step S126 in

[0026] Figure 5 is Figure 1 a schematic diagram of a process of step S130 in

[0027] Figure 6 is Figure 1 another schematic diagram of a process of step S130 in

[0028] Figure 7 is Figure 5 a schematic diagram of a process of step S131 in

[0029] Figure 8 a schematic diagram of a structure of a device for an audio-driven digital human in an embodiment of the present invention;

[0030] Figure 9 a schematic diagram of a structure of a computer device in an embodiment of the present invention;

[0031] Figure 10 another schematic diagram of a structure of a computer device in an embodiment of the present invention. Detailed implementation manners

[0032] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] Please refer to Figure 1 As shown, the embodiments of the present invention provide a flowchart of a method for an audio-driven digital human, including the following steps.

[0034] Step S110, extracting audio features of a target audio.

[0035] Specifically, in this step, the target audio is the audio that the finally synthesized digital human needs to speak, and this audio can be set according to specific needs during application, and no detailed limitation is made here.

[0036] It should be noted that the audio features can be set according to specific requirements during application. For example, the audio features can be prosodic features, spectral features, timbre features, etc., and no specific limitations are imposed here.

[0037] Furthermore, the extraction method of the audio features of the target audio can be any achievable method. For example, when extracting prosodic features, the short-time average magnitude difference function can be used for extraction. When extracting spectral features, the fast Fourier transform can be used to convert the time-domain audio signal to the frequency domain to obtain spectral information.

[0038] As a specific example, in the field of medical and health, during the process of insurance claim assistance, extracting the audio features of the voice of the claimant helps to quickly evaluate the claim situation. When the claimant describes the accident process and losses, information such as pauses and intonation changes in the audio features can reflect the authenticity and complexity of the event. For example, if the audio features show that the claimant has frequent pauses or unnatural tones when describing key details, the claims adjuster can conduct further in-depth investigations to prevent insurance fraud and protect the legitimate rights and interests of the insurance company and other policyholders.

[0039] In some embodiments of the present invention, as Figure 2 shown, step S110 includes the following steps.

[0040] Step S111, generating the pitch contour feature and the Mel frequency cepstral coefficient feature of the target audio;

[0041] Step S112, splicing the pitch contour feature and the Mel frequency cepstral coefficient feature to obtain the audio feature of the target audio.

[0042] It should be noted that the pitch contour feature is a feature obtained by discretizing the fundamental frequency value of the target audio, and the Mel frequency cepstral coefficient feature is a set composed of multiple features of the audio signal, which can be used to describe the overall shape of the spectral envelope.

[0043] Specifically, in step S111, the method for generating the pitch contour feature and the Mel frequency cepstral coefficient feature of the target audio can be any achievable method.

[0044] For example, when generating the pitch contour feature of the target audio, the target audio can be framed. For each frame of audio, the autocorrelation method or the method based on the short-time Fourier transform can be used to calculate the pitch period. After obtaining the pitch period of each frame, it can be converted into a pitch value. The pitch value is inversely proportional to the pitch period. Arranging the pitch values of all frames in chronological order forms the pitch contour feature of the target audio, which can reflect the pitch change situation of the audio in the time dimension;

[0045] When generating the Mel-frequency cepstral coefficient (MFCC) features of the target audio, the target audio can first be framed, and each frame can be pre-emphasized to enhance the energy of the high-frequency part. Further, windowing can be performed to reduce spectral leakage and make the spectral analysis more accurate. Then, the fast Fourier transform (FFT) can be applied to each windowed frame to convert the time-domain signal into a frequency-domain signal, obtaining the spectrum. Further, the spectrum can be filtered through a Mel filter bank to convert the linear frequency into Mel frequency. Taking the logarithm of the result after filtering by the Mel filter bank and then performing the discrete cosine transform can obtain the MFCC features. These MFCC features can contain the main feature information of the speech signal.

[0046] Specifically, in step S112, the manner of splicing the pitch contour features and the MFCC features can be any achievable manner.

[0047] In the above manner, all frames are spliced, and finally an audio feature sequence of the complete target audio is obtained. Since the pitch contour features can reflect the mapping from the audio to the motion steps, the pitch contour features can be used as a useful facial motion indicator to improve the quality of the facial movements of the finally generated digital human. And since the audio feature sequence contains the feature information of both the pitch contour features and the MFCC features, it can describe the characteristics of the target audio more comprehensively.

[0048] As a specific example, when a customer calls the customer service hotline of a medical and health insurance company to consult issues such as insurance product details and claim settlement procedures, in step S111, the pitch contour features and the MFCC features can be generated by processing the customer's speech. If the customer has doubts about some key information, such as the calculation method of the claim amount, the pitch contour of their speech may show obvious fluctuations in relevant words. For example, when mentioning "claim amount", the pitch may increase and the speech rate may accelerate. These changes in emotions and semantic focuses can be captured through the pitch contour features. At the same time, the MFCC features can identify information such as the customer's language habits and accents. For example, customers from different regions have different accents, and their MFCC features will have certain differences. In step S112, the obtained features can be spliced to obtain the audio features of the target audio.

[0049] Step S120: Input the audio features into a first preset model to obtain the target face 3D key points corresponding to the target audio.

[0050] It should be noted that the target face 3D key points corresponding to the target audio are the face 3D key points corresponding to the facial movements generated when a person speaks the target audio.

[0051] Specifically, the first preset model in this step can be a model that can obtain the target face 3D key points corresponding to the audio features when any audio features are input. For example, the first preset model can be a regression model based on a deep model. This regression model can be built based on a multi-layer perceptron. The multi-layer perceptron consists of an input layer, multiple hidden layers, and an output layer. The previously extracted and concatenated audio features, such as the vector composed of the pitch contour feature and the mel-frequency cepstral coefficient feature, are input into the input layer of the multi-layer perceptron. The hidden layer performs non-linear transformation on the input features through a large number of neurons to learn the complex mapping relationship between the audio features and the 3D key points. Finally, the 3D key point coordinates of the target face are output in the output layer. For example, the first preset model can also be a long short-term memory network. The audio features can be input into the long short-term memory network in chronological order. The long short-term memory network can effectively process the time series information in the audio features through a gating mechanism (input gate, forget gate, output gate) and capture the dependence relationship between the speech at different moments and the facial movements. Since the facial movements are continuously changing during speech, the long short-term memory network can remember the information at the previous moment, thereby better predicting the 3D key points corresponding to the current moment. For example, when speaking continuous sentences, the facial movements corresponding to the previous phoneme will affect the movements of the next phoneme, and the long short-term memory network can use this time correlation for accurate prediction.

[0052] In some embodiments of the present invention, as Figure 3 shown, the training process of the first preset model includes the following training steps:

[0053] Step S121, obtaining a data pair composed of historical person image information and historical audio information, where the historical audio information is the audio information in the historical person image information;

[0054] Step S122, extracting the audio features of the historical audio information;

[0055] Step S123, generating the historical face 3D key points in the historical person image information;

[0056] Step S124, inputting the audio features of the historical audio information into the first preset model to be trained, and outputting the predicted face 3D key points corresponding to the historical audio information;

[0057] Step S125, calculating the distance error value between the predicted face 3D key points and the historical face 3D key points;

[0058] Step S126, in the case where the distance error value is less than the first preset value, determining the currently trained first preset model as the final first preset model.

[0059] It should be noted that the historical figure image information in step S121 is the image information formed when the person is speaking the historical audio information.

[0060] It should be noted that the historical face 3D key points in the historical figure image information generated in step S123 can be achieved in any way. For example, a large number of face 3D models with different postures and expressions or face data with accurate annotations can be collected, and these data can be statistically analyzed to construct a general face geometry template. This template can contain information such as the average position, shape, and relative relationship of each key part of the face. Further, the face template can be parameterized. For example, some shape parameters and posture parameters can be defined so that the template can adapt to the face features of different people by adjusting these parameters. Further, the constructed 2D face template can be matched with the figure image. For example, a template matching algorithm, such as a matching algorithm based on cross-correlation, can be used to find the region in the image that is most similar to the template and determine the approximate position and posture of the face in the image. Further, according to the matching result, the shape and parameters of the template can be adjusted to make it better fit the face in the image. By deforming and adjusting the template, the positions of each part of the face are determined, and thus the 2D key points of the face are obtained. Further, using the known 3D structure information of the face template and the projection model of the camera, the 2D key points can be back-projected into the 3D space to obtain the corresponding 3D key points. By calculating the projection of the 3D points on the template in the camera coordinate system and corresponding them with the 2D key points in the image, the coordinates of the 3D key points are determined. Further, the generated 3D key points can be optimized and adjusted. Considering the physiological structure and movement law of the face, some obviously unreasonable points are corrected. For example, according to prior knowledge such as the symmetry of the face and the movement range of the joints, the positions of the key points are slightly adjusted to make them more conform to the real face shape. It can be understood that the historical face 3D key points of the historical figure image information generated in step S123 are the real face 3D key points corresponding to the historical audio information.

[0061] Specifically, the predicted face 3D key points obtained in step S124 are the face 3D key points corresponding to the historical audio information predicted by the first preset model.

[0062] In this way, further, by calculating the error value between the predicted face 3D key points of the historical audio information and the real face 3D key points of the historical audio information, when the error value is less than the first preset value, the first preset model to be trained is determined as the final first preset model, which can make the finally obtained first preset model accurately output the face 3D key points corresponding to the input audio information when the audio information is input.

[0063] Specifically, the distance error between the predicted 3D face key points and the historical 3D face key points can be any distance error. For example, it can be the error of the Euclidean distance or the Mahalanobis distance. The method for calculating the distance error value between the predicted 3D face key points and the historical 3D face key points in step S125 can be any achievable method. For example, the predicted 3D face key points and the historical 3D face key points need to calculate the distance error value based on the same facial feature definition, that is, the semantics of each key point on the face is consistent, such as corresponding to the same positions like the tip of the nose, the inner corner of the left eye, etc. Further, for each pair of corresponding key points, the distance error value between the predicted 3D face key points and the historical 3D face key points can be calculated according to the selected distance metric formula.

[0064] As a specific example, in the field of medical and health, when a user applies for medical insurance, the business personnel of the insurance company can receive the user's need to apply for medical insurance, specifically ask the user which medical insurance to choose and introduce the characteristics of each medical insurance to the user. During this process, the video information and audio information of the business personnel's response to the user can be collected, and then these video information and audio information can be used as a data pair consisting of historical person video information and historical audio information to train the first preset model.

[0065] In some embodiments of the present invention, as Figure 4 shown, step S126 includes the following steps.

[0066] Step S1261, when the distance error value is less than the first preset value, input the predicted 3D face key points and the audio features of the historical audio information into the audio-visual synchronization discriminator to obtain the first synchronization probability of the predicted 3D face key points and the historical audio information;

[0067] Step S1262, calculate the second synchronization probability of the historical 3D face key points and the historical audio information;

[0068] Step S1263, calculate the cross entropy of the first synchronization probability and the second synchronization probability;

[0069] Step S1264, when the value of the cross entropy is less than the second preset value, determine the current first preset model to be trained as the final first preset model.

[0070] It can be understood that the first synchronization probability output by the audio-visual synchronization discriminator in step S1261 can be used to judge the synchronization degree between the historical audio information and the predicted 3D face key points, and essentially it is to judge the synchronization degree between the facial movements when a real person says the historical audio information and the predicted 3D face key points.

[0071] It can be understood that the second synchronization probability between the historical facial 3D key points calculated in step S1262 and the historical audio information is used to characterize the synchronization degree between the historical facial 3D key points and the historical audio information, that is, the synchronization degree between the real audio information and the facial 3D key points corresponding to the audio information. In this way, by calculating the cross-entropy between the first synchronization probability and the second synchronization probability, the difference degree between the first synchronization probability and the second synchronization probability can be measured. And it can be understood that the second synchronization probability characterizes the synchronization degree between the real audio information and the facial 3D key points corresponding to the audio information. Therefore, the smaller the value of the cross-entropy between the first synchronization probability and the second synchronization probability, the closer the predicted facial 3D key points are to the real facial 3D key points. So, in the case where the value of the cross-entropy is less than the second preset value in step S1264, determining the currently to-be-trained first preset model as the final first preset model can make the final obtained first preset model predict the facial 3D key points corresponding to the audio information accurately enough.

[0072] It can be understood that through steps S1261 - S1624, it is possible to make the predicted facial 3D key points finally output by the first preset model close to the historical facial 3D key points (by making the distance error value between the predicted facial 3D key points and the historical facial 3D key points less than the first preset value), and it is also possible to make the synchronization degree between the predicted facial 3D key points finally output by the first preset model and the input audio high enough (by making the value of the cross-entropy between the first synchronization probability and the second synchronization probability less than the second preset value). In this way, the synchronization degree between the finally output facial 3D key points and the input audio can be high enough.

[0073] Step S130: Fuse the target facial 3D key points and the portrait information of the target person to obtain fused facial information.

[0074] Specifically, in this step, the method of fusing the 3D key points of the target face and the portrait information of the target person to obtain the fused face information can be any achievable method. For example, the 3D key points can be used as the basis for texture mapping. Taking the portrait image as the texture, according to the facial geometric structure defined by the 3D key points, the texture is mapped onto the surface of the 3D model. For example, a triangular mesh is constructed using the 3D key points, and then according to the topological structure of the triangular mesh, the pixel values in the portrait image are corresponding to each triangular patch of the 3D model, so as to present the facial texture of the target person on the 3D model. It is also possible to fuse the extracted portrait image features with the features of the 3D key points. The coordinate information of the 3D key points can be encoded as a feature vector, and then concatenated or other forms of fusion operations are performed with the feature vector extracted from the portrait image. For example, the three-dimensional coordinate values of the 3D key points are used as additional feature dimensions and concatenated with the extracted portrait feature vector to form a new fused feature vector. This fused feature vector contains both the geometric structure information of the face (provided by the 3D key points) and the appearance feature information of the face (provided by the portrait image).

[0075] As a specific example, when a customer submits a claim application, they can submit voice to input into the first preset model. Then, the first preset model will output the 3D key points of the face corresponding to the voice. Furthermore, the portrait information of the customer pre-stored in the server can be retrieved, and the portrait information of the customer and the 3D key points of the face corresponding to the voice submitted by the customer are fused together to obtain the fused face information. Furthermore, the mood and facial movements of the user when speaking the voice can be judged, enabling the claims adjuster to judge the credibility of the situation described by the customer, assisting in the review of the claim application, and reducing the fraud risk.

[0076] In some embodiments of the present invention, as Figure 5 shown, step S130 includes the following steps.

[0077] Step S131, rendering the 3D key points of the target face into a 3D mask;

[0078] Step S132, obtaining the lower half face information of the target person;

[0079] Step S133, replacing the lower half face information of the 3D mask with the lower half face information of the target person to obtain the first fused face information.

[0080] Specifically, in step S131, a basic 3D face model can be constructed based on the target face 3D key points. These key points define the key feature positions of the face, such as the contour and position information of parts like eyes, nose, mouth, cheeks, etc. By connecting these key points, a preliminary 3D mesh model can be generated using a surface fitting algorithm to outline the general shape of the face. Additionally, according to parameters such as the set lighting conditions and viewing angles, the 3D model with added materials and textures can be rendered, and finally a realistic 3D mask can be generated. This 3D mask retains the facial shape and features defined by the target face 3D key points and can serve as the basis for subsequent fusion operations.

[0081] It can be understood that the 3D mask is constructed based on 3D key points, can accurately simulate the geometric shape and dynamic features of the human face, and combined with high-quality material and texture rendering, also provides a highly realistic facial appearance. Incorporating the real lower half-face information of the target person further enhances the realism and detail expressiveness of the face, making the fused face more vivid and natural, and applicable to various scenarios that require highly realistic face images, such as virtual character creation.

[0082] In some embodiments of the present invention, as Figure 6 shown, after step S133, the following steps are further included.

[0083] Step S134, obtaining the face information of the target person;

[0084] Step S135, splicing the first fused face information and the face information of the target person to obtain the second fused face information.

[0085] It should be noted that in step S135, the manner of splicing the first fused face information and the face information of the target person to obtain the second fused face information can be any achievable manner. For example, the first fused face information and the face information of the target person can be spliced using a weighted average method based on their pixels.

[0086] It can be understood that by specifically obtaining the comprehensive face information of the target person and splicing it with the first fused face information, various features of the face can be further supplemented and refined. This includes not only the facial details not covered in the previous steps but also the facial expressions under different conditions, making the final second fused face information more complete and accurate in reflecting the facial characteristics of the target person, and enhancing the authenticity and recognition of the face. Enhancing personalization and uniqueness: Each person's face has unique features. Obtaining and integrating more face information of the target person himself can strengthen the personalization of the fused face. Whether it is unique facial textures, subtle expression features or special facial contours, they can all be reflected in the splicing process, making the generated second fused face information more unique and meeting the needs of restoring a specific person's image or personalized creation.

[0087] In some embodiments of the present invention, as Figure 7 shown, step S131 includes the following steps.

[0088] Step S1311, triangulate the 3D key points of the target face to obtain and connect multiple triangles to form a face information with a triangular mesh structure;

[0089] Step S1312, based on the preset camera parameters, preset lighting parameters, and preset angle parameters, render the face information with the triangular mesh structure to obtain a 3D mask corresponding to the 3D key points of the target face.

[0090] It should be noted that the preset camera parameters, preset lighting parameters, and preset angle parameters can be set according to the specific needs during application and will not be specifically limited here.

[0091] Specifically, the preset camera parameters can determine the viewing angle and imaging method of observing a human face. The main parameters may include the position of the camera (coordinates in 3D space), the orientation of the camera (represented by a rotation matrix), the focal length of the camera, etc. The camera position determines where to observe the human face. For example, when observing from the front, the camera position is in front of the human face. The camera orientation determines the direction of the camera lens. The focal length affects the scaling ratio and depth-of-field effect of the image. A shorter focal length produces a wide-angle effect, making the image look broader, while a longer focal length magnifies details and makes the image more focused. The preset lighting parameters can be used to simulate the lighting conditions in the scene, affecting the light and dark and color distribution on the human face surface. For example, the preset lighting parameters may include the position of the light source (coordinates in 3D space), the intensity of the light source (determining the brightness of the lighting), the color of the light source (such as white, yellow, etc., affecting the color reflection of the object surface), and the lighting model. Different lighting models are based on different physical principles and are used to calculate the light intensity received by the object surface and the reflection effect. The preset angle parameters are the rotation angles of the human face relative to the camera and the light source. By adjusting these angles, the human face can be observed from different directions and the illumination angle of the light on the human face can be changed. For example, rotating the human face around a certain axis by a certain angle can show the side or oblique side of the human face; changing the angle of the light can simulate the natural light at different times or the artificial light effect in a specific scene.

[0092] Specifically, triangulation is the process of connecting the 3D key points of the target human face into multiple triangles to construct a triangular mesh structure. Its core goal is to find a reasonable connection method based on the given set of 3D key points, so that these triangles can accurately approximate the surface shape of the human face. The triangulation algorithm can be any algorithm that can be implemented. For example, it can be the Delaunay triangulation algorithm, which has the property of an empty circumcircle, that is, there are no other key points inside the circumcircle of each triangle. This helps to generate a triangular mesh with better quality and avoid the appearance of overly long and narrow or irregular triangles.

[0093] It can be understood that the triangular mesh structure obtained through triangulation in step S1311 is an efficient and general geometric representation method. It can accurately describe the surface shape of the human face and provide a good foundation for subsequent rendering and animation processing. The triangular mesh structure is easy to process and operate and is widely used in computational geometry and graphics. Many algorithms and technologies can be developed based on this structure, such as collision detection and deformation animation, which makes it more flexible and efficient when dealing with face-related applications. Rendering based on the preset camera parameters, lighting parameters, and angle parameters in step S1312 can generate a highly realistic 3D mask. The camera parameters control the viewing angle, enabling the face to be observed from different angles to meet the requirements of different application scenarios. For example, in virtual reality or animation production, it may be necessary to display the character from multiple perspectives. The setting of the lighting parameters simulates the lighting effects in the real world, giving the surface of the 3D mask light and dark variations and a three-dimensional sense, enhancing the visual realism. The angle parameters further enrich the diversity of rendering, allowing the simulation of human faces in different poses and providing more possibilities for animation production or simulation scenarios.

[0094] Step S140: Input the fused face information into a second preset model to obtain the video information of the target person speaking the target audio.

[0095] It should be noted that the second preset model in step S140 can be any model that can obtain the video information of the target person speaking the target audio (i.e., the video information of the digital human speaking the target audio) when inputting the fused face information of the person. For example, the second preset model can be a generative adversarial network.

[0096] Specifically, after inputting the fused face information into the second preset model, a series of neural network layers inside the model process the input fused face information. The convolutional layer may be used to extract the spatial features of the face and capture the details and structure of the face; the recurrent layer processes the temporal dependence between the audio and facial movements and predicts the actions and expressions that the target person's face should present at each time point of the target audio. Through the complex operations and feature extraction of these layers, the model gradually generates the video information of the target person speaking the target audio, and this video information may be presented in the form of a sequence of video frames, with each video frame showing the facial state of the target person speaking at a specific moment.

[0097] Optionally, when the second preset model is a generative adversarial network, the loss function of the model can include the L1 distance loss function, the Visual Geometry Group (VGG) perceptual loss function, and the adversarial loss function. The value of the final loss function can be the value obtained by setting weights for the above three loss function values and weighting the above three loss function values.

[0098] It is understandable that, by combining the L1 distance loss, the VGG perceptual loss, and the adversarial loss, the generated image information is optimized from different dimensions. The L1 distance loss ensures pixel-level similarity, making the generated image close to the real image in terms of basic brightness, color, etc.; the VGG perceptual loss focuses on semantic and structural features, improving the visual perception quality of the image; the adversarial loss further improves the authenticity and fidelity of the image through the adversarial training of the generator and the discriminator. The combination of the three comprehensively improves the quality of the generated image. The VGG perceptual loss utilizes the image semantic information learned by the pre-trained network, and the adversarial loss generates images by simulating the real data distribution, making the generated image information of the target person speaking the target audio more in line with the visual effects and human perception in the real scene, and being able to more accurately present details such as the facial expressions and movements of the target person when speaking, enhancing the realism and credibility of the image.

[0099] As a specific example, assume that a medical insurance institution wants to create a virtual health advisor to answer questions about medical insurance policies, health management knowledge, etc. for customers. First, a large number of videos of real doctors or professionals explaining relevant content are collected as real image data, and at the same time, the face data of customers is collected and processed to obtain fused face information. For example, the face of customer Zhang is processed through the above steps such as extracting 3D key points and fusing portrait information, and the fused face information containing Zhang's facial features is obtained. The fused face information of Zhang is input into the second preset model, and the goal of the model is to generate the image information of Zhang as a virtual health advisor explaining the medical insurance reimbursement process. The neural network layers inside the model process the fused face information, use the convolutional layer to extract the spatial features of Zhang's face, and the recurrent layer processes the temporal dependence between the audio (the audio of the medical insurance reimbursement process explanation) and the facial movements.

[0100] In this way, in the embodiment of the present application, the audio features of the target audio are input into the first preset model to obtain the 3D key points of the target face corresponding to the target audio, so as to obtain the facial movement situation of the person speaking the target audio. Further, the fused face information obtained by fusing the 3D key points of the target face and the portrait information of the target person is input into the second preset model to obtain the final image information of the target person speaking the target audio, thereby further integrating the character image finally required by the digital human on the basis of obtaining the facial movement situation of the target audio. It can be understood from the above that, since the facial movement situation of the target audio and the image information for displaying the character image are output through different preset models, the problems brought by using an end-to-end model to output the speaker's video in the prior art can be avoided.

[0101] It should be understood that the sequence numbers of the steps in the above embodiments do not indicate the order of execution, and the execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention. The non-company software tools or components that appear in the embodiments of this application are only introduced by way of example and do not represent actual use.

[0102] In one embodiment, a device for driving a digital human by audio is provided, and the device for driving a digital human by audio corresponds one-to-one to the method for driving a digital human by audio in the above embodiment. As Figure 8 shown, the device for driving a digital human by audio includes an extraction module 810, a 3D key point output module 820, a fusion module 830, and an image information output module 840. The detailed description of each functional module is as follows:

[0103] The extraction module 810 is used to extract the audio features of the target audio;

[0104] The 3D key point output module 820 is used to input the audio features into a first preset model to obtain the target face 3D key points corresponding to the target audio;

[0105] The fusion module 830 is used to fuse the target face 3D key points and the portrait information of the target person to obtain the fused face information;

[0106] The image information output module 840 is used to input the fused face information into a second preset model to obtain the image information of the target person speaking the target audio.

[0107] In one embodiment, the training process of the first preset model includes the following training steps:

[0108] Obtain a data pair composed of historical person image information and historical audio information, where the historical audio information is the audio information in the historical person image information;

[0109] Extract the audio features of the historical audio information;

[0110] Generate the historical face 3D key points in the historical person image information;

[0111] Input the audio features of the historical audio information into the first preset model to be trained, and output the predicted face 3D key points corresponding to the historical audio information;

[0112] Calculate the distance error value between the predicted face 3D key points and the historical face 3D key points;

[0113] In the case that the distance error value is less than the first preset value, determine the currently trained first preset model as the final first preset model.

[0114] In one embodiment, when the distance error value is less than a first preset value, determining the current first preset model to be trained as the final first preset model includes:

[0115] When the distance error value is less than the first preset value, inputting the predicted 3D face key points and the audio features of the historical audio information into a lip-sync discriminator to obtain a first synchronization probability of the predicted 3D face key points and the historical audio information;

[0116] Calculating a second synchronization probability of the historical 3D face key points and the historical audio information;

[0117] Calculating the cross-entropy of the first synchronization probability and the second synchronization probability;

[0118] When the value of the cross-entropy is less than a second preset value, determining the current first preset model to be trained as the final first preset model.

[0119] In one embodiment, the extraction module 810 is specifically configured to:

[0120] Generate a pitch contour feature and a Mel-frequency cepstral coefficient feature of the target audio;

[0121] Concatenate the pitch contour feature and the Mel-frequency cepstral coefficient feature to obtain the audio feature of the target audio.

[0122] In one embodiment, the fusion module 830 is specifically configured to:

[0123] Render the target 3D face key points into a 3D mask;

[0124] Obtain the lower half face information of the target person;

[0125] Replace the lower half face information of the 3D mask with the lower half face information of the target person to obtain first fused face information.

[0126] In one embodiment, the fusion module 830 is further configured to:

[0127] Obtain the face information of the target person;

[0128] Concatenate the first fused face information and the face information of the target person to obtain second fused face information.

[0129] In one embodiment, the fusion module 830 is further configured to:

[0130] Perform triangulation on the target 3D face key points to obtain and connect a plurality of triangles to form face information with a triangular mesh structure;

[0131] Render the face information of the triangular mesh structure based on preset camera parameters, preset lighting parameters, and preset angle parameters to obtain a 3D mask corresponding to the target face 3D key points.

[0132] The present invention provides a device for an audio-driven digital human. The audio features of a target audio are input into a first preset model to obtain the target face 3D key points corresponding to the target audio, so as to obtain the facial movement situation of a person speaking the target audio. Further, the fused face information obtained by fusing the target face 3D key points and the portrait information of the target person is input into a second preset model to obtain the video information of the final target person speaking the target audio, so as to further incorporate the character image finally required by the digital human on the basis of obtaining the facial movement situation of the target audio. It can be understood from the above that since the facial movement situation of the target audio and the video information for displaying the character image are output through different preset models, the problems brought about by using an end-to-end model to output the speaker's video in the prior art can be avoided.

[0133] For the specific limitations of the device for an audio-driven digital human, reference can be made to the limitations of the method for an audio-driven digital human in the above text, which will not be elaborated here. Each module in the above device for an audio-driven digital human can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor in the computer device in hardware form or be independent of the processor, or can be stored in the memory in the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.

[0134] Based on the above method for an audio-driven digital human, as Figure 9 shown, an embodiment of the present invention further provides a schematic structural diagram of a device for an audio-driven digital human. The device includes a processor 91 and a memory 92 coupled to the processor 91. The memory 92 stores a computer program. When the computer program is executed by the processor 91, the processor 91 is caused to execute the steps of the method for an audio-driven digital human in the above embodiment.

[0135] For other details of the implementation of the above technical solution by the processor 91 in the device for an audio-driven digital human, reference can be made to the description in the method for an audio-driven digital human provided in the above embodiment of the present invention, which will not be elaborated here.

[0136] Among them, the processor 91 can also be referred to as a CPU (Central Processing Unit), and the processor 91 may be an integrated circuit chip with signal processing capabilities; the processor 91 can also be a general-purpose processor, a DSP (Digital Signal Process), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. Among them, the general-purpose processor can be a microprocessor or the processor 91 can also be any conventional processor, etc.

[0137] As Figure 10 shown, the embodiment of the present invention also provides a structural schematic diagram of a computer-readable storage medium, on which a readable computer program 101 is stored; among them, the computer program 101 can be stored in the above storage medium in the form of a software product, including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, magnetic disks, or optical discs, ROM (Read-Only Memory), RAM (Random Access Memory), etc., or terminal devices such as computers, servers, mobile phones, and tablets.

[0138] In several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or module can be in an electrical, mechanical, or other form.

[0139] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical module, that is, it may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0140] In addition, in each embodiment of the present invention, the functional modules can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.

[0141] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0142] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium (such as an SSD (solid state disk)).

[0143] The technical solutions provided by the present invention have been introduced in detail above. Specific examples are used in the present invention to illustrate the principles and implementation manners of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation of the present invention.

[0144] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) that contain computer-usable program code.

[0145] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0146] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0147] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, such that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process, so that the instructions executed on the computer or other programmable devices provide steps for implementing the functions specified in one Figure 1 flow or multiple flows and / or blocks Figure 1 or multiple blocks.

[0148] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations. The non-company software tools or components that appear in the embodiments of the present invention are only introduced by way of example and do not represent actual use.

Claims

1. A method for driving a digital human using audio, characterized in that: include: Extracting audio features of the target audio; Input the audio features into a first preset model to obtain 3D key points of the target face corresponding to the target audio; Fusing the target person's face 3D key points and the target person's portrait information to obtain fused face information; The fused face information is input into a second preset model to obtain image information of the target person speaking the target audio.

2. The method for driving digital human by audio according to claim 1, characterized in that: The training process of the first preset model includes the following training steps: Acquire a data pair consisting of image information of a historical figure and historical audio information, wherein the historical audio information is the audio information in the image information of the historical figure; Extracting audio features of the historical audio information; Generate 3D key points of historical faces in the image information of the historical figures; Inputting the audio features of the historical audio information into a first preset model to be trained, and outputting the predicted 3D key points of the face corresponding to the historical audio information; Calculating the distance error between the predicted 3D facial key points and the historical 3D facial key points; When the distance error value is less than the first preset value, the current first preset model to be trained is determined as the final first preset model.

3. The method for driving digital human by audio according to claim 2, characterized in that: When the distance error value is less than the first preset value, determining the current first preset model to be trained as the final first preset model includes: When the distance error value is less than a first preset value, inputting the predicted 3D facial key points and the audio features of the historical audio information into a lip synchronization discriminator to obtain a first synchronization probability of the predicted 3D facial key points and the historical audio information; Calculating a second synchronization probability between the historical face 3D key points and the historical audio information; Calculating the cross entropy of the first synchronization probability and the second synchronization probability; When the cross entropy value is less than the second preset value, the current first preset model to be trained is determined as the final first preset model.

4. The method for driving digital human by audio according to claim 1, characterized in that: The extracting audio features of the target audio includes: Generate pitch contour features and Mel frequency cepstral coefficient features of the target audio; The pitch contour feature and the Mel-frequency cepstral coefficient feature are concatenated to obtain the audio feature of the target audio.

5. The method for driving digital human by audio according to claim 1, characterized in that: The step of fusing the target face 3D key points and the target person's portrait information to obtain fused face information includes: Rendering the target face 3D key points into a 3D mask; Get the lower half of the target person's face information; The lower half face information of the 3D mask is replaced with the lower half face information of the target person to obtain the first fused face information.

6. The method for driving digital human by audio according to claim 5, characterized in that: After replacing the lower half face information of the 3D mask with the lower half face information of the target person to obtain the first fused face information, the method further includes: Get the target person's facial information; The first fused face information and the face information of the target person are spliced ​​together to obtain second fused face information.

7. The method for driving digital human by audio according to claim 5, characterized in that: The step of rendering the target face 3D key points into a 3D mask comprises: Triangulate the target face 3D key points to obtain and connect multiple triangles to form face information with a triangular mesh structure; Based on preset camera parameters, preset lighting parameters and preset angle parameters, the face information of the triangular mesh structure is rendered to obtain a 3D mask corresponding to the 3D key points of the target face.

8. An audio-driven digital human device, characterized in that: include: An extraction module, used for extracting audio features of the target audio; A 3D key point output module, used to input the audio features into a first preset model to obtain a target face 3D key point corresponding to the target audio; A fusion module is used to fuse the target face 3D key points and the portrait information of the target person to obtain fused face information; The image information output module is used to input the fused face information into a second preset model to obtain the image information of the target person speaking the target audio.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method for audio-driven digital human according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method for audio-driven digital human according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Task scheduling method and system based on hybrid cloud digital human processing architecture

    CN122179477A