Digital human generation method and device, program product and electronic equipment

By extracting features and fusing feature change vectors in digital human generation methods, and using a preset image decoder to generate target digital humans, the problem of insufficient matching degree between digital human expressions and movements is solved, and high-quality lip-sync and naturalness are achieved.

CN121883673APending Publication Date: 2026-04-17INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INDUSTRIAL AND COMMERCIAL BANK OF CHINA
Filing Date
2025-12-22
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Current digital avatars have poor matching accuracy in facial expressions, movements, and audio, making it difficult to achieve perfect matching and lip-sync.

Method used

By receiving a reference image and a target video, the video is split into driving video and audio. Feature extraction is performed, and a preset motion-driven network is used to determine the feature change vector. Combined with the audio feature vector, a preset image decoder is used to generate the target digital human, including a fitting module and eye and mouth repositioning modules to optimize the changes in feature points.

Benefits of technology

It achieves perfect matching of digital human facial expressions and movements, as well as lip-syncing, reducing generation costs and improving the naturalness and flexibility of digital humans.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121883673A_ABST
    Figure CN121883673A_ABST
Patent Text Reader

Abstract

The invention discloses a digital human generation method and device, a program product and electronic equipment, and relates to the field of artificial intelligence or other related fields, and the generation method comprises the steps: receiving a reference image and a target video, splitting the target video into a driving video and an audio, carrying out the feature extraction of the reference image and the driving video, and carrying out the feature extraction of the audio; obtaining an original feature vector of the reference image and a driving feature vector of the driving video, carrying out feature extraction on the audio to obtain an audio feature vector of the audio, determining a feature change vector by adopting a preset action driving network based on the original feature vector and the driving feature vector, and fusing the audio feature vector and the driving feature vector to obtain an audio feature vector of the audio. And inputting the original feature vector, the feature change vector and the fused feature vector into a preset image decoder to obtain a target digital human. According to the invention, the technical problem that the matching degree of the generated digital human in terms of expressions, actions and audios is poor in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and more specifically, to a method, apparatus, program product, and electronic device for generating digital humans. Background Technology

[0002] Digital humans are virtual characters built using computer graphics, artificial intelligence (AI), and animation technologies, encompassing image processing, AI speech synthesis, and natural language interaction. Through lightweight design and low-cost production, digital humans are adaptable to multiple platforms, including mobile and web browsers, and are widely used in scenarios such as virtual anchors, virtual teachers in online education, AI customer service for enterprises, and virtual brand ambassadors. The advantages of digital humans lie in their flexible styles (supporting anime, realistic, and other art styles), real-time interaction (voice dialogue and facial expression feedback), and efficient deployment capabilities, making them a preferred tool for enterprises to reduce costs and increase efficiency, and for individual content creators.

[0003] In related technologies, generative adversarial networks (GANs) are used to extract features from single or multiple images to generate digital human avatars. For example, virtual anchors can quickly adapt to different facial movements by transferring facial expressions, but this relies on high-quality training data. Parametric face models are used to reconstruct 3D face topology from 2D images, and texture mapping is combined to generate a drivable digital human, but this requires accurate keypoint detection and registration algorithms. Training neural radiation fields through multi-view images can generate high-fidelity 3D digital humans that support dynamic expressions and lighting changes, but this has strict requirements on the number and angle of the input images.

[0004] Therefore, digital human generation in related technologies faces problems such as high data dependence (requiring a large number of labeled or multi-view images), high computational cost (such as time-consuming model training and high inference latency of diffusion models), and insufficient generalization ability (such as decreased cross-style performance). Furthermore, current algorithms struggle to achieve perfect matching of expressions and movements, as well as to produce smooth facial expressions and movements while achieving lip-sync.

[0005] There is currently no effective solution to the above problems. Summary of the Invention

[0006] This invention provides a method, apparatus, program product, and electronic device for generating digital humans, to at least solve the technical problem of poor matching of digital humans in terms of facial expressions, movements, and audio in related technologies.

[0007] According to one aspect of the present invention, a method for generating a digital human is provided, comprising: receiving a reference image and a target video, and splitting the target video into a driving video and audio; performing feature extraction on the reference image and the driving video to obtain an original feature vector of the reference image and a driving feature vector of the driving video, and performing feature extraction on the audio to obtain an audio feature vector of the audio; determining a feature change vector using a preset action driving network based on the original feature vector and the driving feature vector; fusing the audio feature vector and the driving feature vector to obtain a fused feature vector; and inputting the original feature vector, the feature change vector, and the fused feature vector into a preset image decoder to obtain the target digital human.

[0008] Furthermore, the preset action-driven network includes at least a fitting module, an eye redirection module, and a mouth redirection module. The step of determining the feature change vector using the preset action-driven network based on the original feature vector and the driving feature vector includes: using the fitting module to determine the change amount of each feature point based on the original feature vector and the driving feature vector, obtaining an initial feature change vector. The original feature vector includes multiple feature points and an initial value for each feature point, and the driving feature vector includes multiple feature points and a target value for each feature point. The eye redirection module updates the target values ​​of all eye feature points, and the mouth redirection module updates the target values ​​of all mouth feature points. Based on the updated target values ​​of all eye feature points and all mouth feature points, the initial feature change vector is updated to obtain the feature change vector.

[0009] Furthermore, the step of updating the target values ​​of all eye feature points using an eye retargeting module includes: obtaining the target values ​​of all eye feature points from the driving feature vector; and updating the target values ​​of all eye feature points based on the driving eye opening coefficient, wherein the driving eye opening coefficient is a module coefficient obtained by training the eye retargeting module based on eye training data, and the eye training data includes at least: multiple historical eye action videos and the range of eye action changes in each historical eye action video.

[0010] Furthermore, the step of updating the target values ​​of all mouth feature points using a mouth redirection module includes: obtaining the target values ​​of all mouth feature points from the driving feature vector; and updating the target values ​​of all mouth feature points based on the driving mouth opening coefficient, wherein the driving mouth opening coefficient is a module coefficient obtained by training the mouth redirection module based on mouth training data, and the mouth training data includes at least: multiple historical mouth action videos and the range of mouth action changes in each historical mouth action video.

[0011] Furthermore, the step of fusing the audio feature vector and the driving feature vector to obtain the fused feature vector includes: aligning the audio feature vector and the driving feature vector to obtain aligned audio feature vector and driving feature vector; and performing a linear transformation on the aligned audio feature vector and driving feature vector to obtain the fused feature vector.

[0012] Furthermore, before inputting the original feature vector, feature change vector, and fused feature vector into the preset image decoder to obtain the target digital human, the process includes: acquiring multiple historical reference images and multiple historical videos; establishing a correlation between each historical reference image and each historical video for each historical reference image; annotating the historical reference images and historical videos with correlation to obtain annotated images, wherein the annotated images are dynamic images with consistent facial expressions and actions, and consistent actions and sounds; and training the preset image decoder using multiple historical reference images, multiple historical videos, and the annotated images corresponding to the historical reference images and historical videos with correlation to obtain the trained preset image decoder.

[0013] Furthermore, before extracting features from the reference image and the driving video to obtain the original feature vector of the reference image and the driving feature vector of the driving video, the process also includes: extracting each video frame from the driving video; and adjusting the proportion of people in each video frame based on the proportion of people in the reference image.

[0014] According to another aspect of the present invention, a digital human generation apparatus is also provided, comprising: a receiving unit, configured to receive a reference image and a target video, and to split the target video into a driving video and audio; an extraction unit, configured to extract features from the reference image and the driving video to obtain an original feature vector of the reference image and a driving feature vector of the driving video, and to extract features from the audio to obtain an audio feature vector of the audio; a determining unit, configured to determine a feature change vector based on the original feature vector and the driving feature vector using a preset action driving network; a fusion unit, configured to fuse the audio feature vector and the driving feature vector to obtain a fused feature vector; and an input unit, configured to input the original feature vector, the feature change vector, and the fused feature vector into a preset image decoder to obtain a target digital human.

[0015] Furthermore, the preset action-driven network includes at least a fitting module, an eye redirection module, and a mouth redirection module. The determining unit includes: a first determining module, used to determine the change amount of each feature point based on the original feature vector and the driving feature vector using the fitting module to obtain an initial feature change vector, wherein the original feature vector includes: multiple feature points and the initial value of each feature point, and the driving feature vector includes: multiple feature points and the target value of each feature point; a first updating module, used to update the target values ​​of all eye feature points using the eye redirection module, and used to update the target values ​​of all mouth feature points using the mouth redirection module; and a second updating module, used to update the initial feature change vector based on the updated target values ​​of all eye feature points and the updated target values ​​of all mouth feature points to obtain a feature change vector.

[0016] Furthermore, the first update module includes: a first acquisition submodule, used to acquire the target values ​​of all eye feature points from the driving feature vector; and a first update submodule, used to update the target values ​​of all eye feature points based on the driving eye opening coefficient, wherein the driving eye opening coefficient is a module coefficient obtained by training the eye repositioning module based on eye training data, and the eye training data includes at least: multiple historical eye action videos and the range of eye action changes in each historical eye action video.

[0017] Furthermore, the first update module also includes: a second acquisition submodule, used to acquire the target values ​​of all mouth feature points from the driving feature vector; and a second update submodule, used to update the target values ​​of all mouth feature points based on the driving mouth opening coefficient, wherein the driving mouth opening coefficient is a module coefficient obtained by training the mouth repositioning module based on mouth training data, and the mouth training data includes at least: multiple historical mouth action videos and the range of mouth action changes in each historical mouth action video.

[0018] Furthermore, the fusion unit includes: a first alignment module, used to align the audio feature vector and the driving feature vector to obtain aligned audio feature vector and driving feature vector; and a first linear module, used to perform linear transformation on the aligned audio feature vector and driving feature vector to obtain fused feature vector.

[0019] Furthermore, the generation device also includes: a first acquisition module, used to acquire multiple historical reference images and multiple historical videos before inputting the original feature vector, feature change vector, and fused feature vector into a preset image decoder to obtain the target digital human; a first establishment module, used to establish the association relationship between each historical reference image and each historical video; a first annotation module, used to annotate the historical reference images and historical videos with the association relationship to obtain annotated images, wherein the annotated images are dynamic images with consistent expressions and actions and consistent actions and sounds; and a first training module, used to train the preset image decoder using multiple historical reference images, multiple historical videos, and the annotated images corresponding to the historical reference images and historical videos with the association relationship to obtain the trained preset image decoder.

[0020] Furthermore, the generating apparatus also includes: a first extraction module, used to extract each video frame in the driving video before performing feature extraction on the reference image and the driving video to obtain the original feature vector of the reference image and the driving feature vector of the driving video; and a first adjustment module, used to adjust the proportion of people in each video frame based on the proportion of people in the reference image.

[0021] According to another aspect of the present invention, a computer program product is also provided, including a non-volatile computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-described methods for generating a digital human.

[0022] According to another aspect of the present invention, an electronic device is also provided, including one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement any of the above-described methods for generating a digital human.

[0023] In this invention, a reference image and a target video are received, and the target video is split into driving video and audio. Feature extraction is performed on the reference image and the driving video to obtain the original feature vector of the reference image and the driving feature vector of the driving video. Feature extraction is also performed on the audio to obtain the audio feature vector. Based on the original feature vector and the driving feature vector, a preset action driving network is used to determine the feature change vector. The audio feature vector and the driving feature vector are fused to obtain the fused feature vector. The original feature vector, the feature change vector, and the fused feature vector are input into a preset image decoder to obtain the target digital human. This solves the technical problem of poor matching degree of digital human in expression, action and audio in related technologies.

[0024] In this invention, feature extraction is performed on reference images and driving videos, combined with audio feature processing, and a preset action-driven network is used to determine the feature change vector. Then, a fused feature vector is obtained by fusing audio and video features. Finally, a preset image decoder is used to generate the target digital human, ensuring the consistency between facial expressions and actions, and between actions and audio. This reduces the generation cost, improves the naturalness and flexibility of the digital human, and realizes end-to-end digital human generation based on images. The generated digital human can achieve perfect matching of facial expressions and actions, and can make smooth facial expressions and actions while achieving lip-sync. Attached Figure Description

[0025] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings:

[0026] Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for generating digital humans is shown.

[0027] Figure 2 This is a flowchart of the digital human generation method according to Embodiment 1 of the present invention;

[0028] Figure 3 This is a schematic diagram of an optional image-based end-to-end digital human generation according to an embodiment of the present invention;

[0029] Figure 4 This is a schematic diagram of an optional digital human generation device according to an embodiment of the present invention;

[0030] Figure 5 This is a structural block diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0031] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0032] It should be noted that the terms "first," "second," etc., used in this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0033] It should be noted that all relevant information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, and displayed data) collected and involved in this invention are information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of this data comply with the relevant laws, regulations, and standards of the relevant regions, necessary confidentiality measures have been taken, and it does not violate public order and good morals. Corresponding operation entry points are provided for users to choose to authorize or refuse. For example, this system has an interface with relevant users or organizations. Before obtaining relevant information, a request to obtain the information needs to be sent to the aforementioned user or organization through the interface. The relevant information is obtained only after receiving consent from the aforementioned user or organization. If the user chooses to refuse, the process proceeds to an expert decision-making process.

[0034] In this invention, to address the issues of facial expression and motion matching, as well as inconsistencies between lip movements and audio, a motion-driven network and a lip-syncing network are employed for digital human generation. The network's inputs are a reference image, driving video, and audio. Then, through an encoder, feature extraction, feature fusion, and a decoder, lip-syncing and facial expression / motion consistency are achieved in the digital human generation process.

[0035] The present invention will now be described in detail with reference to various embodiments.

[0036] Example 1

[0037] According to an embodiment of this application, an embodiment of a method for generating a digital human is also provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0038] The method embodiment provided in Embodiment 1 of this application can be executed on a mobile terminal, computer terminal, or similar computing device. Figure 1 A hardware structure block diagram of a computer terminal (or mobile device) for implementing a method for generating digital humans is shown. Figure 1 As shown, computer terminal 10 (or mobile device) may include one or more ( Figure 1 The processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions may also be included. In addition, it may include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera, wherein the network interface can be connected to wired and / or wireless networks. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.

[0039] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).

[0040] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the digital human generation method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the aforementioned digital human generation method. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0041] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.

[0042] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).

[0043] Under the aforementioned operating environment, this application provides the following: Figure 2 The method for generating digital humans is shown. Figure 2 This is a flowchart of the digital human generation method according to Embodiment 1 of the present invention, as follows: Figure 2 As shown, the method includes the following steps:

[0044] Step S201: Receive the reference image and the target video, and split the target video into driving video and audio.

[0045] In this embodiment of the invention, a reference image uploaded by the user can be received first, and a target video selected by the user (e.g., selected from a pre-set video library) or uploaded can be received. Then, the target video is split into two parts: driving video and audio.

[0046] Here, the reference image refers to the original static image provided by the user that the digital human is expected to imitate; it can be a facial image of a specific person, and this image will serve as the basic appearance template for the generated digital human. The target video refers to video footage containing the actions and expressions the user wants the digital human to imitate, including dynamically changing facial expressions and movements. The driving video is the video portion separated from the target video, used to drive the digital human's actions and facial expression changes. The audio is the audio portion separated from the target video, used to drive the digital human's speech and lip-sync movements.

[0047] Step S202: Perform feature extraction on the reference image and the driving video to obtain the original feature vector of the reference image and the driving feature vector of the driving video, and perform feature extraction on the audio to obtain the audio feature vector.

[0048] In this embodiment of the invention, a feature extractor can be used to extract features such as facial expressions, head rotation, translation, and feature point deformation of a person in a reference image and a driving video, to obtain the original feature vector of the reference image (i.e., a vector composed of a set of feature points, each feature point corresponding to a feature value) and the driving feature vector of the driving video (i.e., a vector composed of a set of feature points, each feature point corresponding to a feature value). Furthermore, a pre-trained speech recognition encoder can be used to extract features from a sequence of audio segments (i.e., audio) to obtain the audio feature vector.

[0049] For example, deep learning models, such as convolutional neural networks, are used to extract visual features from images and videos, and recurrent neural networks or specific audio feature extractors are used to extract speech features from audio.

[0050] Here, the raw feature vector extracted from the reference image includes multidimensional numerical representations of facial expressions, head pose, translation, scaling, and other information. The driving feature vector extracted from the driving video also includes the aforementioned facial and head motion information, but also contains the motion sequence that changes over time. The audio feature vector extracted from the audio is a numerical vector used to represent information such as speech content, intonation, rhythm, and phonemes.

[0051] Step S203: Based on the original feature vector and the driving feature vector, a preset action driving network is used to determine the feature change vector.

[0052] In this embodiment of the invention, the transformation from the reference image to the target expression and action (i.e., feature change vector) can be calculated through a preset action-driven network based on the original feature vector and the driving feature vector, so as to provide dynamic guidance for digital human generation.

[0053] Here, the action-driven network is a pre-trained deep learning model used to analyze and predict the differences in facial expressions and movements in the driving video relative to the reference image, i.e., feature change vectors. Feature change vectors represent the differences in expressions and movements between the driving video and the reference image, and are key data driving changes in the digital human's expressions and movements.

[0054] Step S204: Fuse the audio feature vector and the driving feature vector to obtain the fused feature vector.

[0055] In this embodiment of the invention, a multi-scale fusion feature network can be used to fuse audio features (i.e., audio feature vectors) and visual features (i.e. driving feature vectors) at different scales to obtain a fused feature vector.

[0056] Here, the fusion feature vector refers to combining the audio feature vector and the driving feature vector to form a comprehensive vector containing audio and video information. This vector is used to guide lip-sync and facial expression matching during digital human generation. This ensures the consistency of audio and video features, enabling the generated digital human to achieve lip-sync and improve the naturalness of the digital human.

[0057] Step S205: Input the original feature vector, feature change vector, and fused feature vector into the preset image decoder to obtain the target digital human.

[0058] In this embodiment of the invention, a pre-trained preset image decoder can be used to receive multiple feature vectors (i.e., original feature vectors, feature variation vectors, and fused feature vectors) to generate a target digital human. This allows for accurate reflection of changes in the original reference image and the driving video, while maintaining synchronization with the audio, thereby generating a digital human with consistent facial expressions and lip-sync.

[0059] Here, the image decoder is a pre-trained neural network model used to receive feature vectors and generate images of the corresponding digital human faces, realizing the transformation from feature vectors to the visual appearance of the digital human.

[0060] In summary, by extracting features from reference images and driving videos, combining audio feature processing, determining feature change vectors using a pre-defined action-driven network, obtaining fused feature vectors by fusing audio and video features, and finally generating the target digital human using a pre-defined image decoder, the consistency between facial expressions and movements, and between movements and audio is ensured. This reduces generation costs, improves the naturalness and flexibility of the digital human, and achieves end-to-end digital human generation based on images. The generated digital human can achieve perfect matching of facial expressions and movements, and can make smooth facial expressions and movements while achieving lip-sync.

[0061] Optionally, the preset action-driven network includes at least: a fitting module, an eye redirection module, and a mouth redirection module. To improve the accuracy of determining the feature change vector, in the digital human generation method provided in Embodiment 1 of this application, based on the original feature vector and the driving feature vector, the fitting module determines the change amount of each feature point to obtain an initial feature change vector. The original feature vector includes multiple feature points and the initial value of each feature point, and the driving feature vector includes multiple feature points and the target value of each feature point. The eye redirection module updates the target values ​​of all eye feature points, and the mouth redirection module updates the target values ​​of all mouth feature points. Based on the updated target values ​​of all eye feature points and the updated target values ​​of all mouth feature points, the initial feature change vector is updated to obtain the feature change vector.

[0062] In this embodiment of the invention, the preset motion-driven network includes: a fitting module, an eye redirection module, and a mouth redirection module. In fitting and redirection, the input to the fitting module is the facial expression and motion features of the reference image (i.e., the original feature vector), and the facial expression and motion features of the driving video (driving feature vector), enabling it to estimate the amount of feature change based on the input data. The input to the eye redirection module is the eye keypoint features of the reference image, the eye opening condition (feature values ​​of the eye keypoint features in the initial feature change vector), and the driving eye opening coefficient (parameters after training the eye redirection module), thereby estimating the deformation change of the eye keypoints. The eye opening condition represents the eye opening ratio; a larger value indicates a greater degree of eye opening. The input to the mouth redirection module is the mouth keypoint features of the reference image, the mouth opening condition (feature values ​​of the mouth keypoint features in the initial feature change vector), and the driving mouth opening coefficient (parameters after training the mouth redirection module), thereby estimating the deformation change of the mouth keypoints. Next, the driving keypoint features (eye keypoint features and mouth keypoint features) are updated by the deformation changes corresponding to the eyes and mouth, respectively. Then, during the expression and motion generation process, the encoded features of the reference image (i.e., the original feature vector) and the changes in the expression features, eye and mouth features of the driving video are input into the image decoder to generate a consistent image.

[0063] Specifically, a fitting module can be used to refine the changes in feature points to obtain an initial feature change vector. That is, the fitting module compares the initial and target values ​​of the same feature points in the original feature vector and the driving feature vector to accurately calculate the change in each feature point under the target expression. For example, by fine-tuning the feature points, the driving expression is ensured to fit naturally on the reference image, laying the foundation for generating realistic digital human expressions, thus obtaining the initial feature change vector. However, this initial change vector does not fully consider the special dynamic requirements of the eyes and mouth, so further optimization is needed.

[0064] Here, the original feature vector contains multiple feature points and their initial values ​​from the reference image. These feature points cover key areas of facial expression, such as the position and shape information of eyebrows, eyes, nose, and mouth. The driving feature vector comes from the driving video and contains the aforementioned feature points, but records the target values ​​of these feature points between video frames, reflecting the expressions and movements that change with the video content.

[0065] Then, the eye retargeting module receives features and conditions of key eye points from the reference image, along with a driving eye opening coefficient, to update the target values ​​of all eye feature points. Here, the conditional tuple defines the proportion of eye opening under different conditions (such as blinking, staring, etc.) (which can be obtained from the initial feature change vector), and the coefficient is used to adjust the intensity and subtlety of the eye dynamics. By calculating the target value changes of the driving eye feature points in the video, the eye retargeting module updates the target values ​​of all eye feature points, ensuring the accuracy and smoothness of the generated digitizer's eye expressions.

[0066] Furthermore, the mouth redirection module receives the features and conditions of key mouth points in the reference image, along with a driving mouth opening coefficient, to update the target values ​​of all mouth feature points, thereby achieving a higher level of lip-sound synchronization. The driving mouth opening coefficient can be adjusted according to the intensity and rhythm of the audio, ensuring that mouth movements match the sound and enhancing the realism of the audiovisual experience.

[0067] In this embodiment of the invention, the initial feature change vector needs to be adjusted accordingly based on the updated target values ​​of eye and mouth feature points. At this point, the updated target values ​​of the eye and mouth feature points can be integrated, the change in each feature point can be re-evaluated and optimized, and finally the initial feature change vector is updated to form a feature change vector. This feature change vector comprehensively reflects the precise guidance from the reference image to the changes driving video facial expressions, including more nuanced eye and mouth facial dynamics, thus improving the quality of digital human generation.

[0068] In this embodiment, the coherence and naturalness of facial expressions during the digital human generation process are enhanced, especially in the eyes and mouth, two key areas that best reflect the richness and realism of expressions. Not only are the facial expressions more closely matched with the actions in the driving video, but lip-syncing is also achieved, improving the interactivity and immersion between the digital human and the audience.

[0069] In order to accurately update the target values ​​of all eye feature points, in the digital human generation method provided in Embodiment 1 of this application, the target values ​​of all eye feature points are obtained from the driving feature vector; the target values ​​of all eye feature points are updated based on the driving eye opening coefficient, wherein the driving eye opening coefficient is a module coefficient obtained by training the eye repositioning module based on eye training data, and the eye training data includes at least: multiple historical eye action videos and the range of eye action changes in each historical eye action video.

[0070] In this embodiment of the invention, target values ​​for all eye feature points can be extracted from the driving feature vector. This driving feature vector contains dynamic information about the facial expressions and movements of a person in the driving video, and the target values ​​of its eye feature points reflect details such as eye opening and closing, blinking frequency, and eye movement in the driving video. This ensures that the generated digital human can accurately mimic the changes in the eyes and facial expressions of a person in the video.

[0071] Furthermore, to more precisely control the dynamic changes in the digital human's eye expressions, an eye redirection module was introduced. This module is trained based on a set of eye training data, which includes at least several historical eye movement videos and the range of eye movement variations within each video. The range of eye movement variations was obtained by meticulously annotating the opening and closing, direction of movement, and speed of the eyes in the videos using specialized eye movement annotation tools, providing learning samples for the eye redirection module. Through training, the eye redirection module learns to associate eye movements in videos with specific ranges of variation and can extract the driving eye opening coefficient, which reflects the intensity and detail of the eye movements in the video. In practical applications, the driving eye opening coefficient is used to adjust and optimize the target values ​​of eye feature points, ensuring that the generated digital human's eye expressions not only match the video but also naturally express the changes in the subject's gaze and emotions.

[0072] Subsequently, based on the driving eye opening coefficient, the target values ​​of all eye feature points can be updated. This update process can be implemented through an eye retargeting module, which dynamically adjusts the position, shape, and trajectory of each feature point during the generation of the digital human based on the coefficient, achieving more refined and natural changes in the digital human's eye expressions. For example, a higher driving eye opening coefficient might indicate to the module the need to increase the frequency of blinking or amplify the degree of eye opening and closing, making the generated digital human's eyes more vivid and more likely to attract the viewer's attention.

[0073] In this embodiment, the accuracy and naturalness of the eye expressions generated by the digital human are improved. The introduction of the eye redirection module not only enhances the ability to imitate the eye movements of people in videos, but also allows for flexible adjustment of the intensity and subtlety of eye expressions according to the video content, achieving a high-fidelity synchronization effect between expressions and movements.

[0074] In order to accurately update the target values ​​of all mouth feature points, in the digital human generation method provided in Embodiment 1 of this application, the target values ​​of all mouth feature points are obtained from the driving feature vector; the target values ​​of all mouth feature points are updated based on the driving mouth opening coefficient, wherein the driving mouth opening coefficient is a module coefficient obtained by training a mouth repositioning module based on mouth training data, and the mouth training data includes at least: multiple historical mouth action videos and the range of mouth action changes in each historical mouth action video.

[0075] In this embodiment of the invention, target values ​​for all mouth feature points can be extracted from the driving feature vector. These target values ​​are based on information about the dynamic changes of a person's mouth in the driving video, including the degree of mouth opening and closing, lip shape changes, and the positions of key points related to pronunciation.

[0076] To enable the generated digital human to more naturally mimic the mouth movements of people in videos, especially achieving lip-sync, a mouth repositioning module was introduced and trained using mouth training data. This training data includes at least several historical mouth movement videos, along with detailed records of the range of mouth movement variations for each video. The variation range data includes changes in mouth shape from a static position to different degrees of opening and closing, as well as lip shape variations in different contexts, providing rich learning samples. Subsequently, through deep learning, the mouth repositioning module learned how to map the dynamic changes in the mouth in the video to target values, thereby extracting a mouth opening coefficient. This coefficient quantifies the intensity of mouth movement changes. This coefficient is the result of the module's self-adjustment and optimization during training, guiding how to more accurately update the target values ​​of mouth feature points.

[0077] In this embodiment of the invention, after obtaining the optimized driving mouth opening coefficient, the target values ​​of all mouth feature points can be updated. That is, the target values ​​are adjusted by driving the mouth opening coefficient to ensure that the generated digital human can make accurately matching mouth movements according to changes in audio content. For example, if the audio features indicate laughter or a specific pronunciation, the driving mouth opening coefficient will guide the module to make corresponding mouth opening degrees and lip shape changes, thereby achieving natural synchronization of sound and lips and realistic expression.

[0078] In this embodiment, the generated digital human can precisely control its mouth movements, achieving perfect lip-verb synchronization with the audio content. By training the mouth redirection module and updating the target values ​​of mouth feature points using a driving mouth opening coefficient, not only is the naturalness of the digital human enhanced, but its user experience and communication efficiency are also improved in application scenarios such as virtual anchors, online education, and enterprise AI customer service, especially in the communication and service fields of financial institutions.

[0079] In order to accurately obtain the fused feature vector, in the digital human generation method provided in Embodiment 1 of this application, the audio feature vector and the driving feature vector are aligned to obtain the aligned audio feature vector and the driving feature vector; the aligned audio feature vector and the driving feature vector are linearly transformed to obtain the fused feature vector.

[0080] In this embodiment of the invention, the audio feature vector and the driving feature vector can be precisely aligned in the time dimension. This alignment operation is performed by comparing the timestamps of the two vectors, matching each frame feature in the audio feature vector with the facial expression and action features at the corresponding time points in the driving feature vector, to ensure a one-to-one correspondence between the audio content and the time points of facial expression changes. For example, cross-correlation analysis or other time series matching algorithms can be used to find the best correspondence between the feature changes in the two vectors, ensuring that the audio features and facial movements are strictly synchronized on the time axis.

[0081] Here, the audio feature vector, extracted from the audio segment by the audio encoder, contains information such as pitch, rhythm, and timbre of the speech, and is key to achieving lip-sync. The driving feature vector is facial dynamic information extracted from the driving video, including features such as expression changes and head movements, and is the foundation for driving digital human animation.

[0082] In this embodiment of the invention, after alignment, the aligned audio feature vector and driving feature vector can be linearly transformed to unify the information of the two modalities into a common space for subsequent fusion operations. Linear transformation typically involves the transformation and standardization of the feature vector space, ensuring that the two vectors are represented on the same scale, facilitating subsequent processing and analysis. For example, a linear transformation matrix can be used to achieve the linear transformation of the two vectors. This matrix is ​​optimized during the alignment process to minimize the difference between the two vectors at the aligned time point, thereby obtaining the optimal representation of the aligned audio feature vector and driving feature vector.

[0083] Next, a fused feature vector is generated by weighted summing of the two aligned and transformed vectors. This fused feature vector contains integrated information from the audio and driving video, reflecting both the rhythm and intonation of the audio content and the details of facial expression changes in the driving video. The fused feature vector will serve as input to the subsequent digital human image decoder, guiding the generation of lip movements in the digital human to achieve lip-sync.

[0084] In this embodiment, a high degree of synchronization and consistency between audio features and driving video features is ensured, and the generated fused feature vector provides a precise driving signal for the generation of the digital human. This improves the performance quality of the digital human in voice interaction scenarios, enabling more natural and fluid lip movements and facial expressions, whether in automated online exhibition hall navigation or as a virtual announcer in financial institution communications, thus enhancing user immersion.

[0085] To accurately train the image decoder, in the digital human generation method provided in Embodiment 1 of this application, before inputting the original feature vector, feature change vector, and fused feature vector into the preset image decoder to obtain the target digital human, multiple historical reference images and multiple historical videos are acquired; for each historical reference image, an association relationship is established between the historical reference image and each historical video; the historical reference images and historical videos with association relationships are annotated to obtain annotated images, wherein the annotated images are dynamic images with consistent expressions and actions and consistent actions and sounds; the preset image decoder is trained using multiple historical reference images, multiple historical videos, and the annotated images corresponding to the historical reference images and historical videos with association relationships to obtain the trained preset image decoder.

[0086] In this embodiment of the invention, multiple historical reference images and multiple historical videos can be acquired, forming the basis of the training dataset. Historical reference images typically include different angles and expressions of the target digital human prototype, while historical videos are video clips containing mouth movements, eye movements, and changes in facial expressions. This data covers a wide range of facial expressions and movements, providing ample information for subsequent deep learning model training.

[0087] In this embodiment of the invention, for each historical reference image, a correlation can be established between the historical reference image and each historical video. For example, this can be achieved by analyzing the similar features of people in the image and video, such as facial structure and hair color, ensuring that each image can find a video segment that matches its facial expressions and movements. Thus, during training, the model can learn the consistency between specific images and dynamic changes in videos, i.e., how to generate dynamic images consistent with the expressions and movements in videos starting from static images.

[0088] Then, related historical reference images and videos are annotated to generate annotated images that match facial expressions and movements, as well as movements and sounds. For example, detailed annotations are made of facial expressions, head postures, mouth and eye movements of people in the video, while audio content is also time-series aligned to ensure that the dynamic changes of each video frame match the pronunciation time of the corresponding audio segment. The generation of annotated images relies on the annotation results, and through a series of image processing techniques, such as image synthesis and deformation, the historical reference images can display facial expressions and movements consistent with those of the people in the video, as well as mouth movements synchronized with the audio content.

[0089] Next, the preset image decoder can be trained using multiple historical reference images, multiple historical videos, and corresponding labeled images. The image decoder is the core component for generating digital humans, used to convert encoded feature vectors into video frames. Through this training process, the model learns how to generate dynamic images with consistent facial expressions, movements, and lip-phonetic synchronization based on static reference images and dynamic driving features. The trained image decoder will be able to generate high-quality digital humans based on the input feature vectors, achieving natural transitions in facial expressions and precise lip-phonetic synchronization.

[0090] In this embodiment, a training dataset containing annotations on the consistency between facial expressions and actions, and the consistency between actions and voices was constructed. The image decoder was trained using deep learning, which improved the naturalness and accuracy of the generated digital human in terms of facial expressions, head movements, and mouth movements, and enabled lip-syncing with the audio content.

[0091] To ensure that the facial proportions of people in images and videos are consistent, in the digital human generation method provided in Embodiment 1 of this application, before extracting features from the reference image and the driving video to obtain the original feature vector of the reference image and the driving feature vector of the driving video, each video frame in the driving video is extracted; based on the proportion of people in the reference image, the proportion of people in each video frame is adjusted.

[0092] In this embodiment of the invention, the driving video can be preprocessed, that is, each video frame in the driving video can be extracted to decompose the video into a series of static images. Each video frame contains the facial expressions and action information of the people in the video at a specific point in time. For example, each video frame in the driving video can be extracted through video processing techniques such as frame rate selection and image cropping.

[0093] Then, based on the proportions of the person in the reference image, the proportions of the person in each video frame are adjusted to ensure that the dynamic changes driving the video are visually consistent with the digital human prototype in the reference image, avoiding visual inconsistencies caused by different proportions. For example, operations such as image scaling, cropping, and deformation are typically performed using image processing algorithms such as affine transformation and perspective transformation. Based on specific feature points in the reference image (such as eyes, nose tip, and corners of the mouth), the person image in the video frame is scaled and deformed until its proportions match those in the reference image.

[0094] In this embodiment, visual consistency between the character's movements in the driving video and the digital human prototype is ensured, especially the consistency of the character's proportions. This improves the visual effect of the generated digital human in different scenarios, enabling it to accurately mimic the expressions and movements of the character in the video while maintaining a high degree of similarity to the original character in applications such as virtual anchors, online education, and AI customer service, thus enhancing user immersion. Furthermore, consistent proportions reduce visual distortion caused by proportion mismatches in the generated video, improving the production efficiency and quality of digital human videos.

[0095] Figure 3 This is a schematic diagram of an optional image-based end-to-end digital human generation according to an embodiment of the present invention, such as... Figure 3 As shown, the system can receive reference images, driving videos, and audio clips. An image editor then processes the reference image and driving video to extract facial and motion features. Following this, a fitting module, a mouth redirection module, and an eye redirection module determine the changes in each feature point on the reference image. Simultaneously, an audio encoder processes the audio clips to extract audio features, fusing the features extracted from the driving video with the audio features to obtain fused features. Finally, an image decoder processes the features extracted from the reference image, the fused features, and the changes to generate a dynamic digital human with consistent facial expressions and lip-sync.

[0096] The digital human generation method provided in this application extracts features from reference images and driving videos, processes audio features, determines feature change vectors using a preset action-driven network, obtains fused feature vectors by fusing audio and video features, and finally generates a target digital human using a preset image decoder. This ensures consistency between expressions and actions, and between actions and audio, reduces generation costs, improves the naturalness and flexibility of the digital human, realizes end-to-end digital human generation based on images, and enables the generated digital human to achieve perfect matching of expressions and actions, as well as to make smooth facial expressions and actions while achieving lip-sync.

[0097] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0098] Example 2

[0099] This application also provides a digital human generation apparatus. It should be noted that the digital human generation apparatus of this application can be used to execute the digital human generation method provided in this application. The digital human generation apparatus provided in this application will be described below.

[0100] According to an embodiment of this application, an apparatus for implementing the above-described digital human generation method is also provided. Figure 4 This is a schematic diagram of an optional digital human generation device according to an embodiment of the present invention, such as... Figure 4 As shown, the generating device may include: a receiving unit 40, an extraction unit 41, a determining unit 42, a fusion unit 43, and an input unit 44.

[0101] The receiving unit 40 is used to receive the reference image and the target video, and to split the target video into driving video and audio.

[0102] The extraction unit 41 is used to extract features from the reference image and the driving video to obtain the original feature vector of the reference image and the driving feature vector of the driving video, and to extract features from the audio to obtain the audio feature vector.

[0103] The determination unit 42 is used to determine the feature change vector based on the original feature vector and the driving feature vector, and to use a preset action driving network to determine the feature change vector.

[0104] Fusion unit 43 is used to fuse audio feature vectors and driving feature vectors to obtain fused feature vectors;

[0105] Input unit 44 is used to input the original feature vector, feature change vector and fused feature vector into a preset image decoder to obtain the target digital human.

[0106] The digital human generation apparatus provided in this application extracts features from reference images and driving videos, processes audio features, determines feature change vectors using a preset action-driven network, obtains fused feature vectors by fusing audio and video features, and finally generates a target digital human using a preset image decoder. This ensures consistency between expressions and actions, and between actions and audio, reduces generation costs, improves the naturalness and flexibility of the digital human, realizes end-to-end digital human generation based on images, and enables the generated digital human to achieve perfect matching of expressions and actions, as well as to make smooth facial expressions and actions while achieving lip-sync.

[0107] Optionally, the preset action-driven network includes at least a fitting module, an eye redirection module, and a mouth redirection module. The determining unit includes: a first determining module, used to determine the change amount of each feature point based on the original feature vector and the driving feature vector using the fitting module to obtain an initial feature change vector, wherein the original feature vector includes: multiple feature points and the initial value of each feature point, and the driving feature vector includes: multiple feature points and the target value of each feature point; a first updating module, used to update the target values ​​of all eye feature points using the eye redirection module, and used to update the target values ​​of all mouth feature points using the mouth redirection module; and a second updating module, used to update the initial feature change vector based on the updated target values ​​of all eye feature points and the updated target values ​​of all mouth feature points to obtain a feature change vector.

[0108] Optionally, the first update module includes: a first acquisition submodule, used to acquire the target values ​​of all eye feature points from the driving feature vector; and a first update submodule, used to update the target values ​​of all eye feature points based on the driving eye opening coefficient, wherein the driving eye opening coefficient is a module coefficient obtained by training the eye repositioning module based on eye training data, and the eye training data includes at least: multiple historical eye movement videos and the range of eye movement changes in each historical eye movement video.

[0109] Optionally, the first update module further includes: a second acquisition submodule, used to acquire the target values ​​of all mouth feature points from the driving feature vector; and a second update submodule, used to update the target values ​​of all mouth feature points based on the driving mouth opening coefficient, wherein the driving mouth opening coefficient is a module coefficient obtained by training the mouth repositioning module based on mouth training data, and the mouth training data includes at least: multiple historical mouth action videos and the range of mouth action changes in each historical mouth action video.

[0110] Optionally, the fusion unit includes: a first alignment module for aligning the audio feature vector and the driving feature vector to obtain aligned audio feature vector and driving feature vector; and a first linear module for linearly transforming the aligned audio feature vector and driving feature vector to obtain fused feature vector.

[0111] Optionally, the generation device further includes: a first acquisition module, used to acquire multiple historical reference images and multiple historical videos before inputting the original feature vector, feature change vector, and fused feature vector into a preset image decoder to obtain the target digital human; a first establishment module, used to establish a correlation between each historical reference image and each historical video; a first annotation module, used to annotate the historical reference images and historical videos with correlation to obtain annotated images, wherein the annotated images are dynamic images with consistent expressions and actions and consistent actions and sounds; and a first training module, used to train the preset image decoder using multiple historical reference images, multiple historical videos, and the annotated images corresponding to the historical reference images and historical videos with correlation to obtain the trained preset image decoder.

[0112] Optionally, the generating apparatus further includes: a first extraction module, configured to extract each video frame in the driving video before performing feature extraction on the reference image and the driving video to obtain the original feature vector of the reference image and the driving feature vector of the driving video; and a first adjustment module, configured to adjust the proportion of people in each video frame based on the proportion of people in the reference image.

[0113] The aforementioned generating apparatus may further include a processor and a memory. The receiving unit 40, the extraction unit 41, the determining unit 42, the fusion unit 43, the input unit 44, etc., are all stored in the memory as program units, and the processor executes the aforementioned program units stored in the memory to realize the corresponding functions.

[0114] The aforementioned processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured, and by adjusting kernel parameters, the original feature vector, feature transformation vector, and fused feature vector are input to a preset image decoder to obtain the target digital human.

[0115] The aforementioned memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.

[0116] It should be noted that the receiving unit 40, extraction unit 41, determining unit 42, fusion unit 43, and input unit 44 correspond to steps S201 to S205 in Embodiment 1. The instances and application scenarios implemented by the above units and corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1. It should be noted that the above units can be hardware or software components stored in memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above units can also be part of a device and run in the computer terminal 10 provided in Embodiment 1.

[0117] Example 3

[0118] Embodiments of this application may provide an electronic device. Figure 5 This is a structural block diagram of an electronic device according to an embodiment of the present invention. Figure 5 As shown, the electronic device may include: one or more ( Figure 5 (Only one is shown) processor 502, memory 504, memory controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.

[0119] The memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the digital human generation method and apparatus in this application embodiment. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, thereby realizing the aforementioned digital human generation method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0120] The processor can access information and applications stored in the memory via a transmission device to perform the following steps: receiving a reference image and a target video, and splitting the target video into driving video and audio; extracting features from the reference image and driving video to obtain the original feature vector of the reference image and the driving feature vector of the driving video, and extracting features from the audio to obtain the audio feature vector; determining the feature change vector using a preset action driving network based on the original feature vector and the driving feature vector; fusing the audio feature vector and the driving feature vector to obtain the fused feature vector; and inputting the original feature vector, the feature change vector, and the fused feature vector into a preset image decoder to obtain the target digital human.

[0121] The processor can access information and applications stored in memory via a transmission device to execute the following steps: Based on the original feature vector and the driving feature vector, a fitting module is used to determine the amount of change of each feature point to obtain an initial feature change vector. The original feature vector includes multiple feature points and the initial value of each feature point, and the driving feature vector includes multiple feature points and the target value of each feature point. An eye redirection module is used to update the target values ​​of all eye feature points, and a mouth redirection module is used to update the target values ​​of all mouth feature points. Based on the updated target values ​​of all eye feature points and the updated target values ​​of all mouth feature points, the initial feature change vector is updated to obtain the feature change vector.

[0122] The processor can access the information and application stored in the memory via the transmission device to perform the following steps: obtain the target values ​​of all eye feature points from the driving feature vector; update the target values ​​of all eye feature points based on the driving eye opening coefficient, wherein the driving eye opening coefficient is a module coefficient obtained by training an eye repositioning module based on eye training data, and the eye training data includes at least: multiple historical eye movement videos and the range of eye movement changes in each historical eye movement video.

[0123] The processor can access the information and application stored in the memory via the transmission device to perform the following steps: obtain the target values ​​of all mouth feature points from the driving feature vector; update the target values ​​of all mouth feature points based on the driving mouth opening coefficient, wherein the driving mouth opening coefficient is the module coefficient obtained by training the mouth repositioning module based on mouth training data, and the mouth training data includes at least: multiple historical mouth action videos and the range of mouth action changes in each historical mouth action video.

[0124] The processor can access information and applications stored in memory via a transmission device to perform the following steps: aligning the audio feature vector and the driving feature vector to obtain aligned audio feature vector and driving feature vector; and performing a linear transformation on the aligned audio feature vector and driving feature vector to obtain a fused feature vector.

[0125] The processor can access the information and application programs stored in the memory via the transmission device to perform the following steps: acquiring multiple historical reference images and multiple historical videos; establishing a correlation between each historical reference image and each historical video for each historical reference image; annotating the historical reference images and historical videos with correlation to obtain annotated images, wherein the annotated images are dynamic images with consistent facial expressions and actions, and actions and sounds; training a preset image decoder using multiple historical reference images, multiple historical videos, and the annotated images corresponding to the historical reference images and historical videos with correlation to obtain the trained preset image decoder.

[0126] The processor can access information and applications stored in memory via a transmission device to perform the following steps: extract each video frame from the driving video; adjust the proportions of people in each video frame based on the proportions of people in the reference image.

[0127] This application provides a scheme for generating digital humans. By extracting features from reference images and driving videos, combined with audio feature processing, a preset action-driven network is used to determine feature change vectors. Then, a fused feature vector is obtained by fusing audio and video features. Finally, a preset image decoder is used to generate the target digital human. This ensures consistency between facial expressions and movements, and between movements and audio, reducing generation costs and improving the naturalness and flexibility of the digital human. It achieves end-to-end digital human generation based on images, and the generated digital human can achieve perfect matching of facial expressions and movements, as well as smooth facial expressions and movements while achieving lip-sync.

[0128] Those skilled in the art will understand that Figure 5 The structure shown is for illustrative purposes only. Electronic devices can also be terminal devices such as smartphones, tablets, PDAs, and mobile internet devices (MIDs). Figure 5 This does not limit the structure of the aforementioned electronic device. For example, electronic devices may also include components that are more... Figure 5 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 5 The different configurations shown.

[0129] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0130] Example 4

[0131] Embodiments of this application also provide a storage medium. Optionally, in this embodiment, the storage medium can be used to store the program code executed by the digital human generation method provided in Embodiment 1.

[0132] Optionally, in this embodiment, the storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.

[0133] This application also provides a computer program product that, when executed on a data processing device, is suitable for performing steps of a method for generating a digital human.

[0134] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0135] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0136] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0137] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0138] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0139] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0140] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for generating a digital human, characterized by, include: Receive a reference image and a target video, and split the target video into driving video and audio; Feature extraction is performed on the reference image and the driving video to obtain the original feature vector of the reference image and the driving feature vector of the driving video; and feature extraction is performed on the audio to obtain the audio feature vector of the audio. Based on the original feature vector and the driving feature vector, a preset action-driven network is used to determine the feature change vector; The audio feature vector and the driving feature vector are fused to obtain the fused feature vector; The original feature vector, the feature change vector, and the fused feature vector are input into a preset image decoder to obtain the target digital human.

2. The generation method of claim 1, wherein, The preset action-driven network includes at least an adhesion module, an eye redirection module, and a mouth redirection module. The step of determining the feature change vector using the preset action-driven network based on the original feature vector and the driving feature vector includes: Based on the original feature vector and the driving feature vector, the fitting module determines the change amount of each feature point to obtain an initial feature change vector. The original feature vector includes multiple feature points and an initial value for each feature point. The driving feature vector includes multiple feature points and a target value for each feature point. The target values ​​of all eye feature points are updated using the eye redirection module, and the target values ​​of all mouth feature points are updated using the mouth redirection module. Based on the updated target values ​​of all the eye feature points and the updated target values ​​of all the mouth feature points, the initial feature change vector is updated to obtain the feature change vector.

3. The generation method of claim 2, wherein, The step of updating the target values ​​of all eye feature points using the eye retargeting module includes: Obtain the target value of all the eye feature points from the driving feature vector; Based on the driving eye opening coefficient, the target value of all the eye feature points is updated, wherein the driving eye opening coefficient is a module coefficient obtained by training the eye repositioning module based on eye training data, and the eye training data includes at least: multiple historical eye movement videos and the range of eye movement changes in each historical eye movement video.

4. The generation method of claim 2, wherein, The step of updating the target values ​​of all mouth feature points using the mouth retargeting module includes: Obtain the target value of all mouth feature points from the driving feature vector; Based on the driving mouth opening coefficient, the target value of all mouth feature points is updated, wherein the driving mouth opening coefficient is a module coefficient obtained by training the mouth repositioning module based on mouth training data, and the mouth training data includes at least: multiple historical mouth action videos and the range of mouth action variation for each historical mouth action video.

5. The generation method of claim 1, wherein, The step of fusing the audio feature vector and the driving feature vector to obtain the fused feature vector includes: Align the audio feature vector and the driving feature vector to obtain the aligned audio feature vector and driving feature vector; The aligned audio feature vector and the driving feature vector are linearly transformed to obtain the fused feature vector.

6. The generation method according to claim 1, characterized in that, Before inputting the original feature vector, the feature transformation vector, and the fused feature vector into a preset image decoder to obtain the target digital human, the process further includes: Collect multiple historical reference images and multiple historical videos; For each historical reference image, establish the association between the historical reference image and each historical video; The historical reference images and historical videos that are related are annotated to obtain an annotated image, wherein the annotated image is a dynamic image with consistent facial expressions and actions as well as consistent actions and sounds; The preset image decoder is trained by using multiple historical reference images, multiple historical videos, and the labeled images corresponding to the historical reference images and historical videos that are related, to obtain the trained preset image decoder.

7. The generation method according to claim 1, characterized in that, Before performing feature extraction on the reference image and the driving video to obtain the original feature vector of the reference image and the driving feature vector of the driving video, the method further includes: Extract each video frame from the driving video; Based on the proportions of the people in the reference image, the proportions of the people in each video frame are adjusted.

8. A device for generating a digital human, characterized in that, include: A receiving unit is configured to receive a reference image and a target video, and to split the target video into driving video and audio. The extraction unit is used to extract features from the reference image and the driving video to obtain the original feature vector of the reference image and the driving feature vector of the driving video, and to extract features from the audio to obtain the audio feature vector of the audio. The determining unit is used to determine the feature change vector based on the original feature vector and the driving feature vector, using a preset action driving network. A fusion unit is used to fuse the audio feature vector and the driving feature vector to obtain a fused feature vector; The input unit is used to input the original feature vector, the feature change vector, and the fused feature vector into a preset image decoder to obtain the target digital human.

9. A computer program product, characterized in that, The method includes a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the digital human generation method according to any one of claims 1 to 7.

10. An electronic device, characterized in that, It includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the digital human generation method according to any one of claims 1 to 7.