Digital human driving method and device, electronic equipment and storage medium

By inputting target images, lines and emotions into the pre-trained digital human video generation model, the problem of lack of naturalness and low efficiency of digital human video generation in the prior art is solved, and a more natural and emotionally rich video generation effect is achieved.

CN120075551APending Publication Date: 2025-05-30DUKE KUNSHAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510224297.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing digital human video generation technology is difficult to generate natural and realistic dynamic effects and achieve accurate emotional control, resulting in the generated video lacking nature and low efficiency.

Method used

By determining the target image, target lines, and target emotions, these are input into the pre-trained digital human video generation model to generate the target video. This model uses 3D specification key points characteristics, expression parameters and head posture, and combines self-attention mechanisms and cross-attention mechanisms to model inter-frame timing information and fusion to generate more natural and emotionally rich videos.

Benefits of technology

The efficiency of digital human video generation is improved, so that the generated digital humans can express emotions more naturally, and the naturalness and emotional expression of the video are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075551A_ABST
    Figure CN120075551A_ABST
Patent Text Reader

Abstract

The invention discloses a digital human driving method and device, electronic equipment and a storage medium. The method comprises the following steps: determining a target image, a target line and a target emotion; the target image is an image of a digital human; the target lines are lines taught by a digital person; the target emotion is an emotion expressed when a digital person teaches a target line; and inputting the target image, the target line and the target emotion into a pre-trained digital person video generation model to obtain a target video, the target video being the target line taught by the digital person in the target image by adopting the target emotion. According to the technical scheme, the target image, the target line and the target emotion are determined; and inputting the target image, the target line and the target emotion into the pre-trained digital human video generation model to obtain the target video, so that the generation efficiency of the target video is improved, and the generated digital human can better express the emotion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital human video generation, and in particular, to a digital human driving method, device, electronic device, and storage medium. Background Art

[0002] In recent years, creating digital portraits that can talk and exhibit facial expressions, namely "digital humans", has become a very active research direction.

[0003] With the increasing demand for higher-quality and higher-efficiency video generation, digital human technology still faces many problems to be solved urgently, especially in generating natural and realistic dynamic effects and achieving precise emotional control. For example, it is difficult to maintain the consistency of the appearance between frames in the video, or the generated video often loses its authenticity due to overly exaggerated expressions, making the final output lack naturalness. Summary of the Invention

[0004] The present invention provides a digital human driving method, device, electronic device, and storage medium to solve the problem that the generated digital human video lacks naturalness and has low efficiency.

[0005] According to one aspect of the present invention, there is provided a digital human driving method, the method comprising:

[0006] Determine a target image, a target line, and a target emotion; the target image is an image of a digital human; the target line is the line that the digital human is going to speak; the target emotion is the emotion expressed by the digital human when speaking the target line;

[0007] Input the target image, the target line, and the target emotion into a pre-trained digital human video generation model to obtain a target video, where the target video is the digital human in the target image speaking the target line with the target emotion.

[0008] According to another aspect of the present invention, there is provided a digital human driving device, the device comprising:

[0009] A target data acquisition module, configured to determine a target image, a target line, and a target emotion; the target image is an image of a digital human; the target line is the line that the digital human is going to speak; the target emotion is the emotion expressed by the digital human when speaking the target line;

[0010] A target video generation module, configured to input the target image, the target line, and the target emotion into a pre-trained digital human video generation model to obtain a target video, where the target video is the digital human in the target image speaking the target line with the target emotion.

[0011] According to another aspect of the present invention, there is provided an electronic device, the electronic device comprising:

[0012] At least one processor; and

[0013] a memory communicatively connected to the at least one processor; wherein,

[0014] the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the digital human driving method according to any embodiment of the present invention.

[0015] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the digital human driving method according to any embodiment of the present invention when executed.

[0016] The technical solution of the embodiment of the present invention determines a target image, a target line, and a target emotion; inputs the target image, the target line, and the target emotion into a pre-trained digital human video generation model to obtain a target video, improving the generation efficiency of the target video and enabling the generated digital human to better express emotions.

[0017] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0019] Figure 1 is a flowchart of a digital human driving method according to Embodiment 1 of the present invention;

[0020] Figure 2 is a flowchart of an exemplary digital human driving method provided in Embodiment 2 of the present invention;

[0021] Figure 3 is a schematic structural diagram of a digital human driving device according to Embodiment 3 of the present invention;

[0022] Figure 4 is a schematic structural diagram of an electronic device for implementing the digital human driving method of the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] To enable those skilled in the art to better understand the solution of the present invention, the following will clearly and completely describe the technical solution in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0024] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0025] Embodiment 1

[0026] Figure 1 The following is a flowchart of a digital human driving method provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of generating digital human videos. This method can be executed by a digital human driving device, which can be implemented in the form of hardware and / or software, and the digital human driving device can be configured in an electronic device with data processing capabilities. As Figure 1 shown, the method includes:

[0027] S110. Determine the target image, target lines, and target emotion.

[0028] The target image is an image of the digital human; the target lines are the lines spoken by the digital human; the target emotion is the emotion shown by the digital human when speaking the target lines.

[0029] In the past, when generating digital human videos, the digital human needed to be manually dubbed, and the image was adjusted according to the dubbed content, which made the generation efficiency of the generated digital human videos relatively low. In addition, lines can also be generated in advance and then generated according to the lines and the digital human image, but the digital human in the generated video often lacks or is difficult to express a more appropriate emotion.

[0030] Therefore, the present application proposes a new digital human driving method to achieve the rapid generation of digital human videos and improve the emotion expression of digital humans.

[0031] In this application, it is necessary to first determine the target image, target lines, and target emotion. Among them, the target image is the image of the digital human to be generated, including but not limited to orange cats, rural dogs, and human figures, etc. In addition, it is also necessary to determine the lines that the digital human needs to speak in the target video, that is, the target lines. For example, "Put the red apple on the left". In addition, it is also necessary to determine the emotion of the digital human when speaking the target lines, such as "happy", "sad", etc.

[0032] S120. Input the target image, target lines, and target emotion into a pre-trained digital human video generation model to obtain the target video.

[0033] The target video is the digital human in the target image speaking the target lines with the target emotion.

[0034] To ensure the efficiency of the generated target video and at the same time ensure that the digital human in the target video can accurately speak the target lines with the target emotion, this application pre-trains a digital human video generation model. Inputting the target image, target lines, and target emotion into the pre-trained digital human video generation model can directly generate the target video, thereby improving the video generation efficiency and enhancing its performance ability in specific emotional scenarios.

[0035] Optionally, the training process of the digital human video generation model includes:

[0036] Determine the sample video, sample audio, and sample emotion; the sample audio is the audio of the sample video; the sample emotion is the emotion expressed by the digital human in the sample audio;

[0037] Input the sample video, sample audio, and sample emotion into a diffusion model for training to obtain the digital human video generation model.

[0038] For the sample video and sample audio, they can be obtained from the VoxCeleb2 benchmark dataset. The sample emotion can be obtained from the MEAD dataset.

[0039] During the training process, by obtaining the current frame image of the sample, predicting the next frame, and adjusting the prediction process based on the next frame image of the video.

[0040] In an optional solution, inputting the sample video, sample audio, and sample emotion into a diffusion model for training includes steps A1 - A9:

[0041] Step A1. Extract 3D canonical key point features, expression parameters, and head pose from the current frame image of the sample video.

[0042] Step A2. Concatenate the 3D canonical key point features, expression parameters, and head pose to generate the first feature vector.

[0043] Step A3: Add Gaussian noise to the first feature vector to obtain the second feature vector.

[0044] Step A4: Determine the audio features of the sample audio and the correspondence between the audio features and time to obtain the audio conditional features.

[0045] Step A5: Send the audio conditional features and the second feature vector to the residual network to generate the third feature vector.

[0046] Step A6: Through the self-attention mechanism and the third feature vector, perform inter-frame temporal information modeling to generate the fourth feature vector.

[0047] Step A7: Determine the sample emotion features of the sample emotion.

[0048] Step A8: Based on the cross-attention mechanism, fuse the sample emotion features with the fourth feature vector, and after fusion, obtain the predicted image of the next frame of the sample video through the sample emotion embedding.

[0049] Step A9: Adjust the predicted image of the next frame of the sample video according to the actual image of the next frame of the sample video.

[0050] For the current frame image I of the sample video 0 , extract the 3D canonical key point features x c , expression parameters (s 0 , δ 0 , t 0 ) and the head pose R 0 .

[0051] During extraction, it can be done through a feature extractor, that is, f extract (I 0 ) = x c , (s 0 , δ 0 , t 0 ), R 0 ;

[0052] where f extract is the feature extractor.

[0053] Concatenate x c , s 0 , δ 0 , t 0 , R 0 to obtain the first feature vector h 0 :

[0054] h 0 = concat(x c , s 0 , δ0 , t 0 , R 0 )。

[0055] Since the human characteristics of the digital human in each frame of the target video should be the same, for each frame, the feature vectors are the same.

[0056] The first feature vector h 0 Add Gaussian noise ∈ t Construct the second feature vector h at a specific time step t :

[0057]

[0058] where α t is the noise scheduling parameter at time step t, is Gaussian noise. The label of time step t generates the time feature t through the time encoding network f time (t). i .

[0059] While performing the above steps, the input sample audio a is passed through the pre-trained encoder f audio to obtain the audio hidden feature m = f audio (a). The audio feature m i is concatenated with the time embedding to obtain the audio conditional feature, c t = concat(m i , t i ).

[0060] After that, h t and c t are jointly input into the residual network to generate the third feature vector H t .

[0061] Then, self-attention mechanism is introduced to H t to model the inter-frame temporal information, Q t , K t , V t = H t W Q , H t W K , H t W V , and then calculate the attention weights:

[0062]

[0063] Finally, the fourth feature vector H' t ;

[0064] where H' t = Attention · Vt 。

[0065] Determine the sample emotion features according to the sample lines and sample emotions. Based on the cross-attention mechanism, fuse the sample emotion features with the fourth feature vector, and perform another sample line embedding after fusion. γ, β = tanh(FC 2 (ReLU(FC 1 (e text ))))), and perform another scaling and offset on the key points. The training objective of the above training is to minimize the mean square error between the predicted noise ∈ θ and the true noise ∈:

[0066] Meanwhile, the model can separate the pose loss and expression loss according to the index: and

[0067] In an alternative solution, for the current frame image of the sample video, extract 3D canonical key point features, expression parameters, and head pose, including:

[0068] Extract the 3D canonical key point features, expression parameters, and head pose of the current frame image of the sample video through a real-time portrait animation generation system.

[0069] When extracting 3D canonical key point features, expression parameters, and head pose, to ensure the accuracy of the extraction results, the real-time portrait animation generation system (LivePortrait) can be selected for extraction.

[0070] In an alternative solution, determine the correspondence between the audio features and time to obtain the audio conditional features, including:

[0071] According to the correspondence between the audio features and time, splice the audio features and time embeddings to obtain the audio conditional features.

[0072] To ensure that the audio features can correspond to time, splice the audio feature m i with the time embedding to obtain the conditional feature, c t = concat(m i , t i ).

[0073] In an alternative solution, send the audio conditional features and the second feature vector to the residual network to generate the third feature vector, including:

[0074] Map the audio conditional features to a scale vector and an offset vector through the residual network;

[0075] Scale and offset the second feature vector according to the ratio vector and the offset vector to obtain the adjusted second feature vector;

[0076] Concatenate the adjusted second feature vector and the second feature vector to obtain the third feature vector.

[0077] The second feature vector h t and the audio conditional feature c t are jointly input into the residual network. Among them, the audio conditional feature c t is mapped to the ratio vector γ and the offset vector β. After scaling and offsetting the second feature vector, a residual connection is made with the second feature vector to update the feature, thereby obtaining the third feature vector, denoted as H t .

[0078] Through the above method, the generation of the target video no longer depends on the existing emotional digital human generation technology for additional video input. It can not only generate more natural head and facial movements, but also dynamically adjust facial expressions according to the changes in speech content and intonation, making the generation result more context-appropriate, significantly improving the fidelity and consistency of emotional expression, and making the finally generated video have higher naturalness and emotional expressiveness.

[0079] According to the technical solution of the embodiment of the present invention, by determining the target image, the target line, and the target emotion; inputting the target image, the target line, and the target emotion into the pre-trained digital human video generation model to obtain the target video, the generation efficiency of the target video is improved, and the generated digital human can better express emotions.

[0080] Embodiment 2

[0081] Figure 2 This is a flowchart of an example of a digital human driving method provided by an embodiment of the present invention. As Figure 2 shown, the digital human driving method of this embodiment may include the following steps:

[0082] Given a sample video, set the batch size to 128, the model analyzes 64 frames at a time, and extracts 70 3D key point features for each frame. For the current frame I 0 , extract the key point feature x c , the expression parameters (s 0 , δ 0 , t 0 ) and the head pose R 0 : x c , (s 0 , δ 0 , t 0 ), R 0 = f extract (I 0 ), where f extractIt is a feature extractor.

[0083] Concatenate the above features into a first feature vector. Considering that the human features should be consistent, for any frame, the feature vectors are the same. Since the three RGB dimensions contain redundant information, the model will only learn for the first dimension. Therefore,

[0084] In the model, by adding Gaussian noise to the first feature vector H 0 Construct the input for a specific time step: t where α is the noise scheduling parameter at time step t, t is Gaussian noise. The token at time step t generates time features through the time encoding network f (t). time Generate time features

[0085] Meanwhile, the input sample audio a passes through the pre-trained encoder f audio to obtain audio hidden features. Concatenate the audio feature m i with the time embedding to obtain conditional features. For aligning audio and keypoint processing, the model performs zero-padding on the first two frames of the audio feature and finally obtains

[0086] For a 64-frame processing window, the model adds two additional frames of input and target keypoint information before the 64 frames to make the generated result both meet the reference conditions and maintain the consistency of the time series. After that, the keypoint information expands the hidden feature dimension through mapping and finally inputs

[0087] After that, x a-input and x h-input are jointly input into a residual network with both input and output hidden dimensions of 768. Where c t is mapped to the scale vector γ and the offset vector β, and the features are scaled and offset and then connected with the original input through a residual connection to update the features, denoted as H t .

[0088] Then, introduce the self-attention mechanism to model the inter-frame temporal information for H t , Q t , K t , V t = H t W Q , H t W K , H t W V ;

[0089] Furthermore, calculate the attention weights Finally, output H′ t = Attention·V t . When the Attention module is initialized, the total dimension of the input features is 768, there are 12 attention heads in total, and the feature dimension of each attention head is 64.

[0090] Encode the pre-constructed emotion-driven text using the openai / clip-vit-base-patch3 encoder, By using H′ t as the query, and e text providing the method of keys and values, fuse the text emotion features with H′ t through the cross-attention mechanism. The attention mechanism has the introduction of residual connections.

[0091] In addition to the cross-attention mechanism, this study uses an emotion enhancer to further expand the emotion performance. That is, after output 6, perform text information embedding again, γ, β = tanh(FC 2 (ReLU(FC 1 (e text ))), perform another scaling and offset on the key points.

[0092] The training objective of the diffusion model is to minimize the mean square error between the predicted noise ∈ θ and the real noise ∈: At the same time, the model can separate the pose loss and the expression loss according to the index: and

[0093] Embodiment III

[0094] Figure 3 This embodiment of the present invention provides a structural block diagram of a digital human driving device, and this embodiment is applicable to the situation of generating digital human videos. The digital human driving device can be implemented in the form of hardware and / or software, and the digital human driving device can be configured in an electronic device with data processing capabilities. As Figure 3 shown, the digital human driving device of this embodiment may include: a target data acquisition module 310 and a target video generation module 320. Among them:

[0095] The target data acquisition module 310 is used to determine the target image, the target line, and the target emotion; the target image is an image of the digital human; the target line is the line spoken by the digital human; the target emotion is the emotion shown by the digital human when speaking the target line;

[0096] The target video generation module 320 is configured to input the target image, target lines, and target emotion into a pre-trained digital human video generation model to obtain a target video, where the target video is the digital human in the target image telling the target lines with the target emotion.

[0097] Based on the above embodiments, optionally, the training process of the digital human video generation model includes:

[0098] Determine a sample video, sample audio, and sample emotion; the sample audio is the audio of the sample video; the sample emotion is the emotion expressed by the digital human in the sample audio.

[0099] Input the sample video, sample audio, and sample emotion into a diffusion model for training to obtain the digital human video generation model.

[0100] Based on the above embodiments, optionally, inputting the sample video, sample audio, and sample emotion into a diffusion model for training includes:

[0101] Extract 3D canonical key-point features, expression parameters, and head pose for the current frame image of the sample video.

[0102] Concatenate the 3D canonical key-point features, expression parameters, and head pose to generate a first feature vector.

[0103] Add Gaussian noise to the first feature vector to obtain a second feature vector.

[0104] Determine the audio features of the sample audio and determine the correspondence between the audio features and time to obtain audio conditional features.

[0105] Send the audio conditional features and the second feature vector to a residual network to generate a third feature vector.

[0106] Perform inter-frame temporal information modeling through a self-attention mechanism and the third feature vector to generate a fourth feature vector.

[0107] Determine the sample emotion features of the sample emotion.

[0108] Based on a cross-attention mechanism, fuse the sample emotion features with the fourth feature vector and, after fusion, pass through a sample emotion embedding to obtain a predicted image for the next frame of the sample video.

[0109] Adjust the predicted image for the next frame of the sample video according to the actual image of the next frame of the sample video.

[0110] Based on the above embodiments, optionally, extracting 3D canonical key-point features, expression parameters, and head pose for the current frame image of the sample video includes:

[0111] Extract the 3D canonical key-point features, expression parameters, and head pose of the current frame image of the sample video through a real-time portrait animation generation system.

[0112] Based on the above embodiments, optionally, determine the correspondence between the audio features and time to obtain audio conditional features, including:

[0113] According to the correspondence between the audio features and time, splice the audio features and time embeddings to obtain audio conditional features.

[0114] Based on the above embodiments, optionally, send the audio conditional features and the second feature vector to a residual network to generate a third feature vector, including:

[0115] Map the audio conditional features to a scale vector and an offset vector through the residual network;

[0116] Scale and offset the second feature vector according to the scale vector and the offset vector to obtain an adjusted second feature vector;

[0117] Splice the adjusted second feature vector and the second feature vector to obtain a third feature vector.

[0118] The digital human driving device provided by the embodiments of the present invention can execute the digital human driving method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.

[0119] Embodiment 4

[0120] Figure 4 FIG. shows a schematic structural diagram of an electronic device 10 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present invention described herein and / or claimed.

[0121] Such as Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory communicatively connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc. Among them, the memory stores a computer program executable by the at least one processor. The processor 11 can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. The input / output (I / O) interface 15 is also connected to the bus 14.

[0122] Multiple components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disc, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0123] The processor 11 can be various general-purpose and / or dedicated processing components with processing and computing capabilities. Some examples of the processor 11 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the digital human driving method.

[0124] In some embodiments, the digital human driving method can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the digital human driving method described above can be executed. Alternatively, in other embodiments, the processor 11 can be configured to execute the digital human driving method in any other appropriate way (for example, by means of firmware).

[0125] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems-on-chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0126] The computer programs for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer programs can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0127] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0128] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0129] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.

[0130] A computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.

[0131] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.

[0132] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A digital human driving method, characterized in that: include: Determine a target image, a target line, and a target emotion; the target image is an image of a digital human; The target lines are the lines narrated by the digital human; The target emotion is the emotion expressed by the digital human when speaking the target lines; The target image, target lines and target emotions are input into a pre-trained digital human video generation model to obtain a target video, wherein the target video is a digital human in the target image speaking the target lines with the target emotions.

2. The method according to claim 1, characterized in that The training process of the digital human video generation model includes: Determine a sample video, a sample audio, and a sample emotion; the sample audio is the audio of the sample video; the sample emotion is the emotion expressed by the digital human in the sample audio; The sample video, sample audio and sample emotion are input into the diffusion model for training to obtain the digital human video generation model.

3. The method according to claim 2, characterized in that Input sample videos, sample audio, and sample emotions into the diffusion model for training, including: For the current frame image of the sample video, extract 3D standard key point features, expression parameters and head posture; The 3D standard key point features, expression parameters and head posture are spliced ​​to generate the first feature vector; Adding Gaussian noise to the first eigenvector to obtain a second eigenvector; Determine the audio features of the sample audio, and determine the corresponding relationship between the audio features and time to obtain audio condition features; Sending the audio condition feature and the second feature vector to a residual network to generate a third feature vector; Through the self-attention mechanism and the third eigenvector, the inter-frame temporal information is modeled to generate the fourth eigenvector; determining sample emotion characteristics of the sample emotion; Based on the cross attention mechanism, the sample emotion feature is fused with the fourth eigenvector, and after fusion, the sample emotion is embedded to obtain the predicted image of the next frame of the sample video; The predicted image of the next frame of the sample video is adjusted according to the actual image of the next frame of the sample video.

4. The method according to claim 3, characterized in that For the current frame image of the sample video, extract the 3D standard key point features, expression parameters and head posture, including: Through the real-time portrait animation generation system, the 3D standard key point features, expression parameters and head posture of the current frame image of the sample video are extracted.

5. The method according to claim 3, characterized in that: The expression of the second eigenvector is: Where h0 represents the first eigenvector; α t is the noise scheduling parameter at time step t; is Gaussian noise; the label of time step t is passed through the temporal encoding network f time (t) Generate time feature t i .

6. The method according to claim 3, characterized in that Determine the corresponding relationship between audio features and time, and obtain audio condition features, including: According to the correspondence between audio features and time, the audio features are embedded and concatenated with time to obtain audio conditional features.

7. The method according to claim 3, characterized in that Sending the audio condition feature and the second feature vector to a residual network to generate a third feature vector includes: Mapping the audio condition feature into a scale vector and an offset vector through a residual network; Scaling and offsetting the second feature vector according to the scale vector and the offset vector to obtain an adjusted second feature vector; The adjusted second eigenvector is concatenated with the second eigenvector to obtain a third eigenvector.

8. A digital human driving device, characterized in that: include: A target data acquisition module is used to determine a target image, a target line, and a target emotion; the target image is an image of a digital human; The target lines are the lines spoken by the digital person; the target emotions are the emotions expressed by the digital person when speaking the target lines; The target video generation module is used to input the target image, target lines and target emotions into the pre-trained digital human video generation model to obtain the target video, which is the digital human in the target image speaking the target lines with the target emotions.

9. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the digital human driving method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the digital human driving method according to any one of claims 1 to 7 when executed.