A text-driven AIGC video generation method and device

Through the text-driven AIGC video generation method, the existing digital human speech video generation problems are solved, and fast and realistic video generation is achieved, reducing costs and enhancing user experience.

CN119255064BActive Publication Date: 2025-05-09SHENZHEN AIMALL TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411770572.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-05-09
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

The current digital human speaking videos are generated slowly, have poor production results, and are costly. They are easily affected by factors such as background noise, accent differences and quick speech, resulting in identification errors.

Method used

Through the text-driven AIGC video generation method, driver text and character images are obtained, target speech features and image features are generated, these features are fused to generate multi-frame video images, and finally corresponding speaking videos are generated.

Benefits of technology

It realizes the rapid generation of digital human talking videos, improves the generation effect, makes digital human performance more realistic and natural expressions, reduces generation costs, and enhances user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119255064B_ABST
    Figure CN119255064B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and in particular to a method for generating AIGC videos driven by text, the method comprising: obtaining driving text and a character image; generating a target voice feature according to the driving text; obtaining the image feature and facial feature of the character image according to the character image; fusing the target voice feature, the image feature and the facial feature to obtain a plurality of video frames; generating a speaking video corresponding to the character image according to the plurality of video frames, wherein the speaking video is an AIGC video, and the speaking content of the speaking video is the content of the driving text. The method uses the driving text as input, so that the generation speed of the digital human speaking video is relatively fast, and semantic information can be mined through the text, so that the digital human speaking video generation effect is excellent, the digital human is realistic, the digital human expression is natural, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a text-driven AIGC video generation method and device. Background Art

[0002] In recent years, with the advancement of deep learning technology, the technology for generating speaking videos using voice-driven 3D digital human faces has rapidly developed and has been applied in various fields. These speaking videos are AIGC (Artificial Intelligence Generated Content) videos. Existing digital human speaking videos typically use driving voice audio as input, which increases the cost of producing these videos. Furthermore, these digital human speaking videos using driving voice as input require voice audio recognition. During the voice recognition process, errors can occur due to factors such as background noise, accent differences, or rapid speech. This results in slow and poorly generated digital human speaking videos. Summary of the Invention

[0003] The embodiments of the present application provide a text-driven AIGC video generation method and device, which solves the technical problems of slow generation speed and poor generation effect of digital human speaking videos in the prior art. It uses driving text as input to increase the generation speed of digital human speaking videos, and can also mine semantic information through text, so that the digital human speaking video generation effect is excellent, the digital character is realistic, the digital human expression is natural, and the user experience is improved.

[0004] In a first aspect, an embodiment of the present invention provides a method for generating an AIGC video driven by text, comprising: obtaining driving text and a character image;

[0005] generating target speech features according to the driving text;

[0006] Obtaining image features and facial features of the person image according to the person image;

[0007] fusing the target speech features, the image features, and the facial features to obtain multiple frames of video images;

[0008] A speaking video corresponding to the character image is generated based on multiple frames of the video image, wherein the speaking video is an AIGC video, and the speaking content of the speaking video is the content of the driving text.

[0009] Preferably, generating target speech features according to the driving text includes:

[0010] According to the driving text, obtaining voice and text features corresponding to the driving text;

[0011] Obtaining speech features according to the speech, wherein the speech features are feature vectors containing semantic features;

[0012] The target speech feature is obtained according to the speech feature and the text feature.

[0013] Preferably, obtaining the target speech feature according to the speech feature and the text feature includes:

[0014] Extracting audio features from the speech features using an LSTM network, and extracting text features from the text features using a text extractor;

[0015] Performing a residual connection between the audio feature and the text feature through a variance adapter to obtain a latent attribute feature;

[0016] The latent attribute features and the speech features are concatenated through an encoder to obtain the target speech features.

[0017] Preferably, obtaining the text features according to the driving text includes:

[0018] The driving text is encoded by a clip encoder to obtain the text features.

[0019] Preferably, obtaining speech features according to the speech includes:

[0020] Mapping the speech to a latent space through a convolutional network to obtain latent features of the speech in the latent space;

[0021] The latent features are encoded through a Transformer network to obtain the speech features.

[0022] Preferably, fusing the target voice features, the image features, and the facial features to obtain multiple frames of video images includes:

[0023] The target speech features, the image features and the facial features are fused through a diffusion model until the posture constraint conditions of the motion estimation matrix are met, thereby obtaining one frame of the video image, and then obtaining multiple frames of the video image.

[0024] Preferably, the motion estimation matrix is:

[0025] M = Mt,t,E[||E - Et(Gt,t,C)|| 2 ];

[0026] Wherein, M is the motion estimation matrix, t is the time step, C is the speech feature, E is the multi-layer perceptron, Gt is Gaussian noise, Mt is the motion space matrix under t time step, and Et is the multi-layer perceptron linear operation.

[0027] Preferably, the posture constraint condition of the motion estimation matrix is ​​a target posture of a target feature obtained through the motion estimation matrix, and a condition for adjusting the target feature from a current posture to the target posture, wherein the target feature is a specified feature of the image feature and / or a specified feature of the facial feature.

[0028] Preferably, generating a speaking video corresponding to the character image based on multiple frames of the video image includes:

[0029] Repairing the multiple frames of the video image using a face repair model to obtain multiple frames of repaired video images;

[0030] The multiple frames of the repaired video images are sequentially encoded to obtain the speaking video.

[0031] Based on the same inventive concept, in a second aspect, the present invention also provides an AIGC video generation device driven by text, comprising: a memory, a processor, and a computer program stored in the memory and runnable on the processor, wherein when the processor executes the program, the steps of the AIGC video generation method driven by text in the first aspect are implemented.

[0032] One or more technical solutions in the embodiments of the present invention have at least the following technical effects or advantages:

[0033] In this embodiment of the present invention, after acquiring the driving text and a person image, target speech features are derived from the driving text, and image features and facial features are derived from the person image. Using the driving text as input for the digital human speaking video significantly reduces the cost of generating the speaking video and improves both efficiency and speed. Furthermore, the target speech features derived from the driving text contain high-level semantic information, facilitating realistic speech video generation.

[0034] The target speech features, image features, and facial features are then fused to produce multiple frames of video. Here, the target speech features are embedded and fused with the image and facial features to generate video images frame by frame. Based on the target speech features, which contain voice information, the digital human in the video image is rendered lifelike, with natural expressions and lip movements. This allows the video image to reflect the emotions of the speaker and allows for personalized video images and speaking videos. Then, based on the multiple frames of video, a corresponding speaking video is generated. This results in efficient and effective speaking video generation, enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. Throughout the drawings, the same reference figures denote the same components. In the drawings:

[0036] Figure 1 A schematic diagram showing the steps of a text-driven AIGC video generation method in an embodiment of the present invention is shown;

[0037] Figure 2 A schematic diagram of a module for obtaining target speech features in an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0038] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. Rather, these embodiments are provided to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art.

[0039] Example 1

[0040] The first embodiment of the present invention provides a method for generating AIGC video driven by text, such as Figure 1 Shown, including:

[0041] S101, obtaining driving text and character images;

[0042] S102, generating target speech features based on the driving text;

[0043] S103, obtaining image features and facial features of the person image based on the person image;

[0044] S104, fusing the target speech features, image features, and facial features to obtain multiple frames of video images;

[0045] S105 , generating a speaking video corresponding to the character image based on the multiple frames of video images, wherein the speaking video is an AIGC video, and the speaking content of the speaking video is the content of the driving text.

[0046] It should be noted that the speech video generated in this embodiment is an AIGC (Artificial Intelligence Generated Content) video.

[0047] In this embodiment, after acquiring the driving text and a person image, target speech features are derived from the driving text, while image features and facial features are derived from the person image. Using the driving text as input for the digital human speaking video significantly reduces the cost of generating the speaking video and improves both efficiency and speed. Furthermore, the target speech features derived from the driving text contain high-level semantic information, facilitating realistic speech video generation.

[0048] The target speech features, image features, and facial features are then fused to produce multiple frames of video. Here, the target speech features are embedded and fused with the image and facial features to generate video images frame by frame. Based on the target speech features, which contain voice information, the digital human in the video image is rendered lifelike, with natural expressions and lip movements. This allows the video image to reflect the emotions of the speaker and allows for personalized video images and speaking videos. Then, based on the multiple frames of video, a corresponding speaking video is generated. This results in efficient and effective speaking video generation, enhancing the user experience.

[0049] Next, combine Figure 1 The specific implementation steps of the text-driven AIGC video generation method provided in this embodiment are described in detail:

[0050] First, step S101 is executed to obtain a driving text and a character image. Specifically, the driving text is obtained through a device such as a keyboard. The character image is obtained through a device such as a camera, a mobile phone with a camera, or a tablet computer.

[0051] Next, step S102 is executed to generate target speech features according to the driving text.

[0052] Specifically, the process of generating target speech features is as follows: (1) Based on the driving text, obtain the speech and text features corresponding to the driving text. (2) Based on the speech, obtain speech features, where speech features are feature vectors containing semantic features. (3) Based on the speech features and text features, obtain the target speech features.

[0053] In this embodiment, corresponding voice and text features are generated based on the driving text. The voice generated by the driving text can avoid the voice influence caused by factors such as background noise, accent differences or fast speaking, so that clear and accurate speech content can be obtained during the generation process of the speaking video. Then, voice features are obtained based on the voice, so that the voice features have semantic information, which facilitates the speaking video to have semantic information and achieves excellent generation effect of the speaking video. The voice features and text features are combined to obtain the target voice features. This can further make the target voice features have more precise and accurate voice information, so that the dynamic speaking process of the digital human can be generated more efficiently and accurately during the generation process of the speaking video, thereby making the digital human realistic, the digital human's expression and lip shape natural, and improving the user experience.

[0054] The specific process of step (1) is to generate a speech of corresponding content from the driving text through TTS (Text To Speech) technology. Here, accurate speech is obtained by driving the text, so that the speaking video has accurate speech. In addition, the driving text is encoded by the clip (Contrastive Language-Image Pretraining) encoder to obtain text features. The text features are used to give the calibration speech features, so that the speech features have more precise and accurate semantic information, so that the speaking video has rich emotional information, and the digital person in the speaking video has real human emotional expression. The clip encoder includes an image encoder and a text encoder. The text encoder is used to convert the input driving text into a string of one-dimensional vectors based on the Transformer architecture, and then obtain text features. The clip text encoder has great potential in the field of multimodal learning due to its efficient performance, strong generalization ability and unique cross-modal understanding ability. It can perform multiple natural language processing and quickly and efficiently obtain accurate text features, laying a solid foundation for speaking videos.

[0055] The specific process of step (2) is to first map the speech to the latent space through the convolutional network to obtain the latent features of the speech in the latent space. Among them, the convolutional network includes convolutional networks such as AlexNet, VGG, and ResNet. Then, the latent features are encoded through the Transformer network to obtain speech features. Specifically, the local features are extracted from the input speech through the convolutional network to generate low-dimensional features Z. Then, the context information is extracted from the feature Z through the Transformer network to generate high-level features C, i.e., speech features. Speech features are feature vectors containing semantic features (i.e., semantic information). In the process of obtaining speech features based on speech, combining the local feature extraction capability of the convolutional network and the global context modeling capability of the Transformer network can more effectively process the audio signal of the speech and obtain accurate speech features, so that the speech video has semantic information, making the generation effect of the speech video excellent and the speech video vivid and flexible.

[0056] It's important to note that the latent space refers to the feature representation space of the outputs of one or more layers within a convolutional network. Each point in this space corresponds to an abstract representation of the input data, which typically contains key features and structural information about the input data. The latent space of a convolutional network is a multidimensional feature representation space. Through layer-by-layer feature extraction and transformation, it transforms the input data into an abstract representation suitable for the task. This representation not only contains the structural information of the input data but also has a certain semantic meaning, which is the foundation for the convolutional network to effectively handle complex tasks.

[0057] The specific process of step (3) is as follows: Figure 2 As shown, an LSTM (Long Short-Term Memory) network is used to extract audio features from speech features, and a text extractor is used to extract text features from text features. The audio and text features are then concatenated to produce joint features. A variance adapter is then used to adjust the joint features to produce adjusted joint features. The adjusted joint features are then encoded using an encoder to produce encoded joint features. These encoded joint features are then residually connected with the speech features to produce the target speech features.

[0058] Specifically, audio features are extracted from speech features using an LSTM network. The LSTM network is a special type of recurrent neural network (RNN) that effectively captures long-term dependencies in time series data. Using the LSTM network, temporal features (i.e., audio features) can be extracted from speech signals. These features contain information about the prosody, rhythm, and emotion of the speech, enabling further contextual reorganization of the speech features to achieve better audio features. Furthermore, a text extractor is used to extract text features from text features. Text features contain semantic information. Text extractors include the Word2Vec model and the BERT model. The audio and text features are concatenated to generate joint features. This concatenation is performed additively.

[0059] The joint features are then adjusted using a variance adapter to produce an adjusted joint feature. This adjustment adjusts the variance of the latent feature space, aligning the features of different joint features. This, in turn, ensures consistent features across different video frames during the diffusion model phase, preventing significant discrepancies in the generated video images caused by large changes. This prevents excessive and rapid changes in the video image, which can create abrupt visual artifacts and negatively impact the visual experience. The adjusted joint features are then encoded using an encoder to produce an encoded joint feature. The encoder adjusts the number of channels in the joint feature, aligning the encoded joint feature and speech features to a consistent dimension, thus addressing dimensional mismatches during the subsequent residual connection process. The encoder uses a basic linear network and activation function to adjust the number of channels. The encoded joint feature is then residually connected with the speech features to produce the target speech feature. Residual connections preserve the key information of the input features (the speech features and the encoded joint feature) while also introducing additional feature enhancement. In deep networks, features decay as the number of layers increases, losing their original semantic information. Adding pre- and post-processing features effectively prevents feature degradation. Therefore, a residual connection is created between speech features and encoded joint features to enhance their semantic information and effectively prevent feature degradation. Furthermore, the speech features and encoded joint features are fused to generate richer target speech features.

[0060] In the process of generating target speech features based on the driving text, the semantic information of the features is enhanced. The combination of text and audio improves video playback quality and user experience. By extracting and fusing multimodal features, the quality and naturalness of the generated video are improved, making the digital human speaking video more realistic and vivid.

[0061] Next, step S103 is executed to obtain the image features and facial features of the character image based on the character image. Specifically, image features and facial features are extracted from the character image through target detection and recognition models, such as YOLOv5, RT-DETR, CenterNet and other models. Facial features can play a role in identifying the identity of the character and keeping the identity unchanged, avoiding the situation where the identity of the digital person changes during the generation of the speaking video. Among them, the image features are the features corresponding to all targets in the character image, and the facial features are the relevant features of the character's face. For example, the character image is a photo of the upper body of a girl. The image features of the character image include hair features, facial features, shoulder features, arm features and palm features. The facial features of the character image are the facial features of the girl, including eyebrow features, eye features, nose features, mouth features, cheek features, forehead features and ear features.

[0062] The facial feature detection process uses an object detection model to detect five key points of the face and a face detection frame, then crops the face region based on the face detection frame. The five key points of the face include the eyes (left and right), the nose, and the mouth (left and right corners of the mouth). These five key points are then used to perform face correction. A face recognition feature extraction model is then used to extract a rich set of facial feature vectors, or facial features. This allows the image to be generated from a frontal view, even if the face is viewed from the side. This optimizes the image and efficiently extracts both image and facial features.

[0063] Then, step S104 is executed to fuse the target speech features, image features and facial features to obtain multiple frames of video images.

[0064] Specifically, the diffusion model fuses the target speech, image, and facial features until the pose constraints of the motion estimation matrix are met, resulting in a single frame of video image, and subsequently multiple frames of video image imagery. The diffusion model process also facilitates image denoising. Furthermore, the motion estimation matrix is ​​required during the diffusion model fusion process to enforce pose constraints.

[0065] The motion estimation matrix is: M = Mt,t,E[||E - Et(Gt,t,C)|| 2 ];

[0066] Where M is the motion estimation matrix, t is the time step, C is the speech feature, E is the multilayer perceptron, Gt is Gaussian noise, Mt is the motion space matrix at time step t, and Et is the multilayer perceptron linear operation. The multilayer perceptron (MLP) is a basic feedforward neural network. In this embodiment, it consists of an input layer, a hidden layer (Linear layer + activation function ReLU), and an output layer. M is the result of the operation based on the motion space matrix Mt for each different time step. It should be noted that in the diffusion model process, the motion estimation matrix must be initialized before subsequent operations on the motion estimation matrix.

[0067] The motion estimation matrix utilizes the latent semantics of the speech signal (i.e., speech features) to generate a matrix for controlling motion. During the diffusion model fusion process, the motion estimation matrix constrains the pose of each generated video frame. Furthermore, the motion estimation matrix estimates the motion relationship between each frame, generating a motion estimation trajectory. This constrains the pose of the generated video image, ensuring smooth transitions between frames and effectively reducing jitter.

[0068] The diffusion model works as follows: During the generation of each video frame, image features, facial features, and target voice features are sampled once to obtain the sampled features. The sampled features are then iterated once to generate a video frame. The diffusion model process samples the image, facial features, and target voice features multiple times, and performs multiple denoising operations on the person image. This allows the image to gradually align with the semantic expression of the feature description, resulting in a similar semantic expression between the features and the generated image. This allows the image to gradually approach the target image described by the features, thereby generating the video image. The sampling effect is that multiple sampling is more effective than single sampling, and the denoising process is gradual. This ensures that the generated video image contains both text and voice information while preserving the identity of the face. Furthermore, during the video generation process, the ControlNet module of the diffusion model runs a motion estimation matrix, allowing the diffusion model to fuse the target voice features, image features, facial features, and the motion estimation matrix to generate the video image. The ControlNet module also uses the motion estimation matrix to constrain the pose of the generated video image, ensuring smooth transitions and natural motion between frames, effectively reducing jitter. In this way, the ControlNet module is used to precisely control facial landmarks (eyes, nose, lips), and assist in predicting and generating a motion expression matrix of facial expressions and lip shapes that are synchronized with speech, namely the motion estimation matrix.

[0069] The diffusion model works on each frame, generating a frame-by-frame video image, and then generating multiple frames. The diffusion model extracts and fuses multimodal features, combined with generative adversarial networks and motion estimation techniques, to generate high-quality digital human speaking videos.

[0070] The pose constraint of the motion estimation matrix is ​​a condition for obtaining the target pose of the target feature through the motion estimation matrix, and for adjusting the target feature from the current pose to the target pose. The target feature is a specified feature of the image feature and / or a specified feature of the facial feature. For example, the character image is a photo of the upper body of a girl. In the first frame of the video image, the girl's left pupil is located in the center of the left eye. In the process of generating the second frame of the video image, the image features of the character image are sampled multiple times, and the target speech features and the facial features of the character image are embedded in the iterative process until the girl's left pupil is located in the left position of the left eye according to the motion estimation matrix. The diffusion model iteration ends and the second frame of the video image is generated.

[0071] In this embodiment, a diffusion model combined with a motion estimation matrix is ​​used to generate video images frame by frame. This ensures that each generated video frame retains the unique features of the input face, such as the eyes, nose, and mouth, ensuring a high degree of similarity between the generated digital human and the input face. By combining text and voice features, the generated speaking video not only accurately matches the voice but also better expresses emotion and intent, enhancing the video's realism and naturalness, achieving the relevant effects of multimodal information fusion. In each sampling round, the diffusion model generates a new video frame based on the current features and pose constraints, and further optimizes it in the next sampling round. This iterative process helps gradually approach the ideal video image, improving the overall quality of the generated video. The diffusion denoising process allows for faster convergence to the target image (i.e., the video image), reducing the number of sampling rounds and speeding up the generation time.

[0072] Finally, step S105 is executed to generate a speech video corresponding to the character image based on the multiple frames of video images, wherein the speech video is an AIGC video, and the speech content of the speech video is the content of the driving text. Specifically, the multiple frames of video images are repaired using a face repair model to obtain multiple frames of repaired video images. The face repair model includes models such as GFPGAN, GPEN, Restoreformer, and Codeformer. Each frame of the repaired video image is a video image with high-definition enhancement of the facial area and overall clarity. The multiple frames of the repaired video images are sequentially encoded to obtain a speech video. The term "sequentially" refers to the image sequence of the video images.

[0073] In this way, the face restoration model can significantly improve the resolution and clarity of the generated video images, eliminate facial pixel noise, and provide a better visual experience. The generated facial images are more realistic and natural, improving the quality of the speaking video. In addition, the visual effect of the generated speaking video is significantly improved, enhancing the user experience and ensuring the quality of the speaking video.

[0074] One or more technical solutions in the embodiments of the present invention have at least the following technical effects or advantages:

[0075] In this embodiment, after acquiring the driving text and a person image, target speech features are derived from the driving text, while image features and facial features are derived from the person image. Using the driving text as input for the digital human speaking video significantly reduces the cost of generating the speaking video and improves both efficiency and speed. Furthermore, the target speech features derived from the driving text contain high-level semantic information, facilitating realistic speech video generation.

[0076] The target speech features, image features, and facial features are then fused to produce multiple frames of video. Here, the target speech features are embedded and fused with the image and facial features to generate video images frame by frame. Based on the target speech features, which contain voice information, the digital human in the video image is rendered lifelike, with natural expressions and lip movements. This allows the video image to reflect the emotions of the speaker and allows for personalized video images and speaking videos. Then, based on the multiple frames of video, a corresponding speaking video is generated. This results in efficient and effective speaking video generation, enhancing the user experience.

[0077] Example 2

[0078] Based on the same inventive concept, the second embodiment of the present invention also provides an AIGC video generation device driven by text, including a memory, a processor, and a computer program stored in the memory and runnable on the processor. When the processor executes the program, the steps of any one of the above-mentioned AIGC video generation methods driven by text are implemented.

[0079] It should be noted that this device includes but is not limited to virtual devices, electronic devices, electronic chips, etc.

[0080] Those skilled in the art will appreciate that embodiments of the present invention may be provided as methods, apparatuses, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROMs, optical storage, etc.) containing computer-usable program code.

[0081] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.

[0082] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.

Claims

1. A text-driven AIGC video generation method, characterized in that: include: Get driving text and character images; Generating target speech features according to the driving text includes: According to the driving text, obtaining voice and text features corresponding to the driving text; According to the speech, a speech feature is obtained, wherein the speech feature is a feature vector including a semantic feature, and the speech feature is an audio signal; Obtaining the target speech feature according to the speech feature and the text feature; The step of obtaining the target speech feature according to the speech feature and the text feature includes: Extracting audio features from the speech features through an LSTM network, and extracting text features from the text features through a text extractor, wherein the audio features include time sequence features, and prosody, rhythm, and emotional information of the speech; Concatenating the audio feature and the text feature to obtain a joint feature; The joint feature is adjusted by a variance adapter to obtain an adjusted joint feature; Encoding the adjusted joint feature through an encoder to obtain an encoded joint feature, and performing a residual connection between the encoded joint feature and the speech feature to obtain the target speech feature; According to the character image, obtaining image features and facial features of the character image; fusing the target speech feature, the image feature and the facial feature to obtain multiple frames of video images; A speaking video corresponding to the character image is generated according to the multiple frames of the video image, wherein the speaking video is an AIGC video, and the speaking content of the speaking video is the content of the driving text.

2. The method according to claim 1, characterized in that According to the driving text, the text feature is obtained, including: The driving text is encoded by a clip encoder to obtain the text features.

3. The method according to claim 1, characterized in that The obtaining of speech features according to the speech includes: Mapping the speech to a latent space through a convolutional network to obtain latent features of the speech in the latent space; The latent features are encoded through a Transformer network to obtain the speech features.

4. The method according to claim 1, characterized in that The step of fusing the target voice feature, the image feature and the facial feature to obtain multiple frames of video images includes: The target speech features, the image features and the facial features are fused through a diffusion model until the posture constraint condition of the motion estimation matrix is ​​reached, thereby obtaining a frame of the video image, and then obtaining multiple frames of the video image.

5. The method according to claim 4, characterized in that The posture constraint condition of the motion estimation matrix is ​​a target posture of a target feature obtained through the motion estimation matrix, and a condition for adjusting the target feature from a current posture to the target posture, wherein the target feature is a specified feature of the image feature and / or a specified feature of the facial feature.

6. The method according to claim 1, characterized in that Generating a speaking video corresponding to the character image according to the multiple frames of the video image, including: Repairing the multiple frames of the video image using a face repair model to obtain multiple frames of repaired video images; The multiple frames of the repaired video images are sequentially encoded to obtain the speaking video.

7. A text-driven AIGC video generation device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the text-driven AIGC video generation method as described in any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Real-time driving interaction method and system of virtual digital human

    CN118963561A