Video generation method, device, computer equipment and storage medium

By predicting the posture of the target user's facial image and audio features and adjusting the facial key points, the problem of insufficient anthropomorphism in existing technologies is solved, and more anthropomorphic video generation is achieved.

CN115100325BActive Publication Date: 2025-09-26PING AN TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210816012.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-12
Publication Date
2025-09-26
Estimated Expiration
2042-07-12

AI Technical Summary

Technical Problem

Among existing video generation technologies, the degree of anthropomorphism of voice-driven face synthesis is low, especially under different user style characteristics, where changes in head and facial postures lead to insufficient expressiveness of video generation.

Method used

By obtaining the target user's facial image and audio features, the target posture prediction model is used to predict the posture, obtain the character posture offset vector, and adjust the facial key points to generate a more humanized video.

Benefits of technology

It improves the anthropomorphism of video generation, enhances the expressiveness of user videos, and adapts to the personalized head and facial posture changes of different users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115100325B_ABST
    Figure CN115100325B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence technology, and discloses a video generation method, apparatus, computer equipment, and storage medium. The method extracts features from a target user's acquired facial image and audio to be processed, obtaining corresponding target facial key points and target audio features, and then predicts and adjusts the target facial key points based on the target audio features. A target pose prediction model is then used to perform pose prediction on the target audio features and target facial key points to obtain a target person pose offset vector for pose adjustment. After adjusting the target facial key points based on the target person pose offset vector, a target video corresponding to the target user with a higher degree of anthropomorphism is generated. The target pose prediction model is then used to perform pose prediction on the target audio features, thereby improving the degree of anthropomorphism in the generated user video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a video generation method, device, computer equipment and storage medium. Background Art

[0002] Facial animation synthesis video technology plays a vital role in filmmaking, virtual human synthesis, and simulated reality. Although facial synthesis technology has made many breakthroughs, there are still many technical difficulties that need to be addressed. For example, the posture control of lip synchronization still requires a lot of manual participation. This is because facial posture control relies on high-dimensional manifolds, making it difficult to find a function that maps speech and lip posture one to one.

[0003] Currently, voice-driven face synthesis typically synthesizes images of lip gestures after each frame is cropped, or uses GANs and CNN encoder-decoders to synthesize an entire image. By synchronizing speech and lip sounds, a corresponding user video is generated. However, due to the different styles of different users, their head and facial postures change when they speak. Existing technologies only modify partial facial postures, resulting in a slightly less expressive user video and a low degree of anthropomorphism in the generated video. Summary of the Invention

[0004] The embodiments of the present invention provide a video generation method, apparatus, computer equipment and storage medium to solve the problem of low degree of anthropomorphism in existing video generation.

[0005] An embodiment of the present invention provides a video generation method, including:

[0006] Obtain the target user's corresponding facial image and audio to be processed;

[0007] Performing feature extraction on the face image to be processed to obtain key points of the target face;

[0008] Extracting features from the audio to be processed to obtain target audio features;

[0009] Using a target pose prediction model, perform pose prediction on the target audio features and the target facial key points to obtain a target person pose offset vector;

[0010] According to the target person's posture offset vector, the target face key points are posture-controlled to generate a target video corresponding to the target user.

[0011] An embodiment of the present invention further provides a video generation device, comprising:

[0012] The target user confirmation module obtains the target user's corresponding facial image and audio to be processed;

[0013] A target face key point acquisition module is used to extract features from the face image to be processed and obtain target face key points;

[0014] A target audio feature acquisition module extracts features from the audio to be processed to obtain target audio features;

[0015] A target person posture offset vector acquisition module uses a target posture prediction model to perform posture prediction on the target audio features and the target facial key points to obtain the target person posture offset vector;

[0016] The target video acquisition module performs posture control on the target person's facial key points according to the target person's posture offset vector to generate a target video corresponding to the target user.

[0017] An embodiment of the present invention further provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned video generation method when executing the computer program.

[0018] An embodiment of the present invention further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned video generation method is implemented.

[0019] The above-mentioned video generation method, device, computer equipment and storage medium extract features of the acquired target user's corresponding facial image to be processed and the audio to be processed to obtain the corresponding target facial key points and target audio features, so as to make predictions based on the target audio features and adjust the target facial key points; by adopting a target posture prediction model, the target audio features and target facial key points are subjected to posture prediction to obtain the target person posture offset vector for posture adjustment; after adjusting the target facial key points based on the target person posture offset vector, a target video corresponding to the target user with a higher degree of anthropomorphism is generated, and by utilizing the target posture prediction model, the target audio features are subjected to posture prediction to improve the degree of anthropomorphism of the generated user video. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0021] Figure 1 is a schematic diagram of an application environment of a video generation method according to an embodiment of the present invention;

[0022] Figure 2 is a flow chart of a video generation method according to an embodiment of the present invention;

[0023] Figure 3 is another flow chart of a video generation method according to an embodiment of the present invention;

[0024] Figure 4 is another flow chart of a video generation method according to an embodiment of the present invention;

[0025] Figure 5 is another flow chart of a video generation method according to an embodiment of the present invention;

[0026] Figure 6 is a schematic diagram of a video generating device according to an embodiment of the present invention;

[0027] Figure 7 FIG. 1 is a schematic diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0029] The video generation method provided by the embodiment of the present invention can be applied in Figure 1 In the application environment shown. Figure 1 As shown, the client (computer device) communicates with the server via the network. The client, also known as the user end, refers to the program that corresponds to the server and provides local services to the client. The client (computer device) includes but is not limited to various personal computers, laptops, smartphones, tablets, cameras, and portable wearable devices. The server can be a standalone server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0030] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0031] The video generation method provided by the embodiment of the present invention can be applied as follows: Figure 1 Specifically, the video generation method is applied in a video generation system, which includes the following: Figure 1 The client and server shown communicate with each other through a network to generate and process the facial image and audio to be processed corresponding to the target user, so as to obtain a target video corresponding to the target user with a higher level of humanization.

[0032] In one embodiment, if Figure 2 As shown, a video generation method is provided, which is applied in Figure 1 The server in the example is used as an example, and the steps are as follows:

[0033] S201: Obtaining a face image and audio to be processed corresponding to a target user;

[0034] S202: Extract features from the face image to be processed to obtain key points of the target face;

[0035] S203: Extract features of the audio to be processed to obtain target audio features;

[0036] S204: Using a target pose prediction model, perform pose prediction on the target audio features and target facial key points to obtain a pose offset vector of the target person;

[0037] S205: Performing posture control on the target person's facial key points according to the target person's posture offset vector to generate a target video corresponding to the target user.

[0038] The target user is the user whose voice signal corresponds to the audio. The user's voiceprint is identified based on the difference in voiceprint information extracted from the voice signal. Because human vocal organs vary in size, shape, and function, these subtle differences lead to changes in the airflow, resulting in differences in sound quality and timbre. By identifying the corresponding voiceprint based on sound quality and timbre, different people can be distinguished based on their voices.

[0039] As an example, in step S201, after receiving the facial image to be processed and the audio file to be processed corresponding to the target user provided by the user through the client, the server processes the facial image to be processed according to the audio file to be processed to generate a target video corresponding to the target user.

[0040] In this example, the facial images to be processed include but are not limited to the head portrait photo and ID photo of the target user, and the audio files to be processed include but are not limited to the voice audio file and singing audio file of the target user.

[0041] The formats of facial images to be processed include, but are not limited to, bmp, jpg, png, tif, gif, pcx, tga, fpx, svg, psd, cdr, pcd, dxf, eps, ai, raw, wmf, webp, avif, and apng. The formats of audio files to be processed include, but are not limited to, CDA, WAV, MP3, MP3PRO, APE, FLAC, AAC, RealMedia, and Windows Media.

[0042] As an example, in step S202, the server uses facial key point detection technology to perform key point recognition processing on the acquired facial image to be processed, thereby obtaining target facial key points reflecting the facial features of the facial image to be processed. These facial key points can be used for facial recognition and facial editing. In this example, facial key point detection technology is used to obtain 68*3-dimensional matrix information corresponding to the facial image to be processed, which serves as the target facial key points. This is used to perform posture control on the facial image to obtain facial frame images with different expressions and postures.

[0043] Among them, the English name corresponding to the facial key point technology is face alignment. Facial key points include important feature points of various parts of the face, usually contour points and corner points, which can reflect the facial features of various parts.

[0044] As an example, in step S203, the server performs feature extraction on the obtained audio to be processed to obtain target audio features that can be recognized by the target posture prediction model. In this example, feature extraction is performed on the audio to be processed using voice conversion technology, and the corresponding user posture is predicted based on the obtained target audio features.

[0045] Among them, the audio conversion technology can be AutoVC technology, which is a zero-shot audio conversion technology based on audio coding loss and a many-to-many non-parallel audio conversion framework with better performance in many-to-many audio conversion tasks.

[0046] As an example, in step S204, after confirming the target facial key points and target audio features, the server uses the trained target pose prediction model to predict the pose of the facial image to be processed corresponding to the target audio features based on the target audio features and target facial key points, thereby obtaining a target pose offset vector for modifying the target facial key points. In this example, the pose of the facial image to be processed corresponding to the target audio features is predicted by analyzing the pose changes corresponding to the target audio features. Combined with the target facial key points corresponding to the facial image to be processed, the target pose offset vector for controlling the pose of the target facial key points corresponding to the facial image to be processed is determined.

[0047] Among them, the target person's posture offset vector is the amount of change that controls the change of the face image to be processed in two dimensions. According to the person's posture offset vector, the target face key points corresponding to the face image to be processed are controlled to change, so that the posture of the face image to be processed changes slightly and the face frame image with the required posture is achieved.

[0048] As an example, in step S205, after obtaining the target person's posture offset vector, the server controls the posture of the target face key points through the target person's posture offset vector, so that the face image to be processed makes a corresponding speaking posture according to the audio to be processed, thereby generating a target video corresponding to the target user, and completing the processing of generating the target video corresponding to the target user based on the face image to be processed and the audio to be processed corresponding to the target user.

[0049] Among them, through the unpaired image-to-image translation technology, the face image to be processed after the target face key point posture control is converted into the corresponding face frame image, and the target video corresponding to the final target user is obtained based on the generated face frame image.

[0050] In this example, feature extraction is performed on the acquired target user's corresponding facial image to be processed and audio to be processed to obtain the corresponding target facial key points and target audio features, so as to make predictions based on the target audio features and adjust the target facial key points; by adopting a target posture prediction model, posture prediction is performed on the target audio features and target facial key points to obtain a target person posture offset vector for posture adjustment; after adjusting the target facial key points based on the target person posture offset vector, a target video corresponding to the target user with a higher degree of anthropomorphism is generated, and by utilizing the target posture prediction model, posture prediction is performed on the target audio features to improve the degree of anthropomorphism of the generated target user video.

[0051] In one embodiment, if Figure 3 As shown, step S204, using the target posture prediction model, performs posture prediction on the target audio features and the target facial key points to obtain the target person's posture offset vector, including:

[0052] S301: Based on the target audio features, obtain target audio content features and target audio personality features;

[0053] S302: Performing posture prediction on the target audio content features and the target facial key points to obtain the target mouth posture offset vector;

[0054] S303: performing posture prediction on the target audio personality features and target facial key points to obtain a target head posture offset vector;

[0055] S304: Obtaining a target person's posture offset vector based on the target mouth posture offset vector and the target head posture offset vector.

[0056] As an example, in step S301, after confirming the target audio features, the server uses the features corresponding to the audio content in the target audio features as the target audio content features, and the features corresponding to the audio quality and timbre in the target audio features as the target audio personality features. In this example, since different users have different head and facial postures during speech, the head posture corresponding to each user's audio features is predicted based on the target audio personality features of each user, thereby improving the anthropomorphism of the user's video.

[0057] As an example, in step S302, after acquiring the target audio content features, the server performs mouth posture prediction based on the target audio content features and target facial key points. The server predicts changes in the target user's mouth posture when the corresponding audio content appears in the audio to be processed, corresponding to the target audio content features. This prediction uses the target mouth posture offset vector as the target user's corresponding facial key points. In this example, because the mouth postures of users speaking audio content share common characteristics, a target mouth posture offset vector is generated for each mouth posture change in the user's video based on the target audio content features corresponding to the audio to be processed.

[0058] As an example, in step S303, after obtaining the target audio personality features, the server performs head posture prediction on the target audio personality features and target facial key points. The prediction uses the target head posture offset vector as the target head posture offset vector, based on the changes in the target user's head posture when the target audio corresponding to the target user appears in the processed audio corresponding to the target audio personality features. In this example, the head posture includes the posture of the entire facial expression and the posture of the head movement. Because different users have different speaking postures, the target audio personality features are used to predict the user's corresponding speaking posture personality.

[0059] As an example, in step S304, after confirming the target mouth posture offset vector and the target head posture offset vector, the server obtains the predicted target person posture offset vector based on the target mouth posture offset vector for adjusting the target user's mouth posture and the target head posture offset vector for adjusting the target user's head posture, and adjusts the target facial key points of the target user.

[0060] In this example, by confirming the target mouth posture offset vector corresponding to the target audio feature for adjusting the target user's mouth posture, and the target head posture offset vector for adjusting the target user's head posture, a predicted target person posture offset vector is obtained, and the target facial key points of the target user are adjusted to improve the degree of anthropomorphism of the user video.

[0061] In one embodiment, step S302: performing posture prediction on target audio content features and target facial key points to obtain a target mouth posture offset vector includes:

[0062] S3021: Extracting time series features from target audio content features to obtain target first time series features;

[0063] S3022: Acquire a target first latent feature based on the target first temporal feature and the target audio content feature;

[0064] S3023: Perform posture prediction on the target facial key points and the target first latent feature to obtain the target mouth posture offset vector.

[0065] As an example, in step S3021, after extracting the target audio content features, the server performs time series feature extraction on the target audio content features to obtain the target first time series features. In this example, the LSTM network in the target pose prediction model is used to extract the corresponding target first time series features based on the target audio content features, thereby increasing the dimensionality of the target audio content features and improving their saliency.

[0066] Temporal features extract the temporal dimension of the target audio content based on a preset frame length, increasing the correlation between gesture changes and improving the smoothness of video generation. The preset frame length can be 0.3 seconds per frame, but can also be set based on actual needs.

[0067] As an example, in step S3022, after extracting the target first time series feature, the server fuses the target first time series feature and the target audio content feature to obtain the target first latent feature with the target first time series feature, so as to use the target first latent feature with more significant features to perform corresponding posture prediction.

[0068] As an example, in step S3023, after obtaining the target first latent feature, the server performs pose prediction based on the target facial key points and the target first latent feature, obtaining a target mouth pose offset vector for adjusting the target facial key points of the target user's target facial image at the corresponding time sequence, thereby changing the target user's mouth pose. In this example, the MLP network in the target pose prediction model is used to perform pose prediction based on the target facial key points and the target first latent feature to obtain the target mouth pose offset vector.

[0069] In this example, by extracting the time series features of the target audio content features and fusing the obtained target first time series features with the target audio content features, a more significant target first latent feature is obtained, and posture prediction is performed based on the target facial key points and the target first latent feature. A target mouth posture offset vector with a higher correlation between the posture changes is obtained, thereby improving the smoothness of the video generation.

[0070] In one embodiment, step S303: performing posture prediction on the target audio personality features and target facial key points to obtain the target head posture offset vector includes:

[0071] S3031: Extracting time series features from the target audio personality features to obtain target second time series features;

[0072] S3032: Acquire a second latent feature of the target based on the second temporal feature of the target and the individual audio feature of the target;

[0073] S3033: Perform posture prediction on the target facial key points and the target second latent features to obtain the target mouth posture offset vector.

[0074] As an example, in step S3031, after extracting the target audio individual features, the server performs time series feature extraction on the target audio individual features to obtain the target second time series features. In this example, the LSTM network in the target posture prediction model is used to extract the corresponding target second time series features based on the target audio individual features, thereby increasing the dimensionality of the target audio individual features and improving their salience.

[0075] The temporal feature extracts the time dimension of the target audio's individual characteristics based on a preset frame length to improve smoothness during user generation. The preset frame length can be 0.3 seconds per frame, or it can be set based on actual needs.

[0076] As an example, in step S3032, after extracting the target second time series feature, the server fuses the target second time series feature and the target audio personality feature to obtain the target second latent feature with the target second time series feature, so as to use the target second latent feature with more significant features to perform corresponding posture prediction.

[0077] As an example, in step S3033, after obtaining the target second latent feature, the server performs pose prediction based on the target facial key points and the target second latent feature, obtaining a target head pose offset vector for adjusting the target facial key points of the target user's target facial image at the corresponding time sequence, thereby changing the target user's head pose. In this example, the MLP network in the target pose prediction model is used to perform pose prediction based on the target facial key points and the target second latent feature to obtain the target head pose offset vector.

[0078] In this example, by extracting the time series features of the target audio personality features and fusing the obtained target second time series features with the target audio personality features, a target second latent feature with more significant features is obtained. Posture prediction is performed based on the target facial key points and the target second latent feature, and a target head posture offset vector with a higher correlation between the posture changes is obtained, thereby improving the smoothness of video generation.

[0079] In one embodiment, if Figure 4 As shown, step S204, based on the target person's posture offset vector, performs posture control on the target person's facial key points to generate a target video corresponding to the target user, including:

[0080] S401: performing posture control on the target person's facial key points according to the target person's posture offset vector to generate at least one target face frame image corresponding to the target user;

[0081] S402: performing splicing processing on each target face frame image to obtain a target video corresponding to the target user.

[0082] As an example, in step S401, after obtaining the target person's pose offset vector, the server controls the pose of the target person's facial key points based on the target person's pose offset vector, generating at least one target face frame image corresponding to the target user. In this example, the facial expression and head pose of the target person's facial key points are adjusted using the target head pose offset vector, and the mouth of the target face key points is adjusted using the target mouth pose offset vector. At least one target face frame image corresponding to the target user is generated. The target face frame images are then sorted based on their corresponding temporal features and frame lengths to ensure the order of the generated video.

[0083] As an example, in step S402, after obtaining at least one target face frame image corresponding to the target user, the server splices the target face frame images according to the corresponding temporal features and frame length to obtain the target video corresponding to the target user. In this example, the corresponding video generation can be performed through CGANs. Conditional Generative Adversarial Networks (CGANs) are an improvement on Generative Adversarial Networks (GANs). By adding additional conditional information to the original GAN's generator and discriminator, a conditional generation model is implemented. The advantage of CGANs is that the image generation process is controllable, ensuring the accuracy of the generated image.

[0084] In this example, the target person's facial key points are controlled using the target person's pose offset vector to generate at least one target face frame image corresponding to the target user. These target face frames are then spliced ​​together according to their corresponding temporal features and frame lengths to obtain a target video corresponding to the target user. The target person's facial expressions and head pose are adjusted using the target head pose offset vector, and the mouth of the target facial key points is adjusted using the target mouth pose offset vector, thereby enhancing the anthropomorphism of the generated target video corresponding to the target user.

[0085] In another embodiment, if Figure 5 As shown, before obtaining the face image to be processed and the audio to be processed corresponding to the target user in step S201, the process further includes:

[0086] S501: Obtain training user video;

[0087] S502: Decoupling the training user video to obtain at least one training posture frame and training audio features corresponding to the training posture frame;

[0088] S503: performing posture capture on each training posture frame according to the training audio features, and obtaining a training character posture offset vector corresponding to the training audio features;

[0089] S504: Training a posture prediction model according to at least one training audio feature and a training character posture offset vector corresponding to the training audio feature to obtain a target posture prediction model.

[0090] As an example, in step S501, the server obtains a training user video input by the user, and the training user video is used to train the posture prediction model. The training user video can be obtained by using broadcast video, short video, and recording data to achieve a wide range of training.

[0091] As an example, in step S502, after obtaining the training user video, the server decouples the training user video to obtain at least one training pose frame and the audio corresponding to the training pose frame. Feature extraction is then performed on the audio corresponding to the training pose frame to obtain training audio features. In this example, the training user video is decoupled by frame length, which can be 0.3 seconds. This frame length can be set as needed and is not a limiting condition.

[0092] As an example, in step S503, after obtaining the training posture frame images and the training audio features corresponding to the training posture frame images, the server performs posture capture on each training posture frame image according to the training audio features, obtains the training character posture offset vector corresponding to the facial key points in the current training posture frame image, and associates it with the training audio features for training the posture prediction model.

[0093] As an example, in step S504, the server inputs a posture prediction model based on at least one training audio feature and a training character posture offset vector corresponding to the training audio feature to train the posture prediction model and obtain a more accurate target posture prediction model. In this example, the posture prediction model utilizes a long short-term memory (LSTM) network and a multilayer perceptron (MLP) network. Classification algorithms may also be used, including but not limited to logistic regression, naive Bayesian, nearest neighbor, decision tree, and random forest algorithms.

[0094] In this example, by decoupling the training user video, at least one training posture frame and the training audio features corresponding to the training posture frame are obtained. The posture prediction model is trained using the training audio features and the training character posture offset vector after posture capture of the training posture frame to obtain a more accurate target posture prediction model.

[0095] In one embodiment, step S503, performing posture capture on each training posture frame according to the training audio features, and obtaining a training character posture offset vector corresponding to the training audio features, includes:

[0096] S5031: Based on the training audio features, obtain the training audio content features and the training audio personality features;

[0097] S5032: Capture the content and posture of each training posture frame according to the training audio content features to obtain the training mouth posture offset vector;

[0098] S5033: Capturing individual postures of each training posture frame according to the individual characteristics of the training audio, and obtaining a training head posture offset vector;

[0099] S5034: Obtain a training character posture offset vector based on the training mouth posture offset vector and the training head posture offset vector.

[0100] As an example, in step S5031, after confirming the training audio features, the server uses the features corresponding to the audio content in the training audio features as the training audio content features, and uses the features corresponding to the sound quality and timbre of the audio in the training audio features as the training audio personality features.

[0101] As an example, in step S5032, the server performs posture capture on each training posture frame according to the acquired training audio content features, and obtains the training mouth posture offset vector of the training posture frame corresponding to the current training audio content features and the facial key points, so as to determine the influence of different training audio content features on the mouth posture and improve the accuracy of the model posture prediction.

[0102] As an example, in step S5033, the server performs posture capture on each training posture frame according to the acquired training audio personality feature, and obtains the training head posture offset vector of the training posture frame corresponding to the current training audio personality feature and the facial key point, so as to determine the influence of different training audio personality features on the head posture and improve the accuracy of the model posture prediction.

[0103] As an example, in step S5034, after confirming the training mouth posture offset vector and the training head posture offset vector, the server obtains the training character posture offset vector based on the training mouth posture offset vector for predicting the user's mouth posture and the training head posture offset vector for predicting the user's head posture for subsequent posture prediction model training.

[0104] In this example, a training mouth pose offset vector for predicting the user's mouth pose and a training head pose offset vector for predicting the user's head pose are used to obtain a training character pose offset vector for subsequent pose prediction model training to improve the degree of anthropomorphism of the user video generated by the pose prediction model.

[0105] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0106] In one embodiment, a video generation device is provided, which corresponds to the video generation method in the above embodiment. Figure 6 As shown, the video generation device includes a target user confirmation module 601, a target face key point acquisition module 602, a target audio feature acquisition module 603, a target person posture offset vector acquisition module 604 and a target video acquisition module 605. The functional modules are described in detail as follows:

[0107] The target user confirmation module 601 obtains the target user's corresponding facial image and audio to be processed;

[0108] The target face key point acquisition module 602 performs feature extraction on the face image to be processed to obtain the target face key points;

[0109] The target audio feature acquisition module 603 performs feature extraction on the audio to be processed to obtain target audio features;

[0110] The target person posture offset vector acquisition module 604 uses a target posture prediction model to perform posture prediction on the target audio features and target facial key points to obtain the target person posture offset vector;

[0111] The target video acquisition module 605 performs posture control on the target person's facial key points according to the target person's posture offset vector to generate a target video corresponding to the target user.

[0112] In one embodiment, the target person posture offset vector acquisition module 604 includes:

[0113] A target audio content feature and target audio personality feature acquisition unit, which acquires the target audio content feature and the target audio personality feature based on the target audio feature;

[0114] A target mouth posture offset vector acquisition unit performs posture prediction on target audio content features and target facial key points to obtain a target mouth posture offset vector;

[0115] A target head posture offset vector acquisition unit performs posture prediction on the target audio personality features and target facial key points to obtain the target head posture offset vector;

[0116] The target person posture offset vector acquisition unit acquires the target person posture offset vector based on the target mouth posture offset vector and the target head posture offset vector.

[0117] In one embodiment, the target mouth posture offset vector acquisition unit includes:

[0118] A target first time series feature acquisition subunit extracts time series features from target audio content features to obtain target first time series features;

[0119] A target first latent feature acquisition subunit, which acquires the target first latent feature based on the target first temporal feature and the target audio content feature;

[0120] The target mouth posture offset vector acquisition subunit predicts the posture of the target face key points and the target first latent feature to obtain the target mouth posture offset vector.

[0121] In one embodiment, the target head posture offset vector acquisition unit includes:

[0122] The target second time series feature acquisition subunit is obtained to perform time series feature extraction on the target audio personality feature to obtain the target second time series feature;

[0123] The target second latent feature acquisition subunit acquires the target second latent feature based on the target second temporal feature and the target audio personality feature;

[0124] The target head posture offset vector acquisition subunit predicts the posture of the target face key points and the target second latent features to obtain the target head posture offset vector.

[0125] In one embodiment, the target video acquisition module 605 includes:

[0126] The target face frame image acquisition unit performs posture control on the target face key points according to the target person's posture offset vector to generate at least one target face frame image corresponding to the target user;

[0127] The target video acquisition unit performs splicing processing on each target face frame image to obtain the target video corresponding to the target user.

[0128] In another embodiment, the video generating apparatus further includes:

[0129] A training user video acquisition module acquires training user videos;

[0130] A training user video decoupling module decouples the training user video to obtain at least one gesture frame and audio features corresponding to the gesture frame;

[0131] The character posture offset vector acquisition module captures the posture of each posture frame according to the audio features and obtains the character posture offset vector corresponding to the audio features;

[0132] The target posture prediction model acquisition module trains a posture prediction model based on at least one audio feature and a character posture offset vector corresponding to the audio feature to acquire a target posture prediction model.

[0133] In one embodiment, a video generating device includes:

[0134] An audio content feature and audio personality feature acquisition unit, which acquires audio content features and audio personality features based on the audio features;

[0135] The mouth posture offset vector acquisition unit captures the content posture of each posture frame according to the audio content characteristics and obtains the mouth posture offset vector;

[0136] The head posture offset vector acquisition unit captures the individual posture of each posture frame according to the individual characteristics of the audio and obtains the head posture offset vector;

[0137] The character posture offset vector acquisition unit acquires the character posture offset vector based on the mouth posture offset vector and the head posture offset vector.

[0138] For the specific definition of the video generation device, please refer to the definition of the video generation method above and will not be repeated here. Each module in the above-mentioned video generation device can be implemented in whole or in part through software, hardware, or a combination thereof. Each of the above-mentioned modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each of the above modules.

[0139] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 7As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to execute data used or generated during the video generation method. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a video generation method is implemented.

[0140] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the video generation method in the above embodiment is implemented, for example Figure 2 S201-S205 shown, or Figures 3 to 5 Alternatively, when the processor executes the computer program, the functions of each module / unit in the embodiment of the video generating device are realized, for example, Figure 6 The functions of the target user confirmation module 601, the target facial key point acquisition module 602, the target audio feature acquisition module 603, the target person posture offset vector acquisition module 604 and the target video acquisition module 605 are not described here in detail to avoid repetition.

[0141] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the video generation method in the above embodiment is implemented, for example Figure 2 S201-S205 shown, or Figures 3 to 5 Alternatively, when the computer program is executed by a processor, the functions of each module / unit in the embodiment of the video generating device are realized, for example, Figure 6 The functions of the target user confirmation module 601, the target facial key point acquisition module 602, the target audio feature acquisition module 603, the target person posture offset vector acquisition module 604 and the target video acquisition module 605 are not described here in detail to avoid repetition.

[0142] Those skilled in the art will appreciate that all or part of the processes in the above-described embodiment methods can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above-described embodiment methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided by the present invention can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0143] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0144] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A video generation method, characterized in that: include: Obtain the target user's corresponding facial image and audio to be processed; Performing feature extraction on the face image to be processed to obtain key points of the target face; Extracting features from the audio to be processed to obtain target audio features; Using a target pose prediction model, perform pose prediction on the target audio features and the target facial key points to obtain a target person pose offset vector; According to the target person's posture offset vector, the posture of the target facial key points is controlled to generate a target video corresponding to the target user; the target person's posture offset vector is the amount of change that controls the change of the face image to be processed in two dimensions. According to the target person's posture offset vector, the target facial key points corresponding to the face image to be processed are controlled to change so that the posture of the face image to be processed changes and the face frame image with the required posture is achieved.

2. The video generation method according to claim 1, wherein: The step of performing posture prediction on the target audio features and the target facial key points to obtain a target person's posture offset vector includes: Based on the target audio features, obtaining target audio content features and target audio personality features; Performing posture prediction on the target audio content features and the target facial key points to obtain a target mouth posture offset vector; Performing posture prediction on the target audio personality features and the target facial key points to obtain a target head posture offset vector; A target person posture offset vector is obtained based on the target mouth posture offset vector and the target head posture offset vector.

3. The video generation method according to claim 2, wherein: The step of performing posture prediction on the target facial key points and the target audio content features to obtain a target mouth posture offset vector includes: Performing time series feature extraction on the target audio content feature to obtain a target first time series feature; Acquire a first target latent feature based on the first target time series feature and the target audio content feature; Perform posture prediction on the target facial key points and the target first latent feature to obtain a target mouth posture offset vector.

4. The video generation method according to claim 2, wherein: The step of performing posture prediction on the target facial key points and the target audio personality features to obtain a target head posture offset vector includes: Performing time series feature extraction on the target audio individual feature to obtain a target second time series feature; Acquire a second latent feature of the target based on the second time series feature of the target and the individual audio feature of the target; Perform posture prediction on the target face key points and the target second latent features to obtain a target head posture offset vector.

5. The video generation method according to claim 1, wherein: The step of performing posture control on the target person's facial key points according to the target person's posture offset vector to generate a target video corresponding to the target user includes: Performing posture control on the target person's facial key points according to the target person's posture offset vector to generate at least one target person's facial frame image corresponding to the target user; The target face frame images are spliced ​​together to obtain a target video corresponding to the target user.

6. The video generation method according to claim 1, wherein: Before obtaining the face image to be processed and the audio to be processed corresponding to the target user, the video generation method further includes: Get training user videos; Decoupling the training user video to obtain at least one training posture frame and a training audio feature corresponding to the training posture frame; Performing posture capture on each of the training posture frames according to the training audio features to obtain a training character posture offset vector corresponding to the training audio features; The posture prediction model is trained according to at least one of the training audio features and the training character posture offset vector corresponding to the training audio feature to obtain a target posture prediction model.

7. The video generation method according to claim 6, wherein: The step of performing posture capture on each posture frame according to the audio feature and obtaining a character posture offset vector corresponding to the audio feature includes: Based on the training audio features, obtaining training audio content features and training audio personality features; According to the training audio content features, performing content gesture capture on each of the training gesture frames to obtain a training mouth gesture offset vector; According to the individual characteristics of the training audio, individual posture capture is performed on each of the training posture frames to obtain a training head posture offset vector; A training character posture offset vector is obtained based on the training mouth posture offset vector and the training head posture offset vector.

8. A video generating device, characterized in that: include: The target user confirmation module obtains the target user's corresponding facial image and audio to be processed; A target face key point acquisition module is used to extract features from the face image to be processed and obtain target face key points; A target audio feature acquisition module extracts features from the audio to be processed to obtain target audio features; A target person posture offset vector acquisition module uses a target posture prediction model to perform posture prediction on the target audio features and the target facial key points to obtain the target person posture offset vector; The target video acquisition module controls the posture of the target facial key points according to the target person's posture offset vector to generate a target video corresponding to the target user; the target person's posture offset vector is the amount of change that controls the change of the face image to be processed in two dimensions. According to the target person's posture offset vector, the target facial key points corresponding to the face image to be processed are controlled to change so that the posture of the face image to be processed changes and the face frame image with the posture requirement is achieved.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the video generation method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the video generation method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Audio processing method and device, computer equipment and storage medium

    CN111444382A

  • Voice conversion model training method and device, electronic equipment and medium

    CN113689867A

  • Real-time voice-driven photo-level realistic face portrait video generation method

    CN114639374A