A method, device and storage medium for generating a target person's video driven by voice
By training a multi-dimensional mapping between speech content, audio, and keypoint coordinates, the method addresses the unnatural integration of head and upper body movements in existing video generation methods, resulting in more realistic speaking videos.
Patent Information
- Application Number
- CN202111466434.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-30
AI Technical Summary
In the target person's speech video generated by the prior art, the linkage between the head and the upper body is unnatural and lacks a sense of nature.
By obtaining voice data and frontal images of the character's upper body, extracting the coordinate matrix of key points of the initial head and upper body, separating the voice content and audio information, training the multi-dimensional mapping relationship, generating a video image frame sequence, and finally splicing it into the target person's voice video.
The head movements and upper body movements in the generated video are naturally coordinated, enhancing the realism of the video.
Smart Images

Figure CN114202604B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer information technology, and in particular, to a method, device, and storage medium for generating a target person's video driven by speech. Background Art
[0002] The research on generating a face video driven by speech is an important research direction in human-computer interaction. Given any input speech, a video of a target person speaking with synchronized speech and lip movements is generated. Existing methods mainly use an end-to-end trained encoding-decoding framework. Chen et al. trained facial features using an attention mechanism to improve video accuracy. Mittal et al. proposed separating speech content and emotional information from speech and controlling the movement of the face and head using different feature dimensions. Edwards et al. performed multi-dimensional mapping of speech and the face to control facial movement. Qian et al. proposed the AutoVC method, a few-shot voice conversion method that separates speech into speech content and identity information. Zhou et al. proposed separating speech content and target person information, where the speech content information strongly controls the movement of the lips and adjacent facial positions, while the target person information determines the details of facial expressions and other dynamics of the target head. Nian et al. proposed a method for facial features and mouth feature key points, using the key points of the facial contour and the lips of the human face to represent the head movement information and lip movement information of the target person respectively. Although existing methods have achieved good results, they mainly focus on facial expressions and lip movements, ignoring head movements and the linkage of the upper body. The generated video of the target person speaking basically only has facial movements and a small number of head movements, and the linkage between the head and the upper body is not natural. Summary of the Invention
[0003] The present invention provides a method, device, and storage medium for generating a target person's video driven by speech to solve the problem that the generated video of the target person speaking basically only has facial movements and a small number of head movements, and the linkage between the head and the upper body is not natural in the prior art.
[0004] In a first aspect, a method for generating a target person's video driven by speech is provided, including:
[0005] Obtaining speech data and a front image of the upper body of a person;
[0006] Extracting an initial head key point coordinate matrix and an initial upper body key point coordinate matrix based on the obtained front image of the upper body of the person;
[0007] Separating speech content information and audio information based on the obtained speech data;
[0008] Training a multi-dimensional mapping relationship between the speech content information, audio information, initial head key point coordinate matrix, and initial upper body key point coordinate matrix;
[0009] Generate a video image frame sequence based on a multi-dimensional mapping relationship;
[0010] Concatenate the video image frame sequence with language data to obtain a target person's speech video.
[0011] Further, the extracting of the initial head key point coordinate matrix and the initial upper body key point coordinate matrix based on the obtained front image of the upper body of the person includes:
[0012] Segment the front image of the upper body of the person into a head image and an upper body image;
[0013] Use a head key point extraction model to extract the initial head key point coordinate matrix in the head image;
[0014] Use an upper body key point extraction model to extract the initial upper body key point coordinate matrix in the upper body image.
[0015] Further, the separating of the speech content information and the audio information based on the obtained speech data includes:
[0016] Extract the Mel spectrogram feature matrix of the speech data;
[0017] Input the Mel spectrogram feature matrix into an LSTM-based speech feature extraction network to obtain a speech feature matrix;
[0018] For each frame of the image, input the speech feature matrix within the corresponding preset speech frame window into a speech content encoder to obtain a speech content matrix;
[0019] For each frame of the image, input the speech feature matrix within the corresponding preset speech frame window into an audio encoder to obtain an audio matrix.
[0020] Further, the training of the multi-dimensional mapping relationship between the speech content information, the audio information, the head key point coordinate matrix, and the upper body key point coordinate matrix specifically includes:
[0021] Input the speech content matrix and the initial head key point coordinate matrix into a first multi-layer perceptron to predict the displacement of the head key point coordinate matrix in each frame of the image;
[0022] Based on the initial head key point coordinate matrix and the displacement of the head key point coordinate matrix in each frame of the image, obtain the predicted head position coordinates for each frame of the image;
[0023] Based on a self-attention network, fuse the speech content matrix and the audio matrix to obtain a self-reconstructed audio movement matrix;
[0024] Input the self - reconstructing audio movement matrix, the initial head key - point coordinate matrix, and the initial upper - body key - point coordinate matrix into the second multi - layer perceptron, and predict the overall displacement of the head key - point coordinate matrix and the upper - body key - point coordinate matrix for each frame of the image.
[0025] Based on the predicted coordinates of the head position in each frame of the image and the overall displacement of the head key - point coordinate matrix and the upper - body key - point coordinate matrix for each frame of the image, obtain the overall predicted coordinates for each frame of the image.
[0026] Furthermore, during the process of training the mapping relationship between the speech content matrix and the predicted head position coordinates, the speech content encoder and the first multi - layer perceptron used can be pre - trained synchronously based on the collected video data. The minimized loss function L used during the training process is c as follows:
[0027]
[0028] In the formula, P i,t represents the predicted coordinate position of the i - th head key - point in the t - th frame of the image, represents the actual coordinate position of the i - th head key - point in the t - th frame of the image, λ c represents the weight coefficient, represents the graph Laplacian coordinate of P i,t of, represents graph Laplacian coordinate, N represents the total number of head key - points, T represents the total number of frames of the image, and ||*||2 represents the L2 norm.
[0029] Furthermore, the generation of the video image frame sequence based on the multi - dimensional mapping relationship includes:
[0030] Input the initial head key - point coordinate matrix, the predicted coordinates of the head position in each frame of the image, and the overall predicted coordinates of each frame of the image into the video frame reconstruction model, and fuse them to obtain the reconstructed video frame sequence.
[0031] Furthermore, the minimized loss function L used during the training of the video frame reconstruction model is a as follows:
[0032]
[0033] In the formula, stc represents the number of frames of the source video, trg represents the number of frames of the target video, Q trg represents the reconstructed video frame, represents the real video frame, λ a represents the weight coefficient, λ a = 1, φ(Q trg ) and They respectively represent that the reconstructed video frames and the real video frames are pre-trained using the ResNet-34 network, and ||*||1 represents the L1 norm.
[0034] Further, the splicing of the video image frame sequence and the language data to obtain the target person's voice video includes:
[0035] The ffmpeg method is used to splice the video image frame sequence and the language data to obtain the target person's voice video, and the number of frames per second of the target person's voice video is a preset value.
[0036] In a second aspect, a voice-driven target person video generation device is provided, including:
[0037] A data acquisition module for acquiring voice data and a front image of the upper body of a person;
[0038] A key point extraction module for extracting an initial head key point coordinate matrix and an initial upper body key point coordinate matrix based on the acquired front image of the upper body of the person;
[0039] A voice data separation module for separating content information and audio information based on the acquired voice data;
[0040] A voice-image mapping module for training a multi-dimensional mapping relationship between the content information, the audio information, the head key point coordinates, and the upper body key point coordinates based on the content information, the audio information, the initial head key point coordinate matrix, and the initial upper body key point coordinate matrix;
[0041] A video frame generation module for generating a video image frame sequence based on the multi-dimensional mapping relationship;
[0042] A video generation module for splicing the video image frame sequence and the language data to obtain the target person's voice video.
[0043] In a third aspect, a computer-readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the voice-driven target person video generation method as described above.
[0044] Beneficial effects
[0045] The present invention provides a method, apparatus, and storage medium for generating a target person's video driven by voice. The input is a front image of the upper body of the target person and voice data, and the output is a front video of the target person's upper body speaking. The technical solution of the present invention first separates the voice content information and audio information in the voice data, controls the head movement through the voice content information, and controls the natural swing of the head and upper body through the audio information. The voice content information and audio information are subjected to multi-dimensional mapping training with the head key points and upper body key points to generate a mapping relationship, obtain a video image frame sequence, and finally generate a front voice video of the target person's upper body. The linkage between the head movement and the upper body is fully considered, and the generated video is natural and has a strong sense of reality. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0047] Figure 1 It is a flowchart of the method for generating a target person's video driven by voice provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will describe the technical solutions of the present invention in detail. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other implementation manners obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0049] Embodiment 1
[0050] This embodiment provides a method for generating a target person's video driven by voice, including:
[0051] S1: Obtain voice data and a front image of the upper body of a person; during implementation, the front image of the upper body of the person includes the head and the upper body part, and can be one or more.
[0052] S2: Extract an initial head key point coordinate matrix and an initial upper body key point coordinate matrix based on the obtained front image of the upper body of the person. More specifically, it includes:
[0053] S21: Preprocess the front image of the upper body of the person to obtain a unified size, such as 768*768. Crop and segment it into a head image and an upper body image by regular sub-framing. When cropping, only the coordinates of two points, the upper left corner and the lower right corner, are needed to determine the cropping position of the image. For example:
[0054] S211: The upper left corner coordinates of the head are (256, 0), and the lower right corner coordinates are (512, 256). The head image img is obtained through formula (1) head .
[0055] img head = Image[256:512,0:256,] (1)
[0056] S212: The upper body image does not focus on the head. Therefore, only the y coordinates of the width of the picture from 256 to 768 and all the x coordinate pixels are needed. The upper body image img is obtained through formula (2) body .
[0057] img body = Image[:,256:768,] (2)
[0058] S22: Use the head key point extraction model to extract the initial head key point coordinate matrix in the head portrait. In this embodiment, the open-source face detection and key point extraction tool DLIB library can be directly used to obtain the 68 initial head key point coordinate matrices q of the head image img head , and the lip movement information and head movement information of the human face are contained in the head key point coordinates. head
[0059] q head = DLIB(img head ) (3)
[0060] S23: Use the upper body key point extraction model to extract the initial upper body key point coordinate matrix in the upper body image. In this embodiment, the open-source Inception_v4 model is used to obtain the 128-dimensional initial key point coordinate matrix q of the upper body image img body body .
[0061] q body = Inception_v4(img body ) (4)
[0062] S3: Based on the speech content information, audio information, initial head key point coordinate matrix, and initial upper body key point coordinate matrix, train the multi-dimensional mapping relationship between the speech content information, audio information, head key point coordinates, and upper body key point coordinates.
[0063] A piece of voice data can be divided into voice content information and audio information. Different models are used to obtain the voice content information and audio information. The voice content information mainly determines the movement of the lips and the nearby area in general cases, and the audio information determines the head movement and the slight movement of the body. For example, when people are speaking, there will be changes in the head and facial expressions and no breathing movement will occur. While during a short pause, the changes in the head and facial expressions will decrease and there will be obvious breathing movements. The specific steps are as follows.
[0064] S31: Voice Content Information - Image Training
[0065] The coordinate mapping of the key points of the head features is obtained through the training of the voice content signal, mainly the coordinate mapping of the lips and the surrounding facial key points.
[0066] S311: Read the voice data using the librosa library in Python.
[0067] S312: Extract the Mel spectrogram feature matrix Audio of the voice data through the python_speech_features method mfcc .
[0068] S313: Input the Mel spectrogram feature Audio of the voice data mfcc into the LSTM-based voice feature extraction network to obtain the voice feature matrix A t ∈R M×D , where M is the total number of frames of the input voice, and D is the dimension of the voice feature matrix.
[0069] A t = LSTM(Audio mfcc ) (5)
[0070] S314: The sampling frequency of the voice data is much higher than that of the video frames. Therefore, there are multiple frames of voice data corresponding to the time window of one frame of the image. Thus, for each frame of the image, the voice feature matrix within the preset number of voice frames (set to 18 frames in this embodiment) corresponding to it is input into the AutoCV.E c voice content encoder to obtain the voice content matrix c t .
[0071] c t = AutoCV.E c (A t ; w lstm,C ) (6)
[0072] S315: Combine the voice content matrix c t and the initial head key point coordinate matrix q headInput the first multi-layer perceptron (MLP c ), and predict the displacement Δq of the head key point coordinate matrix in each frame of the image t ; where q head ∈R 68×3 .
[0073] Δq t = MLP c (c t , q head ; w mlp,C ) (7)
[0074] Where {w lstm,C , w mlp,C} are the learnable parameters of the AutoCV.E c and MLP c networks respectively. The LSTM has three layers of units, and each layer of units has an internal hidden state vector of size 256. The decoder (the first multi-layer perceptron) MLP c network has three layers, and the sizes of the internal hidden state vectors are 512, 256, and 204 (68×3) respectively.
[0075] S316: Based on the initial head key point coordinate matrix q head and the displacement Δq of the head key point coordinate matrix in each frame of the image t to obtain the predicted coordinate P of the head position in each frame of the image t .
[0076] P t = q head + Δq t (8)
[0077] During implementation, when training the mapping relationship between the speech content matrix and the predicted head position coordinates, the speech content encoder and the first multi-layer perceptron used can be pre-trained synchronously based on the collected video data. The minimized loss function L c adopted during the training process evaluates the distance between the actual coordinate position and the predicted coordinate position, as well as the distance between their respective image Laplacian coordinates. It promotes the correct placement of the coordinates relative to each other and preserves the head details. The formula of this loss function is as follows:
[0078]
[0079] In the formula, P i,t represents the predicted coordinate position of the i-th head key point in the t-th frame of the image, represents the actual coordinate position of the i-th head key point in the t-th frame of the image, and λ c represents the weight coefficient. In this embodiment, λ c takes 1. Denote \(P\) i,t as the graph Laplacian coordinates of denote the graph Laplacian coordinates, \(N\) represents the total number of head key points, \(T\) represents the total number of image frames, and \(\|\cdot\|_2\) represents the \(L_2\) norm. Among them, it is calculated by the following formula:
[0080]
[0081] where \(N(p\) i ) includes the coordinates adjacent to \(p\) located on different faces i connected. Eight facial parts are adopted, which contain a subset of coordinates predefined for the facial template. The calculation method of
[0082] S32: Speech content information and audio information - image training
[0083] It can maximize the similarity of audio information between different utterances of the same target person speaking and minimize the similarity between different target speakers. Through the training of speech audio information, the mapping relationship between the audio information corresponding head and the upper body is obtained. The movement of the head or the subtle correlation between the upper bodies is also a key factor in generating a realistic speaking target person.
[0084] S321: Use the librosa library in Python to read the speech data.
[0085] S322: Extract the Mel spectrogram feature matrix Audio of the speech data through the python_speech_features method mfcc .
[0086] S323: Input the Mel spectrogram feature Audio of the speech data mfcc into the LSTM-based speech feature extraction network to obtain the speech feature matrix \(A\) t \(\in\mathbb{R}\) M×D , where \(M\) is the total number of frames of the input speech and \(D\) is the dimension of the speech feature matrix.
[0087] \(A\) t \(= \text{LSTM}( \text{Audio}\) mfcc ) (11)
[0088] S324: At each frame of the image, input the speech feature matrix within the window of its corresponding preset number of speech frames (set to 18 frames in this embodiment) into the utoCV.E s audio encoder to obtain the audio matrix
[0089]
[0090] Among them, w lstm,S is the learnable parameter of the autoCV.E s audio encoder.
[0091] S325: Reduce the dimension of from 256 to 128 through a single-layer MLP, obtaining which can improve the generalization ability of facial videos, especially for target persons not observed during training.
[0092]
[0093] S326: To generate coherent head and upper body movements, compared with speech content actions, it requires capturing longer temporal correlations. While speech audio information typically lasts for dozens of milliseconds, head movements (such as the head swinging from left to right) and upper body movements (such as breathing movements) may last for one second or several seconds, even several orders of magnitude longer. To capture such long-term and structured dependencies, a self-attention network is used to calculate the output. Input the speech content matrix and the audio matrix into the decoder to obtain the self-reconstructed audio movement matrix h t , w attn , s represents the trainable parameters of the self-attention network (self-attention encoder).
[0094]
[0095] S327: The transformation of the speech content matrix is concatenated with the audio matrix and two initial coordinates. The weights assigned to each frame are calculated by a compatibility function that compares all frame representations in the window. In all experiments, the window size is set to τ = 256 frames (4 seconds). Input the self-reconstructed audio movement matrix h t and the initial head key point coordinate matrix q head and the initial upper body key point coordinate matrix q boby into the second multi-layer perceptron MLP s to generate the target speaker perception coordinate displacement. This MLP s predicts the overall displacement Δp t , w mlp,S represents the trainable parameters of the second multi-layer perceptron MLP s (MLP s decoder).
[0096] Δp t = MLP s (h t , q head , q body; w mlp,s ) (15)
[0097] S328: Predict the coordinate P based on the head position of each frame of image t and the overall displacement Δp of the head key point coordinate matrix and the upper body key point coordinate matrix of each frame of image t , to obtain the overall predicted coordinate y of each frame of image t .
[0098] y t = p t + Δp t (16)
[0099] During implementation, during the process of pre-training each network model, in addition to capturing the key point positions, it is also necessary to match the head movement and upper body movement of the target speaker. For this purpose, a discriminator network is created. The goal of the discriminator is to find out whether the temporal dynamics of the speaker's facial coordinates look "real" or fake. It takes the overall predicted coordinate sequence within the same window used in the generator, as well as the speech content matrix and the audio matrix as inputs, and returns a representation r t :
[0100]
[0101] During training, use the LSGAN loss function L gan to train the discriminator parameter w attn,d , regard the training coordinates as "real", and regard the generated coordinates as "fake" for each frame:
[0102]
[0103] where represents the output of the generator when the training coordinate y is used as its input, and r t represents the output in the discriminator.
[0104] To train the parameter w attn,s , the "authenticity" of the output will be maximized, and the loss function L s will be minimized, taking into account the training distances of the absolute position and the Laplacian coordinates:
[0105]
[0106] where λ s = 1 and μ s = 0.001 by maintaining the validation settings, y i,t represents the predicted coordinate position of the i-th head or upper body key point in the t-th frame of image, represents the actual coordinate position of the i-th head or upper body key point in the t-th frame of image, Denote y i,t as the graph Laplacian coordinates of Denote the graph Laplacian coordinates. Alternate training between the generator and the discriminator to improve each other as done in the GAN method.
[0107] S4: Generate a video image frame sequence based on the multi-dimensional mapping relationship. Specifically, it includes:
[0108] Input the initial head key point coordinate matrix q head , the predicted head position coordinates P for each frame of the image t and the overall predicted coordinates y for each frame of the image t into the video frame reconstruction model, and fuse them to obtain the reconstructed video frame sequence Q trg .
[0109]
[0110] The minimized loss function L used during the training of the video frame reconstruction model a is as follows:
[0111]
[0112] In the formula, stc represents the number of frames of the source video, trg represents the number of frames of the target video, Q trg represents the reconstructed video frame, represents the real video frame, λ a represents the weight coefficient, λ a = 1, φ(Q trg) and respectively represent that the reconstructed video frame and the real video frame are pre-trained using the ResNet-34 network, and ||*||1 represents the L1 norm.
[0113] The video frame reconstruction model is an encoder-decoder network, and this encoder-decoder network performs image conversion to generate video frame images. The encoder uses 6 convolutional layers, each of which contains a 2-step convolution, followed by two residual blocks, and generates a bottleneck, and then decodes through a symmetric upsampling decoder. Skip connections are used between the symmetric layers of the encoder and the decoder to generate each frame of the image. Since the coordinates change smoothly over time, the output image formed as an interpolation of these coordinates exhibits temporal coherence.
[0114] S5: Concatenate the video image frame sequence with the language data to obtain the target person's speech video. Specifically, it includes:
[0115] Use the ffmpeg method to concatenate the video image frame sequence and the language data to obtain the target person's speech video. In this embodiment, the number of frames per second of the target person's speech video is 29.
[0116] Example 2
[0117] This embodiment provides a voice-driven target person video generation device, including:
[0118] A data acquisition module, configured to acquire voice data and a front image of the upper body of a person;
[0119] A key point extraction module, configured to extract an initial head key point coordinate matrix and an initial upper body key point coordinate matrix based on the acquired front image of the upper body of the person;
[0120] A voice data separation module, configured to separate content information and audio information based on the acquired voice data;
[0121] A voice-image mapping module, configured to train a multi-dimensional mapping relationship between the content information, the audio information, the head key point coordinates, and the upper body key point coordinates based on the content information, the audio information, the initial head key point coordinate matrix, and the initial upper body key point coordinate matrix;
[0122] A video frame generation module, configured to generate a video image frame sequence based on the multi-dimensional mapping relationship;
[0123] A video generation module, configured to splice the video image frame sequence with the language data to obtain a target person voice video.
[0124] Example 3
[0125] This embodiment provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the voice-driven target person video generation method as described above.
[0126] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0127] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce means for implementing the specified functions in the Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the specified functions in a block or multiple blocks.
[0128] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the specified functions in the Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the specified functions in a block or multiple blocks.
[0129] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the specified functions in the Figure 1 one or more flows and / or blocks Figure 1 or means for implementing the specified functions in a block or multiple blocks.
[0130] It can be understood that the same or similar parts in the above embodiments can be referred to each other, and the content not detailed in some embodiments can be seen in the same or similar content of other embodiments.
[0131] Any process or method description in the flowchart or described in other ways herein can be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process, and the scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the technical field to which the embodiments of the present invention belong.
[0132] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for generating a target person's video driven by voice, characterized in that, Including: Obtain voice data and a front image of the upper body of a person; Extract an initial head key point coordinate matrix and an initial upper body key point coordinate matrix based on the obtained front image of the upper body of the person; Separate voice content information and audio information based on the obtained voice data; Train a multi-dimensional mapping relationship between voice content information, audio information, and head key point coordinates and upper body key point coordinates based on the voice content information, audio information, initial head key point coordinate matrix, and initial upper body key point coordinate matrix; Generate a video image frame sequence based on the multi-dimensional mapping relationship; Stitch the video image frame sequence with language data to obtain a target person's voice video; The training of the multi-dimensional mapping relationship between voice content information, audio information, and head key point coordinates and upper body key point coordinates based on the voice content information, audio information, head key point coordinate matrix, and upper body key point coordinate matrix specifically includes: Input the voice content matrix and the initial head key point coordinate matrix into a first multi-layer perceptron to predict the displacement of the head key point coordinate matrix in each frame of the image; Based on the initial head key point coordinate matrix and the displacement of the head key point coordinate matrix in each frame of the image, obtain the predicted head position coordinates in each frame of the image; Based on a self-attention network, fuse the voice content matrix and the audio matrix to obtain a self-reconstructed audio movement matrix; Input the self-reconstructed audio movement matrix, the initial head key point coordinate matrix, and the initial upper body key point coordinate matrix into a second multi-layer perceptron to predict the overall displacement of the head key point coordinate matrix and the upper body key point coordinate matrix in each frame of the image; Based on the predicted head position coordinates in each frame of the image and the overall displacement of the head key point coordinate matrix and the upper body key point coordinate matrix in each frame of the image, obtain the overall predicted coordinates in each frame of the image; During the process of mapping the relationship between the training speech content matrix and the predicted head position coordinates, the speech content encoder and the first multi-layer perceptron used can be pre-trained synchronously based on the collected video data, and the loss function minimized during the training process L c is as follows: ; In the formula, represents the predicted coordinate position of the i th head key point in the t th frame image, represents the actual coordinate position of the i th head key point in the t th frame image, represents the weight coefficient, represents the graph Laplacian coordinates of represents the graph Laplacian coordinates, N represents the total number of head key points, T represents the total number of frames of the image, represents the L2 norm.
2. The method for generating a target person video driven by voice according to claim 1, wherein The extraction of the initial head key point coordinate matrix and the initial upper body key point coordinate matrix based on the obtained front image of the upper body of the person includes: Segment the front image of the upper body of the person into a head image and an upper body image; Use a head key point extraction model to extract the initial head key point coordinate matrix in the head image; Use an upper body key point extraction model to extract the initial upper body key point coordinate matrix in the upper body image.
3. The method for generating a target person video driven by voice according to claim 1, wherein The separation of voice content information and audio information based on the obtained voice data includes: Extract the Mel spectrogram feature matrix of the voice data; Input the Mel spectrogram feature matrix into a voice feature extraction network based on LSTM to obtain a voice feature matrix; For each frame of the image, input the voice feature matrix within the corresponding preset number of voice frames window into a voice content encoder to obtain a voice content matrix; For each frame of the image, input the voice feature matrix within the corresponding preset number of voice frames window into an audio encoder to obtain an audio matrix.
4. The method for generating a target person video driven by voice according to claim 1, wherein The generation of the video image frame sequence based on the multi-dimensional mapping relationship includes: Input the initial head key point coordinate matrix, the predicted head position coordinates in each frame of the image, and the overall predicted coordinates in each frame of the image into a video frame reconstruction model to fuse and obtain a reconstructed video frame sequence.
5. The method for generating a target person video driven by voice according to claim 4, wherein, The loss function minimized during the training of the video frame reconstruction model is as follows: ; Wherein, stc represents the number of frames of the source video, trg represents the number of frames of the target video, represents the reconstructed video frame, represents the real video frame, represents the weight coefficient, , and respectively represent that the reconstructed video frame and the real video frame are pre-trained using the ResNet-34 network, represents the L1 norm.
6. The method for generating a target person video driven by voice according to claim 1, wherein Splicing the video image frame sequence and the language data to obtain the target person's voice video, including: Using the ffmpeg method to splice the video image frame sequence and the language data to obtain the target person's voice video, where the number of frames per second of the target person's voice video is a preset value.
7. A voice-driven target person video generation device for implementing the voice-driven target person video generation method as described in claim 1, characterized in that, Including: A data acquisition module, configured to acquire voice data and a front image of the upper body of a person; A key point extraction module, configured to extract an initial head key point coordinate matrix and an initial upper body key point coordinate matrix based on the acquired front image of the upper body of the person; A voice data separation module, configured to separate content information and audio information based on the acquired voice data; A voice image mapping module, configured to train a multi-dimensional mapping relationship between the content information, the audio information, the initial head key point coordinate matrix, and the initial upper body key point coordinate matrix; A video frame generation module, configured to generate a video image frame sequence based on the multi-dimensional mapping relationship; A video generation module, configured to splice the video image frame sequence and the language data to obtain the target person's voice video.
8. A computer-readable storage medium storing a computer program, characterized in that, The computer program is executed by a processor to implement the voice-driven target person video generation method according to any one of claims 1 to 6.
Citation Information
Patent Citations
High-quality face voice driving method based on neural radiation field
CN112887698A
Video generation method and device, electronic equipment and storage medium
CN113507627A