Voice-driven speaking portrait generation method and electronic equipment
Through the audio-space encoding module, dynamic fusion module and expression optimization module, the problems of lip synchronization, facial expression distortion and identity consistency in voice-driven portrait synthesis are solved, and high-quality and stable voice-driven portrait generation is achieved.
Patent Information
- Application Number
- CN202510241072.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-03
AI Technical Summary
In the existing speech-driven speaking portrait synthesis methods, insufficient lip synchronization accuracy, facial expression distortion and instability, and identity consistency problems have not been effectively solved, resulting in poor visual quality and user experience of generated videos.
The audio-space encoding module, audio-guided dynamic fusion module and facial key points are adopted to optimize the fusion process of audio and visual features by accurately capturing the alignment of audio features and spatial features, and fusing key point information through weighted KNN strategy to generate high-quality and highly synchronous video.
It achieves a high degree of consistency between lip movement and audio changes, avoids facial expression distortion and instability, maintains facial identity consistency, and improves the authenticity and fluency of generated videos.
Smart Images

Figure CN120298548A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of video processing, and particularly to a method for generating a talking portrait driven by voice and an electronic device. Background Art
[0002] Synthesis of talking portraits driven by voice is an important research direction in computer vision and graphics, aiming to generate high-quality facial images that are precisely synchronized with the given audio. With the continuous development of technology, the research applications in this field have gradually expanded, mainly focusing on synthesizing the lip movements of real or virtual characters driven by audio, providing high-quality facial synthesis effects for movies, animations, games, virtual reality, and various other applications.
[0003] In recent years, methods based on Neural Radiance Field (NeRF) have been widely applied to the task of synthesizing talking portraits driven by voice due to their significant advantages in image rendering. NeRF generates high-quality images by constructing a three-dimensional radiance field, especially outstanding in the details, naturalness, and lighting processing of facial images, breaking through the limitations of traditional methods in resolution. Although these methods have achieved some results, they still face a series of technical challenges, especially in improving the lip synchronization accuracy and image fidelity.
[0004] Currently, many existing NeRF-based methods pay more attention to the optimization of lip synchronization accuracy, but often ignore the overall coherence of facial expressions. This approach can improve the lip synchronization effect to a certain extent, but may lead to the distortion of facial expressions, thus greatly reducing the visual quality and user experience of the final generated video. In addition, existing methods usually fail to comprehensively consider the details of other areas of the face, making the overall effect still not coordinated enough to meet the requirements of high-quality video generation.
[0005] In general, the existing technologies mainly have the following technical problems: 1) Insufficient lip synchronization accuracy: Although the existing methods can achieve lip synchronization to a certain extent, due to the incomplete learning of the mapping from audio to image, there is still a significant deviation in the synchronization of audio and lips in the synthesis results. This deviation causes the generated image to be out of sync with the audio, which affects the naturalness and realism of the synthesized portrait. 2) Facial expression distortion and instability: Current methods usually only focus on the optimization of lip synchronization accuracy, but ignore the expression changes in other areas of the face. Since the change of facial expression involves the coordination and smooth transition of multiple areas, inaccurate facial key point information may cause facial expression distortion or jump, especially when the expression transition between two adjacent frames is more drastic. This distortion and jump makes the generated portrait show unnatural expression fluctuations, further destroying the authenticity and smoothness of the visual effect, and reducing the overall effect and viewing of the generated video. 3) Identity consistency: In the existing speech-driven portrait synthesis methods, cross-modal alignment of audio and visual features is still a technical problem. Directly applying audio features to the generation process often leads to changes in facial features and cannot maintain consistent identity information, thereby reducing the generalization ability and practicality of the synthesized portrait between different identities.
[0006] Therefore, how to balance the naturalness of lip synchronization and overall facial expressions and improve the overall authenticity and stability of synthesized videos remains a key issue that needs to be urgently addressed in the current field of speech-driven speaking portrait synthesis. Summary of the invention
[0007] In order to solve at least one of the technical problems existing in the prior art to a certain extent, the purpose of the present invention is to provide a high-quality, highly synchronized voice-driven speaking portrait generation method, electronic device and medium.
[0008] The first technical solution adopted by the present invention is:
[0009] A method for generating a speech-driven speaking portrait comprises the following steps:
[0010] Get audio features f based on input audio a , obtain the spatial feature f according to the input image x , obtain image features F according to the reference video frame;
[0011] The audio feature f a and spatial characteristics f x The audio-spatial encoding module is input to align the audio features with the spatial features by dynamically adjusting the spatial position of the audio features and the facial area, thus obtaining the spatially perceived lip features f lip ;
[0012] The audio feature f aThe audio-guided dynamic fusion module takes the image feature F and uses a dynamic linear layer to calculate the channel attention and spatial attention of audio perception, obtaining the dynamic fusion feature F″;
[0013] Obtain the reference key points L ref , and input the reference key points L ref and the audio feature f a into the expression optimization module, fuse the reference key points and the audio reference points, and obtain the key point feature;
[0014] Input the lip perception feature, the dynamic fusion feature and the key point feature into the trained generation network to generate video frames.
[0015] Furthermore, the working method of the audio-spatial encoding module is as follows:
[0016] Repeatedly expand the audio feature to make its dimension consistent with that of the spatial feature;
[0017] Connect the repeatedly expanded audio feature with the spatial feature, and obtain the attention weight W after processing aud_att ;
[0018] According to the attention weight W aud_att reweight each channel of the audio feature to obtain the spatially aware lip feature f lip .
[0019] Furthermore, the expression of the attention weight W aud_att is:
[0020] h = ReLU(W1[f a ; f x + b1)
[0021] W aud_att = σ(W2h + b2)
[0022] The expression of the lip feature f lip is:
[0023] f lip = f a ⊙W aud_att
[0024] where (W1, b1) and (W2, b2) are the weight matrix and bias vector of the multi-layer perceptron MLP; ⊙ represents the Hadamard product.
[0025] Furthermore, the working method of the dynamic fusion module is as follows:
[0026] Input the image feature F into the dynamic fusion module to calculate the channel attention weight A ch ;
[0027] Multiply the channel attention weight A ch element-wise with the image feature F to obtain the channel-refined visual feature F′;
[0028] Input the visual feature F′ into a dynamic linear layer to calculate and obtain the spatial attention map A sp
[0029] Multiply the spatial attention map A sp with the visual feature F′ to obtain the dynamically fused feature F″.
[0030] Furthermore, the dynamic fusion module includes a dynamic perceptron DyMLP; the dynamic perceptron DyMLP consists of two dynamic linear layers with a LeakyReLU activation function in the middle;
[0031] Input the image feature F into the dynamic perceptron DyMLP to calculate and obtain the channel attention weight A ch , and the calculation formula is as follows:
[0032] A ch = Sigmoid(DyMLP(F))
[0033] DyMLP(x) = DyLinear1(LeakyReLU(DyLinear2(x)))
[0034] In the formula, Sigmoid represents the Sigmoid function, DyLinear1 is the first dynamic linear layer, and DyLinear2 is the second dynamic linear layer;
[0035] The calculation formula of the visual feature F′ is as follows:
[0036]
[0037] In the formula, represents element-wise multiplication;
[0038] The calculation formula of the spatial attention map A sp is as follows:
[0039]
[0040] The calculation formula of the dynamically fused feature F″ is as follows:
[0041]
[0042] where DyLinear3 is the third dynamic linear layer.
[0043] Furthermore, the dynamic linear layer is used to adjust the visual feature and calculate the dynamic audio perception attention, and its expression is:
[0044] x out = DyLinear θ (x in )
[0045] In the formula, and respectively represent the input vector and output vector of the dynamic linear layer;
[0046] The trainable parameters of the dynamic linear layer are represented as Matrix factorization is used to reduce the number of parameters to reduce the training cost, that is, the parameter matrix θ is decomposed into and where M a is the parameter matrix after the audio feature f a undergoes a dynamic reshaping operation, which transforms the input from to S is a static learnable matrix, and K is the dimension of the parameter matrix reconstructed from the audio feature;
[0047] The parameters of the dynamic linear layer are reconstructed by multiplying M a by S, and the expression is as follows:
[0048] θ = M a S
[0049] M a = Reshape(f a )
[0050] where the parameters are not shared among different dynamic linear layers.
[0051] Furthermore, the working mode of the expression optimization module is as follows:
[0052] Taking the reference key points L ref and the audio feature f a as the input of the key point encoder, and outputting the generated key points L audio after audio correction and alignment;
[0053] Using a weighted KNN strategy based on distance and regional weights to perform weighted processing on the similarity and spatial distance between key points, and fusing the key point information from the reference video and the prediction model to generate the final key points;
[0054] Obtaining key point features according to the generated key points.
[0055] Furthermore, the step of using a weighted KNN strategy based on distance and regional weights to perform weighted processing on the similarity and spatial distance between key points, and fusing the key point information from the reference video and the prediction model to generate the final key points includes:
[0056] For the reference key point L ref and the predicted key point L audio perform regional division to obtain the lip key point set R lip , the eye key point set R eye , the nose key point set R nose , the eyebrow key point set R brow and the facial contour key point set R face ;
[0057] For each key point i, calculate the Euclidean distance d i :
[0058] d i = ||L audio,i - L ref,i ||2
[0059] According to the Euclidean distance d i , quantify the deviation between the predicted key point and the reference key point to guide the weighted fusion process;
[0060] Set an initial weight w r for each region, and normalize all weights to obtain the normalized weights The weight of each key point i is determined by the region R r(i) to which it belongs:
[0061]
[0062] Fuse the reference key point and the audio reference point, and the expression of the generated key point is:
[0063]
[0064] where L final is the finally generated fused key point set, represents the distance-based weight, the smaller the distance, the larger the weight, and w i represents the regional weight of each key point;
[0065] Encode the fused key point set L final into a key point feature representation.
[0066] Furthermore, the total loss during the training of the generation network is expressed as follows:
[0067]
[0068] where The MSE loss for each pixel value is specifically expressed as:
[0069]
[0070] In the formula, C represents the color value of each pixel in the generated image, and C gt represents the color value of each pixel in the reference image, and I represents the set of pixels in the image;
[0071] is the L2 loss between the true key points and the predicted key points, serving as the overall reconstruction constraint:
[0072]
[0073] In the formula, T represents the number of frames of the image, K represents the number of face key points in each frame of the image, and L ref ={l1, l2, l3, …, l T} represents the extracted true face key points, represents the predicted face key points;
[0074] is the reconstruction loss between the optimized key points and the reference key points:
[0075]
[0076] is the temporal consistency loss, which is used to enhance the coherence between frames, avoid the situation where the key point positions are different due to a large time span, thereby improving the naturalness and smoothness of the generated results. The temporal consistency loss is defined as follows:
[0077]
[0078] In the formula, t and t + 1 respectively represent the moments of two adjacent frames.
[0079] The second technical solution adopted by the present invention is:
[0080] An electronic device, the electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement a voice-driven talking portrait generation method as described above.
[0081] The third technical solution adopted by the present invention is:
[0082] A computer-readable storage medium stores at least one instruction, at least one program, a code set, or an instruction set, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the above-mentioned method for generating a voice-driven talking portrait.
[0083] The fourth technical solution adopted by the present invention is:
[0084] A computer program product or a computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the computer device to execute the above-mentioned method.
[0085] The beneficial effects of the present invention are as follows: By introducing an audio-spatial encoding module, the present invention achieves high-precision alignment between audio features and lip movements; through an audio-guided dynamic fusion module, the fusion process of audio and visual features is optimized, avoiding the problems of facial expression distortion and instability; the audio-guided dynamic fusion module proposed by the present invention can accurately fuse audio and visual features, ensuring that the influence of audio features on lip and facial expressions is effectively guided while maximizing the retention of facial identity features. Description of the Drawings
[0086] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the accompanying drawings of the relevant technical solutions in the embodiments of the present invention or the prior art. It should be understood that the accompanying drawings in the following introduction are only for conveniently and clearly expressing some embodiments of the technical solutions in the present invention, and those skilled in the art can also obtain other drawings based on these drawings without creative efforts.
[0087] Figure 1 is the overall framework diagram of a method for generating a high-quality and high-synchronization voice-driven talking portrait in an embodiment of the present invention;
[0088] Figure 2 is the structural diagram of an audio-guided dynamic fusion module in an embodiment of the present invention;
[0089] Figure 3 is the structural diagram of an expression optimization module based on facial key points in an embodiment of the present invention;
[0090] Figure 4 is the step flow chart of a method for generating a voice-driven talking portrait in an embodiment of the present invention. Detailed Embodiments
[0091] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention. For the step numbers in the following embodiments, they are only set for the convenience of description and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0092] In the description of the present invention, it should be understood that for the orientation description, such as the orientation or positional relationship indicated by up, down, front, back, left, right, etc., is based on the orientation or positional relationship shown in the accompanying drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the present invention.
[0093] In the description of the present invention, the meaning of several is one or more, the meaning of multiple is two or more, greater than, less than, exceeding, etc. are understood as not including the present number, and above, below, within, etc. are understood as including the present number. If the first and second are described only for the purpose of distinguishing technical features, they should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or the sequence of the indicated technical features.
[0094] In the description of the present invention, unless otherwise clearly defined, words such as set, install, connect, etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meaning of the above words in the present invention in combination with the specific content of the technical solution.
[0095] Term Explanation:
[0096] NeRF: Neural Radiance Field, neural radiance field.
[0097] Embodiment 1
[0098] As Figure 4 shown, this embodiment provides a method for generating a high-quality and highly synchronous voice-driven talking portrait, including the following steps:
[0099] S1. Obtain the audio feature f a according to the input audio, obtain the spatial feature f x according to the input image, and obtain the image feature F according to the reference video frame;
[0100] S2. Combine the audio feature f a and the spatial feature f xInput audio-spatial encoding module, by dynamically adjusting the spatial position of audio features and facial regions, aligns the audio features with spatial features to obtain spatially-aware lip features f lip ;
[0101] S3. Input the audio feature f a and the image feature F into the audio-guided dynamic fusion module, and use the dynamic linear layer to calculate the audio-aware channel attention and spatial attention to obtain the dynamically fused feature f″;
[0102] S4. Obtain the reference key points L ref , input the reference key points L ref and the audio feature f a into the expression optimization module, fuse the reference key points and audio reference points to obtain the key point features;
[0103] S5. Input the lip perception feature, the dynamically fused feature and the key point feature into the trained generation network to generate video frames.
[0104] To solve the technical problems such as insufficient lip synchronization accuracy, facial expression distortion and instability, and identity consistency in the existing speech-driven lip portrait synthesis methods, this embodiment proposes a novel framework, as Figure 1 shown. The framework combines three core modules: the audio-spatial encoding module, the expression optimization module based on facial key points, and the audio-guided dynamic fusion module. The following will explain these three core modules in detail in combination with the accompanying drawings and specific implementation manners.
[0105] (1) Audio-spatial encoding module
[0106] To solve the problem of insufficient lip synchronization accuracy, the audio-spatial encoding module is introduced. This module accurately captures the audio-driven lip movement features and maps them into spatial coordinates, thereby generating a high-quality lip movement representation.
[0107] The audio-spatial encoding module takes spatial features and audio features as inputs. First, the input audio undergoes operations such as resampling, framing, feature extraction, and encoding, and finally is encoded into the audio feature f a . At the same time, the NeRF framework maps the input image into the three-dimensional space and samples on the ray with the camera optical center as the origin to obtain the coordinates (x, y, z) of a series of sampling points and their corresponding directions. These low-dimensional spatial coordinate information is then encoded and converted into a high-dimensional spatial feature vector f x. Existing methods usually simply concatenate audio features and spatial features and then input them into a multi-layer perceptron (MLP) to control the synthesis of lip movements. However, this simple feature concatenation and encoding method fails to fully capture the complex associations between audio features and spatial features, resulting in poor audio synchronization of lip movements in the synthesis results. Therefore, this embodiment proposes an audio-spatial encoding module aimed at further enhancing the correlation between audio features and spatial features in lip movement synthesis.
[0108] First, the audio features are repetitively extended to have the same dimension as the spatial features. Then, the repetitively extended audio features and spatial features are concatenated through a multi-layer perceptron, and the attention weight W is obtained after processing aud_att , as shown in the formula:
[0109] h = ReLU(W1[f a ; f x + b1)
[0110] W aud_att = σ(W2h + b2)
[0111] The attention weight W aud_att re-weights each channel of the audio features to obtain the spatially aware lip features f lip :
[0112] f lip = f a ⊙W aud_att
[0113] where (W1, b1) and (W2, b2) are the weight matrices and bias vectors of the MLP. ⊙ represents the Hadamard product.
[0114] This module can achieve high-precision alignment of audio features and spatial features by dynamically adjusting the spatial positions of audio features and the facial region, making the generated lip movements highly consistent with audio changes and avoiding the lip asynchrony problems in traditional methods.
[0115] (2) Audio-guided dynamic fusion module
[0116] For the problems of facial feature distortion and identity consistency, this embodiment designs an audio-guided dynamic fusion module, as Figure 2 shown.
[0117] Existing audio-driven talking head synthesis methods usually rely on static multi-layer perceptrons. However, since the parameters of these MLPs are randomly initialized, the driving audio cannot provide additional guiding information during the generation process. In addition, the influence of the audio sequence on visual information changes dynamically, and the synthesized portrait is only associated with the dynamic features of the driving audio in the lip region. Therefore, appropriate adjustments must be made to the visual features to adapt to the changes in the audio sequence. To solve this problem, a dynamic linear layer is first introduced, which learns a new linear transformation to adjust the visual features and calculate the dynamic audio-aware attention. The expression of the dynamic linear layer can be represented as:
[0118] x out = DyLinear θ (x in )
[0119] where and represent the input vector and output vector of the dynamic linear layer respectively. The trainable parameters of the dynamic linear layer are represented as To reduce the training cost, matrix factorization is used to reduce the number of parameters. Specifically, the parameter matrix θ is decomposed into and M a is the parameter matrix after the dynamic reshaping operation of the audio feature f a , which transforms the input from to S is a static learnable matrix. The parameters of the dynamic linear layer can be reconstructed by multiplying M a with S, and can be represented as:
[0120] θ = M a S
[0121] M a = Reshape(f a )
[0122] where different dynamic linear layers do not share parameters.
[0123] The key to audio-guided fusion is the dynamic perceptron DyMLP, which consists of two dynamic linear layers with a LeakyReLU activation function in the middle. The present invention uses the dynamic linear layer to calculate the channel attention and spatial attention of audio perception, so as to obtain the fusion features. Feature extraction is performed on the reference video frame to obtain the image feature F. The image feature is input into the dynamic MLP, and the output value is normalized to [0,1] through the Sigmoid function. The calculation formula of the channel attention weight A ch is as follows:
[0124] A ch= Sigmoid(DyMLP(F))
[0125] DyMLP(x) = DyLinear1(LeakyReLU(DyLinear2(x)))
[0126] Multiply A ch element - by - element with F to obtain the channel - refined visual feature F′:
[0127]
[0128] The purpose of channel attention is to activate the visual feature channels related to driving audio.
[0129] To generate the spatial attention map, another dynamic linear layer is used to reduce the channel dimension and learn the region of interest of the audio. The calculation process is as follows:
[0130] A sp = Sigmoid(DyLinear3(F′))
[0131]
[0132] where A sp is the spatial attention map, F″ is the activated dynamic fusion feature, and the goal of spatial attention is to activate the spatial image information related to driving audio.
[0133] This module ensures high identity consistency while generating refined lip movements by guiding the network's dual attention on the lips and the overall face. Specifically, the model fully learns the dynamic mapping relationship between audio and lip movements through the refinement and fusion of channel - domain and spatial - domain features. Guided by the audio features, this module can efficiently fuse two - dimensional visual features, achieving excellent cross - modal alignment. This not only ensures the precise synchronization of the lips but also maximally retains the visual features unrelated to the driving audio, thus effectively guaranteeing the identity consistency in the generated video and significantly improving the overall quality and stability of the generated video.
[0134] (3) Facial expression optimization module based on facial key points
[0135] To solve the problem of facial expression distortion, the present invention proposes a facial expression optimization module based on facial key points. It is used to extract the underlying facial key point information from the input audio, as Figure 3 shown. First, key points are extracted from the face image to obtain a set L ref of 68 reference key points covering the main facial features. These key points include facial regions such as eyes, mouth, nose, and contour. These key points are then combined with the audio feature f aTogether, as the input of the key point encoder, the generated key points L after audio correction and alignment are output audio .
[0136] The predicted facial key points can relatively accurately reflect the dynamic characteristics of the lips changing with the audio. However, due to the existence of certain noise in the audio signal, the predicted facial key points will inevitably generate some errors, and these errors may cause the generated facial expressions to be distorted or have unnatural jumps. In order to maintain the stability of facial features and identity consistency while enhancing the dynamic expression ability, the present invention proposes a key point fusion strategy based on Weighted K-Nearest Neighbors (WKNN). This strategy optimizes the final facial expression generation effect by weighting the similarity and spatial distance between key points and fusing the key point information from the reference video and the prediction model. Specifically, by dynamically assigning weights to each facial key point, weighting according to the relative position between key points and their importance to the final result, and adjusting their contribution to the final generated result. The relatively stable facial regions (such as eyes, nose, etc.) have higher weights to maintain the consistency of individual features, while regions such as the lips that change greatly with speech are appropriately adjusted in weight to more accurately match the dynamic changes driven by the audio. This method can not only effectively suppress the key point offset caused by noise, but also improve the stability and accuracy of facial features while ensuring natural and smooth expressions, making the generated facial expressions more realistic and consistent.
[0137] First, divide the reference key points L ref ∈R 68×2 and the predicted key points L audio ∈R 68×2 into regions, and respectively obtain the lip key point set R lip , the eye key point set R eye , the nose key point set R nose , the eyebrow key point set R brow and the facial contour key point set R face , and the region labels can be expressed by the formula:
[0138]
[0139] For each key point i, calculate the Euclidean distance d i between the reference key point and the predicted key point:
[0140] d i =||L audio,i -L ref,i ||2
[0141] Through this distance metric, the deviation between the predicted key points and the reference key points can be quantified, thus further guiding the weighted fusion process.
[0142] Next, an initial weight w is set for each region r (where r ∈ {lip, eye, nose, brow, face}), and the weight is dynamically adjusted according to the distance d i and the importance of each region during the training process to achieve an adaptive weighted process. For example, neighboring key points have a greater impact on the prediction result, while those far from the target point contribute less, so as to ensure the smooth transition of local features. In this generation task, the goal of the model is to keep the facial regions except the lips stable, while ensuring that the lip key points have a higher influence and can flexibly adapt to audio-driven changes. Therefore, the weight w of other regions except the lips will be increased others to ensure their stability during the optimization process and maintain the consistency of portrait features under different audio inputs. For the lip region, since its key points need to highly adapt to the dynamic changes driven by speech, its weight w is appropriately reduced lip , making it more flexible during optimization, thus avoiding being restricted by other facial regions and ensuring the naturalness and accuracy of lip movement. During training, the regional weight coefficient is set as a learnable hyperparameter, enabling the model to automatically optimize the weight distribution during training to adapt to different audio-driven facial dynamic features. To ensure the stability of the weight optimization process, all weights are first normalized, and the specific formula is as follows:
[0143]
[0144] The normalized weight satisfies and ensures that the weights are always positive. The weight of each key point i is determined by the region R r(i) to which it belongs:
[0145]
[0146] Using the weighted KNN strategy based on distance and regional weights, the reference key points and the audio reference points are fused to generate the final key points. The key point fusion formula is as follows:
[0147]
[0148] where L final is the set of finally generated fused key points, represents the distance-based weight, the smaller the distance, the larger the weight, and w i represents the regional weights of each key point. The fused key point set L finalIt is encoded into key-point feature representations through a series of operations such as convolutional layers and fully connected layers.
[0149] The key-point fusion strategy realizes precise control of the key-point optimization process by dynamically adjusting the weights of different facial regions, ensuring that the optimized key-points can not only meet the dynamic changes driven by audio, but also reasonably distribute weights between different regions to balance local flexibility and overall stability. The expression optimization module based on facial key-points effectively reduces the expression jump problem caused by key-point errors in the generated video, avoids unnecessary jitter or distortion, makes the facial movement smoother and more coherent, thus improving the realism and quality of the synthesis result.
[0150] Finally, the lip perception feature, the dynamic fusion feature and the key-point feature jointly participate in modulating the portrait generation and training the overall network. The total loss is expressed as follows:
[0151]
[0152] where is the MSE loss of each pixel value, and the specific expression is:
[0153]
[0154] is the L2 loss between the real key-points and the predicted key-points, serving as the overall reconstruction constraint:
[0155]
[0156] T represents the number of frames of the image, K represents the number of facial key-points in each frame of the image, and L ref ={l1, l2, l3, …, l T} represents the extracted real facial key-points, represents the predicted facial key-points.
[0157] is the reconstruction loss between the optimized key-points and the reference key-points:
[0158]
[0159] is the temporal consistency loss, which is used to enhance the coherence between frames, avoid the situation where the key-point positions are different due to a large time span, and thus improve the naturalness and smoothness of the generated result. The temporal consistency loss is defined as follows:
[0160]
[0161] where t and t + 1 represent the moments of two adjacent frames respectively.
[0162] In summary, by solving the problems of insufficient lip synchronization accuracy, facial expression distortion and instability, and identity consistency, the method of the present invention achieves a higher-quality, more natural and stable speech-driven portrait synthesis effect, greatly improving the visual realism and application value of the synthesized video. Compared with the prior art, it has at least the following advantages and beneficial effects:
[0163] 1) Improve lip synchronization accuracy: By introducing an audio-spatial encoding module, high-precision alignment between audio features and lip movements is achieved.
[0164] 2) Optimize facial expressions: Through an audio-guided dynamic fusion module, the fusion process of audio and visual features is optimized, avoiding the problems of facial expression distortion and instability.
[0165] 3) Maintain identity consistency: The audio-guided dynamic fusion module proposed by the present invention can accurately fuse audio and visual features, ensuring that the influence of audio features on lip and facial expressions is effectively guided while maximizing the retention of facial identity features.
[0166] Embodiment 2
[0167] The embodiment of the present invention further provides an electronic device, which includes a processor and a memory. The memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement Figure 4 a speech-driven talking portrait generation method as shown.
[0168] It can be understood that the memory may include a random access memory (RAM), and may also include a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory can be used to store instructions, programs, codes, code sets or instruction sets. The memory may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above various method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.
[0169] The processor may include one or more processing cores. The processor uses various interfaces and circuits to connect various parts within the entire server. By running or executing instructions, programs, code sets, or instruction sets stored in the memory, and by invoking data stored in the memory, it performs various functions of the server and processes data. Optionally, the processor may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor may integrate a combination of one or more of a central processing unit (CPU) and a modem, etc. Among them, the CPU mainly processes the operating system, application programs, etc.; the modem is used to process wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor and may be implemented separately through a single chip.
[0170] Since this electronic device is the electronic device corresponding to a voice-driven talking portrait generation method according to an embodiment of the present invention, and the principle by which this electronic device solves problems is similar to that of this method, the implementation of this electronic device can refer to the implementation process of the above method embodiment, and repeated parts will not be elaborated.
[0171] Embodiment 3
[0172] An embodiment of the present invention further provides a computer-readable storage medium. At least one instruction, at least one segment of program, code set, or instruction set is stored in the storage medium. The at least one instruction, the at least one segment of program, the code set, or the instruction set is loaded and executed by a processor to implement Figure 4 a voice-driven talking portrait generation method as shown.
[0173] Those of ordinary skill in the art will understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. This program can be stored in a computer-readable storage medium, which includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc memories, magnetic disc memories, tape memories, or any other computer-readable medium that can be used to carry or store data.
[0174] Since this storage medium is the storage medium corresponding to a method for generating a voice-driven talking portrait according to an embodiment of the present invention, and the principle of solving problems by this storage medium is similar to that of this method, the implementation of this storage medium can refer to the implementation process of the above method embodiment, and the repeated parts will not be described again.
[0175] Embodiment 4
[0176] In some possible implementation manners, various aspects of the method according to an embodiment of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of a method for generating a voice-driven talking portrait according to various exemplary implementation manners described in the present specification above. Among them, the executable computer program code or "code" for executing each embodiment can be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0177] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits with logic gate circuits for implementing logic functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0178] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0179] The above embodiments are only for illustrating the technical concept and features of the present invention, and their purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered within the protection scope of the present invention.
Claims
1. A method for generating a speech-driven talking portrait, characterized in that, It includes the following steps: Obtain the audio feature f based on the input audio a , obtain the spatial feature f based on the input image x , obtain the image feature F based on the reference video frame; Input the audio feature f a and the spatial feature f x into the audio-spatial encoding module. By dynamically adjusting the spatial position of the audio feature and the facial region, align the audio feature with the spatial feature to obtain spatially-aware lip features; Input the audio feature f a and the image feature F into the audio-guided dynamic fusion module, and use the dynamic linear layer to calculate the audio-aware channel attention and spatial attention to obtain the dynamic fusion feature; Obtain the reference key point L ref , and input the reference key point L ref and the audio feature f a into the expression optimization module, fuse the reference key point and the audio reference point to obtain the key point feature; Input the lip perception features, dynamic fusion features, and key point features into the trained generation network to generate video frames.
2. The method for generating a speech-driven talking portrait according to claim 1, wherein The working mode of the audio-spatial encoding module is as follows: Repeatedly expand the audio features to make their dimensions consistent with those of the spatial features; Connect the repeatedly extended audio features with the spatial features, and obtain the attention weight W after processing aud_att ; According to the attention weight W aud_att Reweight each channel of the audio features to obtain the spatially aware lip feature f lip .
3. A method for generating a speech-driven talking portrait according to claim 2, characterized in that, The attention weight W aud_att has the following expression: h = ReLU(W1[f a ; f x + b1) W aud_att = σ(W2h + b2) The lip feature f lip is expressed as: f lip = f a ⊙W aud_att Among them, (W1, b1) and (W2, b2) are the weight matrix and bias vector of the multi-layer perceptron MLP; ⊙ represents the Hadamard product.
4. A method for generating a speech-driven talking portrait according to claim 1, characterized in that, The working mode of the dynamic fusion module is as follows: Input the image feature F into the dynamic fusion module to calculate the channel attention weight A ch ; Multiply the channel attention weight A ch element-wise with the image feature F to obtain the channel-refined visual feature F′; Input the visual feature F′ into a dynamic linear layer to calculate and obtain the spatial attention map A sp Multiply the spatial attention map A sp by the visual feature F′ to obtain the dynamically fused feature F".
5. A method for generating a speech-driven talking portrait according to claim 4, characterized in that, The dynamic fusion module includes a dynamic perceptron DyMLP; the dynamic perceptron DyMLP consists of two dynamic linear layers with a LeakyReLU activation function in the middle; Input the image feature F into the dynamic perceptron DyMLP to calculate and obtain the channel attention weight A ch , and the calculation formula is as follows: A ch = Sigmoid(DyMLP(F)) DyMLP(x) = DyLinear1(LeakyReLU(DyLinear2(x))) In the formula, Sigmoid represents the Sigmoid function, DyLinear1 is the first dynamic linear layer, and DyLinear2 is the second dynamic linear layer; The calculation formula of the visual feature F′ is as follows: In the formula, represents element multiplication; Spatial attention map A sp The calculation formula is as follows: A sp = Sigmoid(DyLinear3(F′)) The calculation formula of the dynamic fusion feature F″ is as follows: Among them, DyLinear3 is the third dynamic linear layer.
6. A method for generating a voice-driven talking portrait according to claim 1 or 4, characterized in that, The dynamic linear layer is used to adjust the visual feature and calculate the dynamic audio perception attention, and its expression is: x out = DyLinear θ (x in ) In the formula, and respectively represent the input vector and the output vector of the dynamic linear layer; The trainable parameters of the dynamic linear layer are represented as Matrix factorization is used to reduce the number of parameters to lower the training cost, that is, the parameter matrix θ is decomposed into and where M a is the parameter matrix after the audio feature f a has undergone a dynamic reshaping operation, S is a static learnable matrix, and K represents the dimension of the parameter matrix reconstructed from the audio feature; The parameters of the dynamic linear layer are reconstructed by multiplying M a by S, and the expression is as follows: θ = M a S M a = Reshape(f a ) Among them, the parameters are not shared between different dynamic linear layers.
7. A method for generating a speech-driven talking portrait according to claim 1, characterized in that The working mode of the expression optimization module is as follows: Take the reference key point L ref and the audio feature f a as the input of the key point encoder, and output the generated key point L after audio correction and alignment audio ; Use a weighted KNN strategy based on distance and regional weights to perform weighted processing on the similarity and spatial distance between key points, fuse the key point information from the reference video and the prediction model, and generate the final key points; Obtain the key point features according to the generated key points.
8. A method for generating a speech-driven talking portrait according to claim 7, characterized in that, The use of a weighted KNN strategy based on distance and regional weights to perform weighted processing on the similarity and spatial distance between key points, fuse the key point information from the reference video and the prediction model, and generate the final key points includes: For the reference key point L ref and the predicted key point L audio perform region division to obtain the lip key point set R lip , the eye key point set R eye , the nose key point set R nose , the eyebrow key point set R brow and the face contour key point set R face ; For each key point i, calculate the Euclidean distance d between the reference key point and the predicted key point i : d i = ||L audio,i -L ref,i ||2 According to the Euclidean distance d i , the deviation between the quantization prediction key point and the reference key point is quantified to guide the weighted fusion process; Set an initial weight w for each region r , and normalize all the weights to obtain the normalized weights The weight of each key point i is determined by the region R to which it belongs r(i) : Fuse the reference key points and the audio reference points, and the expression of the generated key points is: Among them, L final is the finally generated set of fusion key points, represents the distance-based weight, where the smaller the distance, the greater the weight, and w i represents the regional weight of each key point; Encode the fused key point set K final into a key point feature representation.
9. A method for generating a speech-driven talking portrait according to claim 1, characterized in that, The total loss during the training of the generation network has the following expression: Among them, is the MSE loss for each pixel value, and the specific expression is: where C represents the color value of each pixel in the generated image, C gt represents the color value of each pixel in the reference image, and I represents the set of pixels of the image; The L2 loss between the true key points and the predicted key points serves as the overall reconstruction constraint: Wherein, T represents the number of frames of the image, K represents the number of facial key points in each frame of the image, and L ref ={l1, l2, l3, …, l T} represents the extracted real facial key points, represents the predicted facial key points; To optimize the reconstruction loss between the key point and the reference key point: is the temporal consistency loss, defined as follows: In the formula, t and t + 1 respectively represent the moments of two adjacent frames.
10. An electronic device, characterized in that, The electronic device includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Method, system and device for driving image by voice and storage medium
CN113192162A
2D digital human generation method and system
CN118411454A
Three-dimensional facial animation generation method and apparatus based on audio driving, and medium
WO2024164748A1
Cited By
Facial animation generation method and device, electronic equipment and storage medium
CN116912375A
Image fusion method and device, storage medium and program product
CN120543395A