A multi-scale audio-guided face synthesis method and system thereof
Patent Information
- Application Number
- CN202610978354.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-02
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-02
AI Technical Summary
[0004]常规处理方式仅在瓶颈单点注入时,随上采样与卷积重建,在基础同步上可用,但在高分辨率细节阶段容易出现音频信息衰减、口型细节不稳定、闭合不完整等问题,且转置卷积上采样可能带来棋盘格伪影
[0047] 1. Audio information is not only injected at the bottom of the network, but continues to play a role from low resolution to high resolution as the decoding process progresses. The low resolution stage provides coarser audio constraints that are more helpful for the overall lip shape trend, while the high resolution stage provides finer audio constraints that are more helpful for the local lip shape details. The network can not only grasp the overall speaking rhythm and the general direction of lip opening and closing, but also better align with the lip shape corresponding to specific phonemes in detail, thereby avoiding the problem of audio information gradually weakening due to a single injection.
Smart Images

Figure CN122493881B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of computer vision and deep learning, and in particular to a method and system for synchronous synthesis of multi-scale audio-guided faces. Background Technology
[0002] The human audiovisual perception system naturally possesses the ability to correlate auditory and visual information across modalities. When a time shift occurs between the speaker's lip movements and the speech produced in a video (e.g., lip movements lag behind speech, or actions and sound effects are misaligned), audiovisual asynchrony phenomena such as lip-phoneme timing mismatch and incorrect phoneme correspondence occur. Therefore, in applications such as short video generation, virtual digital human driving, video dubbing, and multilingual translation, it is necessary to drive and align the input speech signal with the facial video of the target person, ensuring that the lip movement trajectory and speech phonemes are synchronized in terms of time sequence and morphological features.
[0003] The conventional approach employs a conditional generation framework consisting of an audio encoder, a face encoder, and a decoder. This involves extracting a Mel-spectrogram from the speech to obtain Mel spectral segments, and then using the audio encoder to output a fixed-scale audio feature. The lip region of the current frame is occluded, and the occluded frame is concatenated with a reference frame before being input into the face encoder for feature extraction. At the bottleneck, the features corresponding to the Mel spectral segments are fused with the facial features once, and then the decoder performs progressive upsampling to output the face / lip result.
[0004] Conventional processing methods are only applicable to basic synchronization when injecting at a single bottleneck point, along with upsampling and convolutional reconstruction. However, they are prone to problems such as audio information attenuation, unstable lip-sync details, and incomplete closure at high-resolution detail stages. Furthermore, transposed convolutional upsampling may introduce checkerboard artifacts. Summary of the Invention
[0005] In order to ensure that lip movements and speech phonemes are synchronized in time and form at the high-resolution detail stage of video and audio, and to reduce phenomena such as lip lag and phoneme misalignment, this application provides a multi-scale audio-guided synchronous synthesis method and system for face.
[0006] Firstly, this application provides a method for synchronously synthesizing multi-scale audio-guided faces, employing the following technical solution:
[0007] A method for synchronously synthesizing a human face using multi-scale audio guidance, comprising the following steps:
[0008] Acquire face encoding data and corresponding audio encoding data, input the face encoding data into the face encoder for multi-layer feature extraction, obtain multiple skip connection features of different resolutions, and use the deepest skip connection feature as the bottleneck face feature.
[0009] The audio encoded data is subjected to time-frequency processing, and the processed audio encoded data is subjected to multi-level convolutional sampling to construct multi-scale audio features. The multi-scale audio features represent audio feature tensors with different time-frequency resolutions and / or different feature map sizes.
[0010] The deepest multi-scale audio features are fused with the bottleneck face features to obtain the initial decoding features, which serve as the fusion starting point for face synchronization synthesis during the decoding process.
[0011] During the decoding process, the spatial resolution of the decoding features is gradually increased. In at least two decoding layers with different resolutions, the multi-scale audio features are injected layer by layer into the current decoding layer features and the corresponding skip connection features in the form of feature maps and then fused to obtain the decoding features of each decoding layer.
[0012] A global audio embedding is obtained based on the summarization of the multi-scale audio features. Scaling parameters and bias parameters are generated according to the global audio embedding and applied to the fused decoding features in a channel-level affine modulation manner to output the synthetic face image corresponding to the face encoding data.
[0013] By employing the above technical solutions, multi-scale audio features with differentiated time-frequency resolution and / or feature map size are constructed. These multi-scale audio features cover multiple levels of feature maps, ranging from high resolution to low resolution. Shallow feature maps possess high time-frequency resolution to preserve local short-term phoneme changes (e.g., the instantaneous spectrum of plosives and vowel transitions); deep feature maps, on the other hand, have low resolution and a large receptive field, enabling them to cover longer speech segments and aggregate macroscopic information such as overall energy envelope, speech rate, and intonation.
[0014] During the decoding process, the multi-scale audio features are injected layer by layer into the current decoding layer features and their corresponding skip connection features in the form of feature maps. Specifically, high-resolution audio features are injected into the high-resolution decoding layer, which, with its high temporal localization accuracy, directly constrains the high-resolution reconstruction stage of face synchronization synthesis in the decoding layer, thereby obtaining local lip details (such as the moment the lips close and the subtle movements corresponding to the friction between the teeth); low-resolution audio features are injected into the low-resolution decoding layer to constrain the approximate opening and closing amplitude and movement contour of the low-resolution stage in the decoding layer, thereby avoiding short-term jitter caused by relying solely on local information.
[0015] Based on the correspondence between the resolutions of multi-scale audio features and decoding features, audio features of different scales are injected into the corresponding decoding layers. Combined with the spatial details preserved from the jump connection features, the lip opening and closing and rounding changes corresponding to local phonemes can still be directly constrained in the high-resolution reconstruction stage, thereby reducing phenomena such as lip lag and phoneme misalignment.
[0016] In one embodiment, multi-level convolutional sampling is performed on the processed audio encoded data to construct multi-scale audio features, including the following steps:
[0017] The audio encoded data is converted into a time-frequency feature map, and multi-scale audio features with different time-frequency resolutions and / or different feature map sizes are obtained through multi-layer two-dimensional convolution and stepwise downsampling.
[0018] In one embodiment, obtaining a global audio embedding based on the multi-scale audio feature aggregation includes the following steps:
[0019] Based on the multi-scale audio features, the deepest audio feature tensor is determined, and pooling and linear mapping are performed on the deepest audio feature tensor to obtain the global audio embedding.
[0020] In one embodiment, the deepest skip connection feature is used as the bottleneck face feature, including the following steps:
[0021] The target face frame and the reference face frame are determined based on the face encoding data. The target face frame includes an unprocessed face frame or a face frame after weakening, occluding, or masking the lip area. The reference face frame includes a face frame containing the target's identity appearance information.
[0022] A hierarchical convolutional network structure is adopted to perform multi-layer convolution and sampling on the target face frame and the reference face frame to gradually expand and compress the spatial resolution. Multi-scale jump connection features from shallow local texture to deep high semantic structure are extracted, and the output of the deepest jump connection features is used as the bottleneck face feature.
[0023] By employing the above technical solution, a hierarchical convolutional network structure is used to perform multi-layer convolution and sampling on the target face frame and the reference face frame, gradually expanding and compressing the spatial resolution. This allows for the extraction of local texture features from shallow layers and high semantic structural features from deep layers, enabling the final output bottleneck face features to integrate multi-scale visual information. This preserves key identity and structural information of the face while compressing spatial resolution, enhancing the robustness of features to changes in pose and lighting. Hierarchical convolution extracts deep features from the target face frame and the reference face frame, allowing the bottleneck face features to carry the identity and appearance information in the reference face frame, while mitigating local interference caused by weakening, occlusion, or masking of the lip region in the current target face frame. This improves the stable expression ability of face features in identity preservation and dynamic changing scenarios. Gradually compressing the spatial resolution and outputting the deepest visual features as the bottleneck face features maps high-dimensional input images into compact low-dimensional feature representations, reducing data redundancy while preserving key semantic information. This facilitates efficient processing of subsequent face recognition, comparison, or reconstruction tasks.
[0024] In one embodiment, fusing the deepest multi-scale audio features with the bottleneck face features to obtain initial decoding features includes the following steps:
[0025] The deepest multi-scale audio features are channel-mapped, and then the channel-mapped deepest multi-scale audio features are summed or concatenated with the bottleneck face features to determine the initial features for decoding.
[0026] In one embodiment, the multi-scale audio features are injected layer by layer into the current decoding layer features and the corresponding skip connection features in the form of feature maps and then fused to obtain the decoding features of each decoding layer. The injection method of the multi-scale audio features includes the following steps:
[0027] The multi-scale audio features to be injected are subjected to channel mapping processing, and the processed multi-scale audio features are then interpolated, mapped, or rearranged for size alignment.
[0028] After alignment, coarse-scale multi-scale audio features are injected into the low-resolution decoding layer, while fine-scale multi-scale audio features are injected into the high-resolution decoding layer.
[0029] In one embodiment, scaling and bias parameters are generated based on the global audio embedding, and channel-level affine modulation is applied to the fused decoded features to output a synthetic face image corresponding to the face encoding data, including the following steps:
[0030] The global audio is embedded into the multilayer perceptron to generate scaling / bias parameters that match the number of channels in the decoding layer;
[0031] The decoding features of the decoding layer are multiplied one channel at a time by the scaling parameters corresponding to the number of channels of the decoding layer, and the bias parameters are added one channel at a time to determine the modulation features;
[0032] The modulation features are used as decoding constraints for the decoding layer to output a synthetic face image.
[0033] In one embodiment, embedding the global audio into an input multilayer perceptron includes the following steps:
[0034] The multilayer perceptron comprises a small network with at least two fully connected layers, whose input is a global audio embedding and whose output dimensions correspond to scaling parameters and bias parameters, respectively.
[0035] The decoding layer consists of multiple upsampling layers with different resolutions in the decoder, and each decoding layer independently generates its own modulation parameters;
[0036] The channel-level affine transformation is performed after the skip connection fusion and / or scale-aligned audio injection of the decoding layer is completed, and before entering the next decoding layer or before output.
[0037] In one embodiment, before injecting the multi-scale audio features layer by layer into the current decoding layer features and the corresponding skip connection features in the form of feature maps and fusing them, the following steps are also included:
[0038] The current decoding layer features are interpolated and upsampled to obtain an upsampled feature map. The upsampled feature map is then subjected to at least one convolutional layer to obtain decoding layer features with learnable transformations of channel data and spatial features.
[0039] Secondly, this application provides a multi-scale audio-guided synchronous face synthesis system, employing the following technical solution:
[0040] A multi-scale audio-guided face synchronization synthesis system, performing the multi-scale audio-guided face synchronization synthesis method described in the first aspect, includes:
[0041] The face processing module acquires face encoding data and corresponding audio encoding data, inputs the face encoding data into the face encoder for multi-layer feature extraction, and obtains multiple skip connection features of different resolutions. The deepest skip connection feature is used as the bottleneck face feature.
[0042] An audio processing module performs time-frequency processing on the audio encoded data and performs multi-level convolution sampling on the processed audio encoded data to construct multi-scale audio features. The multi-scale audio features represent audio feature tensors with different time-frequency resolutions and / or different feature map sizes.
[0043] The bottleneck fusion module is used to fuse the deepest multi-scale audio features with the bottleneck face features to obtain the initial decoding features, which serve as the fusion starting point for face synchronization synthesis during the decoding process.
[0044] The feature decoding module progressively increases the spatial resolution of the decoding features during the decoding process. In at least two decoding layers with different resolutions, the multi-scale audio features are injected layer by layer into the current decoding layer features and the corresponding skip connection features in the form of feature maps and then fused to obtain the decoding features of each decoding layer.
[0045] The feature modulation module obtains a global audio embedding based on the summarization of the multi-scale audio features, generates scaling parameters and bias parameters based on the global audio embedding, and applies them to the fused decoded features in a channel-level affine modulation manner to output the synthetic face image corresponding to the face encoding data.
[0046] In summary, this application includes at least one of the following beneficial technical effects:
[0047] 1. Audio information is not only injected at the bottom of the network, but continues to play a role from low resolution to high resolution as the decoding process progresses. The low resolution stage provides coarser audio constraints that are more helpful for the overall lip shape trend, while the high resolution stage provides finer audio constraints that are more helpful for the local lip shape details. The network can not only grasp the overall speaking rhythm and the general direction of lip opening and closing, but also better align with the lip shape corresponding to specific phonemes in detail, thereby avoiding the problem of audio information gradually weakening due to a single injection.
[0048] 2. Global audio embedding applies feature-wise linear modulation (FiLM) to multi-level decoding features, providing speech semantic consistency constraints, suppressing cross-scale inconsistencies caused by local audio injection, and reducing jitter, incomplete closure, and other phenomena.
[0049] 3. Interpolation upsampling + convolution is used instead of transposed convolution to reduce checkerboard artifacts and improve the clarity of lip boundaries and textures. Attached Figure Description
[0050] Figure 1 This is a flowchart illustrating the synchronous synthesis method for multi-scale audio-guided faces provided in this application embodiment;
[0051] Figure 2 This is a block diagram of a multi-scale audio-guided synchronous face synthesis method provided in an embodiment of this application. Detailed Implementation
[0052] To better understand the purpose, technical solutions, and advantages of this application, it has been described and illustrated below with reference to the accompanying drawings and embodiments. However, those skilled in the art should understand that this application can be implemented without these details. In some cases, to avoid obscuring various aspects of this application due to unnecessary description, well-known methods, processes, systems, components, and / or circuits already described at a higher level will not be elaborated upon. It will be apparent to those skilled in the art that various modifications can be made to the embodiments disclosed in this application, and the general principles defined in this application can be applied to other embodiments and application scenarios without departing from the principles and scope of this application. Therefore, this application is not limited to the illustrated embodiments, but conforms to the broadest scope consistent with the scope of protection claimed in this application.
[0053] refer to Figure 1 This application discloses a method for synchronous synthesis of multi-scale audio-guided faces, applied to a multi-scale audio-guided face synchronous synthesis system. This system incorporates an encoder-decoder generation network, with the encoder including a face encoder and an audio pyramid encoder. Target face frames and reference face frames are input to the face encoder and processed using a U-Net encoding structure to output skip connection features. Mel spectrum segments are input to the audio pyramid encoder for processing, outputting multi-scale audio features and global audio embeddings. Bottleneck fusion is performed on the multi-scale audio features and skip connection features, followed by U-Net-style multi-level decoding, ultimately outputting a 256×256 synthesized face image. It should be noted that the target face frame, reference face frame, and Mel spectrum segments are all obtained by processing video sequences.
[0054] Combination Figure 2 A method for synchronous synthesis of faces guided by multi-scale audio, comprising the following steps:
[0055] S100: Obtain face encoding data and corresponding audio encoding data, input the face encoding data into the face encoder for multi-layer feature extraction, obtain multiple skip connection features of different resolutions, and use the deepest skip connection feature as the bottleneck face feature.
[0056] The synchronous synthesis system receives face encoding data and inputs it into a face encoder for multi-layer feature extraction, obtaining multiple skip connection features at different resolutions. The face encoding data includes target face frames and reference face frames. Target face frames include unprocessed face frames or face frames after weakening, occluding, or masking the lip region. Reference face frames have complete mouths, providing prior information on lip appearance, teeth / oral boundaries, and the speaker's own lip texture. The face encoder can be a hierarchical convolutional network, such as a U-Net or ResNet-style encoder. The face encoder integrates and compresses these two visual information components into a bottleneck face feature F at the bottleneck layer. B And retain the skip connection features of the remaining layers.
[0057] During the encoding process, the synchronous synthesis system gradually expands the receptive field and compresses the spatial resolution through multi-layer convolution and downsampling, extracting jump connection features from shallow local textures to deep semantic structures, and using the deepest jump connection features as the bottleneck face feature F. B .
[0058] It's important to note that skip connection features refer to multi-scale visual feature maps extracted from the intermediate layers of the Face Encoder and directly passed to the corresponding layers of the Decoder to supplement spatial details and high-resolution information. During the encoding process, the skip connection features output at each downsampling stage (e.g., sizes of 128×128, 64×64, 32×32, 16×16, etc.) are preserved.
[0059] In one embodiment, the deepest skip connection feature is used as the bottleneck face feature, including the following steps:
[0060] S110, determine the target face frame and the reference face frame based on the face encoding data. The target face frame includes an unprocessed face frame or a face frame after being weakened, occluded, or masked by the lip region. The reference face frame includes a face frame containing the target's identity and appearance information.
[0061] S120 employs a hierarchical convolutional network structure, performing multi-layer convolution and sampling on the target face frame and reference face frame to progressively expand and compress the spatial resolution, extracting multi-scale jump connection features from shallow local textures to deep high semantic structures, with the deepest jump connection feature output as the bottleneck face feature.
[0062] A multi-scale audio-guided face synthesis system reads each frame of a video sequence in chronological order (e.g., 30fps), performs face detection on each frame (e.g., MTCNN, RetinaFace), and crops and aligns them to obtain a standardized face image (e.g., 256×256).
[0063] The first frame or several frames of the video sequence are detected, and a frame with a complete and unobstructed mouth and a frontal posture is selected as the reference face frame. The reference face frame contains the complete identity appearance of the person (lip texture, skin color, facial structure) and initial posture.
[0064] In one embodiment, the target face frame is obtained by using an occlusion detection module (such as a mouth visibility classifier based on key points or pixel variance analysis) to determine whether the mouth region of the current frame is occluded (e.g., by a mask, hand, or water cup). If occlusion exists, the frame is used as the target face frame. If no occlusion exists, the unprocessed current frame is directly used as the target face frame. To prevent feature misalignment between the reference face frame and the current target face frame due to large differences in pose and lighting, a sliding window or confidence update strategy can be used: if the mouth region of the current target face frame is completely unoccluded and the pose difference between it and the current reference face frame is less than a set threshold (e.g., Euler angle change < 15°), then the current frame is set as the new reference face frame.
[0065] In another embodiment, a reference frame queue is constructed to store the highest quality frame (unobstructed, high-resolution, frontal pose) among several frames. When the current frame is determined to be the target face frame, the latest valid reference frame in the queue is used as the corresponding reference face frame. For the current frame that is not occluded, it is also paired with a reference face frame during the training phase or in tasks requiring reconstruction and enhancement. In this case, the reference face frame can be the current frame itself or the previous frame.
[0066] In another embodiment, the face frame to be synthesized or redrawn is directly used as the target face frame, and the target face frame includes an unprocessed face frame or a face frame after being weakened, occluded, or masked by the lip area.
[0067] Before inputting the target face frame and the reference face frame into the hierarchical convolutional network structure (i.e., the encoder), an affine transformation is performed on both based on facial key points (such as 68 points or 106 points) to align the positions of the eyes and the tip of the nose to the same standard template, thereby making the identity texture in the reference face frame correspond spatially to the pose in the target face frame.
[0068] When processing video sequences in real-time or offline using a face-encoded data processing system, the occlusion status of the mouth region is first determined for each new frame: if the current frame is occluded, it is taken as the target face frame, and the most recently available reference face frame is read from the cache (this reference frame may be from several seconds ago). Then, the target face frame and the reference face frame are fed together into a hierarchical convolutional network structure; if the current frame is unoccluded and its quality (such as no occlusion, high definition, frontal pose) is higher than the quality of the current reference face frame, then the current frame is updated as a new reference face frame.
[0069] The hierarchical convolutional network structure performs multi-layer convolution and downsampling operations on the input target face frame and reference face frame, gradually compressing the spatial resolution, thereby extracting multi-scale jump connection features from shallow local texture to deep high semantic structure. The deepest jump connection feature output is used as the bottleneck face feature for subsequent face synthesis or redrawing tasks.
[0070] In a face-encoded data processing system, when a current frame is detected to be unobstructed and its quality (e.g., unobstructed, high-resolution, frontal pose) is higher than that of existing reference frames stored in the reference frame queue, the reference frame queue is updated so that subsequent occluded target face frames can obtain better reference information. During real-time or offline processing, the reference face frame may be taken from a historical frame several seconds prior, but due to the introduction of affine alignment operations and hierarchical convolutional feature extraction mechanisms, stable fusion of multi-scale visual features can still be guaranteed.
[0071] S200 performs time-frequency processing on the audio encoded data and performs multi-level convolutional sampling on the processed audio encoded data to construct multi-scale audio features.
[0072] The synchronous synthesis system performs time-frequency processing on the encoded audio data to generate corresponding Mel spectrum segments. These Mel spectrum segments are then input into an audio pyramid encoder for processing, where K-level multi-scale audio features {A} are obtained through multi-level convolutional downsampling. 1 ,...,A K}, where A K The deepest features are represented by multi-scale audio features, which are audio feature tensors with different time-frequency resolutions and / or different feature map sizes.
[0073] It should be noted that multi-scale audio features encompass feature maps at multiple levels, from high resolution to low resolution. As long as the resolution of the feature maps is different and the corresponding range of information they cover is different, they can be referred to as different scales.
[0074] In one embodiment, multi-level convolutional sampling is performed on the processed audio encoded data to construct multi-scale audio features, including the following steps:
[0075] S210 converts the audio encoded data into a time-frequency feature map, and obtains multiple multi-scale audio features with different time-frequency resolutions and / or different feature map sizes through multi-layer two-dimensional convolution and step-by-step downsampling.
[0076] The audio encoded data (e.g., Pulse Code Modulation (PCM) signal) corresponding in time to the video frame of the face to be synthesized is obtained. This audio encoded data is then converted to a frequency domain representation using a Short Time Fourier Transform (STFT), and further mapped to a Mel spectrogram via a Mel filter bank, resulting in a time-frequency feature map. This time-frequency feature map preserves the local time-frequency structure of the speech while conforming to the processing format of convolutional networks.
[0077] For example, the time-frequency feature map can be a Mel spectrogram, where the horizontal axis represents the time frame and the vertical axis represents the Mel frequency bin.
[0078] In another embodiment, the original audio waveform is pre-emphasized, framed (frame length 25ms, frame shift 10ms), and windowed (Hamming window). Then the power spectrum is calculated, and then the waveform is passed through a set of triangular Mel filters (e.g., 80 filters) to obtain 80-dimensional Mel features, finally forming a time-frequency feature map with shape T×80, where T is the number of audio time frames.
[0079] The audio pyramid encoder consists of multiple stacked two-dimensional convolutional blocks and downsampling layers. The decoder part contains multiple upsampling blocks, each with a feature map of a specific spatial size (e.g., H). j ×W j This corresponds to the width and height of the image space, such as 16×16, 32×32, etc. Each level of feature map has a different temporal-frequency resolution (corresponding to different spatial dimensions).
[0080] It's important to note that after the time-frequency feature map is downsampled by the audio pyramid encoder through convolution, what changes is the resolution of the entire time-frequency feature map. From the perspective of lip-syncing tasks, short-term phoneme and rhythmic changes are indeed reflected in the time dimension; high time resolution corresponds to small scales, while low time resolution corresponds to large scales.
[0081] Fine-scale multi-scale audio features refer to features with high time-frequency resolution and large feature map size, which can retain more short-term details. Coarse-scale multi-scale audio features refer to features with lower time-frequency resolution and smaller feature map size, covering a longer time range.
[0082] Furthermore, the distribution characteristics of the frequency dimension in the time-frequency feature map and the scale variation of the overall receptive field must be considered. The frequency axis of the time-frequency feature map is not equally weighted; the low-frequency region usually carries core speech information such as the fundamental frequency and formant structure, while the high-frequency region contributes more to details such as fricatives. The scale variation of the overall receptive field determines the trade-off between features for long-term context and short-term transients. By co-designing a non-uniform sampling strategy for the frequency dimension and a layer-by-layer expansion scale of the receptive field, multi-scale audio features can achieve effective matching in the three dimensions of time, frequency, and semantic depth, thereby improving the accuracy and robustness of phoneme-lip mapping.
[0083] For example, in speech signals, low-frequency regions (such as the fundamental frequency and the first and second formants) carry the main segment and prosodic information, while high-frequency regions (such as the high-frequency noise segments of fricatives and plosives) provide fine consonant distinguishing features. If the frequency distribution exhibits non-uniform sampling (e.g., logarithmic or Mel-scale), the low-frequency band will receive a denser representation, thereby enhancing the encoding ability of core information such as vowels and tones; conversely, if a linear uniform distribution is used, more high-frequency details are preserved, but it may cause blurring of key low-frequency features. Therefore, the distribution characteristics of the frequency dimension directly affect the discriminative power and noise resistance of the time-frequency feature map for different phoneme categories.
[0084] Time-frequency maps with small receptive fields (e.g., tens of milliseconds) can preserve the instantaneous jumps in short-term spectra, making them suitable for capturing local events such as plosives and transient fricatives. Time-frequency maps with large receptive fields (e.g., hundreds of milliseconds) aggregate spectral statistical characteristics over longer time windows, reflecting macroscopic information such as speech rate, intonation envelope, and intensity contours. By progressively expanding the receptive field, the time-frequency feature map in the deep network gradually transitions from a fine local spectral structure to an overall energy distribution pattern, thus balancing short-term phoneme recognition with long-term prosodic synchronization.
[0085] The S300 fuses the deepest multi-scale audio features with bottleneck facial features to obtain initial decoding features, and uses these initial decoding features as the starting point for face-synchronous synthesis.
[0086] The (lip) bottleneck facial feature F B With the deepest audio feature A K The initial features D are obtained by fusion. B The fusion method can be either convolution after splicing or summation. The bottleneck here is the facial feature F. BIt is a type of lip-related facial structure and texture feature. Bottleneck facial features include the facial contour of the current target face frame, the local geometric relationship around the mouth, the texture of the visible area, and supplementary appearance information provided by the complete lip area in the reference face frame. The initial decoding features provide a visual starting point for the decoder. When the synchronous synthesis system performs subsequent decoding, it lets the decoding network know what the person's face should look like and what the area around the mouth should look like. Then, it combines multi-scale audio features to generate synchronized lip movements.
[0087] Specifically, the face encoder extracts multi-scale visual features from the input image, including the bottleneck face feature F. B It is a high-semantic visual representation of the deepest output of the encoder, which includes at least structural information of the lip region and its neighborhood, local facial texture information, and identity appearance and pose context information provided by the current frame and the reference frame.
[0088] In one embodiment, the deepest multi-scale audio features are fused with bottleneck face features to obtain initial decoding features, including the following steps:
[0089] S310 performs channel mapping on the deepest multi-scale audio features, and then sums or concatenates the channel-mapped deepest multi-scale audio features with the bottleneck face features to determine the initial features for decoding.
[0090] After obtaining the initial decoding features, the decoder reconstructs the face image using a step-by-step upsampling method. The decoder consists of multiple decoder blocks stacked together. Each decoder block performs the following operations: first, it concatenates the skip features from the corresponding layer of the face encoder; then, it concatenates the audio features, which have undergone 1×1 convolution and maintain the same spatial resolution as the current decoder block, onto the skip features to obtain the initial decoding features.
[0091] S400 progressively increases the spatial resolution of the decoding features during the decoding process. In at least two decoding layers with different resolutions, multi-scale audio features are injected layer by layer into the current decoding layer features and corresponding skip connection features in the form of feature maps and then fused to obtain the decoding features of each decoding layer.
[0092] When the decoder upsamples step by step, at least two decoding layers with different resolutions, the skip connection features of the corresponding layer of the encoder (with the same spatial resolution) are directly copied and spliced onto the features of the current decoding layer. Multi-scale audio features are injected layer by layer into the features of the current decoding layer and the corresponding skip connection features in the form of feature maps and fused to obtain the decoding features of each decoding layer.
[0093] It should be noted that in this embodiment, the face encoder outputs multiple skip connection features with different spatial resolutions during multi-layer feature extraction. The deepest skip connection feature (lowest resolution, strongest semantics) is used as the bottleneck face feature, combined with audio features only during the bottleneck fusion stage. The skip connection features of other layers (e.g., from high resolution to medium resolution) are retained and, according to resolution correspondence, multi-scale audio features are injected layer by layer in the form of feature maps into the current decoding layer features and the corresponding skip connection features, and then fused to obtain the decoding features of each decoding layer. The retained skip connection features are used to supplement the spatial details corresponding to the synthesized face image during the fusion process. This design avoids information redundancy caused by the repeated injection of the deepest global semantic features, while ensuring that shallow detail features can effectively guide high-resolution lip reconstruction.
[0094] As the decoding spatial resolution is gradually improved, audio features of the corresponding resolution are injected into each decoding level, thereby ensuring that multi-scale audio information is fully utilized in the multi-level decoding process.
[0095] At each decoding level, a feature-linear modulation method is employed, utilizing global embedding to perform channel-level scaling and bias adjustment on the skip connection features of the current level. Before adjustment, the skip connection features from the corresponding level of the face encoder are concatenated along the channels with the audio features of the same resolution after 1×1 convolution processing to form a fused feature. Then, the modulation parameters generated by global embedding are applied to this fused feature.
[0096] The spatial resolution is restored step by step according to the upsampling hierarchy, and the features of the current decoding layer are obtained at each step. Its spatial dimensions are To achieve the selection of audio features corresponding to the encoder layer and scale in S400, the feature map is not simply selected from a fixed level of the audio pyramid. Instead, the selection of which level of audio feature map to use is determined based on the resolution stage of the current decoding layer and the coarseness of the content it is mainly responsible for recovering.
[0097] Specifically, when the decoding layer is at a low resolution stage (e.g., in the early stages of upsampling), the network mainly recovers the approximate shape of the mouth, its opening and closing trends, and the overall movement contour. At this time, coarse-scale audio feature maps (from deeper layers of the audio encoder) with a larger time coverage and more generalized semantics should be selected. As the decoding layer gradually increases to a higher resolution stage, the network further recovers details such as lip edges, closure states, and local lip shape changes. At this time, fine-scale audio features that preserve short-term fine structures should be selected, and these fine-scale audio features come from shallower layers of the audio encoder.
[0098] After selecting an audio feature map that matches the semantics of the current decoding layer, an injection operation is performed. First, channel mapping is performed on the audio feature map (for example, adjusting the number of channels using a 1×1 convolution to match the channel dimension of the decoding layer features). If the spatial size of the audio feature map at this time matches the channel dimension of the decoding layer features... If there is inconsistency, size alignment is performed using bilinear interpolation or nearest neighbor interpolation to obtain the processed audio features. Finally, Injecting decoding layer features The injection method can be selected as convolution after splicing along the channel. ←Conv([ , ]), or simply add them together ( ← + This injection operation occurs at least in two different resolution decoding layers and is repeated in multiple decoding layers, allowing the audio information to continue to function from low resolution to high resolution, avoiding the layer-by-layer attenuation of audio information caused by injecting only once in the bottleneck layer.
[0099] S500 obtains a global audio embedding based on multi-scale audio feature aggregation, generates scaling and bias parameters based on the global audio embedding, and applies them to the fused decoding features in a channel-level affine modulation manner to output a synthetic face image corresponding to the face encoding data.
[0100] While obtaining the initial decoding features, the synchronous synthesis system aggregates multi-scale audio features of different scales along the time dimension to obtain a global audio embedding vector. This global embedding is used to adjust the features of each channel during subsequent decoding. This global embedding, as a holistic semantic representation of the speech content, is input into a multilayer perceptron. The multilayer perceptron outputs two sets of parameters for each decoding layer (e.g., the j-th layer), namely scaling parameters. and bias parameters Both dimensions are the same as the number of channels of the feature in the decoding layer.
[0101] Channel-level affine modulation refers to independently scaling and translating each channel along the feature channel dimension to achieve global semantic constraints on the decoding layer features. Channel-level affine modulation can be applied to one or more decoding layers and works in conjunction with the step-by-step injection of audio features in step S400. Step S400 is responsible for injecting local audio features layer by layer into the decoding features according to spatial resolution, ensuring the accuracy of detail recovery. This embodiment adjusts the overall channel response of the decoding features based on global audio embedding, ensuring that the decoding results at different scales remain consistent in speech semantics.
[0102] After the above modulation, the decoded features are further restored to the spatial dimensions of the original face image through upsampling and convolution operations, and finally a synthetic face image corresponding to the input face encoding data is output, and the mouth movement of the image is synchronized with the audio signal.
[0103] In one embodiment, a global audio embedding is obtained based on multi-scale audio feature aggregation, including the following steps:
[0104] S510 determines the deepest audio feature tensor based on multi-scale audio features, and performs pooling and linear mapping on the deepest audio feature tensor to obtain the global audio embedding.
[0105] The audio pyramid encoder first converts the encoded audio data into a two-dimensional Mel spectrum segment, which forms a two-dimensional feature map with time as the horizontal axis and frequency as the vertical axis. Then, multiple convolutional layers and progressive downsampling are performed on this spectrum map, with the size of the output feature map decreasing layer by layer, resulting in a set of multi-scale audio features. The shallow feature maps retain high time-frequency resolution and rich detail, while the deeper feature maps are smaller in size but cover a larger receptive field.
[0106] For the deepest feature tensor, adaptive pooling (e.g., global average pooling) is first performed to compress the spatial dimension to 1×1. Then, a fully connected layer is used for linear mapping, ultimately outputting a fixed-dimensional global audio embedding vector. The function of this global embedding is not to replace the local audio conditions required by each layer of the decoder, but to provide a unified global modulation signal for each decoding layer.
[0107] During the decoding process, each DecoderBlock not only concatenates skip connection features from the corresponding level of the face encoder, but also concatenates audio feature maps of the same resolution (after adjusting the number of channels through 1×1 convolution). Simultaneously, it uses the aforementioned global audio embedding to generate linear modulation parameters (including scaling factors and bias terms) to adjust the channel dimension of the fused features at the current level.
[0108] Local audio features are injected by feature map stitching, while global audio embedding is injected by modulation signal. The two types of information provide fine spatial conditions and global unified control, respectively, avoiding the problem of losing details due to injecting an audio signal only once at the bottleneck layer.
[0109] In one embodiment, scaling and bias parameters are generated based on global audio embedding, and applied to the fused decoded features using channel-level affine modulation to output a synthetic face image corresponding to the face encoding data, including the following steps:
[0110] The S520 embeds global audio into the input multilayer perceptron, generating scaling / bias parameters that match the number of decoding layer channels.
[0111] S530 performs a channel-by-channel dot product of the decoding features of the decoding layer with the scaling parameters corresponding to the number of channels in the decoding layer, and adds an offset parameter to each channel to determine the modulation features.
[0112] S540 uses modulation features as decoding constraints for the decoding layer to output a synthetic face image.
[0113] For example, taking the j-th decoding layer as an example, the decoding features output by this layer contain multiple channels. Each channel can be regarded as a type of functional response. For example, some channels are sensitive to the edge of the lips, some are sensitive to the degree of mouth opening and closing, and some are sensitive to changes in texture or lighting. After the global audio embedding g is mapped by a multilayer perceptron, two sets of parameters with the same number of channels as the decoding layer are generated, and the scaling parameters are... and bias parameters .in, Each element controls the magnification or reduction factor of the corresponding channel. Each element controls the overall offset of the corresponding channel.
[0114] Will and Decoding features applied to the current decoding layer ,right For each channel, the modulation feature is obtained by first multiplying it point-by-point with the scaling parameter corresponding to that channel, and then adding the corresponding bias parameter to each channel. This modulation applies differentiated adjustments to different channels within the same decoding layer, selectively enhancing or suppressing each channel according to the overall content of the current speech—channels that match the semantics of the current speech are amplified, while those that do not are suppressed, supplemented by offset adjustments when necessary.
[0115] After modulation, the modulation features are used as decoding constraints for the current decoding layer, and upsampling and subsequent convolution operations are continued. The final decoding layer outputs a synthesized face image, preferably 256×256 pixels, but can also output an image of the lip region, deformation parameters, or intermediate feature representations depending on task requirements. This channel-level affine modulation can be applied to one or more decoding layers, allowing the audio signal to participate in constraining the decoding process from local fine registration to global semantic consistency.
[0116] In one embodiment, embedding global audio into an input multilayer perceptron includes the following steps:
[0117] The S521, a multilayer perceptron, is a small network consisting of at least two fully connected layers. Its input is a global audio embedding, and its output dimensions correspond to the scaling parameter and the bias parameter, respectively.
[0118] S522, the decoding layer consists of multiple upsampling layers with different resolutions in the decoder, and each decoding layer independently generates its own modulation parameters.
[0119] S523, channel-level affine transformation is performed after the skip connection fusion and / or scale-aligned audio injection of the decoding layer, and before entering the next decoding layer or output.
[0120] Combination Figure 1 In this embodiment, the multilayer perceptron employs a small network consisting of at least two fully connected layers. The input is a global audio embedding g, and the output dimensions correspond to the scaling parameter γ and the bias parameter β, respectively. The decoder contains multiple upsampling layers (i.e., multiple decoding layers) with different resolutions, and each decoding layer independently generates a set of modulation parameters (…). , ), used to perform channel-level affine transformations on the features of this layer.
[0121] Within each decoding layer, skip connection fusion is first performed, concatenating or adding the skip connection features from the corresponding layer of the face encoder with the decoding features of that layer. Next, scale-aligned audio injection is performed, selecting an audio feature map that matches the resolution of the decoding layer, mapping the number of channels through a 1×1 convolution, adjusting the spatial size using bilinear interpolation if necessary, and then injecting it into the current decoding features. After these two fusion steps, a channel-level affine transformation is performed, utilizing the features independently generated by that decoding layer. and The fused features are scaled (multiplied by channel) and offset (added by channel) channel by channel to obtain the modulation features. These modulation features are then used as the output of the current decoding layer and fed into the next decoding layer or directly into the final output layer.
[0122] Synchronization evaluation networks such as SyncNet and lip key point stability indicators can be used to evaluate audio-visual synchronization and lip movement stability.
[0123] In another embodiment, multi-scale audio features are injected layer by layer into the current decoding layer features and corresponding skip connection features in the form of feature maps and then fused to obtain the decoding features of each decoding layer. The injection method of multi-scale audio features includes the following steps:
[0124] The multi-scale audio features to be injected are processed by channel mapping, and the processed multi-scale audio features are interpolated, mapped or rearranged for size alignment.
[0125] After alignment, coarse-scale multi-scale audio features are injected into the low-resolution decoding layer, while fine-scale multi-scale audio features are injected into the high-resolution decoding layer.
[0126] To conditionally inject audio feature maps into the corresponding scale of the decoding layer, the time-frequency two-dimensional dimensions of the audio feature maps must be aligned to the spatial dimensions of the decoding feature maps. This alignment can be flexibly achieved through interpolation, mapping, or rearrangement projection to feature layers of different resolutions in the decoder. These feature maps, ranging from fine-scale (high time-frequency resolution, suitable for providing local phoneme constraints) to coarse-scale (low time-frequency resolution, providing global speech semantics), form an audio pyramid, laying the foundation for subsequent scale-aligned audio injection and global audio embedding extraction.
[0127] This embodiment provides the following three alignment methods, including interpolation, mapping, or rearrangement projection. In this embodiment, these three methods can be used individually or in combination. The specific implementation of these three methods adopts existing methods, which will not be elaborated on here.
[0128] In the actual decoding process, coarse-scale audio feature maps are injected into the low-resolution decoding layer (corresponding to the recovery of the general trend of lip movements), while fine-scale audio feature maps are injected into the high-resolution decoding layer (corresponding to the recovery of lip details), ensuring that phoneme-level information still has direct constraints during the detail reconstruction stage. The temporal window of the audio features can cover the speech content before and / or after the current moment, providing temporal context for decoding. Upsampling in the decoding layer preferentially adopts an interpolation-based upsampling combined with convolution, instead of pure transposed convolution, to reduce checkerboard artifacts introduced by upsampling. The face encoder and decoder can use a U-Net or ResNet-style hierarchical convolutional network, as long as the feature combination of multi-scale aligned injection and global audio modulation is satisfied.
[0129] It should be noted that when the decoder contains multi-level upsampling layers from low resolution to high resolution, coarse-scale audio feature maps can be injected into the low-resolution decoding layer, and fine-scale audio feature maps can be injected into the high-resolution decoding layer, so that phoneme-level information still has direct constraints in the detail reconstruction stage.
[0130] In another embodiment, before injecting multi-scale audio features layer by layer into the current decoding layer features and the corresponding skip connection features in the form of feature maps and fusing them, the following steps are also included:
[0131] The current decoding layer features are interpolated and upsampled to obtain an upsampled feature map. The upsampled feature map is then subjected to at least one convolutional layer to obtain decoding layer features with learnable transformations of channel data and spatial features.
[0132] During the decoding process, at least two decoding layers with different resolutions in step S400 need to undergo interpolation upsampling and convolution processing on the features of the current decoding layer to improve its spatial resolution. Interpolation upsampling combined with convolution is used instead of transposed convolution during decoding to reduce artifacts introduced by upsampling. Specifically, for the feature map of the current decoding layer... (Spatial size H×W), the spatial size is enlarged to a target multiple (usually 2 times) using bilinear interpolation or nearest-neighbor interpolation, resulting in an upsampled feature map of 2H×2W. .
[0133] Next, the convolution performs feature smoothing and adaptive adjustment on the upsampled feature map. Apply at least one convolutional layer (e.g., 3×3 convolution + BNReLU) to perform a learnable transformation on its channel count and spatial features. Then use the processed decoding layer features to complete step S400, which involves injecting multi-scale audio features as feature maps layer by layer into the current decoding layer features and corresponding skip connection features, and then fusing them.
[0134] Interpolation is a parameter-free deterministic transformation that does not introduce any learnable weights, filling new pixels solely based on a weighted average (bilinear) or copy (nearest neighbor) of surrounding pixels. Because the interpolation kernel is fixed and smooth, it does not produce checkerboard overlap artifacts.
[0135] After interpolation, resampling and convolution are performed. The convolution kernel can further adjust the local structure, enhance details, or suppress residual interpolation blur based on the smooth transition after upsampling. Since convolution is performed on a regular grid without missing pixels, the uneven overlap problem of transposed convolution is not reproduced.
[0136] In another embodiment, face detection / alignment, audio Mel spectrum segment extraction, and sliding window slicing can be directly reused from open-source implementations for comparison. The face encoder / decoder can be any hierarchical convolutional network (such as U-Net / ResNet style) to provide multi-scale features. The number of channels, layers, and injection layer positions can be configured according to engineering needs, but should at least satisfy the necessary feature combination for multi-scale alignment injection and global audio modulation.
[0137] This application also discloses a multi-scale audio-guided synchronous face synthesis system. The multi-scale audio-guided synchronous face synthesis system includes a face processing module, an audio processing module, a bottleneck fusion module, a feature decoding module, and a feature modulation module.
[0138] The face processing module acquires face encoding data and corresponding audio encoding data, inputs the face encoding data into the face encoder for multi-layer feature extraction, and obtains multiple skip connection features with different resolutions. The deepest skip connection feature is used as the bottleneck face feature.
[0139] The audio processing module performs time-frequency processing on the audio encoded data and performs multi-level convolution sampling on the processed audio encoded data to construct multi-scale audio features. The multi-scale audio features represent audio feature tensors with different time-frequency resolutions and / or different feature map sizes.
[0140] The bottleneck fusion module is used to fuse the deepest multi-scale audio features with bottleneck face features to obtain the initial decoding features, which serve as the fusion starting point for face synchronous synthesis during the decoding process.
[0141] During the decoding process, the feature decoding module progressively increases the spatial resolution of the decoded features. In at least two decoding layers with different resolutions, multi-scale audio features are injected layer by layer into the current decoding layer features and corresponding skip connection features in the form of feature maps and then fused to obtain the decoded features of each decoding layer.
[0142] The feature modulation module obtains a global audio embedding based on the summarization of multi-scale audio features, generates scaling and bias parameters based on the global audio embedding, and applies them to the fused decoded features in a channel-level affine modulation manner to output a synthetic face image corresponding to the face encoding data.
[0143] The other functions performed in the face processing module, audio processing module, bottleneck fusion module, feature decoding module, and feature modulation module, as well as the technical details of each function, are the same as or similar to the corresponding features in the multi-scale audio-guided face synchronization synthesis method described above, so they will not be repeated here.
[0144] The implementation principle is as follows:
[0145] In each Decoder Block, the skip connection features from the corresponding layer of the face encoder are first concatenated (or added) with the current decoded features to inject visual details at that resolution.
[0146] By using a global embedding vector, feature linear modulation (FiLM) is applied to the concatenated feature map for channel-by-channel scaling and bias adjustment, thereby enabling the audio information to adaptively adjust the response of each channel.
[0147] Meanwhile, an additional bottleneck fusion module is set up before the lowest-level decoding block. This module takes the initial decoding features (i.e., the result obtained by adding or splicing the deepest audio features after channel mapping and the bottleneck face features) as input and further fuses them with the low-resolution features inside the decoding block to ensure that the audio signal is continuously transmitted from coarse to fine granular throughout the entire decoding network.
[0148] It should be understood that although the steps in the flowcharts in the accompanying drawings are shown sequentially as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise expressly stated herein, there is no strict order in which these steps are performed, and they may be performed in other orders.
[0149] The above are all preferred embodiments of this application, and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.
Claims
1. A method for synchronous synthesis of multi-scale audio-guided faces, characterized in that, Includes the following steps: Acquire face encoding data and corresponding audio encoding data, input the face encoding data into a face encoder for multi-layer feature extraction, obtain multiple skip connection features of different resolutions, and determine the target face frame and reference face frame based on the face encoding data. The target face frame includes an unprocessed face frame or a face frame after weakening, occluding, or masking the lip region. The reference face frame includes a face frame containing the target's identity and appearance information. A hierarchical convolutional network structure is adopted to perform multi-layer convolution and sampling on the target face frame and the reference face frame to gradually expand and compress the spatial resolution. Multi-scale jump connection features from shallow local texture to deep high semantic structure are extracted, and the deepest jump connection feature output is used as the bottleneck face feature. The audio encoded data is subjected to time-frequency processing, and the processed audio encoded data is subjected to multi-level convolutional sampling to construct multi-scale audio features. The multi-scale audio features represent audio feature tensors with different time-frequency resolutions and / or different feature map sizes. The deepest multi-scale audio features are fused with the bottleneck face features to obtain the initial decoding features, which serve as the fusion starting point for face synchronization synthesis during the decoding process. During the decoding process, the spatial resolution of the decoding features is gradually increased. In at least two decoding layers with different resolutions, the multi-scale audio features are injected layer by layer into the current decoding layer features and the corresponding skip connection features in the form of feature maps and then fused to obtain the decoding features of each decoding layer. A global audio embedding is obtained based on the summarization of the multi-scale audio features. Scaling parameters and bias parameters are generated according to the global audio embedding and applied to the fused decoding features in a channel-level affine modulation manner to output the synthetic face image corresponding to the face encoding data.
2. The method for synchronous synthesis of multi-scale audio-guided faces according to claim 1, characterized in that, The processed audio encoded data is subjected to multi-level convolutional sampling to construct multi-scale audio features, including the following steps: The audio encoded data is converted into a time-frequency feature map, and multi-scale audio features with different time-frequency resolutions and / or different feature map sizes are obtained through multi-layer two-dimensional convolution and stepwise downsampling.
3. The method for synchronous synthesis of multi-scale audio-guided faces according to claim 2, characterized in that, The global audio embedding is obtained based on the summarization of the multi-scale audio features, including the following steps: Based on the multi-scale audio features, the deepest audio feature tensor is determined, and pooling and linear mapping are performed on the deepest audio feature tensor to obtain the global audio embedding.
4. The method for synchronous synthesis of multi-scale audio-guided faces according to claim 1, characterized in that, The deepest multi-scale audio features are fused with the bottleneck face features to obtain the initial decoding features, including the following steps: The deepest multi-scale audio features are channel-mapped, and then the channel-mapped deepest multi-scale audio features are summed or concatenated with the bottleneck face features to determine the initial features for decoding.
5. The method for synchronous synthesis of multi-scale audio-guided faces according to claim 4, characterized in that, The multi-scale audio features are injected layer by layer into the current decoding layer features and corresponding skip connection features in the form of feature maps and then fused to obtain the decoding features of each decoding layer. The injection method of the multi-scale audio features includes the following steps: The multi-scale audio features to be injected are subjected to channel mapping processing, and the processed multi-scale audio features are then interpolated, mapped, or rearranged for size alignment. After alignment, coarse-scale multi-scale audio features are injected into the low-resolution decoding layer, while fine-scale multi-scale audio features are injected into the high-resolution decoding layer.
6. The method for synchronous synthesis of multi-scale audio-guided faces according to claim 5, characterized in that, Scaling and bias parameters are generated based on the global audio embedding, and channel-level affine modulation is applied to the fused decoded features to output a synthetic face image corresponding to the face encoding data, including the following steps: The global audio is embedded into the multilayer perceptron to generate scaling / bias parameters that match the number of channels in the decoding layer; The decoding features of the decoding layer are multiplied one channel at a time by the scaling parameters corresponding to the number of channels of the decoding layer, and the bias parameters are added one channel at a time to determine the modulation features; The modulation features are used as decoding constraints for the decoding layer to output a synthetic face image.
7. The method for synchronous synthesis of multi-scale audio-guided faces according to claim 6, characterized in that, Embedding the global audio into the multilayer perceptron includes the following steps: The multilayer perceptron comprises a small network with at least two fully connected layers, whose input is a global audio embedding and whose output dimensions correspond to scaling parameters and bias parameters, respectively. The decoding layer consists of multiple upsampling layers with different resolutions in the decoder, and each decoding layer independently generates its own modulation parameters; The channel-level affine transformation is performed after the skip connection fusion and / or scale-aligned audio injection of the decoding layer is completed, and before entering the next decoding layer or before output.
8. The method for synchronous synthesis of multi-scale audio-guided faces according to claim 1, characterized in that, Before injecting the multi-scale audio features into the current decoding layer features and corresponding skip connection features layer by layer in the form of feature maps and fusing them, the following steps are also included: The current decoding layer features are interpolated and upsampled to obtain an upsampled feature map, and the upsampled feature map is subjected to at least one convolutional layer to obtain decoding layer features with channel data and spatial features that can be learned and transformed.
9. A multi-scale audio-guided synchronous face synthesis system, characterized in that, The method for synchronous synthesis of a multi-scale audio-guided face according to any one of claims 1-8 includes: The face processing module acquires face encoding data and corresponding audio encoding data, inputs the face encoding data into the face encoder for multi-layer feature extraction, and obtains multiple skip connection features of different resolutions. The deepest skip connection feature is used as the bottleneck face feature. An audio processing module performs time-frequency processing on the audio encoded data and performs multi-level convolution sampling on the processed audio encoded data to construct multi-scale audio features. The multi-scale audio features represent audio feature tensors with different time-frequency resolutions and / or different feature map sizes. The bottleneck fusion module is used to fuse the deepest multi-scale audio features with the bottleneck face features to obtain the initial decoding features, which serve as the fusion starting point for face synchronization synthesis during the decoding process. The feature decoding module progressively increases the spatial resolution of the decoding features during the decoding process. In at least two decoding layers with different resolutions, the multi-scale audio features are injected layer by layer into the current decoding layer features and the corresponding skip connection features in the form of feature maps and then fused to obtain the decoding features of each decoding layer. The feature modulation module obtains a global audio embedding based on the summarization of the multi-scale audio features, generates scaling parameters and bias parameters based on the global audio embedding, and applies them to the fused decoded features in a channel-level affine modulation manner to output the synthetic face image corresponding to the face encoding data.
Citation Information
Patent Citations
Face image generation method and device, equipment, medium and product
CN115761075A
Conversation head animation synthesis method and related device
CN121837463A