Face synthesis method, device and equipment based on multi-modal information interaction

Through the multimodal information interaction method, audio and face features are extracted and aligned, target voice and video are reconstructed, which solves the problem of poor synchronization in voice-driven face synthesis and improves the quality and synchronization of video synthesis.

CN120163907APending Publication Date: 2025-06-17ACADEMY OF BROADCASTING SCI STATE ADMINISTATION OF PRESS PUBLICATION RADIO FILM & TELEVISION
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510191950.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-17

AI Technical Summary

Technical Problem

The existing voice-driven face synthesis technology is difficult to achieve good synchronization between video frames and audio, affecting the quality and practicality of the synthesis effect.

Method used

Using a method based on multimodal information interaction, the audio timing and semantic features are extracted by receiving audio clips and face images, the facial features are aligned with the bidirectional cross attention algorithm and deep learning algorithm, the multimodal features are fused and the target voice and video are decoded and reconstructed, and the synchronization is optimized by combining multimodal identification and dynamic time regularization algorithm.

Benefits of technology

It significantly improves the synchronization between video frames and audio, improves the quality of synthetic videos, and enables the generated video to accurately match the rhythm, semantics and emotions of the audio.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120163907A_ABST
    Figure CN120163907A_ABST
Patent Text Reader

Abstract

The invention relates to a face synthesis method based on multi-modal information interaction, and relates to the technical field of video processing. The multi-modal information interaction-based face synthesis method comprises the steps of receiving audio clips and face images corresponding to the audio clips respectively; based on the frequency information of the audio clips, extracting audio time sequence features corresponding to the audio clips, and extracting audio semantic features from the audio clips through a machine learning algorithm; fusing the audio time sequence features and the audio semantic features through a bidirectional cross attention algorithm to obtain phonetic sequence semantic features; extracting facial features corresponding to the face image through a deep learning algorithm, and aligning the facial features with the audio semantic features to obtain facial tone semantic features; and fusing the phonetic sequence semantic feature and the facial tone semantic feature to obtain a joint feature, decoding and reconstructing the joint feature, and converting the joint feature into a target voice video. By implementing the method provided by the invention, the synchronism between the video frame and the audio can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The disclosed embodiments relate to the field of video processing technology, and more specifically, to a face synthesis method, device, electronic device and storage medium based on multimodal information interaction. Background Art

[0002] In the field of computer vision and graphics, face synthesis technology has a long and fruitful history. In the early days, face synthesis mainly relied on geometry-based models, based on two-dimensional images, and used simple geometric features to generate faces, such as the application of principal component analysis (PCA). With the advancement of technology, the rise of texture mapping and three-dimensional modeling technology, the research focus has shifted to 3D face modeling, which can more accurately reconstruct facial details.

[0003] In recent years, the technology of converting audio into face video has received widespread attention. It is of great significance in reducing video encoding and transmission bandwidth and helping people with hearing loss to obtain audio information. Face synthesis technology is divided into image-driven, facial key point-driven and voice-driven methods according to different driving sources. However, the existing voice-driven face synthesis technology has obvious defects. It is difficult to achieve good synchronization between the generated video frames and the input audio, which seriously affects the quality and practicality of the synthesis effect. It has become a key problem that needs to be solved in the current field of face synthesis technology. Summary of the invention

[0004] An objective of the embodiments of the present disclosure is to provide a new technical solution for a face synthesis method based on multimodal information interaction.

[0005] According to a first aspect of the present disclosure, a face synthesis method based on multimodal information interaction is provided, characterized in that it includes: receiving an audio clip and a face image corresponding to the audio clip respectively; extracting audio timing features corresponding to the audio clip based on the frequency information of the audio clip, and extracting audio semantic features from the audio clip by a machine learning algorithm, the audio timing features being used to represent the law of change of the audio clip over time; fusing the audio timing features and the audio semantic features by a bidirectional cross-attention algorithm to obtain phonetic semantic features; extracting facial features corresponding to the face image by a deep learning algorithm, aligning the facial features with the audio semantic features to obtain facial-voice semantic features; fusing the phonetic semantic features and the facial-voice semantic features to obtain joint features, decoding and reconstructing the joint features, and converting the joint features into a target voice video.

[0006] Optionally, after fusing the phonetic sequence semantic features and the facial-audio semantic features to obtain joint features, decoding and reconstructing the joint features, and converting the joint features into a target speech video, the method further includes: calculating a video difference between the target speech video and a sample video through a multimodal discrimination algorithm, and determining whether the target speech video meets a preset video display standard according to the calculated video difference; calculating a phase difference between the target speech video and a sample corresponding to the sample video through a dynamic time warping algorithm; if the target speech video does not meet the preset video display standard and / or the phase difference is greater than a preset phase difference, inputting the video difference and / or the phase difference into a preset loss function, and calculating an adjustment parameter through the preset loss function to update the target speech video through the adjustment parameter.

[0007] Optionally, aligning the facial features and the audio semantic features to obtain the facial-audio semantic features includes: determining whether the audio semantic features and the facial features come from the same semantic space through a cross-modal feature matching and classification algorithm; if they come from the same semantic space, adjusting the feature distribution of the facial features through a cross-modal feature alignment encoder algorithm to align the feature distribution of the facial features with the feature distribution of the audio semantic features, thereby obtaining the facial-audio semantic features.

[0008] Optionally, the bidirectional cross-attention algorithm is: where Softmax represents a preset activation function, Q represents a query vector generated according to the audio temporal features, K represents a key vector generated according to the audio semantic features, V represents a value vector generated according to the audio semantic features, QK T represents the matrix multiplication of the transpose of Q and K, represents a preset scaling factor, and Attn(Q, K, V) represents the phonetic sequence semantic features.

[0009] Optionally, the preset loss function is: L Total = λ1L LSGAN + λ2L CGAN + λ3L align + λ4L DTW.phase + λ5L Feat , where LTotal represents the total loss value, L LSGAN represents the difference calculated by the algorithm for calculating the facial features and the sample facial features, λ1 represents the weight value corresponding to L LSGAN corresponding, L CGAN represents the difference calculated by the algorithm for matching the facial features based on the audio temporal features, λ2 represents the weight value corresponding to L CGAN corresponding, Lalign represents the loss value calculated by the algorithm for aligning the facial features and the audio semantic features, λ3 represents the weight value corresponding to Lalign, L DTW.phase represents the difference calculated by the algorithm for calculating the audio-visual phase difference, λ4 represents the weight value corresponding to L DTW.phase corresponding, LFeat represents a preset constraint feature, and λ5 represents the weight value corresponding to Lt. Fea

[0010] Optionally, Lalign = E[logDalign(Fsemantic,Fface)] + E[log(1 - Dalign(Fsemantic,Fface.fake))], where E represents a preset mathematical expectation, Dalign represents a cross-modal feature matching classification algorithm, Fsemantic represents audio semantic features, Fface represents facial features, and Fface.fake represents face-audio semantic features.

[0011] Optionally, before receiving the audio segment and the face image corresponding to the audio segment, the method further includes: obtaining a voice audio and the face image corresponding to the voice audio; removing background noise of the voice audio through a denoising algorithm to obtain the voice audio, where the background noise is interfering sound of non-target voice signals mixed in during the audio acquisition process.

[0012] According to a second aspect of the present disclosure, there is also provided a face synthesis device based on multi-modal information interaction, characterized in that the device includes: a receiving module for receiving an audio segment and the face images respectively corresponding to the audio segment; a feature extraction module for extracting audio temporal features corresponding to the audio segment based on the frequency information of the audio segment, and extracting audio semantic features from the audio segment through a machine learning algorithm, where the audio temporal features are used to represent the law of change of the audio segment over time; a feature fusion module for fusing the audio temporal features and the audio semantic features through a bidirectional cross-attention algorithm to obtain audio-sequence semantic features; a feature alignment module for extracting facial features corresponding to the face image through a deep learning algorithm, aligning the facial features with the audio semantic features to obtain face-audio semantic features; and a video generation module for fusing the audio-sequence semantic features and the face-audio semantic features to obtain combined features, decoding and reconstructing the combined features, and converting the combined features into a target voice video.

[0013] According to a third aspect of the present disclosure, there is also provided an electronic device including a memory and a processor, where the memory is used to store a computer program; the processor is used to execute the computer program to implement the method according to the first aspect of the present disclosure.

[0014] According to a fourth aspect of the present disclosure, there is also provided a computer-readable storage medium having a computer program stored thereon, where the computer program, when executed by a processor, implements the method according to the first aspect of the present disclosure.

[0015] ​One beneficial effect of the embodiments of the present disclosure is that various features are extracted from audio segments and face images respectively to obtain key information of speech and face. The bidirectional cross-attention algorithm is used to fuse audio temporal and semantic features to strengthen the correlation within the audio modality; the facial features and audio semantic features are aligned to prevent the generated face from deviating from the input identity due to audio dominance, improving the semantic consistency between facial movements and speech content. Then, the multi-modal features are fused and decoded to reconstruct the target speech video, enabling the synthesized video to accurately match the rhythm, semantics, and emotions of the audio, improving the synchronization between video frames and audio, and significantly enhancing the quality of the synthesized video.

[0016] Other features and advantages of the embodiments of the present disclosure will become clear through the following detailed description of the exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings incorporated in and constituting a part of this specification illustrate embodiments of the present disclosure and, together with the description, are used to explain the principles of the embodiments of the present disclosure.

[0018] Figure 1 is an interaction schematic diagram of a face synthesis system capable of applying the multi-modal information interaction-based system according to one embodiment;

[0019] Figure 2 is a flowchart of face synthesis based on multi-modal information interaction according to one embodiment;

[0020] Figure 3 is a flowchart of face synthesis based on multi-modal information interaction according to another embodiment;

[0021] Figure 4 is a schematic diagram of the principle of a face synthesis device based on multi-modal information interaction according to one embodiment;

[0022] Figure 5 is a schematic diagram of the hardware structure of an electronic device according to one embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] Now, various exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings. It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and values set forth in these embodiments do not limit the scope of the present disclosure.

[0024] The following description of at least one exemplary embodiment is merely illustrative in nature and in no way limits the present disclosure, its application, or its use.

[0025] Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be considered as part of the specification.

[0026] In all examples shown and discussed herein, any specific values should be construed as merely exemplary, and not as limitations. Thus, other examples of the exemplary embodiments may have different values.

[0027] It should be noted that like reference numerals and letters refer to like items in the following figures, and thus, once an item is defined in one figure, further discussion thereof is not required in subsequent figures.

[0028] Face synthesis algorithms are an important research direction in computer vision and computer graphics, aiming to generate, synthesize, or modify face images through computer technology. With the rapid development of artificial intelligence and deep learning, face synthesis has gradually evolved from traditional geometric modeling and image processing to high-quality image generation using advanced generative models and deep neural networks. The research on face synthesis not only promotes innovative applications in fields such as virtual reality, film production, and social networks, but also plays an important role in multiple fields such as biometrics, security monitoring, and psychological research.

[0029] The research on face synthesis technology can be traced back to the 1980s. With the initial development of the computer vision field, researchers began to explore how to automatically generate or modify face images using computers. In the initial stage, the synthesis methods mainly relied on geometry-based models, such as simple geometric modeling and feature extraction of faces. The algorithms of this period were mainly based on two-dimensional images, and faces were generated through simple geometric features (such as the positions of eyes, nose, and mouth). A typical example is the application of PCA in face recognition and synthesis. PCA represents and synthesizes faces by extracting the main feature components of faces (such as "eigenfaces"). With the development of computer hardware and the continuous progress of image processing technology, between the 1990s and 2000s, texture mapping and 3D modeling technologies developed rapidly. People began to use more complex technologies for facial feature synthesis and 3D face modeling. At this time, the research focus shifted to synthesis through 3D face models, and 3D face modeling allows for more accurate reconstruction of facial details, including factors such as expressions, lighting, and perspective changes.

[0030] However, the real technological breakthroughs occurred in the 2010s, especially with the rise of deep learning technologies. The application of deep learning models, particularly convolutional neural networks and generative adversarial networks, has significantly improved the quality of face synthesis technology. Through deep neural networks, computers can learn the complex features of human faces from a large amount of real face data, and the generated synthetic images have reached unprecedented levels in terms of details, realism, and diversity. With the development of these technologies, face synthesis in virtual characters, virtual reality applications, and social platforms has become increasingly common and the effects are becoming more and more realistic.

[0031] In recent years, researchers have proposed technical methods for converting audio into face videos. This technology is of great scientific significance and also has a wide range of practical applications. By converting audio into high-quality videos, the bandwidth for video encoding and transmission can be effectively reduced, which is extremely important for most of the content that occupies the Internet transmission bandwidth. In addition, this technology can also help those with hearing impairments to achieve post-reading, thus making it easier for them to obtain audio information and improving their quality of life. Post-reading is a task of decoding text from the mouth movements of the speaker, which is very helpful for people with hearing impairments. Face synthesis technology can be mainly divided into the following three categories according to the different driving sources: picture-driven methods, facial key-point-driven methods, and speech-driven methods.

[0032] Picture-driven methods (Image-Driven Methods, IDM) are methods for face synthesis based on existing image data. These methods generate or modify face images by analyzing and utilizing a large amount of picture data. Different from model-based methods (such as 3D face models), picture-driven methods focus on directly extracting features or styles from images and then generating or transforming the target face image. Its main feature is that it directly depends on the dataset of real images for training and generation. The basic idea of picture-driven methods is to use a large amount of real image data to learn various features of human faces, such as the basic structure of the face, expression changes, texture details, etc. Then, the system can automatically generate new images according to the input conditions (such as the target age, gender, expression, etc.). Picture-driven methods are commonly used in tasks such as face expression synthesis, style transfer, and image conversion. However, in existing methods, the correspondence between the generated video frames and the input audio cannot be well synchronized.

[0033] Facial Landmark Driven Methods (FLD) are a class of technologies based on facial landmark detection and analysis, mainly used for tasks such as facial image processing, expression synthesis, facial recognition, and pose estimation. This method drives the subsequent image synthesis or transformation process by detecting and utilizing the key points of the human face (such as the feature points of the eyes, nose, mouth, chin, etc.). Facial key points usually refer to the key areas on the human face with obvious positions and shapes, which are usually represented by two-dimensional coordinates in the image. These key points are very important for understanding facial expressions, postures, age, gender, and other characteristics, so they play a crucial role in many computer vision applications. The basic idea of the facial landmark driven method is to accurately detect the positions of the key points of the human face in the image, and then guide the subsequent image generation, deformation, or analysis tasks. Common application scenarios include: facial expression synthesis, facial animation, virtual makeup, facial beautification, and image enhancement, etc.

[0034] Voice-driven Methods (VDM) are a type of technology that drives or controls other tasks or generates targets based on voice input (i.e., sound, audio signal). In multiple fields such as computer vision, image generation, virtual humans, and natural language processing, voice-driven methods achieve multimodal interaction and generation by converting voice signals into corresponding actions, expressions, or images, etc. In short, voice-driven methods refer to the technology of controlling, generating, or modifying other media (such as images, animations, virtual characters, etc.) through voice signals. Voice-driven methods usually involve combining the content of audio signals (such as the audio features, intonation, emotions, etc. of speech) with other types of data (such as facial expressions, actions, image styles, etc.) to achieve an association with the voice content, and then drive certain changes. This method combines technologies such as speech recognition, speech emotion analysis, sound feature extraction, and computer vision, allowing computer systems to interact with users more naturally and intuitively. The core of voice driving includes: speech feature extraction, multimodal learning, generation models, and speech synchronization.

[0035] Time-frequency features and their analysis are important concepts in signal processing, widely used in fields such as speech processing, audio signal analysis, vibration monitoring, and biomedical signal processing. The core idea of time-frequency analysis is to study the features of a signal in both the time domain and the frequency domain simultaneously, aiming to reveal the time-varying law and frequency components of the signal, so as to provide richer information than traditional time-domain or frequency-domain analysis. The goal of time-frequency feature extraction is to extract representative time-frequency information from the original signal, which is usually used for subsequent classification, recognition, reconstruction, or analysis tasks. Compared with traditional time-domain or frequency-domain methods, time-frequency analysis can effectively solve the problem of processing non-stationary signals, that is, the problem that the spectrum of the signal changes with time.

[0036] Time-frequency analysis methods transform signals in two dimensions, time and frequency, to reveal the local characteristics of signals in time and frequency. Common time-frequency analysis tools include the Short-Time Fourier Transform (STFT), Wavelet Transform (WT), Hilbert-Huang Transform (HHT), and other methods based on local analysis. Each method has its unique advantages and application scenarios. STFT is one of the most commonly used time-frequency analysis methods. STFT divides the signal into multiple short time segments and performs Fourier transform on each segment separately to obtain the spectral information of the signal within each time window. Although STFT can provide a joint representation of time and frequency, its time-frequency resolution is fixed, which means its accuracy is limited when analyzing non-stationary signals. Especially for signals with rapidly changing frequencies or transient characteristics, the effect of STFT is often not ideal. WT is another widely used time-frequency analysis method. Compared with STFT, wavelet transform has adjustable time and frequency resolutions, which makes it more flexible in dealing with non-stationary signals. Wavelet transform convolves the signal with mother wavelet functions of different scales, which can adapt to different frequency changes. It can provide higher time resolution for the high-frequency part of the signal and higher frequency resolution for the low-frequency part. Wavelet transform is especially suitable for signals with mutation, pulse, or transient characteristics, such as biomedical signals and seismic signals. HHT is a time-frequency analysis method based on Empirical Mode Decomposition (EMD). The EMD method decomposes the signal into multiple Intrinsic Mode Functions (IMFs), and each IMF represents different scale components of the signal. Then, the instantaneous frequency of each IMF is calculated through Hilbert transform, and finally the time-frequency distribution of the signal is obtained. The HHT method is very effective in dealing with non-linear and non-stationary signals and can accurately capture the instantaneous change characteristics of the signal. Therefore, it has a wide range of applications in vibration analysis, biological signal processing, and other fields. In addition to these common time-frequency analysis methods, there are many other time-frequency analysis algorithms, such as Wigner-Ville distribution, Choi-Williams distribution, etc. These methods can play a role in high-resolution time-frequency analysis, but often face the "cross-term" problem, that is, unnecessary pseudo-frequency components are generated in high-resolution analysis, which affects the accuracy of the analysis results. In time-frequency feature extraction, the commonly used algorithms usually rely on the results of the above time-frequency analysis methods and combine machine learning and pattern recognition techniques to extract and classify the features of the signal. Time-frequency features usually include instantaneous frequency, band energy, spectral entropy, amplitude spectrum, etc., which can reflect various characteristics of the signal in time and frequency.These features have been widely applied in applications such as speech recognition, audio signal classification, biomedical signal diagnosis, and mechanical fault monitoring. In signal processing tasks, the goal of time-frequency feature extraction is to extract the local features of a signal in time and frequency domains, and then use them for signal analysis, classification, and prediction. Through time-frequency analysis, not only can the frequency components of a signal be obtained, but also the patterns of frequency variation over time can be captured. Therefore, when dealing with non-stationary signals (such as speech, music, seismic signals, electrocardiograms, etc.), time-frequency analysis methods demonstrate their unique advantages. With the development of advanced algorithms such as deep learning, time-frequency feature extraction technology has also been continuously improving, capable of more accurately extracting multi-level and multi-scale features of signals, providing a powerful tool for fields such as signal analysis, fault detection, and intelligent diagnosis.

[0037] <System Embodiment>

[0038] Figure 1 Shown is a schematic diagram of a face synthesis system based on multi-modal information interaction. As Figure 1 shown, the system includes a speech-video generation model and a speech-video discrimination model. Among them, the speech-video generation model is used to implement a model for generating corresponding face videos based on multi-modal information interaction. Through a series of processing and transformation of the relevant features of the input audio and face images, the target speech-video is finally output; the speech-video discrimination model is used to evaluate the quality of the target speech-video. By comparing with sample videos and other methods, it determines whether the target speech-video meets the preset standards. Specifically, the system can receive an audio clip and the corresponding face image as input data. The audio clip is used to extract audio temporal features and audio semantic features, while the face image is used to extract facial features; the audio temporal features and audio semantic features are fused to form audio-sequence semantic features, and the facial features are aligned with the audio semantic features to obtain face-audio semantic features; the audio-sequence semantic features and face-audio semantic features are further fused into joint features, and the joint features are decoded and reconstructed to generate the target speech-video; then, the target speech-video is compared and evaluated with the sample video through multi-modal discrimination algorithms and dynamic time warping algorithms, and the difference information obtained from the evaluation is fed back to the previous processing flow to optimize the entire speech-driven face synthesis process.

[0039] For example, a face synthesis system based on multimodal information interaction receives an audio clip describing an exciting scene and a face image of a person with a neutral expression. The audio clip is input into an audio sequence encoding algorithm to extract audio temporal features, and at the same time, a machine learning algorithm is used to extract audio semantic features. Facial features are extracted from the face image through a deep learning algorithm. First, the audio temporal features and audio semantic features are fused into audio sequence semantic features through a bidirectional cross-attention algorithm. Then, the facial features and audio semantic features are processed through a cross-modal feature matching classification algorithm and a cross-modal feature alignment encoder algorithm to obtain face-audio semantic features. Finally, the audio sequence semantic features and face-audio semantic features are fused to form joint features. Decoding and reconstruction operations are performed on the joint features to generate a target speech video. A multimodal discrimination algorithm is used to compare the target speech video with a sample video, calculate the video difference, and determine whether it meets the preset video display standard. The phase difference between the target speech video and the sample video is calculated through a dynamic time warping algorithm to evaluate the synchronization of audio and video. If the target speech video does not meet the preset video display standard or the phase difference is greater than the preset value, these differences are fed back into the system to adjust and optimize the generated target speech video, such as fine-tuning the amplitude of facial expressions and improving the synchronization of audio and video, until the requirements are met.

[0040] <Method Embodiment>

[0041] Figure 2 is a schematic flowchart of a face synthesis method based on multimodal information interaction according to an embodiment, which can be applied to any terminal device capable of executing this method.

[0042] As Figure 2 shown, the face synthesis method based on multimodal information interaction in this embodiment may include the following steps S210 - S250:

[0043] Step S210, receive an audio clip and a face image corresponding to the audio clip respectively.

[0044] The audio clip is used to represent a sound signal containing speech content. For example, the audio clip includes sound components with different frequencies that can reflect the intonation changes of speech and the specific content expressed by the speech.

[0045] The face image is used to represent the initial visual information of the face in the synthesized video. For example, the face image contains the geometric structure features and emotional states of the face, such as the contour shape of the face, the relative positions, proportions, and expression features of the facial features.

[0046] In one embodiment, before receiving the audio clip and the face image corresponding to the audio clip, the method further includes: obtaining the speech audio and the face image corresponding to the speech audio; removing the background noise of the speech audio through a denoising algorithm to obtain the speech audio, where the background noise is the interfering sound of non-target speech signals mixed in during the audio acquisition process.

[0047] The speech audio is used to represent a continuous sound signal containing language information, including sound frequency, pitch change, rhythm characteristics, semantic content, etc. Optionally, the audio clip can be segmented into multiple audio clips according to a preset duration to obtain multiple audio clips, where each audio clip in the multiple audio clips has overlapping audio with the previous adjacent audio clip in chronological order. That is, the audio clip is used for

[0048] The face image is used to represent a set of image collections that have a one-to-one correspondence with the speech audio. Each face image in it precisely matches the speech audio in a specific time period, and this correspondence is constructed based on the time dimension. For example, in a 1-minute speech audio, each second of the speech has a corresponding face image.

[0049] The denoising algorithm is used to represent a technical means that can identify and remove the interfering sound of non-target speech signals mixed in during the audio acquisition process. Optionally, the denoising algorithm can be any algorithm that can remove the interfering sound. In this disclosure, the dynamic threshold denoising algorithm is taken as an example for illustration.

[0050] Exemplarily, assume that speech is collected in a noisy conference room, and the collected speech audio is mixed with background noises such as surrounding conversations and the operation sound of the air conditioner. After obtaining the speech audio containing these background noises, the dynamic threshold denoising algorithm is used to analyze the time-domain signal of the speech audio, dynamically determine the threshold according to the energy distribution of the audio signal, for the signal part below the threshold, judge it as a possible noise signal and suppress or remove it; while for the signal part above the threshold, consider it as a useful speech signal and retain it, so as to obtain the speech audio and the face image.

[0051] In this embodiment, first obtaining the speech audio and the corresponding face image set, and then using the denoising algorithm to remove the background noise can improve the quality of the speech audio and make the synthesized target speech video more coordinated in the matching of speech and facial movements.

[0052] Optionally, in the present disclosure, after collecting the speech audio, the speech audio can be segmented into multiple audio segments, or after removing the noise signal of the speech audio, the speech audio can be segmented into multiple audio segments. It should be understood that after the speech audio is segmented into multiple audio segments, each audio segment among the multiple audio segments has its corresponding face image. For example, a 1-minute speech audio is segmented into 3 segments. The first segment of audio has a duration of 0-21 seconds, and its corresponding face image is the expression state image of the speaker at the starting moment within these 21 seconds. As the speech content progresses, the facial expression will be adjusted accordingly according to the audio semantics and emotions. The second segment of audio has a duration of 20-41 seconds, and the corresponding face image starts to change from the facial state of the speaker at 20 seconds, which is closely related to the speech rhythm and semantic expression of this segment of audio. For example, when the speaker emphasizes a certain point of view, the facial expression will be more vivid, and the movements of parts such as the eyes, eyebrows, and mouth are synchronized with the audio. The third segment of audio has a duration of 40-60 seconds, and the corresponding face image will also show corresponding changes according to the content of this segment of audio. If the speech expresses an exciting emotion at the end, the speaker in the face image may show a more passionate expression, such as an upward-curved mouth and a firm look in the eyes.

[0053] Hereinafter, the present application takes processing a segment of audio as an example.

[0054] Step S220: Based on the frequency information of the audio segment, extract the audio time-series features corresponding to the audio segment, and extract the audio semantic features from the audio segment through a machine learning algorithm. The audio time-series features are used to represent the law of change of the audio segment over time.

[0055] The frequency information is used to represent the distribution characteristics of the audio signal at different frequencies and the change of frequency over time.

[0056] Audio time-series features: Features used to characterize the law of change of an audio segment over time. For example, it can be the dynamic characteristics of the audio signal in the time dimension, such as the change of pitch, rhythm, etc. over time.

[0057] Audio semantic features: Features extracted from an audio segment that represent the information at the audio semantic level, covering the content of speech expression, emotional tendency, etc., and can reflect the semantic information such as the meaning and emotion conveyed by the audio.

[0058] In the present disclosure, audio temporal features corresponding to an audio clip are extracted through Mel-Wavelet Transform (MWT) and Gated Recurrent Unit (GRU); the machine learning algorithm for extracting audio semantic features from the audio clip is One-Dimensional Convolutional Neural Network (1D-CNN).

[0059] Exemplarily, assume that we have an audio clip containing a conversation between people, and the content is that two people are discussing their weekend plans. One person says, "Let's go hiking this weekend. I heard the scenery on the mountain is especially beautiful." When extracting audio temporal features, MWT and GRU are adopted. MWT transforms the audio signal from the time domain to the Mel frequency domain and the wavelet scale domain to obtain multi-scale and multi-resolution time-frequency features. For example, MWT can decompose signals with different frequency components in the audio, such as the fundamental frequency and harmonic components of speech, at different scales, obtaining a series of wavelet coefficients reflecting the local features of different frequencies and times; these wavelet coefficients are used as the input of GRU, and GRU will process these sequence data to capture the dynamic change law of the audio signal in the time dimension. In the above-mentioned dialogue audio, when the word "weekend" is pronounced, GRU can learn the changing trend of the frequency over time during the pronunciation of this word, such as the pitch first rising and then falling, as well as the rhythm and duration of this change, etc. These information combined constitute the audio temporal features; when extracting audio semantic features, 1D-CNN performs a convolution operation on the audio clip through the convolutional layer, and the convolutional kernel slides in the time dimension of the audio signal to automatically extract local features in the audio. For example, the convolutional kernel can capture the features of keywords such as "hiking" and "beautiful scenery", as well as the positions and mutual relationships of these keywords in the audio. Then, through the pooling layer, the features are compressed and the dimensionality is reduced, reducing the computational amount while retaining important features. Finally, after the processing of the fully connected layer, a feature vector representing the audio semantics is output. This vector contains the content expressed by the audio, that is, the weekend plan to go hiking and the description of the scenery on the mountain, as well as semantic information such as the possible expected emotional tendency of the speaker.

[0060] Step S230: Fuse the audio temporal features and the audio semantic features through a bidirectional cross-attention algorithm to obtain audio sequence semantic features.

[0061] In one embodiment, the bidirectional cross-attention algorithm is: where Softmax represents a preset activation function, Q represents a query vector generated according to the audio temporal features, K represents a key vector generated according to the audio semantic features, V represents a value vector generated according to the audio semantic features, and QK T represents the matrix multiplication of Q and the transpose of K, It represents a preset scaling factor, and Attn(Q, K, V) represents the phonetic sequence semantic features.

[0062] Exemplarily, the attention from time series to semantics: When saying "climbing a mountain", the audio time series features (such as increasing frequency and accelerating speech rate) generate the query vector Q, and the audio semantic features (including the semantic information of "climbing a mountain") generate the key vector K and the value vector V. Calculate the matrix multiplication QK of Q and the transpose of K T , and then divide it by the preset scaling factor After being processed by the Softmax activation function, multiply it by the value vector V. The semantic features related to "climbing a mountain" obtain higher weights in the weighted calculation, highlighting the semantics of "climbing a mountain" to obtain time series-enhanced semantic features; and, taking the semantic features containing "climbing a mountain" as the query vector Q, the audio time series features related to "climbing a mountain" generate the key vector K and the value vector V, and continue to calculate according to the above formula, so that the weights of the audio time series features closely related to the semantics of "climbing a mountain" are increased, such as making the features of increasing frequency and accelerating speech rate more prominent when "climbing a mountain", to obtain semantics-guided time series features; integrate the features after fusing the two directions to obtain the phonetic sequence semantic features.

[0063] Based on the formula of the above-mentioned bidirectional cross-attention algorithm, the attention scores can be scaled to an appropriate range, enabling the model to more accurately capture the correlations between features when fusing audio time series and semantic features, effectively avoiding the weight imbalance caused by dimensionality problems, and optimizing the synthesis effect of the target speech and video.

[0064] In another embodiment, it is also possible to calculate the cosine similarity between the audio time series feature vector and the audio semantic feature vector. The higher the similarity, the stronger the correlation between the two features. Correlate the features whose similarity exceeds the preset threshold to generate the phonetic sequence semantic features. Taking the audio of "going climbing on the weekend" as an example, assign weights to the semantic features according to the similarity. The semantic features of "climbing a mountain" have a greater weight in the weighted summation because of their high similarity with the corresponding time series features, thereby achieving feature fusion and generating the phonetic sequence semantic features.

[0065] Step S240: Extract the facial features corresponding to the face image through a deep learning algorithm, and align the facial features with the audio semantic features to obtain the face-audio semantic features.

[0066] Facial features: Used to represent the information in the face image. For example, the facial features can include the geometric structure information of the face (such as the facial contour, the positions of the facial features, etc.) and the emotional state information (such as the joys, sorrows, anger, and pleasures conveyed by the expression), which are used to describe the appearance and emotional expression of the face.

[0067] The deep learning algorithm can be any deep learning algorithm. In this disclosure, a residual network is taken as an example for illustration.

[0068] In one embodiment, step S240 includes determining whether the audio semantic feature and the facial feature come from the same semantic space through a cross-modal feature matching and classification algorithm; if they come from the same semantic space, the feature distribution of the facial feature is adjusted through a cross-modal feature alignment encoder algorithm to align the feature distribution of the facial feature with the feature distribution of the audio semantic feature, thereby obtaining a face-audio semantic feature.

[0069] Exemplarily, assume that in a video call scenario, the audio clip is "We won this competition!", and at the same time, the face image of the speaker is obtained. The facial features corresponding to the face image are extracted using a residual network to obtain the geometric structure information of the face, such as the contour lines of the face, the positions and shapes of the eyes, nose, and mouth, as well as the captured emotional state information, such as the happy expression shown by the speaker at this moment due to winning the competition, features such as the upturned corners of the mouth and narrowed eyes; these facial features are aligned with the previously extracted audio semantic features (the semantic content of "winning the competition" and the excited emotional tendency), and a cross-modal feature matching and classification algorithm is used to determine whether the audio semantic feature and the facial feature come from the same semantic space. Since the happy facial expression of the speaker is consistent with the excited semantics of "winning the competition", the feature distribution of the facial feature is adjusted through a cross-modal feature alignment encoder algorithm to align the feature distribution of the facial feature with the feature distribution of the audio semantic feature. For example, the feature weight regarding the happy expression in the facial feature is increased to better match the excited semantic feature in the audio, thereby obtaining a face-audio semantic feature.

[0070] In this embodiment, by using a deep learning algorithm to extract facial features and align them with audio semantic features, the facial features in the synthesized video can be accurately matched with the audio semantics, enhancing the expressiveness of the video content.

[0071] In another embodiment, based on a feature extraction algorithm for geometric shapes and textures, geometric features such as the contour of the face and the proportion of facial features, as well as skin texture features, are obtained. The audio semantic features are represented using a bag-of-words model, and then by calculating the Euclidean distance or cosine similarity between the two feature vectors, etc., the most matching feature combination is found to align the facial features and the audio semantic features.

[0072] Step S250: Fuse the phonetic sequence semantic feature and the face-audio semantic feature to obtain a combined feature, perform decoding and reconstruction on the combined feature, and convert the combined feature into a target speech video.

[0073] Decoding and reconstruction: The operation of processing the combined feature, the specific process of converting the combined feature into a target speech video, through which the abstract feature information is converted into visual face video content.

[0074] Target speech video: The final output of the speech video generation model is a face video corresponding to a person's speech, which is generated after a series of processes based on the relevant features of the input audio and face image.

[0075] Exemplarily, after obtaining the phonetic sequence semantic features corresponding to the audio "We won this competition!" and the face phonetic semantic features of the speaker's face image, and fusing these two features, the semantics of "winning the competition" in the phonetic sequence semantic features, the audio rhythm and frequency change information are combined with the features related to the happy expression of the face and the face geometric structure features in the face phonetic semantic features to form joint features. The decoder decodes and reconstructs the joint features. The decoder can be a structure based on a convolutional neural network, which converts the abstract information in the joint features into visual image elements. For example, according to the facial expression related information in the joint features, details such as the opening degree of the mouth and the changes in the eyes under the corresponding expression are generated; then, according to the speech rhythm and semantic information, the amplitude and frequency of the face movements are adjusted, such as the slight shaking of the head when speaking and the change of the expression with the key points of the speech. Finally, these elements are combined into a series of coherent pictures to form the target speech video. In a video call, what the receiving party sees is a face video that matches the speech content and emotion of the speaker, with the speaker's face beaming with joy and the mouth opening and closing in sync with the speech.

[0076] Multiple features are extracted from the audio segment and the face image respectively, comprehensively obtaining the key information of the speech and the face. The bidirectional cross-attention algorithm is used to fuse the audio time series and semantic features, strengthening the internal modality correlation of the audio and providing more reliable audio features for subsequent processing. The facial features and audio semantic features are aligned to prevent the generated face from deviating from the input identity due to audio dominance, improving the semantic consistency between the lip movement and the speech content. Finally, the multi-modal features are fused and decoded and reconstructed into the target speech video, enabling the synthesized video to accurately match the rhythm, semantics and emotion of the audio, greatly optimizing the synchronization between the video frames and the audio, significantly improving the quality of the synthesized video, and enhancing the practicality of the face synthesis technology in practical applications such as virtual reality, film and television production, and assisting communication for hearing-impaired people.

[0077] The face synthesis method based on multi-modal information interaction provided by the embodiments of the present disclosure extracts multiple features from the audio segment and the face image respectively, obtains the key information of the speech and the face, uses the bidirectional cross-attention algorithm to fuse the audio time series and semantic features, strengthening the internal modality correlation of the audio; aligns the facial features and audio semantic features to prevent the generated face from deviating from the input identity due to audio dominance, improving the semantic consistency between the facial movement and the speech content, and then fuses the multi-modal features and decodes and reconstructs them into the target speech video, enabling the synthesized video to accurately match the rhythm, semantics and emotion of the audio, improving the synchronization between the video frames and the audio, and significantly improving the quality of the synthesized video.

[0078] In one embodiment, after step S250, it may further include: calculating a video difference between the target voice video and the sample video through a multimodal discrimination algorithm, and determining whether the target voice video meets a preset video display standard according to the calculated video difference; calculating a phase difference between the target voice video and the sample corresponding to the sample video through a dynamic time warping algorithm; if the target voice video does not meet the preset video display standard and / or the phase difference is greater than a preset phase difference, inputting the video difference and / or the phase difference into a preset loss function, and calculating an adjustment parameter through the preset loss function to update the target voice video through the adjustment parameter.

[0079] Multimodal discrimination algorithm: An algorithm used to calculate the video difference between the target voice video and the sample video, which compares the two from multiple modalities (such as vision, audio, etc.) to determine whether the target voice video meets the preset video display standard.

[0080] Dynamic time warping algorithm: In a voice video discrimination model, an algorithm used to calculate the phase difference between the target voice video and the sample corresponding to the sample video, which measures the synchronization of audio and video in time.

[0081] Sample video: A video used as a reference standard, which is used to compare with the target voice video to evaluate the quality of the target voice video and determine whether it meets the preset video display standard and audio-video synchronization standard, etc.

[0082] Exemplarily, the generated target voice video is "Let's go on a vacation by the sea". After generating the target voice video, a multi-modal discrimination algorithm is used to calculate the video difference between the target voice video and the sample video. From the visual modality, the multi-modal discrimination algorithm will compare the facial expressions, action postures, etc. of the two. For example, in the sample video, when the person says "by the sea", the corners of the mouth slightly turn up and the eyes show anticipation, while in the target voice video, the expression of the person when saying "by the sea" is relatively plain, which will generate a certain visual difference; from the audio modality, the intonation, speech rate, etc. of the audio will be compared. If the intonation of the pronunciation of "vacation" in the sample video is lighter, while the intonation in the target voice video is relatively plain, an audio difference will also be generated. Combining the comparison results of these multiple modalities, a comprehensive video difference is obtained, and based on this difference, it is judged whether the target voice video meets the preset video display standard; the dynamic time warping algorithm is used to calculate the phase difference between the target voice video and the corresponding sample of the sample video. For example, in the sample video, when the person opens their mouth to say "Let's", the audio also starts to emit the sound of "Let's" synchronously, while in the target voice video, the mouth-opening action is a short time (0.3 seconds) ahead of the audio pronunciation. The dynamic time warping algorithm will calculate this time difference, that is, the phase difference, to measure the synchronization of the audio and video in time; if after judgment, it is found that the facial expression of the person in the target voice video is too plain and does not meet the preset video display standard, and the phase difference between the audio and video is greater than the preset phase difference (such as the time difference between the mouth movement and the audio pronunciation exceeds 0.2 seconds), then the video difference (difference information in aspects such as facial expression, action, audio intonation, etc.) and the phase difference are input into the preset loss function. The preset loss function calculates adjustment parameters based on these input values. For example, adjust the relevant parameters for generating facial expressions to make the corners of the mouth of the person in the target voice video turn up more when saying "vacation by the sea", and at the same time adjust the synchronization parameters of the audio and video to reduce the time difference between the mouth movement and the audio pronunciation. The target voice video is updated through these adjustment parameters to make it more in line with the preset standard.

[0083] In this embodiment, after generating the target voice video, the quality is judged by the multi-modal discrimination algorithm, the audio-visual synchronization is evaluated by the dynamic time warping algorithm. When the video does not meet the standard or the phase difference is too large, the preset loss function is used to calculate the adjustment parameters to update the video, which can improve the quality of the target voice video.

[0084] In an alternative embodiment, the preset loss function is: L Total = λ1L LSGAN + λ2L CGAN + λ3L align + λ4L DTW.phase + λ5L Fea t, where L Total represents the total loss value, L LSGANrepresents the difference calculated by the algorithm for calculating the facial features and the sample facial features, and λ1 represents L LSGAN the corresponding weight value, L CGAN represents the difference calculated by the algorithm for matching facial features based on audio temporal features, and λ2 represents L CGAN the corresponding weight value, L align represents the loss value calculated by the algorithm for aligning facial features and audio semantic features, and λ3 represents L align the corresponding weight value, L DTW.phase represents the difference calculated by the algorithm for calculating the audio-visual phase difference, and λ4 represents L DTW.phase the corresponding weight value, L Feat represents the preset constraint feature, and λ5 represents L Feat the corresponding weight value.

[0085] Exemplarily, the target speech-video content is "Let's go on a vacation by the sea". When calculating the facial feature differences, in the sample video, when the person says "by the sea on vacation", the corners of the mouth turn up significantly and the eyes sparkle with anticipation. In the target speech-video, the corners of the person's mouth turn up less and the expression in the eyes is not as vivid. Through the algorithm for calculating the facial features and the sample facial features, the difference L1 between the two is obtained. Assume the set weight value λ1 = 0.3; Based on the audio temporal features, in the sample video, when saying "Let's go", the speaking speed is slightly faster, and at the same time, there are corresponding excited changes in the facial expression, such as the eyebrows slightly raising. However, in the target speech-video, when the audio rhythm changes, the facial movements do not keep up synchronously. The difference L2 is calculated through the algorithm for matching facial features based on audio temporal features. Assume the set weight value λ2 = 0.2; In terms of aligning facial features and audio semantic features, the semantics of "by the sea on vacation" in the sample video convey emotions of joy and yearning, and the facial expression highly matches it. In the target speech-video, the facial expression does not convey this emotion well enough. The loss value L3 is calculated through the algorithm. Assume the set weight value λ3 = 0.2; For the audio-visual phase difference, in the sample video, the opening and closing movements of the mouth are completely synchronized with the audio pronunciation. In the target speech-video, when saying "Let's" with the mouth open, it is slightly later than the audio pronunciation. The difference L4 is obtained through the algorithm for calculating the audio-visual phase difference. Assume the set weight value λ4 = 0.2; The preset constraint feature L5, for example, requires that the facial contour of the human face maintains a certain proportion and shape in the video. Here, it is found that the facial contour of the person in the target speech-video is slightly distorted at certain angles. The corresponding result is obtained according to the algorithm for the preset constraint feature, and the weight value λ5 = 0.1. The total loss value LTotal = 0.3×0.5 + 0.2×0.3 + 0.2×0.4 + 0.2×0.2 + 0.1×0.1 = 0.34.

[0086] In this embodiment, the preset loss function comprehensively considers factors such as facial features, audio temporal features, audio semantic features, and audio-visual phase differences. By setting different weight values, it can flexibly balance the influence of each factor on the total loss value, enabling comprehensive and accurate problem discovery during the evaluation and optimization of the target speech video, thereby more effectively calculating adjustment parameters and improving the quality of the target speech video.

[0087] Optionally, Lalign = E[logDalign(Fsemantic,Fface)] + E[log(1 - Dalign(Fsemantic,Fface.fake))], where E represents the preset mathematical expectation, Dalign represents the cross-modal feature matching classification algorithm, Fsemantic represents the audio semantic feature, Fface represents the facial feature, and Fface.fake represents the face-audio semantic feature.

[0088] Combined with the above example, the audio semantic feature has been extracted from the audio, which contains semantic information such as the relaxed and pleasant emotional information and specific activity content conveyed by "beach vacation". At the same time, the facial feature has been extracted from the corresponding face image, such as the amplitude of the mouth corner rising and the expression of slightly squinted eyes when the person says this sentence. Set the preset mathematical expectation E as a measurement standard. For example, it is expected that the audio semantic feature and the facial feature can be highly matched to a certain extent to generate a high-quality face-audio semantic feature. Use the cross-modal feature matching classification algorithm Dalign, take the audio semantic feature Fsemantic and the facial feature Fface as inputs, and the algorithm analyzes these two different-modal features to determine the matching relationship between the audio semantic feature Fsemantic and the facial feature Fface. For example, when the "pleasant" emotion is conveyed in the audio semantic feature, the amplitude of the mouth corner rising in the facial feature has a high correlation with it, and the face-audio semantic feature Fface.fake is obtained. Then, compare the obtained result with the preset mathematical expectation E. If the difference between the face-audio semantic feature Fface.fake and the preset mathematical expectation E is small, it indicates that the matching effect between the audio semantic feature and the facial feature is good, and the generated face-audio semantic feature meets the expectation; if the difference is large, the algorithm or input features need to be adjusted to improve the matching effect.

[0089] In this embodiment, by introducing the preset mathematical expectation to provide a measurement standard for the matching of the audio semantic feature and the facial feature, the generated face-audio semantic feature is more predictable and controllable. And by processing the audio and facial features of two different modalities through the cross-modal feature matching classification algorithm, a face-audio semantic feature with more semantic consistency can be generated, thereby improving the fit between the audio and the picture in the target speech video, enabling the facial expression of the person to more accurately reflect the semantics and emotions conveyed by the audio.

[0090] In another embodiment, there is no need to preset the mathematical expectation and the cross-modal feature matching classification algorithm. The audio semantic features and the facial features are used as the input to the autoencoder model. The encoder encodes these two features and compresses them into a low-dimensional feature representation. In the decoder part, by decoding the low-dimensional feature representation, the original audio semantic features and facial feature information are restored, and at the same time, the face-audio semantic features are generated. By adjusting the parameters of the autoencoder, the difference between the reconstructed features and the original input features is minimized.

[0091] In another example, a loss function based on perceptual similarity is used to replace the preset loss function. For the target speech video "Let's go on a vacation by the sea", using a pre-trained perceptual model (such as the VGG network pre-trained on large-scale image and video data, etc.), the target speech video and the sample video are respectively input into the perceptual model to extract features at different levels, such as obtaining basic features such as edges and textures from the shallow layer, and obtaining abstract features related to semantics from the deep layer. Then, the differences between the features extracted from different levels are calculated, such as calculating the Euclidean distance or cosine similarity between the feature vectors, etc., as the loss value.

[0092] To facilitate the understanding of the face synthesis method based on multi-modal information interaction provided by the present disclosure, the following is combined with Figure 3 for illustration.

[0093] Downloaded a 60 - second speech video of participant A on a shared video platform. The speech video includes A's voice audio and the corresponding sequence of face images. Due to the presence of noise in the actual environment, such as the environmental noise in the meeting room and the slight conversations of others, first, a denoising algorithm, such as the dynamic threshold denoising algorithm, is used to process the voice audio to remove background noise and improve the voice quality. After that, the denoised voice audio is segmented, divided into segments of 15 seconds each, and there is a 5 - second overlap between adjacent audio segments, resulting in 4 audio segments: Segment 1 (0 - 15 seconds, where 0 - 5 seconds overlaps with Segment 2), Segment 2 (10 - 25 seconds, where 10 - 15 seconds overlaps with Segment 1 and 20 - 25 seconds overlaps with Segment 3), Segment 3 (20 - 35 seconds, where 20 - 25 seconds overlaps with Segment 2 and 30 - 35 seconds overlaps with Segment 4), Segment 4 (30 - 45 seconds). Each audio segment has a corresponding sequence of face images. For example, Segment 1 corresponds to a series of face images of A within 0 - 15 seconds, and these images reflect A's facial expressions and pose changes during this period and are closely related to the audio content. Use MWT and GRU to extract the audio temporal features of each audio segment. For example, in Segment 1, A says "Hello everyone, today we are going to discuss an important project". MWT converts the audio signal to the Mel frequency domain and the wavelet scale domain to obtain multi - scale time - frequency features, and GRU captures the variation rules of speech rate and intonation in this sentence. For example, the intonation of "Hello everyone" is relatively flat, and the intonation of "important project" rises and the speech rate slows down. These variation information constitute the audio temporal features. At the same time, use 1D - CNN to extract audio semantic features, identify key semantic information such as "important project" and A's emotional tendency when speaking. Then, through a deep learning algorithm (such as a residual network), extract facial features from the face images corresponding to each audio segment. These features include geometric structure information such as A's facial contour and the positions of facial features, as well as emotional state information conveyed by expressions, such as smiling indicating friendliness and frowning indicating thinking when A is speaking. Use the bidirectional cross - attention algorithm to fuse the audio temporal features and audio semantic features. Taking the example that A mentions "The key to the project lies in the innovative solution" in Segment 2, from time series to semantics, the audio temporal features (such as the intonation change when emphasizing "key") generate the query vector Q, and the audio semantic features ("key", "innovative solution" and other semantic information) generate the key vector K and the value vector V. After calculation according to the formula, the semantic features related to "key" and "innovative solution" are highlighted to obtain time - series enhanced semantic features. From semantics to time series, these semantic features are used as the query vector Q, and the relevant audio temporal features generate the key vector K and the value vector V. After calculation, the weights of the relevant audio temporal features are increased. Finally, the audio - sequence semantic features are integrated.Subsequently, a cross-modal feature matching classification algorithm is used to determine whether the audio semantic features and facial features come from the same semantic space. If A shows an excited expression when mentioning the "innovative solution" and the semantics of the two match, then the feature distribution of the facial features is adjusted through the cross-modal feature alignment encoder algorithm to align the facial features with the audio semantic features, obtaining face-audio semantic features. The audio sequence semantic features and face-audio semantic features are fused to obtain joint features. The joint features are decoded and reconstructed, and converted into a target speech video to generate a series of coherent pictures based on the joint features, completing the accurate matching of A's facial expressions, lip movements, and speech content. For example, when saying "innovation", the lip movement of the mouth is synchronized with the audio, and the excited facial expression also matches the semantics. After generating the target speech video, the video difference between the target speech video and the sample video (such as the standard video of a high-quality video conference) is calculated through a multi-modal discrimination algorithm, and comparisons are made from multiple aspects such as the visual modality (such as the naturalness and clarity of facial expressions) and the audio modality (such as the clarity of speech and the rationality of intonation) to determine whether the target speech video meets the preset video display standard. At the same time, the dynamic time warping algorithm is used to calculate the phase difference between the target speech video and the sample video to evaluate the synchronization of audio and video. If it is found that a certain expression of A in the target speech video is not natural enough, or there is a slight asynchrony between the lip movement and the speech, these differences are input into a preset loss function (such as L; Total = λ1L LSGAN + λ2L CGAN + λ3L align + λ4L DTW.phase + λ5L Feat ) to calculate the adjustment parameters and update the target speech video to make the video quality higher and more in line with the display requirements of the target speech video.

[0094] <Example of Device One>

[0095] Figure 4 is a schematic block diagram of an electronic device according to an embodiment. As Figure 4As shown, the electronic device 400 may include a receiving module 410 for receiving an audio segment and a face image corresponding to the audio segment respectively; a feature extraction module 420 for extracting an audio temporal feature corresponding to the audio segment based on the frequency information of the audio segment, and extracting an audio semantic feature from the audio segment through a machine learning algorithm, where the audio temporal feature is used to represent the law of change of the audio segment over time; a feature fusion module 430 for fusing the audio temporal feature and the audio semantic feature through a bidirectional cross-attention algorithm to obtain a phonetic sequence semantic feature; a feature alignment module 440 for extracting a facial feature corresponding to the face image through a deep learning algorithm, aligning the facial feature with the audio semantic feature to obtain a face-audio semantic feature; and a video generation module 450 for fusing the phonetic sequence semantic feature and the face-audio semantic feature to obtain a combined feature, decoding and reconstructing the combined feature, and converting the combined feature into a target speech video.

[0096] In an optional implementation manner of the embodiment of the present disclosure, the video generation module 450 is further configured to calculate a video difference between the target speech video and a sample video through a multimodal discrimination algorithm, and determine whether the target speech video meets a preset video display standard according to the calculated video difference; calculate a phase difference between the target speech video and a sample corresponding to the sample video through a dynamic time warping algorithm; if the target speech video does not meet the preset video display standard and / or the phase difference is greater than a preset phase difference, input the video difference and / or the phase difference into a preset loss function, and calculate an adjustment parameter through the preset loss function to update the target speech video through the adjustment parameter.

[0097] In an optional implementation manner of the embodiment of the present disclosure, the feature alignment module 440 is further configured to determine whether the audio semantic feature and the facial feature come from the same semantic space through a cross-modal feature matching classification algorithm; if they come from the same semantic space, adjust the feature distribution of the facial feature through a cross-modal feature alignment encoder algorithm to align the feature distribution of the facial feature with the feature distribution of the audio semantic feature to obtain a face-audio semantic feature.

[0098] In an optional implementation manner of the embodiment of the present disclosure, the receiving module 410 is further configured to obtain a voice audio and a face image corresponding to the voice audio; remove background noise of the voice audio through a denoising algorithm to obtain the voice audio, where the background noise is an interfering sound of a non-target voice signal mixed in during the audio acquisition process.

[0099] <Device Embodiment II>

[0100] Figure 5 It is a schematic hardware structure diagram of an electronic device according to another embodiment.

[0101] As Figure 5As shown, the electronic device 500 includes a processor 510 and a memory 520. The memory 520 is used to store executable computer programs, and the processor 510 is used to execute the methods of any of the above method embodiments under the control of the computer programs.

[0102] In some embodiments, the processor 510 may be used to control the overall operation of the electronic device 500. For example, the processor 510 may execute instructions to implement all or part of the steps of the methods in any of the foregoing embodiments of the present disclosure, so as to implement one or more of operations such as voice communication, data communication, database operation, display control, component control, multimedia processing, etc. Among them, the above components may include sensors, cameras, headphones, input / output devices, etc. The components may be internal components of the electronic device itself or external components connected to the electronic device wirelessly or wiredly. The above multimedia may include one or more of voice, image, video, and text.

[0103] The electronic device 500 may be any terminal device capable of executing the method.

[0104] Each module of the above electronic device 500 may be implemented by the processor 510 executing the computer program stored in the memory 510, or may be implemented by other structures, which is not limited herein.

[0105] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for causing a processor to implement various aspects of the present disclosure.

[0106] The computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. The computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (non-exhaustive list) of the computer-readable storage medium include: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), static random access memory (SRAM), portable compact disc read-only memory (CD-ROM), digital versatile disc (DVD), memory stick, floppy disk, mechanical encoding device, such as punched cards or raised structures in grooves having instructions stored thereon, and any suitable combination of the foregoing. The computer-readable storage medium used herein is not construed as an instantaneous signal itself, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission medium (e.g., optical pulses through an optical fiber cable), or electrical signals transmitted through wires.

[0107] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device through a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.

[0108] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state-setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages - such as Smalltalk, C++, etc., and conventional procedural programming languages - such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network - including a local area network (LAN) or a wide area network (WAN) - or, alternatively, may be connected to an external computer (e.g., through the Internet using an Internet service provider). In some embodiments, by using the state information of the computer-readable program instructions to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA), the electronic circuit can execute the computer-readable program instructions to implement various aspects of the present disclosure.

[0109] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer-readable program instructions.

[0110] These computer-readable program instructions may be provided to a processor of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that the instructions, when executed by the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable medium storing the instructions comprises a manufacture including instructions that implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0111] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process such that the instructions executed on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.

[0112] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of code, or a portion of an instruction, and the module, segment of code, or portion of an instruction includes one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may, in fact, be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box of the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions. As will be apparent to those skilled in the art, implementation via hardware, implementation via software, and implementation via a combination of software and hardware are equivalent.

[0113] The embodiments of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and variations are obvious to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The choice of terms used herein is intended to best explain the principles of the embodiments, the practical application, or the improvement of technologies in the market, or to enable other ordinary skilled persons in the art to understand the embodiments disclosed herein. The scope of the present disclosure is defined by the appended claims.

Claims

1. A face synthesis method based on multimodal information interaction, characterized in that: include: Receiving audio clips and facial images corresponding to the audio clips; Extracting audio time sequence features corresponding to the audio clip based on frequency information of the audio clip, and extracting audio semantic features from the audio clip by using a machine learning algorithm, wherein the audio time sequence features are used to represent a regularity of changes in the audio clip over time; The audio time sequence feature and the audio semantic feature are fused by a bidirectional cross attention algorithm to obtain a sound sequence semantic feature; Extracting facial features corresponding to the face image through a deep learning algorithm, aligning the facial features with the audio semantic features, and obtaining facial and audio semantic features; The phonetic sequence semantic features and the facial sound semantic features are integrated to obtain a joint feature, the joint feature is decoded and reconstructed, and the joint feature is converted into a target voice video.

2. The method according to claim 1, characterized in that After fusing the phonetic sequence semantic features and the facial sound semantic features to obtain a joint feature, decoding and reconstructing the joint feature, and converting the joint feature into a target voice video, the method further includes: Calculating the video difference between the target voice video and the sample video through a multimodal identification algorithm, and judging whether the target voice video meets the preset video display standard according to the calculated video difference; Calculate the phase difference between the target voice video and the sample corresponding to the sample video by a dynamic time warping algorithm; If the target voice video does not meet the preset video display standard and / or the phase difference value is greater than the preset phase difference value, the video difference value and / or the phase difference value are input into the preset loss function, and the adjustment parameter is calculated by the preset loss function to update the target voice video by the adjustment parameter.

3. The method according to claim 1, characterized in that: The step of aligning the facial features with the audio semantic features to obtain the facial and audio semantic features includes: Determining whether the audio semantic features and the facial features are from the same semantic space by a cross-modal feature matching classification algorithm; If they are from the same semantic space, the feature distribution of the facial features is adjusted through a cross-modal feature alignment encoder algorithm to align the feature distribution of the facial features with the feature distribution of the audio semantic features to obtain the facial-audio semantic features.

4. The method according to claim 1, characterized in that The bidirectional cross attention algorithm is: Wherein, Softmax represents a preset activation function, Q represents a query vector generated according to the audio temporal features, K represents a key vector generated according to the audio semantic features, V represents a value vector generated according to the audio semantic features, QK T represents the transpose of Q and K for matrix multiplication, represents a preset scaling factor, and Attn(Q, K, V) represents the semantic feature of the phonetic sequence.

5. The method according to claims 2 and 3, characterized in that The preset loss function is: L Total =λ1L LSGAN +λ2L CGAN +λ3L align +λ4L DTW.phase +λ5L Feat Among them, L Total Represents the total loss value, L LSGAN represents the difference between the facial feature and the sample facial feature algorithm, λ1 represents L LSGAN The corresponding weight value, L CGAN represents the difference calculated by the facial feature algorithm based on the audio timing feature matching, λ2 represents L CGAN The corresponding weight value, L align represents the loss value calculated by the algorithm for aligning the facial features and the audio semantic features, and λ3 represents L align The corresponding weight value, L DTW.phase represents the difference calculated by the sound and picture phase difference algorithm, λ4 represents L DTW.phase The corresponding weight value, L Feat represents the preset constraint feature, λ5 represents L Feat The corresponding weight value.

6. The method according to claim 5, characterized in that L align =E[logD align (F semantic ,F face )]+E[log(1-D align (F semantic ,F face.fake ))] Among them, E represents the preset mathematical expectation, D align represents the cross-modal feature matching classification algorithm, F semantic represents the audio semantic features, F face represents the facial features, F face.fake Represents the semantic features of the facial sound.

7. The method according to claim 1, characterized in that Before receiving the audio segment and the facial image corresponding to the audio segment, the method further includes: Acquire the voice audio and a face image corresponding to the voice audio; The background noise of the speech audio is removed by a denoising algorithm to obtain the speech audio, wherein the background noise is the interference sound of the non-target speech signal mixed in during the audio collection process.

8. A face synthesis device based on multimodal information interaction, characterized in that: The device comprises: A receiving module, used for receiving audio clips and face images corresponding to the audio clips; A feature extraction module, configured to extract audio time series features corresponding to the audio segment based on frequency information of the audio segment, and extract audio semantic features from the audio segment by a machine learning algorithm, wherein the audio time series features are used to represent a regularity of changes of the audio segment over time; A feature fusion module, used to fuse the audio time sequence feature and the audio semantic feature through a bidirectional cross attention algorithm to obtain a sound sequence semantic feature; A feature alignment module, used to extract facial features corresponding to the face image through a deep learning algorithm, align the facial features with the audio semantic features, and obtain facial and audio semantic features; The video generation module is used to fuse the phonetic sequence semantic features and the facial sound semantic features to obtain a joint feature, decode and reconstruct the joint feature, and convert the joint feature into a target voice video.

9. An electronic device, comprising a memory and a processor, wherein the memory is used to store a computer program; and the processor is used to execute the computer program to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Spoken English pronunciation correction auxiliary system based on speech recognition

    CN120412648A

  • Dynamic video real-time generation system based on voice features and natural language processing

    CN121037651A