Multi-view-angle-based mode missing robust audio and video self-supervised learning method and device
Through three-dimensional portrait reconstruction and self-supervised multi-view consistency training, combined with a unified modality adapter, the performance degradation problem of audio and video speech recognition systems under multi-view changes and video modality loss is solved, and efficient and stable speech recognition is achieved in complex environments.
Patent Information
- Application Number
- CN202510791502.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-12
AI Technical Summary
The performance of existing audio and video speech recognition systems drops sharply under multi-view changes and missing video modalities, making it difficult to maintain robustness and generalization capabilities in complex environments.
Multi-view lip movement video data is generated through 3D portrait reconstruction. Self-supervised multi-view consistency and domain alignment training are combined, and a unified modality adapter is introduced to expand the perspective diversity of the training dataset. Multi-view consistency loss and feature domain alignment loss are introduced into the model to improve the model's adaptability to perspective changes and modality loss.
The model's robustness and generalization capabilities under conditions of multi-view changes and missing video modalities have been significantly improved, and it can maintain efficient and stable speech recognition performance in complex scenarios.
Smart Images

Figure CN120635264A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of three-dimensional head portrait reconstruction and speech self-supervised learning, and specifically to a self-supervised learning technical solution for generating lip movement videos of a speaker from different perspectives and learning multimodal features of audio and video under different perspectives and video modality missing conditions. Background Art
[0002] In recent years, Automatic Speech Recognition (ASR) technology has achieved significant breakthroughs in closed, quiet environments, with recognition accuracy approaching that of human hearing. However, in real-world applications, complex factors such as ambient noise, high reverberation, and multi-source interference still severely impact system performance, leading to a significant decrease in speech recognition accuracy. This is particularly true in complex scenarios like the "cocktail party effect," where multiple sound sources are simultaneously emitting sound. Traditional single-modal audio recognition systems struggle to effectively extract key information from the target speech, leading to recognition confusion.
[0003] To enhance the robustness of speech recognition systems, multimodal audio-visual speech recognition (AVSR) technology has been widely researched in recent years. This technology aims to combine audio and lip movement video information to improve system stability and accuracy. Audio and video modalities are complementary, with visual information inherently resistant to noise interference. Consequently, AVSR systems outperform traditional audio recognition systems in noisy environments. However, most current AVSR methods assume a stable video input with a fixed viewpoint, typically a frontal perspective. In practical applications, however, video input often exhibits non-ideal viewpoints due to camera angle variations and posture deviations, severely impairing the quality and consistency of lip movement information. Recent advances in self-supervision and large-scale pre-training strategies, such as the AV-HuBERT model (1 Bowen Shi, Wei-Ning Hsu, Kushal Lakhotia, and Abdelrahman Mohamed, “Learning audio-visualspeech representation by masked multimodal cluster prediction,” arXivpreprint arXiv:2201.02184, 2022) have improved the performance of AVSR on standard benchmarks. However, these models still struggle with multi-view modalities. Comprehensive training data reflecting diverse head poses is also lacking. Therefore, addressing the challenge of multi-view variability is crucial for deploying AVSR systems in real-world settings.
[0004] Furthermore, in some scenarios, the video modality may be completely missing or partially unavailable, for example due to occlusion, signal interruption, or low-bandwidth transmission. This further degrades the performance of existing audio and video fusion systems, even outperforming audio-only systems in the absence of a modality. Therefore, ensuring the robustness and generalization of audio and video speech recognition systems in real-world environments with multiple viewpoints and missing modalities has become a critical technical challenge that needs to be addressed. Summary of the Invention
[0005] The present invention aims to address the defects in the existing technology and proposes a new technology for robust self-supervised learning of audio and video with modal omissions based on multi-perspectives, which solves the problem of performance degradation of multi-modal audio and video speech recognition systems in complex scenarios where multi-perspectives and lip movement videos are missing.
[0006] As multimodal technologies continue to develop and audio and lip-movement video data sources become increasingly abundant, existing multimodal audio-visual recognition models based on joint audio and video modeling still suffer from significant performance degradation when faced with varying viewpoints and when video modalities are missing. To address this technical bottleneck, this paper proposes a multi-view, self-supervised learning scheme for robust audio and video with modality missingness.
[0007] The technical solution of the present invention provides a multi-perspective modality-missing robust audio and video self-supervised learning method, which includes the following process: Generate multi-view lip movement video data through 3D head portrait reconstruction; The audio and video features after multi-view and modality loss processing are input into the encoder to extract multi-modal features that are consistent with multi-view and adapted to modality loss; The decoder receives multimodal features and applies them to downstream speech-related tasks.
[0008] Moreover, the generation of multi-perspective lip movement video data through three-dimensional head portrait reconstruction includes estimating the head posture of the original monocular video, reconstructing an animated three-dimensional lip movement video based on a three-dimensional deformation model, rotating the three-dimensional model within a preset angle range, and rendering to generate a multi-perspective training data set.
[0009] Moreover, the method for extracting multi-view consistent and modality-missing adapted multimodal features includes using multi-view consistency loss to constrain feature alignment between real and synthetic views, and using feature domain alignment loss to achieve cross-view representation space alignment through contrastive learning.
[0010] Moreover, the extraction of multi-modal features that are consistent with multiple views and adapt to modality loss includes reconstructing audio and video joint features based on audio features when the video modality is missing.
[0011] Moreover, the reconstruction of audio and video joint features based on audio features includes achieving modality loss adaptation through hierarchical alignment, aligning the correlation structure between audio and audio and video at the early layer, and aligning the specific numerical representation of features at the deep layer.
[0012] Moreover, the decoder adopts an end-to-end Transformer decoding structure, integrating multi-head self-attention and encoder-decoder attention mechanisms.
[0013] Moreover, the robustness is improved by randomly masking the video modality during the training phase.
[0014] On the other hand, the present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and runnable on the processor, wherein when the processor executes the program, the multi-perspective-based modality-missing robust audio and video self-supervised learning method is implemented as described above.
[0015] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the multi-perspective-based modality-missing robust audio and video self-supervised learning method as described above.
[0016] On the other hand, the present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned multi-perspective-based modality-missing robust audio and video self-supervised learning method.
[0017] The present invention synthesizes multi-perspective lip movement video data through three-dimensional head portrait reconstruction technology, expands the perspective diversity of the training data set, and improves the model's ability to model lip movement features under different observation angles. At the same time, the present invention designs a self-supervised multi-perspective representation learning mechanism, adopts multi-perspective consistency loss and feature domain alignment loss, and prompts the model to learn perspective-independent and domain-independent audio and video joint feature expressions. On this basis, a unified modality adapter module is further introduced. When encountering a situation where the video modality is partially missing or completely missing, the recognition mode can be adaptively and smoothly degraded to an audio single-modal recognition mode, thereby significantly improving the robustness and stability of the model in the case of modality loss. Compared with the traditional method of self-supervised learning based on mask prediction, the solution of the present invention shows stronger robustness and better generalization ability under the conditions of multi-perspective changes and video modality loss, and has good prospects for promotion and application.
[0018] Compared with the prior art, the present invention has the following advantages and beneficial technical effects: To address the problem of drastic performance degradation of multimodal audio and video speech recognition systems in the presence of multiple viewpoints and missing video modalities, this paper systematically addresses the robustness of the model in complex scenarios by introducing 3D head portrait reconstruction to generate multi-view samples, combining self-supervised multi-view consistency with domain alignment training, and introducing a unified modality adapter. Furthermore, the present invention proposes a method for simultaneously jointly training multimodal mask prediction loss and multi-view consistency loss, so that the encoder has the ability to adapt to perspective changes and modal failures while learning multimodal context.
[0019] In general, the multi-perspective modality-missing robust audio and video self-supervised learning method proposed in this invention provides an efficient, stable and scalable new approach for speech processing related tasks in actual complex application environments.
[0020] The solution of the present invention is simple and convenient to implement and has strong practicality. It solves the problems of low practicality and inconvenience in actual application existing in related technologies, can improve user experience, and has important market value. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 This is a diagram of the overall model structure of the multi-perspective modality-missing robust audio and video self-supervised learning method of an embodiment of the present invention.
[0022] Figure 2 Schematic diagram of a multi-view data generation strategy according to an embodiment of the present invention.
[0023] Figure 3 This is a graph of the speech recognition performance of the model in an embodiment of the present invention under multi-view changes and real conditions.
[0024] Figure 4 It is a flow chart of an embodiment of the present invention. DETAILED DESCRIPTION
[0025] The following will further illustrate the concept, specific structure and technical effects of the present invention in conjunction with the accompanying drawings and embodiments, so as to fully understand the purpose, characteristics and effects of the present invention.
[0026] The present invention selects a self-supervised model framework to implement a multi-view modality-missing robust audio and video self-supervised learning method, adopts multi-view data synthesis technology to learn viewpoint-invariant lip reading capabilities, and introduces a self-supervised multi-view representation learning paradigm. This paradigm integrates multi-view consistency loss and representation domain alignment loss to ensure that the learned embedding is robust to viewpoint offset and domain change features. Its core advantages are reflected in: the powerful multimodal representation capability provided by the pre-trained model, the robustness to multi-view and modality missing features, and the stable performance in cross-task transfer.
[0027] Example 1 See also Figure 1-4 The embodiment of the present invention discloses a multi-perspective modality-missing robust audio and video self-supervised learning method, comprising: Step 1: Generate multi-view lip movement video data through 3D head portrait reconstruction; Step 2: Input the audio and video features after multi-view and modality loss processing into the encoder, and output multi-modal features that are consistent with multi-view and adapted to modality loss; Step 3: The decoder receives the multi-view and modality-loss-robust audio and video features output by the encoder and applies them to downstream speech-related tasks.
[0028] like Figure 1 The present invention provides a multi-perspective modality-missing robust audio and video self-supervised learning method, and the overall framework is designed as follows: first, the original monocular video sequence is processed using three-dimensional head portrait reconstruction technology, and the head rotation angle information is obtained through facial posture estimation, and the three-dimensional deformable face model and neural rendering method are combined to reconstruct the complete three-dimensional head portrait of the speaker. By performing posture adjustment and texture mapping on the reconstructed model, lip movement videos from multiple perspectives are synthesized to construct rich multi-perspective training data, providing a data basis for subsequent multi-perspective consistency modeling. Then, the video modalities generated by the above multi-perspectives and the original audio data are input into the audio and video encoder based on the AV-HuBERT architecture to extract time-aligned multi-modal features. Figure 1 As shown in (a), the encoder of the AV-HuBERT architecture extracts audio and video modality features respectively and fuses them, while introducing the mask prediction task to construct the basic self-supervised learning objective. Figure 1 As shown in (b), the Multi-View Representation Learning (MVL) module is introduced to achieve cross-view feature representation consistency by sharing encoder parameters. This module contains two key self-supervised optimization objectives: the first is the Multi-View Consistency Loss (MVC Loss), which minimizes the feature differences between the real and synthetic perspectives of the same semantic segment, including the mean square error loss and the correlation matrix difference; the second is the Representation Domain Alignment Loss (RDA Loss), which uses a contrastive learning strategy to guide the model to map samples of the same speaker but different perspectives to a similar representation space, while pulling away non-target samples to improve cross-view discrimination capabilities. Finally, in scenarios where the visual modality is partially missing or completely unavailable, the model uses Figure 1The Unified Modality Adapter (UMA) module shown in (c) performs dynamic adjustments. The UMA module takes audio features as input and uses a multi-layer Transformer structure to model and reconstruct these features layer by layer. Its shallow layers achieve semantic alignment by comparing the correlation structures of audio and video features, while its deeper layers directly align the numerical representations of audio and audio-video fusion features, thereby reconstructing a near-complete modality representation of information within the context of audio alone. This mechanism enables the model to maintain high speech recognition performance even in the absence of the visual modality, achieving smooth degradation and dynamic adaptation.
[0029] Specifically, the main purpose of step one is to expand the diversity of the training dataset and improve the model's robustness to changes in facial perspective. During implementation, head pose estimation can be performed using monocular RGB video data to extract the corresponding Euler angle information. Next, video clips with significant rotational changes are selected as source material for multi-view simulation. Subsequently, a 3D head portrait reconstruction method is used to generate realistic, animated 3D lip-movement videos that are consistent with the original video's timing using a 3D deformable model and texture optimization techniques. Finally, the 3D model is rotated horizontally and vertically within a preset angle range (-25° to 25°) to render video data from multiple different perspectives, thus forming an expanded multi-view training dataset.
[0030] Furthermore, the three-dimensional portrait reconstruction technology is implemented in the following ways: Generate detailed, animatable 3D avatar models based on monocular RGB video; The video sequences are roughly aligned using a 3D deformable model, and the geometric structure and visual details of the lips are maintained through geometric alignment and texture refinement.
[0031] In one embodiment, to improve the model's adaptability to real-world application scenarios, a dataset containing multi-view and modality-missing perturbations is constructed to train a model with multi-view and modality-adaptive capabilities, specifically including: Perform three-dimensional rotation of the horizontal (Yaw) and vertical (Pitch) angles in the range of (-25°, 25°) with a step size of 5° to generate video samples with different perspectives; In the generated video samples, the video frames are randomly masked or the video modality is completely masked to simulate the video missing situation, and the masking ratio varies from 10% to 100%; Through the above data expansion, the amount of training set data has increased by about three to four times, fully covering the real distribution of perspective changes and modal missingness, and improving the practical application robustness of the model.
[0032] The main purpose of step 2 is to obtain a multimodal audio and video feature representation that can adapt to different perspective changes and is domain-independent through a self-supervised learning strategy. Figure 1 As shown in (a)(b), the present invention improves the existing AV-HuBERT encoder structure, introduces a multi-view representation learning module (MVL) and a multimodal consistency optimization mechanism, and constructs an audio and video joint representation model with view invariance and modal robustness.
[0033] First, on the audio and video front-ends, linear projection layers and ResNet-18 structures are used to extract features from the original audio signal and lip movement video, respectively, to obtain aligned modal features. This structure is consistent with the audio and video input module in the AV-HuBERT model.
[0034] The extracted features are then input into the AV-HuBERT encoder composed of a 12-layer Transformer for multimodal feature fusion, and a masked multi-modal prediction loss (MMP loss) is introduced to achieve multimodal context modeling and local feature recovery.
[0035] On this basis, the present invention proposes a multi-view representation learning module (MVL), such as Figure 1 As shown in (b), this module applies two self-supervised losses to the encoder output features: Multi-view consistency loss (MVC Loss): This includes feature-level mean squared error loss (MSE Loss) and feature correlation alignment loss (Corr Loss). It is used to constrain the alignment of real-view samples and simulated-view samples of the same speaker in the representation space, ensuring that changes in viewpoint do not lead to semantic shifts. These two losses are combined through weighted coefficients to form an overall consistency loss function. Feature domain alignment loss (RDA Loss): This uses a contrastive learning mechanism to use real-view samples and multi-view samples generated by 3D head portrait reconstruction as positive sample pairs, and other different speakers or different videos as negative samples. By narrowing the feature distance between positive samples and pushing the features of negative samples further apart, consistent alignment across viewpoints and domains is achieved.
[0036] In addition, in order to improve the robustness of the model to modality loss, the video modalities of some input samples are randomly masked during the training phase, and a unified modality adapter (UMA) module is introduced based on the encoder (e.g. Figure 1(c) shows the module, which is a new structure proposed by the present invention. It aims to restore a nearly complete joint feature representation of audio and video through the audio modality alone. In the case of missing or partially unavailable video modality, it ensures that the system can smoothly degenerate to using only the audio modality for recognition while maintaining overall stable performance.
[0037] Specifically, the UMA module comprises a multi-layer Transformer structure, with each layer receiving the audio encoder output features and the UMA output of the previous layer, semantically enhancing the audio representation layer by layer. During training, a Feature Similarity Loss (FS Loss) is defined to align the structural consistency of the UMA output with the joint features of normal audio and video at different semantic layers. This involves two steps: correlation matrix alignment and feature vector MSE reconstruction. Ultimately, the MVC loss, RDA loss, mask prediction loss (MMP Loss), and UMA alignment loss (UMA Loss) are combined to form a complete joint training objective, making the encoder robust to multi-viewpoint changes and modality loss.
[0038] The main purpose of step three is for the decoder to receive the multi-view and modality-loss-robust audio and video features output by the encoder and apply them to downstream speech-related tasks. The decoder is used to receive the joint audio and video feature representation after multi-view consistency modeling and modality adaptation processing in step two, further extract temporal context information, and complete downstream tasks such as speech recognition.
[0039] In order to achieve further temporal modeling and semantic analysis of the multi-view consistent and modality-loss robust audio and video joint features obtained in step 2 through the decoder, and output the final results for downstream tasks such as speech recognition. The decoder preferably adopted in the embodiment is based on the Transformer architecture design, has strong temporal dependency modeling capabilities and contextual semantic extraction capabilities, and can effectively adapt to multi-modal input and audio and video fusion features under complex viewing conditions. Specifically, the decoder includes 6 layers of Transformer decoding units, each of which includes a multi-head self-attention module (Multi-HeadSelf-Attention), an encoder-decoder attention module (Encoder-Decoder Attention), a feed-forward fully connected network (Feed-Forward Network), a residual connection (Residual Connection) and layer normalization (LayerNormalization). The above structure allows the decoder to capture contextual information at each time step, enhancing the temporal consistency and semantic continuity of the speech content.
[0040] On the input side, the decoder receives the fused audio and video features output by the encoder and uses a positional encoding mechanism to preserve frame-level sequential information. In each layer's encoder-decoder attention module, the decoder can directly access the cross-modal representations modeled in the encoder, enabling efficient information exchange and semantic completion.
[0041] To improve the decoder's adaptability to inputs of varying speech rates and semantic spans, the present invention further proposes the introduction of a multi-scale attention mechanism within the multi-layer decoding unit. By setting different receptive fields in different attention heads, this allows for the simultaneous modeling of short-term pronunciation features and long-term contextual dependencies. Furthermore, residual connections and layer normalization can be employed to enhance the training stability of deep networks and mitigate the vanishing gradient problem.
[0042] In practice, a linear transformation layer and a softmax classification layer can be added at the end of the decoder to normalize the feature representation at each time step and output the corresponding recognition label or text result. The output can be word-level or phoneme-level units, and the specific form can be flexibly adjusted according to the task requirements.
[0043] During the inference phase, the decoder is adaptive to modality loss: when the video modality is completely unavailable, the system can rely solely on the audio enhancement features generated by the Unified Modality Adapter (UMA) for speech recognition. When the video modality is partially available, the decoder integrates the available video information to optimize output quality. This mechanism ensures robust recognition performance despite complex conditions such as viewpoint changes, frame loss, and occlusion.
[0044] The decoder structure described in the present invention not only retains the advantages of the AV-HuBERT model in sequence modeling, but also achieves robust speech recognition capabilities in complex perspectives and modality loss scenarios by introducing a multi-scale fusion path and an embedded residual structure. It is suitable for a variety of practical application scenarios such as natural dialogue recognition, multi-perspective meeting recording, and human-computer interaction.
[0045] Example 2 Based on the multi-view modality loss robust audio and video self-supervised learning process provided in Example 1, a preferred solution is further proposed, wherein the multi-view data simulation process in step 1 is as follows: Figure 2As shown in the figure, the head pose is first estimated for the original monocular RGB video, and the Euler angle information is extracted by solving the PnP problem based on the two-dimensional facial key points; then, video clips with large rotation changes are screened out as multi-view simulation source materials; then, a three-dimensional head portrait reconstruction method based on a three-dimensional deformation model is used to generate an animated three-dimensional lip movement video representation; finally, by adjusting the three-dimensional model posture parameters, it is rotated horizontally and vertically within the range of (-25°, 25°) to render video data from different perspectives, effectively expanding the training dataset and improving the training coverage and model robustness under perspective changes.
[0046] In the embodiment, head pose estimation is first performed on each video clip. This process solves the Perspective-n-Point (PnP) problem through 2D facial key points to extract Euler angles. Then, video clips with significant rotation changes are selected as the main materials for multi-view simulation. In order to generate additional perspectives, a 3D avatar reconstruction method is used, instead of relying on multiple cameras or complex acquisition equipment. A neural avatar model is used to reconstruct a detailed animatable 3D representation from the RGB video sequence, using a 3D deformation model (such as FLAME) as the geometric basis. Given the shape parameters , expression parameters and posture parameters , the baseline model provides a coarse mesh:
[0047] in, is the number of vertices, R is the field of real numbers, is represented by a coarse three-dimensional mesh.
[0048] In order to capture details that cannot be represented by the 3D deformable model, a geometric refinement function is used to obtain a fine mesh:
[0049] in, is a fine grid representation, is the geometry refinement function. Through coarse geometry alignment and fine texture correction, it ensures that the lip contour remains highly consistent with the original frame.
[0050] In the preferred solution, the multi-view representation learning in step 1 adopts a self-supervised training strategy to capture the inherent audio and video correlation and learn audio and video embeddings that are view-invariant and domain-aligned. Figure 1 As shown, assuming a For batches of samples, is a real audio and video sequence, is the corresponding synthetic multi-view sequence, iis the sample number, represents the audio and video input sample of the i-th real perspective, Represents The corresponding synthetic multi-view audio and video samples are generated by 3D head portrait reconstruction. Construct an embedding function , mapping each sample into a representation matrix with a time series length of T and a feature dimension of D per frame , so that the mapped features are invariant to changes in viewing angles. Define the mean square error (MSE) term , to ensure that the real and synthetic embeddings match at the feature level:
[0051] However, feature-level matching alone is not enough to ensure semantic alignment. To this end, we introduce correlation alignment loss. Let is the normalized embedding matrix, and the Frobenius norm is used to measure the difference in its correlation matrix:
[0052] in, is the normalized embedding of the i-th real sample, is the normalized embedding of the i-th synthetic sample.
[0053] The final multi-view consistency loss function for:
[0054] in, is a weighting coefficient used to balance the importance of eigenvalue alignment and structural correlation alignment.
[0055] In the preferred solution, the audio and video features described in step 2 are input to the encoder for alignment, and contrast loss is used to learn domain-invariant features. Synthetic samples that are close to the front (or with minimal rotation) are selected as positive samples and paired with the corresponding real samples. Negative samples come from other data in the same batch. Given an embedding function , for each sample number , let its true perspective sample be , the sample closest to the front in the synthetic perspective is , then the two constitute a positive sample pair; Select the embedding results of other speakers or videos from the same batch as negative samples, numbered , the corresponding sample is ,in . Positive sample similarity The definition is as follows:
[0056] The similarity of negative samples The calculation method is:
[0057] Final representation domain alignment loss for:
[0058] in, is the temperature parameter, is the number of negative samples.
[0059] In order to align with downstream phoneme or lip shape categories, a mask multimodal prediction loss is further introduced The final multi-view representation learning MVL training objectives are as follows:
[0060] in, are the weight hyperparameters that control the proportion of multi-view consistency loss, domain alignment loss, and mask prediction loss in the overall optimization objective.
[0061] In the preferred solution, the process of learning multi-view feature consistency, multi-modal mask prediction and modality loss adaptation is combined to obtain an encoder through overall training. The model structure includes: The audio and video front-ends take as input the original audio waveform signal and the lip movement video frame sequence, respectively, and extract unimodal feature vectors. The audio front-end architecture uses a simple linear layer for feature dimensionality reduction, while the video front-end architecture uses a lightweight ResNet-18 network to extract lip movement dynamic features. The encoder structure is a 12-layer Transformer stack. The input is the feature representation extracted by the audio and video front-end, and the output is a joint audio and video feature vector that integrates multi-view consistency and modal robustness features. The multi-view representation module and feature domain alignment module are implemented based on the Transformer sublayer respectively. The structure of each sub-module is consistent with that of the encoder, and parameter sharing is adopted to ensure that the model uniformly learns multi-view and multi-domain adaptive features during training.
[0062] In step 2, the encoder is trained to learn not only multi-view invariant features, but also audio and video features that are robust to modality loss. The unified modality adapter consists of multiple Transformer layers, each of which receives two inputs: the output of the multi-view consistency encoder (audio features only) and the output of the unified modality adapter of the previous layer. This architecture ensures that the audio features are gradually aligned with the audio and video features.
[0063] set up represents the audio enhancement feature embedding output by UMA, represents the encoder joint embedding feature corresponding to the complete audio and video input, is the number of time steps, is the feature dimension for each time step.
[0064] The total loss function of UMA is defined as:
[0065] in, is the weight coefficient used to control the balance between mask prediction loss and feature alignment loss; Masking the multimodal prediction loss encourages the model to recover context even when parts of the video modality are missing. is the feature similarity loss.
[0066] To further improve the alignment effect, the embodiment divides the Transformer layer in UMA into two stages: the early layer set : Used to align the correlation structure between audio and video; deep collection : Specific numerical representation used for alignment features.
[0067] The feature similarity loss function is defined as follows:
[0068] in: 、 Indicates UMA 、 Embedded features of the complete modality of the reference audio and video in the layer; 、 Indicates UMA 、 Layer audio enhancement features; is a function for calculating the normalized inter-feature correlation matrix; is the Frobenius norm; is the alignment weight coefficient of each layer; ,in is the total number of UMA layers.
[0069] Through the above-mentioned loss function design, the present invention achieves effective alignment of audio modality and audio-video joint features at two levels: structural correlation and specific value representation. This enables the model to rely on audio-enhanced representation to achieve high-performance speech recognition when the video modality is unavailable or partially missing.
[0070] Example 3 Based on the multi-view modality loss robust audio and video self-supervised learning process provided in the previous embodiment, a preferred solution is further proposed. In step 2, the encoder is trained so that it can simultaneously learn the audio and video multi-view features and the modality loss robustness features. The specific training steps include: S1. Prepare a multi-view audio and video self-supervised training dataset. The LRS3 and OuluVS2 audio and video datasets were used. LRS3 contains a standard 30-hour subset (30 hours) and a full 433-hour subset (433 hours). To simulate multi-view scenarios, during 3D head portrait reconstruction, geometric offset optimization was trained for 150 epochs, texture optimization for 80 epochs, and combined geometric and texture optimization for 50 epochs to maximize the preservation of lip movement detail. Sampling angles were set within the range (-25°, 25°), with sampling every 5°. Additionally, slight angle variations were randomly sampled within the range (-10°, 10°) to enhance data diversity. The resulting datasets consisted of a mixture of 30 hours of real data and 30% multi-view synthetic data (mv30h training data), and a mixture of 433 hours of real data and 40% multi-view synthetic data (mv433h training data). In addition, the OuluVS2 dataset contains videos shot at five different viewing angles (0°, 30°, 45°, 60°, and 90°). To address modality-missing scenarios, the video modality is masked from the input at a random ratio (10% to 100%) during training.
[0071] S2. During self-supervised pre-training, the base model obtained from the fifth iteration of AV-HuBERT training was used as initialization, and the pseudo-labels generated by it were used for training the sixth iteration of the present invention. The encoder consists of a 12-layer Transformer module, and the visual front-end uses a ResNet-18-based architecture. The input lip movement video has a uniform resolution of 88×88 and is a single-channel grayscale image. During training, the weights of different loss terms are dynamically adjusted.
[0072] S3. During the training of the unified modality adapter model, the encoder parameters are kept frozen, and the audio unimodal input is processed jointly by the encoder and the unified modality adapter. During training, the feature similarity guidance weight γ = 0.8 and the modality adaptation loss weight λ are set. p =0.2 to promote the audio features to move closer to the complete audio and video feature space and achieve natural degradation under modality loss.
[0073] During fine-tuning, the multi-view learning model is optimized using a sequence-to-sequence cross-entropy loss based on the attention mechanism, using only visual input. During fine-tuning of the unified modality adapter, both audio and video inputs are used, and the video modality is randomly masked during training (with masking ratios ranging from {0.0, 0.1, 0.2, …, 1.0}). All training hyperparameters (such as learning rate, optimizer type, and batch size) remain consistent with the fifth iteration of AV-HuBERT training, ensuring robustness and superiority under multi-view variations and modality loss.
[0074] In the preferred solution, the hyperparameter settings of step three are consistent with the parameter settings of the fine-tuning step in document 1. This ensures model performance and transfer adaptability. Specifically, the encoder is a multimodal encoder based on the AV-HuBERT structure. Its input is the audio and video feature sequence obtained after three-dimensional head portrait reconstruction and multi-view simulation processing in step one. After joint training with multi-view consistency loss (MVC Loss), feature domain alignment loss (RDA Loss), mask multimodal prediction loss (MMP Loss), feature similarity loss (FS Loss) in step two, it outputs audio and video fusion embedding features with perspective invariance and modality loss robustness.
[0075] The decoder is an end-to-end Transformer decoding structure, consisting of six layers of Transformer decoding units. Each layer consists of modules such as Multi-Head Self-Attention, Encoder-Decoder Attention, Feed-Forward Network, Residual Connection, and Layer Normalization, and is capable of modeling long-term contextual information and extracting semantic dependencies.
[0076] The decoder takes as input the joint audio and video feature representation output by the AV-HuBERT encoder and, after multi-layer temporal modeling, outputs the corresponding text sequence or phoneme label. This decoder is widely adaptable to a variety of downstream tasks, including but not limited to audio-video speech recognition (AVSR), visual speech recognition (VSR) with video input only, and robust ASR (audio speech recognition) in scenarios where video is missing or noisy.
[0077] During training, the decoder not only models speech recognition using a standard sequential cross-entropy loss but also shares collaborative optimization objectives with upstream modules, such as masked multimodal prediction loss and audio and video feature alignment loss, to form a multi-task joint training framework. This strategy enables the model to simultaneously learn inter-modal alignment, context modeling, and ultimately recognition task capabilities during training, thereby improving the robustness and generalization of the overall recognition system under multi-viewpoint variations and modality loss.
[0078] Ultimately, the decoder obtained through training can adaptively use the audio, video or joint features output by the encoder according to the completeness of the input modality during the inference stage to achieve high-precision, low-latency speech recognition prediction, which is suitable for a variety of complex speech scenarios such as natural conversations, human-computer interactions, and conference transcription.
[0079] To facilitate understanding of the technical effects of the present invention, the method proposed in the present invention is described and verified through specific examples below.
[0080] The embodiment of the present invention provides a speech recognition method in multi-view and modality-deficient scenarios. Figure 1 The program consists of the following three steps: Step 1: Generate multi-view lip movement video data through 3D head portrait reconstruction.
[0081] In the multi-view data generation module, 3D head portrait reconstruction technology is used to generate lip movement video data from different viewpoints. The original audio and video signals are first fed into the audio encoder and video encoder, respectively, to generate audio and video feature representations. After multi-view simulation, the video features are rotated according to the sampled rotation angle (ranging from -25° to 25°, sampled every 5°, and additionally introducing small random variations from -10° to 10°), thus forming a multi-view extended dataset. During training, the features before and after the viewpoint change are aligned for consistency, and a multi-view consistency loss is used to guide the model to learn a view-invariant multimodal feature representation.
[0082] Step 2: Input the audio and video features after multi-view and modality loss processing into the encoder, and output multi-modal features that are consistent with multi-view and adapted to modality loss.
[0083] The expanded multi-view data is fed into a self-supervised learning model. The encoder utilizes a 12-layer Transformer architecture, the visual front-end uses a ResNet-18 network, and the audio front-end uses a simple linear projection layer. The input video signal is normalized to an 88×88 grayscale image. During training, the model simultaneously optimizes the mask prediction loss (for learning a joint representation of audio and video) and the feature domain alignment loss (for improving domain independence). A unified modality adapter is used to randomly mask the video modality during training (with a masking ratio ranging from 10% to 100%), ensuring stable performance and robustness even when the video modality is missing. During training, the weights of different loss terms are dynamically adjusted for the multi-view and masked data.
[0084] In step 3, the decoder receives the multi-view and modality-loss-robust audio and video features output by the encoder and applies them to downstream speech-related tasks.
[0085] After completing self-supervised training, the trained encoder weights are loaded, and a decoder is added to the encoder's backend. The decoder architecture utilizes a sequence-to-sequence (seq2seq) cross-entropy optimization framework based on the attention mechanism. During the fine-tuning phase, varying degrees of masking are applied to the video modality based on different application scenarios, while maintaining the same training hyperparameters as the fifth iteration of AV-HuBERT to ensure fair comparison. Ultimately, the model outputs aligned and robust multimodal feature representations of audio and video for practical application in speech recognition tasks.
[0086] The performance of this invention is calculated using the Word Error Rate (WER) as the core indicator. For the challenging scenarios of multi-view and modality loss, the test was conducted on the LRS3 subsets grouped by different yaw angles (yaw angle <10°, 10°-30°, >30°) and the OuluVS2 dataset. Figure 3As shown in the figure, the test data is divided into three subsets according to the face yaw angle, corresponding to small angle, medium angle and large angle yaw respectively; the recognition performance of the AV-HuBERT baseline model (base30h, base433h) and the method of the present invention (mv30h, mv433h) are compared on each subset; OuluVS2 represents the results of fine-tuning and testing on the OuluVS2 dataset with large perspective changes; the performance of mv30h decreases slightly on some samples with unchanged pitch angle, but the model maintains strong recognition ability on the OuluVS2 dataset of real scenes, which meets the requirements of real application scenarios. In particular, mv433h further demonstrates better multi-perspective change robustness, verifying the generalization performance of the present invention under cross-domain conditions. Compared with traditional self-supervised models based on mask prediction loss (such as AV-HuBERT), the present invention can effectively maintain recognition performance under various viewpoint changes and modality loss conditions, demonstrating significant multi-view adaptability and modality robustness, and further verified its excellent cross-domain generalization capability in the OuluVS2 dataset test.
[0087] AV-HuBERT is used as the baseline model, and the LRS3 audio, video and speech dataset is used as training and test data, including 30 hours and 433 hours of corpus from the LRS3 dataset. To expand the diversity of training data and enhance the model's robustness to multi-viewpoint changes, a multi-view extended dataset is generated through a perspective synthesis method. Among them, the ms30h dataset refers to a multi-view extension set constructed based on the original 30 hours of data, with a ratio of 30% synthetic data and 70% real data; the ms433h dataset is constructed based on 433 hours of data, with a ratio of 40% synthetic data and 60% real data. For modality loss scenarios, in order to evaluate the robustness in the case of incomplete video modalities, the video modality is randomly masked, with masking ratios ranging from 10% to 100%, to simulate different degrees of video loss in real applications.
[0088] As shown in Table 1, the recognition performance of the proposed method is compared with that of the existing AV-HuBERT method under various test conditions. "0" corresponds to the standard LRS3 test subset, serving as the baseline test condition without viewpoint perturbation or modality loss. In the visual speech recognition (VSR) task, synthetic multi-view test data with horizontal yaw (Yaw) and vertical pitch (Pitch) angles of ±5° and ±10° was constructed to evaluate the model's adaptability to varying viewpoints. In the audio-video speech recognition (AVSR) task, random masking ratios of 10% and 30% were applied to the video modality to simulate real-world scenarios with incomplete or partially missing modalities, testing the model's robustness to modality loss. Without introducing any viewpoint rotation or modality perturbation, the baseline model achieved a word error rate (WER) of 5.4% in the AVSR task. Building on this foundation, the proposed multi-view self-supervised learning model (MVL) significantly improves the system's modeling capabilities in lip reading tasks by introducing multi-view consistency loss (MVC Loss), feature domain alignment loss (RDA Loss), and masked multimodal prediction loss (MMP Loss). Using the same training data, the WER for lip reading tasks is reduced to 43.2%.
[0089] Table 1 Performance comparison between the present invention and the prior art (unit: WER / %)
[0090] Furthermore, by introducing the Unified Modality Adapter (UMA) module based on the MVL model, the system's WER in the AVSR task with complete audio and video modal input was further reduced to 5.0%. More importantly, when the video modality is partially missing, the system still maintains relatively stable recognition performance, demonstrating its robust adaptability to modality loss scenarios.
[0091] Overall, the proposed audio and video speech recognition model achieved a relative improvement of approximately 12.3% in recognition performance compared to the existing self-supervised audio and video recognition model (AV-HuBERT) under multimodal and multi-view input conditions. This result fully demonstrates the significant advantages of the proposed method in improving the robustness, generalization, and adaptability of speech recognition systems to complex practical application scenarios, demonstrating its promising engineering application value and promotion prospects.
[0092] In specific implementation, the method proposed in the technical solution of the present invention can be automatically run by those skilled in the art using computer software technology. System devices that implement the method, such as computer-readable storage media that store the corresponding computer program of the technical solution of the present invention and computer equipment that runs the corresponding computer program, should also be within the scope of protection of the present invention.
[0093] The following embodiment describes the electronic device established by the multi-perspective modal missing robust audio and video self-supervised learning method provided by the present invention. The electronic device established by the multi-perspective modal missing robust audio and video self-supervised learning method described below and the multi-perspective modal missing robust audio and video self-supervised learning method described above can be referenced to each other.
[0094] The electronic device may include: a processor, a communications interface, a memory, and a communications bus, wherein the processor, the communications interface, and the memory communicate with each other via the communications bus. The processor may call logic instructions in the memory to execute a multi-view modality-missing robust audio and video self-supervised learning method, which mainly includes the software processing portion of the above steps.
[0095] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0096] In some possible embodiments, a multi-perspective modality-missing robust audio and video self-supervised learning system is provided, comprising the following modules: The first module is used to label the target category of the training sample set to form a corresponding relationship between sample data and labels; The second module is used to establish a detection framework consisting of a backbone network, a neck network, and a detection network. The backbone network extracts features by stacking convolutional layers and cross-stage local layers, introduces a spatial pyramid pooling module to enhance multi-scale expression, and strengthens attention to fine-grained disease areas by enhancing the spatial attention mechanism. The neck network integrates a dynamic multi-scale convolution module, extracts features through parallel multi-scale convolution kernels, and fuses high- and low-level features with upsampling layers and cross-stage local layers. The detection network is equipped with several parallel detection heads, each of which performs disease detection on feature maps of different scales. The third module is used to train the established detection framework using a training sample set, and realize intelligent road disease detection based on the trained detection framework.
[0097] In some possible embodiments, a non-transitory computer-readable storage medium is provided, including a readable storage medium on which a computer program is stored. When the computer program is executed, a multi-perspective-based modality-missing robust audio and video self-supervised learning method is implemented as described above.
[0098] In some possible embodiments, a computer program product is provided, including a computer program, which, when executed by a processor, implements the multi-perspective-based modality-missing robust audio and video self-supervised learning method as described above.
[0099] The specific embodiments described herein are merely illustrative of the spirit of the present invention. Persons skilled in the art may make various modifications, additions, or substitutions to the described specific embodiments without departing from the spirit of the present invention or exceeding the scope of the appended claims.
Claims
1. A multi-view modality-missing robust audio and video self-supervised learning method, characterized by: Including the following process, Generate multi-view lip movement video data through 3D head portrait reconstruction; The audio and video features after multi-view and modality loss processing are input into the encoder to extract multi-modal features that are consistent with multi-view and adapted to modality loss; The decoder receives multimodal features and applies them to downstream speech-related tasks.
2. The method according to claim 1, wherein: The method of generating multi-perspective lip movement video data through three-dimensional head portrait reconstruction includes estimating the head posture of the original monocular video, reconstructing an animated three-dimensional lip movement video based on a three-dimensional deformation model, rotating the three-dimensional model within a preset angle range, and rendering to generate a multi-perspective training data set.
3. The method according to claim 1, wherein: The method of extracting multi-view consistent and modality-missing adapted multi-modal features includes using multi-view consistency loss to constrain feature alignment between real and synthetic views, and using feature domain alignment loss to achieve cross-view representation space alignment through contrastive learning.
4. The method according to claim 1, wherein: The method of extracting multi-modal features that are consistent with multiple views and adapt to modality loss includes reconstructing audio and video joint features based on audio features when the video modality is missing.
5. The method according to claim 4, characterized in that: The method of reconstructing the joint audio and video features based on audio features includes achieving modality loss adaptation through hierarchical alignment, aligning the correlation structure between audio and audio and video at an early layer, and aligning the specific numerical representation of features at a deep layer.
6. The method according to claim 1, wherein: The decoder adopts an end-to-end Transformer decoding structure, integrating multi-head self-attention and encoder-decoder attention mechanisms.
7. The method according to claim 1, wherein: Robustness is improved by randomly masking the video modality during the training phase.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, it implements the multi-perspective modality-missing robust audio and video self-supervised learning method as described in any one of claims 1 to 7.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the multi-perspective modality-missing robust audio and video self-supervised learning method as described in any one of claims 1 to 7.
10. A computer program product comprising a computer program, characterized in that: When the computer program is executed by a processor, it implements the multi-perspective modality-missing robust audio and video self-supervised learning method as described in any one of claims 1 to 7.