Double-layer missing resistance method for incomplete multi-modal data
Through the IADR and IEDR modules, the adaptive attention mechanism and shared feature prediction are used to solve the problems of single-modal and cross-modal missing in multimodal learning, and improve the robustness and generalization ability of the model.
Patent Information
- Application Number
- CN202510653092.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-19
AI Technical Summary
Existing multimodal learning methods are difficult to dynamically adjust to the interference of low-quality filler data, and pay insufficient attention to the internal missingness of a single modality, which affects the robustness and generalization ability of the model.
The intra-modality deletion-resistance module (IADR) and the inter-modality deletion-resistance module (IEDR) are adopted to handle single-modality internal deletion and cross-modal data deletion through adaptive attention mechanism, shared feature prediction and ability-aware scoring mechanism.
The robustness and generalization ability of multimodal data are improved, ensuring the effectiveness and prediction accuracy of the model in missing data environments.
Smart Images

Figure CN120671067A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multimodal learning, and in particular to a double-layer deletion resistance method for incomplete multimodal data. Background Art
[0002] With the development of artificial intelligence, multimodal learning, due to its ability to integrate multiple types of information such as language, vision, and audio, has shown broad application prospects in tasks such as sentiment analysis and human-computer interaction. However, the incompleteness of multimodal data, including the loss of internal information within a single modality and the lack of cross-modal data, seriously affects the robustness and generalization ability of the model.
[0003] Existing research mainly uses strategies such as generative methods, multimodal joint learning, and knowledge distillation to address the problem of missing data. For example, generative methods use generative adversarial networks or denoising autoencoders to fill in missing data, multimodal joint learning improves overall performance by sharing information between different modalities, and knowledge distillation methods transfer cross-modal knowledge to the student network through the teacher network, so that the student model can still effectively learn the shared features between modalities when some modalities are missing, thereby improving the final feature representation ability. However, these methods still have certain limitations: (1) It is difficult to dynamically adjust during model training to adapt to the interference of low-quality filling data; (2) Research focuses on cross-modal missing data, while paying insufficient attention to the more common single-modal internal missing data. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this paper proposes a two-layered deletion resistance method for incomplete multimodal data, aiming to simultaneously address the challenges posed by both intra-modal and cross-modal data deletions. This method, consisting of an intra-modal deletion resistance module (IADR) and an inter-modal deletion resistance module (IEDR), improves robustness and generalization to incomplete multimodal data by introducing an adaptive attention mechanism, shared feature prediction, and a capability-aware scoring mechanism.
[0005] The object of the present invention is achieved through the following technical solution: a two-layer deletion resistance method for incomplete multimodal data, comprising:
[0006] Fine-grained intra-modal missing data resistance module: uses a dynamic attention compensation mechanism to handle local missing data within a single modality;
[0007] Cross-modal shared feature prediction module: Constructs a multimodal shared semantic space based on feature consistency constraints and uses the collaborative information of available modalities to reconstruct the joint representation of missing modalities;
[0008] Capability-aware scorer module: This module generates contribution metrics for each modality through a task-driven pseudo-quality assessment network and implements nonlinear mapping of modality weights in combination with dynamic scaling factors.
[0009] Coarse-grained inter-modal fusion module: Based on the weights of each modality and combined with normalized weight preprocessing technology, it realizes the adaptive fusion of multimodal features.
[0010] Furthermore, the fine-grained intra-modality deletion resistance module specifically includes:
[0011] Window shift and block preprocessing is performed on the digitized two-dimensional matrix data of audio and video modalities to enhance the global nature of the data. An attention compensation mechanism is constructed, a compensation component is added to the standard attention calculation, and a global information compensation term with learnable parameter control is introduced. The balance between data attention and global information retention strength can be achieved by dynamically adjusting the coefficient.
[0012]
[0013] Among them, Q, K, V come from the input x i A linear transformation of , adds a compensation component U based on the original attention calculation, and introduces a learnable parameter ∈ to dynamically adjust the degree of supplementation of the global ability; d k represents the dimension of the key vector; M is the mask information.
[0014] Furthermore, the cross-modal shared feature prediction module specifically includes:
[0015] Establish a feature distribution consistency constraint mechanism based on KL divergence to force the alignment of shared feature spaces of different modalities;
[0016] Residual connections are designed to connect the original features with the extracted shared features to enhance semantic expression capabilities. The original features are extracted from the corresponding encoders of each modal data, and the shared features are extracted from the original features by the shared feature extractor.
[0017] Modal co-compensation is implemented for completely missing modalities, generating alternative representations through weighted fusion of available modal features.
[0018] Furthermore, the capability-aware scorer module specifically includes:
[0019] Construct a single-modality isolation test network and calculate the pseudo quality score of each modality using the zero-filling method;
[0020] Based on the above pseudo quality score, the ability of the scorer to evaluate the modal quality is optimized through the mean square error;
[0021] A dynamic scaling factor is introduced to adjust the influence of weights on the final cross-modal fusion and enhance the supervisory signal of high-quality modalities.
[0022] Furthermore, the coarse-grained inter-modal fusion module specifically includes:
[0023] Construct a segmented attention adjustment matrix and perform differentiated processing on positive and negative attention scores according to the weight values;
[0024] A multi-layer cross-attention mechanism is used to achieve deep interaction and semantic alignment of cross-modal features, ensuring that higher-quality modalities are given higher weights in the fusion stage.
[0025] The present invention also provides an electronic device comprising a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned double-layer deletion resistance method for incomplete multimodal data.
[0026] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned double-layer deletion resistance method for incomplete multimodal data.
[0027] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned double-layer deletion resistance method for incomplete multimodal data.
[0028] Compared with the prior art, the present invention has the following beneficial effects:
[0029] 1. IADR mainly addresses the problem of fine-grained missing information within a single modality. Due to information loss, noise interference, or uneven sampling that may occur during data acquisition, incomplete fragments may exist within a single modality. Traditional methods often use hard masking or interpolation filling strategies, but these methods may weaken the global learning ability of the model. To solve this problem, IADR adopts an adaptive attention mechanism (Intra-Attn) to dynamically adjust the attention distribution during model training to balance the utilization of valid data and the impact of missing information. This mechanism not only reduces the interference of missing parts on model predictions, but also ensures that the model's feature extraction ability is improved without losing global information.
[0030] 2. To address the problem of missing cross-modal data, IEDR (including a cross-modal shared feature prediction module, a capability-aware scorer module, and a coarse-grained inter-modal fusion module) is optimized through shared feature prediction (SFP) and capability-aware scoring mechanism (CAS). During the multimodal learning process, since different modal data may come from heterogeneous sensors or different acquisition time points, some modal data may be unavailable or poorly synchronized. Traditional methods usually rely on simple feature splicing or missing modality completion, but these methods often fail to fully utilize the information of available modalities. To this end, SFP adopts a cross-modal shared feature extraction strategy, using information from known modalities to predict shared features of missing modalities, thereby maintaining reasonable modal expression in the absence of missing modalities. On the other hand, CAS ensures the effectiveness of information fusion by calculating the contribution of different modalities to downstream tasks and dynamically weighting them based on task requirements, thereby avoiding the negative impact of low-quality features on the final prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0032] Figure 1 is a flow chart of the present invention;
[0033] Figure 2 It is a flow chart of the present invention for the missing modality scenario;
[0034] Figure 3 It is a flow chart of the fusion module of the present invention;
[0035] Figure 4 This is a schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0036] The present invention will be described in detail below with reference to the accompanying drawings. Unless there is any conflict, the features of the following embodiments and implementations may be combined with each other.
[0037] The present invention provides a double-layer deletion resistance method for incomplete multimodal data, such as Figure 1 As shown in the figure, it includes a fine-grained intra-modality deletion resistance module, a cross-modal shared feature prediction module, a capability-aware scorer module, and a coarse-grained inter-modality fusion module.
[0038] The fine-grained intra-modality missingness resistance module includes an Intra-Attn module, which handles the impact of missing data in a single modality by introducing a dynamic Intra-Attn mechanism (dynamic attention compensation mechanism). The specific implementation method is that when processing each modality data, the Intra-Attn module is first applied. This module can dynamically adjust the attention mechanism according to the location of the missing data, focusing on the available data while avoiding excessive suppression of the missing data area. By adding a compensation component to the calculation and introducing a global information compensation term controlled by a learnable parameter to ensure that the model maintains global capabilities, while selectively updating features in local missing areas, the specific formula is as follows:
[0039]
[0040] Where Q, K, V come from the input x i A linear transformation of , adding a compensation component based on the original attention calculation A learnable parameter ∈ is introduced to dynamically adjust the degree of global capability complementation; M is the mask information, the missing part is -∞, and the non-missing part is 0.
[0041] The cross-modal shared feature prediction module uses shared feature prediction (SFP) technology to reconstruct missing modalities using collaborative information between available modalities. In practical applications, existing modal data is used to infer shared features between modalities to fill the gaps in missing modalities. Specifically, the KL divergence is used to ensure the consistency of shared features, ensuring that the shared features extracted from the existing modal data can effectively fill the missing modalities, thereby reducing the information loss caused by the missing modalities.
[0042] The implementation of the capability-aware scorer generates modal contribution metrics through a task-driven pseudo-quality assessment network. The quality of each modal data is assessed based on the task requirements, and the assessment results are used as dynamic weights for further weight distribution. In actual implementation, by calculating the pseudo-quality of each modality (for example, by using the cross-entropy loss function CE to evaluate the task performance of a specific modality) and combining it with a dynamic scaling factor, the modal weight can have different degrees of influence on the final modal fusion based on the different requirements of the task, thereby enhancing the accuracy of multimodal feature fusion.
[0043]
[0044] Where W i is the scoring weight of the capability-aware scorer for modality i, and λ is the dynamic scaling factor that affects the final modality fusion as the normalized weight.
[0045] The coarse-grained inter-modal fusion module includes the Inter-Attn module, which combines the features of each modality based on the weight of each modality.
[0046] Specifically, when performing modal fusion, the contribution of each modality to the current task is dynamically evaluated by calculating the capability-aware weight. This process is completed by the Capability-Aware Scorer (CAS), see Figure 3 CAS uses the pseudo-quality of each modality to train the network and achieve modality quality assessment capabilities. Specifically, pseudo-quality is achieved by constructing a single-modality isolation test network, inputting the features of a certain modality into the decoder, retaining the features of this modality and setting the features of other modalities to zero, and then evaluating the impact of this modality on the final task. By calculating the cross-entropy loss (CE) and combining it with task-specific objectives, CAS generates a weight for each modality, reflecting its importance in the current task.
[0047] Next, during the normalized weight preprocessing process, the weights of each modality are standardized to ensure that the contributions of all modalities can be compared on the same scale. Then, the Inter-Attn module is combined for adaptive fusion. The Inter-Attn module dynamically adjusts the weights of each modality output by the capability-aware scorer, implementing weight adjustment in the attention score calculation part of the Attention-based fusion stage. High-quality modal features receive more attention and weight, while lower-quality modalities are appropriately suppressed. This weighted fusion method based on modality quality ensures that even in the absence of a modality, the model can still effectively utilize the features of other modalities for inference, maximizing information utilization efficiency.
[0048] In this way, the coarse-grained inter-modal fusion module achieves adaptive fusion of inter-modal features, and can dynamically adjust the contribution of each modality in the final decision according to the quality and task requirements of each modality, thereby optimizing the information fusion in the multimodal learning process and ensuring that high-quality feature representations can be generated even when some modalities are missing. This method improves the robustness and prediction accuracy of the model in the absence of data. The method of the present invention can be applied to various tasks, such as multimodal sentiment classification tasks, see Figure 2 , obtain video data, audio data and text data, process them with the method of the present invention to obtain fused features, input them into the decoder, and perform emotion classification.
[0049] The modules interact with each other, and the fine-grained intra-modal missingness resistance module ensures that the data within each modality can retain valid information to the greatest extent possible, avoiding the negative impact of missing areas on overall performance. The cross-modal shared feature prediction module uses data from available modalities to supplement missing modalities by sharing features, ensuring that the model can still make effective predictions in the presence of cross-modal missingness. The ability-aware scorer dynamically adjusts the weight of each modality in the fusion process based on its task performance, ensuring that key modal features are fully utilized. Finally, the coarse-grained inter-modal fusion module precisely adjusts the contribution of each modality, enabling multimodal features to be adaptively fused, thereby improving the robustness and prediction accuracy of the entire model in missing data environments. These modules work together to form a powerful method that can effectively address missing data in multimodal data and enhance the generalization ability of the model.
[0050] Experimental results show that this method outperforms existing methods on multiple datasets. Through testing on CMU-MOSI and CMU-MOSEI, two benchmark datasets widely used for multimodal sentiment analysis, good performance can be achieved in different missing scenarios. In addition, experiments on unimodal tasks (such as text classification, image classification, and audio classification) further verified the effectiveness of IADR in handling fine-grained missing data, while IEDR demonstrated excellent adaptability when faced with large-scale modal missing data. Overall, this method improves the robustness and generalization ability of the model in multimodal learning tasks, and provides a new solution for processing incomplete multimodal data in the future.
[0051] Figure 4 This is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Figure 4 The electronic device provided in this embodiment includes: a memory and a processor, wherein the memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the program instructions are loaded and executed by the processor, a double-layer deletion resistance method for incomplete multimodal data of the present invention is implemented.
[0052] It should be noted that, in addition to Figure 4 In addition to the memory and processor shown, the electronic device may also include other hardware according to its actual functions, which will not be described in detail.
[0053] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned double-layer deletion resistance method for incomplete multimodal data.
[0054] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned double-layer deletion resistance method for incomplete multimodal data.
[0055] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0056] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0057] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0058] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
Claims
1. A two-layer deletion resistance method for incomplete multimodal data, characterized by: include: Fine-grained intra-modal missing data resistance module: uses a dynamic attention compensation mechanism to handle local missing data within a single modality; Cross-modal shared feature prediction module: Constructs a multimodal shared semantic space based on feature consistency constraints and uses the collaborative information of available modalities to reconstruct the joint representation of missing modalities; Capability-aware scorer module: This module generates contribution metrics for each modality through a task-driven pseudo-quality assessment network and implements nonlinear mapping of modality weights in combination with dynamic scaling factors. Coarse-grained inter-modal fusion module: Based on the weights of each modality and combined with normalized weight preprocessing technology, it realizes the adaptive fusion of multimodal features.
2. The method according to claim 1, characterized in that The fine-grained modality deletion resistance module specifically includes: Window shift and block preprocessing is performed on the digitized two-dimensional matrix data of audio and video modalities to enhance the global nature of the data. An attention compensation mechanism is constructed, a compensation component is added to the standard attention calculation, and a global information compensation term with learnable parameter control is introduced. The balance between data attention and global information retention strength can be achieved by dynamically adjusting the coefficient. Among them, Q, K, V come from the input x i A linear transformation of , adds a compensation component U based on the original attention calculation, and introduces a learnable parameter ∈ to dynamically adjust the degree of supplementation of the global ability; d k represents the dimension of the key vector; M is the mask information.
3. The method according to claim 1, characterized in that The cross-modal shared feature prediction module specifically includes: Establish a feature distribution consistency constraint mechanism based on KL divergence to force the alignment of shared feature spaces of different modalities; Residual connections are designed to connect the original features with the extracted shared features to enhance semantic expression capabilities. The original features are extracted from the corresponding encoders of each modal data. Modal co-compensation is implemented for completely missing modalities, generating alternative representations through weighted fusion of available modal features.
4. The method according to claim 1, wherein The capability-aware scorer module specifically includes: Construct a single-modality isolation test network and calculate the pseudo quality score of each modality using the zero-filling method; Based on the above pseudo quality score, the ability of the scorer to evaluate the modal quality is optimized through the mean square error; A dynamic scaling factor is introduced to adjust the influence of weights on the final cross-modal fusion and enhance the supervisory signal of high-quality modalities.
5. The method according to claim 1, wherein The coarse-grained inter-modal fusion module specifically includes: Construct a segmented attention adjustment matrix and perform differentiated processing on positive and negative attention scores according to the weight values; A multi-layer cross-attention mechanism is used to achieve deep interaction and semantic alignment of cross-modal features, ensuring that higher-quality modalities are given higher weights in the fusion stage.
6. An electronic device comprising a memory and a processor, characterized in that: The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the double-layer deletion resistance method for incomplete multimodal data as described in any one of claims 1-5 above.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the program implements a double-layer deletion resistance method for incomplete multimodal data as described in any one of claims 1 to 5.
8. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the double-layer deletion resistance method for incomplete multimodal data as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Multi-modal fusion method based on attention mechanism and adversarial neural network
CN116452935A
Mode-missing-oriented fine-grained multi-mode element learning identification method
CN117009875A
Named entity recognition method based on comparative learning and multi-modal semantic interaction
CN117574904A
Multi-modal pre-training model acquisition method and apparatus, electrnonic device and storage medium
EP3940580A1