Redundant Adaptive Multimodal Robust Fusion Learning Method and System

Through the redundant adaptive multimodal robust fusion learning method, the vulnerability problem of traditional multimodal models under imperfect data is solved, and higher tolerance and robust processing of multimodal data is achieved.

CN116992396BActive Publication Date: 2025-06-24SHANGHAI JIAOTONG UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310981766.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-08-04
Publication Date
2025-06-24
Estimated Expiration
2043-08-04

AI Technical Summary

Technical Problem

When traditional multimodal models face imperfect multimodal data, their performance will be severely affected, especially when partial modalities are damaged or completely lost, they cannot effectively utilize redundant information, resulting in increased vulnerability.

Method used

A redundant adaptive multimodal robust fusion learning method is adopted to construct a single-modal Gaussian probability distribution through single-modal feature extraction, encoding, sparseness and dynamic weight allocation, and a robust multimodal eigenvector is generated through multimodal fusion and feature prediction.

Benefits of technology

It improves the tolerance and robustness of the model for imperfect multimodal data, can dynamically identify and utilize lossless information in each single mode, and improves network performance and overall processing capabilities of multimodal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116992396B_ABST
    Figure CN116992396B_ABST
Patent Text Reader

Abstract

The present invention provides a redundant adaptive multi-modal robust fusion learning method and system, including: extracting single-modal initial features using a pre-trained single-modal feature extraction network; encoding each single-modal initial feature into a probability distribution; performing regularization constraints on each single-modal probability distribution; assigning element-level feature weights to each single-modal mean; generating multi-modal features using the single-modal means after weight assignment; sampling each single-modal distribution to generate corresponding single-modal feature vectors; and obtaining the probability prediction distribution of the corresponding features using each single-modal and multi-modal feature vector. The present invention considers the influence of redundancy among multi-modal data on the robustness of the model, prompting the model to dynamically identify the lossless information therein for fusion while capturing all single-modal information, so as to achieve more robust and accurate multi-modal prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal processing, and specifically, to a redundant adaptive multimodal robust fusion learning method and system. Background Art

[0002] In recent years, with the wide popularization of multimedia devices, multimodal data describing the same or related objects has grown exponentially in Internet scenarios, and multimodal data has become the main carrier of information resources in the new era. Multimodal learning algorithms proposed for multimodal data study how to comprehensively and effectively extract and screen multimodal information by using the correlation relationships between data to obtain a multimodal deep learning model with better performance.

[0003] Traditional multimodal models improve the algorithm's performance by aggregating complementary task clues provided by different modalities. However, in the real world, multimodal models may encounter imperfect multimodal data, that is, there is data with some modalities damaged or completely lost. When encountering such data, the performance of traditional multimodal models trained on clean and modality-complete data may be severely affected, and may even be worse than the performance of models trained only on the remaining undamaged partial modalities. This is because the redundant information existing in different modalities is unlikely to be captured by the neural network simultaneously. Therefore, when some modalities are damaged, multimodal models trained on clean and modality-complete data cannot utilize the redundant information contained in the remaining undamaged modalities, which makes them more vulnerable to imperfect data.

[0004] Patent document CN115983280A (application number: 202310081044.0) discloses a two-modal clustering method and system with missing data. The invention is based on an autoencoder, maps two modalities to a common space through a cross-modal contrast learning loss to learn modal consistency representations, and predicts the missing modality through a cross-modal dual prediction loss to eliminate inconsistent information between modalities and further improve the representation consistency. However, the loss function designed by this patent mainly emphasizes the consistency between modalities and ignores the learning of complementarity between modalities, which limits the overall performance of robust multimodal learning; secondly, when implementing feature fusion, this patent does not consider the possible changes in the quality of different sample data, and can only handle the case of modality missing in imperfect multimodal data, and is insufficient in dealing with the case of data damage. Summary of the Invention

[0005] Aiming at the defects in the prior art, the purpose of the present invention is to provide a redundant adaptive multimodal robust fusion learning method and system.

[0006] The redundant adaptive multimodal robust fusion learning method provided by the present invention includes:

[0007] Single-modal feature extraction step: Pre-train the single-modal feature extraction network, and extract single-modal initial features of a preset dimension from various input modal data respectively;

[0008] Single-modal feature encoding step: Use different single-modal feature encoding networks to encode the extracted single-modal initial features respectively, generate different single-modal mean and variance vector combinations with the same dimension, and construct a single-modal Gaussian probability distribution;

[0009] Single-modal feature sparsification step: Regularize the single-modal probability distributions according to the obtained mean and variance vectors of each single-modal;

[0010] Dynamic weight allocation step: Compare the obtained variance vectors of each single-modal, and allocate element-level feature weights to each single-modal mean vector;

[0011] Multi-modal fusion step: Sum the single-modal mean vectors after weight allocation to generate a multi-modal feature vector;

[0012] Single-modal probability distribution sampling step: Perform a reparameterization operation on each single-modal Gaussian probability distribution composed of different mean and variance combinations to generate corresponding single-modal feature vectors;

[0013] Single-modal and multi-modal feature prediction step: Input the obtained single-modal and multi-modal feature vectors into a class prediction network composed of a multi-layer perceptron to obtain the probability prediction distribution of the corresponding features.

[0014] Preferably, the single-modal feature extraction step includes: fixing the parameters of various pre-trained single-modal feature extraction networks, and mapping the corresponding single-modal data to initial features x1, x2, …, x M , where M is the total number of modalities; different input data types use different feature extraction networks. Use the large-scale text pre-training model BERT-large to extract the input text modal data into text initial features of T×1024 dimensions, where T is the text sequence length; use the visual feature encoding network ResNet-18 composed of deep convolution to extract the input single-image modal data into visual initial features of 512 dimensions.

[0015] Preferably, the single-modal feature encoding step includes: using different single-modal feature encoding networks to encode x1, x2, …, x M extracted respectively, and then generate corresponding single-modal mean vectors μ1, μ2, …, μ of D dimensions through two linear mapping modules M , variance vectors σ1, σ2, …, σ M , and construct a single-modal Gaussian probability distribution Different unimodal initial features should use different feature encoding networks. Use the text feature encoding network composed of TextCNN to encode the initial features of serialized text; use the feature encoding network composed of a multi-layer perceptron to encode the non-serialized initial features. The specific encoding process is as follows:

[0016]

[0017]

[0018] Among them, are the mean and variance vectors of the Gaussian probability distribution of modality m respectively; f (·) is the unimodal feature encoder of modality m; m (·) is the unimodal feature encoder of modality m; and are two linear mapping modules for calculating the mean and variance vectors respectively.

[0019] Preferably, the unimodal feature sparsification step includes: according to the obtained mean vectors μ1, μ2,..., μ M and variance vectors σ1, σ2,..., σ M of each unimodal, perform regularization constraints on each unimodal probability distribution, and train the multi-modal network until the loss function converges. The calculation formula of the loss function is as follows:

[0020]

[0021]

[0022] Among them, ‖·‖1 represents l1 regularization, and ⊙ represents element-wise scale product.

[0023] Preferably, the dynamic weight allocation step includes: comparing the obtained variance vectors of each unimodal, and assigning element-level feature weights to each unimodal mean vector μ1, μ2,..., μ M as follows:

[0024]

[0025]

[0026] Among them, δ m ∈{0,1} indicates whether modality m is missing. If the data of modality m is completely missing, then δ m =0, otherwise, δ m =1.

[0027] Preferably, the multi-modal fusion step includes: combining the obtained unimodal weights ω1, ω2,..., ω MSum the element-wise products of the corresponding unimodal mean vectors μ1, μ2, …, μ M and then sum them up to generate a multimodal feature vector h. The specific process is as follows:

[0028]

[0029] Preferably, the unimodal probability distribution sampling step includes: sampling z from the standard Gaussian distribution and then, after element-wise multiplying z m with σ m and adding the result to μ m to obtain the corresponding unimodal feature h m . The specific process is as follows: m h

[0030] h m = z m ⊙ σ m + μ m

[0031] where

[0032] Preferably, the unimodal and multimodal feature prediction step includes: inputting the obtained unimodal feature vectors h m and the multimodal feature vector h into the same class prediction network composed of a multi-layer perceptron to obtain the probability prediction distribution of the corresponding features, and using the given classification label to supervise the probability prediction distribution, calculating the loss function to train the multimodal network until the loss function converges. The calculation formula of the loss function is as follows:

[0033]

[0034]

[0035] where y is the classification label corresponding to the multimodal data x1, x2, …, x M ; l(·) represents the cross-entropy function; f(·) represents the class prediction network composed of a multi-layer perceptron that shares unimodal and multimodal features.

[0036] According to the redundant adaptive multimodal robust fusion learning system provided by the present invention, it includes:

[0037] Unimodal feature extraction module: pre-train the unimodal feature extraction network, and extract the unimodal initial features of the preset dimension from various input modal data respectively;

[0038] Single-modal feature encoding module: Use different single-modal feature encoding networks to encode the extracted initial single-modal features respectively, generate different combinations of single-modal mean and variance vectors with the same dimension, and construct a single-modal Gaussian probability distribution;

[0039] Single-modal feature sparsification module: Regularize the probability distributions of each single-modal according to the obtained mean and variance vectors of each single-modal;

[0040] Dynamic weight allocation module: Compare the obtained variance vectors of each single-modal and assign element-level feature weights to each single-modal mean vector;

[0041] Multi-modal fusion module: Sum the single-modal mean vectors after weight allocation to generate a multi-modal feature vector;

[0042] Single-modal probability distribution sampling module: Perform a reparameterization operation on each single-modal Gaussian probability distribution composed of different mean and variance combinations to generate corresponding single-modal feature vectors;

[0043] Single-modal and multi-modal feature prediction module: Input the obtained single-modal and multi-modal feature vectors into a class prediction network composed of a multi-layer perceptron to obtain the probability prediction distribution of the corresponding features.

[0044] Preferably, the single-modal feature extraction module includes: Fix the parameters of various pre-trained single-modal feature extraction networks, and map the corresponding single-modal data to initial features x1, x2, …, x M , where M is the total number of modalities; Different input data types use different feature extraction networks. Use the large-scale text pre-training model BERT-large to extract the input text modal data into text initial features of T×1024 dimensions, where T is the text sequence length; Use the visual feature encoding network ResNet-18 composed of deep convolution to extract the input single-image modal data into visual initial features of 512 dimensions;

[0045] The single-modal feature module includes: Use different single-modal feature encoding networks to encode x1, x2, …, x extracted respectively M , and then generate corresponding single-modal mean vectors μ1, μ2, …, μ of D dimensions through two linear mapping modules respectively M , variance vectors σ1, σ2, …, σ M , and construct a single-modal Gaussian probability distribution Different single-modal initial features should use different feature encoding networks. Use the text feature encoding network TextCNN to encode the serialized text initial features; Use the feature encoding network composed of a multi-layer perceptron to encode the non-serialized initial features. The specific encoding process is as follows:

[0046]

[0047]

[0048] Among them, are the mean and variance vectors of the Gaussian probability distribution of modality m respectively; f (·) is the unimodal feature encoder of modality m; m (·) is the unimodal feature encoder of modality m; and are two linear mapping modules for calculating the mean and variance vectors respectively;

[0049] The unimodal feature sparsification module includes: according to the obtained mean vectors μ1, μ2, …, μ M of each unimodal, variance vectors σ1, σ1, …, σ M , perform regularization constraints on the unimodal probability distributions, and thus train the multimodal network until the loss function converges. The calculation formula of the loss function is as follows:

[0050]

[0051]

[0052] Among them, ‖·‖1 represents l1 regularization, and ⊙ represents element-wise scale product;

[0053] The dynamic weight allocation module includes: comparing the obtained variance vectors of each unimodal, and assigning element-level feature weights to each unimodal mean vector μ1, μ2, …, μ M as follows:

[0054]

[0055]

[0056] Among them, δ m ∈{0, 1} indicates whether modality m is missing. If the data of modality m is completely missing, then δ m = 0, otherwise, δ m = 1;

[0057] The multimodal fusion module includes: multiplying the obtained unimodal weights ω1, ω2, …, ω M with the corresponding unimodal mean vectors μ1, μ2, …, μ M element-wise and then summing to generate a multimodal feature vector h. The specific process is as follows:

[0058]

[0059] The single-modal probability distribution sampling module includes: sampling z from a standard Gaussian distribution to obtain z m , and after performing an element-wise scaling product of z m with σ m and adding the result to μ m to obtain the corresponding single-modal feature h m . The specific process is as follows:

[0060]

[0061] The single-modal and multi-modal feature prediction module includes: inputting the obtained single-modal feature vectors h m and the multi-modal feature vector h into the same class prediction network composed of a multi-layer perceptron to obtain the probability prediction distribution of the corresponding features, and using the given classification label to supervise the probability prediction distribution, calculating the loss function to train the multi-modal network until the loss function converges. The calculation formula of the loss function is as follows:

[0062]

[0063]

[0064] where y is the classification label corresponding to the multi-modal data x1, x2, …, x M ; l(·) represents the cross-entropy function; f(·) represents the class prediction network composed of a multi-layer perceptron shared by single-modal and multi-modal features.

[0065] Compared with the prior art, the present invention has the following beneficial effects:

[0066] (1) The present invention uses a probability modeling form in the latent space to represent each modality. This probability modeling form encodes the distribution of possible values of each single-modal feature, rather than just a deterministic vector, making the present invention more tolerant to minor perturbations in single-modal data; in addition, the variance of the probability distribution provides an opportunity to estimate the quality of the single-modal element scale, which is crucial for subsequent dynamic weight allocation;

[0067] (2) The present invention maximally learns all useful information of each single-modal data by simultaneously optimizing the classification loss independent of each single-modal and the sparsity constraint loss imposed on the distribution of each single-modal feature, so as to achieve lossless capture of redundant information on each single-modal feature. This is a prerequisite for robust multi-modal fusion;

[0068] (3) The present invention assigns weights to the element scales for each modality by comparing the variances of the single-modal probability distributions, enabling the network to dynamically identify the lossless information in each single modality for fusion, thereby improving the network performance. In addition, the present invention also uses a shared classifier to constrain the single-modal and multi-modal features to the same common space, making the variances more comparable. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Other features, objects, and advantages of the present invention will become more apparent by reading the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0070] Figure 1 It is a flowchart of the method in an embodiment of the present invention;

[0071] Figure 2 It is a schematic diagram of the system in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0072] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all fall within the protection scope of the present invention.

[0073] Embodiment 1:

[0074] As Figure 1 shown, the present invention provides a redundant adaptive multi-modal robust fusion learning method. Taking the video classification data of image-text pairs as an example, this method includes:

[0075] Single-modal feature extraction step: Use a suitable pre-trained single-modal feature extraction network to extract the single-modal initial features of a preset dimension from various input modality data. Among them, the text data uses the Bert-lagre model to extract 1024-dimensional features, and the image data uses the Resnet18 model to extract 512-dimensional features;

[0076] Single-modal feature encoding step: Use different single-modal feature encoding networks to encode the extracted single-modal initial features respectively, generating different single-modal mean and variance vector combinations with the same dimension, and constructing a single-modal Gaussian probability distribution. Among them, both the text features and the image features use a 2-layer MLP to encode a Gaussian distribution containing 128-dimensional mean and variance;

[0077] Single-modal feature sparsification step: According to the obtained mean and variance vectors of each single modality, perform L1-norm regularization constraints on each single-modal probability distribution;

[0078] Dynamic weight allocation step: Compare the variance vectors of each single modality obtained, and assign element-level feature weights to each single modality mean vector;

[0079] Multi-modal fusion step: Sum the single modality mean vectors after weight allocation to generate a multi-modal feature vector;

[0080] Single modality probability distribution sampling step: Perform reparameterization operations on each single modality Gaussian probability distribution composed of different means and variances to generate corresponding single modality feature vectors;

[0081] Single modality and multi-modal feature prediction step: Input the obtained single modality and multi-modal feature vectors into a class prediction network composed of a multi-layer perceptron to obtain the probability prediction distribution of the corresponding features.

[0082] Specifically, the single modality feature extraction step includes: Fix the parameters of various pre-trained single modality feature extraction networks, and map the corresponding single modality data to initial features x1, x2, …, x M , where M is the total number of modalities. Different input data types use different feature extraction networks. For example, for text-image pair food classification data, use the large-scale text pre-trained model BERT-large to extract the input text modality data into 1024-dimensional text initial features; use the visual feature encoding network ResNet-18 composed of deep convolution to extract the input single image modality data into 512-dimensional visual initial features.

[0083] Specifically, the single modality feature encoding step includes: Use different single modality feature encoding networks to encode the extracted x1, x2, …, x M respectively, and then generate corresponding single modality mean vectors μ1, μ2, …, μ M of dimension D and variance vectors σ1, σ2, …, σ M through two linear mapping modules, and construct a single modality Gaussian probability distribution Different single modality initial features should use different feature encoding networks. For example, for food classification, use a feature encoding network composed of a multi-layer perceptron to encode the text and image initial features. The specific encoding process is as follows:

[0084]

[0085]

[0086] Among them, are respectively the mean and variance vectors of the Gaussian probability distribution of modality m, and f m (·) is the single modality feature encoder of modality m, and They are two linear mapping modules for calculating the mean and variance vectors respectively.

[0087] Specifically, the unimodal feature sparsification step includes: according to the obtained mean vectors μ1, μ2, …, μ M and variance vectors σ1, σ2, …, σ M , perform regularization constraints on each unimodal probability distribution, and train the multimodal network until the loss function converges. The loss function is calculated as follows:

[0088]

[0089]

[0090] where ‖·‖1 represents l1 regularization, and ⊙ represents element-wise scale product.

[0091] Specifically, the dynamic weight allocation step includes: comparing the obtained variance vectors of each unimodal, and assigning element-level feature weights to each unimodal mean vector μ1, μ2, …, μ M as follows:

[0092]

[0093]

[0094] where δ m ∈{0, 1} indicates whether modality m is missing. If the data of modality m is completely missing, then δ m = 0, otherwise, δ m = 1.

[0095] Specifically, the multimodal fusion step includes: performing element-wise scale product on the obtained unimodal weights ω1, ω2, …, ω M and the corresponding unimodal mean vectors μ1, μ2, …, μ M and then summing them to generate a multimodal feature vector h. The specific process is as follows:

[0096]

[0097] Specifically, the unimodal probability distribution sampling step includes: sampling z from the standard Gaussian distribution m , and after performing element-wise scale product on z m and σ m and then adding it to μ m to obtain the corresponding unimodal feature h m , and the specific process is as follows:

[0098]

[0099] Specifically, the single-modal and multi-modal feature prediction steps include: inputting the obtained single-modal feature vectors h m and the multi-modal feature vector h into the same class prediction network composed of a multi-layer perceptron to obtain the probability prediction distribution of the corresponding features, and using the given classification label to supervise the probability prediction distribution, and the loss function can be calculated to train the multi-modal network until the loss function converges. The loss function is calculated as follows:

[0100]

[0101]

[0102] where y is the classification label corresponding to the multi-modal data x1, x2,..., x M l(·) represents the cross-entropy function, and f(·) represents the class prediction network composed of a multi-layer perceptron sharing single-modal and multi-modal features.

[0103] Embodiment 2:

[0104] As Figure 2 , the present invention provides a redundant adaptive multi-modal robust fusion learning system, including:

[0105] Single-modal feature extraction module: Using a suitable pre-trained single-modal feature extraction network, respectively extract the single-modal initial features of a preset dimension from various input modal data; among them, the text data uses the Bert-lagre model to extract 1024-dimensional features, and the image data uses the Resnet18 model to extract 512-dimensional features;

[0106] Single-modal feature encoding module: Using different single-modal feature encoding networks, respectively encode the extracted single-modal initial features to generate different single-modal mean and variance vector combinations with the same dimension, and construct a single-modal Gaussian probability distribution; among them, both text features and image features use a 2-layer MLP to encode a Gaussian distribution containing 128-dimensional mean and variance;

[0107] Single-modal feature sparsification module: According to the obtained mean and variance vectors of each single-modal, perform L1-norm regularization constraints on each single-modal probability distribution;

[0108] Dynamic weight allocation module: Compare the obtained variance vectors of each single-modal, and allocate element-level feature weights to each single-modal mean vector;

[0109] Multi-modal fusion module: Sum the single-modal mean vectors after weight allocation to generate a multi-modal feature vector;

[0110] Unimodal Probability Distribution Sampling Module: Perform reparameterization operations on each unimodal Gaussian probability distribution composed of different combinations of means and variances to generate corresponding unimodal feature vectors;

[0111] Unimodal and Multimodal Feature Prediction Module: Input the obtained unimodal and multimodal feature vectors into a class prediction network composed of a multi-layer perceptron to obtain the probability prediction distribution of the corresponding features.

[0112] Specifically, the unimodal feature extraction module includes: fixing the parameters of various pre-trained unimodal feature extraction networks, and mapping the corresponding unimodal data to initial features x1, x2,..., x M , where M is the total number of modalities. Different input data types use different feature extraction networks. For example, for text-image pair food classification data, use the large-scale text pre-trained model BERT-large to extract the input text modality data into 1024-dimensional text initial features; use the visual feature encoding network ResNet-18 composed of deep convolution to extract the input single-image modality data into 512-dimensional visual initial features.

[0113] Specifically, the unimodal feature module includes: using different unimodal feature encoding networks to encode the extracted x1, x2,..., x M respectively, and then generating D-dimensional corresponding unimodal mean vectors μ1, μ2,..., μ M and variance vectors σ1, σ2,..., σ M through two linear mapping modules respectively, and constructing a unimodal Gaussian probability distribution Different unimodal initial features should use different feature encoding networks. For example, use the text feature encoding network composed of TextCNN to encode the serialized text initial features; use the feature encoding network composed of a multi-layer perceptron to encode the non-serialized initial features. The specific encoding process is as follows:

[0114]

[0115]

[0116] where, are the mean and variance vectors of the Gaussian probability distribution of modality m respectively, f (·) is the unimodal feature encoder of modality m, m and and are two linear mapping modules for calculating the mean and variance vectors respectively.

[0117] Specifically, the unimodal feature sparsification module includes: according to the obtained unimodal mean vectors μ1, μ2,..., μM , variance vectors σ1, σ2, …, σ M , regularize each unimodal probability distribution, and train the multimodal network until the loss function converges. The loss function is calculated as follows:

[0118]

[0119]

[0120] where, ‖·‖1 represents l1 regularization, and ⊙ represents element-wise scale product.

[0121] Specifically, the dynamic weight allocation module includes: comparing the obtained variance vectors of each unimodal, and assigning element-level feature weights to each unimodal mean vector μ1, μ2, …, μ M as follows:

[0122]

[0123]

[0124] where, δ m ∈{0, 1} indicates whether modality m is missing. If the data of modality m is completely missing, then δ m = 0, otherwise, δ m = 1.

[0125] Specifically, the multimodal fusion module includes: performing element-wise scale product of the obtained unimodal weights ω1, ω2, …, ω M and the corresponding unimodal mean vectors μ1, μ2, …, μ M and then summing them to generate a multimodal feature vector h. The specific process is as follows:

[0126]

[0127] Specifically, the unimodal probability distribution sampling module includes: sampling z from the standard Gaussian distribution m , and after performing element-wise scale product of z m and σ m and then adding it to μ m to obtain the corresponding unimodal feature h m , and the specific process is as follows:

[0128]

[0129] Specifically, the unimodal and multimodal feature prediction module includes: the obtained unimodal feature vectors h mThe multi-modal feature vector h and the like are input into the same class prediction network composed of a multi-layer perceptron to obtain the probability prediction distribution of the corresponding features, and the given classification label is used to supervise the probability prediction distribution, and the loss function can be calculated to train the multi-modal network until the loss function converges. The loss function is calculated as follows:

[0130]

[0131]

[0132] where y is the classification label corresponding to the multi-modal data x1, x2, …, x M l(·) represents the cross-entropy function, and f(·) represents the class prediction network composed of a multi-layer perceptron that shares single-modal and multi-modal features.

[0133] In summary, the present invention represents each modality in the form of probability modeling in the latent space. This form of probability modeling encodes the distribution of possible values of each single-modal feature, rather than just a deterministic vector, making the present invention more tolerant to small perturbations in the single-modal data; in addition, the present invention maximally learns all useful information of each single-modal data by simultaneously optimizing the classification loss independent of each single-modal and the sparsity constraint loss imposed on the distribution of each single-modal feature, so as to achieve lossless capture of redundant information on each single-modal feature; moreover, the present invention also assigns weights of element scales to each modality by comparing the variances of the probability distributions of each single-modal, enabling the network to dynamically identify the lossless information in each single-modal for fusion, thereby improving the network performance. In order to make the variances more comparable, the present invention also uses a shared classifier to constrain the single-modal and multi-modal features to the same common space.

[0134] Those skilled in the art know that in addition to implementing the system, device and its various modules provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system, device and its various modules provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. to achieve the same program. Therefore, the system, device and its various modules provided by the present invention can be regarded as a kind of hardware component, and the modules included therein for implementing various programs can also be regarded as the structure within the hardware component; the modules for implementing various functions can also be regarded as both software programs for implementing the method and the structure within the hardware component.

[0135] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. A redundant adaptive multi-modal robust fusion learning method, characterized in that Including: Single-modal feature extraction step: Pre-train the single-modal feature extraction network, and extract the single-modal initial features of a preset dimension from various input modal data respectively; Single-modal feature encoding step: Use different single-modal feature encoding networks to encode the extracted single-modal initial features respectively, generate different single-modal mean and variance vector combinations with the same dimension, and construct a single-modal Gaussian probability distribution; Single-modal feature sparsification step: Regularize the single-modal probability distributions according to the obtained mean and variance vectors of each single-modal; Dynamic weight allocation step: Compare the obtained variance vectors of each single-modal, and allocate element-level feature weights to each single-modal mean vector; Multi-modal fusion step: Sum the single-modal mean vectors after weight allocation to generate a multi-modal feature vector; Single-modal probability distribution sampling step: Perform a reparameterization operation on each single-modal Gaussian probability distribution composed of different mean and variance combinations to generate corresponding single-modal feature vectors; Single-modal and multi-modal feature prediction step: Input the obtained single-modal and multi-modal feature vectors into a class prediction network composed of a multi-layer perceptron to obtain the probability prediction distribution of the corresponding features; The single-modal feature extraction steps include: fixing the parameters of various pre-trained single-modal feature extraction networks, and mapping the corresponding single-modal data into initial features x1, x2, …, x M , where M is the total number of modalities; different feature extraction networks are used for different input data types. The large-scale text pre-trained model BERT-large is used to extract the input text modal data into text initial features of T×1024 dimensions, where T is the text sequence length; the visual feature encoding network ResNet-18 composed of deep convolution is used to extract the input single-image modal data into visual initial features of 512 dimensions.

2. The redundant adaptive multi-modal robust fusion learning method according to claim 1, wherein The single-modal feature encoding step includes: using different single-modal feature encoding networks to separately encode the extracted x1, x2, …, x M and then generating corresponding single-modal mean vectors μ1, μ2, …, μ of dimension D through two linear mapping modules respectively M , variance vectors σ1, σ2, …, σ M , and constructing a single-modal Gaussian probability distribution Different single-modal initial features should use different feature encoding networks. Use the text feature encoding network composed of TextCNN to encode the serialized text initial features; use the feature encoding network composed of a multi-layer perceptron to encode the non-serialized initial features. The specific encoding process is as follows: Among them, are the mean and variance vectors of the Gaussian probability distribution of mode m, respectively ; f m (·) is the unimodal feature encoder of mode m; and are two linear mapping modules for calculating the mean and variance vectors, respectively.

3. The redundant adaptive multi-modal robust fusion learning method according to claim 2, wherein, The single-modal feature sparsification step includes: according to the obtained mean vectors μ1, μ2, …, μ M , variance vectors σ1, σ2, …, σ M , perform regularization constraints on each single-modal probability distribution, and train the multi-modal network with this until the loss function converges. The calculation formula of the loss function is as follows: Where, ‖·‖1 represents l1 regularization, and ⊙ represents element-wise scale product.

4. The redundant adaptive multi-modal robust fusion learning method according to claim 1, characterized in that The dynamic weight allocation step includes: comparing the obtained variance vectors of each single modality, and assigning element-level feature weights to each single modality mean vector μ1, μ2, …, μ according to the following formula: M Assign element-level feature weights: where, δ m ∈ {0, 1} indicates whether the modality m is missing. If the data of the modality m is completely missing, then δ m = 0, otherwise, δ m = 1.

5. The redundant adaptive multi-modal robust fusion learning method according to claim 4, wherein The multi-modal fusion step includes: multiplying each obtained single-modal weight ω1, ω2, …, ω M with the corresponding single-modal mean vector μ1, μ2, …, μ M and then summing them after element-wise scale multiplication to generate a multi-modal feature vector h. The specific process is as follows:

6. The redundant adaptive multi-modal robust fusion learning method according to claim 2, characterized in that The single-modal probability distribution sampling step includes: sampling from a standard Gaussian distribution to obtain z m . Then, after element-wise scaling multiplication of z m and σ m , adding the result to μ m to obtain the corresponding single-modal feature h m . The specific process is as follows: h m = z m ⊙σ m + μ m Among them, 7. The redundant adaptive multi-modal robust fusion learning method according to claim 5 or 6, characterized in that, The single-modal and multi-modal feature prediction steps include: inputting the obtained single-modal feature vectors h m and the multi-modal feature vector h into the same class prediction network composed of a multi-layer perceptron to obtain the probability prediction distribution of the corresponding features, and using the given classification labels to supervise the probability prediction distribution, calculating the loss function to train the multi-modal network until the loss function converges. The calculation formula of the loss function is as follows: where y is the classification label corresponding to the multi-modal data x1, x2, …, x M ; l(·) represents the cross-entropy function; f(·) represents the class prediction network composed of multi-layer perceptrons that share single-modal and multi-modal features.

8. A redundant adaptive multi-modal robust fusion learning system, characterized in that, Including: Single-modal feature extraction module: Pre-train the single-modal feature extraction network, and extract the single-modal initial features of a preset dimension from various input modal data respectively; Single-modal feature encoding module: Use different single-modal feature encoding networks to encode the extracted single-modal initial features respectively, generate different single-modal mean, variance vector combinations with the same dimension, and construct a single-modal Gaussian probability distribution; Single-modal feature sparsification module: Regularize the single-modal probability distributions according to the obtained mean, variance vectors of each single-modal; Dynamic weight allocation module: Compare the obtained variance vectors of each single-modal, and allocate element-level feature weights to each single-modal mean vector; Multi-modal fusion module: Sum the single-modal mean vectors after weight allocation to generate a multi-modal feature vector; Single-modal probability distribution sampling module: Perform a reparameterization operation on each single-modal Gaussian probability distribution composed of different mean, variance combinations to generate corresponding single-modal feature vectors; Single-modal and multi-modal feature prediction module: Input the obtained single-modal and multi-modal feature vectors into a class prediction network composed of a multi-layer perceptron to obtain the probability prediction distribution of the corresponding features; The single-modal feature extraction module includes: fixing the parameters of various pre-trained single-modal feature extraction networks, and mapping the corresponding single-modal data into initial features x1, x2, …, x M , where M is the total number of modalities; different feature extraction networks are used for different input data types. The large-scale text pre-trained model BERT-large is used to extract the input text modal data into text initial features of T×1024 dimensions, where T is the text sequence length; the visual feature encoding network ResNet-18 composed of deep convolution is used to extract the input single-image modal data into visual initial features of 512 dimensions; The single-modal feature module includes: using different single-modal feature encoding networks to respectively encode the extracted x1, x2, …, x M After encoding, two linear mapping modules are used to respectively generate D-dimensional corresponding single-modal mean vectors μ1, μ2, …, μ M , variance vectors σ1, σ2, …, σ M , and construct a single-modal Gaussian probability distribution Different single-modal initial features should use different feature encoding networks. A text feature encoding network composed of TextCNN is used to encode the serialized text initial features; a feature encoding network composed of a multi-layer perceptron is used to encode the non-serialized initial features. The specific encoding process is as follows: Among them, are the mean and variance vectors of the Gaussian probability distribution of mode m, respectively; f m (·) is the unimodal feature encoder of mode m; and are two linear mapping modules for calculating the mean and variance vectors, respectively; The single-modal feature sparsification module includes: according to the obtained mean vectors μ1, μ2, …, μ M , variance vectors σ1, σ2, …, σ M , regularize the probability distributions of each single modality to train the multi-modal network until the loss function converges. The calculation formula of the loss function is as follows: Where, ‖·‖1 represents l1 regularization, and ⊙ represents element-wise scale product; The dynamic weight allocation module includes: comparing the variance vectors of each single modality obtained, and allocating element-level feature weights to each single modality mean vector μ1, μ2, …, μ according to the following formula: M ​ where, δ m ∈ {0, 1} indicates whether the modality m is missing. If the data of the modality m is completely missing, then δ m = 0, otherwise, δ m = 1; The multimodal fusion module includes: multiplying the obtained unimodal weights ω1, ω2, …, ω M element-wise with the corresponding unimodal mean vectors μ1, μ2, …, μ M and then summing them to generate a multimodal feature vector h. The specific process is as follows: The single-modal probability distribution sampling module includes: sampling z from a standard Gaussian distribution and then, after element-wise scaling multiplication of z m with σ m and adding the result to μ m to obtain the corresponding single-modal feature h m . The specific process is as follows: m ​ The single-modal and multi-modal feature prediction module includes: inputting the obtained single-modal feature vectors h m and the multi-modal feature vector h into the same class prediction network composed of a multi-layer perceptron to obtain the probability prediction distribution of the corresponding features, and using the given classification labels to supervise the probability prediction distribution, calculating the loss function to train the multi-modal network until the loss function converges. The calculation formula of the loss function is as follows: where y is the classification label corresponding to the multi-modal data x1, x2, …, x M ; l(·) represents the cross-entropy function; f(·) represents the class prediction network composed of multi-layer perceptrons that share single-modal and multi-modal features.

Citation Information

Patent Citations

  • Multi-modal sentiment analysis method and system for uncertain modal missing

    CN115983280A

  • Multimodal sentiment analysis methods and systems for addressing uncertain modalities

    CN115983280B

  • Content classification model training method, content classification method and device

    CN114462539A

  • Multi-modal prediction method and device, equipment and storage medium

    CN115186062A