Multi-modal large model learning method, system and storage medium based on instantaneous detection and rebalancing
By using instantaneous detection and rebalancing, the modal feature fusion weights of a multimodal large model are dynamically adjusted using KL divergence and cross-entropy loss function, which solves the modal imbalance problem and improves the model's generalization ability and recognition accuracy.
Patent Information
- Application Number
- CN202511788140.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-01
- Publication Date
- 2026-03-20
- Estimated Expiration
- 2045-12-01
AI Technical Summary
During the learning process, multimodal large models tend to favor strong modes and ignore weak modes due to mode imbalance, which affects the overall performance. Existing methods cannot adjust the modal contribution in real time, resulting in performance degradation.
By employing an instantaneous detection and rebalancing method, the fusion weights of modal features are calculated using KL divergence, and the contribution between modalities is dynamically adjusted to achieve real-time perception and update of feature fusion weights. The model parameters are then optimized by combining the cross-entropy loss function.
It significantly improves the generalization ability and recognition accuracy of multimodal large models in complex tasks, avoids the performance loss of traditional lag strategies, and ensures the stability and adaptability of the model during training.
Smart Images

Figure CN121234026B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to a multimodal large model, in particular to a multimodal large model learning method, system and storage medium based on instantaneous detection and rebalancing. BACKGROUND
[0002] Multimodal large models aim to use multi-source heterogeneous data for deep learning to improve the generalization ability and performance of the model, and have greater advantages compared to traditional single-modal methods. Multimodal large models have shown wide application value in practical applications such as offensive expression recognition, fake news detection, and sentiment analysis. However, multimodal large models often face the problem of modal imbalance during learning (training), that is, the information contribution of different modalities is different, which leads the model to be biased towards strong modalities while ignoring the information of weak modalities, thereby affecting the overall performance. Existing methods usually evaluate the modal strength and learn at the same time in an unbalanced state, and this lagging rebalancing strategy cannot immediately regulate the modal contribution, resulting in a decrease in model performance. SUMMARY
[0003] The purpose of the present application is to provide a multimodal large model learning method, system and storage medium based on instantaneous detection and rebalancing, which can instantly perceive and dynamically adjust the fusion weight between modalities to solve the problem of modal imbalance in model training, thereby improving the accuracy and robustness of multimodal large models.
[0004] Technical solution: The multimodal large model learning method based on instantaneous detection and rebalancing provided by the present application comprises the following steps:
[0005] S1, constructing a multimodal large model matched with the input multimodal data sample set;
[0006] S2, sequentially selecting multiple samples of each modality from the sample set and inputting them into the multimodal large model encoder to extract single-modal features, and inputting the single-modal features into the multimodal large model classifier to obtain single-modal prediction probabilities;
[0007] S3, performing feature fusion on all single-modal features to form initial multimodal features, and inputting the initial multimodal features into the multimodal large model classifier to obtain initial multimodal prediction probabilities;
[0008] S4, calculating an instantaneous intensity coefficient based on the difference between the single-modal prediction probabilities and the initial multimodal prediction probabilities, and calculating a rebalancing feature fusion weight for each single-modal feature through the instantaneous intensity coefficient;
[0009] S5, using the rebalancing feature fusion weight to re-weight and fuse all single-modal features to obtain balanced multimodal features, and inputting the balanced multimodal features into the multimodal large model classifier to obtain balanced multimodal prediction probabilities;
[0010] S6, calculating cross-entropy loss based on balanced multi-modal prediction probability, and updating model parameters through back propagation;
[0011] S7, updating initial weight of feature fusion in next round based on instantaneous intensity coefficient of current round, and returning to step S2 until all samples in sample set are input into multi-modal big model.
[0012] The steps S3 and S4 are the first forward propagation. Step S3 generates initial multi-modal features and prediction probability by extracting single-modal features and performing weighted fusion, which lays a foundation for subsequent rebalancing of feature fusion weight. Step S4 obtains rebalanced feature fusion weight based on the difference (such as KL divergence) between single-modal prediction probability and initial multi-modal prediction probability. Through the first forward propagation, the contribution of weak modal is identified and weakened, and the dominance of strong modal is strengthened. Steps S5 and S6 are the second forward propagation. Step S5 forms balanced multi-modal features and calculates balanced multi-modal prediction probability by performing weighted fusion of all single-modal features again through rebalanced feature fusion weight. Step S6 updates model parameters through back propagation of the model by using cross-entropy loss as loss function. In summary, the first step of double forward propagation is mainly used to intensify the influence of strong and weak modal, and the second step uses the intensified features for training, so as to balance the contribution of different modal and accelerate the convergence of the model. Step S7 updates the initial weight of feature fusion in next round training based on the rebalanced feature fusion weight of current round training, so that the model continuously adapts to the change of inter-modal contribution in the whole training process. This method breaks through the limitation of traditional lagging rebalancing strategy, realizes the closed-loop training of “intensification-rebalancing-optimization”, and significantly improves the generalization ability and recognition accuracy of the model in complex multi-modal tasks.
[0013] As a preferred, the formula for calculating the instantaneous intensity coefficient in step S4 is:
[0014] ,
[0015] wherein, is the instantaneous intensity coefficient corresponding to the modal feature of the t-th input sample, is the KL divergence between the modal prediction probability of the t-th input sample and the initial multi-modal prediction probability .
[0016] The formula for calculating the instantaneous intensity coefficient is:
[0017] ,
[0018] wherein, For the t-th input sample, for The corresponding number in the middle The probability of each prediction outcome. for The corresponding number in the middle The probability of each predicted outcome, where K is the total number of possible outcomes;
[0019] The formula for calculating the rebalancing feature fusion weights is as follows:
[0020] ,
[0021] in, Indicates the first In the second input sample The rebalancing feature fusion weights corresponding to modal features.
[0022] KL divergence accurately reflects the strength of each modality in the current training state. The larger the divergence, the more inconsistent the single modality is with the overall multimodal representation, meaning it needs to be weakened. By normalizing the KL divergence into weight update values, the system achieves automatic suppression of weak modalities and enhancement of strong modalities. This mechanism ensures the objectivity and adaptability of the rebalancing process and effectively avoids subjective bias from manually setting weights.
[0023] Preferably, the formula for calculating the cross-entropy loss in step S6 is as follows:
[0024]
[0025] in, For the first Input sample data, For the first The judgment result of the input sample data; As an indicator function, its possible values are as follows: , ; For the first The predicted probability of balanced multimodal features for the next input sample. for The corresponding number in the middle The probability of each prediction result is calculated. A multi-class cross-entropy loss function is introduced as the loss function of this method. By minimizing the negative log-likelihood, the model parameters are optimized, which can efficiently measure the differences in probability distributions, achieve stable model convergence, and improve the applicability of the method.
[0026] Preferably, in step S7, the feature fusion weights for the next round are updated according to the following formula:
[0027] in, For the first the next input sample the initial weight of the modal feature, is a hyperparameter, is set to wherein is the number of all modal types contained.
[0028] The initial weight is dynamically updated by using the exponential moving average (EMA) method, and the EMA is balanced by the hyperparameter The historical weight and the current updated weight are balanced to make the weight change smooth and stable, avoiding shocks or mutations in the training process. This strategy not only improves the stability of the training, but also enables the model to "remember" the historical balance state and gradually converge to the optimal weight configuration. Especially in long-term training, the EMA mechanism can effectively prevent overfitting and improve the generalization performance of the model.
[0029] The multi-modal large model learning system based on instantaneous detection and rebalancing according to the present application comprises:
[0030] A model construction module is used to construct a multi-modal large model matched with the input multi-modal data sample set;
[0031] A single modal extraction and classification module is used to sequentially select multiple sample inputs of each modal from the sample set, input the multi-modal large model encoder to extract single modal features, and input the single modal features into the multi-modal large model classifier to obtain single modal prediction probabilities;
[0032] A multi-modal fusion and classification module is used to fuse all single modal features to form initial multi-modal features, and input the initial multi-modal features into the multi-modal large model classifier to obtain initial multi-modal prediction probabilities;
[0033] A weight updating module is used to calculate the instantaneous intensity coefficient based on the difference between the single modal prediction probability and the initial multi-modal prediction probability, and calculate the rebalanced feature fusion weight of each single modal feature through the instantaneous intensity coefficient;
[0034] A multi-modal re-fusion module is used to re-weight and fuse all single modal features using the rebalanced feature fusion weight to obtain balanced multi-modal features, and input the balanced multi-modal features into the multi-modal large model classifier to obtain balanced multi-modal prediction probabilities;
[0035] A parameter updating module is used to calculate the cross-entropy loss based on the balanced multi-modal prediction probability, and update the model parameters by back propagation;
[0036] A cycle module is used to update the initial weight of the feature fusion in the next round based on the instantaneous intensity coefficient of the current round, and return to the single modal extraction and classification module until all samples in the sample set have been input into the multi-modal large model.
[0037] The computer-readable storage medium storing one or more programs of the present application includes one or more programs including instructions that, when executed by a computing device, cause the computing device to perform any of the above methods.
[0038] Beneficial effects: By fusing the single-modal features extracted from the samples to obtain the initial multi-modal features, and updating the initial weight of feature fusion according to the difference of the prediction probabilities, the contribution of the modes is instantaneously regulated, the performance loss of the traditional lag strategy is avoided, and the accuracy and robustness of the model in multi-modal data processing tasks such as aggressive expression recognition and sentiment analysis are significantly improved. Secondly, the calculation of the rebalanced feature fusion weight is based on the KL divergence to ensure the scientificity and adaptability of the rebalancing, so that the updated weight can suppress the weak modal contribution with poor prediction results and enhance the strong modal contribution with good prediction results, thereby improving the final accuracy and robustness of the model. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figure 1 The overall flowchart of the method of the present application is shown in the figure;
[0040] Figure 2 The network framework diagram of the method of the present application is shown in the figure. DETAILED DESCRIPTION
[0041] As shown in the figure, the multi-modal large model learning method based on instantaneous detection and rebalancing of the present application includes the following steps:
[0042] S1, constructing a multi-modal large model matched with the input multi-modal data sample set;
[0043] The multi-modal data sample can have two single-modal data or more single-modal data, and a corresponding multi-modal large model is constructed according to the number of modes in the input data, such as input data including image and text data, then the constructed multi-modal large model can recognize and process these two kinds of data.
[0044] The multi-modal large model includes an encoder for extracting single-modal features in multi-source data and a classifier for classifying and discriminating the features.
[0045] S2, sequentially selecting multiple samples of each mode from the sample set and inputting them into the multi-modal large model encoder to extract single-modal features, and inputting the single-modal features into the multi-modal large model classifier to obtain single-modal prediction probabilities;
[0046] The expression of the encoder extracting features is:
[0047]
[0048] wherein, is an encoder function, is a model parameter of the encoder, is a feature representing an input sample, and is an output result of the encoder; is an input sample.
[0049] The classifier obtains a prediction probability expression as:
[0050]
[0051] wherein, is a classifier function, is a model parameter of the classifier module, is a classification function, is a prediction probability of the feature extracted for the input sample, is an output result of the classifier, which is a vector including multiple components, each component representing a prediction probability of a prediction result, assuming that there are K results in total, then contains K components.
[0052] S3, all single-modal features are fused to form initial multi-modal features, and the initial multi-modal prediction probability is obtained by inputting the initial multi-modal features into a multi-modal large model classifier;
[0053] The expression of the fusion of all single-modal features is:
[0054] ,
[0055] wherein, is the initial multi-modal feature of the tth input sample, and are the normalized values of the and modal features in the tth input sample, and are the initial weights of the and modal features in the tth input sample, and the sum of the initial weights of all single-modal features is 1; and are the serial numbers of the elements in , and and represent two different modal types in (every specific modal is a single modal, such as the and modal are two single-modals), represents the universal set of all modal types, all elements in For feature fusion function;
[0056] The expression is:
[0057]
[0058] in, .
[0059] The meaning of the calculation formula is: First, we need to start from the first... In the input sample, all possible combinations are selected in pairs (the combinations are unordered and non-repeating). Then, the pairs of normalized single-modal features corresponding to these combinations are used as input and fused through a feature fusion function to obtain the initial multimodal features.
[0060] No. Secondary input sample Normalized values of modal features The expression is:
[0061]
[0062] in, For the first In the second input sample Modal features, Its normalized value, Represents the L2 norm. This indicates the modality type to which the input sample data belongs.
[0063] S4. Calculate the instantaneous intensity coefficient based on the difference between the single-modal prediction probability and the initial multimodal prediction probability, and calculate the rebalancing feature fusion weight of each single-modal feature through the instantaneous intensity coefficient;
[0064] The formula for calculating the instantaneous intensity coefficient is:
[0065] ,
[0066] in, For the t-th input sample The instantaneous intensity coefficient corresponding to the modal characteristics For the first In the second input sample Modal prediction probability Compared with the initial multimodal prediction probability KL divergence; For the sequence number in S modality, Indicates that S excluding All modes other than the modal.
[0067] 、 The expression of the formula is respectively:
[0068]
[0069]
[0070] The calculation formula of the expression is:
[0071] ,
[0072] Among them, is the tth input sample, is the probability corresponding to the tth prediction result in the formula. is the probability corresponding to the tth prediction result in the formula, and K is the total number of all possible results.
[0073] Take the instantaneous intensity coefficient as the rebalanced feature fusion weight
[0074] ,
[0075] Among them, indicates the rebalanced feature fusion weight corresponding to the modal feature in the tth input sample.
[0076] S5, use the rebalanced feature fusion weight to reweight and fuse all single-modal features to obtain balanced multi-modal features, and input the balanced multi-modal features into the multi-modal large model classifier to obtain balanced multi-modal prediction probability;
[0077] The expression of the balanced multi-modal feature is:
[0078]
[0079] The expression of the balanced multi-modal prediction probability is:
[0080]
[0081] S6, calculate the cross-entropy loss based on the balanced multi-modal prediction probability, and update the model parameters by back propagation;
[0082] The cross-entropy loss calculation formula is
[0083]
[0084] Among them, is the tth input sample, Input sample data, For the first The judgment result of the input sample data; As an indicator function, its possible values are as follows: , ; For the first The input sample balances the multimodal prediction probability. for The corresponding number in the middle The probability of a prediction outcome.
[0085] S7. Update the initial weights for the next round of feature fusion based on the instantaneous intensity coefficients of the current round, and return to step S2 until all samples in the sample set have been input into the multimodal large model.
[0086] Specifically, the initial weights for the next round of feature fusion are updated using an exponential moving average, and the update expression is as follows:
[0087]
[0088] in, For the first In the second input sample Initial weights for modal features For superparameters, Set as ,in for The number of all modal types included is equal to the initial weights at the very beginning when no samples were input for weight updates.
[0089] The multimodal large model learning system based on instantaneous detection and rebalancing described in this invention includes:
[0090] Model building module: Constructs a large multimodal model that matches the input multimodal data sample set;
[0091] Single-modal extraction and classification module: used to sequentially select multiple samples of each modality from the sample set and input them into the multimodal large model encoder to extract the features of each single modality, and input the single-modal features into the multimodal large model classifier to obtain the prediction probability of each single modality;
[0092] Multimodal fusion and classification module: This module fuses all single-modal features to form initial multimodal features, and then inputs these features into a large multimodal model classifier to obtain initial multimodal prediction probabilities.
[0093] Weight update module: used to calculate the instantaneous intensity coefficient based on the difference between the single-modal prediction probability and the initial multimodal prediction probability, and to calculate the rebalancing feature fusion weight of each single-modal feature through the instantaneous intensity coefficient;
[0094] Multi-modal re-fusion module: used for re-weighting and fusing all single-modal features using rebalanced feature fusion weights to obtain balanced multi-modal features, and inputting the balanced multi-modal features into a multi-modal large model classifier to obtain balanced multi-modal prediction probabilities;
[0095] Parameter updating module: used for calculating cross-entropy loss based on the balanced multi-modal prediction probabilities, and updating model parameters through back propagation;
[0096] The computer-readable storage medium storing one or more programs includes one or more programs including instructions, which when executed by a computing device, cause the computing device to perform the above method.
[0097] In order to better illustrate the method of the present application, a specific example is further described below:
[0098] S1, taking a pair of image-text multi-modal data entities containing aggressive expressions as original samples and constructing a multi-modal large model.
[0099] The pair of image-text multi-modal data entities containing aggressive expressions are initialized, and the specific form is:
[0100]
[0101]
[0102] Wherein, the data set The number of elements in ; represents all modal data of the th sample, represents the label of each data, 0 represents no aggressive expression, and 1 represents aggressive expression; the sample data contains image modal data and text modal data , that is, the multi-modal data sample set in this example has only two single modalities, which are image modality and text modality; the text modality may contain language content with aggression, and the image modality may carry visual information associated with the semantic information.
[0103] The overall architecture of the multi-modal large model mainly consists of two parts of encoder and classifier modules. The encoder module focuses on encoding processing of image and text pair multi-modal input, in which the image modality adopts VisionTransformer (ViT), and the text modality uses BERT model based on Transformer architecture. The classifier module is composed of a plurality of linear transformation layers and a nonlinear activation function, which is responsible for discriminating and classifying single-modal features and fused multi-modal features. Encoder module and classifier module of modal are respectively denoted as and , and corresponding model parameters of feature extraction module and decision module of modal are denoted as and ; corresponding parameters of decision module and multi-modal are respectively denoted as and .
[0104] S2, input the training sample sequence into the model, extract each single modal feature, and calculate each single modal prediction probability.
[0105] S2-1: the number of training samples input each time is , the input sample data , that is is a vector containing components, and each sample data contains modal (image) and modal (text) data, where t is the batch input, because each batch contains samples, so the serial number of the first sample in the t-th batch is ;
[0106] Input the data into the multi-modal large model encoder module to extract single modal features, and the specific form of the features is:
[0107]
[0108] wherein represents the input sample modal feature (that is, single modal feature), is the modal data in the t-th input sample.
[0109] S2-2: calculate each single modal prediction probability, and the specific form is:
[0110]
[0111] wherein represents the prediction probability of the input sample modal data (that is, the modal prediction probability), and specifically each component corresponds to each element (data) in is the probability of being an offensive expression and the probability of not being an offensive expression, for example corresponds the first input sample data is the probability of being an aggressive expression and the probability of not being an aggressive expression, that is is a vector; it should be noted that the parameters in bold in this article are vectors. Here, is equivalent to The specific form of , including the prediction probability formula involved below, will be replaced by its specific form according to the specific circumstances of this example.
[0112] S3, all single-modal features are fused to form initial multi-modal features, and the initial multi-modal prediction probability is obtained by inputting the multi-modal large model classifier;
[0113] S3-1: normalize the extracted single-modal features for subsequent processing, and the specific form is
[0114]
[0115] wherein, is the normalized value of the modal feature in the th input sample, denotes the norm.
[0116] S3-2: apply feature fusion weight to fuse all single-modal features to form initial multi-modal features , the specific form is:
[0117]
[0118] wherein, is the initial multi-modal feature of the th input sample, and are the normalized values of the modal (image) and modal (text) features in the th input sample, denotes the initial weight of the image modal feature fusion, denotes the initial weight of the text modal feature fusion, and the two initial weights satisfy , ; is a feature fusion function; in this example, because there are only two modalities, there is no process of selecting all possible combinations from two groups, and there is no and the serial number in the set, but the normalized values of the only two single-modal features are directly substituted into for calculation.
[0119] S3-3: Calculate the initial multimodal prediction probability The specific form is as follows:
[0120]
[0121] in, For the first The predicted probability of the initial multimodal features of the input sample data.
[0122] S4. Calculate the instantaneous intensity coefficient based on the difference between the single-modal prediction probability and the initial multimodal prediction probability, and calculate the rebalancing feature fusion weight of each single-modal feature using the instantaneous intensity coefficient;
[0123] S4-1: To measure the learning status of different modalities, KL divergence is used to measure the prediction difference between single-modal features and multi-modal features, in the following form:
[0124]
[0125] in, For the first In the second input sample The KL divergence between the modal prediction probabilities and the initial multimodal prediction probabilities; for The corresponding number in the middle The probability of each prediction outcome. for The corresponding number in the middle The probability of each predicted outcome, since the predicted outcome here There are only two possible outcomes, therefore This represents the probability that the input data is a non-aggressive representation. This represents the probability that the input data is an aggressive expression.
[0126] S4-2: Calculate the instantaneous intensity coefficient based on the prediction difference between the single-modal features and the initial multimodal features.
[0127] The weights are updated based on the KL divergence between the single-modal prediction probabilities and the initial multimodal prediction probabilities. A smaller KL divergence value indicates that the two distributions are relatively close. Therefore, the mode with a larger KL divergence is considered a weak mode, and its corresponding weight should be weakened during the mixed training process, as shown in the following form:
[0128]
[0129] Based on the above formula for calculating the updated weights, for Modality and Modality, regardless of which is a weak mode, for example If the mode is weak, then its corresponding KL divergence value will necessarily be relatively large. The KL divergence value corresponding to each mode will be relatively small. Since the denominator of the update weights for both modes is the same, the numerator is the KL divergence of the other mode. The molecule of a mode as a weak mode is The KL divergence of the modes (with smaller values), thus The update weights for each modality will be smaller, weakening its role in feature fusion. The modality, as a strong modality, has a larger update weight value, which increases its contribution to feature fusion.
[0130] Because the above-mentioned weighting formula accurately captures the degree of imbalance between different modes, in the th... In the next iteration, the instantaneous intensity coefficient is directly used as the rebalancing feature fusion weight for multimodal feature fusion:
[0131]
[0132] For the first In the second input sample The rebalancing of modal features and the feature fusion weights (that is, the weights after the current round update).
[0133] S5. Use the rebalancing feature fusion weights to re-weight and fuse all single-modal features to obtain balanced multimodal features, and input them into the multimodal large model classifier to obtain balanced multimodal prediction probabilities;
[0134] S5-1: Based on the obtained rebalanced feature fusion weights, the single-modal features are re-fused to obtain balanced multimodal features, in the following form:
[0135]
[0136] in, For the first Balanced multimodal features of the input samples.
[0137] S5-2: Calculate the balanced multimodal prediction probability, in the following form:
[0138]
[0139] in, For the first Balanced multimodal prediction probabilities for each input sample.
[0140] S6. Calculate the cross-entropy loss based on the balanced multimodal prediction probability and perform backpropagation to update the model parameters;
[0141] The cross-entropy loss function is expressed as follows:
[0142]
[0143] in, For the first Input sample data, For the first The judgment result of the input sample data; For indicator functions ( , ); for The corresponding number in the middle The probability of each predicted outcome is also due to the predicted outcome here. There are only two possible outcomes, therefore This represents the probability that the input data is a non-aggressive representation. This represents the probability that the input data is an aggressive expression.
[0144] S7. Based on the feature fusion weights and instantaneous intensity coefficients of the current training round, use exponential moving average to update the initial weights of the next round of feature fusion, and return to step S2 until all samples in the sample set have been input into the multimodal large model.
[0145] The expression for adjusting the initial weights for the next round of feature fusion is as follows:
[0146]
[0147] in, For the first In the second input sample Initial weights for modal features For the first In the second input sample Initial weights for modal features For superparameters, It is initially set to 0.5 (because there are only two modes, so the value is 0.5; if there are three modes, the value is 1 / 3).
[0148] The method of this invention employs a double forward propagation mechanism, which can be detailed in [link to documentation]. Figure 2 Specifically, the approach first utilizes unimodal features to construct initial multimodal features through fusion, and then detects the intensity of each unimodality. Subsequently, the initial weights of the unimodals are adjusted in real-time based on the instantaneous intensity coefficients to optimize learning under balanced conditions. This strategy is dynamically applicable throughout the entire training process, effectively alleviating the lag problem of existing methods in synchronous learning under unbalanced conditions.
Claims
1. A multimodal large model learning method based on instantaneous detection and rebalancing, characterized in that, Includes the following steps: S1. Construct a large multimodal model that matches the multimodal data sample set to be input. The multimodal data sample set includes image and text data. S2. Select multiple samples of each modality sequentially from the sample set and input them into the multimodal large model encoder to extract the features of each single modality. Then input the single modality features into the multimodal large model classifier to obtain the prediction probability of each single modality. S3. Fuse all single-modal features to form initial multimodal features, and input them into a multimodal large model classifier to obtain initial multimodal prediction probabilities; S4. Calculate the instantaneous intensity coefficient based on the difference between the single-modal prediction probability and the initial multimodal prediction probability, and calculate the rebalancing feature fusion weight of each single-modal feature through the instantaneous intensity coefficient; The formula for calculating the instantaneous intensity coefficient is: , in, For the t-th input sample The instantaneous intensity coefficient corresponding to the modal characteristics For the first In the second input sample Modal prediction probability Compared with the initial multimodal prediction probability The KL divergence, where S is the universal set of all unimodal types; The calculation formula is: , in, For the t-th input sample, for The corresponding number in the middle The probability of each prediction outcome. for The corresponding number in the middle The probability of each predicted outcome, where K is the total number of possible outcomes; The formula for calculating the rebalancing feature fusion weights is as follows: , in, Indicates the first In the second input sample Rebalancing feature fusion weights corresponding to modal features; S5. Use the rebalancing feature fusion weights to re-weight and fuse all single-modal features to obtain balanced multimodal features, and input them into the multimodal large model classifier to obtain balanced multimodal prediction probabilities; S6. Calculate the cross-entropy loss based on the balanced multimodal prediction probability and perform backpropagation to update the model parameters; S7. Update the initial weights for the next round of feature fusion based on the instantaneous intensity coefficients of the current round, and return to step S2 until all samples in the sample set have been input into the multimodal large model.
2. The method according to claim 1, characterized in that: The multimodal large model includes an encoder for extracting single-modal features from multi-source data and a classifier for classifying and discriminating features to derive predicted probabilities.
3. The method according to claim 1, characterized in that: In step S3, before feature fusion, each single-modal feature needs to be normalized, which is specifically achieved through the following formula: , in, For the first In the second input sample Modal features, Its normalized value, Describing the L2 norm, This indicates the modality type to which the input sample data belongs. .
4. The method according to claim 1, characterized in that: The expression for feature fusion of all single-modal features in step S3 is as follows: , in, Let be the initial multimodal features of the t-th input sample. and In the t-th input sample and Normalized values of modal features and In the t-th input sample and Initial weights for modal features and for The index of the element; For feature fusion function; The expression is , in, .
5. The method according to claim 1, characterized in that: The expression for balancing the multimodal features in step S5 is: 。 6. The method according to claim 1, characterized in that: The formula for calculating cross-entropy loss in step S6 is as follows: , in, For the first Input sample data, For the first The judgment result of the input sample data; As an indicator function, its possible values are as follows: , ; For the first The predicted probabilities of balanced multimodal features for the next input sample. for The corresponding number in the middle The probability of a prediction outcome.
7. The method according to claim 1, characterized in that: In step S7, the feature fusion weights for the next round are updated according to the following formula: , in, For the first In the second input sample Initial weights for modal features For superparameters, Set as ,in for The number of all modal types included.
8. A multimodal large model learning system based on instantaneous detection and rebalancing, characterized in that, include: Model building module: Constructs a large multimodal model that matches the input multimodal data sample set, which includes image and text data; Single-modal extraction and classification module: used to sequentially select multiple samples of each modality from the sample set and input them into the multimodal large model encoder to extract the features of each single modality, and input the single-modal features into the multimodal large model classifier to obtain the prediction probability of each single modality; Multimodal fusion and classification module: This module fuses all single-modal features to form initial multimodal features, and then inputs these features into a large multimodal model classifier to obtain initial multimodal prediction probabilities. Weight update module: used to calculate the instantaneous intensity coefficient based on the difference between the single-modal prediction probability and the initial multimodal prediction probability, and to calculate the rebalancing feature fusion weight of each single-modal feature through the instantaneous intensity coefficient; The formula for calculating the instantaneous intensity coefficient is: , in, For the t-th input sample The instantaneous intensity coefficient corresponding to the modal characteristics For the first In the second input sample Modal prediction probability Compared with the initial multimodal prediction probability The KL divergence, where S is the universal set of all unimodal types; The calculation formula is: , in, For the t-th input sample, for The corresponding number in the middle The probability of each prediction outcome. for The corresponding number in the middle The probability of each predicted outcome, where K is the total number of possible outcomes; The formula for calculating the rebalancing feature fusion weights is as follows: , in, Indicates the first In the second input sample Rebalancing feature fusion weights corresponding to modal features; Multimodal re-fusion module: Used to re-weight and fuse all single-modal features again using rebalanced feature fusion weights to obtain balanced multimodal features, and input them into a multimodal large model classifier to obtain balanced multimodal prediction probabilities; Parameter update module: used to calculate cross-entropy loss based on balanced multimodal prediction probabilities and update model parameters through backpropagation; The loop module is used to update the initial weights of the next round of feature fusion based on the instantaneous intensity coefficient of the current round, and return to the single-modal extraction and classification module until all samples in the sample set have been input into the multimodal large model.
9. A computer-readable storage medium for storing one or more programs, characterized in that: It includes one or more programs, each program including instructions that, when executed by a computing device, cause the computing device to perform any of the methods according to claims 1 to 7.
Citation Information
Patent Citations
Redundant adaptive multi-modal robust fusion learning method and system
CN116992396A
Method for multimodal emotion classification based on modal space assimilation and contrastive learning
US20240119716A1