Model optimization method for dynamically adjusting multi-modal weight based on SHAP

Through the method of dynamically adjusting multimodal weights, the model performance limitations caused by fixed weights in traditional multimodal data fusion are solved, and the model's adaptive optimization and performance improvement under multimodal data are achieved.

CN120561547APending Publication Date: 2025-08-29ZHEJIANG UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510421023.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-08-29

AI Technical Summary

Technical Problem

In the traditional multimodal data fusion method, the fixed weight strategy fails to effectively balance the contributions of different modes, resulting in limited model performance and lack of continuous monitoring of modal prediction capabilities, affecting the accuracy of model prediction.

Method used

The method of dynamically adjusting multimodal weights is adopted to optimize the model performance by SHAP through confidence and feature contribution calculations, modal weights are adaptively adjusted, combined with neural network model optimization, and AUC indicators and SHAP values ​​are used to optimize the model performance.

Benefits of technology

It improves the prediction accuracy and interpretability of the model, adapts to different task scenarios, overcomes the limitations of traditional weighting methods, and improves the robustness and generalization ability of the model under multimodal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561547A_ABST
    Figure CN120561547A_ABST
Patent Text Reader

Abstract

The invention discloses a model optimization method for dynamically adjusting a multi-modal weight based on SHAP. The method comprises the following steps: step 1, cleaning and preprocessing multi-modal data; 2, constructing a multi-modal neural network model, and carrying out feature extraction on multi-modal data; 3, calculating the confidence coefficient of each modal data based on an AUC index; 4, calculating the contribution degree of each modal feature to model output by using an SHAP method, and performing adaptive adjustment on input weights of different modals in combination with the confidence coefficient; and 5, carrying out training optimization on the multi-modal neural network of which the weight is adaptively adjusted, and verifying the classification performance of the model on a test set. According to the method, the weights of different modalities are adaptively optimized along with the training process, so that the precision and robustness of the classification task are improved, and the method can be widely applied to machine learning tasks needing fusion of multi-source data, such as the fields of medical diagnosis, recommendation systems, industrial detection, intelligent manufacturing, data-driven decision making and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence and multimodal machine learning, and in particular to a model optimization method for dynamically adjusting multimodal weights based on SHAP. Background Art

[0002] Traditional classification tasks typically focus on using structured data for prediction. However, this approach fails to fully consider the importance of unstructured data, which often contains rich information and can provide models with more context and deeper features. With the development of artificial intelligence technology, the application of data fusion has gradually gained attention, especially in the field of multimodal data fusion. By combining structured and unstructured data, the predictive performance of machine learning models can be significantly improved.

[0003] Multimodal data fusion involves information from multiple sources, and these data types usually have different expressions and statistical characteristics. However, the information provided by each modality contributes differently to the task. Therefore, how to effectively balance the impact of different modalities on the final prediction results has become a major challenge in the data fusion process. Traditional multimodal data fusion methods usually use fixed weights or simple linear weighting to integrate the features of different modalities. This method assumes that the contribution of all modalities to the model is constant, and fails to take into account that the importance of each modal feature changes dynamically in different tasks or scenarios. This method often leads to performance limitations of the model, especially when the relationship between multimodal information is more complex. The fixed weight method may not be able to fully tap the potential of each modality, thereby affecting the prediction accuracy of the model. In addition, the existing technology lacks a continuous monitoring mechanism for modal prediction capabilities, resulting in model updates lagging behind changes in data distribution.

[0004] SHAP, a game-theoretic method, has garnered widespread attention in recent years. SHAP can explain the importance of individual features in model predictions, revealing the contribution of different input features to the final prediction. In traditional applications, SHAP is primarily used for feature selection or model interpretability analysis, helping researchers understand the model's decision-making process. However, existing research has focused less on using SHAP values ​​for dynamic weight adjustment to optimize overall model performance. Summary of the Invention

[0005] In order to solve the above technical problems existing in the prior art, the present invention proposes a method for dynamically adjusting multimodal weights based on SHAP, which is suitable for classification tasks of multimodal data. The specific technical solution is as follows:

[0006] A model optimization method for dynamically adjusting multimodal weights based on SHAP includes the following steps:

[0007] Step 1: Clean and preprocess the multimodal data;

[0008] Step 2: construct a multimodal neural network model, perform feature extraction on the multimodal data, and obtain multimodal features;

[0009] Step 3: Calculate the confidence of each modal data based on the AUC index;

[0010] Step 4: Use the SHAP method to calculate the contribution of each modal feature to the model output, and adaptively adjust the input weights of different modalities based on the confidence level;

[0011] Step 5: Train and optimize the multimodal neural network with adaptive weight adjustment, and verify the model classification performance on the test set.

[0012] Furthermore, in step one, the multimodal data includes structured data, and data cleaning is performed on the structured data, specifically: missing values ​​are filled and One-Hot encoding is performed on the user comment type feature column to generate a structured feature vector, feature importance is calculated using LightGBM, and the top 60% important features are selected.

[0013] Furthermore, in step one, the multimodal data also includes unstructured data, and data cleaning is performed on the unstructured data, specifically: using a pre-trained BERT model to encode the text data of user comments, generating a text feature vector, using PCA to reduce the dimension of the feature vector, and using VADER sentiment analysis to calculate the sentiment score of each comment.

[0014] Furthermore, in step 1, the multimodal data is preprocessed, specifically: Z-score normalization is performed on the structured feature vector, L2 normalization is performed on the text feature vector, and the calculated sentiment score is added to the structured data.

[0015] Furthermore, in step 2, the multimodal neural network model includes a multilayer perceptron with a structured data branch and a text data branch. The features of the structured data and the text data are extracted respectively through the structured data branch and the text data branch. The modal weights are initialized and the confidence storage unit is initialized using a parameter learning method. The Sigmoid function is used to constrain the weight range, and a loss function is set to optimize the model classification task.

[0016] Furthermore, in step 3, the confidence calculation adopts the formula: C new =λC old +(1-λ)AUC; where C is the confidence and λ∈[0.7,0.9] is the momentum parameter.

[0017] Furthermore, in step 4, the SHAP method is used to calculate the contribution of each modal feature to the model output, and the SHAP mean of the high-contribution features of each modality is extracted to obtain the contribution value of each modality. The contribution value of each modality is normalized into a training parameter to obtain the initial weight of each modality. The calculation method is as follows:

[0018] base α =S / (S+T),

[0019] base β =T / (S+T),

[0020] Among them, S is the contribution value of the structured feature vector, T is the contribution value of the text feature vector, base α 、base β represent the initial weights of the structured modality and text modality respectively.

[0021] Furthermore, in step 4, the modal weight is calculated by combining the confidence and contribution values ​​as follows:

[0022]

[0023] Among them, α and β represent the final weights of structured modality and text modality respectively, and C struct 、C text They represent the confidence of structured mode and text mode respectively, and ∈ is a numerical stability term;

[0024] The values ​​of α and β are adjusted through smooth update optimization so that they are dynamically optimized to the updated parameter values ​​during the training process, and the model training and adjustment process is repeated.

[0025] Furthermore, in step five, the BCEWithLogitsLoss loss function and the Adam optimizer are used to train and optimize the multimodal neural network.

[0026] Compared with the prior art, the present invention has the following beneficial effects:

[0027] 1. Adaptive feature weight adjustment: Dynamically adjust the weights of different modal input data through confidence and feature contribution values ​​calculated by SHAP, avoiding information loss that may be caused by fixed weight strategies.

[0028] 2. Improve model interpretability: SHAP can quantify the impact of each input feature on the prediction result, making the model's decision-making process more transparent and facilitating subsequent optimization.

[0029] 3. Applicable to various multimodal tasks: This method is suitable for fusing different types of data and improving the adaptability of the model in complex tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 This is a schematic diagram of the main flow of a model optimization method for dynamically adjusting multimodal weights based on SHAP in this embodiment;

[0031] Figure 2 This is a flow chart of the method of this embodiment for using SHAP to dynamically adjust weights for model training. DETAILED DESCRIPTION

[0032] In order to make the purpose, technical solution and technical effect of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and embodiments.

[0033] like Figure 1 and Figure 2 As shown, a model optimization method for dynamically adjusting multimodal weights based on SHAP in this embodiment includes the following steps:

[0034] Step 1: Data preprocessing: Clean and preprocess multimodal data containing structured data and unstructured data.

[0035] Specifically, the processing of the structured data is as follows: missing values ​​are filled and One-Hot encoding is performed on the user comment type feature column to generate a structured feature vector, feature importance is calculated using LightGBM, and the top 60% important features are selected.

[0036] Text data processing: Clean user comment data, encode user comments using the pre-trained BERT model, generate text feature vectors, use PCA to reduce the dimensionality of the feature vectors, and use VADER sentiment analysis to calculate the sentiment score of each comment.

[0037] Data standardization and merging: Z-score standardization is performed on the structured feature vector, L2 normalization is performed on the text feature vector, and the sentiment score calculated from the text feature vector is added to the structured data.

[0038] Due to the uneven distribution of sample categories in the dataset, stratified sampling is used in the test set to ensure that there are at least 500 samples in the positive and negative classes. SMOTE oversampling is used in the training set to synthesize minority class samples, adjust the imbalanced data, and improve the generalization ability of the model.

[0039] Step 2: Model construction: Construct a multimodal neural network model. The model includes a multi-layer perceptron with structured data branches and a text data branch, and performs feature extraction on data of different modalities respectively.

[0040] Specifically, the features of structured data and text data are extracted through the dual-branch channel of the neural network, the modal weights are initialized and the confidence storage unit is initialized through parameter learning, the Sigmoid function is used to constrain the weight range, and the loss function is set to optimize the model classification task to prevent model overfitting.

[0041] Step 3: Confidence calculation: Calculate the confidence of each modal data based on the AUC indicator.

[0042] Specifically, the confidence calculation is based on the AUC score of the modal data and uses the formula:

[0043] C new =λC old +(1-λ)AUC.

[0044] Among them, C is the confidence level and λ∈[0.7,0.9] is the momentum parameter.

[0045] The confidence of each mode is calculated by AUC, and the confidence is smoothly updated in combination with the momentum parameter.

[0046] Step 4: SHAP interpretation and weight adjustment: Use the SHAP method to calculate the contribution of each modal feature to the model output, and adaptively adjust the input weights of different modes based on the confidence level.

[0047] Specifically, after every 20 rounds of training, the MLP neural network model randomly samples 5,000 training samples to calculate the SHAP value. The SHAP mean of the top 10% high-contribution features of each modality is extracted to obtain the contribution value of each modality. The contribution value of each modality is normalized to the training parameter to obtain the initial weight, which is calculated as follows:

[0048] base α =S / (S+T),

[0049] base β =T / (S+T),

[0050] Among them, S is the contribution value of the structured feature vector, T is the contribution value of the text feature vector, base α 、base β Represent the initial weights of structured modality and text modality respectively;

[0051] The modal weight is calculated by combining the confidence and contribution values ​​as follows:

[0052]

[0053] Among them, α and β represent the final weights of structured modality and text modality respectively, and C struct 、C textRespectively represent their confidence, ∈ is a numerical stability term;

[0054] The values ​​of α and β are adjusted through smooth update optimization so that they can be dynamically optimized to the updated parameter values ​​during the training process. The training and adjustment process is repeated until the weight change Δα is less than 0.01.

[0055] Step 5: Model training and optimization: Train the multimodal neural network with adaptive weights and verify the model classification performance on the test set.

[0056] Specifically, the BCEWithLogitsLoss loss function is adopted, and the loss calculation method with positive sample weights is used to improve the adaptability to category-imbalanced data; the Adam optimizer is used for model training, and the learning rate is dynamically adjusted to improve training efficiency; performance is evaluated on the test set, and the accuracy, precision, recall rate, F1 score and AUC-ROC indicators are used to measure model performance.

[0057] When training models in classification tasks, using weights adjusted with confidence and SHAP improves accuracy and recall compared to fixed weights, thereby improving model performance.

[0058] In summary, the present invention adopts a neural network model, introduces a modal confidence adaptive adjustment mechanism, and dynamically updates the weights of different modalities using historical confidence values ​​and SHAP analysis results. In addition, by introducing a momentum smoothing mechanism, the model weight adjustment process is made more stable, thereby improving the robustness and generalization ability of the model classification task. By dynamically adjusting the weights, the model can adaptively adjust the influence of each modality in different tasks, effectively overcoming the limitations of traditional weighting methods. This method can be widely applied to machine learning tasks that require the integration of multi-source data, such as medical diagnosis, recommendation systems, industrial detection, intelligent manufacturing, and data-driven decision-making.

[0059] The above description is only a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the implementation process of the present invention is described in detail above, it is still possible for those familiar with the art to modify the technical solutions described in the above examples or to replace some of the technical features therein with equivalents. Any modifications, equivalent replacements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.

Claims

1. A model optimization method for dynamically adjusting multimodal weights based on SHAP, characterized in that: The following steps are involved: Step 1: Clean and preprocess the multimodal data; Step 2: construct a multimodal neural network model, perform feature extraction on the multimodal data, and obtain multimodal features; Step 3: Calculate the confidence of each modal data based on the AUC index; Step 4: Use the SHAP method to calculate the contribution of each modal feature to the model output, and adaptively adjust the input weights of different modalities based on the confidence level; Step 5: Train and optimize the multimodal neural network with adaptive weight adjustment, and verify the model classification performance on the test set.

2. The method according to claim 1, characterized in that In step one, the multimodal data includes structured data, and data cleaning is performed on the structured data, specifically: missing values ​​are filled and One-Hot encoding is performed on the user comment type feature column to generate a structured feature vector, and feature importance is calculated using LightGBM to select the top 60% important features.

3. The method according to claim 2, characterized in that In step one, the multimodal data also includes unstructured data, and data cleaning is performed on the unstructured data, specifically: the pre-trained BERT model is used to encode the text data of user comments, generate text feature vectors, PCA is used to reduce the dimension of the feature vectors, and VADER sentiment analysis is used to calculate the sentiment score of each comment.

4. The method according to claim 3, characterized in that In step 1, the multimodal data is preprocessed, specifically: the structured feature vector is Z-score normalized, the text feature vector is L2 normalized, and the calculated sentiment score is added to the structured data.

5. The method according to claim 4, characterized in that In step 2, the multimodal neural network model includes a multilayer perceptron with a structured data branch and a text data branch. The features of the structured data and text data are extracted respectively through the structured data branch and the text data branch. The modal weights are initialized and the confidence storage unit is initialized using a parameter learning method. The Sigmoid function is used to constrain the weight range, and a loss function is set to optimize the model classification task.

6. The method according to claim 5, characterized in that In step 3, the confidence calculation adopts the formula: C new =λC old +(1-λ)AUC; where C is the confidence and λ∈[0.7,0.9] is the momentum parameter.

7. The method according to claim 5, characterized in that In step 4, the SHAP method is used to calculate the contribution of each modal feature to the model output. The SHAP mean of the high-contribution features of each modality is extracted to obtain the contribution value of each modality. The contribution value of each modality is normalized to the training parameter to obtain the initial weight of each modality. The calculation method is as follows: base α =S / (S+T), base β =T / (S+T), Among them, S is the contribution value of the structured feature vector, T is the contribution value of the text feature vector, base α 、base β represent the initial weights of structured modality and text modality respectively.

8. The method according to claim 7, characterized in that In step 4, the modal weight is calculated by combining the confidence and contribution values ​​as follows: Among them, α and β represent the final weights of structured modality and text modality respectively, and C struct 、C text They represent the confidence of structured mode and text mode respectively, and ∈ is a numerical stability term; The values ​​of α and β are adjusted through smooth update optimization so that they are dynamically optimized to the updated parameter values ​​during the training process, and the model training and adjustment process is repeated.

9. The method according to claim 1, characterized in that In step five, the BCEWithLogitsLoss loss function and Adam optimizer are used to train and optimize the multimodal neural network.

Citation Information

Cited By

  • Shale gas reservoir characteristic prediction method

    CN121030225A

  • Method and system for detecting geometric accuracy of refrigeration valve seat core based on multi-modal sensor

    CN121207042A

  • Pressure vessel heat treatment process parameter optimization method based on interpretable neural network

    CN121413467A