Interactive multi-modal learning method based on flat gradient modification and false news detection system
By introducing perceptual matrix and sharpness minimization optimization methods into multimodal models, the gradients of each modal are adjusted to achieve flat information exchange, which solves the problem of modal imbalance in multimodal model training and improves model performance.
Patent Information
- Application Number
- CN202510205875.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-30
AI Technical Summary
There is a problem of modal imbalance during multimodal model training, which leads to a skew in model learning and ignores learning of weak modes, which ultimately leads to the failure of model performance to achieve optimal performance.
An interactive multimodal learning method based on flat gradient modification is designed to sense the flatness in the current gradient direction of the modal through the perception matrix, and use this matrix to adjust the gradients of each modal, thereby promoting flat information exchange between modals, thereby achieving modal equalization.
Through the interactive multimodal learning method with flat gradient modification, the problem of modal imbalance is alleviated, the performance of multimodal deep learning models is improved, and the performance of multimodal deep learning models is shown in particular in tasks such as fake news detection.
Smart Images

Figure CN120067861A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning, and particularly to an interactive multi-modal learning method based on flat gradient modification and a fake news detection system. Background Art
[0002] Multi-modal models aim to train deep learning models through multi-modal data to achieve higher accuracy than traditional single-modal models. Multi-modal provides strong support for practical application scenarios such as fake news detection and offensive expression detection. However, during the training process of multi-modal models, a modal imbalance problem will occur. The main reason lies in the heterogeneity of modal data. The strong modality trains fast, and the weak modality trains slowly, resulting in the model learning being skewed, ignoring the learning of the weak modality, and ultimately the performance of the model not reaching the optimal. Traditional rebalancing methods only focus on balancing the training speed between different modalities, such as accelerating the learning speed of the weak modality and weakening the learning speed of the strong modality, but fail to consider the imbalance in the learning optimization level of multi-modal models. Therefore, an interactive multi-modal learning method based on flat gradient modification is proposed to alleviate and improve the modal imbalance problem, thereby improving the model performance. Summary of the Invention
[0003] The purpose of the present invention is to provide an interactive multi-modal learning method based on flat gradient modification and a fake news detection system. The present invention designs a perception matrix to perceive the flatness in the current gradient direction of the modality, and uses this matrix to adjust the gradients of each modality, promoting the exchange of flat information between modalities, thereby achieving modal balance and further improving the performance of multi-modal deep learning models.
[0004] The technical solution to achieve the purpose of the present invention is as follows: In the first aspect, the present invention provides an interactive multi-modal learning method based on flat gradient modification, including the following steps:
[0005] Step 1, process the original samples into an original sample sequence. The original samples are paired graphic-text multi-modal data entities of fake news, and construct a multi-modal deep learning model;
[0006] Step 2, input the original sample sequence into the multi-modal deep learning model to obtain the fake news prediction result and calculate the mean and variance;
[0007] Step 3, calculate the perception matrix through the mean and variance, and perform singular value decomposition to find the flat direction;
[0008] Step 4, introduce the sharpness minimization optimization method to flatten the optimization objective of the multi-modal deep learning model;
[0009] Step 5, finally calculate the loss based on the false news prediction result of the multimodal deep learning model, and update the parameters using the gradient modified by the perception matrix to train the model.
[0010] In a second aspect, the present invention provides a false news detection system based on flat gradient modification for implementing the method described in the first aspect. The system includes:
[0011] A first module that processes the original sample into an original sample sequence. The original sample is a paired graphic and text multimodal data entity of false news, and constructs a multimodal deep learning model;
[0012] A second module that inputs the original sample sequence into the multimodal deep learning model to obtain a false news prediction result and calculates the mean and variance;
[0013] A third module that calculates a perception matrix through the mean and variance, and performs singular value decomposition to find the flat direction;
[0014] A fourth module that introduces a sharpness minimization optimization method to flatten the optimization objective of the multimodal deep learning model;
[0015] A fifth module that finally calculates the loss based on the false news prediction result of the multimodal deep learning model, and updates the parameters using the gradient modified by the perception matrix to train the model.
[0016] In a third aspect, the present invention provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method described in the first aspect.
[0017] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the method described in the first aspect.
[0018] In a fifth aspect, the present invention provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the steps of the method described in the first aspect.
[0019] Compared with the prior art, the significant advantages of the present invention are: The present invention proposes a novel multimodal model training method. Considering the learning optimization of the model and the imbalance of the perception modality, a flat direction perception matrix is designed to perceive the learning optimization state of each modality for the flat direction, and exchanges the flat direction information between modalities through an interactive multimodal learning method. And a sharpness minimization optimization strategy is introduced to further smooth the learning objective. This system can be applied to many scenarios such as false news detection and offensive expression detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1It is the overall flowchart of the interactive multimodal learning method based on flat gradient modification.
[0021] Figure 2 It is the network framework diagram of the interactive multimodal learning method based on flat gradient modification.
[0022] Figure 3 It is the sub-flowchart of the perception matrix calculation and decomposition steps.
[0023] Figure 4 It is the sub-flowchart of the sharpness minimization optimization method. Specific implementation manner
[0024] Combined with Figures 1 to 4 , the present invention provides an interactive multimodal learning method based on flat gradient modification, including the following steps:
[0025] (1) Process the original samples into an original sample sequence. The original samples are paired text-image multimodal data entities of fake news, and construct a multimodal deep learning model.
[0026] Initialize the paired text-image multimodal data entities of fake news, and the specific form is:
[0027]
[0028] Define the size of the dataset D as n, and it contains m modalities; X (j) represents all the data of the j-th modality, represents the i-th data of the j-th modality; define y i ∈{0,1} c to represent the label of each data, where c represents the total number of categories.
[0029] Constructing the multimodal deep learning model includes constructing a backbone network and a classification network. Use the BERT model based on transformer and the Vision Transformer model as the backbone networks for processing paired text-image data. The classification network is composed of the simplest linear layer and the non-linear activation function ReLU. Denote the backbone network and the classification network of the j-th modality as and g (j) , and the corresponding parameters are denoted as θ (j) and W (j) .
[0030] (2) Input the original sample sequence into the multimodal deep learning model to obtain the fake news prediction result and calculate the mean and variance.
[0031] ① Input the randomly sampled small batch of data into the multimodal deep learning model to obtain the prediction result. represents the feature of the j-th modality of the i-th sample, It represents the prediction probability, and softmax represents the normalization function. The specific forms of its features and prediction probabilities are as follows:
[0032]
[0033] ② Given a total of n b data in the t-th batch and obtaining the prediction results of the multi-modal deep learning model for it, calculate the mean and variance of the features of this batch of data. The specific form is as follows, where represents the model input data and the obtained features after represents the mean of the features, and is the feature variance:
[0034]
[0035]
[0036] ③ To estimate the learning situation of the model for the overall data set, the variance calculation is generalized to the cumulative variance. The specific form is as follows, where represents the cumulative variance:
[0037]
[0038] (3) Calculate the sensing matrix through the mean and variance, and perform singular value decomposition to find the flat direction.
[0039] ① For the cumulative variance of the k-th modality obtained after calculation perform singular value decomposition on it to obtain the decomposed matrix. The specific form is as follows:
[0040]
[0041] where represents the singular value matrix, represents the singular value, U (k) and respectively represent the left and right singular vectors, represents the component direction of the singular vector.
[0042] ② Considering the geometric characteristics of the direction indicated by the singular vector if a perturbation vector with a magnitude of is applied to the direction represented by the singular vector of the feature , the change in the final output can be expressed by the following formula:
[0043]
[0044] ③ While keeping the perturbation vector unchanged, a singular vector The flatness property indicating the direction is determined by the magnitude of its singular value. The larger the singular value, the greater the change in the final output. That is, the change in the singular vector with a larger singular value will be greater. The flat direction can be represented by the singular vector with a smaller singular value. Therefore, a flat gradient modification matrix, namely the modality perception matrix, is defined as follows:
[0045] T (k) = V (k) Σ (k) [V (k) T
[0046]
[0047] where τ is a scaling parameter, I represents the identity matrix, represents the largest singular value, represents the smallest singular value.
[0048] (4) Introduce the sharpness minimization optimization method to further flatten the optimization objective of the multi-modal deep learning model, including the following steps:
[0049] ① For the multi-modal learning objective L(θ), define a perturbation ∈ and apply it to the model parameters. The specific form of the optimization objective function for sharpness minimization is as follows:
[0050]
[0051] where ρ restricts the magnitude of the perturbation of the model parameters within the p-norm of θ, y i represents the label of the i-th sample, θ represents the parameters of the model structure, p i represents the prediction result of the model for the i-th sample, ||∈|| p represents the p-norm of the perturbation, and l represents the loss function of a single sample.
[0052] ② To find the optimal perturbation ∈ * , that is, after applying this perturbation, the performance of the model drops most severely. Therefore, the following maximization problem is constructed and the optimal perturbation is approximately solved, where ∈ T represents the transpose of the perturbation vector, L represents the loss function of the entire dataset, and argmax represents the argument for finding the maximum value:
[0053]
[0054] ③ Differentiating the above equation gives the specific form of the gradient update of the model:
[0055]
[0056] where \(o(\theta)\) represents the second-order expansion of \(\theta\). Finally, the calculated optimal perturbation is applied to the model parameters, and the specific form is as follows: denotes the model parameters after adding the perturbation:
[0057]
[0058] (5) Calculate the loss based on the false news prediction results of the multimodal deep learning model, and use the gradient modified by the sensing matrix to update the parameters to train the model.
[0059] ① Taking the cross-entropy loss as an example, the empirical risk minimization of the multimodal deep learning model is expressed as follows:
[0060]
[0061] ② Combining the sensing matrix and the sharpness minimization optimization method, the specific form of the final model parameter update is as follows, where \(\eta\) represents the learning rate hyperparameter of the model: denotes the model at the \(t\)-th update, denotes the model after the latest training update:
[0062]
[0063] Based on the same inventive concept, the present invention also proposes a false news detection system for implementing the above-mentioned interactive multimodal learning method based on flat gradient modification and detecting false news for news. The system includes:
[0064] The first module processes the original sample into an original sample sequence. The original sample is a paired graphic and text multimodal data entity of false news, and constructs a multimodal deep learning model;
[0065] The second module inputs the original sample sequence into the multimodal deep learning model to obtain the false news prediction results and calculates the mean and variance;
[0066] The third module calculates the sensing matrix through the mean and variance, and performs singular value decomposition to find the flat direction;
[0067] The fourth module introduces the sharpness minimization optimization method to further flatten the optimization objective of the multimodal deep learning model;
[0068] The fifth module finally calculates the loss based on the false news prediction results of the multimodal deep learning model, and uses the gradient modified by the sensing matrix to update the parameters to train the model.
[0069] The specific implementation manners of the first to fifth modules are the same as the foregoing method steps and will not be elaborated here.
[0070] Based on a multi-modal deep learning model, the present invention is committed to alleviating the modal imbalance problem in the training process. Considering from the perspective of the learning and optimization of the model, a perception matrix for perceiving the learning situation of each modality in the platform direction is proposed, and the flat information between modalities is exchanged through a multi-modal interactive learning method to achieve modal balance. In multi-modal classification tasks, it has excellent performance and can be further applied to tasks such as fake news detection and aggressive expression detection.
[0071] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. An interactive multimodal learning method based on flat gradient modification, characterized in that: The steps include: Step 1: Process the original samples into original sample sequences, where the original samples are paired image-text multimodal data entities of fake news, and build a multimodal deep learning model; Step 2: Input the original sample sequence into the multimodal deep learning model to obtain the fake news prediction results and calculate the mean and variance; Step 3, calculate the perception matrix through mean and variance, and perform singular value decomposition to find the flat direction; Step 4, introduce the sharpness minimization optimization method to flatten the optimization target of the multimodal deep learning model; Step 5, finally, the loss is calculated based on the fake news prediction results of the multimodal deep learning model, and the modified gradient of the perception matrix is used to update the parameters to train the model.
2. The interactive multimodal learning method based on flat gradient modification according to claim 1, characterized in that: In step 1, obtaining the original sample sequence and constructing a multimodal deep learning model includes the following steps: Initialize the fake news paired multimodal data entities of images and texts. The specific form is: Define the data set D to be of size n and contain m modes; X (j) represents the entire data of the jth mode, Represents the i-th data of the j-th mode; define y i ∈{0,1} c To represent the label of each data, where c represents the total number of categories; Building a multimodal deep learning model includes building a backbone network and a classification network; using the transformer-based BERT model and the Vision Transformer model as the backbone network for processing paired image and text data; the classification network is composed of a linear layer and a nonlinear activation function ReLU; the backbone network and the linear classification layer of the jth modality are denoted as and g (j) , and the corresponding model parameter is recorded as θ (j) and W (j) .
3. The interactive multimodal learning method based on flat gradient modification as claimed in claim 2, characterized in that: In step 2, the original sample sequence is input into the multimodal deep learning model to obtain the false news prediction results and calculate the mean and variance, which specifically includes: ① Input randomly sampled small batches of data into the multimodal deep learning model to obtain prediction results; represents the characteristics of the jth mode of the ith sample, It represents its predicted probability, and softmax represents the normalization function; its features and predicted probabilities are in the following specific forms: ② Given the tth batch of n b Data And get the prediction results of the multimodal deep learning model, calculate the mean and variance of the batch data features, the specific form is as follows: in Represents model input data The features obtained later, represents the mean of the feature, is the characteristic variance; ③ In order to estimate the learning of the model for the entire data set, the variance calculation is generalized to the cumulative variance; the specific form is as follows, where Denotes the cumulative variance:
4. The interactive multimodal learning method based on flat gradient modification as claimed in claim 3, characterized in that: In step 3, the perception matrix is calculated by the mean and variance, and singular value decomposition is performed to find the flat direction, which specifically includes: ① For the cumulative variance of the kth mode obtained after calculation Use singular value decomposition to get the decomposed matrix, the specific form is as follows: in represents the singular value matrix, represents singular value, U (k) and denote the left and right singular vectors respectively, represents the component directions of singular vectors; ② Consider singular vectors The geometric characteristics of the direction indicated, if the feature The singular vectors of The direction indicated imposes a magnitude of The disturbance vector is , where γ is a random number between 0 and 1. The final output change is expressed by the following formula: ③ While keeping the perturbation vector unchanged, a singular vector The flatness of the indicated direction is determined by the size of its singular value. The larger the singular value, the greater the change in the final output. The flat direction is represented by a singular vector with a small singular value. Therefore, the flat gradient modification matrix, that is, the modal perception matrix, is defined. The specific form is as follows: T (k) =V (k) Σ (k) [V (k) ] T Where τ is the scaling parameter, I represents the identity matrix, represents the maximum singular value, Represents the smallest singular value.
5. The interactive multimodal learning method based on flat gradient modification as claimed in claim 4, characterized in that: In step 4, a sharpness minimization optimization method is introduced to flatten the optimization target of the multimodal deep learning model, including the following steps: ① For the multimodal learning objective L(θ), define a perturbation ∈ and apply it to the model parameters; the specific form of the optimization objective function for minimizing sharpness is as follows: Where ρ limits the perturbation size of the model parameters to be within the p-norm of θ, and y i The label θ of the i-th sample represents the parameters of the model structure, p i Represents the model's prediction result for the i-th sample, ||∈|| p It means to find the p-norm of the disturbance, and l represents the loss function of a single sample; ② To find the optimal perturbation ∈ * , construct the following maximization problem and approximately solve the optimal perturbation, where ∈ T represents the transpose of the perturbation vector, L represents the loss function of the entire data set, and argmax represents the independent variable for which the maximum value is sought: ③ Differentiating the above formula, we can get the specific form of the gradient update of the model: Where o(θ) represents the second-order expansion of θ; finally, the calculated optimal perturbation is applied to the model parameters, and the specific form is as follows: Represents the model parameters after adding disturbance.
6. The interactive multimodal learning method based on flat gradient modification according to claim 5, characterized in that: In step 5, the loss is calculated according to the fake news prediction result of the multimodal deep learning model, and the gradient modified by the perception matrix is used to update the parameters to train the model, including the following steps: ①For cross entropy loss, the empirical risk minimization of the multimodal deep learning model is expressed as follows: ② Combining the perception matrix and the sharpness minimization optimization method, the final model parameter update is as follows: Where η represents the learning rate hyperparameter of the model, represents the model updated for the tth time, Represents the model after the latest training update.
7. A fake news detection system based on flat gradient modification, characterized in that: For implementing the method described in any one of claims 1 to 6, the system comprises: The first module processes the original samples into original sample sequences, where the original samples are paired image-text multimodal data entities of fake news, and builds a multimodal deep learning model; In the second module, the original sample sequence is input into the multimodal deep learning model to obtain the false news prediction results and calculate the mean and variance; The third module calculates the perception matrix by the mean and variance, and performs singular value decomposition to find the flat direction; The fourth module introduces the sharpness minimization optimization method and the optimization target of the flat multimodal deep learning model; In the fifth module, the loss is finally calculated based on the fake news prediction results of the multimodal deep learning model, and the modified gradient of the perception matrix is used to update the parameters to train the model.
8. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the steps of the method described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of any method described in claims 1-6 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method described in any one of claims 1 to 6 are implemented.