Multi-modal rumor detection method based on modal alternating optimization

Through the multimodal rumor detection method optimized by modal alternating optimization, the modal laziness problem is solved, the balanced fusion of three-modal features of images, text and social is achieved, and the accuracy and robustness of multimodal rumor detection is improved.

CN120296665AActive Publication Date: 2025-07-11GUANGDONG UNIV OF TECH

Patent Information

Application Number
CN202510393064.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-11
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

There is a problem of modal inertia in the existing multimodal rumor detection technology, which leads to excessive dominance of some modalities and other modal features being ignored or degraded, seriously weakening the reliability and robustness of the detection.

Method used

The multimodal rumor detection method with modal alternating optimization is adopted, and the hyperparameters are dynamically adjusted through phased training strategies and parameter freezing mechanisms to ensure that each modality has an opportunity for balanced optimization, and dynamic balance and collaborative optimization between modals is achieved through adaptive learning rate adjustment and Fisher's information matrix regularization method.

Benefits of technology

Effectively breaking the characterization degradation caused by gradient coupling improves the comprehensive performance of the multimodal rumor detection model and enhances the accuracy and robustness of the detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296665A_ABST
    Figure CN120296665A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multi-modal pattern recognition, in particular to a multi-modal rumor detection method based on modal alternating optimization, which comprises the following steps: constructing and utilizing a training data set DS to train a modal alternating optimization multi-modal rumor detection model, and adopting a staged training strategy for the multi-modal rumor detection model, comprising the steps of only performing parameter updating on a neural network path of a target optimization mode in a specific training stage, freezing corresponding network parameters of other modes, dynamically adjusting hyper-parameters to regulate and control training progresses of different modes, and adjusting the training progress of the different modes when the specific training stage and the non-specific training stage are switched. Historical task key parameters in the fixed multi-modal fusion classifier are evaluated through parameter importance; and inputting a to-be-detected sample into the trained multi-modal rumor detection model, and outputting authenticity probability distribution of the corresponding sample. According to the method, the problem of modal inertia in the existing multi-modal rumor detection technology can be solved, so that balanced optimization of different modal contribution degrees is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimodal pattern recognition, and in particular to a multimodal rumor detection method based on modal alternation optimization. Background Art

[0002] In the era of information explosion, the rapid popularization of social media and the Internet has greatly accelerated the speed of information dissemination, while also exacerbating the spread of rumors. According to a research report from the MIT Media Lab, on the social media Twitter, false information spreads 6 times faster than true information, and the probability of false news being forwarded is 70% higher than that of true news. The rapid spread of such rumors not only disrupts people's daily lives, but may also cause social panic, cause economic losses, and even threaten national security. Especially during emergencies or public health crises, the harmfulness of rumor spread is more prominent. Therefore, how to effectively detect and suppress the spread of rumors has become an important issue that needs to be solved urgently.

[0003] At present, methods for multimodal rumor detection that combine multiple information sources such as text, images, and videos have gradually attracted widespread attention from the industry and academia. However, these methods generally face the problem of modality laziness (ML), that is, in multimodal learning, some modalities may be overly dominant, causing the features of other modalities to be ignored or degraded, thus forming "pseudo-multimodal" decisions. Once the information of the dominant modality is biased or missing, the detection accuracy of the entire model will drop sharply, seriously weakening the reliability and robustness of rumor detection. Summary of the invention

[0004] The purpose of the present invention is to propose a multimodal rumor detection method based on modal alternation optimization to solve the modal inertia problem existing in the existing multimodal rumor detection technology proposed in the background technology, thereby achieving balanced optimization of the contribution of different modalities.

[0005] To achieve this object, the present invention adopts the following technical solutions:

[0006] A multimodal rumor detection method based on modal alternation optimization comprises the following steps:

[0007] Step A: Obtain multimodal datasets published on social media platforms, including images, texts, and social information, annotate the authenticity of the multimodal datasets and perform data preprocessing operations to form a training dataset DS;

[0008] Step B: Construct and use the training dataset DS to train a multimodal rumor detection model with alternating mode optimization. The multimodal rumor detection model adopts a phased training strategy, including optimizing only the target mode M in a specific training phase. tUpdate the parameters of the neural network pathway, freeze the corresponding network parameters of the remaining modalities, and dynamically adjust the hyperparameters to control the training progress of different modalities. When switching between specific training phases and non-specific training phases, fix the key parameters of the historical tasks in the multi-modal fusion classifier through parameter importance evaluation;

[0009] Step C: Input the sample to be detected into the trained multi-modal rumor detection model, and output the authenticity probability distribution of the corresponding sample.

[0010] Preferably, the specific steps of step B are as follows:

[0011] Step B1: Determine the target optimization modality M at the current training stage t t , read the image data I in the training data set DS t , text data T t and social data S t as the input of the multi-modal rumor detection model. The multi-modal rumor detection model includes an image encoder DeiT, a text encoder BERT, and a sentiment analyzer ALBERT. The image encoder DeiT is used to extract the image features in the image data I t The text encoder BERT is used to extract the text features in the text data T The sentiment analyzer ALBERT first analyzes to obtain the sentiment features of the text data T t in

[0012] Step B2: The sentiment analyzer ALBERT first analyzes to obtain the sentiment features of the text data T t Then count the number of special symbols in the text data T to form text statistical features Concatenate the sentiment features t text statistical features and social data S to obtain social features and social data S t to obtain social features

[0013] Step B3: Use the cross-modal cross-attention mechanism to deeply fuse the image features text features and social features of the three modalities to obtain multi-modal features Use the multi-modal fusion classifier to classify the multi-modal features to obtain the multi-modal classification result At the same time, use the independent image classifier, text classifier, and social classifier of the three modalities to classify the image features text features and social features Classify the three modal features to obtain the unimodal classification results and

[0014] Step B4: According to the multimodal classification results Unimodal classification results and Calculate the classification loss and Evaluate the importance of each parameter in the multimodal fusion classifier for historical tasks and determine the regularization term of the loss function accordingly Combine the classification loss and the regularization term of the loss function To obtain the final loss function L t ;

[0015] Step B5: Based on the multimodal classification results Unimodal classification results and Dynamically calculate the learning rate α t of the target optimization modality M t , update the parameters of the neural network path of the target optimization modality M t , freeze the corresponding network parameters of the remaining modalities, and enter the next training stage after the update to repeat the above steps B1 - B4.

[0016] Preferably, the specific steps of the said Step B1 include the following steps:

[0017] Step B11: Determine the target optimization modality M t of the current training stage t through modular arithmetic, expressed as follows:

[0018]

[0019] where T represents the total number of training steps, mod represents the modulo operation, and I, T, and S represent the image modality, text modality, and social modality respectively;

[0020] Step B12: Read a batch of multimodal data from the training dataset DS as the input of the multimodal rumor detection model in the current training stage, including image data I t , text data T t and social data S t , where the social data S t includes the number of user fans, follows, authentication status, post likes, forwards, comments, and geographical location;

[0021] Step B13: Load the pre-trained image encoder DeiT, send the image data I t into the image encoder DeiT, and output the image features It is expressed as follows:

[0022]

[0023] Step B14: Translate the text data T t The WordPiece algorithm is used to split the word into subword units and map them into integer ID sequences, which are expressed as follows:

[0024]

[0025]

[0026] in is the WordPiece word segmentation function, represents sequence concatenation, L=512 represents the maximum input length of the BERT model, Truncate(·,L) represents truncating the sequence to the first L-2 words, |S| represents the length of the sentence, V -1 (·) represents the inverse mapping function of the BERT model vocabulary, and its corresponding integer ID can be obtained through subword lookup table;

[0027] When the length of the ID sequence is greater than L, it is truncated to L bits, otherwise it is padded with 0 elements to L bits. The processed sequence is input into the pre-trained text encoder BERT, and the text features are output after processing. It is expressed as follows:

[0028]

[0029] Where pad_id represents the filler. represents sequence splicing, X 1:L It means to cut off 1 to L bits of X, and |X| means the length of X.

[0030] Preferably, the step B2 specifically comprises the following steps:

[0031] Step B21: Referring to step B14, use the pre-trained sentiment analyzer ALBERT to analyze the text data T t Emotional characteristics It is expressed as follows:

[0032]

[0033] Step B22: Define a special symbol set S, which is expressed as follows:

[0034] S={! ,? ,@,#, / ,url,(),"",...} (8)

[0035] Among them,!,?, @, #, / , url, (), "", and... represent exclamation mark, question mark, user mention symbol, hash sign, slash, website URL, parentheses, double quotes, and ellipsis respectively;

[0036] Initialize the character counter Traverse each character t in the text data T t in it i , if t i is in the set S, increment the corresponding counter, which is expressed as follows:

[0037] C[t i += 1 if t i ∈S (9)

[0038] Vectorize the character counter C to obtain the text statistical feature

[0039] Step B23: Combine the sentiment feature text statistical feature and the social data S t to splice and obtain the social feature which is expressed as follows:

[0040]

[0041] where Concat(·) represents the tensor splicing operation.

[0042] Preferably, the specific steps of step B3 are as follows:

[0043] Step B31: Generate the query, key, and value of the image feature text feature and social feature . For each modality i ∈ {I, T, S}, define:

[0044]

[0045] where j represents other modalities different from modality i, and Q, K, and V represent the query vector, key vector, and value vector respectively, and W q , W k and W v represent the projection matrices of the query vector, key vector, and value vector respectively;

[0046] Step B32: Obtain the enhanced features E I , E T and E S of the three modalities:

[0047] Calculate the attention scores of modality i ∈ {I, T, S}. Taking the image modality I as an example, it is expressed as follows:

[0048]

[0049] where d K represents the dimension of the key vector, softmax represents normalizing the dot product result, and α I represents the degree of attention of the image modality I to each feature in modalities T and S. represents tensor concatenation;

[0050] Extract the context information of other modalities and concatenate it residually with the original modality features. Taking the image modality I as an example, it is expressed as follows:

[0051]

[0052] E I = V I + C I (14)

[0053] where C I represents the context information extracted by the image modality I after paying attention to the important features of the text modality T and the social modality S, V I represents the features of the original modality I, and E I represents the feature representation of the enhanced modality I;

[0054] Step B33: Concatenate the enhanced features E I 、E T and E S of the three modalities i and project them to the target dimension to obtain the final multi-modal fusion feature which is expressed as follows:

[0055]

[0056] where W o represents the projection matrix, b o represents the bias, represents tensor concatenation;

[0057] Step B34: Use the multi-modal fusion classifier to classify the multi-modal fusion feature which is expressed as follows:

[0058]

[0059] where MLP M represents the multi-modal fusion classifier, which consists of a multi-layer perceptron, and softmax represents converting the classifier result into a probability distribution. represents the probability distribution of the authenticity of the rumor obtained by classifying the multi-modal fusion feature at the current training stage t;

[0060] Step B35: Use three modality-independent classifiers to classify and the three modality features respectively to obtain unimodal classification results and which are expressed as follows:

[0061]

[0062] where MLP I , MLP T and MLP S represent the independent classifiers for the image, text, and social modalities respectively, which are composed of multi-layer perceptrons, and softmax represents converting the classifier results into probability distributions. and represent the probability distributions of the authenticity of rumors obtained by classifying the three unimodal features at the current training stage t respectively.

[0063] Preferably, step B4 specifically includes the following steps:

[0064] Step B41: Use the cross-entropy loss function to calculate the loss values of the multimodal classification result the unimodal classification results and respectively, which are expressed as follows:

[0065]

[0066] where N represents the number of samples, R represents the probability distribution of the authenticity of rumors calculated in the previous step, represents the true label, and represent the multimodal classification loss of the multimodal classification result at the current training stage t, the unimodal classification results and and their corresponding loss values;

[0067] Step B42: Calculate the Fisher information matrix F M of the parameters in the multimodal fusion classifier MLP i , which is expressed as follows:

[0068]

[0069] where θ i represents the i-th parameter of the multimodal fusion classifier MLP M , L t-1 represents the loss function of the multimodal fusion classifier MLP M in the previous training stage, represents the derivative of the loss function with respect to the parameter θi The gradient, D represents the data set in the previous training phase, N represents the number of samples in the data set D, and F i represents the Fisher information matrix corresponding to the parameter θ i ;

[0070] Step B43: Use the Fisher information matrix F i to calculate the regularization term of the loss function in the multi-modal rumor detection model which is expressed as follows:

[0071]

[0072] where θ i represents the i-th parameter of the multi-modal fusion classifier MLP M in the current training phase, and represents the i-th parameter of the multi-modal fusion classifier MLP M in the previous training phase, and λ represents the hyperparameter that balances the learning in the new and old training phases;

[0073] Step B44: Combine the multi-modal classification loss the target optimization modal classification loss and the regularization term to obtain the final loss function L in the current training phase t t , which is expressed as follows:

[0074]

[0075] where α is the hyperparameter that balances and the two losses, and the target optimization modal classification loss in the current training phase t and the target optimization modal M t is calculated and determined by formula (1).

[0076] Preferably, the step B5 specifically includes the following steps:

[0077] Step B51: Dynamically calculate the learning rate α of the target optimization modal M in the current training phase t based on the single-modal classification result and t which is expressed as follows: t

[0078]

[0079]

[0080] where α base = 0.001 represents the base learning rate, ​Denote the modality M defined by formula (28) t 's classification accuracy rate Denote non-M t modality's accuracy rate Denote the total number of samples of modality M t , where 1(A) is an indicator function that is 1 when A is true and 0 otherwise Denote that the i-th sample is correctly classified;

[0081] Step B52: After blocking the gradients of the non-target modality neural network path, perform directional backpropagation, that is, only backpropagate the gradient signal through the target modality path and update and iterate its exclusive parameter W t as follows:

[0082] W t = SGD(W t-1 , L t , α t ) (29)

[0083] where SGD represents updating the parameter using the stochastic gradient descent method, and W t-1 represents the parameter updated and completed in the previous training stage;

[0084] Step B53: Enter the next training stage and repeat Steps B1 to B5 until the preset stop condition is met.

[0085] One of the above technical solutions has the following beneficial effects:

[0086] (1) Aiming at the modality inertia problem in multi-modal rumor detection, the present invention innovatively proposes an alternating gradient update mechanism for modality-specific parameters. By establishing a cyclic modality optimization sequence, it ensures that each modality encoder obtains an equal optimization opportunity. This mechanism effectively breaks the problem of representation degradation caused by gradient coupling in traditional joint training, enabling the discriminative features of the image, text, and social three modalities to be fully integrated into the multi-modal decision space and improving the comprehensive performance of the multi-modal rumor detection model.

[0087] (2) The present invention designs an adaptive learning rate adjustment strategy based on the comparison of classification accuracy rates. By monitoring the performance differences of each modality independent classifier in real time, a learning rate reinforcement mechanism is implemented for the weak modality. This negative feedback adjustment method breaks through the limitations of the traditional fixed learning rate, forms a dynamic balance between modalities during the parameter update process, promotes the collaborative optimization of cross-modal representations, and further alleviates the modality inertia problem.

[0088] (3) The present invention constructs an elastic constraint mechanism during the conversion of the modal training phase by introducing an incremental regularization method based on the Fisher information matrix. This method quantitatively evaluates the importance of the parameters of the multi-modal fusion classifier for historical training tasks, and applies elastic constraints to key parameters when optimizing the current mode, thereby effectively preventing the multi-modal fusion classifier from forgetting the learned knowledge during the training of the switching training phase. Description of the Drawings

[0089] Figure 1 is a schematic flow chart of the steps of a multi-modal rumor detection method based on modal alternating optimization according to the present invention;

[0090] Figure 2 is a schematic diagram of the multi-modal rumor detection model structure and the phased training process in a multi-modal rumor detection method based on modal alternating optimization according to the present invention. Detailed Embodiments

[0091] The technical solutions of the present invention will be further described below with reference to the drawings and through specific embodiments.

[0092] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by the terms "center", "longitudinal", "lateral", "length", "width", "thickness", "upper", "lower", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", "axial", "radial", "circumferential", etc. are based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as limiting the present invention.

[0093] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "plurality" is two or more.

[0094] In the description of the present invention, it should be noted that unless otherwise clearly defined and limited, the terms "installed", "connected", and "connected" should be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or an integral connection; it may be directly connected, or indirectly connected through an intermediate medium, and it may be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0095] Such as Figure 1As shown in the figure, this embodiment provides a multi-modal rumor detection method based on modal alternating optimization, including the following steps:

[0096] Step A: Obtain a multi-modal dataset published on a social media platform, including images, text, and social information. Perform annotation and data preprocessing operations on the authenticity of the multi-modal dataset to form a training dataset DS.

[0097] Step B: Construct and use the training dataset DS to train a multi-modal rumor detection model with modal alternating optimization. The multi-modal rumor detection model adopts a phased training strategy, including only updating the parameters of the neural network path of the target optimization modality M t in a specific training stage, freezing the corresponding network parameters of the remaining modalities, and dynamically adjusting hyperparameters to control the training progress of different modalities. When switching between specific training stages and non-specific training stages, fix the key parameters of historical tasks in the multi-modal fusion classifier through parameter importance evaluation.

[0098] Step C: Input the sample to be detected into the trained multi-modal rumor detection model, and output the authenticity probability distribution of the corresponding sample.

[0099] Figure 2 It is the structural diagram of the multi-modal rumor detection model and the schematic diagram of the phased training process based on modal alternating optimization in this embodiment.

[0100] For further explanation, the specific implementation steps of step B are as follows:

[0101] Step B1: Determine the target optimization modality M of the current training stage t t , read the image data I in the training dataset DS t , text data T t and social data S t as the input of the multi-modal rumor detection model. The multi-modal rumor detection model includes an image encoder DeiT, a text encoder BERT, and a sentiment analyzer ALBERT. The image encoder DeiT is used to extract the image features t in the image data I The text encoder BERT is used to extract the text features t in the text data T

[0102] In this embodiment, step B1 specifically includes the following steps:

[0103] Step B11: Determine the target optimization modality M of the current training stage t through modulo operation t , expressed as follows:

[0104]

[0105] Where T represents the total number of training steps, mod represents the modulo operation, and I, T, and S represent the image modality, text modality, and social modality, respectively;

[0106] Step B12: Read a batch of multimodal data from the training dataset DS as the input of the multimodal rumor detection model in the current training stage, including image data I t , text data T t and social data S t , where the social data S t includes the number of user fans, the number of follows, the authentication status, the number of likes, forwards, comments on posts, and the geographical location;

[0107] Step B13: Load the pre-trained image encoder DeiT, and send the image data I t into the image encoder DeiT. After processing, the output image features are represented as follows:

[0108]

[0109] Step B14: Segment the text data T t into sub-word units according to the WordPiece algorithm and map them to an integer ID sequence, which is represented as follows:

[0110]

[0111]

[0112] where is the WordPiece tokenization function, represents sequence concatenation, L = 512 represents the maximum input length of the BERT model, Truncate(·, L) represents truncating the sequence to the first L - 2 words, |S| represents the length of the sentence, and V -1 (·) represents the inverse mapping function of the BERT model vocabulary, and its corresponding integer ID can be obtained by looking up the sub-word table;

[0113] When the length of the ID sequence is greater than L, it is truncated to L bits. Otherwise, it is padded with 0 elements to L bits. The processed sequence is input into the pre-trained text encoder BERT. After processing, the output text features FT t are represented as follows:

[0114]

[0115]

[0116] where pad_id represents the padding symbol, Denotes sequence concatenation, X 1:L Denotes truncating the 1st to Lth bits of X, and |X| represents the length of X.

[0117] Step B2: The sentiment analyzer ALBERT first analyzes to obtain the sentiment features of the text data T t Then count the number of special symbols in the text data T t to form text statistical features The sentiment features Text statistical features and the social data S t are concatenated to obtain social features

[0118] In this embodiment, the specific steps of step B2 include the following steps:

[0119] Step B21: Referring to step B14, use the pre-trained sentiment analyzer ALBERT to analyze and obtain the sentiment features of the text data T t as follows:

[0120]

[0121] Step B22: Define a set of special symbols S, which is expressed as follows:

[0122]

[0123] Among them,!,?, @, #, / , url, (), "" and... represent exclamation mark, question mark, user mention symbol, hash sign, slash, website URL, parentheses, double quotes and ellipsis respectively;

[0124] Initialize the character counter Traverse each character t t in the text data T i , if t i is in the set S, then increment the corresponding counter, which is expressed as follows:

[0125] C[t i += 1 if t i ∈S (9)

[0126] Vectorize the character counter C to obtain text statistical features

[0127] Step B23: The sentiment features Text statistical features and the social data S t are concatenated to obtain social features as follows: ​​

[0128]

[0129] Among them, Conca t (·) represents the tensor splicing operation.

[0130] Step B3: Deeply fuse the image features using the cross-modal cross-attention mechanism Text features and social features The three modal features are used to obtain the multi-modal features Use the multi-modal fusion classifier to classify the multi-modal features to obtain the multi-modal classification result At the same time, use the image classifier, text classifier, and social classifier independent of the three modalities to classify the image features Text features and social features The three modal features are used to obtain the single-modal classification result and

[0131] In this embodiment, the specific steps of step B3 are as follows:

[0132] Step B31: Generate the image features Text features and social features The queries, keys, and values of are defined for each modality i ∈ {I, T, S}:

[0133]

[0134] where j represents other modalities different from modality i, and Q, K, and V represent the query vector, key vector, and value vector respectively, and W q , W k and W v represent the projection matrices of the query vector, key vector, and value vector respectively;

[0135] Step B32: Obtain the enhanced features E I , E T and E S of the three modalities:

[0136] Calculate the attention scores of modality i ∈ {I, T, S}. Taking the image modality I as an example, it is expressed as follows:

[0137]

[0138] where d k represents the key vector dimension, softmax represents normalizing the dot product result, and α IIndicates the degree of attention of image modality I to each feature in modalities T and S. Indicates tensor concatenation;

[0139] Extract the context information of other modalities and concatenate it with the original modality feature residuals. Taking image modality I as an example, it is expressed as follows:

[0140]

[0141] Where C I Indicates the context information extracted by image modality I after paying attention to the important features of text modality T and social modality S, V I Indicates the feature of the original modality I, E I Indicates the feature representation of the enhanced modality I;

[0142] Step B33: Concatenate and project the enhanced features E I 、E T and E S of the three modalities to the target dimension to obtain the final multi-modal fusion feature It is expressed as follows:

[0143]

[0144] Where W o Indicates the projection matrix, b o Indicates the bias, Indicates tensor concatenation;

[0145] Step B34: Use the multi-modal fusion classifier to classify the multi-modal fusion feature It is expressed as follows:

[0146]

[0147] Where MLP M Indicates the multi-modal fusion classifier, which consists of a multi-layer perceptron. Softmax indicates converting the classifier result into a probability distribution. Indicates the probability distribution of the authenticity of the rumor obtained by classifying the multi-modal fusion feature at the current training stage t;

[0148] Step B35: Use the classifiers independent of the three modalities to classify and the three modality features respectively to obtain the single-modal classification results and It is expressed as follows:

[0149]

[0150] Where MLP I 、MLPT and MLP S respectively represent independent classifiers for the three modalities of image, text, and social, which are composed of multi-layer perceptrons. Softmax represents converting the classifier results into a probability distribution. and respectively represent the probability distributions of the authenticity of rumors obtained by classifying with three single-modal features at the current training stage t.

[0151] Step B4: According to the multi-modal classification results single-modal classification results and calculate the classification loss and evaluate the importance of each parameter in the multi-modal fusion classifier for historical tasks and determine the regularization term of the loss function accordingly Combine the classification loss and the regularization term of the loss function to obtain the final loss function L t ;

[0152] In this embodiment, the specific steps of Step B4 are as follows:

[0153] Step B41: Use the cross-entropy loss function to calculate the loss values of the multi-modal classification results single-modal classification results and respectively, which are expressed as follows:

[0154]

[0155]

[0156]

[0157]

[0158] where N represents the number of samples, R represents the probability distribution of the authenticity of rumors calculated in the previous step, represents the true label, and respectively represent the multi-modal classification loss of the multi-modal classification results at the current training stage t and the loss values corresponding to the single-modal classification results and ;

[0159] Step B42: Calculate the Fisher information matrix F M of the parameters in the multi-modal fusion classifier MLP i , which is expressed as follows:

[0160]

[0161] where θ i represents the i-th parameter of the multi-modal fusion classifier MLP M , L t-1 represents the loss function of the multi-modal fusion classifier MLP M in the previous training stage, represents the gradient of the loss function with respect to the parameter θ i , D represents the data set in the previous training stage, N represents the number of samples in the data set D, and F i represents the Fisher information matrix corresponding to the parameter θ i ;

[0162] Step B43: Calculate the regularization term of the loss function in the multi-modal rumor detection model using the Fisher information matrix F i as follows: It is expressed as:

[0163]

[0164] where θ i represents the i-th parameter of the multi-modal fusion classifier MLP M in the current training stage, represents the i-th parameter of the multi-modal fusion classifier MLP M in the previous training stage, and λ represents the hyperparameter that balances the learning in the new and old training stages;

[0165] Step B44: Combine the multi-modal classification loss the target optimization modal classification loss and the regularization term to obtain the final loss function L t in the current training stage t, which is expressed as:

[0166]

[0167] where α is the hyperparameter that balances and the two losses, the target optimization modal classification loss in the current training stage t and the target optimization modal M t is determined by calculating according to formula (1).

[0168] Step B5: Dynamically calculate the learning rate α of the single-modal classification result and the target optimization modal M t based on the multi-modal classification result t , and for the target optimization modal M tThe neural network pathway of the training state is updated, and the corresponding network parameters of the remaining modes are frozen. After the update is completed, the next training phase is entered to repeat the above steps B1-B4.

[0169] In this embodiment, step B5 specifically includes the following steps:

[0170] Step B51: Based on the unimodal classification results and Dynamically calculate the target optimization mode M of the current training stage t t The learning rate α t , which is expressed as follows:

[0171]

[0172] where α base =0.001 represents the baseline learning rate, represents the mode M defined by formula (28) t The classification accuracy of Indicates non-M t The accuracy of the modality, Indicates mode M t The total number of samples, 1(A) is the indicator function, the value is 1 when A is true, otherwise it is 0, Indicates that the i-th sample is correctly classified;

[0173] Step B52: After blocking the gradient of the non-target modality neural network pathway, perform directed back propagation, that is, only return the gradient signal through the target modality pathway and adjust its exclusive parameter W t Perform update iterations, as follows:

[0174] W t =SGD(W t-1 ,L t ,α t ) (29)

[0175] Where SGD means using stochastic gradient descent to update parameters, W t-1 Indicates the parameters updated in the previous training phase;

[0176] Step B53: Enter the next training phase and repeat steps B1 to B5 until a preset stop condition is met.

[0177] The technical principle of the present invention is described above in conjunction with specific embodiments. These descriptions are only for explaining the principle of the present invention and cannot be interpreted as limiting the scope of protection of the present invention in any way. Based on the explanations herein, those skilled in the art can associate other specific embodiments of the present invention without creative work, and these equivalent variations or substitutions are all included in the scope defined by the claims of this application.

Claims

1. A multi-modal rumor detection method based on modal alternating optimization, characterized in that It includes the following steps: Step A: Obtain a multi-modal data set published on a social media platform, including images, text, and social information. Perform annotation and data preprocessing operations on the authenticity of the multi-modal data set to form a training data set DS; Step B: Construct and use the training dataset DS to train a multimodal rumor detection model with alternating optimization of modalities. The multimodal rumor detection model adopts a phased training strategy, including updating the parameters of the neural network pathway of the target optimization modality M only in a specific training stage, freezing the corresponding network parameters of the remaining modalities, and dynamically adjusting hyperparameters to control the training progress of different modalities. When switching between specific training stages and non-specific training stages, fix the key parameters of the historical tasks in the multimodal fusion classifier through parameter importance evaluation; t During the conversion between specific training stages and non-specific training stages, fix the key parameters of the historical tasks in the multimodal fusion classifier through parameter importance evaluation; Step C: Input the sample to be detected into the trained multi-modal rumor detection model, and output the authenticity probability distribution of the corresponding sample.

2. The multimodal rumor detection method based on modal alternation optimization according to claim 1, wherein The specific steps of Step B include the following steps: Step B1: Determine the target optimization modality M for the current training stage t t , read the image data I in the training dataset DS t , text data T t and social data S t as the input of the multimodal rumor detection model, where the multimodal rumor detection model includes an image encoder DeiT, a text encoder BERT, and a sentiment analyzer ALBERT, and the image encoder DeiT is used to extract the image features in the image data I t The text encoder BERT is used to extract the text features in the text data T t ​​ Step B2: The sentiment analyzer ALBERT first analyzes to obtain the sentiment features of the text data T t of the text data T Then, it counts the number of special symbols in the text data T t to form text statistical features The sentiment features text statistical features and the social data S t are concatenated to obtain social features Step B3: Deeply fuse image features using a cross-modal cross-attention mechanism Text features and social features to obtain multi-modal features from the three modal features Use a multi-modal fusion classifier to classify the multi-modal features to obtain a multi-modal classification result At the same time, use three modality-independent image classifiers, text classifiers, and social classifiers to classify the image features Text features and social features from the three modal features to obtain single-modal classification results and Step B4: Based on the multi-modal classification results single-modal classification results and calculate the classification loss and evaluate the importance of each parameter in the multi-modal fusion classifier for historical tasks and determine the regularization term of the loss function accordingly Combine the classification loss and the regularization term of the loss function to obtain the final loss function L t ; Step B5: Based on the multi-modal classification results single-modal classification results and dynamically calculate the learning rate α t for the target optimization modality M t , update the parameters of the neural network path of the target optimization modality M t , freeze the corresponding network parameters of the remaining modalities, and enter the next training phase to repeat the above steps B1 - B4 after the update is completed.

3. A multimodal rumor detection method based on modal alternation optimization according to claim 2, characterized in that, The specific steps of Step B1 include the following steps: Step B11: Determine the target optimization modality M for the current training stage t through modular arithmetic t , which is expressed as follows: Where T represents the total number of training steps, mod represents the modulo operation, and I, T, and S represent the image modality, text modality, and social modality respectively; Step B12: Read a batch of multimodal data from the training dataset DS as the input of the multimodal rumor detection model for the current training stage, including image data I t , text data T t and social data S t , where the social data S t includes the number of user fans, the number of follows, the authentication status, the number of likes, reposts, comments on posts, and the geographical location; Step B13: Load the pre-trained image encoder DeiT and input the image data I t into the image encoder DeiT. After processing, the image features are output which are represented as follows: Step B14: Take the text data T t and segment it into sub-word units according to the WordPiece algorithm and map it into a sequence of integer IDs, which is shown as follows: where s i ∈S (4) Among them is the WordPiece tokenization function, represents sequence concatenation. L = 512 represents the maximum input length of the BERT model. Truncate(·, L) represents truncating the sequence to the first L - 2 words. |S| represents the length of the sentence, and V -1 (·) represents the inverse mapping function of the BERT model vocabulary, and its corresponding integer ID can be obtained by looking up the sub-word table; When the length of the ID sequence is greater than L, truncate it to L bits; otherwise, pad it with 0 elements to L bits. Input the processed sequence into the pre-trained text encoder BERT, and output text features after processing. It is represented as follows: Where pad_id represents the filler. represents sequence splicing, X 1:L It means to cut off 1 to L bits of X, and |X| means the length of X.

4. A multi-modal rumor detection method based on modal alternation optimization according to claim 3, characterized in that, The specific steps of Step B2 include the following steps: Step B21: Referring to Step B14, use the pre-trained sentiment analyzer ALBERT to analyze and obtain the sentiment features of the text data T t as follows: which are shown as follows: Step B22: Define a set of special symbols S, which is expressed as follows: S = {!,?, @, #, / , url, (), "",...} (8) Where!,?, @, #, / , url, (), "", and... represent the exclamation mark, question mark, user mention symbol, hash sign, slash, website address, parentheses, double quotes, and ellipsis respectively; Initialize the character counter Traverse the text data T t for each character t i in it. If t i is in the set S, increment the corresponding counter as follows: C[t i + = 1 if t i ∈S (9) Vectorize the character counter C to obtain text statistical features Step B23: Combine the sentiment feature text statistical feature and social data S t to obtain the social feature which is expressed as follows: Where Concat(·) represents the tensor concatenation operation.

5. A multimodal rumor detection method based on modal alternation optimization according to claim 4, characterized in that The specific steps of Step B3 include the following steps: Step B31: Generate image features Text features and social features For the queries, keys, and values of each modality \(i\in\{I, T, S\}\), define: where j represents other modalities different from modality i, and Q, K, and V respectively represent the query vector, key vector, and value vector, and W q , W k , and W v respectively represent the projection matrices of the query vector, key vector, and value vector; Step B32: Obtain the enhanced features E of three modalities i I , E T and E S : Calculate the attention score of modality i ∈ {I, T, S}. Taking the image modality I as an example, it is expressed as follows: where d k represents the dimension of the key vector, softmax represents normalizing the dot product result, and α I represents the degree of attention of the image modality I to each feature in modalities T and S, represents tensor concatenation; Extract the context information of other modalities and perform residual connection with the original modality features. Taking the image modality I as an example, it is expressed as follows: E I = V I + C I (14) Among which C I represents the context information extracted from the image modality I by focusing on the important features of the text modality T and the social modality S, V I represents the features of the original modality I, E I represents the feature representation of the enhanced modality I; Step B33: Concatenate the features E enhanced by the three modalities i I , E T and E S , project them onto the target dimension, and obtain the final multi-modal fusion feature which is expressed as follows: Among which W o represents a projection matrix, b o represents a bias, represents tensor concatenation; Step B34: Classify the multi-modal fusion features using a multi-modal fusion classifier as follows: Among them, MLP M represents a multi-modal fusion classifier, which is composed of a multi-layer perceptron. Softmax represents converting the classifier results into a probability distribution. represents the probability distribution of rumor authenticity obtained by classifying using multi-modal fusion features at the current training stage t; Step B35: Use three modality-independent classifiers to classify and the three modality features respectively to obtain unimodal classification results and as follows: Among them, MLP I , MLP T and MLP S respectively represent independent classifiers for the three modalities of image, text, and social, which are composed of multi-layer perceptrons. Softmax represents converting the classifier results into a probability distribution. and respectively represent the probability distributions of rumor authenticity obtained by classifying using the three single-modal features at the current training stage t.

6. The multimodal rumor detection method based on modal alternation optimization according to claim 5, characterized in that The specific steps of Step B4 include the following steps: Step B41: Calculate the multi-modal classification results using the cross-entropy loss function respectively Unimodal classification results and The loss values are expressed as follows: where \(N\) represents the number of samples, and \(R\) represents the probability distribution of rumor authenticity calculated in the previous step, represents the true label, and respectively represent the multi-modal classification result at the current training stage \(t\), the multi-modal classification loss, the single-modal classification result and the corresponding loss value; Step B42: Calculate the Fisher information matrix F of the parameters in the multi-modal fusion classifier MLP M in the multi-modal fusion classifier MLP i , which is expressed as follows: where θ i represents the i-th parameter of the multi-modal fusion classifier MLP M , L t-1 represents the loss function of the multi-modal fusion classifier MLP M in the previous training stage, denotes the gradient of the loss function with respect to the parameter θ i , D represents the data set in the previous training stage, N represents the number of samples in the data set D, and F i represents the Fisher information matrix corresponding to the parameter θ i ; Step B43: Utilize the Fisher information matrix F i to calculate the regularization term of the loss function in the multi-modal rumor detection model which is expressed as follows: where θ i represents the i-th parameter of the multi-modal fusion classifier MLP M at the current training stage, represents the i-th parameter of the multi-modal fusion classifier MLP M at the previous training stage, and λ represents the hyperparameter that balances the learning of the new and old training stages; Step B44: Combine the multi-modal classification loss The objective is to optimize the modal classification loss and the regularization term to obtain the final loss function L for the current training stage t t , which is expressed as follows: where α is the hyperparameter for balancing and the two losses, and the objective optimization modal classification loss at the current training stage t and the objective optimization modality M t is determined by calculation according to formula (1).

7. A multi-modal rumor detection method based on modal alternation optimization according to claim 6, characterized in that, The specific steps of Step B5 include the following steps: Step B51: According to the single-modal classification result and dynamically calculate the learning rate α of the target optimization modality M in the current training stage t t , which is expressed as follows: t , as shown below: where α base = 0.001 represents the base learning rate, represents the classification accuracy of modality M defined by formula (28), t the classification accuracy, represents the accuracy of non-M t modalities, represents the total number of samples of modality M t 1(A) is an indicator function that is 1 when A is true and 0 otherwise, indicates that the i-th sample is correctly classified; Step B52: After blocking the gradients of the non-target modality neural network pathways, perform directed backpropagation, that is, only backpropagate the gradient signals through the target modality pathways and update and iterate its exclusive parameter W t as follows: W t = SGD(W t-1 , L t , α t ) (29) Among them, SGD represents updating the parameters using the stochastic gradient descent method, and W t-1 represents the parameters updated at the end of the previous training stage; Step B53: Enter the next training stage and repeat Steps B1 to B5 until the preset stop condition is met.

Citation Information

Patent Citations

  • Multi-modal pre-training model migration method based on self-supervised learning

    CN118097685A

  • Multi-modal rumor detection method and system fusing multi-granularity features

    CN119166907A

  • Method, electronic device, and computer program product for generating cross-modality encoder

    US20240371358A1

Cited By

  • Multi-modal large model learning method and system based on instantaneous detection and rebalance, and storage medium

    CN121234026A