Error information detection method and system based on cross-modal difference classification

By extracting the modal features of images and text and calculating the difference characteristics between modals, generating prototype features and calculating similarity scores, the problem of the false information detection method in the prior art weakening modal differences in the information fusion process is solved, and stable detection and high accuracy are achieved under the condition of modal loss.

CN120045945AActive Publication Date: 2025-05-27NANKAI UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510048104.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-27
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

Existing multimodal false information detection methods may weaken the differential expression between image and text modality during information fusion, resulting in inaccurate predictions when detecting false content such as sarcastic and misleading information.

Method used

Using an error information detection method based on cross-modal difference classification, the training data is acquired and preprocessed, the modal features of images and text are extracted, and the differential features between modalities are calculated. Then, prototype features are generated and similarity scores are calculated, and prototype features are adjusted according to the real results of the training data to improve the accuracy of detection.

Benefits of technology

It realizes stable detection under modal loss conditions, improves the model's ability to interpret different modal features, and significantly improves the accuracy and robustness of false information detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045945A_ABST
    Figure CN120045945A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of error information detection, and particularly relates to an error information detection method and system based on cross-modal difference classification. Comprising the following steps: acquiring training data and preprocessing the data to obtain data set data; carrying out image-text modal feature extraction on the data of the data set to obtain an image modal and a text modal of the data of the data set; performing inter-modal difference feature extraction to obtain difference information; generating prototype features according to the image modality, the text modality and the difference information, and splicing the image modality, the difference information and the text modality to obtain sample features; according to the prototype features and the sample features, performing training to obtain judgment prototype features; and judging whether the prototype feature judgment information is correct or not according to the to-be-detected data. The technical problems that a classifier method has high requirements for modal integrity of data and prediction is inaccurate are solved, and the technical effect that the modal integrity needs to be low and prediction is accurate is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of error information detection, and particularly relates to an error information detection method and system based on cross-modal difference classification. Background Art

[0002] With the rapid development of the Internet and social media, the spread of false information (such as fake news, satirical content, etc.) on online platforms has become more extensive and rapid, having a significant negative impact on society and individuals. Traditional false information detection methods mainly rely on information in a single modality (such as text), but with the explosion of multimedia content, multi-modal false information detection has gradually become a research hotspot. Existing multi-modal false information detection methods usually map multi-modal data such as images and texts into a shared embedding space and perform classification through information fusion between modalities. However, in the process of information fusion, these methods may weaken the differential expression between the image and text modalities, and this inter-modal difference is often of great significance in detecting false content such as satire and misleading information. Summary of the Invention

[0003] The present invention aims to solve at least one of the technical problems existing in the related art. For this purpose, the present invention provides an error information detection method and system based on cross-modal difference classification, aiming to solve the technical problems that the MLP classifier method has high requirements for the modal integrity of data and inaccurate prediction.

[0004] The present invention provides an error information detection method based on cross-modal difference classification, including the following steps: S1: Obtain training data and preprocess the data to obtain dataset data; extract image-modal and text-modal features from the dataset data to obtain the image modality and text modality of the dataset data; S2: Extract inter-modal difference features from the image modality and the text modality to obtain difference information; S3: Generate prototype features according to the image modality, text modality, and difference information, splice the image modality, difference information, and text modality to obtain sample features; S4: Calculate a similarity score according to the prototype features and the sample features, and adjust the prototype features according to the true results of the training data to obtain judgment prototype features; S5: Extract the image modality of the data to be detected, the text modality of the data to be detected, and the difference information of the data to be detected, calculate the similarity score of the data to be detected according to the image modality of the data to be detected, the text modality of the data to be detected, the difference information of the data to be detected, and the judgment prototype features, and judge whether the information is correct according to the similarity score of the data to be detected.

[0005] According to an error information detection method based on cross-modal difference classification provided by the present invention, the text data of the data set is encoded by a deep learning model encoder to obtain a text modality; the text data of the data set is encoded by a deep learning model encoder to obtain a text modality; the image data of the data set is encoded by a convolutional neural network encoder to obtain an image modality: wherein, is the image modality; is the text modality; is the encoding by the convolutional neural network encoder; is the encoding by the deep learning model encoder; is the image data; is the text data.

[0006] According to an error information detection method based on cross-modal difference classification provided by the present invention, the deep learning model encoder is a BERT (Bidirectional Encoder Representations from Transformers, a bidirectional encoder using Transformer) encoder, and the convolutional neural network encoder is a ResNet34 encoder.

[0007] According to an error information detection method based on cross-modal difference classification provided by the present invention, step S2 includes: S21: Input the image modality and the text modality into a joint encoder to obtain the image modality in the embedding space and the text modality in the embedding space : wherein, is the joint encoder processing function; S22: Input the image modality and the text modality into a joint encoder to obtain the image modality in the embedding space and the text modality in the embedding space : wherein, is the joint encoder processing function; S23: Calculate The difference in modal information in the embedding space and Modal difference information in embedding space: in, for The modal difference information in the embedding space, for Modal difference information in embedding space; S24: Calculate modal difference information : ; S25: Use a multi-layer perceptron to align the modal difference information to obtain the difference information : .

[0008] According to a method for detecting error information based on cross-modal difference classification provided by the present invention, The joint encoder includes u linear layers and v activation layers. The activation layer of the joint encoder uses the ReLU function; Said The joint encoder contains m linear layers and n activation layers. The activation layer of the joint encoder uses the ReLU function.

[0009] According to the present invention, a method for detecting erroneous information based on cross-modal difference classification, step S23, comprises: like When Normalize using the infinity norm: in, is the comparison threshold, is the infinite norm; like When Normalize using the infinity norm: ; like When Normalize using the infinity norm: ; like When Normalize using the infinity norm: .

[0010] Step S3 of an error information detection method based on cross-modal difference classification provided by the present invention includes: S31: Generate initial prototype features using Gaussian initialization according to the vector lengths of the image modality, text modality, and difference information , wherein, is the initial image prototype feature, is the initial difference prototype feature, is the initial text prototype feature; S32: Perform Schmidt orthogonalization on the initial prototype features so that the initial image prototype feature, initial difference prototype feature, and initial text prototype feature are orthogonal to each other to obtain prototype features : wherein, is the image prototype feature, is the difference prototype feature, is the text prototype feature; S33: Concatenate the image modality, the text modality, and the difference information to obtain the sample feature : .

[0011] Step S4 of an error information detection method based on cross-modal difference classification provided by the present invention includes: S41: Calculate the similarity score : S42: Adjust the prototype features using an optimizer according to the true results of the training data to obtain judgment prototype features.

[0012] Step S41 of an error information detection method based on cross-modal difference classification provided by the present invention includes: When the training data lacks the image modality, the similarity score calculation formula is: When the training data lacks the text modality, the similarity score calculation formula is: .

[0013] The present invention also provides an error information detection system based on cross-modal difference classification, including: Preprocessing module: Obtain training data and preprocess the data to obtain dataset data; extract graphic-modal features from the dataset data to obtain the image modality and text modality of the dataset data; extract inter-modal difference features from the image modality and the text modality to obtain difference information; generate prototype features according to the image modality, text modality and difference information, and splice the image modality, difference information and text modality to obtain sample features; Judgment prototype feature training module: Calculate a similarity score according to the prototype feature and the sample feature, and adjust the prototype feature according to the true result of the training data to obtain a judgment prototype feature; Prediction module: Extract the image modality of the data to be detected, the text modality of the data to be detected and the difference information of the data to be detected, calculate the similarity score of the data to be detected according to the image modality of the data to be detected, the text modality of the data to be detected, the difference information of the data to be detected and the judgment prototype feature, and judge whether the information is correct according to the similarity score of the data to be detected.

[0014] One or more of the above technical solutions in the embodiments of the present invention have at least one of the following technical effects: An error information detection method and system based on cross-modal difference classification provided by the present invention. By utilizing the information difference between the image and text modalities and combining a prototype-based classifier, stable detection under the condition of modality missing is achieved, and the model's ability to interpret different modality features is improved.

[0015] The additional aspects and advantages of the present invention will be partly given in the following description, partly will become obvious from the following description, or will be understood through the practice of the present invention. Description of the Drawings

[0016] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a flowchart of the error information detection method based on cross-modal difference classification provided by the present invention.

[0018] Figure 2 It is a structural block diagram of the error information detection device based on cross-modal difference classification provided by the present invention.

[0019] Reference Signs: 101. Preprocessing module; 102. Judgment prototype feature training module; 103. Prediction module. Detailed implementation manners

[0020] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention. The following embodiments are used to illustrate the present invention, but cannot be used to limit the scope of the present invention.

[0021] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the embodiments of the present invention. In this specification, the schematic descriptions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0022] Embodiment The following combines Figure 1 and Figure 2 to describe the present invention.

[0023] As Figure 1 shown, the present invention provides a method for detecting error information based on cross-modal difference classification, and the steps are as follows: S1: Obtain training data and preprocess the data to obtain dataset data; extract image-modal and text-modal features from the dataset data to obtain the image modality and text modality of the dataset data; S2: Extract inter-modal difference features from the image modality and the text modality to obtain difference information; S3: Generate prototype features according to the image modality, text modality and difference information, and splice the image modality, difference information and text modality to obtain sample features; S4: Calculate a similarity score according to the prototype features and the sample features, and adjust the prototype features according to the true results of the training data to obtain judgment prototype features; S5: Extract the image modality, text modality, and difference information of the data to be detected. Calculate the similarity score of the data to be detected based on the image modality, text modality, difference information of the data to be detected, and the judgment prototype features, and determine whether the information is correct according to the similarity score of the data to be detected.

[0024] To verify the effectiveness and robustness of the present invention, three publicly available multi-modal datasets are used for training in the embodiments of the present invention, including: 1. Weibo dataset: The Weibo dataset is a publicly available dataset for detecting Chinese multi-modal false information, containing image-text pairs on social media platforms. The dataset is divided into two categories: rumor samples and non-rumor samples. The training set contains 3,749 rumor samples and 3,783 non-rumor samples.

[0025] 2. Pheme dataset: The Pheme dataset is an English multi-modal false information detection dataset, focusing on the detection of social media rumors. The dataset contains 590 rumor samples and 1,428 non-rumor samples, and is suitable for evaluating the false information detection ability of the model in a multi-modal environment. The Pheme dataset contains different modal missing situations. For example, about 36.75% of the samples lack the image modality. The present invention uses the Pheme dataset to verify the stability of the prototype-based classifier in the case of modal missing, and proves the effectiveness of the dynamic masking mechanism in the scenario of incomplete modality.

[0026] 3. Sarcasm dataset: The Sarcasm dataset is an English multi-modal sarcasm detection dataset, containing image-text pairs of sarcastic content and non-sarcastic content. The training set contains 8,642 sarcastic samples and 11,174 non-sarcastic samples, and the test set contains 959 sarcastic samples and 1,450 non-sarcastic samples.

[0027] To ensure the comparability of the experimental results, the embodiments of the present invention need to first perform data partitioning on the three datasets. The Weibo dataset and the Sarcasm dataset use the partitioning method of LogicDM (Logical Data Model); the Pheme dataset uses the partitioning method of MFAN (Multi-modal Feature-enhanced Attention Network) to obtain the training data.

[0028] After that, the present invention preprocesses the training data to obtain the dataset data; extracts the image modality and text modality of the dataset data by performing graphic and text modal feature extraction on the dataset data; Encode the text data of the dataset data using a deep learning model encoder to obtain a text modality; encode the image data of the dataset data using a convolutional neural network encoder to obtain an image modality: Among them, is the image modality; is the text modality; is the encoding by the convolutional neural network encoder; is the encoding by the deep learning model encoder; is the image data; is the text data. In the embodiment of the present invention, the deep learning model encoder is a BERT encoder, and in the embodiment of the present invention, the convolutional neural network encoder is a ResNet34 encoder.

[0029] Specifically, in order to obtain the differences between image and text features in the embodiment of the present invention, the embodiment of the present invention first maps them into the same semantic space to a certain extent through a joint encoder, so as to facilitate subsequent processing of different modality features and obtain difference information. The specific steps are as follows: Input the image modality and the text modality into the joint encoder to obtain the image modality in the embedding space and the text modality in the embedding space : Among them, is the joint encoder processing function; Input the image modality and the text modality into the joint encoder to obtain the image modality in the embedding space and the text modality in the embedding space : Among them, is the joint encoder processing function. In particular, in the embodiment of the present invention, the joint encoder includes u layers of linear layers and v layers of activation layers; the joint encoder includes m layers of linear layers and n layers of activation layers. In the embodiment of the present invention, the joint encoder and the joint encoder have 4 layers of linear layers and 1 layer of activation layers. The Joint encoder and The activation layer of the joint encoder uses the ReLU function.

[0030] Calculate The modal information difference in the embedding space and The modal difference information in the embedding space: Among them, is the modal difference information in the embedding space, is the modal difference information in the embedding space; Calculate the modal difference information : ; Use a multi-layer perceptron to perform an alignment operation on the modal difference information to obtain the difference information : .

[0031] Since for different samples and different data sets, the model may have different degrees of bias for the features of different modalities, resulting in inconsistent weight sizes of the generated features of different modalities. When the difference in weight distribution between modalities is very large, it may have an adverse effect on the calculation of the difference information, resulting in the final difference information not really obtaining the difference features between different modalities well, but relying too much on one of the modalities. Therefore, in the embodiments of the present invention, features with too large differences between modalities are all constrained by the ∞ norm, and the numerical value range of features with too large differences is constrained to between. The specific processing method is: If when, Use the infinite norm for normalization processing: ; If when, Use the infinite norm for normalization processing: ; If when, Use the infinite norm for normalization processing: ; If when, Use the infinite norm for normalization processing: ; Among them, is the comparison threshold. In the embodiments of the present invention , is the infinity norm.

[0032] Generate prototype features according to the image modality, text modality and difference information, splice the image modality, difference information and text modality to obtain sample features; calculate a similarity score according to the prototype features and the sample features, and adjust the prototype features according to the true results of the training data to obtain judgment prototype features.

[0033] Generate initial prototype features using Gaussian initialization according to the vector lengths of the image modality, text modality and difference information , Among them, is the initial image prototype feature, is the initial difference prototype feature, is the initial text prototype feature; Perform Schmidt orthogonalization on the initial prototype features so that the initial image prototype feature, initial difference prototype feature and initial text prototype feature are orthogonal to each other to obtain prototype features : Among them, is the image prototype feature, is the difference prototype feature, is the text prototype feature; Splice the image modality, the text modality and the difference information to obtain the sample features : .

[0034] Calculate the similarity score : Adjust the prototype features according to the true results of the training data using an optimizer to obtain judgment prototype features. Adjust the similarity score calculation formula using a prototype classifier: When the training data lacks the image modality, the similarity score calculation formula is: When the training data lacks the text modality, the similarity score calculation formula is: .

[0035] Specifically, the method for determining whether the data to be detected is correct in the embodiments of the present invention is as follows: extract the image modality of the data to be detected, the text modality of the data to be detected, and the difference information of the data to be detected, calculate the similarity score of the data to be detected according to the image modality of the data to be detected, the text modality of the data to be detected, the difference information of the data to be detected, and the judgment prototype features. When the similarity score is less than the judgment threshold, it is judged as false information; when the similarity score is greater than the judgment threshold, it is judged as true information.

[0036] The effectiveness and advantages of the embodiments of the present invention are demonstrated below: Table 1 Performance comparison table of different detection methods on Weibo and Pheme datasets

[0037] Table 2 Performance comparison table of different detection methods on Sarcasm datasets

[0038] Among them, EANN (Event adversarial neural networks), MVAE (Multimodal variational autoencoder), Spotfake, Spotfake+, MCAN (Multimodal fusion with co-attention network), BMR (Bootstrapping multi-view representations), LogicDM+ (Interpretable multimodal misinformation detection with logic reasoning+), FSRU (Frequency Spectrum Representation and fUsion network), MFAN+ (Multi-modal Feature-enhanced Attention Network+), MMCAN-Res (Multimodal Matching-aware Co-Attention Network-Res), Bert (Bidirectional Encoder Representations from Transformers), ViT (Vision Transformer), HFM (Hierarchical Fusion Model), D&R Net (Decomposition and Relation Network), Att-Bert, InCrossMGs (in-modal and cross-modal graphs), HCM (Hierarchical Congruity Modeling), and LogicDM (Interpretable multimodal misinformation detection with logic reasoning) are all other misinformation detection techniques.

[0039] As shown in Table 1 and Table 2: On the Weibo, Pheme, and Sarcasm datasets, the embodiments of the present invention outperform existing advanced methods in multiple metrics such as accuracy, precision, recall, and F1-score, demonstrating excellent detection capabilities. For example, on the Weibo dataset, the F1-score of the embodiments of the present invention reaches 92.5%, which is a 1% improvement compared to the traditional MLP classifier. On the Pheme dataset and the Sarcasm dataset, the F1-scores of the embodiments of the present invention are improved by 1.7% and 4.4% respectively, and especially show strong generalization ability in the sarcasm detection task, verifying the adaptability of the model in different scenarios and language environments.

[0040] Since the embodiments of the present invention have a very strong decoupling ability for features, they also have a very strong effect on sample classification in the case of missing sample modalities. Therefore, the embodiments of the present invention detect the adaptability and better performance of the model of the embodiments of the present invention in the environment of missing sample modalities from two aspects respectively.

[0041] First, the embodiments of the present invention conduct tests from the level of the number of features used. The embodiments of the present invention test the capabilities of the model in three datasets under different usage conditions for three types of features: images, texts, and information difference features. On the one hand, the feature decoupling ability of the embodiments of the present invention can be tested, and on the other hand, the dependence degree of the model on different types of features and the effectiveness of the features can also be reflected. The experimental results are shown in Table 3.

[0042] Table 3 Performance comparison table for the environment of missing modalities and different features

[0043] Among them, indicates the existence of this feature type.

[0044] Table 3 shows the results of the feature ablation experiments of the model of the present invention on three datasets, Weibo, Pheme, and Sarcasm, to analyze the contributions of image features, text features, and inter-modal difference features to the classification performance. The experiments show that when only using a single feature, the inter-modal difference feature has better classification performance than the image feature or the text feature on all datasets, indicating that the inter-modal difference feature can effectively capture the contradictory information in multi-modal data and plays an important role. In the joint feature experiments, the combination of any two features can significantly improve the model performance, especially the combination of the difference feature and the text or image feature, further verifying the importance of the difference feature in enhancing the expression of multi-modal information. When the image, text, and inter-modal difference features are used jointly, the classification accuracy and F1 score of the model reach the highest on all datasets, fully demonstrating the advantages of multi-modal feature fusion. The results also show that the model of the present invention can effectively decouple the feature contributions through the dynamic masking mechanism, thus showing high classification performance and robustness under different feature combinations.

[0045] Secondly, to prove that the embodiment of the present invention has a strong feature decoupling ability, in the embodiment of the present invention, while keeping other structures of the model the same, the final prototype classifier is replaced with an MLP classifier. On the one hand, it can test the effective utilization of the ability based on the inter-modal information difference of the embodiment of the present invention. On the other hand, it can better show that the classification method based on the prototype of the embodiment of the present invention can better decouple different category features and obtain better results in the case of modal missing. The final results are shown in Table 4.

[0046] Table 4 Comparison table of modal missing performance between MLP structure and prototype classifier

[0047] Table 4 shows the experimental results of the modality missing scenarios of the embodiments of the present invention on the Weibo dataset, which are used to evaluate the classification performance of the model in the case of missing image or text modalities. In the full modality scenario, the highest accuracy and F1 score are achieved based on the prototype classifier, which outperforms the MLP-based classifier, verifying the advantages of the prototype classifier in processing multi-modal inputs. In the single modality scenario, the embodiments of the present invention still show strong adaptability. Especially when there is only text input, the classification accuracy only drops by 1.4% compared to the full modality scenario, which is significantly better than the scenario with only image input, indicating that the text modality provides more discriminative information in false information detection. In addition, compared with the traditional MLP classifier, the performance improvement of the embodiments of the present invention in the single modality scenario is more significant. When there is only image input, the accuracy increases by 12%, further proving the robustness and applicability of the dynamic masking mechanism and the prototype classifier in the modality missing scenario. Generally speaking, the experimental results show that the model of the present invention can effectively handle the scenario of incomplete modalities. Especially when the text modality is retained, the detection performance is close to the full modality scenario, demonstrating great practical application value.

[0048] Table 5 Comparative Table of Ablation Performance of Model Structures

[0049] Among them, -w / o is the abbreviation of without, β is the β joint encoder, and ∞-norm is the infinity norm normalization.

[0050] Through the ablation experiment analysis in Table 5, each module plays a significant role in improving the overall performance of the model. The bidirectional difference calculation (β joint encoder) effectively captures the inconsistency of information between modalities, and the ∞-norm normalization constraint balances the numerical range of modality features. The prototype-based classifier is significantly superior to the traditional MLP classifier in terms of classification accuracy and adaptability. The synergistic effect of these modules enables the model to exhibit excellent performance in both full modality and modality missing cases.

[0051] As Figure 2 shown, the present invention also provides a system including: Preprocessing module 101: Obtain training data and preprocess the data to obtain dataset data; extract image and text modality features from the dataset data to obtain the image modality and text modality of the dataset data; extract inter-modal difference features from the image modality and the text modality to obtain difference information; generate prototype features according to the image modality, text modality and difference information, and splice the image modality, difference information and text modality to obtain sample features; Judgment prototype feature training module 102: Calculate a similarity score based on the prototype feature and the sample feature, and adjust the prototype feature according to the true result of the training data to obtain a judgment prototype feature; Prediction module 103: Extract the image modality of the data to be detected, the text modality of the data to be detected, and the difference information of the data to be detected, calculate the similarity score of the data to be detected based on the image modality of the data to be detected, the text modality of the data to be detected, the difference information of the data to be detected, and the judgment prototype feature, and judge whether the information is correct based on the similarity score of the data to be detected.

[0052] Finally, it should be noted that: The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: They can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting false information based on cross-modal difference classification, characterized in that: The following steps are involved: S1: Acquire training data and preprocess the data to obtain data set data; extract image and text modality features from the data set data to obtain image modality and text modality of the data set data; S2: extracting difference features between the image modality and the text modality to obtain difference information; S3: Generate a prototype feature according to the image modality, text modality and difference information, and concatenate the image modality, difference information and text modality to obtain a sample feature; S4: Calculate a similarity score based on the prototype feature and the sample feature, and adjust the prototype feature according to the actual result of the training data to obtain a judgment prototype feature; S5: extract the image modality of the data to be detected, the text modality of the data to be detected and the difference information of the data to be detected, calculate the similarity score of the data to be detected according to the image modality of the data to be detected, the text modality of the data to be detected, the difference information of the data to be detected and the judgment prototype feature, and judge whether the information is correct according to the similarity score of the data to be detected.

2. The method for detecting false information based on cross-modal difference classification according to claim 1, characterized in that: The text data of the data set data is encoded using a deep learning model encoder to obtain a text modality; the image data of the data set data is encoded using a convolutional neural network encoder to obtain an image modality: in, is the image mode; For text mode; Encode a convolutional neural network encoder for processing image modality; Encoding a deep learning encoder for processing text modalities; is the image data; For text data.

3. The error information detection method based on cross-modal difference classification according to claim 2 is characterized in that: The deep learning model encoder is a BERT encoder, and the convolutional neural network encoder is a ResNet34 encoder.

4. The error information detection method based on cross-modal difference classification according to claim 2 is characterized in that: Step S2 includes: S21: Input the image mode and the text mode into Combined encoder, we get Image modality in embedding space and Text modal with embedded space : in, for Joint encoder processing function; S22: Input the image mode and the text mode into Combined encoder, we get Image modality in embedding space and Text modal with embedded space : in, for Joint encoder processing function; S23: Calculation The difference in modal information in the embedding space and Modal difference information in embedding space: in, for The modal difference information in the embedding space, for Modal difference information in embedding space; S24: Calculate modal difference information : ; S25: Use a multi-layer perceptron to align the modal difference information to obtain the difference information : 。 5. The method for detecting false information based on cross-modal difference classification according to claim 4, characterized in that: Said The joint encoder includes u linear layers and v activation layers. The activation layer of the joint encoder uses the ReLU function; Said The joint encoder contains m linear layers and n activation layers. The activation layer of the joint encoder uses the ReLU function.

6. The method for detecting false information based on cross-modal difference classification according to claim 4, characterized in that: Step S23 includes: like When Normalize using the infinity norm: in, is the comparison threshold, is the infinite norm; like When Normalize using the infinity norm: ; like When Normalize using the infinity norm: ; like When Normalize using the infinity norm: 。 7. The method for detecting false information based on cross-modal difference classification according to claim 6, characterized in that: Step S3 includes: S31: Generate initial prototype features using Gaussian initialization according to the image modality, text modality and vector length of difference information , in, is the initial image prototype feature, is the initial difference prototype feature, It is the initial text prototype feature; S32: Perform Schmidt orthogonalization on the initial prototype features so that the initial image prototype features, the initial difference prototype features and the initial text prototype features are mutually orthogonal to obtain the prototype features : in, is the image prototype feature, is the difference prototype feature, It is the prototype feature of the text; S33: combining the image modality, the text modality and the difference information to obtain the sample feature : 。 8. The method for detecting false information based on cross-modal difference classification according to claim 7, characterized in that: Step S4 includes: S41: Calculating the similarity score : S42: Using an optimizer to adjust the prototype features according to the actual results of the training data to obtain judgment prototype features.

9. The method for detecting false information based on cross-modal difference classification according to claim 8, characterized in that: Step S41 includes: When the training data lacks image modality, the similarity score calculation formula is: When the training data lacks text modality, the similarity score calculation formula is: 。 10. An error information detection system based on cross-modal difference classification, used to execute the error information detection method based on cross-modal difference classification as claimed in any one of claims 1 to 9, characterized in that: include: Preprocessing module: obtain training data and preprocess the data to obtain data set data; Performing image-text modality feature extraction on the data set data to obtain the image modality and text modality of the data set data; performing inter-modality difference feature extraction on the image modality and the text modality to obtain difference information; generating prototype features according to the image modality, text modality and difference information, and splicing the image modality, difference information and text modality to obtain sample features; A judgment prototype feature training module: a similarity score is calculated based on the prototype feature and the sample feature, and the prototype feature is adjusted based on the actual result of the training data to obtain a judgment prototype feature; Prediction module: extract the image modality of the data to be detected, the text modality of the data to be detected and the difference information of the data to be detected, calculate the similarity score of the data to be detected based on the image modality of the data to be detected, the text modality of the data to be detected, the difference information of the data to be detected and the judgment prototype features, and judge whether the information is correct based on the similarity score of the data to be detected.

Citation Information

Patent Citations

  • Visual question and answer method and device, electronic equipment and storage medium

    CN115861995A

  • CLIP guidance-based multi-scale multi-mode false information detection method and device, electronic equipment and storage medium

    CN117216709A

  • Method, device and equipment for detecting news containing misleading information

    CN118152594A

  • Novel irony detection method and system based on comparative learning and modal mutual assistance

    CN118643369A

  • Method and apparatus for collecting, detecting and visualizing fake news

    US20210089579A1