Multi-modal rumor identification method and system, electronic equipment and storage medium
By employing a multimodal rumor identification method, text and image features are processed separately. The probability of a rumor is determined using a language model and a multimodal large language model. Combined with image-text consistency judgment, this method solves the problem of insufficient accuracy in multimodal rumor identification and improves the identification capability.
Patent Information
- Application Number
- CN202511093456.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-05
- Publication Date
- 2025-11-21
AI Technical Summary
In multimodal rumor identification, existing technologies, particularly traditional spatial domain feature fusion, have limited ability to identify rumors involving misaligned images and text or discrepancies between images and text, and the accuracy of identification needs to be improved.
A multimodal rumor identification method is adopted, which extracts text and image features respectively, uses language models and multimodal large language models to determine the probability of rumors, and uses the image-text consistency discrimination module to detect the dissimilarity probability, and finally comprehensively determines the rumor judgment result.
It improves the accuracy of identifying multimodal rumors, especially the ability to identify fake news using old images in new ways and rumors with contradictory text and images, and enhances the comprehensiveness and reliability of the judgment results.
Smart Images

Figure CN120994825A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of data processing, and particularly relates to a multi-modal rumor identification method and system, an electronic device and a storage medium. BACKGROUND
[0002] With the iterative upgrade of Internet technology and the full popularity of mobile terminals, information dissemination has broken through the time and space limitations of traditional media, forming a stereoscopic dissemination network all day long and across regions. According to statistics, global social media users generate more than 500 million multi-modal contents per day, among which the combination of video, picture and text accounts for 78%, such a composite information carrier not only meets the diversified expression needs of users, but also provides a technical breeding ground for the covert dissemination of rumors, resulting in multi-modal rumors.
[0003] For multi-modal rumors, although progress has been made in research in multiple directions, the traditional spatial domain feature fusion has limited ability to identify rumors with misaligned text and images or inconsistent text and images, and the recognition accuracy of multi-modal rumors still needs to be improved. SUMMARY
[0004] Therefore, the present disclosure proposes a multi-modal rumor identification scheme.
[0005] According to an aspect of the present disclosure, a multi-modal rumor identification method is provided, comprising: extracting text and images in a to-be-identified message; determining a first probability that the to-be-identified message is a rumor for the text; determining a second probability that the to-be-identified message is a rumor for the image; detecting consistency of the text and the image to obtain a dissimilarity probability; and obtaining a rumor determination result representing whether the to-be-identified message is a rumor based on the first probability, the second probability and the dissimilarity probability.
[0006] In a possible implementation, the rumor identification system includes a language model, and the determination of the first probability that the to-be-identified message is a rumor for the text includes: extracting key factual information in the text; determining a first message related to the key factual information based on the key factual information; determining a related message related to the to-be-identified message based on a similarity between the to-be-identified message and the first message, a credibility of the first message and a timeliness indicator of the first message; and determining the first probability and a reason based on the to-be-identified message and the related message using the language model.
[0007] In a possible implementation, the rumor identification system comprises a multi-modal large language model, and the determining, for the image, a second probability that the message to be identified is a rumor comprises: identifying, based on the multi-modal large language model, a fake category to which the image belongs, and extracting image features; and obtaining the second probability and reasons according to the fake category and the image features.
[0008] In a possible implementation, the identifying, based on the multi-modal large language model, the fake category to which the image belongs and the extracting of the image features comprise: dividing the image to obtain a plurality of sub-images; identifying, for each of the sub-images, a first fake category of the sub-image and obtaining first image features corresponding to the sub-image; and fusing each of the first fake categories according to positions of the sub-images in the image to obtain the fake category, and fusing each of the first image features according to the positions of the sub-images in the image to obtain the image features.
[0009] In a possible implementation, the rumor identification system comprises a text-image consistency discrimination module, and the detecting the consistency of the text and the image to obtain a dissimilarity probability comprises: inputting the text and the image into the text-image consistency discrimination module to obtain the dissimilarity probability; and the text-image consistency discrimination module is trained by using a cross-loss entropy function and a contrast loss function as a loss function.
[0010] In a possible implementation, the obtaining, based on the first probability, the second probability and the dissimilarity probability, a rumor determination result representing whether the message to be identified is a rumor comprises: fusing the first probability, the second probability and the dissimilarity probability to obtain a rumor probability; in a case where the rumor probability is not greater than a first threshold value and each of the first probability, the second probability and the dissimilarity probability is not greater than a second threshold value, the determination result represents that the message to be identified is not a rumor; and in a case where the rumor probability is greater than the first threshold value or any of the first probability, the second probability and the dissimilarity probability is greater than the second threshold value, the determination result represents that the message to be identified is a rumor.
[0011] In a possible implementation, the method further comprises: in a case where the determination result represents that the message to be identified is a rumor, modifying the text and / or the image to obtain a plurality of message samples; constructing text-image mismatch samples based on the message samples; and updating the rumor identification system based on the text-image mismatch samples and the message samples.
[0012] According to another aspect of the present disclosure, a multi-modal rumor identification device is provided, which is applied to a rumor identification system, and the device comprises:
[0013] extracting a text and an image in a message to be identified;
[0014] a first probability determining module configured to determine a first probability that the message to be identified is a rumor for the text;
[0015] a second probability determining module configured to determine a second probability that the message to be identified is a rumor for the image;
[0016] a dissimilarity probability determining module configured to detect consistency of the text and the image to obtain a dissimilarity probability;
[0017] a rumor judgment result determining module configured to obtain a rumor judgment result representing whether the message to be identified is a rumor based on the first probability, the second probability and the dissimilarity probability.
[0018] In a possible implementation, the rumor identification system includes a language model, and the first probability determining module is further configured to:
[0019] extract key factual information in the text;
[0020] determine a first message related to the key factual information based on the key factual information;
[0021] determine a related message related to the message to be identified based on similarity between the message to be identified and the first message, credibility of the first message and timeliness index of the first message;
[0022] determine the first probability and a reason based on the message to be identified and the related message using the language model.
[0023] In a possible implementation, the rumor identification system includes a multi-modal large language model, and the second probability determining module is further configured to:
[0024] identify a fake category to which the image belongs and extract image features based on the multi-modal large language model;
[0025] obtain the second probability and a reason according to the fake category and the image features.
[0026] In a possible implementation, the second probability determining module is further configured to:
[0027] divide the image to obtain a plurality of sub-images;
[0028] identify a first fake category of each of the sub-images and obtain a first image feature corresponding to each of the sub-images;
[0029] The first fake categories are fused according to positions of the subgraphs in the image to obtain the fake category;
[0030] The first image features are fused according to positions of the subgraphs in the image to obtain the image feature.
[0031] In a possible implementation, the rumor identification system comprises a text-image consistency discrimination module and the dissimilarity probability determination module is further configured to:
[0032] input the text and the image into the text-image consistency discrimination module to obtain the dissimilarity probability; the text-image consistency discrimination module is trained by using a cross-entropy loss function and a contrastive loss function as a loss function.
[0033] In a possible implementation, the rumor determination result determination module is further configured to:
[0034] fuse the first probability, the second probability and the dissimilarity probability to obtain a rumor probability;
[0035] in a case where the rumor probability is not greater than a first threshold value and the first probability, the second probability and the dissimilarity probability are all not greater than a second threshold value, the determination result represents that the to-be-identified message is a non-rumor;
[0036] in a case where the rumor probability is greater than the first threshold value or any of the first probability, the second probability and the dissimilarity probability is greater than the second threshold value, the determination result represents that the to-be-identified message is a rumor.
[0037] In a possible implementation, the apparatus is further configured to:
[0038] in a case where the determination result represents that the to-be-identified message is a rumor, modify the text and / or the image to obtain a plurality of message samples;
[0039] construct text-image mismatch samples based on the message samples;
[0040] update the rumor identification system based on the text-image mismatch samples and the message samples.
[0041] According to another aspect of the present disclosure, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement the steps of the above method.
[0042] According to another aspect of the present disclosure, a non-volatile computer readable storage medium is provided, which stores a computer program, and the computer program is executed by a processor to implement the steps of the above method.
[0043] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, or a non-transitory computer-readable storage medium carrying the computer program, which, when executed by a processor, implements the steps of the above method.
[0044] In the embodiments of the present disclosure, whether the to-be-identified message is a rumor is evaluated from three aspects of consistency of text, image, and text and image respectively. The first probability and the second probability are obtained for the text and the image respectively, without fusing the text and the image, and the two do not interfere with each other, thereby improving the objectivity and accuracy of the first probability and the second probability. Although the first probability and the second probability are determined by separately processing the text and the image, the present method does not ignore the connection between the two, but detects the consistency of the text and the image. The comprehensiveness of the process of determining the rumor determination result is improved. Finally, the determination result is determined based on the first probability, the second probability, and the dissimilarity probability, thereby improving the identification capability for the old image new use false news and the image-text contradictory rumor, and further improving the accuracy of the multi-modal rumor identification.
[0045] Other features and aspects of the present disclosure will become apparent from the following detailed description of exemplary embodiments, taken in conjunction with the accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS
[0046] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate exemplary embodiments, features, and aspects of the present disclosure and serve to explain the principles of the present disclosure.
[0047] Figure 1 A flowchart of a multi-modal rumor identification method provided by an embodiment of the present disclosure.
[0048] Figure 2 A structure diagram of a multi-modal rumor identification device provided by an embodiment of the present disclosure.
[0049] Figure 3 A structure diagram of an electronic device for multi-modal rumor identification provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0050] Various exemplary embodiments, features, and aspects of the present disclosure will be described in detail below with reference to the accompanying drawings. The same reference numbers in the drawings represent the same or similar elements having the same or similar functions. Although various aspects of the embodiments are illustrated in the drawings, the drawings are not necessarily drawn to scale unless specifically indicated.
[0051] As used herein, the terms "comprise", "comprising", "have", "having", "include", "including", "contain", "containing", or variants thereof, are open-ended, and include one or more stated features, integers, elements, steps, components, or functions but do not preclude the presence or addition of one or more other features, integers, elements, steps, components, functions, or groups thereof.
[0052] When an element is referred to as “connected,” “coupled,” “responding,” or a variation thereof relative to another element, it may be directly connected, coupled, or responding to another element, or there may be an intermediate element present.
[0053] Although the terms first, second, third, etc., may be used herein to describe various elements / operations, these elements / operations should not be limited by these terms. These terms are only used to distinguish one element / operation from another. Therefore, without departing from the teachings of the inventive concept, a first element / operation in some embodiments may be referred to as a second element / operation in other embodiments.
[0054] The term “exemplary” as used herein means “serving as an example, embodiment, or illustration.” Any embodiment illustrated herein as “exemplary” is not necessarily to be construed as superior to or better than other embodiments.
[0055] Furthermore, to better illustrate this disclosure, numerous specific details are set forth in the following detailed description. Those skilled in the art will understand that this disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of this disclosure.
[0056] Figure 1 A flowchart illustrating the multimodal rumor identification method provided in this disclosure. The method includes:
[0057] S11, extract the text and images from the message to be recognized.
[0058] In this embodiment, text can be extracted from a message to be recognized. When there are multiple text segments in the message, these segments can be arranged in the order they appear in the message to generate new text, reducing the probability of misalignment and preserving the semantic meaning of the text. Additionally, images can be extracted from the message. Images can include both static and dynamic images.
[0059] In this embodiment, the image can be preprocessed by scaling it to a preset resolution to eliminate the impact of size differences on the accuracy of subsequent image processing. Furthermore, the image is mapped from the RGB color space to the YCbCr color space to improve the sensitivity of subsequent image processing.
[0060] S12, for the text, determine the first probability that the message to be identified is a rumor.
[0061] The first probability characterizes the likelihood that the text in the message to be identified is a rumor. The value of the first probability is positively correlated with this likelihood.
[0062] In the embodiments of the present disclosure, the text can be analyzed using natural language processing technology. The overall semantics of the text can be analyzed, and the local semantics can be analyzed according to the keywords extracted from the text. In addition, the sentiment tendency of the text can be analyzed. The first probability can be determined based on the overall semantics, the local semantics, and the sentiment tendency. For example, the first probability can be determined using a language model. For another example, a known rumor database can be obtained, and the first matching degree of the overall semantics, the local semantics, and the sentiment tendency with the data in the known rumor database can be determined. The first matching degree can be taken as the first probability. The above is only an example, and the embodiments of the present disclosure do not limit the method for determining the first probability.
[0063] S13, determining, for the image, a second probability that the message to be identified is a rumor.
[0064] The second probability can represent the possibility that the image in the message to be identified is a rumor. The numerical value of the second probability is positively correlated with the possibility.
[0065] In the embodiments of the present disclosure, the image can be analyzed using computer vision technology, the semantics of the image can be determined, and the third probability that the image is a fake image and the fake type can be identified. The second probability can be determined based on the semantics of the image, the third probability, and the fake type. For example, the second probability can be determined using a classification model. For another example, the second matching degree of the semantics of the image, the third probability, and the fake type with the data in the known rumor database can be determined. The second matching degree can be taken as the second probability. The above is only an example, and the embodiments of the present disclosure do not limit the method for determining the first probability.
[0066] S14, detecting the consistency of the text and the image to obtain a dissimilarity probability.
[0067] In the embodiments of the present disclosure, whether the semantics of the text is consistent with the semantics of the image can be detected. The similarity of the characters, places, times, events, etc. described by the text and the image can be detected to obtain the dissimilarity probability. For example, the dissimilarity probability can be determined based on a contrast learning technology or a twin network.
[0068] S15, obtaining a rumor determination result representing whether the message to be identified is a rumor based on the first probability, the second probability, and the dissimilarity probability.
[0069] The rumor determination result can be a binary classification determination result. The two branches of the binary classification determination result can be that the message to be identified is a rumor or that the message to be identified is not a rumor.
[0070] In the embodiments of the present disclosure, the first probability, the second probability and the dissimilarity probability can be weighted and summed to obtain a weighted result. The determination result is determined according to the weighted result. The numerical intervals corresponding to the above two branches can be set in advance, and the rumor determination result is determined according to the numerical interval in which the weighted result falls.
[0071] In the embodiments of the present disclosure, whether the to-be-identified message is a rumor is evaluated from three aspects of consistency of text, image and both text and image. The first probability and the second probability are obtained for the text and the image respectively, without fusing the text and the image, and the two do not interfere with each other, thereby improving the objectivity and accuracy of the first probability and the second probability. Although the first probability and the second probability are determined by separately processing the text and the image, the present method does not ignore the relationship between the two, but detects the consistency of the text and the image. The comprehensiveness of the process of determining the rumor determination result is improved. Finally, the determination result is determined based on the first probability, the second probability and the dissimilarity probability, which improves the identification ability of the old picture new use false news and the image-text contradiction rumor, and further improves the accuracy of the multi-modal rumor identification.
[0072] In a possible implementation, the rumor identification system includes a language model, and the determining, for the text, the first probability that the to-be-identified message is a rumor includes: extracting key factual information in the text; determining a first message related to the key factual information based on the key factual information; determining a related message related to the to-be-identified message based on a similarity between the to-be-identified message and the first message, a credibility of the first message and a timeliness index of the first message; and determining the first probability and a reason based on the to-be-identified message and the related message using the language model.
[0073] The key factual information can be a keyword representing the content of the to-be-identified message. The key factual information can include words representing time, place, event, and person in the to-be-identified message.
[0074] Exemplarily, a Bidirectional Encoder Representations from Transformers-Conditional Random Field (BERT-CRF) model can be used to represent the message to be identified, and a Boundary Annotation System (BIO) can be used to annotate the message to be identified to identify the key factual information. The BERT-CRF model includes a BERT layer and a CRF layer. The BERT layer uses a Bert-Base-Chinese model as an encoder, and the CRF layer uses a transition matrix to model the dependency relationship as a decoder. The CRF layer maps each word, the position of each word in the sentence, and the paragraph to which each word belongs in the message to be identified to a 768-dimensional vector. In this way, each word corresponds to a 768-dimensional vector. The CRF layer decodes the vector and identifies the key factual information. The loss function of the BERT-CRF model is defined as the negative log-likelihood of the true path score and the sum of all path scores. The accuracy and comprehensiveness of extracting key factual information in the scenario of the present disclosure are improved.
[0075] The first message can be a message on the Internet other than the message to be identified. The first message can include a message published by a government website, a message published by an authoritative news website, and a message published by another website. There is at least one first vocabulary in the first message, and the first similarity between the at least one first vocabulary and the at least one key factual information in the message to be identified is greater than a preset first similarity threshold. The first vocabulary can include a vocabulary representing time, place, event, and person in the first message.
[0076] Exemplarily, the maximum first similarity in the first similarity corresponding to the first message can be used as the similarity between the message to be identified and the first message.
[0077] Exemplarily, an Okapi BM25 algorithm can be used to determine the similarity between the message to be identified and the first message by comprehensively considering the word frequency, word importance, and word-document deviation.
[0078] The credibility represents the authority of the channel of the first message. For example, the credibility can be determined according to the ranking of the domain name of the first message in the PageRank and manual verification. The timeliness indicator can represent the effect of the first message on determining the first probability in the time dimension. For example, the time interval between the time of publishing the first message and the time of determining the first probability can be used as the timeliness indicator.
[0079] The relevant message can be a message published on the Internet, the same or similar in content to the to-be-identified message, and reliable and effective. The relevance of the to-be-identified message and the first message can be determined based on the similarity of the to-be-identified message and the first message, the credibility of the first message, and the timeliness index of the first message. The first message with a relevance greater than a first relevance threshold is taken as the relevant message.
[0080] To facilitate understanding, the process of determining the relevance is represented by formula (1).
[0081] Relevance(W,K)=T+α+β×e -λΔt (1)
[0082] Wherein, W represents the to-be-identified message; K represents the first message; Relevance(W,K) represents the relevance; T represents the similarity of the to-be-identified message and the first message; a represents the credibility, for example, the a corresponding to the first message with a webpage level of 6 and above can be 0.3, the a corresponding to the first message with a webpage level below 6 can be 0.2, and the a corresponding to the first message on social media is reduced by 50%; b represents the timeliness index; e represents the base of the natural logarithm; l represents a constant; and At represents the time interval between the publishing time of the first message and the time of determining the first probability, and b is forced to be zero when At is greater than 365.
[0083] In the embodiments of the present disclosure, a language model can be used to construct a first prompt word based on the to-be-identified message and the relevant message; and the language model can determine the first probability and the reason based on the first prompt word.
[0084] In the embodiments of the present disclosure, first, the first messages are screened based on the key factual information, so that the first messages are all related to the to-be-identified message at least in part. Then, more rigorous screening is performed, and the relevant messages are selected from the first messages according to the similarity, the credibility, and the timeliness, thereby improving the effectiveness of the relevant messages in determining the first probability. Further, the accuracy of the first probability is improved. Moreover, the language model is used to obtain the first probability and the reason, thereby improving the explainability of the first probability.
[0085] In a possible implementation, the rumor identification system includes a multi-modal large language model, and the determining, for the image, the second probability that the to-be-identified message is a rumor includes: identifying a fake category to which the image belongs and extracting image features based on the multi-modal large language model; and obtaining the second probability and a reason according to the fake category and the image features.
[0086] In the embodiments of the present disclosure, the fake category of the image can include: modification of an object in the image, replacement of a face of a person in the image, and artificial intelligence generated content. A multi-modal large language model can be developed using the FakeShield framework. The multi-modal large language model is used to process the image to obtain a fake category to which the image belongs and an image feature corresponding to the image.
[0087] A prompt word is constructed based on the fake category, and the image feature is used as input data. The multi-modal large language model is used to map the fake category and the image feature to the same latent space to obtain a first vector. The first vector is analyzed to obtain a second probability and a reason for the second probability.
[0088] The same image feature has different meanings in different fake categories. In the embodiments of the present disclosure, the association between the fake category and the image feature is fully utilized to determine the second probability. Moreover, the multi-modal large language model is used to obtain the second probability and the reason, thereby improving the explainability of the second probability.
[0089] In a possible implementation, the multi-modal large language model is used to identify the fake category to which the image belongs and extract an image feature, including: dividing the image to obtain a plurality of sub-images; identifying a first fake category of each sub-image and obtaining a first image feature corresponding to each sub-image; fusing the first fake categories according to the positions of the sub-images in the image to obtain the fake category; and fusing the first image features according to the positions of the sub-images in the image to obtain the image feature.
[0090] A sub-image can represent a part of an image. A single sub-image can correspond to its position in the image. The position can be represented as a row number and a column number of the sub-image in the image. In the embodiments of the present disclosure, the multi-modal large language model can include a residual network 50 (ResNet50) and a classifier.
[0091] For example, the classifier can be used to identify the fake category of each sub-image. For ease of description, the fake category of the sub-image is named as a first fake category. The first fake categories can be fused according to the positions of the sub-images in the image to obtain the fake category of the image. The fusion of the first fake categories can be directly splicing the first fake categories according to the positions of the sub-images in the image. The fusion of the first fake categories can be performing bitwise operations on vectors corresponding to the first fake categories of sub-images in the same row or column to obtain fake category vectors, and then splicing the fake category vectors in row or column order. The above is only an example, and the fusion manner of the first fake categories in the embodiments of the present disclosure is not limited.
[0092] For another example, the residual network 50 can extract image features of each subgraph. For ease of description, the image features of the subgraph are named as first image features. The first image features can be fused according to the positions of the subgraphs in the image to obtain image features of the image. The fusion of the first image features can be that the first image features are directly spliced according to the positions of the subgraphs in the image. The fusion of the first image features can be that the first image features of the subgraphs in the same row or the same column are operated bit by bit to obtain image feature vectors, and then the image feature vectors are spliced in the order of rows or columns. The above is only an example, and the fusion manner of the first image features is not limited in the embodiments of the present disclosure.
[0093] In the embodiments of the present disclosure, the image features can include the first image features of each subgraph, and the fake category can include the first fake category of each subgraph. In this way, the fake area can be accurately located. In the case that there are multiple fake categories on the same image, the second probability corresponds to a more fine and reliable reason. In addition, even if the fake area is very small, it will not be submerged, improving the accuracy of the second probability.
[0094] In a possible implementation, the rumor identification system includes a text-image consistency discrimination module. The detection of the consistency of the text and the image to obtain the dissimilarity probability includes: inputting the text and the image into the text-image consistency discrimination module to obtain the dissimilarity probability. The text-image consistency discrimination module is obtained by training using a cross-entropy loss function and a contrastive loss function as a loss function.
[0095] The input data of the text-image consistency discrimination module is the text and the image in the same to-be-identified message. The output data is the dissimilarity probability of the text and the image in the same to-be-identified message.
[0096] In the embodiments of the present disclosure, the cross-entropy loss function and the contrastive loss function are used as a loss function. The loss function is used in the process of training the text-image consistency discrimination module. The cross-entropy loss function can refine the granularity of cross-modal retrieval. The contrastive loss function can capture the association between the image and the text threshold. Considering that the news semantics are relatively abstract, this abstraction increases the difficulty of determining the association between the image and the text. In the embodiments of the present disclosure, the cross-entropy loss function and the contrastive loss function are combined, which can also improve the accuracy of the dissimilarity probability in the case of processing the to-be-identified message with abstract semantics.
[0097] In a possible implementation, the rumor determination result of whether the to-be-identified message is a rumor is obtained based on the first probability, the second probability and the dissimilarity probability, including: fusing the first probability, the second probability and the dissimilarity probability to obtain a rumor probability; in a case where the rumor probability is not greater than a first threshold value, and the first probability, the second probability and the dissimilarity probability are all not greater than a second threshold value, the determination result indicates that the to-be-identified message is a non-rumor; in a case where the rumor probability is greater than the first threshold value, or any of the first probability, the second probability and the dissimilarity probability is greater than the second threshold value, the determination result indicates that the to-be-identified message is a rumor.
[0098] In the embodiment of the present disclosure, the first probability, the second probability and the dissimilarity probability can be weighted and summed to obtain the rumor probability.
[0099] In a case where the rumor probability is greater than the first threshold value, the to-be-identified message is a rumor; in a case where any of the first probability, the second probability and the dissimilarity probability is greater than the second threshold value, the to-be-identified message is a rumor.
[0100] In a case where the rumor probability is not greater than the first threshold value, and the first probability, the second probability and the dissimilarity probability are all not greater than the second threshold value, the to-be-identified message is a non-rumor.
[0101] In determining the rumor determination result, not only the consistency of the image, the text and the image-text is considered, that is, the rumor probability is used, but also the possibility that the image and the text cause the to-be-identified message to be a rumor is considered, and the possibility that the inconsistency between the image and the text causes the to-be-identified message to be a rumor is considered. Therefore, the rumor determination result determined by using the method of the embodiment of the present disclosure is more reasonable, accurate and reliable. Also, in a case where the to-be-identified message only contains one-sided false content (image false, or text false, or image-text inconsistency), the rumor can also be determined without missing.
[0102] In a possible implementation, the method further includes: in a case where the determination result indicates that the to-be-identified message is a rumor, modifying the text and / or the image to obtain a plurality of message samples; constructing an image-text inconsistency sample based on each of the message samples; and updating the rumor identification system based on the image-text inconsistency sample and the message sample.
[0103] The modification of the text can include synonym replacement of words in the text, transformation of word or sentence order in the text, fuzzy processing of semantics of the text, and replacement of keywords in the text. The modification of the image can include adding Gaussian noise to the image, and fine-tuning of elements or local parts in the image. In addition, the modified image and text can be mismatched.
[0104] In the embodiments of the present disclosure, the images and texts of the to-be-identified messages determined as rumors can be modified to obtain message samples. The rumor identification system is subjected to adversarial training using the message samples. The number of message samples subjected to text modification in the message samples can be controlled to be about 20%, the number of message samples subjected to image modification in the message samples can be controlled to be about 15%, and the number of message samples with incorrect matching of texts and images can be controlled to be about 10%.
[0105] In this way, in the face of rumor governance needs in a sudden major public event, the rumor identification system can be fine-tuned even in the case of insufficient available training samples, so that the rumor identification system can quickly and accurately complete the identification task of rumors related to the sudden major public event. Efficient and accurate technical support is provided for government regulation, media review, and platform content governance.
[0106] The rumor identification method is described below in an embodiment.
[0107] The to-be-identified message is a first news. The text W and the image P are extracted from the first news. The text W is that a serious explosion event occurred in a XXX university in Beijing on April 12, 2024, and the fire caused multiple casualties. For ease of understanding, the image P is described in words in the embodiments of the present disclosure. The image P is an image that looks like a scene photo, showing obvious fire and casualty situations in the picture.
[0108] The BERT-CRF model is used to extract the key fact information from the text W: April 12, 2024 (time), XXX university in Beijing (location), explosion (event). Based on the key fact information, the first message is searched on the Internet. Based on the similarity of the first message and the first news, the credibility of the first message, and the timeliness index of the first message, five related messages are determined. The five related messages are as follows:
[0109] Related message 1: 36 cases of domestic laboratory safety accidents in recent years.
[0110] Related message 2: On December 10, 2024, a high-falling accident occurred in a university in Beijing.
[0111] Related message 3: On January 9, 2024, an explosion occurred in an experimental building of a certain university in Henan.
[0112] Related message 4: On December 26, 2018, an explosion occurred in a laboratory of Xiamen University.
[0113] Related message 5: In 2015, an explosion occurred in a chemical laboratory of a certain university in Beijing.
[0114] The language model is used to construct a first prompt word based on the to-be-identified message and the related message. The language model can determine a first probability and a reason for obtaining the first probability based on the first prompt word. The first probability is 0.89.
[0115] The image P is scaled to a preset resolution. The image P is mapped from an RGB color space to a YCbCr color space. An image feature of the image P is extracted using the residual network 50. A fake category of the image P is determined using a classifier.
[0116] A multimodal large language model is used to determine a second probability and a reason based on the fake category and the image feature. The second probability is 0.13. It can be basically judged that the image P is not fake.
[0117] The image-text consistency discrimination module calls the aforementioned BERT-CRF model to encode the text W to obtain a text vector. The image-text consistency discrimination module also calls the aforementioned residual network 50 to encode the image P to obtain an image vector. The text vector and the image vector are mapped to the same embedding space, and a dissimilarity probability of the text vector and the image vector is determined to be 0.55.
[0118] Based on the first probability (0.85), the second probability (0.13), and the dissimilarity probability (0.55), a rumor probability is determined to be 0.546. The first threshold value and the second threshold value are both 0.55. Although the rumor probability is lower than the first threshold value, the first probability is higher than the second threshold value. Therefore, it can be determined that the first news is a rumor.
[0119] Next, the text W and the image P can be modified to generate a message sample, and the modified text and the modified image can be mismatched to generate a message sample. The rumor identification system is fine-tuned using the message sample. The rumor identification model can quickly respond to the identification task of identifying rumors similar to the first news.
[0120] Figure 2 A structural schematic diagram of a multimodal rumor identification device provided by the embodiments of the present disclosure is provided. The device is applied to a rumor identification system. The device 20 includes:
[0121] An extraction module 21 is configured to extract a text and an image in a to-be-identified message.
[0122] A first probability determination module 22 is configured to determine a first probability that the to-be-identified message is a rumor based on the text.
[0123] A second probability determination module 23 is configured to determine a second probability that the to-be-identified message is a rumor based on the image.
[0124] A dissimilarity probability determination module 24 is configured to detect consistency of the text and the image to obtain a dissimilarity probability.
[0125] The rumor determination result determination module 25 is configured to determine a rumor determination result of the to-be-identified message based on the first probability, the second probability, and the dissimilarity probability.
[0126] In a possible implementation, the rumor identification system comprises a language model, and the first probability determination module 22 is further configured to:
[0127] extract key factual information in the text;
[0128] determine a first message related to the key factual information based on the key factual information;
[0129] determine a related message related to the to-be-identified message based on a similarity between the to-be-identified message and the first message, a credibility of the first message, and a timeliness index of the first message;
[0130] determine the first probability and a reason based on the to-be-identified message and the related message by using the language model.
[0131] In a possible implementation, the rumor identification system comprises a multi-modal large language model, and the second probability determination module 23 is further configured to:
[0132] identify a fake category to which the image belongs and extract image features based on the multi-modal large language model;
[0133] determine the second probability and a reason based on the fake category and the image features.
[0134] In a possible implementation, the second probability determination module 23 is further configured to:
[0135] divide the image into a plurality of sub-images;
[0136] identify a first fake category of each of the sub-images and obtain a first image feature corresponding to each of the sub-images;
[0137] fuse the first fake categories according to positions of the sub-images in the image to obtain the fake category;
[0138] fuse the first image features according to the positions of the sub-images in the image to obtain the image features.
[0139] In a possible implementation, the rumor identification system comprises a text-image consistency discrimination module, and the dissimilarity probability determination module 24 is further configured to:
[0140] input the text and the image into the text-image consistency discrimination module to obtain the dissimilarity probability; the text-image consistency discrimination module is obtained by training using a cross loss entropy function and a contrast loss function as a loss function.
[0141] In a possible implementation, the rumor determination result determination module 25 is further configured to:
[0142] fuse the first probability, the second probability, and the dissimilarity probability to obtain a rumor probability;
[0143] in a case where the rumor probability is not greater than a first threshold value, and the first probability, the second probability, and the dissimilarity probability are all not greater than a second threshold value, the determination result represents that the to-be-identified message is a non-rumor;
[0144] in a case where the rumor probability is greater than the first threshold value, or any of the first probability, the second probability, and the dissimilarity probability is greater than the second threshold value, the determination result represents that the to-be-identified message is a rumor.
[0145] In a possible implementation, the apparatus 20 is further configured to:
[0146] in a case where the determination result represents that the to-be-identified message is a rumor, modify the text and / or the image to obtain a plurality of message samples;
[0147] construct text-image mismatch samples based on the message samples;
[0148] update the rumor identification system based on the text-image mismatch samples and the message samples.
[0149] In some embodiments, the apparatus provided by the embodiments of the present disclosure has functions or includes modules that can be used to perform the methods described in the above method embodiments, and the specific implementation can refer to the description of the above method embodiments. For brevity, details are not repeated here.
[0150] The embodiments of the present disclosure also provide an electronic device, including a memory, a processor, and a computer program stored in the memory, and the processor executes the computer program to implement the steps of the above method.
[0151] The embodiments of the present disclosure also provide a non-volatile computer readable storage medium having a computer program stored thereon, and the computer program is executed by a processor to implement the steps of the above method.
[0152] The embodiments of the present disclosure also provide a computer program product, including a computer program or a non-volatile computer readable storage medium carrying a computer program, and the computer program is executed by a processor to implement the steps of the above method.
[0153] Figure 3 A structural diagram of an electronic device for multi-modal rumor identification is provided for embodiments of the present disclosure. For example, the electronic device 1900 can be provided as a server or a terminal device. Referring to Figure 3 , the electronic device 1900 includes a processing component 1922, further including one or more processors, and a memory resource represented by a memory 1932, for storing instructions executable by the processing component 1922, such as an application program. The application program stored in the memory 1932 can include one or more than one module each corresponding to a set of instructions. In addition, the processing component 1922 is configured to execute the instructions to perform the above method.
[0154] The electronic device 1900 can also include a power supply component 1926 configured to perform power management of the electronic device 1900, a wired or wireless network interface 1950 configured to connect the electronic device 1900 to a network, and an input / output interface 1958 (I / O interface). The electronic device 1900 can operate based on an operating system stored in the memory 1932, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM or the like.
[0155] In an exemplary embodiment, a non-volatile computer readable storage medium, such as the memory 1932 including computer program instructions executable by the processing component 1922 of the electronic device 1900 to complete the above method is also provided.
[0156] Computer readable storage media can be any media that can be read by a machine. Such media can include, but is not limited to, optical discs, magnetic discs, magnetic tapes, electronic memories, and / or any combination thereof. Computer readable storage media can be non-transitory, in that it can be a tangible medium. In some embodiments, computer readable storage media can be non-transitory, in that it can not be a signal per se. In other embodiments, computer readable storage media can be a transitory medium, in that it can be a signal. In some embodiments, computer readable storage media can be non-transitory, in that it can not be a signal per se, but can be a tangible medium. In other embodiments, computer readable storage media can be a transitory medium, in that it can be a signal. In some embodiments, computer readable storage media can be non-transitory, in that it can not be a signal per se, but can be a tangible medium. In other embodiments, computer readable storage media can be a transitory medium, in that it can be a signal.
[0157] The computer programs (or computer readable program instructions) described herein can be downloaded from a computer readable storage medium to respective computing / processing devices or external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions to storage media in respective computing / processing devices for execution.
[0158] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate array (FPGA), or programmable logic array (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0159] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0160] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include random access memory (RAM), read only memory (ROM), electrically erasable programmable read only memory (EEPROM), flash memory or other data storage device. When the computer readable program instructions are loaded into the computer and other programmable data processing apparatus, a series of operational steps are implemented that provide processes such that the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0161] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0162] The flow diagrams and the block diagrams in the drawings are presented to illustrate the architecture, functionality, and operation of possible implementations of systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks can sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and
[0163] Embodiments of the present disclosure have been described above, and the description is intended to be illustrative of the embodiments and not restrictive. Many modifications and variations of the described embodiments are possible and are within the scope of the disclosure. The selection of terms is intended to best describe the principles of the embodiments, practical application, or technical improvements in the art, or to enable others skilled in the art to understand the embodiments disclosed herein.
Claims
1. A multimodal rumor identification method, characterized in that, The method is applied to a rumor identification system, and the method comprises: extracting text and images in a message to be identified; determining a first probability that the message to be identified is a rumor for the text; determining a second probability that the message to be identified is a rumor for the images; detecting consistency of the text and the images to obtain a dissimilarity probability; obtaining a rumor determination result representing whether the message to be identified is a rumor based on the first probability, the second probability and the dissimilarity probability.
2. The method of claim 1, wherein, The rumor identification system comprises a language model, and the determination of the first probability that the message to be identified is a rumor for the text comprises: extracting key factual information in the text; determining a first message related to the key factual information based on the key factual information; determining a related message related to the message to be identified based on a similarity between the message to be identified and the first message, a credibility of the first message and a timeliness index of the first message; determining the first probability and a reason based on the message to be identified and the related message using the language model.
3. The method of claim 1, wherein, The rumor identification system comprises a multi-modal large language model, and the determination of the second probability that the message to be identified is a rumor for the images comprises: identifying a fake category to which the images belong and extracting image features based on the multi-modal large language model; obtaining the second probability and a reason according to the fake category and the image features.
4. The method of claim 3, wherein, The identification of the fake category to which the images belong and the extraction of the image features based on the multi-modal large language model comprise: dividing the images to obtain a plurality of sub-images; identifying a first fake category of each of the sub-images and obtaining first image features corresponding to each of the sub-images; fusing each of the first fake categories according to positions of the sub-images in the images to obtain the fake category; fusing each of the first image features according to the positions of the sub-images in the images to obtain the image features.
5. The method of claim 1, wherein, The rumor identification system comprises a text-image consistency discrimination module, and the detection of the consistency of the text and the images to obtain the dissimilarity probability comprises: inputting the text and the images into the text-image consistency discrimination module to obtain the dissimilarity probability; the text-image consistency discrimination module is obtained by training using a cross-entropy loss function and a contrast loss function as a loss function.
6. The method of claim 1, wherein, The obtaining of the rumor determination result representing whether the message to be identified is a rumor based on the first probability, the second probability and the dissimilarity probability comprises: fusing the first probability, the second probability and the dissimilarity probability to obtain a rumor probability; in a case where the rumor probability is not greater than a first threshold value and the first probability, the second probability and the dissimilarity probability are all not greater than a second threshold value, the determination result represents that the message to be identified is not a rumor; in a case where the rumor probability is greater than the first threshold value or any of the first probability, the second probability and the dissimilarity probability is greater than the second threshold value, the determination result represents that the message to be identified is a rumor.
7. The method of claim 1, wherein, The method further comprises: In a case where the determination result represents that the message to be identified is a rumor, the text and / or the image are modified to obtain a plurality of message samples; Based on each of the message samples, a text-image mismatch sample is constructed; Based on the text-image mismatch sample and the message sample, the rumor identification system is updated. 8.A multi-modal rumor identification device, comprising: The device is applied to a rumor identification system, and the device comprises: a text and image extraction module configured to extract text and an image in a message to be identified; a first probability determination module configured to determine a first probability that the message to be identified is a rumor with respect to the text; a second probability determination module configured to determine a second probability that the message to be identified is a rumor with respect to the image; a dissimilarity probability determination module configured to detect consistency of the text and the image to obtain a dissimilarity probability; a rumor determination result determination module configured to obtain a rumor determination result representing whether the message to be identified is a rumor based on the first probability, the second probability, and the dissimilarity probability.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory, wherein the computer program comprises instructions that, when executed by the processor, cause the electronic device to perform the method of any one of claims 1-8. The processor executes the computer program to implement the steps of the method of any one of claims 1 to 7.
10. A non-transitory computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.