Multi-scale semantic alignment false news detection method and system
Through modal perception processing and multi-scale semantic aggregation mechanism, the problem of unmodeled and dynamic adjustment of semantic differences between modalities in existing technologies is solved, efficient detection of fake news is achieved, and the fusion capability and detection accuracy of cross-modal features are enhanced.
Patent Information
- Application Number
- CN202510587717.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-08
- Publication Date
- 2025-09-19
AI Technical Summary
Existing multimodal fake news detection methods fail to fully model the semantic differences and dynamic adjustments between modalities, ignore the importance of modality, and show limitations in processing multi-scale semantic information, and cannot effectively capture complex cross-modal relationships.
A modality-aware processing strategy is adopted, and a large-scale pre-trained model is used to extract the self-attention coarse-grained representation of each modality. Through a multi-scale semantic aggregation mechanism, residual connections and convolution kernels of different scales are introduced to gradually fuse the fine-grained representations between modalities to form a joint representation space. Finally, the classification layer is used to judge the authenticity of the news.
It realizes joint modeling and feature learning of heterogeneous data, enhances the ability to distinguish fake news, can dynamically allocate modal weights, comprehensively capture complex cross-modal relationships, and improves the accuracy and comprehensiveness of detection.
Smart Images

Figure CN120671064A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of false information detection, and in particular to a false news detection method and system based on multi-scale semantic alignment. Background Art
[0002] With the rapid development of the internet, online fake news has become a serious global problem. Fake news primarily refers to news reports that deliberately spread false information to mislead the public and achieve a specific purpose. Fake news is characterized by poor authenticity and high deceptiveness, not only misleading public judgment but also endangering social stability. Therefore, detecting fake news is of great significance.
[0003] At present, the mainstream method of multimodal fake news detection usually uses a multi-branch network to independently represent the data of different modalities in the news, and combines these modal features through simple splicing, weighting or direct projection in the feature fusion stage. Although this method is intuitive, it has several significant problems. First, the semantic differences between modalities are not fully modeled. Different modalities (such as text and images) carry different information dimensions and forms of expression. Simple fusion methods may lead to insufficient alignment of cross-modal features in a unified representation space. In addition, the feature fusion process often ignores the dynamic adjustment of the importance of modalities. In fact, in different news scenarios, different modalities do not contribute equally to the final semantic judgment. For example, images may play a key role in some news, while in other scenarios, text may be the dominant modality.
[0004] These methods also exhibit limitations when processing multi-scale semantic information. Multimodal fake news detection requires not only understanding the global semantics and local details within the modalities but also capturing the spatial and semantic dependencies between modalities. However, existing methods typically use fixed-scale feature processing during the fusion process, failing to effectively model multi-scale semantic information, especially when long-range dependencies or heterogeneous features exist between modalities. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a multi-scale semantic alignment fake news detection method and system that can solve the technical problems of the existing technology in detecting fake news, such as the inability to dynamically assign modal weights and comprehensively capture complex cross-modal relationships.
[0006] A multi-scale semantic alignment fake news detection method of the present invention comprises the following steps:
[0007] S10: Obtain multimodal news raw data including text, audio, and images from different information sources, integrate and store them as structured data using a predefined format;
[0008] S20: using a modality-aware processing strategy to pre-process text, audio, and image modality data respectively;
[0009] S30: Extract self-attention coarse-grained representations of each modality using a large-scale pre-trained model;
[0010] S40: Introducing a multi-scale semantic aggregation mechanism in the fusion process, using residual connections and convolution kernels of different scales to gradually fuse the fine-grained representations between modalities to form a joint representation space;
[0011] S50: Using a classification layer, the joint representation space is converted into a probability distribution of true or false categories, and whether the current news is false news is determined based on the probability distribution.
[0012] Furthermore, the step of preprocessing the text includes:
[0013] Clean the text to remove irrelevant characters, emoticons, URLs, punctuation, unify capitalization, and remove stop words;
[0014] Perform word segmentation and part-of-speech tagging on the cleaned text to identify nouns, verbs, and adjectives;
[0015] The Word2Vec word vector model is used to obtain the distributed semantic representation of the text and merge it into sequence features.
[0016] Furthermore, the step of preprocessing the image includes:
[0017] Normalize, resize, and Gaussian blur the image;
[0018] Vision-Transformer (ViT) is used to extract the visual features of the image and obtain image features in the form of multi-channel vectors.
[0019] Furthermore, the step S30 includes extracting semantic features of the image, and the step of extracting semantic features of the image includes:
[0020] Obtain text in the image as recognized text through optical character recognition (OCR);
[0021] The recognized text is processed using Transformer to extract image semantic features in vector form.
[0022] Furthermore, the use of a large-scale pre-trained model to extract self-attention coarse-grained representations of each modality includes:
[0023] Use the Transformer model to extract the sequence features of the text and obtain the text features in vector form;
[0024] Use Conv-TasNet to extract audio features and obtain audio features in vector form.
[0025] Furthermore, the fusion process of the multi-scale semantic aggregation mechanism includes:
[0026] (1) Add the image semantic features and text features element by element to obtain the first joint feature, and multiply the audio features and image features element by element to obtain the second result;
[0027] (2) Connecting the image feature, the audio feature and the second result to obtain a first joint feature;
[0028] (3) connecting the first joint feature and the second joint feature to obtain a third joint feature;
[0029] (4) Perform multi-scale convolution on the third joint feature to obtain a final multi-scale joint representation.
[0030] Furthermore, the processing steps of the classification layer include:
[0031] Performing a nonlinear transformation on the joint representation through a linear transformation layer, a batch normalization layer, and a ReLU activation function to obtain the original category score;
[0032] The raw category scores are processed using a fully connected layer with a softmax activation function to obtain the category probability distribution of real news and fake news.
[0033] A multi-scale semantic alignment fake news detection system, including:
[0034] a. Input layer, used to perform data acquisition and preprocessing:
[0035] Obtain multimodal news raw data from different information sources, including text, audio, and images;
[0036] Perform word segmentation, stop word removal, cleaning, and vectorization on text data;
[0037] Perform normalization, resampling, and framing on audio data;
[0038] Perform pixel normalization, size cropping, and Gaussian blur preprocessing on image data;
[0039] b. Feature representation layer, used to extract features of each modality:
[0040] Text feature extraction: Use the Transformer model to extract the sequence features of the text and combine it with Word2Vec to generate distributed semantic representations;
[0041] Image feature extraction: Vision-Transformer (ViT) is used to extract visual features of images, and Optical Character Recognition (OCR) combined with Transformer is used to extract semantic features of text in images.
[0042] Audio feature extraction: Use Conv-TasNet to extract the time domain features of audio;
[0043] c. Multi-scale semantic alignment layer for inter-modal feature fusion:
[0044] Using residual connections and convolution kernels of different scales, text features, image visual features, image semantic features, and audio features are fused stage by stage.
[0045] Generate a joint representation space containing multi-scale semantic information;
[0046] d. Output layer, used for classification decision:
[0047] The joint representation is nonlinearly transformed through a linear transformation layer, a batch normalization layer, and a ReLU activation function to obtain the original category score;
[0048] Use a fully connected layer with a softmax activation function to output the class probability distribution of true / false news;
[0049] The input layer, feature representation layer, multi-scale semantic alignment layer, and output layer are connected in sequence to form a multi-scale semantic alignment network, realizing end-to-end processing from raw data input to fake news discrimination.
[0050] Furthermore, the image feature extraction of the feature representation layer includes:
[0051] Recognize text content in the image through an optical character recognition (OCR) subunit;
[0052] The text recognized by OCR is processed by the Transformer semantic extraction sub-unit, and the image semantic features in vector form are output.
[0053] A computer-readable storage medium, characterized in that the computer-readable storage medium stores program instructions, and when the program instructions are executed, it is used to execute the multi-scale semantic alignment fake news detection method according to any one of claims 1 to 7
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] The method of the present invention realizes joint modeling and feature learning of heterogeneous data through a carefully designed modal representation layer and a multi-scale modal fusion layer. The text uses the Transformer model to extract semantic features, the image uses VIT to extract visual features, and the audio uses a joint representation of the time and frequency domain to obtain audio features. Furthermore, an inter-modal multi-scale semantic alignment network is designed to learn the intrinsic associations between different modalities and guide the interactive fusion of cross-modal features. Among them, the present invention uses large-scale pre-training methods such as BERT and VIT for modal feature extraction. Compared with traditional methods, it can obtain more abstract and semantic feature representations, thereby enhancing the ability to distinguish fake news. The use of a multi-scale semantic alignment network can learn the long-term and short-term distance dependencies between different modalities and guide the effective fusion of information. This association modeling method enables the fusion features to contain richer inter-modal context, which is not available in current methods. Therefore, the technical solution of the present invention solves the technical problem that the existing technology cannot dynamically allocate modal weights and fully capture complex cross-modal relationships when detecting fake news. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is a flow chart of the method of the present invention;
[0057] Figure 2 Schematic diagram of the multimodal feature fusion network structure of the present invention;
[0058] Figure 3 Schematic diagram of the VIT model structure details of the present invention;
[0059] Figure 4 Schematic diagram of the multi-scale semantic alignment network structure of the present invention; DETAILED DESCRIPTION
[0060] The following embodiments of the present invention are described in further detail with reference to the accompanying drawings and examples. The following examples are used to illustrate the present invention but are not intended to limit the scope of the present invention.
[0061] like Figures 1 to 4 As shown, the present invention provides a multi-scale semantic alignment fake news detection method and system,
[0062] Example 1:
[0063] 1. Overview of the overall technical solution
[0064] This embodiment maps the method flow to the system architecture to build an end-to-end fake news detection solution. The core includes:
[0065] Method steps: data acquisition → modal preprocessing → multimodal feature extraction → multi-scale semantic fusion → classification decision.
[0066] System architecture: input layer (supporting data preprocessing), feature representation layer (supporting feature extraction), multi-scale semantic alignment layer (supporting modal fusion), and output layer (supporting classification).
[0067] 2. Combining the method implementation steps with the system architecture
[0068] 1. Multimodal data acquisition and preprocessing (input layer function)
[0069] Method steps:
[0070] Obtain original news data containing text, audio, and images from different information sources, integrate them into structured data (such as a table containing news ID and data paths of each modality) in a predefined format, and store them in the system data layer.
[0071] System input layer:
[0072] Data access module: supports acquiring multimodal data from social media, news platforms, etc., and outputs raw data to the preprocessing unit;
[0073] Modality perception preprocessing unit:
[0074] Text: cleaning (removing irrelevant characters and stop words), word segmentation, vectorization (such as Word2Vec), and generating text sequence features;
[0075] Image: size cropping, pixel standardization, noise reduction, and output of an image matrix of uniform size;
[0076] Audio: Resampling, framing, and amplitude normalization to generate audio frame sequences that the model can process.
[0077] 2. Multimodal feature extraction (feature representation layer function)
[0078] Method steps:
[0079] Utilize large-scale pre-trained models to extract self-attention coarse-grained representations of each modality:
[0080] The text is processed through Transformer to extract sequence features;
[0081] The image is processed through Vision-Transformer (ViT) to extract visual features, and OCR+Transformer is used to extract semantic features of the text in the image.
[0082] The audio is processed through Conv-TasNet to extract time domain features.
[0083] System feature representation layer:
[0084] Text feature extraction unit: inputs the preprocessed text sequence and encodes it into vector text features through Transformer;
[0085] Image feature extraction unit:
[0086] Vision branch: ViT processes the image matrix and outputs image visual features in the form of multi-channel vectors;
[0087] Semantic branch: OCR recognizes image text and encodes it into image semantic features through Transformer;
[0088] Audio feature extraction unit: Conv-TasNet processes audio frame sequences and outputs audio features in vector form.
[0089] 3. Multi-scale semantic fusion (multi-scale semantic alignment layer function)
[0090] Method steps:
[0091] Residual connections and convolution kernels of different scales are introduced to gradually fuse the fine-grained representations between modalities to form a joint representation space:
[0092] The image semantic features and text features are added element by element, and the audio features and image visual features are multiplied element by element to generate preliminary interaction features;
[0093] It concatenates multimodal features and captures semantic dependencies of different granularities through multi-scale convolution kernels (such as 3, 5, and 7), and outputs multi-scale joint representations.
[0094] System multi-scale semantic alignment layer:
[0095] Residual Interaction Module: implements element-by-element addition / multiplication of cross-modal features, preserving original feature differences and enhancing interactions;
[0096] Multi-scale convolution module: Use convolution kernels of different scales in parallel to process the concatenated features, fuse multi-scale information through residual connections, and generate a joint representation containing cross-modal fine-grained semantics.
[0097] 4. Classification decision (output layer function)
[0098] Method steps:
[0099] The joint representation is converted into a probability distribution of true / false categories through the classification layer, and the authenticity of the news is judged based on the probability:
[0100] Perform nonlinear transformation on the joint representation (linear layer + batch normalization + ReLU);
[0101] The binary probability distribution is output through the fully connected layer with softmax, and a threshold (such as 0.5) is set to determine the result.
[0102] System output layer:
[0103] Nonlinear transformation unit: performs dimension transformation and activation processing on the joint representation to generate the original category score;
[0104] Classification unit: The fully connected layer combines with the softmax function to output the probability distribution of true news and false news, supporting real-time classification decisions.
[0105] Example 2:
[0106] This method comprises the following steps:
[0107] S10. Obtain multimodal news raw data from different information sources, including text, audio, and images, integrate them using a predefined format, and store them as structured data;
[0108] In this step, the system acquires multimodal raw data containing news content from multiple heterogeneous information sources, including but not limited to social media platforms, news portals, video sharing platforms, and blog forums. The acquired data covers multiple modalities, including text (such as headlines, body text, and comments), audio (such as voice broadcasts and interview recordings), and images (such as illustrations, screenshots, and news posters).
[0109] To ensure uniformity and effectiveness in subsequent processing, the system uses predefined data format specifications to organize and align the aforementioned multimodal data. Specifically, the system aggregates and matches different modal data from the same news event based on timestamps, unique identifiers (such as URLs, IDs), or contextual information. Subsequently, the system performs structured parsing of the meta-information (such as source, release time, author, etc.) and content information of each modality according to a predefined data template, converting it into a unified data representation to facilitate subsequent preprocessing and feature extraction operations.
[0110] S20, using a modality-aware processing strategy to pre-process text, audio, and image modality data;
[0111] In this step, the system preprocesses the multimodal news data acquired and structured in step S10. This step corresponds to the system's input layer, and its primary function is to standardize and cleanse the data content across different modalities to ensure data quality and consistency.
[0112] First, for text modal data, the system uses a natural language standardization module to standardize the format of the raw text, including punctuation, removal of invalid spaces, and encoding conversion. Once standardization is complete, a text cleaning process is performed, removing HTML tags, script code, special symbols, and meaningless terms to extract the semantically valid text. Once cleansed, the system incorporates an automated text correction mechanism to identify and correct content deviations such as spelling errors, internet slang abbreviations, and traditional Chinese character conversions, resulting in standardized, high-quality text data.
[0113] Secondly, for audio modal data, the system uses noise detection and signal enhancement technology to reduce noise and clean the original speech content. This includes processing steps such as silence segment identification, background noise suppression, and speech segment reconstruction to extract clear and recognizable speech segments, ensuring the robustness and accuracy of the speech feature extraction model.
[0114] Finally, the system performs image regularization and standardization on the image modal data. It first identifies and removes image samples that affect feature extraction due to extreme or unbalanced sizes (such as extreme stretching or compression). The remaining images are then uniformly resized to a preset size to accommodate the input requirements of image recognition models and OCR (optical character recognition) tools. Furthermore, image format conversion, edge cropping, and resolution adjustment can be performed to improve the expressive integrity of the image content and enhance model processing efficiency.
[0115] After completing the above preprocessing, the system uses the normalized text data, cleaned audio data, and standardized image data as multimodal inputs and sends them to the subsequent feature extraction module for semantic modeling and feature fusion.
[0116] S30, using a large-scale pre-trained model to extract self-attention coarse-grained representations of each modality;
[0117] In this step, the preprocessed data from each modality is input into the feature representation layer to extract a coarse-grained representation of each modality. The feature representation layer consists of four main parts: image feature extraction, text feature extraction, audio feature extraction, and image semantic feature extraction.
[0118] In this step, to effectively model news image data and obtain semantic feature representations within the images, the present invention uses the Vision Transformer (ViT), a visual model based on the Transformer architecture, as an image feature extraction module to perform deep representation learning on news images. The ViT model can capture long-range dependencies and spatial semantic structure within image content, improving the discriminative power of image modalities in fake news detection tasks.
[0119] Specifically, the system first divides the input image into fixed-size patches (patches), using convolution operations with a stride of 16 to achieve equidistant image segmentation. Each patch is then flattened and mapped to a low-dimensional embedding space through a linear transformation to form a patch embedding. This process, similar to word embedding in natural language processing, effectively compresses high-dimensional information in localized image regions while preserving their semantic characteristics.
[0120] After the embedding vectors are generated, the system further adds a learnable position encoding vector to each block embedding to preserve the spatial position of each image block within the original image. This position encoding mechanism enables the model to understand the relative layout relationships between different regions in the image, helping to build global semantic understanding of the image.
[0121] The embedded sequence is then fed into a multi-layer stacked Transformer encoder, which includes a multi-head self-attention mechanism and a feedforward neural network to capture the deep semantic connections between image patches. At the output stage, the system extracts a classification identity vector representing the entire image information and maps it to the final image feature representation through a fully connected head:
[0122] f img =MLP_Head(x class )
[0123] In the above formula, f img Represents the extracted global feature vector of the image, which is an important input of the multimodal fusion module. class It is the classification vector in Transformer, which aggregates the semantic information of the entire image.
[0124] After extracting the visual features of the image, the present invention also introduces an image semantic feature extraction mechanism to further explore the potential semantic information contained in the image. This mechanism fully utilizes the text and visual context embedded in the image, enhancing the expressive power of the image modality from multiple perspectives. Specifically, this part extracts image semantic information from two dimensions: first, it uses an image description generation model to perform a semantic summary of the entire image to obtain the image description content at the natural language level; second, it uses an optical character recognition (OCR) model to detect and identify the text areas embedded in the image, extracting the actual text information in the image, such as titles, slogans, screenshots, etc.
[0125] The second step is text feature extraction. Text is the main content of the news text language, which contains the main descriptive information of the news. The pre-trained language model represented by BERT, because it is trained on large-scale data, also contains a lot of grammatical knowledge and implicit information between sentences, which greatly enhances the semantic representation ability of the model. Therefore, the present invention uses the BERT pre-trained model as a feature extractor for news text modality. For the input text sentence S = [s1, s2, ..., s n ] to encode and obtain the token list T, where T = [[CLS], token1, token2, ..., token n , [SEP]]. Then the token list T is fed into the BERT model to obtain the word embedding vector W = [w cls ,w1,...,w n , w sep ], where W = BERT(T). Finally, the embedding vectors of all words are fed into the masked attention network. By calculating the weight of each word, its vector representation is weighted to highlight the influence of more important words on the sentence. The semantic feature representation e of the news text sentence with a dimension of 768×1 is obtained. t , the formula is expressed as e t =Mask_attention(W).
[0126] Finally, audio feature extraction is performed. This section focuses on modeling speech audio in news data and uses the Conv-TasNet fully convolutional time-domain audio separation network. First, the original speech waveform is converted into a time-domain representation through a linear encoder, retaining its structural features. Subsequently, a mask estimation module is introduced to perform weighted processing on the encoded representation to enhance the target speech content, suppress background interference, and achieve effective decoupling of the sound subject. Finally, a linear decoder is used to restore and extract high-quality speech audio features. This method can achieve end-to-end feature modeling without the need for frequency domain transformation, making it suitable for audio processing tasks in multi-speaking environments and providing stable audio modal input for fake news detection systems.
[0127] The feature fusion layer is primarily responsible for effectively integrating feature representations from multiple modalities, such as text, images, and audio. It uses a fusion strategy to semantically align and complement the features of each modality, ultimately outputting a unified joint representation. This joint representation retains the key semantic information of each modality while enhancing cross-modal semantic relevance, providing a more comprehensive and accurate semantic foundation for subsequent classification and discrimination.
[0128] S40: The fusion process introduces a multi-scale semantic aggregation mechanism, which uses residual connections and convolution kernels of different scales to gradually fuse the fine-grained representations within the modality to form a joint representation space.
[0129] In this step, the image semantic features and text features are first concatenated to form the overall text features. Subsequently, the system introduces a scale semantic alignment network to achieve interactive fusion between feature vectors from the three modalities.
[0130] In terms of naming, let T0 represent the semantic feature vector of the text sentence, T p Represents the semantic feature vector of the text in the image, T a The semantic feature vector representing the entire text; the visual feature vector of the image modality is denoted as I0. The system uses element-by-element addition (+) and element-by-element multiplication (⊙) to construct the semantic interaction features between text and image, then performs multi-scale one-dimensional convolution (1DConv) and maximum pooling (MaxPool) processing on them, and combines cosine similarity for feature alignment. The calculation formula is as follows:
[0131]
[0132] In the above formula, ε is a minimum constant that prevents the denominator from being zero. The product of the two constitutes the joint representation vector e after the fusion of image and text p :
[0133] e p =f cmp (I0, T a )*cosine_similarity(I0,T a )
[0134] Next, to introduce audio modal information, the system combines the speech audio feature vector V0 with the image and text to represent the p The semantic fusion of similar processes is performed again, that is, one-dimensional convolution, maximum pooling and cosine similarity calculation are used to enhance its multimodal collaborative expression ability. The calculation of cosine similarity and the final multimodal fusion representation is as follows:
[0135]
[0136] e f =f cmp (e p ,V0)·cosine_similarity(e p ,V0)
[0137] The final fusion vector e f It is a multi-scale joint semantic representation containing trimodal information of text, image and audio.
[0138] S50. Use the classification layer to convert the joint representation space into a probability distribution of true or false categories, and determine whether the current news is false news based on the probability distribution.
[0139] In this step, the output layer first includes two sets of linear transformation layers, batch normalization layers and ReLU activation functions to f Perform nonlinear transformation to obtain e` f , and then use a fully connected layer with a softmax activation function to output e` f Represents the probability of true news and false news. The specific formula is y = softmax(W'×e` f )+b. Where W' is the weight parameter of the fully connected layer and b is the bias term.
[0140] This problem is a binary classification problem. The present invention uses the binary cross entropy loss function (BCELoss) as the loss function. The formula is: where y i represents the true label value, p i Represents the model prediction results.
[0141] The present invention proposes a fake news detection model based on multi-scale semantic alignment. The model is composed of an input layer, a feature representation layer, a multi-scale semantic alignment layer, and an output layer in sequence. It has end-to-end modeling capabilities and can realize full-process intelligent processing from raw multimodal data input to fake news determination.
[0142] Among them, the input layer is used to complete the preprocessing of multimodal data, including obtaining original modal data such as news text, images and audio from different information sources, and performing corresponding standardized processing operations on each modal data, such as text segmentation, cleaning, and vectorization, image size normalization and noise reduction, audio framing and resampling, etc., thereby providing a unified and structured input format for subsequent feature extraction.
[0143] The feature representation layer is based on the current mainstream deep neural network architecture. It uses the Transformer network to extract the contextual semantic representation of news text, uses OCR technology and the Vision-Transformer model to obtain text information and visual features in images, and uses the Conv-TasNet network to decouple and model the voice information in the audio modality, ultimately generating deep semantic representations of the three modalities.
[0144] The multi-scale semantic alignment layer is the model's core innovation. Its design incorporates multi-scale modeling and inter-modal semantic interaction mechanisms to fully exploit the complementary relationships between text, image, and audio modalities at different semantic levels. Through semantic alignment and residual fusion strategies, a cross-modal semantically consistent joint representation vector is constructed, effectively alleviating inter-modal semantic heterogeneity and significantly enhancing the ability to discriminate complex fake news structures.
[0145] The output layer uses a classifier to map the fused multimodal semantic representations into corresponding probability distributions, enabling the true / false distinction of news content. The model not only outputs a binary "true" / false" classification result but also exhibits excellent interpretability and scalability, adapting to multimodal news data input from diverse sources and genres.
[0146] A second aspect of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores program instructions, and when the program instructions are executed, they are used to execute the above-mentioned multi-scale semantic alignment fake news detection method and system.
[0147] A third aspect of the present invention provides a multimodal false information detection system, which includes the above-mentioned computer-readable storage medium.
[0148] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the technical principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A multi-scale semantic alignment fake news detection method, characterized by: The following steps are involved: S10: Obtain multimodal news raw data including text, audio, and images from different information sources, integrate and store them as structured data using a predefined format; S20: using a modality-aware processing strategy to pre-process text, audio, and image modality data respectively; S30: Extract self-attention coarse-grained representations of each modality using a large-scale pre-trained model; S40: Introducing a multi-scale semantic aggregation mechanism in the fusion process, using residual connections and convolution kernels of different scales to gradually fuse the fine-grained representations between modalities to form a joint representation space; S50: Using a classification layer, the joint representation space is converted into a probability distribution of true or false categories, and whether the current news is false news is determined based on the probability distribution.
2. The multi-scale semantic alignment fake news detection method according to claim 1, characterized in that: The steps of preprocessing the text include: Clean the text to remove irrelevant characters, emoticons, URLs, punctuation, unify capitalization, and remove stop words; Perform word segmentation and part-of-speech tagging on the cleaned text to identify nouns, verbs, and adjectives; The Word2Vec word vector model is used to obtain the distributed semantic representation of the text and merge it into sequence features.
3. The multi-scale semantic alignment fake news detection method according to claim 1, characterized in that: The step of preprocessing the image comprises: Normalize, resize, and Gaussian blur the image; Vision-Transformer is used to extract the visual features of the image and obtain image features in the form of multi-channel vectors.
4. The multi-scale semantic alignment fake news detection method according to claim 1, characterized in that: The step S30 includes extracting semantic features of the image, and the step of extracting semantic features of the image includes: Obtain text in the image as recognized text through optical character recognition; The recognized text is processed using Transformer to extract image semantic features in vector form.
5. The multi-scale semantic alignment fake news detection method according to claim 1, characterized in that: The method of using a large-scale pre-trained model to extract self-attention coarse-grained representations of each modality includes: Use the Transformer model to extract the sequence features of the text and obtain the text features in vector form; Use Conv-TasNet to extract audio features and obtain audio features in vector form.
6. The multi-scale semantic alignment fake news detection method according to claim 1, characterized in that: The fusion process of the multi-scale semantic aggregation mechanism includes: (1) Add the image semantic features and text features element by element to obtain the first joint feature, and multiply the audio features and image features element by element to obtain the second result; (2) Connecting the image feature, the audio feature and the second result to obtain a first joint feature; (3) connecting the first joint feature and the second joint feature to obtain a third joint feature; (4) Perform multi-scale convolution on the third joint feature to obtain a final multi-scale joint representation.
7. The multi-scale semantic alignment fake news detection method according to claim 1, characterized in that: The processing steps of the classification layer include: Performing a nonlinear transformation on the joint representation through a linear transformation layer, a batch normalization layer, and a ReLU activation function to obtain the original category score; The raw category scores are processed using a fully connected layer with a softmax activation function to obtain the category probability distribution of real news and fake news.
8. A multi-scale semantic alignment fake news detection system, characterized by: The system is used to implement the method according to any one of claims 1 to 7, comprising: a. Input layer, used to perform data acquisition and preprocessing: Obtain multimodal news raw data from different information sources, including text, audio, and images; Perform word segmentation, stop word removal, cleaning, and vectorization on text data; Perform normalization, resampling, and framing on audio data; Perform pixel normalization, size cropping, and Gaussian blur preprocessing on image data; b. Feature representation layer, used to extract features of each modality: Text feature extraction: Use the Transformer model to extract the sequence features of the text and combine it with Word2Vec to generate distributed semantic representations; Image feature extraction: Vision-Transformer is used to extract the visual features of the image, and optical character recognition combined with Transformer is used to extract the semantic features of the text in the image; Audio feature extraction: Use Conv-TasNet to extract the time domain features of audio; c. Multi-scale semantic alignment layer for inter-modal feature fusion: Using residual connections and convolution kernels of different scales, text features, image visual features, image semantic features, and audio features are fused stage by stage. Generate a joint representation space containing multi-scale semantic information; d. Output layer, used for classification decision: The joint representation is nonlinearly transformed through a linear transformation layer, a batch normalization layer, and a ReLU activation function to obtain the original category score; Use a fully connected layer with a softmax activation function to output the class probability distribution of true / false news; The input layer, feature representation layer, multi-scale semantic alignment layer, and output layer are connected in sequence to form a multi-scale semantic alignment network, realizing end-to-end processing from raw data input to fake news discrimination.
9. The multi-scale semantic alignment fake news detection system according to claim 8, characterized in that: The image feature extraction of the feature representation layer includes: Recognizing text content in the image by an optical character recognition subunit; The text recognized by OCR is processed by the Transformer semantic extraction sub-unit, and the image semantic features in vector form are output.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program instructions, which, when executed, are used to execute the multi-scale semantic alignment fake news detection method according to any one of claims 1 to 7.
Citation Information
Cited By
Q-Former mechanism-based emotion semantic multi-modal fusion detection method
CN122221120A