A fake news detection inference method and system based on large model and pre-training
Through the large language model and multimodal language alignment pre-training method, combined with multimodal feature fusion and enhancement module, the problem of difficult detection of forged news in the existing technology is solved, and higher detection accuracy and reasoning capabilities are achieved.
Patent Information
- Application Number
- CN202510314705.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-03-18
AI Technical Summary
The prior art is difficult to effectively detect and identify fake news, especially in multimodal information processing and potentially manipulated reasoning.
The large language model and multimodal language alignment pre-training method are adopted to perform multimodal features fusion and enhancement of real and forged news data, and the multimodal cross attention module and multimodal feature enhancement module are used to perform forged news detection and reasoning in combination with the large language visual model.
Improve the accuracy and reasoning ability of fake news detection, and can more effectively capture nuances and potential signs of manipulation in multimodal information, enhancing the robustness and generalization of the model.
Smart Images

Figure CN119851078B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of natural language processing, image processing, deep learning and deep fake face detection, and in particular to a method and system for detecting and inferring fake news using a large language model and multimodal language alignment pre-training. Background Art
[0002] In recent years, the rapid development of the Internet has promoted the widespread popularization of digitalization and the breakthrough improvement of computing power, and the era of artificial intelligence has quietly arrived. However, at the same time, with the development of media technology and the changes in social ecology, as a product accompanying mass communication, fake news, which inevitably exists to a certain extent, has shown new characteristics and trends in the context of network communication and even the era of intelligent communication. Fake news specifically describes and reflects various false news, fake news events, fake news reports, news hype, and public relations news. It includes both false reproduction of real events and fabrication and fiction. As a companion phenomenon of news work, fake news has a certain negative impact on news media and the public. Governing fake news is also an inevitable and important task to maintain the current network content ecology. Some people disguise themselves as authoritative news media by forging news studio scenes, imitating professional hosts, abusing artificial intelligence virtual anchors, etc., or use cutting and pasting, patchwork and other means to fabricate fake news related to hot topics such as social events and international current affairs. The continuous innovation of counterfeiting methods also reflects the new changes in the expression, dissemination mode, and identification difficulty of fake news under the background of artificial intelligence technology application, which puts forward higher requirements and challenges for the governance of online fake news. Therefore, how to detect such fake news has become an urgent problem to be solved.
[0003] In the direction of fake news detection based on deep fake technology, there are three main challenges: (1) Visual deep fake models can edit faces and generate high-fidelity images and videos. With large language models such as BERT and GPT, vocabulary replacement and editing can be easily performed to modify semantics and facts. Media manipulation increases the difficulty of detection. (2) With the development of deep generative models, media manipulation has expanded to multiple modalities. For example, visual deep fake models can edit faces and generate high-fidelity images and videos. This requires fake news detection methods to be able to process and analyze multimodal information to identify and reason about potential manipulations. (3) Most existing forgery detectors can only perform binary classification of authenticity and cannot reason about potential manipulations. This is far from enough for analyzing forgeries or identifying forgery mechanisms, limiting the breadth of practical applications.
[0004] With the development of visual language pre-training models, the number of modal alignments has expanded from the initial visual and language modalities to other modal information alignments, and this pre-training method has been cited in many works. Since the language modality contains rich semantic information, this helps the model better understand and reason about the authenticity of news content. By using language as a binding between different modalities, the data of multiple modalities are directly mapped into a semantically shared embedding space through a contrastive learning mechanism to achieve direct and accurate semantic alignment, which helps the model to show excellent performance in multimodal downstream tasks. Large language models are usually pre-trained on large-scale text data and have a rich knowledge reserve. This enables them to use this knowledge for deeper reasoning and analysis when processing multimodal information, and they are designed with good adaptability and generalization capabilities, and can handle multiple types of multimodal data and tasks. This ability enables the model to maintain high reasoning performance when facing new or unseen multimodal information. By fusing multiple modal information and using a large language model for inference, the subtle differences and potential associations between different modalities can be captured, thereby improving the accuracy of reasoning.
[0005] Therefore, the methods of multimodal pre-training alignment, multimodal information fusion and large oracle model inference can fully explore the subtle differences and potential signs of forgery in fake news, and improve the accuracy of fake news detection and reasoning ability. Summary of the invention
[0006] The purpose of the present invention is to provide a fake news detection inference method and system based on a large model and pre-training, which can improve the accuracy, robustness and reasoning ability of fake news detection. To achieve the above purpose, the present invention provides the following solutions:
[0007] The present invention provides a fake news detection and inference method based on a large model and pre-training, characterized in that the specific steps are as follows:
[0008] Step 1: Collect original real information from real-world news media, including video, audio and corresponding text information;
[0009] Step 2: Use the collected real image information and text information to forge images, texts, and face frame images in the video using a forging method to obtain a forged data set; combine the real data and the forged data into a real and forged data set as a training data set;
[0010] Step 3: Use Vision Transformer to encode image and audio information, and Transformer to encode text information. Use contrastive learning technology to achieve multimodal semantic alignment, and obtain the required encoder through training.
[0011] Specifically, for image and audio information, Vision Transformer is used as an encoder to encode and obtain a series of vectors, which are then input into the Transformer model for feature extraction;
[0012] For text information, the word segmenter is used to split the word into common subwords, each subword corresponds to a unique token, and the Transformer model is used to encode the text data according to the token;
[0013] The contrastive learning technique is used to bind non-text modal information, i.e., image and audio information, to text information, increase the similarity of paired information, bring them to the same semantic space, and minimize the similarity of unpaired information, projecting the non-text modal information into the same semantic space as the text information.
[0014] Using real audio, image and text information to train an encoder for processing text, image and audio information;
[0015] Step 4: Use the encoder for processing text, image and audio information obtained in step 3 to encode the text, image and audio information in the real and forged data sets, and the image information contains the face image information;
[0016] Step 5: The text features extracted in step 4 are fused with the image features using a multimodal cross attention module to obtain text-image fusion features; the text-image fusion features, face image features, and audio features are then processed using a multimodal feature enhancement module to obtain enhanced features that capture subtle differences and potential signs of manipulation in the news; a multimodal feature fusion and enhancement model is obtained by training on real and fake data sets;
[0017] The multimodal cross-attention module is expressed as:
[0018]
[0019] Among them, cross attention is performed on the image embedding by taking Q as the image embedding and K and V as the text embedding; represents the dimension of the embedding vector, represents the input image, Represents input text, represents the image encoder, Represents a text encoder; Represents the final multimodal cross-attention features;
[0020] The multimodal feature enhancement module is expressed as:
[0021]
[0022]
[0023] in represents a linear transformation, represents the attention map obtained by the attention mechanism, Represents multimodal features, namely facial features and audio features, represents the linear transformation operation for multimodal cross-attention features, represents the vector product operation, Represents a dimension transformation operation, and represents the activation function;
[0024] Step 6: The original information obtained is subjected to the multimodal feature fusion and enhancement model extraction obtained in step 5 to achieve enhanced features. The enhanced features obtained are input into the pre-trained large language vision model as prompt words. The model is used to detect and infer the truth of the news and obtain the inference results.
[0025] Preferably, the method for forging the face frame image in step 2 is:
[0026] Extracting face frame images from news videos, and then determining the face bounding box based on the results of face key point positioning, and cropping the area containing the face based on the bounding box;
[0027] Then, the cropped face frame image is subjected to data processing, including image denoising, contrast enhancement, image scaling, image rotation, and image horizontal and vertical mirror transformation, to obtain a two-dimensional face image;
[0028] Finally, the two-dimensional face image is forged by using face synthesis, face identity exchange and facial expression modification.
[0029] Preferably, the method of text forgery in step 2 is text word replacement and / or text entity replacement;
[0030] Among them, text word replacement is achieved by replacing the original words with words with attacking semantics; text entity replacement is to extract entities from the text through a named entity recognition model, retain only the extracted entity character names, discard the rest of the entities, and then replace the entity extracted character names with randomly selected character names to obtain forged text information.
[0031] Preferably, in step 3, for the audio information, when Vision Transformer is used as an encoder for encoding, the audio is first cut into blocks of fixed time length, and the insufficient part is supplemented by repetition and zero padding to obtain a single-channel spectrogram, and then the single-channel spectrogram is copied to three channels, divided into patches, and position encoding is added, and then these patches are linearly embedded into a series of vectors.
[0032] Preferably, in step 3, the Vision Transformer encoder freezes the encoder weight matrix during training and uses a low-rank adaptation method to fine-tune the model parameters.
[0033] The present invention also provides a fake news detection and inference system using the fake news detection and inference method based on a large model and pre-training, characterized in that the system comprises: a forged data generation unit, an encoder unit for text, image and audio information, a multimodal feature fusion and enhancement model unit and a large language model detection and inference unit;
[0034] The forged data generating unit is used to perform forging operations on images, texts and face frame images in videos to obtain real and forged data sets;
[0035] The encoder unit for text, image and audio information encodes the text, image and audio information, and the image information includes face image information;
[0036] The multimodal feature fusion and enhancement model unit consists of a multimodal cross attention fusion unit and a multimodal information enhancement unit.
[0037] Among them, the multimodal cross attention fusion unit is used to fuse the text features and image features extracted by the encoder unit of text, image and audio information to obtain text-image fusion features;
[0038] The multimodal information enhancement unit is used to process the text image fusion features, face image features and audio features to obtain enhanced features;
[0039] The large language model detection and inference unit passes the obtained features as prompt words to the large language vision model, and uses its powerful knowledge base to detect and infer the authenticity of the news content.
[0040] Preferably, the forged data generating unit comprises: a face frame image acquiring unit, a face image acquiring unit, an image preprocessing unit, a face image forging unit and a text forging unit.
[0041] Wherein, the face frame image acquisition unit is used to extract the face frame image from the video;
[0042] A face image acquisition unit, used for extracting a face frame image from a face video;
[0043] An image preprocessing unit, used to perform data processing on the frame image, including image denoising, contrast enhancement, image scaling, image rotation, and image horizontal and vertical mirror transformation to obtain a processed image;
[0044] A face image forging unit for forging the obtained face image, wherein the forging methods include face synthesis, face identity exchange and face expression modification;
[0045] The text forging unit forges the obtained text information according to the following two methods: text word replacement and text entity replacement.
[0046] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0047] The invention discloses a fake news detection inference method and system based on a large model and pre-training. The method comprises extracting text and related images from news data; preprocessing the text and images, including text cleaning, feature extraction, image enhancement and corresponding data falsification, to obtain a multimodal feature representation; using a large visual language model as a basis, combined with a multimodal alignment technology, to perform deep fusion and semantic understanding on the multimodal feature representation; through multimodal feature fusion, integrating text and image features to form a unified feature representation for fake news detection; through multimodal feature enhancement, taking facial image features and audio information as feature enhancement information, and guiding the learning of more robust and information-rich multimodal features by performing multimodal fusion features and facial features and audio feature enhancement; finally, using a large language model to analyze the fused features to identify the authenticity of the news and infer potential manipulation behaviors.
[0048] The present invention can enhance the model's ability to perceive subtle differences in news content, while suppressing attention to irrelevant or misleading information, by sequentially performing multimodal feature alignment, extraction, attention feature fusion, and feature enhancement, thereby improving the accuracy of fake news detection and the ability to reason about fake behavior. In addition, by introducing a large language model, the present invention can detect and adapt to different types of fake news and perform reasoning, improving the generalization and robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 The figure is a flowchart of the fake news detection and reasoning method based on a large model and pre-training in the present invention.
[0050] Figure 2 The figure is a schematic diagram of the flow of data processing of the image according to the present invention.
[0051] Figure 3 The figure is a schematic diagram of the process of forging the text data according to the present invention.
[0052] Figure 4 The figure is a schematic diagram of the process of forging the face according to the present invention.
[0053] Figure 5 This is a schematic diagram of the multimodal alignment pre-training process of the present invention.
[0054] Figure 6 The figure is a schematic diagram of the reasoning process based on a large language model of the present invention. DETAILED DESCRIPTION
[0055] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0056] The purpose of the present invention is to provide a fake news detection and inference method based on a large model and pre-training, which can detect and adapt to different types of fake news and perform inference, thereby improving the generalization and robustness of the model.
[0057] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0058] Example
[0059] like Figure 1 As shown, this embodiment provides a face forgery detection method based on reconstruction learning and hybrid expert mode, including:
[0060] Step 100: Collecting original real information from real-world news media, including video, audio, and corresponding text information;
[0061] Step 200: The collected real image data and text data are forged using a forging method to obtain forged data, and the real data and the forged data are combined into a real and forged data set as a training data set;
[0062] like Figure 2 As shown, the present invention extracts face frame images from news videos; wherein, the face frame image extraction is performed by using dlib to locate the key points of the face for each detected face, and then determining the bounding box of the face according to the result of the facial key point location, and cropping the image according to the bounding box to crop out the area containing the face;
[0063] Performing data processing on the cropped face frame image, including image denoising, contrast enhancement, image scaling, image rotation, and image horizontal and vertical mirror transformation, to obtain a two-dimensional face image;
[0064] Among them, image denoising processing includes three operations: dilation, erosion, and median filtering;
[0065] The obtained face image is forged, and the forging methods are face synthesis, face identity exchange and facial expression modification. In this embodiment, Deepfake is used for face synthesis, Faceswap is used for face identity exchange, and a GAN-based method is used for facial expression modification, so as to finally achieve the forgery of face information;
[0066] like Figure 3 As shown, the obtained text information is forged according to the following two methods to obtain the data set used for training:
[0067] a. Text word replacement, by replacing the original word with a word with offensive semantics. In the present invention, offensive semantics is defined as language with insulting and discriminatory meanings such as abuse, slander, contempt, and ridicule. The words in the input text are modified to reverse the global semantics of the text;
[0068] b. Text entity replacement: A named entity recognition model is started for the input text to extract the names of the main entities. The existing tool Hanlp is used to extract entities, and only the extracted entity character names are retained, and the rest of the entities are discarded. The entity extracted character names are then replaced with randomly selected character names.
[0069] Step 300: Figure 4 As shown in Figure 1, the Vision Transformer is used as an encoder to encode the image in blocks, dividing the image into multiple patches of fixed size, and then linearly embedding these patches into a series of vectors. These vectors are then input into the Transformer model.
[0070] The Vision Transformer model is composed of several separate ViT modules. Each ViT block contains a layer normalization (LN) module, a feed-forward network (FFN) layer, and a residual connection, which can effectively extract local and global semantic features of the image.
[0071] Vision Transformer is used as an encoder to convert audio information into a spectrogram for block encoding. First, the audio is cut into blocks of fixed time length. The insufficient part is supplemented by repetition and zero padding to obtain a single-channel spectrogram. Next, the single-channel spectrogram is copied to three channels, divided into patches, and position encoding is added. These patches are then linearly embedded into a series of vectors. These vectors are then input into the Transformer model. Finally, features are extracted through the self-attention mechanism.
[0072] The mask operation is used to selectively transfer the unmasked features to the encoder of the next level of frozen parameters for fine-tuning. The Vision Transformer encoder freezes the encoder weight matrix during training and uses the Low-Rank Adaptation (LoRA) method to fine-tune the model parameters.
[0073] The forward propagation process in the Low-Rank Adaptation (LoRA) method is defined as follows:
[0074]
[0075] in, represents the weight matrix of a modality encoder, which remains unchanged during training. represents the trainable low-rank matrix added by LoRA, Represents the input data of the model.
[0076] The Transformer model is used as the encoder of text data. The Transformer model is initialized using a pre-trained model. For a given text, the Byte-Pair Encoding (BPE) tokenizer is first used to segment words into relatively common subwords. Each subword corresponds to a unique token, which is encoded by the language encoder.
[0077] Contrastive learning technology is used to bind each modality to language, increase the similarity of paired data, bring them to the same semantic space, and minimize the similarity of unpaired data, projecting non-text modal information into the same semantic space as text information, which can be expressed as:
[0078]
[0079]
[0080] in represents the i-th modal data, represents the jth text, is the batch size, is the temperature parameter that aligns each mode M directly with the language T.
[0081] Step 400: The encoders for processing language, image, and audio in the Vision Transformer trained in step 300 are used as feature extraction encoders in the second stage, and feature extraction is performed on text, image, audio, and face image respectively;
[0082] Step 500: The multimodal features extracted in step 400 are fused and enhanced using the multimodal information fusion module and the multimodal feature enhancement module. The extracted text features and image features are used as low-level features, and feature fusion is performed using the multimodal cross attention module (e.g. Figure 6 as shown).
[0083] The multimodal cross-attention fusion module is able to fuse information from different modalities by applying a cross-attention mechanism. This process enhances the model's attention to the important parts of the input features while suppressing unimportant feature representations. This approach allows the model to effectively share and fuse information between different modalities to improve the ability to understand complex tasks. This process can be expressed as:
[0084]
[0085] Among them, cross attention is performed on the image embedding by taking Q as the image embedding and K and V as the text embedding; represents the dimension of the embedding vector, represents the input image, Represents input text, represents the image encoder, Represents a text encoder; Represents the final multimodal cross-attention features.
[0086] The multimodal feature enhancement module obtains an attention map, i.e., enhanced features, by performing fusion features and cross-attention between face image features and audio image features to capture subtle differences and potential signs of manipulation in news, thereby improving detection accuracy;
[0087] The multimodal feature enhancement module can be expressed as:
[0088]
[0089]
[0090] in represents a linear transformation, represents the attention map obtained by the attention mechanism, Represents multimodal features, namely facial features and audio features, represents the linear transformation operation for multimodal cross-attention features, represents the vector product operation, Represents a dimension transformation operation, and represents the activation function;
[0091] Step 600: Using a pre-trained large language vision model (such as ChatGPT), the enhanced features obtained in the previous stage are input into the model as prompt words, and its powerful knowledge base is used to enhance the model's ability to understand and infer the authenticity of news content.
[0092] In addition, the present invention also provides a fake news detection inference system based on a large model and pre-training, including:
[0093] An image acquisition unit, used for extracting frame images from a video;
[0094] A face image acquisition unit, used for extracting a face image from a video;
[0095] An image preprocessing unit, used for performing data processing on the image, including image denoising, contrast enhancement, image scaling, image rotation, and image horizontal and vertical mirror transformation, to obtain an image;
[0096] The text forging unit forges the obtained text information according to the following two methods: text word replacement and text entity replacement;
[0097] The face image forging unit forges the obtained face image, and the forging methods include face synthesis, face identity exchange and facial expression modification.
[0098] For its pre-trained units, multimodal language uses Vision Transformer to encode images and audio, and uses ordinary Transformer to encode text. The encoded image and audio features are masked according to a certain ratio and then encoded again using the pre-trained encoder. At the same time, its parameters are frozen, and LoRA is used for parameter fine-tuning. Contrastive learning is used to bring them to the same semantic space while minimizing the similarity of unpaired data.
[0099] The multimodal cross-attention fusion unit uses the facial image features and audio information obtained in the previous stage as feature enhancement information, and uses the features based on the attention mechanism to enhance the information of the fused features. By performing cross-attention between multimodal fusion features and facial image features and audio image features, an attention map is obtained, thereby guiding the learning of more robust and information-rich multimodal features;
[0100] The multimodal information enhancement unit uses the facial image features and audio information obtained in the previous stage as feature enhancement information, and uses the features based on the attention mechanism to enhance the information of the fused features. By performing cross-attention between multimodal fusion features and facial image features and audio image features, an attention map is obtained, thereby guiding the learning of more robust and information-rich multimodal features;
[0101] The large language model detection and inference unit passes the obtained feature vector as a prompt word into the large language visual model, and uses its powerful knowledge base to enhance the model's ability to understand and infer the authenticity of news content. In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the embodiments can be referred to each other.
[0102] This article uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only used to help understand the core idea of the present invention. At the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.
Claims
1. A fake news detection inference method based on a large model and pre-training, characterized in that: The specific steps are as follows: Step 1: Collect original real information from real-world news media, including video, audio and corresponding text information; Step 2: Use the collected real image information and text information to forge images, texts, and face frame images in the video using a forging method to obtain forged data, and combine the real data and the forged data to form a real and forged data set as a training data set; Step 3: Use Vision Transformer to encode image and audio information, and Transformer to encode text information. Use contrastive learning technology to achieve multimodal semantic alignment, and obtain the required encoder through training. Step 4: Use the encoder for processing text, image and audio information obtained in step 3 to encode the text, image and audio information in the real and forged data sets, and the image information contains the face image information; Step 5: The text features extracted in step 4 are fused with the image features using a multimodal cross attention module to obtain text-image fusion features; the text-image fusion features, face image features, and audio features are then processed using a multimodal feature enhancement module to obtain enhanced features that capture subtle differences and potential signs of manipulation in the news; a multimodal feature fusion and enhancement model is obtained by training on real and fake data sets; The multimodal cross-attention module is expressed as: Among them, by Q as image embedding, K and V as text embedding, the image embedding is cross-attentioned, where represents the dimension of the embedding vector, represents the input image, Represents input text, represents the image encoder, represents a text encoder, Represents the final multimodal cross-attention features; The multimodal feature enhancement module is expressed as: in represents a linear transformation, represents the attention map obtained by the attention mechanism, Represents multimodal features, namely facial features and audio features, represents the linear transformation operation for multimodal cross-attention features, represents the vector product operation, Represents a dimension transformation operation, and represents the activation function; Step 6: The original information obtained is subjected to the multimodal feature fusion and enhancement model extraction obtained in step 5 to achieve enhanced features. The enhanced features obtained are input into the pre-trained large language vision model as prompt words. The model is used to detect and infer the truth of the news and obtain the inference results.
2. The fake news detection and inference method based on large models and pre-training according to claim 1 is characterized in that: The method for forging the face frame image in step 2 is: Extracting face frame images from news videos, and then determining the face bounding box based on the results of face key point positioning, and cropping the area containing the face based on the bounding box; Then, the cropped face frame image is subjected to data processing, including image denoising, contrast enhancement, image scaling, image rotation, and image horizontal and vertical mirror transformation, to obtain a two-dimensional face image; Finally, the two-dimensional face image is forged by using face synthesis, face identity exchange and facial expression modification.
3. The fake news detection and inference method based on large models and pre-training according to claim 1 is characterized in that: The method of text forgery in step 2 is text word replacement and / or text entity replacement; Among them, text word replacement is achieved by replacing the original words with words with attacking semantics; text entity replacement is to extract entities from the text through a named entity recognition model, retain only the extracted entity character names, discard the rest of the entities, and then replace the entity extracted character names with randomly selected character names to obtain forged text information.
4. The fake news detection and inference method based on large models and pre-training according to claim 1 is characterized in that: In step 3, the image and audio information are encoded using Vision Transformer as an encoder to obtain a series of vectors, which are then input into the Transformer model for feature extraction. For text information, the word segmenter is used to split the word into common subwords, each subword corresponds to a unique token, and the Transformer model is used to encode the text data according to the token; The contrastive learning technique is used to bind non-text modal information, i.e., image and audio information, to text information, increasing the similarity of paired information and bringing them into the same semantic space, while minimizing the similarity of unpaired information and projecting non-text modal information into the same semantic space as text information. The encoder for processing text, image and audio information is trained using real audio, image and text information.
5. The fake news detection and inference method based on large models and pre-training according to claim 4 is characterized in that: In step 3, for the audio information, when using Vision Transformer as the encoder for encoding, the audio is first cut into blocks of fixed time length, and the insufficient part is supplemented by repetition and zero padding to obtain a single-channel spectrogram. The single-channel spectrogram is then copied to three channels, split into patches, and position encoding is added. These patches are then linearly embedded into a series of vectors.
6. The fake news detection inference method based on large model and pre-training according to claim 5, characterized in that: In step 3, the Vision Transformer encoder freezes the encoder weight matrix during training and uses a low-rank adaptation method to fine-tune the model parameters.
7. A fake news detection and inference system using the fake news detection and inference method based on a large model and pre-training as described in claim 1, characterized in that: The system includes: a forged data generation unit, an encoder unit for text, image and audio information, a multimodal feature fusion and enhancement model unit, and a large language model detection inference unit; The forged data generating unit is used to perform forging operations on images, texts and face frame images in videos to obtain real and forged data sets; The encoder unit for text, image and audio information encodes the text, image and audio information, and the image information includes face image information; The multimodal feature fusion and enhancement model unit consists of a multimodal cross attention fusion unit and a multimodal information enhancement unit. Among them, the multimodal cross attention fusion unit is used to fuse the text features and image features extracted by the encoder unit of text, image and audio information to obtain text-image fusion features; The multimodal information enhancement unit is used to process the text image fusion features, face image features and audio features to obtain enhanced features; The large language model detection and inference unit passes the obtained features as prompt words to the large language vision model, and uses its powerful knowledge base to detect and infer the authenticity of the news content.
8. The fake news detection and inference system according to claim 7, characterized in that: The forged data generating unit comprises: a face frame image acquiring unit, a face image acquiring unit, an image preprocessing unit, a face image forging unit and a text forging unit. Wherein, the face frame image acquisition unit is used to extract the face frame image from the video; A face image acquisition unit, used for extracting a face frame image from a face video; An image preprocessing unit, used to perform data processing on the frame image, including image denoising, contrast enhancement, image scaling, image rotation, and image horizontal and vertical mirror transformation to obtain a processed image; A face image forging unit for forging the obtained face image, wherein the forging methods include face synthesis, face identity exchange and face expression modification; The text forging unit forges the obtained text information according to the following two methods: text word replacement and text entity replacement.
Citation Information
Patent Citations
Video file detection method and device and computing equipment
CN118587625A
Risk content identification method based on multi-modal large model
CN119339419A