Multi-modal false news detection method based on wavelet frequency perception and high-order interaction
The multimodal fake news detection method based on wavelet frequency sensing and high-order interaction solves the problems of insufficient utilization of frequency domain information and low-order interaction, realizes stronger frequency anomaly description and high-order correlation expression, and improves the accuracy and robustness of fake news detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-22
- Publication Date
- 2026-03-27
AI Technical Summary
Existing multimodal fake news detection technologies suffer from insufficient utilization of frequency domain information, lack of cross-granularity interaction modeling, lack of mandatory constraints in frequency anomaly statistics, and low-order interactions on the fusion side, making it difficult to fully express high-order correlations between modalities.
We employ a wavelet frequency sensing and high-order interaction approach, using multi-level two-dimensional wavelet decomposition and cross-frequency three-path gating fusion, combined with a multiplicative coupling high-order interaction network, to construct frequency anomaly descriptors and multimodal fusion features. We also introduce strong constraints such as cross-scale phase consistency statistics for end-to-end optimization.
The model's ability to represent consistency between text and image content and visual detail anomalies has been improved, enhancing detection accuracy and robustness. It can better identify forged details and phase structure damage caused by forgery, and output explanatory hints.
Smart Images

Figure CN121743591A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of network and information security technology, specifically relating to a multimodal fake news detection method based on wavelet frequency perception and high-order interaction. Background Technology
[0002] With the development of social media and generative content technologies, fake news is showing a trend of multimodality and diversified forgery methods. It is characterized by the coexistence of textual narrative inducement and image detail forgery, which makes detection methods that rely solely on single textual semantics or conventional visual semantics prone to failure in the face of new forgeries.
[0003] Reference 1 (SHEN X, HUANG M, HU Z, et al. Multimodal Fake News Detection with Contrastive Learning and Optimal Transport[J]. Frontiers in ComputerScience, 2024, 6: 1473457.) proposes a multimodal fake news detection framework MCOT that integrates cross-modal attention, contrastive learning, and optimal transport. First, cross-modal attention is used to enhance the interaction between text and image features. Then, contrastive learning is used to align the cross-modal embedding space. Finally, optimal transport is combined to optimize the consistency of multimodal feature distribution, thereby improving detection performance. Reference 2 (QIAO J, LI X, GAO C, et al. Improving multimodal fakenews detection by leveraging cross-modal content correlation[J]. InformationProcessing & Management, 2025, 62(5): 104120.) proposes the cross-modal content correlation network C3N, which calculates the correlation between text and image content at both micro and macro levels, starting from the correspondence between text and images. Reference 3 (LIU J, WANG J, ZHANG P, et al. Multi-scale Wavelet Transformer for Face ForgeryDetection[C] / / Computer Vision – ACCV 2022. Springer, 2022: (52-68.) proposed the Multi-Scale Wavelet Transformer (MSWT), which uses discrete wavelet transform to extract multi-level frequency domain representations and guides the spatial feature extractor to focus on the forged region through a frequency domain-guided spatial attention module. At the same time, it adopts cross-modal attention to fuse frequency domain and spatial domain information to achieve effective fusion of spatial domain and frequency domain information.
[0004] Patent 1 (Liu Mingming, Liu Mengying, Wu Yike, et al. A method for detecting fake news by fusing evidence credibility [P]. Tianjin: CN116579337B, 2023-10-10.) captures and evaluates the credibility of multi-source evidence related to news, and integrates credibility features with text and social features for fake news detection. This scheme alleviates the misjudgment caused by "only looking at single text content" to a certain extent, but its modeling of image modalities is still biased towards conventional visual semantics, and it lacks stronger constraints on the multi-scale frequency domain structure within the image and the interaction between different frequency granularities. Patent 2 (Chen Ao, Huang Qi, Luo Wenbing, et al. A method for detecting multimodal fake news by fusing emotion and common attention network [P]. Jiangxi: CN117391051B, 2024-03-08.) proposes to combine emotion or multi-layer feature fusion to improve the effect of multimodal fake news detection, but its fusion side is mostly weighted or attention-based mechanisms, which are easily classified as "common fusion / gating" in retrieval comparison. / Attention" is replaced by equivalent; at the same time, the constraints on the composition of frequency anomaly statistics are not "mandatory" enough; Patent 3 (Wang Xiaoqiang, Xia Xu, Qi Chengrui. A multimodal fake news detection method and system based on text sentiment features and multi-level fusion [P]. Inner Mongolia Autonomous Region: CN118673165B, 2025-03-07.) proposes multi-level semantic fusion to improve detection performance, but the fine-grained structural constraints and cross-scale statistical stability characterization on the frequency domain side are still weak, and it lacks more difficult-to-replace structural constraints such as "one-to-one correspondence of cross-frequency three-path gating, gating competition normalization and dominance constraints, and gating generation input must contain statistical summary".
[0005] In summary, existing technologies generally suffer from the following shortcomings: First, the frequency domain side often remains at the level of "extracting or enhancing high frequencies," lacking structural constraints that make it more difficult to be equivalently replaced in retrieval and comparison; Second, the composition of frequency anomaly statistics is mostly an "optional set," lacking "mandatory" tightening for core statistics, making it easy for comparison documents to "flatten" the data by replacing statistical items; Third, the fusion side is mostly first-order weighting or attention, which is difficult to fully express the higher-order statistical correlations between multimodalities.
[0006] Therefore, there is an urgent need for a multimodal fake news detection method that can simultaneously provide more difficult equivalent substitutions on the frequency domain side, tighten the core statistical composition of frequency anomaly descriptors by making them "mandatory", and introduce multiplicative coupling high-order interactive fusion on the fusion side to enhance the expression of high-order statistical correlation between modalities, so as to make it more difficult to be "smoothed out" in retrieval and comparison. Summary of the Invention
[0007] This invention aims to overcome the problems in existing multimodal fake news detection technologies, such as insufficient utilization of image frequency domain information, inadequate modeling of frequency domain directionality differences and cross-granularity interaction, and multimodal fusion often remaining at low-order interaction and lacking scalable high-order collaborative modeling. It proposes a multimodal fake news detection method (WFHN) based on wavelet frequency perception and high-order interaction to improve the model's ability to represent consistency / inconsistency of image and text content and visual detail anomalies, thereby enhancing detection accuracy and robustness.
[0008] 1. A multimodal fake news detection method based on wavelet frequency sensing and higher-order interaction, characterized by comprising the following steps: S1. Obtain multimodal news samples containing text and image content, and optionally obtain dissemination metadata associated with the news samples, and encode the metadata features. ; S2. Input the text into a text encoder to obtain text modal features. The image is input into a visual encoder to obtain visual spatial modal features. and at least one layer of intermediate visual features ; S3. For the image or the intermediate visual features Perform multi-level two-dimensional wavelet decomposition to obtain the low-frequency subband of each level. and directional high-frequency subband , , Based on text modal features The directional high-frequency subbands are evaluated for directional importance and normalized using Softmax to obtain directional weights. Competitive weighted fusion is then performed on the directional high-frequency subbands to obtain directional-aware high-frequency features. ; S4. The low-frequency sub-band With the aforementioned direction-sensing high-frequency features Inputting a cross-frequency three-path gated fusion unit, the frequency domain feature map is obtained, and then aggregated and projected to obtain the frequency domain modal features. The cross-frequency three-path gating fusion unit includes at least a high-frequency to low-frequency update path, a low-frequency to high-frequency update path, and a same-frequency update path, and generates gating items for each. , , The gated item is composed of the and Statistical summarization and text modal features The joint representation is generated, and a normalized competition constraint is applied to the gating term to make the three paths adaptively dominant in a competitive manner; at the same time, the frequency anomaly descriptor is calculated. , wherein It must include at least cross-scale phase consistency statistics and subband correlation statistics, and may optionally include at least one of subband energy ratio, spectral flatness, kurtosis or skewness; S5, regarding the above , , After dimensional alignment, multimodal fusion features are obtained through a multi-order iterative multiplicative coupling high-order interaction network. , will the With the and optional The input classifier outputs the true prediction result and confidence score, along with explanatory information; a classification loss is constructed. Loss due to cross-domain logical consistency constraints under label conditions With frequency anomaly regularization loss End-to-end joint optimization is performed by weighted summation of the three factors.
[0009] 2. The multimodal fake news detection method based on wavelet frequency sensing and higher-order interaction according to claim 1, characterized in that, the 1 includes the following steps: S11. Construct a propagation representation from the propagation metadata. The propagation characterization includes at least one of the following: publishing source, time series features, or user interaction features; S12. Optional execution of evidence retrieval and semantic adjudication: retrieve evidence from a trusted corpus and calculate evidence consistency scores. ; S13, the above and or the aforementioned The evidence consistency score is used to calibrate the confidence level of the predicted authenticity or to trigger secondary discrimination of suspected contradictory samples; It can be obtained by measuring the similarity between textual representations and evidence representations, for example: ; in For the retrieved set of evidence, For evidence aggregation function, This is a similarity measurement function.
[0010] 3. The multimodal fake news detection method based on wavelet frequency sensing and higher-order interaction according to claim 1, characterized in that, the 3 includes the following steps: S31. Select the decomposition form of the multi-level two-dimensional wavelet decomposition, wherein the decomposition form is any one or a combination of fixed wavelet decomposition, learnable wavelet decomposition, dual-tree complex wavelet decomposition or wavelet packet decomposition. S32. Decompose the above according to the decomposition form. By performing step-by-step decomposition, low-frequency subbands at each level are obtained. and directional high-frequency subband , , ; S33. When using learnable wavelet decomposition, apply approximate orthogonality constraints and / or reconstruction consistency constraints to the wavelet filter bank parameters, and train them together with the detection model parameters to improve the ability of frequency domain decomposition to distinguish different forgery methods. S34. Utilize the text modality features The directional importance of the high-frequency subbands is evaluated and normalized to obtain directional weights, and then directional competitive weighted fusion is performed to obtain... ; where directional competitive fusion can be represented as: ; in For the reason The direction score is obtained by condition guidance.
[0011] 4. A multimodal fake news detection method based on wavelet frequency sensing and high-order interaction according to claim 1, characterized in that, the Includes the following steps: S41, based on the above and Computational Statistical Summary The statistical summary It should include at least a subband energy distribution summary, a directional response dispersion summary, and a cross-scale stability summary; S42. Generate three candidate outputs on the high-frequency to low-frequency update path, the low-frequency to high-frequency update path, and the same-frequency update path, respectively. , , ; S43, Based on the statistical summary Text modal features Generate three gating items respectively , , Furthermore, normalization competition constraints and sparse dominance constraints are applied to the three gating terms, so that the three candidate outputs exhibit competitive dominance fusion. S44, according to the above , , Regarding the , , Weighted fusion is performed to obtain the frequency domain feature map, which is then aggregated and projected to obtain... ; S45, Calculate the frequency anomaly descriptor , wherein It includes at least cross-scale phase consistency statistics and subband correlation statistics, and may optionally further incorporate at least one of subband energy ratio, spectral flatness, kurtosis, or skewness; wherein the normalization competition of the three-way gating terms can be characterized in the following form: ; This leads to a fused frequency domain representation of the three candidate outputs: ; in Generate mappings for gating. For aggregation and projection operations.
[0012] 5. A multimodal fake news detection method based on wavelet frequency sensing and higher-order interaction according to claim 1, characterized in that, the Includes the following steps: S51, regarding the above , , Perform feature alignment and normalization to reduce the impact of modal scale differences on the interaction; S52. The fused features are generated through at least two iterations of multiplicative coupling high-order interaction updates. Each iteration includes linear mapping of interactive inputs, nonlinear transformation, multiplicative coupling and residual update, and gating noise reduction is introduced to reduce spurious interactions. S53, the above With the and optional The input classifier outputs a true prediction result and confidence score, and generates explanatory information, which includes at least one of frequency significance hints and cross-modal contradiction hints; S54, Construct the above This makes the representation of real samples more consistent in the text and visual spatial domains, as well as in the text and frequency domains, and introduces an interval parameter. This causes the fake samples to exhibit inconsistency in at least one of the two pathways mentioned above; S55, Construct the above Used for measurement Compared with the reference distribution The difference, wherein the difference measure is any one of KL divergence, Wasserstein distance, or maximum mean difference (MMD). Obtained through sliding window statistics, exponential moving average, or category-based prototype updates, and the... , , Joint optimization based on weighted coefficients; a single update of the multiplicative coupling high-order interaction can be illustrated in the following form: ; in This indicates element-wise multiplication. and For two modal representations or intermediate fusion representations involved in the interaction.
[0013] 1. Stronger structural differences in the frequency domain and more difficult to replace with equivalent components. Through multi-level wavelet decomposition and text-guided directional competitive high-frequency fusion, a gate competition structure with three candidate outputs and three gate terms corresponding one-to-one is adopted in the cross-frequency fusion stage. Normalized competitive constraints and sparse dominance constraints are applied to the gates, so that the contributions of high-frequency to low-frequency, low-frequency to high-frequency and same-frequency updates can be adaptively selected and dominated. This makes it more stable to capture cross-scale and cross-directional forgery details and frequency domain anomalies, and improves robustness to different forgery types.
[0014] 2. The frequency anomaly descriptor is more discriminative and interpretable. The composition of the frequency anomaly descriptor has been tightened from an "optional set of statistics" to "must include cross-scale phase consistency statistics and subband correlation statistics," while allowing optional parameters such as energy ratio, spectral flatness, kurtosis, or skewness. Simultaneously, this descriptor participates in the calibration input of gating generation to suppress noise-induced spurious gating biases. This enhances sensitivity to phase structure and subband correlation disruptions caused by covert tampering and forgery, and can output frequency significance indicators, improving interpretability.
[0015] 3. Cross-modal fusion provides more comprehensive expression and more reliable discrimination. A multi-level iterative multiplicative coupling high-order interaction network is introduced in the fusion stage to perform high-order association modeling of text, visual spatial domain, and frequency domain features. A gated, induced dual-adjudication logic consistency constraint is constructed, ensuring that genuine samples receive consistent support in both the text-spatial and text-frequency domains simultaneously, while fake samples are induced to have low consistency support in at least one path. Combined with classification loss and frequency anomaly regularization for joint optimization, the accuracy and reliability of identifying difficult cases such as "semantically consistent but detail-forged" or "weak text-image contradictions" are improved. Attached Figure Description
[0016] Figure 1 This is a diagram illustrating the overall framework of a multimodal fake news detection method based on wavelet frequency sensing and high-order interaction used in the implementation of this invention. Figure 2This is a schematic diagram of the structure of the high-order interaction fusion module in a multimodal fake news detection method based on wavelet frequency sensing and high-order interaction used in the implementation of this invention. Figure 3 The training method is implemented for a multimodal fake news detection method based on wavelet frequency sensing and high-order interaction used in the implementation of this invention. Figure 4 This is a visualization of the fusion features of a multimodal fake news detection method based on wavelet frequency sensing and high-order interaction used in the implementation of this invention. Detailed Implementation
[0017] To provide a clearer understanding of the technical features, objectives, and beneficial effects of the present invention, an embodiment of the invention will be further described in conjunction with the accompanying drawings. This embodiment is only for further illustration of the invention and should not be construed as limiting the scope of protection of the invention. Non-essential improvements and adjustments made by those skilled in the art based on the content of the invention also fall within the scope of protection of the present invention.
[0018] Example 1: Model Training Implementation Method Obtain a multimodal news sample set, where each sample includes at least news text and an accompanying image, and optionally dissemination metadata. Divide the samples into training, validation, and test sets.
[0019] The text is cleaned, segmented into words or sub-words, and a maximum length is set, with truncation or padding to obtain the input text sequence. Images are scaled to a uniform size and normalized; random cropping and flipping can be used for enhancement during training. Discrete fields are embedded and encoded, while continuous fields are normalized and concatenated to form a metadata vector. Input a text sequence into a text encoder and output a text representation. And obtain the global text vector through pooling. Input the image into the visual encoder and output spatial visual features. Simultaneously, intermediate visual features are extracted from the intermediate layers of the visual encoder. It is used for frequency domain modeling.
[0020] Frequency domain modeling, multi-level wavelet decomposition, and direction competition fusion for intermediate visual features Perform multi-level two-dimensional Haar wavelet decomposition to obtain the low-frequency subband at each level. With three-directional high-frequency sub-band , , To strengthen the coupling between high-frequency directional information and text semantics, a directional importance evaluation branch is constructed, using text vectors. The system takes the aggregated statistics of high-frequency subbands in each direction as input, outputs scores for the three directions, and uses Softmax normalization to obtain directional weights. Then, it performs weighted fusion of the high-frequency subbands in the three directions to obtain directional-aware high-frequency features. Its normalization can be illustrated in the following form: Cross-frequency three-path gating competition fusion and frequency domain feature generation, three-path candidate update: at each scale, three candidate update paths are constructed around the "bidirectional information flow of high frequency and low frequency", and three candidate outputs are generated respectively. , , .in: Used to compensate for the low-frequency structure by reflecting high-frequency details; Used to demonstrate the constraint of low-frequency structure on high-frequency texture; For robust updates within the same scale; the above candidate outputs can be implemented by a combination of convolutional layers / linear mapping layers / normalization layers, and the three output dimensions are kept consistent for subsequent fusion.
[0021] Gated input construction, and statistical summarization for each scale. Preferably, it includes at least: a subband energy distribution summary, a directional response dispersion summary, and a cross-scale stability summary; simultaneously, a frequency anomaly descriptor is constructed. It must include cross-scale phase consistency statistics and subband correlation statistics; on this basis, statistics such as energy ratio, spectral flatness, kurtosis or skewness can be added to enhance the coverage.
[0022] Three-way gating generation and competition normalization, using statistical summaries Text vectors and frequency anomaly descriptor The combination of these terms serves as the input to the gating generator network, outputting three gating terms. , , Furthermore, a competitive normalization process using Softmax with temperature parameters is employed to create a competitive constraint among the three gating methods. The normalization can be illustrated as follows: Simultaneously, a gated dominance constraint (such as entropy penalty or sparsity regularization) is introduced to encourage the three-way gated selection to exhibit a dominant choice, thereby making the cross-frequency information flow more stable and interpretable. For the frequency domain feature output, the three candidate outputs are weighted and fused using the gated terms to obtain the scale-domain features, and the multi-scale frequency domain features are aggregated and projected to obtain the frequency domain modal features. In addition, Used as a calibration input for gating generation to suppress noise-induced spurious gating biases.
[0023] The high-order interactive fusion and classification output method aligns the dimensions of text representation, spatial visual representation, and frequency domain representation to construct a multiplicative coupled high-order interactive fusion network. Through multiple rounds of iterative interaction, it achieves deep collaborative modeling among the three to obtain fused features. .Will and and optional Input a classifier, output true prediction results and confidence scores; optional output explanation information may include: directional saliency hints (derived from directional weights), cross-frequency path hints (derived from gated dominant paths), and frequency anomaly hints (derived from...). The degree of abnormality was obtained.
[0024] The training objective and optimization employ classification loss as the primary objective, while introducing two types of auxiliary constraints; cross-domain logical consistency constraints measure the consistency of the "text-spatial domain" and "text-frequency domain" pathways respectively, and combine gating terms with... Adaptive determination of the importance of two paths ensures that real samples tend to be consistent across both paths, while fake samples exhibit low consistency across at least one path; frequency anomaly regularization constrains... The differences between the model and the reference statistical distribution or category prototype improve the stability of outlier statistics and can be used for confidence calibration; gated dominance regularization can also be added to enhance gating selectivity. End-to-end training is performed using common optimizers (such as Adam / AdamW), and the validation set is used for early stopping or selection of optimal model parameters.
[0025] Example 2: Detection Implementation Method The system inputs the news text and accompanying images to be detected, performs the same preprocessing as training, text and visual encoding, frequency domain modeling, three-path gating fusion and high-order interaction fusion to obtain the authenticity prediction results and confidence level. When the confidence level is lower than the preset threshold, an optional secondary verification strategy can be triggered, and an explanation prompt can be output simultaneously to support manual review.
[0026] Simulation Experiment This invention was evaluated against various comparative models on two widely used multimodal fake news detection datasets: Weibo and Twitter. Accuracy, precision, recall, and F1 score—common metrics in classification tasks—were selected as evaluation metrics. As shown in Table 1, the proposed WFHN model outperforms other comparative models in accuracy on both the Weibo and Twitter datasets, and also achieves outstanding performance on the other evaluation metrics. These results demonstrate that WFHN possesses good effectiveness and robustness in fake news detection tasks, fully reflecting its significant advantages in multimodal information fusion and feature representation.
[0027] To further verify the effectiveness of the feature fusion layer, this invention uses the t-SNE algorithm to reduce the dimensionality of the multimodal features obtained from the complete model WFHN and the ablation model WFHN-HF (which removes the wavelet cross-granularity frequency domain sensing module and the high-order interaction module) on the Weibo test set to a two-dimensional space and then visualizes them. For example... Figure 4 As shown, the WFHN-HF model exhibits significant feature overlap across different labels, making it difficult to distinguish between the two classes of samples in regions where real and fake news are mixed. In contrast, the features extracted by WFHN are more compactly distributed, with clearer boundaries between different labels, significantly reducing inter-class overlap. Therefore, the wavelet cross-granularity frequency domain sensing and high-order interactive multimodal feature fusion method proposed in this invention can learn more discriminative feature representations, contributing to improved fake news detection performance.
[0028] Table 1. Model comparison results on Weibo and Twitter datasets. The above description of the method of the present invention provides a clear understanding of how those skilled in the art can implement the method based on this description. Other embodiments obtained by those skilled in the art based on the above description of the present invention without inventive effort should also fall within the scope of protection of the present invention.
Claims
1. A multimodal fake news detection method based on wavelet frequency sensing and high-order interaction, characterized in that, Includes the following steps: S1. Obtain multimodal news samples containing text and image content, and optionally obtain dissemination metadata associated with the news samples, and encode the metadata features. ; S2. Input the text into a text encoder to obtain text modal features. The image is input into a visual encoder to obtain visual spatial modal features. and at least one layer of intermediate visual features ; S3. For the image or the intermediate visual features Perform multi-level two-dimensional wavelet decomposition to obtain the low-frequency subband of each level. and directional high-frequency subband , , Based on text modal features The directional high-frequency subbands are evaluated for directional importance and normalized using Softmax to obtain directional weights. Then, directional competitive weighted fusion is performed on the directional high-frequency subbands to obtain directional-aware high-frequency features. ; S4, based on the above and Forming three cross-frequency candidate outputs , , ,in Corresponding to the high-frequency to low-frequency update path. Corresponding to the low-frequency to high-frequency update path Corresponding to the same frequency update path; based on the and Statistical summary and text modal features Generate gated items , , The three candidate outputs are normalized and competitively fused using Softmax competitive gating with temperature parameters. A sparse dominance constraint is applied to the gating terms to encourage dominant selection among the three gating paths. The fusion result yields a frequency domain feature map, which is then aggregated and projected to obtain frequency domain modal features. ; Simultaneously calculate the frequency anomaly descriptor , wherein It must include at least cross-scale phase consistency statistics and subband correlation statistics, and may optionally further include at least one of subband energy ratio, spectral flatness, kurtosis, or skewness; and the stated Calibration inputs used for gating term generation to suppress spurious gating biases caused by noise; S5, regarding the above , , After dimensional alignment, multimodal fusion features are obtained through a multi-order iterative multiplicative coupling high-order interaction network. , will the With the and optional The input classifier outputs the true prediction result and confidence score, along with explanatory information; a classification loss is constructed. Gating-induced dual-decision logic consistency loss With frequency anomaly regularization loss End-to-end joint optimization is performed by weighted summation of the three factors; wherein... The decision weight is determined by the gating item. , , With the frequency anomaly descriptor The joint determination ensures that genuine samples receive consistent support in both the text and visual spatial domains, as well as in both the text and frequency domains, while fake samples are gated and induced to have low consistency support in at least one of the two decision paths; Used to constrain the The difference between the confidence level and the reference distribution is used to perform consistency calibration on the confidence level.
2. The multimodal fake news detection method based on wavelet frequency sensing and high-order interaction according to claim 1, characterized in that, S1 includes the following steps: S11. Construct a propagation representation from the propagation metadata. The propagation characterization includes at least one of the following: publishing source, time series features, or user interaction features; S12. Optional execution of evidence retrieval and semantic adjudication: retrieve evidence from a trusted corpus and calculate evidence consistency scores. ; S13, the above and or the aforementioned Used to calibrate the confidence level of the authenticity prediction results or to trigger secondary discrimination of suspected contradictory samples.
3. The multimodal fake news detection method based on wavelet frequency sensing and high-order interaction according to claim 1, characterized in that, S3 includes the following steps: S31. Select the decomposition form of the multi-level two-dimensional wavelet decomposition, wherein the decomposition form is any one or a combination of fixed wavelet decomposition, learnable wavelet decomposition, dual-tree complex wavelet decomposition or wavelet packet decomposition. S32. Decompose the above according to the decomposition form. By performing step-by-step decomposition, low-frequency subbands at each level are obtained. and directional high-frequency subband , , ; S33. When using learnable wavelet decomposition, apply approximate orthogonality constraints and / or reconstruction consistency constraints to the wavelet filter bank parameters, and train them together with the detection model parameters to improve the ability of frequency domain decomposition to distinguish different forgery methods. S34. Utilize the text modality features The directional importance of the high-frequency subbands is evaluated and normalized to obtain directional weights, and then directional competitive weighted fusion is performed to obtain... .
4. The multimodal fake news detection method based on wavelet frequency sensing and high-order interaction according to claim 1, characterized in that, S4 includes the following steps: S41, based on the above and Computational Statistical Summary The statistical summary It should include at least a subband energy distribution summary, a directional response dispersion summary, and a cross-scale stability summary; S42. Generate three candidate outputs on the high-frequency to low-frequency update path, the low-frequency to high-frequency update path, and the same-frequency update path, respectively. , , ; S43, Based on the statistical summary Text modal features Generate three gating items respectively , , Furthermore, normalization competition constraints and sparse dominance constraints are applied to the three gating terms, so that the three candidate outputs exhibit competitive dominance fusion. S44, according to the above , , Regarding the , , Weighted fusion is performed to obtain the frequency domain feature map, which is then aggregated and projected to obtain... ; S45, Calculate the frequency anomaly descriptor , wherein It includes at least cross-scale phase consistency statistics and subband correlation statistics, and may optionally further incorporate at least one of subband energy ratio, spectral flatness, kurtosis or skewness.
5. The multimodal fake news detection method based on wavelet frequency sensing and high-order interaction according to claim 1, characterized in that, S5 includes the following steps: S51, regarding the above , , Perform feature alignment and normalization to reduce the impact of modal scale differences on the interaction; S52. The fused features are generated through at least two iterations of multiplicative coupling higher-order interaction updates. Each iteration includes linear mapping of interactive inputs, nonlinear transformation, multiplicative coupling and residual update, and gating noise reduction is introduced to reduce spurious interactions. S53, the above With the and optional The input classifier outputs a true prediction result and confidence score, and generates explanatory information, which includes at least one of frequency significance hints and cross-modal contradiction hints; S54, Construct the above ,make The weights of the two adjudication pathways are determined by , , and Jointly determine and form a dual-adjudication target for gating inducement based on the authenticity label of the sample; S55, Construct the above Used for measurement Compared with the reference distribution The difference, wherein the difference measure is any one of KL divergence, Wasserstein distance, or maximum mean difference (MMD). Obtained through sliding window statistics, exponential moving average, or category-based prototype updates, and the... , , Joint optimization based on weighted coefficients.
Citation Information
Patent Citations
False news detection method fusing evidence credibility
CN116579337A
Emotion-fused common attention network multi-mode false news detection method
CN117391051A
Multi-modal false news detection method and system based on text emotion features and multi-level fusion
CN118673165A