A network news false information propagation screening method

By using historical data to filter and evaluate dimensions and weights, and combining forwarding volume and dissemination path to assess online news misinformation, this approach solves the problems of insufficient comprehensiveness and efficiency in existing technologies, and achieves multi-dimensional accurate screening and rapid response.

CN121092809BActive Publication Date: 2026-05-08NANJING FORESTRY UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING FORESTRY UNIV
Filing Date
2025-08-15
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies fail to effectively capture dynamic dissemination characteristics when detecting false information in online news, resulting in insufficient detection comprehensiveness, difficulty in distinguishing dissemination risk levels, and inadequate effectiveness of assessment dimensions and screening efficiency.

Method used

By selecting key evaluation dimensions based on historical monitoring data and determining their weighting percentages, and combining the increase in forwarding volume, text content, and dissemination path, a dissemination map is constructed to assess the source's authority, content consistency, and dissemination anomalies, thereby achieving multi-dimensional collaborative evaluation and switching the evaluation process between normal and emergency modes.

Benefits of technology

It has improved the adaptability and accuracy of the assessment system, enabling rapid response and accurate identification of false information, and ensuring efficient screening results in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121092809B_ABST
    Figure CN121092809B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of false information transmission screening, and specifically discloses a network news false information transmission screening method, which comprises the following steps: screening key evaluation dimensions and determining weights based on historical data, generating evaluation rules, determining a regular or emergency monitoring mode according to the transmission amount increase amplitude of the information to be screened and the number of matched words in a preset keyword library, identifying an initial publishing source through a transmission path, evaluating source authority, content consistency and transmission abnormality in the regular mode and outputting a comprehensive credibility, quickly evaluating and outputting an emergency credibility in the emergency mode, and finally matching the credibility with a warning interval to obtain a warning level; the source authority is evaluated based on source information, the content consistency is determined based on text content, and the transmission abnormality is evaluated based on transmission data, and then the comprehensive credibility is outputted in combination with the evaluation dimension rules, so that multi-dimensional collaborative evaluation of the source, the content and the transmission is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of screening technology for the spread of false information, and relates to a method for screening the spread of false information in online news. Background Technology

[0002] False information in online news refers to news content disseminated through online platforms that is inconsistent with objective facts or deliberately distorted. Typical manifestations include, but are not limited to, false statements, misleading expressions, fabricated data, and malicious alteration of core facts. The proliferation of false information in online news can mislead public perception, disrupt the order of information dissemination, and may even trigger social panic, damage public trust, and incite mass incidents or economic losses. Therefore, it is necessary to screen the dissemination of false information in online news.

[0003] For example, Chinese invention patent CN113032525A discloses a method, device, electronic device and storage medium for detecting fake news. The method obtains the text content, corresponding comment information and user information of the news to be detected, uses an encoding module to extract features from the text content and user comment information respectively, and then outputs the detection results through a joint attention module. The aim is to improve the detection accuracy through multi-source data fusion.

[0004] The existing technologies mentioned above have the following shortcomings: 1. The current focus is mainly on the analysis of static features such as the text content, user comments and user information of the news to be detected, without considering the dynamic features of news in the process of online dissemination. It is difficult to capture the typical risks of false information expanding its influence through rapid spread and abnormal dissemination paths, thus resulting in insufficient comprehensiveness of detection.

[0005] 2. The current detection logic does not distinguish between routine transmission and emergency transmission, and cannot initiate differentiated screening processes for different risk levels. At the same time, it does not select key assessment dimensions and dynamically set weight ratios based on historical detection data, resulting in insufficient effectiveness of assessment dimensions and difficulty in achieving targeted and rapid assessment based on the urgency of information dissemination, thus leading to insufficient screening efficiency and accuracy of risk response. Summary of the Invention

[0006] In view of this, in order to solve the problems mentioned in the background technology, a method for screening the spread of false information in online news is proposed.

[0007] The objective of this invention can be achieved through the following technical solution: This invention provides a method for screening the spread of false information in online news, including: S1, screening key evaluation dimensions based on historical monitoring data, determining the weight ratio of each evaluation dimension, and generating evaluation dimension rules accordingly.

[0008] S2. Obtain the increase in forwarding volume and text content of the information to be screened within a preset period, and determine the monitoring mode based on whether the increase in forwarding volume and text content match the number of words in the preset keyword library that exceed the preset threshold.

[0009] S3. Construct a propagation graph based on the propagation path of the information to be screened to identify the initial release source.

[0010] S4. Under the normal monitoring mode, the source authority and content consistency are assessed based on the source information of the information to be screened and the initial release source, as well as the text content. The dissemination anomaly assessment is conducted through the dissemination data of the information to be screened, and the comprehensive credibility is output in combination with the assessment dimension rules.

[0011] In S5, under the emergency monitoring mode, the information to be screened is rapidly assessed based on the evaluation dimension rules, and the urgency credibility is output.

[0012] S6. Match the overall credibility or emergency credibility with the credibility range corresponding to each warning level to obtain the warning level of the information to be screened.

[0013] Compared with the prior art, the beneficial effects of the present invention are as follows: (1) The present invention improves the effectiveness and pertinence of the evaluation dimensions by screening key evaluation dimensions based on historical monitoring data. At the same time, it determines the weight ratio by normalizing the evaluation consistency, realizes dynamic adaptation and data-driven weight allocation, ensures that the evaluation system is continuously optimized as historical monitoring data is updated, and thus enhances the adaptability to complex propagation scenarios.

[0014] (2) This invention triggers an emergency mode by using the dual conditions of the increase in forwarding volume exceeding the threshold and the number of keyword matching exceeding the threshold, thereby focusing on key evaluation dimensions to achieve rapid response. At the same time, the two modes can be switched on demand, ensuring the response speed in emergency scenarios while taking into account the needs of in-depth analysis, thus breaking through the technical contradiction that efficiency and accuracy cannot be achieved simultaneously in the traditional single mode.

[0015] (3) This invention assesses the authority of the source based on source information, determines the consistency of the content based on text content, and assesses the abnormality of dissemination based on dissemination data. Then, it outputs comprehensive credibility by combining the evaluation dimension rules. This achieves multi-dimensional collaborative evaluation of source, content, and dissemination, and provides a comprehensive and quantitative credibility basis for the accurate identification of false information.

[0016] (4) This invention filters key evaluation dimensions based on evaluation dimension rules, quantifies and evaluates key evaluation dimensions, and then calculates them by weighting their weight ratios to output emergency credibility. This achieves targeted and rapid evaluation under emergency monitoring mode, thereby ensuring evaluation efficiency. At the same time, it relies on rule-based processes to ensure the reliability and scenario adaptability of emergency credibility. Attached Figure Description

[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram showing the connections between the steps of the method of the present invention.

[0019] Figure 2 This is a schematic diagram showing the connection steps of the key evaluation dimension screening process in this invention.

[0020] Figure 3 This is a schematic diagram illustrating the steps involved in the authoritative assessment of the source of this invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Please see Figure 1 As shown, the present invention provides a method for screening the spread of false information in online news. The method includes: S1, screening key evaluation dimensions based on historical monitoring data, determining the weight ratio of each evaluation dimension, and generating evaluation dimension rules accordingly.

[0023] It should be added that the evaluation dimension rules include key evaluation dimensions and the weight percentage of each evaluation dimension.

[0024] Please see Figure 2 As shown, for example, the key evaluation dimensions for screening include: the final judgment results of false information for each historical dissemination information obtained from historical monitoring data and the judgment results of false information for each evaluation dimension.

[0025] The result of the false information determination is compared with the final result of the corresponding false information determination.

[0026] The number of times the false information judgment results of each evaluation dimension are consistent with the total number of times, and the ratio of the two is used as the evaluation consistency of each evaluation dimension. Evaluation dimensions with an evaluation consistency greater than the preset evaluation consistency threshold are selected as candidate evaluation dimensions. The evaluation consistency is the percentage of times the judgment result of the evaluation dimension is consistent with the final result.

[0027] It should be added that the preset evaluation consistency threshold is a critical value used to screen whether the evaluation dimension has the ability to effectively determine false information. That is, it is the minimum standard that the consistency between the false information determination result of the evaluation dimension and the final false information determination result must reach. The specific setting method is as follows: extract the historical evaluation consistency of each evaluation dimension from historical monitoring data, calculate the average of them, and use the calculation result as the preset evaluation consistency threshold.

[0028] The number of candidate evaluation dimensions is counted. If the number of candidate evaluation dimensions is 0, a second screening of key evaluation dimensions is triggered.

[0029] Furthermore, the triggering of secondary screening of key evaluation dimensions includes: combining each evaluation dimension in pairs, obtaining the false information judgment results of each evaluation dimension group from historical monitoring data, and comparing them with the corresponding false information final judgment results.

[0030] The number of times the false information judgment results of each evaluation dimension group are consistent is counted, and the ratio of this number to the total number of times is used as the evaluation consistency of each evaluation dimension group.

[0031] The evaluation consistency of the evaluation dimension group is compared with the preset evaluation consistency threshold. If there is an evaluation dimension group with an evaluation consistency greater than the preset evaluation consistency threshold, the evaluation dimension groups with an evaluation consistency greater than the preset evaluation consistency threshold are sorted from largest to smallest, and then the evaluation dimension in the evaluation dimension group with the highest ranking is selected as the key evaluation dimension.

[0032] If there is no evaluation dimension group with an evaluation consistency greater than the preset evaluation consistency threshold, then the evaluation consistency of each evaluation dimension is sorted from largest to smallest, and the evaluation dimension with the highest ranking is selected as the key evaluation dimension.

[0033] It should be added that by combining dimensions in pairs for evaluation and ranking based on evaluation consistency, effective evaluation dimensions can still be discovered even in extreme scenarios where there are 0 candidate evaluation dimensions. This avoids the screening mechanism from being paralyzed due to the failure of a single dimension. At the same time, by using the logic of prioritizing combination and downgrading as a fallback, it ensures that key evaluation dimensions with reference value can still be output when historical data is insufficient or the effectiveness of dimensions is insufficient. This improves the fault tolerance and adaptability of the evaluation system and provides a reliable foundation for subsequent weight allocation and credibility calculation.

[0034] If there are 3 candidate evaluation dimensions, sort the evaluation consistency of each evaluation dimension from largest to smallest, and select the top two evaluation dimensions as key evaluation dimensions.

[0035] For other candidate evaluation dimensions, evaluation dimensions with an evaluation consistency greater than the preset evaluation consistency threshold are selected as key evaluation dimensions.

[0036] For example, determining the weight percentage of each evaluation dimension includes summing the evaluation consistency of each evaluation dimension to obtain the total evaluation consistency.

[0037] Calculate the ratio of the consistency of each evaluation dimension to the overall consistency of the evaluation, and use it as the weight of the corresponding evaluation dimension. The sum of the weights of each evaluation dimension is 1.

[0038] S2. Obtain the increase in forwarding volume and text content of the information to be screened within a preset period, and determine the monitoring mode based on whether the increase in forwarding volume and text content match the number of words in the preset keyword library that exceed the preset threshold.

[0039] It should be added that the preset keyword library refers to a set of words that frequently appear in historical false information and are strongly associated with risky content. This includes typical characteristic words of false information, words from sensitive fields, and emotionally suggestive words, used to quickly identify the content risk characteristics of the information to be screened. The library is obtained by extracting all text content judged as false information from historical monitoring data, filtering the top 30% of words by frequency statistics and keyword extraction algorithms to form an initial candidate keyword library, and then having domain experts verify the candidate keyword library, removing words without risk association, and supplementing it with known risky words within the industry.

[0040] For example, determining the monitoring mode includes: calculating the number of matching words between the text content of the information to be screened and a preset keyword library, comparing the number of matching words with a preset matching word threshold, and comparing the increase in the number of forwards of the information to be screened within a preset period with a preset increase in the number of forwards threshold.

[0041] It should be added that the preset forwarding volume increase threshold is a critical value used to determine whether the spread speed of the information to be screened is abnormal. Specifically, it is the upper limit of the forwarding volume growth rate of the information to be screened within a preset period. Exceeding this threshold indicates that the information has the risk of explosive spread. The specific setting method is as follows: extract the forwarding volume increase data of similar information during the spread process from historical monitoring data, analyze the distribution characteristics of the forwarding volume increase during the explosive spread of false information in history and the normal spread increase range of true information, determine the critical value that can effectively distinguish between normal spread and explosive abnormal spread, and use it as the preset forwarding volume increase threshold.

[0042] If the increase in forwarding volume exceeds the preset threshold for the increase in forwarding volume, and the number of matched words is greater than the preset threshold for the number of matched words, then it is determined to be in emergency monitoring mode; otherwise, it is determined to be in regular monitoring mode.

[0043] It should be added that the preset matching word count threshold is a critical value used to determine the risk level of the information content to be screened. Specifically, it is the upper limit of the number of keywords in the text to be screened that match the preset keyword library. Exceeding this threshold indicates that the information content involves a high-risk area. It is obtained by: extracting all information identified as false information from historical monitoring data, counting the number of keyword matches for each piece of information, performing frequency distribution analysis on the number of matches, calculating the number of matches with a cumulative probability reaching a preset proportion, and using this as the preset matching word count threshold.

[0044] S3. Construct a propagation graph based on the propagation path of the information to be screened to identify the initial release source.

[0045] For example, identifying the initial publishing source includes: collecting node information of the information to be screened during the propagation process; based on the forwarding relationship of each account in the node information, forming a propagation path network from the source to subsequent propagation nodes with the publishing account as the node and the forwarding behavior as the directed edge; and marking the publishing time of each node. The node information includes the publishing account, the publishing time, and the forwarding relationship.

[0046] The earliest candidate starting node in the propagation path network is traced, and its timing logic is verified through the forwarding relationship to determine whether it is a propagation link without a preceding node and consistent with the timing logic of the downstream forwarding nodes. This ensures that it meets the conditions of no preceding node and no timing contradiction, and it is used as the initial starting point.

[0047] It should be added that the determination of the initial starting point includes: based on the directed edge structure of the propagation path network, verifying whether the candidate starting node has no preceding node, and checking whether the forwarding time of all first-level nodes that directly forward the content of the node is later than the publication time of the candidate starting node and whether the time difference conforms to the normal propagation delay law.

[0048] If a candidate starting node meets the conditions of having no preceding node and no conflict in downstream forwarding timing, it is confirmed as the initial starting point. If it does not meet the conditions, the node with the second earliest release time is selected as the new candidate starting node, and the above verification steps are repeated until an initial starting point that meets the conditions is determined.

[0049] S4. Under the normal monitoring mode, the source authority and content consistency are assessed based on the source information of the information to be screened and the initial release source, as well as the text content. The dissemination anomaly assessment is conducted through the dissemination data of the information to be screened, and the comprehensive credibility is output in combination with the assessment dimension rules.

[0050] Please see Figure 3 As shown, for example, the assessment of source authority includes: obtaining the organization type of the initial publishing source from the source information, matching the organization type with the organization type corresponding to each type weight, and obtaining the type weight of the initial publishing source.

[0051] It should be added that the institutional types corresponding to the weights of each type refer to a pre-defined institutional classification system that is associated with the historical occurrence rate of false information. Each category corresponds to a unique type weight, which is used to quantify the differences in the credibility of information sources of different institutions. The institutional types include, but are not limited to, government and public institutions, authoritative news media, official accounts of commercial enterprises, ordinary self-media, anonymous or unqualified accounts, etc.

[0052] The analysis and acquisition method for the institution types corresponding to the weights of each type is as follows: extract the attribute information of all information publishing entities from historical monitoring data, and cluster and classify them according to characteristics such as administrative attributes, professional qualifications, and operational norms to form an institution type system.

[0053] For each type of organization, the proportion of its historical published information that was judged to be false information was calculated, where the occurrence rate is the ratio of the number of false information to the total number of published information, thereby clarifying the risk differences of different types, such as the significantly higher occurrence rate of anonymous accounts compared to government departments.

[0054] The mapping logic of higher credibility gives higher weight is adopted. The credibility of the institution type is defined as 1 minus the occurrence rate. Then, the weight is calculated through normalization. Specifically, the weight of each type is the credibility of each institution type divided by the sum of the credibility of all institution types, ensuring that the sum of the weights of all institution types is 1, so that the weights can be directly used for subsequent scoring calculations.

[0055] The reason for assigning type weights based on the type of the initial publishing source is that, as the origin of information dissemination, the type of the initial publishing source directly relates to the original credibility of the information. Different types of institutions exhibit significant differences in their information publishing standards, professional qualifications, and historical rates of false information. Quantifying these differences through type weights transforms institutional credit characteristics into calculable evaluation indicators, providing an objective, data-driven basis for scoring the source's authority and ensuring that the evaluation results are directly linked to the risk level at the source.

[0056] The source key elements of the information to be screened are obtained from the source information, including but not limited to the name of the publishing entity, qualification certificate documents, information release time, contact information of the publishing entity, and information source traceability mark.

[0057] Verify the completeness status of the key source elements and calculate the source labeling completeness score based on their completeness percentage.

[0058] It should be added that the calculation of the source labeling integrity score includes: based on the total number of source key elements, such as the above-mentioned total number of source key elements being 5, the basic score of each element is determined, such as when the total score is 1 point, the basic score of each element is 0.2 points.

[0059] The completeness of key elements from each source is assessed: if the elements are presented in full, such as a clear name of the publishing entity and verifiable qualification documents, full marks are awarded; if some elements are missing, such as only the name of the entity is indicated without qualification documents, points are deducted according to the proportion of the missing elements; if all elements are missing, 0 marks are awarded.

[0060] Summarize the scores of all elements and calculate their ratio to the total score. Use this ratio as the source labeling integrity score.

[0061] The source authority score is generated by multiplying the type weight of the initial release source with the source labeling integrity score.

[0062] For example, the content consistency assessment includes: performing word segmentation on the text content of the information to be screened and the initial publication source, generating text feature vectors based on the term frequency-inverse document frequency algorithm, and calculating the cosine similarity value between the feature vectors as the overall semantic matching degree.

[0063] It should be added that the generated text feature vector includes: preprocessing the text content of the information to be screened and the initial publishing source separately, removing stop words without actual semantic meaning, and then using a word segmentation tool to decompose the text content into independent valid words, forming a vocabulary set for the two text segments. The vocabulary sets of the two text segments are merged and deduplicated to form a unified vocabulary list covering all unique words. The number of words in this vocabulary list is the dimension of the feature vector, with each word corresponding to one dimension of the vector. The standard TF-IDF algorithm is used to calculate the weight of each word. Finally, the TF-IDF weights of each word in the two text segments are arranged in vocabulary order to form their respective text feature vectors.

[0064] It should be added that the calculation of the cosine similarity value between feature vectors includes: the text feature vector of the information to be screened and the text feature vector of the original text are constructed based on the same vocabulary, ensuring that the two vectors have the same number of dimensions and that each dimension corresponds to the same words.

[0065] Multiply the TF-IDF weights of the two vectors one by one along their corresponding dimensions, and then sum all the product results to obtain the dot product of the two vectors. This value reflects the degree of overlap between the two vectors in each dimension.

[0066] The TF-IDF weights of each dimension of the two vectors are squared, summed, and the square root is taken to obtain the vector magnitudes. The magnitudes reflect the overall length of the vectors or the total size of the weights.

[0067] Dividing the dot product by the product of the magnitudes of the two vectors yields the cosine similarity value, which quantifies the overall semantic matching degree between the two texts.

[0068] The core keywords are extracted from the text content of the information to be screened and the initial release source, and the core keywords of the two are compared. The key information variation types are identified by combining the preset tampering rule base.

[0069] It should be added that core keywords refer to words extracted from the text content of the information to be screened and the initial release source that can directly reflect the core meaning, key facts or core conclusions of the text, including but not limited to key data, subject names, core qualitative statements and important limiting conditions.

[0070] The pre-defined rule base for tampering refers to the pre-defined criteria for judging various key information variations, covering specific characteristics of scenarios such as missing keywords, semantic conflicts, data tampering, and the addition of irrelevant key expressions. Among them, the absence of core data in the information to be screened is judged as missing keywords; the opposite of the core qualitative expression of the information to be screened from the original text is judged as semantic conflict; the modification of factual information such as specific values, time, and location in the information to be screened and inconsistent with the original text is judged as data tampering; and the addition of inflammatory, misleading expressions or irrelevant topic keywords unrelated to the core theme in the information to be screened is judged as the addition of irrelevant key expressions.

[0071] The identification of key information variation types includes: comparing the core keywords of the information to be screened with the core keywords of the original text word by word, matching the rules in the preset tampering rule base according to the comparison results, and determining the corresponding variation type if it meets the feature description of a certain type of variation.

[0072] The mutation types are matched with the mutation types corresponding to each mutation score, and then the corresponding scores are matched based on the identified mutation types to generate the key information variability.

[0073] It should be added that the variation types corresponding to each variation score refer to the categories divided according to different situations of variation in key information text, and each category is assigned a corresponding scoring standard. The specific setting method is as follows: Based on the analysis of false information cases and research on dissemination characteristics in historical monitoring data, the variation type of key information is determined, and the corresponding score is set according to the degree of damage to the authenticity of the information and the level of misleading risk caused by the variation type:

[0074] Variations that directly alter the core facts of the information, seriously affect its authenticity, and are likely to cause public misunderstanding are given a higher score; variation types that do not directly change the core facts but may interfere with the understanding of the information are given a medium score; and variation types that have no substantial impact on the core facts but have a potential tendency to mislead are given a lower score.

[0075] Based on the cosine similarity value and the variability of key information, a content consistency score is generated through weighted fusion calculation.

[0076] For example, the assessment of propagation anomalies includes: extracting the increase in the number of forwards of the information to be screened in each monitoring time period from the propagation data, extracting the maximum increase in the number of forwards from it, and calculating the ratio of it to a preset threshold for the increase in the number of forwards to obtain the propagation rate mutation rate.

[0077] Traverse all nodes in the propagation path network, identify the number of abnormal nodes and the total number of nodes, and use the ratio of the two as the propagation anomaly degree, where the propagation anomaly degree is the proportion of the number of abnormal nodes in the propagation path to the total number of nodes.

[0078] It should be added that the identification of the number of abnormal nodes includes: based on the node hierarchy relationship marked in the propagation path network, clarifying the theoretical forwarding level of each node. If a node does not forward according to the hierarchical progression, but directly skips at least one intermediate level to forward the content of a higher-level node, it is determined to be a hierarchical skipping abnormal node.

[0079] In the early stages of propagation, if a node suddenly participates in propagation without any associated nodes guiding it, and the forwarding volume of a single node exceeds the preset forwarding volume, and at the same time, the node has no association with the preceding nodes in the propagation path, it is determined to be an abnormal node that suddenly intervenes without association.

[0080] If an account is a newly registered, inactive zombie account, or has a history of spreading illegal content, or has engaged in batch operations of the same content in a short period of time, or has a fan activity level that is far lower than the average of similar accounts, it is determined to be an abnormal node in terms of account attributes. The number of abnormal nodes is then calculated by combining the abnormal nodes of hierarchical jump type, the abnormal nodes of unrelated sudden intervention type, and the abnormal nodes in terms of account attributes.

[0081] The propagation rate mutation rate is multiplied by the propagation anomaly degree, and the result is normalized to generate a propagation anomaly score.

[0082] It should be added that the normalization process includes: extracting the historical maximum and minimum values ​​of the product of the propagation rate mutation rate and the propagation anomaly degree of similar information from historical monitoring data, determining the historical value range of the product result, and then mapping the current product result to the [0,1] interval through a linear normalization formula to obtain the propagation anomaly score.

[0083] It should be added that the comprehensive credibility output by combining the evaluation dimension rules includes: obtaining the source authority score, the content consistency score, and the dissemination anomaly score.

[0084] Based on the weighting of each evaluation dimension, the scores of each evaluation dimension are weighted and integrated with their corresponding weightings to obtain the overall credibility.

[0085] In S5, under the emergency monitoring mode, the information to be screened is rapidly assessed based on the evaluation dimension rules, and the urgency credibility is output.

[0086] For example, the output urgency credibility includes: quantitatively evaluating the key evaluation dimensions according to the key evaluation dimensions in the evaluation dimension rules, and obtaining the score of each key evaluation dimension.

[0087] The weight percentage of key assessment dimensions is obtained from the weight percentage of each assessment dimension. The scores of each key assessment dimension are weighted and integrated with their corresponding weight percentages to obtain the urgency credibility.

[0088] S6. Match the overall credibility or emergency credibility with the credibility range corresponding to each warning level to obtain the warning level of the information to be screened.

[0089] It should be added that the credibility intervals corresponding to each warning level refer to the multiple warning levels divided according to the degree of risk of spreading false information, such as low risk, medium risk, and high risk.

[0090] Specifically, this interval is determined by the credibility distribution characteristics of previously identified false and true information in historical monitoring data: For information clearly identified as false in historical monitoring data, the concentrated distribution range of its overall credibility or urgent credibility is statistically analyzed, and this range is set as the interval corresponding to the high-risk warning level. For cases historically confirmed as true information, the concentrated distribution range of their credibility is statistically analyzed, and this range is set as the interval corresponding to the low-risk warning level. The intermediate interval is divided based on the natural transition characteristics of false information risk from low to high in historical data, combined with the degree of overlap in the credibility distribution of true and false information and the risk progression pattern, forming intervals corresponding to other warning levels such as medium risk, ensuring the matching of each interval with the actual risk level. When the overall credibility or urgent credibility of the information to be screened falls into a certain interval, it matches the warning level corresponding to that interval.

[0091] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, in the form of a computer program product.

[0092] Those skilled in the art will recognize that the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0093] In addition, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.

[0094] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0095] Finally, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for screening the spread of false information in online news, characterized in that: The method includes: S1. Select key evaluation dimensions based on historical monitoring data, determine the weight ratio of each evaluation dimension, and generate evaluation dimension rules accordingly. S2. Obtain the increase in forwarding volume and text content of the information to be screened within a preset period, and determine the monitoring mode based on whether the increase in forwarding volume and text content match the number of words in the preset keyword library exceeding the preset threshold. S3. Construct a propagation graph based on the propagation path of the information to be screened to identify the initial release source; S4. Under the normal monitoring mode, the source authority and content consistency are evaluated based on the source information of the information to be screened and the initial release source, as well as the text content. The dissemination anomaly is evaluated through the dissemination data of the information to be screened. Based on the weight ratio of each evaluation dimension, the scores of each evaluation dimension are weighted and integrated with the corresponding weight ratio to obtain the comprehensive credibility. S5. In emergency monitoring mode, based on the evaluation dimension rules, the information to be screened is evaluated quickly and targeted, and the urgency credibility is output. S6. Match the comprehensive credibility or emergency credibility with the credibility intervals corresponding to each warning level to obtain the warning level of the information to be screened. The determined monitoring mode includes: Calculate the number of matching words between the text content of the information to be screened and the preset keyword library, and compare the number of matching words with the preset matching word threshold. At the same time, compare the increase in the forwarding volume of the information to be screened within the preset period with the preset forwarding volume increase threshold. If the increase in forwarding volume exceeds the preset threshold for the increase in forwarding volume, and the number of matched words is greater than the preset threshold for the number of matched words, then it is determined to be in emergency monitoring mode; otherwise, it is determined to be in regular monitoring mode. The output urgency credibility includes: Based on the key evaluation dimensions in the evaluation dimension rules, each key evaluation dimension is quantitatively evaluated to obtain a score for each key evaluation dimension; The weight percentage of key assessment dimensions is obtained from the weight percentage of each assessment dimension. The scores of each key assessment dimension are weighted and integrated with their corresponding weight percentages to obtain the urgency credibility.

2. The method for screening the spread of false information in online news according to claim 1, characterized in that: The key evaluation dimensions for screening include: The final determination results of false information in each historical dissemination information and the determination results of false information in each assessment dimension are obtained from historical monitoring data. The result of the false information determination is compared with the final result of the corresponding false information determination; The number of times the false information judgment results are consistent across all evaluation dimensions and the total number of times are counted. The ratio of the two is used as the evaluation consistency of each evaluation dimension. Evaluation dimensions with an evaluation consistency greater than the preset evaluation consistency threshold are selected as candidate evaluation dimensions. The number of candidate evaluation dimensions is counted. If the number of candidate evaluation dimensions is 0, a second screening of key evaluation dimensions is triggered. If there are 3 candidate evaluation dimensions, sort the evaluation consistency of each evaluation dimension from largest to smallest, and select the top two evaluation dimensions as key evaluation dimensions. For other candidate evaluation dimensions, evaluation dimensions with an evaluation consistency greater than the preset evaluation consistency threshold are selected as key evaluation dimensions.

3. The method for screening the spread of false information in online news according to claim 2, characterized in that: The secondary screening of the key evaluation dimensions that triggers the process includes: The evaluation dimensions are combined in pairs, and the false information judgment results of each evaluation dimension group are obtained from historical monitoring data and compared with the corresponding false information final judgment results. The number of times the false information judgment results of each evaluation dimension group are consistent is counted, and the ratio of the number of times the result is consistent to the total number of times is used as the evaluation consistency of each evaluation dimension group. The evaluation consistency of the evaluation dimension group is compared with the preset evaluation consistency threshold. If there is an evaluation dimension group with an evaluation consistency greater than the preset evaluation consistency threshold, the evaluation dimension groups with an evaluation consistency greater than the preset evaluation consistency threshold are sorted from largest to smallest, and the evaluation dimension in the evaluation dimension group with the first position in the sort is selected as the key evaluation dimension. If there is no evaluation dimension group with an evaluation consistency greater than the preset evaluation consistency threshold, then the evaluation consistency of each evaluation dimension is sorted from largest to smallest, and the evaluation dimension with the highest ranking is selected as the key evaluation dimension.

4. The method for screening the spread of false information in online news according to claim 2, characterized in that: The determination of the weighting percentage for each evaluation dimension includes: The overall evaluation consistency is obtained by summing the evaluation consistency of each evaluation dimension. Calculate the ratio of the consistency of each evaluation dimension to the overall consistency of the evaluation, and use it as the weight of the corresponding evaluation dimension. The sum of the weights of each evaluation dimension is 1.

5. The method for screening the spread of false information in online news according to claim 1, characterized in that: The identification of the initial release source includes: Collect node information of the information to be screened during the dissemination process. Based on the forwarding relationship of each account in the node information, take the publishing account as the node and the forwarding behavior as the directed edge to form a propagation path network from the source to the subsequent propagation nodes, and mark the publishing time of each node. The earliest candidate starting node in the propagation path network is traced, and its timing logic is verified through the forwarding relationship to determine whether it is a propagation link without a preceding node and consistent with the timing logic of the downstream forwarding nodes. This ensures that it meets the conditions of no preceding node and no timing contradiction, and it is used as the initial starting point.

6. The method for screening the spread of false information in online news according to claim 1, characterized in that: The authority of the assessment sources includes: The organization type of the initial release source is obtained from the source information, and the organization type is matched with the organization type corresponding to each type weight to obtain the type weight of the initial release source. Obtain the key source elements of the information to be screened from the source information; Verify the completeness status of the key source elements and calculate the source labeling completeness score based on their completeness percentage. The source authority score is generated by multiplying the type weight of the initial release source with the source label integrity score.

7. The method for screening the spread of false information in online news according to claim 1, characterized in that: The content consistency includes: The text content of the information to be screened and the initial publication source are segmented into words, and text feature vectors are generated based on the term frequency-inverse document frequency algorithm. The cosine similarity value between the feature vectors is calculated as the overall semantic matching degree. The core keywords are extracted from the text content of the information to be screened and the initial release source, and the core keywords of the two are compared. The key information variation types are identified by combining the preset tampering rule base. The mutation types are matched with the mutation types corresponding to each mutation score, and then the corresponding scores are matched based on the identified mutation types to generate the key information variability. Based on the cosine similarity value and the variation of key information, a content consistency score is generated through weighted fusion calculation.

8. The method for screening the spread of false information in online news according to claim 1, characterized in that: The assessment of propagation anomalies includes: The increase in the forwarding volume of the information to be screened during each monitoring time period is extracted from the propagation data. The maximum increase in forwarding volume is then extracted and its ratio with a preset threshold for the increase in forwarding volume is calculated to obtain the propagation rate mutation rate. Traverse all nodes in the propagation path network, identify the number of abnormal nodes and the total number of nodes, and use the ratio of the two as the propagation anomaly degree. The propagation rate mutation rate is multiplied by the propagation anomaly degree, and the result is normalized to generate a propagation anomaly score.

Citation Information

Patent Citations

  • False news detection method and device, electronic equipment and storage medium

    CN113032525A

  • Multi-platform collaborative new media content monitoring management system based on big data

    CN113177164A

  • Public information issuing platform and method

    CN120256722A