Public health information authenticity prediction method and device based on key information tracing
By extracting strings of public health information and keyword phrases, combining official document and announcement retrieval and semantic similarity calculation, the problem of popularization of public health information is solved, and rapid and accurate identification of false information is achieved.
Patent Information
- Application Number
- CN202510313505.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-07-29
AI Technical Summary
The existing methods for predicting the authenticity of public health information are poorly popularized, cannot be widely promoted, and cannot effectively identify tampered public health information.
By obtaining public health information, extracting strings through text or picture processing, extracting key information phrases, and performing content search and matching in official documents and announcements, and calculating and judging the authenticity of the information based on semantic similarity.
It realizes rapid and accurate judgment of public health information, can identify tampered information, and the method is easy to promote and apply.
Smart Images

Figure CN120388727A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field related to the prediction of information authenticity, and particularly to a method and device for predicting the authenticity of public health information based on the traceability of key information. Background Art
[0002] Public health information refers to data and information related to population health, including data in aspects such as disease surveillance, health behaviors, and environmental factors.
[0003] If some lawbreakers or unscrupulous persons tamper with and release public health information, it is very likely to cause public panic. Therefore, it is necessary to predict the authenticity of public health information.
[0004] Most of the existing authenticity predictions of public health information are still manually compared with the relevant content on official websites. This method has poor popularity and cannot be promoted for public use. Summary of the Invention
[0005] The purpose of the present invention is to at least solve one of the deficiencies of the prior art, and provide a method for predicting the authenticity of public health information based on the traceability of key information.
[0006] To achieve the above purpose, the present invention adopts the following technical solutions:
[0007] Specifically, a method for predicting the authenticity of public health information based on the traceability of key information is proposed, including the following:
[0008] Obtain the public health information to be judged;
[0009] Extract the text of the public health information to be judged to obtain an ordered string;
[0010] Extract key information from the string to find the phrases corresponding to the key information;
[0011] Based on the phrases, perform content retrieval and matching in the corresponding official documents and announcements to find content-matching paragraphs. If no content-matching paragraphs are found, directly determine that the public health information to be judged is false information;
[0012] Calculate the similarity between the string and the content-matching paragraphs based on semantic matching. If the similarity between any content-matching paragraph and the string is higher than the similarity threshold, determine that the public health information to be judged is true information. If there is no similarity between any content-matching paragraph and the string that is higher than the similarity threshold, determine that the public health information to be judged is false information.
[0013] Further, specifically, extracting the text of the public health information to be judged to obtain an ordered string includes,
[0014] If the type of the public health information to be judged is text data, directly perform string extraction on it to obtain an ordered string;
[0015] If the type of the public health information to be judged is image data, perform string extraction through OCR recognition to obtain an ordered string.
[0016] Furthermore, specifically, performing string extraction on text data to obtain an ordered string includes,
[0017] Performing string extraction on text data based on the NLTK library of NLP to obtain an ordered string.
[0018] Furthermore, specifically, performing string extraction on image data through OCR recognition to obtain an ordered string includes,
[0019] Performing preprocessing on the image data, including denoising, binarization, and skew correction, to obtain a preprocessed image;
[0020] Obtaining a contour image from the preprocessed image through an edge detection algorithm;
[0021] Performing text region segmentation on the contour image to obtain multiple characters in order;
[0022] Extracting the feature information of each character and performing character content recognition through a pre-trained classifier to obtain a rough recognition result;
[0023] Correcting and optimizing the rough recognition result to obtain an ordered string.
[0024] Furthermore, specifically, extracting key information in the string to find the phrases corresponding to the key information includes,
[0025] Extracting key information such as location, person, date, and numerical indicators in the string through a preset word template, and the phrases formed by the above key information are the phrases corresponding to the key information.
[0026] Furthermore, specifically, calculating the similarity between the string and the content matching paragraph based on semantic matching includes,
[0027] Converting the entire string and the content matching paragraph into vector representations of a fixed length through a pre-trained sentence vector model Sentence-BERT, and then calculating the cosine similarity between these vectors. The obtained cosine similarity is the similarity between the string and the content matching paragraph.
[0028] The present invention also provides a public health information authenticity prediction device based on key information traceability, including the following:
[0029] A data acquisition module for acquiring the public health information to be judged;
[0030] A string extraction module for extracting text from the public health information to be judged to obtain an ordered string;
[0031] A key information extraction module for extracting key information from the string to find the corresponding phrases of the key information;
[0032] A content retrieval module for retrieving and matching the content in the corresponding official documents and announcements based on the phrases to find the content matching paragraphs. If no content matching paragraphs are found, it is directly determined that the public health information to be judged is false information;
[0033] A false information judgment module for calculating the similarity between the string and the content matching paragraphs based on semantic matching. If the similarity between any content matching paragraph and the string is higher than the similarity threshold, it is determined that the public health information to be judged is true information. If there is no similarity between any content matching paragraph and the string that is higher than the similarity threshold, it is determined that the public health information to be judged is false information.
[0034] The beneficial effects of the present invention are as follows:
[0035] Considering the issues of public use, the public health information authenticity prediction method based on key information traceability proposed by the present invention can directly obtain public health information in text or picture form, process these public health information separately to obtain the corresponding text strings, then extract key information from the text strings to find the corresponding phrases of the key information, and then perform content retrieval and matching of official documents and announcements through these phrases. If no matching content is found at this time, it can be directly judged as false information. When matching content is found, the similarity between the string and the matching content is calculated through semantic matching. If the similarity is not high, it is also considered false information. The public health information authenticity prediction method proposed by the present invention has high popularity and is easy to promote. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] By elaborating on the embodiments shown in conjunction with the drawings, the above and other features of the present disclosure will become more obvious. The same reference numerals in the drawings of the present disclosure represent the same or similar elements. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0037] Figure 1 The figure shows a flowchart of the public health information authenticity prediction method based on key information traceability of the present invention. Detailed implementation manners
[0038] The following will clearly and completely describe the concept, specific structure and technical effects generated by the present invention in combination with embodiments and drawings, so as to fully understand the purpose, solution and effects of the present invention. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The same reference numerals used throughout the drawings indicate the same or similar parts.
[0039] Embodiment 1, referring to Figure 1 , the present invention proposes a public health information authenticity prediction method based on key information traceability, including the following:
[0040] Step 110: Obtain the public health information to be judged;
[0041] Step 120: Extract the text of the public health information to be judged to obtain an ordered string;
[0042] Step 130: Extract key information from the string to find the phrases corresponding to the key information;
[0043] Step 140: Based on the phrases, perform content retrieval and matching in the corresponding official documents and announcements to find the content matching paragraphs (that is, if the phrase matches the corresponding word in the official document and announcement, the sentence where the word is located is recorded as the content matching paragraph). If no content matching paragraphs are found, directly determine that the public health information to be judged is false information;
[0044] Step 150: Calculate the similarity between the string and the content matching paragraphs based on semantic matching. If the similarity between any content matching paragraph and the string is higher than the similarity threshold, determine that the public health information to be judged is true information. If there is no similarity between any content matching paragraph and the string that is higher than the similarity threshold, determine that the public health information to be judged is false information.
[0045] In this Embodiment 1, considering the issue of public use, the method for predicting the authenticity of public health information based on the traceability of key information proposed by the present invention can directly obtain public health information in the form of text or pictures. The corresponding text strings are obtained by separately processing these public health information, and then key information extraction is carried out in the text strings to find the phrases corresponding to the key information. After that, content retrieval and matching of official documents and announcements are carried out through these phrases. If no matching content is found at this time, it can be directly determined as false information. When matching content is found, the similarity between the string and the matching content is calculated through semantic matching. If the similarity is not high, it is also considered false information. The method for predicting the authenticity of public health information of the present invention has a high popularity and is easy to promote.
[0046] As a preferred embodiment of the present invention, specifically, text extraction is performed on the public health information to be judged to obtain an ordered string, including
[0047] If the type of the public health information to be judged is text data, then string extraction is directly performed on it to obtain an ordered string;
[0048] If the type of the public health information to be judged is picture data, then string extraction is performed by means of OCR recognition to obtain an ordered string.
[0049] As a preferred embodiment of the present invention, specifically, string extraction is performed on text data to obtain an ordered string, including
[0050] String extraction is performed on text data based on the NLTK library of NLP to obtain an ordered string.
[0051] The following is an example of string recognition using the NLTK library
[0052] import nltk
[0053] from nltk.tokenize import word_tokenize
[0054] nltk.download('punkt')
[0055] text="Hello,my name is John Doe.I am a software engineer."
[0056] tokens=word_tokenize(text)
[0057] print(tokens)# Output: ['Hello', ',','my', 'name', 'is', 'John', 'Doe', '.', 'I', 'am', 'a','software',
[0058] 'engineer', '.']。
[0059] As a preferred embodiment of the present invention, specifically, string extraction is performed on the picture-type data by means of OCR recognition to obtain an ordered string, including,
[0060] Preprocessing the picture-type data by denoising, binarization, and skew correction to obtain a preprocessed image;
[0061] Obtaining a contour image from the preprocessed image through an edge detection algorithm;
[0062] Performing text region segmentation on the contour image to obtain a plurality of characters arranged in order;
[0063] Extracting the feature information of each character and performing character content recognition through a pre-trained classifier to obtain a rough recognition result;
[0064] Correcting and optimizing the rough recognition result to obtain an ordered string.
[0065] In this preferred embodiment, considering that the optical character recognition (OCR) technology is used to convert the text information in the image into editable and processable text data. Its main process includes the following steps:
[0066] Image preprocessing: Performing operations such as denoising, binarization, and skew correction on the input image to improve the image quality for subsequent processing.
[0067] Text region detection: Using image processing techniques (such as edge detection, contour analysis, etc.) to identify the regions in the image that may contain text.
[0068] Character segmentation: Segmenting the detected text region into individual characters to prepare for subsequent recognition.
[0069] Feature extraction and character recognition: Extracting the feature information of each character and using a classifier (such as a machine learning algorithm, a deep learning model, etc.) for recognition.
[0070] Post-processing: Correcting and optimizing the recognition result to improve the overall recognition accuracy.
[0071] As a preferred embodiment of the present invention, specifically, key information extraction is performed on the string to find the phrases corresponding to the key information, including,
[0072] The key information such as location, person, date, digital indicator, etc. in the string is extracted through a preset word template, and the phrases formed by the above key information are the phrases corresponding to the key information.
[0073] In this preferred embodiment, considering that the key information such as location, person, date, digital indicator, etc. in official document announcements are obvious features and are easy to detect once tampered with, so this is regarded as key information
[0074] As a preferred embodiment of the present invention, specifically, the similarity between the string and the content matching paragraph is calculated based on the semantic matching method, including,
[0075] The entire string and the content matching paragraph are converted into vector representations of a fixed length through a pre-trained sentence vector model Sentence-BERT, and then the cosine similarity between these vectors is calculated. The obtained cosine similarity is the similarity between the string and the content matching paragraph.
[0076] In this preferred embodiment, a pre-trained sentence vector model (such as BERT, Sentence-BERT) is used to convert the entire string and the paragraph into vector representations of a fixed length. Then, the cosine similarity between these vectors is calculated. This method can better capture the semantic information and context relationship of the sentences.
[0077] Example 2, the present invention also proposes a public health information authenticity prediction device based on key information traceability, including the following:
[0078] A data acquisition module for acquiring the public health information to be judged;
[0079] A string extraction module for performing text extraction on the public health information to be judged to obtain an ordered string;
[0080] A key information extraction module for performing key information extraction on the string to find the phrases corresponding to the key information;
[0081] A content retrieval module for performing content retrieval and matching in the corresponding official document announcements based on the phrases to find the content matching paragraphs. If no content matching paragraphs are found, it is directly determined that the public health information to be judged is false information;
[0082] The false message judgment module is used to calculate the similarity between the string and the content matching paragraph based on semantic matching. If the similarity between any content matching paragraph and the string is higher than the similarity threshold, it is determined that the public health information to be judged is a true message. If there is no similarity between any content matching paragraph and the string that is higher than the similarity threshold, it is determined that the public health information to be judged is a false message.
[0083] In this Embodiment 2, it is consistent with the public health information authenticity prediction method based on key information traceability proposed by the present invention. Considering the problem of public use, the public health information authenticity prediction method based on key information traceability proposed by the present invention can directly obtain public health information in text or picture form, process these public health information separately to obtain corresponding text strings, and then extract key information in the text strings to find the phrases corresponding to the key information. After that, the content of official documents and announcements is retrieved and matched through these phrases. If no matching content is found at this time, it can be directly judged as a false message. When matching content is found, the similarity between the string and the matching content is calculated through semantic matching. If the similarity is not high, it is also considered a false message. The authenticity prediction method of the public health information of the present invention has a high popularity and is easy to promote.
[0084] Although the description of the present invention has been quite detailed and especially describes several of the described embodiments, it is not intended to be limited to any of these details or embodiments or any particular embodiment, but should be regarded as providing a broad possible interpretation of these claims in light of the prior art by reference to the appended claims, so as to effectively cover the intended scope of the present invention. In addition, the present invention is described above with embodiments foreseeable by the inventor for the purpose of providing a useful description, and those non-substantive modifications to the present invention that are not currently foreseeable may still represent equivalent modifications of the present invention.
[0085] As mentioned above, it is only a preferred embodiment of the present invention. The present invention is not limited to the above-mentioned implementation manners. As long as it achieves the technical effects of the present invention by the same means, it should fall within the protection scope of the present invention. Within the protection scope of the present invention, various different modifications and changes can be made to its technical solutions and / or implementation manners.
Claims
1. A method for predicting the authenticity of public health information based on tracing key information, characterized in that including the following: Obtain the public health information to be judged; Perform text extraction on the public health information to be judged to obtain an ordered string; Extract key information from the string to find the phrases corresponding to the key information; Based on the phrases, perform content retrieval and matching in the corresponding official documents and announcements to find the content matching paragraphs. If no content matching paragraphs are found, directly determine that the public health information to be judged is false information; Calculate the similarity between the string and the content matching paragraphs based on semantic matching. If the similarity between any content matching paragraph and the string is higher than the similarity threshold, determine that the public health information to be judged is true information. If there is no similarity between any content matching paragraph and the string that is higher than the similarity threshold, determine that the public health information to be judged is false information.
2. The method for predicting the authenticity of public health information based on the traceability of key information according to claim 1, wherein Specifically, performing text extraction on the public health information to be judged to obtain an ordered string includes If the type of the public health information to be judged is text data, directly perform string extraction on it to obtain an ordered string; If the type of the public health information to be judged is image data, perform string extraction through OCR recognition to obtain an ordered string.
3. The method for predicting the authenticity of public health information based on tracing key information according to claim 2, wherein Specifically, performing string extraction on text data to obtain an ordered string includes Perform string extraction on text data based on the NLTK library of NLP to obtain an ordered string.
4. The method for predicting the authenticity of public health information based on tracing key information according to claim 1, characterized in that, Specifically, performing string extraction on image data through OCR recognition to obtain an ordered string includes Perform preprocessing of denoising, binarization, and skew correction on the image data to obtain a preprocessed image; Obtain a contour image from the preprocessed image through an edge detection algorithm; Perform text region segmentation on the contour image to obtain a plurality of characters arranged in order; Extract the feature information of each character, and perform character content recognition through a pre-trained classifier to obtain a rough recognition result; Correct and optimize the rough recognition result to obtain an ordered string.
5. The method for predicting the authenticity of public health information based on tracing key information according to claim 4, wherein, Specifically, extracting key information from the string to find the phrases corresponding to the key information includes Extract key information such as location, person, date, numerical indicators, etc. from the string through a preset word template, and the phrases formed by the above key information are the phrases corresponding to the key information.
6. The method for predicting the authenticity of public health information based on tracing key information according to claim 1, wherein, Specifically, calculating the similarity between the string and the content matching paragraphs based on semantic matching includes Convert the entire string and the content matching paragraphs into fixed-length vector representations through the pre-trained sentence vector model Sentence-BERT, and then calculate the cosine similarity between these vectors. The obtained cosine similarity is the similarity between the string and the content matching paragraphs.
7. A public health information authenticity prediction device based on tracing key information, characterized in that, including the following: A data acquisition module for obtaining the public health information to be judged; A string extraction module for performing text extraction on the public health information to be judged to obtain an ordered string; A key information extraction module, which is used to extract key information from the string and find out the phrases corresponding to the key information; A content retrieval module, which is used to perform content retrieval and matching in the corresponding official documents and announcements based on the phrases, and find out the content matching paragraphs. If no content matching paragraphs are found, it is directly determined that the public health information to be judged is false information; A false information judgment module, which is used to calculate the similarity between the string and the content matching paragraphs based on semantic matching. If the similarity between any content matching paragraph and the string is higher than the similarity threshold, it is determined that the public health information to be judged is true information. If there is no similarity between any content matching paragraph and the string that is higher than the similarity threshold, it is determined that the public health information to be judged is false information.