Ffact verification system based on large model and semantic implication analysis

By combining large language model and web search technology, using the BERT model and HITS algorithm to evaluate sentence similarity and logical relationships, the problems of insufficient depth and timeliness in professional fields in the existing technology are solved, and the factual verification of high accuracy and timeliness is achieved.

CN120373459APending Publication Date: 2025-07-25CENTRAL UNIVERSITY OF FINANCE AND ECONOMICS +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510458042.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing fact verification technology is insufficient in handling complex expression verification in professional fields, and its timeliness and applicability are insufficient, making it difficult to meet the timeliness requirements for risk prevention and control.

Method used

Combining the natural language reasoning ability of large language models and web page retrieval technology, query strings are generated through the input module, sentence similarity is evaluated using the BERT model, and the HITS algorithm calculates the authority of the web page, and logical relationship judgment is performed through the dual encoder to generate trusted fact verification results.

Benefits of technology

It has achieved high accuracy verification of complex expressions in professional fields, has good timeliness and applicability, and can detect and intervene in the early stages of false information transmission, making up for the insufficient timeliness of the propagation path method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373459A_ABST
    Figure CN120373459A_ABST
Patent Text Reader

Abstract

The invention relates to the field of network information fact verification research, in particular to a fact verification system relating to webpage content matching and natural language reasoning, which extracts key information from user input, retrieves and screens webpage content by combining the natural language reasoning capability of a large language model and a webpage retrieval enhancement technology, and finally verifies the webpage content. According to the method, webpage authority and semantic similarity are calculated, finally, the reasoning result of the large model and general knowledge judgment are synthesized, credible and related output content is generated, complex expression verification of the professional field can be met, and timeliness is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the research field of online information fact-checking, and is a fact-checking method and system involving web page content matching and natural language reasoning. Background Art

[0002] Fact-checking technology plays a crucial role in maintaining social information order, safeguarding the public's right to know, and supporting scientific decision-making. In the digital age of information explosion, unverified statements may lead to group cognitive biases. Especially in the field of public health, false health claims may cause the public to ignore professional medical advice; in the financial market, false rumors may induce irrational investment behaviors and cause abnormal asset fluctuations. A timely and accurate fact-checking mechanism can effectively block the spread chain of false information, provide a reliable basis for public decision-making, and is of great value for building a healthy information ecosystem.

[0003] The current fact-checking technology system mainly includes three types of methodologies: content feature-based analysis technology, authoritative knowledge base-based verification technology, and dissemination trend-based monitoring technology. Content feature-based technology realizes preliminary screening by parsing the grammatical structure, semantic features, and sentiment polarity of text, and has the advantage of real-time processing of a large amount of information flow; knowledge base-based technology relies on professional databases constructed by authoritative institutions (such as the WHO medical knowledge graph and the central bank financial database), and realizes high-precision fact determination through cross-verification of multi-source data; dissemination trend-based technology can effectively identify abnormal dissemination patterns and locate key dissemination nodes by analyzing features such as information diffusion rate and dissemination network topology structure.

[0004] However, the existing technical solutions still have significant limitations: Although content feature-based methods can quickly process text data, they lack a deep understanding of implicit semantic logic and context associations, and are difficult to handle the verification of complex expressions in professional fields; Although knowledge base-based technology has high accuracy in specific vertical fields, its effectiveness highly depends on the completeness of the knowledge base, there are knowledge blind spots in emerging fields or interdisciplinary scenarios, and it faces the timeliness challenge brought by lagging data updates; Although dissemination trend-based technology can effectively identify suspicious information with large-scale dissemination, there is an obvious lag in detection response, usually triggering warnings only after abnormal dissemination forms a scale, and it is difficult to meet the timeliness requirements of risk prevention and control. Summary of the Invention

[0005] To at least solve some of the above problems existing in the prior art, the technical object of the present disclosure is to propose a fact-checking system involving web content matching and natural language reasoning. By combining the natural language reasoning ability of large language models and web retrieval enhancement technology, this system extracts key information from user input, retrieves and filters web content, calculates web authority and semantic similarity, and finally generates credible and relevant output content by integrating the inference results of large models and general knowledge judgment. It can not only meet the verification of complex expressions in professional fields but also has timeliness.

[0006] To achieve the above technical object, a fact-checking system proposed by the present disclosure includes an input module, a query retrieval module, a first calculation module, a second calculation module, and a third calculation module; wherein: the input module is configured to obtain a corresponding query string based on the image or text of the user query information; the query retrieval module is configured to obtain web pages in text form related to the query string; the first calculation module is configured to perform sentence splitting on the obtained web pages to obtain a sentence sequence, and then evaluate the similarity between each sentence and the query string; the second calculation module is configured to calculate the authority score of the reference web pages using the HITS algorithm for the top K% of the reference web pages with the highest similarity; the third calculation module is configured to use the user query information as a hypothesis sentence and the reference web page content as a premise sentence, and calculate the logical relationship between the hypothesis sentence and the premise sentence as "support", "irrelevant", or "oppose" through conditional probability calculation as the inference result, and obtain the support degree corresponding to the inference result. Then, based on the authority score of each web page and the support degree corresponding to each web page, the credibility of the fact-checking result is obtained.

[0007] In an implementation manner of the above technical solution, the system further includes a knowledge base, and the knowledge base is configured to include a first data table. The data stored in the first data table includes the user query information and its corresponding inference result, as well as relevant information of the web page set that meets the conditions. The relevant information includes the web page title, the web page body content, and the sentence most matching the query information. The web page set that meets the conditions belongs to the top K% of the web page sets with the highest similarity, and the product of the authority score of the web page in the web page set that meets the conditions and the support degree corresponding to the web page is greater than a set threshold.

[0008] In an implementation manner of the above technical solution, the knowledge base is further configured to include a second data table, and the second data table stores other web page information that is not stored in the first data table in the set W, including the web page title and the web page body content.

[0009] In one implementation of the above technical solution, the system further includes a result output module, which is configured to output the credibility of the fact-checking result, the source of the web page inference, and the general result supplement. The source of the web page inference is the web page text content of the reference web page and the sentence that best matches the query information. The general result supplement is the result of using a large language model to judge the query string based on the general knowledge it stores.

[0010] In one implementation of the above technical solution, the input module obtains the query string by: performing structured semantic parsing on the image content to obtain the descriptive text; for the text content, using a word segmentation tool to extract keywords, and then generating the query string through semantic focusing.

[0011] In one implementation of the above technical solution, the query retrieval module obtains the web page by: constructing a distributed web crawler engine based on the query string, extracting the title, source, URL, and text content of each search result from the crawled web pages, skipping the web pages that lack text content and have videos, removing the tables and pictures in the remaining web pages, and and Extract the text content from the tags and merge it into the final web page text.

[0012] In one implementation of the above technical solution, the first computing module evaluates the similarity between each sentence and the query string as follows: Based on n sentences in the web page, obtain the sentence sequence S = {s1, s2,..., s n}; Use the BERT model to convert the query string q and each sentence s i into semantic vectors q and s with a vector dimension of d respectively, that is i , namely Calculate the cosine similarity between the vector q and each sentence vector s i to evaluate the similarity between each sentence and the query string.

[0013] In one implementation of the above technical solution, the third computing module obtains the credibility of the fact-checking result as follows: Take the query string as the hypothesis sentence s1, and take the content of the reference web page as the sentence s2 with a deterministic premise. Use a dual encoder to encode s1 and s2 to obtain semantic representations φ hypo (s1) and φ prem (s2); Generate a relationship representation h θ (s1, s2) = W · [φ hypo (s1); φ prem (s2); φ hypo (s1) ⊙ φ prem (s2)] + b, where and are learnable parameters, and ⊙ represents element-wise multiplication; Calculate the conditional probability and determine the logical relationship of "support", "irrelevant", or "oppose" between the hypothesis and the premise by maximizing the conditional probability as the reasoning result. Then, obtain the support degree of the reference web page according to the support degree set by the reasoning result; Calculate the weight of each web page according to the authority score of each web page, and use the weight to weight the support degrees corresponding to all web pages to obtain the credibility of the fact-checking result.

[0014] To achieve the above technical purpose, a computer-readable storage medium proposed by the present disclosure stores a computer program that can be loaded and executed by a processor for any of the above systems.

[0015] To achieve the above technical purpose, a fact-checking method proposed by the present disclosure includes:

[0016] Step 1: Generate a unified query string q based on the image or text of the user query information;

[0017] Step 2: Obtain the web pages in text form related to the query string q;

[0018] Step 3: Sentence processing is performed on each acquired web page to obtain a sentence sequence, and then the similarity between each sentence and the query string is evaluated;

[0019] Step 4: The web pages ranked in the top K% of similarity are used as reference web pages, and the authority scores of the reference web pages are calculated using the HITS algorithm;

[0020] Step 5: Take the user query information as the hypothesis sentence and the webpage content as the premise sentence, and determine the logical relationship between the hypothesis sentence and the premise sentence as "support", "irrelevant" or "oppose" through conditional probability calculation as the inference result;

[0021] Step 6: Obtain the support corresponding to the inference result, and obtain the credibility of the fact verification result based on the authority score of each web page and the support corresponding to each web page.

[0022] Beneficial technical effects:

[0023] The disclosed technical solution can delve into the semantics and contextual relationships of the text. At the same time, the system uses the general knowledge of the large language model to solve the problem of knowledge-based methods relying on domain experts or knowledge bases, and avoids the limitations of cold start. In addition, the system searches and filters web content in real time through user input, and combined with the reasoning ability of the large language model, it can assist users in completing fact verification, detect and intervene in the early stages of the spread of false information, and make up for the shortcomings of the propagation path-based method in terms of timeliness. The disclosed technical solution is used for fact verification with high accuracy, good timeliness and applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0025] Figure 1 , one A structural diagram of a fact verification system in an implementation manner.

[0026] Figure 2 , one Schematic diagram of the query retrieval module processing flow in this implementation manner.

[0027] Figure 3 , one Schematic diagram of the processing flow of the first computing module in this implementation manner. DETAILED DESCRIPTION

[0028] The following will clearly and completely describe how to implement the technical solution of this case in conjunction with the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this case, rather than all the embodiments. Based on the embodiments in this case, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope protected by this application.

[0029] As Figure 1 shown, a fact verification system, the system includes an input module, a query retrieval module, a first calculation module, a second calculation module, a third calculation module, and a result output module. The specific module introduction is as follows.

[0030] The input module is configured to accept text or image input provided by the user, establish a unified multimodal representation framework, including performing structured semantic parsing on the image, extracting keywords from the text, and generating a unified query string. The input module is used to complete multimodal input parsing and semantic focusing.

[0031] Specifically, for the image content, the DeepSeek-V3 lightweight version is used to perform structured semantic parsing on the picture and convert the picture into a descriptive text. For the text content, the jieba word segmentation tool is used to extract keywords. The system synthesizes this information to focus the semantics to generate a more accurate search string and obtain an accurate user query string q.

[0032] The query retrieval module is configured to build a distributed web crawler engine according to the query string, obtain web content through an asynchronous IO architecture, extract text using an HTML parsing tool, and filter pages containing videos, pictures or tables, see Figure 2 。

[0033] Specifically, a distributed web crawler engine is built based on the query string q, and an asynchronous IO architecture is used to improve the data acquisition efficiency. The requests library is used to send HTTP requests to obtain web content, BeautifulSoup is used to parse HTML and extract text, and chardet is used to detect the encoding to avoid garbled problems. The title, source, URL and content of each result are also extracted. To ensure that the extracted content is in text form and relevant to the user's query, the system will screen the web content and filter out irrelevant content (such as videos, tables and pictures). The system first detects and skips pages containing videos (such as <video>page with a label), and remove the table in the web page( and Labels) and images (labels) to avoid dealing with structured data and non-text information. Then from The text content is extracted from the tag and merged into the final web page text. If the web page lacks valid text, it is marked as "no text content found" and the web page is skipped.

[0034] The query generation and web page retrieval module ensures that only information that can be used for text analysis is retained.

[0035] The first computing module is configured to use a pre-trained lightweight BERT model to convert the user query q and each sentence in the web page content into a semantic vector with a dimension of d, and determine the sentence in the web page that best matches the query and the overall relevance score by calculating the cosine similarity, take the highest similarity as the overall similarity score of the web page, and sort the web pages by score.

[0036] Specifically, see Figure 3 , perform sentence processing on the crawled web page content and obtain the sentence sequence S = {s1, s2, ..., s n }, where n is the total number of sentences contained in the webpage content. The BERT model is used to combine the user query q and each sentence s i Converted into semantic vectors q and s with vector dimension d respectively i ,Right now Calculate the user query vector q and each sentence vector s i The cosine similarity of is used to evaluate the semantic relevance of each sentence to the user query. The calculation formula is:

[0037]

[0038] Among them, q·s i Represents the dot product operation of two vectors, ||q|| and ||s i || respectively represent the Euclidean norm of the two vectors. Finally, the highest similarity score among all sentences is selected as the overall similarity score of the web page content, and the web page content is arranged in descending order from high to low according to the similarity score.

[0039] score(S)=max sim(q, s i )

[0040] The second calculation module is configured to obtain the web pages ranked in the top K% of the similarity ranking results and calculate the authority scores of the web pages.

[0041] Specifically, firstly, for the top-ranked web pages after screening, the authoritative navigation pages and their external links are given initial weights. Let A = {p1, p2, p3, ..., p m } is the set of authoritative navigation pages, m is the number of authoritative navigation pages, including all navigation pages that are recognized as authoritative. i ∈A, let L(p i ) is the set of all external links in the authoritative navigation page p i , which is defined as:

[0042]

[0043] Assign a weight of 1 to the web pages on the authoritative navigation page and the external links in the authoritative pages. That is, at the initial stage, it is considered that the web pages in the set L4∪A are highly authoritative.

[0044] Secondly, for the pages not in the set L A ∪A, classify them as non-authoritative pages, and use the HITS algorithm to iteratively calculate the Authority value and Hub value of the web pages to measure their authority:

[0045] For each page p in the authoritative navigation page and its external link set A∪L A , its initial Authority value and Hub value are both set to 1. Construct a directed graph containing all web pages and their external links, with web pages as nodes and external links between web pages as directed edges. For each page p, its external link set L(p) contains links pointing to other pages, and page p may be linked by other web pages to form backlinks. Use the HITS algorithm for iterative calculation to update the Authority value and Hub value of each page respectively.

[0046] The update of the Authority value and Hub value of each page includes:

[0047] 4.1 Iterative calculation: For each page p, its Authority value is the sum of the Hub values of all web pages pointing to this page; for each page p, its Hub value is the sum of the Authority values of the pages pointed to by all external links in page p, that is:

[0048]

[0049] 4.2 Normalization processing: After each iteration, perform L2 normalization on the Authority value and Hub value to ensure the stability of the calculation results.

[0050]

[0051]

[0052] 4.3 Convergence judgment: After multiple iterations, judge whether the Authority value and Hub value converge. If the difference between the results of two iterations is less than the preset threshold, it is considered that the calculation results have converged, the iteration stops, and the Authority value of each page is output.

[0053] The Authority value of the relevant page will participate in the weighted calculation of credibility in the NLI and credibility calculation module.

[0054] The third computing module is configured to include a natural language reasoning unit and a credibility computing unit.

[0055] The natural language inference (NLI) unit is configured to use user query information as a hypothesis sentence and web page content as a premise sentence, adopt a dual encoder architecture and interactive operations to generate relationship representations, and calculate conditional probabilities based on softmax to determine whether the hypothesis and the premise have a logical relationship of "support", "irrelevant" or "opposition", and assign support values to different logical relationships.

[0056] The dual encoder structure semantically encodes the hypothesis sentence (query string) and the premise sentence (web page content) to obtain corresponding semantic representations.

[0057] The above interactive operations include vector dot products and element-by-element multiplications.

[0058] Specifically, given two sentences s1 and s2, where s1 is a hypothesis sentence and s2 is a premise sentence, the natural language inference task (NLI) determines the relationship between the two sentences r∈{Support,Contradiction,Neutral} by judging the logic between the hypothesis and the premise, and is classified as "support", "contradiction", or "irrelevant". That is:

[0059] f NLI (s1,s2)→{Support,Contradiction,Neutral}

[0060] The user query q is used as the hypothesis s1 to be verified, and the top-ranked web page content results (content(P)) in step 3 are used as known and deterministic premises s2. The system proposes a framework based on probability and semantic coding.

[0061] First, the hypothesis sentence and premise sentence are semantically encoded through the dual encoder architecture to obtain their semantic representations:

[0062]

[0063] Generate relational representations through interactive operations:

[0064] h θ (s1, s2) = W·[φ hypo (s1);φ prem (s2); φ hypo (s1)⊙φ prem (s2)]+b

[0065] Among them and are learnable parameters, and ⊙ represents element-wise multiplication. The conditional probability is calculated based on the softmax function:

[0066]

[0067] The final inference result is determined by maximizing the conditional probability:

[0068]

[0069] When the results are "Support", "Neutral", and "Contradiction" respectively, the support degree values of the inference results are set to That is:

[0070]

[0071] The credibility calculation unit is configured to use the web page authority score Auth(p) as a weight to weight the support degree values of the inference results of the natural language inference of each web page content for weighted calculation, and the calculation result is used as the credibility of the query. Specifically, the weighted average method is used to calculate the overall fact-checking credibility, and the formula is as follows, to obtain the final fact-checking credibility score Cred(q):

[0072]

[0073] The above system can further build a knowledge base.

[0074] The knowledge base is configured to screen, sort the search results according to the calculated similarity scores, and filter the relevant web page content based on a preset threshold K. For the top K% web page set W before sorting, store its relevant information in the database. Among them, the stored data includes but is not limited to: user query information query(q), inference result result(p) of the natural language inference (NLI) task, web page title title(p), web page text content content(p), and the sentence best_match(p) that best matches the query information.

[0075] Specifically, for the web pages to be stored in the database, a judgment value is obtained based on the absolute value of the product of the web page authority and the inference result support degree calculated in the above credibility calculation unit, and a threshold L (L < 1) is set. If the judgment value is greater than L, it is considered that the web page has a more significant effect on the support of the inference result, and the relevant data of these web pages are stored in the first data table. The data stored in the first data table includes user query information query(q), inference result result(p) of the natural language inference (NLI) task, web page title title(p), web page body content content(p), and the sentence best_match(p) that best matches the query information.

[0076] In addition, for the other web pages in the set W that are not stored in the first data table, their web page titles title(p) and web page body contents content(p) will be stored in the second data table.

[0077] In the subsequent application process, the first data table can be used to quickly match user queries to retrieve the content for which inference judgments have been completed, thereby improving the inference efficiency and reducing the consumption of computing resources. The web page content in the second data table can be used as additional supporting evidence to avoid repeated web page crawling and improve the system response speed.

[0078] Taking the constructed knowledge base as the knowledge base of the large language model, in subsequent use, user queries can be input into the large language model, and the large language model uses the general knowledge in the knowledge base to make judgments, generating judgment results and judgment reasons with limited word count, forming a complementary judgment basis with the web page content inference results.

[0079] The result output module is configured such that the result presented to the user is supported by two parts. The first part is the credibility calculation result obtained through the inference task on the web page content, and the second part is the general knowledge judgment result of the large language model. Finally, the system presents to the user:

[0080] The final credibility score Cred(q), which intuitively quantifies and examines the credibility of the query statement;

[0081] The source of web page inference, including the reference web pages and their matching sentences involved in the calculation;

[0082] General knowledge supplement, that is, the knowledge-based explanation of the large language model, including the general knowledge judgment result generated by the large language model, to provide additional judgment basis.

[0083] The system intuitively displays the final output result to the user in graphical and text forms, facilitating the user to understand the credibility of the query statement.

[0084] In summary, the system of the present disclosure can deeply understand the semantics and context relationships of text. At the same time, the system utilizes the general knowledge of the large language model to solve the problem of dependence on domain experts or knowledge bases in knowledge-based methods and avoid the limitations of cold start. In addition, the system can search and filter web content in real time based on user input, and combine the reasoning ability of the large language model to assist users in fact-checking, detect and intervene in the early stage of the spread of false information, making up for the lack of timeliness in the method based on the propagation path. Moreover, the system is superior to traditional methods in terms of accuracy, timeliness and applicability, providing a new solution for fact-checking.

[0085] Through the description of the above embodiments, those skilled in the art can clearly understand that based on the above system, a fact-checking method based on a large model and semantic entailment analysis can be obtained, and the steps include:

[0086] Step 1: Receive the text or image data input by the user, perform structured semantic parsing on the image, extract keywords from the text, and generate a unified query string q.

[0087] Step 2: Build a distributed web crawling engine based on the query string q, obtain web content through an asynchronous IO architecture, extract text using an HTML parsing tool, and filter pages containing videos, pictures or tables to obtain web pages in text form related to the query string.

[0088] Step 3: Perform sentence splitting on each obtained web page to obtain a sentence set S = {s1, s2,..., s n}(n is the total number of sentences contained in the web page), and use the pre-trained BERT model to convert the user query q and each sentence s i into semantic vectors with a dimension of d, and calculate the cosine similarity:

[0089]

[0090] For each web page, select the highest similarity score as the overall similarity score of the web page, and set the corresponding sentence as the best matching sentence best_match(p) of the web page.

[0091] Step 4: For the top K% of the web pages, assign initial weights to the authoritative navigation pages and their external links, then construct a directed graph of web page links, and iteratively calculate the Authority value and Hub value of each web page through the HITS algorithm to measure the authority of the web page. Among them, the Hub value is the sum of the Authority values of the pages pointed to by all external links in page p, and the Authority value is the sum of the Hub values of all web pages pointing to page p. The Authority value is an authoritative score for page p and can be denoted as Auth(p).

[0092] Step 5: Use the user query information as the hypothesis sentence and the web page content as the premise sentence, adopt a dual-encoder architecture and interactive operations to generate a relationship representation, and calculate the conditional probability based on softmax to judge the logical relationship between the hypothesis and the premise as "support", "irrelevant", or "oppose".

[0093] Step 6: According to the authoritative score of each page p and the support degree of the inference result in the natural language inference result, calculate the overall fact-checking credibility by using the weighted average method.

[0094] Step 7: Build a knowledge base. In the knowledge base, for the set of web pages W with the highest similarity, calculate the product of the web page authority and the support degree of the inference result. The web pages whose absolute value exceeds the specified threshold will have relevant information. The relevant information includes the user query, web page title, body content, natural language inference result, and the best matching sentence. Store the relevant information in the first data table, and at the same time store the web page content in the set W that has not been stored in the first data table in the second data table.

[0095] So far, those skilled in the art can clearly know that Steps 1 to 7 implement a method for building a knowledge base.

[0096] Step 8: Input the user query into the large language model to generate an answer with a limited number of words and the reasons for judgment. The result output by the large model is used as a general knowledge judgment to supplement the web page inference result and provide additional reference basis for users.

[0097] Taking the user query "Will frozen steamed buns grow Aspergillus flavus?" as an example, describe the specific implementation of the above method steps.

[0098] In Step 10, for the image content, use the lightweight version of DeepSeek-V3 to perform structured semantic parsing on the picture and convert the image into text; for the text content, use the jieba word segmentation tool to extract keywords. For example, the keywords extracted in this embodiment are: "frozen", "steamed bun", "grow", "Aspergillus flavus". Based on semantic focusing, the system combines the extracted keywords to generate a more accurate search string to obtain the user query string q.

[0099] Step 20 constructs a distributed web crawler engine based on the query string q (using an asynchronous IO architecture to improve data acquisition efficiency), and uses the requests library to send HTTP requests to the search engine to obtain web page content. Use BeautifulSoup to parse HTML, and at the same time use chardet to detect the encoding to avoid garbled characters. The system extracts the title, source, URL, and body content of each search result. To ensure that the content is plain text and relevant to the query, the following screening is required:

[0100] Detect whether the web page contains <video>Label, skip this page if any;

[0101] Remove from the web page< / video>

[0102] and and labels; from Extract the text from the tags and combine it into the final body text.

[0103] If the web page lacks valid text, mark it as "No body content found" and skip it.

[0104] Title: Will frozen steamed buns produce aflatoxin after more than two days? Countless netizens are on edge! What's the truth?

[0105] Content: In fact, regarding the statement that "frozen steamed buns will produce aflatoxin after more than two days", experts have refuted it many times. Because the production of aflatoxin requires specific temperature and humidity conditions. The suitable growth temperature is 12°C to 42°C, and the optimal temperature is 33°C. The temperature in the freezer of a refrigerator is generally -18°C, which is much lower than the temperature range for the production of aflatoxin...

[0106] Step 30 performs sentence segmentation on each web page content obtained in Step 20 to obtain a sentence set S = {s1, s2,..., s n}(n is the total number of sentences in the web page). Use the pre-trained BERT model to convert the user query q and each sentence s i into semantic vectors with dimension d. Calculate the cosine similarity:

[0107]

[0108] For each web page, select the highest similarity score as the overall similarity score of the web page, and set the corresponding sentence as the best matching sentence best_match(p) of the web page.

[0109] Step 40 selects the top K% of the web pages sorted in Step 30 for authority calculation. Let A be a preset set of authoritative navigation pages (such as government official websites, authoritative media). For p ∈ A, let Auth(p) = 1; at the same time, for each p ∈ A, denote its external link set as L(p), and define it as:

[0110]

[0111] Initially assign a weight value of 1 to all web pages in the set A ∪ L A . For pages not in this set, they are classified as non-authoritative pages, and the HITS algorithm is used to calculate the authority between web pages. The specific process includes:

[0112] Construct a directed graph with web pages as nodes and external links as directed edges;

[0113] Initially set the Authority and Hub values of each page to 1;

[0114] Iterative update:

[0115]

[0116] After each iteration, the Authority and Hub values are L2 normalized;

[0117] When the difference between two consecutive iteration results is less than the preset threshold, it is considered to have converged and the Authority value of each page is output. The Authority value will be used as a weight in the credibility calculation in subsequent steps.

[0118] Step 50 uses the user query q as the hypothesis sentence s1 to be verified, and the webpage content (content(p)) ranked high in step 30 as the sentence s2 with a deterministic premise. The specific implementation adopts a dual encoder architecture:

[0119] Use the Transformer model (with parameters θ1 and θ2) to encode s1 and s2 respectively, and get the semantic representation φ hypo (s1) and φ prem (s2);

[0120] Generate relational representations through interactive operations:

[0121] h θ (s1, s2) = W·[φ hypo ; φ prem ; φ hypo ⊙φ prem ]+b

[0122] in and is a learnable parameter and ⊙ represents element-wise multiplication.

[0123] Use the softmax function to calculate the conditional probability:

[0124]

[0125] And obtain the inference result by maximizing the conditional probability

[0126] During the reasoning process, set the following system prompt: "You are a reasoning assistant. The user will provide a question and content, and you need to perform reasoning based on the question and content. Take the result of the reasoning as the premise and the question provided by the user as the hypothesis, and give a clear judgment. If the subject of the hypothesis is inconsistent with that of the premise, output 'irrelevant'; if the hypothesis can be deduced from the premise, output'supported'; if the premise contradicts the hypothesis, output 'opposed'. For example, for the input 'Content: It is sunny today, and it will be cloudy tomorrow with a temperature of 18 - 23 °C; Question: Will it be sunny tomorrow?' Since it will be cloudy tomorrow, the output should be 'opposed'. Note that the output is limited to 'irrelevant','supported', 'opposed'." According to the reasoning result, if the output is "supported", "irrelevant", or "opposed", assign the reasoning result support degree y = 1, 0, -1 respectively.

[0127] Step 60 uses the authority values Auth(p) of each web page obtained in Step 40 as weights to perform weighted calculation on the reasoning result support degrees y obtained from natural language reasoning of each web page in Step 50, to obtain the final credibility score of the user query q.

[0128] Step 70 filters and sorts the search results according to the similarity scores calculated in Step 30. The top K% web pages form a set W. Set a threshold L based on the product of the web page authority and the reasoning result support degree. For web pages whose absolute value of the product exceeds L, store their relevant information in the first data table of the database. The stored content includes:

[0129] User query information query(q)

[0130] The reasoning result result(p) of the natural language reasoning task

[0131] Web page title title(p)

[0132] Web page body content content(p)

[0133] The sentence best_match(p) most matching the query

[0134] For other web pages in set W that are not stored in the first data table, store their titles and body contents in the second data table. The first data table is used to quickly match user queries and retrieve the content for which reasoning judgments have been completed; the second data table serves as supplementary evidence to avoid repeated crawling and improve the response speed.

[0135] Table 1 Example of the content stored in the first data table

[0136]

[0137] Table 2 Content stored in the second data table

[0138]

[0139] Step 8: Use a large language model to take the user query q as input and generate an answer with a limited number of words and the reasons for judgment. The result output by this large model is used as a general knowledge judgment to supplement the web page reasoning result and provide an additional reference basis for the user.

[0140] Step 9: The system finally comprehensively displays the two parts of the results to the user: The first part is the credibility score Cred(q) obtained through the web page content reasoning task, as well as the reference web pages and the best matching sentences involved in the calculation; The second part is the general knowledge judgment result of the large model and its brief explanatory note. Finally, the user will intuitively obtain the credibility value of the query q and detailed reference basis, so as to achieve fact verification and judgment.

[0141] Through the description of the above embodiments, those skilled in the art can clearly understand that the method or system of the present disclosure can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits, etc. However, in more cases for the present disclosure, software program implementation is a better embodiment.

[0142] Although the embodiments of the present disclosure have been described above in conjunction with the accompanying drawings, the present disclosure is not limited to the above specific embodiments and application fields. The above specific embodiments are merely illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present disclosure, and these all belong to the scope of protection of the present disclosure. < / video>

Claims

1. A fact-checking system, characterized in that, The system includes an input module, a query retrieval module, a first calculation module, a second calculation module, and a third calculation module; wherein: The input module is configured to obtain a corresponding query string based on the image or text of the user query information; The query retrieval module is configured to obtain web pages in text form related to the query string; The first calculation module is configured to perform sentence segmentation processing on the obtained web pages to obtain a sentence sequence, and then evaluate the similarity between each sentence and the query string; The second calculation module is configured to calculate the authority score of the reference web pages using the HITS algorithm for the top K% of the reference web pages with the highest similarity; The third calculation module is configured to use the user query information as a hypothesis sentence and the content of the reference web page as a premise sentence, and judge the logical relationship between the hypothesis sentence and the premise sentence as "support", "irrelevant", or "opposition" through conditional probability calculation as the reasoning result, and obtain the support degree corresponding to the reasoning result. Furthermore, based on the authority score of each web page and the support degree corresponding to each web page, the credibility of the fact-checking result is obtained.

2. The fact verification system according to claim 1, wherein The system further includes a knowledge base, and the knowledge base is configured to include a first data table. The data stored in the first data table includes the user query information and its corresponding reasoning result, as well as relevant information of the web page set that meets the conditions. The relevant information includes the web page title, the web page text content, and the sentence that best matches the query information. The web page set that meets the conditions belongs to the top K% of the web page sets with the highest similarity, and the product of the authority score of the web page in the web page set that meets the conditions and the support degree corresponding to the web page is greater than a set threshold.

3. The fact verification system according to claim 2, wherein The knowledge base is further configured to include a second data table, and the second data table stores other web page information that is not stored in the first data table in set W, including the web page title and the web page text content.

4. The fact verification system according to claim 2, wherein The system further includes a result output module, and the result output module is configured to output the credibility of the fact-checking result, the web page reasoning source, and the general result supplement. The web page reasoning source is the web page text content of the reference web page and the sentence that best matches the query information. The general result supplement is the result of using the large language model to judge the query string based on the general knowledge it stores.

5. The fact verification system according to claim 1, wherein The input module obtaining the query string includes: Performing structured semantic parsing on the image content to obtain a description text; For the text content, using a word segmentation tool to extract keywords, and then generating a query string through semantic focusing.

6. The fact verification system according to claim 1, characterized in that, The query and retrieval module obtains web pages as follows: constructing a distributed web crawler engine based on the query string, extracting the title, source, URL, and body content of each search result from the crawled web pages, skipping web pages that lack body content and have videos, removing tables and pictures in the remaining web pages, and from And Extracting the text content from the tags and merging it into the final web page text.

7. The fact verification system according to claim 1, wherein The first calculation module evaluating the similarity between each sentence and the query string includes: Based on n sentences in a web page, obtain the sentence sequence S = {s1, s2, …, s n}; Use the BERT model to convert the query string q and each sentence s i into semantic vectors q and s with a vector dimension of d respectively i , that is Calculate the cosine similarity between the vector q and each sentence vector s i to evaluate the similarity between each sentence and the query string.

8. The fact verification system according to claim 1, wherein The third calculation module obtaining the credibility of the fact-checking result includes: Take the query string as the hypothetical sentence s1, and take the content of the reference web page as the sentence s2 with a deterministic premise. Use a dual encoder to encode s1 and s2 to obtain the semantic representation φ hypo (s1) and φ prem (s2); Generate the relational representation h through interactive operations θ (s1, s2) = W · [φ hypo (s1); φ prem (s2); φ hypo (s1) ⊙ φ prem (s2)] + b, where known are learnable parameters, and ⊙ represents element-wise multiplication; Calculate the conditional probability And determine the logical relationship of "support", "irrelevant", or "opposition" between the hypothesis and the premise by maximizing the conditional probability as the reasoning result, and then obtain the support degree of the reference web page according to the support degree set by the reasoning result; Calculating the weight of each web page according to the authority score of each web page, and weighting the support degree corresponding to all web pages using the weight to obtain the credibility of the fact-checking result.

9. A computer-readable storage medium, characterized in that: A computer program is stored that can be loaded and executed by a processor for a system according to any one of claims 1 to 8.

10. A fact-checking method, characterized in that, The method includes the following steps: Step 1, generating a unified query string q based on the image or text of the user query information; Step 2: Obtain web pages in text form related to the query string q; Step 3: Perform sentence splitting on each obtained web page to obtain a sentence sequence, and then evaluate the similarity of each sentence to the query string; Step 4: Use the top K% of the web pages ranked by similarity as reference web pages, and calculate the authority scores of the reference web pages using the HITS algorithm; Step 5: Use the user query information as the hypothetical sentence and the web page content as the premise sentence, and calculate the conditional probability to determine the logical relationship of "support", "irrelevant", or "opposition" between the hypothetical sentence and the premise sentence as the reasoning result; Step 6: Obtain the support degree corresponding to the reasoning result, and based on the authority score of each web page and the support degree corresponding to each web page, obtain the credibility of the fact-checking result.

Citation Information

Patent Citations

  • Question and answer method and system based on credibility perception and retrieval enhanced language model

    CN118153690A

  • Multi-modal fact checking method based on retrieval enhancement generation

    CN119128739A

  • Method, system and equipment for large language model fact verification based on retrieval enhancement and medium

    CN119204025A