Method and device for evaluating output of large language model, medium and product
Through the iterative tree analysis method, key views are extracted from long medical text and adaptive verification is solved, and the problem that the existing technology is difficult to effectively verify views in long medical text is improved, and the quality of medical applications and verification accuracy are improved.
Patent Information
- Application Number
- CN202510072893.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has difficulty effectively verifying the views in long medical texts, especially in the medical field, which makes existing automation methods difficult to cope with.
Using an iterative tree analysis (ITA)-based approach, key perspectives are extracted from long medical texts and each perspective is verified through an adaptive tree-like reasoning process, combining top-down task splitting and bottom-up evidence integration.
It improves the accuracy and reliability of long medical text verification, and achieves accurate verification of complex medical views through detailed mechanism-level reasoning, which enhances the quality of medical fact verification.
Smart Images

Figure CN120012926A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence, and more specifically, to a method, device, medium, and product for evaluating the output of a large language model. Background Art
[0002] Large Language Models (LLMs) have been widely used in various fields and have performed well in various tasks. However, their application in the medical field brings unique challenges, especially in the generation of hallucinations (i.e., erroneous or fictitious outputs).
[0003] Hallucinations in open-ended long medical texts manifest as misleading key claims that are difficult to verify for two reasons. First, key claims are often deeply entangled in the text and cannot be extracted based on surface-level presentation alone. Second, verifying these claims is challenging because surface-level token retrieval often lacks precise or specific evidence, and these claims cannot be verified without deeper mechanism-based analysis.
[0004] With the rapid development of LLM-based QA systems, many solutions have been proposed for generating accurate answers and opinions. However, factual verification remains a major challenge, especially in fields such as medicine. Medical fact verification is particularly challenging due to the complex relationships between medical concepts, symptoms, and treatments, which are often not verifiable through simple queries. This process requires understanding implicit causal relationships and constructing detailed chains of evidence. Existing automated methods, such as fact checking and knowledge-based systems, typically compare opinions with static databases or predefined reference answers. While these methods are effective for simple opinions, they have difficulty coping with the complexity and interrelatedness of long medical texts.
[0005] Therefore, a new approach is needed to address the above issues to improve the reliability and accuracy of verification of long medical texts and ultimately improve the quality of medical applications. Summary of the invention
[0006] In response to the above problems, the present disclosure provides a method for evaluating the output of a large language model based on iterative tree analysis (ITA). The ITA method provided by the present disclosure aims to extract implicit ideas from long medical texts and verify each idea through an iterative and adaptive tree-like reasoning process. The process of the ITA method combines top-down task splitting and bottom-up evidence integration to achieve accurate verification of complex medical ideas through detailed mechanism-level reasoning.
[0007] The disclosed embodiment provides a method for evaluating the output of a large language model, comprising: extracting one or more key viewpoints from the output of the large language model, each of the key viewpoints comprising a self-consistent statement of fact; constructing an adaptive mind tree structure based on the extracted key viewpoints to split each key viewpoint into a plurality of verifiable sub-viewpoints; retrieving evidence data associated with the plurality of sub-viewpoints from one or more external information sources, and verifying the plurality of sub-viewpoints based on the retrieved evidence data; in response to the verification of the plurality of sub-viewpoints having been completed, integrating the verification results of the plurality of sub-viewpoints, and determining the evaluation results of the corresponding key viewpoints based on the integrated verification results of the plurality of sub-viewpoints; and outputting the evaluation results of each of the key viewpoints and the corresponding evidence data.
[0008] According to an embodiment of the present disclosure, the output of the large language model includes output text generated in response to an input query, and the output text includes a long medical text.
[0009] According to an embodiment of the present disclosure, the key viewpoints are independent of each other, and the root node in the tree structure graph is the key viewpoint, the child nodes of the tree structure are the child viewpoints, and the edges in the tree structure graph indicate the relationship between the nodes.
[0010] According to an embodiment of the present disclosure, the method also includes: in response to the verification of the sub-viewpoint indicating that the granularity of the sub-viewpoint is low, iteratively further splitting the sub-viewpoint into multiple atomic viewpoints with higher granularity; retrieving evidence data associated with the multiple atomic viewpoints from one or more external information sources, and verifying the multiple atomic viewpoints based on the retrieved evidence data.
[0011] According to an embodiment of the present disclosure, in response to the verification of the plurality of sub-views being completed, it also includes: determining that the plurality of sub-views do not need to be further split or determining that the sub-views have reached a maximum granularity.
[0012] According to an embodiment of the present disclosure, integrating the verification results of the multiple sub-viewpoints and determining the evaluation results of the corresponding key viewpoints based on the integrated verification results of the multiple sub-viewpoints also includes: integrating the verification results of the multiple atomic viewpoints corresponding to each sub-viewpoint of the multiple sub-viewpoints, and determining the verification result of each corresponding child node based on the integrated verification results of the multiple atomic viewpoints; determining the evaluation results of the corresponding key viewpoint based on the integrated verification results of the multiple sub-viewpoints.
[0013] According to the embodiment of the present disclosure, the initial parameters included in each node include a viewpoint and a scalar value indicating whether the viewpoint is acceptable. Integrating the verification results of the multiple sub-viewpoints, and determining the evaluation results of the corresponding key viewpoints based on the integrated verification results of the multiple sub-viewpoints also includes: starting from the leaf node of the tree structure graph, passing the scalar value of the verified sub-viewpoint and the evidence data for providing judgment to the parent node; integrating the scalar values of multiple child nodes to modify the state of the parent node; and using the updated state of the parent node to modify the state of the root node, and determining the evaluation results of the key viewpoint based on the state of the root node.
[0014] According to an embodiment of the present disclosure, the large language model includes a set of pre-set retrieval tools, each retrieval tool is configured to retrieve a specific type of information or call an external calculator for evaluation, and for each sub-viewpoint, an appropriate retrieval tool is selected to retrieve evidence data associated with the multiple sub-viewpoints.
[0015] According to an embodiment of the present disclosure, selecting an appropriate retrieval tool to retrieve the evidence data associated with the multiple sub-viewpoints includes: using the viewpoint of the parent node of the sub-viewpoint as context to determine an appropriate retrieval tool; generating a retrieval query task to start retrieving the evidence data associated with the multiple sub-viewpoints; and sorting and selecting the retrieved evidence data based on the relevance to the viewpoint and the origin and type of the information source.
[0016] According to an embodiment of the present disclosure, the method further includes: systematically evaluating key viewpoints using a checklist score based on LLM-as-a-Judge.
[0017] According to an embodiment of the present disclosure, the method further includes: constructing a test data set, wherein the benchmark test data set includes a set of correct opinions, and a set of false opinions, factual texts, and non-factual texts; and using the constructed test data set to evaluate the evaluation results of each of the key opinions.
[0018] According to an embodiment of the present disclosure, constructing a test data set also includes: extracting a series of atomic opinions from a sentence in a given medical guide text as correct opinions; adding random errors to the correct opinions to generate fake opinions; generating a paraphrase of the original text based on the correct opinions and the fake opinions as factual text; and creating an alternative version of the factual text as non-factual text by merging the fake opinions.
[0019] According to an embodiment of the present disclosure, the categories of the test data in the test data set include the following pathophysiology, drug therapy, diagnosis, symptoms, treatment and prevention.
[0020] According to another embodiment of the present disclosure, a device for evaluating the output of a large language model is provided, including: a processor, and a memory, wherein the memory stores computer executable instructions, and when the computer executable instructions are executed by the processor, the processor is prompted to perform the method described above.
[0021] According to another embodiment of the present disclosure, a computer-readable recording medium is provided, storing computer-executable instructions, wherein when the computer-executable instructions are executed by a processor, the processor is prompted to perform the method described above.
[0022] According to another embodiment of the present disclosure, a computer program product is provided, comprising computer executable instructions, wherein when the computer executable instructions are executed by a processor, the processor is prompted to perform the method described above.
[0023] According to the ITA method of the embodiment of the present disclosure, medical fact verification can be enhanced by using adaptive mind tree reasoning. The ITA method can effectively extract atomic viewpoints from the original text and construct an evidence tree to support true and false judgments, thereby improving the accuracy and reliability of viewpoint verification. The ITA method can also provide a unified and universal framework for medical viewpoint detection by generating and merging subtrees using retrieved external reference information, which allows comprehensive description and verification of different important viewpoints in the input query text, thereby facilitating a more meticulous and detailed analysis of medical information. In addition, the embodiment of the present disclosure also constructs a verification data set, including a fine-grained checklist that can support the evaluation of verification tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some exemplary embodiments of the present disclosure, and a person of ordinary skill in the art can obtain other drawings based on these drawings without creative work.
[0025] Figure 1 is a schematic diagram showing an ITA framework;
[0026] Figure 2 is an example diagram illustrating the verification phase of ITA;
[0027] Figure 3 is a flow chart illustrating a method 300 of evaluating the output of a large language model according to an embodiment of the present disclosure;
[0028] Figure 4 is a schematic diagram showing the comparison between manual evaluation and ITA evaluation;
[0029] Figure 5(a) and Figure 5 (b) is a comparison chart showing the performance of various models evaluating long medical texts;
[0030] Figure 6 A block diagram of an apparatus 600 for evaluating the output of a large language model according to an embodiment of the present disclosure is shown; and
[0031] Figure 7 A schematic diagram of a recording medium 700 according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical solution and advantages of the present disclosure more obvious, the exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, rather than all the embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described here.
[0033] In this specification and the accompanying drawings, substantially the same or similar steps and elements are represented by the same or similar reference numerals, and repeated descriptions of these steps and elements will be omitted. At the same time, in the description of the present disclosure, the terms "first", "second", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance or ranking.
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which the present disclosure pertains. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0035] To facilitate description of the present disclosure, concepts related to the present disclosure are introduced below.
[0036] The data processing related to the large language model (LLM) mentioned in the present disclosure can be based on artificial intelligence (AI). Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. For example, for a large language model based on artificial intelligence, it can learn the subtasks, combined tasks and training data to be introduced next in a similar way to humans based on their acquired knowledge and use the learned content to handle subsequent similar problems.
[0037] Large language models (LLMs) have performed well on a variety of tasks. However, in high-stakes domains such as evidence-based medicine, they face serious challenges that limit their practical reliability. Although LLMs have performed well on standardized medical benchmarks, such evaluations often ignore the nuances and complexities of real-world clinical scenarios. To address this gap, it is expected that future evaluations should go beyond the traditional multiple-choice question format and incorporate open-ended, long-form evaluations. These evaluations should rigorously test the model's ability to verify complex and nuanced information. Such evaluations should incorporate complex and implicit medical knowledge, challenging LLMs to demonstrate deeper understanding and reasoning capabilities. Accurately verifying responses is particularly important in medical applications, as LLMs often produce hallucinations - erroneous or fictitious outputs. By emphasizing detailed factual verification, researchers can better understand the strengths and limitations of LLMs, guide their improvements, and increase their reliability in critical applications. Many solutions have been proposed for generating accurate answers and opinions for LLM-based QA systems. However, factual verification remains a major challenge, especially in fields such as medicine.
[0038] Current approaches typically split factual verification into two tasks: (i) determining whether a claim is factually correct, and (ii) evaluating whether the claim is supported by evidence.
[0039] These methods typically rely on surface-level input descriptions, which are insufficient for complex, long-form medical statements. Medical fact verification is particularly challenging due to the complex relationships between medical concepts, symptoms, and treatments, which are often not verifiable through simple queries. This process requires understanding implicit causal relationships and constructing detailed chains of evidence. Existing automated methods, such as fact-checking and knowledge-based systems, typically compare opinions against static databases or predefined reference answers. While these methods are effective for simple opinions, they struggle to cope with the complexity and interconnectedness of long-form medical statements, highlighting the need for more advanced investigation processes and dynamic evaluation systems.
[0040] In this disclosure, a novel framework, Iterative Tree Analysis (ITA), is proposed to identify and correct factual errors in medical opinions expressed in natural language. ITA innovatively solves the main task by implementing a verification process that systematically manages specific sub-opinions. This process involves recursively verifying sub-opinions and building a verification tree. The tree structure is developed based on the sub-opinions and the currently available information retrieved for each sub-tree. The ITA approach achieves complex verification goals by merging verification trees from the bottom up. This divide-and-conquer strategy simplifies the task into manageable sub-opinions and facilitates the step-by-step exploration of challenging verification tasks by integrating more reliable external references.
[0041] Next, the method provided by the present disclosure will be described in detail with reference to the accompanying drawings.
[0042] First, refer to Figure 1 The overall process of the ITA method is explained. Figure 1 In the text, medical text output by a large language model is used as an example. It should be noted that medical text is only used as an example, and other types of text can also be used.
[0043] Overview of the ITA Method
[0044] like Figure 1 As shown in the figure, the overall process of the ITA method consists of two stages: the splitting stage and the merging stage. In the tree generation stage (i.e., the splitting stage), individual opinions are extracted from the medical text, ensuring that each opinion is independent by integrating information from other parts of the text. The verification subtasks are recursively distributed to the leaf nodes. Then, in the integration stage, deeper knowledge insights into each individual opinion are used to determine whether to accept or reject its parent opinion. The framework outputs the verification status of each individual opinion and its supporting references.
[0045] In the splitting stage, all verifiable opinions are extracted from the medical text. It should be noted that because the medical text is a long medical text generated by a large language model for query input, it may be necessary to extract content from multiple different locations in the medical text to generate an opinion. Figure 1 Figure 3 shows three viewpoints extracted from a medical text. Figure 1 The three viewpoints in the figure are merely examples, and the number of viewpoints that can be extracted is not limited to three, and can be one, two, three, four or more.
[0046] Then, a verification tree T is constructed based on these viewpoints. The medical viewpoints under review often require complex chains of reasoning that rely on fine-grained knowledge points to accurately identify different viewpoints. In some cases, the verification process requires indicator calculation or the merging of multiple layers of viewpoints, especially when the combination of medical concepts does not exist or is difficult to retrieve in external references. In this case, it is crucial to collect sufficient evidence from external sources (such as recent research articles) to construct the implicit chain of evidence necessary to determine whether the viewpoint is supported. Figure 1 As shown in the figure, ? indicates that the view has not been verified, and √ indicates that the view has been verified.
[0047] Therefore, a complete ITA approach includes the following components:
[0048] Large language model M, opinion variable V:={vi}, evidence relation E:={(vi,vj)}, verification tree T:={V,E}, input query q, information source D and retriever R.
[0049] The verification task Q involves determining the factuality of a statement by constructing a logical sequence of thoughts supported by external references. Given an input query q, the goal is to construct an optimal thought tree T that adaptively generates verification subtrees and retrieves relevant document sources D as evidence.
[0050] Specifically, for a given input query q, a large language model (LLM) M is prompted to iteratively expand the mind tree T, ensuring that each node and relationship in the tree is based on the retrieved evidence:
[0051]
[0052] Where A is a set of verified opinions ai, each with judgment flags STATE and REASON. The ultimate goal is to build a logical sequence to determine whether all opinions in the query can be verified by evidence and follow specific rules. These rules cover a variety of sources, including Wikipedia, textbooks, expert explanations, and accurate calculators. By adopting this broader definition, a unified framework can be established to solve the problem of factual accuracy.
[0053] While textual content can be easily divided into sentences and brief representations, identifying precise atomic claims requires a higher level of granularity. Often, information that is not immediately apparent is critical to substantiating these claims. For example, some claims require multi-hop verification, while others may be susceptible to spurious relationships and require careful scrutiny to avoid misleading confounders. In a tree-based representation, these claims can be constructed as a graph with each node representing a minimal claim and edges representing relationships between claims. Simply extracted claims often lack sufficient reasoning details for proper verification. By systematically refining claims with the support of external fact-checking tools and rigorously analyzing the entire tree, the accuracy of individual claims and the integrity of the extended tree can be ensured.
[0054] Verification using spanning tree
[0055] The core idea of ITA is to explore when the answer is uncertain and integrate when external information is sufficient. After repeatedly referring to external resources, deep reasoning tracking can be used to verify the insights extracted from the initial query.
[0056] Splitting process
[0057] The process of generating a validation tree begins by identifying key ideas from the initial sentences of the medical text, which serve as the basis for exploration and exploitation. Depending on the specific medical question, ideas can take various forms, such as a few words (e.g., a drug name), a line of an equation (e.g., a biomedical indicator), or a sentence expressing a causal relationship (e.g., symptom inference). Typically, sub-ideas should be sufficiently "atomic" to enable the language model (LM) to generate valid queries. Given a query containing task-specific information, the language model Mgenerate is then prompted to generate ideas that need to be validated:
[0058]
[0059] Here, v′ i ={CLAIM,STATE} is initially set to "verify". To ensure thorough fact-checking, the proposed view is revised to be self-consistent according to predefined guidelines. Once node v′ i Once verified, it will be supplemented with reference information, judgment reasons and updated STATE. This updated node is represented as Once the initial opinions are generated, the verification tree T is built by iteratively extending the tree and adding new opinions based on the information retrieved through the retriever R:
[0060]
[0061] Here, v′ cur is the root node of the current SUBTREE, and δ indicates the generation termination flag (e.g., δ:=supported / not supported indicates termination, and δ:=undecidable indicates splitting). Once the large language model (LLM) M determines that no further expansion is required or the maximum exploration condition is reached, the generated SUBTREE will only contain node v′ cur . This will trigger a bottom-up node validation merge process.
[0062] The splitting method of the ITA disclosed in the present invention involves generating subsequent queries based on a subtree structure originating from a root node. The root node contains meta-information, such as a summary of the input topic, and splits the extracted entities into leaf nodes. Each node is evaluated to maintain and update the current meta-information, thereby facilitating a thorough verification process.
[0063] Merger Process
[0064] The tree merging process starts from the leaf nodes in a bottom-up manner and gradually integrates the verified opinions with the help of external information. Each generated subtree corresponds to a meta-opinion, with the retrieved documents and related leaf opinions attached. The state evaluator Mcons evaluates the progress of problem solving and acts as an opinion merging mechanism. In the merging step, RAG-assisted verification starts from the leaf nodes and passes the STATE of the processed opinions (indicating whether the opinions are supported) and the REASON that provides the judgment to the parent node. This process modifies the parent node with the reasoning results of its child nodes:
[0065]
[0066] where v′ cur is the root node of the SUBTREE currently being verified. The hint reasoning subtrees are merged to produce a scalar value for STATE (e.g., a score from 1 to 10) or a confidence category (e.g., accepted, rejected, or unconfirmed), which can be mapped to a numeric value. The specific category may vary depending on the problem or the reasoning steps involved.
[0067] Intuitively, when v′ cur When expressing symptoms and medication, CHILD (v′ cur ) will retain fragments of information such as organs, tissues, and drug treatments. By deeply studying the CHILD of CHILD, we can explore the relationship between cells, proteins, and molecular mechanisms, such as Figure 1 shown.
[0068] Independent Views and Searches
[0069] Medical texts often contain key ideas embedded in the main narrative, like intricate patterns woven into a complex textile. The surface presentation of these texts does not always transparently convey the underlying logical structure. To accurately identify and interpret key ideas, this paper designs a multi-hop idea refinement agent and refines its reasoning trajectory into a large language model (LLM) and fine-tunes it. This process allows the extraction of important ideas that may remain obscure in a broader context. Figure 2 As shown in Figure 1, long texts often contain a lot of information that can be either mixed or ambiguous. A key component of successful verification is to ensure that each SUBTREE is independent. This requires that each SUBTREE contains clear entity names, rather than pronouns or vague references, which may hinder the verification process. The present disclosure constructs self-consistent factual statements from the original text to enable more fine-grained evaluation, such as Figure 2 Each factual statement is a stand-alone sentence that conveys information about a disease, condition, medication, treatment, therapy, diagnosis, indication, and side effect.
[0070] Query Generation
[0071] For a given point v i , the retriever R, with the help of a large language model (LLM) Mquery, determines the appropriate external sources to consult. The search space of the retriever R initially consists of a set of pre-built tools R = {R1, R2, ...}, where each tool retriever Ri is customized to retrieve a specific type of information or call an external calculator to evaluate the metric results. The selection process is guided by the following equation:
[0072]
[0073] In this context, Ri is determined using the meta-opinion of its parent node as context, while s represents the generated search query or encoding task. This helps in distributed verification of the opinion v′ i , such as the calculation of specific biomedical indicators that depend on the parent node meta-viewpoint. In addition, some molecular mechanisms associated with the parent node meta-viewpoint may require external information updates for accurate verification.
[0074] Retrieve and select.
[0075] After the information is retrieved, it is processed by the retriever R to select relevant documents and extract key information. The external sources used can range from search engines and medical textbooks to specialized calculators. Despite the variety of these sources, the retrieval process always aims to provide evidence supporting the point of view. Therefore, the retrieved information is re-ranked and selected according to its relevance to the point of view:
[0076]
[0077] Here, D represents the retrieved information set, and each d i Represents pre-processed content extracted from a document. The RERANK function can be rule-based or learned. For example, it can be configured to prioritize scientific sources over less reliable sources such as advertisements. This curated information is essential to validate ideas and provide strong evidence for the integration process.
[0078] Next, we will refer to Figure 3 A method of evaluating the output of a large model according to an embodiment of the present disclosure is described. Figure 3 is a flow chart illustrating a method 300 of evaluating the output of a large model according to an embodiment of the present disclosure. The method 300 may be performed by a suitable processor or processing unit in any system.
[0079] Reference Figure 3 In step S310, one or more key ideas are extracted from the output of the large language model, each of the key ideas comprising a self-consistent factual statement.
[0080] In some embodiments, the output of the large language model includes output text generated in response to the input query.
[0081] For example, a user may use a large language model and input a query "What are the causes of primary childhood nephrotic syndrome?" In response to the input query, the large language model may output an output text including a lot of content. The present disclosure takes the output content in the medical field as an example, and therefore, the output text includes a long medical text.
[0082] An example of a long medical text output by a large language model in response to an input query is shown below:
[0083] Medical Text 1:
[0084] "Although idiopathic or unknown disorders are the most common causes of primary nephrotic syndrome, researchers have linked certain diseases and certain gene changes that damage the kidneys to primary nephrotic syndrome.\n\nSecondary nephrotic syndrome is caused by an underlying disease or infection. This underlying disease or infection is called the primary disease, and it causes changes in kidney function that lead to secondary nephrotic syndrome.\n\nCongenital disorders Disorders that are present at birth can also cause nephrotic syndrome.\n\nPrimary nephrotic syndrome\n\nThe following disorders are different types of idiopathic nephrotic syndrome:\ n\n-Minimal change disease involves damage to the glomeruli that can only be seen with an electron microscope. This microscope can show tiny details better than any other microscope. Scientists do not know the exact cause of minimal change disease. Minimal change disease is the most common cause of idiopathic childhood nephrotic syndrome. 1-Focal segmental glomerulosclerosis is scarring in discrete areas of the kidney: -Focal means only some of the glomeruli are scarred. -Segmental means the damage affects only part of a single glomerulus. -Membranoproliferative glomerulonephritis is a group of diseases that involve the deposition of antibodies within the glomeruli, causing thickening and damage to the glomeruli. Antibodies are proteins produced by the immune system that Proteins are used to protect the body from foreign substances such as bacteria or viruses.\n\nSecondary Nephrotic Syndrome in Children\n\nSome common diseases that may cause secondary nephrotic syndrome in children include\n\n-Diabetes, a disease that occurs when the body cannot use glucose (a type of sugar) properly-Henchollera purpura, a disease that causes small blood vessels in the body to become inflamed and leaky-Hepatitis, inflammation of the liver caused by a virus-Human immunodeficiency virus (HIV), a virus that changes the immune system-Lupus, an autoimmune disease that occurs when the body attacks its own immune system-Malaria, a blood disease spread by mosquitoes-Streptococcal infection, An infection that results when bacteria that cause strep throat or skin infections go untreated\n\nOther causes of secondary nephrotic syndrome in children are medications and environmental. Common medications like acetaminophen, paracetamol, or other common pain relievers, and exposure to chemicals like mercury and lithium.\n\nCongenital Disorders and Nephrotic Syndrome in Children\n\nCongenital nephrotic syndrome is rare and usually affects babies within the first 3 months of life. 2 This type of nephrotic syndrome, sometimes called infantile nephrotic syndrome, can be caused by\n\n- An inherited gene defect, which is a problem that is passed down from parent to child through the genes - An infection during birth. ”
[0085] Then, one or more key ideas can be extracted from the above output long medical text, for example, by a trained model. Multiple key ideas extracted from the same long medical text are independent of each other. In this embodiment, the key ideas include self-consistent factual statements. An example of the extracted key ideas is shown below.
[0086] Key points to take away: Certain medications, such as acetaminophen, paracetamol, or other common pain relievers, may cause secondary nephrotic syndrome in children.
[0087] In step S320: construct an adaptive mindset tree based on the extracted key ideas to split each key idea into multiple verifiable sub-ideas.
[0088] For example, in the case where the key viewpoint above is used as the root node, there may be no effective way to directly verify whether the viewpoint is true (acceptable). Therefore, an adaptive thinking tree can be constructed, with the key viewpoint as the root node, and split into multiple verifiable sub-viewpoints.
[0089] For example, the following shows the sub-views obtained by splitting:
[0090] Sub-viewpoint 1: Acetaminophen can cause secondary nephrotic syndrome in children.
[0091] Sub-viewpoint 2: Paracetamol can cause secondary nephrotic syndrome in children.
[0092] Sub-viewpoint 3: Other common painkillers can cause secondary nephrotic syndrome in children.
[0093] Compared with the extracted key ideas, the split sub-ideas have higher granularity and are easier to verify.
[0094] As referenced above Figure 2 As described above, the root node in the structure graph of the thinking tree is the key viewpoint, the child nodes of the tree structure are the child viewpoints, and the edges in the tree structure graph indicate the relationship between the nodes.
[0095] At step 330 , evidence data associated with the plurality of sub-viewpoints are retrieved from one or more external information sources, and the plurality of sub-viewpoints are verified based on the retrieved evidence data.
[0096] As referenced above Figure 1 As described, a large language model (LLM) may include a set of pre-set search tools (e.g., search engines such as Google and Baidu, or various database search tools, etc.), each of which is configured to retrieve a specific type of information or call an external calculator for evaluation.
[0097] For each sub-viewpoint, an appropriate search tool is selected to search for evidence data associated with the plurality of sub-viewpoints. For example, Baidu can be used to search for evidence data associated with "acetaminophen can cause secondary nephrotic syndrome in children".
[0098] In some embodiments, the viewpoint of the parent node of the sub-viewpoint can be used as context to determine an appropriate retrieval tool. Then, a retrieval query task is generated to start retrieving evidence data associated with the plurality of sub-viewpoints; and the retrieved evidence data is sorted and selected according to the relevance to the viewpoint and the source and type of the information source.
[0099] It should be noted that in some cases, the sub-views obtained by splitting from the root node may not be able to be effectively verified, so the sub-views need to be further split for verification.
[0100] That is, in response to the verification of the sub-viewpoint indicating that the granularity of the sub-viewpoint is low, the sub-viewpoint is iteratively further decomposed into multiple atomic viewpoints with higher granularity. On this basis, evidence data associated with the multiple atomic viewpoints is retrieved from one or more external information sources, and the multiple atomic viewpoints are verified based on the retrieved evidence data.
[0101] For example, for the differential sub-viewpoint 3 above, the sub-viewpoints obtained by splitting from the root node may not be able to be effectively verified, so the sub-viewpoints need to be further split for verification.
[0102] Therefore, subview 3 is further differentiated to obtain the following atomic view:
[0103] Atomic Viewpoint 1: Investigate whether there are clinical trials or meta-analyses reporting a direct causal relationship between NSAID use and secondary nephrotic syndrome in children
[0104] Atomic Viewpoint 2: Explore case studies or medical records documenting children who developed nephrotic syndrome after taking acetaminophen
[0105] In step S340, in response to the verification of the plurality of sub-viewpoints being completed, the verification results of the plurality of sub-viewpoints are integrated, and the evaluation results of the corresponding key viewpoints are determined according to the integrated verification results of the plurality of sub-viewpoints.
[0106] In some cases, it is determined that the validation of the multiple sub-views has ended when it is determined that the multiple sub-views do not need to be further decomposed or when it is determined that the sub-views have reached the highest granularity.
[0107] At this point, the ITA method enters the integration phase, in which the verification results of the multiple atomic viewpoints corresponding to each of the multiple sub-viewpoints are integrated, and the verification result of each corresponding sub-node is determined based on the verification results of the integrated multiple atomic viewpoints.
[0108] Integrating the verification results of multiple sub-viewpoints, and determining the evaluation results of the corresponding key viewpoints based on the integrated verification results of the multiple sub-viewpoints also includes: integrating the verification results of multiple atomic viewpoints corresponding to each sub-viewpoint of the multiple sub-viewpoints, and determining the verification result of each corresponding sub-node based on the integrated verification results of the multiple atomic viewpoints; determining the evaluation results of the corresponding key viewpoint based on the integrated verification results of the multiple sub-viewpoints.
[0109] Specifically, as mentioned above Figure 1 As described, each node includes initial parameters including a viewpoint and a scalar value (CLAIM, STATE) indicating whether the viewpoint is acceptable.
[0110] For each key viewpoint, starting from the leaf node of the mind tree structure graph, the scalar value (STATE) of the verified sub-viewpoint and the evidence data (REASON) for providing judgment are passed to the parent node. Then, the scalar values of multiple child nodes are integrated to modify the state of the parent node. Finally, the state of the root node is modified using the updated state of the parent node, and the evaluation result of the key viewpoint is determined based on the state of the root node.
[0111] In step S350, the evaluation result of each of the key viewpoints and the corresponding evidence data are output.
[0112] The following shows an example of the extracted key viewpoints, the evaluation results of the split sub-viewpoints 1, 2, and 3, and the corresponding evidence data:
[0113] Extracted key insights
[0114] Certain medications, such as acetaminophen, paracetamol, or other common pain relievers, may cause secondary nephrotic syndrome in children.
×
[0115] Split sub-viewpoint 1
[0116] Acetaminophen can cause secondary nephrotic syndrome in children.
×
[0117] Analysis reasons for the split sub-viewpoint 1
[0118] To determine if the statement "Acetaminophen can cause secondary nephrotic syndrome in children" is supported by the given knowledge, this disclosure requires analysis of the information provided:\n\n1. **Case Report Summary:** This report details a 17-year-old female who experienced acute alcohol intoxication and acetaminophen overdose, followed by rapid onset of renal insufficiency and was diagnosed with acute interstitial nephritis rather than nephrotic syndrome. Renal biopsy findings and clinical presentation (e.g., elevated serum creatinine and BUN levels, eosinophil staining, lack of evidence of glomerulonephritis) were consistent with acute interstitial nephritis but not nephrotic syndrome.\n\n2. Nephrotic Syndrome Information in Children:\n-Nephrotic syndrome involves proteinuria, edema, and the cause is often unknown. Common causes include minimal change nephropathy, which is often associated with infection, allergic reactions, and overdose of medications, including acetaminophen and ibuprofen.\n-It primarily affects children between 1.5 and 5 years of age and is characterized by symptoms such as swelling, proteinuria, and possible glomerular damage. \n\n3. Relevance to the statement\n- The knowledge mentions that overdose of acetaminophen may be associated with minimal change disease, which is one of the causes of nephrotic syndrome in children. However, it does not clearly state the direct pathogenic role of acetaminophen in inducing secondary nephrotic syndrome in children. \n- The case study does not mention that nephrotic syndrome is a result of taking acetaminophen; it focuses on acute interstitial nephritis. \n\n4. Conclusion: \n- Although the knowledge suggests that overdose of acetaminophen may lead to diseases that affect kidney function, such as minimal change disease and potential nephrotic syndrome, there is no conclusive evidence that acetaminophen alone can cause secondary nephrotic syndrome in children. \n- The case study provided describes a disease that is different from nephrotic syndrome (acute interstitial nephritis). \n\n Given this reasoning, the statement: \n\n\"Acetaminophen can cause secondary nephrotic syndrome in children. \n\n Based on the knowledge points provided and evaluated, the statement is ultimately determined to be [Not Supported].
[0119] References
[0120] 1.https: / / www.nejm.org / doi / 10.1056 / NEJMcpc2115851? url_ver=Z39.88-2003&rfr_id=ori:rid:crossref.org&rfr_dat=cr_pub%20%200pubmed
[0121] 2.https: / / www.nhs.uk / conditions / nephrotic-syndrome /
[0122] 3.https: / / www.niddk.nih.gov / health-information / kidney-disease / children / nephrotic-syndrome-children
[0123] 4.https: / / pmc.ncbi.nlm.nih.gov / articles / PMC10182827 /
[0124] 5.…
[0125] Split sub-viewpoint 2:
[0126] Paracetamol can cause secondary nephrotic syndrome in children.
×
[0127] Analysis reasons for the split sub-viewpoint 2
[0128] To determine whether the statement "paracetamol can cause secondary nephrotic syndrome in children" is supported by the knowledge provided, the disclosure needs to review and analyze relevant information. \n\n1. Understanding nephrotic syndrome:\n-Nephrotic syndrome in children is characterized by excessive protein in the urine, swelling, and other symptoms caused by kidney problems. \n-It is divided into two forms: primary and secondary. Primary nephrotic syndrome is caused directly by kidney problems, while secondary nephrotic syndrome is related to other systemic diseases or factors. \n\n2. Causes of secondary nephrotic syndrome:\n-Secondary nephrotic syndrome can be caused by systemic diseases, infections, drugs, and other factors. \n-Specific causes include diseases such as lupus, infections such as hepatitis B, certain drugs, etc. \n\n3. The role of acetaminophen, also known as paracetamol or acetaminophen, in nephrotic syndrome:\n-The knowledge provided mainly discusses acetaminophen in the context of asthma, indicating a potential link between acetaminophen and asthma exacerbations. \n-There is a mention that acetaminophen use during pregnancy can lead to increased respiratory problems in offspring, but there is no direct mention of secondary nephrotic syndrome as an outcome. \n\n4. Assessment based on current knowledge: \n-This knowledge does not directly or indirectly suggest that acetaminophen can cause secondary nephrotic syndrome. \n-While acetaminophen has been linked to various health issues, particularly with respiratory health and oxidative stress, the knowledge provided does not list nephrotic syndrome as one of these. \n\n5. Conclusion: \n-As the claim does not appear to be supported by knowledge about acetaminophen and its health effects, particularly with nephrotic syndrome, it indicates a lack of connection or evidence in the context provided. \n\nRestatement: Acetaminophen can cause secondary nephrotic syndrome in children. \n\nFinal answer: [Not supported]
[0129] References
[0130] 1.https: / / pubmed.ncbi.nlm.nih.gov / 9176846 /
[0131] 2.https: / / pmc.ncbi.nlm.nih.gov / articles / PMC5306275 / 3.
[0133] https: / / www.singhealth.com.sg / patient-care / conditions-treatments / nephrotic-syndrome-children
[0134] 4.…
[0135] Split sub-viewpoint 3:
[0136] Other common painkillers can cause secondary nephrotic syndrome in children.
×
[0137] Split Atomic View 1
[0138] To investigate whether there are clinical trials or meta-analyses reporting a direct causal relationship between NSAID use and secondary nephrotic syndrome in children [×]
[0139] Analysis of the split atomic view 1
[0140] To determine if the statement is supported by the given knowledge, let this disclosure go through the reasoning process step by step:\n\n1. Understand the statement: The inquiry is to determine if there are clinical trials or meta-analyses that have established a direct causal relationship between the use of nonsteroidal anti-inflammatory drugs (NSAIDs) and the development of secondary nephrotic syndrome in children.\n\n2. Summarize the knowledge: The knowledge provided includes a lot of information about various treatments and studies on nephrotic syndrome in children, with a special focus on immunosuppressants and their efficacy, safety, and acceptability. These documents involve descriptions of clinical trials and meta-analyses of different drugs (such as rituximab, cyclophosphamide, etc.) for the treatment of conditions such as frequently relapsing or steroid-dependent nephrotic syndrome. However, there is no direct mention or analysis of NSAIDs or their effects related to nephrotic syndrome in children.\n\n3. Analyze the knowledge: Most of the knowledge focuses on immunosuppressants rather than NSAIDs. The content discusses the effectiveness of drugs such as rituximab, cyclophosphamide, and levamisole for nephrotic syndrome in children. It includes a comparison of different treatment options, but does not specifically focus on the role of NSAIDs. \n\n4. Supporting Evidence: As part of the knowledge base provided, there is no direct mention of NSAID use, nor any clinical trial reports or meta-analyses linking NSAID use to secondary nephrotic syndrome in children. \n\n5. Concluding Reasoning: As the knowledge base does not mention or refer to NSAIDs, nor does it provide data or findings related to NSAIDs causing secondary nephrotic syndrome in children, it cannot be said that the information provided here supports the statement. \n\nRestate the statement: \nInvestigate whether there are clinical trials or meta-analyses reporting a direct causal relationship between NSAID use and secondary nephrotic syndrome in children. \n\nFinal Answer: Based on the knowledge provided, there is no information to support or substantiate the statement regarding NSAID use and its causal relationship with secondary nephrotic syndrome in children. [Not Supported]
[0141] References
[0142] 1.https: / / pmc.ncbi.nlm.nih.gov / articles / PMC6709258 /
[0143] 2.https: / / pubmed.ncbi.nlm.nih.gov / 25224196 /
[0144] …
[0145] Split atomic view 2:
[0146] Search for case studies or medical records documenting nephrotic syndrome in children after taking acetaminophen [unconfirmed]
[0147] Analytical rationale for the split atomic view 2
[0148] The input statement suggested exploring case studies or medical records of children who developed nephrotic syndrome after taking acetaminophen. However, all sub-claims provided regarding pharmacokinetics, renal toxicity profiles, and association with nephrotic syndrome in the pediatric population were labeled “unsubstantiated.” This indicates a lack of evidence or definitive findings to support these individual assessments. Therefore, there was insufficient evidence to directly link acetaminophen use to nephrotic syndrome in children, rendering the primary statement unsubstantiated.
[0149] References
[0150] 1.https: / / pmc.ncbi.nlm.nih.gov / articles / PMC89104 /
[0151] 2.https: / / pmc.ncbi.nlm.nih.gov / articles / PMC4377508 /
[0152] 3.https: / / pmc.ncbi.nlm.nih.gov / articles / PMC7007944 /
[0153] 4.…
[0154] Split atomic view 3:
[0155] Review of pharmacovigilance databases for reports of nephrotic syndrome in pediatric patients associated with over-the-counter analgesic use [×]
[0156] Analytical reasons for the split atomic view 3
[0157] The input statement requested an examination of the effects of common analgesics on renal function in children and the identification of mechanisms by which they may induce nephrotic syndrome. The sub-statement provided limited evidence regarding the effects of NSAIDs on renal function, suggesting that renal hemodynamics may change, but this was not clearly linked to nephrotic syndrome. In addition, the sub-statement discussed the known effects of acetaminophen, but did not find any association with nephrotic syndrome through oxidative stress or tubular injury. Furthermore, there was no evidence that ibuprofen had an effect on proteins associated with nephrotic syndrome in children. Therefore, the sub-statement did not support the idea that these analgesics induce nephrotic syndrome through any of the implicated mechanisms.
[0158] References
[0159] 1.https: / / www.kidney.org / kidney-topics / nephrotic-syndrome 2.
[0161] https: / / www.kidneyfund.org / all-about-kidneys / other-kidney-problems / nephrotic-syndrome-treatments-causes-symptoms
[0162] 3.https: / / pubmed.ncbi.nlm.nih.gov / 29169714 /
[0163] 4.https: / / pmc.ncbi.nlm.nih.gov / articles / PMC3163400 /
[0164] 5.…
[0165] Analysis reasons for the split sub-viewpoint 3
[0166] The subclaims provided do not support the input claim that commonly used analgesics can cause secondary nephrotic syndrome in children. The subclaims overall fail to establish a causal relationship between the use of commonly used analgesics such as NSAIDs or acetaminophen and the development of secondary nephrotic syndrome in children. There is no direct evidence, documented cases, or clinical studies to demonstrate such a link. In addition, there are no detailed reports from pediatric nephrologists indicating that the incidence of nephrotic syndrome has been observed in children after the use of these medications. Therefore, the claim should be rejected.
[0167] Analytical rationale for the extracted key insights
[0168] The entered statement, “Certain medications, such as acetaminophen, paracetamol, or other common pain relievers, can cause secondary childhood nephrotic syndrome” is not supported by the evidence provided. While some sub-claims suggest a possible link between certain medications and kidney-related problems, such as minimal change disease or acute interstitial nephritis, there is no solid evidence or direct cause-and-effect relationship for secondary childhood nephrotic syndrome. Additionally, no sub-claim provides clinical studies, case reports, or written evidence that definitively links acetaminophen, paracetamol, or other common pain relievers to this condition in children. Therefore, the statement should be rejected.
[0169] Example 2 of the ITA method is shown below:
[0170] Medical Text 2:
[0171] Aplastic anemia is a disease in which the bone marrow cannot produce enough blood cells. Various factors can increase your risk of developing aplastic anemia:\n\n1. **Genetic disorders**: People with certain genetic disorders, such as Fanconi anemia and dyskeratosis congenita, are at higher risk.\n\n2. **Exposure to toxins**: Long-term exposure to industrial chemicals such as benzene can increase your risk.\n\n3. **Radiation and chemotherapy**: People who receive radiation or chemotherapy to treat cancer may experience bone marrow damage.\n\n4. **Use of certain medications**: Certain medications, including antibiotics and nonsteroidal anti-inflammatory drugs (NSAIDs), have been linked to the development of aplastic anemia.\n\n5. **Autoimmune disorders**: In diseases such as lupus or rheumatoid arthritis, the body's immune system attacks its own cells, so the bone marrow can also be affected.\n\n6. **Viral infections**: Certain viruses, such as hepatitis, Epstein-Barr virus, and HIV, are associated with an increased risk.\n\n7. **Unknown factors**: In many cases, the exact cause is unknown, a condition called idiopathic aplastic anemia. \n\n8. **Age and Gender**: Although aplastic anemia can occur at any age, the highest incidence is in adolescents and older adults, with a slightly higher incidence in males. \n\nPeople concerned about their risk for aplastic anemia should consult a healthcare provider for a comprehensive evaluation based on their personal health history and potential exposures.
[0172] Extracted key insights
[0173] Use of certain medications, such as antibiotics and nonsteroidal anti-inflammatory drugs (NSAIDs), can increase your risk of developing aplastic anemia.
[0174] Split sub-viewpoint 1
[0175] Antibiotic use increases the risk of developing aplastic anemia.
[0176] Analysis of reasons
[0177] To determine whether the statement "antibiotic use increases the risk of developing aplastic anemia" is supported by the knowledge provided, let this disclosure break down the information about aplastic anemia and antibiotic use:\n\n1.**Aplastic anemia and infection**: Aplastic anemia (AA) is a disease characterized by pancytopenia and bone marrow cytopenia. It is associated with exposure to drugs, chemicals and toxins, as well as certain viral infections.\n\n2.**Medications and risk factors**: There is evidence that certain drugs are associated with the development of agranulocytosis and aplastic anemia. Specifically, some anti-infective drugs have been shown to be associated with these diseases.\n\n3.**Knowledge point - Risks associated with drugs**: A study conducted in Europe and Israel found that the use of anti-infective drugs that were significantly associated with agranulocytosis (a symptom of aplastic anemia involving a severe decrease in neutrophils) included trimethoprim / sulfamethoxazole, with a relative risk of 12, and macrolides with an uncertain relative risk. The relative risk was also elevated for sulfonamides without trimethoprim, although it was not statistically significant (3.6; 0.7 to 18). \n\n4. **Causes and Drug Exposure**: Another snippet in the knowledge states that AA may be triggered by exposure to specific drugs, although many cases remain idiopathic or have no known cause. \n\n5. **Knowledge Conclusion**: The knowledge suggests that anti-infective drugs such as trimethoprim / sulfamethoxazole, and possibly other sulfonamides, are associated with an increased risk of agranulocytosis, a condition that may co-occur with or cause aplastic anemia. This association highlights a direct link between certain antibiotics and the risk of aplastic anemia. \n\nIn summary, the knowledge points reveal a clear association between certain antibiotics, particularly anti-infective drugs such as trimethoprim / sulfamethoxazole and macrolides, and an increased risk of conditions that are part of or may cause aplastic anemia. \n\n**Recapitulation**: Antibiotic use increases the risk of aplastic anemia. \n\n**Final Answer**: Based on the discussed evidence that specific antibiotics are associated with an increased risk of aplastic anemia-related disorders, this statement is [Supported].
[0178] References
[0179] https: / / ashpublications.org / blood / article / 140 / Supplement%201 / 10784 / 490087 / Infectious-Disease-Antimicrobial-Prophylaxis
[0180] https: / / academic.oup.com / cid / article / 52 / 6 / 726 / 363030
[0181] https: / / www.ncbi.nlm.nih.gov / pmc / articles / PMC6280474 /
[0182] https: / / jamanetwork.com / journals / jamainternalmedicine / fullarticle / 611577
[0183] https: / / pmc.ncbi.nlm.nih.gov / articles / PMC6443538 /
[0184] https: / / ashpublications.org / blood / article / 140 / Supplement%201 / 5037 / 491342 / Fever-and-Neutropenia-Management-and-Infectious
[0185] https: / / pmc.ncbi.nlm.nih.gov / articles / PMC3138728 /
[0186] https: / / pubmed.ncbi.nlm.nih.gov / 2785785 /
[0187] https: / / www.sciencedirect.com / science / article / pii / S0037196309000638
[0188] https: / / www.researchgate.net / publication / 359347353_Antimicrobial_Therapy_An_Important_Risk_for_Acquired_Aplastic_Anemia
[0189] Split sub-viewpoint 2
[0190] Use of nonsteroidal anti-inflammatory drugs (NSAIDs) increases the risk of developing aplastic anemia.
[0191] Analysis of reasons
[0192] **Step-by-step reasoning:**\n\n1.**Understand the statement:** The statement claims that the use of NSAIDs increases the risk of developing aplastic anemia. Aplastic anemia is a disease characterized by pancytopenia (decreased red blood cells, white blood cells, and platelets) and bone marrow cytopenia, which may be caused by drug exposure, among other reasons.\n\n2.**Summarize key knowledge points**\n-Knowledge discusses various blood system diseases caused by drugs, including aplastic anemia.\n-Mentions that aplastic anemia can be caused by drugs, but can also be idiopathic (without a known cause).\n-Drugs that cause aplastic anemia include chloramphenicol, gold, NSAIDs, and others.\n-Knowledge discusses that many NSAIDs, such as diclofenac and sulindac, can cause various blood system syndromes, including aplastic anemia, although the direct evidence in individual case reports is not always clear.\n-NSAIDs have been associated with certain blood system diseases, including aplastic anemia, as noted in the discussion of drug-induced blood diseases. \n-There are studies and case reports exploring the association between NSAIDs and hematologic disorders, although the number of cases directly associated with NSAIDs is likely insignificant compared to their widespread use. \n\n3.**Analysis Support for Statement**\n-Aplastic anemia is listed as one of the potential hematologic complications caused by NSAIDs. \n-There are some case reports and references suggesting that NSAIDs may cause aplastic anemia, although most of the evidence is based on case reports or observational reviews. \n\n4.**Assessment of Prevalence and Impact**\n-While there are indications that NSAIDs may be associated with aplastic anemia, knowledge suggests that this is not a common finding given their widespread use and low reported incidence. \n-The association appears to exist but may not be strong for all NSAIDs or conclusive in every case. \n\n**Statement Restatement**\nThe use of nonsteroidal anti-inflammatory drugs (NSAIDs) is associated with an increased risk of aplastic anemia. \n\n**Final Decision**\nBased on the knowledge provided, there is evidence that NSAIDs may be associated with aplastic anemia, although it is not a widespread adverse effect given the high prevalence of NSAID use and the relatively low incidence of reported cases. \n\n**Final Answer**\n[Support]
[0193] References
[0194] https: / / www.shebaonline.org / patient-knowledge-base / question / patient-knowl edge-base-what-happens-if-aplastic-anemia-goes-untreated /
[0195] https: / / www.ncbi.nlm.nih.gov / pmc / articles / PMC1663360 /
[0196] https: / / pmc.ncbi.nlm.nih.gov / articles / PMC2778502 /
[0197] https: / / www.ncbi.nlm.nih.gov / books / NBK534212 /
[0198] https: / / jamanetwork.com / journals / jamainternalmedicine / fullarticle / 617829
[0199] https: / / pmc.ncbi.nlm.nih.gov / articles / PMC3201839 /
[0200] https: / / emedicine.medscape.com / article / 816117-overview
[0201] https: / / www.uptodate.com / contents / nonselective-nsaids-overview-of-advers e-effects
[0202] https: / / e-century.us / files / ijcem / 11 / 6 / ijcem0068128.pdf
[0203] https: / / www.sciencedirect.com / science / article / abs / pii / S0009279702000522
[0204] Analytical rationale for the extracted key insights
[0205] The input argument states that the use of certain medications, such as antibiotics and nonsteroidal anti-inflammatory drugs (NSAIDs), may increase the risk of developing aplastic anemia. An evaluation of the individual subarguments supports this argument. Subargument 1 states that specific antibiotics, particularly trimethoprim / sulfamethoxazole and macrolide antibiotics, are associated with an increased risk of aplastic anemia or related hematologic disorders, which has been confirmed by clinical studies and epidemiological data. Subargument 2 shows that NSAIDs are also associated with aplastic anemia, and although this risk is lower, case reports and clinical reviews suggest a potential risk. Overall, the evidence collectively supports the input argument that the use of antibiotics and NSAIDs may indeed increase the risk of developing aplastic anemia.
[0206] The final output content can be found in Figure 1 The right side of the display mode lists the extracted key viewpoints one by one, and indicates whether the analysis result of each key viewpoint is accepted or rejected. Accordingly, the evidence data is output for each key viewpoint.
[0207] The method for evaluating large language model (LLM) output according to an embodiment of the present disclosure may also use a LLM-as-a-Judge based checklist score to systematically evaluate key viewpoints.
[0208] For example, the method also includes constructing a test dataset (e.g., Med-Critics), wherein the benchmark test dataset includes a set of correct opinions, and a set of false opinions, factual texts, and non-factual texts; and using the constructed test dataset to evaluate the evaluation results of each of the key opinions.
[0209] Specifically, constructing a test data set also includes:
[0210] Extract a series of atomic viewpoints as correct viewpoints from a sentence in a given medical guide text;
[0211] Add random errors to correct ideas to generate fake ideas;
[0212] Generate paraphrases of the original text based on the correct and falsified views as factual text;
[0213] Creating an alternative version of a factual text by incorporating falsified opinions as a non-factual text.
[0214] In certain embodiments, the categories of test data in the test data set include the following: pathophysiology, medication, diagnosis, symptoms, treatment, and prevention.
[0215] Specifically, long-form factuality assessment is challenging due to the difficulty in defining a set of deterministic facts. To address this issue, we develop a fine-grained benchmark with specific viewpoint modifications to evaluate the performance of ITA methods.
[0216] For example, the test dataset Med-Critics is constructed by systematically extracting opinions from medical guideline texts, forging subsets of these opinions, and generating factual and non-factual text pairs. A subset of the MedQuAD dataset can be used to select 20-30 sentence fragments from medical text. MedQuAD is derived from various websites of the National Institutes of Health and includes real-world medical QA pairs covering 37 question types on topics including treatment, diagnosis, and side effects. The construction of Med-Critics involves the following steps:
[0217] 1. Opinion Extraction: Given a passage in a medical text, LLM extracts a series of atomic opinions, breaking down the content into basic factual components.
[0218] 2. Opinion falsification: In order to introduce controlled inaccuracies, one of the extracted opinions is intentionally falsified. This involves adding random errors, such as misleading opinions or distortions of key information, to simulate real misinformation scenarios.
[0219] 3. Text Paraphrase: Using two sets of opinions (factual and falsified), LLM generates paraphrases of the original text that maintain the factual integrity of the opinions.
[0220] 4. Alternative text generation: Creating alternative versions of texts by incorporating fabricated viewpoints, thereby producing non-factual narratives for evaluation purposes.
[0221] This approach facilitates the development of a comprehensive dataset designed to evaluate the ability of models to distinguish factual and non-factual information in medical texts. To make the benchmark more effective in finding factual errors across multiple dimensions, the test data can be divided into six main categories:
[0222] Pathophysiology: covers the biological and physiological processes behind disease or injury, providing a framework for understanding disease mechanisms.
[0223] ·Pharmacotherapy: Focuses on drug therapy, including drug interactions, side effects, and therapeutic effects.
[0224] Diagnosis: involves the identification and classification of diseases, with an emphasis on diagnostic criteria and methods.
[0225] Symptoms: Describe the clinical presentation of the disease, detailing the symptoms and their relevance to the specific medical condition.
[0226] Treatment: involves medical and surgical interventions aimed at controlling or curing disease.
[0227] Prevention: Explore strategies and measures to prevent disease onset or recurrence, including lifestyle changes and preventive treatments.
[0228] Table 1 shows the statistics of the Med-Critics benchmark. The average length is measured in terms of text, while the positivity rate is defined by the proportion of opinions that are factually correct.
[0229]
[0230] Evaluation Setup
[0231] This paper adopts a fine-grained evaluation method to evaluate the factuality of long medical texts, which involves evaluating the factuality of each atomic-level fact in the input text and reporting the overall performance of the method on a test dataset.
[0232] Baselines
[0233] This disclosure evaluates the performance of the ITA method of the present disclosure against several state-of-the-art fact consistency assessment baselines.
[0234] Model performance.
[0235] The disclosure also evaluates the performance of the standard baseline LLM in terms of reliability accuracy when processing long medical texts. The disclosure's evaluation includes experiments on a subset of the Med-Critics open medical QA challenge. These datasets are specifically designed to test the model's ability to handle complex medical information, handle subtle reasoning, and generate accurate, context-sensitive responses.
[0236] index
[0237] To evaluate the performance of the disclosed method in assessing the authenticity of long medical texts, the disclosed method uses several key metrics. Accuracy measures the difference between the facts and the factual verification, providing a basic accuracy assessment. The disclosed method also uses the F1@K metric, which evaluates precision and recall. This metric provides a comprehensive understanding of the factual accuracy of the model.
[0238] Key results
[0239] To emphasize the reliability of the ITA method proposed in this disclosure, this disclosure focuses on two key aspects: accurately identifying all incorrect opinions and correctly evaluating each opinion. However, the inherent unpredictability of generative models, coupled with the lack of clear rules for determining the number and nature of opinions, makes this task particularly challenging. This ambiguity complicates the assessment of whether the judging system operates fairly. To address this issue, this disclosure utilizes the checklist scores of LLM-as-a-Judge to systematically evaluate opinions. In addition, this disclosure addresses these challenges by introducing the Med-Critics benchmark, which provides a set of pre-defined incorrect opinions with varying numbers for rigorous evaluation.
[0240] ITA method consistency with predefined truth
[0241] In Table 2, the present disclosure presents the fact verification accuracy for different categories of facts. The results show that ITA achieves the highest accuracy in all categories. The key performance gap can be attributed to two main factors: ITA's ability to extract independent opinions and its ability to verify these opinions through comprehensive analysis. Baseline methods usually focus on decomposing opinions into atomic units for verification. However, the results show that many failures stem from the lack of sufficient contextual information. In the medical field, facts out of context may appear accurate, but in fact they are not. The main advantage of ITA lies in the balance between contextual basis and atomicity.
[0242] During the opinion verification process, ITA emphasizes the context of the complete subtree, thereby achieving balanced granularity and deeper analysis of external scientific knowledge. This approach allows evaluators to determine whether any sub-opinion is false and assess whether the original opinion can still be supported. By maintaining this balance, ITA significantly improves the overall performance of fact-checking.
[0243]
[0244] Table 2 shows the performance comparison of the actual verification on the test dataset Med-Critics.
[0245] In addition, this disclosure also verifies the impact of opinion extraction.
[0246] In order to rigorously evaluate the impact of initial opinion extraction on the ITA method validation process, the present disclosure conducted a series of controlled experiments. In these experiments, the extracted opinions were fixed while rerunning the ITA subtree validation. This setup enables the present disclosure to isolate and analyze the specific impact of opinion extraction on the overall performance and accuracy of the system. As shown in Table 3, the present disclosure tested three baseline opinion extraction methods: ATOMIC, DECONTEXT, and MED-DECONTEXT. The results show that the fine-tuned opinion extractor in the ITA method consistently outperforms these baselines when validating the target medical text.
[0247]
[0248]
[0249] Table 3 Opinions on cancelling financing
[0250] Human evaluation
[0251] To evaluate the reliability of ITA in processing long medical texts, we conducted a manual evaluation study involving three medical experts. The experts evaluated the reliability of the ITA-labeled opinion verification based on the Med-Critics benchmark. We randomly selected a sample of 100 opinions for evaluation, which were classified as "accepted" or "rejected" by ITA. Each expert evaluated the authenticity of the opinion and was allowed to conduct an internet search to verify its accuracy. Figure 4 As shown, the results showed a high degree of agreement between the experts’ assessments and the ITA validation records. Although some opinions could not be fully verified due to the ambiguity of medical information or lack of consensus in the academic community, the number of unverified opinions was small. These findings highlight the robustness of ITA in accurately assessing factuality in complex medical texts.
[0252] Performance of Large Language Models (LLM) on Long Medical Texts
[0253] The present disclosure provides a more comprehensive evaluation toolbox for the performance of large language models (LLM) on long medical texts and fine-grained medical dimensions of medical texts, as shown in Table 1. For these reasons, the present disclosure benchmarks four model families (LLMModel 1 (e.g., GPT-3.5), LLM Model (e.g., GPT-4) 2, Claude, and Qwen models). The present disclosure evaluates each model on the same random subset of 240 prompts in the Med-Critics benchmark and reports the validation result statistics in Table 4. The results show that LLM Model 2 and Claude tend to speak more and have higher accuracy.
[0254] Figure 5The radar chart in a visually compares the performance of the models on different dimensions of medical text. The results show that LLM Model 2 exhibits the highest precision and recall, followed by Claude, Qwen, and LLM Model 1, especially in the symptom and diagnosis dimensions. Figure 5 As shown in (b), LLM Model 2 consistently outperforms other models in the F1@K metric on the Med-Critics benchmark.
[0255] Therefore, according to the ITA method of the embodiment of the present disclosure, medical fact verification can be enhanced by using adaptive mind tree reasoning. The ITA method can effectively extract atomic viewpoints from the original text and construct an evidence tree to support true and false judgments, thereby improving the accuracy and reliability of viewpoint verification. The ITA method can also provide a unified and universal framework for medical viewpoint detection by generating and merging subtrees using retrieved external reference information, which allows comprehensive description and verification of different important viewpoints in the input query text, thereby facilitating a more meticulous and detailed analysis of medical information. In addition, the embodiment of the present disclosure also constructs a verification data set, including a fine-grained checklist that can support the evaluation of verification tasks.
[0256] Figure 6 is a block diagram illustrating a device 600 according to an embodiment of the present disclosure.
[0257] See also Figure 6 , the device 600 may include a processor 601 and a memory 602. The processor 601 and the memory 602 may be connected via a bus 603.
[0258] The processor 601 can perform various actions and processes according to the program stored in the memory 602. Specifically, the processor 601 can be an integrated circuit chip with signal processing capabilities. The above-mentioned processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components. The various methods, steps and logic block diagrams disclosed in the embodiments of the present application can be implemented or executed. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc., which can be an X86 architecture or an ARM architecture.
[0259] The memory 602 stores computer instructions, which are executed by the processor 601 to realize the above combination. Figure 1 Any of the above methods, combined Figure 2Any of the methods described or methods for training a large language model. Memory 602 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. Volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous connection dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory of the method described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0260] According to yet another embodiment of the present disclosure, a computer-readable recording medium is also provided. Figure 7 A schematic diagram of a recording medium 700 according to an embodiment of the present disclosure is shown.
[0261] like Figure 7 As shown, the computer recording medium 700 stores a computer executable instruction 710. When the computer executable instruction 710 is executed by the processor, the above-mentioned combination Figure 1 Any of the methods 7 described. The computer-readable recording medium in the disclosed embodiments may be a volatile memory or a nonvolatile memory, or may include both volatile and nonvolatile memories. The nonvolatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous connection dynamic random access memory (SLDRAM), and direct memory bus random access memory (DR RAM). It should be noted that the memory of the method described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0262] Another embodiment of the present disclosure further provides a computer program product, which includes computer executable instructions. When the computer executable instructions are executed by a processor, the processor is prompted to perform the above-mentioned combination Figure 1 Any of the methods 7.
[0263] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the methods, devices, media, equipment and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of a code contains at least one executable instruction for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.
[0264] In general, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Certain aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that may be executed by a controller, microprocessor, or other computing device. When various aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flow charts, or using some other graphical representation, it will be understood that the blocks, devices, systems, techniques, or methods described herein may be implemented in hardware, software, firmware, dedicated circuits or logic, general purpose hardware or controllers or other computing devices, or some combination thereof as non-limiting examples.
[0265] The exemplary embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. It should be understood by those skilled in the art that various modifications and combinations may be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A method for evaluating the output of a large language model, comprising: extracting one or more key ideas from an output of a large language model, the output of the large language model comprising a long medical text generated in response to an input query, and each of the key ideas comprising a self-consistent factual statement; Construct an adaptive thinking tree based on the extracted key ideas to split each key idea into multiple verifiable sub-ideas; retrieving evidence data associated with the plurality of sub-viewpoints from one or more external information sources, and verifying the plurality of sub-viewpoints based on the retrieved evidence data; In response to the verification of the plurality of sub-viewpoints being completed, integrating the verification results of the plurality of sub-viewpoints, and determining the evaluation results of the corresponding key viewpoints according to the integrated verification results of the plurality of sub-viewpoints; as well as The evaluation results of each of the key viewpoints and the corresponding evidence data are output.
2. The method according to claim 1, wherein: The key points are independent of each other, and The root node in the structure graph of the thinking tree is the key concept, the child nodes of the tree structure are the child concepts, and the edges in the tree structure graph indicate the relationship between the nodes.
3. The method according to claim 1, further comprising: In response to the verification of the sub-concept indicating that the sub-concept has a low granularity, iteratively further splitting the sub-concept into a plurality of atomic concepts with higher granularity; Evidence data associated with the plurality of atomic viewpoints is retrieved from one or more external information sources, and the plurality of atomic viewpoints is verified based on the retrieved evidence data.
4. The method according to claim 3, wherein: In response to the verification of the plurality of sub-viewpoints being completed, the method further includes: It is determined that the plurality of sub-views do not need to be further split or that the sub-views have reached a maximum granularity.
5. The method according to claim 3, wherein: Integrating the verification results of the plurality of sub-viewpoints, and determining the evaluation results of the corresponding key viewpoints according to the integrated verification results of the plurality of sub-viewpoints further comprises: Integrate the verification results of the multiple atomic viewpoints corresponding to each of the multiple sub-viewpoints, and determine the verification result of each corresponding sub-node according to the integrated verification results of the multiple atomic viewpoints; The evaluation result of the corresponding key viewpoint is determined according to the verification results of the integrated multiple sub-viewpoints.
6. The method according to claim 2, wherein: Each node includes initial parameters including a viewpoint and a scalar value indicating whether the viewpoint is acceptable, and Integrating the verification results of the plurality of sub-viewpoints, and determining the evaluation results of the corresponding key viewpoints according to the integrated verification results of the plurality of sub-viewpoints further comprises: Starting from a leaf node of the tree structure graph, the scalar value of the verified sub-viewpoint and the evidence data providing the judgment are passed to the parent node; Combining the scalar values of multiple child nodes to modify the state of the parent node; and The state of the root node is modified using the updated state of the parent node, and the evaluation result of the key view is determined based on the state of the root node.
7. The method according to claim 1, wherein: The large language model includes a set of pre-set search tools, each of which is configured to retrieve a specific type of information or call an external calculator for evaluation, and For each sub-viewpoint, an appropriate search tool is selected to search for evidence data associated with the plurality of sub-viewpoints.
8. The method according to claim 7, wherein: Selecting an appropriate search tool to retrieve the evidence data associated with the plurality of sub-viewpoints includes: Use the child view's parent's view as context to determine the appropriate search tool; generating a search query task to initiate a search for evidence data associated with the plurality of sub-viewpoints; and The retrieved evidential data are ranked and selected based on their relevance to the viewpoint and the origin and type of information sources.
9. The method according to claim 1, further comprising: Key ideas were systematically evaluated using a checklist score based on the LLM-as-a-Judge.
10. The method according to claim 9, further comprising: Constructing a test data set, wherein the benchmark test data set includes a set of correct opinions, and a set of forged opinions, factual texts, and non-factual texts; The evaluation results of each of the key viewpoints are evaluated using the constructed test dataset.
11. The method according to claim 10, wherein: Constructing the test data set also includes: Extract a series of atomic viewpoints as correct viewpoints from a sentence in a given medical guide text; Add random errors to correct ideas to generate fake ideas; Generate paraphrases of the original text based on the correct and falsified views as factual text; Creating an alternative version of a factual text by incorporating falsified opinions as a non-factual text.
12. The method according to claim 11, wherein: The categories of test data in the test data set include the following pathophysiology, drug therapy, diagnosis, symptoms, treatment, and prevention.
13. An apparatus for evaluating the output of a large language model, comprising: processor, and A memory storing computer executable instructions, which, when executed by a processor, cause the processor to perform the method according to any one of claims 1 to 12.
14. A computer-readable recording medium storing computer-executable instructions, wherein: The computer executable instructions, when executed by a processor, cause the processor to perform the method of any one of claims 1-12.
15. A computer program product comprising computer executable instructions, wherein: The computer executable instructions, when executed by a processor, cause the processor to perform the method of any one of claims 1-12.
Citation Information
Cited By
Circuit design verification method, device, equipment, medium and product
CN121351730A
Circuit design verification method, apparatus, device, medium, and product
CN121351730B