Method of extracting data from written articles and associated system
A multi-stage LLM process with a user interface linking answers to evidence addresses the reliability issues of LLMs, enhancing data extraction efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2026-04-09
AI Technical Summary
Existing computer-assisted data extraction systems, particularly those using large language models (LLMs), lack reliability and require extensive manual validation to ensure accuracy, which undermines their efficiency in high-stakes environments.
A multi-stage process involving a first LLM for query expansion and a second LLM constrained to a focused context, combined with a user interface that links answers to their evidence, enhances reliability and verifiability.
This approach improves the reliability and efficiency of data extraction by providing transparent, instantly verifiable answers, bridging the gap between automated and manual validation.
Smart Images

Figure CA2025051292_09042026_PF_FP_ABST
Abstract
Description
METHOD OF EXTRACTING DATA FROM WRITTEN ARTICLES AND ASSOCIATED SYSTEMTechnical Field
[0001] The present disclosure relates to the field of research articles review, and in particular to the extraction of data and evidence from the research articles.Background
[0002] Computer-assisted extraction of data from articles is undergoing promising advances in concert with advances being made in artificial intelligence (Al) and large language models (LLMs). While accuracy is important in an answer-suggestion system, users must validate every answer suggestion because of regulatory constraints, lack of trust in AI / LLM systems, and / or internal policies. When manual validation is time-consuming, for example, when it requires reading the entire document, the value of AI / LLM data extractions systems is limited, and in cases where an incorrect answer is provided, can be net-negative. On the other hand, when an incorrect answer can be quickly identified as incorrect in can help speed up the time to finding the correct answer. Regardless of the advances being made, any reliable method of extracting data from articles still needs a human to verify the accuracy of the extracted data by reading a portion, or the entirety of the article being reviewed.
[0003] Therefore, advances in methods and systems for extracting data from written articles are desirable, particularly those that reduce the manual validation burden.Summary
[0004] In a first aspect of the disclosure, there is provided a method of extracting data from an article, the method comprising: providing a question to a large language model (LLM) system, the question being in relation to the article; obtaining, from the LLM system, an answer to the question; and displaying a user interface, the user interface showing: the question; the answer to the question; and an indication of where in the article there is evidence for the answer. Advantageously, this aspect is user efficiency and trust when interacting with an Al (LLM) data extraction tool. By displaying the Al-generated answer and a direct, interactive link to its supporting evidence within a single user interface, the invention solves the critical problem of validating Al outputs. This integration allows a user to instantly verify the factual basis of an answer without needing to manually search the entire source document, thereby transforming the system into a more effective and trustworthy data review tool.
[0005] In an embodiment, the user interface may display a control that, when actuated, causes the user interface to show where in the article there is the evidence for the answer.
[0006] In an embodiment, the user interface may show the indication of the section of the article where there is evidence for the answer by performing at least one of: displaying an as-published version the article, at a section of the article that supports the answer; displaying a section of the article that supports the answer; and displaying coordinates of where in the article there is support for the answer.
[0007] In an embodiment, the coordinates may include a page number of the article.
[0008] In an embodiment, the user interface may be configured to show the page number as part of a page number button that, when actuated, causes the user interface to show a page associated with the page number or a section of the page associated with the page number.
[0009] In an embodiment, displaying the as-published version the article, at a section of the article that supports the answer may include marking one or more portions of the section that supports the answer; and displaying the section of the article that supports the answer may include marking one or more portions of the section that supports the answer.
[0010] In an embodiment, marking the one or more portions may include highlighting the one or more portions.
[0011] In an embodiment, the user interface may include a confirmation button that when actuated, causes the answer to the question and the evidence for the answer to be attached to the question.
[0012] In an embodiment, the user interface may show a plurality of indications of a respective plurality of sections of the article where there is evidence for the answer.
[0013] In an embodiment, the user interface may include a confirmation button that when actuated, causes the answer to the question and the sections of the plurality of sections where there is evidence for the answer to be attached to the question.
[0014] In an embodiment, the user interface may include a confirmation button that when actuated, subsequent verification that at least one section of the plurality of sections of the article supports the answer, tags the at least one section as being verified evidence of the answer to the question.
[0015] In an embodiment, the user interface may include an answer entry box; and a confirmation button that when actuated, subsequent verification that at least one section of the plurality of sections of the article supports the answer, copies the answer to answer entry box to display the answer in the answer entry box.
[0016] In an embodiment, the user interface may includes a confirmation button that, when actuated, subsequent verification that at least one section of the plurality of sections of the article supports the answer causes the at least one section to be saved as being verified evidence of the answer to the question. The at least one section may be saved to an evidence storage.
[0017] In an embodiment, the user interface may include a storage button that, when actuated, causes the plurality of indications of the respective plurality of sections of the article that supports the answer to be saved to an evidence storage and tagged as being associated with the question.
[0018] In another aspect of the disclosure, there is provided a method of extracting data from an article, the method comprising: providing a question to a large language model (LLM) system, the question being in relation to the article; obtaining, from the LLM system, an answer to the question, the answer being an LLM answer; obtaining, from curated evidence associated with the written article, a curated answer to the question; and displaying a user interface, the user interface showing: the question; and at least one of: the LLM answer and an indication of where in the article there is evidence for the LLM answer; and the curated answer and an indication of where in the article there is evidence for the curated answer. An advantage of this aspect is that it provides a comprehensive data review environment that bridges the gap between fully automated Al (LLM) systems and manually verified data. By allowing a user to view and compare Al-generated suggestions alongside previously validated, human-curated answers, the invention supports workflows that require maintaining a record of information while still leveraging Al for new insights or faster initial reviews. This functionality allows for managing and validating Al outputs in a regulated or high-stakes environment where both speed and verified accuracy are important.
[0019] In an embodiment, the user interface may include: a question panel configured to display the question; a suggestions panel having a first configuration where the answer panel displays the LLM answer and the indication of where in the article there is evidence for the LLM answer, the evidence for the LLM answer being suggested evidence; and a second configuration where the answer panel displays the curated answer and the indication of where in the article there is evidence for the curated answer.
[0020] In an embodiment, the user interface may display a toggle button that, when actuated, toggles the answer panel between the first configuration and the second configuration.
[0021] In an embodiment, the user interface may include an evidence panel, the indication of where in the article there is evidence for the LLM answer includes a first control that, whenactuated, causes the evidence for the LLM answer to be displayed in the evidence panel, and the indication of where in the article there evidence for the curated answer is includes a second control that, when actuated, cause the evidence for the curated answer to be displayed in the evidence panel.
[0022] In an embodiment, the user interface may include an evidence management panel configured to display the question, the LLM answer to the question, and the indication of where in the article there is evidence for the LLM answer, the evidence management panel may include a manage evidence button that, when actuated, causes the user interface to display an evidence drawer that includes: the evidence for the LLM answer and an accept button that when activated accepts the evidence for the LLM answer as verified evidence; and confirmed evidence for the LLM answer.
[0023] In another aspect of the disclosure, there is provided a computer-implemented method of extracting data from a written article, the method comprising: obtaining a plurality of chunks from the written article; generating a respective chunk numerical representation for each of the plurality of chunks; generating, by a first large language model (LLM) in response to an initial question, one or more secondary questions related to the initial question; generating a respective question numerical representation for the initial question and for each of the one or more secondary questions; selecting, from the plurality of chunks, a subset of two or more chunks based on a comparison between the chunk numerical representations and the question numerical representations; generating a context from the subset of two or more chunks; generating, by a second LLM based on the context and the initial question, an answer to the initial question; and identifying, within the subset of two or more chunks, one or more portions of text that support the generated answer. An advantage of this aspect is the improved technical accuracy and reliability of the back-end process used for generating answers. The disclosure’s specific, unconventional architecture — using a first LLM for query expansion to better understand user intent and a second LLM for grounded answering constrained to a focused context — directly addresses the core technical problem of Al "hallucination." This results in a more factually accurate answer being generated before it is ever presented to the user, thereby improving the fundamental quality and trustworthiness of the data extraction process.
[0024] In an embodiment, generating the context may comprise concatenating the text of the subset of two or more chunks.
[0025] In an embodiment, the chunk numerical representations and the question numerical representations may be embeddings.
[0026] In an embodiment, obtaining the plurality of chunks from the written article may comprise performing a semantic search and retrieval operation.
[0027] In an embodiment, generating the one or more secondary questions may comprise providing the initial question to the first LLM using one-shot prompting.
[0028] In an embodiment, the first LLM and the second LLM may be the same LLM.
[0029] In an embodiment, the method may further comprise, prior to generating the answer by the second LLM, providing the context to at least one annotation model to generate an annotated context; and wherein generating the answer by the second LLM is based on the annotated context and the initial question.
[0030] In an embodiment, the at least one annotation model may be an expert model configured to detect and annotate instances of predetermined expressions of an expert technical field.
[0031] In an embodiment, the method may further comprise providing an instruction to the second LLM specifying an answer type, wherein the answer type is selected from the group consisting of a single answer type, a multiple answer type, and a free text answer type.
[0032] In an embodiment, identifying the one or more portions of text that support the generated answer may comprise at least one of: the second LLM reproducing the one or more portions of text; and the second LLM providing coordinates of the one or more portions of text with respect to a layout of the written article.
[0033] In another aspect of the disclosure, there is provided an intelligent document processing system, comprising: a processing module; a user interface; a tangible computer-readable medium having instructions recorded thereon, the instructions to be carried out by the processing module to perform any of the methods mentioned above.
[0034] In another aspect of the disclosure, there is provided a non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processor, cause the one or more processors to perform any of the methods mentioned herein.
[0035] Embodiments have been described above in conjunctions with aspects of the present disclosure upon which they can be implemented. Those skilled in the art will appreciate that embodiments may be implemented in conjunction with the aspect with which they are described, but may also be implemented with other embodiments of that aspect. When embodiments are mutually exclusive, or are otherwise incompatible with each other, it will be apparent to those skilled in the art. Some embodiments may be described in relation to one aspect, but may also be applicable to other aspects, as will be apparent to those of skill in the art.Brief Description of the Figures
[0036] Further features and advantages of the present invention will become apparent from the following detailed description, taken in combination with the appended drawings, in which:
[0037] FIG. 1 shows a flowchart of a method, in accordance with an embodiment of the present disclosure.
[0038] FIG. 2 shows a flowchart of a method of annotating context, in accordance with an embodiment of the present disclosure.
[0039] FIG. 3 shown an embodiment of an intelligent document processing system in accordance with the present disclosure.
[0040] FIG. 4 shows an embodiment of a user interface in accordance with the present disclosure.
[0041] FIG. 5 shows another embodiment of a user interface in accordance with the present disclosure.Detailed Description
[0042] Modern large language models (LLMs) show great promise for automating data extraction from documents, but they can from a lack of inherent reliability. An LLM can generate a fluent, convincing answer that is factually incorrect, a phenomenon often called "hallucination." In fields requiring high accuracy, a user must manually verify every piece of LLM-generated data by reading the source article, which can be so time-consuming that it negates the efficiency benefits of using the LLM in the first place. The core technical challenge, therefore, is to configure a computer system that can leverage the power of LLMs for data extraction while fundamentally improving the reliability and verifiability of the output.
[0043] The present disclosure addresses this challenge through a specific, multi-stage process that improves the computer's functionality. Rather than performing a simple search with a user's question, the system first employs an LLM process to analyze the initial question and generate a set of related, more granular secondary questions. This technique creates a richer and more comprehensive query. The system then converts both the source article and this expanded set of questions into numerical representations and compares them to select a small, highly relevant subset of text chunks from the article.
[0044] This selected subset of chunks creates a controlled, focused context. A second LLM process is then tasked with generating an answer, but it is technically constrained to use only the information within this provided context. This grounding step is critical as it prevents the LLM from relying on its general, and potentially incorrect, knowledge base. Finally, the system identifies the exact portions of text within that context that support the generated answer and presents them to the user through a purpose-built interface that provides a direct, interactive link between the answer and its supporting evidence. This transforms the computer from a simple questionanswering machine into a more effective and trustworthy data review tool, as the system is specifically architected to make the basis for its conclusions transparent and instantly verifiable by a human user.
[0045] The present disclosure provides a method that can facilitate the review of research articles (e.g., medical research articles) by suggesting answers to a question and showing the reviewer where in the articles there is support for the suggested answers. The present disclosure includes a system and method that presents to a reviewer of an article, a user interface that shows to the reviewer: (1) a review question, (2) a suggested answer, and (3) sections of the article that support the suggested answer. This allows the reviewer to establish quickly whether the answer is factual (reliable, true, correct, etc.) or not. Upon determining that the answer is factual, the user interface allows the answer to be tagged as a valid answer to the question and the sections of the articles that support the suggested answer to be tagged as verified evidence. Within the context of the present disclosure, a user interface may be referred to as a user experience (UX) interface.
[0046] FIG. 1 shows a flowchart of a method of extracting data from a written article, in accordance with an embodiment of the present disclosure. At action 20, a written article loaded into an intelligent document processing system is analyzed according to predetermined analysis criteria and one or more chunks of the written article are obtained. As a non-limiting example, the written article may be analyzed using any suitable semantic search and retrieval operation where the retrieval aspect provides the one or more chunks. In another non-limiting example, the written article may be subjected to any suitable similarity search operation to identify the one or more chunks. In another non-limiting example, the written article may be stored in a vector database and obtaining the one or more chunks of the written article may include performing a similarity search between a vector representation of the written article and a vector representation of analysis criteria. For example, the one or more chunks may correspond to logical divisions already present in the written article, such as paragraphs or numbered sections. The 'predetermined analysis criteria' may be one or more initial questions provided by a user.The 'similarity search operation' then serves to identify which of these logical chunks are most semantically relevant to the initial questions by comparing their respective vector representations within the vector database.
[0047] Action 20 may help identify portions of written article that may include answers to questions a reviewer of the written article would like answered. Alternatively, action 20 may be seen as a filter that ignores portions to the written article that are not likely to include the answers to the questions of the reviewer.
[0048] At action 22, a numerical representation (NR) of each of the one or more chunks is generated. A NR of text or of chunks of text may be generated using any suitable technique or method. As a non-limiting example, the NRs may be generated by generating embeddings of the written article. As another non-limiting example, the NRs may be generated by generating features of the written article. As a non-limiting example, the 'embeddings' may be generated using publicly known transformer-based models that are trained to produce vectors representing the semantic meaning of text. In some embodiments, the 'features' could be based on lexical statistics derived from the text of the one or more chunks.
[0049] At action 24, an initial question is provided to a first large language model (LLM), which can be any suitable LLM. The initial question is provided to the LLM for the LLM to generate secondary questions based on the initial question. The initial question will generally relate to the subject matter of the written article. As a non-limiting example, if the written article is related to a medical study, the initial question may be: “What tests were conducted in the study?” In some embodiments, the initial question may be provided to the first LLM using one-shot prompting where the prompt in the one-shot prompting includes an example of an initial question and of the resulting secondary questions. As an example, when the initial question is: Which of the following sources were obtained to help inform the risk of bias assessment?, the secondary questions generated by the first LLM may be:1. What tools were used to assess the risk of bias in the studies?2. Were any guidelines followed in the risk of bias assessment?3. What types of sources were used to inform the risk of bias assessment?4. Were any specific criteria used to select the sources for the risk of bias assessment?5. How were the sources used to inform the risk of bias assessment?
[0050] At action 26, secondary questions generated by the first LLM are obtained and at action 28, the NR of the initial question and the NRs of each of the secondary questions are obtained. At action 30, in accordance with the NR of the initial question and the NRs of secondary questions, two or more chunks that are relevant to the initial question and / or the secondary questions are selected from the one or more chunks obtained at action 20. In some embodiments, the selection of the two or more chunks may be based on a comparison of the initial question NR, the secondary questions NRs, and the NRs of the one or more chunks. In some embodiments, selecting the two or more chunks NRs may include ranking the two or more chunks NRs according to any suitable ranking criterion.
[0051] Advantageously, action 30 may increase the number of chunks among which the answer to the initial question may be found.
[0052] Action 30 may be seen as limiting the total number of chunks potentially relevant to all the questions the reviewer of the written article wishes to answer. Action 30 may be seen as expanding the number of chunks in which the answer to the initial question may be found.
[0053] At action 32, a context based on the two or more chunks associated with the two or more chunks NRs selected at action 30 is generated. In some embodiments, the context may be a concatenation of the two or more chunks.
[0054] At action 34, the initial question and the context are provided to a second LLM, which is configured to generate an answer to the initial question. The initial question and the context may be provided to the second LLM in the form of a prompt that includes the initial question, the context, the type of answer to be returned by the second LLM, and instructions to the second LLM. For instance, a portion of the prompt could include instructions such as: 'Based only on the provided context, generate an answer to the following question. Your answer should preferably be derived directly from the provided context.' In some embodiments, the type of answer may be a single answer type, a multiple answer type or a free text answer type or any other answer type the application offers.
[0055] When the type of answer is a single answer type, the instructions to the second LLM may include instructions to select a single answer from a group of possible answers. The group of possible answers may also be provided to the second LLM as part of the instructions.
[0056] When the type of answer is a multiple answer type, the instructions to the second LLM may include instructions to select a one or more answer from a group of possible answers. The group of possible answers may also be provided to the second LLM as part of the instructions.
[0057] When the type of answer is a free text answer type, the instructions to the second LLM may include instructions to generate, in reply to the initial question, a free text answer.
[0058] In some embodiments, the first LLM and the second LLM may be distinct LLMs. In other embodiments, the first LLM and the second LLM may be the same LLM.
[0059] At action 36, the chunks or portions of chunks that support the answer provided at action 34 are identified. Advantageously, this allows the user to verify rapidly if the answer provided at action 34 is correct, incorrect, or partially correct. In some embodiments, an indication of where in the written article there is support for the answer to the initial question includes at least one of: the second LLM reproducing a portion of the article that includes the support basis, the second LLM providing coordinates, with respect to a layout of the written article, of one or more excerpts of the written article that include the support basis, and the second LLM reproducing a section of the article and highlighting one or more parts of the section that include the support basis.
[0060] In some embodiments, action 20 may include obtaining the one or more chunks of the written article by using the initial question and the secondary questions. As a non-limiting example, the written article may be analyzed using any suitable semantic search and retrieval operation based on the initial question and on the secondary questions. In another non-limiting example, the written article may be subjected to any suitable similarity search operation based on the initial question and the secondary questions to identify the one or more chunks.
[0061] FIG. 2 shows flowchart of a method of annotating the context generated at action 32 of FIG. 1. In the embodiment shown in FIG. 2, the context is provided, at action 38, to an annotation model, which may be part of a class of extractive models. The annotation model may be configured to detect instances of predetermined expressions in the context and to annotate the context in accordance with the predetermined expressions. The annotated context may be obtained from the annotation model at action 40. The annotated context may then be provided, in lieu of non-annotated context, to the second LLM, at action 34 of FIG. 1. The instructions provided to the second LLM may be revised to specify that the context (the annotated context) includes annotations and that the second LLM is to consider the annotations when generatingthe response to the initial question. In some embodiments, the annotations may be in the form of tags such as, for example, <adverse_events>, <drug>, etc. The tags will depend on the expert models’ area of expertise.
[0062] As an example of how the method shown in the example of FIG. 2 may be applied, we can consider the following paragraphs (source: Sadykov et al. Biomed Pap Med Fac Univ Palacky Olomouc Czech Repub. 2024 Jun;168(2):147-155. doi: 10.5507 / bp.2023.026. Epub 2023 Jul 17), in their original form:Background: The aim of our study was to find a possible association between retinal microvascular abnormality and major depression in a non-geriatric population.Method: The participants with major depression were hospitalised at the University Hospital in Hradec Kralove, Department of Psychiatry. Retinal images were obtained using a stationary Fundus camera FF450 by Zeiss and a hand-held camera by oDocs.Results: Fifty patients (men n=18, women n=32) aged 16 to 55 (men's average age 33.7±9.9 years, women's average age 37.9±11.5 years) were compared with fifty mentally healthy subjects (men n=28, women n=22) aged 18 to 61 (men's average age 35.3±9.2 years, women's average age 36.6±10.6 years) in a cross-sectional design. The patients were diagnosed with a single depressive episode (n=26) or a recurrent depressive disorder (n=24) according to the ICD-10 classification. Our results confirmed significant microvascular changes in the retina in patients with depressive disorder in comparison to the control group of mentally healthy subjects, with significantly larger arteriolar (P<0.0001) as well as venular (P<0.001-0.0001) calibres in major depression.Conclusion: According to the literature, acute and chronic neuroinflammation is associated with changes in microvascular form and function. The endothelium becomes a major participant in the inflammatory response damaging the surrounding tissue and its function. Because the retina and brain tissue share a common embryonic origin, we suspect similar microvascular pathology in the retina and in the brain in major depression. Our results may contribute to a better understanding of depression etiopathogenesis and to its personalized treatment.
[0063] When the above text is processed in accordance with the embodiment of the method shown at FIG. 2, with an annotation model, or multiple annotation models, configured to detect and annotate instances of Adverse Events, Institutes, Patient Sizes, Male and Female Patient Ratios, Patient Ages, Men’s Average Ages, Women’s Average Ages, Study Types, Disease Classification Codes, and Pathological Findings, the resulting annotated text may be as follows:Background: The aim of our study was to find a possible association between retinal microvascular abnormality and <Adverse Event> major depression < / Adverse Event> in a non-geriatric population.Method: The participants with major depression were hospitalised at the <lnstitute> University Hospital in Hradec Kralove < / lnstitute>, Department of Psychiatry. Retinal images were obtained using a stationary Fundus camera FF450 by Zeiss and a hand-held camera by oDocs.Results: <Patient Size> Fifty patients <Male and Female Patient Ratio (men n=18, women n=32) < / Male and Female Patient Ratio < / Patient Size> aged <Patient Age> 16 to 55 < / Patient Age> (men's average age <Mens' Average Age> 33.7±9.9 years < / Mens' Average Age>, women's average age <Womens' Average Age> 37.9±11.5 years < / Womens' Average Age>) were compared with fifty mentally healthy subjects (men n=28, women n=22) aged 18 to 61 (men's average age 35.3±9.2 years, women's average age 36.6±10.6 years) in a <Study Type> cross-sectional < / Study Type> design. The patients were diagnosed with a single depressive episode (n=26) or a recurrent depressive disorder (n=24) according to the <Disease Classification Code> ICD-10 < / Disease Classification Code> classification. Our results confirmed significant microvascular changes in the retina in patients with depressive disorder in comparison to the control group of mentally healthy subjects, with significantly larger arteriolar (P<0.0001) as well as venular (P<0.001-0.0001) calibres in major depression.<Pathological Findings> Conclusion: < / Pathological Findings> According to the literature, acute and chronic neuroinflammation / s associated with changes in microvascular form and function. The endothelium becomes a major participant in the inflammatory response damaging the surrounding tissue and its function.Because the retina and brain tissue share a common embryonic origin, we suspect similar microvascular pathology in the retina and in the brain in major depression. Our results may contribute to a better understanding of depression etiopathogenesis and to its personalized treatment.
[0064] In some embodiments, the annotation model may include an expert model configured to detect instances of predetermined expressions of an expert technical field and to annotate those expressions. In some embodiments, the 'expert model' may be an 'extractive model' configured to perform tasks such as Named Entity Recognition (NER), a known technique for detecting and classifying specific entities in text, which aligns with the model's described function of detecting 'instances of predetermined expressions.' In some embodiments, there may be multiple annotation models that can annotate the original context or the annotated context received from a previous annotation model. In some embodiments, the multiple annotation models may be configured to annotate the original context in parallel.
[0065] As an example, when the written article is a medical research article, a first expert model may be configured to detect instances of drug names and to tag predetermined instances of drug names in the context, and a second expert model may be configured to detect instances of adverse effects caused by drugs and to tag predetermined instances of adverse effects in the context.
[0066] FIG. 3 shows an embodiment of a system in accordance with the present disclosure. The system of FIG. 3 includes a database 42 in which a written article may be stored, a processing module 44 coupled to the database 42, an annotation model 41 (or multiple annotation models) coupled to processing module 44, a user interface 46 coupled to the processing module 44, a first LLM 48 and a second LLM 50 both coupled to the processing module 44, and a memory 52 couple to the processing module 44. In this embodiment, the User Interface 46, the database 42, the annotation model 41 , the memory 52, the first LLM 48, and the second LLM 50 are shown as interconnected to each other via the processing module 44. In other embodiments, the user interface 46, the database 42, the annotation model 41, the memory 52, the first LLM 48, and the second LLM 50 may be interconnected to each other directly, or through the processing module, or through any other suitable circuitry.
[0067] The processing module 44 may be configured to receive commands from the user interface 46 and to perform the actions of the methods described herein.
[0068] FIG. 4 shows an embodiment of a user experience (LIX) interface 100, in accordance with the present disclosure. As described below, the features of the LIX interface 100 may reduce the time needed for manual validation of suggestions (suggested answers) provided by a data extraction system, eliminate the need for a reviewer wanting to validate the suggestions to read the entire article from which the data is extracted, demonstrate to the user where in the document the data extraction system found the suggested answer (e.g., with up to 5 citations that support the suggested answer), allow the user to automatically jump to a cited location (citation) with a single click, have the citation highlighted in the article to support quick visual validation, and allow the user to copy all citations as evidence with a single click, thereby removing the need for users to manually transcribe responses or interact with form elements (input form elements). The LIX interface 100 may also be coupled to a curated evidence storage and be configured to allow the user to distinguish or compare or visualize evidence extracted by an AI / LLM data extraction system and evidence stored in the curated evidence storage (the latter being evidence validated by a human and / or extracted from the article by a human). The LIX interface 100 may also be configured to consolidate the experience as part of a general “leverage suggestions” framework where users can toggle between human curated evidence and suggested answers obtained by the AI / LLM data extraction system. The UX interface 100 may be configured to include a question panel 102 that displays questions that are to be answered as part of a review of an article. In the present example, the article being reviewed is entitled: “Gut feelings: A randomised, triple-blind, placebo-controlled trial of probiotics for depressive symptoms”, by Bahia Chahwan, Sophia Kwan, Ashling Isik, Saskia van Hemert, Catherine Burke, and Lynette Roberts. Journal of Affective Disorders 253, (2019), pages 317-326.
[0069] The questions panel may also be referred to as, for example, a form or as a question and answer entry form. The questions panel 102 may include any number of questions. For example, the questions panel 102 may include a first question Q1: Study Type, the answer to which is to be selected by the data extraction system between the possible answers: Prospective Cohort Study (PCS), Randomized Control Trial (RCT), and Observational. The question panel 102 may also include a second question Q2: Describe the study design factors. Q2 has a text box 103 that can be filled with an answer to the question.
[0070] The UX interface 100 may also include a suggestions panel (answer panel, suggested answer panel) 104 that is configured to display suggestions (suggested answers) to the questions displayed in the questions panel 102. In the present example, the suggestions panel 104 may display a suggested answer S1 to the first question Q1 , the suggested answer being, in the present example, S1 : Randomized Control Trial (RCT). The answer panel may also display a suggested answer S2 to the second question S2, the second question S2 being a question that requires a text answer, the second question S2 being: “Use of M.I.N.I. and the analysis of stools samples so that the results could be paired with actual gut microbiome differences, and not just whether a probiotic was taken.”.
[0071] The suggested answer S1 displayed in the suggestions panel 104 may include an article identifier (ID) 105 that identifies the article being reviewed. The article ID 105 may be configured to display the reference to the article when a user moves a pointer 106 of the UX interface 100 over the article ID 105 and / or clicks on the article ID 105. The suggested answer S1 displayed in the suggestions panel 104 may include an indication 108 that indicates where in the article there is evidence for the suggested answer S1. The indication 108 may also be referred to as, for example, a button, a control, a displayed item, an icon, etc. In the present example, the indication 108 is in the form of a page number (Page 2). In other embodiments, the indication 108 could be a paragraph number, a page number plus a line number, a figure number, or any other suitable indication of where in the article there is support for the suggested answer S1.
[0072] The UX interface 100 may include a source material (or evidence) panel 110 configured to display evidence 112 associated with the indication 108 of a suggestion displayed in the suggestions panel 104, when the cursor 106 hovers over or otherwise actuates (clicks) the indication 108. In the present embodiment, the evidence 112 that supports the suggested answer S1 is a reproduction of a portion of the article (as published) being reviewed. The evidence 112 may include highlighted text 114 that may highlight precisely where, in the portion of the article, there is support for the answer. Any other suitable means of showing or marking where in the portion of the article there is support for the answer, is to be considered within the scope of the present disclosure.
[0073] In some embodiments, the suggested answer S1 may include more than one indication 108, when there is more than one section of the article that supports the suggested answer S1. In other embodiments, the suggested answer S1 may include an indication 108 that, when actuated, causes the UX interface 100 to display a list of instances of evidence, within the article,of where the article supports the suggested answer S1. The UX interface 100 may be configured such that when an instance of the list is actuated (or clicked), the source material panel 110 displays the particular instance of support for the suggested answer S1.
[0074] The suggestions panel 104 may include, with the suggested answer S1 , a confirmation button 116 that may be configured such that when actuated, causes the suggested answer S1 to be attached to the first question Q1 or to be accepted as the confirmed answer to the question Q1. In turn, the LIX interface 100 may be configured to check the box 118 next to the Randomized Control Trial (RCT) line 120 in the question panel 102. In some embodiments, clicking on the confirmation button 116 may also cause all the evidence (all the instances of support) for the suggested answer S1 (now the confirmed answer) to be saved to an evidence storage (memory).
[0075] Referring now to the second question Q2: “Describe the study design factors.”, the Suggestions panel 104 displays a suggested answer S2 which includes: “Use of M.I.N.I. and the analysis of stools samples so that the results could be paired with actual gut microbiome differences, and not just whether a probiotic was taken.” The suggested answer S2 may also include the article ID 105, an indication 122 that indicates where in the article there is a first instance of support for the suggested answer S2, and an indication 123 that indicates where in the article there is a second instance of support for the suggested answer S2. Clicking on the indication 122 may cause the source material panel 110 to display where, in the article, there is the first instance of support for the answer S2. Clicking on the indication 123 may cause the source material panel 110 to display where, in the article, there is the second instance of support for the answer S2.
[0076] The suggestions panel 104 may also include a confirmation button 124 that may be configured such that, when actuated, causes the suggested answer S2 to be attached to the second question Q2 or to be accepted as the confirmed answer and to be copied into the text box 103 of the question panel. In some embodiments, clicking on the confirmation button 124 may also cause all the evidence (all the instances of support) for the suggested answer S2 (now the confirmed answer) to be saved to an evidence storage (memory).
[0077] In some embodiments, instead of showing suggested answers obtained from a data extraction system, the suggestions panel 104 can show suggestions obtained from curated evidence previously obtained by a previous review of the same article. In some embodiments, the UX interface 100 may have a toggle button allowing a user to toggle views of the suggestionspanel 104 between a first configuration where answers obtained from reviewing the article using the data extraction system are displayed, and a second configuration where curated answers to the same questions are displayed.
[0078] In alternative embodiments, the features of the user interface 100 may be presented differently. For example, rather than a separate suggestions panel 104, a suggested answer could appear as a 'ghost text' directly within an answer field in the questions panel 102. The evidence 112 could be displayed in a pop-up window or dialog box that appears when a user hovers over or clicks on the suggested answer, rather than in a persistent source material panel 110. Such variations are contemplated and fall within the scope of the present disclosure.
[0079] The LIX interface 100 may include an evidence management panel configured to manage the evidence obtained for the answers to the questions that are part of the review of the article. The evidence management panel may be accessed through a control button displayed on the LIX interface 100, or through any other suitable command input means. The evidence management window may be displayed automatically when an answer to a question of the questions panel 102 is answered and / or accepted. FIG. 5 shows an embodiment of an evidence management panel 130. The evidence management panel 130 may be a modified version of the questions panel 102. The embodiment of FIG. 5 relates to a different article than the one discussed in relation to FIG. 4. The article reviewed in relation to FIG. 5 has a questions panels (not shown) that includes, as a first question: “Does the article relate to ultraviolet treatments?” and, as a second question: “What interventions are mentioned in the article?” FIG. 5 shows that the answer to the first question is “Yes”, and that page 1 of the article includes two instances of support for the answer to the first question, page 2 of the article includes one instance of support for the answer, and page 6 the article also includes one instance of support for the answer. See “Pages 1 | 1 | 2 | 6” at reference number 109.
[0080] FIG. 5 also shows, in the text box 105, that the answer to the second question is “The article mentions various interventions for inducing re-pigmentation in vitiligo patients, including ultraviolet B (UVB)-based phototherapy, dermabrasion, microneedling, ablative fractional CO2 lasers and punch grafting.” FIG. 5 also shows that page 1 of the article includes one instance of support for the answer to the second question, page 2 of the article includes two instances of support for the answer, and page 6 the article also includes one instance of support for the answer. See “Pages 1 | 2 | 2 | 6” at reference number 125. FIG. 5 also shows the article ID 107 for the article being reviewed.
[0081] FIG. 5 shows that the evidence management panel 130 includes a manage evidence control (button) 132, which is configured to manage the evidence related to the answer to question 1 , and a manage evidence control (button) 134, which is configured to manage the evidence related to the answer to question 2.
[0082] The evidence management panel 130 may be configured such that when the confirmation button associated with the first question is actuated, all the evidence supporting the answer to the first question is copied (accepted) to the form (question and entry form) that is part of the questions panel. Actuation of the confirmation button associated with the second question may cause all the evidence supporting the answer to the second question to be copied (accepted) to the form (question and entry form) that is part of the questions panel.
[0083] Subsequent to the evidence of an answer to a question being copied to the form that is part of the questions panel, the evidence management panel 130 may, upon actuation of the manage evidence button 132 or the manage evidence button 134, cause the opening of an evidence drawer 136. As an example, upon actuation of the manage evidence button 134, the evidence drawer 136 may be configured to display, under a heading 140, existing evidence for the answer to the second question. Existing evidence is to be understood as meeting evidence obtained from previous reviews and confirmed by a human. The existing evidence may be part of a curated evidence storage.
[0084] Below the heading 140 are displayed the article ID 107 and the page number 125 for each instance of support for the answer to the second question. Also displayed, for each pair of article ID and page number, is a phrase 127 that is part of the respective instance of support for the answer.
[0085] The evidence drawer 136 may be configured to allow a user to review each of the questions (e.g., first question, second question, etc.), answers, and all the evidence for the respective answers. FIG. 5 shows that the evidence drawer 136 of the present embodiment also displays a heading 160, which reads: “Suggestions for Evidence for answer to Q2”. Below the heading 160 are displayed the article ID 107 and the page number 125 for each instance of support for the answer to the second question. Also displayed, for each pair of article ID and page number, is a phrase 127 that is part of the respective instance of support for the answer. The evidence drawer 136 of the example of FIG. 5 also shows, for each instance of support of the answer to the second question either an indication 129 that the instance has been accepted, or abutton 130 that may be actuated to accept the particular instance of support for the answer to the second question.
[0086] The evidence drawer 126 may be configured such that moving the cursor 106 over an instance of support for the answer to the question and / clicking on the instance in question causes a source material window (source material panel) to be displayed, the source material window showing where in the article is the particular support (evidence) for the answer to the question.
[0087] Upon confirmation of an answer to a question and confirmation of evidence for the answer, the question, the answer and evidence may be saved to a curated evidence storage for the article reviewed.
[0088] Each action or operation of the method described herein may be executed on any computing device, such as a personal computer, server, PDA, or the like and pursuant to one or more, or a part of one or more, program elements, modules or objects generated from any programming language, such as C++, Java, or the like. In addition, each operation, or a file or object or the like implementing each said operation, may be executed by special purpose hardware or a circuit module designed for that purpose.
[0089] Through the descriptions of the preceding embodiments, the present invention may be implemented by using hardware only or by using software and a necessary universal hardware platform. The present invention may also be cloud-based. Based on such understandings, the technical solution of the present invention may be embodied in the form of a software product. The software product may be stored in a non-volatile or non-transitory storage medium, which can be a compact disk read-only memory (CD-ROM), USB flash disk, or a removable hard disk. The software product may include a number of instructions that enable a computer device (personal computer, server, or network device) to execute the methods provided in the embodiments of the present invention. For example, such an execution may correspond to a simulation of the logical operations as described herein. The software product may additionally or alternatively include number of instructions that enable a computer device to execute operations for configuring or programming a digital logic apparatus in accordance with embodiments of the present invention.
[0090] The word “a” or “an” when used in conjunction with the term “comprising” or “including” in the claims and / or the specification may mean “one”, but it is also consistent with the meaning of “one or more”, “at least one”, and “one or more than one” unless the content clearly dictatesotherwise. Similarly, the word “another” may mean at least a second or more unless the content clearly dictates otherwise.
[0091] The terms “coupled”, “coupling” or “connected” as used herein can have several different meanings depending on the context in which these terms are used. For example, as used herein, the terms coupled, coupling, or connected can indicate that two elements or devices are directly connected to one another or connected to one another through one or more intermediate elements or devices via an electronic element depending on the particular context. The term “and / or” herein when used in association with a list of items means any one or more of the items comprising that list.
[0092] Although a combination of features is shown in the illustrated embodiments, not all of them need to be combined to realize the benefits of various embodiments of this disclosure. In other words, a system or method designed according to an embodiment of this disclosure will not necessarily include all features shown in any one of the figures or all portions schematically shown in the Figures. Moreover, selected features of one example embodiment may be combined with selected features of other example embodiments.
[0093] Although the present invention has been described with reference to specific features and embodiments thereof, it is evident that various modifications and combinations can be made thereto without departing from the invention. The specification and drawings are, accordingly, to be regarded simply as an illustration of the invention as defined by the appended claims, and are contemplated to cover any and all modifications, variations, combinations or equivalents that fall within the scope of the present invention.
Claims
CLAIMS:1 . A method of extracting data from an article, the method comprising: providing a question to a large language model (LLM) system, the question being in relation to the article; obtaining, from the LLM system, an answer to the question; and displaying a user interface, the user interface showing: the question; the answer to the question; and an indication of where in the article there is evidence for the answer.
2. The method of claim 1 , wherein the user interface displays a control that, when actuated, causes the user interface to show where in the article there is the evidence for the answer.
3. The method of claim 2, wherein the user interface shows the indication of the section of the article where there is evidence for the answer by performing at least one of: displaying an as-published version the article, at a section of the article that supports the answer; displaying a section of the article that supports the answer; and displaying coordinates of where in the article there is support for the answer.
4. The method of claim 3, wherein the coordinates include a page number of the article.
5. The method of claim 4, wherein the user interface is configured to show the page number as part of a page number button that, when actuated, causes the user interface to show a page associated with the page number or a section of the page associated with the page number.
6. The method of claim 3, wherein: displaying the as-published version the article, at a section of the article that supports the answer includes marking one or more portions of the section that supports the answer; and displaying the section of the article that supports the answer includes marking one or more portions of the section that supports the answer.
7. The method of claim 6, wherein marking the one or more portions includes highlighting the one or more portions.
8. The method of claim 1, wherein the user interface includes a confirmation button that when actuated, causes the answer to the question and the evidence for the answer to be attached to the question.
9. The method of claim 1, wherein the user interface shows a plurality of indications of a respective plurality of sections of the article where there is evidence for the answer.
10. The method of claim 9, wherein the user interface includes a confirmation button that when actuated, causes the answer to the question and the sections of the plurality of sections where there is evidence for the answer to be attached to the question.
11. The method of claim 9, wherein the user interface includes a confirmation button that when actuated, subsequent verification that at least one section of the plurality of sections of the article supports the answer, tags the at least one section as being verified evidence of the answer to the question.
12. The method of claim 9, wherein: the user interface includes an answer entry box; and a confirmation button that when actuated, subsequent verification that at least one section of the plurality of sections of the article supports the answer, copies the answer to answer entry box to display the answer in the answer entry box.
13. The method of claim 9, wherein the user interface includes a confirmation button that, when actuated, subsequent verification that at least one section of the plurality of sections of the article supports the answer causes the at least one section to be saved as being verified evidence of the answer to the question, the at least one section to be saved to an evidence storage.
14. The method of claim 13, wherein the user interface includes a storage button that, when actuated, causes the plurality of indications of the respective plurality of sections of the articlethat supports the answer to be saved to an evidence storage and tagged as being associated with the question.
15. A method of extracting data from an article, the method comprising: providing a question to a large language model (LLM) system, the question being in relation to the article; obtaining, from the LLM system, an answer to the question, the answer being an LLM answer; obtaining, from curated evidence associated with the written article, a curated answer to the question; and displaying a user interface, the user interface showing: the question; and at least one of: the LLM answer and an indication of where in the article there is evidence for the LLM answer; and the curated answer and an indication of where in the article there is evidence for the curated answer.
16. The method of claim 15, wherein the user interface includes: a question panel configured to display the question; a suggestions panel having a first configuration where the answer panel displays the LLM answer and the indication of where in the article there is evidence for the LLM answer, the evidence for the LLM answer being suggested evidence; and a second configuration where the answer panel displays the curated answer and the indication of where in the article there is evidence for the curated answer.
17. The method of claim 15, wherein the user interface displays a toggle button that, when actuated, toggles the answer panel between the first configuration and the second configuration.
18. The method of claim 15, wherein: the user interface further includes an evidence panel, the indication of where in the article there is evidence for the LLM answer includes a first control that, when actuated, causes the evidence for the LLM answer to be displayed in the evidence panel, andthe indication of where in the article there evidence for the curated answer is includes a second control that, when actuated, cause the evidence for the curated answer to be displayed in the evidence panel.
19. The method of claim 15, wherein: the user interface includes an evidence management panel configured to display the question, the LLM answer to the question, and the indication of where in the article there is evidence for the LLM answer, the evidence management panel include a manage evidence button that, when actuated, causes the user interface to display an evidence drawer that includes: the evidence for the LLM answer and an accept button that when activated accepts the evidence for the LLM answer as verified evidence; and confirmed evidence for the LLM answer.
20. A computer-implemented method of extracting data from a written article, the method comprising: obtaining a plurality of chunks from the written article; generating a respective chunk numerical representation for each of the plurality of chunks; generating, by a first large language model (LLM) in response to an initial question, one or more secondary questions related to the initial question; generating a respective question numerical representation for the initial question and for each of the one or more secondary questions; selecting, from the plurality of chunks, a subset of two or more chunks based on a comparison between the chunk numerical representations and the question numerical representations; generating a context from the subset of two or more chunks; generating, by a second LLM based on the context and the initial question, an answer to the initial question; and identifying, within the subset of two or more chunks, one or more portions of text that support the generated answer.
21. The method of claim 20, wherein generating the context comprises concatenating the text of the subset of two or more chunks.
22. The method of claim 20, wherein the chunk numerical representations and the question numerical representations are embeddings.
23. The method of claim 20, wherein obtaining the plurality of chunks from the written article comprises performing a semantic search and retrieval operation.
24. The method of claim 20, wherein generating the one or more secondary questions comprises providing the initial question to the first LLM using one-shot prompting.
25. The method of claim 20, wherein the first LLM and the second LLM are the same LLM.
26. The method of claim 20, further comprising: prior to generating the answer by the second LLM, providing the context to at least one annotation model to generate an annotated context; and wherein generating the answer by the second LLM is based on the annotated context and the initial question.
27. The method of claim 26, wherein the at least one annotation model is an expert model configured to detect and annotate instances of predetermined expressions of an expert technical field.
28. The method of claim 26, further comprising: providing an instruction to the second LLM specifying an answer type, wherein the answer type is selected from the group consisting of a single answer type, a multiple answer type, and a free text answer type.
29. The method of claim 20, wherein identifying the one or more portions of text that support the generated answer comprises at least one of: the second LLM reproducing the one or more portions of textIO; and the second LLM providing coordinates of the one or more portions of text with respect to a layout of the written article.
30. An intelligent document processing system, comprising:a processing module; a user interface; a tangible computer-readable medium having instructions recorded thereon, the instructions to be carried out by the processing module to perform the method of any one of claims 1 to 29.
31. A non-transitory computer-readable medium having stored thereon instructions that, when executed by one or more processor, cause the one or more processors to perform the method of any one of claims 1 to 29.