Critical content extraction method and system

The method and system address the challenge of extracting critical content from electronic documents by using a large language model to identify and filter confounding factors, improving the understanding and direction of future research through unique content extraction.

WO2026013540A1PCT designated stage Publication Date: 2026-01-15LITMAP LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/056850
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-08
Filing Date
2025-07-07
Publication Date
2026-01-15

AI Technical Summary

Technical Problem

Existing search systems fail to effectively locate and extract critical, qualitative information such as confounding factors or negative outcomes in electronic documents, which are crucial for comprehensive understanding and future research directions.

Method used

A method and system that utilizes a large language model (LLM) to identify confounding factors by posing questions to electronic documents, comparing them semantically, and filtering based on frequency of appearance thresholds to extract unique or peculiar critical content.

Benefits of technology

Enables the identification and presentation of unique critical content, guiding future research by highlighting confounding factors and negative outcomes, thereby enhancing the comprehension and interpretation of search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025056850_15012026_PF_FP_ABST
    Figure IB2025056850_15012026_PF_FP_ABST
Patent Text Reader

Abstract

This invention relates generally to critical content extraction. In particular, the invention is a method and a system to locate, extract and map critical content from database content especially electronic documents such as electronically published books, articles, and results that are a record human knowledge.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CRITICAL CONTENT EXTRACTION METHOD AND SYSTEM

[0002] Field of the Invention

[0003] This invention relates generally to critical content extraction. In particular, the invention is a method and a system to locate, extract and map critical content from database content especially electronic documents such as electronically published books, articles, and results that are a record human knowledge.

[0004] It is an advantage of the present invention that it extracts content missed by the logic of search systems.

[0005] Background

[0006] The volume of information available today in many domains confounds the retrieval of critical, qualitative information.

[0007] We use the term ‘critical’ to mean incidental information, indicators or confounding information that are important for meaningful comprehension and interpretation of search results. Specifically, critical content may be for example components of research papers that mention limitations of a study condition or ideas mentioned in discussions that provide directions for meaningful next steps.

[0008] Many systems are known whereby a user can retrieve information and documents that meet preset criteria and thereby analyse the documents, categorize portions of the analysed documents, present images of the documents for a portion of the categories, extracts thereof, or indicia of a portion of an image of the document related to the category.

[0009] Sub-domains of interest, data-centric search refinement, and systems that analyse the linguistic and syntactic structure of a search query are also known, for example, using machine learning models.

[0010] Systems are known to provide simultaneous search of multiple data sources, retrieving information from various content locations with a single query and search interface. Methods are known whereby searching uses ranked concept markers to refine the relevancy of content associated with a query. Refining includes for example, re-ranking documents based on a combination of the original query and selected concept markers. However, asking questions in an unstructured way to Large Language Models like OpenAI’s GPT 4o or GPT o3 is known to have limitations of reliability and search scope.

[0011] Unstructured data can be analysed to generate a structured data output e.g. in a predefined or customisable template. The predefined template includes a plurality of fields: each field corresponds to a field of the structured report and defines the extraction rules for each field of the predefined template, which in turn define parameters for identifying unstructured data relevant to the associated field.

[0012] Also known are relevancy scoring algorithms which determine and represent the relevance of a search term to the text contained in the document. The scores, including an aggregate score, may be normalized and based on relevancy scoring. Terms are ranked and further processed.

[0013] Relevancy may be determined based on the occurrences and locations of the search terms within the document.

[0014] More recently, use of flexible natural language interfaces is common, which compromises receiving an initiated user question at a graphical user interface and generating automatically one or more suggested completed questions in response to the receipt of the initiated user question.

[0015] However, known search means are biased by the predicate of ‘positive’ outcome or positive correlation. For example, if a clinical trial succeeds, it will be published and searched via terms relating to objective and positive outcome. Here ‘positive’ refers to criteria of a study that supports, enables, or defines results. Whereas ‘negative’ means the confounding or ‘critical’ criteria that may in some way be contrary to a result or conclusion of the study.

[0016] ‘Negative’, fault or flaw discriminating against, or confounding information, i.e. ‘critical information’ may be couched as hypotheses or exclusions, neither of which are anticipated. Such negative or confounding information, i.e. ‘critical information’, is crucial for meaningful comprehension, interpretation and development in the field. For example, in scientific, medical, legal, and academic literature negative and confounding aspects are often presented in a ‘discussion section’: “we only tested on males of age 30-40” or “the drug was ineffective on premenopausal women” or “BMI was obtained for a select group of patients so we ignore BMI, other studies have found BMI to be a critical confounder in clinical trials”.

[0017] The present invention provides means to locate, extract and map critical content. Critical content is components of research papers that mention limitations of a study condition or ideas mentioned in discussions or conclusions or both from theoretical analysis, results or experiments, and surveys. Critical content is critical information and it can provide directions for next steps. of the Invention

[0018] According to a first aspect of the invention there is a method of critical content extraction (CCE), comprising: inputting into a (CCE) system a plurality of electronic documents comprising a plurality of sections of text, finding confounding factors between the plurality of sections by posing at least one question to a large language model (LLM) or a large reasoning model (LRM) about each of the plurality of sections of text, comparing each one of the confounding factors to at least of other confounding factors using a semantic similarity, and deriving unique or peculiar critical content by filtering each one of the confounding factors using a frequency of appearance threshold.

[0019] The method of provides a user unique or peculiar critical content in one or more retrieved electronic documents or sections of the retrieved electronic documents or both.

[0020] According to a second aspect of the invention there is a critical content extraction CCE system comprising a computer, an input receiving device via which the computer receives electronic documents, an interactive medium by which the computer provides the unique or peculiar critical content from the retrieved electronic documents to a user, and a computer program directly loadable into a memory of the computer for controlling the method when the program is run on the computer. Preferably the interactive medium is further configured to provide a source of the unique or peculiar content from at least one of the retrieved electronic documents. The CCE method and the CCE system provide means to locate and extract critical content from databases, for example databases of literature or results of experiments and surveys or both. Preferably the CCE method and the CCE system are also configured to present a map of the critical content.

[0021] Critical content extraction CCE refersto extraction of content that is critical information. Critical information may be contrary to and critical of positive information or peculiar in its own way. This may be accomplished by all or some of several steps of a method.

[0022] The method of CCE may comprise formulating the at least one question before electronically reading the plurality of sections. The method may comprise after finding the confounding factors, storing the confounding factors in a relational database.

[0023] The method of CCE may comprise identifying in the plurality of sections at least one specific section and posing the at least one question about the at least one specific section. The least one specific section may comprise one or more of a results section, a discussion section, and a conclusion section.

[0024] The method of CCE may comprise filtering by comparing the frequency of appearance threshold to a first frequency of appearance of each of the confounding factors in each of the electronic documents. The method of CCE may comprise comparing the frequency of appearance threshold to a second frequency of appearance of each of the confounding factors in each of the sections in each of the electronic documents.

[0025] The method of CCE may comprise assigning a first uniqueness to each of the electronic documents for which the first frequency of appearance of any of the confounding factors is less than the frequency of appearance threshold. The method of CCE may comprise assigning a second uniqueness to each of the electronic documents for which the first frequency of appearance of all the confounding factors in total is less than the frequency of appearance threshold. The method of CCE may comprise determining a user selection by prompting a user to select at least one of the confounding factors and assigning a third uniqueness to each of the electronic documents for which the first frequency of appearance of the confounding factors in the user selection is less than the frequency of appearance threshold. The steps of the method may comprise: a user defining an area of investigation; the user selecting a search range using a search methodology such as key words or citations and retrieving from the database one or more electronically published books, articles, results of experiments, surveys, and so forth that are in the search range ~ these are retrieved electronic documents; parsing text in the retrieved electronic documents to extract or formulate a section of results, a section of conclusion, and a section of discussion for each retrieved electronic document; composing at least one question about the section of results, conclusion, and discussion for each retrieved electronic document; posing the at least one question to a Large Language Model (LLM) or a Large Reasoning Model (LRM) to extract or formulate critical content from each of the sections of each of the retrieved electronic documents, comparing critical content between two or more or all of the retrieved electronic documents to determine which of the critical content is unique or peculiar critical content; collating the unique or distinct critical content; displaying the unique or distinct critical content to the user; providing the user access to a source of the unique or peculiar critical content in one or more of the retrieved electronic documents or sections or both from which the unique or peculiar critical content was extracted or formulated.

[0026] The unique or peculiar critical content may be determined according to a maximum frequency threshold.

[0027] The unique or peculiar critical content may be provided to an interactive medium so that the user may see or hear it and input a response . The interactive medium may comprise one or more of a microphone, a speaker, a video screen, a touch screen, a keyboard, and a mouse and another device by which the user may provide feedback to the CCE system and vice versa.

[0028] By the CCE method and system key future research is identified from confounding factors or critical information in electronic documents that have been published. “Experts” in certain fields may be identified. A user is enabled to identify overview timelines key points by looking for uniqueness, not similarity. Qualitative content can be categorized via quantity of critical content.

[0029] The invention will now be described, by way of example only, with reference to the accompanying figures in which: Brief Description of the Figures

[0030] Figure 1 shows a first flow diagram of a method of critical content extraction according to the invention;

[0031] Figure 2 shows a second flow diagram of initial steps in the method;

[0032] Figure 3 shows a third flow diagram of intermediate steps in the method;

[0033] Figure 4 shows a fourth flow diagram of terminal steps in the method;

[0034] Figure 5 shows three critical factors each mentioning a particular electronic document individually; and

[0035] Figure 6 shows an example of text and sections in an electronic document.

[0036] Detailed Description of the Invention

[0037] Referring to the Figures, there is shown in Figure 1 a first flow diagram of a method of critical content extraction 100. In a first step 10 of the method 100, a user defines an area of investigation. In a second step 20, the user selects a search range using a collection of key words. In a third step 30 and a fourth step 40, results are retrieved, and full text of these results are fed to the system.

[0038] In a fifth step 50, full text for example PDF is converted into linear full text in reading order. For example, structured text is extracted using PDF extraction framework GROBID, although other extraction techniques are also available. Figure 2 shows a second flow diagram of a part of the method 150 beginning with the fifth step 50 of extracting the raw text of the electronic document.

[0039] In a sixth step 60 shown in Figure 1 and in Figure 2, questions are posed to an LLM or LRM about the Results, Conclusion, and Discussion sections to extract the critical confounding factors. In an embodiment shown in Figure 1 and Figure 2 the questions are formulated based upon initial information about the retrieved documents provided by the user. In another embodiment which may be executed alone or with the information provided by the user, the questions are formulated from initial information gleaned from the retrieved documents. In a seventh step 70 shown in Figure 1 , each confounding factor is compared by semantic similarity and filtered by a frequency of appearance threshold. As shown in a step 72 in Figure 2, confounding factors are found and stored in a relational database.

[0040] Then in an eighth step 80, confounding factors that qualify according to semantic similarity and frequency of appearance threshold criteria in step 70 are qualifying critical content factors. The confounding factors that qualify are ranked by inverted frequency and presented with links to each source document. As shown in a step 82 and step 84 in Figure 3, a critical factor that is mentioned in a number of the retrieved electronic documents less than a maximum threshold e.g. 3 times then qualifies for uniqueness or peculiarity. In this way unique or peculiar critical content is determined according to the maximum frequency threshold.

[0041] As shown Figure 4 in a step 86, the confounding factors that qualify for uniqueness or peculiarity are ordered by inverted frequency. In an embodiment inverted frequency is from highest number of occurrences of the confounding factors that qualify in a retrieved electronic document to the lowest.

[0042] Confounding factors may qualify by uniqueness or peculiarity per each retrieved electronic document or by uniqueness or peculiarity accounting for all the retrieved electronic documents. For example, Figure 5 illustrates three separate electronic documents: paper one 92, paper two 94, and paper three 96. There is a confounding factor that qualifies by uniqueness or peculiarity associated with each of these electronic documents.

[0043] This aids the user in further lines of research 90 since the confounding factors that qualify are critical content factors that provide guidance for and signs of novel paths of endeavor.

[0044] A first illustrative example of steps in the method are as follows. Questions are posed to an LLM or LRM about the Results, Conclusion, and Discussion sections to extract the critical confounding factors. Using for example OpenAI o3 pro prompts may comprise:

[0045] (1) results Section Prompt: “DO NOT HALLUCINATE PLEASE. Can you extract confounding factors from an electronic document such as a research paper. This paper is in the medical domain. This is the results section. Can you find the factors that were the limitations of this study? This is so future researchers can find areas to investigate.

[0046] Here are the results extracted from the PDF:” and / or

[0047] (2) discussion Section Prompt: “DO NOT HALLUCINATE PLEASE. Can you extract confounding factors from an electronic document that is a research paper. This paper is in the medical domain. This is the discussion section. Can you find the factors that were the limitations of this study? This is so future researchers can find areas to investigate. Here is the discussion extracted from the PDF:” and / or resulting critical content passages are linked to each paper and stored in a PostgreSQL relational database; and / or each electronic document in the input set is compared with each other for the frequency of its confounding factors, where a threshold is defined, for example, a threshold of less than 3 other papers mention the same confounding factor to qualify; and / or critical content comparison is done via semantic similarity as defined by the similarity between embedding vector directions generated by OpenAI’s text-embedding-3-large model. The cosine similarity approach is used to measure similarity, where the overall vector direction is compared with a custom threshold for example, a custom threshold of 0.85; and / or the critical content is collated and shown to the user with unique confounding factors are highlighted and emphasized with references to the full text content that the confounding factor was extracted from. The presentation is a timeline of confounding factors with links to papers. The user then selects critical content to read and further lines of research.

[0048] In another example of the method steps are as follows.

[0049] 1. a user (human or automated) selects a range of academic papers using a system

[0050] 2. full texts are retrieved

[0051] 3. the LLM or LRM, for example, OpenAI o3 Pro extracts critical content using pre-defined questions of the Results, Conclusion, or Discussion sections

[0052] 4. critical content from each electronic document is embedded, where the text is converted into a vector form using an embedding model. Using these embedding vectors, each critical content piece is compared using cosine similarity of embedding vectors, if a critical factor is the same it is grouped together. The maximum threshold for grouping is 3.

[0053] 5. the unique critical content is collated and shown to the user, who can access the source information via clickable links to each paper behind the critical content

[0054] 6. the user then selects critical content of purpose.

[0055] In another embodiment, the method comprises steps as follows.

[0056] A user searches for academic papers relating to the effect of patient characteristics on the performance of an Al algorithm to interpret digital breast tomosynthesis studies in breast cancer screening. The citations include a paper entitled ‘Patient characteristics impact performance of Al algorithm in interpreting negative screening digital breast tomosynthesis studies’ with a published ‘Conclusion’ that patient characteristics influenced the cause and risk scores of FDA approved Al algorithm analysing negative screening DBT examinations.

[0057] Along the examples is a case assigned a false-positive score of 96 in a 59 year old woman of African descent. The mammogram showed scattered fibroglandular breast density and vascular calcifications.

[0058] A conclusion that the user was looking for is found, i.e. evidence of bias in an Al algorithm was found.

[0059] A further step retrieves critical information: the false positive indicative of more false positives and the potential research imperative that women of African descent themselves are more prone to breast arterial calcification and thus heart disease. According to the present method, although in the context of this paper, a false positive rate in these women was incidental, but alerts the user that there may be a biological explanation and better detection of early cardiac disease.

[0060] In another embodiment, the method comprises steps as follows.

[0061] A user searches publications of advanced technology in a form of electronic documents, for example as shown in Figure 6, relating to the mammographic density in women of different ethnicities. The citations include and a published electronic document entitled ‘Association of area and volumetric-mammographic density and breast cancer in women of Asian descent: A case control study’. The publication shows that in Asian women, two types of breast density measurement both reflected increased breast cancer risk, with the strongest associations postmenopausal. Within the text is the observation that larger samples in premenopausal women are required to detect a significant association with breast cancer risk.

[0062] In accordance with the present invention, the critical information can be deduced: the difficulty of measuring breast density in particular women due to the level of density present measurement techniques have more error and larger sample sizes are needed.

[0063] The invention has been described by way of examples only. Therefore, the foregoing is considered as illustrative only of the principles of the invention. Further, since numerous modifications and changes will readily occur to those skilled in the art, it is not desired to limit the invention to the exact construction and operation shown and described, and accordingly, all suitable modifications and equivalents may be resorted to, falling within the scope of the claims.

Claims

Claims1 . A method of critical content extraction CCE, comprising: inputting into a CCE system a plurality of electronic documents comprising a plurality of sections of text, finding confounding factors between the plurality of sections by posing at least one question to a large language model (LLM) or large reasoning model (LRM) about each of the plurality of sections of text, comparing each one of the confounding factors to at least of other confounding factors using a semantic similarity, and deriving unique or peculiar critical content by filtering each one of the confounding factors using a frequency of appearance threshold.

2. A method of CCE according to claim 1 comprising formulating the at least one question before electronically reading the plurality of sections.

3. A method of CCE according to claim 1 or 2 comprising after finding the confounding factors, storing the confounding factors in a relational database.

4. A method of CCE according to any preceding claim comprising identifying in the plurality of sections at least one specific section and posing the at least one question about the at least one specific section.

5. A method of CCE according to claim 4 wherein the least one specific section comprises one or more of a results section, a discussion section, and a conclusion section.

6. A method of CCE according to any preceding claim comprising filtering by comparing the frequency of appearance threshold to a first frequency of appearance of each of the confounding factors in each of the electronic documents.

7. A method of CCE according to any preceding claim comprising by comparing the frequency of appearance threshold to a second frequency of appearance of each of the confounding factors in each of the sections in each of the electronic documents.

8. A method of CCE according to claim 6 comprising assigning a first uniqueness to each of the electronic documents for which the first frequency of appearance of any of the confounding factors is less than the frequency of appearance threshold.

9. A method of CCE according to claim 6 comprising assigning a second uniquenessto each of the electronic documents for which the first frequency of appearance of all the confounding factors in total is less than the frequency of appearance threshold.

10. A method of CCE according to claim 6 comprising determining a user selection by prompting a user to select at least one of the confounding factors and assigning a third uniqueness to each of the electronic documents for which the first frequency of appearance of the confounding factors in the user selection is less than the frequency of appearance threshold.

11. A CCE system to execute the method of CCE according to any of claims 1 to 10, the CCE system comprising a computer, an input receiving device via which the computer receives the electronic documents, an interactive medium by which the computer provides the unique or peculiar critical content from the retrieved electronic documents to a user, and a computer program directly loadable into a memory of the computer for controlling the method when the program is run on the computer.

Citation Information

Patent Citations

  • Heuristic Domain Targeted Table Detection and Extraction Technique

    US20190171704A1

  • Method and system for the computer-assisted implementation of radiology recommendations

    US20230343454A1

  • Framework for Evaluation of Document Summarization Models

    US20240078380A1