Biomedical literature analysis method and system

By reordering and deduplicating medical literature data, dividing it into sentence-level evidence blocks and binding them with unique identifiers, and constructing a multi-agent collaborative architecture, the problem of insufficient accuracy and verifiability in the processing of medical literature data in existing technologies is solved, and multi-dimensional intelligent analysis is realized.

CN121659936APending Publication Date: 2026-03-13ZHEJIANG LAB
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-21
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing technologies lack quality screening and fine-grained control over source data in medical literature data processing, resulting in insufficient accuracy and verifiability of the generated conclusions. Furthermore, their functions are limited and cannot meet the complex and ever-changing analytical needs in scientific research practice.

Method used

By retrieving data from several data sources and performing reordering and hierarchical deduplication based on preset indicators, the literature data is divided into sentence-level evidence blocks, and a unique identifier is bound to each sentence-level evidence block. The system responds to literature analysis requests by performing review, question answering, and extraction tasks, and constructs a unified evidence backbone and multi-agent collaborative architecture.

Benefits of technology

It enables high-quality and reliable multi-dimensional processing of medical literature data, ensuring the traceability and verifiability of results. It can seamlessly acquire multi-dimensional analysis results from macro reviews to micro data, meeting the diverse analytical needs of researchers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121659936A_ABST
    Figure CN121659936A_ABST
Patent Text Reader

Abstract

The invention relates to a biomedical literature analysis method and system.The method comprises the steps that a plurality of data sources are retrieved to obtain literature data and unique identifiers corresponding to the data sources; the method comprises the following steps: reordering literature data on the basis of a preset index, performing hierarchical de-duplication processing, dividing the literature data into sentence-level evidence blocks, and binding a corresponding unique identifier for each sentence-level evidence block; in response to a literature analysis request, selecting the corresponding sentence-level evidence block based on the unique identifier, and executing one or more of the following tasks: executing a review task to generate a structured review text; executing the question and answer task to generate an answer text; executing the extraction task to identify and output a structured parameter table; and presenting all results of the executed task. The problem that intelligent multi-dimensional processing of the biomedical literature data cannot be realized on the basis of accurate tracing of the biomedical literature data in related technologies can be effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of biomedical literature processing technology, and in particular to a biomedical literature analysis method and system. Background Technology

[0002] As biomedical research enters the era of big data and artificial intelligence, utilizing intelligent analysis tools for literature review, literature review writing, and data mining has gradually become a mainstream trend in scientific research. Existing technologies have provided implementation methods such as question answering based on large language models, automated literature reviews, and information extraction.

[0003] However, existing technical solutions have limitations in constructing a foundation of evidence related to medical literature data, and their overemphasis on the processing of medical literature data leads to insufficient reliability and verifiability of the generated content. Current technical solutions involve constructing a biological entity database and a literature abstract knowledge graph database, using large models for user intent recognition and logical chain instruction generation, and finally integrating information to generate query results. This approach lacks quality screening and fine-grained control over source data, making it difficult to guarantee the accuracy and verifiability of the generated conclusions. Furthermore, its functionality focuses on question answering and retrieval, with a fixed functional architecture, making it unable to respond to comprehensive literature analysis requests, and unable to intelligently select and coordinate the execution of multiple tasks such as review generation, question answering, and data extraction, thus failing to meet the complex and ever-changing analytical needs of scientific research practice.

[0004] In view of this, existing technologies do not yet have a complete solution to the technical problem of how to achieve intelligent multi-dimensional processing of medical literature data based on accurate traceability of medical literature data. Summary of the Invention

[0005] This application provides a biomedical literature analysis method and system to solve the problem in related technologies that it is impossible to achieve intelligent multi-dimensional processing of biomedical literature data based on accurate traceability of biomedical literature data.

[0006] In a first aspect, embodiments of this application provide a biomedical literature analysis method, comprising: retrieving several data sources to obtain literature data and unique identifiers corresponding to the several data sources; reordering the literature data based on preset indicators and then performing hierarchical deduplication; dividing the hierarchically deduplicated literature data into sentence-level evidence blocks and binding a corresponding unique identifier to each sentence-level evidence block; responding to a literature analysis request, selecting the corresponding sentence-level evidence block based on the unique identifier and performing one or more of the following tasks: performing the review task to generate structured review text; performing the question-and-answer task to generate answer text; performing the extraction task to identify and output a structured parameter table; and presenting all results of the executed tasks.

[0007] In some embodiments, the execution process of the review task includes: generating the structured review text based on the selected sentence-level evidence blocks via an outline generation stage, a section draft generation stage, a section summary generation stage, and a final review generation stage.

[0008] In a further embodiment, the execution process of the review task further includes establishing a reflection mechanism, which is applied to the outline generation stage, the section draft generation stage, and the final review generation stage. The reflection mechanism includes at least: calculating the keyword matching degree and medical subject term coverage of the sentence-level evidence blocks; performing logic and evidence comparison tests on the sentence-level evidence blocks to calculate the logical consistency of the sentence-level evidence blocks; calculating the journal impact factor of the sentence-level evidence blocks and performing citation format verification on the sentence-level evidence blocks to calculate the authority of the sentence-level evidence blocks; evaluating the terminology consistency of the sentence-level evidence blocks based on a medical terminology thesaurus; establishing a comprehensive evaluation result based on the keyword matching degree, the medical subject term coverage, the logical consistency, the authority, and the terminology consistency; correcting the corresponding sentence-level evidence blocks based on the comprehensive evaluation result; and outputting the corrected structured review text.

[0009] In some other embodiments, the execution process of the question-answering task includes: calculating the source reliability, overall consistency, and number of the sentence-level evidence blocks to calculate the confidence score of the sentence-level evidence blocks; comparing the confidence score with a preset confidence threshold, and outputting a low-confidence prompt when the confidence score is lower than the confidence threshold; and performing a domain-specific question answer based on the sentence-level evidence blocks and outputting the answer text when the confidence score is not lower than the confidence threshold. Specifically, the conflict of the sentence-level evidence blocks is detected based on the source reliability and overall consistency corresponding to the sentence-level evidence blocks, and the corresponding sentence-level evidence blocks are listed and marked when a conflict occurs.

[0010] In some other embodiments, the extraction task includes: tracing the literature data corresponding to the unique identifier bound to the sentence-level evidence block; extracting the required biomedical parameters based on the sentence-level evidence block or the corresponding literature data; processing the biomedical parameters in combination with a set rule template, regular expression matching and entity recognition algorithm to obtain the processing result; performing standardization processing on the processing result and outputting the structured parameter table.

[0011] In some other embodiments, presenting all results of the executed task includes: obtaining all results of the executed task and outputting them based on a streaming response; limiting the context length of all results based on a lexical management mechanism, and adjusting the output progress and output format of all results.

[0012] In some other embodiments, a single literature analysis request may include one or more of a review task request, a question-and-answer task request, and an extraction task request; when the review task request or the question-and-answer task request is identified in a single literature analysis request, the structured review text or the answer text is output accordingly; then, it is identified whether the extraction task request is present in the single literature analysis request.

[0013] In some other embodiments, the step of re-ranking the literature data based on preset indicators includes: establishing preset indicators, including at least journal impact factor, time factor, and type factor; setting weights for the journal impact factor, the time factor, and the type factor respectively, and then performing weighted ranking on the literature data.

[0014] In some other embodiments, the hierarchical deduplication process includes performing deduplication on the reordered document data at least based on the Digital Object Unique Identifier (DOI) level, the Uniform Resource Locator (URL) level, and the title level.

[0015] Secondly, embodiments of this application provide a biomedical literature analysis system, comprising: an evidence backbone module, used to retrieve several data sources to obtain literature data and unique identifiers corresponding to the several data sources, and to perform hierarchical deduplication processing on the literature data after reordering it based on preset indicators, dividing the hierarchically deduplicated literature data into sentence-level evidence blocks, and binding a corresponding unique identifier to each sentence-level evidence block; a review proxy module, used to select the corresponding sentence-level evidence block based on the unique identifier to perform a review task to generate structured review text; a question-and-answer proxy module, used to select the corresponding sentence-level evidence block based on the unique identifier to perform a question-and-answer task to generate answer text; an extraction proxy module, used to select the corresponding sentence-level evidence block based on the unique identifier to perform an extraction task to identify and output a structured parameter table; and an interaction module, used to respond to literature analysis requests, and to communicate with the review proxy module, the question-and-answer proxy module, and the extraction proxy module based on the literature analysis requests, to perform one or more of the review task, the question-and-answer task, and the extraction task, and to present all execution results of the executed tasks.

[0016] Compared with existing technical solutions, the technical solutions provided in this application have the following advantages:

[0017] First, unlike existing technical solutions that rely on pre-built, static knowledge bases as sources of evidence, this application's embodiments dynamically construct a high-quality collection of documents that has been filtered for authority and timeliness by "searching several data sources" and "re-sorting the literature data based on preset indicators and then performing hierarchical deduplication processing," thereby taking into account both the multi-source nature and reliability of the literature data.

[0018] Secondly, unlike existing technical solutions that lack fine-grained control over evidence, the embodiments of this application divide the processed document data into sentence-level evidence blocks and bind a corresponding unique identifier to each sentence-level evidence block, thereby achieving precise source tracing at the sentence level and binding the sentence-level evidence blocks to the original documents, significantly enhancing their verifiability and traceability.

[0019] Third, unlike existing technical solutions that are overly focused and singular in their processing of medical literature data, the embodiments of this application "respond to the literature analysis request, select the corresponding sentence-level evidence block based on the unique identifier, and perform one or more tasks among the review task, question-and-answer task, and extraction task." This multi-agent collaborative architecture based on a unified evidence backbone can break down functional silos while achieving sentence-level evidence block tracing, enabling users to seamlessly obtain multi-dimensional analysis results from macro reviews to micro data in one process.

[0020] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description

[0021] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0022] Figure 1 This is a flowchart of a biomedical literature analysis method provided in one embodiment of this application;

[0023] Figure 2 This is an architecture diagram of a biomedical literature analysis system provided in one embodiment of this application;

[0024] Figure 3 This is a flowchart of the overview task execution process provided in an embodiment of this application;

[0025] Figure 4 This is a flowchart illustrating the execution process of the reflection mechanism provided in one embodiment of this application;

[0026] Figure 5 This is a flowchart of the question-and-answer task execution process provided in an embodiment of this application;

[0027] Figure 6 This is a flowchart illustrating a specific method for analyzing biomedical literature provided in one embodiment of this application;

[0028] Figure 7 This is a specific timing diagram of a biomedical literature analysis method provided in one embodiment of this application. Detailed Implementation

[0029] To better understand the purpose, technical solution, and advantages of this application, the application is described and illustrated below in conjunction with the accompanying drawings and embodiments.

[0030] Unless otherwise defined, the technical or scientific terms used in this application shall have the general meaning understood by one of ordinary skill in the art to which this application pertains. Words such as “a,” “an,” “an,” “the,” “the,” and “these” used in this application do not indicate quantitative limitation and may be singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include steps or modules (units) not listed, or may include other steps or modules (units) inherent to these processes, methods, products, or devices. Words such as “connected,” “linked,” and “coupled” used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. Normally, the character " / " indicates that the objects before and after it are in an "or" relationship. The terms "first," "second," "third," etc., used in this application are merely used to distinguish similar objects.

[0031] The method embodiments provided in this example can be executed in a terminal, computer, or similar computing device. For example, when running on a terminal, the terminal may include one or more processors and a memory for storing data, wherein the processor may include, but is not limited to, a microprocessor (MCU) or a programmable logic device (FPGA). The terminal may also include transmission devices for communication functions and input / output devices, and may include more or fewer components, or have different configurations.

[0032] The memory can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the biomedical literature analysis method in this embodiment. The processor executes various functional applications and data processing by running the computer program stored in the memory, thereby implementing the above-described method. The memory may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory may further include memory remotely located relative to the processor, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0033] Please see the appendix Figure 1 In one exemplary embodiment, a biomedical literature analysis method is provided, the method comprising the following steps:

[0034] Step S101: Retrieve several data sources to obtain literature data and unique identifiers corresponding to each data source. In this step, retrieving data from several data sources can comprehensively cover authoritative databases, open network resources, and user-personalized data, ensuring the comprehensiveness and diversity of literature collection. Obtaining unique identifiers corresponding to data sources can establish a reliable foundation for subsequent evidence tracing, ensuring that each data source can be uniquely identified and cited.

[0035] Step S102: After re-ranking the literature data based on preset indicators, hierarchical deduplication is performed. The deduplicated literature data is divided into sentence-level evidence blocks, and a unique identifier is assigned to each sentence-level evidence block. In this step, re-ranking the literature data prioritizes the display of high-quality, highly relevant literature, enhancing the authority and timeliness of subsequent analysis results. Hierarchical deduplication effectively eliminates redundant information, preventing duplicate evidence from interfering with the analysis results. Importantly, dividing the processed literature data into sentence-level evidence blocks enables fine-grained organization of evidence, providing a foundation for accurate citation and evidence comparison. Assigning a unique identifier to each sentence-level evidence block establishes a precise link from the evidence to the original literature, ensuring the traceability and verifiability of the analysis process.

[0036] Step S103: In response to the literature analysis request, select the corresponding sentence-level evidence block based on the unique identifier and execute one or more of the following tasks: execute a review task to generate a structured review text; execute a question-and-answer task to generate answer text; execute an extraction task to identify and output a structured parameter table; and present all results of the executed tasks. In this step, the literature analysis request is generally initiated by the user, containing the task request the user intends to execute. Therefore, it clearly indicates the type of task to be executed. Thus, in response, the corresponding task can be activated and executed by analyzing the content of the literature analysis request, and the results obtained. The unique identifier, as the unique identity credential of the sentence-level evidence block, supports the rapid location, filtering, and retrieval of sentence-level evidence blocks related to the current task. Based on the above technical means, this step can construct an execution system that integrates a unified evidence backbone and multi-task agent collaboration, enabling the tracing of sentence-level evidence blocks while breaking down functional silos.

[0037] Therefore, based on the above-mentioned technical means and the beneficial effects they produce, this embodiment effectively solves the problem in related technologies that it is impossible to achieve intelligent multi-dimensional processing of biomedical literature data on the basis of accurate tracing of biomedical literature data.

[0038] In some embodiments, several data sources may be selected from the PubMed (United States Library of Medicine) database, filtered open online resources, and user-constructed biomedical vector databases. Other professional databases or user-uploaded data resources may also be chosen. It is understood that data sources may also be user-uploaded data, such as PDF documents, Word documents, CSV / TSV spreadsheet files, or ZIP archives containing research data. This improves the hybrid data source system constructed in this embodiment, which includes both authoritative publicly available data and caters to specific user research needs, thereby enhancing the flexibility and relevance of the method application.

[0039] In other embodiments, the process of re-ranking literature data based on preset indicators may further include: establishing preset indicators, including at least journal impact factor, time factor, and type factor; assigning weights to the journal impact factor, time factor, and type factor respectively, and then performing a weighted ranking of the literature data. The preset indicators include the aforementioned factors, which respectively serve to assess the authority, timeliness, and type relevance of the literature. Furthermore, by assigning weights to the aforementioned factors and performing a weighted ranking of the literature data based on the factors and their corresponding weights, the ranking strategy can be flexibly adjusted according to different analytical needs. For example, the weight of the time factor can be increased when quickly understanding the research frontier, and the weight of the impact factor can be increased when conducting in-depth reviews.

[0040] It should be understood that, in further embodiments, the preset indicators may also include factors such as citation count and research institution influence, which can further refine the dimensions of literature quality assessment.

[0041] It is understood that in other further embodiments, the weights may be specifically configured as follows: influence factor weight 0.4-0.5, time factor weight 0.3-0.4, and type factor weight 0.1-0.2.

[0042] In other embodiments, the hierarchical deduplication process for the reordered document data may further include: performing deduplication on the reordered document data at least based on the Digital Object Identifier (DOI) level, the Uniform Resource Locator (URL) level, and the title level. The DOI level is the unique digital identifier level for documents; performing deduplication at this level can accurately identify and remove completely identical document records. The URL level is the document's web address level; performing deduplication at this level can identify and remove the same document from different sources. Furthermore, performing deduplication at the title level can identify and remove highly similar documents based on semantic similarity, further purifying the evidence set.

[0043] It should be understood that, in further embodiments, the hierarchical deduplication processing can also be adapted to author combinations, publishing institutions, etc., which can improve the deduplication mechanism from different dimensions and ensure the purity and representativeness of the evidence set.

[0044] It should be noted that a literature analysis request contains the user's specific analytical intent and target information. In some embodiments, the literature analysis request is typically initiated by the user. After breaking down the literature analysis request, it becomes clear that it requires the execution of one or more of the three tasks mentioned above. Therefore, it is understandable that the appropriate combinations include executing the review task, the question-and-answer task, or the extraction task individually, or choosing two or all of them. For example, in some specific embodiments, the review task can be executed alone, suitable for scenarios requiring systematic literature review. In other specific embodiments, the question-and-answer task can be executed alone, suitable for scenarios requiring rapid acquisition of specific knowledge points. In other specific embodiments, the extraction task can be executed alone, suitable for scenarios requiring structured data extraction. Furthermore, in other specific embodiments, a combination of review and question-and-answer tasks can be executed to meet the dual needs of in-depth analysis and immediate Q&A. In other specific embodiments, a combination of question-and-answer and extraction tasks can be executed to achieve simultaneous output of answers and supporting data. In other specific embodiments, a combination of review and extraction tasks can be executed to provide a complete chain of evidence, encompassing both macro-level discussion and micro-level data. In other specific embodiments, all three tasks can be executed simultaneously to complete an end-to-end analysis process from background knowledge to data support. In summary, the specific content of the literature analysis request needs to be analyzed, and the order of processing the three tasks can be specifically designed during the implementation process.

[0045] For example, in a specific embodiment, the task flow can be processed according to the principle of narrative priority and extraction on demand. That is, the literature analysis request and / or question-and-answer request in the literature analysis request are responded to first. After the literature analysis request includes the literature review task and / or question-and-answer task request, and the processing results are obtained, if the literature analysis request includes an extraction task request, then the extraction task is executed, the processing results are obtained, and the results of all executed tasks are presented in a comprehensive manner. The advantage of this method compared to the prior art is that it can achieve a natural transition from qualitative analysis to quantitative mining, which meets the thinking habit of researchers to "understand the whole picture first, and then dig deeper," while ensuring that the structured data table and the narrative content are based on the same evidence foundation, avoiding the disconnect between data and conclusions. Users can dynamically decide whether to start deep data extraction based on the previous results, improving the system's flexibility and practicality. It should be understood that this embodiment is only an example and not a limitation on other implementation methods.

[0046] In some other embodiments, the review task execution process includes: generating a biomedical review text based on sentence-level evidence blocks, through an outline generation stage, a section draft generation stage, a section summary generation stage, and a final review generation stage. Specifically, the outline generation stage determines the overall logical framework and main content modules of the review; the section draft generation stage fills in the specific content of each module based on the evidence blocks; the section summary generation stage extracts the core viewpoints and evidence of each section; and the final review stage integrates and optimizes the content of each chapter to form a coherent and professional academic text. Sentence-level evidence blocks provide accurate and traceable evidence, and the aforementioned stages sequentially construct the framework, enrich the content, extract the essence, and optimize the integration.

[0047] It should be understood that, in one specific embodiment, in response to a literature analysis request for a review, such as "the relationship between psoriasis and intestinal barrier dysfunction," 10–20 subqueries are generated, and a multi-source search is performed. The final review text is obtained through stages including generating an outline, generating draft sections, extracting section abstracts, and generating the final review. Each stage is based on sentence-level evidence blocks to ensure that every argument in the review is supported by solid literature and is precisely cited using unique identifiers.

[0048] In a further embodiment, please refer to the appendix. Figure 4 In the process of reviewing tasks, a reflection mechanism can also be introduced. The process is as follows: when the review task is executed and the judgment of failing the reflection mechanism is triggered, the corresponding sentence-level evidence block is retrieved, and at least the keyword matching degree, medical subject term coverage, calculation logic consistency, calculation authority, and calculation terminology consistency are checked. The above contents are combined to establish a comprehensive evaluation result and generate correction suggestions. Finally, the corresponding sentence-level evidence block is corrected.

[0049] In a further embodiment, the reflection mechanism is applied to the outline generation stage, the section draft generation stage, and the final review generation stage. Specifically, the keyword matching degree and medical subject term coverage of sentence-level evidence blocks are calculated; logical and evidence comparison tests are performed on sentence-level evidence blocks to calculate the logical consistency of the sentence-level evidence blocks; the journal impact factor of sentence-level evidence blocks is calculated, and citation format checks are performed on sentence-level evidence blocks to calculate the authority of the sentence-level evidence blocks; the terminology consistency of sentence-level evidence blocks is evaluated based on a medical terminology thesaurus; a comprehensive evaluation result is established based on keyword matching degree, medical subject term coverage, logical consistency, authority, and terminology consistency; the corresponding sentence-level evidence blocks are corrected based on the comprehensive evaluation result, and the corrected structured review text is output.

[0050] The calculation of keyword matching degree and medical subject term coverage aims to ensure that the review comprehensively covers core concepts and domain knowledge systems, avoiding the omission of important topics. The calculation of the logical consistency of sentence-level evidence blocks aims to detect and eliminate conflicts between different pieces of evidence, ensuring the logical rigor of the argument. The reflection mechanism is applied to the outline generation stage, the section draft generation stage, and the final review generation stage to achieve full-process quality control, promptly identifying and correcting problems at each key node of content construction. Verifying the authority of sentence-level evidence blocks based on journal impact factor and citation format checks prioritizes the adoption of high-quality research results while ensuring standardized citation formats, thus enhancing the academic credibility of the review. It should be understood that the implementation of the reflection mechanism may also include other evaluation dimensions that can affect the quality and reliability of sentence-level evidence blocks, such as the rigor of research methods and the adequacy of sample size.

[0051] In a further embodiment, with the introduction of a reflection mechanism, the execution flow of this method can be found in the appendix. Figure 3 After selecting the corresponding sentence-level evidence block as input parameter based on the unique identifier, an outline is generated in the outline generation stage. This outline then undergoes a reflection mechanism for evaluation. If the reflection mechanism passes, the process proceeds to the section draft generation stage; otherwise, the corresponding sentence-level evidence block is corrected based on the reflection mechanism, and the process returns to the outline generation stage. In the section draft generation stage, a section draft is first generated, which also undergoes a reflection mechanism for evaluation. If the reflection mechanism passes, the process proceeds to the section summary generation stage; otherwise, the corresponding sentence-level evidence block is corrected, and the process returns to the outline generation stage. After generating the section summary in the section summary generation stage, the process proceeds to the final review generation stage, which again requires a reflection mechanism for evaluation. If the reflection mechanism passes, the structured review text is output; otherwise, it is corrected again, and the process returns to the outline generation stage. It should be understood that the stages restricted by the reflection mechanism in the above process are merely illustrative and do not necessarily require strict adherence to the stage names. In practical applications, reflection mechanisms can be established at different required stages, as long as the triggering conditions of the reflection mechanism are met; such designs are considered reasonable.

[0052] In other embodiments, the execution of the question-answering task needs to be based on an answerable gate control mechanism, specifically including: calculating the source reliability, overall consistency, and number of sentence-level evidence blocks to calculate the confidence score of the sentence-level evidence blocks; comparing the confidence score with a preset confidence threshold, and outputting a low-confidence prompt when the confidence score is lower than the confidence threshold; and performing a question answer within the domain and outputting the answer text based on the sentence-level evidence blocks when the confidence score is not lower than the confidence threshold. Specifically, the conflict of sentence-level evidence blocks is detected based on the source reliability and overall consistency of the sentence-level evidence blocks, and the corresponding sentence-level evidence blocks are listed and marked when a conflict occurs.

[0053] The overall consistency score assesses the degree of consistency among different sentence-level evidence blocks across multiple dimensions, including objective statements, data support, and conclusion derivation. Its calculation comprehensively considers logical consistency, data consistency, and conclusion consistency. Furthermore, the calculation of source reliability, overall consistency, and quantity comprehensively evaluates the credibility of the answer from three dimensions: evidence quality, inter-evidence consistency, and evidence sufficiency. This allows the confidence score to reflect the reliability of the current evidence set in supporting the answer to the question. Setting a confidence threshold constrains the confidence score, indirectly controlling the threshold for outputting definitive answers and preventing misleading content when evidence is insufficient. Moreover, detecting conflicts among sentence-level evidence blocks can promptly identify and alert users to points of contention in their research. The use of source reliability and overall consistency considers both the authority of individual pieces of evidence and the degree of consistency among multiple pieces of evidence in terms of objective statements, data support, and conclusion derivation, forming a comprehensive evaluation system. In addition, inline citations refer to citation markers that are directly embedded in the answer text and point to the unique identifier of the sentence-level evidence block. Their purpose is to enable precise tracing of the source of each argument in the answer and enhance the verifiability of the results.

[0054] In a further embodiment, please refer to the appendix. Figure 5 The execution flow of the answering task can be set as follows: calculate the source reliability, overall consistency, and number of sentence-level evidence blocks to calculate the confidence score of the sentence-level evidence blocks, and simultaneously determine whether there is evidence conflict. If evidence conflict exists, detect the conflict status of sentence-level evidence blocks, list and mark the corresponding sentence-level evidence blocks. If no evidence conflict exists, determine whether the confidence score is lower than the confidence threshold. If yes, output a low confidence prompt; otherwise, perform question answering within the domain based on the sentence-level evidence blocks and output the answer text. It should be noted that in this embodiment, the order of evidence conflict judgment and low confidence threshold judgment is specified; however, in other embodiments, different execution orders can be set.

[0055] For example, in some other embodiments, the determination of confidence score and confidence threshold can be performed first, followed by the determination of evidence conflict. In other embodiments, the determination of confidence threshold and determination of evidence conflict can be performed simultaneously; any of these methods are acceptable as long as they conform to the implementation logic.

[0056] In one specific embodiment, in response to a document analysis request containing content such as "Does glutamine help improve psoriasis symptoms?", 5–10 subqueries are generated, and a multi-source search is performed. A concise answer is generated based on the evidence block, and inline citations are inserted into the answer. A confidence level is calculated for the evidence block using an answerability gating mechanism: "Uncertain" is output when evidence is insufficient, and conflicting evidence is output with a warning to the user to interpret it with caution when conflicting evidence exists.

[0057] In other specific embodiments, inline citations are generated by automatically inserting the unique identifier of the sentence-level evidence block used as the basis for the answer after the corresponding statement in a standard citation format when generating the answer. In one specific embodiment, the answerability gate control mechanism specifically includes: outputting an "uncertain" prompt when the confidence score is lower than the confidence threshold; outputting the answer text when the confidence score is not lower than the confidence threshold; and listing and marking conflicting sentence-level evidence blocks when conflicts occur, with a prompt of "interpret with caution".

[0058] In other embodiments, the extraction task includes: tracing the corresponding literature data based on the unique identifier bound to the sentence-level evidence block; extracting the required biomedical parameters based on the sentence-level evidence block or the corresponding literature data; processing the biomedical parameters using a set rule template, regular expression matching, and entity recognition algorithm to obtain the processing result; performing standardization processing on the processing result and outputting a structured parameter table. It should be explained that: the rule template is used to define the extraction pattern for specific types of parameters; regular expression matching is used to accurately capture numerical values ​​and units; and the entity recognition algorithm is used to accurately identify relevant biomedical entities. These three work together to ensure the accuracy and completeness of parameter extraction.

[0059] In one specific embodiment, please refer to the appendix. Figure 6 The unique identifiers associated with sentence-level evidence blocks are traced back to user-uploaded PDF files or selected documents, whose content relates to a set of relevant documents selected by the user through review and question-answering tasks. Furthermore, based on the responding document analysis request, an extraction task is required. Therefore, the aforementioned PDF files are converted into structured text through text parsing or structured processing, and a task template is loaded. The following steps are performed: entity recognition, parameter value extraction, unit normalization, record alignment, and tabular output to obtain the corresponding structured parameter table.

[0060] In the entity recognition process, information such as enzymes, substrates, and conditions can be identified. For example, in enzyme kinetics tasks, the Michaelis constant Km, catalytic constant kcat, and specificity constant kcat / Km can be identified and output as standardized spreadsheet files (CSV or Excel), which can be directly imported into a database. In some specific application scenarios, the extraction task execution process in this embodiment is applicable not only to enzyme kinetic parameters but also to the extraction of biomedical data such as drug dosage, gene mutations, and disease indicators.

[0061] In other embodiments, in response to a literature analysis request, one or more tasks, including a review task, a question-answering task, and an extraction task, are performed based on sentence-level evidence blocks. All execution results of the performed tasks are then presented. This complete process may further include: identifying the content of the literature analysis request and obtaining all execution results of the performed tasks; outputting all execution results based on a streaming response; and limiting the context length of all execution results based on a token management mechanism, while also adjusting the output progress and format. Specifically, when processing long text generation, the system dynamically adjusts the output pace by calculating the number of tokens in the generated content in real time, ensuring high-quality content generation is completed within the model context window constraints.

[0062] It should be understood that, in further embodiments, the interaction module can implement streaming responses based on Server-SentEvents (SSE) technology, allowing users to receive results gradually during the generation process. Furthermore, the system controls the context length through a lexical management mechanism to avoid exceeding limits in large-scale interactions.

[0063] Please see the appendix Figure 2 In one exemplary embodiment, a biomedical literature analysis system is provided, the system comprising:

[0064] The evidence backbone module is used to retrieve several data sources to obtain literature data and unique identifiers corresponding to several data sources. It is also used to perform hierarchical deduplication after reordering the literature data based on preset indicators, divide the literature data after hierarchical deduplication into sentence-level evidence blocks, and bind a corresponding unique identifier to each sentence-level evidence block.

[0065] The review proxy module is used to select the corresponding sentence-level evidence block based on a unique identifier to perform the review task and generate a structured review text.

[0066] The question-and-answer proxy module is used to select the corresponding sentence-level evidence block based on a unique identifier to perform the question-and-answer task and generate the answer text;

[0067] The extraction proxy module is used to select the corresponding sentence-level evidence block based on the unique identifier to perform the extraction task, so as to identify and output the structured parameter table;

[0068] The interaction module is used to respond to literature analysis requests and communicate with the review proxy module, question answering proxy module, and extraction proxy module based on the literature analysis requests to execute one or more of the review task, question answering task, and extraction task, and present all execution results of the executed tasks.

[0069] In some specific embodiments, the appendix Figure 2The labeled summary proxy module, question-answering proxy module, and extraction proxy module can be configured to belong to the same intelligent proxy system and be uniformly scheduled and coordinated by the intelligent proxy. In some more specific embodiments, the intelligent proxy can also integrate a task allocation module and a result fusion module to provide task parsing and result synthesis functions. This allows it to work with the summary proxy module, question-answering proxy module, and extraction proxy module to execute complex multi-round analysis tasks, achieving evidence sharing and conclusion cross-verification among tasks. This architecture design enables the intelligent proxy to automatically plan task execution paths based on the complexity of user requests, and to cross-verify and organically integrate the results after multiple tasks are completed, ultimately outputting a consistent and comprehensive analysis conclusion.

[0070] It should be understood that the biomedical literature analysis system provided in this embodiment belongs to the same application concept as the biomedical literature analysis method provided in the above embodiments of this application. It can execute the method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects of the method. Technical details not described in detail in this embodiment can be found in the specific processing content of the biomedical literature analysis method provided in the above embodiments of this application, and will not be repeated here.

[0071] In a further embodiment, the system integrates a primary and backup large language model switching module, that is, when the primary model fails to be called, it automatically switches to the backup model to continue generation, so as to ensure the high availability of the system in biomedical literature analysis tasks.

[0072] In other embodiments, please refer to the appendix. Figure 7 The diagram illustrates the workflow of a narrative-first, on-demand extraction approach, with the specific execution process as follows: First, the user initiates a literature analysis request to either the review proxy module or the question-and-answer proxy module. These modules identify the review task request or the answer task request, and then stream the narrative results containing evidence citations via SSE. Simultaneously, based on the user's literature analysis request, sentence-level evidence and source literature data are filtered using unique identifiers. The review proxy module and the question-and-answer proxy module execute the corresponding tasks and provide the execution results. After obtaining the execution results of the aforementioned tasks, the extraction proxy module identifies the extraction task request from the literature analysis request and executes the corresponding extraction task to output a structured parameter table.

[0073] It is understood that in some specific embodiments, if the literature analysis request initiated by the user only includes an extraction task request, then only the extraction task needs to be executed and the corresponding structured parameter table output needs to be generated. Similarly, in other specific embodiments, if the literature analysis request initiated by the user includes both a review task request and a question-answering task request, then the tasks can be executed simultaneously or sequentially. The order and timing of task execution can be coordinated by the interaction module or other related modules. It is clear that the above embodiments are merely exemplary, used to express the possible order of execution for different tasks, and are not strictly limited, nor will they affect the specific implementation of other embodiments. Other implementation methods should still be performed according to the actual situation.

[0074] It should be understood that, in some other embodiments, the present invention can also be extended to the fields of clinical medicine and drug discovery. In some embodiments of drug reuse scenarios, an extraction task can be performed to extract the correspondence between drugs, diseases, and genes, and generate a list of potential drugs. In some embodiments of clinical research scenarios, an extraction task can be performed to extract parameters such as sample size, dosage, and efficacy from clinical trial literature, forming a clinical data table. In molecular biology scenarios, the system can extract the correspondence between gene mutations, protein function parameters, and disease phenotypes, and obtain corresponding data tables.

[0075] The functions implemented by each module in the above biomedical literature analysis system can be implemented by the same or different processors, and this application embodiment does not limit this.

[0076] In addition to the methods and systems described above, embodiments of this application may also be computer program products or storage media. These computer program products include computer program instructions that, when executed by a processor, cause the processor to perform the steps of the biomedical literature analysis method described in any of the above embodiments of this specification. The computer program product can be written in any combination of one or more programming languages ​​to perform the operations of the embodiments of this application. The programming languages ​​include object-oriented programming languages ​​and conventional procedural programming languages, such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. The storage medium stores the computer program product, which, when executed by a processor, implements the biomedical literature analysis method of any of the above embodiments of this specification.

[0077] It should be understood that the specific embodiments described herein are merely illustrative of the application and not intended to limit it. All other embodiments derived by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.

[0078] Obviously, the accompanying drawings are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar situations based on these drawings without any creative effort. Furthermore, it is understood that although the work done in this development process may be complex and lengthy, for those skilled in the art, certain design, manufacturing, or production modifications made based on the technical content disclosed in this application are merely conventional technical means and should not be considered as insufficient disclosure of this application.

[0079] The term "embodiment" in this application refers to a specific feature, structure, or characteristic described in connection with an embodiment that may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily imply the same embodiment, nor does it imply that it is mutually exclusive with or independent of other embodiments. It will be clearly or implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.

[0080] The embodiments described above are merely examples of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these modifications and improvements all fall within the scope of protection of this application.

Claims

1. A biomedical literature analysis method, characterized in that, include: Search several data sources to obtain literature data and unique identifiers corresponding to the several data sources; After re-sorting the literature data based on preset indicators, hierarchical deduplication is performed. The document data after hierarchical deduplication is divided into sentence-level evidence blocks, and each sentence-level evidence block is bound with a corresponding unique identifier. In response to the literature analysis request, the corresponding sentence-level evidence block is selected based on the unique identifier, and one or more of the following tasks are performed: Perform review tasks to generate structured review texts; Perform a question-and-answer task to generate answer text; Perform the extraction task to identify and output a structured parameter table; Display all results of the tasks that have been performed.

2. The biomedical literature analysis method according to claim 1, characterized in that, The execution process of the review task includes: Based on the selected sentence-level evidence blocks, the structured review text is generated through the outline generation stage, the section draft generation stage, the section summary generation stage, and the final review generation stage.

3. The biomedical literature analysis method according to claim 2, characterized in that, The execution process of the review task also includes establishing a reflection mechanism, and the reflection mechanism is applied to the outline generation stage, the section draft generation stage, and the final review generation stage. The reflection mechanism includes at least the following: Calculate the keyword matching degree and medical subject term coverage of the sentence-level evidence block; The sentence-level evidence block is subjected to a logical and evidence comparison test to calculate the logical consistency of the sentence-level evidence block. Calculate the journal impact factor of the sentence-level evidence block and perform citation format validation on the sentence-level evidence block to calculate the authority of the sentence-level evidence block; The terminology consistency of the sentence-level evidence blocks was assessed based on a medical terminology thesaurus. A comprehensive evaluation result is established based on the keyword matching degree, the medical subject term coverage, the logical consistency, the authority, and the terminology consistency. The corresponding sentence-level evidence blocks are then corrected based on the comprehensive evaluation result, and the corrected structured review text is output.

4. The biomedical literature analysis method according to claim 1, characterized in that, The execution process of the question-and-answer task includes: The source reliability, overall consistency, and number of the sentence-level evidence blocks are calculated to determine the confidence score of the sentence-level evidence blocks. The confidence score is compared with a preset confidence threshold, and a low confidence prompt is output when the confidence score is lower than the confidence threshold. When the confidence score is not lower than the confidence threshold, the domain-specific question answering is performed based on the sentence-level evidence block and the answer text is output. Specifically, the system detects conflicts between corresponding sentence-level evidence blocks based on the source reliability and the overall consistency, and lists and marks the corresponding sentence-level evidence blocks when conflicts occur.

5. The biomedical literature analysis method according to claim 1, characterized in that, The execution process of the extraction task includes: Based on the unique identifier bound to the sentence-level evidence block, trace the corresponding literature data, extract the required biomedical parameters based on the sentence-level evidence block or the corresponding literature data, and process the biomedical parameters in combination with the set rule template, regular expression matching and entity recognition algorithm to obtain the processing result; The processing result is standardized, and the structured parameter table is output.

6. The biomedical literature analysis method according to claim 1, characterized in that, The presentation of all results of the executed tasks includes: Retrieve all results from executed tasks and output them based on streaming responses; The context length of all results is limited by a lexical management mechanism, and the output progress and format of all results are adjusted.

7. The biomedical literature analysis method according to claim 1, characterized in that, Also includes: A single literature analysis request may include one or more of the following: review task request, question answering task request, and extraction task request; When the review task request or the question-and-answer task request is identified in a single literature analysis request, the structured review text or the answer text is output accordingly. Then identify whether the extraction task request is included in a single document analysis request.

8. The biomedical literature analysis method according to claim 1, characterized in that, The re-ranking of the literature data based on preset indicators includes: Establish pre-defined indicators, including at least journal impact factor, time factor, and type factor; After assigning weights to the journal impact factor, the time factor, and the type factor, the literature data is then sorted using weighted methods.

9. The biomedical literature analysis method according to claim 1, characterized in that, The hierarchical deduplication process includes: At least based on the unique identifier level of digital objects, the Uniform Resource Locator (URL) level, and the title level, deduplication is performed on the literature data that is being reordered.

10. A biomedical literature analysis system, characterized in that, include: The evidence backbone module is used to retrieve several data sources to obtain literature data and unique identifiers corresponding to the several data sources, and to perform hierarchical deduplication processing after reordering the literature data based on preset indicators, dividing the literature data after hierarchical deduplication processing into sentence-level evidence blocks, and binding a corresponding unique identifier to each sentence-level evidence block. The review proxy module is used to select the corresponding sentence-level evidence block based on the unique identifier to perform the review task and generate a structured review text. The question-and-answer agent module is used to select the corresponding sentence-level evidence block based on the unique identifier to perform the question-and-answer task and generate the answer text; The extraction proxy module is used to select the corresponding sentence-level evidence block based on the unique identifier to perform the extraction task, so as to identify and output a structured parameter table; The interaction module is used to respond to literature analysis requests, and based on the literature analysis requests, communicate with the review proxy module, the question-answering proxy module, and the extraction proxy module to execute one or more of the review task, the question-answering task, and the extraction task, and present all execution results of the executed tasks.