Intelligent processing and analysis system, method and terminal for medical literature based on large model
The intelligent medical literature processing system based on a large model solves the problems of cumbersome information acquisition and complex analysis in medical literature management and analysis tools, and achieves efficient intelligent processing and analysis of literature, thereby improving scientific research efficiency.
Patent Information
- Application Number
- CN202411871645.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2026-02-17
- Estimated Expiration
- 2044-12-18
AI Technical Summary
Existing medical literature management and analysis tools cannot fully acquire and process non-open access literature, resulting in poor information mining depth and knowledge extraction effect, as well as high computational resource consumption.
A medical literature intelligent processing system based on a large model is adopted, including a user interaction module, a biological entity module, and a literature abstract knowledge graph module. It performs intelligent analysis by recognizing user intent, extracting biological entities, generating logical chain instructions, and combining the literature abstract knowledge graph database.
It has improved the efficiency of information acquisition and analysis accuracy of medical literature, reduced the cost of manual analysis, and increased the work efficiency of researchers.
Smart Images

Figure CN119807435B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical literature processing, and particularly relates to a medical literature intelligent processing and analysis system and method based on a large model and a terminal. BACKGROUND
[0002] With the rapid development of scientific research, the annual publication volume of scientific literature is growing exponentially. This growth trend brings great challenges to researchers, especially in terms of efficiently obtaining, reading and analyzing a large number of literatures. Traditional literature retrieval and reading methods are often inefficient and difficult to cope with the information overload problem brought by massive literature.
[0003] There are various literature management and analysis tools in the current market, such as Aminor and Elicit, etc. These tools use large language models to realize intelligent retrieval, automatic abstract generation and interactive literature question and answer functions. However, these tools have significant limitations in medical literature information acquisition: the information sources are mainly from the metadata of the literature (such as title, abstract) or the PDF full text uploaded by the user manually, and cannot fully acquire and process the complete content including non-open access literature. This limitation directly affects the depth of literature content mining and the effect of knowledge extraction. In addition, for PDF full text analysis, existing tools usually use the method of directly inputting full text information into a large language model. However, considering the current limitation of large language models in the length of context, this processing method inevitably introduces a large amount of redundant information. This not only reduces the accuracy and efficiency of analysis, but also significantly increases the consumption of computing resources. SUMMARY
[0004] In view of the above-mentioned shortcomings of the prior art, the purpose of the present application is to provide a medical literature intelligent processing and analysis system and method based on a large model and a terminal, which solves the technical problems of complicated literature acquisition, time-consuming reading and complex analysis of the existing medical literature management and analysis tools.
[0005] To achieve the above object and other related objects, the present application provides a medical literature intelligent processing and analysis system based on a large model, which comprises: a user interaction module for receiving a query request of a user through a user search interface; wherein the types of the query request include a literature keyword query request or a literature related question query request; a user intention recognition module connected to the user interaction module, for recognizing the user intention based on a large model according to the query request of the user, extracting biological entities and generating logical chain instructions; a biological entity module connected to the user intention recognition module, for retrieving matching corresponding entity background information from a pre-constructed biological entity database based on the extracted biological entities; a literature abstract knowledge graph module connected to the biological entity module, for retrieving matching corresponding related literature information from a pre-constructed literature abstract knowledge graph database based on a large model according to the extracted biological entities and the generated logical chain instructions; and a query result acquisition module connected to the user interaction module, the user intention recognition module, the biological entity module and the literature abstract knowledge graph module, for integrating the logical chain instructions, the entity background information and the related literature information to generate a query result corresponding to the query request of the user based on a large model and displaying the query result by the user interaction module.
[0006] In an embodiment of the present application, the user interaction module comprises: a user search interface for receiving a literature keyword or a literature related question input by a user; a user guidance and prompt interface for showing the main functions and use methods of the system to a user who uses the system for the first time; and an information display interface comprising a data statistics display area, a graph visualization display area and an intelligent dialogue area for displaying a literature query result for a literature keyword query or an answer result for a question feedback.
[0007] In an embodiment of the present application, the biological entity module comprises: a biological entity database construction unit for integrating multiple data sources to obtain multi-source biomedical entity information and constructing a biological entity database; and an entity information retrieval unit connected to the biological entity database construction unit for retrieving matching corresponding entity background information from the biological entity database based on the extracted biological entities.
[0008] In an embodiment of the present application, the literature abstract knowledge graph module comprises: a literature collection unit for dynamically updating literature and extracting core information of the literature; a literature analysis unit connected to the literature collection unit for deep analysis and systematic summary of the core information of each literature based on a large model to obtain multi-dimensional literature analysis results; a knowledge graph construction unit connected to the literature analysis unit for extracting biological entities, relationship information and evidence levels of each entity relationship from the core information based on the literature analysis results and constructing a literature abstract knowledge graph database; and a literature information retrieval unit connected to the knowledge graph construction unit for retrieving corresponding relevant literature information in the literature abstract knowledge graph database based on the extracted biological entities and generated logical chain instructions through semantic similarity based on a large model, and sorting based on relevance, literature citation volume and publication time.
[0009] In an embodiment of the present application, the literature collection unit comprises: an incremental crawler update subunit for crawling literature in an incremental crawler mode based on a set crawling frequency and concurrency; a literature grouping and storage subunit connected to the incremental crawler update subunit for classifying and storing the crawled literature according to fields; and a core information extraction subunit connected to the literature grouping and storage subunit for adaptive structure analysis of each classified literature to extract core paragraph information and charts; wherein the core paragraph information comprises: Title, HighLight, Abstract, Introduction, Results, legends and titles of Figure and Table, Discussion and Conclusion.
[0010] In an embodiment of the present application, the literature analysis unit comprises: a translation subunit for calling a large model for performing a translation task to perform Chinese translation of the extracted core information of the literature in context; and a structured abstract generation subunit connected to the translation subunit for processing the core information of each literature based on a frame parser to obtain corresponding structured abstracts and chart analysis results; wherein the structured abstracts comprise: research highlights, research abstracts, research problems, research methods, research results, research deficiencies and research conclusions.
[0011] In an embodiment of the present application, the knowledge graph construction unit comprises: an entity and relationship extraction subunit configured to extract biological entities and relationship information based on a large model performing a biological entity and relationship extraction task according to core information of each literature, and determine the evidence level of each entity relationship based on the research method in the structured abstract of each literature; and a knowledge graph generation subunit connected to the entity and relationship extraction subunit and configured to generate a literature abstract knowledge graph database based on the extracted biological entities, relationship information, and evidence level of each entity relationship, and the structured abstract of each literature.
[0012] In an embodiment of the present application, based on a framework parser, the core information of each literature is processed to obtain a corresponding structured abstract and chart parsing result, which includes: based on a large model performing a chart parsing task, the charts in the core information of the literature are parsed to obtain a chart parsing result; based on the core paragraph information in the core information of the literature and the chart parsing result, the structured abstract of the literature is obtained, which includes: taking the HighLight paragraph of the literature as the research highlight of the literature; based on a large model performing a literature abstract task, the Abstract paragraph, the last paragraph of Introduction, and the first paragraph of Discussion of the literature are used to obtain the research abstract of the literature; based on a large model performing a literature abstract task, all paragraphs of Introduction of the literature except the last paragraph are used to generate the research question of the literature; based on a large model performing a literature abstract task, the legends and titles of Figure of the literature are used to generate the research method of the literature; based on a large model performing a literature abstract task, the title of Results, the legends and titles of Figure, and the chart parsing result of the literature are used to generate the research result of the literature; based on a large model performing a literature abstract task, all paragraphs of Discussion of the literature except the first paragraph are used to generate the research deficiency of the literature; and the Conclusion paragraph of the literature is taken as the research conclusion of the literature.
[0013] To achieve the above object and other related objects, the present application provides a medical literature intelligent processing and analysis method based on a large model, which comprises: receiving a query request of a user through a user search interface; wherein the types of the query request include a literature keyword query request or a literature related question query request; based on a large model, performing user intent recognition according to the query request of the user, extracting biological entities, and generating a logical chain instruction; based on the extracted biological entities, retrieving matching corresponding entity background information from a pre-constructed biological entity library; based on a large model, retrieving matching corresponding related literature information from a pre-constructed literature abstract knowledge graph database according to the extracted biological entities and the generated logical chain instruction; and based on a large model, integrating the logical chain instruction, the entity background information, and the related literature information to generate and display a query result corresponding to the query request of the user.
[0014] To achieve the above and other related objectives, the present invention provides an electronic terminal, comprising: one or more memories and one or more processors; the one or more memories are used to store computer programs; the one or more processors are connected to the memories and are used to run the computer programs to execute the intelligent processing and analysis method for medical literature based on large models.
[0015] As described above, this invention is a medical literature intelligent processing and analysis system, method, and terminal based on a large model, which has the following beneficial effects: First, the user's query request is obtained through the user search interface; then, the large model is invoked to identify the user's intent, extract relevant biological entities, and generate corresponding logical chain instructions; based on the extracted biological entities, background information of relevant entities is matched in a pre-constructed biological entity database; and according to the logical chain instructions, matching is performed in a literature abstract knowledge graph database to obtain the relevance ranking of the literature; finally, the logical chain instructions, biological entity background information, and relevant literature information are integrated to generate and display the corresponding query results. This invention solves the problems faced by medical researchers, such as cumbersome literature acquisition, time-consuming reading, and complex analysis, by using a large model-based method for intelligent literature processing and analysis, accelerating the literature reading process, reducing manual analysis costs, and improving research efficiency. Attached Figure Description
[0016] Figure 1 The diagram shown is a structural schematic of a medical literature intelligent processing and analysis system based on a large model, according to an embodiment of the present invention.
[0017] Figure 2 The diagram shown is a structural schematic of a medical literature intelligent processing and analysis system based on a large model, according to an embodiment of the present invention.
[0018] Figure 3 The diagram shown is a structural schematic of a document summary knowledge graph module in one embodiment of the present invention.
[0019] Figure 4 The diagram shown illustrates the analysis of the core information structure of a document by a frame parser in one embodiment of the present invention.
[0020] Figure 5 The diagram shown is a flowchart of a method for intelligent processing and analysis of medical literature based on a large model, according to an embodiment of the present invention.
[0021] Figure 6 The diagram shown is a flowchart of a method for intelligent processing and analysis of medical literature based on a large model, according to an embodiment of the present invention.
[0022] Figure 7The diagram shown is a flowchart illustrating a method for constructing a document summary knowledge graph database according to an embodiment of the present invention.
[0023] Figure 8 The diagram shown is a structural schematic of an electronic terminal according to an embodiment of the present invention. Detailed Implementation
[0024] The following specific examples illustrate the implementation of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that, unless otherwise specified, the following embodiments and features described therein can be combined with each other.
[0025] It should be noted that in the following description, reference is made to the accompanying drawings, which illustrate several embodiments of the invention. It should be understood that other embodiments may also be used, and changes in mechanical composition, structure, electrical system, and operation may be made without departing from the spirit and scope of the invention. The following detailed description should not be considered limiting, and the scope of the embodiments of the invention is defined only by the claims of the published patents. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the invention. Spatially related terms, such as “upper,” “lower,” “left,” “right,” “below,” “below,” “lower part,” “above,” “upper part,” etc., may be used herein to illustrate the relationship between one element or feature shown in the figures and another element or feature.
[0026] Throughout this specification, when it is said that a part is "connected" to another part, this includes not only "direct connection" but also "indirect connection" by placing other elements in between. Furthermore, when it is said that a part "includes" a certain constituent element, unless otherwise stated otherwise, this does not exclude other constituent elements, but rather means that other constituent elements may also be included.
[0027] The terms "first," "second," and "third," etc., used herein are for the purpose of describing various parts, components, regions, layers, and / or segments, but are not limiting. These terms are used only to distinguish one part, component, region, layer, or segment from others. Therefore, the "first part," "component," "region," "layer," or "segment" described below may refer to a "second part," "component," "region," "layer," or "segment" without departing from the scope of this invention.
[0028] Furthermore, as used herein, the singular forms “a,” “an,” and “the” are intended to include the plural forms as well, unless the context indicates otherwise. It should be further understood that the terms “comprising,” “including,” indicate the presence of the stated feature, operation, element, component, item, kind, and / or group, but do not preclude the presence, occurrence, or addition of one or more other features, operations, elements, components, items, kinds, and / or groups. The terms “or” and “and / or” as used herein are interpreted as inclusive, or mean any one or any combination thereof. Thus, “A, B, or C” or “A, B, and / or C” means “any one of: A; B; C; A and B; A and C; B and C; A, B, and C.” Exceptions to this definition arise only when combinations of elements, functions, or operations are inherently mutually exclusive in some manner.
[0029] This invention provides a large-scale model-based intelligent processing and analysis system for medical literature. First, it obtains the user's query request through a user search interface. Then, it invokes the large-scale model to identify the user's intent, extract relevant biological entities, and generate corresponding logical chain instructions. Based on the extracted biological entities, it matches relevant entity background information in a pre-constructed biological entity database. Next, it matches the logical chain instructions against a literature abstract knowledge graph database to obtain a relevance ranking of the literature. Finally, it integrates the logical chain instructions, biological entity background information, and relevant literature information to generate and display the corresponding query results. This invention solves the problems faced by medical researchers, such as cumbersome literature acquisition, time-consuming reading, and complex analysis. By employing a large-scale model-based method for intelligent literature processing and analysis, it accelerates the literature reading process, reduces manual analysis costs, and improves research efficiency.
[0030] The present invention will now be described in detail with reference to the accompanying drawings, so that those skilled in the art can readily implement it. The present invention can be embodied in many different forms and is not limited to the embodiments described herein.
[0031] like Figure 1 This diagram illustrates the structure of a medical literature intelligent processing and analysis system based on a large model, as described in an embodiment of the present invention.
[0032] The system includes:
[0033] User interaction module 1 is used to receive user query requests through a user search interface. The types of query requests include: keyword query requests and document-related question queries. Specifically, users input specific keywords, and the system retrieves relevant literature from the database based on these keywords. This type of query is typically used to quickly find literature containing specific terms or topics. Users can also ask document-related questions, and the system needs to understand these questions and provide corresponding answers or literature recommendations. This type of query focuses more on resolving user doubts and may require more sophisticated natural language processing capabilities. While asking questions, users can also upload files, ask targeted questions, and receive and view the system's answers.
[0034] User intent recognition module 2, connected to user interaction module 1, is used to recognize user intent based on a large model and user query requests, extract biological entities, and generate logical chain instructions that guide subsequent modules on how to retrieve and process information.
[0035] The biological entity module 3 is connected to the user intent recognition module 2 and is used to retrieve matching entity background information from a pre-built biological entity library based on the extracted biological entities.
[0036] The literature abstract knowledge graph module 4, connected to the biological entity module 3, is used to retrieve relevant literature information matching the pre-constructed literature abstract knowledge graph database based on the generated logical chain instructions.
[0037] The query result acquisition module 5 is connected to the user intent recognition module 2, the biological entity module 3, and the literature abstract knowledge graph module 4. It is used to generate query results corresponding to the user's query request based on the large model, integrating the logical chain instructions, entity background information, and relevant literature information, and then display them by the user interaction module 1.
[0038] By combining natural language processing, large-scale model generation, retrieval enhancement, and deep learning technologies, this system provides users with an intelligent and systematic literature analysis solution. It should be able to automatically identify and extract key information from literature, such as research abstracts, background, results, and conclusions, while filtering out non-critical content. This significantly improves the efficiency and accuracy of literature analysis, enabling efficient processing and in-depth analysis of medical literature. Furthermore, the system possesses cross-document knowledge association and reasoning capabilities, helping researchers quickly grasp research frontiers and discover potential research opportunities. This intelligent literature analysis system will greatly enhance the work efficiency of researchers and accelerate the process of scientific discovery.
[0039] In one embodiment, the system has dynamic adjustment capabilities, enabling it to generate prompt words for large models adapted to different task types, and select the appropriate type of large model for parsing.
[0040] Specifically, it includes:
[0041] The large model that performs the translation task uses an open-source large language model, such as the GPT series, deployed on the server side, to translate Chinese into context based on prompts that adapt to the input.
[0042] Large models for performing literature summarization tasks use open-source large language models, such as BERT or BART, deployed on the server side to generate summaries based on prompts adapted to the model from the input.
[0043] The large model that performs the chart parsing task uses an open-source multimodal large language model, such as ChartLlama, deployed on the server side, to parse the chart based on prompts adapted to the model from the input.
[0044] This is a large model for extracting biological entities and relationships, adapted to the model based on input prompts. For example, it uses commercial large model interfaces, including those from OpenAI, Google, Baidu, Microsoft, Alibaba, Tencent, and others; the interfaces support automatic updates to ensure the large models are the latest versions from each company.
[0045] A large model for performing intent recognition tasks is used to perform graph recognition based on input prompts adapted to the model, extract biological entities, and generate logical chain instructions.
[0046] In one embodiment, such as Figure 2 The user interaction module 1 includes:
[0047] The user guidance and prompt interface is designed to teach users how to use the system, providing specific examples so that users can learn the main functions and usage methods of the system when they use it for the first time.
[0048] The user search interface receives keywords or related questions from users. Users can enter keywords, questions, or specific queries. These can be keywords related to biological entities such as drugs, diseases, proteins, and genes, or specific questions related to medical literature. The system provides auto-completion and keyword suggestions to help users quickly complete their queries. User queries support complex combinations using Boolean logic, such as "A AND B" or "A OR B". Further configuration can be achieved through the search settings module. Search settings include journal scope, time span, document type, and impact factor, allowing filtering of abstracts and knowledge graph databases in specific fields based on selection criteria.
[0049] The information display interface includes: a data statistics display area, a graph visualization display area, and an intelligent dialogue area, which are used to display literature search results for literature keywords or answers to questions.
[0050] The data statistics display area is used to display the list of literature results for the query, and supports sorting by time, citations and relevance; it also supports filtering by time, literature type and impact factor; the map visualization display area is used to support interactive operations of the map, including zooming, panning, local magnification, node dragging and shortest path analysis; the intelligent dialogue area uses large model retrieval enhancement technology to support multi-turn conversational understanding based on PDF documents.
[0051] In one embodiment, such as Figure 2 The user intent recognition module 2 includes:
[0052] This PDF text recognition unit can receive locally uploaded PDF documents, supporting batch and drag-and-drop uploads. It provides users with convenient file management functions, including file preview, management, and deduplication. Furthermore, utilizing the Alibaba OCR interface, it can efficiently recognize text content in PDF files, supporting multilingual recognition and layout analysis. It also features chart recognition, formula recognition and conversion, and document structure recognition capabilities, providing users with a one-stop PDF document processing service.
[0053] The user intent parsing unit is used to understand and parse the query request input by the user, extract the key entities and query intent of the query request; specifically, the system generates prompt words based on the query content input by the user, inputs them into a pre-trained large model based on the execution intent recognition task, performs intent analysis on the query content input by the user, extracts biological entities and identifies the core query target, and combines them according to a certain method to generate logical chain instructions, which are intended to guide the subsequent retrieval and information matching process.
[0054] In one embodiment, such as Figure 2The biological entity module includes:
[0055] The biological entity library construction unit is used to integrate biomedical entity information from multiple data sources to build a biological entity library. Specifically, it integrates biomedical entity information from multiple authoritative biological entity databases, including but not limited to gene databases (such as NCBI Gene, Ensembl), protein databases (such as UniProt), disease databases (such as OMIM), and drug databases (such as DrugBank, ChEMBL). These data sources cover various types of biological entities, such as genes, proteins, diseases, drugs, and molecular pathways. Through entity data integration, this unit acquires, cleans, and standardizes biological entity information from multi-source data, such as NCBI Gene, Ensembl, and UniProt, and constructs the biological entity library based on the processed entity background information.
[0056] The entity information retrieval unit, connected to the biological entity database construction module, is used to retrieve matching background information for extracted biological entities from the biological entity database. Specifically, this unit matches corresponding entities in a pre-built biological entity database using fuzzy search. The biological entity information database covers information on most diseases, drugs, proteins, and genes. After the retrieval is complete, detailed background information related to the entity is returned, such as: the indications and clinical trial data of a drug; the function, mutation data, and disease-related association of a gene; and the pathological characteristics and treatment plans of a disease. This unit provides an efficient and accurate biological entity search engine, supporting rapid retrieval and access to biological entity information in the database.
[0057] In one embodiment, such as Figure 2 The document abstract knowledge graph module 4 includes:
[0058] The literature acquisition unit is used to dynamically update literature and extract its core information. This unit efficiently and accurately updates academic literature dynamically and extracts the core paragraphs and figures from the literature.
[0059] The document analysis unit, connected to the document acquisition unit, is used to perform in-depth analysis and systematic summarization of the core information of each document based on a large model, and to obtain multi-dimensional document analysis results.
[0060] The knowledge graph construction unit, connected to the document parsing unit, is used to extract biological entities, relational information, and the evidence level of each entity's relation from the core information based on the large model and the document parsing results, and to construct a document summary knowledge graph database.
[0061] The literature information retrieval unit, connected to the knowledge graph construction unit, is used to retrieve and match relevant literature information in the literature abstract knowledge graph database based on the large model, according to the extracted biological entities and the generated logical chain instructions, and can sort the literature based on relevance, citation count, and publication time.
[0062] In one specific embodiment, such as Figure 3 The document acquisition unit includes:
[0063] The incremental crawler update subunit is used to crawl documents incrementally based on the set crawling frequency and concurrency. Specifically, the incremental crawler update subunit supports integrating API interfaces of various public document sources to obtain document metadata, including but not limited to databases such as PubMed and PubTator3. Document metadata includes title, author, keywords, abstract, DOI number, PMID number, PMCID number, and URL. By maintaining document index records and deduplicating collected documents, it can periodically or in real-time monitor updates to the target database and automatically identify new documents. Furthermore, it dynamically adjusts the document crawling frequency and concurrency based on factors such as server load and network bandwidth to ensure efficient data collection using the incremental crawler method under legal conditions.
[0064] The document grouping and storage subunit, connected to the incremental crawler update subunit, is used to classify and store the crawled documents according to their domains.
[0065] The core information extraction subunit connects the document grouping and storage subunits and is used to perform adaptive structural analysis on the classified documents, process different journal page structures, and extract core paragraph information and charts.
[0066] The core paragraph information includes: Title, Highlight, Abstract, Introduction, Results, Figure and Table legends and titles, Discussion, and Conclusion. In addition to text information, the core information extraction subunit can also extract chart information, providing screenshots of the Figures and Tables.
[0067] In one embodiment, such as Figure 3 The document parsing unit includes:
[0068] The translation subunit is used to retrieve the large model to perform translation tasks and translate the core information of the extracted documents into Chinese within the context. If the document contains English content, the large model will automatically translate it into Chinese so that users can read and understand it.
[0069] A structured abstract generation subunit, connected to the translation subunit, is used to process the core information of each document based on a frame parser, obtaining corresponding structured abstracts and figure / table analysis results. The structured abstract includes: research highlights, research summary, research questions, research methods, research results, research limitations, and research conclusions. A structured abstract is an information organization method that presents the key content of a document in a predefined format, allowing users to quickly grasp the main points. It helps users quickly identify the value and relevance of documents, especially when conducting literature reviews or rapidly screening a large number of documents.
[0070] In one embodiment, based on a frame parser, the core information of each document is processed to obtain the corresponding structured abstract and chart parsing results, including:
[0071] Based on a large model for performing chart parsing tasks, the charts in the core information of the literature are parsed to obtain chart parsing results; specifically, screenshots of Figure and Table are used as prompt words and input into the multimodal large model to generate interpretations of the charts and obtain chart parsing results.
[0072] Structured summaries of documents are obtained based on core paragraph information and the results of figure and table analysis. The methods include:
[0073] The entire highlighted section of the literature is considered a research highlight.
[0074] Based on the large model of performing literature abstracting tasks, a research abstract of a literature is obtained by taking the entire Abstract paragraph, the last paragraph of the Introduction, and the first paragraph of the Discussion.
[0075] Based on a large model for performing document summarization tasks, the research questions of a document are generated from all paragraphs of the document's introduction except for the last paragraph.
[0076] A research method for generating literature based on the legends and titles of figures in a document, using a large model for performing document abstracting tasks.
[0077] Based on a large model for performing literature abstracting tasks, the research results of the literature are generated according to the titles of the Results, the legends and titles of the Figures, and the analysis results of the figures.
[0078] There is a lack of research on generating literature from all paragraphs except the first paragraph of the Discussion section based on a large model for performing literature summarization tasks.
[0079] The entire Conclusion section of the literature is taken as the research conclusion of the literature.
[0080] For example, such as Figure 4 The frame parser performs structural analysis on the input HTML document, identifies different parts of the document, extracts its content, and tags and classifies it. The frame parser identifies the specific structure of the document by performing a large-scale model of a document summarization task. The processing methods based on the core information of each document include:
[0081] The frame parser directly marks the entire Highlighted section of the literature as a "research highlight", which is very important for quickly understanding the core contributions of the research.
[0082] The frame parser combines the entire Abstract, the last paragraph of the Introduction, and the first paragraph of the Discussion section of the document, generates prompt words, inputs them into a large model performing the document summarization task to generate a summary, and uses this summary as the document's research abstract; the document abstract is a concise summary of the entire study. The frame parser extracts this part from the document's HTML and marks it as the "research abstract." This part typically includes the research background, objectives, methods, results, and conclusions, and is key for readers to quickly understand the document's content.
[0083] The frame parser combines all paragraphs of the introduction (excluding the last paragraph) to generate prompts, inputs them into a large model performing a document summarization task to generate a summary, and uses this summary as the research question of the document. In the introduction section, the document typically discusses the research background, existing problems, and research motivation. The frame parser extracts the introduction section and marks it as the "research question," which is the starting point of the document and helps understand why the authors conducted this research.
[0084] The frame parser combines the legends and titles of the figures in the literature to generate prompts, which are then input into a large model for performing a literature abstracting task to generate an abstract, and this abstract is used as the research method of the literature. The legends and titles of the figures in the literature describe in detail the methods and experimental procedures used in the research. The frame parser extracts relevant content from the figure legends and titles and marks it as "research method." This part is crucial for understanding how the research was conducted.
[0085] The frame parser combines the titles of the Results section, the legends and titles of the Figures, and the analysis results of the figures and tables to generate prompts. These prompts are then input into a large model performing a literature abstracting task to generate an abstract, which is then presented as the literature's research results. The titles of the Results section, the legends and titles of the Figures, and the analysis results of the figures and tables are typically used to display the literature's research data and statistical results. The frame parser extracts this content from the HTML document and marks it as "Research Results" for subsequent analysis and presentation.
[0086] The frame parser generates prompt words based on all paragraphs of the document's Discussion section except the first paragraph. These prompt words are then input into a large model that performs the document summarization task to generate a summary, which is used as the basis for identifying the document's shortcomings. The Discussion section of a document typically mentions the limitations and shortcomings of the research. The frame parser extracts these points from the Discussion section and marks them as "shortcomings in the research." This represents the researcher's analysis of the limitations of the experiment and directions for future improvement.
[0087] The frame parser treats the entire Conclusion section of a document as its research conclusion. The Conclusion section summarizes the research findings and provides the final conclusion. The frame parser extracts key content from the conclusion section of the document and marks it as "Research Conclusion" to summarize the final findings of the research.
[0088] In one embodiment, such as Figure 3 The knowledge graph construction unit includes:
[0089] The entity and relation extraction subunit is used to extract biological entity and relation information based on the core information of each literature, using a large model that performs the task of extracting biological entities and relations. It then determines the level of evidence for each entity relation based on the research methods in the structured abstracts of each literature. This step requires evaluating the reliability and effectiveness of the research methods to determine the strength of evidence for the entity relations, taking into account factors such as the rigor of the research design, sample size, and experimental reproducibility. Specifically, it generates cue words based on the core information of each literature, uses the large model to identify biomedical entities, including but not limited to diseases, phenotypes, drugs, proteins, genes, and side effects; it uses the large model that performs the task of extracting biological entities and relations to identify semantic relationships between the extracted entities, such as "treatment," "cause," "inhibition," and "interaction"; and it uses the research methods in the analysis results to provide the level of evidence for each entity relation.
[0090] The knowledge graph generation subunit, connected to the entity and relation extraction subunit, is used to generate a literature summary knowledge graph database based on the extracted biological entities, relation information, evidence levels of each entity's relations, and structured summaries of each document. Specifically, the identified entity and relation information and evidence levels are converted into JSON format, stored in the Neo4j database, and visualized as a knowledge graph. The knowledge graph will be stored as an independent dataset, forming a visualized knowledge base. Through continuous updates and expansion, the knowledge graph will cover more biomedical entities and their complex relationships, providing users with rich reference materials. Based on the structured summaries of each document, they are stored in a domain-specific literature summary database.
[0091] In one embodiment, the query result acquisition module is used to integrate the logical chain instructions, entity background information, and relevant literature information to generate prompt words, input them into the large model to generate query results corresponding to the user's query request, and the generated query results may be text, charts, or other forms of data. These results directly correspond to the user's query request and are finally displayed through the user interaction module.
[0092] Similar to the principles of the above embodiments, the present invention provides a medical literature intelligent processing and analysis system based on a large model.
[0093] The following specific embodiments are provided in conjunction with the accompanying drawings:
[0094] like Figure 5 This is a flowchart illustrating an intelligent processing and analysis method for medical literature based on a large model, as shown in an embodiment of the present invention.
[0095] The method includes:
[0096] Step S1: Receive user query requests through the user search interface; wherein, the types of query requests include: document keyword query requests or document-related question query requests;
[0097] Step S2: Based on the large model, identify user intent according to the user's query request, extract biological entities, and generate logical chain instructions;
[0098] Step S3: Based on the extracted biological entities, retrieve the corresponding entity background information from the pre-constructed biological entity database;
[0099] Step S4: Based on the large model, according to the extracted biological entities and the generated logical chain instructions, retrieve the corresponding relevant literature information from the pre-constructed literature summary knowledge graph database;
[0100] Step S5: Based on the large model, integrate the logical chain instructions, entity background information and relevant literature information to generate and display the query results corresponding to the user's query request.
[0101] Since the implementation principle of this intelligent processing and analysis method for medical literature based on a large model has been described in the foregoing embodiments, it will not be repeated here.
[0102] To better describe the intelligent processing and analysis method for medical literature based on a large model, the following specific embodiments are provided for illustration.
[0103] Example 1: A method for intelligent processing and analysis of medical literature based on a large model. For example... Figure 6 This is a flowchart illustrating the intelligent processing and analysis method for medical literature based on a large model in this embodiment.
[0104] The method includes:
[0105] Step 1: First, obtain the user's query request through the user search interface. Users can enter keywords, questions, or specific query content. This can be keywords related to biological entities such as drugs, diseases, proteins, and genes, or specific questions related to medical literature.
[0106] Step 2: Invoke a pre-trained large language model to perform intent analysis on the user's query input. The model will parse the user's natural language input, identify the core query target, and generate logical chain instructions according to a certain method, aiming to guide the subsequent retrieval and information matching process.
[0107] Step 3: Based on the biological entities extracted in Step 2, match the corresponding entities in the pre-constructed biological entity database using fuzzy search. The biological entity information database covers most diseases, drugs, proteins, and gene biological entity information. After the search is completed, detailed background information related to the entity will be returned, such as: the indications and clinical trial data of a certain drug; the function, mutation data, and correlation with diseases of a certain gene; the pathological characteristics and treatment plans of a certain disease, etc.
[0108] Step 4: Based on the entities and logical chain instructions extracted in Step 3, match them in a domain literature summary knowledge graph database using semantic similarity matching. This knowledge graph database is constructed based on semantic analysis and knowledge extraction from a large number of medical documents. Through semantic association, return the structured summary of relevant documents and corresponding knowledge graph information. Based on relevance, citation count, publication time, and other dimensions, the system will sort the query results, prioritizing the display of the top 10 high-quality, highly relevant documents.
[0109] Step 5: Integrate logical chain instructions, biological entity background information, literature abstracts, and knowledge graphs to generate further prompts, input them into the large model, output query results, and display the results to the user.
[0110] Example 2: A method for constructing a document abstract knowledge graph database. For example... Figure 7 This is a flowchart illustrating the method for constructing the document abstract knowledge graph database.
[0111] The method includes:
[0112] Step 1: Download relevant literature metadata from different databases or literature resource platforms (such as Pubtator3), and use an incremental crawler script to keep up with publicly available literature databases in real time. The literature metadata includes title, author, keywords, abstract, DOI number, PMID number, PMCID number, and URL.
[0113] Step 2: Filter literature in specific biomedical fields through the large language model pipeline based on basic setting prompts. Based on user or system needs, the model automatically determines the specific field described in the literature through semantic understanding analysis, such as the field of oncology, the field of autoimmune diseases, etc.
[0114] Step 3: For each article in the selected professional field, the system will access its corresponding URL to obtain the HTML source code of the article page. The acquisition strategy involves establishing API partnerships with major publishers and using a headless browser to automatically simulate clicking on the article page to obtain its HTML source code.
[0115] The obtained HTML source code is used to extract key paragraph and chart information based on the XPath parsing template, depending on the publisher. This includes: Title, Highlight, Abstract, Introduction, Results titles, Figure, Table screenshots, legends and titles, Discussion, and Conclusion.
[0116] Step 4: Input the extracted information into the large model pipeline for in-depth analysis, generating a series of outputs: If the literature contains English content, the large model will automatically translate it into Chinese so that users can read and understand it. The large model will automatically generate concise literature summaries based on different paragraphs, condensing the research content and allowing users to quickly understand the main findings and contributions of the literature. For figures and tables in the literature, the large model can generate interpretations of the figures and tables by calling the multimodal analysis in the large model pipeline, helping users understand the data and research methods in the figures and tables.
[0117] Step 5: The system identifies biomedical entities in the literature through a large model interface. These entities may include diseases, drugs, genes, proteins, molecular pathways, etc. The large model automatically extracts these key terms through semantic understanding. The system not only extracts individual entities but also identifies the relationships between them. For example, how a certain drug treats a certain disease, or how a gene mutation affects the occurrence of a certain disease. These relationships are the core of knowledge graph construction. Based on the extracted entities and relationships, the system begins to construct the knowledge graph. The knowledge graph displays the connections between entities in the form of nodes and edges and can be visualized. Through the knowledge graph, users can intuitively understand the research content in the literature and its relevance.
[0118] Step 6: After the system generates document translations, abstracts, and graph analyses, this information will be stored in a domain-specific document abstract database. The database will remove the original text data of each document, containing only the secondary abstract version to prevent the leakage of non-public document data. The knowledge graph will be stored as an independent dataset, forming a visualized knowledge base. Through continuous updates and expansion, the knowledge graph will cover more biomedical entities and the complex relationships between them, providing users with rich reference materials. Document abstracts and the knowledge graph will be further categorized by domain, ensuring that users can quickly find relevant research results in specific fields.
[0119] The intelligent medical literature processing and analysis system based on a large model provided in this invention can be implemented on the terminal side or the server side. Regarding the hardware structure of the electronic terminal, please refer to... Figure 8 This is a schematic diagram of an optional hardware structure of an electronic terminal 1000 provided in an embodiment of the present invention. The terminal 1000 can be a mobile phone, computer device, tablet device, etc. The terminal 1000 includes: at least one processor 1001, a memory 1002, at least one network interface 10010, and a user interface 1009. The various components in the device are coupled together through a bus system 1005. It is understood that the bus system 1005 is used to realize the connection and communication between these components. In addition to a data bus, the bus system 1005 also includes a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 8 The general will label all buses as bus systems.
[0120] The user interface 1009 may include a display, keyboard, mouse, buttons, keypad, or touchscreen.
[0121] It is understood that memory 1002 can be volatile memory or non-volatile memory, or both. Non-volatile memory can be read-only memory (ROM) or programmable read-only memory (PROM), which serves as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM) and synchronous static random access memory (SSRAM). The memories described in the embodiments of this invention are intended to include, but are not limited to, these and any other suitable categories of memory.
[0122] In this embodiment of the invention, the memory 1002 is used to store various categories of data to support the operation of the terminal 1000. Examples of this data include: any executable program for operation on the terminal 1000, such as the operating system 10021 and application program 10022; the operating system 10021 contains various system programs, such as the framework layer, core library layer, driver layer, etc., for implementing various basic services and handling hardware-based tasks. The application program 10022 may contain various applications, such as media players, browsers, etc., for implementing various application services. The intelligent processing and analysis system for medical literature based on a large model provided in this embodiment of the invention can be included in the application program 10022.
[0123] The methods disclosed in the above embodiments of the present invention can be applied to or implemented by the processor 1001. The processor 1001 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 1001 or by instructions in the form of software. The processor 1001 may be a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 1001 can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor 1001 may be a microprocessor or any conventional processor, etc. The steps of the accessory optimization method provided in the embodiments of the present invention can be directly reflected as being executed by a hardware decoding processor, or being executed by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, which is located in a memory. The processor reads the information in the memory and combines it with its hardware to complete the steps of the aforementioned method.
[0124] In an exemplary embodiment, the terminal 1000 may be used to execute the aforementioned method by one or more application-specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), or complex programmable logic devices (CPLDs).
[0125] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented using computer program-related hardware. The aforementioned computer program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0126] In the embodiments provided in this application, the computer-readable and writable storage medium may include read-only memory, random access memory, EEPROM, CD-ROM or other optical disc storage devices, disk storage devices or other magnetic storage devices, flash memory, USB flash drive, portable hard drive, or any other medium capable of storing desired program code in the form of instructions or data structures and accessible by a computer. Additionally, any connection may be appropriately referred to as a computer-readable medium. For example, if instructions are transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of the medium. However, it should be understood that computer-readable and writable storage media and data storage media do not include connections, carrier waves, signals, or other transient media, but are intended for non-transient, tangible storage media. The disks and optical discs used in the application include compact discs (CDs), laser discs, optical discs, digital multifunction discs (DVDs), floppy disks, and Blu-ray discs, where disks typically copy data magnetically, while optical discs use lasers to copy data optically.
[0127] In summary, the intelligent medical literature processing and analysis system, method, and terminal based on a large model of the present invention first obtains the user's query request through a user search interface, then calls the large model to identify the user's intent, extract relevant biological entities, and generate corresponding logical chain instructions; based on the extracted biological entities, it matches the background information of relevant entities in a pre-constructed biological entity database; and according to the logical chain instructions, it matches the literature abstract knowledge graph database to obtain the relevance ranking of the literature; finally, it integrates the logical chain instructions, biological entity background information, and relevant literature information to generate and display the corresponding query results. This invention solves the problems faced by medical researchers, such as cumbersome literature acquisition, time-consuming reading, and complex analysis. By adopting a large model-based method for intelligent literature processing and analysis, it accelerates the literature reading process, reduces manual analysis costs, and improves research efficiency. Therefore, this invention effectively overcomes the various shortcomings of existing technologies and has high industrial application value.
[0128] The above embodiments are merely illustrative of the principles and effects of the present invention and are not intended to limit the invention. Any person skilled in the art can modify or alter the above embodiments without departing from the spirit and scope of the present invention. Therefore, all equivalent modifications or alterations made by those skilled in the art without departing from the spirit and technical concept disclosed in the present invention should still be covered by the claims of the present invention.
Claims
1. A medical literature intelligent processing and analysis system based on a large model, characterized in that, The system includes: The user interaction module is used to receive user query requests through the user search interface; wherein, the types of query requests include: document keyword query requests or document-related question query requests; The user interaction module includes: The user search interface is used to receive keywords or questions related to literature that users want to search for. The user guidance and prompt interface is used to demonstrate the system's main functions and usage methods to first-time users; The information display interface includes: a data statistics display area, a graph visualization display area, and an intelligent dialogue area, which are used to display literature search results for literature keyword searches or answer results for question feedback. The user intent recognition module, connected to the user interaction module, is used to recognize user intent based on a large model and the user's query request, extract biological entities, and generate logical chain instructions. The logical chain instructions are used to guide subsequent modules to retrieve and process information. The biological entity module, connected to the user intent recognition module, is used to retrieve matching entity background information from a pre-built biological entity library based on the extracted biological entities. The literature summary knowledge graph module, connected to the biological entity module, is used to retrieve relevant literature information matching the pre-constructed literature summary knowledge graph database based on the large model, according to the extracted biological entities and the generated logical chain instructions; The document abstract knowledge graph module includes: a document acquisition unit, a document parsing unit, a knowledge graph construction unit, and a document information retrieval unit; wherein... The document parsing unit includes: The translation subunit is used to retrieve the large model to perform the translation task and translate the core information of the extracted document into Chinese within the context. A structured summary generation subunit, connected to the translation subunit, is used to process the core information of each document based on the frame parser to obtain the corresponding structured summary and chart parsing results; wherein, the structured summary includes: research highlights, research abstract, research questions, research methods, research results, research limitations, and research conclusions; The query result acquisition module connects the user interaction module, user intent recognition module, biological entity module, and literature abstract knowledge graph module. It is used to generate query results corresponding to the user's query request based on the large model, integrating the logical chain instructions, entity background information, and relevant literature information. The generated query results are displayed by the user interaction module. The generated query results are text, charts, or other forms of data that directly correspond to the user's query request.
2. The intelligent medical literature processing and analysis system based on a large model as described in claim 1, characterized in that, The biological entity module includes: The biological entity library construction unit is used to integrate multiple data sources to obtain multi-source biomedical entity information and construct a biological entity library. The entity information retrieval unit, connected to the biological entity database construction unit, is used to retrieve matching entity background information from the biological entity database based on the extracted biological entities.
3. The intelligent medical literature processing and analysis system based on a large model as described in claim 1, characterized in that, The document abstract knowledge graph module includes: The document acquisition unit is used to dynamically update documents and extract their core information. The document analysis unit, connected to the document acquisition unit, is used to perform in-depth analysis and systematic summarization of the core information of each document based on a large model, and to obtain multi-dimensional document analysis results. The knowledge graph construction unit, connected to the document parsing unit, is used to extract biological entities, relational information, and the evidence level of each entity's relation from the core information based on the large model and the document parsing results, and to construct a document summary knowledge graph database. The literature information retrieval unit, connected to the knowledge graph construction unit, is used to search and match relevant literature information in the literature abstract knowledge graph database based on the large model, according to the extracted biological entities and the generated logical chain instructions, and sort them based on relevance, citation volume, and publication time.
4. The intelligent medical literature processing and analysis system based on a large model as described in claim 3, characterized in that, The document acquisition unit includes: The incremental crawler update subunit is used to crawl documents using an incremental crawler method based on the set crawling frequency and concurrency. The document grouping and storage subunit, connected to the incremental crawler update subunit, is used to classify and store the crawled documents according to their domains. The core information extraction subunit connects to the document grouping and storage subunit, and is used to perform adaptive structural analysis on each classified document to extract core paragraph information and figures; wherein, the core paragraph information includes: Title, Highlight, Abstract, Introduction, Results, legends and titles of Figures and Tables, Discussion, and Conclusion.
5. The intelligent medical literature processing and analysis system based on a large model as described in claim 1, characterized in that, The knowledge graph construction unit includes: The entity and relation extraction subunit is used to extract biological entity and relation information based on the core information of each document, and to determine the evidence level of each entity relation based on the research methods in the structured abstracts of each document, based on the large model that performs the biological entity and relation extraction task. The knowledge graph generation subunit connects the entity and relation extraction subunits and is used to generate a document summary knowledge graph database based on the extracted biological entities, relation information, evidence levels of each entity's relations, and structured summaries of each document.
6. The intelligent medical literature processing and analysis system based on a large model as described in claim 1, characterized in that, Based on the frame parser, the core information of each document is processed to obtain the corresponding structured abstracts and figure / table parsing results, including: Based on a large model for performing chart parsing tasks, the charts in the core information of the document are parsed to obtain the chart parsing results; Structured summaries of documents are obtained based on core paragraph information and the results of figure and table analysis. The methods include: The entire highlighted section of the literature is considered a research highlight. Based on the large model of performing literature abstracting tasks, a research abstract of a literature is obtained by taking the entire Abstract paragraph, the last paragraph of the Introduction, and the first paragraph of the Discussion. Based on a large model for performing document summarization tasks, the research questions of a document are generated from all paragraphs of the document's introduction except for the last paragraph. A research method for generating literature based on the legends and titles of figures in a document, using a large model for performing document abstracting tasks. Based on a large model for performing literature abstracting tasks, the research results of the literature are generated according to the titles of the Results, the legends and titles of the Figures, and the analysis results of the figures. There is a lack of research on generating literature from all paragraphs except the first paragraph of the Discussion section based on a large model for performing literature summarization tasks. The entire Conclusion section of the literature is taken as the research conclusion of the literature.
7. A method for intelligent processing and analysis of medical literature based on a large model, characterized in that, Applied to claim 1, the method comprises: The system receives user query requests through a user search interface; the types of query requests include: document keyword query requests or document-related question query requests. Based on the large model, user intent is identified according to the user's query request, and biological entities are extracted and logical chain instructions are generated. Based on the extracted biological entities, the corresponding entity background information is retrieved from a pre-constructed biological entity database; Based on the large model, relevant literature information matching the extracted biological entities and generated logical chain instructions is retrieved from a pre-constructed literature summary knowledge graph database. Based on the large model, the logical chain instructions, entity background information, and relevant literature information are integrated to generate and display the query results corresponding to the user's query request.
8. An electronic terminal, characterized in that, include: One or more memories and one or more processors; The one or more memories are used to store computer programs; The one or more processors are connected to the memory and are used to run the computer program to perform the method as described in claim 7.
Citation Information
Patent Citations
Medical literature retrieval method and device, electronic equipment and storage medium
CN112885478A