Literature extraction method, system and equipment based on large language model and storage medium

By constructing a document extraction template and using a large language model for automated analysis, the accuracy and efficiency of document extraction are solved, standardized and standardized document information extraction is realized, adapted to multiple document formats, and system compatibility and extraction efficiency are improved.

CN120409450AActive Publication Date: 2025-08-01PEKING UNIV

Patent Information

Application Number
CN202510905584.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-01
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

The accuracy and efficiency of literature extraction in the prior art are poor, especially in the systematic review, the probability of artificial extraction of missing questions or wrong questions is high, and there is a lack of standardized and standardized information extraction schemes.

Method used

The literature extraction method based on the large language model is adopted, and the literature extraction template is constructed, the extraction rules and format requirements are clarified, and the literature is read and analyzed using the large language model, and key information such as basic research information, research design types, participant characteristics, intervention measures, etc. are automatically extracted, and converted into structured extraction results.

Benefits of technology

It improves the accuracy and efficiency of document extraction, reduces manual errors, realizes a standardized and standardized information extraction process, adapts to multiple document formats, and improves the compatibility and versatility of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409450A_ABST
    Figure CN120409450A_ABST
Patent Text Reader

Abstract

The invention provides a literature extraction method, system and device based on a large language model and a storage medium, and relates to the technical field of medical research, and the method comprises the steps: constructing a literature extraction template according to a literature extraction demand; wherein the literature extraction requirements comprise extraction rules, format requirements and pre-extraction fields, and the pre-extraction fields comprise research basic information, research design types, research groups, participant characteristics, intervention measures of each research group and background medication of each research group; obtaining a to-be-extracted literature, and converting the to-be-extracted literature into a processable text format to form a first text; and inputting the first text and the literature extraction template into a large language model, performing reading analysis on the first text by applying the large language model, and extracting the first text according to the literature extraction template to obtain a literature extraction result corresponding to the to-be-extracted literature. According to the scheme, the accuracy and efficiency of document extraction are improved, and the requirements of system review can be better met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical research technology, and in particular to a document extraction method, system, device and storage medium based on a large language model. Background Art

[0002] A literature review summarizes, analyzes, and organizes existing research findings in a specific field or academic question. It's a crucial tool for researchers to fully grasp the current state of research and formulate research questions. Almost all research is based on a literature review, helping researchers understand the current state of their field, avoid duplication of effort, and provide a theoretical foundation for subsequent research.

[0003] In addition to general literature reviews, systematic reviews, as a more rigorous and standardized form of review, aim to comprehensively and systematically collect, evaluate, and synthesize all relevant studies on a specific research topic to draw comprehensive conclusions. They are widely used in fields such as medicine, social sciences, and education, and are a key step in evidence-based research and decision-making. In practice, systematic reviews require the most extensive and comprehensive database searches possible, often requiring the reading and extraction of a large number of articles. This undoubtedly increases the researcher's workload and consumes considerable time and labor.

[0004] In existing technologies, manual extraction is prone to omissions and errors. Currently, there is no better solution besides two-person extraction, which increases time costs. Manual extraction is primarily based on the extractor's own understanding of the literature research, which is highly proactive. The format of the extracted information is also based on the extractor's own understanding, making it difficult to standardize. Furthermore, systematic reviews require standardized and regularized extracted information. Currently, most literature review technologies are designed for general review writing and lack standardization for information extraction, making them inadequate for the needs of systematic reviews.

[0005] Therefore, the accuracy and efficiency of literature extraction in existing technical solutions are poor. Summary of the Invention

[0006] The present invention provides a document extraction method, system, device and storage medium based on a large language model, which are used to solve the defects of poor accuracy and efficiency of document extraction in the prior art, and realize the improvement of the accuracy and efficiency of document extraction based on the large language model.

[0007] The present invention provides a document extraction method based on a large language model, comprising the following steps: Construct a document extraction template based on document extraction requirements; wherein the document extraction requirements include extraction rules, format requirements, and pre-extraction fields, and the pre-extraction fields include: basic study information, study design type, study group, participant characteristics, intervention measures for each study group, and background medication for each study group; Obtain the document to be extracted, and convert the document to be extracted into a processable text format to form a first text; Input the first text and the document extraction template into a large language model, apply the large language model to read and analyze the first text, and extract the first text according to the document extraction template to obtain the document extraction result corresponding to the document to be extracted.

[0008] According to a document extraction method based on a large language model provided by the present invention, the document extraction result is an extraction table; the applying the large language model to read and analyze the first text, and extract the first text according to the document extraction template to obtain the document extraction result corresponding to the document to be extracted includes: Divide the first text according to a preset structured rule to obtain a plurality of text segments to form a second text; Apply the large language model to read and analyze the second text, and extract the information corresponding to the pre-extracted field according to the extraction rule and the format requirement to form a third text; Convert the third text into a table form to obtain the extraction table corresponding to the document to be extracted.

[0009] According to a document extraction method based on a large language model provided by the present invention, the applying the large language model to read and analyze the second text, and extract the information corresponding to the pre-extracted field according to the extraction rule and the format requirement to form a third text includes: Receive the second text and deeply analyze the text information corresponding to each subtitle; Apply the large language model to extract the information corresponding to the pre-extracted field one by one from the text information corresponding to each subtitle according to an extraction table based on the PICO format to form the third text.

[0010] According to a document extraction method based on a large language model provided by the present invention, the applying the large language model to extract the information corresponding to the pre-extracted field one by one from the text information corresponding to each subtitle according to an extraction table based on the PICO format to form the third text includes: Extract the basic research information and research design type from the text information corresponding to the title and abstract subtitle; wherein, the basic research information includes: the first author of the research, the publication year, and the country of the participants; the research design type includes at least one of the following: randomized controlled trial, cohort study, case-control experiment; Extract the research grouping from the text information corresponding to the abstract, method, and baseline information table subtitle; Extract participant characteristics from the text information corresponding to the title, the abstract, and the subtitle of the method; Extract the intervention measures for each study group from the text information corresponding to the abstract, the method, and the subtitle of the baseline information table; the intervention measures for each study group include: drug trade name, manufacturer, dose and dosing frequency, dosing time, and dosing method; Extract the background medications for each study group from the text information corresponding to the subtitle of the method.

[0011] According to a literature extraction method based on a large language model provided by the present invention, the research design type is a randomized controlled trial; the method further includes: Extract the experimental registration platform and registration number from the text information corresponding to the subtitle of the method.

[0012] According to a literature extraction method based on a large language model provided by the present invention, after obtaining the literature extraction result corresponding to the literature to be extracted, the method further includes: Perform normalization processing on the literature extraction result, and the normalization processing includes at least one of the following: deleting useless descriptive information, unifying the information format, and adding marks to information that cannot be extracted; Store the literature extraction result after normalization processing.

[0013] According to a literature extraction method based on a large language model provided by the present invention, there are multiple literatures to be extracted; after obtaining the literature extraction results corresponding to the literatures to be extracted, the method further includes: Determine whether all the literatures to be extracted have been extracted; if all the literatures to be extracted have been extracted, integrate and output the literature extraction results corresponding to all the literatures to be extracted; otherwise, return to execute the step of obtaining the literatures to be extracted.

[0014] The present invention also provides a literature extraction system based on a large language model, including the following modules: A construction module for constructing a literature extraction template according to the literature extraction requirements; wherein, the literature extraction requirements include extraction rules, format requirements, and pre-extracted fields, and the pre-extracted fields include: basic research information, research design type, research groups, participant characteristics, intervention measures for each research group, and background medications for each research group; An acquisition module for acquiring the literatures to be extracted and converting the literatures to be extracted into a processable text format to form a first text; An extraction module for inputting the first text and the document extraction template into a large language model, applying the large language model to read and analyze the first text, and extracting the first text according to the document extraction template to obtain a document extraction result corresponding to the document to be extracted.

[0015] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and running on the processor. When the processor executes the computer program, the method for extracting documents based on a large language model as described in any one of the above is implemented.

[0016] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for extracting documents based on a large language model as described in any one of the above is implemented.

[0017] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the method for extracting documents based on a large language model as described in any one of the above is implemented.

[0018] The method, system, device, and storage medium for extracting documents based on a large language model provided by the present invention, by constructing a document extraction template, clarify the specific requirements for document extraction, including extraction rules, format requirements, and pre-extracted fields, etc. This makes the entire extraction process have a clear goal and direction, and the construction of the template standardizes and regularizes the extraction process. The large language model can directly operate according to the template without redefining extraction rules and formats every time extraction is performed, thus greatly improving the extraction efficiency. Further, by converting the document to be extracted into a processable text format to form the first text, it is convenient for subsequent processing, improving the compatibility and versatility of the system to be able to process various document formats. Further, inputting the first text and the document extraction template into the large language model, applying the large language model to read and analyze the first text, and extracting the first text according to the document extraction template to obtain a document extraction result corresponding to the document to be extracted, reduces the cumbersome process of manual extraction, improves the extraction efficiency, and moreover, the large language model can understand complex language structures and context relationships and can extract information according to the document extraction template, improving the accuracy and consistency of the extracted information. Therefore, the solution of the present application improves the accuracy and efficiency of document extraction. Description of the Drawings

[0019] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0020] Figure 1 It is one of the schematic flowcharts of the literature extraction method based on the large language model provided by the present invention.

[0021] Figure 2 It is another schematic flowchart of the literature extraction method based on the large language model provided by the present invention.

[0022] Figure 3 It is the schematic structural diagram of the literature extraction system based on the large language model provided by the present invention.

[0023] Figure 4 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed implementation manners

[0024] To make the objectives, technical solutions and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some embodiments of the present invention, rather than all embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0025] It should be noted that the brief description of the terms in this application is only for facilitating the understanding of the following described embodiments, rather than intending to limit the embodiments of this application. Unless otherwise specified, these terms should be understood in their ordinary and general meanings.

[0026] The terms "first", "second", etc. in the description, claims and the above drawings of this application are used to distinguish similar or like objects or entities, and do not necessarily mean to limit a specific order or sequence, unless otherwise indicated (Unless otherwise indicated). It should be understood that such terms can be interchanged under appropriate circumstances, for example, it is possible to implement in an order other than those given in the illustration or description of the embodiments of this application.

[0027] In addition, the terms "comprising" and "having" and any variations thereof are intended to cover inclusion but not exclusivity. For example, a product or device comprising a series of components need not be limited to those components clearly listed, but may include other components not clearly listed or inherent to such products or devices. The term "module" as used in this application refers to any known or later-developed hardware, software, firmware, artificial intelligence, fuzzy logic, or a combination of hardware and / or software code that can perform functions related to that element.

[0028] The following will specifically describe the technical solutions of this application and how the technical solutions of this application solve the above technical problems through specific embodiments. These several specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The following will describe the method for extracting documents based on a large language model of the present invention in conjunction with Figure 1 and Figure 2 Describe the method for extracting documents based on a large language model of the present invention.

[0029] Exemplarily, the following will specifically describe the execution subject of the method for extracting documents based on a large language model as the system for extracting documents based on a large language model.

[0030] Figure 1 FIG. 14 is one of the schematic flowcharts of the method for extracting documents based on a large language model provided by the present invention. As Figure 1 shown, the method includes steps 101 to 103.

[0031] Step 101: Construct a document extraction template according to the document extraction requirements.

[0032] Step 102: Obtain the document to be extracted, and convert the document to be extracted into a processable text format to form a first text.

[0033] Step 103: Input the first text and the document extraction template into the large language model, apply the large language model to read and analyze the first text, and extract the first text according to the document extraction template to obtain the document extraction result corresponding to the document to be extracted.

[0034] In practical applications, the execution subject of the method for extracting documents based on a large language model can be the system for extracting documents based on a large language model. There are various implementation methods for the system for extracting documents based on a large language model. For example, it can be implemented through a computer program, such as an application software, etc.; or, for example, a chip, etc. It can also be implemented as a medium storing relevant computer programs, such as a USB flash drive, a cloud disk, etc.; or, it can also be implemented through an entity device integrated or installed with relevant computer programs, such as a server, a smart device, etc.

[0035] Specifically, step 101 includes: constructing a literature extraction template according to the literature extraction requirements.

[0036] Among them, the literature extraction requirements include extraction rules, format requirements, and pre-extracted fields. The pre-extracted fields include: basic research information, research design type, research grouping, participant characteristics, intervention measures for each research grouping, and background medications for each research grouping.

[0037] In this embodiment, the extraction rules refer to the specific instructions or criteria used to guide the large language model to identify and extract specific information from the literature during the literature extraction process. These rules are usually formulated based on the structure of the literature, content characteristics, and research requirements.

[0038] In this embodiment, the format requirements refer to the specific specifications for the organization, presentation, and storage methods of information when extracting information. These requirements ensure that the extracted information has a unified structure and format, facilitating subsequent analysis and comparison. Exemplarily, the format requirements may include the format of the literature extraction results.

[0039] Exemplarily, the basic research information includes but is not limited to: the first author of the research, the publication year, and the country of the participants. The research design types include but are not limited to: randomized controlled trials, cohort studies, case-control experiments, cross-sectional studies, ecological studies, case reports, and case series.

[0040] In practical applications, in medical research, the research grouping (Study Groups) refers to the assignment of research subjects (such as patients, subjects, samples, etc.) to different groups according to the research design in order to compare the effects of different treatments or intervention measures. The research grouping is the core part of the experimental design, especially in research types such as randomized controlled trials, cohort studies, and case-control studies.

[0041] Specifically, according to the text information corresponding to the abstract, method, and subtitle of the baseline information table, the research grouping is clarified. On the basis of fully retaining the original statement, the grouping is identified and extracted according to population characteristics (Population Characteristics), intervention methods (Intervention Methods), etc. The control group is extracted according to the original description, such as placebo, control, etc.

[0042] Exemplarily, participant characteristics include but are not limited to: age, health status, common surgeries or medical treatments, occupation, gender, and other population characteristics clearly mentioned in the text. Specifically, according to the text information corresponding to the title, abstract, and method subtitle, extract the participant characteristics of each group, refer to the inclusion criteria in the reference and other relevant content, extract the common characteristics of the participants in each group according to the original description, and integrate them into six elements: age, health status, common surgeries or medical treatments, occupation, gender, and other population characteristics clearly mentioned in the text.

[0043] Among them, an intervention, in medical research, refers to the treatment or measure applied to the research subjects, aiming to evaluate its impact on health outcomes. Exemplarily, the interventions for each research group include but are not limited to: drug trade name, manufacturer, dosage and administration frequency, administration time, and administration method.

[0044] Specifically, according to the text information corresponding to the abstract, method, and baseline information table subtitle, extract the interventions for each research group, including drug trade name, manufacturer, dosage and administration frequency, administration time, and administration method.

[0045] Among them, background medication refers to the drugs or treatment measures commonly used by all research groups (including experimental groups and control groups) during the medical research process, that is, the drugs used simultaneously with the intervention drugs and used in all groups.

[0046] Specifically, according to the text information corresponding to the method subtitle, extract the background medication for each research group, that is, the drugs used simultaneously with the intervention drugs and used in all groups. In practice, the literature extraction system based on the large language model will judge whether it is the background medication for each research group through the large language model, and output the drug name or drug type according to the original description, such as metformin, antidiabetic drugs, etc.

[0047] It can be understood that by constructing a literature extraction template, the specific requirements for literature extraction are clarified, including extraction rules, format requirements, and pre-extracted fields, etc. This makes the entire extraction process have a clear goal and direction, and the construction of the template standardizes and normalizes the extraction process. The large language model can directly operate according to the template without redefining the extraction rules and format every time, thus greatly improving the extraction efficiency.

[0048] Furthermore, step 102 includes: obtaining the literature to be extracted and converting the literature to be extracted into a processable text format to form a first text.

[0049] In practical applications, the document to be extracted can be provided by the researcher. For example, the document extraction system based on the large language model is provided with an input interface or an input port, and the researcher can directly input the document to be extracted. For another example, the document extraction system based on the large language model is connected to an intelligent device, and the researcher sends the document to be extracted to the document extraction system based on the large language model through the intelligent device. For another example, the document extraction system based on the large language model is provided with a human-computer interaction interface. The researcher can input keywords to search for relevant documents. The researcher can select the document to be extracted from the relevant documents found. After the researcher selects the document to be extracted, the document extraction system based on the large language model obtains the document to be extracted and starts to extract the document.

[0050] In this embodiment, documents in multiple formats can be extracted. Exemplarily, the document to be extracted can be a PDF file, a Word document, an HTML file, an XML file, a plain text file, an Excel file, etc.

[0051] Among them, the text format that can be processed refers to the text format that is convenient for subsequent natural language processing and information extraction. As an example, the text format that can be processed can be a plain text format, a structured text format (such as HTML, XML), a tokenized text format, a JSON format, etc. In practice, the text format that can be processed can be determined according to the actual document extraction requirements and actual application scenarios, and no specific limitation is made here.

[0052] It can be understood that by supporting multiple document formats and converting them into a text format that can be processed, different input requirements can be better adapted, and the flexibility and efficiency of document extraction can be improved. Obtain the document to be extracted and convert the document to be extracted into a text format that can be processed to form the first text, which is convenient for subsequent processing, improves the compatibility and versatility of the system, and enables the processing of multiple document formats.

[0053] Specifically, step 103 includes: inputting the first text and the document extraction template into the large language model, applying the large language model to read and analyze the first text, and extracting the first text according to the document extraction template to obtain the document extraction result corresponding to the document to be extracted.

[0054] In practical applications, the first text and the document extraction template are input into the large language model. The large language model applies its powerful language understanding ability to read and analyze the first text, can accurately identify and extract the key information in the document, and reduces the errors and omissions that may occur in manual extraction. Further, the large language model extracts the first text according to the document extraction template, and a document extraction result that meets the document extraction requirements can be obtained.

[0055] In this embodiment, the format of the literature extraction result is not specifically limited. As an example, the format of the literature extraction result includes at least one of the following: an extraction table, a hierarchical diagram, a tree diagram, and a mind map.

[0056] Optionally, in a possible implementation manner, Figure 2 is the second flowchart of the literature extraction method based on a large language model provided by the present invention. On the basis of the above embodiment, the above literature extraction result is an extraction table; the above step 103 includes: Step 201: Input the first text and the literature extraction template into the large language model; Step 202: Divide the first text according to a preset structured rule to obtain multiple text segments and form a second text; Step 203: Apply the large language model to read and analyze the second text, and extract the information corresponding to the pre-extracted fields according to the extraction rule and format requirement to form a third text; Step 204: Convert the third text into a table form to obtain the extraction table corresponding to the literature to be extracted.

[0057] In this application, the manner of inputting the first text and the literature extraction template into the large speech model is not specifically limited. In one example, the first text and the literature extraction template are directly input into the large speech model. In another example, a prompt word is generated according to the first text and the literature extraction template, and the prompt word is input into the large speech model.

[0058] Among them, the preset structured rule refers to the division logic predefined according to the common format and content structure of the literature. Exemplarily, the preset structured rule can be based on structural information such as titles, paragraphs, and keywords in the literature. For example, the paragraphs in the literature are divided into multiple text segments. For example, according to the order of the paragraphs in the literature, the literature is divided into multiple text segments, and each text segment includes 5 paragraphs.

[0059] Optionally, the literature to be extracted can be divided according to the subtitle. In some examples, the above step 202 specifically includes: Divide the first text according to the subtitle in the literature to be extracted to obtain the text information corresponding to each subtitle and form a second text; where the subtitle includes: title, abstract, background introduction, method, and baseline information table.

[0060] Specifically, divide the first text according to the subtitle in the literature to be extracted to obtain the text information corresponding to each subtitle and form a second text. Among them, the text information corresponding to each subtitle is used as a text segment, and each text segment corresponds to a logical part in the literature.

[0061] In practical applications, the divided text fragments are organized into a second text, and each text fragment is accompanied by a clear identifier (such as a subtitle name) for facilitating the processing of subsequent steps. The second text is a structured data set. Exemplarily, the second text can be stored in the form of a list or a dictionary.

[0062] It can be understood that, on the one hand, through structured division, the system can quickly locate the area where the key information is located, reducing the processing time of invalid information. The automated division reduces the time for manual reading and dividing the literature, significantly improving the extraction efficiency. On the other hand, the structured division enables the system to extract information more accurately, reducing errors caused by unclear information positions. Each fragment has a clear identifier, facilitating targeted processing in subsequent steps. On yet another hand, the divided second text is a structured data set, facilitating subsequent information extraction and analysis. Through the preset structured rules, the system can adapt to documents in different formats, improving the versatility and flexibility of the system.

[0063] Optionally, step 203 described above includes: Receiving the second text and deeply analyzing the text information corresponding to each subtitle; Applying a large language model to extract the information corresponding to the pre-extracted fields one by one from the text information corresponding to each subtitle according to the extraction table based on the PICO format to form a third text.

[0064] Among them, the PICO (Population Intervention Comparison Outcome) format is a framework widely used in evidence-based medicine and medical literature research for clarifying and organizing the key elements of clinical problems or research questions. PICO is the initials abbreviation of four English words, representing Population, Intervention, Comparison, and Outcome respectively. This format helps researchers and clinicians clearly define research questions, improving the efficiency and accuracy of literature retrieval, research design, and result analysis.

[0065] In practical applications, in the literature extraction system based on a large language model, an extraction table formed through extensive discussions among experts in the field of evidence-based medicine based on the PICO format is used to extract key information such as basic research information, research design type, research grouping, participant characteristics, intervention measures for each research grouping, and background medications for each research grouping one by one from the second text according to the original content.

[0066] Optionally, in a possible implementation, the above application of the large language model extracts the information corresponding to the pre-extracted fields one by one from the text information corresponding to each subtitle according to the extraction table based on the PICO format, forming the third text, including: Extract the basic research information and research design type from the text information corresponding to the title and abstract subtitles; among them, the basic research information includes: the first author of the research, the publication year, and the country of the participants; the research design type includes at least one of the following: randomized controlled trial, cohort study, case-control experiment, cross-sectional study, ecological study, case report, and case series; Extract the research grouping from the text information corresponding to the abstract, method, and baseline information table subtitles; Extract the participant characteristics from the text information corresponding to the title, abstract, and method subtitles; Extract the intervention measures for each research grouping from the text information corresponding to the abstract, method, and baseline information table subtitles; the intervention measures for each research grouping include: drug trade name, manufacturer, dose and administration frequency, administration time, and administration method; Extract the background medications for each research grouping from the text information corresponding to the method subtitle.

[0067] Specifically, extract the first author of the research, the publication year, the country of the participants, and the research design type from the text information corresponding to the title and abstract subtitles. If it is a randomized controlled trial (RCT) study, the system will further extract the experimental registration platform and registration number.

[0068] Specifically, clarify the research grouping according to the text information corresponding to the abstract, method, and baseline information table subtitles. On the basis of fully retaining the original statement, identify and extract the grouping according to population characteristics, intervention methods, etc. The control group is extracted according to the original description, such as placebo, control, etc.

[0069] Specifically, extract the participant characteristics of each group according to the text information corresponding to the title, abstract, and method subtitles, refer to the inclusion criteria in the reference and other relevant content, extract the common characteristics of the participants in each group according to the original description, and integrate them into six major elements: age, health status, common surgery or drug treatment, occupation, gender, and other population characteristics clearly mentioned in the text.

[0070] Specifically, according to the text information corresponding to the abstract, method, and subtitle of the baseline information table, extract the intervention measures for each study group in each subgroup, including the drug trade name, manufacturer, dose, administration frequency, administration time, and administration method.

[0071] Specifically, according to the text information corresponding to the subtitle of the method, extract the background medications for each study group, that is, the medications that are used simultaneously with the intervention drugs and are used in all subgroups. In practice, the literature extraction system based on the large language model will determine whether it is the background medication for each study group through the large language model, and output the drug name or drug type according to the original description, such as metformin, antidiabetic drugs, etc.

[0072] It can be understood that by applying the large language model to read and analyze the second text, and extracting the information corresponding to the preset fields according to the extraction rules and format requirements to form the third text, the cumbersome process of manual extraction is reduced, the extraction efficiency is improved, and moreover, the large language model can understand complex language structures and context relationships, and extract information according to the preset rules, improving the accuracy and consistency of the extracted information.

[0073] Optionally, in some possible implementation manners, the above research design type is a randomized controlled trial, and the above method further includes: Extract and obtain the experimental registration platform and registration number from the text information corresponding to the subtitle of the method.

[0074] Furthermore, step 204 includes: converting the third text into a table form to obtain the extraction table corresponding to the literature to be extracted.

[0075] It can be understood that converting the third text into a table form to obtain and store the extraction table corresponding to the literature to be extracted is convenient for subsequent analysis and comparison, improves the readability and usability of the information, and is convenient for researchers to conduct further research. Therefore, the solution of this embodiment improves the accuracy and efficiency of literature extraction.

[0076] It should be noted that in addition to being able to be stored in a table form, the third text (i.e., the extracted information set) can also be stored in a variety of other ways, specifically depending on the subsequent usage requirements and application scenarios. As an example, the third text can be in JSON format, databases (such as MySQL, MongoDB), XML format, etc.

[0077] In addition, in order to further improve the accuracy of literature extraction, the literature extraction results can be normalized. In one possible implementation manner, after the above step 103, the above method further includes: Normalize the literature extraction results. The normalization process includes at least one of the following: deleting useless descriptive information, unifying the information format, and adding marks to the information that cannot be extracted; Store the normalized literature extraction results.

[0078] In practical applications, deleting useless descriptive information makes the extraction results more concise and clear, avoids interference from irrelevant information in subsequent analysis, improves data quality, reduces the complexity of subsequent processing, and saves time and computing resources. For example, after extracting the intervention measures for each study group in each group, normalize the information and remove marks such as "®" and "TM" of the trade name. It can be understood that unifying the information format ensures the uniformity of the extracted information format, facilitates subsequent comparison and analysis, and the unified format makes the extraction results easier to understand and use. For example, after extracting the intervention measures for each study group in each group, summarize the dose and administration frequency into the "dose / time" format, such as 0.1mg / d; simplify the administration time to numbers and units, such as 14 days; summarize the administration route into inhalation, oral, injection, etc.

[0079] It can be understood that adding marks to the information that cannot be extracted can clearly mark the information that cannot be extracted, enable researchers to understand the integrity of the data, provide a reference for subsequent research or data supplementation, and avoid missing important information. Further, by marking the missing information, the reliability and credibility of the data are improved. For example, check the information corresponding to the pre-extracted fields obtained by extraction. If the relevant information is not mentioned in the original text, output "NA".

[0080] Optionally, the normalization process may further include deleting redundant information, data cleaning, standardizing terms, unifying the date and numerical format, unifying the logical structure, marking uncertain information, etc., which are not specifically limited herein.

[0081] In practical applications, researchers need to study multiple research literatures. The method of the present invention can perform batch information extraction on multiple research literatures. Optionally, in a possible implementation manner, there are multiple literatures to be extracted; after the above step 103, the above method further includes: Determine whether all the literatures to be extracted have been extracted; if all the literatures to be extracted have been extracted, integrate and output the literature extraction results corresponding to all the literatures to be extracted; otherwise, return to execute the above step 102.

[0082] In this embodiment, it is supported to extract multiple research documents, ensuring that all documents have gone through a complete extraction process, and integrating the extraction results into a comprehensive full table of document information. This method not only improves the integrity and efficiency of data processing, but also provides high-quality data support for subsequent comprehensive analysis. Through automatic judgment and loop processing, the system can efficiently process a large number of documents, reduce human errors, and improve the reliability and consistency of data.

[0083] The document extraction method based on the large language model provided in this embodiment constructs a document extraction template, which clarifies the specific requirements for document extraction, including extraction rules, format requirements, pre-extracted fields, etc. This makes the entire extraction process have clear goals and directions, and the construction of the template standardizes and normalizes the extraction process. The large language model can directly operate according to the template without redefining the extraction rules and formats every time extraction is performed, thus greatly improving the extraction efficiency. Further, by converting the document to be extracted into a processable text format to form the first text, it is convenient for subsequent processing, improves the compatibility and versatility of the system, and enables it to process various document formats. Further, inputting the first text and the document extraction template into the large language model, applying the large language model to read and analyze the first text, and extracting the first text according to the document extraction template to obtain the document extraction result corresponding to the document to be extracted, reduces the cumbersome process of manual extraction, improves the extraction efficiency, and moreover, the large language model can understand complex language structures and context relationships, and can extract information according to the document extraction template, improving the accuracy and consistency of the extracted information. Therefore, the solution of this embodiment improves the accuracy and efficiency of document extraction.

[0084] The document extraction system based on the large language model provided by the present invention will be described below. The document extraction system based on the large language model described below can be mutually referred to and corresponding to the document extraction method based on the large language model described above.

[0085] Figure 3 is a schematic structural diagram of the document extraction system based on the large language model provided by the present invention, as Figure 3 shown, the above-mentioned document extraction system based on the large language model includes: a construction module 31, an acquisition module 32, and an extraction module 33.

[0086] The construction module 31 is used to construct a document extraction template according to document extraction requirements.

[0087] The acquisition module 32 is used to acquire the document to be extracted and convert the document to be extracted into a processable text format to form the first text.

[0088] The extraction module 33 is configured to input the first text and the literature extraction template into the large language model, apply the large language model to read and analyze the first text, and extract the first text according to the literature extraction template to obtain the literature extraction result corresponding to the literature to be extracted.

[0089] In practical applications, there are various implementation methods for the literature extraction system based on the large language model. For example, it can be implemented through computer programs, such as application software, etc.; or, for example, chips, etc. It can also be implemented as a medium storing relevant computer programs, such as USB flash drives, cloud disks, etc.; or, it can also be implemented through an entity device integrated or installed with relevant computer programs, such as servers, intelligent devices, etc.

[0090] Specifically, the construction module 31 is configured to: construct a literature extraction template according to the literature extraction requirements.

[0091] Among them, the literature extraction requirements include extraction rules, format requirements, and pre-extracted fields. The pre-extracted fields include: basic research information, research design type, research grouping, participant characteristics, intervention measures for each research grouping, and background medications for each research grouping.

[0092] In this embodiment, the extraction rule refers to specific instructions or criteria used to guide the large language model to identify and extract specific information from the literature during the literature extraction process. These rules are usually formulated based on the structure, content characteristics, and research requirements of the literature.

[0093] In this embodiment, the format requirement refers to the specific specifications for the organization, presentation, and storage methods of information when extracting information. These requirements ensure that the extracted information has a unified structure and format, facilitating subsequent analysis and comparison. Exemplarily, the format requirement may include the format of the literature extraction result.

[0094] Exemplarily, the basic research information includes but is not limited to: the first author of the research, the publication year, and the country of the participants. The research design types include but are not limited to: randomized controlled trials, cohort studies, case-control experiments, cross-sectional studies, ecological studies, case reports, and case series.

[0095] In practical applications, in medical research, the research grouping (Study Groups) refers to the allocation of research subjects (such as patients, subjects, samples, etc.) to different groups according to the research design in order to compare the effects of different treatments or intervention measures. The research grouping is the core part of the experimental design, especially in research types such as randomized controlled trials, cohort studies, and case-control studies.

[0096] Specifically, based on the text information corresponding to the abstract, method, and subtitle of the baseline information table, clarify the research groups. On the basis of fully retaining the original statements, identify and extract the groups according to population characteristics, intervention methods, etc. The control group is extracted according to the original description, such as placebo, control, etc.

[0097] Exemplarily, the participant characteristics include but are not limited to: age, health status, common surgeries or drug treatments, occupation, gender, and other population characteristics clearly mentioned in the text. Specifically, according to the text information corresponding to the title, abstract, and subtitle of the method, extract the participant characteristics of each group, refer to the inclusion criteria in the reference and other relevant content, extract the common characteristics of the participants in each group according to the original description, and integrate them into six major elements: age, health status, common surgeries or drug treatments, occupation, gender, and other population characteristics clearly mentioned in the text.

[0098] Among them, the intervention in medical research refers to the treatment or therapy applied to the research subjects, aiming to evaluate its impact on health outcomes. Exemplarily, the intervention measures for each research group include but are not limited to: drug trade name, manufacturer, dose and administration frequency, administration time, administration method.

[0099] Specifically, according to the text information corresponding to the abstract, method, and subtitle of the baseline information table, extract the intervention measures for each research group, including drug trade name, manufacturer, dose and administration frequency, administration time, administration method.

[0100] Among them, background medication refers to the drugs or treatment measures commonly used by all research groups (including experimental groups and control groups) during the medical research process, that is, the drugs used simultaneously with the intervention drugs and used in all groups.

[0101] Specifically, according to the text information corresponding to the subtitle of the method, extract the background medication for each research group, that is, the drugs used simultaneously with the intervention drugs and used in all groups. In practice, the literature extraction system based on the large language model will judge whether it is the background medication for each research group through the large language model, and output the drug name or drug type according to the original description, such as metformin, antidiabetic drugs, etc.

[0102] It can be understood that by constructing a literature extraction template, the specific requirements for literature extraction are clarified, including extraction rules, format requirements, pre-extracted fields, etc. This makes the entire extraction process have clear goals and directions, and the construction of the template standardizes and regularizes the extraction process. The large language model can directly operate according to the template without redefining extraction rules and formats every time it extracts, thus greatly improving the extraction efficiency.

[0103] Furthermore, the acquisition module 32 is used to: acquire the literature to be extracted and convert the literature to be extracted into a processable text format to form a first text.

[0104] In practical applications, the literature to be extracted can be provided by researchers. For example, the literature extraction system based on the large language model is provided with an input interface or an input port, and the researcher can directly input the literature to be extracted. Another example is that the literature extraction system based on the large language model is connected to an intelligent device, and the researcher sends the literature to be extracted to the literature extraction system based on the large language model through the intelligent device. Another example is that the literature extraction system based on the large language model is provided with a human-computer interaction interface, and the researcher can input keywords to search for relevant literature. The researcher can select the literature to be extracted from the relevant documents found. After the researcher selects the literature to be extracted, the literature extraction system based on the large language model acquires the literature to be extracted and starts to extract the literature.

[0105] In this embodiment, literatures in multiple formats can be extracted. Exemplarily, the literature to be extracted can be a PDF file, a Word document, an HTML file, an XML file, a plain text file, an Excel file, etc.

[0106] Among them, the processable text format refers to a text format that is convenient for subsequent natural language processing and information extraction. As an example, the processable text format can be a plain text format, a structured text format (such as HTML, XML), a tokenized text format, a JSON format, etc. In practice, the processable text format can be determined according to the actual literature extraction requirements and actual application scenarios, and no specific limitation is made here.

[0107] It can be understood that by supporting multiple literature formats and converting them into processable text formats, different input requirements can be better adapted, and the flexibility and efficiency of literature extraction can be improved. Acquiring the literature to be extracted and converting the literature to be extracted into a processable text format to form a first text is convenient for subsequent processing, improving the compatibility and generality of the system so as to be able to process multiple literature formats.

[0108] Specifically, the extraction module 33 is used to: input the first text and the literature extraction template into the large language model, apply the large language model to read and analyze the first text, and extract the first text according to the literature extraction template to obtain the literature extraction result corresponding to the literature to be extracted.

[0109] In practical applications, the first text and the literature extraction template are input into the large language model. The large language model applies its powerful language understanding ability to read and analyze the first text, can accurately identify and extract key information in the literature, and reduces errors and omissions that may occur in manual extraction. Further, the large language model extracts the first text according to the literature extraction template, and a literature extraction result that meets the literature extraction requirements can be obtained.

[0110] In this embodiment, the format of the literature extraction result is not specifically limited. As an example, the format of the literature extraction result includes at least one of the following: extraction table, hierarchical diagram, tree diagram, mind map.

[0111] Optionally, in a possible implementation manner, on the basis of the above embodiment, the above literature extraction result is an extraction table; the above extraction module 33 includes: An input unit, configured to input the first text and the literature extraction template into the large language model; A partitioning unit, configured to partition the first text according to a preset structured rule to obtain a plurality of text segments, forming a second text; An extraction unit, configured to apply the large language model to read and analyze the second text, and extract information corresponding to pre-extracted fields according to extraction rules and format requirements, forming a third text; A format conversion unit, configured to convert the third text into a table form to obtain an extraction table corresponding to the literature to be extracted.

[0112] In this application, the manner of inputting the first text and the literature extraction template into the large voice model is not specifically limited. In one example, the input unit directly inputs the first text and the literature extraction template into the large voice model. In another example, the input unit generates a prompt word according to the first text and the literature extraction template, and inputs the prompt word into the large voice model.

[0113] The preset structured rule refers to a partitioning logic predefined according to the common format and content structure of the literature. Exemplarily, the preset structured rule may be based on structural information such as titles, paragraphs, keywords, etc. in the literature. For example, the paragraphs in the literature are partitioned into a plurality of text segments. For example, according to the order of paragraphs in the literature, the literature is partitioned into a plurality of text segments, and each text segment includes 5 paragraphs.

[0114] Optionally, the extracted documents can be divided according to the subheadings. In some examples, the above division unit is specifically used for: Dividing the first text according to the subheadings in the document to be extracted, obtaining the text information corresponding to each subheading, and forming a second text; wherein the subheadings include: title, abstract, background introduction, method, baseline information table.

[0115] Specifically, divide the first text according to the subheadings in the document to be extracted, obtain the text information corresponding to each subheading, and form a second text. Among them, the text information corresponding to each subheading is used as a text segment, and each text segment corresponds to a logical part in the document.

[0116] In practical applications, the divided text segments are organized into a second text, and each text segment has a clear identifier (such as the subheading name), which is convenient for processing in subsequent steps. The second text is a structured data set. Exemplarily, the second text can be stored in the form of a list or a dictionary.

[0117] It can be understood that, on the one hand, through structured division, the system can quickly locate the area where the key information is located, reducing the processing time of invalid information. The automated division reduces the time for manual reading and dividing the documents, significantly improving the extraction efficiency. On the other hand, the structured division enables the system to extract information more accurately, reducing errors caused by unclear information positions. Each segment has a clear identifier, which is convenient for targeted processing in subsequent steps. On the other hand, the divided second text is a structured data set, which is convenient for subsequent information extraction and analysis. Through the preset structured rules, the system can adapt to documents in different formats, improving the versatility and flexibility of the system.

[0118] Optionally, the above extraction unit is specifically used for: Receiving the second text and deeply analyzing the text information corresponding to each subheading; Applying a large language model, extracting the information corresponding to the pre-extracted fields one by one from the text information corresponding to each subheading according to the extraction table based on the PICO format, and forming a third text.

[0119] Among them, the PICO (Population Intervention Comparison Outcome) format is a framework widely used in evidence-based medicine and medical literature research to clarify and organize the key elements of clinical or research questions. PICO is an acronym for four English words, representing Population, Intervention, Comparison, and Outcome respectively. This format helps researchers and clinicians clearly define research questions and improve the efficiency and accuracy of literature retrieval, research design, and result analysis.

[0120] In practical applications, in the literature extraction system based on large language models, an extraction form based on the PICO format and widely discussed by experts in the field of evidence-based medicine is used to extract key information such as basic research information, research design type, research grouping, participant characteristics, intervention measures for each research grouping, and background medications for each research grouping from the second text one by one according to the original content.

[0121] Optionally, in a possible implementation manner, the above extraction unit is used to apply a large language model and extract the information corresponding to the pre-extracted fields one by one from the text information corresponding to each subtitle according to the extraction form based on the PICO format. When forming the third text, it is specifically used for: Extract the basic research information and research design type from the text information corresponding to the title and abstract subtitles; among them, the basic research information includes: the first author of the research, the publication year, and the country of the participants; the research design type includes at least one of the following: randomized controlled trial, cohort study, case-control experiment, cross-sectional study, ecological study, case report, and case series; Extract the research grouping from the text information corresponding to the abstract, methods, and baseline information table subtitles; Extract the participant characteristics from the text information corresponding to the title, abstract, and methods subtitles; Extract the intervention measures for each research grouping from the text information corresponding to the abstract, methods, and baseline information table subtitles; the intervention measures for each research grouping include: drug trade name, manufacturer, dose and administration frequency, administration time, and administration method; Extract the background medications for each research grouping from the text information corresponding to the methods subtitle.

[0122] Specifically, extract the first author of the research, the publication year, the country of the participants, and the research design type from the text information corresponding to the title and abstract subtitles. If it is a randomized controlled trial (RCT) study, the system will further extract the experimental registration platform and registration number.

[0123] Specifically, based on the text information corresponding to the abstract, method, and subtitle of the baseline information table, the research groups are clearly defined. On the basis of fully retaining the original statements, the groups are identified and extracted according to population characteristics, intervention methods, etc. The control group is extracted according to the original description, such as placebo, control, etc.

[0124] Specifically, according to the text information corresponding to the title, abstract, and subtitle of the method, the participant characteristics of each group are extracted, including the inclusion criteria in the reference and other relevant content. The common characteristics of the participants in each group are extracted according to the original description and integrated into six major elements: age, health status, common surgeries or drug treatments received, occupation, gender, and other population characteristics clearly mentioned in the text.

[0125] Specifically, according to the text information corresponding to the abstract, method, and subtitle of the baseline information table, the intervention measures of each research group are extracted, including the drug trade name, manufacturer, dose and administration frequency, administration time, and administration method.

[0126] Specifically, according to the text information corresponding to the subtitle of the method, the background medications of each research group are extracted, that is, the medications used simultaneously with the intervention drugs and used in all groups. In practice, the literature extraction system based on the large language model will judge whether it is the background medication of each research group through the large language model and output the drug name or drug type according to the original description, such as metformin, antidiabetic drugs, etc.

[0127] It can be understood that by applying the large language model, reading and analyzing the second text, and extracting the information corresponding to the preset fields according to the extraction rules and format requirements to form the third text, the cumbersome process of manual extraction is reduced, the extraction efficiency is improved, and moreover, the large language model can understand complex language structures and context relationships and extract information according to the preset rules, improving the accuracy and consistency of the extracted information.

[0128] Optionally, in some possible implementation manners, the above research design type is a randomized controlled trial, and the above extraction unit is further used for: Extracting the experimental registration platform and registration number from the text information corresponding to the subtitle of the method.

[0129] Furthermore, the format conversion unit is used for: converting the third text into a table form to obtain the extraction table corresponding to the literature to be extracted.

[0130] It can be understood that converting the third text into a tabular form to obtain and store the extraction table corresponding to the literature to be extracted facilitates subsequent analysis and comparison, improves the readability and usability of information, and is convenient for researchers to conduct further research. Therefore, the solution of this embodiment improves the accuracy and efficiency of literature extraction.

[0131] It should be noted that in addition to being converted into a tabular form for storage, the third text (i.e., the extracted information set) can also be stored in a variety of other ways, specifically depending on subsequent usage requirements and application scenarios. As an example, the third text can be in JSON format, databases (such as MySQL, MongoDB), XML format, etc.

[0132] In addition, in order to further improve the accuracy of literature extraction, the literature extraction results can be normalized. In one possible implementation manner, the above system further includes: A normalization module for normalizing the literature extraction results. The normalization process includes at least one of the following: deleting useless descriptive information, unifying the information format, and adding marks to information that cannot be extracted; A storage module for storing the literature extraction results after normalization processing.

[0133] In practical applications, deleting useless descriptive information makes the extraction results more concise and clear, avoids interference from irrelevant information in subsequent analysis, improves data quality, reduces the complexity of subsequent processing, and saves time and computing resources. For example, after extracting the intervention measures of each research group in each group, the information is normalized to remove marks such as "®" and "TM" of the trade name. It can be understood that unifying the information format ensures the uniformity of the extracted information format, facilitates subsequent comparison and analysis, and the unified format makes the extraction results easier to understand and use. For example, after extracting the intervention measures of each research group in each group, the dose and administration frequency are summarized into the "dose / time" format, such as 0.1mg / d; the administration time is simplified to numbers and units, such as 14 days; the administration method is summarized into inhalation, oral, injection, etc.

[0134] It can be understood that adding marks to information that cannot be extracted can clearly mark the information that cannot be extracted, enabling researchers to understand the integrity of the data, providing a reference for subsequent research or data supplementation, and avoiding missing important information. Further, by marking the missing information, the reliability and credibility of the data are improved. For example, check the information corresponding to the pre-extracted fields obtained by extraction. If the relevant information is not mentioned in the original text, output "NA".

[0135] Optionally, the normalization process may also include deleting redundant information, data cleaning, standardizing terms, unifying date and numerical formats, unifying logical structures, marking uncertain information, etc., which are not specifically limited herein.

[0136] In practical applications, researchers need to study multiple research documents, and the method of the present invention can perform batch information extraction on multiple research documents. Optionally, in a possible implementation manner, there are multiple documents to be extracted; the above system further includes: A processing module, configured to determine whether all the documents to be extracted have been completely extracted; if all the documents to be extracted have been completely extracted, integrate and output the document extraction results corresponding to all the documents to be extracted; otherwise, return to execute the step of obtaining the documents to be extracted described above.

[0137] In this embodiment, it supports the extraction of multiple research documents, ensures that all documents have gone through a complete extraction process, and integrates the extraction results into a comprehensive full table of document information. This method not only improves the integrity and efficiency of data processing, but also provides high-quality data support for subsequent comprehensive analysis. Through automatic judgment and loop processing, the system can efficiently process a large number of documents, reduce human errors, and improve the reliability and consistency of data.

[0138] The literature extraction system based on a large language model provided in this embodiment, the construction module clarifies the specific requirements for literature extraction, including extraction rules, format requirements, pre-extracted fields, etc. by constructing a literature extraction template. This makes the entire extraction process have a clear goal and direction, and the construction of the template standardizes and normalizes the extraction process. The large language model can directly operate according to the template without redefining the extraction rules and formats every time extraction is performed, thus greatly improving the extraction efficiency. Further, the acquisition module converts the documents to be extracted into a processable text format to form a first text, which is convenient for subsequent processing and improves the compatibility and versatility of the system to be able to process various document formats. Further, the extraction module inputs the first text and the literature extraction template into the large language model, applies the large language model to read and analyze the first text, and extracts the first text according to the literature extraction template to obtain the document extraction results corresponding to the documents to be extracted, reducing the cumbersome process of manual extraction and improving the extraction efficiency. Moreover, the large language model can understand complex language structures and context relationships and can extract information according to the literature extraction template, improving the accuracy and consistency of the extracted information. Therefore, the solution of this embodiment improves the accuracy and efficiency of literature extraction.

[0139] Figure 4 is a schematic structural diagram of the electronic device provided by the present invention, as Figure 4As shown, the electronic device may include: a processor 410, a communications interface 420, a memory 430, and a communication bus 440. Among them, the processor 410, the communications interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logical instructions in the memory 430 to execute a literature extraction method based on a large language model. The method includes: constructing a literature extraction template according to the literature extraction requirements; where the literature extraction requirements include extraction rules, format requirements, and pre-extracted fields, and the pre-extracted fields include: basic research information, research design type, research grouping, participant characteristics, intervention measures for each research grouping, and background medications for each research grouping; obtaining the literature to be extracted and converting the literature to be extracted into a processable text format to form a first text; inputting the first text and the literature extraction template into the large language model, applying the large language model to read and analyze the first text, and extracting the first text according to the literature extraction template to obtain a literature extraction result corresponding to the literature to be extracted.

[0140] In addition, when the logical instructions in the above-mentioned memory 430 can be implemented in the form of software functional units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.

[0141] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the literature extraction method based on a large language model provided by the above-mentioned various methods. The method includes: constructing a literature extraction template according to the literature extraction requirements; wherein, the literature extraction requirements include extraction rules, format requirements, and pre-extracted fields. The pre-extracted fields include: basic research information, research design type, research grouping, participant characteristics, intervention measures for each research grouping, and background medications for each research grouping; obtaining the literature to be extracted, and converting the literature to be extracted into a processable text format to form a first text; inputting the first text and the literature extraction template into the large language model, applying the large language model to perform reading analysis on the first text, and extracting the first text according to the literature extraction template to obtain a literature extraction result corresponding to the literature to be extracted.

[0142] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it is implemented to execute the literature extraction method based on a large language model provided by the above-mentioned various methods. The method includes: constructing a literature extraction template according to the literature extraction requirements; wherein, the literature extraction requirements include extraction rules, format requirements, and pre-extracted fields. The pre-extracted fields include: basic research information, research design type, research grouping, participant characteristics, intervention measures for each research grouping, and background medications for each research grouping; obtaining the literature to be extracted, and converting the literature to be extracted into a processable text format to form a first text; inputting the first text and the literature extraction template into the large language model, applying the large language model to perform reading analysis on the first text, and extracting the first text according to the literature extraction template to obtain a literature extraction result corresponding to the literature to be extracted.

[0143] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0144] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A literature extraction method based on large language models, characterized in that, Including: Construct a literature extraction template according to the literature extraction requirements. Among them, the literature extraction requirements include extraction rules, format requirements, and pre-extracted fields. The pre-extracted fields include: basic research information, research design type, research grouping, participant characteristics, intervention measures for each research grouping, and background medications for each research grouping. Obtain the literature to be extracted and convert the literature to be extracted into a processable text format to form a first text. Input the first text and the literature extraction template into a large language model, apply the large language model to read and analyze the first text, and extract the first text according to the literature extraction template to obtain the literature extraction result corresponding to the literature to be extracted.

2. The method for extracting documents based on a large language model according to claim 1, wherein The literature extraction result is an extraction table. Applying the large language model to read and analyze the first text and extracting the first text according to the literature extraction template to obtain the literature extraction result corresponding to the literature to be extracted includes: Divide the first text according to preset structured rules to obtain multiple text fragments to form a second text. Apply the large language model to read and analyze the second text, and extract the information corresponding to the pre-extracted fields according to the extraction rules and the format requirements to form a third text. Convert the third text into a table form to obtain the extraction table corresponding to the literature to be extracted.

3. The literature extraction method based on a large language model according to claim 2, wherein Applying the large language model to read and analyze the second text, and extracting the information corresponding to the pre-extracted fields according to the extraction rules and the format requirements to form a third text includes: Receive the second text and deeply analyze the text information corresponding to each subtitle. Apply the large language model to extract the information corresponding to the pre-extracted fields one by one from the text information corresponding to each subtitle according to the extraction table based on the PICO format to form the third text.

4. The literature extraction method based on a large language model according to claim 3, characterized in that, Applying the large language model to extract the information corresponding to the pre-extracted fields one by one from the text information corresponding to each subtitle according to the extraction table based on the PICO format to form the third text includes: Extract the basic research information and research design type from the text information corresponding to the title and abstract subtitles. Among them, the basic research information includes: first author of the research, publication year, and country of participants. The research design type includes at least one of the following: randomized controlled trial, cohort study, case-control experiment, cross-sectional study, ecological study, case report, and case series. Extract the research grouping from the text information corresponding to the abstract, method, and baseline information table subtitles. Extract the participant characteristics from the text information corresponding to the title, the abstract, and the method subtitles. Extract the intervention measures for each research grouping from the text information corresponding to the abstract, the method, and the baseline information table subtitles. The intervention measures for each research grouping include: drug trade name, manufacturer, dose and administration frequency, administration time, and administration method. Extract the background medications for each study group from the text information corresponding to the subtitle of the method.

5. The method for literature extraction based on a large language model according to claim 4, characterized in that, The research design type is a randomized controlled trial; the method further includes: Extract the experimental registration platform and registration number from the text information corresponding to the subtitle of the method.

6. The literature extraction method based on a large language model according to any one of claims 1-5, characterized in that, After obtaining the literature extraction result corresponding to the literature to be extracted, the method further includes: Normalize the literature extraction result, where the normalization process includes at least one of the following: deleting useless descriptive information, unifying the information format, and adding marks to information that cannot be extracted; Store the literature extraction result after normalization.

7. The method for literature extraction based on a large language model according to any one of claims 1-5, characterized in that, There are multiple literatures to be extracted; after obtaining the literature extraction result corresponding to the literature to be extracted, the method further includes: Determine whether all the literatures to be extracted have been completely extracted; if all the literatures to be extracted have been completely extracted, integrate and output the literature extraction results corresponding to all the literatures to be extracted; otherwise, return to execute the step of obtaining the literature to be extracted.

8. A literature extraction system based on large language models, characterized in that, Include: A construction module for constructing a literature extraction template according to the literature extraction requirements, where the literature extraction requirements include extraction rules, format requirements, and pre-extracted fields, and the pre-extracted fields include: basic research information, research design type, study groups, participant characteristics, intervention measures for each study group, and background medications for each study group; An acquisition module for acquiring the literature to be extracted and converting the literature to be extracted into a processable text format to form a first text; An extraction module for inputting the first text and the literature extraction template into a large language model, applying the large language model to read and analyze the first text, and extracting the first text according to the literature extraction template to obtain the literature extraction result corresponding to the literature to be extracted.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the literature extraction method based on a large language model according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the literature extraction method based on a large language model according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Scientific and technical literature block function identification method and device based on multi-level semantics

    CN116050395A

  • Knowledge graph construction method and device, knowledge graph query method and device, electronic equipment and storage medium

    CN116775897A

  • Automatic selection and recommendation system for evidence in traditional Chinese medicine literature information

    CN118821759A

  • Literature screening method and device, electronic equipment and storage medium

    CN119226432A

  • Drug safety large model method and system based on medical knowledge graph

    CN119541902A

Cited By

  • Index simplification method and system for water ecological environment toughness evaluation

    CN122134178A

  • Biomedical literature elaborated content generation system and generation method

    CN122153053A

  • Material paper parameter extraction and model evaluation method based on RagFlow

    CN122197874A