Scientific literature experimental data extraction method and device based on large language model

Through a method based on a large language model, scientific literature is automatically acquired and analyzed, and experimental data is extracted and cleaned, which solves the problems of low data extraction efficiency and insufficient accuracy in the existing technology, and realizes efficient and accurate data extraction to meet cross-domain needs.

CN120011542APending Publication Date: 2025-05-16GUANGDONG SHUIMU QINGYU TECHNOLOGY CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510106016.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When the prior art processes a large number of scientific literature, data extraction efficiency is low, insufficient accuracy, and poor flexibility, which cannot meet cross-domain needs.

Method used

Using a method based on a large language model, relevant literature in the target field is obtained through script code, and the literature is analyzed based on preset prompt words through the trained large language model and experimental data is extracted. This method includes data cleaning and verification, and finally exported as a structured data file.

Benefits of technology

It realizes automatic, efficient and accurate extraction of experimental data from scientific literature, improves the efficiency and accuracy of data processing, can meet the needs of scientific researchers to quickly obtain key information, and improves the overall efficiency of literature analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011542A_ABST
    Figure CN120011542A_ABST
Patent Text Reader

Abstract

The invention relates to a scientific literature experimental data extraction method and device based on a large language model. The method comprises the following steps: acquiring related literatures in a target field according to a preset keyword through a script code; analyzing the related literature according to a preset cue word through the trained large language model, and extracting experimental data in the related literature; performing data cleaning and verification on the extracted experimental data to obtain cleaned data; and exporting the cleaned data in a preset export format. By adopting the method, the experimental data can be automatically, efficiently and accurately extracted from scientific literatures.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of scientific document processing, and in particular to a method and device for extracting scientific document experimental data based on a large language model. Background Art

[0002] In the field of scientific research, literature analysis is an important step in obtaining key data and guiding experimental design. However, with the rapid growth of the number of papers in various disciplines, manually screening and extracting experimental data and key information from papers has become extremely time-consuming and inefficient.

[0003] At present, traditional literature analysis methods mainly include the following two methods: one is manual screening based on keyword retrieval, that is, researchers use keywords to retrieve relevant papers through literature databases, and extract experimental data by manually reading the paper content. This method relies on the experience and energy of researchers, and is prone to missing key data due to omissions; at the same time, facing massive literature, the processing efficiency is low and the cost is high. The other is rule-based automated data mining, such as literature analysis software based on natural language processing, which can automatically extract information from tables, images or experimental paragraphs from papers in PDF or HTML format through predefined rules and algorithms. However, this type of method cannot effectively parse complex natural language expressions and ambiguity problems, resulting in inaccurate data extraction; at the same time, the regularized process has limited adaptability to domain knowledge and lacks flexibility. It is often only applicable to certain specific fields or specific formats of literature, and cannot handle cross-domain or emerging research content; and the ability to extract unstructured data is insufficient, resulting in the frequent loss of key experimental information.

[0004] In summary, the existing technology has low efficiency, insufficient accuracy and poor flexibility in data extraction when processing a large amount of literature, and cannot meet cross-domain needs. Summary of the invention

[0005] Based on this, it is necessary to provide a method and device for extracting experimental data from scientific literature based on a large language model to address the above technical problems, which can automatically, efficiently and accurately extract experimental data from scientific literature.

[0006] A method for extracting experimental data from scientific literature based on a large language model, the method comprising:

[0007] Use script code to obtain relevant literature in the target field based on preset keywords;

[0008] Analyze relevant literature based on preset prompt words through the trained large language model, and extract experimental data from the relevant literature;

[0009] Perform data cleaning and verification on the extracted experimental data to obtain cleaned data;

[0010] Export the cleaned data in the preset export format.

[0011] In one embodiment, relevant documents in the target field are obtained according to preset keywords through script code, including: searching in a preset scientific literature database according to preset keywords through script code to determine relevant documents in the target field; downloading and storing relevant documents in local storage or cloud server through script code.

[0012] In one of the embodiments, relevant documents are parsed according to preset prompt words through a trained large language model, and experimental data in the relevant documents are extracted, including: converting the acquired relevant documents into a preprocessed format through a format conversion tool; parsing the relevant documents in the preprocessed format according to preset prompt words through a trained large language model, and extracting experimental data in the relevant documents.

[0013] In one embodiment, before parsing relevant documents according to preset prompt words through a trained large language model and extracting experimental data from the relevant documents, the method also includes: obtaining preset prompt words input by a user; and before obtaining relevant documents in the target field according to preset keywords through a script code, the method also includes: obtaining preset keywords input by a user.

[0014] In one embodiment, the extracted experimental data is cleaned and verified to obtain cleaned data, including: performing data cleaning and verification on the extracted experimental data according to preset rules to obtain cleaned data.

[0015] In one embodiment, the method further includes: when performing data cleaning and verification on the extracted experimental data, if abnormal data is found, re-entering the step of parsing relevant literature according to preset prompt words through the trained large language model and extracting experimental data from the relevant literature, or outputting and displaying the abnormal data for manual verification.

[0016] In one of the embodiments, exporting the cleaned data in a preset export format includes: processing the cleaned data to generate a data result in a preset export format, where the data result is a structured data file; and exporting the data result.

[0017] A scientific literature experimental data extraction device based on a large language model, the device comprising:

[0018] The document acquisition module is used to acquire relevant documents in the target field according to preset keywords through script code;

[0019] The data extraction module is used to parse relevant literature according to preset prompt words through the trained large language model and extract experimental data from the relevant literature;

[0020] The data cleaning and verification module is used to clean and verify the extracted experimental data to obtain the cleaned data;

[0021] The result export module exports the cleaned data in a preset export format.

[0022] A computer device comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the following steps when executing the computer program: obtaining relevant documents in a target field according to preset keywords through script code; parsing relevant documents according to preset prompt words through a trained large language model, and extracting experimental data from the relevant documents; performing data cleaning and verification on the extracted experimental data to obtain cleaned data; and exporting the cleaned data in a preset export format.

[0023] A computer-readable storage medium stores a computer program, which implements the following steps when executed by a processor: obtaining relevant documents in a target field according to preset keywords through script code; parsing relevant documents according to preset prompt words through a trained large language model, and extracting experimental data from the relevant documents; performing data cleaning and verification on the extracted experimental data to obtain cleaned data; and exporting the cleaned data in a preset export format.

[0024] The above-mentioned scientific literature experimental data extraction method, device, computer equipment and storage medium based on the large language model can obtain relevant literature in the target field according to preset keywords through script code, that is, use script code to automatically search and download relevant literature in the target field, which can effectively save time for manual screening and data acquisition, and improve the efficiency of obtaining relevant literature; then use the trained large language model to parse the relevant literature according to preset prompt words, and extract experimental data from the relevant literature. In this step, the large language model automatically parses the literature, which can accurately extract experimental data, conditions and results, avoid errors in manual extraction, and users can design preset prompt words according to field requirements to ensure that the large language model can efficiently extract information related to the target data; the extracted experimental data is cleaned and verified to obtain cleaned data to ensure data accuracy and consistency; the cleaned data is exported in a preset export format to obtain data results in a unified format, which is convenient for subsequent analysis and application.

[0025] The above method forms a complete set of automated experimental data extraction processes by combining automated script code, large language model analysis, data cleaning and verification, and standardized data output, which greatly improves the efficiency and accuracy of data processing. It can automatically, efficiently and accurately extract experimental data from scientific literature to meet the needs of scientific researchers to quickly obtain key information and improve the overall efficiency of literature analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 A schematic diagram of a flow chart of a method for extracting experimental data from scientific literature based on a large language model in one embodiment;

[0027] Figure 2 is a structural block diagram of a scientific literature experimental data extraction device based on a large language model in one embodiment;

[0028] Figure 3 FIG. 4 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0030] In one embodiment, Figure 1 As shown, a method for extracting experimental data from scientific literature based on a large language model is provided. This embodiment takes the method applied to a terminal as an example for explanation. It can be understood that the method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers, and portable wearable devices, and the server can be implemented as an independent server or a server cluster consisting of multiple servers.

[0031] In this embodiment, the method may include the following steps:

[0032] Step 102, obtaining relevant documents in the target field according to preset keywords through script code.

[0033] Among them, the script code is a small program used to implement specific automation tasks, usually written in programming languages ​​such as Python and Shell. The script code in the above steps is used to automatically search and download relevant literature in the target field; the keywords are pre-set by the user and can be the name of the target field, etc.; the target field is the subject field corresponding to the preset keywords, such as catalyst research, electrochemical analysis, etc.; the relevant literature is the scientific literature related to the preset keywords in the target field, such as scientific papers, which are the literature carriers that record scientific research results and experimental processes, usually including background introduction, experimental methods, results and discussion, etc.

[0034] Specifically, the terminal can use script code to perform automatic retrieval in a preset scientific literature database according to preset keywords, determine relevant literature in the target field, and download relevant literature; in specific implementation, the scientific literature database can be a currently public academic database, such as PubMed, Web of Science, etc.; the terminal can use script code to perform keyword retrieval in one or more scientific literature databases; depending on the preset keywords, the number of retrieved relevant documents may be one or more. When the number of relevant documents is large, multiple relevant documents can be downloaded in batches through script code, and the downloaded relevant documents can be stored; further, in order to facilitate the parsing of the large language model, all the acquired relevant documents can be converted into a unified preprocessing format, such as PDF format.

[0035] It can be seen that step 102 realizes automatic retrieval and downloading of scientific literature through script code, realizes automation of literature retrieval and acquisition, and improves the efficiency of literature retrieval and acquisition.

[0036] Step 104: parse the relevant documents according to the preset prompt words through the trained large language model, and extract the experimental data in the relevant documents.

[0037] Among them, the large language model is a large-scale natural language processing model based on deep learning technology, which can understand and generate natural language text. Its functions include text generation, information extraction, language translation, etc., such as ChatGPT, GPT-4, etc.; the trained large language model has the ability to automatically parse relevant literature and extract experimental data from it; prompt words can be set by the user according to the research focus of the target field, which is used to guide the large language model to extract experimental data from relevant literature. Each target field has its corresponding prompt words, and effective prompt words may include extraction target information and key problem information. Extraction target information is used to clarify the extraction target, such as experimental conditions, catalytic activity data, sample preparation methods, etc.; key problem information is used to define key problems, which can guide the model to focus on specific paragraphs, tables or numerical information.

[0038] Specifically, the terminal can call a trained large language model and import all the acquired relevant documents into the trained large language model. The large language model can parse the content of the relevant documents based on preset prompt words, and automatically locate the position of the experimental data in the relevant documents, such as the paragraph or table where the experimental data is located, and then extract the experimental data from the relevant documents; experimental data is quantitative or qualitative information obtained through experiments in the scientific research process, including but not limited to tabular data, experimental parameters, result descriptions, etc. The experimental data in the above steps refers to the data related to the preset prompt words in each relevant document, that is, the part of the experimental data required by the user; in specific implementation, the number of relevant documents is usually more than one, and the trained large language model can parse the relevant documents one by one, or parse multiple relevant documents at the same time.

[0039] Step 106, performing data cleaning and verification on the extracted experimental data to obtain cleaned data.

[0040] Specifically, the extracted experimental data are cleaned and verified according to preset rules and logics to correct and verify the extracted data to obtain cleaned data.

[0041] The above step 106 can check the validity of the extracted data and complete the preliminary standardization process by cleaning and verifying the extracted experimental data, thereby ensuring the integrity and consistency of the cleaned data.

[0042] Step 108: Export the cleaned data in a preset export format.

[0043] The preset export format may be a commonly used structured data file format.

[0044] Specifically, the terminal can organize the cleaned data according to the preset export format, generate data results that conform to the preset export format, and then export the data results, which are structured data files, such as CSV files, etc., so that scientific researchers can perform subsequent processing and analysis. In specific implementation, the cleaned data can be organized according to field names, unit standardization, etc.; finally, the data results can be used to construct a data set, or the data results can be stored. Among them, CSV (Comma-Separated Values, comma-separated value file) is a file format that stores tabular data in plain text, which is easy to import into various data analysis software or databases.

[0045] In specific implementation, the terminal or server is used to run the script code, obtain prompt words, and call the trained large language model, and the memory of the terminal or server is used to store the downloaded scientific literature data (such as PDF format files) and the generated structured data files (such as CSV files).

[0046] In the above-mentioned scientific literature experimental data extraction method based on the large language model, the script code can be used to obtain relevant literature in the target field according to preset keywords, that is, the script code is used to automatically search for and download relevant literature in the target field, which can effectively save the time of manual screening and data acquisition, and improve the efficiency of obtaining relevant literature; then the trained large language model is used to parse the relevant literature according to the preset prompt words, and the experimental data in the relevant literature is extracted. In this step, the large language model is used to automatically parse the literature, which can accurately extract experimental data, conditions and results, avoid errors in manual extraction, and users can design preset prompt words according to field requirements to ensure that the large language model can efficiently extract information related to the target data; the extracted experimental data is cleaned and verified to obtain cleaned data to ensure data accuracy and consistency; the cleaned data is exported in a preset export format to obtain data results in a unified format, which is convenient for subsequent analysis and application. In summary, the present invention forms a complete set of automated experimental data extraction processes by combining automated script code, large language model analysis, data cleaning and verification, and standardized data output, which greatly improves the efficiency and accuracy of data processing. It can automatically, efficiently and accurately extract experimental data from scientific literature to meet the needs of scientific researchers to quickly obtain key information and improve the overall efficiency of literature analysis.

[0047] In one embodiment, step 102 includes: searching a preset scientific literature database according to preset keywords through a script code to determine relevant literature in the target field;

[0048] Download relevant documents and store them in local storage or cloud server through script code.

[0049] In the above embodiment, using script code to automatically search for and download documents in the target field can effectively save time for manual screening and data acquisition, and improve the efficiency of acquiring relevant documents.

[0050] In one embodiment, step 104 includes: converting the acquired relevant documents into a pre-processing format using a format conversion tool;

[0051] The trained large language model is used to parse the relevant literature in the preprocessing format according to the preset prompt words, and the experimental data in the relevant literature is extracted.

[0052] The preprocessing format can be a PDF file, word document, or other formats. Since the PDF format can retain the typesetting and format of the original document, it ensures that the format and layout of the document content remain consistent when viewed on different devices. In actual application scenarios, in order to facilitate the parsing and data extraction of the trained large language model, the preprocessing format is preferably a PDF file, and the format conversion tool is a PDF processing program.

[0053] The above embodiment converts the relevant documents into a preprocessing format before the large language model is parsed, which makes it easier for the large language model to perform parsing and data extraction, and helps to improve the parsing and extraction efficiency. At the same time, it utilizes the large language model's efficient parsing capability for complex semantics, which can effectively avoid errors caused by human understanding bias and effectively improve the accuracy of data extraction.

[0054] In one embodiment, before step 104, the method further includes: obtaining preset prompt words input by the user; and before obtaining relevant documents in the target field according to the preset keywords through the script code, the method further includes: obtaining preset keywords input by the user.

[0055] In the above embodiments, both keywords and prompt words can be designed by users themselves, which can achieve the customization of scientific field knowledge and the high flexibility of large language model extraction structure, making this method suitable for various scientific research fields and can be widely used in research work in multiple disciplines such as materials science, chemistry, and life sciences.

[0056] In one embodiment, step 106 includes: performing data cleaning and verification on the extracted experimental data according to preset rules to obtain cleaned data.

[0057] Among them, the preset rules can be pre-set by the user according to needs, and can be rule restrictions on the numerical range and unit specifications of experimental data; in specific implementation, the preset rules may include data verification rules and data cleaning rules.

[0058] In this embodiment, the extracted experimental data are cleaned and verified according to preset rules, which effectively avoids errors caused by format confusion and ensures the integrity and consistency of the cleaned data, thereby improving the availability of the extracted results.

[0059] In one embodiment, the method further comprises:

[0060] When the extracted experimental data are cleaned and verified, if abnormal data is found, the process reenters step 104, or the abnormal data is output and displayed for manual verification.

[0061] In the above embodiment, during data verification and cleaning, if abnormal data is found, it can be corrected by repeated extraction or manual verification to reduce the error rate.

[0062] In one embodiment, step 108 includes: processing the cleaned data to generate data results in a preset export format, where the data results are structured data files; and exporting the data results.

[0063] Among them, the preset export format can be a commonly used structured data file format, such as a table in a relational database, an XML document, a CSV file, JSON data, etc.; in a specific implementation, the preset export format is preferably a CSV file, which is convenient for analysis or importing into a database.

[0064] The above-mentioned embodiment realizes the standardization of the output data results by generating data results in a preset export format, ensuring that all data conform to the unified format requirements, so that users do not need to perform additional data sorting and can directly use the exported results for subsequent analysis, thereby greatly saving time and energy.

[0065] In each of the above embodiments, before step 102, the method further includes a model training step, that is, training a preset large language model in advance to obtain a trained large language model.

[0066] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0067] In one embodiment, Figure 2 As shown, a scientific literature experimental data extraction device based on a large language model is provided, comprising: a literature acquisition module 202, a data extraction module 204, a data cleaning and verification module 206 and a result export module 208, wherein:

[0068] The document acquisition module 202 is used to acquire relevant documents in the target field according to preset keywords through script code;

[0069] The data extraction module 204 is used to parse the relevant documents according to the preset prompt words through the trained large language model, and extract the experimental data in the relevant documents;

[0070] The data cleaning and verification module 206 is used to clean and verify the extracted experimental data to obtain cleaned data;

[0071] The result export module 208 exports the cleaned data in a preset export format.

[0072] In one embodiment, the document acquisition module 202 includes: a document retrieval unit, which is used to search in a preset scientific document database according to preset keywords through script code to determine relevant documents in the target field; and a document downloading unit, which is used to download and store relevant documents to a local storage or a cloud server through script code.

[0073] In one embodiment, the data extraction module 204 includes: a format conversion unit, which is used to convert the acquired relevant documents into a preprocessing format through a format conversion tool; a document parsing unit, which is used to parse the relevant documents in the preprocessing format according to preset prompt words through a trained large language model; and a data extraction unit, which is used to extract experimental data from the relevant documents.

[0074] In one embodiment, the device further includes: a prompt word acquisition module, used to acquire a preset prompt word input by a user; and a keyword acquisition module, used to acquire a preset keyword input by a user.

[0075] In one embodiment, the data cleaning and verification module 206 is specifically used to perform data cleaning and verification on the extracted experimental data according to preset rules to obtain cleaned data.

[0076] In one embodiment, the device also includes: an abnormal data processing module, which is used to re-enter the step of parsing relevant documents according to preset prompt words through a trained large language model and extracting experimental data from the relevant documents if abnormal data is found during data cleaning and verification of the extracted experimental data, or output and display the abnormal data for manual verification.

[0077] In one embodiment, the result export module 208 includes: a data processing unit, which is used to process the cleaned data and generate data results in a preset export format, where the data results are structured data files; and a data export unit, which is used to export the data results.

[0078] For the specific definition of the scientific literature experimental data extraction device based on the large language model, please refer to the definition of the scientific literature experimental data extraction method based on the large language model above, which will not be repeated here. Each module in the above-mentioned scientific literature experimental data extraction device based on the large language model can be implemented in whole or in part by software, hardware and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.

[0079] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 3As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for extracting experimental data from scientific literature based on a large language model is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a key, trackball or touchpad set on the computer device housing, or an external keyboard, touchpad or mouse, etc.

[0080] Those skilled in the art will understand that Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0081] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: obtaining relevant documents in a target field according to preset keywords through a script code; parsing the relevant documents according to preset prompt words through a trained large language model, and extracting experimental data from the relevant documents; performing data cleaning and verification on the extracted experimental data to obtain cleaned data; and exporting the cleaned data in a preset export format.

[0082] In one embodiment, when the processor executes the computer program, the following steps are also implemented: searching in a preset scientific literature database according to preset keywords through script code to determine relevant literature in the target field; downloading and storing relevant literature to local storage or cloud server through script code.

[0083] In one embodiment, when the processor executes the computer program, the following steps are also implemented: converting the acquired relevant documents into a preprocessing format through a format conversion tool; parsing the relevant documents in the preprocessing format according to preset prompt words through a trained large language model, and extracting experimental data from the relevant documents.

[0084] In one embodiment, when the processor executes the computer program, the following steps are also implemented: obtaining a preset prompt word input by the user; and obtaining a preset keyword input by the user.

[0085] In one embodiment, when the processor executes the computer program, the following steps are also implemented: data cleaning and verification are performed on the extracted experimental data according to preset rules to obtain cleaned data.

[0086] In one embodiment, when the processor executes the computer program, the following steps are also implemented: when performing data cleaning and verification on the extracted experimental data, if abnormal data is found, re-entering the step of parsing relevant literature according to preset prompt words through the trained large language model and extracting experimental data from the relevant literature, or outputting and displaying the abnormal data for manual verification.

[0087] In one embodiment, when the processor executes the computer program, the following steps are also implemented: processing the cleaned data to generate data results in a preset export format, where the data results are structured data files; and exporting the data results.

[0088] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining relevant documents in a target field according to preset keywords through a script code; parsing the relevant documents according to preset prompt words through a trained large language model, and extracting experimental data from the relevant documents; performing data cleaning and verification on the extracted experimental data to obtain cleaned data; and exporting the cleaned data in a preset export format.

[0089] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: searching in a preset scientific literature database according to preset keywords through script code to determine relevant literature in the target field; downloading and storing relevant literature to local storage or cloud server through script code.

[0090] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: converting the acquired relevant documents into a preprocessing format through a format conversion tool; parsing the relevant documents in the preprocessing format according to preset prompt words through a trained large language model, and extracting experimental data from the relevant documents.

[0091] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: obtaining a preset prompt word input by the user; obtaining a preset keyword input by the user.

[0092] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: data cleaning and verification are performed on the extracted experimental data according to preset rules to obtain cleaned data.

[0093] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: when the extracted experimental data is cleaned and verified, if abnormal data is found, the step of re-entering the step of parsing relevant literature according to preset prompt words through the trained large language model and extracting experimental data from the relevant literature, or outputting and displaying the abnormal data for manual verification.

[0094] In one embodiment, when the computer program is executed by the processor, the following steps are also implemented: processing the cleaned data to generate data results in a preset export format, where the data results are structured data files; and exporting the data results.

[0095] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0096] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0097] The above-mentioned embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the attached claims.

Claims

1. A method for extracting experimental data from scientific literature based on a large language model, the method comprising: Use script code to obtain relevant literature in the target field based on preset keywords; Parsing the relevant documents according to preset prompt words through a trained large language model, and extracting experimental data from the relevant documents; Performing data cleaning and verification on the extracted experimental data to obtain cleaned data; The cleaned data is exported in a preset export format.

2. The method according to claim 1, characterized in that The script code is used to obtain relevant documents in the target field according to preset keywords, including: Through script code, search the preset scientific literature database according to preset keywords to identify relevant literature in the target field; The relevant documents are downloaded and stored in a local storage or a cloud server through the script code.

3. The method according to claim 2, characterized in that The method of parsing the relevant documents according to preset prompt words through the trained large language model and extracting experimental data from the relevant documents includes: Convert the acquired relevant documents into a preprocessing format using a format conversion tool; The relevant documents in the preprocessing format are parsed according to preset prompt words through the trained large language model, and the experimental data in the relevant documents are extracted.

4. The method according to claim 3, characterized in that Before parsing the relevant documents according to the preset prompt words through the trained large language model and extracting the experimental data in the relevant documents, the method further includes: obtaining the preset prompt words input by the user; Furthermore, before obtaining relevant documents in the target field according to preset keywords through the script code, the method further includes: obtaining preset keywords input by the user.

5. The method according to claim 1, characterized in that: The step of performing data cleaning and verification on the extracted experimental data to obtain cleaned data includes: The extracted experimental data are cleaned and verified according to preset rules to obtain cleaned data.

6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: When the extracted experimental data is cleaned and verified, if abnormal data is found, the step of parsing the relevant documents according to preset prompt words through the trained large language model and extracting the experimental data in the relevant documents is re-entered, or the abnormal data is output and displayed for manual verification.

7. The method according to any one of claims 1 to 5, characterized in that: The step of exporting the cleaned data in a preset export format includes: Processing the cleaned data to generate a data result in a preset export format, wherein the data result is a structured data file; The data results are exported.

8. A scientific literature experimental data extraction device based on a large language model, characterized in that: The device comprises: The document acquisition module is used to acquire relevant documents in the target field according to preset keywords through script code; A data extraction module, used to parse the relevant documents according to preset prompt words through a trained large language model, and extract experimental data from the relevant documents; A data cleaning and verification module is used to clean and verify the extracted experimental data to obtain cleaned data; The result export module exports the cleaned data in a preset export format.

9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Method for automatically extracting ecological parameters and driving factors of large language model

    CN120910562A

  • Experimental scheme automatic generation system based on knowledge graph and large language model

    CN121638472A

  • Corrosion inhibitor field literature text mining and data cleaning method based on large language model

    CN121835660A