Scientific literature data extraction method and device, computer equipment and storage medium
By obtaining target literature from the literature database and extracting and analyzing experimental data using machine learning models, the problems of low efficiency and poor accuracy of literature data extraction in the existing technology are solved, and efficient and accurate data extraction and future experimental results prediction are achieved.
Patent Information
- Application Number
- CN202510197399.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
Existing literature data extraction methods are inefficient and have poor accuracy, making it difficult to meet the needs of different fields or emerging research fields.
By obtaining the target literature from the preset literature database, extracting experimental data using the trained machine learning model, and analyzing it through regression algorithms, the predicted experimental results are obtained. At the same time, data cleaning and verification of experimental data is carried out to ensure the accuracy and consistency of the extracted data.
It has achieved efficient and accurate extraction of experimental data, predicted future experimental results, improved the efficiency and accuracy of data extraction, and significantly improved the adaptability.
Smart Images

Figure CN120124009A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of scientific literature data processing, and particularly to a method, device, computer device, and storage medium for extracting scientific literature data. Background Art
[0002] With the continuous development of science and technology, the number of literature has increased sharply, and scientific research personnel are facing the pressure of screening and analyzing a large amount of literature data. The existing literature data extraction methods mainly rely on manual screening or rule-based automated processing.
[0003] However, manual processing of literature data not only consumes a lot of time, but also easily misses important information, resulting in low efficiency and high error rates; while rule-based automated processing methods are difficult to process complex natural language texts, and the extraction results are prone to errors, leading to poor accuracy of data extraction; at the same time, the above two methods have poor adaptability to different fields or emerging research fields, and it is difficult to meet the literature data extraction needs of different fields.
[0004] In summary, the existing literature data extraction methods have problems such as low efficiency, poor accuracy, and poor flexibility. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a method, device, computer device, and storage medium for extracting scientific literature data, which can efficiently and accurately extract experimental data and predict future experimental results.
[0006] A method for extracting scientific literature data, the method comprising:
[0007] Obtaining a target document from a preset literature database;
[0008] Extracting experimental data from the target document through a trained machine learning model;
[0009] Analyzing the experimental data through a regression algorithm to obtain a predicted experimental result;
[0010] Performing data cleaning and verification on the experimental data to obtain valid data;
[0011] Outputting the valid data and the predicted experimental result in a preset format to obtain a target file.
[0012] In one embodiment, obtaining a target document from a preset literature database includes: retrieving relevant documents in the target field in the preset literature database through script code; preprocessing the relevant documents through natural language processing technology to obtain a target document, and the target document is in a format that can be processed by a machine learning model.
[0013] In one embodiment, experimental data is extracted from a target document by a trained machine learning model, including: performing semantic analysis on the target document by the trained machine learning model and extracting the experimental data in the target document; the experimental data includes at least one of experimental conditions, experimental parameters, and experimental results.
[0014] In one embodiment, before extracting experimental data from a target document by a trained machine learning model, the method further includes: obtaining a training data set, the training data set containing multiple labeled training data; training a preset machine learning model based on the training data set to obtain a trained machine learning model.
[0015] In one embodiment, the experimental data is analyzed by a regression algorithm to obtain a predicted experimental result, including: inputting the experimental data into the regression algorithm to construct a regression prediction model; predicting based on the experimental data and historical experimental data through the regression prediction model to obtain the predicted experimental result corresponding to the experimental data.
[0016] In one embodiment, the experimental data is cleaned and verified to obtain valid data, including: cleaning the experimental data; and performing data verification on the experimental data through the regression prediction model, and identifying and correcting abnormal data points in the experimental data through an anomaly detection algorithm.
[0017] In one embodiment, the method further includes:
[0018] Performing visualization processing on the target file to generate a visualization result and displaying the visualization result.
[0019] A scientific literature data extraction device, the device includes:
[0020] A literature acquisition module for acquiring a target document from a preset literature database;
[0021] A data extraction module for extracting experimental data from a target document by a trained machine learning model;
[0022] An experimental result prediction module for analyzing the experimental data by a regression algorithm to obtain a predicted experimental result;
[0023] A data cleaning and verification module for cleaning and verifying the experimental data to obtain valid data;
[0024] A target file output module for outputting the valid data and the predicted experimental result in a preset format to obtain a target file.
[0025] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: obtaining a target document from a preset literature database; extracting experimental data from the target document through a trained machine learning model; analyzing the experimental data through a regression algorithm to obtain a predicted experimental result; performing data cleaning and verification on the experimental data to obtain valid data; and outputting the valid data and the predicted experimental result in a preset format to obtain a target file.
[0026] A computer-readable storage medium stores a computer program thereon. When the computer program is executed by a processor, the following steps are implemented: obtaining a target document from a preset literature database; extracting experimental data from the target document through a trained machine learning model; analyzing the experimental data through a regression algorithm to obtain a predicted experimental result; performing data cleaning and verification on the experimental data to obtain valid data; and outputting the valid data and the predicted experimental result in a preset format to obtain a target file.
[0027] For the above scientific literature data extraction method, device, computer device, and storage medium, by obtaining a target document from a preset literature database, automated retrieval and acquisition of the target document can be achieved, improving the efficiency of literature retrieval and acquisition. Then, by extracting experimental data from the target document through a trained machine learning model, errors in manual extraction can be avoided, improving the accuracy and efficiency of data extraction. By analyzing the experimental data through a regression algorithm to obtain a predicted experimental result, experimental trend prediction and data optimization support can be provided for scientific researchers. Performing data cleaning and verification on the experimental data to obtain valid data ensures the accuracy and consistency of the extracted data. Outputting the valid data and the predicted experimental result in a preset format to obtain a target file facilitates storage and subsequent analysis.
[0028] In summary, by combining machine learning and a regression prediction model, the above method can automatically extract experimental data from a large amount of scientific literature, perform regression analysis on the extracted experimental data, and predict future experimental results. It can not only improve the efficiency, accuracy, and adaptability of data extraction, but also provide predicted experimental results, thus providing a more efficient and accurate data extraction and analysis tool for scientific researchers. Description of the Drawings
[0029] Figure 1 It is a schematic flowchart of the scientific literature data extraction method in an embodiment;
[0030] Figure 2 It is a schematic flowchart of the scientific literature data extraction method in another embodiment;
[0031] Figure 3Structural block diagram of a scientific literature data extraction device in an embodiment;
[0032] Figure 4 Structural block diagram of a scientific literature data extraction device in another embodiment;
[0033] Figure 5 Internal structure diagram of a computer device in an embodiment. Detailed implementation manners
[0034] To make the objectives, technical solutions and advantages of this application clearer, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application, but not to limit this application.
[0035] In one embodiment, as Figure 1-2 shown, a scientific literature data extraction method is provided. In this embodiment, this method is described by taking its application to a terminal as an example. It can be understood that this method can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. Among them, the terminal can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers and portable wearable devices, and the server can be implemented by an independent server or a server cluster composed of multiple servers.
[0036] In this embodiment, the method may include the following steps:
[0037] Step 102, obtain a target document from a preset document database.
[0038] Among them, according to the actual needs of the user, the number of document databases can be one or more. To increase the number of target documents that can be obtained, generally multiple document databases are preset, such as databases like PubMed, IEEE Xplore, Web of Science, etc., that is, the terminal supports multiple document database interfaces; the target document can be a scientific document related to the target field required by the user, and its format is a format suitable for input into a machine learning model, such as JSON, CSV, etc.
[0039] Specifically, the terminal automatically obtains the target document from one or more document databases; it should be noted that the documents in the document database may be in a format that can be processed by the machine learning model, or may be a format that is not suitable for input into the machine learning model. For the former, the terminal can directly obtain the target document from the document database; but for the latter, the terminal needs to first convert the format of the retrieved document to obtain the target document.
[0040] The above-mentioned step 102 can achieve the automated retrieval and acquisition of target documents, improving the efficiency of document retrieval and acquisition.
[0041] Step 104: Extract experimental data from the target document through the trained machine learning model.
[0042] Among them, the machine learning model is an algorithm or mathematical function that can transform input data into specific outputs. They are the core components of machine learning. Through training data, the machine learning model can automatically learn and improve itself, making it more accurate and effective when processing new data. In step 104, the trained machine learning model can output the extracted experimental data according to the input target document.
[0043] Specifically, the terminal calls the trained deep learning model to automatically analyze and extract data from the content of the target document, obtaining the experimental data in the target document; the experimental data can include key information such as experimental conditions, experimental parameters, and experimental results. In actual application scenarios, the number of target documents is usually more than one. The machine learning model can either extract experimental data one by one or extract experimental data from multiple target documents simultaneously.
[0044] The above-mentioned step 104 uses the trained machine learning model to extract experimental data, which can avoid the errors of manual extraction and improve the accuracy and efficiency of data extraction.
[0045] Step 106: Analyze the experimental data through a regression algorithm to obtain the predicted experimental results.
[0046] Among them, the regression algorithm is a machine learning algorithm used to predict numerical values based on input data and can be used for accurate prediction of new, unseen data; specifically in implementation, the above-mentioned regression algorithm can adopt linear regression, support vector regression, random forest regression, decision tree regression, etc.
[0047] Specifically, the terminal uses the regression algorithm to analyze and predict the experimental data extracted in step 104 to obtain the predicted experimental results, which can provide experimental trend prediction and data optimization support for scientific researchers.
[0048] Step 108: Clean and validate the experimental data to obtain valid data.
[0049] Among them, data cleaning can remove redundant, repeated, or invalid data in the experimental data and ensure the unity and consistency of the experimental data; data validation can be achieved by using the regression algorithm and anomaly detection algorithm. The regression algorithm is used to verify the experimental data, and the anomaly detection algorithm is used to identify and correct abnormal data points in the experimental data; valid data is the experimental data after data cleaning and validation.
[0050] Specifically, the terminal performs data cleaning, duplicate removal, and outlier detection on the experimental data, which can ensure the quality of the extracted data.
[0051] Step 110: Output the valid data and the predicted experimental results in a preset format to obtain a target file.
[0052] Among them, the preset format is a structured format, such as CSV, Excel, etc.; the target file is a data file in a structured format.
[0053] Specifically, the terminal outputs the valid data and the predicted experimental results as a structured data file, which is convenient for storage and subsequent analysis.
[0054] In the above scientific literature data extraction method, by obtaining the target literature from a preset literature database, the automatic retrieval and acquisition of the target literature can be realized, improving the efficiency of literature retrieval and acquisition; then, by using the trained machine learning model to extract experimental data from the target literature, the error of manual extraction can be avoided, improving the accuracy and efficiency of data extraction; by using the regression algorithm to analyze the experimental data to obtain the predicted experimental results, it can provide experimental trend prediction and data optimization support for scientific researchers; by performing data cleaning and verification on the experimental data to obtain valid data, ensuring the accuracy and consistency of the extracted data; by outputting the valid data and the predicted experimental results in a preset format to obtain a target file, which is convenient for storage and subsequent analysis. In summary, by combining machine learning and regression prediction models, the present invention can automatically extract experimental data from a large amount of scientific literature, perform regression analysis on the extracted experimental data, and predict future experimental results, which can not only improve the efficiency, accuracy, and adaptability of data extraction, but also provide predicted experimental results, thus providing a more efficient and accurate data extraction and analysis tool for scientific researchers.
[0055] In one embodiment, step 102 includes:
[0056] Retrieve relevant literature in the target field in the preset literature database through script code;
[0057] Preprocess the relevant literature through natural language processing technology to obtain the target literature, and the target literature is in a format that can be processed by the machine learning model.
[0058] Among them, the target field can be determined by the preset keywords input by the user. The script code can be an automated script written in the Python language, which is used to implement the functions of literature retrieval and download. Natural Language Processing (NLP) is the science and applied technology of using computers to process, understand, and generate human languages. It can use computers to replace manual labor in processing large-scale natural language information. Here, it is used for the preprocessing of relevant literature, that is, automatically converting relevant literature into a format suitable for input into a machine learning model, such as JSON, CSV, etc., that is, the target literature.
[0059] Specifically, the terminal determines the target field according to the preset keywords input by the user, and uses the automated script code to retrieve and download relevant literature in the preset various literature databases; then, through natural language processing technology, the relevant literature is converted into the target literature in a format that can be processed by the machine learning model.
[0060] In the above embodiment, by using the script code to retrieve and obtain relevant literature, while realizing the automation of literature retrieval, it also improves the efficiency and accuracy of literature retrieval; then, using natural language processing technology for preprocessing to obtain the target literature that can be input into the machine learning model is more convenient for the machine learning model to extract experimental data and helps improve the data extraction efficiency.
[0061] In one embodiment, step 104 includes: performing semantic analysis on the target literature through the trained machine learning model, and extracting the experimental data in the target literature; the experimental data includes at least one of experimental conditions, experimental parameters, and experimental results.
[0062] Among them, the machine learning model can adopt a deep learning model with the ability to automatically extract features, so as to extract experimental data more efficiently and accurately; the deep learning model can adopt common deep learning models in the prior art, such as BERT, GPT, T5, etc.
[0063] In one embodiment, before step 104, the method further includes:
[0064] Obtaining a training data set, where the training data set contains multiple labeled training data;
[0065] Based on the training data set, training the preset machine learning model to obtain the trained machine learning model.
[0066] Among them, the training data set contains a large number of labeled training data, and these training data can come from the same field or different fields.
[0067] In specific implementation, by adding training data from different fields to the training dataset and further training the machine learning model, the trained machine learning model can process documents from different fields and complex natural language descriptions during the extraction of experimental data, so as to more accurately extract experimental data, meet the requirements for extracting literature data in different fields, and improve flexibility.
[0068] In one embodiment, step 106 includes:
[0069] Input the experimental data into a regression algorithm to construct a regression prediction model;
[0070] Based on the experimental data and historical experimental data, make a prediction through the regression prediction model to obtain the predicted experimental result corresponding to the experimental data.
[0071] Among them, the historical experimental data is the experimental data extracted previously. Specifically, the terminal inputs the experimental data into a preset regression algorithm to construct a regression prediction model; and extracts the experimental conditions in the experimental data. The regression prediction model makes a prediction based on the experimental conditions and historical experimental data to obtain the predicted experimental result corresponding to the experimental conditions; based on the predicted experimental result, scientific researchers can optimize the experimental design or predict the future experimental result.
[0072] In one embodiment, step 108 includes:
[0073] Perform data cleaning on the experimental data;
[0074] And perform data verification on the experimental data through the regression prediction model, and identify and correct abnormal data points in the experimental data through an anomaly detection algorithm.
[0075] Among them, the abnormal data points can be error values or inconsistencies in the experimental data.
[0076] Through the above-mentioned embodiment, by performing data cleaning on the experimental data, redundant, repeated or invalid data in the experimental data can be removed, and the unity and consistency of the experimental data can be ensured; during the data cleaning process, by performing data verification on the experimental data through the regression prediction model, more accurate verification can be achieved. At the same time, by identifying and correcting abnormal data points in the experimental data through the anomaly detection algorithm, high-quality data can be obtained.
[0077] In one embodiment, as Figure 2 shown, the method further includes:
[0078] Step 112, perform visualization processing on the target file to generate a visualization result and display the visualization result.
[0079] Among them, the visualization tool can be Matplotlib, Seaborn, etc.; the visualization result can be in the form of charts, such as a prediction trend chart, a regression curve chart, etc.
[0080] Specifically, the terminal performs visualization processing on the target file through a visualization tool to generate visualization results such as a prediction trend chart and a regression curve chart, which can help scientific researchers quickly understand the data and its trends.
[0081] In the above embodiments, through the visualization of the target file, a trend chart of valid data and predicted experimental results is displayed, which is convenient for scientific researchers to more intuitively understand the data and predicted experimental results.
[0082] It should be understood that although Figure 1-2 the steps in the flowchart of Figure 1-2 are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover,
[0083] In one embodiment, as Figure 3 shown, a scientific literature data extraction device is provided, including: a literature acquisition module 202, a data extraction module 204, an experimental result prediction module 206, a data cleaning and verification module 208, and a target file output module 210, where:
[0084] The literature acquisition module 202 is used to acquire target literature from a preset literature database;
[0085] The data extraction module 204 is used to extract experimental data from the target literature through a trained machine learning model;
[0086] The experimental result prediction module 206 is used to analyze the experimental data through a regression algorithm to obtain predicted experimental results;
[0087] The data cleaning and verification module 208 is used to perform data cleaning and verification on the experimental data to obtain valid data;
[0088] The target file output module 210 is used to output the valid data and the predicted experimental results in a preset format to obtain a target file.
[0089] In one embodiment, the literature acquisition module 202 includes: a retrieval unit for retrieving relevant literature in the target field from a preset literature database through script code; a preprocessing unit for preprocessing the relevant literature through natural language processing technology to obtain target literature, and the target literature is in a format that can be processed by a machine learning model.
[0090] In one embodiment, the data extraction module 204 is specifically configured to perform semantic analysis on the target literature through a trained machine learning model, and extract experimental data from the target literature; the experimental data includes at least one of experimental conditions, experimental parameters, and experimental results.
[0091] In one embodiment, the device further includes: a model training module for obtaining a training data set, where the training data set contains a plurality of labeled training data; and based on the training data set, training a preset machine learning model to obtain a trained machine learning model.
[0092] In one embodiment, the experimental result prediction module 206 includes: a model construction unit for inputting experimental data into a regression algorithm to construct a regression prediction model; and a result prediction unit for predicting through the regression prediction model based on the experimental data and historical experimental data to obtain a predicted experimental result corresponding to the experimental data.
[0093] In one embodiment, the data cleaning and verification module 208 includes: a data cleaning unit for cleaning the experimental data; and a data verification unit for verifying the experimental data through the regression prediction model, and identifying and correcting abnormal data points in the experimental data through an anomaly detection algorithm.
[0094] In one embodiment, as Figure 4 shown, the device further includes: a visualization module 212 for performing visualization processing on the target file, generating a visualization result, and displaying the visualization result.
[0095] For the specific limitations of the scientific literature data extraction device, reference can be made to the limitations on the scientific literature data extraction method in the above text, which will not be elaborated here. Each module in the above scientific literature data extraction device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above modules.
[0096] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 5As shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, it realizes a method for extracting scientific literature data. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, a touchpad, or a mouse, etc.
[0097] Those skilled in the art can understand that Figure 5 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0098] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are realized: obtaining a target document from a preset document database; extracting experimental data from the target document through a trained machine learning model; analyzing the experimental data through a regression algorithm to obtain a predicted experimental result; performing data cleaning and verification on the experimental data to obtain valid data; and outputting the valid data and the predicted experimental result in a preset format to obtain a target file.
[0099] In one embodiment, when the processor executes the computer program, the following steps are further realized: retrieving relevant documents in the target field in the preset document database through script code; preprocessing the relevant documents through natural language processing technology to obtain a target document, and the target document is in a format that can be processed by the machine learning model.
[0100] In one embodiment, when the processor executes the computer program, the following steps are further realized: performing semantic analysis on the target document through a trained machine learning model, and extracting experimental data from the target document; the experimental data includes at least one of experimental conditions, experimental parameters, and experimental results.
[0101] In one embodiment, when the processor executes the computer program, the following steps are further implemented: obtaining a training data set, where the training data set includes a plurality of labeled training data; based on the training data set, training a preset machine learning model to obtain a trained machine learning model.
[0102] In one embodiment, when the processor executes the computer program, the following steps are further implemented: inputting experimental data into a regression algorithm to construct a regression prediction model; making a prediction based on the experimental data and historical experimental data through the regression prediction model to obtain a predicted experimental result corresponding to the experimental data.
[0103] In one embodiment, when the processor executes the computer program, the following steps are further implemented: performing data cleaning on the experimental data; and performing data verification on the experimental data through the regression prediction model, and identifying and correcting abnormal data points in the experimental data through an anomaly detection algorithm.
[0104] In one embodiment, when the processor executes the computer program, the following steps are further implemented: performing visualization processing on a target file to generate a visualization result, and displaying the visualization result.
[0105] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining a target document from a preset literature database; extracting experimental data from the target document through the trained machine learning model; analyzing the experimental data through a regression algorithm to obtain a predicted experimental result; performing data cleaning and verification on the experimental data to obtain valid data; outputting the valid data and the predicted experimental result in a preset format to obtain a target file.
[0106] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: retrieving relevant documents in a target field in the preset literature database through script code; performing preprocessing on the relevant documents through natural language processing technology to obtain a target document, where the target document is in a format that can be processed by a machine learning model.
[0107] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing semantic analysis on the target document through the trained machine learning model, and extracting experimental data from the target document; the experimental data includes at least one of experimental conditions, experimental parameters, and experimental results.
[0108] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: obtaining a training data set, where the training data set includes a plurality of labeled training data; based on the training data set, training a preset machine learning model to obtain a trained machine learning model.
[0109] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: inputting experimental data into a regression algorithm to construct a regression prediction model; and predicting based on the experimental data and historical experimental data through the regression prediction model to obtain a predicted experimental result corresponding to the experimental data.
[0110] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing data cleaning on the experimental data; and performing data verification on the experimental data through the regression prediction model, and identifying and correcting abnormal data points in the experimental data through an anomaly detection algorithm.
[0111] In one embodiment, when the computer program is executed by a processor, the following steps are further implemented: performing visualization processing on a target file to generate a visualization result and displaying the visualization result.
[0112] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0113] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0114] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A method for extracting scientific literature data, the method comprising: Obtain target documents from the preset document database; Extracting experimental data from the target literature through a trained machine learning model; Analyzing the experimental data by a regression algorithm to obtain a predicted experimental result; Performing data cleaning and verification on the experimental data to obtain valid data; The valid data and the prediction experiment results are output in a preset format to obtain a target file.
2. The method according to claim 1, characterized in that: The step of obtaining the target document from a preset document database includes: Retrieve relevant literature in the target field from the preset literature database through script code; The relevant documents are preprocessed by natural language processing technology to obtain target documents, and the target documents are in a format that can be processed by the machine learning model.
3. The method according to claim 2, characterized in that The step of extracting experimental data from the target document by using a trained machine learning model includes: The target document is semantically analyzed by a trained machine learning model, and experimental data in the target document is extracted; the experimental data includes at least one of experimental conditions, experimental parameters and experimental results.
4. The method according to claim 1, characterized in that: Before extracting experimental data from the target document by using the trained machine learning model, the method further includes: Obtain a training data set, where the training data set includes a plurality of labeled training data; Based on the training data set, a preset machine learning model is trained to obtain the trained machine learning model.
5. The method according to any one of claims 1 to 4, characterized in that: The step of analyzing the experimental data by a regression algorithm to obtain a predicted experimental result includes: Inputting the experimental data into a regression algorithm to construct a regression prediction model; The regression prediction model is used to perform prediction based on the experimental data and historical experimental data to obtain predicted experimental results corresponding to the experimental data.
6. The method according to claim 5, characterized in that The data cleaning and verification of the experimental data to obtain valid data includes: Performing data cleaning on the experimental data; Furthermore, the experimental data is verified by using the regression prediction model, and abnormal data points in the experimental data are identified and corrected by using an abnormality detection algorithm.
7. The method according to claim 6, characterized in that The method further comprises: Perform visualization processing on the target file, generate a visualization result, and display the visualization result.
8. A scientific literature data extraction device, characterized in that: The device comprises: The document acquisition module is used to acquire target documents from a preset document database; A data extraction module, used to extract experimental data from the target literature through a trained machine learning model; An experimental result prediction module is used to analyze the experimental data through a regression algorithm to obtain a predicted experimental result; A data cleaning and verification module is used to clean and verify the experimental data to obtain valid data; The target file output module is used to output the valid data and the prediction experiment results in a preset format to obtain a target file.
9. A computer device comprising a memory, a processor and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.