Material scientific research text data extraction method based on large language model
Through the material scientific research text data extraction method based on the large language model, the problems of low efficiency and poor integrity of scientific research text data extraction are solved, and efficient and accurate extraction and verification of different types of scientific research texts are achieved, which improves the adaptability and standardization of the data and supports material design.
Patent Information
- Application Number
- CN202510809536.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2045-06-17
AI Technical Summary
When faced with complex and changeable scientific research texts, existing technologies have problems such as low data extraction efficiency, poor integrity, susceptibility to human factors and lack of universality. In particular, when processing unstructured or semi-structured data of different text styles, information loss or erroneous extraction is prone to occur.
A materials research text data extraction method based on a large language model is adopted. Through a classification-specific fine-tuning model and a header extraction-specific fine-tuning model, combined with an adaptive header template and a multi-layer verification mechanism, structured data is extracted from the materials research text and verified. This includes building a classification model, a header extraction model and an adaptive header template, using the cross-entropy loss function for model fine-tuning, and combining a standard dictionary of materials science terminology and a language model for semantic reasoning and data correction.
It significantly improves the adaptability, completeness and accuracy of material research text data extraction, can adapt to different types of scientific research texts, ensure the logical consistency and standardization of data, and provide high-quality data to support material design.
Smart Images

Figure CN120654671A_ABST
Abstract
Description
Technical Field
[0001] The present invention generally relates to the field of material data management, and more specifically, to a method and device for extracting material scientific research text data based on a large language model. Background Art
[0002] With the development of technologies like artificial intelligence and big data, the research and development and application of new materials are gradually shifting towards digitalization and intelligence. Data has become a core resource for new materials research and development. As materials science research continues to deepen, a large number of scientific research documents, including scientific papers, patents, and experimental reports, continue to emerge. These documents contain a wealth of experimental data, material properties, structural information, and application cases. The key challenge currently in need of resolution is how to efficiently, accurately, and completely extract materials data from these diverse research documents.
[0003] Traditional data extraction relies primarily on manual screening and entry. While this method can obtain relatively complete data, it is inefficient and susceptible to human error. Faced with massive amounts of scientific research literature, patents, and experimental reports, manual extraction is not only time-consuming but also prone to data omissions and extraction errors. In recent years, natural language processing (NLP) technology has been applied to the automated extraction of material data, significantly improving extraction efficiency.
[0004] However, existing rule-based or traditional machine learning methods often struggle with complex and diverse scientific research texts due to their inability to understand context. As a result, the extracted material data often only contains material composition and performance data, resulting in poor data integrity. Furthermore, when processing unstructured or semi-structured data of varying text styles, existing methods are less universal and prone to information loss or incorrect extraction. Summary of the Invention
[0005] According to the present invention, a method for extracting materials research text data based on a large language model is provided. This method is applicable to unstructured or semi-structured data of various text styles, has strong universality, improves the integrity of extracted information, and reduces the probability of erroneous extraction.
[0006] In a first aspect of the present invention, a method for extracting material research text data based on a large language model is provided. The method comprises: Acquire material research text information, and filter out research texts containing target data based on the material research text information to obtain a target data set; A classification fine-tuning dataset is obtained according to the target dataset, a classification model is constructed, the classification model is fine-tuned based on the classification fine-tuning dataset using a standard binary classification cross entropy loss function to obtain a classification-specific fine-tuning model, and the material research text information is classified by the classification-specific fine-tuning model to obtain a dataset to be extracted; Obtaining an extraction and fine-tuning dataset based on the target dataset, constructing a header extraction model, obtaining a sample set based on the extraction and fine-tuning dataset and preset prompt words, fine-tuning the header extraction model based on the sample set using a cross-entropy loss function to obtain a header extraction-specific fine-tuning model, performing header extraction on the dataset to be extracted using the header extraction-specific fine-tuning model to obtain a header dataset, and then combining the header dataset with a preset header template to form an adaptive header template; According to the adaptive header template, structured data is extracted from the material scientific research text information, and then the structured data is verified to obtain the verified structured data.
[0007] In a second aspect of the present invention, an electronic device is provided. The electronic device comprises at least one processor; and a memory communicatively connected to the at least one processor; the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of the first aspect of the present invention.
[0008] Compared with the prior art, the present invention has the following beneficial technical effects: The present invention comprehensively improves the adaptability, completeness, accuracy and intelligence of the system in scientific research texts of different types of materials.
[0009] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The above and other features, advantages and aspects of the embodiments of the present invention will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings. In the accompanying drawings, the same or similar reference numerals represent the same or similar elements, wherein: Figure 1 A flowchart of a method for extracting material scientific research text data based on a large language model according to an embodiment of the present invention is shown; Figure 2 A flowchart of structured data verification according to an embodiment of the present invention is shown; Figure 3 A block diagram of an exemplary electronic device is shown in which embodiments of the present invention can be implemented.
[0011] Among them, 300 is an electronic device, 301 is a computing unit, 302 is a ROM, 303 is a RAM, 304 is a bus, 305 is an I / O interface, 306 is an input unit, 307 is an output unit, 308 is a storage unit, and 309 is a communication unit. DETAILED DESCRIPTION
[0012] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0013] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.
[0014] In this invention, the materials research text information is classified using a specialized fine-tuning model for classification to obtain a dataset to be extracted. A specialized fine-tuning model for header extraction is then used to extract headers from the dataset to obtain a header dataset. This header dataset is then combined with a preset header template to form an adaptive header template. Structured data is extracted from the materials research text information and verified using the adaptive header template. This approach comprehensively improves the system's adaptability, completeness, accuracy, and intelligence across different types of materials research text.
[0015] Example 1 Figure 1 A flowchart of a method for extracting material scientific research text data based on a large language model according to an embodiment of the present invention is shown.
[0016] The method includes: S101. Acquire material scientific research text information, and filter out scientific research texts containing target data based on the material scientific research text information to obtain a target data set, including: if the material scientific research text information contains the target data, mark it as True; if the material scientific research text information does not contain the target data, mark it as False; wherein all material scientific research text information marked as True constitutes the target data set.
[0017] Specifically, the target dataset is:
[0018] in, Fine-tuning dataset for classification tasks; For the The input features of the sample, that is, the text field to be classified; is the total number of training samples; is the training sample index; Indicates the The labels of samples, Indicates that the condition is true (True), Indicates that the condition is not met (False).
[0019] This screening method can efficiently identify and extract valid text containing target data from large-scale scientific research texts, significantly reducing the redundancy of subsequent data processing and improving the accuracy and efficiency of data extraction. The label-based binary classification mechanism can provide a high-quality data source for downstream tasks, preventing irrelevant information from interfering with the extraction model. It also facilitates model optimization and training under a supervised framework, improving the overall system's data utilization value and automation level in materials research scenarios.
[0020] S102. Obtain a classification fine-tuning dataset based on the target dataset, construct a classification model, fine-tune the classification model based on the classification fine-tuning dataset using a standard binary classification cross entropy loss function to obtain a classification-specific fine-tuning model, and classify the material scientific research text information through the classification-specific fine-tuning model to obtain a dataset to be extracted.
[0021] In this embodiment, the binary cross entropy loss function includes:
[0022]
[0023] in, is the binary cross entropy loss function; is the total number of training samples; is the training sample index; is the parameter set of the model; For the The true label of the training sample is 0 or 1; For the The input feature vector of training samples; is the predicted probability, specifically the model's samples belong to the positive class (i.e. ) can be calculated by identifying the first word of the generated sequence (token).
[0024] Fine-tuning the classification model using the standard binary cross-entropy loss function effectively optimizes the model's ability to distinguish between positive and negative samples, enabling it to more accurately determine whether a material research document contains the target data, with samples containing the target data being considered positive and samples not containing the target data being considered negative. The standard binary cross-entropy loss function offers excellent numerical stability and gradient-directedness, helping to improve the model's robustness and generalization capabilities in practical tasks.
[0025] S103. An extraction and fine-tuning dataset is obtained based on the target dataset, and a header extraction model is constructed. A sample set is obtained based on the extraction and fine-tuning dataset and preset prompt words. The header extraction model is fine-tuned based on the sample set using a cross-entropy loss function to obtain a header extraction-specific fine-tuning model. Headers are extracted from the dataset to be extracted using the header extraction-specific fine-tuning model to obtain a header dataset. The header dataset is then combined with a preset header template to form an adaptive header template.
[0026] Specifically, obtaining an extraction and fine-tuning dataset based on the target dataset refers to extracting the elemental composition, preparation method, hot and cold processing, heat treatment, and test-related experimental parameter names and corresponding units of the material, expressed in the format of "parameter name [unit]". The extraction and fine-tuning dataset is:
[0027] in, To extract the fine-tuning dataset; The original input text and prompt words; It is the preset header template; is the total number of training samples; is the training sample index.
[0028] As some optional implementations of this embodiment, the original input text and the prompt word , the default header template .
[0029] In this embodiment, the cross entropy loss function includes:
[0030]
[0031] in, is the cross entropy loss function; is the total number of training samples; is the training sample index; For the The total number of sample output sequences; For the The index of the sample output sequence; For the samples The predicted probability of a token; The predicted probability of the token; Header Sequence The joint probability of For input text; For the Sample input text; is the header sequence; For the In the sample The previous header sequence; For the In the sample A header sequence; are model parameters; For the A header sequence; For the The previous header sequence.
[0032] Specifically, For a given Sample input text And the generated header token sequence Under the condition of The predicted probability of a token; For any input text And the token sequence generated previously Under the conditions of The predicted probability of a token; For a given input text Generate the header sequence under the condition The joint probability of .
[0033] Fine-tuning the header extraction model using the cross-entropy loss function significantly improved the model's ability to generate standardized headers. This cross-entropy loss function aims to maximize the probability of generating true header sequences within the model. This strengthens the model's contextual understanding and sequence generation capabilities for diverse scientific research texts, ensuring it accurately outputs header items with standard semantics and unit formats.
[0034] Specifically, header extraction is performed on the dataset to be extracted using a dedicated fine-tuning model for header extraction, including:
[0035] in, The header data set includes standardized material composition, experimental parameter names and units; are the model parameters after fine-tuning; The original input text and prompt words; The header sequence.
[0036] In this embodiment, the header dataset is combined with a preset header template to form an adaptive header template to adapt to the data extraction requirements of different types of scientific research texts, specifically including:
[0037]
[0038] in, It is an adaptive header template; This is a header template, such as "Journal", "DOI", "Material Name", etc. This is the header data set; The qth extracted field (such as "heat treatment temperature [℃]", "heat treatment time [h]", "yield strength [MPa]", etc.).
[0039] S104. Extract structured data from the material research text information according to the adaptive header template, verify the structured data, obtain the verified structured data, and save it.
[0040] In this embodiment, structured data is extracted from material research text information according to an adaptive header template, including: extracting structured data from material research text information with the help of prompt words and an adaptive header template, and outputting it in the format specified by the prompt words.
[0041] Specifically, the prompt word construction follows the following principles: (1) Clear objectives: The prompt word should clearly express the task objective (such as "list fields", "extract experimental parameters", "structured output") to avoid vague instructions.
[0042] (2) Contextual supplementation: Provide sufficient contextual information to help the model accurately understand the content of the document paragraph.
[0043] (3) Information extraction control: limit the output field type through adaptive header template.
[0044] (4) Format constraints: The model output is explicitly required to be in a structured format and output table.
[0045] (5) Hallucination constraints: Add constraints such as “If there is no information, write ‘none’” and “Please do not make up” to the prompt words to avoid fictitious output.
[0046] In this embodiment, if Figure 2As shown, the structured data is verified to obtain verified structured data, including: S201. Build a standard dictionary of materials science terms, including:
[0047]
[0048] in, A collection of standard dictionaries of materials science terms; For the A standard field definition item, including its standard name and common spellings, abbreviations, variations and other synonyms; For the The standardized name of the item; for A collection of common spellings, abbreviations, or variant expressions of .
[0049] Specifically, the standard dictionary set of materials science terminology is used to standardize the naming of text fields such as materials, processes, tests, units, etc.
[0050] As some optional implementations of this embodiment, the standard dictionary of material science terms includes information such as field names, standard terms, abbreviations, synonyms, etc. For example, for the field "hot rolling", its standard expressions may include variants such as "hot rolling processing, hot rolling, hot rolling treatment, hot deformation, hot forming", etc. Its synonyms can be expressed as: : {"hot rolling": {"synonyms": ["hot-rolling", "hot rolled", "hot roll", "hotrollng", "hotrolling"]}.
[0051] S202. Perform text fuzzy matching on the structured data according to a standard dictionary set of material science terms. If the match is successful, automatically correct the structured data. If the match fails, correct the structured data, and use the corrected data and the automatically corrected data as a text correction data set.
[0052] In this embodiment, text fuzzy matching is performed on the structured data according to a standard dictionary set of material science terms. If the match is successful, the structured data is automatically corrected. If the match fails, the structured data is corrected, including: (1) Calculation of candidate standard items:
[0053] in, w To extract fields (text to be corrected); is a standard item; For fields With standard items The semantic similarity between For The most similar candidate criteria item.
[0054] (2) Setting is the similarity matching threshold, if , then it is considered a successful match and jump to step (3), otherwise it is considered a failed match and jump to step (4).
[0055] (3) Change the original field Replace with .
[0056] (4) Introducing language models for contextual semantic reasoning, including: 1) Calculate candidate standard items Generation probability under context conditions:
[0057] in, To extract fields and its context information, the language model generates standard terms The conditional probability of To extract fields The context fragment in which it is located; is the candidate with the highest conditional probability.
[0058] 2) If , Replace with Otherwise, keep the original field , and mark it as an abnormal field and delete it.
[0059] As some optional implementations of this embodiment, it is assumed that the extracted field The sentence "hotrollng" is clearly misspelled, so a fuzzy match is performed. When the match exceeds the set threshold, the match is considered successful and automatically normalized to the standard field, replacing "hotrollng" with "hot rolling." If the match is below the threshold but a high-confidence candidate exists, the match is considered unsuccessful, and a suggested field is generated based on the context. For example, based on the sentence "the alloy was subjected to tohomogenization at 1100°C," the standard term is inferred to be "homogenization heat treatment." A high-confidence candidate is one whose match is below the threshold but with a small deviation from the threshold. For example, if the threshold is 0.85, a high-confidence candidate is considered to exist when 0.7 < match < 0.85.
[0060] Text fuzzy matching is highly automated and semantically aware, and can effectively eliminate data pollution problems caused by mixed terminology, spelling errors, and ambiguous expressions.
[0061] S203 , based on a standard dictionary set of material science terms, using regular expression rules to perform standardized recognition on the units in the text correction dataset, performing unit conversion on the text correction dataset according to the standardized recognition result, and using the converted data as the unit correction dataset.
[0062] Specifically, unit conversion is performed using a unit conversion mapping table, which is:
[0063] in, To set the value From the original unit Convert to target standard units Conversion function of is the original value; is the conversion factor from the source unit to the target unit.
[0064] As some optional implementations of this embodiment, by looking up the unit conversion mapping table to match the unit conversion relationship, the numerical conversion is performed using the set target unit system (such as the International System of Units SI). For example, to uniformly convert 500 MPa to 0.5 GPa, the conversion formula is as follows:
[0065] S204. Set the numerical range of each type of parameter, filter the unit correction data set by the numerical range, treat the data whose numerical range is not within the numerical range as an abnormal data set, score the abnormal data set by the outlier scoring function, if the score is not greater than the threshold, mark it as normal data, if the score is greater than the threshold, mark it as abnormal data, and correct the abnormal data by the correction rule, and use the corrected data and normal data as the verified structured data.
[0066] In this embodiment, the numerical range of each parameter is set, including:
[0067] in, For the The reasonable value range of each field; For the The minimum reasonable value of the field; For the k The maximum reasonable value for the field.
[0068] As some optional implementations of this embodiment, taking the hardness field of a high entropy alloy as an example, a reasonable value range of this field is constructed as follows: ; When the extraction value If the data is not within the set range, for example, when the material hardness of a document is 985 HV, and it is found to exceed the maximum threshold of 800 after range judgment, it is considered abnormal and marked as "abnormal data".
[0069] Specifically, the outlier scoring function is:
[0070] in, is the probability density estimate of the field value; is the anomaly scoring function (the larger the value, the more abnormal).
[0071] As some optional implementation methods of this embodiment, taking the hardness field of high entropy alloy as an example, Gaussian kernel density estimation is used to construct the probability density function of the hardness field , assuming the preset threshold , for the extracted hardness value of 985 HV, it is calculated that , significantly higher than the preset threshold , so it is judged as an outlier and marked as “abnormal data”.
[0072] In this embodiment, the abnormal data is corrected using correction rules, including: Correcting the atomic percentage sums by the component sum consistency rule; and / or Correction of yield strength data using physical constraints on mechanical properties; and / or The phase transition onset temperature is corrected by thermodynamic constraints.
[0073] Specifically, the correction of the atomic percentage sum by the component sum consistency rule includes: determining whether the atomic percentage sum meets the percentage threshold condition; if so, the correction is deemed successful; if not, the data set with physical laws or constraints is corrected by automatic normalization.
[0074] The percentage threshold condition is:
[0075] Wherein, the automatic normalization is:
[0076] in, is the number of element types; is the element type index; For the The original atomic percentage of the element; is the allowable error threshold; is the normalized atomic percentage.
[0077] As some optional implementations of this embodiment, a data set with physical laws or constraints between multiple fields, for example, the extraction result of the component field of a certain high entropy alloy is: Table 1 element Extracted value (original) [at%] Al 15.3 Co 24.8 Cr 20.2 Fe 23.7 Ni 19.1 sum 103.1 Specifically, assuming the sum deviates from the theoretical value of 100, the error exceeds the threshold , triggering normalization correction, and the corrected value is calculated as:
[0078] Correcting the atomic percentage sums using component sum consistency rules can automatically fix component sum anomalies caused by OCR errors, incomplete extraction, or numerical drift.
[0079] Specifically, the yield strength is corrected by the physical constraint rules of mechanical properties, including: when the yield strength is not less than the ultimate strength, judging the data abnormality; if the data abnormality is a field mapping abnormality, exchanging the yield strength data and the ultimate strength data; if the abnormality is a unit abnormality, correcting the unit of the yield strength to a standard unit.
[0080] Specifically, the phase change starting temperature is corrected by thermodynamic constraint rules, including: when the phase change starting temperature is not less than the melting point temperature of the material, judging the data abnormality; if the data abnormality is a field mapping abnormality, the phase change starting temperature and the material melting point temperature are exchanged; if the data abnormality is a unit abnormality, the unit of the phase change starting temperature is corrected to a standard unit.
[0081] In this embodiment, the material field data can also be verified according to the following rules: Table 2 Rule Type Verify logical relationships Applicable fields Hardness-strength constraint rules Yield Strength-Hardness Data Thermal conductivity–temperature relationship Thermal conductivity decreases as temperature increases (metals) Thermal Conductivity-Temperature Data Resistivity-temperature relationship Resistivity increases with temperature (metals) Resistivity-temperature data set in, is the yield strength; is hardness; is the fitting coefficient of different metal systems.
[0082] For specialized materials, such as high-entropy alloys, superconducting materials, or functional ceramics, specific physical or chemical constraints can be introduced as auxiliary verification criteria. By integrating physical laws, automated, highly reliable calibration of complex material data can be achieved, significantly improving data quality and usability.
[0083] By establishing component sum consistency rules, physical constraints for mechanical properties, and thermodynamic constraints, this approach automatically identifies and corrects data anomalies that arise during the extraction process, such as inconsistent component sums, incorrect mechanical property field mappings, and conflicting physical parameters. This method fully integrates the physical laws of materials science, enabling a complete process from numerical detection to semantic correction. This effectively improves the accuracy, rationality, and scientific credibility of structured data, making it particularly suitable for repairing and enhancing the consistency of multi-field coupled data in complex material systems such as high-entropy alloys.
[0084] The multi-layered validation mechanism constructed through steps S201–S204 in this invention covers key aspects such as terminology standardization, text ambiguity correction, unit conversion, and outlier detection. It can systematically identify and automatically repair structured data issues such as mixed terminology, spelling errors, inconsistent units, and numerical anomalies. This method combines domain knowledge dictionaries, language model semantic reasoning, and a probability density function scoring mechanism to achieve efficient cleaning and standardization of materials research data, significantly improving data standardization, comparability, and analytical applicability, providing high-quality data support for subsequent machine learning modeling and materials design.
[0085] As some optional implementations of this embodiment, the verified structured data is saved as an Excel file, and the processing status of each text is recorded. All processing modules will aggregate the processing logs, change records, and exception labels to form a final data quality score, which facilitates subsequent manual or model screening and priority processing.
[0086] Example 2 Taking the mechanical property data extraction of high entropy alloys in the literature as an example, the specific steps include: (1) Data Collection: Using scientific research literature databases, including but not limited to the Elsevier journal database and the Springer Nature journal database, we searched for and downloaded literature related to high-entropy alloys using "High-entropy alloys" or "Multi-principal elemental alloys" as search keywords. We collected a wide range of literature and obtained a high-entropy alloy literature dataset. The PDF documents were converted into text string format. In this example, PyMuPDF (Python MuPDFWrapper) was used to parse and read the PDF file, output the text characters, and use the text characters as the target dataset.
[0087] (2) Classify the data: 1) Training data preparation stage: Some documents are selected for manual classification to determine whether they contain target data such as yield strength and strain. If they do, they are marked as True, otherwise False. A classification fine-tuning dataset is established. The data format is as follows: {"messages": [{"role":"user","content":"{paper_int}"}, {"role":"assistant","content":"False"}]}, where "paper_int" refers to the text characters parsed from the PDF documents. The classification fine-tuning dataset is saved in JSONL format and contains 30 sets of training examples and 10 sets of validation data, totaling 2,339,000 tokens.
[0088] 2) In this example, GPT-4O-Mini is used as the basic large model. The fine-tuning dataset obtained in 1) is used to fine-tune the large model. The fine-tuning hyperparameters for the classification task are shown in Table 2, and the classification-specific fine-tuning model GPT-4O-Mini-Classify is obtained.
[0089] Table 3 Batch Size Learning Rate Multiplier Epochs Seed 1 1 3 1869035027 3) Use the classification task fine-tuning model gpt-4o-mini-classify to classify the text obtained in (1), and use the data containing target data such as yield strength and strain as the dataset to be extracted.
[0090] (3) Obtain a dedicated fine-tuning model for header extraction through training: 1) Training data preparation stage: Data extraction is performed on the dataset to be extracted obtained in (2), mainly extracting the elemental composition, preparation method, hot and cold processing, heat treatment, and test-related experimental parameter names and corresponding units of the materials, which are expressed in the format of "parameter name [unit]". The prompt words extracted from the header, the text characters of the document, the parameter names and the corresponding units are combined to form a header extraction fine-tuning dataset. The data format is as follows [{"role":"system","content":"{Prompt}"}, {"role":"user","content":"{paper_int}"}]}, where "Prompt1" is the prompt word that guides the large model to extract the header, and "paper_int" is used to refer to the text characters parsed from the PDF document.
[0091] Specifically, "Prompt1" mainly guides the large model to extract data through the following aspects: 1. Extract the elemental composition of the material and express it in the format of "element name [unit]", such as "Al [wt.%]" or "Hf [at.%]".
[0092] 2. Extraction material preparation method and parameters. Parameters include temperature, time, pressure, medium, and size, etc., expressed in the format of "parameter name [unit]".
[0093] 3. Extract the hot and cold processing and heat treatment parameters of the material. The parameters include temperature, time, pressure, medium and size, etc., and are expressed in the format of "parameter name [unit]", such as "Annealing Temperature [K]".
[0094] 4. Extract the tensile or compression test parameters of the material. The parameters may include equipment, temperature, time, pressure, medium, and size, etc., and are expressed in the format of "parameter name [unit]".
[0095] 5. Combine the parameters extracted from steps 1 to 4 to obtain all parameters.
[0096] 2) In this example, GPT-4O-Mini is used as the basic large model. A sample set is obtained based on the extraction fine-tuning dataset and preset prompt words. The sample set is used to fine-tune the header extraction model. The fine-tuning includes 39 sets of training examples and 10 sets of validation data, totaling 1,694,000 tokens. The fine-tuning hyperparameters for the header extraction task are shown in Table 3. Through fine-tuning, the GPT-4O-Mini-Header model specifically for header extraction is obtained.
[0097] Table 4 Batch Size Learning Rate Multiplier Epochs Seed 1 0.5 3 460188267 (4) Extract the header data set: The header extraction fine-tuning model gpt-4o-mini-header is used to perform header extraction on the text obtained in (2). The material composition, experimental parameter names, and units are extracted and output in a standard format to obtain a header dataset. The material composition and experimental procedures of different documents vary greatly. The purpose of this step is to enable the large language model to dynamically extract header data based on different documents to adapt to documents with different styles.
[0098] (5) Build an adaptive header template: The preset header template and the header data set extracted in (4) form a complete adaptive header template.
[0099] Specifically, the journal, DOI, paper title, material, and target extracted data (e.g., yield strength) are used as preset header templates and expressed in the format of "parameter name [unit]". The preset header template is combined with the dynamic header dataset extracted in (4) to form a complete adaptive header template.
[0100] (6) Extracting structured data: The prompt word "Prompt2" is constructed to guide the large language model to extract parameter values according to the adaptive header. The prompt project guides the large language model to analyze the document content, extract experimental data, and output it according to the specified format. In this embodiment, the large language model selected is the o1 model.
[0101] Specifically, "Prompt2" mainly guides the large model to extract data through the following aspects: 1) Guide the large language model to extract experimental parameter values and performance data from a specified paragraph.
[0102] 2) Instruct large language models to pay attention to problems with abbreviations or synonyms.
[0103] 3) Guide the large language model to output in a standardized format. For example, the sample size is output in the format of "A×B×C", non-existent data is represented by "-", and the room temperature or ambient temperature is saved as 25℃ or 298K.
[0104] 4) Guide the large language model to extract data of different parameters, and each row of material composition and experimental parameters corresponds to unique performance data.
[0105] 5) Finally, the structured data is output in CSV format.
[0106] (8) Storage of structured data: The xlwt library is used to store structured data in Excel files, automatically naming and storing them according to the source. The processing status of each document, including the number of input and output tokens, is recorded for subsequent analysis and saved in txt format. The extracted data is accurate and complete, including the journal, DOI, document title, preparation method, equipment information, material composition, process flow, sample dimensions, and performance data, which are beyond the capabilities of existing materials literature data extraction methods.
[0107] (9) The final step of Example 2 is to encapsulate the operations (1)-(8) above into a function. The function takes the text as input and applies the prompt word to extract the text information. In this way, an automated process can be built to process documents in batches, improving the efficiency of data extraction.
[0108] Example 3 Taking the performance data extraction of 3D printed tungsten alloy in patent literature as an example, the specific steps are as follows: (1) Data collection includes searching for and downloading relevant literature on tungsten alloy 3D printing from a patent database using "tungsten alloy 3D printing" as a search keyword, and extensively collecting patents. The collected patents are in PDF format, and the PDF is parsed using the same method as in Example 2 to output text characters.
[0109] (2) Classify the data: In this embodiment, the Deepseek R1 model is used, the searched patents are Chinese patents, and the prompt words are designed in Chinese, specifically "Determine whether the main content of the article includes hardness, tensile strength or strain parameters. If yes, return 'true', otherwise return: 'false'.
[0110] (3) Extract the header data set: The Deepseek R1 model was used to extract the header data set. The specific prompt was: "You are a researcher specializing in data extraction in the field of materials science. Please extract information in the following order: 1. Material elements: Extract the elements contained in the material in the format of 'element name [unit]', such as 'Al [wt.%]' or 'Hf [at.%]'. If multiple materials are involved, list them together without distinction. 2. Preparation method: Key experimental parameters and units related to material preparation, such as temperature, time, pressure, medium, sample size, equipment, etc. The format is 'parameter name [unit]'. If there is no unit, only the parameter name is listed. Arrange them in the order of the experiments. 3. Processing method: Extract the key experimental parameters and units of the processing. The format is the same as above. If the same process step is repeated, '(nth time)' must be indicated after the parameter name, such as 'Annealing temperature (2nd time) [K]'. Arrange them in the order of the experiments, ignoring the differences between different materials. 4. Test parameters: Extract the key experimental parameters and units of the compression or tensile test itself. The format is the same as above. Only the experimental parameters are extracted, not the performance results. 5. Final output: Integrate the results of the above four steps into a line of text, separated by a semicolon ';'. Example format: 'Al [%]; Hf [%]; Mo [%]; Preparation method; ...; Hot isostatic pressing temperature [°C]; Annealing temperature [K]; ...; Test method; ...'. Only the final integration result is output. ” As some optional implementation methods of this embodiment, the extracted dynamic data headers are as follows: "W [at.%]; Ni [at.%]; Fe [at.%]; Co [at.%]; Al [at.%]; V [at.%]; Nb [at.%]; Preparation method; Ball-to-material ratio; Ball milling speed [rpm]; Ball milling time [min]; First mixing speed [rpm]; First mixing time [h]; Second mixing speed [rpm]; Second mixing time [min]; Laser power [W]; Laser rate [mm / min]; Single layer thickness [mm]; Scanning spacing [mm]; Test method; Test standard; Tensile rate [mm / min]".
[0111] (4) Build an adaptive header template: Combine the preset header and the dynamic header data set extracted from the large model to form a complete adaptive header template.
[0112] Specifically, the patent source, patent number, title, material, and target extraction data are used as preset headers and combined with the dynamic header extracted in (3) to form a complete adaptive header template, as follows: "Source, Patent number, Name, Material, W [at. %]; Ni [at. %]; Fe [at. %]; Co [at. %]; Al [at. %]; V [at. %]; Nb [at. %]; Preparation method; Ball-to-material ratio; Milling speed [rpm]; Milling time [min]; First mixing speed [rpm]; First mixing time [h]; Second mixing speed [rpm]; Second mixing time [min]; Laser power [W]; Laser rate [mm / min]; Single layer thickness [mm]; Scanning spacing [mm]; Test method; Test standard; Tensile rate [mm / min]; Hardness [HV], Tensile strength [Mpa], Strain [%]".
[0113] (6) Extracting structured data: Construct prompt words to guide the large language model to extract parameter values according to the adaptive header and output them in the specified format. In this embodiment, the Deepseek R1 model is selected. The specific prompt words are: "Please extract material composition, experimental parameters and performance data from the provided text according to the following requirements, and return them in tabular form. The column headers are: {column_headers}. 1. Try to extract clear values. Experimental parameters and performance data are usually located in the examples section. 2. If a piece of data is missing, use '-' to indicate it. 3. Pay attention to abbreviations. 'Bal.' means 'balance' and is used to describe element content. 4. Sample size format: rectangular samples are 'AxBxC', A, B, C are the dimensions of the three sides (such as '4x4x8'); cylindrical ... For rod-shaped samples, use 'dDxL' or 'rRxL', where D is the diameter and R is the radius (e.g., 'd4x10' indicates a sample with a diameter of 4 and a length of 10). 5. For samples with different experimental parameters and different material compositions, please extract and record them separately in separate rows. 6. Ensure that the extracted data is accurate and complete, and maintain a consistent relationship between the experimental parameters and performance data in each row. 7. The default room temperature / ambient temperature is 25°C or 298K. 8. The returned results are in CSV format. No prefixes or suffixes may be added, and the header must be retained. Please start processing the following text: "{paper_int}" Among them, "column_headers" represents the dynamic header text, and "paper_int" represents the text characters used to extract data patents.
[0114] (6) Storage of structured data: Use the xlwt library to store structured data in Excel files.
[0115] According to the embodiments of the present invention, the present invention has the following advantages compared with the prior art: (1) The present invention utilizes a large language model and realizes the accurate extraction of multi-source heterogeneous structured data in materials research texts through a multi-layer data extraction design, thereby obtaining an adaptive header that is highly adaptable to the style of the document. It can adapt to scientific research texts with diverse styles, structures and sources, including scientific research documents, patents, reports, etc., significantly improving the generalization ability of the model in different fields and documents from different sources, and also improving the accuracy and standardization of the extraction results.
[0116] (2) The present invention relies on the powerful contextual understanding and reasoning capabilities of the large language model to accurately and completely extract key information from various scientific research texts, including source information (such as journal name, DOI, document title or patent number, etc.), material preparation methods, equipment information, material composition, process flow, sample size and performance data, etc. The extracted data is automatically associated based on the text context relationship to achieve corresponding matching of material composition, preparation process and performance parameters, ensuring the logical consistency and integrity of the data. In addition, the present invention constructs a full-process information extraction mechanism for material scientific research tasks, relying on the long text understanding and multi-layer structured output capabilities of the large language model to achieve chain extraction from document information, material composition, process parameters, experimental conditions to mechanical properties, effectively solving the problems of incomplete field coverage, upstream and downstream disconnection, and loose information granularity in existing methods, and improving the integrity and reusability of scientific research data.
[0117] (3) In order to improve the accuracy, standardization and usability of the results of material research text data extraction, the present invention further proposes a data post-processing and verification enhancement module. This module is used after the large language model extraction results, and mainly includes three functions: dictionary-based text field anomaly detection and correction, unit consistency detection and conversion, and numerical anomaly detection and intelligent correction. It integrates dictionary rules, language model reasoning, statistical modeling and physical constraints to ensure the consistency and standardization of the extracted data in semantic expression, unit standards and physical rationality, and comprehensively improves the integrity, consistency and credibility of structured data.
[0118] (4) This invention significantly improves the integrity, standardization, and credibility of data extraction from materials research literature. The overall approach not only achieves high-precision structured data acquisition, but also provides a solid and reliable data foundation for big data construction and intelligent analysis in the field of materials science. It further promotes the implementation of the materials data-driven R&D paradigm and provides innovative solutions and theoretical support for materials data management, standard setting, and intelligent research.
[0119] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0120] Figure 3 A schematic block diagram of an electronic device 300 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0121] Electronic device 300 includes a computing unit 301, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 302 or a computer program loaded from a storage unit 308 into a random access memory (RAM) 303. RAM 303 may also store various programs and data required for the operation of electronic device 300. Computing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to bus 304.
[0122] Multiple components in the electronic device 300 are connected to the I / O interface 305, including an input unit 306, such as a keyboard, a mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; a storage unit 308, such as a magnetic disk, an optical disk, etc.; and a communication unit 309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the electronic device 300 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0123] The computing unit 301 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 301 performs the various methods and processes described above, such as methods S101 through S104. For example, in some embodiments, methods S101 through S104 can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 300 via the ROM 302 and / or the communication unit 309. When the computer program is loaded into the RAM 303 and executed by the computing unit 301, one or more steps of methods S101 through S104 described above can be performed. Alternatively, in other embodiments, the computing unit 301 may be configured to execute methods S101 to S104 in any other appropriate manner (eg, by means of firmware).
[0124] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0125] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for extracting material scientific research text data based on a large language model, characterized in that: include: Acquire material research text information, and filter out research texts containing target data based on the material research text information to obtain a target data set; A classification fine-tuning dataset is obtained according to the target dataset, a classification model is constructed, the classification model is fine-tuned based on the classification fine-tuning dataset using a standard binary classification cross entropy loss function to obtain a classification-specific fine-tuning model, and the material research text information is classified by the classification-specific fine-tuning model to obtain a dataset to be extracted; Obtaining an extraction and fine-tuning dataset based on the target dataset, constructing a header extraction model, obtaining a sample set based on the extraction and fine-tuning dataset and preset prompt words, fine-tuning the header extraction model based on the sample set using a cross-entropy loss function to obtain a header extraction-specific fine-tuning model, performing header extraction on the dataset to be extracted using the header extraction-specific fine-tuning model to obtain a header dataset, and then combining the header dataset with a preset header template to form an adaptive header template; According to the adaptive header template, structured data is extracted from the material scientific research text information, and then the structured data is verified to obtain the verified structured data.
2. The method according to claim 1, characterized in that The method of screening out scientific research texts containing target data based on the material scientific research text information to obtain a target data set includes: If the material research text information contains target data, it is marked as True; If the material research text information does not contain the target data, it is marked as False; Among them, all material scientific research text information marked as True constitutes the target dataset.
3. The method according to claim 1, characterized in that The binary cross entropy loss function includes: in, is the binary cross entropy loss function; is the total number of training samples; is the training sample index; is the parameter set of the model; For the The true labels of the training samples; For the The input feature vector of training samples; is the predicted probability.
4. The method according to claim 1, wherein The cross entropy loss function includes: in, is the cross entropy loss function; is the total number of training samples; is the training sample index; For the The total number of sample output sequences; For the The index of the sample output sequence; For the In the sample The predicted probability of a token; For the The predicted probability of a token; Header sequence The joint probability of For input text; For the Sample input text; is the header sequence; For the In the sample The previous header sequence; For the In the sample A header sequence; are model parameters; For the A header sequence; For the The previous header sequence.
5. The method according to claim 1, wherein Verifying the structured data to obtain verified structured data includes: Construct a standard dictionary collection of materials science terms; Performing text fuzzy matching on the structured data according to a standard dictionary set of material science terms; if the match is successful, automatically correcting the structured data; if the match fails, correcting the structured data; and using the corrected data and the automatically corrected data as a text correction data set; Based on a standard dictionary set of material science terms, using regular expression rules to perform standardized recognition of units in the text correction dataset, performing unit conversion on the text correction dataset according to the standardized recognition result, and using the converted data as the unit correction dataset; The numerical range of each type of parameter is set, and the unit correction data set is filtered by the numerical range. The data whose numerical range is not within the numerical range is regarded as an abnormal data set. The abnormal data set is scored by the outlier scoring function. If the score is not greater than the threshold, it is marked as normal data. If the score is greater than the threshold, it is marked as abnormal data. The abnormal data is corrected by the correction rule, and the corrected data and normal data are used as the verified structured data.
6. The method according to claim 5, characterized in that Correct abnormal data through correction rules, including: Correcting the atomic percentage sums by the component sum consistency rule; and / or Correction of yield strength data using physical constraints on mechanical properties; and / or The phase transition onset temperature is corrected by thermodynamic constraints.
7. The method according to claim 6, characterized in that The correction of the atomic percentage sum by the component sum consistency rule includes: Determine whether the sum of the atomic percentages meets the percentage threshold condition. If so, the correction is considered successful. If not, the data set with physical laws or constraints is corrected through automatic normalization; The percentage threshold condition is: The automatic normalization is: in, is the number of element types; is the element type index; For the The original atomic percentage of the element; is the allowable error threshold; is the normalized atomic percentage.
8. The method according to claim 6, characterized in that The correction of yield strength by physical constraint rules of mechanical properties includes: When the yield strength is not less than the ultimate strength, the data anomaly is determined. If the data anomaly is a field mapping anomaly, the yield strength data and the ultimate strength data are exchanged; if the anomaly is a unit anomaly, the unit of the yield strength is corrected to the standard unit.
9. The method according to claim 6, characterized in that The correction of the phase change starting temperature by thermodynamic constraint rules includes: When the phase change starting temperature is not less than the material melting point temperature, the data abnormality is determined. If the data abnormality is a field mapping abnormality, the data of the phase change starting temperature and the data of the material melting point temperature are exchanged; if the data abnormality is a unit abnormality, the unit of the phase change starting temperature is corrected to the standard unit.
10. An electronic device comprising at least one processor; and a memory communicatively connected to the at least one processor; characterized in that: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Text structuring method and system based on large model, terminal and storage medium
CN118133784A
Contrast test identification method and device based on large language model and storage medium
CN119670750A
Automatic generation method of traffic safety facility statistical table
CN120124605A