A material scientific research text data extraction method based on a large language model

By using a large language model-based approach, combined with adaptive header templates and a standard dictionary of materials science terminology, the problem of accuracy and completeness in data extraction from scientific research texts was solved. This approach enables efficient and accurate extraction and verification of complex texts, improving the adaptability and intelligence of the data.

CN120654671BActive Publication Date: 2025-11-28辽宁材料实验室
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510809536.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-11-28
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

Existing technologies struggle to efficiently and accurately extract material data from complex and ever-changing scientific research texts, especially in unstructured or semi-structured data where information loss and erroneous extraction occur.

Method used

A large language model-based approach is adopted, which combines a classification fine-tuning model and a header extraction model with an adaptive header template and a standard dictionary of materials science terminology to achieve classification and structured data extraction of materials research texts. Furthermore, a multi-layered verification mechanism is used to improve the accuracy and completeness of the data.

Benefits of technology

It significantly improves the adaptability, accuracy, and intelligence of data extraction from materials science texts, enabling it to adapt to different types of scientific texts, reduce errors, and enhance the standardization and comparability of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654671B_ABST
    Figure CN120654671B_ABST
Patent Text Reader

Abstract

The present application relates to the field of material data management, in particular to a material research text data extraction method and device based on a large language model. The method comprises obtaining a target data set from a material research text; constructing a classification model and fine-tuning it using a standard binary classification cross-entropy loss function, obtaining a data set to be extracted through the fine-tuned model; constructing a table header extraction model and fine-tuning the table header extraction model using a cross-entropy loss function, obtaining a table header data set through the fine-tuned model, combining the table header data set with a preset table header template to form an adaptive table header template; extracting structured data from the material research text information according to the adaptive table header template, and verifying the structured data to obtain verified structured data. In this way, the adaptability, completeness, accuracy and intelligence level of the system in different types of material research text can be comprehensively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of material data management, and more particularly, to a material research text data extraction method and device based on a large language model. BACKGROUND

[0002] With the development of technologies such as artificial intelligence and big data, the research and application of new materials are gradually changing towards digitization and intelligentization, and data has become the core resource for the research and development of new materials. With the deepening of materials science research, a large amount of scientific research literature, patents, experimental reports and other scientific research texts continue to emerge, which contain rich experimental data, material properties, structural information and application cases. How to efficiently, accurately and completely extract material data from various scientific research texts has become a key problem that needs to be solved.

[0003] Traditional data extraction mainly relies on manual screening and manual input, although it can obtain data relatively completely, but the efficiency is low, and it is easily affected by human factors. When facing a large amount of scientific research literature, patents and experimental reports, manual extraction not only takes a long time, but also is prone to data omission and extraction errors. In recent years, natural language processing (NLP) technology has been applied to the automatic extraction of material data, which has significantly improved the extraction efficiency.

[0004] However, the existing methods based on rules or traditional machine learning often lack the ability to understand the context when facing complex and variable scientific research texts, resulting in the extracted material data often only containing material composition and performance data, and the data completeness is poor. In addition, when dealing with unstructured or semi-structured data of different text styles, the existing methods have poor universality and are prone to information loss or incorrect extraction. SUMMARY

[0005] According to the present application, a material research text data extraction scheme based on a large language model is provided. This scheme can be applied to unstructured or semi-structured data of different text styles, has strong universality, improves the completeness of extracted information, and reduces the probability of incorrect extraction.

[0006] In a first aspect of the present application, a material research text data extraction method based on a large language model is provided. The method comprises:

[0007] Obtaining material research text information, screening scientific research texts containing target data from the material research text information to obtain a target data set;

[0008] According to the target data set, a classification fine-tuning data set is obtained, a classification model is constructed, the classification model is fine-tuned based on a standard binary classification cross-entropy loss function based on the classification fine-tuning data set, a classification special fine-tuning model is obtained, and the material research text information is classified through the classification special fine-tuning model to obtain a to-be-extracted data set;

[0009] According to the target data set, an extraction fine-tuning data set is obtained, a table header extraction model is constructed, a sample set is obtained according to the extraction fine-tuning data set and a preset prompt word, the table header extraction model is fine-tuned based on the sample set using a cross-entropy loss function, a table header extraction special fine-tuning model is obtained, and the table header extraction special fine-tuning model is used to extract a table header from the to-be-extracted data set to obtain a table header data set, and then the table header data set is combined with a preset table header template to form an adaptive table header template.

[0010] According to the adaptive table header template, structured data is extracted from the material research text information, and the structured data is verified to obtain verified structured data.

[0011] In a second aspect of the present application, an electronic device is provided. The electronic device comprises at least one processor; and a memory connected to the at least one processor in communication; the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of the first aspect of the present application.

[0012] Compared with the prior art, the present application has the following beneficial technical effects:

[0013] The present application comprehensively improves the adaptability, integrity, accuracy and intelligent level of the system in different types of material research text.

[0014] It should be understood that the content described in the summary section is not intended to limit the key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS

[0015] The above and other features, advantages, and aspects of embodiments of the present application will become more apparent by describing in detail the following embodiments with reference to the attached drawings. In the drawings, the same or similar reference numerals refer to the same or similar elements, and:

[0016] Figure 1 A flowchart of a material research text data extraction method based on a large language model according to an embodiment of the present application is shown;

[0017] Figure 2 A flowchart of structured data verification according to an embodiment of the present application is shown.

[0018] Figure 3 A block diagram of an exemplary electronic device capable of implementing embodiments of the present application is shown.

[0019] In the figure, 300 is an electronic device, 301 is a computing unit, 302 is a ROM, 303 is a RAM, 304 is a bus, 305 is an I / O interface, 306 is an input unit, 307 is an output unit, 308 is a storage unit, and 309 is a communication unit. DETAILED DESCRIPTION

[0020] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some but not all of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.

[0021] In addition, the term "and / or" herein is only to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after it.

[0022] In the present application, the material research text information is classified by a classification special fine-tuning model to obtain a to-be-extracted data set; a table header extraction special fine-tuning model is used to extract the table header of the to-be-extracted data set to obtain a table header data set, and then the table header data set is combined with a preset table header template to form an adaptive table header template. According to the adaptive table header template, structured data is extracted from the material research text information and verified. In this way, the adaptability, integrity, accuracy and intelligent level of the system in different types of material research text can be comprehensively improved.

[0023] Embodiment 1

[0024] Figure 1 A flowchart of a material research text data extraction method based on a large language model according to an embodiment of the present application is shown.

[0025] The method comprises:

[0026] S101. Obtain materials research text information, and filter out research texts containing target data based on the materials research text information to obtain a target dataset, including: if the materials research text information contains target data, mark it as True; if the materials research text information does not contain target data, mark it as False; wherein, all materials research text information marked as True constitutes the target dataset.

[0027] Specifically, the target dataset is:

[0028]

[0029] in, Fine-tuning the dataset for classification tasks; For the first The input features of each sample, i.e., the text field to be classified; The total number of training samples; For training sample index; Indicates the first The labels of each sample, among which This indicates that the condition is true. This indicates that the condition is not met (False).

[0030] By employing the above screening methods, effective text containing target data can be efficiently identified and extracted from large-scale scientific research texts, significantly reducing the redundancy of subsequent data processing and improving the accuracy and efficiency of data extraction. The label-based binary classification mechanism provides a high-quality data source for downstream tasks, avoiding interference from irrelevant information on the extraction model. It also facilitates supervised training of the model, enhancing the overall system's data utilization value and automation level in materials science research scenarios.

[0031] S102. Obtain a classification fine-tuning dataset based on the target dataset, construct a classification model, fine-tune the classification model using the standard binary cross-entropy loss function based on the classification fine-tuning dataset to obtain a classification-specific fine-tuning model, and classify the material research text information using the classification-specific fine-tuning model to obtain the dataset to be extracted.

[0032] In this embodiment, the binary classification cross-entropy loss function includes:

[0033]

[0034]

[0035] in, The cross-entropy loss function is used for binary classification. The total number of training samples; For training sample index; This is the set of parameters for the model; For the first The true label of each training sample, with a value of 0 or 1; For the first The input feature vector of each training sample; To predict probabilities, specifically, the model for the first... Each sample belongs to the positive class (i.e.) The predicted probability of a token can be calculated by identifying the first word of the generated sequence (token).

[0036] Fine-tuning the classification model using the standard binary cross-entropy loss function effectively optimizes its ability to distinguish between positive and negative samples. This allows the model to more accurately determine whether a research text contains the target data; samples containing the target data are considered positive, while those not containing it are considered negative. The standard binary cross-entropy loss function exhibits good numerical stability and gradient-directedness, which helps improve the model's robustness and generalization ability in practical tasks.

[0037] S103. Based on the target dataset, obtain the extracted fine-tuning dataset, construct the header extraction model, obtain the sample set based on the extracted fine-tuning dataset and preset prompt words, fine-tune the header extraction model using the cross-entropy loss function based on the sample set to obtain a dedicated header extraction fine-tuning model, extract the header from the dataset to be extracted using the dedicated header extraction fine-tuning model to obtain the header dataset, and then combine the header dataset with the preset header template to form an adaptive header template.

[0038] Specifically, the extracted fine-tuning dataset obtained from the target dataset refers to extracting the elemental composition, preparation method, cold and hot processing, heat treatment, and test-related experimental parameter names and corresponding units of the material, represented in the format of "parameter name [unit]". The extracted fine-tuning dataset is as follows:

[0039]

[0040] in, To extract the fine-tuning dataset; The original input text and prompt words; This is a preset header template; The total number of training samples; This is the index for the training samples.

[0041] As some optional implementations of this embodiment, the original input text and prompt words Preset header template .

[0042] In the present embodiment, the cross-entropy loss function comprises:

[0043]

[0044]

[0045] wherein, is a cross-entropy loss function; is the total number of training samples; is the index of a training sample; is the total number of output sequences of the th sample; is the index of the th output sequence of the th sample; is the predicted probability of the th token of the th output sequence of the th sample; is the joint probability of the table header sequence is an input text; is the input text of the th sample; is a table header sequence; is the table header sequence before the th in the th sample; is the th table header sequence in the th sample; is a model parameter; is the th table header sequence; is the table header sequence before the th.

[0046] Specifically, is the predicted probability of the th token in the th sample given the input text and the generated table header token sequence ; is the predicted probability of the th token given any input text and the token sequence generated so far ; is the joint probability of generating the table header sequence given the input text .

[0047] The table header extraction model is fine-tuned by a cross-entropy loss function, which can significantly improve the generation capability of the model for standardized table headers. The cross-entropy loss function aims to maximize the generation probability of the real table header sequence in the model, and strengthens the context understanding and sequence generation capability of the model when facing diversified scientific research texts, ensuring that it can accurately output table header items with standard semantics and unit formats.

[0048] Specifically, the table header extraction of the to-be-extracted data set is performed by a table header extraction special fine-tuning model, including:

[0049]

[0050] Among them, is a table header data set, which includes standardized material components, experimental parameter names and units; is the fine-tuned model parameters; is the original input text and prompt word; is the table header sequence.

[0051] In this embodiment, the table header data set is combined with the preset table header template to form an adaptive table header template to adapt to the data extraction requirements in different types of scientific research texts, specifically including:

[0052]

[0053]

[0054] Among them, is an adaptive table header template; is a table header template, such as "journal", "DOI", "material name", etc.; is a table header data set; is the qth extraction field (such as "heat treatment temperature [℃]", "heat treatment time [h]", "yield strength [MPa]", etc.).

[0055] S104, according to the adaptive table header template, extract structured data from the material scientific research text information, and then verify the structured data to obtain the verified structured data and save.

[0056] In this embodiment, the structured data is extracted from the material scientific research text information according to the adaptive table header template, including: extracting the structured data from the material scientific research text information by means of the prompt word and the adaptive table header template, and outputting according to the format specified by the prompt word.

[0057] Specifically, the prompt word construction follows the following principles:

[0058] (1) Clear goal: prompt words should clearly express the task goal (such as "list fields", "extract experimental parameters", "structured output"), avoid ambiguous instructions.

[0059] (2) Context supplement: Provide enough context information to help the model accurately understand the content of the literature paragraph.

[0060] (3) Information extraction control: Limit the output field type through adaptive table header templates.

[0061] (4) Format constraint: Clearly require the model output to be structured format, output table.

[0062] (5) Illusion constraint: Add constraints such as "if there is no information to write 'no'", "please do not fabricate" in the prompt words, avoid fictional output.

[0063] In this embodiment, as shown in Figure 2 , the structured data is verified to obtain verified structured data, including:

[0064] S201, a material science term standard dictionary set is constructed, including:

[0065]

[0066]

[0067] Wherein, is a material science term standard dictionary set; is the standard field definition item, which contains its standard name and common spelling, abbreviation, variant and other synonymous expressions; is the standardized name of the item; is a set of common spelling, abbreviation or variant expressions of .

[0068] Specifically, the material science term standard dictionary set is used to standardize the naming of materials, processes, tests, units and other text fields.

[0069] As some optional embodiments of the present embodiment, the material science term standard dictionary includes field name, standard term, abbreviation, synonymous expression and other information. For example, for the "hot rolling" field, its standard expression can include "hot rolling processing, hot rolling, hot rolling treatment, hot deformation, hot forming" and other variants. Its synonyms can be expressed as: dictionary {"hot rolling": {"synonyms": ["hot-rolling", "hot rolled", "hot roll", "hotrollng", "hotrolling"]}。

[0070] S202, according to the material science terminology standard dictionary set, the structured data is fuzzy matched, if the matching is successful, the structured data is automatically corrected, if the matching fails, the structured data is corrected, and the corrected data and the automatically corrected data are used as the text correction data set.

[0071] In this embodiment, according to the material science terminology standard dictionary set, the structured data is fuzzy matched, if the matching is successful, the structured data is automatically corrected, if the matching fails, the structured data is corrected, including:

[0072] (1) Calculate the candidate standard item:

[0073]

[0074] Wherein, w is the extracted field (to be corrected text); is the standard item; is the field The semantic similarity between and the standard item ; is the most similar candidate standard item.

[0075] (2) Set as the similarity matching threshold, if , it is considered as matching success and jump to step (3), otherwise it is considered as matching failure and jump to step (4).

[0076] (3) Replace the original field with .

[0077] (4) Introduce language model for context semantic reasoning, including:

[0078] 1) Calculate the candidate standard item The generation probability under the context condition:

[0079]

[0080] Wherein, is the extracted field and its context information, and the language model generates the conditional probability of the standard item ; extracting fields the context fragment in which the field is located; the candidate with the maximum conditional probability.

[0081] 2) If , is replaced by , otherwise, the original field is kept while being marked as an abnormal field and deleted.

[0082] As some optional implementations of the present embodiment, assuming that the extracted field is "hotrollng", which obviously has a spelling error, fuzzy matching is performed, when the matching degree is higher than a set threshold, it is determined that the matching is successful, and it is automatically normalized to the standard field, "hotrollng" is replaced by "hot rolling"; when it is lower than the threshold but there is a high-confidence candidate, it is determined that the matching fails, and a suggestion field is generated by combining the context, for example, according to the sentence "the alloy was subjected to homogenization at 1100°C", it is inferred that the standard term is "homogenization heat treatment". The high-confidence candidate refers to a candidate whose matching degree is lower than the threshold but deviates from the threshold by a small amount, for example, when 0.7 < matching degree < 0.85, it is considered that there is a high-confidence candidate.

[0083] Text fuzzy matching has high automation and semantic perception ability, and through text fuzzy matching, the data pollution problem caused by mixed use of terms, spelling errors and ambiguous expressions can be effectively eliminated.

[0084] S203, based on the material science term standard dictionary set, using regular expression rules to standardize and identify the units in the text correction data set, and according to the standardization and identification result, the text correction data set is unit converted, and the converted data is taken as a unit correction data set.

[0085] Specifically, unit conversion is performed through a unit conversion mapping table, and the unit conversion mapping table is:

[0086]

[0087] wherein, is a conversion function for converting a numerical value from an original unit to a target standard unit ; is the original numerical value; is the conversion coefficient from the original unit to the target unit.

[0088] As some optional implementations of the present embodiment, the unit conversion relationship is matched by searching a unit conversion mapping table, and numerical conversion is performed using a set target unit system (such as the International System of Units SI). For example, 500 MPa is converted to 0.5 GPa, and the conversion formula is as follows:

[0089]

[0090] S204, set the numerical range of each type of parameter, filter the unit correction data set through the numerical range, take the data not in the numerical range as an abnormal data set, score the abnormal data set through an outlier score function, if the score is not greater than a threshold value, mark it as normal data, if the score is greater than the threshold value, mark it as abnormal data, and correct the abnormal data through a correction rule, and take the corrected data and normal data as the verified structured data.

[0091] In the present embodiment, the numerical range of each type of parameter is set, including:

[0092]

[0093] wherein, is the reasonable value interval of the i-th field; is the minimum reasonable value of the i-th field; is the maximum reasonable value of the i-th field. k

[0094] As some optional implementations of the present embodiment, taking the hardness field of high-entropy alloy as an example, the reasonable value range of the field is constructed as: When the extracted value is not within the set range, for example, when the material hardness in a certain literature is extracted as 985 HV, it is found that it exceeds the maximum threshold value 800, that is, it is considered abnormal and marked as “abnormal data”.

[0095] Specifically, the outlier score function is:

[0096]

[0097] wherein, is the probability density estimation of the field value; is the abnormal score function (the larger the value, the more abnormal).

[0098] As some optional implementations of the present embodiment, taking the hardness field of high-entropy alloy as an example, the Gaussian kernel density estimation is used to construct the probability density function of the hardness field , and the preset threshold value ​​​For the extracted hardness value of 985 HV, it is calculated that it is significantly higher than the preset threshold , so it is determined as an abnormal value, marked as "abnormal data".

[0099] In this embodiment, the abnormal data is corrected by a correction rule, including:

[0100] correcting the atomic percentage sum by a component sum consistency rule; and / or

[0101] correcting the yield strength data by a mechanical property physical constraint rule; and / or

[0102] correcting the phase transition starting temperature by a thermodynamic constraint rule.

[0103] Specifically, the atomic percentage sum is corrected by the component sum consistency rule, including: judging whether the atomic percentage sum satisfies a percentage threshold condition, if yes, it is considered that the correction is successful, if not, the data set with physical law or constraint relationship is corrected by automatic normalization.

[0104] The percentage threshold condition is:

[0105]

[0106] The automatic normalization is:

[0107]

[0108] wherein, is the number of element types; is the element type index; is the original atomic percentage of the th element; is the allowable error threshold; is the normalized atomic percentage.

[0109] As some optional embodiments of the present embodiment, the data group with physical law or constraint relationship between multiple fields, for example, the component field extraction result of a certain high-entropy alloy is:

[0110] Table 1

[0111] Element Extracted Value (raw) [at%] Al 15.3 Co 24.8 Cr 20.2 Fe 23.7 Ni 19.1 Sum 103.1

[0112] Specifically, assuming that the sum deviates from the theoretical value of 100, the error exceeds the threshold , triggering normalization correction, and the corrected value is:

[0113]

[0114] The atomic percentage sum is corrected by the component total consistency rule, so that the component total abnormality problem caused by OCR error, extraction defect or value drift can be automatically repaired.

[0115] Specifically, the yield strength is corrected by the mechanical property physical constraint rule, including: when the yield strength is not less than the ultimate strength, judging the data abnormality condition, if the data abnormality condition is field mapping abnormality, the data of the yield strength and the data of the ultimate strength are exchanged; if the abnormality condition is unit abnormality, the unit of the yield strength is corrected to the standard unit.

[0116] Specifically, the phase transition starting temperature is corrected by the thermodynamic constraint rule, including: when the phase transition starting temperature is not less than the material melting point temperature, judging the data abnormality condition, if the data abnormality condition is field mapping abnormality, the phase transition starting temperature and the material melting point temperature are exchanged; if the data abnormality condition is unit abnormality, the unit of the phase transition starting temperature is corrected to the standard unit.

[0117] In the embodiment, the material field data can also be verified according to the following rules:

[0118] Table 2

[0119] Rule Type Check Logic Relationship Applicable Field Hardness- Strength Constraint Rule Yield Strength- Hardness Data Thermal Conductivity- Temperature Relationship Thermal Conductivity Decreases with Increasing Temperature (Metallic) Thermal Conductivity- Temperature Group Data Resistivity- Temperature Relationship Resistivity Increases with Increasing Temperature (Metallic) Resistivity- Temperature Group Data

[0120] wherein, is the yield strength; is the hardness; is the fitting coefficient of different metal systems.

[0121] For some special field materials, such as high-entropy alloy, superconducting material or functional ceramic, specific physical or chemical constraint conditions can also be introduced as auxiliary verification basis. By fusing the physical law, the automatic high-reliability correction of complex material data can be realized, and the data quality and usability are significantly improved.

[0122] By constructing the component total consistency rule, the mechanical property physical constraint rule and the thermodynamic constraint rule, the abnormal data such as inconsistent component total, mechanical property field mapping error and physical parameter contradiction in the extraction process can be automatically identified and corrected. The method fully fuses the physical law in material science, realizes the whole process from numerical detection to semantic correction, effectively improves the accuracy, rationality and scientific credibility of structured data, and is especially suitable for data repair and consistency enhancement of multi-field coupling in complex material system such as high-entropy alloy.

[0123] The multilayer verification mechanism constructed through S201-S204 steps can cover key links such as term standardization, text fuzzy correction, unit conversion and abnormal value detection, and can systematically identify and automatically repair structured data problems such as mixed use of terms, spelling errors, inconsistent units and abnormal values. The method combines the field knowledge dictionary, language model semantic reasoning and probability density function scoring mechanism to realize efficient cleaning and standardization processing of material research data, significantly improve the standardization, comparability and analysis applicability of the data, and provide high-quality data support for subsequent machine learning modeling and material design.

[0124] As some optional embodiments of the present embodiment, the verified structured data is saved as an Excel file, and the processing of each text is recorded. All processing modules uniformly summarize the processing logs, change records and abnormal labels to form the quality score of the final data, facilitating subsequent manual or model screening and priority processing.

[0125] Embodiment 2

[0126] Taking the extraction of mechanical property data of high-entropy alloys in literature as an example, the following steps are included:

[0127] (1) Collect data: Through scientific research literature databases including but not limited to Elsevier journal database, SpringerNature journal database, etc., use "High-entropy alloys" or "Multi-principal elementalloys" as the retrieval keyword, retrieve high-entropy alloy related literature and download, widely collect literature, obtain high-entropy alloy literature dataset. Convert the PDF format literature into text string format, this embodiment uses PyMuPDF (Python MuPDF Wrapper) to parse and read PDF files, output text characters, and take the text characters as the target dataset.

[0128] (2) Classify the data:

[0129] 1) Training data preparation phase: select part of the literature for manual classification to determine whether the literature contains yield strength, strain, etc. Target data, if it contains, mark as True, otherwise False, establish classification fine-tuning dataset, data format as follows {"messages": [{"role":"user","content":"{paper_int}"}, {"role":"assistant","content":"False"}]}}, Where "paper_int" is used to refer to the text characters parsed from the PDF literature, the classification fine-tuning dataset is saved in JSONL format, containing 30 training examples, 10 validation data, a total of 2339000 tokens.

[0130] 2) In this embodiment, gpt-4o-mini is used as the basis of the large model, and the fine-tuning data set obtained in 1) is used to fine-tune the large model, and the classification task fine-tuning hyperparameters are as shown in Table 2, and the classification special fine-tuning model gpt-4o-mini-classify is obtained.

[0131] Table 3

[0132] Batch Size Learning Rate Multiplier Epochs Seed 1 1 3 1869035027

[0133] 3) Use the classification task fine-tuning model gpt-4o-mini-classify to classify the text obtained in (1), and the data containing yield strength, strain, etc. Target data as the data set to be extracted.

[0134] (3) Through training, a table header extraction special fine-tuning model is obtained:

[0135] 1) Training data preparation phase: data extraction is performed on the extracted data set obtained in (2), mainly extracting the elemental composition of the material, the preparation method, cold and hot processing, heat treatment, and the experimental parameter name and corresponding unit related to testing, in the format of "parameter name [unit]". Combine the prompt words for table header extraction, text characters and parameter names and corresponding units of the literature to form a table header extraction fine-tuning dataset, the data format is as follows [{"role":"system","content":"{Prompt}"}, {"role":"user","content":"{paper_int}"}]}}, Where "Prompt1" is the prompt word for guiding the large model to extract the table header, and "paper_int" is used to refer to the text characters parsed from the PDF literature.

[0136] Specifically, "Prompt1" mainly guides the large model to perform data extraction in the following aspects:

[0137] 1. Extract the material element composition in the format of "element name [unit]", such as "Al [wt.%]" or "Hf [at.%]".

[0138] 2. Extract the material preparation method and parameters, including temperature, time, pressure, medium, and size, in the format of "parameter name [unit]".

[0139] 3. Extract the cold and hot processing and heat treatment parameters of the material, including temperature, time, pressure, medium, and size, in the format of "parameter name [unit]", such as "Annealing Temperature [K]".

[0140] 4. Extract the tensile or compression test parameters of the material, which may include equipment, temperature, time, pressure, medium, and size, in the format of "parameter name [unit]".

[0141] 5. Combine the parameters extracted in steps 1 to 4 to obtain all parameters.

[0142] 2) In this embodiment, gpt-4o-mini is used as the base large model, and the sample set is obtained according to the extracted fine-tuning data set and the preset prompt word. The sample set is used to fine-tune the table header extraction model, which contains 39 groups of training examples, 10 groups of validation data, a total of 1,694,000 tokens, and the table header extraction task fine-tuning hyperparameters are as shown in Table 3. The table header extraction special fine-tuning model gpt-4o-mini-header is obtained by fine-tuning.

[0143] Table 4

[0144] Batch Size Learning Rate Multiplier Epochs Seed 1 0.5 3 460188267

[0145] (4) Extract the table header data set:

[0146] Use the table header extraction fine-tuning model gpt-4o-mini-header to perform table header extraction tasks on the text obtained in (2), extract material composition, experimental parameter name and unit, and output in standard format to obtain the table header data set. The material composition and experimental steps of different documents are different, and the purpose of this step is to enable the large language model to dynamically extract the table header data according to different documents to adapt to documents with different styles.

[0147] (5) Construct an adaptive table header template:

[0148] The preset table header template and the table header data set extracted in (4) are combined to form a complete adaptive table header template.

[0149] Specifically, the journal, DOI, paper title, materials, and target extraction data (such as yield strength) are taken as the preset header template, represented in the format of "Parameter name [unit]". The preset header template and the dynamic header data set extracted in (4) are combined to form a complete adaptive header template.

[0150] (6) Extraction of structured data:

[0151] The construction prompt "Prompt2" guides the large language model to extract parameter values according to the adaptive header. Through the prompt engineering, the large language model is guided to analyze the literature content, extract experimental data, and output in the specified format. The large language model selected in this embodiment is the o1 model.

[0152] Specifically, "Prompt2" guides the large model to extract data in the following aspects:

[0153] 1) Guide the large language model to extract experimental parameter values and performance data in the specified paragraph.

[0154] 2) Guide the large language model to pay attention to abbreviations or synonyms.

[0155] 3) Guide the large language model to output in a standard format, such as sample size in the format of "A x B x C", non-existent data represented by "-", and room temperature or environmental temperature saved as 25°C or 298K.

[0156] 4) Guide the large language model to extract data for different parameters, with each row of material composition, experimental parameter corresponding to unique performance data.

[0157] 5) Finally, output the structured data in csv format.

[0158] (8) Storage of structured data:

[0159] The structured data is stored in an Excel file using the xlwt library, and is automatically named and stored according to the literature source. The processing of each literature is recorded, including the number of input and output Tokens, which facilitates subsequent analysis and is saved in txt format. The extracted data is accurate and complete, including the journal, DOI, literature title, preparation method, equipment information, material composition, process flow, sample size, and performance data, which is not achieved by existing material literature data extraction methods.

[0160] (9) The last step of Example 2 is to encapsulate the operations of (1)-(8) into a function, which takes text as input and applies prompts to extract information from the text. In this way, an automated process can be constructed to batch process literature and improve the efficiency of data extraction.

[0161] Example 3

[0162] Take the performance data extraction of 3D printed tungsten alloy in patent documents as an example, the specific steps are as follows:

[0163] (1) Collect data through patent database, including using "3D printing of tungsten alloy" as the retrieval keyword, retrieve tungsten alloy 3D printing related literature and download, and collect patents widely. The collected patents are in PDF format, and the same method as in example 2 is used to parse the PDF and output the text characters.

[0164] (2) Classify the data:

[0165] In this embodiment, the Deepseek R1 model is used, the retrieval patent is Chinese patent, and the prompt word is designed in Chinese, which is "determine the main content of this paper is including hardness, tensile strength or strain parameter. If yes, return 'true', otherwise return: 'false'.

[0166] (3) Extract the table header dataset:

[0167] The table header dataset is obtained by using the Deepseek R1 model for table header extraction, and the specific prompt word is: "You are a researcher specializing in data extraction in the field of materials science. Please extract the information in the following order: 1. Material elements: Extract the elements contained in the material, format 'element name [unit]', such as 'Al [wt.%]' or 'Hf [at.%]', if multiple materials are involved, combine and list them without distinction. 2. Preparation method: Key experimental parameters and units related to material preparation, such as temperature, time, pressure, medium, sample size, equipment, etc. Format 'parameter name [unit]', if no unit, only parameter name, arrange in experimental order. 3. Processing method: Extract the key experimental parameters and units of processing, format same as above, if the same process step is repeated, need to mark '(nth)' after parameter name, such as 'annealing temperature (2nd) [K]'. Arrange in experimental order, ignore differences between different materials. 4. Test parameters: Extract the key experimental parameters and units of compression or tensile test itself, format consistent with the above, only extract experimental parameters, do not include performance results. 5. Final output: Integrate the results of the above four steps into one line of text, separated by semicolons ';', format example: 'Al [%]; Hf [%]; Mo [%]; Preparation method;...; Hot isostatic pressing temperature [°C]; Annealing temperature [K];...; Test method;...'. Only output the final integrated result."

[0168] As some optional implementations of the present embodiment, the extracted dynamic data table header is as follows: "W [at.%]; Ni[at.%]; Fe [at.%]; Co [at.%]; Al [at.%]; V [at.%]; Nb [at.%]; preparation method; ball-to-powder ratio; ball milling speed [rpm]; ball milling time [min]; first mixing speed [rpm]; first mixing time [h]; second mixing speed [rpm]; second mixing time [min]; laser power [W]; laser speed [mm / min]; single layer thickness [mm]; scanning interval [mm]; test method; test standard; tensile speed [mm / min]".

[0169] (4) Construct an adaptive table header template:

[0170] Combine the preset table header and the dynamic table header dataset extracted by the large model to form a complete adaptive table header template.

[0171] Specifically, the patent source, patent number, title, material, and target extraction data are combined with the dynamic table header extracted in (3) to form a complete adaptive table header template, which is specifically as follows: "Source, Patent No., Name, Material, W [at.%]; Ni[at.%]; Fe [at.%]; Co [at.%]; Al [at.%]; V [at.%]; Nb [at.%]; preparation method; ball-to-powder ratio; ball milling speed [rpm]; ball milling time [min]; first mixing speed [rpm]; first mixing time [h]; second mixing speed [rpm]; second mixing time [min]; laser power [W]; laser speed [mm / min]; single layer thickness [mm]; scanning interval [mm]; test method; test standard; tensile speed [mm / min]; hardness [HV], tensile strength [Mpa], strain [%]".

[0172] (6) Extract structured data:

[0173] The construction prompt word guides the large language model to extract parameter values according to the adaptive table header, and outputs according to the specified format, and in the embodiment, the Deepseek R1 model is selected. The specific prompt word is: "Please extract material components, experimental parameters and performance data from the provided text according to the following requirements and return them in table form. The column headers are: {column_headers}. 1. Try to extract clear numerical values, experimental parameters and performance data are usually located in the example section. 2. If some data is missing, use '-' to represent it. 3. Pay attention to processing abbreviations, 'Bal.' represents 'balance' for element content description. 4. The sample size format is: 'AxBxC' for rectangular samples, A, B and C are three edge sizes (such as '4x4x8'); 'dDxL' or 'rRxL' for cylindrical or rod-shaped samples, D is the diameter and R is the radius (such as 'd4x10' represents a sample with a diameter of 4 and a length of 10). 5. For samples with different experimental parameters and different material components, please extract and record them in different rows respectively. 6. Ensure that the extracted data is accurate and complete, and keep the correspondence between experimental parameters and performance data consistent in each row. 7. The default room temperature / environmental temperature is 25°C or 298K. 8. The returned result is in CSV format, do not add prefixes or suffixes, and the table headers must be preserved. Please start processing the following text: {paper_int}."

[0174] Wherein, "column_headers" represents a dynamic table header text, and "paper_int" represents a text character used for extracting data patents.

[0175] (6) Storage of structured data: structured data is stored into an Excel file by using the xlwt library.

[0176] According to the embodiments of the present application, compared with the prior art, the present application has the following advantages:

[0177] (1) The present application uses a large language model to realize the accurate extraction of multi-source heterogeneous structured data in material research text through multi-layer data extraction design, and further obtains an adaptive table header highly adapted to the style of literature, which can adapt to research texts with diverse styles, different structures and multiple sources, including scientific research literature, patents, reports, etc., significantly improving the generalization ability of the model in different fields and different sources of literature, and also improving the accuracy and standardization of the extraction results.

[0178] (2) The present application relies on the powerful context understanding and reasoning ability of the large language model, can accurately and completely extract key information in various scientific research texts, covering source information (such as journal name, DOI, literature title or patent number, etc.), material preparation method, equipment information, material composition, process flow, sample size and performance data, etc. The extracted data is automatically associated based on the context relationship of the text, realizing the corresponding matching of material composition, preparation process and performance parameters, ensuring the logical consistency and integrity of the data. In addition, the present application constructs a full-process information extraction mechanism for material research tasks, relying on the long text understanding and multi-layer structured output capability of the large language model, realizes the chain extraction from literature information, material composition, process parameters, experimental conditions to mechanical properties, effectively solves the problems of incomplete field coverage, upstream and downstream disconnection and coarse information granularity in existing methods, and improves the integrity and reusability of scientific research data.

[0179] (3) In order to improve the accuracy, standardization and usability of the material scientific research text data extraction result, the present application further proposes a data post-processing and verification enhancement module, which is used after the large language model extraction result, mainly including three parts: text field abnormality detection and correction based on dictionary, unit consistency detection and conversion, numerical abnormality detection and intelligent correction, combining dictionary rules, language model reasoning, statistical modeling and physical constraints, to ensure the consistency and standardization of the extracted data in semantic expression, unit standard and physical rationality, and comprehensively improve the integrity, consistency and reliability of the structured data.

[0180] (4) The present application significantly improves the integrity, standardization and reliability of material scientific literature data extraction. The overall method not only realizes high-precision structured data acquisition, but also provides a solid and reliable data foundation for large data construction and intelligent analysis in the field of materials science, further promotes the landing of material data-driven research and development paradigm, and provides innovative solutions and theoretical support for material data management, standard development and intelligent research.

[0181] It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily required by the present application.

[0182] Figure 3A schematic block diagram of an electronic device 300 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular telephones, smartphones, wearable devices, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not intended to limit the inventiveness in the present document as described and / or claimed.

[0183] The electronic device 300 includes a computing unit 301 that can perform various appropriate actions and processes in accordance with a computer program stored in a read-only memory (ROM) 302 or a computer program loaded into a random access memory (RAM) 303 from a storage unit 308. Various programs and data required for the operation of the electronic device 300 can also be stored in the RAM 303. The computing unit 301, the ROM 302, and the RAM 303 are connected to each other through a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.

[0184] Various components in the electronic device 300 are connected to the I / O interface 305, including an input unit 306, such as a keyboard, a mouse, etc.; an output unit 307, such as various types of displays, speakers, etc.; the storage unit 308, such as a magnetic disk, an optical disk, etc.; and a communication unit 309, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 309 allows the electronic device 300 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0185] The computing unit 301 can be various general and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 301 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 301 performs various methods and processes described above, such as the methods S101-S104. For example, in some embodiments, the methods S101-S104 can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 308. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 300 via the ROM 302 and / or the communication unit 309. When the computer program is loaded onto the RAM 303 and executed by the computing unit 301, one or more steps of the methods S101-S104 described above can be performed. Alternatively, in other embodiments, the computing unit 301 can be configured to perform the methods S101-S104 by any other appropriate means, such as by means of firmware.

[0186] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0187] The specific embodiments discussed above do not constrain the scope of the present application. Those skilled in the art will readily understand that various modifications, combinations, sub-combinations, and alternatives of the specific embodiments discussed above can be made in light of design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present application should be included in the scope of the present application.

Claims

1. A method for extracting material research text data based on a large language model, characterized in that, include: Obtain materials research text information, and filter out research texts containing target data based on the materials research text information to obtain the target dataset; Based on the target dataset, a classification fine-tuning dataset is obtained, a classification model is constructed, and the classification model is fine-tuned using the standard binary cross-entropy loss function based on the classification fine-tuning dataset to obtain a classification-specific fine-tuning model. The material research text information is classified using the classification-specific fine-tuning model to obtain the dataset to be extracted. Based on the target dataset, an extraction and fine-tuning dataset is obtained, and a header extraction model is constructed. A sample set is obtained based on the extraction and fine-tuning dataset and preset prompt words. The header extraction model is fine-tuned using the cross-entropy loss function based on the sample set to obtain a dedicated header extraction fine-tuning model. The header is extracted from the dataset to be extracted using the dedicated header extraction fine-tuning model to obtain a header dataset. The header dataset is then combined with a preset header template to form an adaptive header template. Based on the adaptive header template, structured data is extracted from the materials research text information, and then the structured data is verified to obtain the verified structured data. The step of validating the structured data to obtain validated structured data includes: Construct a standard dictionary set of materials science terminology; The structured data is subjected to text fuzzy matching based on the standard dictionary set of materials science terminology. If the match is successful, the structured data is automatically corrected. If the match fails, the structured data is corrected. The corrected data and the automatically corrected data are used as the text correction dataset. Based on the standard dictionary set of materials science terminology, the units in the text correction dataset are standardized and identified using regular expression rules. The units in the text correction dataset are converted according to the standardized identification results, and the converted data is used as the unit correction dataset. Set a numerical range for each type of parameter, filter the unit calibration dataset by the numerical range, and identify the data that is outside the numerical range as the abnormal dataset. Use an outlier scoring function to score the abnormal dataset. If the score is not greater than the threshold, it is marked as normal data. If the score is greater than the threshold, it is marked as abnormal data. Correct the abnormal data using correction rules, and use the corrected data and normal data as the verified structured data.

2. The method according to claim 1, characterized in that, The step of filtering out research texts containing target data based on the material research text information to obtain the target dataset includes: If the material research text information contains the target data, mark it as True; If the material research text information does not contain the target data, mark it as False; The target dataset consists of all the research text information on materials marked as True.

3. The method according to claim 1, characterized in that, The binary classification cross-entropy loss function includes: ; ; in, The cross-entropy loss function is used for binary classification. The total number of training samples; For training sample index; The set of parameters for the model; For the first The true labels of each training sample; For the first The input feature vector of each training sample; To predict probabilities.

4. The method according to claim 1, characterized in that, The cross-entropy loss function includes: ; ; in, The cross-entropy loss function; The total number of training samples; For training sample index; For the first The total number of output sequences for each sample; For the first The index of the output sequence of each sample; For the first In the nth sample The predicted probability of each token; For the first The predicted probability of each token; For the header sequence The joint probability; Enter the text; For the first The input text for each sample; This is the header sequence; For the first In the nth sample The previous header sequence; For the first In the nth sample A header sequence; These are model parameters; For the first A header sequence; For the first The previous header sequence.

5. The method according to claim 1, characterized in that, Abnormal data is corrected using correction rules, including: The total atomic percentage is corrected using the consistency rule for total component composition; and / or The yield strength data is corrected using physical constraint rules based on mechanical properties; and / or The phase transition initiation temperature is corrected using thermodynamic constraint rules.

6. The method according to claim 5, characterized in that, The correction of the total atomic percentage using the consistency rule of total component summation includes: Determine whether the total atomic percentage meets the percentage threshold condition. If it does, the correction is considered successful. If it does not, the dataset with physical laws or constraints is corrected through automatic normalization. The percentage threshold condition is: ; The automatic normalization is as follows: ; in, The number of element types; Index for element types; For the first The percentage of original atoms of an element; This is the allowable error threshold; This represents the normalized percentage of atoms.

7. The method according to claim 5, characterized in that, The correction of yield strength data through physical constraint rules of mechanical properties includes: When the yield strength is not less than the ultimate strength, the data anomaly is identified. If the data anomaly is due to a field mapping anomaly, the yield strength data and the ultimate strength data are swapped. If the anomaly is due to a unit anomaly, the unit of the yield strength is corrected to the standard unit.

8. The method according to claim 5, characterized in that, The correction of the phase transition initiation temperature using thermodynamic constraint rules includes: When the phase change initiation temperature is not less than the material melting point temperature, the data anomaly is identified. If the data anomaly is due to a field mapping anomaly, the data of the phase change initiation temperature and the data of the material melting point temperature are swapped. If the data anomaly is due to a unit anomaly, the unit of the phase change initiation temperature is corrected to the standard unit.

9. An electronic device, comprising at least one processor; and a memory communicatively connected to said at least one processor; characterized in that, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.

Citation Information

Patent Citations

  • Text structuring method and system based on large model, terminal and storage medium

    CN118133784A

  • Contrast test identification method and device based on large language model and storage medium

    CN119670750A