Alloy structured information acquisition method based on scientific literature
By combining the MatSciBERT model with regular expressions, the data extraction problem in alloy material literature was solved, and structured data acquisition of alloy composition and properties was achieved, supporting material design and optimization and reducing experimental resources.
Patent Information
- Application Number
- CN202510791213.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-12
AI Technical Summary
Existing technologies make it difficult to efficiently extract the chemical composition and property data of alloy materials from massive scientific literature, especially because the writing format and content structure of the literature vary greatly, the data is distributed in the text, tables, charts, etc., and professional terminology in the field of alloy materials is difficult to identify.
The MatSciBERT language model is used for text mining, and the dictionary is expanded by combining regular expressions and semantic similarity. Downstream tasks are developed and model training and fine-tuning are performed to identify chemical elements, alloy components and their related terms to form structured data.
It has achieved the accurate extraction of the relationship between alloy composition and performance from scientific literature, forming structured data with a one-to-one correspondence between "alloy composition-alloy performance", supporting material design and optimization, and reducing experimental resources and costs.
Smart Images

Figure CN120636644A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of text mining and material science, and relates to a method for acquiring alloy structured data based on scientific literature. Background Art
[0002] Alloy materials are widely used in industry and daily life. Since alloys generally have superior physical and chemical properties than pure metals, such as higher strength, corrosion resistance, and wear resistance, they are widely used in the manufacture of aircraft, automobiles, ships, as well as in construction and electronic equipment. In recent years, with the advancement of science and technology and the improvement of manufacturing technology, new high-performance alloy materials have continued to emerge. For example, single-crystal high-temperature alloys have made significant progress in the aerospace field and can maintain stable performance under extreme conditions. Lightweight alloys such as aluminum alloys and titanium alloys are widely used in automobile manufacturing, significantly reducing vehicle weight and improving fuel efficiency. In the field of renewable energy, corrosion-resistant alloys also play an important role in marine wind power equipment and solar energy equipment.
[0003] In recent years, the number of research papers on materials such as alloys has increased significantly. This literature contains a large amount of high-quality experimental data and research results. How to efficiently extract and organize valuable structured data from this massive amount of literature to enable data-driven materials design and analysis has become an urgent problem to be solved.
[0004] Natural Language Processing (NLP), as an important branch of artificial intelligence, can automatically process and analyze large amounts of text data through machine learning and deep learning algorithms, enabling information extraction, classification, and summarization. Currently, some studies have attempted to apply NLP technology to mine and extract data from scientific literature. However, specialized research on materials such as alloys is still relatively rare. Due to the large differences in the writing format and content structure of different scientific literature, data is distributed in different locations such as the text, tables, and charts, making data extraction difficult. At the same time, the field of materials such as alloys contains a large number of chemical elements, alloy components, and their related terms. How to accurately identify and parse these professional terms is also an issue that requires special attention.
[0005] Based on the above background, the present invention proposes a text mining method based on natural language processing technology to automatically extract the chemical composition and property data of alloys from scientific literature. Summary of the Invention
[0006] Aiming at the massive scientific literature data on alloys and other materials, the present invention proposes a method for acquiring alloy structured data. MatSciBERT is based on the BERT language model, which can simultaneously consider the previous and next information in the context and learn deeper semantic representations in the text. Using the MatSciBERT model, derivative downstream tasks are developed and the downstream tasks are trained and fine-tuned. It is possible to more accurately identify chemical elements, alloy components and related terms from the text information of scientific literature, and extract the relationship between chemical composition and material properties to form a one-to-one correspondence of "alloy composition-alloy properties" structured data. In addition, scientific literature also covers some tabular data. By using regular expressions to construct rules and combining semantic similarity to expand the dictionary, specific alloy components and properties can be extracted from the tabular data.
[0007] The technical method of the present invention is a method for obtaining alloy structured data based on scientific literature, comprising the following steps:
[0008] S1 Scientific article acquisition: Using web big data crawler technology, download a large number of scientific articles related to alloy materials and save them in TXT and XML formats.
[0009] S2 article data preprocessing and table parsing. The text in the article is split into sentences and data cleaning is performed simultaneously. For tabular data, specific alloy compositions and properties are extracted.
[0010] S3 Sentence classification: Sentence classification is performed on the large number of sentences generated in step S2 to identify sentences containing specific alloy properties.
[0011] S4 Named Entity Recognition: Based on the classification results of step S3, identify and extract the specific alloy components and corresponding property values in the sentence.
[0012] S5 Entity Relationship Extraction: The alloy composition entities extracted in step S4 are matched one-to-one with the property entities, and are sorted to form a structured database.
[0013] S6 uses the extracted structured data to build a machine learning model to predict the performance indicators of the alloy and evaluate the performance of the alloy.
[0014] Furthermore, step S2 specifically includes using the nltk.sent_tokenize function to implement sentence splitting and table parsing, the process is as follows:
[0015] S2011 first uses TXT documents from various documents to identify common sentence-ending punctuation marks, such as periods (.), question marks (?), and exclamation points (!), and uses these punctuation marks as potential sentence endings. It then uses contextual rules and heuristics to determine whether a punctuation mark actually signifies the end of a sentence. For example, a period followed by a capital letter typically signifies the end of a sentence, but this is not necessarily the case if the period appears within an abbreviation. It then relies on a pre-trained language model and rule base to further handle various complex situations, such as quotations, parentheses, and abbreviations. Through these steps, the entire text is split into individual sentences.
[0016] S2012 deletes and filters out sentences that are too short (<10 bytes).
[0017] S2021 parses the XML files of each document, extracts all the table information mentioned in the document, and saves it locally in the "file DOI number.xlsx" format.
[0018] S2022 creates a dictionary of alloy properties to be extracted and constructs regular expression rules to match the corresponding fields. Data on specific alloy properties is extracted from all table data and organized into Excel spreadsheets.
[0019] Furthermore, the step S3 specifically includes:
[0020] S3011 prepares a training data set in advance, marking sentences that mention the properties of alloys as "relevant" and sentences that do not mention the corresponding content as "irrelevant".
[0021] S3012 encodes the input text after passing it through the word segmenter to obtain word embedding and position encoding.
[0022] S302 uses the MatSciBERT model to load pre-trained weights and obtain the feature representation of each word. The output vector representation of the [CLS] position is denoted as h [CLS] .
[0023] S303 will h [CLS] Pass it to the classifier self.classifier and classify it through the fully connected layer (two classifiers). The first fully connected layer will [CLS] The input layer (of dimension D_in) is mapped to a hidden layer (of dimension H) and a ReLU activation function is applied. The second fully connected layer maps the output of the hidden layer (of dimension H) to the output layer (of dimension D_out = 2), resulting in two scores z0 and z1. The scores are converted to probabilities p0 and p1 using the softmax activation function, and the class with the higher probability is selected as the prediction result.
[0024]
[0025] Here, i is 0 or 1.
[0026] S3041 adopts the focal loss function, introduces a dynamic weight mechanism, and sets the positive sample ratio weight α t And the focusing parameter γ makes the model pay more attention to samples that are difficult to classify. Let p t Represents the probability that the model predicts a positive value, and its formula is:
[0027]
[0028] S3042 uses text classification datasets to train and fine-tune the model, saves the model weights, and applies it to a large unlabeled corpus for prediction.
[0029] Furthermore, the step S4 specifically includes:
[0030] S401 constructs a named entity recognition training dataset, and encodes entities in the sentence using the BIO method, where each character in the sentence corresponds to an entity category.
[0031] The S4021 model uses a Dropout layer and a linear layer after the MatSciBERT encoder for sequence labeling. To further improve the accuracy of entity boundary recognition and label prediction, a Conditional Random Field (CRF) is introduced as a post-processing layer. The model is trained and fine-tuned using the constructed dataset, and the training weights are saved.
[0032] S4022 uses the trained model weights to make predictions, identify alloy materials and property entities from relevant text sentences, and extract the specific alloy components and corresponding property values in the sentences.
[0033] Furthermore, the step S5 specifically includes:
[0034] S501 constructs a relationship extraction training dataset and labels the entities extracted in step S4 in pairs. For example, if a sentence contains an alloy component and an alloy property, [E1] and [ / E1] are added at the beginning and end of the alloy component, and [E2] and [ / E2] are added at the beginning and end of the alloy property, respectively. If a sentence contains a alloy component and b specific alloy property values, a permutation and combination method is used to generate a×b sentences with the identifiers [E1], [ / E1], [E2], and [ / E2]. If the identifiers are added correctly, the sentence is marked as a positive sample; otherwise, it is marked as a negative sample.
[0035] S502 uses the relationship extraction training data set to perform model training and save the optimal training weights.
[0036] S503 performs predictions using the trained model and outputs pairing relationships between different entities.
[0037] The present invention has the following beneficial effects:
[0038] (1) Existing conventional text mining technology is often difficult to achieve good results when applied to the specific field of alloy materials because the text contains a large amount of biochemistry-related professional terms. Starting from the field of alloy materials, the present invention, based on the MatSciBERT model, trains and fine-tunes the model for the corpus in this field. It can more accurately identify relevant professional terms, extract the relationship between chemical composition and material properties, and form a structured database. The method can automatically extract the chemical composition and property data of alloys from scientific literature, forming structured data with a one-to-one correspondence between "alloy composition-alloy properties".
[0039] (2) The present invention develops an automated pipeline that inputs scientific literature and outputs specific material structured data, covering multiple stages such as article downloading, preprocessing, table parsing, text classification, named entity recognition and relationship extraction, to obtain regular structured data, which can effectively realize data-driven material design.
[0040] (3) Text classification, named entity recognition, and relationship extraction are the three core steps, all of which are rule-free algorithmic processes. High-performance data acquisition can be achieved by only building a small amount of training data. This automated information extraction method has low computational overhead and can be executed on lightweight terminal devices.
[0041] (4) By utilizing big data technology, the present invention can extract valuable information from a large amount of experimental data and literature data to assist in material design and optimization. Through machine learning and data mining technology, material properties can be quickly screened and predicted, shortening the R&D cycle.
[0042] (5) This invention can reduce the number of actual alloy R&D experiments, thereby saving experimental resources and costs. Virtual experiments and simulation calculations can replace some actual experiments, reducing R&D costs, and have important value and significance for the practical promotion and application of alloy material text mining. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 is a flow chart of the present invention;
[0044] Figure 2 This is a model architecture diagram for the sentence classification stage of the present invention;
[0045] Figure 3 This is a model architecture diagram of the entity recognition stage and relationship extraction stage of the present invention;
[0046] Figure 4 This is a comparison chart of the training effects of different loss functions in the sentence classification stage;
[0047] Figure 5 This is the training loss change graph and validation set accuracy change graph of the entity recognition stage model;
[0048] Figure 6 This is a comparison chart of the comprehensive performance of the model on three types of tasks;
[0049] Figure 7 This is an example diagram of extracting temperature structured data of single crystal superalloy solute using this method;
[0050] Figure 8 It is a comparative analysis chart of actual values and predicted values drawn based on the support vector regression model;
[0051] Figure 9 is a SHAP value distribution diagram drawn based on the extracted data in Example 2; DETAILED DESCRIPTION
[0052] To more clearly illustrate the use, effects, and advantages of the present invention, the following description will be further illustrated with reference to actual cases and accompanying figures. It should be understood that the specific examples described herein are merely illustrative of the present invention and do not limit the present invention. The present invention can be used to flexibly extract structured information from diverse material fields.
[0053] Example 1
[0054] Taking the extraction of solute temperature of different single crystal high temperature alloys as an example, Figure 1 The data extraction process of the present invention is described. A method for obtaining structural information of single crystal superalloys based on scientific literature includes the following steps:
[0055] S1 Scientific article acquisition stage: Using web big data crawler technology, more than 50,000 scientific articles on single crystal high-temperature alloys (considered as alloy corpus) were obtained from the websites of large journals such as Elsevier, Springer and Nature, and downloaded locally.
[0056] S2 preprocessing stage: Split the entire paragraph of each article into separate sentences, perform table parsing at the same time, download all table data appearing in the article to the local computer, use regular expressions to construct a rule dictionary to match the header information, obtain data containing γ′ solute temperature, and automatically organize it into an Excel table according to the format.
[0057] S3 sentence classification stage: identify sentences associated with γ′ solute temperature from the sentences divided in step S2.
[0058] S4 entity recognition stage: for the sentences associated with the solute temperature extracted in step S3, named entity recognition is performed to extract the name of the single crystal high temperature alloy and the value of the solute temperature.
[0059] S5 entity relationship extraction stage: If multiple pairs of single crystal high-temperature alloy names and solute temperature values can be identified in the sentence, this stage can be used to match the alloy names with the solute temperature values one by one.
[0060] S6 builds a solute temperature prediction model to evaluate the high-temperature properties of unknown single crystal alloys.
[0061] Specifically, step S3 sentence classification stage includes the following steps:
[0062] In step S301, the sentences segmented in step S2 are fed into MatSciBERT, and an external binary classification layer is added to define the output. A dataset of sentences with γ′ melting temperature classification is constructed. If a sentence contains the word "solvus," it is considered relevant; if not, it is considered irrelevant.
[0063] S302 uses the solute temperature sentence classification dataset to train and fine-tune the improved MatSciBERT model, finds the optimal parameters, predicts the relationship between alloy composition and properties, and then uses the fine-tuned weights to make predictions on the alloy corpus, identifying a total of more than 6,000 sentences related to the γ′ dissolution temperature.
[0064] In S303, due to the serious imbalance between the number of relevant texts and irrelevant texts in the actual extraction process, the conventional Cross-Entropy Loss is not used here, but the Focal Loss is redefined.
[0065] Cross entropy loss is often used to measure the difference between the model's predicted probability distribution and the true label distribution. i is the true label of the sample (0 or 1), p i Represents the probability that the model predicts a positive result, and N represents the number of samples. The formula is:
[0066]
[0067] Focus loss is an improvement on cross entropy loss. It introduces a dynamic weight mechanism and sets the positive sample ratio weight α t And the focusing parameter γ makes the model pay more attention to samples that are difficult to classify. Let p t Represents the probability that the model predicts a positive value, and its formula is:
[0068]
[0069] Attachment Figure 4The following graphs compare the effectiveness of two loss functions in the sentence classification stage. Figure (a) shows the curve of the loss function value changing with the number of training epochs, and Figure (b) shows the validation set accuracy for different loss functions. In the early stages of training (approximately the first 10 epochs), both loss values drop rapidly. As training progresses, both loss values tend to stabilize, but the focal loss is almost always lower than the cross-entropy loss, indicating that the focal loss performs better during training and more effectively addresses class imbalance. The MatSciBERT text classification model, fine-tuned using the cross-entropy loss function, achieved a test set accuracy of 94%, while the model fine-tuned using the focal loss function achieved a test set accuracy of 98.27%, with a recall of 1 and an F1 value of 0.9902.
[0070] Specifically, the entity recognition stage in step S4 includes the following steps:
[0071] S401 constructs a named entity recognition dataset and uses LabelStudio annotation software to encode it in BIO mode. For example, in the sentence “Moreover, it is worth noting that the solvus temperature value measured in 30Ni alloy (~1048℃) is comparable to those obtained in Co–30Ni–8Al–12V (~1032℃)
[36] and Co–30Ni–10Al–7.5W (~1050℃)
[40] quaternary alloys.”, 30Ni alloy, Co–30Ni–8Al–12V and Co–30Ni–10Al–7.5W are marked as “B-CHE”, 1048℃, 1032℃ and 1050℃ are marked as “B-TEMP”, and the rest of the sentence is uniformly marked as “O”.
[0072] S402 defines the NER output layer based on the MatSciBERT model, inputs the dataset into the model for fine-tuning, sets the AdamW optimizer, and optimizes for weight decay. At the same time, linear scheduling and warmup are used to achieve dynamic adjustment of the learning rate. Figure 5 The training loss change graph and validation set accuracy change graph of different models in the entity recognition stage.
[0073] BERT, SciBERT, and MatSciBERT are three pre-trained language models based on the Transformer architecture. Due to their different pre-training data and domain adaptability, they perform differently in entity recognition tasks. The loss curve of the entity recognition model fine-tuned based on MatSciBERT decreases most steadily, reaching a minimum after 80 epochs. Its validation accuracy remains consistently higher than that of SciBERT and BERT, reaching convergence after 60 epochs and ultimately stabilizing above 0.92.
[0074] S403 uses the fine-tuned weights to predict the relevant text and obtain the alloy name and melting temperature entities.
[0075] Specifically, step S5, the relationship extraction phase, includes the following steps:
[0076] S501 constructs a relationship extraction dataset. For example, in the above sentence, nine "alloy-temperature" relationship pairs can be formed for the three compounds 30Ni alloy, Co–30Ni–8Al–12V, and Co–30Ni–10Al–7.5W, and the three temperatures 1048°C, 1032°C, and 1050°C. A small number of sentences are manually labeled for binary classification. If a relationship exists between two labeled entity mentions, the sentence is considered positive; if no such relationship exists, the sentence is considered negative.
[0077] S502 uses the relationship extraction dataset to fine-tune the improved MatSciBERT to predict the relationships between other entity pairs.
[0078] The following table compares the performance of the models in the sentence classification, entity recognition, and entity relationship extraction stages. Finetuned-MatSciBERT performed exceptionally well across all tasks, particularly in text classification, achieving an accuracy of 98.27%. In entity recognition, Finetuned-MatSciBERT and Finetuned-SciBERT also performed exceptionally well, achieving accuracies of 93.48% and 92.85%, respectively, and exhibiting high F1 scores. For the relationship classification task, Finetuned-MatSciBERT achieved an accuracy of 90.91%, a recall of 1, and an F1 score of 0.8947.
[0079] Table 1 Comparison of model effects
[0080]
[0081] Attachment Figure 6The model's overall performance across three tasks is shown. Finetuned-SciBERT performs best on text classification, while Finetuned-MatSciBERT slightly outperforms on relation classification. Overall, the fine-tuned MatSciBERT (Finetuned-MatSciBERT) stacked graph has the highest column height and performs well across all three tasks, demonstrating excellent performance.
[0082] Through the above steps, the format of the extracted single crystal superalloy γ′ dissolution temperature dataset is as follows: Figure 7 The DOIs column indicates the source of the data, citing the article; the material column represents the alloy name; and the solvus column shows the alloy's solvus value. The material's specific elemental composition includes the mass percentages of 25 elements, including Co, Al, W, Ti, and Cr.
[0083] Example 2
[0084] For the structured data that has been obtained, further data mining modeling can be performed to construct a machine learning model, taking the construction of a solute temperature prediction model as an example.
[0085] Predicting a material's high-temperature tolerance (γ'-phase solute temperature) based on its elemental composition can be considered a regression problem. This example uses the extracted elemental composition as the feature input and the γ'-phase solute temperature as the target variable. From the extracted data, 200 complete and accurate γ'-phase solute temperature data sets, verified to be complete, were selected for the experiment. The training and test sets were divided into a ratio of 8:2, and a 5-fold cross-validation was used to construct a regression model.
[0086] Specifically, we constructed several machine learning models, including radial basis function kernel support vector regression, linear kernel support vector regression, linear regression, ridge regression, lasso regression, and gradient boosting regression, to predict the solute temperature of an alloy based on its chemical element content. Model performance was measured using three metrics: R², MSE, and rMSE. A comparison of these models is shown in the table below.
[0087] Table 2 Comparison of regression model results
[0088]
[0089] Based on the SVR model, a comparative analysis chart of actual values and predicted values is drawn as shown in the attached figure. Figure 8As shown, the model fitting effect is good. In the GBR model, the calculation of the SHAP value is based on the split contribution of different features in the tree structure. The trained model is parsed by shap.TreeExplainer to obtain the SHAP value of each feature on each sample. In order to evaluate the average impact of the feature in the global scope, this example calculates the average SHAP value of each feature (element) to quantify the contribution of each feature to the prediction result. The SHAP value distribution diagram is shown in the attached figure. Figure 9 shown.
[0090] We randomly selected some solute temperature data extracted using the method described in this invention and verified the accuracy of the extracted data by accessing the corresponding source article website through the link in the DOI column. Furthermore, Example 2 conducted subsequent modeling analysis based on the extracted data, further verifying the effectiveness of the method for obtaining material-related information based on this invention.
[0091] In summary, the present invention aims to develop a rule-free automated method that inputs scientific literature and outputs specific material structured data. It covers multiple stages such as article downloading, preprocessing, text classification, named entity recognition, and relationship extraction, and can effectively realize data-driven material design.
[0092] The above examples are only used to illustrate the specific implementation process of this method. The present invention can be used to obtain structured information in related fields such as different alloy materials. By annotating a small amount of training data, large-scale data acquisition applications can be achieved.
Claims
1. A method for acquiring alloy structured data based on scientific literature, characterized in that: The following steps are involved: S1 Scientific article acquisition: Obtain scientific articles on alloys from journal websites; S2 Data Preprocessing and Table Parsing: Split the entire text in the article into separate sentences, parse the table data in the article, obtain data on specific alloy properties, and organize them into Excel tables; S3 Sentence Classification: Identify sentences related to alloy properties from preprocessed sentences; S4 Entity Recognition: Perform named entity recognition on the sentence classification results to extract the values of alloy composition and corresponding properties; S5 entity relationship extraction: The extracted alloy components and properties are matched one by one, and then organized to form a structured database. S6 Alloy performance prediction based on extracted data: Use the extracted structured data to build a machine learning model to predict the performance indicators of the alloy and evaluate the performance of the alloy.
2. The method according to claim 1, characterized in that The pre-processing step in step S2 specifically includes: S201 Sentence Splitting: Split the entire text into separate sentences using punctuation and contextual rules; S202 Table Parsing: Parse table information in the article and match the header information through regular expressions to obtain data on specific alloy properties.
3. The method according to claim 1, characterized in that The sentence classification step in step S3 includes: S301 inputs the sentence into the MatSciBERT model; S302 loads pre-trained weights; S303 performs binary classification through a fully connected layer; S304 identifies sentences related to specific alloy properties.
4. The method according to claim 1, wherein The entity identification step in step S4 includes: S401 builds a named entity recognition training dataset and encodes entities in sentences using the BIO method; S402 uses the MatSciBERT model to perform sequence labeling tasks and identify entities of alloy composition and properties.
5. The method according to claim 1, wherein The relationship extraction step in step S5 includes: S501 builds a relation extraction dataset and manually labels sentences containing different entities; S501 uses the constructed dataset to train the MatSciBERT model, find the optimal parameters, and predict the relationship between alloy composition and properties.