Method for predicting tensile strength performance of composite material based on knowledge graph constructed by material literatures
By constructing a knowledge graph and machine learning model based on material literature, data dispersion and traditional methods in the tensile strength performance prediction of composite materials are solved, efficient and accurate tensile strength prediction is achieved, and the material R&D cycle is shortened.
Patent Information
- Application Number
- CN202510515187.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-26
AI Technical Summary
There are problems in the tensile strength performance prediction of composite materials such as data dispersion and inefficient integration, defects in traditional information extraction technology and insufficient application of knowledge graphs, which makes it difficult to guarantee data integrity and consistency and cannot achieve accurate performance prediction.
Build a knowledge graph based on material literature, label materials science literature through predefined entities and relationship types, train joint entity and relationship extraction models, extract triple structured information, use machine learning models to predict tensile strength, combine natural language processing and material calculation, and optimize data processing and feature learning.
It realizes efficient prediction of the tensile strength properties of composite materials, shortens the material R&D cycle by more than 50%, and improves data integration efficiency and prediction accuracy.
Smart Images

Figure FDA0005372366960000022
Abstract
Description
Technical Field
[0001] The present invention relates to a method for predicting the tensile strength properties of composite materials based on a knowledge graph constructed from material literature. Background Art
[0002] Predicting the tensile strength of composite materials is crucial for revealing the inherent relationships between material composition, processing technology, and mechanical properties. However, the current field of tensile strength prediction for composite materials faces the following technical bottlenecks:
[0003] 1. Data dispersion and inefficient integration: Data on the composition, process, and performance of composite materials are often scattered in isolated documents. Relying on manual collection leads to inefficiency and error-proneness, making it difficult to ensure data integrity and consistency.
[0004] 2. Deficiencies of traditional information extraction technology: Existing methods mostly use pipeline extraction (entity recognition first, then relationship extraction), which has the problem of error propagation. Errors in the entity recognition stage will directly affect the accuracy of subsequent relationship extraction.
[0005] 3. Insufficient application of knowledge graphs: Research on knowledge graphs in the field of materials science is still in its early stages. Existing graphs (such as MGED-KG and MatKG) mostly focus on the integration of basic terms and lack in-depth modeling of the entire chain relationship of "materials-processes-performance", and cannot support accurate performance prediction. Summary of the Invention
[0006] The purpose of the present invention is to propose a method for predicting the tensile strength properties of composite materials based on a knowledge graph constructed from material literature that can overcome the above problems.
[0007] To solve the above problems, the present invention provides a method for predicting the tensile strength properties of composite materials based on a knowledge graph constructed from material literature, which is characterized by comprising:
[0008] S1. Annotate relevant materials science literature according to predefined entity and relationship types to construct a materials literature dataset for model training.
[0009] S2. Using the material document dataset constructed in step S1, train a joint entity and relationship extraction model;
[0010] S3. Apply the joint extraction model trained in step S2 to process the material literature, extract triple structured information, construct a composite material knowledge graph based on it, and store it in a query-supported graph database;
[0011] S4. Identify input features related to tensile strength based on the knowledge graph constructed in step S3, and use these features to train a machine learning model to predict the tensile strength of the composite material.
[0012] As a further improvement of the present invention, the entity types predefined in step S1 include at least one or more of formula, filler, matrix, molding process, test, performance name, and performance value, and the relationship types predefined in step S1 include at least one or more of composition, process parameters, preparation method, test condition value, and test method.
[0013] As a further improvement of the present invention, the knowledge graph is a graph-based method for representing and organizing knowledge, which describes entities and their relationships within a domain through subject-predicate-object triples, so that data can be presented in the form of a graphical structure. The entities constitute the subject and object of the knowledge graph, and the relationships constitute the predicate of the knowledge graph.
[0014] As a further improvement of the present invention, step S2 includes:
[0015] S2.1. The input text is processed by the pre-trained MatSciBERT model, which maps the text into word embeddings (WE). Subsequently, these word embeddings are further processed by the self-attention mechanism to generate contextual embeddings (CE). This process enhances the semantic representation of each word embedding in the context, ensuring that each embedding more accurately reflects its semantic role in the text. Finally, CE is converted into an output vector x t :x t =MCB(X t ), t=1,2,...,n, where MCB(·) represents the conversion function based on the pre-trained model MatSciBERT;
[0016] S2.2. Vector x t is input into the feature extraction module to extract relevant features of entities and relations. The feature extraction module mainly consists of a partition filter network (PFN) and a projection layer. At each time step, the PFN receives the current input x t and the hidden state h at the previous time step t-1 , thereby generating the hidden state h of the current time step t and memory state c t :h t ,c t =F pfn (x t ,h t-1 ), where F pfn (·) Receive current input x t and the hidden state h at the previous time step t-1 , generate a new hidden state h t and memory state c t , used to extract relevant features of entities and relations;
[0017] S2.3, for c t and h t Perform projection operation to optimize the data structure. First, c t and h t Centralization is performed to reduce systematic bias in the data, which is beneficial for subsequent data processing and feature learning. Then, normalization is performed to eliminate scale differences between features, where: mean(·) represents the mean operation;
[0018] S2.4. After all time steps are completed, the memory state c is continuously updated. t , to generate the corresponding entity feature h et , relationship feature h rt and shared features h st , finally, the complete feature (h e 、h s and h r ) are stacked and input into the feature fusion module, in which the features of different tasks are concatenated and a global feature representation containing rich contextual information is generated using a multi-head attention mechanism: h ge =concat(h e ,h s ),h gr =concat(h r ,h s ), m ge =MH(h ge ),m gr =MH(h gr ), where h ge and h gr denote the global features obtained by fusing shared features and partitioned features, MH(·) denotes the multi-head attention operation, and m ge and m gr It is the global feature obtained by multi-head attention operation.
[0019] S2.5, the global feature m ge and m gr The data is sent to the task unit, which performs table filling operations and combines masking strategies to perform joint entity recognition and relationship extraction to output entity and relationship results.
[0020] As a further improvement of the present invention, step S4 includes:
[0021] S4.1. Identify input features related to the tensile strength by performing a reverse query operation on the knowledge graph constructed in step S3, and the machine learning model is an XGBoost model.
[0022] S4.2. A step of analyzing the machine learning model trained in step S4, wherein the analysis includes at least one or more of feature importance analysis, one-way sensitivity analysis (OAT) and SHAP analysis.
[0023] The beneficial effect of the present invention is that this patent adopts a full-link system of "literature mining-knowledge graph-machine learning", combines natural language processing with material calculation, breaks through the traditional experimental trial and error mode, and shortens the material research and development cycle by more than 50%. DETAILED DESCRIPTION
[0024] The technical solution of the present invention is further illustrated below through specific implementation methods.
[0025] Example 1
[0026] This embodiment includes the following steps:
[0027] S1. Annotate relevant materials science literature according to predefined entity and relationship types to construct a materials literature dataset for model training.
[0028] S2. Using the material document dataset constructed in step S1, train a joint entity and relationship extraction model;
[0029] S3. Apply the joint extraction model trained in step S2 to process the material literature, extract triple structured information, construct a composite material knowledge graph based on it, and store it in a query-supported graph database;
[0030] S4. Identify input features related to tensile strength based on the knowledge graph constructed in step S3, and use these features to train a machine learning model to predict the tensile strength of the composite material.
[0031] This embodiment adopts a full-link system of "literature mining-knowledge graph-machine learning", combining natural language processing with material calculation, breaking through the traditional experimental trial and error model, and shortening the material research and development cycle by more than 50%.
[0032] Example 2
[0033] The difference between this embodiment and embodiment 1 lies in step 2. In this embodiment, step 2 includes:
[0034] S2.1. The input text is processed by the pre-trained MatSciBERT model, which maps the text into word embeddings (WE). Subsequently, these word embeddings are further processed by the self-attention mechanism to generate contextual embeddings (CE). This process enhances the semantic representation of each word embedding in the context, ensuring that each embedding more accurately reflects its semantic role in the text. Finally, CE is converted into an output vector x t :x t =MCB(X t ), t=1,2,...,n, where MCB(·) represents the conversion function based on the pre-trained model MatSciBERT;
[0035] S2.2. Vector x t is input into the feature extraction module to extract relevant features of entities and relations. The feature extraction module mainly consists of a partition filter network (PFN) and a projection layer. At each time step, the PFN receives the current input x t and the hidden state h at the previous time step t-1 , thereby generating the hidden state h of the current time step t and memory state c t :h t ,c t =F pfn (x t ,h t-1 ), where F pfn (·) Receive current input x t and the hidden state h at the previous time step t-1 , generate a new hidden state h t and memory state c t , used to extract relevant features of entities and relations;
[0036] S2.3, for c t and h t Perform projection operation to optimize the data structure. First, c t and h t Centralization is performed to reduce systematic bias in the data, which is beneficial for subsequent data processing and feature learning. Then, normalization is performed to eliminate scale differences between features, where: mean(·) represents the mean operation;
[0037] S2.4. After all time steps are completed, the memory state c is continuously updated. t , to generate the corresponding entity feature h et , relationship feature h rt and shared features h st , finally, the complete feature (h e、h s and h r ) are stacked and input into the feature fusion module, in which the features of different tasks are concatenated and a global feature representation containing rich contextual information is generated using a multi-head attention mechanism: h ge =concat(h e ,h s ),h gr =concat(h r ,h s ), m ge =MH(h ge ),m gr =MH(h gr ), where h ge and h gr denote the global features obtained by fusing shared features and partitioned features, MH(·) denotes the multi-head attention operation, and m ge and m gr It is the global feature obtained by multi-head attention operation.
[0038] S2.5, the global feature m ge and m gr The data is sent to the task unit, which performs table filling operations and combines masking strategies to perform joint entity recognition and relationship extraction to output entity and relationship results.
[0039] In this embodiment, a joint extraction method is adopted, that is, entity recognition (NER) and relationship extraction (RE) tasks are processed simultaneously, which can more effectively utilize the interdependence between entities and relationships. This method helps to reduce error propagation and improve the overall extraction accuracy. Therefore, in order to better integrate NER and RE and optimize the accuracy and efficiency of information extraction. Specifically, this embodiment constructs a two-dimensional matrix in which rows and columns correspond to words in a sentence. By predicting the label of each cell in the matrix, the category of the entity and the relationship between the entities can be determined at the same time. In this way, not only the recognition results of entities and relationships are obtained, but these results can also be directly used to construct a knowledge graph of composite materials.
[0040] Example 3
[0041] The difference between this embodiment and Examples 1 and 2 lies in step 4. In this embodiment, step 4 includes:
[0042] S4.1. Identify input features related to the tensile strength by performing a reverse query operation on the knowledge graph constructed in step S3, and the machine learning model is an XGBoost model.
[0043] S4.2. A step of analyzing the machine learning model trained in step S4, wherein the analysis includes at least one or more of feature importance analysis, one-way sensitivity analysis (OAT) and SHAP analysis.
[0044] In this example, XGBoost tensile strength was selected for prediction because it performed best among all evaluation indicators of various other machine learning models, with an RMSE of 44.26, a MAE of 30.26, and an R 2 The prediction accuracy reached 0.96, showing extremely high prediction accuracy and excellent fitting ability. This excellent performance can be attributed to XGBoost's powerful feature selection ability and its ability to model complex nonlinear relationships, effectively capturing the interactions between input features.
[0045] The technical principles of the present invention have been described above with reference to specific embodiments. These descriptions are intended solely to illustrate the principles of the present invention and are not to be construed in any way as limiting the scope of protection of the present invention. Based on the explanations herein, those skilled in the art will readily conceive of other specific embodiments of the present invention without inventive effort, and such embodiments will fall within the scope of protection of the present invention.
Claims
1. A method for predicting the tensile strength performance of composite materials based on a knowledge graph constructed from material literature, characterized in that: include: S1. Annotate relevant materials science literature according to predefined entity and relationship types to construct a materials literature dataset for model training. S2. Using the material document dataset constructed in step S1, train a joint entity and relationship extraction model; S3. Apply the joint extraction model trained in step S2 to process the material literature, extract triple structured information, construct a composite material knowledge graph based on it, and store it in a query-supported graph database; S4. Identify input features related to tensile strength based on the knowledge graph constructed in step S3, and use these features to train a machine learning model to predict the tensile strength of the composite material.
2. The method for predicting tensile strength properties of composite materials based on a knowledge graph constructed from material literature according to claim 1, characterized in that: The entity types predefined in step S1 include at least one or more of formula, filler, matrix, molding process, test, performance name, and performance value, and the relationship types predefined in step S1 include at least one or more of composition, process parameters, preparation method, test condition value, and test method.
3. The method for predicting the tensile strength performance of composite materials based on a knowledge graph constructed from material literature according to claim 1, characterized in that: The knowledge graph is a graph-based method for representing and organizing knowledge. It describes entities and their relationships within a domain through subject-verb-object triples, so that data can be presented in the form of a graphical structure. The entities constitute the subject and object of the knowledge graph, and the relationships constitute the predicate of the knowledge graph.
4. The method for predicting tensile strength properties of composite materials based on a knowledge graph constructed from material literature according to claim 1, characterized in that: The step S2 comprises: S2.
1. The input text is processed by the pre-trained MatSciBERT model, which maps the text into word embeddings (WE). Subsequently, these word embeddings are further processed by the self-attention mechanism to generate contextual embeddings (CE). This process enhances the semantic representation of each word embedding in the context, ensuring that each embedding more accurately reflects its semantic role in the text. Finally, CE is converted into an output vector x t :x t =MCB(X t ), t=1,2,...,n, where MCB(·) represents the conversion function based on the pre-trained model MatSciBERT; S2.
2. Vector x t is input into the feature extraction module to extract relevant features of entities and relations. The feature extraction module mainly consists of a partition filter network (PFN) and a projection layer. At each time step, the PFN receives the current input x t and the hidden state h at the previous time step t-1 , thereby generating the hidden state h of the current time step t and memory state c t :h t ,c t =F pfn (x t ,h t-1 ), where F pfn (·) Receive current input x t and the hidden state h at the previous time step t-1 , generate a new hidden state h t and memory state c t , used to extract relevant features of entities and relations; S2.3, for c t and h t Perform projection operation to optimize the data structure. First, c t and h t Centralization is performed to reduce systematic bias in the data, which is beneficial for subsequent data processing and feature learning. Then, normalization is performed to eliminate scale differences between features, where: mean(·) represents the mean operation; S2.
4. After all time steps are completed, the memory state c is continuously updated. t , to generate the corresponding entity feature h et , relationship feature h rt and shared features h st , finally, the complete feature (h e 、h s and h r ) are stacked and input into the feature fusion module, in which the features of different tasks are concatenated and a global feature representation containing rich contextual information is generated using a multi-head attention mechanism: h ge =concat(h e ,h s ),h gr =concat(h r ,h s ), m ge =MH(h ge ),m gr =MH(h gr ), where h ge and h gr denote the global features obtained by fusing shared features and partitioned features, MH(·) denotes the multi-head attention operation, and m ge and m gr It is the global feature obtained by multi-head attention operation. S2.5, the global feature m ge and m gr The data is sent to the task unit, which performs table filling operations and combines masking strategies to perform joint entity recognition and relationship extraction to output entity and relationship results.
5. The method for predicting tensile strength properties of composite materials based on a knowledge graph constructed from material literature according to claim 1, characterized in that: The step S4 comprises: S4.
1. Identify input features related to the tensile strength by performing a reverse query operation on the knowledge graph constructed in step S3, and the machine learning model is an XGBoost model. S4.
2. A step of analyzing the machine learning model trained in step S4, wherein the analysis includes at least one or more of feature importance analysis, one-way sensitivity analysis (OAT) and SHAP analysis.