A Steel Potential Knowledge Inference Method and System Based on a Steel Knowledge Graph

By constructing steel knowledge graphs and training inference models, the problem that relational databases cannot integrate steel knowledge is solved, automatic reasoning of steel grades and attributes is realized, and accurate potential knowledge is provided.

CN114860889BActive Publication Date: 2025-07-29UNIV OF SCI & TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210611454.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-31
Publication Date
2025-07-29
Estimated Expiration
2042-05-31

AI Technical Summary

Technical Problem

The existing relational database cannot effectively integrate and explore potential knowledge of steel grades, resulting in incomplete information and difficult to meet the needs of engineers and experts for quickly and accurately deriving metal grades and their attributes.

Method used

Build a system based on steel knowledge graphs, build a knowledge graph, train a knowledge representation model and inference model by acquiring and extracting structured data, and use the knowledge graph to integrate steel field knowledge and conduct potential knowledge inference, including steel substitution grades, mechanical properties or chemical composition.

Benefits of technology

It realizes efficient integration and formal description of steel knowledge, can automatically discover potential relationships between steel, solves the problem of difficulty in exploring potential knowledge of steel grades, and infers accurate potential knowledge.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114860889B_ABST
    Figure CN114860889B_ABST
Patent Text Reader

Abstract

The present invention relates to a method and system for inferring potential knowledge of steel based on a steel knowledge graph, belonging to the fields of knowledge graph and steel materials. The method includes: extracting a structured steel knowledge triple dataset from existing steel data; constructing and storing a steel knowledge graph; training a knowledge representation model using the steel knowledge triples in the steel knowledge graph; training an inference model based on potential relationships based on the steel knowledge graph and the trained knowledge representation model; and performing potential knowledge inference using the trained inference model. The method of the present invention uses a knowledge graph to integrate the knowledge in the steel field and formalize its description. Then, based on the knowledge representation model, it can learn the embedding representation of the entity relationships in the steel knowledge graph in an end-to-end learning manner, thereby further modeling the relationships between known steels, and solving the problem of difficult to mine the potential knowledge of steel grades.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of knowledge graphs and steel materials, and in particular, to a method and system for inferring potential knowledge of steel materials based on a steel material knowledge graph. Background Art

[0002] Complete and accurate knowledge information about the source, properties, and approximate substitutes of metal grades is crucial for material design, reverse engineering, material procurement, processing, machining, and many other practical applications. For thousands of engineers and experts worldwide, this information usually needs to be quickly and accurately derived from the analysis and testing of potential metals, which has proven to be a complex problem in many cases.

[0003] With the continuous expansion of the scale of steel enterprises and the gradual increase of various applications, a vast amount of data and information about general steel grades has been accumulated in the field of steel materials. Most traditional material databases are established for querying basic data for scientific and technological development, material management, and use (material selection). Information such as substitute grades, chemical compositions, structures, properties, and service efficiencies related to material grades usually has imperfections, and relational databases are insufficient to provide complete information. And these data are usually stored in relational databases. In the relational data model, although the relationship between two data tables can be defined by using the primary key, this type of link is implicit rather than explicit. Moreover, the relationships in material information are complex. The substitution relationships, chemical compositions, mechanical properties, physical properties, manufacturing processes, product shapes, classifications, general uses, and other attributes of steel materials should not be isolated but interrelated. Storing steel data through relational databases has not yet unlocked all the knowledge available in the existing data.

[0004] Based on this background, users need a new form of knowledge organization to integrate multi-source heterogeneous data in the steel field and discover the hidden knowledge in steel. A knowledge graph is a data representation model, and its basic unit of composition is a triple composed of entity-relationship-entity, as well as entity and its related attribute value pairs. Entities are interconnected through relationships to form a network-like knowledge structure. By representing the semantic information of entities and relationships as dense low-dimensional real-valued vectors through the knowledge representation model, the semantic associations between entities, relationships, and between them can be efficiently calculated in the low-dimensional space, and new knowledge can be discovered. Therefore, there is an urgent need in the art to propose a method for inferring potential knowledge based on knowledge graph technology to solve the problem of difficult excavation of potential knowledge of steel grades. Summary of the Invention

[0005] The object of the present invention is to provide a method and system for reasoning about potential knowledge of steel based on a steel knowledge graph, so as to integrate the knowledge in the steel field by using the knowledge graph, and thus solve the problem of difficult excavation of potential knowledge of steel grades based on the steel knowledge graph.

[0006] To achieve the above object, the present invention provides the following solutions:

[0007] A method for reasoning about potential knowledge of steel based on a steel knowledge graph, comprising:

[0008] Obtaining existing steel data in the steel field and extracting a structured steel knowledge triple dataset from the existing steel data;

[0009] Constructing and storing a steel knowledge graph by using the structured steel knowledge triple dataset;

[0010] Training a knowledge representation model by using the steel knowledge triples in the steel knowledge graph to obtain a trained knowledge representation model;

[0011] Training an inference model based on potential relationships by using the steel knowledge graph and the trained knowledge representation model to obtain a trained inference model;

[0012] Performing potential knowledge reasoning by using the trained inference model to infer potential knowledge of steel; the potential knowledge of steel includes steel alternative grades, mechanical properties or chemical compositions.

[0013] Optionally, the obtaining existing steel data in the steel field and extracting a structured steel knowledge triple dataset from the existing steel data specifically includes:

[0014] Collecting steel grade data in the steel field from the Internet and literature manuals, and classifying them into structured data and unstructured data according to their degree of structuring, storing the structured data in the form of a two-dimensional form and the unstructured data in the form of text in a local steel database as the existing steel data;

[0015] Mapping the structured data stored in the form of a two-dimensional form in the steel database into row name-column name-data triples according to the rule that the row name of the data is the head entity, the column name is the relationship, and the data itself is the tail entity;

[0016] Extracting corresponding entity-attribute-attribute value triples from the unstructured data in the steel database by using an entity attribute extraction model;

[0017] Performing data cleaning on the row name-column name-data triples and the entity-attribute-attribute value triples to obtain corresponding structured steel knowledge triples to form the structured steel knowledge triple dataset.

[0018] Optionally, constructing and storing a steel knowledge graph by using the structured steel knowledge triple dataset specifically includes:

[0019] Based on the entities and relationships in the structured steel knowledge triple dataset, entity alignment is performed using a text similarity measurement method to eliminate ambiguity, and a steel knowledge triple dataset for constructing a steel knowledge graph is obtained;

[0020] Taking the head and tail entities of each steel knowledge triple in the steel knowledge triple dataset as nodes in the knowledge graph, and taking the relationship between the head and tail entities in the steel knowledge triple dataset as the edge in the knowledge graph, the steel knowledge graph is constructed;

[0021] Storing the steel knowledge graph in a graph database.

[0022] Optionally, training a knowledge representation model by using the steel knowledge triples in the steel knowledge graph to obtain a trained knowledge representation model, specifically including:

[0023] The steel knowledge triples in the steel knowledge graph are existing fact triples, and the head and tail entities of the fact triples are respectively replaced according to a preset probability to generate corresponding negative example triples;

[0024] Using the fact triples and the generated negative example triples to construct and train a knowledge representation model, and the knowledge representation model performs gradient update according to a loss function, and the trained knowledge representation model is obtained after reaching a specified number of training rounds.

[0025] Optionally, training an inference model based on potential relationships by using the steel knowledge graph and the trained knowledge representation model to obtain a trained inference model, specifically including:

[0026] Decomposing all relationship paths in the steel knowledge graph into triple data as a model dataset, and dividing the triple data with potential relationships in the model dataset into a validation set according to a proportion, and the remaining triple data in the model dataset as a training set;

[0027] Using the trained knowledge representation model to obtain the initial vector representation of the entities and relationships in the model dataset in a low-dimensional space;

[0028] Concatenating the initial vector representations of the entities and relationships in the training set into a matrix, using the matrix to train the inference model, and using the validation set to adjust the hyperparameters of the inference model, thereby obtaining a trained inference model.

[0029] Optionally, performing potential knowledge inference by using the trained inference model to infer steel potential knowledge, specifically including:

[0030] Based on the triple to be inferred composed of the target potential relationship and the target steel grade, use the trained inference model to score all entities in the steel knowledge graph, and identify the optimal entity with the target potential relationship with the target steel grade according to the score.

[0031] A steel potential knowledge inference system based on a steel knowledge graph, comprising:

[0032] A triple data acquisition module, configured to acquire existing steel data in the steel field and extract a structured steel knowledge triple data set from the existing steel data;

[0033] A steel knowledge graph construction module, configured to construct a steel knowledge graph using the structured steel knowledge triple data set and store it;

[0034] A knowledge representation model training module, configured to train a knowledge representation model using the steel knowledge triples in the steel knowledge graph to obtain a trained knowledge representation model;

[0035] An inference model training module, configured to train an inference model based on potential relationships based on the steel knowledge graph and the trained knowledge representation model to obtain a trained inference model;

[0036] A potential knowledge inference module, configured to perform potential knowledge inference using the trained inference model to infer steel potential knowledge; the steel potential knowledge includes steel alternative grades, mechanical properties or chemical compositions.

[0037] Optionally, the triple data acquisition module specifically includes:

[0038] A steel data acquisition unit, configured to collect steel grade data in the steel field from the Internet and literature manuals, divide it into structured data and unstructured data according to its degree of structuring, store the structured data in a two-dimensional form and the unstructured data in a text form in a local steel database as existing steel data;

[0039] A rule mapping unit, configured to map the structured data stored in the two-dimensional form in the steel database into row name-column name-data triples according to the rule that the row name of the data is the head entity, the column name is the relationship, and the data itself is the tail entity;

[0040] An entity attribute extraction unit, configured to extract corresponding entity-attribute-attribute value triples from the unstructured data in the steel database by using an entity attribute extraction model;

[0041] A data cleaning unit, configured to clean the row name-column name-data triples and entity-attribute-attribute value triples to obtain corresponding structured steel knowledge triples, which constitute the structured steel knowledge triple dataset.

[0042] Optionally, the steel knowledge graph construction module specifically includes:

[0043] An entity alignment unit, configured to perform entity alignment based on the entities and relationships in the structured steel knowledge triple dataset by using a text similarity measurement method to eliminate ambiguity, and obtain a steel knowledge triple dataset for constructing a steel knowledge graph;

[0044] A graph construction unit, configured to use the head and tail entities of each steel knowledge triple in the steel knowledge triple dataset as nodes in the knowledge graph, and use the relationship between the head and tail entities in the steel knowledge triple dataset as an edge in the knowledge graph to construct the steel knowledge graph;

[0045] A graph storage unit, configured to store the steel knowledge graph in a graph database.

[0046] Optionally, the knowledge representation model training module specifically includes:

[0047] A negative example triple generation unit, configured to use the steel knowledge triples in the steel knowledge graph as existing fact triples, and replace the head and tail entities of the fact triples respectively according to a preset probability to generate corresponding negative example triples;

[0048] A knowledge representation model construction and training unit, configured to construct and train a knowledge representation model by using the fact triples and the generated negative example triples. The knowledge representation model performs gradient update according to a loss function, and obtains the trained knowledge representation model after reaching a specified number of training rounds.

[0049] According to the specific embodiments provided by the present invention, the following technical effects are disclosed by the present invention:

[0050] The present invention provides a method and system for reasoning about latent steel knowledge based on a steel knowledge graph. The method comprises: obtaining existing steel data in the steel field and extracting a structured steel knowledge triple dataset from the existing steel data; constructing and storing a steel knowledge graph using the structured steel knowledge triple dataset; training a knowledge representation model using the steel knowledge triples in the steel knowledge graph to obtain a trained knowledge representation model; training a reasoning model based on latent relationships based on the steel knowledge graph and the trained knowledge representation model to obtain a trained reasoning model; and reasoning about latent knowledge using the trained reasoning model to infer latent steel knowledge; the latent steel knowledge includes alternative steel grades, mechanical properties, or chemical composition. The method of the present invention utilizes a knowledge graph to integrate steel domain knowledge and formalize its description. Subsequently, based on the knowledge representation model, it can learn embedded representations of entity relationships in the steel knowledge graph in an end-to-end learning manner, thereby further modeling relationships between known steels and solving the problem of difficulty in mining latent knowledge of steel grades. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0052] Figure 1 This is a flow chart of a method for reasoning about potential steel knowledge based on a steel knowledge graph according to the present invention;

[0053] Figure 2 Schematic diagram of the principle of a steel material potential knowledge reasoning method based on a steel material knowledge graph according to the present invention;

[0054] Figure 3 A schematic diagram of the process of obtaining triplet data in the steel field according to an embodiment of the present invention;

[0055] Figure 4 Schematic diagram of the steel knowledge graph constructed according to an embodiment of the present invention. DETAILED DESCRIPTION

[0056] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0057] The object of the present invention is to provide a method and system for reasoning about potential knowledge of steel based on a steel knowledge graph, which is applied to the direction of steel substitution knowledge reasoning. By using the knowledge graph to integrate knowledge in the steel field, the problem of difficult to excavate potential knowledge of steel grades is solved based on the steel knowledge graph.

[0058] In order to make the above objects, features and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0059] Figure 1 It is a flowchart of a method for reasoning about potential knowledge of steel based on a steel knowledge graph according to the present invention;

[0060] Figure 2 It is a schematic diagram of the principle of a method for reasoning about potential knowledge of steel based on a steel knowledge graph according to the present invention. Refer to Figure 1 and Figure 2 , a method for reasoning about potential knowledge of steel based on a steel knowledge graph according to the present invention includes:

[0061] Step 1: Obtain existing steel data in the steel field and extract a structured steel knowledge triple dataset from the existing steel data.

[0062] The purpose of this step 1 is to obtain triple data in the steel field, mainly by aggregating information resources related to the steel field, obtaining steel field data and extracting structured steel field knowledge triples.

[0063] Figure 3 It is a schematic diagram of the process of obtaining triple data in the steel field in an embodiment of the present invention. Refer to Figure 3 , the step 1 obtains existing steel data in the steel field and extracts a structured steel knowledge triple dataset from the existing steel data, specifically including:

[0064] Step 1.1: Obtain steel data: Collect steel grade data in the steel field from the Internet and literature manuals, and divide them into structured data and unstructured data according to their structured degree. Store the structured data in the form of a two-dimensional form and the unstructured data in the form of text in the local steel database as existing steel data; among them, the steel grade data in the steel field includes information such as alternative grades, chemical compositions, structures, property performances, service efficiencies related to steel grades, and also includes attributes such as substitution relationships, chemical compositions, mechanical properties, physical properties, manufacturing processes, product shapes, classifications, and general uses of steel materials.

[0065] Step 1.2: Rule mapping: Map the structured data stored in the form of a two-dimensional form in the steel database into row name - column name - data triples according to the rule that the row name is the head entity, the column name is the relationship, and the data itself is the tail entity.

[0066] In a specific embodiment of the present invention, the two-dimensional form is shown in Table 1 below. Using the above rules, three triple data <Y12, standard, GB / T 8731-2008>, <Y12, type, carbon steel>, and <Y12, performance, machinability> can be mapped.

[0067] Table 1 Two-dimensional data form

[0068] Steel grade Standard Type Performance Y12 GB / T8731-2008 Carbon steel Machinability

[0069] Step 1.3: Entity-attribute extraction model: Extract the corresponding entity-attribute-attribute value triples from the unstructured data in the steel database by using the entity-attribute extraction model.

[0070] Step 1.3 mainly includes the following steps:

[0071] Step 1.3.1: Manual annotation: Divide the unstructured data into an annotation candidate set according to a ratio, and manually annotate the entities, attributes, and attribute values it contains to obtain annotation samples. In the embodiment of the present invention, 1 / 5 of the sentences in the corpus are divided into the annotation candidate set, and the sentences in the annotation candidate set are manually annotated with the entities, attributes, and attribute values they contain by using the BIO method (that is, using the letter B to mark the start of the entity, using the letter I to mark the rest, and using the letter O to mark non-entities) to obtain annotation samples. For example, the result after annotating the corpus "The delivery state of C50E steel has a tensile strength of 600 Mpa" is "C / B-G 5 / I-G0 / I-G E / I-G steel / O material / O delivery / O state / O anti / B-P pull / I-P strength / I-P 6 / B-N 0 / I-N 0 / I-NM / I-N p / I-N a / I-N", which is used as an annotation sample. Among them, G, P, and N respectively represent three types of entities: steel name, attribute, and attribute value.

[0072] Step 1.3.2: Construct an entity-attribute extraction model: Split the annotation samples into a training set, a validation set, and a test set, train the entity-attribute extraction model to obtain evaluation indicators. If the indicators do not reach the threshold, continue to add the corpus to the annotation candidate set for manual annotation and retrain the model; when the threshold is reached, use the trained entity-attribute extraction model to predict the unannotated data and extract the corresponding entity-attribute-attribute value triples in the steel data.

[0073] The input of the entity attribute extraction model constructed in the present invention is unlabeled unstructured text data in the steel field, and the output of the model is the steel entities and attributes contained in the text. Specifically, the entities include steel grades, classifications, and uses; the attributes include tensile strength, yield point, elongation, and percentage elongation; and the attribute values are the specific numerical type values of the above attributes. For example, the model input can be "25Cr2MoVA is a medium-carbon alloy structural steel with high strength and toughness at room temperature, the tensile strength is 980 MPA, and it is used for high-temperature bolts of gas turbines", then the corresponding model output is "25Cr2MoVA; medium-carbon alloy structural steel; tensile strength; 980 MPA; high-temperature bolts of gas turbines".

[0074] In the embodiment of the present invention, the constructed entity attribute extraction model is an IDCNN-CRF model. The labeled samples are split into a training set, a validation set, and a test set according to a ratio of 7:1:2. The IDCNN-CRF model is trained, and the accuracy of the trained entity attribute extraction model is 84.9%, the recall rate is 80.55%, and the accuracy reaches the threshold of 80%; then the trained entity attribute extraction model is used to predict the unlabeled data, and the corresponding entity-attribute-attribute value triples in the steel data are extracted.

[0075] Step 1.4: Data cleaning: Clean the row name-column name-data triples and entity-attribute-attribute value triples to obtain the corresponding structured steel knowledge triples to form the structured steel knowledge triple dataset.

[0076] Specifically, format verification is performed on the data obtained in steps 1.2 and 1.3. In particular, the unit of the attribute values of the same attribute is unified, and fuzzy data is manually confirmed. Through this series of operations, the recognizable errors in the data are discovered and corrected, so as to obtain the corresponding structured steel knowledge triples. For example, in the embodiment of the present invention, it is necessary to unify the strength units Mpa and Kpa, and 1 Mpa = 1000 Kpa.

[0077] Step 2: Use the structured steel knowledge triple dataset to construct a steel knowledge graph and store it.

[0078] This step 2 mainly uses the triples obtained in step 1 to construct a steel knowledge graph, and stores and visually displays it through a graph database. Based on the knowledge graph technology, the knowledge entities related to steel grades are abstracted as connected network nodes, and the edges represent the relationships between the entities, which can naturally formalize the knowledge in the steel field. Further, by using a knowledge reasoning model to model the known relationships between steels, potential knowledge between steels can be automatically discovered.

[0079] The above step 2 uses the structured steel knowledge triple dataset to construct a steel knowledge graph and store it, specifically including:

[0080] Step 2.1: Entity alignment: Based on the entities and relationships in the structured steel knowledge triple dataset obtained in Step 1, use the text similarity measurement method for entity alignment to eliminate ambiguity, and obtain the steel knowledge triple dataset S for constructing the steel knowledge graph.

[0081] In the embodiment of the present invention, the Levenshtein distance is used to determine whether two entities are the same entity. Usually, two entities with a Levenshtein distance greater than 0.9 are determined to be the same entity. For example, the Levenshtein ratio of "high-quality alloy steel" and "high-quality alloy steel" is 0.91, and they can be determined to be the same entity.

[0082] Step 2.2: Graph construction: Use the steel knowledge triple dataset S obtained in Step 2.1 to construct the steel knowledge graph; specifically, use the head and tail entities of each steel knowledge triple in the steel knowledge triple dataset S as the nodes in the knowledge graph, and use the relationship between the head and tail entities in the steel knowledge triple dataset S as the edges in the knowledge graph to construct the steel knowledge graph. The scale of the steel knowledge graph constructed in the embodiment of the present invention is shown in Table 2 below:

[0083] Table 2 Scale of the steel knowledge graph

[0084] Total number of subject terms (grades) Total number of nodes Total number of relationships Number of node types Number of relationship types 11881 66849 247784 14 15

[0085] Figure 4 is a schematic diagram of the steel knowledge graph constructed in the embodiment of the present invention. As Figure 4 shown, the steel knowledge graph constructed in the embodiment of the present invention includes 14 types of entities and 15 types of relationships. The entities include: steel grade, steel alias, use, standard, standard conditions, standard description, type, classification basis, macroscopic properties, mechanical properties, product specifications, performance values, chemical composition, element content. The entity relationships are the association relationships between various entities, including: substitution relationship, application relationship, belonging relationship, basis relationship, having other name relationship, detailed description relationship, inclusion relationship, condition relationship, having type relationship, having performance relationship, product condition relationship, performance value relationship, chemical composition relationship, element content relationship, having characteristic relationship.

[0086] Step 2.3: Graph storage: Store the steel knowledge graph in a graph database. In the embodiment of the present invention, the official import tool neo4j-import of neo4j is used to store the graph in the graph database neo4j.

[0087] Step 3: Use the steel knowledge triples in the steel knowledge graph to train the knowledge representation model to obtain a trained knowledge representation model.

[0088] The trained knowledge representation model is used for initializing the vectors of entities and relationships. The knowledge representation learning model is trained using the structured steel knowledge triples in the steel knowledge graph, so as to obtain the initial vector representations of entities and relationships in the low-dimensional space in the steel triples.

[0089] Step 3 uses the steel knowledge triples in the steel knowledge graph to train the knowledge representation model to obtain a trained knowledge representation model, which specifically includes:

[0090] Step 3.1: Generation of negative example triples: The steel knowledge triples in the steel knowledge graph are existing fact triples. The head and tail entities of the fact triples are respectively replaced according to a preset probability to generate corresponding negative example triples (also called damaged triples or incorrect triples).

[0091] Specifically, the triples in the steel knowledge graph are represented in the form of (h, l, t), where h represents the head entity, l represents the relationship, and t represents the tail entity; the numbers of head and tail entities are respectively counted as N h 、N t , and the probability P is obtained. The specific formula is as follows:

[0092]

[0093] The tail entity of the triples in the steel knowledge graph is replaced according to the probability P, and the head entity is replaced according to the probability 1 - P, and it is ensured that the replaced triples are not in the steel knowledge graph, so as to obtain the negative example triple dataset S'. Its formula definition is:

[0094] S' (h,l,t) ={1 - P|(h', l, t)|h' ∈ E} ∪ {P|(h, l, t')|t' ∈ E} (2)

[0095] where E represents the entity dataset, h' and t' are randomly replaced head and tail entities, and S′ (h,l,t) is the negative example triple dataset after replacing the head and tail entities. In the embodiment of the present invention, the probability P = 30%, and 30% of the obtained negative example triples are obtained by replacing the tail entity of the fact triples (also called correct triples), and the other 70% are obtained by replacing the head entity of the fact triples.

[0096] Step 3.2: Construction and training of the knowledge representation model: The fact triples and the generated negative example triples are used to construct and train the knowledge representation model, and this knowledge representation model is based on the existing TransE model. The input of the knowledge representation model established in the present invention is all the triples in the constructed knowledge graph, and the type is in text form. For example, the content is:

[0097] C40 belongs to carbon steel

[0098] C40 C 0.2

[0099] ……

[0100] The output of this knowledge representation model is the representation vectors of entities and relationships in the knowledge graph. For example:

[0101] C40[0.0012,0..34,…]

[0102] Belong to[0.009,0.76,…]

[0103] Carbon steel[0.233,0.443,…]

[0104] C[0.876,0.265,…]

[0105] 0.2[0.35626,0.9173,…]

[0106] ……

[0107] The knowledge representation model performs gradient update according to the loss function, and obtains the trained knowledge representation model after reaching the specified number of training rounds.

[0108] Specifically, randomly initialize a vector E of a specified dimension s for the entities and relationships h, l, t in all triples h , E l , E t ; for the existing factual triples (h, l, t) in the steel knowledge graph, the distance between E h +E l and E t should be as close as possible; for the negative example triples (h, l, t) that do not exist in the steel knowledge graph, it is necessary to make the distance between E h +E l and E t quite far; for the distance measurement between vectors, the L2 norm can be selected, and the specific formula is as follows:

[0109]

[0110] where x i represents the i-th vector in x, and N represents the number of vectors in x.

[0111] Set the loss function of the knowledge representation model of the present invention as follows:

[0112]

[0113] Among them, S represents the steel knowledge triple dataset, (h, l, t) represents the existing fact triple in S, S' represents the negative example triple dataset, and (h′, l, t′) is the negative example triple generated through step 3.1. [x] + is the hinge loss function, which means taking the non - negative part of x. If x ≤ 0, then [x] + = 0. The hyper - parameter γ is a positive number, representing the margin between the scores of correct triples and incorrect triples. The knowledge representation model updates the gradient according to the loss function (3), and finally obtains the vector representations of all entities and relationships in the steel knowledge graph after reaching the specified number of training rounds.

[0114] Step 4: Train an inference model based on potential relationships using the steel knowledge graph and the trained knowledge representation model to obtain a trained inference model.

[0115] This step 4 determines the potential relationships of knowledge inference, and constructs and trains an inference model based on potential relationships using the vector representations of steel entities and relationships obtained in step 3. The inference model based on potential relationships needs to determine that the potential relationship for inference is r. In the embodiment of the present invention, the relationship r is the substitution relationship. The inference model adopts the CapsE model. The CapsE model encodes entities and relationships in the knowledge base using a capsule network, encodes entities and relationships at a deeper level, and can learn more features of triples.

[0116] The step 4 trains an inference model based on potential relationships using the steel knowledge graph and the trained knowledge representation model to obtain a trained inference model, which specifically includes:

[0117] Step 4.1: Construction of a model dataset targeted at potential relationship r: Decompose all relationship paths in the steel knowledge graph into triple data as the model dataset, and divide the triple data with potential relationships in the model dataset into a validation set according to a ratio, and the remaining triple data in the model dataset as the training set.

[0118] Specifically, retrieve all existing potential relationships r in the steel knowledge graph, and retrieve bidirectionally from r to obtain the path E0r0E1r1E2r2...E n rE n+1 where E0, E1, E2, E n ...E n+1 are entity nodes in the steel knowledge graph, E n is the steel grade node, r0, r1, r2... are the relationships between adjacent entity nodes, r is the potential relationship for inference (i.e., the substitution relationship), and n is the steel grade node E nThe length to entity E0 is greater than or equal to 0. All relationship paths E0r0E1r1E2r2...E in the steel knowledge graph n rE n+1 are decomposed into triple data E0r0E1, E1r1E2, E2r2E3…E n rE n+1 as the model dataset; and the triples E n rE n+1 with potential relationship r are divided into a small part as the validation set according to a ratio, and the rest are the training set; the inference model is trained using the training set, and the hyperparameters of the inference model are adjusted using the validation set.

[0119] Step 4.2: Use the trained knowledge representation model to obtain the initial vector representations of the entities and relationships in the model dataset in the low-dimensional space.

[0120] In the embodiment of the present invention, the vector representations of the steel entities and relationships in the model dataset are initialized to the results obtained by the knowledge representation learning model in Step 3.2.

[0121] Step 4.3: Model training stage: Concatenate the initial vector representations of the entities and relationships in the training set into a matrix, use the matrix to train the inference model, and use the validation set to adjust the hyperparameters of the inference model, so as to obtain a trained inference model.

[0122] Specifically, concatenate the initial vector representations of the triples (h, l, t) in the training set into a matrix A, and then perform convolution with 50 filters w to obtain 50 feature maps q. The formula is defined as follows:

[0123] q i = g(w·A i + b) (5)

[0124] where · is the dot product, b is the bias term, g is a non-linear activation function, such as the ReLU function, A i is the i-th row vector of the matrix A , and q i is the i-th feature map in q.

[0125] Concatenate the same dimensions of the many feature maps q obtained at the end of the convolutional layer into the first layer of capsules, and obtain the final output vector s through the dynamic routing process. The formula for the whole process is as follows:

[0126]

[0127] where u i is the capsule vector, W i is the weight matrix, b iis the hyperparameter that the first-layer capsule can learn, and soft max(·) maps the input vector to a real number between 0 and 1.

[0128] The loss function of the inference model is as follows:

[0129]

[0130] Among them,

[0131]

[0132]

[0133] Among them, similar to step 3 above, S represents the model data set targeted at the potential relationship r, and S' is the damaged triple data set generated through step 3.1 based on the model data set targeted at the potential relationship r. ||·|| is an operation of the vector two-norm, and ||·|| 2 is an operation of the square of the vector two-norm, squash(·) is the activation function in the entire capsule network, and t (h,l,t) is an intermediate parameter calculated. The inference model performs gradient update according to the loss function on the training data until the specified number of training rounds, 30, is reached, thereby obtaining the trained inference model, denoted as capsnet(·). The input of this trained inference model is the target steel grade and the target potential relationship to be inferred, and the output is a series of candidate results of the target potential relationship that the target steel grade has, sorted according to the likelihood. For example, if the input of the inference model is "Y12, substitution relationship", the corresponding output is "A576 Gr.1212, 10S20, 10SPb20, SUM21, 9S20".

[0134] Step 5: Use the trained inference model to perform potential knowledge inference and infer the potential knowledge of steel.

[0135] This step 5 uses the inference model obtained in step 4 to perform inference on potential relationships and discover potential knowledge. The potential knowledge of steel described in the present invention includes but is not limited to steel substitution grades, mechanical properties, or chemical compositions. Based on the triple to be inferred composed of the target potential relationship to be inferred and the target steel grade, use the trained inference model to score all entities in the steel knowledge graph, and identify the optimal entity having the target potential relationship with the target steel grade according to the score.

[0136] Specifically, for a given target potential relationship r to be inferred and a target steel grade E n , use the following scoring function to score all entities in the steel knowledge graph:

[0137] score() = capsnet(E n , r, E i ) | E i ∈ E(10)

[0138] where score() is the score calculated by the scoring function, capsnet(·) is the trained inference model, E represents the entity dataset in the steel knowledge graph, and E i is the i-th entity in the entity dataset E.

[0139] According to the descending order of the scores, the rankings of all entities in the steel knowledge graph among the candidate entities are obtained, so as to identify the optimal entity with the target potential relationship r with the target steel grade E n Usually, the entity with the largest score is taken as the optimal entity.

[0140] A steel potential knowledge reasoning method based on a steel knowledge graph provided by the present invention can realize reasoning of potential knowledge based on existing steel knowledge, including but not limited to steel alternative grades, mechanical properties or chemical compositions, by obtaining triple data in the steel field, constructing and storing a steel grade graph, initializing vectors of entities and relationships, establishing an inference model based on potential relationships, and using the trained inference model for potential knowledge reasoning. It is of great significance for mining potential knowledge in the steel field. The mined steel potential knowledge can be further applied to material design, reverse engineering, material procurement, processing, machining and many other practical applications, and has broad application prospects.

[0141] In a specific embodiment of the present invention, the method of the present invention is used to reason about steel alternative grades, that is, the potential relationship of knowledge reasoning is the substitution relationship. The total number of triples in the model dataset obtained through step 4.1 is 98,186, and among them, the triples E n rE n+1 with the substitution relationship r total 7,078. After the model training reaches the specified number of training rounds of 30 times, for any given two target steel grades Y12 and 45 steel, through step 5 for substitution knowledge reasoning, the obtained results are shown in Table 3 below:

[0142] Table 3 Model Inference Results

[0143] Target grade 1 2 3 4 5 Y12 A576Gr.1212 10S20 10SPb20 SUM21 9S20 45 ML45 A576Gr.1045 SWRCH45K C45 080M46

[0144] The results obtained by the inference model in Table 3 can all be found in the approximate comparison table of Chinese and foreign steel materials in the World Steel Handbook, indicating that there is indeed a substitution relationship between them. This shows that the steel knowledge inference model based on potential relationships proposed by the present invention can accurately complete the relationships of steel according to known knowledge, so as to infer potential knowledge.

[0145] Based on the independently constructed steel knowledge graph, the present invention integrates the knowledge in the steel field and constructs a relationship-based reasoning model. Without the need for manual rule design, it learns the embedding representations of entities and relationships in the steel knowledge graph in an end-to-end learning manner. By giving the potential relationships of steel grades, it can automatically complete the relationships according to the existing steel knowledge, and then infer potential knowledge. The potential knowledge in the steel field includes but is not limited to steel alternative grades, unknown mechanical properties or chemical compositions of steel.

[0146] Based on the method provided by the present invention, the present invention also provides a steel potential knowledge reasoning system based on the steel knowledge graph. The system includes:

[0147] A triple data acquisition module for acquiring existing steel data in the steel field and extracting the structured steel knowledge triple data set from the existing steel data;

[0148] A steel knowledge graph construction module for constructing and storing a steel knowledge graph by using the structured steel knowledge triple data set;

[0149] A knowledge representation model training module for training a knowledge representation model by using the steel knowledge triples in the steel knowledge graph to obtain a trained knowledge representation model;

[0150] An inference model training module for training an inference model based on potential relationships based on the steel knowledge graph and the trained knowledge representation model to obtain a trained inference model;

[0151] A potential knowledge inference module for performing potential knowledge inference by using the trained inference model to infer steel potential knowledge; the steel potential knowledge includes steel alternative grades, mechanical properties or chemical compositions.

[0152] Among them, the triple data acquisition module specifically includes:

[0153] A steel data acquisition unit for collecting steel grade data in the steel field from the Internet and literature manuals, classifying them into structured data and unstructured data according to their structured degrees, storing the structured data in the form of a two-dimensional form and the unstructured data in the form of text in a local steel database as existing steel data;

[0154] A rule mapping unit for mapping the structured data stored in the form of a two-dimensional form in the steel database into row name-column name-data triples according to the rule that the row name of the data is the head entity, the column name is the relationship, and the data itself is the tail entity;

[0155] The entity attribute extraction unit is used to extract the corresponding entity-attribute-attribute value triples from the unstructured data in the steel database by using an entity attribute extraction model;

[0156] The data cleaning unit is used to clean the row name-column name-data triples and the entity-attribute-attribute value triples to obtain the corresponding structured steel knowledge triples, which constitute the structured steel knowledge triple dataset.

[0157] The steel knowledge graph construction module specifically includes:

[0158] The entity alignment unit is used to perform entity alignment based on the entities and relationships in the structured steel knowledge triple dataset by using a text similarity measurement method to eliminate ambiguity, and obtain a steel knowledge triple dataset for constructing a steel knowledge graph;

[0159] The graph construction unit is used to construct the steel knowledge graph by taking the head and tail entities of each steel knowledge triple in the steel knowledge triple dataset as nodes in the knowledge graph and taking the relationship between the head and tail entities in the steel knowledge triple dataset as edges in the knowledge graph;

[0160] The graph storage unit is used to store the steel knowledge graph in a graph database.

[0161] The knowledge representation model training module specifically includes:

[0162] The negative example triple generation unit is used to take the steel knowledge triples in the steel knowledge graph as existing fact triples, and replace their head and tail entities respectively according to a preset probability for the fact triples to generate corresponding negative example triples;

[0163] The knowledge representation model construction and training unit is used to construct and train a knowledge representation model by using the fact triples and the generated negative example triples. The knowledge representation model performs gradient update according to a loss function, and obtains the trained knowledge representation model after reaching a specified number of training rounds.

[0164] For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For relevant parts, please refer to the description in the method section. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple. For relevant parts, please refer to the description in the method section.

[0165] In this article, specific examples are used to elaborate on the principles and implementation modes of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation modes and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.

Claims

1. A method for reasoning about potential knowledge of steel based on a steel knowledge graph, characterized in that, Including: Obtain existing steel data in the steel field and extract the structured steel knowledge triple dataset from the existing steel data; The obtaining of existing steel data in the steel field and the extraction of the structured steel knowledge triple dataset from the existing steel data specifically include: Collect steel grade data in the steel field from the Internet and literature manuals, and classify them into structured data and unstructured data according to their degree of structuring. Store the structured data in the form of a two-dimensional form and the unstructured data in the form of text in the local steel database as the existing steel data; among them, the steel grade data in the steel field includes information such as alternative grades, chemical compositions, structures, property performances, and service efficiencies related to steel grades, and also includes attributes such as substitution relationships, chemical compositions, mechanical properties, physical properties, manufacturing processes, product shapes, classifications, and general uses of steel materials; Map the structured data stored in the form of a two-dimensional form in the steel database into row name-column name-data triples according to the rule that the row name of the data is the head entity, the column name is the relationship, and the data itself is the tail entity; Extract the corresponding entity-attribute-attribute value triples from the unstructured data in the steel database by using an entity attribute extraction model; Perform data cleaning on the row name-column name-data triples and the entity-attribute-attribute value triples to obtain the corresponding structured steel knowledge triples that constitute the structured steel knowledge triple dataset; Construct and store a steel knowledge graph by using the structured steel knowledge triple dataset; The constructing and storing of a steel knowledge graph by using the structured steel knowledge triple dataset specifically include: Based on the entities and relationships in the structured steel knowledge triple dataset, use a text similarity measurement method to perform entity alignment to eliminate ambiguity, and obtain a steel knowledge triple dataset for constructing a steel knowledge graph; Construct the steel knowledge graph with the head and tail entities of each steel knowledge triple in the steel knowledge triple dataset as the nodes in the knowledge graph and the relationship between the head and tail entities in the steel knowledge triple dataset as the edges in the knowledge graph; Store the steel knowledge graph in a graph database; Train a knowledge representation model by using the steel knowledge triples in the steel knowledge graph to obtain a trained knowledge representation model; The training of a knowledge representation model by using the steel knowledge triples in the steel knowledge graph to obtain a trained knowledge representation model specifically includes: Step 3.1: Generation of negative example triples: The steel knowledge triples in the steel knowledge graph are existing fact triples. Replace the head and tail entities of the fact triples according to a preset probability respectively to generate corresponding negative example triples; Specifically, the steel knowledge triples in the steel knowledge graph are represented in the form of (h, l, t), where h represents the head entity, l represents the relationship, and t represents the tail entity; the numbers of head and tail entities are respectively counted as N h , N t , and the probability P is obtained. The specific formula is as follows: Replace the tail entity of the steel knowledge triple in the steel knowledge graph according to the probability P, and replace its head entity according to the probability 1 - P, and ensure that the replaced triple is not in the steel knowledge graph to obtain a negative example triple dataset. Its formula is defined as: S' (h,l,t) = {1 - P|(h', l, t)|h' ∈ E} ∪ {P|(h, l, t')|t' ∈ E}; where E represents the entity dataset, h' and t' are randomly replaced head and tail entities, and S' (h,l,t) is the negative example triple dataset after replacing the head and tail entities; Step 3.2: Construction and training of the knowledge representation model: Use the fact triples and the generated negative example triples to construct and train the knowledge representation model. The knowledge representation model performs gradient update according to the loss function, and obtains the trained knowledge representation model after reaching the specified number of training rounds; Specifically, randomly initialize a vector E of a specified dimension s for the entities and relationships h, l, and t in all triples. h , E l , E t ; for the existing fact triples (h, l, t) in the steel knowledge graph, the distance between E h + E l and E t should be as close as possible; for the negative example triples (h, l, t) that do not exist in the steel knowledge graph, the distance between E h + E l and E t should be quite far; for the distance metric between vectors, choose the L2 norm, and the specific formula is as follows: where xi i represents the i-th vector in x, and N represents the number of vectors in x; Set the loss function of the knowledge representation model as follows: Where \(S\) represents the steel knowledge triple dataset, \((h, l, t)\) represents the existing fact triples in \(S\), \(S'\) represents the negative example triple dataset, and \((h', l, t')\) is the negative example triple; \([x]\) + is the hinge loss function, which means taking the non - negative part of \(x\). If \(x\leq0\), then \([x]\) + \( = 0\); the hyperparameter \(\gamma\) is a positive number, representing the margin between the scores of correct triples and incorrect triples. Train an inference model based on latent relationships using the steel knowledge graph and the trained knowledge representation model to obtain a trained inference model; the inference model uses the CapsE model, and the CapsE model encodes entities and relationships in the knowledge base using a capsule network; The training of the inference model based on latent relationships using the steel knowledge graph and the trained knowledge representation model to obtain a trained inference model specifically includes: Step 4.1: Construction of a model dataset targeted at latent relationship r: Decompose all relationship paths in the steel knowledge graph into triple data as the model dataset, and divide the triple data with latent relationships in the model dataset into a validation set according to a ratio, and the remaining triple data in the model dataset as the training set; Step 4.2: Use the trained knowledge representation model to obtain the initial vector representations of entities and relationships in the low-dimensional space of the model dataset; Initialize the vector representations of steel entities and relationships in the model dataset to the results obtained by the knowledge representation learning model in Step 3.2; Step 4.3: Model training stage: Concatenate the initial vector representations of entities and relationships in the training set into a matrix, use the matrix to train the inference model, and use the validation set to adjust the hyperparameters of the inference model to obtain a trained inference model; Specifically, concatenate the initial vector representations of the triple (h, l, t) in the training set into a matrix A, and then perform convolution with 50 filters w to obtain 50 feature maps q. The formula is defined as follows: q i = g(w·A i + b); where · is the dot product, b is the bias term, g is the non-linear activation function, A i is the i-th row vector of matrix A, q i is the i-th feature map in q; Concatenate the same dimensions of many feature maps q obtained at the end of the convolutional layer into the first-layer capsule, and obtain the final output vector s through the dynamic routing process. The formula for the whole process is as follows: where u i is the capsule vector, W i is the weight matrix, b i is the hyperparameter that can be learned by the first-layer capsule, and soft max(·) maps the input vector to a real number between 0 and 1; The loss function of the inference model is as follows: where, Among them, S represents the model data set targeted at the potential relationship r, and S' is the damaged triple data set generated based on the model data set targeted at the potential relationship r through step 3.1; ||·|| is an operation of the vector two-norm, ||·|| 2 is an operation of the square of the vector two-norm, squash(·) is the activation function in the entire capsule network, and t (h,l,t) is an intermediate parameter calculated; the inference model performs gradient update according to the loss function on the training data until the specified number of training rounds, 30, is reached, thereby obtaining the trained inference model, denoted as capsnet(·); the input of the trained inference model is the target steel grade and the target potential relationship to be inferred, and the output is a series of candidate results of the target potential relationship that the target steel grade has, sorted according to the likelihood; Use the trained inference model to perform latent knowledge inference to infer steel latent knowledge; the steel latent knowledge includes steel alternative grades, mechanical properties, or chemical compositions; Based on the triple to be inferred composed of the target latent relationship to be inferred and the target steel grade, use the trained inference model to score all entities in the steel knowledge graph, and identify the optimal entity with the target latent relationship with the target steel grade according to the score; Specifically, for a given potential relationship r of the target to be inferred and the target steel grade E n , the following scoring function is used to score all entities in the steel knowledge graph: score() = capsnet(E n , r, E i ) | E i ∈ E; where score() is the score calculated by the scoring function, capsnet(·) is the trained inference model, E represents the entity dataset in the steel knowledge graph, and E i is the i-th entity in the entity dataset E; Sort in descending order according to the scores to obtain the rankings of all entities in the steel knowledge graph among the candidate entities, so as to identify the entity with the target steel grade E n The optimal entity with the target potential relationship r, and usually the entity with the largest score is taken as the optimal entity.

2. A steel potential knowledge reasoning system based on a steel knowledge graph, characterized in that, including: A triple data acquisition module for acquiring existing steel data in the steel field and extracting a structured steel knowledge triple dataset from the existing steel data; The triple data acquisition module specifically includes: The steel data acquisition unit is used to collect steel grade data in the steel field from the Internet and literature manuals, and classify it into structured data and unstructured data according to its degree of structuring. The structured data is stored in the local steel database in the form of a two-dimensional form, and the unstructured data is stored in the form of text as the existing steel data. Among them, the steel grade data in the steel field includes information such as alternative grades, chemical compositions, structures, property performances, and service efficiencies related to the steel grades, and also includes attributes such as alternative relationships, chemical compositions, mechanical properties, physical properties, manufacturing processes, product shapes, classifications, and general uses of steel materials. The rule mapping unit is used to map the structured data stored in the form of a two-dimensional form in the steel database into row name-column name-data triples according to the rule that the row name of the data is the head entity, the column name is the relationship, and the data itself is the tail entity. The entity attribute extraction unit is used to extract the corresponding entity-attribute-attribute value triples from the unstructured data in the steel database by using an entity attribute extraction model. The data cleaning unit is used to clean the row name-column name-data triples and entity-attribute-attribute value triples to obtain the corresponding structured steel knowledge triples to form the structured steel knowledge triple dataset. The steel knowledge graph construction module is used to construct and store a steel knowledge graph by using the structured steel knowledge triple dataset. The steel knowledge graph construction module specifically includes: The entity alignment unit is used to perform entity alignment based on the entities and relationships in the structured steel knowledge triple dataset by using a text similarity measurement method to eliminate ambiguities, and obtain a steel knowledge triple dataset for constructing a steel knowledge graph. The graph construction unit is used to construct the steel knowledge graph by using the head and tail entities of each steel knowledge triple in the steel knowledge triple dataset as nodes in the knowledge graph and the relationship between the head and tail entities in the steel knowledge triple dataset as edges in the knowledge graph. The graph storage unit is used to store the steel knowledge graph in a graph database. The knowledge representation model training module is used to train a knowledge representation model by using the steel knowledge triples in the steel knowledge graph to obtain a trained knowledge representation model. The knowledge representation model training module specifically includes: The negative example triple generation unit is used for Step 3.1: Negative example triple generation: Taking the steel knowledge triples in the steel knowledge graph as existing fact triples, replacing the head and tail entities of the fact triples respectively according to a preset probability to generate corresponding negative example triples. Specifically, the steel knowledge triples in the steel knowledge graph are represented in the form of (h, l, t), where h represents the head entity, l represents the relationship, and t represents the tail entity; the numbers of head and tail entities are counted as N h and N t respectively, and the probability P is obtained. The specific formula is as follows: Replace the tail entity of the steel knowledge triple in the steel knowledge graph according to the probability P, and replace its head entity according to the probability 1 - P, and ensure that the replaced triple is not in the steel knowledge graph to obtain a negative example triple dataset. Its formula is defined as: S' (h,l,t) ={1 - P|(h', l, t)|h' ∈ E} ∪ {P|(h, l, t')|t' ∈ E}; where E represents the entity dataset, h' and t' are randomly replaced head and tail entities, and S' (h,l,t) is the negative example triple dataset after replacing the head and tail entities; The knowledge representation model construction and training unit is used for Step 3.2: Knowledge representation model construction and training: Using the fact triples and the generated negative example triples to construct and train a knowledge representation model, and the knowledge representation model is updated by gradient according to the loss function, and the trained knowledge representation model is obtained after reaching the specified number of training rounds. Specifically, randomly initialize a vector E of a specified dimension s for the entities and relationships h, l, and t in all triples. h , E l , E t ; for the existing fact triples (h, l, t) in the steel knowledge graph, the distance between E h + E l and E t should be as close as possible; for the negative example triples (h, l, t) that do not exist in the steel knowledge graph, the distance between E h + E l and E t should be quite far; for the distance metric between vectors, choose the L2 norm, and the specific formula is as follows: where x i represents the i-th vector in x, and N represents the number of vectors in x; Set the loss function of the knowledge representation model as follows: Where \(S\) represents the steel knowledge triple dataset, \((h, l, t)\) represents the existing fact triples in \(S\), \(S'\) represents the negative example triple dataset, and \((h', l, t')\) is the negative example triple; \([x]\) + is the hinge loss function, which means taking the non - negative part of \(x\). If \(x\leq0\), then \([x]\) + \( = 0\); The hyperparameter \(\gamma\) is a positive number, representing the margin between the scores of correct triples and incorrect triples. The inference model training module is used to train an inference model based on potential relationships based on the steel knowledge graph and the trained knowledge representation model, and obtain a trained inference model; The inference model adopts the CapsE model, and the CapsE model uses a capsule network to encode entities and relationships in the knowledge base; The training of the inference model based on potential relationships based on the steel knowledge graph and the trained knowledge representation model to obtain a trained inference model specifically includes: Step 4.1: Construction of a model data set targeted at potential relationship r: Decompose all relationship paths in the steel knowledge graph into triple data as the model data set, and divide the triple data with potential relationships in the model data set into a validation set according to a ratio, and the remaining triple data in the model data set are used as the training set; Step 4.2: Use the trained knowledge representation model to obtain the initial vector representations of entities and relationships in the model data set in the low-dimensional space; Initialize the vector representations of steel entities and relationships in the model data set to the results obtained by the knowledge representation learning model in step 3.2; Step 4.3: Model training stage: Concatenate the initial vector representations of entities and relationships in the training set into a matrix, use the matrix to train the inference model, and use the validation set to adjust the hyperparameters of the inference model, so as to obtain a trained inference model; Specifically, concatenate the initial vector representations of the triple (h, l, t) in the training set into a matrix A, and then perform convolution with 50 filters w to obtain 50 feature maps q. The formula is defined as follows: q i = g(w·A i + b); where · is the dot product, b is the bias term, g is the non-linear activation function, A i is the i-th row vector of matrix A, q i is the i-th feature map in q; Concatenate the same dimensions of many feature maps q obtained at the end of the convolutional layer into the first layer of capsules, and obtain the final output vector s through the dynamic routing process. The formula for the whole process is as follows: where u i is the capsule vector, W i is the weight matrix, b i is the hyperparameter that can be learned by the first-layer capsule, and soft max(·) maps the input vector to a real number between 0 and 1; The loss function of the inference model is as follows: Among them, Among them, S represents the model dataset targeted at the potential relationship r, and S' is the damaged triple dataset generated based on the model dataset targeted at the potential relationship r through step 3.1; ||·|| is an operation of vector two-norm, ||·|| 2 is an operation of the square of the vector two-norm, squash(·) is the activation function in the entire capsule network, and t (h,l,t) is an intermediate parameter calculated; the inference model performs gradient update according to the loss function on the training data until the specified number of training rounds 30 is reached, so as to obtain the trained inference model, denoted as capsnet(·); the input of the trained inference model is the target steel grade and the potential relationship to be inferred, and the output is a series of candidate results of the potential relationship to be inferred possessed by the target steel grade, sorted according to the likelihood The potential knowledge inference module is used to perform potential knowledge inference using the trained inference model to infer steel potential knowledge; the steel potential knowledge includes steel alternative grades, mechanical properties or chemical compositions; Based on the triple to be inferred composed of the target potential relationship to be inferred and the target steel grade, use the trained inference model to score all entities in the steel knowledge graph, and identify the optimal entity with the target potential relationship with the target steel grade according to the score; Specifically, for a given potential relationship r of the target to be inferred and the target steel grade E n , the following scoring function is used to score all entities in the steel knowledge graph: score() = capsnet(E n , r, E i ) | E i ∈ E; where score() is the score calculated by the scoring function, capsnet(·) is the trained inference model, E represents the entity dataset in the steel knowledge graph, and E i is the i-th entity in the entity dataset E; Sort in descending order according to the scores to obtain the rankings of all entities in the steel knowledge graph among the candidate entities, so as to identify the entity that has the target potential relationship r with the target steel grade E n The optimal entity with the target potential relationship r. Usually, the entity with the largest score is taken as the optimal entity.

Citation Information

Patent Citations

  • Knowledge reasoning method and system of knowledge graph based on meta-knowledge and medium

    CN111260064A

  • Knowledge graph automatic construction method and system for steel and iron manufacturing enterprise

    CN113868432A