Converter steelmaking knowledge graph construction method and device and computer equipment
By constructing a knowledge graph for converter steelmaking, the problem of knowledge fragmentation in the converter steelmaking field has been solved, systematic integration and efficient utilization of multi-source data have been achieved, the efficiency of information retrieval and knowledge acquisition has been improved, and solid data support has been provided for converter intelligent steelmaking.
Patent Information
- Application Number
- CN202510778865.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-16
AI Technical Summary
The converter steelmaking process suffers from problems of knowledge fragmentation and lack of integration and sharing. Traditional search engines find it difficult to efficiently process and utilize massive amounts of data, resulting in an inability to provide constructive guidance and suggestions, especially when faced with new steel grades.
A knowledge graph of converter steelmaking is constructed by acquiring multi-source heterogeneous data, parsing and constructing an ontology hierarchy, and using the TF-IDF and SBERT similarity hierarchical clustering method for entity alignment and knowledge fusion, which is finally stored in the Neo4j graph database.
It has achieved systematic integration of converter steelmaking knowledge, improved information retrieval and knowledge utilization efficiency, and provided solid data support for converter intelligent steelmaking.
Smart Images

Figure CN120654794A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device and computer equipment for constructing a knowledge graph for converter steelmaking. Background Art
[0002] With the acceleration of global industrialization, the steel industry's core position as a basic material is gradually increasing. As the world's largest steel producer and consumer, the development of China's steel industry plays a vital supporting role in the national economy. Among the various steelmaking processes, converter steelmaking, thanks to its advantages such as high efficiency, energy conservation, and reduced emissions, has gradually become the mainstream process and has attracted widespread attention in the industry. However, in actual production, converter steelmaking often faces numerous problems and difficulties. The process is complex and involves multiple and interrelated operations, such as molten iron requirements, scrap requirements, converter gun position control, bottom blowing mode, and converter tapping temperature. Furthermore, the demand for efficient utilization and rapid response of real-time data in the production process is increasing. Traditional search engines and other methods are unable to effectively process and utilize massive amounts of data. Different steel grades often require different operating methods, and steel grade manuals are complex, requiring retrieval from technical manuals and other information, which affects production efficiency. With the continuous development of new steel grades, existing knowledge bases cannot cover all possible problems and scenarios, resulting in a lack of constructive guidance and reference when facing new situations.
[0003] To address the above challenges, knowledge graphs, as a structured knowledge representation and management tool, can effectively integrate and utilize a large amount of scattered technical data and empirical knowledge in the converter steelmaking process, and understand the possible similar attributes between different steel grades through multi-intent question-answering.
[0004] Most of the research on industrial vertical knowledge graphs focuses on equipment fault repair, workpiece processing, assembly, etc. At present, there is no complete, usable and high-quality knowledge graph in the converter steelmaking vertical field. Summary of the Invention
[0005] Based on this, it is necessary to provide a method, device and computer equipment for constructing a converter steelmaking knowledge graph in order to address the problems in the converter steelmaking field, such as the converter steelmaking process usually involving many complex processes and technical specifications, knowledge fragmentation, lack of integration and sharing.
[0006] A method for constructing a converter steelmaking knowledge graph, the method comprising: Acquire multi-source heterogeneous data in the field of converter steelmaking; multi-source heterogeneous data includes: unstructured data and semi-structured data; unstructured data includes: converter steelmaking instruction manuals and national and enterprise standard documents; semi-structured data includes enterprise databases and public patent information.
[0007] Multi-source heterogeneous data are parsed and a converter steelmaking ontology hierarchy is constructed based on the obtained domain corpus characteristics; the ontology hierarchy includes ontology concept and class attribute definitions, ontology relationship definitions, and ontology instantiation.
[0008] According to the hierarchical structure of converter steelmaking ontology, corresponding methods are used to extract entities, entity attributes and entity relationships from different types of data to obtain a set of knowledge triples.
[0009] According to the knowledge triple set, the similarity hierarchical clustering entity alignment method based on TF-IDF and SBERT is used to perform knowledge fusion.
[0010] The knowledge triple set after knowledge fusion is stored in the Neo4j graph database to complete the construction of the converter steelmaking knowledge graph.
[0011] A converter steelmaking knowledge graph construction device, the device comprising: The multi-source heterogeneous data acquisition module is used to acquire multi-source heterogeneous data in the field of converter steelmaking; multi-source heterogeneous data includes: unstructured data and semi-structured data; unstructured data includes: converter steelmaking guidance manuals and national and enterprise standard documents; semi-structured data includes enterprise databases and public patent information.
[0012] The converter steelmaking ontology hierarchy construction module is used to parse multi-source heterogeneous data and construct the converter steelmaking ontology hierarchy based on the obtained domain corpus characteristics; the ontology hierarchy includes ontology concept and class attribute definitions, ontology relationship definitions, and ontology instantiation.
[0013] The knowledge extraction module is used to extract entities, entity attributes and entity relationships from different types of data using corresponding methods according to the hierarchical structure of the converter steelmaking ontology to obtain a set of knowledge triples.
[0014] The knowledge fusion module is used to perform knowledge fusion based on the knowledge triple set using the similarity hierarchical clustering entity alignment method based on TF-IDF and SBERT.
[0015] The knowledge graph storage module is used to store the knowledge triple set after knowledge fusion using the Neo4j graph database to complete the construction of the converter steelmaking knowledge graph.
[0016] A computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0017] The aforementioned method, apparatus, and computer equipment for constructing a knowledge graph for converter steelmaking include: constructing an ontology hierarchy based on collected multi-source heterogeneous data; then extracting knowledge from different types of data based on the ontology hierarchy; and processing and integrating converter steelmaking knowledge using a hierarchical clustering method based on similarity calculation using TF-IDF and SBERT to ensure the data quality of the knowledge graph. This systematic integration of multi-source knowledge, scattered across technical documents, operating manuals, and enterprise information databases, forms a structured knowledge base, positively impacting information retrieval, knowledge utilization, and acquisition processes for converter steelmaking, providing solid knowledge data support for the subsequent implementation of intelligent converter steelmaking. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 A schematic diagram of a process for constructing a converter steelmaking knowledge graph in one embodiment; Figure 2 Constructing a flow chart for a converter steelmaking knowledge graph in one embodiment; Figure 3 Schematic diagram of the converter steelmaking knowledge graph ontology structure in one embodiment; Figure 4 is an entity relationship structure diagram in another embodiment; Figure 5 is a flow chart of a rule-based knowledge extraction method in another embodiment; Figure 6 Schematic diagram of semi-structured data after rule template extraction in another embodiment; Figure 7 A flowchart of knowledge extraction using a CMeKG tool in another embodiment; Figure 8 1 is a pseudo code diagram of a similarity hierarchical clustering entity alignment algorithm based on TF-IDF and SBERT in another embodiment; Figure 9 A heat map of the similarity matrix based on TF-IDF and SBERT in another embodiment; Figure 10 A schematic diagram of hierarchical clustering results in another embodiment; Figure 11 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0019] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0020] Most of the research on knowledge graphs in industrial vertical fields focuses on equipment fault repair, workpiece processing, assembly, etc. Therefore, constructing a high-quality knowledge graph specifically for converter steelmaking, systematically integrating multi-source knowledge scattered in technical documents, operating manuals, enterprise information databases, etc., to form a structured knowledge base, is of great significance to promoting technological progress and innovation of my country's converters.
[0021] In one embodiment, Figure 1 As shown, a method for constructing a converter steelmaking knowledge graph is provided, which includes the following steps: Step 100: Acquire multi-source heterogeneous data in the field of converter steelmaking; multi-source heterogeneous data includes: unstructured data and semi-structured data; unstructured data includes: converter steelmaking instruction manuals and national and enterprise standard documents; semi-structured data includes enterprise databases and public patent information.
[0022] Specifically, due to the severe fragmentation of knowledge in the converter steelmaking field, relevant knowledge is dispersed across multiple data sources, including unstructured and semi-structured data. Unstructured data primarily includes converter steelmaking manuals and national and enterprise standard documents, while semi-structured data includes enterprise databases and publicly available patent information. This paper employs methods to extract knowledge from these different types of data based on their structural characteristics.
[0023] Step 102: parse the multi-source heterogeneous data and construct a converter steelmaking ontology hierarchical structure based on the obtained domain corpus features; the ontology hierarchical structure includes ontology concept and class attribute definitions, ontology relationship definitions, and ontology instantiation.
[0024] Specifically, the industrial type of converter steelmaking is highly specialized and complex, involving numerous precise technical aspects such as steel grade standards, converter smelting operations, and process flow requirements. This method uses a top-down approach to build the ontology starting from the top-level concepts to ensure the accuracy and consistency of the knowledge graph. The top-down construction method relies on expert knowledge to analyze structured knowledge, building the ontology starting from the top-level concepts, and continuously optimizing concepts and relationships to ultimately form a high-quality knowledge hierarchy. The top-down method can ensure high accuracy and fast construction, and is suitable for the construction of knowledge graphs in vertical fields, reducing noise and ensuring the quality and consistency of knowledge.
[0025] The process of building a knowledge graph for converter steelmaking is as follows: Figure 2 shown.
[0026] The hierarchical structure of converter steelmaking ontology is constructed according to the characteristics of domain corpus, including the definition of concepts / classes, relationships and attributes, and the clarification of knowledge extraction boundaries.
[0027] Ontology hierarchical construction is a crucial step in knowledge engineering. Its core goal is to systematically define and organize domain concepts, attributes, and their relationships, thereby constructing a shared, formalized knowledge model. It not only provides a standardized framework for expressing domain knowledge but also provides a theoretical foundation for knowledge sharing, exchange, and reasoning. The seven-step approach employed in ontology construction allows for a more systematic and comprehensive establishment of an ontology's hierarchical structure.
[0028] Step 104: According to the hierarchical structure of the converter steelmaking ontology, corresponding methods are used to extract entities, entity attributes, and entity relationships from different types of data to obtain a set of knowledge triples.
[0029] Specifically, for semi-structured data in the form of web pages, web crawlers are used to crawl information. For unstructured data, data parsing and cleaning are first performed to remove noise and irrelevant content. Knowledge is then extracted using rule-based knowledge extraction methods or the CMeKG open-source tool. Entities, relationships, and attributes are then extracted based on the constructed ontology layer structure.
[0030] Step 106: Based on the knowledge triple set, a similarity hierarchical clustering entity alignment method based on TF-IDF and SBERT is used to perform knowledge fusion.
[0031] Specifically, entities are divided according to different label types. For entity clusters with different labels, the similarity matrix is calculated by combining TF-IDF and SBERT, and then the entity alignment operation is achieved through hierarchical clustering.
[0032] Step 108: The knowledge triple set after knowledge fusion is stored in the Neo4j graph database to complete the construction of the converter steelmaking knowledge graph.
[0033] Specifically, the triple data after knowledge fusion is stored in the Neo4j graph database to provide data support for the converter steelmaking question-answering system.
[0034] Neo4j, with its native graph storage structure and efficient associative query capabilities, provides an ideal solution for organizing and visualizing complex relational knowledge. Neo4j employs an entity-relationship model of nodes, relationships, and properties, enabling direct mapping to real-world semantic networks. This model is particularly well-suited for rapidly traversing multi-hop relationships within knowledge graphs. Its built-in Cypher query language supports declarative graph pattern matching, significantly simplifying the logic for expressing complex associative queries. Therefore, this method selected Neo4j as a high-performance graph database to store the BOF steelmaking knowledge fusion data. Neo4j's LOAD CSV command was used to batch import structured relationships, while the Python py2neo library was leveraged to dynamically write entity attribute information.
[0035] The aforementioned method for constructing a knowledge graph for converter steelmaking includes constructing an ontology hierarchy based on collected multi-source heterogeneous data, extracting knowledge from different types of data based on the ontology hierarchy, and processing and integrating converter steelmaking knowledge using a hierarchical clustering method based on similarity calculation using TF-IDF and SBERT to ensure the data quality of the knowledge graph. This systematic integration of multi-source knowledge, scattered across technical documentation, operating manuals, and enterprise information databases, forms a structured knowledge base, positively impacting information retrieval, knowledge utilization, and acquisition processes for converter steelmaking, providing solid knowledge data support for the subsequent implementation of intelligent converter steelmaking.
[0036] In one embodiment, the ontology concept and class attribute definition in step 102 include: seven ontology concepts and corresponding class attributes; the seven ontology concepts include: steel type, process parameter class, converter smelting operation class, steel grade standard class, steel grade characteristic class, deoxidation alloy operation class, and refined steel liquid condition class; the class attributes of steel type include: steel grade code, steel grade name, and steel grade brand; the class attributes of process parameter class include: process flow, molten iron requirements, and scrap steel requirements; the class attributes of converter smelting operation class include: converter gun position control, bottom blowing mode, and making Slag system, converter tapping temperature, liquidus temperature; the class attributes of steel grade standard class include: release upper limit, release lower limit, agreement upper limit, agreement lower limit, standard upper limit, standard lower limit, included elements, element targets; the class attributes of steel grade characteristic class include: steel grade definition, steel grade characteristics, steel grade use; the class attributes of deoxidation alloy operation class include: deoxidation alloy system, added alloys, element corresponding alloy addition, precautions; the class attributes of refined steel liquid condition class include: elements for refined steel liquid, range of elements for refined steel liquid, and target of elements for refined steel liquid.
[0037] Specifically, considering the specific characteristics of the converter steelmaking domain, the core knowledge in this field can be abstracted into seven ontological concepts. These concepts cover the key elements and operational processes in the converter steelmaking process, providing a complete and systematic framework for constructing a converter steelmaking knowledge graph. These seven ontological concepts are: steel type, process parameter, converter smelting operation, steel grade standard, steel grade characteristic, deoxidation alloy operation, and molten steel supply conditions. They meet the needs of the converter steelmaking knowledge graph ontology library. Steel type serves as the core node of the converter steelmaking knowledge graph, with its child nodes being process parameter, converter smelting operation, steel grade standard, steel grade characteristic, deoxidation alloy operation, and molten steel supply conditions. Each ontological concept corresponds to multiple class attributes. For example, the converter smelting operation class includes class attributes for converter lance position control, bottom blowing mode, slag formation system, converter tapping temperature, and liquidus temperature, as shown in Table 1.
[0038] Table 1 Ontology concepts and class attribute definitions
[0039] In one embodiment, the ontology relationship definition in step 102 includes: subordinate relationship, inclusion relationship, influence relationship, dependency relationship, definition relationship, and similarity relationship.
[0040] Specifically, properly defined ontology relationships can connect different ontology concepts, providing a more complete and meaningful representation framework for the information in the knowledge base. The relationships between ontology concepts are primarily manifested in several aspects: subclass-of, part-of, influence, dependency, define, and similarity. Table 2 describes the relationships between the core node "Steel Type" and its six other child nodes, providing guidance for the subsequent ontology instantiation process.
[0041] Table 2 Ontology relationship definition
[0042] Finally, the converter steelmaking knowledge graph ontology structure constructed by this method is as follows Figure 4 As shown in Figure 1, it contains a total of 7 core classes, 6 subclasses, 6 object relationships, and 26 class attributes. By constructing ontology concepts, ontology relationships, and subdividing ontology attributes, we can provide a framework for the construction of entities, relationships, and attributes in the subsequent ontology instantiation process.
[0043] In one embodiment, ontology instantiation specifically includes: instantiating the ontology concept into six entities; the six entities include: steel type code entity, process parameter entity, smelting operation entity, steel type standard entity, deoxidation alloy operation entity, and refined steel condition entity; defining three entity relationships under the influence relationship between steel type and process parameter ontology; defining five entity relationships under the dependency relationship between steel type and converter smelting operation ontology; defining fourteen entity relationships under the inclusion relationship between steel type and steel type standard ontology; defining four entity relationships under the influence relationship between steel type and deoxidation alloy operation ontology; and defining five entity relationships under the influence relationship between steel type and refined steel molten condition ontology.
[0044] Specifically, combining the constructed ontology concepts and class attribute characteristics, this method instantiates the ontology concepts into six entities, as shown in Table 3. The steel grade code, as an ontological attribute of the steel type, possesses uniqueness and retrieval properties and is therefore defined as an entity. While the steel grade name and steel grade brand are used to distinguish different steel types, their role in distinguishing different steel types is relatively limited. Therefore, they do not exist as independent entities but as attributes of the steel grade code entity. Furthermore, considering that the steel grade definition, steel grade characteristics, and steel grade usage information in the steel grade feature class primarily describe specific steel grades and are lengthy, they are not suitable as separate entities. Therefore, this information is stored as attributes of the steel grade code entity. Other ontological concepts can be directly instantiated into corresponding entities.
[0045] Table 3 Entity types in the converter steelmaking field
[0046] Table 4 defines and divides the five ontological relationships in a more fine-grained manner. Three entity relationships are defined under the affect-craftwork ontological relationship; five entity relationships are defined under the depend-on-operate ontological relationship; fourteen entity relationships are defined under the has-benchmark ontological relationship; four entity relationships are defined under the affect-deoxyalloy ontological relationship; and five entity relationships are defined under the affect-steelmaking ontological relationship. Through this process, a total of 31 different entity relationships are defined. These refined relationships provide the basis for more precise ontological instantiation and multi-intent parsing, ensuring the comprehensiveness and flexibility of the knowledge graph in the field of converter steelmaking. Finally, the structure of entities and relationships is as follows: Figure 3 shown.
[0047] Table 4 Relationship types in converter steelmaking
[0048] In one embodiment, step 104 includes: according to the hierarchical structure of the converter steelmaking ontology, for semi-structured data in the form of web pages, using web crawler technology to perform entity extraction, entity attribute extraction, and entity relationship extraction; for unstructured data, using a rule-based method or a CMeKG model to perform entity extraction, entity attribute extraction, and entity relationship extraction to obtain a set of knowledge triples.
[0049] In one embodiment, the steps of the rule-based method include: Step S41: Implement the conversion operation between PDF file format and TXT text data through the pdfplumber library.
[0050] Step S42: Filter out meaningless information by defining a regularized expression; meaningless information includes page numbers and comments.
[0051] Step S43: Segment the text data according to the material attributes of different steel grades.
[0052] Step S44: define a rule module to perform preliminary knowledge extraction on different tables and text information to obtain semi-structured data.
[0053] Step S45: If the semi-structured data contains noise, the process proceeds to step S42 to continue the knowledge extraction; if the semi-structured data contains no noise, the knowledge extraction is completed.
[0054] Specifically, converter steelmaking knowledge, such as instruction manuals and process standards, is usually stored in PDF format. In order to make the expression more intuitive and easier for operators to understand, many steel grade attribute information is expressed in table form, which increases the difficulty of knowledge extraction tools and deep learning methods. Therefore, this paper adopts a rule-based method to extract knowledge from converter steelmaking instruction manuals. The specific extraction process is shown in Figure 5 As shown in the figure, we first use the pdfplumber library to convert between PDF file formats and TXT text data. Compared to libraries like pdfminer and PyPDF2, pdfplumber's data parsing method can extract text content from tables while maintaining the integrity of the table's structure, facilitating the design of subsequent rule templates. We then define a regular expression to filter out meaningless information, including page numbers and comments.
[0055] After obtaining the semi-structured data after the initial extraction through the rule template, such as Figure 6As shown, each steel grade attribute needs to be directly linked to the relevant information of the steel grade code entity type. For example, in the steel grade standard and internal control chemical composition table, the standard upper limit value for element C is 0.18. Establishing only a relationship between the "element C" and "standard upper limit value 0.18" entities is meaningless. Only by linking specific attribute values with the relevant steel grade can it be meaningful, such as establishing triples (B251A, code_standard_cap, standard upper limit value for element C is 0.18) and (B251A, code_standard_cap, standard upper limit value for element Mn is 1.50). Secondly, within the steel grade concept class, only the steel grade code is uniquely searchable. A steel grade name or steel grade brand may correspond to multiple steel grade codes, and the specific operations corresponding to each code are different. For example, the steel grade named "carbon structural steel" may contain steel grade codes such as B251A and B163B. In actual queries, one-to-many queries may occur, which is not conducive to the subsequent construction of the converter steelmaking question-answering system.
[0056] In one embodiment, for unstructured data, after preliminary cleaning and parsing using pdfplumber and regular expressions, if there is a lack of clear logical relationship between the texts, the CMeKG model is used to extract entities, entity attributes, and entity relationships.
[0057] Specifically, this paper focuses on the basic idea of entity and relationship extraction in the converter steelmaking field, and adopts the CMeKG (Chinese Medical Knowledge Graph) model for knowledge extraction. The improved model mainly includes two parts: the subject position prediction model (Model4s) and the object and relationship prediction model (Model4po). Figure 7 As shown in Figure 3, Model4po first uses the subject's starting position representation vector and the ending position representation vector to obtain a complete subject feature representation through splicing. This subject feature is then fused with the hidden state features of each position in the text to obtain the contextual feature representation of each position relative to the subject, ultimately achieving the prediction of the subject-object relationship and the object's position.
[0058] For unstructured data, after using pdfplumber and regular expressions for preliminary cleaning and parsing, if there is a lack of clear logical relationship between the texts, it is difficult to directly achieve effective knowledge extraction by defining rule templates. The method based on deep learning usually requires a large amount of labeled data, and the labeling workload is huge. In order to reduce the cost of manual labeling, this embodiment only labels a small amount of data, and uses the knowledge data extracted based on the rule template in the above text to construct a training sample to fit the knowledge expression form in the unstructured text, thereby training the CMeKG model so that the model has a more general knowledge extraction capability. The training set contains 68,640 data, and the test set contains 17,160 data. The final model has an F1 value of 86.69%, an accuracy rate of 95.45%, and a recall rate of 79.41% on the test set, and in order to ensure the quality of the extracted triples, it is also necessary to perform a large model and manual review operations on it.
[0059] In one embodiment, step 106 includes: calculating the TF-IDF vector and SBERT embedding vector of each entity in the knowledge triple set using the TF-IDF method; traversing all entity texts, respectively calculating the TF-IDF similarity and SBERT semantic similarity based on cosine similarity, linearly fusing each TF-IDF similarity and the corresponding SBERT semantic similarity through a weight parameter to generate a comprehensive similarity matrix; converting the comprehensive similarity matrix into a distance matrix; using a hierarchical clustering model with pre-calculated affinity based on the distance matrix, completing entity cluster division with an average link strategy, obtaining an entity cluster set, and implementing an entity alignment operation; the entity cluster set includes multiple entity clusters.
[0060] Specifically, in the bottom blowing mode of converter steelmaking, "adopting full-time argon blowing to switch to bottom blowing mode" and "adopting full-time argon blowing bottom blowing mode" refer to the same entity, but are stored as two entities due to different forms of expression. This entity with repeated meaning will not only cause data redundancy in the knowledge graph, but also cause problems such as semantic confusion. Therefore, when constructing the converter steelmaking knowledge graph, it is necessary to avoid generating duplicate entities and to fuse different expressions of the same knowledge before storing and representing them. In view of the limitations of traditional text similarity calculation methods that cannot fully balance key terms and semantic information in text, this application proposes a similarity hierarchical clustering entity alignment algorithm based on TF-IDF and SBERT. The specific algorithm is as follows: Figure 8 shown.
[0061] After extracting knowledge about converter steelmaking, this paper first performs a more detailed classification of entities based on their categories and operational attributes, such as bottom blowing mode, hot metal requirements, and scrap requirements. TF-IDF is then used to dynamically adjust weights based on entity characteristics, allowing the algorithm to prioritize these key terms during entity alignment. Simultaneously, Sentence-BERT (SBERT) provides rich contextual semantic representations to capture the deep semantic connections within the text. This fusion of the two not only ensures fine-grained attention to relational terms within the text but also fully leverages overall semantic information, enabling dynamic entity alignment across entities of different categories.
[0062] In the knowledge fusion process, we first calculate the TF-IDF vector and SBERT embedding vector for each entity in the text collection; then we traverse all entity texts and calculate the TF-IDF similarity based on cosine similarity. Semantic similarity with SBERT , and through the weight parameter Perform linear fusion to generate a comprehensive similarity matrix, such as Figure 9 The heat map of the similarity matrix between entity texts in the bottom-blowing mode is shown in Figure 2. The similarity matrix is then converted into a distance matrix, and the agglomerative clustering model (pre-calculated affinity) is used to complete the entity clustering with an average link strategy. Finally, the algorithm outputs a set of k entity clusters, thereby achieving entity alignment operations, such as Figure 10 is the clustering result of the bottom blowing mode, where , k=3.
[0063] It should be understood that although Figure 1 The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. In addition, Figure 1 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.
[0064] In one embodiment, a device for constructing a converter steelmaking knowledge graph is provided, comprising: a multi-source heterogeneous data acquisition module, a converter steelmaking ontology hierarchical structure construction module, a knowledge extraction module, a knowledge fusion module, and a knowledge graph storage module, wherein: The multi-source heterogeneous data acquisition module is used to acquire multi-source heterogeneous data in the field of converter steelmaking; multi-source heterogeneous data includes: unstructured data and semi-structured data; unstructured data includes: converter steelmaking guidance manuals and national and enterprise standard documents; semi-structured data includes enterprise databases and public patent information.
[0065] The converter steelmaking ontology hierarchy construction module is used to parse multi-source heterogeneous data and construct the converter steelmaking ontology hierarchy based on the obtained domain corpus characteristics; the ontology hierarchy includes ontology concept and class attribute definitions, ontology relationship definitions, and ontology instantiation.
[0066] The knowledge extraction module is used to extract entities, entity attributes and entity relationships from different types of data using corresponding methods according to the hierarchical structure of the converter steelmaking ontology to obtain a set of knowledge triples.
[0067] The knowledge fusion module is used to perform knowledge fusion based on the knowledge triple set using the similarity hierarchical clustering entity alignment method based on TF-IDF and SBERT.
[0068] The knowledge graph storage module is used to store the knowledge triple set after knowledge fusion using the Neo4j graph database to complete the construction of the converter steelmaking knowledge graph.
[0069] In one embodiment, the ontology concepts and class attribute definitions in the converter steelmaking ontology hierarchy construction module include: seven ontology concepts and corresponding class attributes; the seven ontology concepts include: steel type, process parameter class, converter smelting operation class, steel grade standard class, steel grade feature class, deoxidation alloy operation class, and refined steel liquid condition class; the class attributes of steel type include: steel grade code, steel grade name, and steel grade brand; the class attributes of process parameter class include: process flow, molten iron requirements, and scrap steel requirements; the class attributes of converter smelting operation class include: converter gun position control, bottom blowing, etc. Mode, slag making system, converter tapping temperature, liquidus temperature; the class attributes of steel grade standard class include: release upper limit, release lower limit, agreement upper limit, agreement lower limit, standard upper limit, standard lower limit, included elements, element target; the class attributes of steel grade characteristic class include: steel grade definition, steel grade characteristics, steel grade use; the class attributes of deoxidation alloy operation class include: deoxidation alloy system, alloy addition, element corresponding alloy addition, precautions; the class attributes of refined steel liquid condition class include: elements for refined steel liquid, range of elements for refined steel liquid, and target of elements for refined steel liquid.
[0070] In one embodiment, the ontology relationship definitions in the converter steelmaking ontology hierarchical structure construction module include: subordinate relationship, inclusion relationship, influence relationship, dependency relationship, definition relationship and similarity relationship.
[0071] In one embodiment, ontology instantiation specifically includes: instantiating the ontology concept into six entities; the six entities include: steel type code entity, process parameter entity, smelting operation entity, steel type standard entity, deoxidation alloy operation entity, and refined steel condition entity; defining three entity relationships under the influence relationship between steel type and process parameter ontology; defining five entity relationships under the dependency relationship between steel type and converter smelting operation ontology; defining fourteen entity relationships under the inclusion relationship between steel type and steel type standard ontology; defining four entity relationships under the influence relationship between steel type and deoxidation alloy operation ontology; and defining five entity relationships under the influence relationship between steel type and refined steel molten condition ontology.
[0072] In one embodiment, the knowledge extraction module is also used to perform entity extraction, entity attribute extraction, and entity relationship extraction on semi-structured data in the form of web pages using web crawler technology based on the hierarchical structure of the converter steelmaking ontology, and to perform entity extraction, entity attribute extraction, and entity relationship extraction on unstructured data using a rule-based method or CMeKG model to obtain a set of knowledge triples.
[0073] In one embodiment, the steps of the rule-based method in the knowledge extraction module include: Step S41: Implement the conversion operation between PDF file format and TXT text data through the pdfplumber library.
[0074] Step S42: Filter out meaningless information by defining a regularized expression; meaningless information includes page numbers and comments.
[0075] Step S43: Segment the text data according to the material attributes of different steel grades.
[0076] Step S44: define a rule module to perform preliminary knowledge extraction on different tables and text information to obtain semi-structured data.
[0077] Step S45: If the semi-structured data contains noise, the process proceeds to step S42 to continue the knowledge extraction; if the semi-structured data contains no noise, the knowledge extraction is completed.
[0078] In one embodiment, the knowledge extraction module is also used to perform preliminary cleaning and parsing of unstructured data using pdfplumber and regular expressions. If there is a lack of clear logical relationship between the texts, the CMeKG model is used to perform entity extraction, entity attribute extraction, and entity relationship extraction.
[0079] In one embodiment, the knowledge fusion module is also used to calculate the TF-IDF vector and SBERT embedding vector of each entity in the knowledge triple set using the TF-IDF method; traverse all entity texts, calculate the TF-IDF similarity and SBERT semantic similarity based on cosine similarity, linearly fuse each TF-IDF similarity and the corresponding SBERT semantic similarity through weight parameters to generate a comprehensive similarity matrix; convert the comprehensive similarity matrix into a distance matrix; adopt a hierarchical clustering model with pre-calculated affinity based on the distance matrix, complete the entity cluster division with an average link strategy, obtain an entity cluster set, and implement entity alignment operation; the entity cluster set includes multiple entity clusters.
[0080] For the specific limitations of the converter steelmaking knowledge graph construction device, please refer to the limitations of the converter steelmaking knowledge graph construction method above, which will not be repeated here. The various modules in the above-mentioned converter steelmaking knowledge graph construction device can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0081] In one embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as follows: Figure 11 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for constructing a converter steelmaking knowledge graph is implemented. The display screen of the computer device can be a liquid crystal display or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a key, trackball or touchpad provided on the computer device housing, or an external keyboard, touchpad or mouse.
[0082] Those skilled in the art will understand that Figure 11 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0083] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiment when executing the computer program.
[0084] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0085] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A method for constructing a converter steelmaking knowledge graph, characterized in that: The method comprises: Acquire multi-source heterogeneous data in the field of converter steelmaking; the multi-source heterogeneous data includes: unstructured data and semi-structured data; the unstructured data includes: converter steelmaking instruction manuals and national and enterprise standard documents; the semi-structured data includes enterprise databases and public patent information; Parsing the multi-source heterogeneous data, and constructing a converter steelmaking ontology hierarchical structure based on the obtained domain corpus features; the ontology hierarchical structure includes ontology concept and class attribute definitions, ontology relationship definitions, and ontology instantiation; According to the hierarchical structure of converter steelmaking ontology, corresponding methods are used to extract entities, entity attributes and entity relationships from different types of data to obtain a set of knowledge triples. Based on the knowledge triple set, a similarity hierarchical clustering entity alignment method based on TF-IDF and SBERT is used to perform knowledge fusion; The knowledge triple set after knowledge fusion is stored in the Neo4j graph database to complete the construction of the converter steelmaking knowledge graph.
2. The method for constructing a converter steelmaking knowledge graph according to claim 1, wherein: The ontology concepts and class attributes definitions include: seven ontology concepts and corresponding class attributes; The seven ontology concepts include: steel type, process parameter class, converter smelting operation class, steel grade standard class, steel grade characteristic class, deoxidation alloy operation class, and refined steel supply condition class; The class attributes of the steel type include: steel type code, steel type name, and steel type brand; The class attributes of the process parameter class include: process flow, molten iron requirements, and scrap steel requirements; The class attributes of the converter smelting operation class include: converter gun position control, bottom blowing mode, slag making system, converter tapping temperature, and liquidus temperature; The class attributes of the steel grade standard class include: release upper limit, release lower limit, agreement upper limit, agreement lower limit, standard upper limit, standard lower limit, included elements, and element target; The class attributes of the steel grade feature class include: steel grade definition, steel grade characteristics, and steel grade usage; The class attributes of the deoxidation alloy operation class include: deoxidation alloy system, alloy addition, element corresponding alloy addition, and precautions; The class attributes of the molten steel supply condition class include: molten steel supply element, molten steel supply element range, and molten steel supply element target.
3. The method for constructing a converter steelmaking knowledge graph according to claim 1, wherein: The ontology relationship definition includes: subordinate relationship, inclusion relationship, influence relationship, dependency relationship, definition relationship and similarity relationship.
4. The method for constructing a converter steelmaking knowledge graph according to claim 1, wherein: The ontology instantiation specifically includes: instantiating the ontology concept into six entities; the six entities include: steel grade code entity, process parameter entity, smelting operation entity, steel grade standard entity, deoxidation alloy operation entity, and refined steel condition entity; Three entity relationships are defined under the influence relationship between steel type and process parameter ontology; five entity relationships are defined under the dependency relationship between steel type and converter smelting operation ontology; fourteen entity relationships are defined under the inclusion relationship between steel type and steel standard ontology; four entity relationships are defined under the influence relationship between steel type and deoxidation alloy operation ontology; and five entity relationships are defined under the influence relationship between steel type and refined steel supply conditions ontology.
5. The method for constructing a converter steelmaking knowledge graph according to claim 1, wherein: According to the hierarchical structure of converter steelmaking ontology, corresponding knowledge extraction methods are used to extract entities, entity attributes and entity relationships from different types of data, and a set of knowledge triples is obtained, including: According to the hierarchical structure of converter steelmaking ontology, web crawler technology is used to extract entities, entity attributes and entity relationships for semi-structured data in the form of web pages. Rule-based methods or CMeKG model are used to extract entities, entity attributes and entity relationships for unstructured data to obtain a set of knowledge triples.
6. The method for constructing a converter steelmaking knowledge graph according to claim 5, characterized in that: The steps of the rule-based approach include: Step S61: implementing the conversion operation between the PDF file format and TXT text data through the pdfplumber library; Step S62: filtering out meaningless information by defining a regularized expression; meaningless information includes page numbers and comments; Step S63: Segment the text data according to the material attributes of different steel grades; Step S64: defining a rule module to perform preliminary knowledge extraction on different tables and text information to obtain semi-structured data; Step S65: If the semi-structured data contains noise, the process proceeds to step S62 to continue the knowledge extraction; if the semi-structured data does not contain noise, the knowledge extraction is completed.
7. The method for constructing a converter steelmaking knowledge graph according to claim 5, wherein: For unstructured data, after preliminary cleaning and parsing using pdfplumber and regular expressions, if there is a lack of clear logical relationship between the texts, the CMeKG model is used for entity extraction, entity attribute extraction, and entity relationship extraction.
8. The method for constructing a converter steelmaking knowledge graph according to claim 1, wherein: Based on the knowledge triple set, a similarity hierarchical clustering entity alignment method based on TF-IDF and SBERT is used to perform knowledge fusion, including: The TF-IDF method is used to calculate the TF-IDF vector and SBERT embedding vector of each entity in the knowledge triple set; Traverse all entity texts, calculate the TF-IDF similarity and SBERT semantic similarity based on cosine similarity respectively, linearly fuse each TF-IDF similarity and the corresponding SBERT semantic similarity through weight parameters to generate a comprehensive similarity matrix; Converting the comprehensive similarity matrix into a distance matrix; A hierarchical clustering model with pre-calculated affinity is adopted according to the distance matrix, and entity cluster division is completed with an average link strategy to obtain an entity cluster set and implement an entity alignment operation; the entity cluster set includes multiple entity clusters.
9. A converter steelmaking knowledge graph construction device, characterized in that: The device comprises: A multi-source heterogeneous data acquisition module is used to acquire multi-source heterogeneous data in the field of converter steelmaking; the multi-source heterogeneous data includes: unstructured data and semi-structured data; the unstructured data includes: converter steelmaking instruction manuals and national and enterprise standard documents; the semi-structured data includes enterprise databases and public patent information; A converter steelmaking ontology hierarchical structure construction module is used to parse the multi-source heterogeneous data and construct a converter steelmaking ontology hierarchical structure based on the obtained domain corpus features; the ontology hierarchical structure includes ontology concept and class attribute definitions, ontology relationship definitions, and ontology instantiation; The knowledge extraction module is used to extract entities, entity attributes, and entity relationships from different types of data using corresponding methods according to the hierarchical structure of the converter steelmaking ontology to obtain a set of knowledge triples; A knowledge fusion module is used to perform knowledge fusion based on the knowledge triple set by adopting a similarity hierarchical clustering entity alignment method based on TF-IDF and SBERT; The knowledge graph storage module is used to store the knowledge triple set after knowledge fusion using the Neo4j graph database to complete the construction of the converter steelmaking knowledge graph.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for constructing a converter steelmaking knowledge graph as described in any one of claims 1 to 8 are implemented.
Citation Information
Cited By
Power grid professional knowledge graph construction method for green supply chain platform
CN121766416A
A power grid professional knowledge graph construction method for a green supply chain platform
CN121766416B