A steel manufacturing knowledge graph construction method based on multi-source heterogeneous data

By constructing a knowledge graph of steel manufacturing based on multi-source heterogeneous data, a structured organization and accurate indexing of complex manufacturing processes and failure mechanisms were achieved. This solved the problems of data silos and difficulties in cross-database relational queries, enhanced the semantic completeness of the graph, and provided dynamic quality optimization support for the entire steel manufacturing process.

CN121808732BActive Publication Date: 2026-06-26ZHEJIANG LIYUAN ZHONGGONG SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
ZHEJIANG LIYUAN ZHONGGONG SCI & TECH CO LTD
Filing Date
2026-03-11
Publication Date
2026-06-26

AI Technical Summary

Technical Problem

The lack of a unified semantic standardization mechanism for multi-source heterogeneous data in the steel manufacturing field makes cross-database correlation queries difficult, makes it impossible to understand the topological relationships between data items, and existing retrieval methods cannot effectively identify causal evolution relationships in unstructured text, thus limiting knowledge utilization.

Method used

By using semantic mapping of multi-source data, cross-modal coupling, and topological feature aggregation, a steel manufacturing knowledge graph with a dynamic weighting mechanism is constructed to achieve structured organization and accurate indexing of complex manufacturing processes and failure mechanisms.

Benefits of technology

It solves the problem of difficulty in correlating microscopic morphology and macroscopic mechanism data in the field of steel failure, greatly enhances the semantic completeness of the map, and provides accurate logical support, providing a dynamic mapping relationship between process fluctuations and finished product performance for quality optimization of the entire manufacturing process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808732B_ABST
    Figure CN121808732B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of knowledge graph, in particular to a steel manufacturing knowledge graph construction method based on multi-source heterogeneous data, comprising: converting message data into standardized semantic triples through semantic mapping rules, and using physical and chemical component balance logic to complete missing attributes to form standardized entities; at the same time, using a semantic processor to extract the logic of unstructured report text, and performing cross-modal coupling with SEM image features to generate failure mechanism enhanced entities. Subsequently, collect full-process dynamic parameters and regularize them into equidistant time sequence parameters. Finally, taking the standardized entity as the index feature, taking the mechanism enhanced entity as the correlation constraint, using a topological feature mapping engine to perform semantic fusion, and combining a dynamic probability distribution correction operator to iterate parameter weights, a dynamic weight graph structure is constructed. The present application realizes efficient correlation and semantic organization of multi-modal data in the whole manufacturing process, greatly improving the retrieval accuracy and knowledge discovery ability of steel manufacturing data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of knowledge graph technology, specifically to a method for constructing a knowledge graph for steel manufacturing based on multi-source heterogeneous data. Background Technology

[0002] Currently, data storage and organization in the steel manufacturing industry are relatively isolated. Laboratory reports, management system messages, failure analysis reports, and real-time monitoring parameters are often stored in different heterogeneous databases. The lack of a unified global metadata schema and semantic standardization mechanism makes cross-database queries extremely difficult, resulting in severe data silos. Traditional retrieval methods are mainly based on keyword matching, which is shallow text processing. They cannot understand the topological relationships between data items, struggle to identify complex causal evolutionary relationships in unstructured text, and lack multimodal data indexes, failing to effectively couple microscopic morphological features with macroscopic process logic. Furthermore, existing time-series data processing lacks physicochemical mechanism constraints; the storage dimension of the data model is disconnected from physical logic, failing to dynamically reflect the deep logical mapping between process fluctuations and finished product performance. This leads to low retrieval efficiency and poor consistency in graph databases when handling large-scale complex evolutionary logic, severely limiting the knowledge utilization rate of multi-source heterogeneous data.

[0003] How to achieve semantic standardization and deep logical association of multi-source heterogeneous data throughout the entire steel manufacturing process, in order to construct a structured data map that can dynamically reflect the process mechanism, is an urgent problem to be solved.

[0004] To address this, a method for constructing a knowledge graph of steel manufacturing based on multi-source heterogeneous data is proposed. Summary of the Invention

[0005] This invention aims to provide a method for constructing a knowledge graph of steel manufacturing based on multi-source heterogeneous data. Through semantic mapping, cross-modal coupling and topological feature aggregation of multi-source data, a knowledge graph of steel manufacturing with a dynamic weighting mechanism is constructed to achieve structured organization and accurate indexing of complex manufacturing processes and failure mechanisms.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] A method for constructing a knowledge graph of steel manufacturing based on multi-source heterogeneous data includes:

[0008] Entity recognition and normalization are performed on multi-source heterogeneous data. The chemical components and contents in the message are converted into standardized semantic triples through semantic mapping rules. The missing smelting components are completed by using the physicochemical component balance logic to obtain the standardized triple entity.

[0009] A multi-layer attention semantic processor is used to parse unstructured failure analysis reports and extract text logic through dependency parsing. The text logic includes the causal evolution relationship between deformation-induced martensite, dislocation density and stress corrosion cracking. Simultaneously, a multimodal correlation framework is used to parse SEM fracture scan images, extract cleavage fracture visual feature vectors and perform cross-modal semantic coupling with the text logic to obtain steel failure mechanism enhancement entities.

[0010] Real-time acquisition of dynamic monitoring parameters throughout the entire steel manufacturing process; resampling of these dynamic monitoring parameters into equally spaced time-series parameters; semantic feature fusion of the time-series parameters using the normalized triplet entity as node index features and the steel failure mechanism enhancement entity as global association constraints; and Bayesian update and correction of parameter fluctuation weights within each process interval to generate a dynamic weight graph structure.

[0011] Preferably, the specific steps for converting the data into standardized semantic triples include:

[0012] The multi-source heterogeneous data includes at least original laboratory test reports, LIMS quality management system data, and steel standard manuals. The process involves parsing the multi-source attribute fields in the original laboratory test reports and LIMS quality management system data, and using a pre-defined steel industry identifier mapping matrix to align non-standard grades and element abbreviations from various heterogeneous data sources to the standard ontology concepts defined in the steel standard manual. Character features are extracted from the report text, and combined with the semantic association distance in the identifier mapping matrix, context disambiguation is performed on entities with semantic ambiguity to determine the standard semantic type and process category of the entity. The content values ​​of the chemical components and their corresponding units of measurement are identified, and consistency verification is performed on the content values ​​based on physicochemical logic constraints. Values ​​with different dimensions are uniformly converted into standardized attribute scores. Based on a pre-defined knowledge graph pattern, with the aligned standard ontology concepts as head entities, the chemical components as relational attributes, and the verified attribute scores as tail entities, graph pattern mapping is performed to generate the standardized semantic triples.

[0013] Preferably, the specific steps for obtaining the normalized triplet entity include:

[0014] The percentage values ​​of known chemical components are extracted from the standardized semantic triplet. Based on the furnace association attributes in the multi-source heterogeneous data, the alloy composition constraint range of the corresponding grade in the steel standard manual is retrieved. Element equilibrium boundary conditions are constructed using the alloy composition constraint range. Combined with a pre-set composition evolution distribution model, the attribute probability distribution of missing smelting components under the current process path is calculated. A constraint search algorithm is used to determine the optimal completion value in the attribute probability distribution, and consistency correction is performed on the total sum of all elements in the standardized semantic triplet according to the element equilibrium boundary conditions. The corrected completion value is embedded as a tail entity into the attribute slot of the standardized semantic triplet, generating a normalized triplet entity that satisfies consistency constraints in both the semantic space and the physicochemical value space.

[0015] Preferably, the specific steps for parsing the unstructured failure analysis report and extracting text logic include:

[0016] Unstructured failure analysis reports are vectorized using text feature encoding operators to extract hidden layer feature vectors containing contextual semantic information. Dependency parsing is performed on the failure analysis report to identify the dominance and subordination relationships between various technical terms in the text, and a syntactic tree structure is constructed based on these relationships. The hidden layer feature vectors are mapped to nodes in the syntactic tree structure, and semantic association weights between deformation-induced martensite, dislocation density, and stress corrosion cracking key entities are calculated using node saliency. Based on these semantic association weights, logical trigger words and directional paths in the text are identified, and causal evolution triples with temporal characteristics and logical guidance are extracted to form the textual logic corresponding to the failure analysis report.

[0017] Preferably, the step of extracting text logic and forming causal evolution triples further includes the following semantic analysis steps:

[0018] By utilizing the dominance weights and dependency distances between entities in the syntactic tree structure, path filtering is performed on the initially identified causal relationships to eliminate redundant logical arcs below a preset semantic threshold. Combined with a preset semantic constraint pattern in the steel failure domain, directional verification is performed on the extracted text logic to identify and correct logical conflicts where subject and object are reversed in the causal evolution triples. Modifying semantic features are extracted from the failure analysis report, the logical confidence of the causal evolution triples is calculated, and the logical confidence is added as a meta-attribute to the steel failure mechanism enhancement entity, thus achieving semantic management of text logical weights.

[0019] Preferably, the specific steps for parsing the SEM fracture scan image and performing semantic coupling include:

[0020] A deep spatial feature extractor is used to extract global visual features and local texture features from SEM fracture scan images to generate visual feature vectors representing the morphology of cleavage fracture surfaces. A cross-modal alignment algorithm is used to project these visual feature vectors into the same semantic embedding space as the text logic, and the vector space correlation between the visual feature vectors and the causal evolution relationships in the text logic is calculated. Based on the vector space correlation, the text logic path with the highest correlation strength to the visual feature vectors is identified, and the micro-features of failure implicit in the image are extracted as supplementary logical information. The visual feature vectors are converted into semantic tags and attached as additional attributes to the corresponding causal evolution triples. A graph node enhancement mechanism is then used to generate the steel failure mechanism enhancement entity.

[0021] Preferably, the real-time acquisition and resampling of dynamic monitoring parameters into equally spaced time-series parameters specifically includes:

[0022] The dynamic monitoring parameters include lance height, oxygen pressure, and oxygen supply intensity during the converter blowing stage; argon stirring intensity and heating rate during the refining stage; and rolling force and rolling speed during the rolling stage. The sensor data streams of each stage of steel manufacturing are monitored in real time through the process control system interface, and multidimensional dynamic monitoring parameters with non-uniform sampling periods are extracted. The physical dimension attributes and timestamp accuracy of the multidimensional dynamic monitoring parameters are identified, and a globally unified time axis is established based on the process start point. Using adaptive linear interpolation or cubic spline interpolation algorithms, time-domain normalization is performed on the multidimensional dynamic monitoring parameters of different frequencies according to a preset standard sampling step size, generating multidimensional feature vectors with a uniform time step size. The normalized data is then processed by segmented aggregation to filter out instantaneous noise generated by sensor jitter, and the processed equally spaced time-series parameters are mapped to dynamic attribute sequence tables of corresponding normalized triplet entities.

[0023] Preferably, the specific steps for generating the dynamic weight graph structure include:

[0024] The causal evolution relationship in the steel failure mechanism enhancement entity is mapped as oriented edges of a graph structure. The combination operators and edge type weights of the heterogeneous graph topology feature mapping engine are initialized based on the logical strength of the causal evolution relationship to construct a mechanism-guided graph topology. Using the normalized triplet entity as node metadata features, the mapping engine performs multi-layer neighborhood attribute aggregation operations on the equally spaced time-series parameters in the graph topology to achieve a correlation vector mapping between node semantic features and dynamic monitoring features. Real-time measured values ​​of finished product performance are acquired as feedback signals. A dynamic probability distribution correction operator is used to calculate the posterior probability of generating the feedback signal under the current process parameter fluctuation distribution. The attention association weights between graph nodes are iteratively corrected based on the posterior probability. The edge weight parameters in the graph topology are reconstructed based on the corrected association weights, outputting the dynamic weighted graph structure that reflects the real-time mapping relationship between process fluctuations and finished product performance.

[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0026] 1. By using cross-modal semantic coupling technology, the visual features of SEM fracture surfaces are spatially mapped with the text logic path, which solves the problem of the difficulty in associating microscopic morphology and macroscopic mechanism data in the field of steel failure and greatly enhances the semantic completeness of the map.

[0027] 2. By utilizing semantic alignment rules and physicochemical balance logic to perform data completion and disambiguation, a standardized entity that satisfies physicochemical logic constraints is constructed, enabling the retrieval of steel composition and process parameters to accurately match standard ontology concepts.

[0028] 3. By combining the dynamic probability distribution correction operator and the topology feature mapping engine, real-time monitoring parameters and failure mechanism logic are integrated to construct a dynamic graph structure that reflects the impact weight of performance fluctuations, providing precise logical support for quality optimization throughout the manufacturing process. Attached Figure Description

[0029] Figure 1 This is a flowchart illustrating the steps of a method for constructing a knowledge graph of steel manufacturing based on multi-source heterogeneous data according to the present invention.

[0030] Figure 2 A schematic diagram illustrating the process of constructing a knowledge graph for steel manufacturing in this invention;

[0031] Figure 3 This is a schematic diagram of the data flow and hierarchical processing structure of the present invention. Detailed Implementation

[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0033] Please see Figures 1 to 3 This invention provides a method for constructing a knowledge graph of steel manufacturing based on multi-source heterogeneous data, referring to... Figure 1 The flowchart and technical solution are as follows:

[0034] Entity recognition and normalization are performed on multi-source heterogeneous data. The chemical components and contents in the message are converted into standardized semantic triples through semantic mapping rules. The missing smelting components are completed by using the physicochemical component balance logic to obtain the standardized triple entity.

[0035] A multi-layer attention semantic processor is used to parse unstructured failure analysis reports and extract text logic through dependency parsing. The text logic includes the causal evolution relationship between deformation-induced martensite, dislocation density and stress corrosion cracking. Simultaneously, a multimodal correlation framework is used to parse SEM fracture scan images, extract cleavage fracture visual feature vectors and perform cross-modal semantic coupling with the text logic to obtain steel failure mechanism enhancement entities.

[0036] Real-time acquisition of dynamic monitoring parameters throughout the entire steel manufacturing process; resampling of these dynamic monitoring parameters into equally spaced time-series parameters; semantic feature fusion of the time-series parameters using the normalized triplet entity as node index features and the steel failure mechanism enhancement entity as global association constraints; and Bayesian update and correction of parameter fluctuation weights within each process interval to generate a dynamic weight graph structure.

[0037] Example 1:

[0038] This embodiment is set in the scenario of quality monitoring and process traceability in the steel smelting and refining process. First, entity recognition and normalization processing are performed on multi-source heterogeneous data. Then, the chemical composition and content in the message are converted into standardized semantic triples through semantic mapping rules.

[0039] Reference Figure 2 A flowchart illustrating the process of constructing a knowledge graph for steel manufacturing in China, and Figure 3Schematic diagram of data flow and hierarchical processing in the system. Specifically, the system first obtains the original inspection report message stream encoded in UTF-8 through an industrial Ethernet interface, and undergoes preliminary parsing of text fields using a deep entity recognition model based on the Transformer architecture. This model architecture is specifically composed of 12 layers of fully symmetric transformer encoder units stacked together. Each layer of the encoder integrates 16 parallel self-attention heads and a feed-forward fully connected network with 768 hidden neurons. The Adam optimizer is used to complete weight locking through 2000 rounds of iterative training on a training set containing 1 million historical steel inspection data samples.

[0040] The semantic mapping rule is jointly implemented by a preset text feature extraction operator library and an ontology concept mapping table in the steel field. First, the system uses a regular expression operator containing Chinese names of chemical elements, international standard symbols, and their variant expressions to automatically identify candidate element identifiers and subsequent numerical characters from the original report. Subsequently, through the ontology concept mapping table, non-standard expressions are uniformly aligned to standard chemical element terms. On this basis, according to the predefined generation logic, the identified steel grade is determined as the head entity, the standard element term is determined as the relational attribute, and the content value after dimensionality conversion is determined as the tail entity, thereby constructing a standardized semantic triple.

[0041] Specifically, the Transformer recognition model receives a text sequence of 512 characters as input, injects sequence features into the input character stream using a position encoding operator, calculates the association matrix between characters through a multi-head attention layer, and extracts original semantic fragments such as "C content 0.02", "P amount 50 ppm", and "Q235B".

[0042] Specifically, the preset identifier mapping matrix in the steel field is composed of a semantic vector mapping table trained based on the Word2Vec algorithm, with its dimension set to 1000×1000. Each element value in the matrix is obtained by calculating the cosine similarity of different element abbreviations (such as "C", "carbon", "Carbon") in a three-dimensional feature space. The dot product operation is performed between the extracted non-standard grade "Q235-B" and the standard index in the matrix. When the similarity score exceeds 95%, it is automatically aligned to the "GB-T700-Q235B" standard ontology concept defined in the steel standard manual.

[0043] Specifically, the context disambiguation logic uses a decision operator based on semantic association distance. When the character "S" appears in the report with ambiguity (possibly representing the sulfur element or a certain refining furnace number), the algorithm extracts the context feature vectors of the first 10 characters before and after the current entity and calculates the geodesic distances between this vector and the cluster centers of "chemical elements" and "equipment coding" in the identifier mapping matrix.

[0044] Specifically, the algorithm logic identifies the presence of physicochemical descriptive operators such as "content" and "percentage" in the current context by comparing the magnitude of two distance values, thereby determining that the standard semantic type of the entity is sulfur element and the process category it belongs to is refining component, thus eliminating the semantic overlap and conflict of multi-source attribute fields.

[0045] Specifically, the numerical normalization operator is activated to identify the extracted content value of 150 and its corresponding unit of measurement, ppm. Based on the content threshold table of each element stored in the physicochemical logic constraint library (e.g., the C content must be between 0.01% and 5.0%), 150ppm is converted to 0.015% in percentage form using a unit conversion function. This value is then checked for consistency with the deviation tolerance of 0.005% in the standard manual.

[0046] Specifically, the verified 0.015 is extracted as a standardized attribute score, and a predefined resource description framework storage mode is retrieved. The aligned "GB-T700-Q235B" is encapsulated as a head entity, "contains sulfur" is defined as a relational attribute, and "0.015" is encapsulated as a tail entity. Through the dynamic mapping of the function execution graph mode constructed by the triples, standardized semantic triples with unique physical addressing meaning are generated and written into the persistent partition of the graph database.

[0047] Furthermore, in the context disambiguation step, the co-occurrence frequency score and distance weight of the entity to be identified in the full text of the message are extracted; a local topological subgraph of the entity candidate set is constructed using the semantic association distance, and the information entropy increment of each candidate entity in the subgraph is calculated; the candidate entity with the largest information entropy increment is determined as the final standard semantic type.

[0048] When the abbreviation "Q235" appears in the original test report, the system identifies it as potentially pointing to either a "steel grade" or a "test record number." Extracting keywords from the first and last 10 characters of this code reveals that "carbon content" and "yield strength" co-occur with a frequency exceeding 85.0%. Calculations show that mapping this to the "steel grade" ontology results in the highest improvement in semantic certainty of the local topological subgraph, thus eliminating ambiguity.

[0049] By introducing an information entropy increment evaluation mechanism, the pain point of overlapping meanings of professional abbreviations in the steel industry across different production systems was resolved, the accuracy of entity recognition was improved, and the risk of semantic drift in subsequent related queries was reduced.

[0050] By constructing a Transformer entity recognition architecture and coupling it with a steel industry-specific identifier mapping matrix, a closed-loop transformation of laboratory report data from discrete text to standardized semantic entities was achieved. Under the premise of ensuring physical and chemical consistency constraints, this provides a data benchmark with high physical confidence for subsequent steel quality prediction and process parameter optimization.

[0051] Furthermore, the system utilizes physicochemical component balance logic to perform attribute completion on missing smelting components, resulting in a standardized ternary entity. This physicochemical component balance logic strictly adheres to the law of conservation of mass and metallurgical kinetics. During attribute completion, the system first accumulates the known mass percentages of all elements in the current furnace batch and subtracts a preset element oxidation loss component from the total. This oxidation loss component is dynamically estimated based on process parameters such as oxygen blowing time and oxygen pressure, simulating the gas escape generated by elements like carbon and sulfur during high-temperature smelting. The residual mass percentage after deducting the loss is the total budget for the elements to be completed. The system then initially allocates this budget according to the element content constraint ratios defined in the steel standard manual and introduces a dynamic relaxation factor to ensure that the sum of all element percentages is precisely equal to 100% by fine-tuning the iron content in the base metal.

[0052] Specifically, the mass percentage values ​​of known chemical components are first extracted from the standardized semantic triples generated by the preceding module through the data extraction interface. Specifically, the carbon content in the furnace with the number LC20240105 is 0.18% and the manganese content is 0.55%, and the planned production grade associated with this furnace is Q235B.

[0053] Specifically, based on the identified Q235B grade information, the system automatically retrieves the metadata from the pre-stored steel standard manual PDF in the database after parsing, and extracts the alloy composition constraint range of nickel and chromium under this grade, specifically locking that the mass fraction of nickel must be within a closed range of 0% to 0.30%.

[0054] Specifically, the element equilibrium boundary conditions are constructed using the above-mentioned constraint range. The core logic is to require that the sum of the mass percentages of all elements be strictly equal to 100%, and this operator is injected into the subsequent regression calculation task as a restrictive filtering condition.

[0055] Specifically, a component evolution distribution model based on a gated recurrent unit architecture is constructed to predict the dynamic changes in composition caused by alloying operations during the smelting process. The model architecture includes one input layer for receiving known component vectors, two hidden recurrent layers, each containing 128 neurons, for extracting the temporal features of the process path, and one output layer that predicts the mean and variance of the components. The model uses the Adam optimizer to perform 1500 rounds of training on 50,000 historical smelting sequences with an initial learning rate of 0.001 to lock the parameters.

[0056] Specifically, the input to the component evolution distribution model consists of a feature vector composed of three concatenated parts:

[0057] The first part is a vector of known chemical composition containing the mass percentage of 20 common elements, with unknown elements filled with 0. The second part is a vector of current process parameters containing 10 key process parameters such as refining time, decarburization rate, and alloy addition time. The third part is a coded vector representing 5 furnace types, including ordinary steel, high-strength steel, and stainless steel, using unique thermal coding.

[0058] The above three features are concatenated to form a 35-dimensional input vector. The output of the model is the attribute probability distribution parameters of the missing elements, specifically including the mean and variance. That is, the output layer of the model is a 2-dimensional vector, which is used to indicate that the content of missing elements follows a Gaussian distribution.

[0059] The model's training data comes from 50,000 historical smelting records, each containing complete chemical composition analysis results and process parameters. During training, the system constructs missing scenarios by randomly masking the content value of a certain element. The masked component vector and process parameters are used as input, and the actual element content is used as the supervision label. The mean squared error loss function is used for model training.

[0060] Taking a specific furnace batch as an example, if the carbon content is known to be 0.18%, the manganese content to be 0.55%, and the refining time to be 15 minutes, then the corresponding positions in the model input vector will be filled with these values, while the remaining positions will be filled according to the rules. If the 2D vector output by the model is 0.045 and 0.000004, it means that the missing nickel content follows a Gaussian distribution with a mean of 0.045 and a standard deviation of 0.002.

[0061] Specifically, the component evolution distribution model receives the 15-minute refining time already executed in this heat and the current carbon and manganese content as input vectors. Through the hidden layer neurons, it performs nonlinear feature mapping on the decarburization of molten steel and alloy yield. The output layer calculates the attribute probability distribution of the missing nickel element at the current process point, for example, generating a Gaussian probability density curve with a mean of 0.045 and a variance of 0.002.

[0062] Specifically, a constrained search algorithm based on the Lagrange multiplier method is used to perform numerical optimization in the attribute probability distribution. The optimal complement value that minimizes the residual of the objective function for the conservation of total element mass is searched within a range of 3 times the standard deviation of the mean.

[0063] Specifically, the algorithm continuously adjusts the candidate values ​​of nickel and calculates its contribution to the overall mass balance in real time. When it identifies that the value of 0.048 can make the calculation of the residual iron component most stable and meet the 100% sum requirement, it locks 0.048 as the attribute completion result of nickel.

[0064] Specifically, a consistency correction operation is performed on the total sum of all elements. The values ​​of all chemical components after completion are summed up, and the accuracy of the sum is fine-tuned by adjusting the iron matrix component, which has the largest proportion, to ensure that the numerical residual under floating-point operations is less than 0.00001.

[0065] Specifically, the corrected nickel element value of 0.048 is used as the tail entity, and the triplet encapsulation function is used to embed it into the corresponding standardized semantic triplet attribute slot.

[0066] Specifically, this process transforms the original LC20240105 furnace record containing component gaps into a normalized triplet entity with high credibility in both the semantic space (conforming to the triplet pattern) and the physicochemical numerical space (satisfying 100% mass conservation), and pushes the result to the subsequent mass prediction module in real time.

[0067] Furthermore, in the consistency correction step, a dynamic relaxation factor based on the law of conservation of mass is introduced; the percentage values ​​of each known chemical component are summed with the completed value, and the residual is calculated compared with 100% of the theoretical total value; if the absolute value of the residual is greater than 0.05%, the residual is allocated to the attribute values ​​of each element according to the variance weight of each element in the component evolution distribution model; in the component completion process of a certain furnace smelting, the system predicts that the missing "manganese" content is 1.2%. After calculating the total of all elements, the residual is found to be 0.08%, exceeding the preset threshold of 0.05%. The system identifies that the fluctuation variance of "carbon" and "manganese" is the largest in this process, and therefore the residual is amortized and corrected according to a weight ratio of 6:4 to ensure that the final chemical component logic is completely closed at the physical level.

[0068] By introducing a dynamic relaxation factor and a variance-weighted allocation mechanism, the underlying data of the knowledge graph is forced to conform to the physical mass conservation constraint, avoiding attribute value conflicts caused by model prediction errors and ensuring the consistency of the knowledge graph in the physicochemical numerical dimension.

[0069] Through the collaborative processing of cross-modal semantic mapping and physicochemical balance completion, the refining process quality data was efficiently transformed from fragmented and heterogeneous messages to logically consistent normalized triples. While ensuring the physical confidence of the data, this provided a solid knowledge representation benchmark for the digital quality traceability of the entire steel manufacturing process.

[0070] To address the complex metallurgical evolution logic contained in a large number of unstructured expert failure analysis reports, a semantic logic extraction engine integrated into a materials big data platform is used to perform knowledge extraction tasks.

[0071] Furthermore, text feature encoding operators are used to vectorize the unstructured failure analysis report, extracting hidden layer feature vectors containing contextual semantic information.

[0072] The multi-layer attention semantic processor, serving as a feature enhancement module, consists of three consecutively stacked self-attention processing units. This processor receives a high-dimensional feature vector sequence output from the text feature encoding operator. The first layer calculates the global word association distribution, the second layer captures long-distance logical dependencies in failure descriptions, and the third layer outputs an enhanced contextual semantic vector. This vector is mapped to the corresponding node in the dependency syntax tree, ensuring that each node simultaneously carries syntactic structure information and deep semantic features. The final output text logic records the head entity, causal relationship, tail entity, and corresponding confidence score in structured data format.

[0073] Specifically, the unstructured failure analysis report text stream encoded in UTF-8 is first obtained through the data reading interface. Feature extraction is then performed using a deep text encoding model based on the Transformer architecture. This model architecture consists of 12 layers of fully symmetrical transformer encoder units stacked together. Each encoder layer integrates 12 parallel self-attention heads and a feedforward fully connected network with 768 hidden neurons. The Adam optimizer is used to perform 2000 rounds of iterative training on 300,000 metal failure corpora with an initial learning rate of 0.0001 to lock the weights.

[0074] Specifically, the Transformer encoding model receives a text segment with a maximum length of 512 characters as input, injects relative position information into the character stream through a positional encoding operator, and uses a multi-head self-attention mechanism to calculate the dot product correlation between professional terms with large spans, such as "martensite" and "cracking". Finally, the last encoder outputs a dense hidden layer feature vector with a dimension of 768. This vector encapsulates the semantic features of the macro context and micro metallurgical terminology in the failure report in numerical form.

[0075] Furthermore, dependency parsing is performed on the failure analysis report to identify the dominance and subordination relationships between various technical terms in the text, and a syntactic tree structure is constructed based on the relationships.

[0076] Specifically, the syntactic analysis engine adopts a deep syntactic analysis model based on a dual affine attention mechanism. The model architecture includes an input layer for receiving the feature vectors of the aforementioned hidden layer, two bidirectional long short-term memory networks each containing 512 hidden units, and a scoring layer that outputs a dependency matrix. The model is trained to convergence on a standard annotation set in the steel industry using the Adam optimizer.

[0077] Specifically, the algorithm logic extracts local features of the word sequence through the BiLSTM layer and uses the scoring layer to calculate the probability weight score of each pair of words as "dominant word-subordinate word". For example, it identifies the verb "induce" as the dominant word and "Marseille" as the subordinate relationship of its direct object.

[0078] Specifically, by identifying expressions in the text such as "high dislocation density promotes the initiation of microcracks", the topological linking operator is used to link the nodes of various professional terms according to their dominance relationship, thereby constructing a dynamic syntactic tree structure in computer memory with a subject-verb-object structure as the skeleton and including modifying and restricting semantic branches.

[0079] The beneficial effect of this step is that by capturing the grammatical structure features of the text, the linear text flow is restored into a grid-like topology with engineering logic orientation, thus realizing a hard physical constraint on the semantic framework of the text.

[0080] Furthermore, the hidden layer feature vectors are mapped to the nodes of the syntax tree structure, and the semantic association weights between deformation-induced martensite, dislocation density, and stress corrosion cracking key entities are calculated through node saliency.

[0081] The node saliency is calculated using a weighted degree centrality algorithm. The specific calculation logic is as follows: the saliency of a target node is equal to the sum of the products of all edge weights connected to it and the self-attention scores of its neighboring nodes. The edge weights are determined by the dependency relation type in the syntax tree; for example, the subject-verb relation weight is set to 1.0, the verb-object relation weight to 0.9, and the noun-head relation weight to 0.7. The self-attention scores of neighboring nodes are output by the transformer model, and their values ​​range from 0 to 1.

[0082] For key entities such as "deformation-induced martensite," "dislocation density," and "stress corrosion cracking," the system calculates their node saliency separately. Then, it calculates the semantic association weight between each pair of entities, using the logic of multiplying the two node saliencies by the sum of those two saliencies. This logic ensures that nodes with high saliency receive higher association weights.

[0083] Finally, the system performs normalization on all associated weights to ensure that the sum of all weight values ​​equals 1. Taking the sentence in the syntactic tree about dislocation density promoting microcrack initiation as an example, if the weighted degree centrality of the "dislocation density" node is 0.85, and the weighted degree centrality of the "stress corrosion cracking" node is 0.92, then the semantic association weight obtained through the above logical calculation is 0.44, and the final weight score after normalization is 0.88.

[0084] Specifically, the mapping logic uses the method of filling the 768-dimensional hidden layer feature vectors one by one into the corresponding word nodes of the syntax tree to construct a weighted graph with semantic information and extract the node coordinates of the three core entities: "deformation-induced martensite", "dislocation density" and "stress corrosion cracking".

[0085] Specifically, the node saliency calculation adopts a graph centrality-based evaluation algorithm. By calculating the weighted sum of the number of paths from the core entity node to other semantic nodes in the syntactic tree and the path weights, combined with the probability scores output by the self-attention mechanism, the semantic association weights between the three entities are calculated. For example, the association weight between martensite and dislocation density is 0.88, while the weight between dislocation density and stress corrosion cracking is 0.92.

[0086] Furthermore, based on the semantic association weights, logical trigger words and directional paths in the identified text are extracted, and causal evolution triples with temporal characteristics and logical guidance are extracted to form the text logic corresponding to the failure analysis report.

[0087] Specifically, the logic recognition engine searches for key paths with a weight greater than 0.85 in the syntax tree and uses a keyword matching algorithm to locate logical trigger words on the path, such as "cause", "induce", "cause" or "accompany".

[0088] Specifically, the algorithm process extracts semantic fragments that satisfy the topological structure of "entity A-trigger word-entity B" and generates causal evolution triples with clear temporal orientation, such as (deformation-induced martensite, positive correlation-induced, high dislocation density) and (high dislocation density, main cause of structural degradation, stress corrosion cracking).

[0089] Specifically, the extracted triples are arranged according to the physical time sequence of the failure process to form a complete text logic that reflects the entire chain of "deformation-organizational evolution-failure", and then sent to the subsequent graph reconstruction module in binary JSON format.

[0090] By constructing a multi-layer semantic processing architecture based on the coupling of Transformer and syntax tree, we have achieved white-box extraction of deep causal relationships in complex metal failure reports. While ensuring the closed loop of data flow, this significantly enhances the physical consistency of the logical deconstruction of professional texts in the steel industry.

[0091] Furthermore, the extraction of text logic and formation of causal evolution triples specifically includes the following semantic analysis steps: using the dominance weights and dependency distances between entities in the syntactic tree structure, performing path filtering on the initially identified causal relationships to eliminate redundant logical arcs below a preset semantic threshold; combining the preset semantic constraint patterns in the steel failure domain, performing directional verification on the extracted text logic to identify and correct logical conflicts where the subject and object are reversed in the causal evolution triples; extracting the modifying semantic features from the failure analysis report, calculating the logical confidence of the causal evolution triples, and attaching the logical confidence as a meta-attribute to the steel failure mechanism enhancement entity to achieve semantic management of text logic weights.

[0092] Specifically, the topological relationships between nodes are extracted from the syntactic tree structure generated by the preceding module, and a path filtering algorithm based on graph attention network is launched. The model architecture consists of three stacked graph attention layers, each containing eight parallel attention heads and a hidden layer dimension of 128. The Adam optimizer is used to perform gradient descent on a training set containing 5000 sets of standard semantic paths with an initial learning rate of 0.0001.

[0093] Specifically, the algorithm processing logic calculates the dominance weight score between adjacent entities by performing a multiplication operation between the node feature vector and the edge label weight matrix, and simultaneously measures the shortest edge number of the two core words in the syntax tree as the dependency distance.

[0094] Specifically, a weighted operator is used to fuse the dominance weight with the reciprocal of the distance to generate a logical arc score that represents semantic relevance. When the score of a logical arc is found to be lower than the preset threshold of 0.65, the connection is determined to be a false association caused by interference from long and difficult sentences and physical removal is performed.

[0095] Specifically, the directional verification logic adopts a rule verification operator based on the knowledge ontology of the steel industry. This operator has a pre-built semantic constraint pattern library containing 500 metallurgical causal laws. For example, it stipulates that "dislocation density" can only be used as the inducing source of "microcrack initiation".

[0096] Specifically, the system matches the extracted causal evolution triples (dislocation density, induced, stress corrosion cracking) with the constraint pattern library, and identifies whether there is a subject-object reversal phenomenon by calculating the attribute compatibility scores of the head entity and the tail entity in the triples.

[0097] Specifically, when a logical conflict is detected in the text, such as "stress corrosion cracking causes an increase in dislocation density" due to passive voice, the algorithm automatically performs a mirror inversion correction on the triples based on physical logic to ensure that the logic of the output text conforms to the basic laws of materials science.

[0098] Specifically, the system constructs a logical confidence evaluation model based on a multilayer perceptron. The model architecture includes an input layer for receiving modifying semantic features, two hidden layers containing 32 and 16 neurons respectively, and a single-dimensional output layer.

[0099] Specifically, the model extracts the sentiment polarity features of modifiers such as "significant", "highly likely", and "speculated" from the failure analysis report and converts them into numerical weights. For example, the weight of the term "significantly induced" is set to 0.95, while the weight of the term "may cause" is set to 0.42.

[0100] Specifically, the algorithm logic calculates the logical confidence value of the causal evolution path by performing a nonlinear synthesis operation on the association weights and modification weights of the triples, and encapsulates it into a meta-attribute in JSON format.

[0101] Specifically, by using the attribute mounting function, a logical confidence score of 0.88 is attached to the attribute slot of the "steel failure mechanism enhancement entity", thereby realizing the weighted hierarchical management of failure knowledge from different sources.

[0102] Specifically, the resulting causal evolution sequence with semantic management attributes is written into a distributed graph database, serving as the core logic support for subsequent failure probability prediction and process optimization.

[0103] By constructing a graph attention filtering model and coupling it with domain semantic constraint patterns, a closed-loop transformation of metal failure reports from natural language descriptions to high-confidence causal logic chains was achieved. While ensuring that the logical flow conforms to the laws of metallurgy, the rigor and systematicity of failure analysis knowledge reconstruction were greatly enhanced.

[0104] Furthermore, the SEM fracture scan image is parsed simultaneously using a multimodal correlation framework, the visual feature vector of the cleavage fracture is extracted and cross-modal semantic coupling is performed with the text logic to obtain the steel failure mechanism enhancement entity;

[0105] Specifically, the deep spatial feature extractor adopts a feature recognition model based on a convolutional neural network architecture. This model architecture includes one input convolutional layer, 16 stacked residual blocks, and one global average pooling layer. The convolutional layer is configured with 512 convolutional kernels of size 3×3 with a stride of 1. By receiving 1024×1024 resolution TIFF format SEM fracture scan images transmitted by a scanning electron microscope on a computing server configured with NVIDIA A100 GPU, the local texture features representing cleavage steps, river patterns, and cleavage facets in the image are extracted using multi-layer convolutional operators. Finally, a dense visual feature vector of dimension 512 is output and written to the cache partition.

[0106] Specifically, the cross-modal alignment algorithm constructs a mapping operator based on a dual-path Siamese network. This operator consists of a fully connected mapping layer with an input dimension of 512 and a projection layer with a hidden dimension of 1024. The Adam optimizer is used to train the contrastive loss function on 30,000 pairs of steel failure image-text association samples with an initial learning rate of 0.0001. By calculating the weight distribution of the projection matrix, the 512-dimensional visual feature vector is mapped to the same 768-dimensional semantic embedding space as the aforementioned text logic.

[0107] The physical significance of the cross-modal semantic coupling lies in establishing a statistical correlation mapping between microscopic fracture morphology and macroscopic failure mechanisms. Since river patterns and cleavage steps in scanning electron microscope images are intuitive physical manifestations of failure mechanisms, while textual logic provides a linguistic description of these mechanisms, the system transforms visual feature vectors into the same semantic embedding space as the textual features using a dual-path projection matrix. Subsequently, a similarity determination operator is used to calculate the cosine correlation score between the visual vector and the textual logic vector in the spatial dimension. The system selects the triplet path with the highest correlation strength and attaches the visual features as additional attributes to the corresponding graph nodes.

[0108] Specifically, the linear algebra computation operator is invoked to perform vector dot product operations in a 768-dimensional coordinate space. By calculating the cosine similarity score between the projected visual feature vector and the causal evolution triplet feature vector extracted from the text logic, the spatial correlation between the cleavage morphology features in the SEM image and the dislocation density evolution description in the text is quantified, and it is determined whether the two point to the same physical failure event.

[0109] Specifically, the text logic path with the highest score (e.g., 0.92) is identified based on the calculated association strength score, and the convergence point of the river pattern is extracted as logical supplementary information of the failure micro-features by analyzing the gray-scale gradient distribution of local stress concentration areas in the image.

[0110] Specifically, the visual feature vector is transformed into a semantic label with metallurgical significance using an attribute mapping function. This label is then attached as metadata to the tail entity of the corresponding standardized semantic triple through a graph node enhancement mechanism. Finally, an enhanced entity of steel failure mechanism with graph-text feature coupling is generated in the distributed graph database.

[0111] By constructing a collaborative analysis architecture based on residual convolutional networks and cross-modal projection matrices, we have achieved accurate matching and knowledge enhancement of image visual features and textual logical chains in steel failure analysis. While ensuring the closed loop of multi-source heterogeneous data flow, we have significantly enhanced the accuracy and interpretability of steel quality evaluation in complex environments.

[0112] Furthermore, dynamic monitoring parameters of the entire steel manufacturing process are collected in real time, and these dynamic monitoring parameters are resampled into time-series parameters with equal intervals.

[0113] Specifically, the sensor data stream of the process control system is monitored in real time via the OPC-UA protocol through the industrial Ethernet interface. From this data, original physical quantities with non-uniform sampling characteristics are extracted, such as the gun position height of 1.5m and oxygen pressure of 0.8MPa in the converter process with a sampling period of 200ms, the heating rate of 2.5℃ / min in the refining process with a sampling period of 2000ms, and the rolling force of 15000kN in the hot rolling process with a sampling period of 10ms.

[0114] Specifically, the internal time series processor automatically parses the physical dimension attributes (such as MPa, kN, m) and Unix nanosecond-level timestamps in the header of each parameter packet, determines the initial furnace entry time T0 of the molten steel by retrieving the process work order information, and uses this as the zero point of time to establish a globally unified time axis covering the entire process of blowing, refining and rolling.

[0115] Specifically, for refining heating rate data with a sampling frequency of only 0.5Hz and rolling force data with a sampling frequency of 100Hz, the cubic spline interpolation algorithm is called to perform time-domain resampling. The algorithm logic is to construct a cubic polynomial function between every two adjacent discrete observation points, and by ensuring the continuity of the first and second derivatives at the connection points, numerical mapping is performed according to the preset standard sampling step size of 100ms, thereby transforming the non-uniformly distributed discrete sampling points into multidimensional feature vectors with a uniform time-domain step size.

[0116] The cubic spline interpolation algorithm uses a cubic spline function with natural boundary conditions, and the specific implementation process is described below:

[0117] For time-series parameters with a specific sampling period, the system first extracts a set of discrete sampling points, including the timestamp of each sampling point and its corresponding parameter value. Then, based on a preset standard sampling step size of 100 milliseconds, a set of equally spaced target sampling timestamps is generated on the time axis. The total number of points is determined by dividing the original sampling time span by the sampling step size and rounding up. Next, for each target timetamp, the system automatically finds its position interval within the original sampling points and constructs a corresponding cubic polynomial interpolation function. The coefficients of this function are obtained by solving a tridiagonal linear equation system, and the boundary conditions are set to natural boundaries, i.e., the second derivatives at the first and last sampling points are both zero. Finally, the system calculates and outputs the parameter values ​​corresponding to each target sampling timetamp.

[0118] Taking the refining heating rate data with a sampling frequency of 0.5 Hz (i.e., a period of 2000 milliseconds) as an example, if the original sampling point values ​​at 0 seconds, 2 seconds, and 4 seconds are 2.3, 2.5, and 2.7 degrees Celsius per minute, respectively, then after performing interpolation processing with a standard step size of 100 milliseconds, a sequence of equally spaced sampling points starting from 0 seconds with a step size of 0.1 seconds will be generated, ultimately resulting in a total of 41 smoothly aligned sampling points.

[0119] Specifically, in order to filter out instantaneous noise introduced by mill vibration or electromagnetic interference, a segmented aggregation model based on a 1D-CNN convolutional neural network was constructed. The model architecture includes one input layer, two feature extraction layers each containing 32 3×1 convolutional kernels, and one global mean pooling layer. The Adam optimizer was used to train the model for 500 rounds on a training set containing 1000 sets of typical sensor jitter samples. The convolutional kernels were used to extract the steady-state components of the signal in the time dimension and filter out abnormal pulse points with amplitudes exceeding three times the standard deviation of the mean.

[0120] Specifically, the equally spaced time-series parameters after segmented aggregation are associated with the corresponding normalized triplet entities (such as the LC20240105 heat entity) through a hash mapping function. The normalized rolling force curve and gun position fluctuation sequence are stored in the dynamic attribute sequence table slots of the entity in the form of key-value pairs, realizing the closed-loop flow of physical monitoring data to semantic attributes of knowledge graph.

[0121] By constructing a real-time monitoring mechanism based on the PCS interface and coupling cubic spline resampling with a 1D-CNN denoising algorithm, the structured regularization of heterogeneous dynamic monitoring parameters throughout the entire steel manufacturing process was achieved. While ensuring the time sequence alignment of parameters in each process, the confidence and integrity of the quality traceability data stream were significantly enhanced.

[0122] Furthermore, using normalized triple entities as node metadata features, the causal evolution relationship is mapped as directional edges of a graph structure. Semantic feature fusion is performed on the time-series parameters, and Bayesian updates are used to correct the parameter fluctuation weights within each process interval to generate a dynamic weighted graph structure.

[0123] Specifically, the semantic feature fusion process is as follows: For the graph node corresponding to the normalized triplet entity, its initial semantic feature vector is first obtained; simultaneously, for the equally spaced temporal parameter sequence associated with this node, its local fluctuation features evolving over time are extracted using a one-dimensional convolutional layer to obtain the corresponding temporal feature vector. Then, the contribution weights of the semantic features and temporal features are calculated through an attention fusion mechanism: the two types of feature vectors are concatenated, a linear transformation is performed using a learnable attention weight matrix, and two sets of weight coefficients are output using a normalized exponential function. Finally, a weighted summation is performed on the initial semantic features and the extracted temporal features based on the weight coefficients to obtain the fused node comprehensive feature vector.

[0124] Specifically, the causal evolution triplet data of deformation-induced martensite, dislocation density, and stress corrosion cracking in the previously generated steel failure mechanism enhancement entity are extracted and used as the topological backbone of the graph database. Initial weight scores are configured for relation edges with different semantic strengths in the logic control unit. For example, a strongly associated path with a logical confidence of 0.88 is defined as a directional edge with high priority in the graph structure.

[0125] Specifically, the heterogeneous graph topology feature mapping engine adopts a deep computing model based on the heterogeneous graph attention network HGAT. The model architecture includes a feature linear transformation layer for aligning heterogeneous node dimensions, two stacked heterogeneous graph convolutional layers for cross-type attribute aggregation, and a semantic attention layer with eight parallel attention heads. The hidden layer neuron dimension is uniformly set to 256.

[0126] Specifically, the model parameters were locked using the Adam optimizer through 3,000 iterations of training on a process-quality correlation sample set containing 50,000 historical smelting furnaces. A structure regularization term was introduced into the loss function to ensure that the mapping engine could accurately identify the semantic interactions between different node types.

[0127] Specifically, the mapping engine receives the normalized triplet of heat LC20240105 as the initial node features and simultaneously reads in the processed equally spaced time series parameter sequence, which includes a dynamic vector of real-time rolling force of 15000kN and rolling speed of 150m per minute. Through the convolution operator, it performs nonlinear weight summation based on the first-order neighborhood in the graph topology guided by syntax, thereby mapping and aggregating the physical monitoring features of the process into the corresponding metallurgical mechanism node embedding vector in real time.

[0128] Specifically, the measured values ​​of tensile strength (600 MPa) and yield strength (450 MPa) from the finished product performance testing module are obtained in real time as feedback signals. The conditional likelihood probability of the finished product's strength index is generated by calculating the parameter fluctuation distribution under the current gun position height of 1.5 m and heating rate of 2.5 degrees Celsius per minute.

[0129] Specifically, the workflow of the dynamic probability distribution correction operator is as follows: Initial edge weights between graph nodes are pre-defined, and the measured values ​​of finished product performance and the process parameter fluctuation distribution composed of gun height and heating rate are acquired in real time. The correction process employs a Bayesian update algorithm: a uniform distribution is used as the initial prior distribution of edge weights, and a Gaussian distribution is fitted using historical process and quality correlation data as the likelihood function. By calculating the posterior probability of generating the measured values ​​of the finished product under the current process parameter fluctuation conditions, the system iteratively corrects the attention weights in the heterogeneous graph processing engine.

[0130] The specific correction logic is as follows: based on the original weights, an incremental correction consisting of the deviation between the preset learning rate and the posterior probability is accumulated. This iterative process continues until the absolute value of the weight change between two adjacent iterations is lower than a preset accuracy threshold, or the preset maximum number of iterations is reached. The preset accuracy threshold is 0.001: this threshold represents the convergence criterion for weight changes between two adjacent iterations, and its setting is based on the parameter sensitivity and numerical stability requirements of the physical process in steel manufacturing. The system analyzes the random fluctuation range of historical process data under steady-state conditions and selects a value slightly higher than the sensor's environmental noise level as the threshold. When the weight update amount is lower than this value, the decision graph structure has reached semantic consistency balance, thereby terminating the calculation and reducing redundant overhead.

[0131] The maximum number of iterations is capped at 100: This cap is a hard time constraint set to ensure real-time retrieval response in an online production environment. Its value is jointly determined by the aggregation time of the heterogeneous graph topology feature mapping engine and the millisecond-level latency requirements of production line process switching, ensuring that the weight correction process can be completed within the preset calculation cycle under complex smelting conditions, preventing system blockage due to excessive iteration.

[0132] Specifically, the eight head weight coefficients in the HGAT model's attention mechanism are iteratively corrected based on the calculated posterior probability distribution. The gradient of the attention distribution is biased and adjusted through a hardware accumulator, so that the process interval edge weight parameters, which are more closely related to quality fluctuations, can obtain higher gain scores.

[0133] Specifically, the algorithm logic iteratively updates the association weights between graph nodes, and finally outputs a dynamic weighted graph structure in the real-time analysis partition of the distributed memory. This structure is a multi-dimensional mapping relationship between chemical composition deviations, production process fluctuations and finished product mechanical properties, and is carried by a sparse matrix of 1024×1024 dimensions.

[0134] Furthermore, in the step of constructing the heterogeneous semantic index directory, a set of high-frequency query paths is extracted from the dynamic weighted graph structure; the path set is used to perform physical clustering and reorganization of the graph storage space, and nodes with logical causal relationships are stored near each other in physical disk sectors; a multi-level skip table index based on node topology is established to achieve non-linear accelerated retrieval of cross-process failure paths; analysis of historical query records reveals that "converter endpoint carbon - refining heating rate - finished product cold bending crack" is a high-frequency associated path. The system physically relocates the above three nodes and their dynamic attribute tables, which are distributed in different storage partitions, to contiguous storage blocks. When a user retrieves a similar failure mode, it can be directly located through the three-level skip table index without a full table scan of the graph topology.

[0135] By transforming logical associations into physically proximate storage and combining them with multi-level skip list indexes, the number of disk input / output operations when the graph database performs long-chain retrieval is significantly reduced, effectively reducing the response latency for complex failure tracing.

[0136] The Bayesian update and correction operator is responsible for dynamically maintaining the consistency of graph edge weights. First, based on the similarity between the current steel grade and the process path, the empirical distribution of edge weights is extracted from the historical database as the prior probability. If historical data is scarce, an unbiased uniform distribution is used. Once the measured performance values ​​of the finished product are returned, the operator calculates the likelihood function in conjunction with the current process parameter fluctuation distribution, and then derives the posterior probability. Based on the deviation between the posterior and prior probabilities, the graph edge weights are iteratively corrected using a preset learning rate. This iterative process continues until the weight change is below a preset accuracy threshold or the maximum number of iterations is reached.

[0137] By constructing the HGAT heterogeneous graph computation framework and introducing a Bayesian closed-loop update mechanism, the discrete mechanism knowledge and continuous dynamic parameters in the entire steel production process are digitally integrated. Under the premise of ensuring that the prediction logic has physical interpretability, the real-time performance and accuracy of the finished product performance mapping under complex smelting environment are significantly improved.

[0138] This embodiment integrates Transformer semantic parsing, physically constrained variational autoencoders, heterogeneous graph attention networks, and Bayesian closed-loop update mechanisms at multiple levels to construct a complete technical chain from raw message perception to finished product performance mapping. This not only accurately characterizes the causal evolutionary features in steel manufacturing but also significantly improves the sensitivity and interpretability of finished product performance prediction under complex process fluctuations by generating a dynamic weighted graph structure with real-time feedback correction capabilities.

[0139] Example 2:

[0140] This embodiment is set in a multi-station collaborative quality optimization scenario for the production of ultra-high strength marine steel. To address the problem of nonlinear yield drift caused by oxidation kinetics due to trace alloying elements under extreme high temperature fluctuations (above 1600 degrees Celsius), a collaborative wire feeding scheduling operator based on a multi-agent reinforcement learning architecture is introduced.

[0141] Furthermore, by utilizing multi-agent reinforcement learning operators combined with a dynamic weighted graph structure, and by analyzing the multi-dimensional interaction features in the real-time melting path, a collaborative wire feeding scheduling instruction sequence is generated that can minimize alloy oxidation loss and maximize microstructure uniformity.

[0142] Specifically, the system constructs a multi-agent cooperative scheduling model based on the Proximal Policy Optimization (PPO) algorithm. The model architecture includes a 1024-dimensional input layer for receiving environmental state vectors, three hidden layers with 512, 256, and 128 neurons respectively and using the ReLU activation function, and an output layer that outputs discrete control actions.

[0143] Specifically, the multi-agent collaborative scheduling model obtains the instantaneous temperature of molten steel at 1620 degrees Celsius, the residual oxygen content of 0.012%, and a 128-dimensional normalized feature vector of trace elements such as niobium, vanadium, and titanium output by the dynamic weighted graph structure in real time through the industrial Ethernet interface as the initial state.

[0144] To ensure the convergence and safety of the reinforcement learning algorithm under extreme smelting environments, the system introduces a high-fidelity simulation environment and a progressive training strategy. First, large-scale pre-training is performed within a simulation space based on physicochemical reaction kinetics to avoid direct trial and error on the actual production line. Second, a priority empirical sampling mechanism is introduced to enhance the model's ability to learn from key samples of quality fluctuations. The system sets strict convergence criteria, requiring the reward function to remain stable across multiple consecutive training cycles. Simultaneously, an online monitoring and fault-tolerance mechanism is deployed; if the component deviation generated by the decision exceeds a preset safety standard deviation, the system will immediately and automatically switch to manual intervention mode.

[0145] Specifically, the internal algorithm logic of the model uses the Adam optimizer to perform 2000 rounds of iterative training on a sample set of 100,000 groups containing the order of addition of multiple alloys and the purity of molten steel with an initial learning rate of 0.0001. The model performs deep feature aggregation through the nonlinear competitive adsorption and thermodynamic equilibrium reaction between elements by the hidden layer neurons.

[0146] Specifically, the policy network generates a set of 32-dimensional action mapping vectors in the output layer by calculating the expected reward scores of different action sequences in the digital twin space, i.e., calculating the negative logarithmic function of the difference between the target alloy component and the measured component.

[0147] Specifically, the system parses the action mapping vector and converts it into an automated execution message containing the wire feeding start time of 14 minutes and 25 seconds, the wire feeding speed of 6 meters per second, and the alloy compensation amount of 35.5 kg, which is then sent to the field feeding system via the Profibus control protocol.

[0148] Specifically, this step solves the temporal coupling problem of multi-component alloying sequence in complex smelting environments by using a nonlinear optimization strategy based on reinforcement learning.

[0149] This embodiment achieves intelligent closed-loop compensation for micro-composition fluctuations during the manufacturing process of ultra-high strength steel by introducing a deep coupling between a multi-agent reinforcement learning architecture and a dynamic weight graph structure. Utilizing a deep learning model to extract deep features of the high-temperature thermodynamic evolution law not only overcomes the lag of traditional static empirical models in handling extreme conditions, but also ensures the uniformity and reliability of special steel quality on large-scale complex production lines through a data-driven real-time feedback mechanism.

[0150] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A method for constructing a knowledge graph of steel manufacturing based on multi-source heterogeneous data, characterized in that, include: Entity recognition and normalization are performed on multi-source heterogeneous data. The chemical components and contents in the message are converted into standardized semantic triples through semantic mapping rules. The missing smelting components are completed by using the physicochemical component balance logic to obtain the standardized triple entity. A multi-layer attention semantic processor is used to parse unstructured failure analysis reports. Text logic is extracted through dependency parsing, which includes the causal evolution relationship between deformation-induced martensite, dislocation density, and stress corrosion cracking. Simultaneously, a multimodal association framework is used to parse SEM fracture scan images, extract cleavage fracture visual feature vectors, and perform cross-modal semantic coupling with the text logic to obtain a steel failure mechanism enhancement entity. Causal evolution triplet data about deformation-induced martensite, dislocation density, and stress corrosion cracking are extracted from the steel failure mechanism enhancement entity and used as the topological backbone of a graph database. Initial weight scores are assigned to relation edges with different semantic strengths in the logic control unit. The system collects dynamic monitoring parameters from the entire steel manufacturing process in real time and resamples these parameters into equally spaced time-series parameters. The mapping engine receives normalized triples as initial node features and simultaneously reads in the processed equally spaced time-series parameter sequence. Through convolution operators, it performs nonlinear weight summation in a syntax-guided graph topology, mapping and aggregating the physical monitoring features of the process into corresponding metallurgical mechanism node embedding vectors in real time. Using normalized triple entities as node metadata features, the system maps causal evolution relationships into directional edges of the graph structure. Semantic feature fusion is performed on the time-series parameters, and Bayesian updates are used to correct parameter fluctuation weights within each process interval, generating a dynamic weighted graph structure.

2. The method for constructing a steel manufacturing knowledge graph based on multi-source heterogeneous data according to claim 1, characterized in that, The specific steps for converting the data into standardized semantic triples include: The multi-source heterogeneous data includes at least original laboratory test reports, LIMS quality management system data, and steel standard manuals. The process involves parsing the multi-source attribute fields in the original laboratory test reports and LIMS quality management system data, and using a pre-defined steel industry identifier mapping matrix to align non-standard grades and element abbreviations from various heterogeneous data sources to the standard ontology concepts defined in the steel standard manual. Character features are extracted from the report text, and combined with the semantic association distance in the identifier mapping matrix, context disambiguation is performed on entities with semantic ambiguity to determine the standard semantic type and process category of the entity. The content values ​​of the chemical components and their corresponding units of measurement are identified, and consistency verification is performed on the content values ​​based on physicochemical logic constraints. Values ​​with different dimensions are uniformly converted into standardized attribute scores. Based on a pre-defined knowledge graph pattern, with the aligned standard ontology concepts as head entities, the chemical components as relational attributes, and the verified attribute scores as tail entities, graph pattern mapping is performed to generate the standardized semantic triples.

3. The method for constructing a steel manufacturing knowledge graph based on multi-source heterogeneous data according to claim 1, characterized in that, The specific steps for obtaining the normalized triplet entity include: The percentage values ​​of known chemical components are extracted from the standardized semantic triplet. Based on the furnace association attributes in the multi-source heterogeneous data, the alloy composition constraint range of the corresponding grade in the steel standard manual is retrieved. Element equilibrium boundary conditions are constructed using the alloy composition constraint range. Combined with a pre-set composition evolution distribution model, the attribute probability distribution of missing smelting components under the current process path is calculated. A constraint search algorithm is used to determine the optimal completion value in the attribute probability distribution, and consistency correction is performed on the total sum of all elements in the standardized semantic triplet according to the element equilibrium boundary conditions. The corrected completion value is embedded as a tail entity into the attribute slot of the standardized semantic triplet, generating a normalized triplet entity that satisfies consistency constraints in both the semantic space and the physicochemical value space.

4. The method for constructing a knowledge graph of steel manufacturing based on multi-source heterogeneous data according to claim 1, characterized in that, The specific steps for parsing unstructured failure analysis reports and extracting text logic include: Unstructured failure analysis reports are vectorized using text feature encoding operators to extract hidden layer feature vectors containing contextual semantic information. Dependency parsing is performed on the failure analysis report to identify the dominance and subordination relationships between various technical terms in the text, and a syntactic tree structure is constructed based on these relationships. The hidden layer feature vectors are mapped to nodes in the syntactic tree structure, and semantic association weights between deformation-induced martensite, dislocation density, and stress corrosion cracking key entities are calculated using node saliency. Based on these semantic association weights, logical trigger words and directional paths in the text are identified, and causal evolution triples with temporal characteristics and logical guidance are extracted to form the textual logic corresponding to the failure analysis report.

5. The method for constructing a steel manufacturing knowledge graph based on multi-source heterogeneous data according to claim 4, characterized in that, The extraction of text logic and the formation of causal evolution triples specifically includes the following semantic analysis steps: Using the dominance weights and dependency distances between entities in the syntactic tree structure, path filtering is performed on the initially identified causal relationships to remove redundant logical arcs that are below a preset semantic threshold. By combining the semantic constraint patterns preset in the field of steel failure, the extracted text logic is subjected to directional verification, and logical conflicts in which the subject and object are reversed in the causal evolution triple are identified and corrected. Modifying semantic features are extracted from the failure analysis report, the logical confidence of the causal evolution triple is calculated, and the logical confidence is attached as a meta-attribute to the steel failure mechanism enhancement entity to realize semantic management of text logic weight.

6. The method for constructing a steel manufacturing knowledge graph based on multi-source heterogeneous data according to claim 1, characterized in that, The specific steps for parsing the SEM fracture scan image and performing semantic coupling include: A deep spatial feature extractor is used to extract global visual features and local texture features from SEM fracture scan images to generate visual feature vectors representing the morphology of cleavage fracture surfaces. A cross-modal alignment algorithm is used to project these visual feature vectors into the same semantic embedding space as the text logic, and the vector space correlation between the visual feature vectors and the causal evolution relationships in the text logic is calculated. Based on the vector space correlation, the text logic path with the highest correlation strength to the visual feature vectors is identified, and the micro-features of failure implicit in the image are extracted as supplementary logical information. The visual feature vectors are converted into semantic tags and attached as additional attributes to the corresponding causal evolution triples. A graph node enhancement mechanism is then used to generate the steel failure mechanism enhancement entity.

7. The method for constructing a knowledge graph of steel manufacturing based on multi-source heterogeneous data according to claim 1, characterized in that, The real-time acquisition of dynamic monitoring parameters throughout the entire steel manufacturing process, and the resampling of these dynamic monitoring parameters into equally spaced time-series parameters, specifically includes: The dynamic monitoring parameters include lance height, oxygen pressure, and oxygen supply intensity during the converter blowing stage; argon stirring intensity and heating rate during the refining stage; and rolling force and rolling speed during the rolling stage. The sensor data streams of each stage of steel manufacturing are monitored in real time through the process control system interface, and multidimensional dynamic monitoring parameters with non-uniform sampling periods are extracted. The physical dimension attributes and timestamp accuracy of the multidimensional dynamic monitoring parameters are identified, and a globally unified time axis is established based on the process start point. Using a cubic spline interpolation algorithm, time-domain normalization is performed on the multidimensional dynamic monitoring parameters of different frequencies according to a preset standard sampling step size, generating multidimensional feature vectors with a uniform time step size. The normalized data is then subjected to segmented aggregation processing to filter out instantaneous noise generated by sensor jitter, and the processed equally spaced time-series parameters are mapped to a dynamic attribute sequence table of the corresponding normalized triplet entities.

8. The method for constructing a steel manufacturing knowledge graph based on multi-source heterogeneous data according to claim 1, characterized in that, The specific steps for generating the dynamic weight graph structure include: The causal evolution relationship in the enhanced entity of the steel failure mechanism is mapped as the directional edge of the graph structure. The combination operator and edge type weight of the heterogeneous graph topology feature mapping engine are initialized according to the logical strength of the causal evolution relationship to construct a mechanism-guided graph topology structure. The normalized triple entity is used as the node metadata feature. The mapping engine performs multi-layer neighborhood attribute aggregation operation on the equally spaced time series parameters in the graph topology structure to realize the association vector mapping between node semantic features and dynamic monitoring features. The measured value of finished product performance is obtained in real time as a feedback signal. The posterior probability of generating the feedback signal under the current process parameter fluctuation distribution is calculated using the Bayesian update correction operator. The attention association weight between graph nodes is iteratively corrected according to the posterior probability. The edge weight parameters in the graph topology structure are reconstructed based on the corrected association weights, and the dynamic weight graph structure reflecting the real-time mapping relationship between process fluctuation and finished product performance is output.

Citation Information

Patent Citations

  • Industrial time series data learning fusion and anomaly detection method

    CN120179654A

  • Multi-path logic generation method and system based on aviation accident knowledge graph

    CN121094080A