A process industry computable graph construction method and system

CN122840202APending Publication Date: 2026-09-29UNIV OF JINAN
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611004915.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]本发明提供了一种流程工业可计算图谱构建方法及系统,用以解决传统流程工业(如新型干法水泥烧成系统)热工环境高维耦合、显着时滞带来的建模难题

Benefits of technology

1.本发明打破文本语义与严格符号代数运算的技术壁垒:依靠本发明提出的文本-公式-拓扑(TFT)解耦架构与语义对齐规范化降级层,能够成功消除大语言模型输出的带有格式缺陷和冗余的公式,将核心机理规则转换为符合计算机代数系统语法标准的有效执行模型,规则解析成功率高,且绝对保证保留方程的前向计算可计算度,填补了自然语言知识与数值符号计算之间的鸿沟。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840202A_ABST
    Figure CN122840202A_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for constructing a computable graph for process industries, relating to the fields of industrial control and digital twin technology. First, it extracts mechanistic rules using a large model, eliminating non-algebraic logic and completing operators through semantic degradation and standardization alignment layers, generating an algebraic rule set with forward computability. Second, it abstracts process entities into heterogeneous nodes and constructs an initial heterogeneous computable graph. Subsequently, it introduces a symbolic graph rewriting algorithm based on abstract syntax trees, eliminating redundant nodes while preserving thermodynamic feedback loops, significantly compressing the topology depth. Finally, it constructs a Monte Carlo steady-state inference engine by combining the steady-state mean and variance of industrial big data, outputting confidence intervals for core indicators through asynchronous relaxation iteration, and achieving online calibration. This invention completely breaks down the barriers between textual knowledge and algebraic computation, reducing inference latency and enabling precise quantification of on-site interference, providing high-fidelity intelligent control data for process industries.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of industrial control and digital twin technology, and in particular to a method and system for constructing a computable graph for process industries. Background Technology

[0002] Process industries (such as the cement industry) are accelerating their transformation towards digital twins and intelligent control. The core characteristic of digital twins is the seamless integration of virtual and physical spaces, and a key prerequisite for building high-fidelity digital twins and intelligent control systems is the realization of structured and executable representations of process mechanism knowledge. However, the new dry-process cement production involves a complex heat and mass transfer environment with gas-solid-liquid three-phase coexistence, characterized by high dimensionality, strong coupling, and significant pure time delay.

[0003] Currently, existing technologies for acquiring and constructing industrial process mechanism models face the following significant technical bottlenecks: First, Large Language Models (LLMs) suffer from impedance mismatch and grammatical format defects. While LLMs provide an effective way to extract industrial knowledge, the output mechanism formulas often exhibit format distortions and redundancy at the text level. Due to the differences between natural language mathematics and formal mathematics, the formulas extracted by large models often have implicit format defects such as missing multiplication signs or redundant LaTeX decorators, leading to the failure of computer algebra systems (CAS) to parse them and making it impossible to perform rigorous symbolic computation directly. Second, traditional knowledge graphs (KGs) have poor structural adaptability and low computational efficiency. Traditional directed acyclic graphs (DAGs) are biased towards static relation retrieval and cannot support the multiphysics feedback loops commonly found in cement production. Furthermore, thermodynamic cascades themselves generate topological redundancy, severely hindering the model's convergence speed and preventing traditional graphs from completing millisecond-level closed-loop inference operations. Third, there is a lack of steady-state uncertainty quantification for actual industrial noise. Sensors in industrial distributed control systems (DCS) rarely output purely scalar measurements, and their results are inherently accompanied by high-frequency noise interference. Traditional mechanistic modeling is often based on ideal transient differential equations. Under extremely complex factory conditions, the inherently huge data processing overhead of transient models severely restricts the real-time availability of solvers, and directly discarding interference terms can cause model predictions to deviate from reality, making it difficult to define safe boundaries for industrial production operations. Summary of the Invention

[0004] This invention provides a method and system for constructing computable graphs in process industries, which can solve the modeling problems caused by the high-dimensional coupling and significant time delay of the thermal environment in traditional process industries (such as new dry-process cement calcination systems).

[0005] On the one hand, this invention provides a method for constructing a computable graph for process industries. The method includes the following steps: calling a pre-trained large language model, injecting a qualitative knowledge graph of the process industry as prior knowledge, and generating an initial knowledge set after processing; the initial knowledge set contains process mechanism rules and mathematical formulas; constructing a semantic alignment operator to eliminate non-algebraic logic in the initial knowledge set and generate an explicit algebraic rule set; abstracting process entities of the process industry into three types of heterogeneous nodes, transforming the algebraic rule set into computational binding edges between nodes, and transforming process causal relationships into causal driving edges, thus constructing an initial heterogeneous computable graph; performing topology optimization on the initial heterogeneous computable graph based on an abstract syntax tree, identifying and eliminating redundant intermediate nodes, introducing a strict loop protection mechanism, strictly limiting algebraic fusion to acyclic forward paths, and compressing the graph topology depth while preserving the inherent feedback loop structure of the industrial system; embedding an asynchronous relaxation iterative solver into the optimized computable graph, based on the steady-state mean extracted from the industrial large dataset. With variance Inject Gaussian distributions into the manipulated variable nodes Noise is used to generate the probability distribution and 95% confidence interval of key performance indicators through Monte Carlo sampling and iterative inference.

[0006] In one implementation of the present invention, the method further includes: using real-time collected industrial big data, combined with an orthogonal variance synthesis algorithm, periodically updating and adjusting the node probability distribution, edge association strength and algebraic rule parameters of the computable graph, so as to achieve adaptive synchronization between the graph and the actual production conditions.

[0007] In one implementation of the present invention, the processing of the semantic alignment operator is specifically as follows: constructing a syntax verification mechanism based on an abstract syntax tree, traversing the syntax tree nodes of all formulas in the initial knowledge set; automatically completing all missing binary operators such as multiplication and division; removing all LaTeX format decorators and non-mathematical symbols; eliminating non-algebraic rules containing fuzzy qualitative descriptions and fuzzy scale constraints, and retaining only algebraic equations and inequalities that can be directly used for calculation and execution.

[0008] In one implementation of the present invention, the execution process of the abstract syntax tree specifically includes: using the Tarjan algorithm to identify all directed cycles and strongly connected components in the initial heterogeneous computable graph; utilizing a loop protection mechanism to perform operator fusion on intermediate variable nodes on acyclic forward paths at the abstract syntax tree level, compressing multi-step continuous algebraic operations into single-step composite operations; for strongly connected components containing feedback loops, preserving their loop structure to prevent logic explosion caused by unconstrained recursion; and ensuring that the maximum topological depth of the optimized graph does not exceed 3 layers, while strictly maintaining the thermodynamic equivalence of the original system.

[0009] In one implementation of the present invention, the Monte Carlo processing procedure is as follows: for each Monte Carlo sampling sample, input values ​​are randomly sampled from the Gaussian noise distribution of the manipulated variable node; an asynchronous and lock-free update mechanism independent of the predecessor node is used to sequentially update the state values ​​of all nodes along the graph topology path; iterative calculation is performed until the global graph residual is less than a preset convergence threshold, and the key performance index value corresponding to the sample is output; the calculation results of all samples are statistically analyzed to generate the probability density function and 95% confidence interval of the key performance index in all simulation results.

[0010] In one implementation of the present invention, before calling the pre-trained large language model, the method further includes: collecting real-time DCS operation data, laboratory test data, equipment status data and raw material composition data from the process industry production site, performing data cleaning, outlier removal and time-series alignment, and constructing a standardized industrial big data set.

[0011] In one implementation of the present invention, the process of generating the initial knowledge set is as follows: the qualitative knowledge graph of the process industry accumulated in the previous research is injected as prior knowledge, a pre-trained large language model is called, and the structured prompt word engineering and the self-reflection mechanism of the large model are combined. Through multiple rounds of cross-validation and self-correction by comparing with the prior knowledge graph, the process mechanism rules and mathematical formulas are extracted from the unstructured documents with high precision to generate the initial knowledge set.

[0012] In one implementation of the present invention, the method is implemented using a cloud-edge-device collaborative architecture. The implementation of the cloud-edge-device collaborative architecture specifically includes: deploying a lightweight computable graph inference engine at the edge to receive real-time data from the field DCS and output millisecond-level control commands; deploying a large language model knowledge extraction module and a graph dynamic calibration module at the cloud to periodically push updated graph parameters to the edge; and realizing data interaction and command issuance between the edge, the cloud, and the field control system through the OPC UA industrial internet protocol at the device side.

[0013] In one implementation of the present invention, the dynamic calibration of the graph based on industrial big data specifically involves: collecting the latest industrial production data in a fixed time window; calculating the deviation between the inference results of the computable graph and the actual production data; introducing an orthogonal variance synthesis algorithm and using automatic grid search to dynamically filter out the optimal adjustment factor parameters within a preset range; and using the optimal adjustment factor to update the steady-state mean and variance parameters of the nodes to complete the online adaptive synchronization and benchmark calibration of the graph in response to changes in the underlying data.

[0014] On the other hand, the present invention also provides a process industry computable map construction system, the system comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to complete the aforementioned process industry computable map construction method.

[0015] The present invention provides a method and system for constructing a computable map for process industries, which has the following beneficial effects: 1. This invention breaks through the technical barriers between text semantics and strict symbolic algebra operations: Relying on the text-formula-topology (TFT) decoupling architecture and semantic alignment normalization degradation layer proposed in this invention, it can successfully eliminate the formulas with format defects and redundancy output by large language models, transform the core mechanism rules into an effective execution model that conforms to the syntax standards of computer algebra systems, with a high rule parsing success rate, and absolutely guaranteeing the computability of forward computation of equations, thus bridging the gap between natural language knowledge and numerical symbolic computation.

[0016] 2. This invention significantly reduces the graph topology depth, enabling millisecond-level closed-loop control in industrial applications: Addressing the inherent topological redundancy in process systems, this invention employs a symbolic graph rewriting optimization algorithm and loop protection mechanism based on Abstract Syntax Trees (ASTs). While maintaining complete physical and thermodynamic equivalence, this successfully eliminates invalid intermediate multi-step computation nodes, compressing the maximum topology depth of the graph. This shortens the full-graph inference latency of 2000 parallel Monte Carlo relaxation iterations, transforming high-performance real-time digital twin closed-loop control into an engineering reality.

[0017] 3. This invention accurately quantifies high-frequency sensor noise and defines safe production boundaries: This invention explicitly introduces industrial sensor noise into the steady-state inference framework and incorporates it into the uncertainty probability space modeling. After interval calibration through bias correction and orthogonal variance synthesis (including correction factor grid search), the mean absolute error of the model prediction is reduced, which can effectively suppress the interference of on-site sensor noise on decision-making and provide operators and APC controllers with quantifiable and interpretable safety margins. Attached Figure Description

[0018] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings: Figure 1 A flowchart of a method for constructing a computable map of a process industry, provided by an embodiment of the present invention; Figure 2A schematic diagram of the steady-state fluctuation of the kiln speed, an operational variable, and the input calibration confidence band generated by rolling mean variance compensation over 40 consecutive batches provided in this embodiment of the invention. Figure 3 A schematic diagram of the steady-state fluctuation of the manipulated variable Total Coal FeedRate and the input calibration confidence band for 40 consecutive batches provided in this embodiment of the invention; Figure 4 This is a schematic diagram of the steady-state fluctuation of the total feeding rate (Total FeedingRate) and the input calibration confidence band over 40 consecutive batches provided in this embodiment of the invention. Figure 5 For the embodiments of the present invention, a comparison chart of the Monte Carlo steady-state robustness prediction confidence band of clinker lime saturation coefficient Lab KH output by the HCG engine of the present invention and the coverage of actual laboratory measurement data is provided for 40 consecutive batches. Figure 6 This is a diagram illustrating the composition of a computable graph construction system for process industries, provided as an embodiment of the present invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.

[0020] This invention provides a method and system for constructing computable graphs in process industries, supported by a text-formula-topology (TFT) decoupling architecture. The technical solutions proposed in this invention will be described in detail below with reference to the accompanying drawings.

[0021] Figure 1 This is a flowchart illustrating a method for constructing a computable graph for a process industry, as provided in an embodiment of the present invention. Figure 1 As shown, the method mainly includes the following steps: Step S1: Multi-source industrial big data collection and quasi-steady-state assumption preprocessing.

[0022] This method collects real-time operating data, laboratory test data, equipment status data, and raw material composition data from the distributed control system (DCS) at the production site in process industries (taking a new dry-process cement calcination system as an example). Due to the extreme complexity of actual plant conditions, the primary first-order thermodynamic differential equation for energy evolution within the system is:

[0023] In the formula, Input heat into the system, To output heat to the system, The heat released or absorbed by the material reaction. For material quality, For specific heat capacity, This represents the rate of change of transient temperature over time.

[0024] To eliminate the inherently large data processing overhead of transient differential solvers and ensure the real-time availability of the model, this method implements a data-driven quasi-steady-state simplification: by statistically analyzing continuous DCS operation logs, within a specific sliding time window when the large thermal inertial system is in normal and stable operating conditions (this embodiment uses a standard 30-minute production window, during which the main control parameters such as rotary kiln speed, feed rate, and coal feed rate fluctuation are extremely small), it is assumed that the rate of temperature change tends to zero, i.e., satisfying the condition... The physical constraints are thus filtered out, and the difficult differential terms in formula (1) are successfully transformed into steady-state algebraic constraints that can be directly solved numerically. At the same time, the collected time series data are cleaned and extreme abnormal mutation values ​​that deviate significantly from the normal values ​​are removed to construct a standardized industrial big data set.

[0025] Step S2: Extracting mechanistic knowledge from a large model based on prior knowledge graphs and reflection mechanisms.

[0026] The qualitative knowledge graph of the process industry (covering basic process entity relationships and qualitative causal chains), accumulated from previous research, is injected into the system as a prior knowledge base. A pre-trained large language model (Qwen3.6-plus model is used in this embodiment), combined with structured prompt engineering, is invoked to batch process unstructured text from authoritative process industry literature, guided by the aforementioned qualitative prior knowledge. During the extraction and parsing process, a self-reflection mechanism is specifically introduced: after the system drives the large model to generate preliminary extraction results through prompts, it automatically guides the model to perform multiple rounds of cross-validation and self-correction against the logical constraints of the prior qualitative knowledge graph. Through this hybrid mechanism of "prior guidance + reflective verification," the system extracts material balance formulas, heat balance formulas, clinker mineral composition mechanisms (such as the Bogue formula), and basic engineering constraints such as energy conservation and mass conservation across the entire process flow with high precision, ultimately generating an initial set of mechanistic rules knowledge containing a precise mapping relationship between the original text string and the mechanistic formulas.

[0027] Step S3: Semantic degradation and canonical alignment processing.

[0028] Because the initial formulas extracted by large language models often have implicit formatting defects (such as missing binary arithmetic operators, redundancy) Combinations such as wrapper decorators and punctuation obfuscation cannot be directly parsed using a computer algebra system (CAS). This method constructs a semantic alignment operator. As a strict structural filter, the original rule string... Perform systematic restructuring and downgrade standardization:

[0029] This operator traverses each syntax tree node of the formula in the initial rule, forcibly performing binary operator completion (e.g., executing...). The system automatically and explicitly rewrites rules, completely stripping away all irrelevant textual embellishments. Furthermore, it integrates a syntax verification mechanism based on Abstract Syntax Trees (ASTs), rigorously eliminating any qualitative empirical rules containing fuzzy scale constraints (such as "slightly increased" or "significantly hotter") or non-algebraic qualitative descriptions during the parsing phase. While this hard limit restricts the overall parsing success rate of the final rules to 50.5%, it absolutely guarantees that every algebraic equation or inequality retained and entered into subsequent stages possesses 100% forward algebraic validity and direct computability, thereby generating an explicit algebraic rule set.

[0030] Step S4: Initial construction of heterogeneous computable graph.

[0031] Process industry entities are abstracted and defined as heterogeneous computable graphs (HCGs) with exclusive tensor characteristics, and their mathematical structure is formally expressed as set pairs:

[0032] in, For a heterogeneous node set, This method uses heterogeneous edge sets. The node set is determined based on the specific control attributes of the process industry. Strictly divided into three categories: (1) Manipulated Variable (MV) node: corresponding to the control input (such as the source node carrying the coal powder feed rate, rotary kiln speed, etc.), which carries the real-time observation features and steady-state mean extracted from the underlying big data set. ) and variance ); (2) Process Variable (PV) node: The corresponding intermediate transmission end (such as the transmission node of the outlet temperature of each stage of preheater, local negative pressure, etc.) serves as the intermediate transmission channel responsible for the topological transmission of the driving distribution tensor; (3) Key Performance Indicator (KPI) nodes: corresponding to the quality target aggregation nodes (such as the clinker lime saturation coefficient KH value, tricalcium silicate content, etc. in laboratory tests), and finally the probability distribution is generated through iterative deductive reasoning within the graph.

[0033] Accordingly, edge set It is also divided into two categories: (1) Causal driving edge: responsible for encoding the physical interaction dynamics between process parameters and the forward / backward directional physical driving relationship, faithfully restoring the dynamic heat and mass exchange process of gas-solid-liquid three phases; (2) Computation binding edge: enforces the strict algebraic operation mapping relationship between interconnected nodes, clearly and explicitly defining the input and output routing boundaries of the formula.

[0034] In reasoning and deductive topological calculations, this graph breaks through the limitation of traditional directed acyclic graphs (DAGs) being unable to represent industrial feedback loops. It adopts a relaxed iterative mechanism based on asynchronous updates, and its state evolution equation is as follows:

[0035] In the formula, For the first The node at the th The state feature tensor at the next iteration For all predecessor nodes of this node in the th... The state feature vector at the next iteration. These are algebraic operator functions bound to computational edges. The system eventually converges through multiple iterations to a stable fixed point that accurately reflects the operating characteristics of the rotary kiln or preheater.

[0036] Step S5: AST-guided symbol graph rewriting optimization.

[0037] To address the topological overlap and computational redundancy inherent in multi-level heat conduction and thermodynamic cascades, this method employs the Tarjan algorithm during the preprocessing stage to identify all directed feedback loops and strongly connected components (SCCs) in the initial heterogeneous computable graph. To prevent unconstrained recursion from causing severe logic explosions and deadlocks in computer symbolic systems, a strict loop protection mechanism is implemented: symbol-level algebraic fusion operations are strictly restricted to acyclic forward paths.

[0038] For consecutive operators on a cascaded path and The algorithm performs direct node fusion and simplification rewriting of intermediate redundant variables at the Abstract Syntax Tree (AST) level:

[0039] Using the node fusion technique guided by the aforementioned AST, 140 redundant and invalid intermediate operator nodes were systematically eliminated and compressed in the experiment, significantly reducing the maximum topology cascading depth from 8 layers to 3 layers. This topology simplification process was completed automatically within the computer using symbolic algebra tools, and mathematically, it strictly maintained the first-order thermodynamic equivalence of the original system, thereby effectively reducing the computational complexity of full-graph reasoning.

[0040] Step S6: Building the Monte Carlo steady-state inference engine.

[0041] The hardware and software implementation details of the Monte Carlo steady-state inference engine construction and asynchronous relaxation solution in step S6 are as follows: To achieve highly robust forward mechanism deduction under high-frequency random environmental noise and electromagnetic interference in real-world industrial scenarios, this embodiment constructs a Monte Carlo steady-state inference engine based on the graph probability tensor space. This engine relies on high-parallel computing hardware (preferably an RTX 4080 graphics processor in this embodiment) and utilizes explicitly injected normal noise and an asynchronous lock-free relaxation iterative algorithm based on heterogeneous graph topology for efficient inference. The specific execution steps are as follows: 1) The probabilistic sample expansion and Gaussian noise injection inference engine first reads the industrial DCS steady-state aligned dataset obtained in step S1 and obtains the measurement mean of each manipulated variable (MV, including kiln speed, total coal consumption, and total feed rate) in the current sampling batch. With volatility and variance .like Figure 2 As shown, in Figure 2 The text demonstrates the steady-state fluctuations of kiln speed in a real-world industrial setting and the effectiveness of its confidence band extraction. Figure 3 As shown, in Figure 3 The diagram shows the steady-state fluctuations and confidence bands corresponding to the total coal feedrate (Total Coal Feed Rate); for example... Figure 4 As shown, in Figure 4 The figures above illustrate the steady-state fluctuations and confidence bands of the total feeding rate. These three sets of figures visually demonstrate the inherent high-frequency noise characteristics of the DCS sensor, providing a realistic industrial physical boundary for subsequent probability sample expansion.

[0042] To simulate sensor measurement noise at the software level, the engine incorporates a Gaussian random number generator to perform probabilistic sample expansion. Specifically, for each type of manipulated variable node, a dimension of [missing value] is allocated in the GPU memory using the PyTorch GPU tensor framework. probability sample tensor The sampling scale is set in this plan. The calculation formula is as follows: In the formula, the standard normal random variable satisfies Subscript {} represents the sample index; It is a minimal smoothing constant used to avoid numerical underflow anomalies in square root operations when the variance approaches 0.

[0043] 2) The asynchronous relaxation solution based on heterogeneous topology graphs for solving the thermal-mass balance equations of the cement firing system involves coupled causal loops. Traditional unidirectional topology sorting cannot achieve closed-loop solutions. Therefore, this embodiment designs an asynchronous relaxation solution strategy adapted to heterogeneous graphs: Initialization phase: The MV multidimensional sample tensor injected with Gaussian noise is stored in a global sample state container as the calculation input; Iterative phase: The engine iterates through all calculation rule nodes, with the maximum number of iterations constrained to be 3 times the total number of rule nodes, defined as follows: Single-round computation logic: Verify if all predecessor variable tensors for the current rule are ready; if ready, call the SymPy lambdify pre-compiled vectorization mechanism function to perform 2000 sets of nonlinear operations on the GPU in parallel, generating the current node's output tensor; after computation, directly write to the storage container of all subsequent PV nodes without waiting for other branch operations, achieving asynchronous, lock-free data stream transmission; convergence termination condition: the global change in the full graph output of two adjacent rounds is lower than a preset threshold, or the number of iterations reaches... Stop the relaxation iteration.

[0044] 3) Robust forward numerical cleaning mechanism: To avoid the gradual spread of NaN and ±∞ caused by illegal operations such as division by zero and negative logarithms under extreme conditions during graph iteration, tensor numerical cleaning is performed after each formula calculation: all finite real numbers in the tensor are extracted, and the median is calculated as the filling reference value. If the tensor has no valid finite values, the default padding value is set to 0.0; all invalid values ​​are replaced in batches in the video memory using the built-in nan_to_num operator in PyTorch, ensuring stable convergence of the multi-ring coupled graph iteration throughout the process.

[0045] 4) After the performance relaxation iteration converges in the extraction and inference of confidence intervals for quality indicators, all M groups of Monte Carlo samples from LabKH at the KPI node are extracted, and the 95% two-sided confidence boundary is solved using the quantile operator: Simultaneously output the sample mean and standard deviation to generate confidence bands. For example... Figure 5 As shown, in Figure 5 The figure shows the coverage comparison between the LabKH prediction confidence band of lime saturation coefficient and the actual measured value in the field test. The attached figure intuitively verifies that the output range of the engine can effectively encompass real industrial data, proving that the solution has excellent robustness.

[0046] Leveraging RTX 4080 GPU matrix parallel computing, the total time for a single batch of 2000 sets of full-map Monte Carlo inference is only 15.69 milliseconds; compared with the CPU sample-by-sample serial solution based on NumPy, the overall computational efficiency is improved by 43%. This inference latency is far lower than the 2-5 second control cycle of a cement plant DCS, and can be directly connected to the industrial online closed-loop control system, providing millisecond-level real-time mechanism deduction support.

[0047] Step S7: Dynamic calibration of the map based on industrial big data.

[0048] To compensate for the inability of purely theoretical thermodynamic formulas or empirical formulas to fully cover the complex systematic deviations that occur during actual calcination, such as coal ash adsorption deviations and heat transfer loss drift, this method embeds a data-driven online calibration loop in the outer layer of the spectrum. The system periodically collects the latest big data from industrial production at a fixed sliding time window and dynamically calculates the mean absolute error (MAE) between the inferred mean of the spectrum and the actual measured value (such as the clinker saturation coefficient measured in the laboratory).

[0049] The calibration system incorporates an orthogonal variance synthesis algorithm, rather than manual adjustment, using an automatic grid search algorithm within a defined mathematically and physically valid range (e.g., ...). Within a given range, a global traversal optimization is performed with a fine step size (e.g., 0.01) to accurately calculate and filter out the optimal process adjustment factor that can compensate for the defects in the theoretical mechanism. By utilizing this optimal adjustment factor, the weight matrix and parameters of the algebraic rules within the graph are corrected in real time, enabling online dynamic adaptive updates and benchmark synchronization of the graph. This allows the model to have the ability to self-calibrate in the event of sudden changes in underlying materials, proportions, or environmental conditions.

[0050] The above is an embodiment of the present invention providing a method for constructing a computable graph for a process industry. Based on the same inventive concept, the present invention also provides a system for constructing a computable graph for a process industry, such as... Figure 6 As shown, the device mainly includes: at least one processor 601; and a memory 602 communicatively connected to the at least one processor; wherein the memory 602 stores instructions that can be executed by the at least one processor 601, and the instructions are executed by the at least one processor 601 to enable the at least one processor 601 to complete the aforementioned method for constructing a computable map of a process industry.

[0051] The present invention specifically achieves the following technical effects: 1. Breaking down the technical barriers between text semantics and strict symbolic algebra operations: Relying on the text-formula-topology (TFT) decoupling architecture and semantic alignment normalization degradation layer proposed in this invention, it is possible to successfully eliminate the formulas with format defects and redundancy output by large language models, and transform the core mechanism rules into an effective execution model that conforms to the syntax standards of computer algebra systems. The rule parsing success rate is high, and the computability of the forward computation of the equation is absolutely guaranteed, thus bridging the gap between natural language knowledge and numerical symbolic computation.

[0052] 2. Significantly reduces graph topology depth, enabling millisecond-level closed-loop control in industrial applications: Addressing the inherent topological redundancy in process systems, this invention employs a symbolic graph rewriting optimization algorithm and loop protection mechanism based on Abstract Syntax Trees (ASTs). While maintaining complete physical and thermodynamic equivalence (with an equivalence error strictly of 0.0), it successfully eliminates invalid intermediate multi-step computation nodes, compressing the maximum topology depth of the graph. This shortens the full-graph inference latency of 2000 parallel Monte Carlo relaxation iterations, transforming high-performance real-time digital twin closed-loop control into an engineering reality.

[0053] 3. Precise Quantification of High-Frequency Sensor Noise and Defining of Safe Production Boundaries: This invention explicitly introduces industrial sensor noise into the steady-state inference framework, incorporating it into the uncertainty probability space modeling. After interval calibration through bias correction and orthogonal variance synthesis (including correction factor grid search), the mean absolute error (MAE) of the model prediction decreased from 0.0027 to 0.0010. Based on six consecutive months of historical DCS data from the kiln line, and after windowed statistical dimensionality reduction and alignment with laboratory test samples, the results of 40 consecutive steady-state batch stress tests show that the 95% prediction confidence interval of the model output achieves 92.50% empirical coverage of the dynamic changes in core KPI quality parameters. This result demonstrates that this method can effectively suppress the interference of on-site sensor noise on decision-making and provide operators and APC controllers with quantifiable and interpretable safety margins.

[0054] In addition, embodiments of the present invention also provide a non-volatile computer storage medium for constructing a computable graph for a process industry, which stores computer-executable instructions, which are executed by a processor to implement the aforementioned method for constructing a computable graph for a process industry.

[0055] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0056] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0057] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0058] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0059] The various embodiments in this invention are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0060] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0061] The above description is merely an embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the present invention should be included within the scope of the claims of the present invention.

Claims

1. A method for constructing a computable map of a process industry, characterized in that, The method includes the following steps: A pre-trained large language model is invoked, and a qualitative knowledge graph of the process industry is injected as prior knowledge. After processing, an initial knowledge set is generated. The initial knowledge set contains process mechanism rules and mathematical formulas. A semantic alignment operator is constructed to eliminate non-algebraic logic in the initial knowledge set and generate an explicit algebraic rule set. The process entities of the process industry are abstracted into three types of heterogeneous nodes, and the algebraic rule set is transformed into computational binding edges between nodes. The process causal relationship is transformed into causal driving edges, and an initial heterogeneous computable graph is constructed. The initial heterogeneous computable graph is topologically optimized based on the abstract syntax tree, redundant intermediate nodes are identified and eliminated, a strict loop protection mechanism is introduced, and algebraic fusion is strictly restricted to acyclic forward paths, compressing the graph topology depth while preserving the inherent feedback loop structure of the industrial system. An asynchronous relaxation iterative solver is embedded in the optimized computable graph, based on the steady-state mean extracted from an industrial large dataset. With variance Inject Gaussian distributions into the manipulated variable nodes Noise is used to generate the probability distribution and 95% confidence interval of key performance indicators through Monte Carlo sampling and iterative inference.

2. The method for constructing a computable map of a process industry according to claim 1, characterized in that, The method also includes: using real-time collected industrial big data, combined with an orthogonal variance synthesis algorithm, periodically updating and adjusting the node probability distribution, edge association strength and algebraic rule parameters of the computable graph, so as to achieve adaptive synchronization between the graph and the actual production conditions.

3. The method for constructing a computable map of a process industry according to claim 1, characterized in that, The processing procedure of the semantic alignment operator is as follows: Construct a syntax verification mechanism based on an abstract syntax tree by traversing the syntax tree nodes of all formulas in the initial knowledge set; Automatically completes all missing multiplication and division binary operators; Remove all LaTeX formatting decorators and non-mathematical symbols; Non-algebraic rules containing fuzzy qualitative descriptions and fuzzy scale constraints are removed, and only algebraic equations and inequalities that can be directly used for calculation are retained.

4. The method for constructing a computable map of a process industry according to claim 1, characterized in that, The execution process of the abstract syntax tree is as follows: The Tarjan algorithm was used to identify all directed loops and strongly connected components in the initial heterogeneous computable graph. By utilizing the loop protection mechanism, operator fusion is performed on intermediate variable nodes on the acyclic forward path at the abstract syntax tree level, compressing multi-step continuous algebraic operations into single-step compound operations. For strongly connected components containing feedback loops, preserve their loop structure to prevent logic explosion caused by unconstrained recursion. The maximum topological depth of the optimized map does not exceed 3 layers, and the thermodynamic equivalence of the original system is strictly maintained.

5. The method for constructing a computable map of a process industry according to claim 1, characterized in that, The Monte Carlo processing procedure is as follows: For each Monte Carlo sampling sample, input values ​​are randomly sampled from the Gaussian noise distribution of the manipulated variable node; An asynchronous and lock-free update mechanism independent of the predecessor node is adopted to update the state values ​​of all nodes sequentially along the graph topology path; Iterative calculations continue until the global graph residual is less than the preset convergence threshold, then the key performance index value corresponding to the sample is output. The calculation results of all samples are statistically analyzed, and the probability density function and 95% confidence interval of the key performance indicators in all simulation results are generated.

6. The method for constructing a computable map of a process industry according to claim 1, characterized in that, Before calling the pre-trained large language model, the method also includes: collecting real-time DCS operation data, laboratory test data, equipment status data and raw material composition data from the process industry production site, performing data cleaning, outlier removal and time-series alignment, and constructing a standardized industrial big data set.

7. The method for constructing a computable map of a process industry according to claim 1, characterized in that, The process of generating the initial knowledge set is as follows: the qualitative knowledge graph of the process industry accumulated in the previous research is injected as prior knowledge, the pre-trained large language model is called, and the structured prompt word engineering and the self-reflection mechanism of the large model are combined. Through multiple rounds of cross-validation and self-correction by comparing with the prior knowledge graph, the process mechanism rules and mathematical formulas are extracted from the unstructured documents with high precision to generate the initial knowledge set.

8. The method for constructing a computable map of a process industry according to claim 1, characterized in that, The method is implemented using a cloud-edge-device collaborative architecture, the implementation of which specifically includes: A lightweight, computable graph inference engine is deployed at the edge to receive real-time data from the on-site DCS and output millisecond-level control commands. The knowledge extraction module and dynamic graph calibration module of the large language model are deployed in the cloud, and the updated graph parameters are pushed to the edge on a regular basis. The edge device uses the OPC UA industrial internet protocol to enable data interaction and command issuance between the edge device, cloud, and field control system.

9. The method for constructing a computable map of a process industry according to claim 2, characterized in that, The specific steps of dynamic map calibration based on industrial big data are as follows: The latest industrial production data is collected by sliding within a fixed time window; Calculate the deviation between the inference results of the computable graph and the actual production data; An orthogonal variance synthesis algorithm is introduced, and the optimal adjustment factor parameters are dynamically filtered out within a preset range using an automatic grid search. By using the optimal adjustment factor to update the steady-state mean and variance parameters of the nodes, the map can achieve online adaptive synchronization and benchmark calibration when responding to changes in the underlying data.

10. A computable map construction system for process industries, characterized in that, The system includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a process industry computable map construction method as described in any one of claims 1-9.