An industrial energy state estimation method based on size model cooperation

CN122527596APending Publication Date: 2026-08-07DALIAN UNIV OF TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
DALIAN UNIV OF TECH
Filing Date
2026-07-08
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

在实际工业应用场景中,工业数据普遍存在噪声干扰强、数据采集不稳定、数据分布随工况变化而发生漂移以及数据缺失等问题,使得数据质量与系统运行状态之间存在显著不确定性

Benefits of technology

[0056]本申请提出了一种大小模型协同的工业能源状态估计方法,该方法可将大语言模型的推理能力引入到工业能源系统状态估计任务。在该方法中,在考虑大语言模型推理可靠性的基础上,将大语言模型的推理结果视为高层语义观测,其不确定性通过能反映推理置信度的协方差建模,进而作为状态语义伪观测估计参与状态估计过程。进一步,通过可微卡尔曼滤波结构,在考虑状态数值估计与状态语义伪观测估计不确定性差异的基础上,自适应调节卡尔曼增益,实现两类信息在统一动态系统框架下的合理融合。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122527596A_ABST
    Figure CN122527596A_ABST
Patent Text Reader

Abstract

The application belongs to the field of information technology and relates to an industrial energy state estimation method based on large and small models. The reasoning ability of a large language model is introduced into the industrial energy system state estimation task. In the method, the reasoning result of the large language model is regarded as high-level semantic observation on the basis of considering the reasoning reliability of the large language model, the uncertainty thereof is modeled and measured through covariance that can reflect reasoning confidence, and then the reasoning result is used as state semantic pseudo-observation estimation to participate in the state estimation process. Further, through a differentiable Kalman filter structure, the Kalman gain is adaptively adjusted on the basis of considering the uncertainty difference between state numerical estimation and state semantic pseudo-observation estimation, so that the two types of information are reasonably fused in a unified dynamic system framework.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of information technology and relates to an industrial energy state estimation method that uses a combination of large and small models. Background Technology

[0002] Industrial production is a high-energy-consuming and high-emission process, in which energy media play a crucial role. With the increasing scarcity of primary energy sources such as coal and oil, fully utilizing secondary energy sources in industrial production processes can not only improve energy conservation and emission reduction levels for enterprises but also reduce environmental pollution and carbon emissions. To optimize the operation of industrial energy systems, on-site personnel need to promptly grasp the trends in energy generation, consumption, and temporary storage. Therefore, accurate estimation of energy operating status can timely map the patterns of energy media generation, consumption, and storage changes, providing key support for the rational utilization, coordinated scheduling, and optimized operation of industrial energy systems.

[0003] In existing research, the state estimation methods for industrial energy systems mainly rely on energy time series data modeling. By mining the statistical characteristics and time-series dependencies of historical operating data, the system state change process is learned, thereby achieving state estimation or trend prediction. Examples include prediction interval models based on Gaussian mixture kernel functions (J. Zhao, C. Sheng, W. Wang, W. Pedrycz, and Q. Liu, “Data-based predictive optimization for byproduct gas system in steel industry,” IEEE Trans. Autom. Sci. Eng., vol. 14, no. 4, pp.1761–1770, Oct. 2017.), state estimation models integrating convolutional neural networks and long short-term memory networks (P. Lan, D. Han, X. Xu, Z. Yan, X. Ren, and S. Xia, “Data-driven state estimation of integrated electric-gas energy system,” Energy, vol. 252, p. 124049, 2022.), and data-model interactive prediction methods based on bidirectional gated recurrent units and temporal saliency attention mechanisms (BiGRU-TSAM) (J. Zhang, et al., “A data-model interactive remaining useful lifeprediction approach of lithium-ion batteries based on (PF-BiGRU-TSAM, IEEE Trans. Ind. Informat., vol. 20, no. 2, pp. 1144–1154, Feb. 2024.). The advantage of the above method lies in its ability to learn the dynamic characteristics of the system's operating state and its strong ability to represent the temporal correlations and periodic changes inherent in the data. However, in industrial scenarios, a large amount of production process and equipment information (such as production rhythm and equipment operating load) is interrelated with the operating state of the energy system. Modeling based solely on energy-side data may affect the performance of state estimation and prediction under production-energy coupling conditions.

[0004] To overcome the limitations of single-energy time series modeling in production-energy coupling scenarios, previous studies have further incorporated production factors and process constraints into the industrial energy system model framework to enhance the model's ability to express complex operating mechanisms and changes in operating conditions. Examples include a time-varying dynamic Bayesian network (TVDBN) model that considers the connectivity between gas holders, pressure differential constraints, and interlocking control laws (L. Chen, Y. Yan, J. Zhao, and W. Wang, “Dynamic modeling of multi operating conditions for a three-converter gas holder system in steel industry by using time-varying dynamic Bayesian networks,” IEEE Trans. Instrum. Meas., vol. 74, pp.1–13, 2025.), and a dynamic scheduling and prediction framework for byproduct gas systems that integrates expert knowledge and production plan (T. Wang, J. Zhao, Q. Xu, W. Pedrycz, and W. Wang, “A dynamic scheduling framework for byproduct gas system combining expert knowledge and production plan,” IEEE Trans. Autom. Sci. Eng., vol. 20, no. 1, pp. 541–552, Jan.2023.). The aforementioned methods, by modeling the correlation between production factors such as production plans and equipment operating status and energy fluctuations, integrate key information from the production side into the state estimation or trend prediction framework, which helps improve the accuracy of state estimation. However, they are limited to using a single type of time series data as input, or simply adding mechanistic information as a regularization term to the model's loss function, and lack the ability to integrate unstructured information such as production status and planning data. On the other hand, traditional small models struggle to understand and utilize the logical relationships between multimodal knowledge to achieve complex reasoning tasks.

[0005] With the development of artificial intelligence technology, large language models have demonstrated strong capabilities in complex information modeling, cross-type data understanding, and multi-source information fusion. Existing technologies attempt to map time-series data, multimodal data, or industrial operation information into a unified representation space. By introducing language modeling and inference mechanisms, they can comprehensively analyze and predict the state of complex systems, thereby expanding the modeling capabilities of traditional data-driven methods in complex scenarios. In practical industrial applications, industrial data generally suffers from strong noise interference, unstable data acquisition, data distribution drifting with changing operating conditions, and data gaps, resulting in significant uncertainty between data quality and system operating status. Under these conditions, existing state estimation and inference methods based on large language models are prone to producing inference results inconsistent with the actual engineering operating state, thus affecting the reliability and security of subsequent control decisions and scheduling strategies. Furthermore, large language models are not sensitive to the characteristics of time-series data, thus losing some of the dynamic characteristics of industrial data. Therefore, in industrial scenarios, leveraging the semantic modeling and inference advantages of large language models while supplementing the dynamic characteristics of industrial time series data and ensuring the confidence of inference results becomes a key prerequisite for their engineering application. Summary of the Invention

[0006] The technical solution of this application is as follows:

[0007] In the task of state estimation of industrial energy systems, large language models can be used to integrate multiple types of industrial semantic information, providing high-level semantic understanding and reasoning capabilities for identifying process operation modes. However, since the reasoning process of large language models is not sensitive to the dynamic characteristics of time series, and mostly relies on static statistical features for semantic modeling, it lacks consideration for the dynamic characteristics of time series data. At the same time, due to the fact that large language models are essentially parameterized models based on conditional probability modeling, they are prone to outputting results that do not conform to the rules of industrial mechanisms. To overcome these two shortcomings, this application proposes a collaborative method for industrial energy state estimation using both large and small models. This method compensates for the dynamic characteristics of the output of the large language model by introducing a differentiable state transition matrix, and uses the data statistical laws learned by the differentiable state transition matrix to correct and supplement the semantic output of the large language model. The collaborative optimization of the two will help improve the accuracy of state estimation. The overall method mainly includes three steps: first, constructing a cognitive fusion filtering part driven by semantic confidence; second, establishing a differentiable state transition matrix for state estimation; and third, constructing a Kalman filter structure.

[0008] Step 1: Construct the semantic confidence-driven cognitive fusion filtering component;

[0009] Step 1.1, Construction of semantic pseudo-observations:

[0010] The core of Kalman filtering lies in using observed values ​​to correct state estimates and adjusting the fusion weights of the two through Kalman gain. However, in industrial energy state estimation tasks, the research object is often focused on predicting future operating states. Since real observational data for future moments is unavailable in the prediction phase, it is difficult to directly construct the physical observation terms in the traditional Kalman filtering framework. Although the actual future state is not directly obtainable, real industrial systems still possess a wealth of semantic operating condition information related to future operating trends, such as production plans, operating condition descriptions, scheduling instructions, abnormal maintenance information, and expert experience rules. While this type of information does not belong to traditional numerical observations, it can provide prior constraints and directional judgments on the future evolution trend of the system from a semantic perspective. Therefore, it can be regarded as a kind of "semantic observation" and introduced into the Kalman filter update process, thereby providing "semantic pseudo-observations" for industrial energy state estimation in the absence of real observations.

[0011] Based on the above discussion, this application generalizes the traditional Kalman filter concept into cognitive fusion filtering. First, semantic working condition information such as production plans is input into a large language model. Then, the inference output of the large language model generates three types of complementary information based on cue word engineering according to a preset output template: 1) the final state estimation result, i.e., semantic state estimation. 1) Used to represent the prediction results of the large language model for the target state; 2) Supplementary explanations of the trend of the prediction results, i.e., trend inference text. Its content covers key production information such as production patterns (e.g., peak / off-peak production), maintenance arrangements (planned or emergency maintenance), and abnormal events (e.g., sudden shutdowns, abnormal energy consumption), used to describe future trends in state changes and their causes. 3) Inference confidence This is used to indicate the credibility of the current inference result by the large language model.

[0012] In filtering modeling, semantic state estimation It is regarded as a semantic pseudo-observation, used to replace the observation value in the traditional filtering process, and its formal definition is shown in Equation (1):

[0013] (1)

[0014] in For semantic reasoning mapping of LLM, It is semantic operating condition information.

[0015] Step 1.2, Measurement of semantic observation covariance:

[0016] The filtering update process also relies on covariance estimation, which accurately reflects observation uncertainty. Therefore, after introducing semantic pseudo-observations into the Kalman filter framework, it is necessary to further model and measure the uncertainty (confidence) of semantic pseudo-observations. In the traditional Kalman filter framework, the observation covariance is usually set based on human experience or static priors. However, in semantic reasoning scenarios, fixed covariance is difficult to adapt to the dynamic changes in the confidence of semantic pseudo-observations with operating conditions. In actual industrial scenarios, the system uncertainty corresponding to different semantic contexts varies significantly. For example, in semantic contexts such as "under maintenance," "unit malfunction," or "significant future fluctuations," the system operating status usually has stronger uncertainty and weaker predictability. Therefore, the confidence of the corresponding semantic reasoning results should be lower than that of stable operating conditions such as "normal operation." Furthermore, even under the same semantic context, the reasoning stability of large language models may still be affected by factors such as semantic conflicts, missing context, and incomplete historical information, leading to deviations or inconsistencies in the reasoning results. Therefore, it is necessary to construct a mechanism that can dynamically reflect the reliability of semantic reasoning, so that when the reasoning ability of the large language model is insufficient or the reasoning results are biased, the confidence of the corresponding semantic pseudo-observations can be automatically reduced.

[0017] To address the uncertainty measurement problem of semantic pseudo-observations, this application, based on cue word engineering, guides large language models to simultaneously generate corresponding inference confidence scores while outputting inference results. Subsequently, this confidence level was correlated with the semantic change features of the working conditions (i.e., trend-based inference text). The observation covariance is jointly mapped to the observation covariance and used in the differentiable Kalman filter fusion process. The specific formula for calculating the observation covariance is as follows:

[0018] (2)

[0019] (3)

[0020] (4)

[0021] (5)

[0022] (6)

[0023] in Represents semantic feature enhancement vectors. The mean of the observed covariance, Represents the standard deviation of the observed covariance. and These represent the median value and standardized output value of the observed covariance, respectively. To observe the trainable amplitude parameter of the covariance, This is a semantic feature mapping network used to process trend-based inference text output by large language models. Mapped to semantic feature enhancement vectors that can participate in filter modeling ; This is a mean feature mapping network used to map semantic feature enhancement vectors into intermediate features that participate in the generation of the mean covariance of observations; It is a feature fusion network used to fuse semantic feature enhancement vectors, inference confidence and relevant intermediate features, and generate observation covariance parameters.

[0024] Step 2: Establish a differentiable state transition matrix for state estimation;

[0025] Step 2.1, Design the differentiable state transition matrix:

[0026] After constructing the semantic pseudo-observations in step 1, further state estimates need to be built. In the traditional Kalman filter framework, state estimation typically relies on a pre-defined and fixed state transition matrix to predict the state, and the filtering performance is mainly achieved by weighting and fusing the predicted and observed information using Kalman gain. However, in actual industrial systems, the system operation is often affected by factors such as operating mode switching, process parameter coupling, and complex nonlinear dynamic processes. Relying solely on Kalman gain to adjust the weights of the predicted and observed values ​​is insufficient to fully learn the intrinsic mechanism of system state evolution. If the state transition matrix remains fixed, the filtering process essentially degenerates into a simple weighted fusion of numerical prediction results and semantic inference results, failing to effectively characterize the dynamic changes in the state under different operating conditions. Its model expressive power, adaptability to operating conditions, and robustness in complex scenarios will all be significantly limited. Therefore, to enhance the state estimation model's ability to learn the dynamic characteristics of complex industrial processes, this application breaks through the constraints of the traditional fixed state transition matrix, enabling the state transition process to adaptively adjust with changes in operating conditions, thereby improving the filtering model's ability to characterize the non-stationary dynamic characteristics of industrial systems. This design allows the model to learn state evolution patterns under different operating conditions through data-driven methods, thereby improving its ability to dynamically represent complex production processes. On the other hand, it introduces semantic reasoning results from a large language model as high-level cognitive terms during the training phase, guiding the differentiable state transition matrix to identify and absorb some operating condition evolution features that are difficult to capture using only time-series data. As the training process iterates, the model may gradually internalize some of the process knowledge derived from semantic reasoning into the parameters of the differentiable state transition matrix. This allows for relatively reliable state predictions even when the large language model's reasoning is lacking or biased (especially under complex and special operating conditions such as maintenance and abnormal switching).

[0027] Based on the above analysis, this application constructs the state transition matrix as a differentiable state transition matrix driven by both numerical state and semantic context, and embeds it into a differentiable Kalman filter framework, enabling the parameters of the differentiable state transition matrix and the Kalman gain to be jointly optimized through a backpropagation mechanism. This takes into account trend-based inference text. The data contains a large amount of high-level semantic information related to the evolution of working conditions and production mechanisms. This application further utilizes a semantic feature mapping network to encode it into semantically enhanced features for the conditional differentiable state transition matrix and its uncertainty modeling process. The formula for calculating the differentiable state transition matrix is ​​shown in equation (7):

[0028] (7)

[0029] in For numerical status input, It is a differentiable state transition matrix. This is for estimating the state at time t in the future.

[0030] Step 2.2, semantic-driven covariance generation;

[0031] To measure the uncertainty of the state estimate, this application will use trend inference text. This is integrated into the measurement of uncertainty in state estimates, i.e., it participates in the calculation of process covariance. Elements of process covariance include trend-based inference text. Cumulative prediction bias of differentiable state transition matrix ,in Obtained from the historical state prediction error statistics of the differentiable state transition matrix during the training phase, it is used to characterize the inherent uncertainty of the differentiable state transition matrix. (Trend inference text) The feature processing procedure and the cumulative prediction bias of the differentiable state transition matrix The calculation method and the process of merging the two are shown in the following formula:

[0032] (8)

[0033] (9)

[0034] (10)

[0035] in, This is a semantic feature mapping network used to process trend-based inference text output by large language models. Mapped to semantically enhanced features that can participate in filter modeling; This is a feature fusion network used to fuse semantically enhanced features with the cumulative prediction bias of a differentiable state transition matrix; This represents a semantic feature enhancement vector.

[0036] The mean, variance, and sampling process covariance based on the Gaussian distribution are calculated as follows:

[0037] (11)

[0038] (12)

[0039] (13)

[0040] (14)

[0041] in, These are mean and standard deviation feature mapping networks, used to map fused features to mean and standard deviation; The mean of the covariance of the representative process. Representative process covariance and standard deviation and These represent the intermediate value of the process covariance sampling and the standardized output value, respectively. is the trainable amplitude parameter of the process covariance.

[0042] Step 3: Construct the Kalman filter structure;

[0043] Based on the preparation of the core elements of the Kalman filter in steps 1 and 2, the Kalman filter process is constructed. For example... Figure 1 As shown, the core structure of the Kalman filter is retained during the modeling process, including prior estimation, observation update mechanism, and Kalman gain. Explicit computation. It should be emphasized that in traditional Kalman filtering, the covariance is estimated... The state transition matrix is ​​jointly determined by the propagation term of historical uncertainty and the process noise term. However, since this application uses a neural network to construct a differentiable state transition matrix, its error propagation process is difficult to solve using a traditional linear state-space model. Therefore, this application uses the semantically driven covariance approximation in step 2.2 to estimate the covariance, and adaptively corrects the covariance parameter through end-to-end training, thereby achieving dynamic modeling of state uncertainty. The filtering process formula is as follows:

[0044] (15)

[0045] (16)

[0046] (17)

[0047] (18)

[0048] (19)

[0049] in represent Always looking towards the future The estimated covariance at time, This is the noise covariance generation process.

[0050] Furthermore, the state update cycle of the Kalman filter is consistent with the discrete time step of the state; that is, a complete estimation and update process is performed after each state transition, and the corresponding Kalman gain is calculated and updated synchronously. And with all LLM parameters frozen, global network optimization is performed using Equation (20) as the loss function.

[0051] (20)

[0052] in, This is the estimated value output after Kalman filtering. This represents the actual state value.

[0053] After the above steps, this application outputs the following results: 1) State-filtered estimate of the target industrial energy system 1) Used to represent the final estimation result of the target state; 2) State filtering estimation covariance , used to represent the uncertainty of the state estimation result; where the state of the target industrial energy system can be cabinet position, pressure, flow rate, energy consumption or other industrial process state variables.

[0054] Furthermore, this application adopts an "offline training, online deployment" working mode. During the training phase, a supervised learning dataset is constructed using historical industrial data to train each network module in the large language model and the differentiable Kalman filter framework. Specifically, the large language model is responsible for generating semantic state estimation results, trend-based inference text, and inference confidence; the differentiable Kalman filter framework is responsible for state estimation and covariance learning. During training, real historical state values ​​are used to optimize the model parameters. In the application phase, the trained model parameters remain fixed, and new industrial production information is received in real time as input. The target industrial energy system state filter estimate and state filter estimate covariance are generated according to the same output structure as in the training phase.

[0055] The advantages of this application are:

[0056] This application proposes a method for industrial energy state estimation using a combination of large and small model approaches. This method introduces the reasoning capabilities of large language models into the task of estimating the state of industrial energy systems. In this method, considering the reliability of the large language model's reasoning, the reasoning results are treated as high-level semantic observations. Their uncertainty is modeled using covariance, which reflects the confidence level of the reasoning, and then used as a state semantic pseudo-observation estimate in the state estimation process. Furthermore, through a differentiable Kalman filter structure, the Kalman gain is adaptively adjusted to consider the difference in uncertainty between the numerical state estimation and the state semantic pseudo-observation estimation, achieving a reasonable fusion of the two types of information within a unified dynamic system framework. Attached Figure Description

[0057] Figure 1 This is a structural diagram of an industrial energy state estimation method that combines large and small models.

[0058] Figure 2 This refers to the process by which each estimated value approaches the true value as the number of training iterations increases.

[0059] Figure 3 This shows how the Kalman filter gain changes with the number of training iterations during the training iteration process.

[0060] Figure 4 This paper compares the state estimation curves of the method in this application with those of other mainstream baseline models under conditions of smooth energy fluctuations.

[0061] Figure 5 This paper compares the state estimation curves of the method in this application with those of other mainstream baseline models under conditions of severe energy fluctuations. Detailed Implementation

[0062] A method for estimating industrial energy state using a combination of large and small models includes the following steps:

[0063] Step 1: Construct the semantic confidence-driven cognitive fusion filtering component;

[0064] Step 1.1, Construction of semantic pseudo-observations:

[0065] First, semantic working condition information such as production plans is input into the large language model. Then, the inference output of the large language model generates three types of complementary information according to the preset output template: 1) the final state estimation result, i.e., semantic state estimation. 1) Used to represent the prediction results of the large language model for the target state; 2) Supplementary explanations of the trend of the result, i.e., trend inference text. Its content covers key production information such as production patterns (e.g., peak / off-peak production), maintenance arrangements (planned or emergency maintenance), and abnormal events (e.g., sudden shutdowns, abnormal energy consumption), used to describe future trends in state changes and their causes. 3) Inference confidence This is used to indicate the credibility of the current inference result by the large language model.

[0066] In filtering modeling, semantic state estimation It is regarded as a semantic pseudo-observation, used to replace the observation value in the traditional filtering process, and its formal definition is shown in Equation (1):

[0067] (1)

[0068] in For semantic reasoning mapping of LLM, It is semantic operating condition information.

[0069] In the specific implementation of the large language model inference part, based on the multivariate time-series data of converter gas production-consumption-storage, and integrating unstructured industrial semantic information such as production scheduling plans, abnormal maintenance information, and dispatch instructions, a converter gas holder status estimation task is constructed. The large language model is divided into two categories according to the task: domain-based inference large language model and task-oriented professional inference large language model. The domain-based inference large language model is used to synthesize the outputs of various task-oriented professional inference large language models and complete semantic state inference to obtain semantic pseudo-observations. It adopts the Deepseek-r1:70B large language model, setting the temperature coefficient to 1.0, the kernel sampling parameter top-p to 0.9, and the number of context rounds to 5 during the inference phase, and using a streaming output method, with a maximum of 8192 generated tokens. This module is deployed based on the vLLM framework to support efficient large model inference. The task-oriented professional inference large language model is used to identify multiple types of production data in the industrial field. It adopts the Llama-8B model and performs efficient parameter fine-tuning based on LoRA. The LoRA fine-tuning parameters were set to rank=8 and scaling factor=16, and applied to the query / value projection of the attention layer. During the inference phase, the temperature coefficient was set to 0.2, top-p to 0.9, and the maximum number of generated tokens was 128. Three task-oriented semantic modules were constructed in the experiment, used for production rhythm recognition, production plan recognition, and maintenance scheduling text parsing, respectively.

[0070] Step 1.2, Measurement of semantic observation covariance:

[0071] To address the uncertainty measurement problem of semantic pseudo-observations, this application, based on cue word engineering, guides large language models to simultaneously generate corresponding inference confidence scores while outputting inference results. Subsequently, this confidence level was correlated with the semantic change features of the working conditions (i.e., trend-based inference text). The observation covariance is jointly mapped to the observation covariance and used in the differentiable Kalman filter fusion process. The specific formula for calculating the observation covariance is as follows:

[0072] (2)

[0073] (3)

[0074] (4)

[0075] (5)

[0076] (6)

[0077] in Represents semantic feature enhancement vectors. The mean of the observed covariance, Represents the standard deviation of the observed covariance. and These represent the median value and standardized output value of the observed covariance, respectively. To observe the trainable amplitude parameter of the covariance, This is a semantic feature mapping network used to process trend-based inference text output by large language models. Mapped to semantic feature enhancement vectors that can participate in filter modeling ; This is a mean feature mapping network used to map semantic feature enhancement vectors into intermediate features that participate in the generation of the mean covariance of observations; It is a feature fusion network used to fuse semantic feature enhancement vectors, inference confidence and relevant intermediate features, and generate observation covariance parameters.

[0078] Step 2: Establish a differentiable state transition matrix for state estimation;

[0079] Step 2.1, Design the differentiable state transition matrix:

[0080] The state transition matrix is ​​constructed as a differentiable state transition matrix driven by both numerical state and semantic context, and embedded in a differentiable Kalman filter framework. This allows the parameters of the differentiable state transition matrix and the Kalman gain to be jointly optimized through backpropagation. This takes into account trend-based inference text. The data contains a large amount of high-level semantic information related to the evolution of working conditions and production mechanisms. This application further utilizes a semantic feature mapping network to encode it into semantically enhanced features for the conditional differentiable state transition matrix and its uncertainty modeling process. The formula for calculating the differentiable state transition matrix is ​​shown in equation (7):

[0081] (7)

[0082] in For numerical status input, It is a differentiable state transition matrix. This is for estimating the state at time t in the future.

[0083] Step 2.2, semantically driven process covariance generation;

[0084] To measure the uncertainty of the state estimate, this application will use trend inference text. This is integrated into the measurement of uncertainty in state estimates, i.e., it participates in the calculation of process covariance. Elements of process covariance include trend-based inference text. Cumulative prediction bias of differentiable state transition matrix ,in Obtained from the historical state prediction error statistics of the differentiable state transition matrix during the training phase, it is used to characterize the inherent uncertainty of the differentiable state transition matrix. (Trend inference text) The feature processing procedure and the cumulative prediction bias of the differentiable state transition matrix The calculation method and the process of merging the two are shown in the following formula:

[0085] (8)

[0086] (9)

[0087] (10)

[0088] in, This is a semantic feature mapping network used to process trend-based inference text output by large language models. Mapped to semantically enhanced features that can participate in filter modeling; This is a feature fusion network used to fuse semantically enhanced features with the cumulative prediction bias of a differentiable state transition matrix; This represents a semantic feature enhancement vector.

[0089] The mean, variance, and sampling process covariance based on the Gaussian distribution are calculated as follows:

[0090] (11)

[0091] (12)

[0092] (13)

[0093] (14)

[0094] in, These are mean and standard deviation feature mapping networks, used to map fused features to mean and standard deviation; The mean of the covariance of the representative process. Representative process covariance and standard deviation and These represent the intermediate value of the process covariance sampling and the standardized output value, respectively. is the trainable amplitude parameter of the process covariance.

[0095] During implementation, the overall network was built using the PyTorch framework. In terms of modeling, the differentiable state transition matrix, observation covariance generation network, and process covariance generation network in the Kalman filter module all adopted a multilayer perceptron (MLP) structure.

[0096] Step 3, Kalman filter structure construction:

[0097] Based on the preparation of the core elements of the Kalman filter in steps 1 and 2, the Kalman filter process is constructed. For example... Figure 1 As shown, the core structure of the Kalman filter is retained during the modeling process, including prior estimation, observation update mechanism, and Kalman gain. The explicit computation is performed. The covariance is estimated using the semantically driven covariance approximation from step 2.2, and the covariance parameters are adaptively corrected through end-to-end training to achieve dynamic modeling of state uncertainty. The filtering process formula is as follows:

[0098] (15)

[0099] (16)

[0100] (17)

[0101] (18)

[0102] (19)

[0103] in represent Always looking towards the future The estimated covariance at time, This is the noise covariance generation process.

[0104] Furthermore, the state update cycle of the Kalman filter is consistent with the discrete time step of the state; that is, a complete estimation and update process is performed after each state transition, and the corresponding Kalman gain is calculated and updated synchronously. And with all LLM parameters frozen, global network optimization is performed using Equation (20) as the loss function.

[0105] (20)

[0106] in, This is the estimated value output after Kalman filtering. This represents the actual state value.

[0107] After the above steps, this application outputs the following results: 1) State-filtered estimate of the target industrial energy system 1) Used to represent the final estimation result of the target state; 2) State filtering estimation covariance , used to represent the uncertainty of the state estimation result; where the state of the target industrial energy system can be cabinet position, pressure, flow rate, energy consumption or other industrial process state variables.

[0108] The system adopts an "offline training, online deployment" working mode. During the training phase, a supervised learning dataset is constructed using historical industrial data to train the various network modules within the large language model and the differentiable Kalman filter framework. The large language model is responsible for generating semantic state estimation results, trend-based inference text, and inference confidence; the differentiable Kalman filter framework is responsible for state estimation and covariance learning. Real historical state values ​​are used to optimize the model parameters during training. In the application phase, the trained model parameters remain fixed. New industrial production information is received in real time as input, and the system generates the target industrial energy system state filter estimate and state filter estimate covariance according to the same output structure as in the training phase, achieving online estimation and prediction of the industrial energy system state.

[0109] In the specific implementation process, considering that the state estimation results ultimately serve production scheduling decisions, the state update interval is set to 20 minutes to match the time granularity of the actual scheduling strategy.

[0110] To verify the effectiveness of the proposed method, this application conducted experiments based on real production data from the converter gas system of a domestic steel enterprise. The experimental data covers the period from December 2024 to July 2025, with a sampling interval of 1 minute. The dataset is divided chronologically, with the first 150 days of data used for model training, the middle 30 days of data used for model selection and hyperparameter tuning, and the last 30 days of data used for testing and validation.

[0111] To demonstrate the guiding role of LLM semantic inference estimation on the differentiable state transition matrix during training iterations within a large-scale model collaboration framework, and the ability of Kalman gain to fuse the two, Figure 2 , Figure 3 It demonstrates the process by which the estimated value approximates the true state value, and the corresponding change in the Kalman gain.

[0112] like Figure 2 As shown, in the initial stage of filtering, due to the large process covariance of the differentiable state transition matrix, the filter will highly depend on the semantic pseudo-observation estimate. As the number of iterations increases, the process covariance continuously decreases, and the filter will gradually depend on the differentiable state transition matrix. Figure 3 As shown, Kalman's confidence in the semantic pseudo-observation estimates gradually adjusts from a relatively large value in the early stage to a low value range, reflecting the beneficial effect of the adaptive adjustment of the filter gain on the state estimation process during the training iteration.

[0113] To verify the effectiveness of the proposed method in energy state estimation, this paper selects several representative comparative methods (including the numerical learning-based prediction methods DLinear and PatchTST, the kernel learning-based regression method LSSV, the iterative computation-based state estimation method KalmanNet, the large language model semantic inference method based on LoRA fine-tuning, and the large language model-based semantic alignment prediction method TimeLLM) to compare with the proposed method in terms of state curve fitting and performance metrics. Performance metrics are quantified using standard error indicators, including root mean square error (RMSE), mean absolute error (MAE), mean squared error (MSE), and mean absolute percentage error (MAPE).

[0114] Figure 4 Comparative experimental results are presented under a scenario of relatively stable energy fluctuations. It can be seen that all comparative methods can effectively learn the overall trend of energy change. Among the comparative methods, PatchTST, relying on its strong time-series modeling capabilities, performs better in terms of estimation accuracy. Meanwhile... Figure 5 Under the conditions of drastic energy fluctuations (typically caused by overlapping production schedules of multiple devices or abnormal equipment shutdowns), all comparative methods except DLinear showed significant deviations. Our proposed method achieved better prediction results under both stable and drastic conditions, with a particularly pronounced advantage over other comparative methods in cases of severe fluctuations. Figure 4 , Figure 5The results show that our proposed method excels in both response speed and fitting accuracy at points of significant trend inflection. This may be because LLM, through planning or maintenance information, anticipates such rapid trend changes and guides the numerical model accordingly. Table 1 summarizes the state estimation accuracy metrics under different scenarios, demonstrating that our proposed method improves upon all compared methods.

[0115] Table 1 Comparison of experimental indicators

[0116]

Claims

1. A method for estimating industrial energy state using a combination of large and small models, characterized in that, By introducing a differentiable state transition matrix to compensate for the dynamic characteristics of the output of a large language model, and by utilizing the statistical patterns of data learned from the differentiable state transition matrix, the semantic output of the large language model is corrected and supplemented, thereby improving the accuracy of state estimation. Specifically, the following steps are included: Step 1: Construct the semantic confidence-driven cognitive fusion filtering component; Step 1.1, Construction of semantic pseudo-observations: The traditional Kalman filter concept is generalized to cognitive fusion filtering. First, semantic work status information such as production plans is input into a large language model. Then, the inference output of the large language model generates three types of complementary information based on the cue word engineering according to a preset output template: 1) The final state estimation result, i.e., the semantic state estimation , used to represent the prediction result of the large language model for the target state; 2) Supplementary explanations of the trends in the prediction results, i.e., trend inference text. Its content covers key production information; 3) Inference confidence This is used to indicate the credibility of the current inference result by the large language model; semantic state estimation It is regarded as a semantic pseudo-observation, used to replace the observation value in the traditional Kalman filtering process, and its formal definition is shown in Equation (1): (1) in For semantic reasoning mapping of LLM, It is semantic working condition information; Step 1.2, Measurement of semantic observation covariance: After introducing semantic pseudo-observations into the Kalman filter framework, the uncertainty of semantic pseudo-observations is further modeled and measured, and a mechanism that can dynamically reflect the reliability of semantic reasoning is constructed. When the reasoning ability of the large language model is insufficient or the reasoning results are biased, the confidence of the corresponding semantic pseudo-observation can be automatically reduced. Step 2: Establish a differentiable state transition matrix for state estimation; Step 2.1, Design the differentiable state transition matrix: The state transition matrix is ​​constructed as a differentiable state transition matrix driven by both numerical state and semantic context, and embedded in a differentiable Kalman filter framework, so that the parameters of the differentiable state transition matrix and the Kalman gain can be jointly optimized through the backpropagation mechanism. Step 2.2, Semantic-driven covariance generation: Trend-based reasoning text It is integrated into the process of measuring the uncertainty of the state estimate, that is, it participates in the calculation of the process covariance; Step 3: Construct the Kalman filter structure; retain the core structure of the Kalman filter during the modeling process, including prior estimation, observation update mechanism, and Kalman gain. Explicit computation; Output the following results: 1) State-filtered estimate of the target industrial energy system , used to represent the final estimate of the target state; 2) State filtering estimates covariance , is used to represent the uncertainty of the state estimation result.

2. The industrial energy state estimation method based on large and small model coordination according to claim 1, characterized in that, In step 1.2, based on cue word engineering, the large language model is guided to simultaneously generate the corresponding inference confidence score while outputting the inference result. ; Then, the confidence level of the inference was... With trend-based reasoning text The common mapping is the observation covariance, which is used in the differentiable Kalman filter fusion process; The calculation formula is as follows: (2) (3) (4) (5) (6) in Represents semantic feature enhancement vectors. The mean of the observed covariance, Represents the standard deviation of the observed covariance. and These represent the median value and standardized output value of the observed covariance, respectively. To observe the trainable amplitude parameter of the covariance, This is a semantic feature mapping network used to process trend-based inference text output by large language models. Mapped to semantic feature enhancement vectors that can participate in filter modeling ; This is a mean feature mapping network used to map semantic feature enhancement vectors into intermediate features that participate in the generation of the mean covariance of observations; It is a feature fusion network used to fuse semantic feature enhancement vectors, inference confidence and relevant intermediate features, and generate observation covariance parameters.

3. The industrial energy state estimation method based on large and small model coordination according to claim 2, characterized in that, In step 2.1, the semantic feature mapping network is used to encode it into semantically enhanced features, which are used for the conditional differentiable state transition matrix and its uncertainty modeling process. The formula for calculating the differentiable state transition matrix is ​​shown in equation (7): (7) in For numerical status input, It is a differentiable state transition matrix. This is for estimating the state at time t in the future.

4. The industrial energy state estimation method based on large and small model coordination according to claim 3, characterized in that, In step 2.2, the elements of process covariance include trend-based inference text. Cumulative prediction bias of differentiable state transition matrix ,in It is obtained from the historical state prediction error statistics of the differentiable state transition matrix during the training phase and is used to characterize the inherent uncertainty of the differentiable state transition matrix. Trend-based inference text The feature processing procedure and the cumulative prediction bias of the differentiable state transition matrix The calculation method and the process of merging the two are shown in the following formula: (8) (9) (10) in, This is a semantic feature mapping network used to process trend-based inference text output by large language models. Mapped to semantically enhanced features that can participate in filter modeling; This is a feature fusion network used to fuse semantically enhanced features with the cumulative prediction bias of a differentiable state transition matrix; Represents a semantic feature enhancement vector; The mean, variance, and sampling process covariance based on the Gaussian distribution are calculated as follows: (11) (12) (13) (14) in, These are mean and standard deviation feature mapping networks, used to map fused features to mean and standard deviation; The mean of the covariance of the representative process. Representative process covariance and standard deviation and These represent the intermediate value of the process covariance sampling and the standardized output value, respectively. is the trainable amplitude parameter of the process covariance.

5. The industrial energy state estimation method based on large and small model coordination according to claim 4, characterized in that, In step 3, the covariance is estimated using the semantically driven covariance approximation from step 2.2, and the covariance parameters are adaptively corrected through end-to-end training, thereby achieving dynamic modeling of state uncertainty. The formula for the filtering process is as follows: (15) (16) (17) (18) (19) in represent Always looking towards the future The estimated covariance at time, This refers to the noise covariance generation process; In addition, the state update cycle of the Kalman filter is consistent with the state discrete time step, that is, after each state transition, a complete estimation and update process is performed, and the corresponding Kalman gain is calculated and updated synchronously; and with all LLM parameters frozen, Equation (20) is used as the loss function for global network optimization. (20) in, This is the estimated value output after Kalman filtering. This represents the actual state value.

6. The industrial energy state estimation method based on large and small model coordination according to claim 5, characterized in that, A supervised learning dataset was constructed using historical industrial data to train various network modules in the large language model and differentiable Kalman filter framework. During the training process, the model parameters were optimized using real historical state values. After training, the model parameters remain fixed. New industrial production information is received in real time as input, and the target industrial energy system state filter estimate and state filter estimate covariance are generated according to the same output structure as in the training phase.

7. The industrial energy state estimation method based on large and small model coordination according to claim 6, characterized in that, The large language model is responsible for generating semantic state estimation results, trend inference text, and inference confidence; the differentiable Kalman filter framework is responsible for completing state estimation and covariance learning.