A NO based on a time-series large model x Prediction System

By using a NOx prediction system based on a time-series large model, the problem of inaccurate NOx emission under load fluctuations and fuel characteristic changes in traditional control methods has been solved, achieving high-precision NOx prediction and improving the system's prediction capability and control effect under complex operating conditions.

CN121393633BActive Publication Date: 2026-04-28BEIJING XINYE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING XINYE TECH CO LTD
Filing Date
2025-12-23
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

When faced with load fluctuations and changes in fuel characteristics, the existing denitrification systems of coal-fired power plants cannot adjust the ammonia injection volume in a timely manner using traditional control methods, resulting in excessive fluctuations in NOx emission concentrations, or even exceeding the standard. Furthermore, traditional prediction methods lack sufficient accuracy and real-time performance under complex operating conditions.

Method used

A NOx prediction system based on a time-series large model is adopted to achieve high-precision prediction of NOx emissions from thermal power units through time-domain and frequency-domain feature extraction and LoRA-large language model deep semantic understanding. The system includes data preprocessing, feature mapping, feature extraction and deep prediction modules, and uses LoRA-large language model for deep semantic understanding and context-related reasoning.

Benefits of technology

It significantly improves the accuracy and reliability of NOx emission prediction, can capture short-term transient fluctuations and long-term global coupling characteristics in the combustion process, provides reliable control basis, and supports emission optimization of thermal power units under deep peak shaving.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121393633B_ABST
    Figure CN121393633B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of thermal power plant pollutant prediction, and particularly relates to a NO x Prediction system, comprising: a data preprocessing and feature mapping module is used for collecting historical parameters in a predetermined first time period and planned parameters in a predetermined second time period, obtaining time sequence fusion data; a feature extraction module is used for extracting time domain features and frequency features of the time sequence fusion data, and generating joint features based on the time domain features and the frequency features; a deep prediction module is used for deep semantic understanding of the joint features based on a preset LoRA-large language model and outputting a NO x Concentration prediction sequence. Through time domain and frequency domain feature extraction and LoRA-large language model deep semantic understanding, the present application realizes high-precision prediction of NO x Emission of thermal power units under changes in fuel characteristics and load fluctuations, solving the problem that traditional models are difficult to simultaneously process instantaneous fluctuations and long-range coupling effects of boiler working conditions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of pollutant prediction technology for thermal power plants, specifically to a NO prediction method based on a large time-series model. x Prediction system. Background Technology

[0002] Currently, the proportion of new energy sources such as wind power and photovoltaic power in the power grid is continuously increasing, and thermal power units are gradually transforming from the main power source to a regulating power source. During this transformation, thermal power units face the dual challenges of deep peak shaving and rapid load changes, while simultaneously needing to meet increasingly stringent ultra-low emission requirements, particularly regarding nitrogen oxides (NOx). x The accurate prediction and effective control of [the virus] have placed higher technical demands on the technology.

[0003] Existing denitrification systems in coal-fired power plants generally employ fixed control strategies or simple PID control, which are insufficient to meet the demands of flexible unit operation. Due to the inherent delay in flue gas sampling and analysis, traditional control methods cannot promptly adjust ammonia injection rates when unit load fluctuates drastically, easily leading to NO buildup at the denitrification outlet. x Excessive concentration fluctuations can even lead to problems such as exceeding emission standards or excessive ammonia injection.

[0004] Currently, regarding NO x Emissions prediction methods mainly rely on traditional machine learning algorithms or static neural network models. These methods have shortcomings in prediction accuracy and real-time performance when faced with complex and variable combustion conditions. In particular, the prediction errors are significant when there are sudden load changes or changes in fuel characteristics, which cannot provide a reliable basis for advanced control strategies and limit the emission optimization capabilities of thermal power units under deep peak shaving. Summary of the Invention

[0005] (a) Purpose of the invention

[0006] The purpose of this invention is to provide a NO based on a large time-series model. x The prediction system, through time-domain and frequency-domain feature extraction and LoRA-large language model deep semantic understanding, enables the prediction of NO2 in thermal power units under changes in fuel characteristics and load fluctuations. x The high-precision emission prediction solves the problem that traditional models cannot simultaneously handle the instantaneous fluctuations and long-range coupling effects of boiler operating conditions.

[0007] (II) Technical Solution

[0008] To address the aforementioned problems, this invention provides a NO model based on a large time-series model. x Prediction systems include:

[0009] The module includes data preprocessing and feature mapping, feature extraction, and depth prediction.

[0010] The data preprocessing and feature mapping module is used to collect historical parameters within a predetermined first time period and planned parameters within a predetermined second time period, and to map the historical parameters and planned parameters to a time feature space to obtain time-series fusion data.

[0011] The feature extraction module is used to extract the temporal and frequency features of the time-series fusion data, and generate joint features based on the temporal and frequency features;

[0012] The deep prediction module is used to perform deep semantic understanding on the joint features based on a preset LoRA-Large Language Model and output the NO within a predetermined second time period. x Concentration prediction sequence.

[0013] In another aspect of the present invention, preferably, the data preprocessing and feature mapping module includes an acquisition unit and a preprocessing unit;

[0014] The acquisition unit is used to acquire historical parameters within a predetermined first time period and planned parameters within a predetermined second time period at a minute-level. The historical parameters include: primary air volume, secondary air volume, air-coal ratio, coal feed rate, flue gas oxygen content, unit real-time load, fuel volatile matter content, and fuel calorific value.

[0015] The planned parameters include: load command, ambient temperature, wind speed, solar radiation intensity, and weather condition code;

[0016] The relationship between the predetermined first time period and the predetermined second time period includes: the predetermined second time period immediately follows the predetermined first time period;

[0017] The preprocessing unit is used to perform time-series alignment, missing parameter correction, and normalization on the historical parameters and planned parameters obtained by the acquisition unit.

[0018] In another aspect of the present invention, preferably, the data preprocessing and feature mapping module further includes a linear projection unit and a temporal fusion unit;

[0019] The linear projection unit is used to map the preprocessed historical parameters and the planned parameters respectively to obtain the high-dimensional feature vector sequence of historical parameters and the high-dimensional feature vector sequence of planned parameters;

[0020] The time-series fusion unit is used to fuse the high-dimensional feature vector sequence of historical parameters and the high-dimensional feature vector sequence of planned parameters along the time dimension to obtain time-series fused data.

[0021] In another aspect of the present invention, preferably, the linear projection unit includes a first linear projection layer and a second linear projection layer;

[0022] The first linear projection layer is used to map the preprocessed planning parameters to obtain a high-dimensional feature vector sequence of historical parameters, wherein the dimension of the high-dimensional feature vector sequence of historical parameters is the first dimension.

[0023] The second linear projection layer is used to map the preprocessed historical parameters to obtain a high-dimensional feature vector sequence of planning parameters, wherein the dimension of the high-dimensional feature vector sequence of planning parameters is the second dimension;

[0024] In another aspect of the present invention, preferably, the feature extraction module includes a temporal extraction unit; the temporal extraction unit includes a first LSTM layer and a first Transformer encoder layer;

[0025] The time-domain extraction unit is used to extract time-domain features from the time-series fused data, including:

[0026] The short-term temporal dependencies of the time-series fusion data are extracted using the first LSTM layer to generate primary temporal features, wherein the dimension of the primary temporal features is the third dimension.

[0027] The global correlation of the primary temporal features is extracted using the first Transformer encoder layer based on the self-attention mechanism to generate temporal enhanced features, which have a fourth dimension.

[0028] In another aspect of the present invention, preferably, the feature extraction module includes a frequency domain extraction unit, which includes a Fourier transform layer, a second LSTM layer, and a second Transformer encoder layer.

[0029] The frequency domain extraction unit is used to extract frequency domain features from the time-series fused data, including:

[0030] The time-series fusion data is converted to frequency domain data using the Fourier transform layer, and the frequency domain data has a fifth dimension.

[0031] The short-term time dependencies of the frequency domain data are extracted using the second LSTM layer to generate primary frequency domain features, wherein the primary frequency domain features have a sixth dimension.

[0032] The global correlation of the primary frequency domain features is extracted using the second Transformer encoder layer to generate enhanced frequency domain features, which have a seventh dimension.

[0033] In another aspect of the present invention, preferably, the feature extraction module further includes a nonlinear feature fusion unit;

[0034] The nonlinear feature fusion unit interacts and weights the time-domain enhancement features and frequency-domain enhancement features to obtain joint features, which have a seventh dimension.

[0035] In another aspect of the present invention, preferably, the deep prediction module includes a LoRA-Large Language Model unit;

[0036] The LoRA-Big Language Model Unit performs deep semantic understanding and contextual reasoning on the joint features based on the preset LoRA-Big Language Model to obtain semantic features, the dimension of which is the eighth dimension.

[0037] The preset LoRA-Large Language Model's base weight parameters are frozen during training and do not participate in gradient updates; only its low-rank adapter parameters are trained.

[0038] In another aspect of the present invention, preferably, the depth prediction module includes a prediction unit, the prediction unit being based on a three-layer fully connected network;

[0039] The prediction unit performs regression prediction processing on the semantic features, including:

[0040] The semantic features within a predetermined second time period are determined. Through a multi-layer nonlinear mapping structure of a three-layer fully connected network, the semantic features within the predetermined second time period are subjected to layer-by-layer dimensionality compression and feature mapping, continuously reducing the feature space from the eighth dimension to the ninth dimension, thus obtaining the NO within the predetermined second time period. x Concentration prediction sequence.

[0041] In another aspect, preferably, the invention further includes a training module for training the prediction system;

[0042] The training module is based on a phased progressive training strategy, which controls the freezing and unfreezing process of the linear projection unit, time domain extraction unit, frequency domain extraction unit, nonlinear feature fusion unit and LoRA-large language model unit.

[0043] The term "freeze" means that the corresponding network layer parameters are kept out of backpropagation and gradient update during the current training phase, and the term "unfreeze" means that the parameter freeze state is lifted, allowing the corresponding network layer parameters to participate in backpropagation and gradient update.

[0044] (III) Beneficial Effects

[0045] The above-described technical solution of the present invention has the following beneficial technical effects:

[0046] This invention extracts time-domain and frequency-domain features from historical and planned parameters, and further utilizes the LoRA-Large Language Model for deep semantic understanding to achieve NO of thermal power units. x High-precision emission prediction significantly improves the accuracy and reliability of forecasts. By fusing historical and planned parameters over time, the system can capture short-term transient fluctuations and long-term global coupling characteristics during combustion, enabling the prediction results to accurately reflect emission dynamics under rapid load changes and fuel characteristic fluctuations. Joint features enhance the richness and generalization ability of feature representation. Deep semantic modeling based on the LoRA-large language model gives the system powerful contextual understanding capabilities, allowing for advanced predictions under complex operating conditions. Attached Figure Description

[0047] Figure 1 This is a schematic diagram of the overall structure of one embodiment of the present invention;

[0048] Figure 2 This is a schematic diagram of the overall structure of another embodiment of the present invention. Detailed Implementation

[0049] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments and the accompanying drawings. It should be understood that these descriptions are merely exemplary and not intended to limit the scope of the invention. Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concept of the invention.

[0050] Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0051] In the description of this invention, it should be noted that the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0052] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0053] The invention will now be described in more detail with reference to the accompanying drawings. In the various drawings, the same elements are indicated by similar reference numerals. For clarity, the various parts in the drawings are not drawn to scale.

[0054] Example 1

[0055] A NO based on a time-series large model x Prediction system, Figure 1An overall flowchart of one embodiment of the present invention is shown, as follows: Figure 1 As shown, it includes:

[0056] The module includes data preprocessing and feature mapping, feature extraction, and depth prediction.

[0057] The data preprocessing and feature mapping module is used to collect historical parameters within a predetermined first time period and planned parameters within a predetermined second time period, and map the historical parameters and planned parameters to a time feature space to obtain time-series fused data. The specific content of the historical parameters within the predetermined first time period and the planned parameters within the predetermined second time period is not limited here. Optionally, in this embodiment, the data preprocessing and feature mapping module includes a collection unit, which is used to collect historical parameters within the predetermined first time period and planned parameters within the predetermined second time period at a minute-level granularity. The historical parameters include: primary air volume, secondary air volume, air-coal ratio, coal feed rate, flue gas oxygen content, real-time unit load, fuel volatile matter content, and fuel calorific value. The historical parameters can comprehensively reflect the unit's combustion state, load operating characteristics, and fuel energy characteristics. For NO... x The analysis of generation and emission trends is directly related. The planned parameters include: load command, ambient temperature, wind speed, solar radiation intensity, and weather condition codes, which can reflect the future load scheduling requirements of the unit and external environmental conditions, providing future operating background information for the prediction model; the relationship between the predetermined first time period and the predetermined second time period is that the predetermined second time period immediately follows the predetermined first time period, forming a seamless time series interval.

[0058] Table 1 shows the historical parameter table, and Table 2 shows the planned parameter table. Figure 2 A schematic diagram of the overall structure of another embodiment of the present invention is shown, as follows. Figure 2 As shown, the first time period is predetermined to be 60 minutes, and the second time period is predetermined to be 10 minutes, which serve as the prediction window.

[0059] Table 1 Historical Parameter Table

[0060]

[0061] Table 2 Plan Parameter Table

[0062]

[0063] Furthermore, the data preprocessing and feature mapping module also includes a preprocessing unit, which performs time-series alignment, missing value correction, and normalization on the historical and planned parameters obtained by the acquisition unit. The preprocessing unit first performs time-series alignment on the acquired historical and planned parameters. Since historical and planned parameters may have differences in acquisition frequency, sampling delay, or inconsistencies in timestamps, time-series alignment uses interpolation, resampling, or timestamp correction methods to map various parameters onto a standard time scale, forming a strictly corresponding time series. This ensures that the parameters are comparable and consistent at the same time step during feature extraction. The preprocessing unit corrects for any missing or outlier values ​​in the data. Missing values ​​may be caused by sensor malfunctions, communication delays, or data recording errors. The preprocessing unit uses nearest-neighbor interpolation, mean imputation, or prediction methods based on historical trends to reasonably estimate the missing data. For outliers, the preprocessing unit identifies and corrects them by setting reasonable physical thresholds, statistical anomaly detection, or sliding window analysis to avoid extreme values ​​misleadingly affecting feature extraction and model training. The preprocessing unit normalizes all parameters, mapping parameters with different dimensions and ranges to a unified numerical range, thus eliminating the influence of parameter scale differences.

[0064] Furthermore, in this embodiment, the data preprocessing and feature mapping module further includes a linear projection unit and a temporal fusion unit;

[0065] The linear projection unit is used to map the preprocessed historical parameters and the planned parameters respectively to obtain the high-dimensional feature vector sequence of historical parameters and the high-dimensional feature vector sequence of planned parameters;

[0066] The linear projection unit includes a first linear projection layer and a second linear projection layer;

[0067] The preprocessed planning parameters are mapped using the first linear projection layer to obtain a high-dimensional feature vector sequence of historical parameters, wherein the dimension of the high-dimensional feature vector sequence of historical parameters is the first dimension.

[0068] The preprocessed historical parameters are mapped using the second linear projection layer to obtain a high-dimensional feature vector sequence of planning parameters, where the dimension of the high-dimensional feature vector sequence is the second dimension. The first linear projection layer maps the preprocessed historical parameters, converting the original historical parameter sequence into a corresponding high-dimensional feature vector sequence, where the dimension of the high-dimensional feature vector sequence is defined as the first dimension. The second linear projection layer maps the preprocessed planning parameters, converting them into a high-dimensional feature vector sequence of planning parameters, where the dimension of the sequence is defined as the second dimension. To ensure consistency in data fusion, the first and second dimensions are set to the same dimension at a single time step, thus facilitating subsequent feature fusion operations and joint modeling. Figure 2 As shown, in this embodiment, the historical parameters are 60×8 dimensional parameters, and the planning parameters are 10×5 dimensional parameters. After mapping through the first and second linear projection layers, the dimensions are expanded to 60×128 and 10×128 dimensions respectively. That is, the feature dimensions of the first and second dimensions at a single time step are 128 dimensions. The expansion process is expressed by the following formula:

[0069]

[0070]

[0071] in, For historical parameters, For planning parameters, A sequence of high-dimensional feature vectors representing historical parameters. Let be the high-dimensional feature vector sequence of the planning parameters, and Linear be the linear projection unit.

[0072] The time-series fusion unit is used to fuse the high-dimensional feature vector sequences of historical parameters and the high-dimensional feature vector sequences of planned parameters along the time dimension to obtain time-series fused data. After projection, the high-dimensional feature vector sequences of historical parameters and the high-dimensional feature vector sequences of planned parameters are concatenated and fused along the time dimension to form 70×128-dimensional time-series fused data. The fusion process is represented by the following formula:

[0073] ,in ,

[0074] In the formula: For time-series fusion data, A sequence of high-dimensional feature vectors representing historical parameters. The sequence of high-dimensional feature vectors for the planning parameters.

[0075] The first 60 time steps of the time-series fusion data contain deep feature representations of historical parameters, while the last 10 time steps embed prior knowledge of planned parameters, providing a consistent input foundation for subsequent time-series-frequency domain dual-path analysis. This ensures the comparability of parameters with different physical dimensions and value ranges in the feature space, while preserving the inherent physical relationships between parameters.

[0076] The feature extraction module is used to extract the temporal and frequency features of the time-series fusion data, and generate joint features based on the temporal and frequency features to fully reflect the multi-timescale dynamic characteristics of the combustion process. In this embodiment, the feature extraction module includes a temporal extraction unit; the temporal extraction unit extracts multi-level, multi-timescale dynamic features to fully capture the changing patterns and coupling relationships of the combustion process at different timescales. The temporal extraction unit includes a first LSTM layer and a first Transformer encoder layer, which work together to take into account both local dynamic changes and global time-dependent features.

[0077] The time-domain extraction unit is used to extract time-domain features from the time-series fused data, including:

[0078] The first LSTM layer is used to extract short-term temporal dependencies from the time-series fusion data, generating primary temporal features with a third dimension. The input time-series fusion data is fed into the first LSTM layer (Long Short-Term Memory network layer). This layer uses a gating mechanism (input gate, forget gate, and output gate) to efficiently model short-term dependencies in the time series, capturing rapid dynamic features such as combustion load fluctuations and fuel quality changes. After processing by this layer, the output forms the primary temporal features, defined as a third dimension of 70×256.

[0079] A first Transformer encoder layer based on a self-attention mechanism is used to extract the global correlations of the primary temporal features, generating temporal augmented features with a fourth dimension. By calculating the attention weights between any two time steps, the first Transformer encoder layer can capture long-range correlation information, such as load changes and NO. x The relationship between emission lag effects is investigated to overcome the gradient vanishing problem in long-term dependency modeling of traditional RNN structures. The output generated by this process is a temporal enhanced feature, defined as a fourth dimension. In this embodiment, the fourth dimension is 256-dimensional to enhance the feature representation capability and depth of the model. The temporal extraction unit is represented by the following formula:

[0080]

[0081] in, For time domain features, The timing coding path consists of the first LSTM layer and the first Transformer encoder layer.

[0082] Furthermore, in this embodiment, the feature extraction module includes a frequency domain extraction unit, which includes a Fourier transform layer, a second LSTM layer, and a second Transformer encoder layer; the input time-series fusion data is converted from the time domain to the frequency domain in order to identify potential periodic patterns, oscillation modes, and energy distribution of characteristic frequency bands under different operating conditions.

[0083] The frequency domain extraction unit is used to extract frequency domain features from the time-series fused data, including:

[0084] The Fourier transform layer is used to convert the time-series fused data into frequency domain data, which has a fifth dimension; this achieves a mapping from time signals to spectral signals. This transformation can extract the periodic components, combustion fluctuation patterns, and dominant frequency characteristics hidden in the time series, generating frequency domain data. The fifth dimension of the frequency domain data is defined as 70×128 dimensions.

[0085] The second LSTM layer is used to extract the short-term temporal dependencies of the frequency domain data, generating primary frequency domain features with a sixth dimension. It captures the local dependencies of frequency features over time, such as the short-term fluctuations in spectral energy distribution during boiler load adjustments. Since the frequency domain data retains the sequential structure of the time series, the LSTM gating mechanism effectively extracts the short-term trends of spectral energy changes, outputting primary frequency domain features with a sixth dimension (70×256).

[0086] The second Transformer encoder layer is used to extract the global correlation of the primary frequency domain features, generating enhanced frequency domain features. These enhanced features have a seventh dimension (70×256). A self-attention mechanism is used to calculate the correlation between arbitrary frequency bands, thereby identifying the coupling characteristics and energy distribution patterns between different frequency components. For example, during deep peak shaving or changes in fuel characteristics, the combustion trend reflected by low-frequency components may be correlated with the transient fluctuations reflected by high-frequency components. The frequency domain extraction unit uses the following formula to represent this:

[0087]

[0088] in, For frequency domain characteristics, The frequency domain encoding path consists of the second LSTM layer and the second Transformer encoder layer, and FFT is the Fast Fourier Transform.

[0089] The feature extraction module further includes a nonlinear feature fusion unit. This unit interacts with and weights the time-domain enhanced features and frequency-domain enhanced features to obtain joint features. The joint features have a seventh dimension, as expressed by the following formula:

[0090]

[0091] in, For joint features, It is a nonlinear feature fusion unit based on multi-head attention.

[0092] The deep prediction module is used to perform deep semantic understanding on the joint features based on a preset LoRA-Large Language Model and output the NO within a predetermined second time period. x Concentration prediction sequence. Through global contextual modeling and complex correlation analysis of joint features, the model generates NO concentration prediction sequences within the prediction time window. x Concentration sequence output. In this embodiment, the deep prediction module includes a LoRA-large language model unit;

[0093] The LoRA-Large Language Model unit performs deep semantic understanding and contextual reasoning on the joint features based on the preset LoRA-Large Language Model to obtain semantic features. The semantic features have an eighth dimension, which is 70×4096 dimensions, to carry richer global information and contextual dependencies. Through LoRA (Low-Rank Adaptation of Large Language Models) technology, the large language model can significantly reduce the scale of training parameters and computational complexity while maintaining strong semantic expressive power, thereby improving the model's generalization performance under small sample and complex conditions.

[0094] The pre-defined LoRA-Large Language Model's base weight parameters remain frozen during training and do not participate in gradient updates; only its low-rank adapter parameters are trained. The LoRA-Large Language Model unit is based on a pre-defined Large Language Model (LLM), whose core structure includes a multi-layer Transformer encoder to capture higher-order associations and semantic dependencies of input features. This avoids overfitting or catastrophic forgetting problems under limited data conditions. Simultaneously, only the low-rank adapter parameters (LoRA parameters) inserted into the attention layer and feedforward network layer are fine-tuned. The LoRA parameters adopt a low-rank decomposition form, introducing trainable low-rank increment terms into the weight matrix to achieve rapid adaptation to specific task features. This design allows the model to retain the generalization ability of the original large model while being able to target NO... xThe temporal characteristics and energy distribution patterns of the prediction task are customized and optimized. The preset LoRA-Large Language Model is represented by the following formula:

[0095]

[0096] in, For semantic features, It is LoRA - a large language model.

[0097] Furthermore, in this embodiment, the depth prediction module includes a prediction unit, which is based on a three-layer fully connected network;

[0098] The prediction unit performs regression prediction processing on the semantic features, including:

[0099] The semantic features within a predetermined second time period are determined. Through a multi-layer nonlinear mapping structure of a three-layer fully connected network, the semantic features within the predetermined second time period are subjected to layer-by-layer dimensionality compression and feature mapping, continuously reducing the feature space from the eighth dimension to the ninth dimension (10×1), thus obtaining the NO within the predetermined second time period. x Concentration prediction sequence. The prediction unit is represented by the following formula:

[0100]

[0101] in, This means taking features only from the last 10 time steps (i.e., the next 10 minutes). That is, the predicted next ten minutes NO x Emissions sequence.

[0102] Furthermore, in this embodiment, a training module is also included, which is used to train the prediction system;

[0103] The training module is based on a phased, progressive training strategy to ensure stable model convergence and avoid overfitting. It controls the freezing and unfreezing processes of the linear projection unit, time-domain extraction unit, frequency-domain extraction unit, nonlinear feature fusion unit, and LoRA-large language model unit.

[0104] The term "freeze" means that the corresponding network layer parameters are kept out of backpropagation and gradient update during the current training phase, and the term "unfreeze" means that the parameter freeze state is lifted, allowing the corresponding network layer parameters to participate in backpropagation and gradient update.

[0105] The total training data volume is approximately 20,000 minutes, with each training epoch fully traversing the dataset once. The specific training process is as follows: First, the linear projection unit, temporal extraction unit, and frequency extraction unit are frozen, and only the nonlinear feature fusion unit and prediction unit at the end are trained, enabling the model to initially learn to perform fusion and prediction based on fixed features. This stage runs for approximately 10 training epochs. Next, the temporal and frequency extraction units are unfrozen to adapt them to the temporal pattern of the current task, while the nonlinear feature fusion unit and prediction unit are further fine-tuned. This stage runs for 15 training epochs. Then, the linear projection unit is unfrozen to optimize the mapping relationship from the original data to the feature space and to train it in conjunction with downstream networks. This stage runs for 10 training epochs. Finally, the parameters of the large language model (LLM) are efficiently fine-tuned (only the LoRA adapter parameters are trained) to stimulate its deep semantic understanding capabilities. This stage runs for 20 training epochs. Through the above progressive training sequence from the output end to the input end, the model refines the feature representation layer by layer. With a total of 55 training epochs, the model fully exploits the potential of the data while ensuring training stability, ultimately achieving high-precision NO. x Emissions prediction capability.

[0106] This invention extracts time-domain and frequency-domain features from historical and planned parameters, and further utilizes the LoRA-Large Language Model for deep semantic understanding to achieve NO of thermal power units. x High-precision emission prediction significantly improves the accuracy and reliability of forecasts. By fusing historical and planned parameters over time, the system can capture short-term transient fluctuations and long-term global coupling characteristics during combustion, enabling the prediction results to accurately reflect emission dynamics under rapid load changes and fuel characteristic fluctuations. Joint features enhance the richness and generalization ability of feature representation. Deep semantic modeling based on the LoRA-large language model gives the system powerful contextual understanding capabilities, allowing for advanced predictions under complex operating conditions.

[0107] It should be understood that the specific embodiments described above are merely illustrative or explanatory of the principles of the invention and do not constitute a limitation thereof. Therefore, any modifications, equivalent substitutions, improvements, etc., made without departing from the spirit and scope of the invention should be included within the protection scope of the invention. Furthermore, the appended claims are intended to cover all variations and modifications falling within the scope and boundaries of the appended claims, or equivalent forms of such scope and boundaries.

[0108] The present invention has been described above with reference to embodiments thereof. However, these embodiments are merely illustrative and not intended to limit the scope of the invention. The scope of the invention is defined by the appended claims and their equivalents. Various substitutions and modifications can be made by those skilled in the art without departing from the scope of the invention, and all such substitutions and modifications should fall within the scope of the invention.

[0109] Although embodiments of the present invention have been described in detail, it should be understood that various changes, substitutions, and modifications can be made to the embodiments of the present invention without departing from the spirit and scope of the invention.

[0110] Obviously, the above embodiments are merely illustrative examples for clear explanation and are not intended to limit the implementation. Those skilled in the art will recognize that other variations or modifications can be made based on the above description. It is neither necessary nor possible to exhaustively list all possible implementations here. However, obvious variations or modifications derived therefrom are still within the scope of protection of this invention.

Claims

1. A NO model based on a large time series model x The prediction system is characterized by, include: The module includes data preprocessing and feature mapping, feature extraction, and depth prediction. The data preprocessing and feature mapping module is used to collect historical parameters within a predetermined first time period and planned parameters within a predetermined second time period, and to map the historical parameters and planned parameters to a time feature space to obtain time-series fusion data. The feature extraction module is used to extract the temporal and frequency features of the time-series fusion data, and generate joint features based on the temporal and frequency features; The deep prediction module is used to perform deep semantic understanding on the joint features based on a preset LoRA-Large Language Model and output the NO within a predetermined second time period. x Concentration prediction sequences; The feature extraction module includes a temporal extraction unit; the temporal extraction unit includes a first LSTM layer and a first Transformer encoder layer; The time-domain extraction unit is used to extract time-domain features from the time-series fused data, including: The short-term temporal dependencies of the time-series fusion data are extracted using the first LSTM layer to generate primary temporal features, wherein the dimension of the primary temporal features is the third dimension. The global correlation of the temporal primary features is extracted using the first Transformer encoder layer based on the self-attention mechanism to generate temporal enhanced features, wherein the temporal enhanced features have a fourth dimension. The feature extraction module includes a frequency domain extraction unit, which includes a Fourier transform layer, a second LSTM layer, and a second Transformer encoder layer. The frequency domain extraction unit is used to extract frequency domain features from the time-series fused data, including: The time-series fusion data is converted to frequency domain data using the Fourier transform layer, and the frequency domain data has a fifth dimension. The short-term time dependencies of the frequency domain data are extracted using the second LSTM layer to generate primary frequency domain features, wherein the primary frequency domain features have a sixth dimension. The global correlation of the primary frequency domain features is extracted using the second Transformer encoder layer to generate enhanced frequency domain features, wherein the enhanced frequency domain features have a seventh dimension. The feature extraction module further includes a nonlinear feature fusion unit; The nonlinear feature fusion unit interacts and weights the time-domain enhancement features and frequency-domain enhancement features to obtain joint features, which have a seventh dimension.

2. The NO based on a large time-series model as described in claim 1 x The prediction system is characterized by, The data preprocessing and feature mapping module includes an acquisition unit and a preprocessing unit; The acquisition unit is used to acquire historical parameters within a predetermined first time period and planned parameters within a predetermined second time period at a minute-level. The historical parameters include: primary air volume, secondary air volume, air-coal ratio, coal feed rate, flue gas oxygen content, unit real-time load, fuel volatile matter content, and fuel calorific value. The planned parameters include: load command, ambient temperature, wind speed, solar radiation intensity, and weather condition code; The relationship between the predetermined first time period and the predetermined second time period includes: the predetermined second time period immediately follows the predetermined first time period; The preprocessing unit is used to perform time-series alignment, missing parameter correction, and normalization on the historical parameters and planned parameters obtained by the acquisition unit.

3. The NO based on a large time-series model as described in claim 2 x The prediction system is characterized by, The data preprocessing and feature mapping module also includes a linear projection unit and a temporal fusion unit; The linear projection unit is used to map the preprocessed historical parameters and the planned parameters respectively to obtain the high-dimensional feature vector sequence of historical parameters and the high-dimensional feature vector sequence of planned parameters; The time-series fusion unit is used to fuse the high-dimensional feature vector sequence of historical parameters and the high-dimensional feature vector sequence of planned parameters along the time dimension to obtain time-series fused data.

4. The NO based on a large time-series model as described in claim 3 x The prediction system is characterized by, The linear projection unit includes a first linear projection layer and a second linear projection layer; The first linear projection layer is used to map the preprocessed historical parameters to obtain a high-dimensional feature vector sequence of historical parameters, wherein the dimension of the high-dimensional feature vector sequence of historical parameters is the first dimension; The second linear projection layer is used to map the preprocessed planning parameters to obtain a high-dimensional feature vector sequence of planning parameters. The dimension of the high-dimensional feature vector sequence of planning parameters is the second dimension, and the first dimension and the second dimension have the same feature dimension at a single time step.

5. The NO based on a large time-series model as described in claim 4 x The prediction system is characterized by, The deep prediction module includes the LoRA-Large Language Model Unit; The LoRA-Big Language Model Unit performs deep semantic understanding and contextual reasoning on the joint features based on the preset LoRA-Big Language Model to obtain semantic features, the dimension of which is the eighth dimension. The preset LoRA-Large Language Model's base weight parameters are frozen during training and do not participate in gradient updates; only its low-rank adapter parameters are trained.

6. The NO based on a large time-series model as described in claim 5 x The prediction system is characterized by, The depth prediction module includes a prediction unit, which is based on a three-layer fully connected network. The prediction unit performs regression prediction processing on the semantic features, including: The semantic features within a predetermined second time period are determined. Through a multi-layer nonlinear mapping structure of a three-layer fully connected network, the semantic features within the predetermined second time period are subjected to layer-by-layer dimensionality compression and feature mapping, continuously reducing the feature space from the eighth dimension to the ninth dimension, thus obtaining the NO within the predetermined second time period. x Concentration prediction sequence.

7. The NO based on a large time-series model as described in claim 6 x The prediction system is characterized by, It also includes a training module for training the prediction system; The training module is based on a phased progressive training strategy, which controls the freezing and unfreezing process of the linear projection unit, time domain extraction unit, frequency domain extraction unit, nonlinear feature fusion unit and LoRA-large language model unit. The term "freeze" means that the corresponding network layer parameters are kept out of backpropagation and gradient update during the current training phase, and the term "unfreeze" means that the parameter freeze state is lifted, allowing the corresponding network layer parameters to participate in backpropagation and gradient update.

Citation Information

Patent Citations

  • Mixed deep neural network modeling-based boiler short-term NOx emission prediction method

    CN113947013A

  • Multi-modal large language model inference engine, site restoration method and storage medium

    CN121094213A