Dry quenching boiler inlet temperature prediction method based on DeepSeek large model
Through the prediction method based on the DeepSeek big model, the accuracy of the inlet temperature prediction of dry quenching boiler is solved, efficient energy utilization and automatic control are achieved, and the reliability and energy efficiency of the system are improved.
Patent Information
- Application Number
- CN202510666440.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-29
AI Technical Summary
The prior art is difficult to accurately predict the inlet temperature of the dry quenching boiler, resulting in low waste heat recovery efficiency, low energy utilization and high production costs, and traditional methods cannot effectively capture the time-dependent and nonlinear relationships of temperature changes.
The prediction method based on the DeepSeek big model is adopted, and through data preprocessing, graph convolution module, lightweight architecture adjustment and dynamic model distillation, combined with the prediction-control joint optimization framework, accurate prediction and automatic control of the inlet temperature of the dry quenching boiler is achieved.
It improves prediction accuracy, enhances the adaptability and generalization capabilities of the model, reduces computing resource consumption, and optimizes energy utilization and production costs.
Smart Images

Figure CN120561587A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of dry coke quenching technology, and in particular relates to a dry coke quenching boiler inlet temperature prediction method based on a DeepSeek large model. Background Art
[0002] After the coke matures, it is pushed out of the coke oven and loaded into the dry quenching device, where it is cooled using an inert gas (such as nitrogen). The temperature of the cooled coke drops to an appropriate range, and the cooling gas absorbs the heat of the coke and is sent to the waste heat boiler for cooling, generating steam or electricity, thereby realizing energy recovery and reuse.
[0003] The waste heat boiler (HRSG) inlet temperature is a critical parameter in the CDQ process, directly impacting waste heat recovery efficiency, steam production, and the overall system energy efficiency. Accurately predicting the boiler inlet temperature is crucial for optimizing CDQ operations, improving energy efficiency, reducing production costs, and minimizing environmental pollution.
[0004] The mechanistic modeling approach, based on the principles of thermodynamics and heat transfer, constructs a mathematical model describing boiler inlet temperature variations. By inputting parameters such as the physical properties of coke and the flow and temperature of the cooling gas, the boiler inlet temperature is calculated using the conservation of energy and heat transfer equations. This approach has a strong theoretical foundation and is highly interpretable, reflecting the essential laws of the process.
[0005] Traditional machine learning methods, such as support vector machines (SVM) and random forests (RF), extract features related to boiler inlet temperature through feature engineering, use historical data to train models, establish a mapping relationship between input features and temperature, and achieve prediction.
[0006] The MLP is a simple feedforward neural network that uses multiple layers of neurons to perform nonlinear transformations on input data, predicting boiler inlet temperature. While MLP has some nonlinear fitting capabilities, its processing power for time series data is limited, making it ineffective in capturing the temporal dependence of temperature changes.
[0007] LSTM: By introducing a gating mechanism, it effectively solves the vanishing gradient problem of traditional RNNs when processing long sequences. It can learn long-term dependencies and improve prediction accuracy. However, the LSTM structure is relatively complex, requiring high training time and computing resources.
[0008] Direct application of general large models: Not optimized for industrial time series data, resulting in redundant calculations (for example, the text tokenization mechanism is not applicable); high computing resource consumption and long single inference time.
[0009] In the coke dry quenching (CDQ) process, boiler inlet temperature is a key factor affecting heat recovery efficiency and system safety. Being able to predict boiler inlet temperature in advance allows for more precise control of the entire CDQ process.
[0010] Traditional methods for predicting boiler inlet temperature have the following problems: strong coupling of multiple variables, nonlinear influence of boiler inlet temperature and multiple factors such as circulation, CO, and coke discharge temperature, making it difficult for traditional linear models to accurately model; the prediction results have a lag, and existing methods (such as ARIMA) cannot capture long-term time series characteristics and nonlinear relationships between data, resulting in large prediction errors; industrial scenario data quality is poor, sensor noise is significant (signal-to-noise ratio SNR≤10dB), and there are data missing and asynchronous problems.
[0011] The information in this background technology section is only intended to enhance understanding of the overall background of the application and is not necessarily considered as an admission or any form of suggestion that the information constitutes the prior art already known to a person skilled in the art. Summary of the Invention
[0012] To address the above-mentioned issues, this application provides a method for predicting the inlet temperature of a dry coke quenching boiler based on the DeepSeek large model. The prediction method of this application can capture complex nonlinear relationships and interactions between multiple variables, and handle the strong coupling and nonlinear effects between multiple process parameters in the dry coke quenching process. By fine-tuning on a large amount of historical data, it can automatically learn and adapt to different operating conditions, reducing dependence on manual feature engineering and improving the generalization ability and prediction performance of the model.
[0013] In some embodiments of the present application, a method for predicting the inlet temperature of a dry coke quenching boiler based on a DeepSeek large model is provided, comprising the following steps:
[0014] S1. Data preprocessing, which includes the following steps:
[0015] S11. Data collection: Select a 24-hour sequence in history, and the sampling interval is recorded as T t-24h:t The process variables are the lower temperature of the cooling section T3, the upper temperature of the cooling section T4, the pre-storage chamber temperature T5, the circulating gas flow rate Q g , boiler inlet temperature T6, pre-storage chamber pressure P c , O2 and CO2 concentrations, boiler inlet pressure Pa, coke discharge amount N, coke discharge temperature T12, the above process parameters are collected in real time through corresponding sensors;
[0016] S12. Data cleaning: using time dimension interpolation to handle missing values of process variables;
[0017] S122. Outlier Detection: To avoid data bias, a joint detection of multiple CDQ variables is performed based on an improved DBSCAN algorithm. A dynamic threshold adjustment method is used to set outlier determination criteria, introducing temporal locality constraints.
[0018] S13. Feature engineering, which includes the following steps:
[0019] S131. Use sliding window method to construct time series dataset;
[0020] S14. Data standardization, including the following steps:
[0021] S141. Working condition adaptive segmented normalization,
[0022] In the CDQ process, fluctuations in production parameters under different operating conditions lead to differences in data distribution. Furthermore, the magnitude of various monitoring data varies widely. Therefore, based on the coke pushing plan, a working condition label s is obtained, where s represents the load level, to dynamically divide the data into normalized intervals.
[0023]
[0024] in, is the normalized data, X feat is the original input data, μ s and σ s
[0025] are the mean and standard deviation of the data under the current load condition, ∈=10 -6 ;
[0026] S142. Real-time update mechanism:
[0027] For each working condition tag s, historical data under the corresponding working condition is collected and its mean, standard deviation, maximum value, and minimum value distribution characteristics are calculated. The data normalization interval is dynamically calculated based on the current working condition tag s and these statistical characteristics. As new data is continuously collected, the normalization interval is updated in real time at a cycle of every minute or every five minutes to match the actual distribution of data under the current working condition.
[0028] S2. Model construction, including the following steps:
[0029] S21. Graph convolution module, including building a graph structure, selecting nodes corresponding to different sensor variables, and defining the adjacency matrix A by considering the linear relationship and nonlinear dependency between variables.
[0030] S22. Adjust the lightweight architecture of DeepSeek-7B, including model transformation: select the backbone network for adjustment, retain the first 8 Transformer layers of DeepSeek-7B, with a hidden layer dimension of 768, introduce an exponential decay factor to strengthen data weights, and implement a time-decayed attention mechanism:
[0031]
[0032] Among them, Q is the query matrix; K is the key matrix; V is the value matrix; d k The dimensions of K and Q; Q, K, and V are all derived from the input X in After linear transformation, we get: Δt is the time interval matrix, which expresses the time difference between K and Q;
[0033] A bidirectional GRU layer is added to the adaptation layer to capture local temporal information and obtain the feature representation X gru After the bidirectional GRU layer, a fully connected layer is connected to the X layer through linear transformation and activation function operation. gru Mapped to the target dimension, output the teacher model's prediction result y for the boiler inlet temperature T6 teacher ;
[0034] S23. Dynamic model distillation, which includes the distillation strategy: using the adjusted DeepSeek-7B as the teacher model, a lightweight TCN network with 4 layers, 3 convolution kernels, and 5M parameters as the student model, and joint temperature scaling.
[0035] KL divergence and MSE loss are used as loss functions;
[0036] S24. Prediction-control joint optimization framework, using a penalty function-based reinforcement learning control method; the system state space includes the predicted temperature T_pred and key control variables, the key control variables include the fan speed R f , amplitude A, air conduction flow Q a and bypass opening u bypass , and its penalty function is:
[0037] R t =-(||T pred,t -T set ||+0.3||Δu t ||)-β·max(0,T pred,t -950)
[0038] The strict constraints are: T_pred,t≤980℃, 980℃ is the absolute threshold; T_pred,t is the T6 predicted temperature at time t, T set T6 temperature setting value, i.e. working target value, Δu_t=u t -ut-1 is the range of control variable, β is the over-temperature penalty coefficient, u_t=[R f ,A,Q a ,u bypass ] is the multi-dimensional control input;
[0039] When the T6 predicted temperature exceeds the threshold, the forced override bypass opening is:
[0040] u bypass =max(u bypass RL ,α(T_pred,t-850)) is the bypass recommended by the RL strategy, and the rest of the control quantities are still determined by the RL strategy. α is the bypass adjustment coefficient, such as 0.5, T_pred,t is the predicted temperature at time t, u bypass RL The bypass opening recommended by the RL strategy ensures that the temperature does not exceed the safety limit, while continuously optimizing the control efficiency through RL under normal operating conditions.
[0041] S3. Model training and evaluation, which includes:
[0042] S31. Two-stage training strategy. In the pre-training stage, the historical working condition data of one year is selected and divided into training set and test set in chronological order. The hyperparameters required for model pre-training are set, and the learning rate is 3×10 -5 , Batch Size is 32, training cycle is 50 epochs, early stopping mechanism is added to avoid overfitting, AdamW is selected as the optimizer, and weight decay is 1×10 -4 ;
[0043] S32. Fine-tuning phase: Prepare one month of process parameter data and use the method of taking data once every minute to increase the amount of training data; divide the training set, validation set, and test set into a ratio of 7:2:1; then, use the freezing strategy to freeze the first 6 layers of DeepSeek parameters and only train the GRU and prediction head; finally, set the learning rate of the fine-tuning phase to 1×10 -5 , and adopt cosine annealing scheduling, and keep the batch size at 32;
[0044] S33. Model Forecast Evaluation: During the forecasting process, the forecast value at the current time point is integrated into the forecast data input for the next time period, and the forecast is performed in strict chronological order. Evaluation metrics such as mean absolute error (MAE), root mean square error (Rmse), and coefficient of determination (R²) are used to comprehensively measure the model's forecast performance.
[0045] S34. Model predictive control logic: Call the prediction model every hour to update the temperature forecast for the next 10 hours; use the prediction-control joint optimization framework proposed in S25 to optimize the control variable fan speed Rf , amplitude A, air conduction flow Q a and bypass opening u bypass If the predicted temperature exceeds 980°C, emergency control will be triggered to increase the bypass opening.
[0046] In some embodiments of the present application, the time dimension interpolation adopts the cubic spline interpolation method, and the missing points are filled by interpolation using the adjacent time point data; for the time interval [t k , t k+1 Any time point t within ] a , the interpolation formula is:
[0047] v(t a )=a k +b k ·(t a -t k )+c k ·(t a -t k ) 2 +d k ·(t a -t k ) 3
[0048] Among them, v refers to a variable mentioned above, v(t a ) indicates missing variables, and the interval starting point is a k =var(t k ), b k is the slope of the linear term, c k is the quadratic curvature, d k is the cubic rate of change.
[0049] In some embodiments of the present application, in S122, if a point is determined to be an outlier more than three times within one hour, it is marked as abnormal. The abnormality determination conditions are as follows:
[0050] |x t -μ w |>3σ w (1+0.2 var(X))
[0051] Among them, x t is the observed value of all variables at time t, μ w and σ w are the mean and standard deviation calculated within a sliding window of w = 60 seconds, var(X) is the variance of the global dataset, and (1 + 0.2 var(X)) is a dynamic adjustment term that determines the threshold by adjusting the local density.
[0052] In some embodiments of the present application, the sliding window method used in S131 to construct a time series data set is specifically as follows: assuming that the window size is n time steps, the input of each sample is all sensor variable values of the previous n time steps, including T6, and the predicted output is the T6 value of the next time step; the training set selects the first T train -n time steps of data, where T train Represents the training time step, and the test set selects the subsequent time step data to ensure that the test data does not overlap with the training data to avoid data leakage; the time domain feature extraction is to set the sliding window statistic as the mean μ w , variance σ w , skewness γ w and the first-order difference is Δx t =x t -x t-1 , where the window size is w = 60 seconds; the sliding window statistical features reflect the distribution characteristics and fluctuation of the signal as a whole, while the first-order difference is used to capture the short-term signal change trend; the above feature extraction operation is performed on each variable separately, and finally the features extracted from all variables are integrated to form a unified feature set X feat .
[0053] In some embodiments of the present application, the expression of the graph structure in S21 is as follows:
[0054]
[0055] in, Represent two different and normalized variables, is a variable and variables The Pearson correlation coefficient, which measures linear correlation; For variables and variables Mutual information of , which measures nonlinear statistical dependence; ∈ = 10 -6 , to prevent division by zero;
[0056] By the adjacency matrix A ij and the normalized node feature X norm Perform a specific graph convolution operation to obtain the output X of the graph convolution model graph ;
[0057]
[0058] in, A is the adjacency matrix, and I is the identity matrix used to add self-connections to avoid information loss; for The transition matrix W is the trainable matrix; σ is the Relu activation function; T is the time step, N is the feature dimension, and D is the graph convolution hidden layer dimension;
[0059] Flatten and linearly project the features processed by the spatiotemporal graph convolutional network to adjust the dimension to Get the input X of the Transformer architecture in the deepseek model in .
[0060] In some embodiments of the present application, the loss function in S23 is specifically: temperature-scaled KL divergence introduces MSE loss to measure the difference between the student model prediction value and the true value to optimize the model:
[0061]
[0062] in, is the total loss value; y teacher is the output probability distribution of the teacher model; y student is the output probability distribution of the student model; MSE MSE is the predicted value y of the student model student The mean squared error with the label y.
[0063] Compared with the prior art, this application has at least the following features:
[0064] 1. Precision breakthrough: improving the reliability of industrial systems.
[0065] 2. Adaptability: Rapid migration of multiple working conditions is achieved through MSTGC and meta-learning.
[0066] 3. Energy efficiency optimization: Dynamic distillation technology reduces energy consumption of edge devices.
[0067] 4. This application uses the real-time changing data of the CDQ process, i.e., process variables, to predict the CDQ boiler inlet temperature T6; the key control variables are artificially controlled variables such as the fan speed R f , amplitude A, air conduction flow Q a and bypass opening u bypass By changing the key control variables to control / adjust the changes in process variables, automatic control is achieved by using the prediction-control joint optimization method and reinforcement learning.
[0068] 5. Process variables (such as temperature, pressure, flow, etc.) have time-series correlation (current values are related to historical values), and their changing patterns can be captured through time-series prediction models; key control variables are the leading factors affecting process variables. By reversely deducing the adjustment strategy of the control variables through the prediction results, "predict first, intervene later" feedforward control can be achieved, thereby improving the system response speed.
[0069] The prediction method of the present application can capture complex nonlinear relationships and interactions between multiple variables, and handle the strong coupling and nonlinear effects between multiple process parameters in the dry quenching process; by fine-tuning on a large amount of historical data, it can automatically learn and adapt to different operating conditions, reducing dependence on manual feature engineering and improving the model's generalization ability and prediction performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] The drawings in the specification, which constitute a part of this application, are used to provide further understanding of this application. The illustrative embodiments of this application and their descriptions are used to explain this application and do not constitute improper limitations on this application.
[0071] Figure 1 Schematic diagram of a flow chart of a method for predicting the inlet temperature of a dry coke quenching boiler based on a DeepSeek large model in some embodiments of the present application;
[0072] Figure 2 This is a flowchart of the steps of data preprocessing in some embodiments of the present application;
[0073] Figure 3 This is a flowchart of the steps for model configuration in some embodiments of the present application. DETAILED DESCRIPTION
[0074] Below, specific embodiments of the present invention are described in detail with reference to the accompanying drawings, but are not intended to limit the present invention. To enable those skilled in the art to better understand the technical solutions of the present disclosure, the present disclosure is described in detail below with reference to the accompanying drawings and specific embodiments. Below, embodiments of the present disclosure are further described in detail with reference to the accompanying drawings and specific embodiments, but are not intended to limit the present disclosure.
[0075] All terms (including technical or scientific terms) used in this disclosure have the same meaning as those understood by one of ordinary skill in the art to which this disclosure belongs, unless otherwise specifically defined. It should also be understood that terms defined in, for example, general dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an idealized or highly formal sense, unless explicitly defined herein.
[0076] Technologies, methods, and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the technologies, methods, and equipment should be considered part of the specification.
[0077] In some embodiments of the present application, a method for predicting the inlet temperature of a dry coke quenching boiler based on a DeepSeek large model is provided, comprising the following steps:
[0078] S1. Data preprocessing, which includes the following steps:
[0079] S11. Data collection: Select a 24-hour sequence in history, and the sampling interval is recorded as T t-24h:t The process variables are the lower temperature of the cooling section T3, the upper temperature of the cooling section T4, the pre-storage chamber temperature T5, the circulating gas flow rate Q g , boiler inlet temperature T6, pre-storage chamber pressure P c , O2 and CO2 concentrations, boiler inlet pressure Pa, coke discharge amount N, coke discharge temperature T12, the above process variables are collected in real time through corresponding sensors to ensure the accuracy and timeliness of the data.
[0080] S12. Data cleaning, including:
[0081] S121. Missing Value Handling: During data collection, CDQ sensors are susceptible to sensor failures, poor installation and maintenance, and harsh environments such as high temperatures and dust, which can lead to missing data.
[0082] The time dimension interpolation is used to process the missing values of the above variables; the time dimension interpolation is to use the cubic spline interpolation method to fill the missing points with the adjacent time point data; for the time interval [t k , t k+1 Any time point t within ] a , the interpolation formula is:
[0083] v(t a )=a k +b k ·(t a -t k )+c k ·(t a -t k ) 2 +d k ·(t a -t k ) 3
[0084] Among them, v refers to a variable mentioned above, v(t a ) indicates missing variables, and the interval starting point is a k =var(t k ), b k is the slope of the linear term, c k is the quadratic curvature, d k is the cubic rate of change;
[0085] S122. Outlier Detection: To avoid data bias, a joint detection of CDQ multivariate data is performed based on an improved DBSCAN algorithm. A dynamic threshold adjustment method is used to set anomaly determination criteria, introducing temporal locality constraints. If a point is identified as an outlier more than three times within an hour, it is marked as an anomaly. The anomaly determination criteria are as follows:
[0086] |x t -μ w |>3σ w (1+0.2 var(X))
[0087] Among them, x t is the observed value of all variables at time t, μ w and σ w are the mean and standard deviation calculated within a sliding window of w = 60 seconds, var(X) is the variance of the global dataset, and (1 + 0.2 var(X)) is a dynamic adjustment term that determines the threshold by adjusting the local density.
[0088] S13. Feature engineering, which includes the following steps:
[0089] S131. Use the sliding window method to construct a time series dataset. Assume that the window size is n time steps, the input of each sample is all sensor variable values in the previous n time steps, including T6, and the predicted output is the T6 value of the next time step; the training set is the first T time steps. train -n time steps of data, where T train Represents the training time step, and the test set selects the time step data after that to ensure that the test data does not overlap with the training data to avoid data leakage.
[0090] Time domain feature extraction is to set the sliding window statistics as the mean μ w , variance σ w , skewness γ w and the first-order difference is Δx t =x t -x t-1 , where the window size is w = 60 seconds; the sliding window statistical features reflect the distribution characteristics and fluctuation of the signal as a whole, while the first-order difference is used to capture the short-term signal change trend; the above feature extraction operation is performed on each variable separately, and finally the features extracted from all variables are integrated to form a unified feature set X feat .
[0091] S14. Data standardization, including the following steps:
[0092] S141. Working condition adaptive segmented normalization,
[0093] In the CDQ process, fluctuations in production parameters under different operating conditions lead to significant differences in data distribution. Furthermore, the magnitude of various monitoring data varies widely. Therefore, based on the coke pushing plan, we obtain the operating condition label s, where s represents the load level, to dynamically divide the data into normalized intervals.
[0094]
[0095] in, is the normalized data, X feat is the original input data, μ s and σ s
[0096] are the mean and standard deviation of the data under the current load condition, ∈=10 -6 ;
[0097] S142. Real-time update mechanism:
[0098] For each working condition label s, historical data under the corresponding working condition is collected and its mean, standard deviation, maximum value, and minimum value distribution characteristics are statistically analyzed. The data normalization interval is dynamically calculated based on the current working condition label s and these statistical characteristics. As new data is continuously collected, the normalization interval is updated in real time at a cycle of every minute or every five minutes to match the actual distribution of data under the current working condition.
[0099] S2. Model construction, including the following steps:
[0100] S21. Graph convolution module, including building a graph structure, selecting nodes corresponding to different sensor variables, and defining the adjacency matrix A by comprehensively considering the linear relationship and nonlinear dependency between variables. The expression of the graph structure is as follows:
[0101]
[0102] in, Represent two different and normalized variables, is a variable and variables The Pearson correlation coefficient, which measures linear correlation; For variables and variables Mutual information of , which measures nonlinear statistical dependence; ∈ = 10 -6 , to prevent division by zero;
[0103] By the adjacency matrix A ij and the normalized node feature X norm Perform a specific graph convolution operation to obtain the output X of the graph convolution model graph ;
[0104]
[0105] in, A is the adjacency matrix, and I is the identity matrix used to add self-connections to avoid information loss; for The transition matrix W is the trainable matrix; σ is the Relu activation function; T is the time step, N is the feature dimension, and D is the graph convolution hidden layer dimension;
[0106] Flatten and linearly project the features processed by the spatiotemporal graph convolutional network to adjust the dimension to Get the input X of the Transformer architecture in the deepseek model in ;
[0107] S22. Adjust the lightweight architecture of DeepSeek-7B, including model transformation: select the backbone network for adjustment, retain the first 8 Transformer layers of DeepSeek-7B, with a hidden layer dimension of 768, introduce exponential decay silver to strengthen the weight of recent data, and implement the time-decayed attention mechanism:
[0108]
[0109] Among them, Q is the query matrix; K is the key matrix; V is the value matrix; d k The dimensions of K and Q; Q, K, and V are all derived from the input X in After linear transformation, we get: Δt is the time interval matrix, which expresses the time difference between K and Q;
[0110] A bidirectional GRU layer is added to the adaptation layer to capture local temporal information and obtain the feature representation X gru After the bidirectional GRU layer, a fully connected layer is connected to the X layer through linear transformation and activation function operation. gru Mapped to the target dimension, output the teacher model's prediction result y for the boiler inlet temperature T6 teacher .
[0111] S23. Dynamic model distillation, which includes the distillation strategy: using the adjusted DeepSeek-7B as the teacher model, a lightweight TCN network with 4 layers, 3 convolution kernels, and 5M parameters as the student model, and joint temperature scaling.
[0112] KL divergence and MSE loss are used as loss functions;
[0113] Loss function: Temperature scaling KL divergence introduces MSE loss to measure the difference between the student model prediction value and the true value
[0114] The difference between the two is used to optimize the model:
[0115]
[0116] in, is the total loss value; y teacher is the output probability distribution of the teacher model; y student is the output probability distribution of the student model; MSE MSE is the predicted value y of the student model student Mean squared error with label y;
[0117] S24. Prediction-control joint optimization framework adopts a reinforcement learning control method based on penalty function. The system state space contains the predicted temperature T_pred and key control variables (such as fan speed R f , amplitude A, air conduction flow Q a and bypass opening u bypass ), bypass opening is the opening of the CDQ bypass valve, and amplitude is the amplitude of the vibrating feeder;
[0118] Its penalty function is:
[0119] R t =-(||T pred,t -T set ||+0.3||Δu t ||)-β·max(0,T pred,t -950)
[0120] The strict constraints are: T_pred,t≤980℃, 980℃ is the absolute threshold; T_pred,t is the T6 predicted temperature at time t, T set T6 temperature setting value, i.e. working target value, Δu_t=u t -u t-1 is the range of control variable, β is the over-temperature penalty coefficient, u_t=[R f ,A,Q a ,u bypass ] is the multi-dimensional control input.
[0121] The penalty function of this application can achieve: 1) accurate tracking of target temperature; 2) smooth control action; 3) strict temperature constraint.
[0122] When the T6 predicted temperature exceeds the threshold, the forced override bypass opening is:
[0123] u bypass =max(u bypass RL,α(T_pred,t-850)) is the bypass recommended by the RL strategy, and the rest of the control quantities are still determined by the RL strategy. α is the bypass adjustment coefficient, such as 0.5, T_pred,t is the predicted temperature at time t, u bypass RL is the bypass opening recommended by the RL strategy. This design ensures that the temperature strictly does not exceed the safety limit and can continuously optimize the control efficiency through RL under normal operating conditions.
[0124] S3. Model training and evaluation, which includes:
[0125] S31. Two-stage training strategy: In the pre-training stage, the operating condition data of one year is selected and divided into a training set (e.g., the first 90%) and a test set (e.g., the last 10%) in chronological order. The hyperparameters required for model pre-training are set, and the learning rate is 3×10 -5 , Batch Size is 32, training cycle is 50 epochs, early stopping mechanism is added to avoid overfitting, AdamW is selected as the optimizer, and weight decay is 1×10 -4 ;
[0126] S32. Fine-tuning stage: Prepare nearly one month of process parameter data and use the method of taking data once every minute to increase the amount of training data; divide the training set, validation set, and test set into a ratio of 7:2:1; then, use the freezing strategy to freeze the parameters of the first 6 layers of DeepSeek and only train the GRU and prediction head; finally, set the learning rate of the fine-tuning stage to 1×10 -5 , and adopt cosine annealing scheduling, and keep the batch size at 32;
[0127] S33. Model Forecast Evaluation: During the forecasting process, the forecast value at the current time point is integrated into the forecast data input for the next time period, and the forecast is performed in strict chronological order. Evaluation metrics such as mean absolute error (MAE), root mean square error (Rmse), and coefficient of determination (R²) are used to comprehensively measure the model's forecast performance.
[0128] S34. Model predictive control logic: Call the prediction model every hour to update the temperature forecast for the next 10 hours; use the prediction-control joint optimization framework proposed in S25 to optimize the control variable fan speed R f , amplitude A, air conduction flow Q a and bypass opening u bypass If the predicted temperature exceeds 980°C, emergency control will be triggered to increase the bypass opening.
[0129] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for predicting the inlet temperature of a dry coke quenching boiler based on a DeepSeek large model, characterized in that: The following steps are involved: S1. Data preprocessing, which includes the following steps: S11. Data collection: Select a 24-hour sequence in history, and the sampling interval is recorded as T t-24h:t The process variables are the lower temperature of the cooling section T3, the upper temperature of the cooling section T4, the pre-storage chamber temperature T5, the circulating gas flow rate Q g , boiler inlet temperature T6, pre-storage chamber pressure P c , O2 and CO2 concentrations, boiler inlet pressure Pa, coke discharge volume N, coke discharge temperature T 12 ,The above process variables are collected in real time through the corresponding sensors; S12. Data cleaning: using time dimension interpolation to handle missing values of process variables; S122. Outlier Detection: To avoid data bias, a joint detection of multiple CDQ variables is performed based on an improved DBSCAN algorithm. A dynamic threshold adjustment method is used to set outlier determination criteria, introducing temporal locality constraints. S13. Feature engineering, which includes the following steps: S131. Use sliding window method to construct time series dataset; S14. Data standardization, including the following steps: S141. Adaptive segmented normalization of operating conditions. In the CDQ process, fluctuations in production parameters under different operating conditions lead to differences in data distribution. Furthermore, the magnitude of various monitoring data varies widely. Therefore, based on the coke pushing plan, an operating condition label s is obtained, where s represents the load level, to dynamically divide the data into normalization intervals. in, is the normalized data, X feat is the original input data, μ s and σ s are the mean and standard deviation of the data under the current load condition, ∈=10 -6 ; S142. Real-time update mechanism: For each working condition tag s, historical data under the corresponding working condition is collected and its mean, standard deviation, maximum value, and minimum value distribution characteristics are calculated. The data normalization interval is dynamically calculated based on the current working condition tag s and these statistical characteristics. As new data is continuously collected, the normalization interval is updated in real time at a cycle of every minute or every five minutes to match the actual distribution of data under the current working condition. S2. Model construction, including the following steps: S21. Graph convolution module, including building a graph structure, selecting nodes corresponding to different sensor variables, and defining the adjacency matrix A by considering the linear relationship and nonlinear dependency between variables; S22. Adjust the lightweight architecture of DeepSeek-7B, including model transformation: select the backbone network for adjustment, retain the first 8 Transformer layers of DeepSeek-7B, with a hidden layer dimension of 768, introduce an exponential decay factor to strengthen data weights, and implement a time-decayed attention mechanism: Among them, Q is the query matrix; K is the key matrix; V is the value matrix; d k The dimensions of K and Q; Q, K, and V are all derived from the input X in After linear transformation, we get: Δt is the time interval matrix, which expresses the time difference between K and Q; A bidirectional GRU layer is added to the adaptation layer to capture local temporal information and obtain the feature representation X gru ; After the bidirectional GRU layer, the fully connected layer is connected, and X is transformed through linear transformation and activation function operation. gru Mapped to the target dimension, output the teacher model's prediction result y for the boiler inlet temperature T6 teacher ; S23. Dynamic model distillation, which includes the distillation strategy: using the adjusted DeepSeek-7B as the teacher model, a lightweight TCN network with 4 layers, 3 convolution kernels, and 5M parameters as the student model, and joint temperature scaling. KL divergence and MSE loss are used as loss functions; S24. Prediction-control joint optimization framework, using a penalty function-based reinforcement learning control method; the system state space includes the predicted temperature T_pred and key control variables, the key control variables include the fan speed R f , amplitude A, air conduction flow Q a and bypass opening u bypass , and its penalty function is: R t =-(||T pred,t -T set ||+0.3||Du t ||)-β·max(0,T pred,t -950) The strict constraints are: T_pred,t≤980℃, 980℃ is the absolute threshold; T_pred,t is the T6 predicted temperature at time t, T set T6 temperature setting value, i.e. working target value, Δu_t=u t -u t-1 is the range of control variable, β is the over-temperature penalty coefficient, u_t=[R f ,A,Q a ,u bypass ] is the multi-dimensional control input; When the T6 predicted temperature exceeds the threshold, the forced override bypass opening is: u bypass =max(u bypass RL ,α(T_pred,t-850)) is the bypass recommended by the RL strategy, α is the bypass adjustment coefficient, T_pred,t is the predicted temperature at time t, u bypass RL is the bypass opening recommended by the RL strategy; S3. Model training and evaluation, which includes: S31. Two-stage training strategy: In the pre-training stage, one year of historical operating data is selected and divided into training and test sets in chronological order. The hyperparameters required for model pre-training are set. S32. Fine-tuning phase: Prepare one month of process parameter data, extracting data every minute to increase the amount of training data. Split the training, validation, and test sets into a ratio of 7:2:
1. Then, adopt a freezing strategy. Finally, set the learning rate for the fine-tuning phase. S33. Model prediction evaluation: During the prediction process, the prediction value at the current time point is integrated into the prediction data input for the next time period, and the prediction is performed in chronological order. S34. Model predictive control logic: Call the prediction model every hour to update the temperature forecast for the next 10 hours; use the prediction-control joint optimization framework proposed in S25 to optimize the control variable fan speed R f , amplitude A, air conduction flow Q a and bypass opening u bypass If the predicted temperature exceeds 980°C, emergency control will be triggered to increase the bypass opening.
2. The method for predicting the inlet temperature of a dry coke quenching boiler based on the DeepSeek large model according to claim 1, characterized in that: The time dimension interpolation adopts cubic spline interpolation method, using adjacent time point data to interpolate and fill in the missing points; for the time interval [t k , t k+1 Any time point t within ] a , the interpolation formula is: v(t a )=a k +b k ·(t a -t k )+c k ·(t a -t k ) 2 +d k ·(t a -t k ) 3 Among them, v refers to a variable mentioned above, v(t a ) indicates missing variables, and the interval starting point is a k =var(t k ), b k is the slope of the linear term, c k is the quadratic curvature, d k is the cubic rate of change.
3. The method for predicting the inlet temperature of a dry coke quenching boiler based on the DeepSeek large model according to claim 1, characterized in that: In S122, if a point is identified as an outlier more than three times within one hour, it is marked as abnormal. The abnormality determination conditions are as follows: |x t -m w |>3s w ·(1+0.2·var(X)) Among them, x t is the observed value of all variables at time t, μ w and σ w are the mean and standard deviation calculated within a sliding window of w = 60 seconds, var(X) is the variance of the global dataset, and (1 + 0.2 var(X)) is a dynamic adjustment term that determines the threshold by adjusting the local density.
4. The method for predicting the inlet temperature of a dry coke quenching boiler based on the DeepSeek large model according to claim 1, characterized in that: The specific steps of using the sliding window method to construct the time series data set in S131 are as follows: assuming that the window size is n time steps, the input of each sample is all sensor variable values of the previous n time steps, including T6, and the predicted output is the T6 value of the next time step; the training set selects the first T train -n time steps of data, where T train Represents the training time step, and the test set selects the subsequent time step data to ensure that the test data does not overlap with the training data to avoid data leakage; the time domain feature extraction is to set the sliding window statistic as the mean μ w , variance σ w , skewness γ w and the first-order difference is Δx t =x t -x t-1 , where the window size is w = 60 seconds; the sliding window statistical features reflect the distribution characteristics and fluctuation of the signal as a whole, while the first-order difference is used to capture the short-term signal change trend; the above feature extraction operation is performed on each variable separately, and finally the features extracted from all variables are integrated to form a unified feature set X feat .
5. The method for predicting the inlet temperature of a dry coke quenching boiler based on the DeepSeek large model according to claim 1, characterized in that: The expression of the graph structure in S21 is as follows: in, Represent two different and normalized variables, is a variable and variables The Pearson correlation coefficient, which measures linear correlation; For variables and variables Mutual information of , which measures nonlinear statistical dependence; ∈ = 10 -6 , to prevent division by zero.
6. The method for predicting the inlet temperature of a dry coke quenching boiler based on the DeepSeek large model according to claim 5, characterized in that: By the adjacency matrix A ij and the normalized node feature X norm Perform a specific graph convolution operation to obtain the output X of the graph convolution model graph ; in, A is the adjacency matrix, and I is the identity matrix used to add self-connections to avoid information loss; for The transition matrix W is the trainable matrix; σ is the Relu activation function; T is the time step, N is the feature dimension, and D is the graph convolution hidden layer dimension; Flatten and linearly project the features processed by the spatiotemporal graph convolutional network to adjust the dimension to Get the input X of the Transformer architecture in the deepseek model in .
7. The method for predicting the inlet temperature of a dry coke quenching boiler based on the DeepSeek large model according to claim 1, characterized in that: The loss function in S23 is as follows: Temperature scaling KL divergence introduces MSE loss to measure the difference between the student model prediction value and the true value to optimize the model: in, is the total loss value; y teacher is the output probability distribution of the teacher model; y student is the output probability distribution of the student model; MSE MSE is the predicted value y of the student model student The mean squared error with the label y.
8. The method for predicting the inlet temperature of a dry coke quenching boiler based on the DeepSeek large model according to claim 1, characterized in that: The hyperparameters required for model pre-training are specifically the learning rate 3×10 -5 , Batch Size is 32, training cycle is 50 epochs, early stopping mechanism is added to avoid overfitting, AdamW is selected as the optimizer, and weight decay is 1×10 -4 .
9. The method for predicting the inlet temperature of a dry coke quenching boiler based on the DeepSeek large model according to claim 1, characterized in that: In S33, the mean absolute error (MAE), root mean square error (Rmse), and determination coefficient (R2) are used as evaluation indicators to measure the prediction effect of the model.
10. The method for predicting the inlet temperature of a dry coke quenching boiler based on the DeepSeek large model according to claim 1, characterized in that: In S32, the amount of trainable data is increased; the training set, validation set, and test set are divided into a ratio of 7:2:1; the freezing strategy is to freeze the first 6 layers of DeepSeek parameters and only train the GRU and prediction head; the learning rate of the fine-tuning stage is set to 1×10 -5 , and cosine annealing scheduling is adopted, and the batch size is kept at 32.
Citation Information
Cited By
Improvement method for serial compilation of coke pushing operation of large coke oven
CN117763798A
Improved method for programming a large coke oven pushing operation sequence
CN117763798B
Blast furnace state monitoring method and device based on interpretability enhanced neural network
CN121354737A
Circulating fluidized bed temperature prediction method based on dynamic correlation guided graph space-time learning
CN122015084A