Multi-parameter cooperative temperature control system of energy-saving industrial oven
The stable load data is obtained through high-precision sensors and Kalman filters, combined with support vector regression and deep Q network to optimize heating power and airflow speed, solving the problem of temperature control instability in industrial ovens under load changes, and achieving high-efficiency energy consumption optimization.
Patent Information
- Application Number
- CN202510457552.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing industrial oven temperature control system is difficult to respond quickly to thermal inertia fluctuations in scenarios where load changes frequently, resulting in temperature instability and increased energy consumption, and lacks targeted optimization design for load changes.
High-precision sensors are used to obtain stable load data in combination with Kalman filters, and the temperature recovery time and energy consumption are predicted using the support vector regression model. The heating power and airflow speed are dynamically optimized through the deep Q network, combined with reinforcement learning algorithms to achieve dynamic balance between temperature and energy consumption, and the system strategy is continuously optimized through the closed-loop feedback mechanism.
It significantly improves the temperature control accuracy and energy consumption efficiency of industrial ovens in load changes scenarios, reduces temperature fluctuations and energy consumption waste, and enhances the system's adaptability and long-term stability.
Smart Images

Figure CN120295136A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of temperature control of industrial ovens, and particularly to a multi-parameter collaborative temperature control system for energy-saving industrial ovens. Background Art
[0002] As a core heat treatment equipment in modern manufacturing, energy-saving industrial ovens play an important role in multiple industrial fields, such as baking and drying in food processing, high-temperature curing of chemical materials, and moisture removal in the production of electronic components. With the rapid development of intelligent manufacturing technology, the temperature control system of industrial ovens has gradually evolved towards high efficiency, precision, and adaptability. In the food processing industry, the oven needs to process different batches or types of foods, such as bread, biscuits, or meat products, and there are significant differences in their load characteristics (such as the number of trays, product thickness, and material density); in the chemical industry, the heat absorption characteristics of materials during the curing process vary with batches; in the electronics industry, changes in the size and arrangement of components also pose a need for dynamic adjustment of the heat distribution. In these scenarios, the frequent changes in the load directly affect the thermal inertia of the oven, resulting in a decrease in temperature stability or an increase in energy consumption. For example, a sudden increase in the number of trays during the baking process may extend the temperature recovery time, thereby affecting the color and taste of the food and even causing energy waste. In addition, external environmental factors such as seasonal temperature differences, workshop ventilation conditions, and fluctuations in energy prices further exacerbate the challenges to the operating efficiency and control accuracy of the oven. Therefore, developing an intelligent temperature control system that can dynamically adapt to load changes, optimize temperature control, and improve energy consumption efficiency has become a key requirement for enhancing the performance of industrial ovens.
[0003] In the prior art, the temperature control systems of industrial ovens mostly adopt methods such as PID (Proportional-Integral-Derivative) control or Model Predictive Control (MPC). These technologies can achieve good temperature control effects under stable load conditions. However, in the actual application scenarios with frequent load changes, the limitations of the existing methods are significantly exposed. Taking PID control as an example, its adjustment method based on the feedback mechanism has a certain lag. When the number of trays or the thickness of the material in the oven suddenly changes, the system is difficult to quickly respond to the fluctuations in thermal inertia, resulting in overshoot or oscillation of the temperature, which in turn affects the product quality and increases energy consumption waste. For example, in food baking, a sudden increase in the load may extend the temperature recovery time by 10 - 20 minutes, directly leading to uneven baking quality; in chemical curing, temperature fluctuations may affect the physical properties of the material. Although MPC has a certain degree of forward-looking through the prediction model, its performance highly depends on the accurate description of the system dynamics, and the non-linear thermal inertia fluctuations caused by load changes are often difficult to be accurately modeled, resulting in a decline in the control effect.
[0004] In addition, existing temperature control systems usually only focus on the stability of a single temperature parameter and lack targeted optimization design for thermal inertia fluctuations under load changes. Therefore, how to design an intelligent temperature control method that can quickly respond to thermal inertia fluctuations and achieve precise temperature control in scenarios where the load changes frequently has become the core problem faced by current technologies. Summary of the Invention
[0005] (1) Technical problems to be solved
[0006] In view of the deficiencies of the prior art, the present invention provides a multi-parameter collaborative temperature control system for an energy-saving industrial oven. By using high-precision sensors to collect load data inside the oven in real time, and cleaning and calibrating it with a Kalman filter to filter out environmental noise and drift, smooth and stable data is obtained; based on the support vector regression model to analyze the real-time load, combined with the radial basis function kernel to capture non-linear relationships, the temperature recovery time and expected energy consumption are predicted; the reinforcement learning algorithm is used to analyze the high-dimensional state space through the deep Q network, dynamically optimize the heating power and air flow speed, and output the best control parameters; the control module executes the adjusted settings, monitors the temperature and energy consumption in real time, and feeds back the data to the prediction model and learning strategy, and updates the parameters regularly to adapt to load and environmental changes. The adaptive ability and long-term stability are significantly improved, which is suitable for industrial scenarios with load changes, equipment aging or environmental fluctuations, thus solving the technical problems recorded in the background art.
[0007] (2) Technical solutions
[0008] To achieve the above objectives, the present invention is realized through the following technical solutions: A multi-parameter collaborative temperature control system for an energy-saving industrial oven, including, data collection is triggered by a production batch change or a regular monitoring period, and the original load data inside the oven is obtained relying on a high-precision weight sensor or a vision detection system Through a Kalman filter combined with a load mutation detection mechanism (switching the data source when ) to clean the noise and calibrate, generating filtered load data L fd ;
[0009] Based on the filtered load data Using a support vector regression model (SVR) combined with a pre-trained radial basis function (RBF) kernel function to calculate the similarity between the historical load L i and the current load , outputting the predicted temperature recovery time T r,t and the predicted energy consumption E t , providing forward-looking data support for optimizing dynamic temperature control and energy efficiency management;
[0010] Adjust upon receiving the filtered load data The predicted temperature recovery time T r,t and the current temperature Tpt 、Target temperature Tp tp 、Actual recovery time T r,al,t-1 and actual energy consumption E al,t-1 After that, according to the state space vector S by the Deep Q-Network (DQN) t Select the heating power P t and air flow velocity F t to achieve a dynamic balance between rapid temperature stabilization and energy consumption optimization;
[0011] Execute the output heating power P t and air flow velocity F t , combined with the continuously monitored actual temperature Tp t 、Actual energy consumption E al,t and real-time load L t Through the collaborative deviation index TI t (Integrating temperature deviation, energy consumption ratio, and load fluctuation) judge the response effect, and update the SVR model and DQN strategy when TI t >TI td to ensure system adaptability and continuous optimization;
[0012] Preferably, use a weight sensor to obtain the total load weight data, use a vision detection module to calculate the volume or product thickness of the load using image recognition and depth perception, and set the acquisition frequency dynamically according to the production rhythm, and output the unprocessed original load data L raue , representing the load at time t;
[0013] Preferably, apply a Kalman filter to process the collected original load data L rav to remove environmental noise and sensor drift; where: based on the load estimate value at the previous moment and the system dynamic model, predict the current load
[0014]
[0015] In the formula: A is the state transition matrix; B is the control input matrix; u t is the control input;
[0016] According to the predicted load and the current measured load Calculate the filtered load data
[0017]
[0018] In the formula: K tis the Kalman gain, H is the measurement matrix, the calibration process is based on a preset load-temperature response curve, correct the filtered data, and the output is the filtered and calibrated load data L t ;
[0019] Introduce an anomaly detection mechanism based on the Kalman filter:
[0020] Set the change rate threshold δ, at the current measured load and the filtered value at the previous moment difference; When it is determined that the load has mutated, the filter update is paused when the mutation occurs, and the current measured load is directly used as the filtered load data Subsequently, resume normal filtering;
[0021] Preferably, use the output real-time load data L t and historical data to train the model. The historical data includes the load L t , the corresponding temperature recovery time T r and the record of energy consumption E; Use support vector regression to construct prediction functions respectively:
[0022] T r,t = f(L t ): Predict the temperature recovery time; E t = g(L t ): Predict the energy consumption;
[0023] Support vector regression uses the radial basis function kernel to capture the non-linear relationship between the load and the response. The kernel function is defined as:
[0024] K(L i ,L j ) = exp(-γ||L i -L j || 2 )
[0025] In the formula: L i ,L j are the historical load data point and the current load data point, and γ is the kernel parameter;
[0026] Training process: Solve the support vector regression model through the dual optimization problem to minimize the prediction error and improve the generalization ability of the model, and output the trained support vector regression model.
[0027] Preferably, input the real-time collected and cleaned load data L t into the trained support vector regression model, where:
[0028] The trained support vector regression model is based on the input real-time load data Lt , calculate the similarity with historical data through the kernel function and output the prediction result: the radial kernel function K(L i ,L t ) generates the temperature recovery time T r,t and energy consumption E t ;
[0029] Preferably, define the industrial oven temperature control as the environment of reinforcement learning, and define the state space vector S t , including real-time load data L t , predicted temperature recovery time T r,t , predicted energy consumption E t , current temperature Tp t , target temperature Tp tp , the state vector is expressed as:
[0030] S t = [L t ,T r,t ,E t ,Tp t ,Tp tp
[0031] Define the action space vector A t : includes heating power P t , air flow velocity F t ; the action space vector A t is expressed as:
[0032] A t = [P t ,F t
[0033] Define the reward function R t :
[0034] R t = -(α·|Tp t - Tp tp | + β·E t )
[0035] In the formula: α is the temperature deviation weight.
[0036] Preferably, use a neural network to approximate the Q function, with the input being the state space vector S t , and the output being the expected return of each action space vector A t , where:
[0037] Q(S t ,A t ; θ) ≈ Q * (S t ,At )
[0038] where: θ is the neural network parameter, Q * (S t , A t ) is the theoretically optimal Q value;
[0039] In the state space vector S t , add the actual measurement value as feedback: add the actual temperature recovery time T r,al,t and the actual energy consumption E al,t , and update the state space vector to: S t = [L t , T r,t , E t , Tp t , Tp pt , T r,al,t-1 , E al,t-1 ;
[0040] Generate an experience tuple (S t , A t , R t , S t+1 ) by interacting with the environment, store it in the experience replay buffer, sample a small batch of data regularly, and minimize the following loss function to update θ:
[0041]
[0042] where: γ is the discount factor, with a value range of (0, 1), which balances immediate and future rewards, and θ - is the target network parameter, which is updated synchronously from θ regularly.
[0043] During training, use the ∈-greedy policy:
[0044]
[0045] During application, select the action with the largest Q value:
[0046]
[0047] Generate the optimal action A t = [P t , F t according to the state space vector S t , that is, the adjusted heating power and air flow rate;
[0048] Preferably, apply the heating power P t to the heating element to adjust the heat input of the oven, apply the air flow rate F t to the air flow control, and record the heating power P corresponding to the execution time tt and the air flow velocity F t values to form a time series data set;
[0049] Preferably, the monitoring parameters are collected at fixed time intervals, including the actual temperature Tp t , the actual energy consumption E al,t and the real-time load L t , and stored as a monitoring time series (t, Tp t , E al,t , L t ); generate a collaborative deviation index TI t . When this index exceeds the preset threshold, subsequent feedback optimization will be executed;
[0050] Set the deviation threshold TI td . If TI t ≤TI td , continue real-time monitoring without optimization; if TI t >TI td , trigger the feedback optimization prediction model and the reinforcement learning strategy;
[0051] Preferably, the collected monitoring time series (t, Tp t , E al,t , L t ) is associated with the predicted temperature recovery time T r,t and the expected energy consumption E t , as well as the heating power P t and the air flow velocity F t to construct an updated feedback data set; use the collected updated feedback data set to update the support vector regression model SVR; use the real-time load data L t as the input to calculate the actual temperature recovery time T r,al,t ;
[0052] Construct updated training sets (L t , T r,al,t ) and (L t , E al,t ), and retrain the support vector regression model SVR regularly to optimize its kernel function parameters and hyperparameters;
[0053] Use the updated reinforcement learning (RL) strategy with the updated feedback data set to add the state-action-reward-next state tuple (S t , A t , R al.t , S t+1 ) to the experience replay buffer,
[0054] Sample from the experience replay buffer regularly and update the parameters θ of the deep Q-network by minimizing the loss function. If |Tp t-Tp tp | or E al,t If it continues to be higher than expected, dynamically adjust the weight α of the temperature deviation and the weight β of the energy consumption.
[0055] Preferably,
[0056] (III) Beneficial effects
[0057] The present invention provides a multi-parameter collaborative temperature control system for an energy-saving industrial oven, which has the following beneficial effects:
[0058] By using high-precision sensors combined with a Kalman filter, the real-time and accuracy of the load data are ensured, providing a reliable basis for temperature control. The support vector regression (SVR) model can accurately predict the temperature recovery time and energy consumption, while the deep Q-network (DQN) of reinforcement learning dynamically adjusts the heating power and air flow speed to achieve the optimal balance between temperature and energy consumption.
[0059] The closed-loop feedback mechanism continuously optimizes the system strategy through real-time monitoring and performance evaluation to ensure long-term stability and self-adaptability. This technology integration significantly improves the temperature control accuracy and operation efficiency of the industrial oven.
[0060] Compared with the traditional PID or MPC control methods, it performs better in scenarios with frequent load changes, and can effectively reduce temperature fluctuations and energy consumption waste. Applying reinforcement learning to the dynamic load processing of industrial ovens and introducing the collaborative deviation index TELSDI comprehensively quantify the dynamic and static performance of the system, providing an objective evaluation of the operating state. By setting performance thresholds, the optimization mechanism is triggered to ensure timely adjustment when the performance deviates from the expected value to avoid efficiency decay during long-term operation.
[0061] Using the real-time collected data to update the SVR model and DQN parameters enables the prediction and control strategies to dynamically adapt to load changes, equipment aging, or environmental fluctuations, maintaining the high efficiency of the system. The closed-loop feedback mechanism significantly improves the system's adaptive ability and long-term stability through data-driven continuous learning, reducing the dependence on manual tuning.
[0062] The optimized model and strategy better match the current operating conditions, improving the temperature control accuracy and energy consumption efficiency, extending the service life of the equipment. The feedback process supports cross-scenario knowledge transfer, enhancing the universality of the system in different production environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 It is a schematic structural diagram of the multi-parameter collaborative temperature control system for the energy-saving industrial oven of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0064] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0065] Please refer to Figure 1 , the present invention provides a multi-parameter collaborative temperature control system for an energy-saving industrial oven, including
[0066] Step 1: Data collection is triggered by a production batch change or a regular monitoring cycle. Relying on a high-precision weight sensor or a vision detection system to obtain the original load data in the oven Through the Kalman filter combined with the load mutation detection mechanism (when the data source is switched), clean the noise and calibrate it to generate the filtered load data L fd ;
[0067] The above Step 1 includes the following contents:
[0068] Step 101: Real-time load data collection
[0069] Use a high-precision weight sensor or a vision detection module, such as a 3D laser scanner, to measure the current load in the oven, where:
[0070] The weight sensor obtains data by measuring the total weight of the load, which is suitable for scenarios with a fixed number of trays or regular load forms; the vision detection module uses image recognition and depth perception to calculate the volume or product thickness of the load, which is suitable for cases with complex or irregular load forms;
[0071] The acquisition frequency is dynamically set according to the production rhythm, for example, once per second, to ensure timely capture of load changes; output the unprocessed original load data L raue , whose value changes with time t, representing the load at time t;
[0072] By accurately measuring the load in the oven in real time Provide reliable basic data input for the follow-up, effectively filter out environmental noise (such as interference caused by temperature fluctuations or mechanical vibrations) and sensor drift, and improve the reliability and consistency of the load data ;
[0073] Step 102: Data cleaning and calibration
[0074] For the collected original load data L ravProcess using a Kalman Filter to remove environmental noise (such as vibrations or temperature fluctuations) and sensor drift; the Kalman Filter processing is divided into two stages: prediction and update, where:
[0075] Prediction stage: Based on the load estimate value at the previous moment and the system dynamic model, predict the current load
[0076]
[0077] In the formula: A is the state transition matrix, representing the natural change trend of the load over time (e.g., set to 1 assuming continuous load change);
[0078] B is the control input matrix, considering external disturbances (e.g., set to 0 with no control input); u t is the control input (ignored here);
[0079] Update stage: According to the predicted load and the current measured load Calculate the filtered load data
[0080]
[0081] In the formula: K t is the Kalman gain, dynamically adjusted over time to balance the weights of prediction and measurement, H is the measurement matrix (set to 1, indicating direct measurement of the load), and the calibration process is based on a preset load-temperature response curve to correct the filtered data to ensure its consistency with the actual load conditions. For example, if the sensor reading is too high, adjust the filtered load data according to the curve The output is the filtered and calibrated load data L t ;
[0082] Introduce an anomaly detection mechanism based on the Kalman Filter:
[0083] Set a change rate threshold δ, for the difference between the current measured load and the filtered value at the previous moment ; When it is, it is determined that there is a load mutation. When a mutation occurs, pause the filter update and directly use the current measured load as the filtered load data Subsequently, resume normal filtering; introduce a load mutation detection mechanism (triggered when ), which can quickly identify significant changes in the load and avoid system response delays caused by filtering delays, thus ensuring the real-time performance of temperature control;
[0084] The Kalman filter, through its dynamic prediction and update mechanism, can smooth the load data more efficiently compared to traditional static filtering methods (such as mean filtering). Significantly reduce data fluctuations and improve the stability of subsequent control. The filtered load data with high stability reduces the interference of input noise on the prediction model and control strategy; when the load changes rapidly, it can quickly switch to the direct measurement value to ensure immediate adaptability in dynamic scenarios.
[0085] It should be noted that: The system dynamic model refers to the dynamic relationship that describes the evolution of various parameters (such as temperature, load, heating power, air flow, etc.) in the industrial oven temperature control system through a mathematical form. Here, the load usually refers to the quantity or nature of the materials in the oven, which affects the heat absorption and the temperature response of the system. The core role of the system dynamic model is to capture the interactions between these parameters and predict the behavior of the system under specific conditions. This model can have various forms, depending on the system design and application requirements.
[0086] Step 2: Based on the filtered load data Use the support vector regression model (SVR) combined with the pre-trained radial basis function (RBF) kernel function to calculate the similarity between the historical load L i and the current load , and output the predicted temperature recovery time T r,t and the predicted energy consumption E t , providing forward-looking data support for optimizing dynamic temperature control and energy efficiency management;
[0087] The content of the above Step 2 includes the following:
[0088] Step 201: Construct a support vector regression model
[0089] Use the output real-time load data L t and historical data to train the model. The historical data includes the load quantity L t , the corresponding temperature recovery time T r and the records of energy consumption E; adopt the support vector regression (SVR) method to construct the prediction functions respectively:
[0090] T r,t = f(L t ): Predict the temperature recovery time; E t = g(L t ): Predict the energy consumption;
[0091] Support vector regression (SVR) uses the radial basis function (RBF) kernel to capture the non-linear relationship between the load and the response. The kernel function is defined as:
[0092] K(L i ,L j )=exp(-γ||L i -L j || 2 )
[0093] where: L i ,L j are historical load data points and current load data points, γ is the kernel parameter that controls the sensitivity of the model to local features, and its value range is (0, 1];
[0094] Training process: Solve the support vector regression SVR model through a dual optimization problem, and the prediction function is expressed as:
[0095]
[0096] where: α i , is the Lagrange multiplier of the temperature recovery time model, and its value range is [0, C]; β i , is the Lagrange multiplier of the energy consumption model, and its value range is [0, C]; b and d are bias terms, which are automatically solved through optimization, C is the regularization parameter that controls the balance between model complexity and error, and its value range is [1, 100];
[0097] Use grid search to optimize the kernel parameter γ and the regularization parameter C to minimize the prediction error and improve the model generalization ability, and output the trained support vector regression model SVR. Thus, according to the input real-time load data L t output the predicted temperature recovery time T r,t and the predicted energy consumption E t ;
[0098] The support vector regression model SVR uses the radial basis function (RBF) kernel to accurately capture the complex non-linear relationship between the real-time load data L t and the temperature recovery time T r,t , the predicted energy consumption E t , improving the accuracy and credibility of the prediction results; it performs excellently in the small sample data scenario, and by regularly updating model parameters (such as support vectors and weights), it adapts to factors such as equipment aging and production condition changes to ensure the long-term accuracy and practicality of the prediction results.
[0099] Step 202, Real-time prediction of temperature recovery time and energy consumption
[0100] Input the load data L that has been collected and cleaned in real time into the trained support vector regression model SVR, where: t
[0101] The trained support vector regression model SVR is based on the input real-time load data L t , calculate the similarity with historical data through the kernel function and output the prediction result:
[0102]
[0103] The meanings of the parameters are the same as those in step 201. The radial kernel function K(L i ,L t ) Through RBF kernel calculation, ensure that the prediction takes into account the nonlinear effect of the load and generates a real-time prediction of the temperature recovery time T r,t and energy consumption E t ;
[0104] By real-time prediction of temperature recovery time T r,t and energy consumption E t , can adjust the control strategy in advance and significantly reduce the delay of feedback control based on the current load L t Dynamic prediction can ensure rapid response to load changes and maintain high efficiency and stability of temperature control. The prediction results provide a reference for production plan optimization, such as reasonably arranging heating time, reducing ineffective energy consumption, and improving overall production efficiency.
[0105] Step 3: Adjust the load data after receiving the filter Predicted temperature recovery time T r,t 、Current temperature Tp t 、Target temperature Tp tp , actual recovery time T r,al,t-1 And actual energy consumption E al,t-1 Then, the deep Q network (DQN) is used to calculate the state space vector S t Select the heating power P t and air velocity F t , achieving a dynamic balance between rapid temperature stabilization and energy consumption optimization;
[0106] The step three includes the following contents:
[0107] Step 301: Construct reinforcement learning environment and state space vector
[0108] The temperature control of industrial ovens is defined as a reinforcement learning environment, where the state evolves over time and the control actions directly affect the state change. The state space vector S is defined as t , which consists of the following five variables:
[0109] Real-time load data L t , predict temperature recovery time T r,t , predict energy consumption E t 、Current temperature Tp t 、Target temperature Tptp , the state vector is represented as:
[0110] S t = [L t , T r,t , E t , Tp t , Tp tp
[0111] Define the action space vector A t : It contains two adjustable control variables: heating power P t : The adjustment range is [P min , P max , air flow velocity F t : The adjustment range is [F min , F max , and the action space vector A t is represented as:
[0112] A t = [P t , F t
[0113] Define the reward function R t : The reward function is designed as a weighted penalty term for temperature deviation and energy consumption. The specific form is:
[0114] R t = -(α·|Tp t - Tp tp | + β·E t )
[0115] In the formula: α is the temperature deviation weight, dimensionless, and the value range is [0, 1], reflecting the temperature control priority; β is the energy consumption weight, with the unit of kWh -1 , and the value range is [0, 1], reflecting the energy efficiency optimization priority;
[0116] By using the negative value form, the system is motivated to reduce the temperature deviation (reach the target temperature quickly) and lower the energy consumption, ensuring that the control strategy meets both accuracy and energy conservation requirements. It can avoid using simple statistics (such as standard deviation), but instead reflects the characteristics of multi-objective optimization through a weighted linear combination;
[0117] The state vector S t and the action space vector A t provide the input basis for step 302, and the reward function R t is used to evaluate the action effect;
[0118] Furthermore, the state space vector S t can be defined as: [L t , T r,t , Et , Tp t , Tp tp , T r,al,t-1 , E al,t-1 Integrate real-time measurement and prediction data to comprehensively reflect the system operation status and improve the scientificity and robustness of control strategies; where, T r,al,t-1 is the actual temperature recovery time at the previous time step (t-1), and E al,t-1 is the actual energy consumption at the previous time step (t-1);
[0119] Reward function R t = -(α·|Tp t - Tp tp | + β·E t ); By adjusting the weights α and β, flexibly balance the priorities of temperature accuracy and energy consumption optimization to meet the diverse needs of different production scenarios. The multi-dimensional design of the state space vector provides rich decision-making basis for reinforcement learning, avoids the limitations of single-variable control, and improves the system's ability to handle complex working conditions. Through the introduction of historical data (such as T r,al,t-1 , E al,t-1 ), enhance the algorithm's perception ability of operation trends and optimize the long-term control effect.
[0120] Step 302. Apply the reinforcement learning algorithm to optimize the control strategy
[0121] Adopt the Deep Q-Network (DQN) combining deep neural network and Q-learning, which is suitable for processing high-dimensional state space vector S t ;
[0122] Use the neural network to approximate the Q function, with the input being the state space vector S t , and the output being the expected return of each action space vector A t , where:
[0123] Q(S t , A t ; θ) ≈ Q * (S t , A t )
[0124] In the formula: θ is the neural network parameter, which is automatically updated through training, and Q * (S t , A t ) is the theoretically optimal Q value; The Q function evaluates the long-term return of each action, facilitating the selection of the best heating power P t and the air flow velocity F t ;
[0125] Furthermore, in the state space vector S tAdd the actual measurement value as feedback: Add the actual temperature recovery time T r,al,t and the actual energy consumption E al,t , and update the state space vector to: S t =[L t , T r,t , E t , Tp t , Tp pt , T r,al,t-1 , E al,t-1 , use the actual value at the previous moment to correct the prediction effect, reduce the prediction error, and improve the robustness of the control strategy;
[0126] Generate experience tuples (S t , A t , R t , S t+1 ) by interacting with the environment, store them in the experience replay buffer, sample a small batch of data regularly, update θ, and minimize the following loss function The loss function is based on temporal difference learning to ensure that the Q value gradually approaches the optimal strategy:
[0127]
[0128] In the formula: γ is the discount factor, and its value range is (0, 1), which balances immediate and future rewards, and θ - is the target network parameter, which is synchronously updated from θ regularly.
[0129] During training, use the ∈-greedy strategy:
[0130]
[0131] During application, select the action with the largest Q value:
[0132]
[0133] Generate the optimal action A t =[P t , F t according to the state space vector S t , that is, the adjusted heating power and air flow velocity;
[0134] The deep Q-network (DQN) can efficiently process the high-dimensional state space vector S t , and realize the heating power P t and the air flow velocity F tDynamic optimization ensures rapid temperature stabilization and reduced energy consumption. Through trial-and-error learning and experience replay mechanisms, the reinforcement learning algorithm continuously iteratively optimizes the control strategy, enhancing the system's adaptability and long-term performance under load changes. Dynamically adjusting control parameters avoids the deficiencies of traditional PID control in non-linear scenarios, significantly improving the accuracy and response speed of temperature control.
[0135] Step Four: Execute the output heating power P t and air flow velocity F t , combined with the continuously monitored actual temperature Tp t 、actual energy consumption E al,t and real-time load L t Through the collaborative deviation index TI t (integrating temperature deviation, energy consumption ratio, and load fluctuation) to judge the response effect, and update the SVR model and DQN strategy when TI t >TI td to ensure system adaptability and continuous optimization;
[0136] The above Step Four includes the following content:
[0137] Step 401: Apply the optimized heating and air flow settings
[0138] Receive the output heating power P t and air flow velocity F t as control instructions, apply the heating power P t to the heating element to precisely adjust the heat input of the oven, and apply the air flow velocity F t to the air flow control to optimize the air flow distribution inside the oven, and record the heating power P t and air flow velocity F t values corresponding to the execution time t to form a time series data set;
[0139] Apply the optimized heating power P t and air flow velocity F t in real time, so that the oven temperature Tp t quickly approaches the target temperature Tp tp , effectively reducing temperature fluctuations and overshoot phenomena. The optimized settings improve energy utilization efficiency, reduce unnecessary heating or air flow waste, and lower the energy consumption cost during the production process; The control effect that can respond quickly improves product quality consistency.
[0140] Step 402: Real-time monitoring and data collection
[0141] Collect monitoring parameters at fixed time intervals (such as every second), including the actual temperature Tp t 、actual energy consumption E al,t and real-time load L t, stored as a monitored time series (t, Tp t , E al,t , L t ); and generate a collaborative deviation index TI t , when this index exceeds a preset threshold, subsequent feedback optimization will be executed;
[0142] Among them, within the time interval [t - Δt, t], the collaborative deviation index TI t is defined as:
[0143]
[0144] In the formula: the deviation vector d t is a three-dimensional vector representing the temperature, energy consumption, and load deviations at time t:
[0145]
[0146] is the temperature deviation at the current time t, after being normalized;
[0147] is the ratio of the actual energy consumption E al,t to the expected energy consumption E ed ;
[0148] is the deviation of the actual load L t from the expected load L ed , after being normalized.
[0149] The non-linear transformation f(d t ) processes the deviation vector d t into a non-linear form to highlight the effects of different deviations:
[0150]
[0151] f1(d 1,t ) = exp(d 1,t ) - 1 is the exponential transformation of the temperature deviation; is the square transformation of the energy consumption deviation;
[0152] f3(d 3,t ) = log(1 + d 3,t ) is the logarithmic transformation of the load deviation;
[0153] The weight vector w is used to balance the relative importance of temperature, energy consumption, and load:
[0154] w1 + w2 + w3 = 1
[0155] w1 is the weight of the temperature deviation, w2 is the weight of the energy consumption deviation, and w3 is the weight of the load deviation;
[0156] The dynamic change term calculates the change rate of the temperature deviation through integration:
[0157]
[0158] Δt is the evaluation time window, is the instantaneous change rate of the temperature deviation at time τ; the adjustment parameter λ ≥ 0, with an initial value of 0.5, adjusted according to the impact of temperature fluctuations on product quality;
[0159] Set the deviation threshold TI td , representing the maximum allowable comprehensive deviation, and the threshold can be adjusted according to the actual application scenario;
[0160] If TI t ≤TI td , it indicates that the performance meets the expectation, continue real-time monitoring without optimization; if TI t >TI td , it indicates that the performance does not meet the expectation, trigger feedback to optimize the prediction model and the reinforcement learning strategy;
[0161] Continuously monitor the temperature Tp t the actual energy consumption E al.t and the load L t Provide high-precision real-time data for closed-loop optimization, ensure the controllability of the operating state, the high frequency and accuracy of data collection support the immediate evaluation of the control effect, identify potential deviations and correct them in a timely manner, improve stability, and multi-parameter monitoring provides a basis for fault diagnosis and early warning, such as detecting abnormal temperature fluctuations or sudden increases in energy consumption, reducing production risks. The collected data provides materials for model update and performance optimization, promoting the continuous improvement of the system.
[0162] The collaborative deviation index TI t Quantitatively comprehensively evaluate the dynamic and static performance of the system through the average value within the time window (such as the weighted combination of temperature deviation and energy consumption), and provide an objective evaluation of the operating state. By setting performance thresholds (such as TI t <I td ), trigger the optimization mechanism to ensure timely adjustment when the performance deviates from the expectation and avoid efficiency decay during long-term operation.
[0163] Step 403, Feedback to optimize the prediction model and the reinforcement learning strategy
[0164] The collected monitoring time series (t, Tp t , E al,t , L t ) and the predicted temperature recovery time T r,t and the expected energy consumption Et , and the heating power P t and the air flow velocity F t are associated to construct an updated feedback data set; use the collected updated feedback data set to update the support vector regression model SVR; with the real-time load data L t as the input, calculate the actual temperature recovery time T r,al,t , which is defined as the actual temperature Tp t recovering from the deviation state to the target temperature Tp tp of the time;
[0165] Construct updated training sets (L t , T r,al,t ) and (L t , E al,t ), and retrain the support vector regression model SVR regularly to optimize its kernel function parameters and hyperparameters;
[0166] Use the updated reinforcement learning (RL) policy with the updated feedback data set, where:
[0167] Calculate the actual reward R al.t , evaluate the control effect by weighted combination of temperature deviation and energy consumption, and a negative value indicates that the higher the deviation or energy consumption, the lower the reward. The formula is:
[0168] R al.t = -(α·|Tp t - Tp tp | + β·E al,t )
[0169] In the formula: α is the temperature deviation weight, which controls the priority of temperature accuracy, β is the energy consumption weight, which regulates the degree of emphasis on energy efficiency, and Tp tp is the target temperature, the preset ideal temperature value;
[0170] Add the state-action-reward-next state tuple (S t , A t , R al.t , S t+1 ) to the experience replay buffer, where:
[0171] S t = (L t , Tp t , E al,t ) is the current state, A t = (P t , F t ) is the execution action, and S t+1 is the next moment state;
[0172] Sample from the experience replay buffer regularly (e.g., every hour) and update the parameters θ of the deep Q-network (DQN) by minimizing the loss function, where:
[0173]
[0174] In the formula: γ is the discount factor, balancing immediate and long-term rewards; Q is the Q-value function, predicting the expected return of the state-action pair; θ - is the target network parameter, updated from θ regularly;
[0175] If |Tp t -Tp tp | or E al,t continues to be higher than expected, dynamically adjust the weights α of the temperature deviation and β of the energy consumption to optimize the balance between temperature accuracy and energy efficiency;
[0176] Update the SVR model and DQN parameters using real-time collected data, enabling the prediction and control strategies to dynamically adapt to load changes, equipment aging, or environmental fluctuations, maintaining the high efficiency of the system. The closed-loop feedback mechanism significantly enhances the system's adaptive ability and long-term stability through data-driven continuous learning, reducing the dependence on manual tuning. The optimized model and strategy better match the current operating conditions, improving the temperature control accuracy and energy consumption efficiency, extending the equipment service life. The feedback process supports cross-scenario knowledge transfer, enhancing the system's versatility in different production environments.
[0177] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0178] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0179] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only for some logical function divisions. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0180] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0181] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily think of changes or substitutions, which should all be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. Multi-parameter collaborative temperature control system for energy-saving industrial oven, characterized in that: including Real-time collect the load data inside the oven through high-precision sensors. After cleaning and calibration using a Kalman filter, filter out the influence of environmental noise and drift through a dynamic prediction and update mechanism to obtain smooth and stable load data. Analyze the real-time load data based on a support vector regression model to obtain the predicted temperature recovery time and expected energy consumption. Capture the non-linear relationship between the load and the response through a radial basis function kernel to generate a forward-looking prediction result. Adopt a reinforcement learning algorithm to process the real-time load and prediction information, dynamically optimize the heating power and air flow speed, analyze the high-dimensional state space through a deep Q network, and output the optimal control parameters in real time. Apply the adjusted heating and air flow settings to the control module, monitor the temperature and energy consumption data in real time to evaluate the operation effect, and feedback the monitored data to the prediction model and learning strategy. Continuously adapt to the load and environmental changes by regularly updating the optimization parameters.
2. The multi-parameter collaborative temperature control system of the energy-saving industrial oven according to claim 1, characterized in that Obtain the total load weight data, calculate the volume or product thickness of the load, and the acquisition frequency is dynamically set according to the production rhythm. Output the raw payload data L without processing raue , and process it using a Kalman filter to remove environmental noise and sensor drift.
3. The multi-parameter collaborative temperature control system of the energy-saving industrial oven according to claim 2, characterized in that Introduce an anomaly detection mechanism based on the Kalman filter Set the change rate threshold δ, and for the current measured load and the filtered value at the previous moment The difference is: When it is the case, it is determined that there is a load mutation. When the mutation occurs, the filter update is paused, and the current measured load is directly used as the filtered load data Subsequently, normal filtering is resumed.
4. The multi-parameter collaborative temperature control system of the energy-saving industrial oven according to claim 3, characterized in that Utilize the output real-time load data L t and historical data to train the model, where the historical data includes the load amount L t , the corresponding temperature recovery time T r and records of energy consumption E, and use support vector regression to construct prediction functions respectively; Support vector regression uses a radial basis function kernel to capture the non-linear relationship between the load and the response, solves the support vector regression model through a dual optimization problem, minimizes the prediction error and improves the generalization ability of the model, and outputs the trained support vector regression model.
5. The multi-parameter collaborative temperature control system of the energy-saving industrial oven according to claim 4, characterized in that Input the load data L after real-time acquisition and cleaning t into the trained support vector regression model, where: The trained support vector regression model, based on the input real-time load data L t , calculates the similarity with historical data through a kernel function and outputs the prediction results: the radial kernel function K(L i , L t ) calculates through the radial basis function kernel to generate the real-time predicted temperature recovery time T r,t and energy consumption E t .
6. The multi-parameter collaborative temperature control system of the energy-saving industrial oven according to claim 5, characterized in that Define the industrial oven temperature control as the environment of reinforcement learning and define the state space vector S t , including real-time load data L t , predicted temperature recovery time T r,t , predicted energy consumption E t , current temperature Tp t , target temperature Tp tp ; and define an action space vector A that includes the heating power P t , the air flow velocity F t ; and a reward function R that includes the temperature deviation weight t . t .
7. The multi-parameter collaborative temperature control system of the energy-saving industrial oven according to claim 6, characterized in that Use a neural network to approximate the Q function, with the input being the state space vector S t , and the output being the expected return of each action space vector A t . Generate experience tuples (S t , A t , R t , S t+1 ) by interacting with the environment, store them in the experience replay buffer, sample mini-batch data regularly, and minimize the following loss function Update θ: During training, the ε-greedy strategy is used, and when applying, the action space vector A with the largest Q value is selected t ; According to the state space vector S t Generate the optimal action space vector A t , that is, the adjusted heating power and air flow velocity 8. The multi-parameter collaborative temperature control system of the energy-saving industrial oven according to claim 7, characterized in that Apply the heating power P t to the heating element to adjust the heat input of the oven, and apply the air flow velocity F t to the air flow control; Collect monitoring parameters at fixed time intervals, including the actual temperature Tp t , the actual energy consumption E al,t and the real-time load L t , and store them as a monitoring time series (t, Tp t , E al,t , L t ). Generate a collaborative deviation index TI t from the monitoring time series (t, Tp al,t , E t , L t ; Set the deviation threshold TI td , if TI t ≤TI td , continue real-time monitoring without optimization; if TI t >TI td , trigger the feedback optimization prediction model and the reinforcement learning strategy.
9. The multi-parameter collaborative temperature control system of the energy-saving industrial oven according to claim 8, characterized in that Associate the collected monitoring time series (t, Tp t , E al,t , L t ) with the predicted temperature recovery time T r,t and the expected energy consumption E t , as well as the heating power P t and the air flow velocity F t to construct an updated feedback data set; Update the support vector regression model using the collected update feedback data set with the real-time load data L t as input, and calculate the actual temperature recovery time T r,al,t .
10. The multi-parameter collaborative temperature control system of the energy-saving industrial oven according to claim 9, characterized in that Construct an updated training set, retrain the support vector regression model regularly, and optimize its kernel function parameters and hyperparameters; use the updated reinforcement learning policy of the updated feedback data set, and add the tuple (S t , A t , R al.t , S t+1 ) to the experience replay buffer; Regularly sample from the experience replay buffer, update the parameters θ of the deep Q-network by minimizing the loss function. If |Tp t -Tp tp | or E al,t continues to be higher than expected, dynamically adjust the weights α of the temperature deviation and the weight β of the energy consumption.
Citation Information
Cited By
Industrial robot task cooperative scheduling method for intelligent factory
CN120791794A
Energy-saving ceramic tile firing roller kiln
CN120890263A
Grain quality dynamic evaluation management system based on big data analysis
CN120931152A