Energy and low-carbon cooperative regulation and control method and system based on reinforcement learning
By constructing a reinforcement learning model, real-time collection and optimization of multi-source data in industrial production are achieved, generating precise control strategies. This solves the problem of synergistic optimization of energy and carbon emissions in traditional control methods, and realizes efficient, low-cost, and green development in industrial production.
Patent Information
- Application Number
- CN202511257237.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-04
- Publication Date
- 2025-12-16
AI Technical Summary
In existing industrial production, it is difficult to achieve coordinated optimization of energy and carbon emission control. Traditional methods are inefficient and inaccurate, unable to adapt to complex and ever-changing production processes, and existing technologies are costly or unaffordable.
By collecting multi-source data, constructing a reinforcement learning model, monitoring and dynamically optimizing equipment operating parameters in real time, generating control strategies, and achieving coordinated control of energy and carbon emissions, the reinforcement learning model aims to generate precise control strategies with the weighted optimality of energy efficiency and carbon emission indicators.
It has improved energy efficiency and reduced carbon emissions in industrial production, solved the problems of lag and arbitrariness in traditional control methods, improved control efficiency and accuracy, and adapted to complex working conditions and dynamic changes.
Smart Images

Figure CN121145006A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of energy management and low-carbon control technology, specifically relating to a method and system for coordinated regulation of energy and low carbon based on reinforcement learning. Background Technology
[0002] In the industrial production sector, achieving coordinated and optimized control of energy and carbon emissions is crucial for the sustainable development of enterprises. With increasingly stringent environmental protection requirements and rising energy costs, industrial enterprises urgently need to accurately control energy consumption and carbon emissions to achieve cost reduction, efficiency improvement, and green development.
[0003] However, current industrial energy and carbon emission control faces numerous challenges. Traditional control methods rely on manual experience and simple statistical analysis, which are ill-suited to the complex and ever-changing industrial production processes. Manual operation is not only inefficient and inaccurate, but also unable to respond promptly to unexpected situations in production, leading to energy waste and excessive carbon emissions. For example, in cement production, adjusting equipment parameters based on manual experience cannot accurately match production demands, resulting in persistently high energy consumption. Simultaneously, existing control systems lack the ability to integrate and deeply analyze multi-source data. Different types of energy consumption data, carbon emission data, and equipment operating parameters are collected in a scattered manner with inconsistent formats, making it difficult to form effective data support; the analytical models are simplistic and unable to uncover the underlying patterns in the data, thus failing to provide a scientific basis for control decisions.
[0004] Existing improvement technologies have significant shortcomings. Some technologies optimize only for a single objective of energy or carbon emissions, neglecting the synergy between the two; some new technologies are too costly to apply, requiring large-scale equipment upgrades and maintenance by specialized technicians, which is unaffordable for small and medium-sized enterprises. Therefore, there is an urgent need for a method and system that is efficient, low-cost, and capable of achieving synergistic optimization and control of energy and carbon emissions. Summary of the Invention
[0005] The purpose of this invention is to propose a method and system for coordinated regulation of energy and low-carbon emissions based on reinforcement learning. By collecting and preprocessing multi-source data, a reinforcement learning model is constructed to generate regulation strategies. The effects are monitored in real time and the model is dynamically optimized to achieve coordinated regulation of energy and carbon emissions in industrial production, improve energy utilization efficiency, reduce carbon emissions, contribute to the green and sustainable development of industry, and meet the needs of enterprises for energy conservation, cost reduction, and environmental protection.
[0006] Therefore, the technical solution adopted by the present invention is as follows:
[0007] A reinforcement learning-based method for synergistic regulation of energy and low carbon emissions includes the following steps:
[0008] S1. Collect energy consumption data, carbon emission data and key operating parameters in the industrial production process in real time through sensors to build an initial dataset;
[0009] S2. Preprocess the initial dataset, including missing value imputation, outlier data removal, and dynamic weight adaptive normalization, to generate a standardized dataset.
[0010] S3. Based on the standardized dataset, construct a reinforcement learning optimization model, aiming at the weighted comprehensive optimality of energy efficiency indicators and carbon emission indicators, and dynamically generate control strategies for equipment operating parameters.
[0011] S4. The control strategy is sent to the production equipment for execution, and the changes in energy consumption and carbon emissions caused by the control strategy are monitored in real time using the sensors in step S1, and the control effect deviation is calculated; when the control effect deviation exceeds a preset threshold, the parameters of the reinforcement learning optimization model are updated.
[0012] Furthermore, the energy consumption data mentioned in step S1 includes the power of the kiln main motor and the coal powder flow rate; the carbon emission data is the CO2 emission amount;
[0013] The key operating parameters include the firing zone temperature, kiln rotation speed, clinker f-CaO content, clinker output, and induced draft fan opening.
[0014] Furthermore, the missing value processing in step S2 uses linear interpolation for filling. For data with no more than 3 consecutive missing sampling points, linear interpolation of the valid data before and after is used, as shown below:
[0015]
[0016] Where, x t x represents the missing value to be filled at time point t; prev x represents the valid observation value before the missing time point t; next For the valid observations at the next time point after the missing time point t; t prev For x prev The corresponding time point; t next For x next The corresponding time point;
[0017] The outlier removal method, which is based on statistics, includes the following steps: First, take one hour of data from the initial dataset and calculate the mean μ and standard deviation σ of each parameter in the initial dataset; then remove data points that exceed the range of μ±3σ; finally, mark the missing positions after removal as missing values and fill the missing values using the missing value filling method.
[0018] Furthermore, the dynamic weight adaptive normalization process is specifically formulated as follows:
[0019]
[0020] Where, x' i (t) is the parameter x i The normalized value of x is calculated at time t. i (t) is the value at time t before normalization; dynamic boundary L i (t) and U i (t) is defined as:
[0021]
[0022] Where τ is the sliding time window, with a default of 1 hour; min(x i ) [t-τ,t] For x i Minimum value within one hour; max(x) i ) [t-τ,t] For x i Maximum value in one hour; R i The maximum value is max(x) i ) [t-τ,t] and minimum value min(x) i ) [t-τ,t] The difference.
[0023] Furthermore, the reinforcement learning described in step S3 includes the following steps:
[0024] 1) State space s t Based on the standardized dataset, construct the state vector:
[0025] s t =[P' kiln,t ,Q' coal,t ,C' CO2,t ,T' t ,,R' rpm ,f' CaO,t ,V' IDfan,t ]
[0026] Where t is a time point; P ' kiln,t Q is the power of the main motor of the kiln; ' coal,t C' is the pulverized coal flow rate; CO2,t T' represents the CO2 emissions. t R' is the firing temperature of the zone. rpm f' is the rotational speed of the kiln cylinder; CaO,t V' is the f-CaO content of the clinker; IDfan,t The opening degree of the induced draft fan;
[0027] 2) Action space a t Action space includes specific control strategies.
[0028] a t =[ΔQ coal ,ΔV IDfan ,ΔT set ]
[0029] Where, ΔQ coal The pulverized coal flow rate adjustment amount; ΔV IDfan The adjustment amount of the induced draft fan opening t; ΔT set The adjustment amount for the setpoint of the firing zone temperature is within the following range:
[0030]
[0031] 3) Reward function r t :
[0032] r t =0.6×η EE,t -0.4×ξ CE,t -0.1×||a t || 2
[0033] Among them, the energy efficiency index η EE,t It is expressed as follows:
[0034]
[0035] Among them, C Clinker The clinker production rate; carbon emission index ξ CE,t It is expressed as follows:
[0036]
[0037] Action penalty item || a t || 2 It is expressed as follows:
[0038] ||a t || 2 =(ΔQ) coal ) 2 +(ΔV IDfan ) 2 +(ΔT set ) 2
[0039] 4) Policy function π θ (a t |s t The update formula for the network parameter θ in the policy function is expressed as follows:
[0040]
[0041] Where, θt θ represents the parameter value of the policy network at time step t. t+1 This represents the updated policy network parameter values at time step t+1; α is the learning rate. Here is the gradient operator with respect to θ; π θ (a t |s t )and π θ (a t |s t ) is the current policy network in state s t Take action a t The probability of; It is the policy network θ before the update. old In state s t Take action a t The probability; clip is the clipping function, which will... The value of is restricted to the interval [1-∈, 1+∈]; ∈ is a hyperparameter used to limit the magnitude of policy updates, preventing excessively large policy updates from causing model instability; A t The dominance function is expressed as follows:
[0042]
[0043] Where T is the total length of the entire time series, t is the current time step; γ is the discount factor, ranging from [0,1], which is used to measure the importance of future rewards; λ is used to control the contribution weight of different time steps in the multi-step guided algorithm; δ t The calculation formula is δ t+k =r t+k +γV(s t+k+1 )-V(s t+k ), where r t+k Action a is executed at time step t+k. t+k The immediate reward obtained afterward; V(s) t+k ) is state s t+k The value function; V(s) t+k+1 ) is state s t+k+1 The value function, γV(s) t+k+1 This represents a discounted estimate of the value of the next state.
[0044] 5) Industrial constraint handling: Setting safe ranges for pulverized coal flow and process limit thresholds for temperature.
[0045]
[0046] Among them, Q coal,t Let ΔQ be the pulverized coal flow rate at time point t. coal T is the amount of coal powder flow rate adjustment;set,t The value of the firing zone temperature at time point t; ΔT set The adjustment amount for the setpoint of the firing zone temperature.
[0047] Furthermore, the control strategy mentioned in step S3 is the action in the reinforcement learning action space, including the coal powder flow rate adjustment amount, the induced draft fan opening adjustment amount, and the calcination zone temperature setpoint.
[0048] The coal powder flow rate is controlled by adjusting the belt speed through the frequency converter of the coal feeder;
[0049] The kiln tail induced draft fan controls the air volume by adjusting the output frequency of the frequency converter;
[0050] The setpoint for the firing zone temperature is written into the firing zone temperature controller, and the new setpoint is tracked by a PID algorithm.
[0051] Furthermore, the deviation in the regulation effect includes deviations in energy consumption performance and carbon emission performance, taking into account the deviation in energy consumption performance δ. η and carbon emission implementation deviation δ CE Construct a comprehensive deviation score, the formula is:
[0052] Score = 0.7δ η +0.3δ CE
[0053] When Score > 0.15, the parameter update of the reinforcement learning optimization model is triggered, that is, steps S1, S2 and S3 are re-executed.
[0054] A reinforcement learning-based system for the coordinated optimization and control of industrial energy and carbon emissions includes:
[0055] Data acquisition module: Collects energy consumption data, carbon emission data and key operating parameters in the industrial production process in real time through sensors to build an initial dataset;
[0056] Data preprocessing module: preprocesses the initial dataset, including missing value imputation, outlier removal, and dynamic weight adaptive normalization, to generate a standardized dataset;
[0057] Reinforcement learning optimization module: Based on the standardized dataset, a reinforcement learning optimization model is constructed, aiming at the weighted comprehensive optimality of energy efficiency indicators and carbon emission indicators, and dynamically generating control strategies for equipment operating parameters;
[0058] Control execution and feedback module: The control strategy is sent to the production equipment for execution, and the changes in energy consumption and carbon emissions caused by the control strategy are monitored in real time. The control effect deviation is calculated. When the control effect deviation exceeds a preset threshold, the parameters of the reinforcement learning optimization model are updated.
[0059] Compared with the prior art, the advantages of the present invention are as follows:
[0060] This invention deeply integrates reinforcement learning with industrial production to construct a closed-loop system encompassing data acquisition, model optimization, and precise control, overcoming the limitations of traditional manual experience-based control and single-objective optimization. Compared to traditional methods, this invention can collect multi-source data such as energy consumption and carbon emissions in real time. After dynamic normalization, it utilizes a reinforcement learning model to generate precise control strategies with a weighted optimal balance between energy efficiency and carbon emission indicators. This avoids the lag and arbitrariness of manual operations, improving control efficiency and accuracy.
[0061] This invention dynamically updates model parameters based on the control effect, achieving adaptive optimization of the strategy, which can effectively cope with the complex operating conditions and dynamic changes in industrial production. In contrast, traditional technologies mostly rely on fixed rules or simple analysis, making it difficult to simultaneously optimize energy and carbon emissions. Attached Figure Description
[0062] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0063] Figure 1 This is a flowchart of the method of the present invention;
[0064] Figure 2 This is a flowchart of step S2 of the present invention;
[0065] Figure 3 This is a flowchart of step S3 of the present invention. Detailed Implementation
[0066] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present application.
[0067] Example 1: Please refer to Figure 1As shown, this embodiment is a reinforcement learning-based method for synergistic regulation of energy and low carbon emissions, applied to a novel dry-process cement production line, and includes the following steps:
[0068] S1. Collect energy consumption data, carbon emission data and key operating parameters in the industrial production process in real time through sensors to build an initial dataset;
[0069] In this embodiment, the energy consumption data includes the kiln main motor power: the real-time power of the kiln drive motor is monitored using a smart meter; and the pulverized coal flow rate: the pulverized coal conveying rate is accurately measured using an Emerson Micro Motion Coriolis mass flow meter.
[0070] The carbon emission data includes CO2 emissions: the volume concentration of CO2 in the flue gas at the kiln tail is continuously monitored using a Siemens ULTRAMAT 23 infrared gas analyzer, and the instantaneous CO2 emissions are calculated by combining the flue gas flow meter data.
[0071] The key operating parameters include: firing zone temperature: a FLIR A615 infrared thermometer is installed at the firing zone of the kiln body 15m from the kiln head, using a dual-probe redundancy design to avoid interference from kiln skin ring formation; kiln body rotation speed: the kiln drive shaft rotation speed is monitored by a HEIDENHAIN ECN413 high-precision encoder, this parameter directly affects material residence time and combustion efficiency; clinker f-CaO content: the free calcium oxide content on the clinker conveyor belt is detected using a Malvern PANalytical online X-ray fluorescence analyzer, this parameter is a core indicator for evaluating clinker sintering quality; clinker output: measured by a belt scale; induced draft fan opening: the valve opening is monitored in real time by a potentiometer installed on the induced draft fan actuator.
[0072] S2. Preprocess the initial dataset, including missing value imputation, outlier removal, and dynamic weight adaptive normalization, to generate a standardized dataset, as referenced. Figure 2 As shown;
[0073] In this implementation, the missing value handling uses linear interpolation to fill in the missing values. For data with no more than 3 consecutive missing sampling points, linear interpolation of the valid data before and after is used, as shown below:
[0074]
[0075] Where, x t x represents the missing value to be filled at time point t; prev x represents the valid observation value before the missing time point t; next For the valid observations at the next time point after the missing time point t; t prev For x prev The corresponding time point; tnext For x next The corresponding time point;
[0076] The outlier removal method based on statistics includes the following steps: First, take one hour of data from the initial dataset and calculate the mean μ and standard deviation σ of each parameter in the initial dataset; then remove data points that exceed the range of μ±3σ; finally, mark the missing positions after removal as missing values and fill the missing values using the missing value filling method.
[0077] The dynamic weight adaptive normalization process is specifically formulated as follows:
[0078]
[0079] Where, x' i (t) is the parameter x i The normalized value of x is calculated at time t. i (t) is the value at time t before normalization; dynamic boundary L i (t) and U i (t) is defined as:
[0080]
[0081] Where τ is the sliding time window, with a default of 1 hour; min(x i ) [t-τ,t] For x i Minimum value within one hour; max(x) i ) [t-τ,t] For x i Maximum value in one hour; R i The maximum value is max(x) i ) [t-τ,t] and minimum value min(x) i ) [t-τ,t] The difference;
[0082] The standardized dataset consists of standardized data with values ranging from [-1, 1], while retaining a mapping table between the original values and normalized parameters for reverse conversion; at the same time, when a kiln shutdown signal is detected, the normalized values of all parameters are forcibly set to zero.
[0083] S3. Construct a reinforcement learning optimization model based on the standardized dataset, aiming at the weighted optimal combination of energy efficiency and carbon emission indicators, and dynamically generate control strategies for equipment operating parameters, referencing... Figure 3 As shown;
[0084] In this implementation, the reinforcement learning includes the following steps:
[0085] 1) State space s tBased on the standardized dataset, construct the state vector:
[0086] s t =[P' kiln,t ,Q' coal,t ,C' CO2,t ,T' t ,,R' rpm ,f' CaO,t ,V' IDfan,t ]
[0087] Where t is a time point; P ' kiln,t Q is the power of the main motor of the kiln; ' coal,t C' is the pulverized coal flow rate; CO2,t T' represents the CO2 emissions. t R' is the firing temperature of the zone. rpm f' is the rotational speed of the kiln cylinder; CaO,t V' is the f-CaO content of the clinker; IDfan,t The opening degree of the induced draft fan;
[0088] 2) Action space a t Action space includes specific control strategies.
[0089] a t =[ΔQ coal ,ΔV IDfan ,ΔT set ]
[0090] Where, ΔQ coal The pulverized coal flow rate adjustment amount; ΔV IDfan The adjustment amount of the induced draft fan opening t; ΔT set The adjustment amount for the setpoint of the firing zone temperature is within the following range:
[0091]
[0092] 3) Reward function r t :
[0093] r t =0.6×η EE,t -0.4×ξ CE,t -0.1×||a t || 2
[0094] Among them, the energy efficiency index η EE,t It is expressed as follows:
[0095]
[0096] Among them, CClinker The clinker production rate; carbon emission index ξ CE,t It is expressed as follows:
[0097]
[0098] Action penalty item || a t || 2 It is expressed as follows:
[0099] ||a t || 2 =(ΔQ) coal ) 2 +(ΔV IDfan ) 2 +(ΔT set ) 2
[0100] 4) Policy function τ θ (a t |s t The update formula for the network parameter θ in the policy function is expressed as follows:
[0101]
[0102] Where, θ t θ represents the parameter value of the policy network at time step t. t+1 This represents the updated policy network parameter values at time step t+1; α is the learning rate. Here is the gradient operator with respect to θ; π θ (a t |s t )and π θ (a t |s t ) is the current policy network (parameter θ) in state s t Take action a t The probability of; It is the policy network before the update (parameter is θ) old In state s t Take action a t The probability; clip is the clipping function, which will... The value of is restricted to the interval [1-∈, 1+∈]; ∈ is a hyperparameter used to limit the magnitude of policy updates, preventing excessively large policy updates from causing model instability; A t The advantage function is used to measure the advantage in state s. t Take action a t Compared to the advantage of the average strategy, a larger advantage function value indicates that the action brings greater benefits compared to the average strategy, as shown below:
[0103]
[0104] Where T is the total length of the entire time series, i.e., one hour, and t is the current time step; γ is the discount factor, ranging from [0,1], which is used to measure the importance of future rewards; λ is used to control the contribution weight of different time steps in the multi-step guided algorithm; δ t The calculation formula is δ t+k =r t+k +γV(s t+k+1 )-V(s t+k ), where r t+k Action a is executed at time step t+k. t+k The immediate reward obtained afterward; V(s) t+k ) is state s t+k The value function represents the value of state s. t+k The expected cumulative reward that can be obtained by starting with the current strategy; V(s) t+k+1 ) is state s t+k+1 The value function, γV(s) t+k+1 This represents a discounted estimate of the value of the next state.
[0105] 5) Industrial constraint handling: Setting safe ranges for pulverized coal flow and process limit thresholds for temperature.
[0106]
[0107] Among them, Q coal,t Let ΔQ be the pulverized coal flow rate at time point t. coal T is the amount of coal powder flow rate adjustment; set,t The value of the firing zone temperature at time point t; ΔT set The adjustment amount for the setpoint of the firing zone temperature;
[0108] Based on the construction of the state space, action space, reward function, and policy function, the training process of the reinforcement learning model needs to be realized through a closed loop of data acquisition, policy evaluation, and parameter update. The specific steps are as follows:
[0109] 1) Initialization: Construct the state space S, action space A, reward function R, value function V, and policy π;
[0110] 2) Time step t operation: Get the current state s t Choose action a based on the current strategy π t That is, a t ~π(s) t ); Perform action a t , and obtain the new state s t+1 and reward R(s) t ,at ).
[0111] 3) Policy function update: Update θ by executing the update formula for network parameter θ described in the policy function;
[0112] 4) Training termination condition: Repeat the above steps until convergence.
[0113] S4. The control strategy is sent to the production equipment for execution, and the changes in energy consumption and carbon emissions caused by the control strategy are monitored in real time using the sensors in step S1, and the control effect deviation is calculated; when the control effect deviation exceeds a preset threshold, the parameters of the reinforcement learning optimization model are updated.
[0114] In this embodiment, the standardized actions output by reinforcement learning in step S3 need to be converted into instructions that can be executed by actual equipment; taking pulverized coal flow adjustment as an example, the instruction conversion (DCS side) is implemented through the following pseudocode:
[0115]
[0116] The converted instructions will be executed at the equipment level. Specifically, the pulverized coal flow rate is controlled by adjusting the belt speed via the coal feeder's frequency converter, with a control accuracy of ±0.1 t / h. The kiln tail induced draft fan's air volume is controlled by adjusting the frequency converter's output frequency, with each 1% opening corresponding to 500 m³ / h. 3 / h air volume; the firing zone temperature setpoint is written into the firing zone temperature controller, and the new setpoint is tracked by the PID algorithm;
[0117] In this implementation, the deviation in the control effect includes the deviation in energy consumption performance and the deviation in carbon emission performance, wherein the deviation in energy consumption performance is δ η It is expressed as follows:
[0118]
[0119] Where, η target η represents the target energy consumption value. actual The actual energy consumption calculation method is the energy index calculation formula described above;
[0120] The carbon emission performance deviation δ CE It is expressed as follows:
[0121]
[0122] Among them, C CO2,actual The actual CO2 emissions are measured by the sensor described in step S1; C CO2,target Target carbon emissions;
[0123] Taking into account both energy consumption and carbon emission performance deviations, a comprehensive deviation score is constructed, with the following formula:
[0124] Score = 0.7δ η +0.3δ CE
[0125] When Score > 0.15, the parameter update of the reinforcement learning optimization model is triggered, that is, steps S1, S2 and S3 are re-executed; the parameter update of the reinforcement learning optimization model includes the following steps:
[0126] First, based on the energy consumption data, carbon emission data, and key operating parameters collected in real time by the sensors, an initial dataset is reconstructed; then, preprocessing is performed to generate the standardized dataset; finally, based on the standardized dataset, the state space s is reconstructed. t Then, according to the formula for updating the network parameter θ of the policy function described in step S3, θ is updated, and then the policy function π is updated according to the updated network parameter θ. θ (a t |s t ), where a t The action space is defined; finally, based on the updated reinforcement model parameters, the control strategy is output, and the comprehensive deviation score is further verified to ensure that the value is less than 0.15. If it is not satisfied, the above process is repeated.
[0127] This implementation also includes a reinforcement learning-based industrial energy and carbon emission synergistic optimization and control system, comprising:
[0128] Data acquisition module: Collects energy consumption data, carbon emission data and key operating parameters in the industrial production process in real time through sensors to build an initial dataset;
[0129] Data preprocessing module: preprocesses the initial dataset, including missing value imputation, outlier removal, and dynamic weight adaptive normalization, to generate a standardized dataset;
[0130] Reinforcement learning optimization module: Based on the standardized dataset, a reinforcement learning optimization model is constructed, aiming at the weighted comprehensive optimality of energy efficiency indicators and carbon emission indicators, and dynamically generating control strategies for equipment operating parameters;
[0131] Control execution and feedback module: The control strategy is sent to the production equipment for execution, and the changes in energy consumption and carbon emissions caused by the control strategy are monitored in real time, and the control effect deviation is calculated; when the control effect deviation exceeds a preset threshold, the parameters of the reinforcement learning optimization model are updated.
[0132] All the above formulas are dimensionless and use only numerical values for calculation. These formulas are derived from a large amount of data and through software simulation, aiming to approximate reality as closely as possible. The preset parameters in the formulas can be adjusted by those skilled in the art according to specific needs.
[0133] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0134] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.
Claims
1. A method for coordinated regulation of energy and low-carbon processes based on reinforcement learning, characterized in that, Includes the following steps: S1. Collect energy consumption data, carbon emission data and key operating parameters in the industrial production process in real time through sensors to build an initial dataset; S2. Preprocess the initial dataset, including missing value imputation, outlier data removal, and dynamic weight adaptive normalization, to generate a standardized dataset. S3. Based on the standardized dataset, construct a reinforcement learning optimization model, aiming at the weighted comprehensive optimality of energy efficiency indicators and carbon emission indicators, and dynamically generate control strategies for equipment operating parameters. S4. The control strategy is sent to the production equipment for execution, and the changes in energy consumption and carbon emissions caused by the control strategy are monitored in real time using the sensors in step S1, and the control effect deviation is calculated; when the control effect deviation exceeds a preset threshold, the parameters of the reinforcement learning optimization model are updated.
2. The method according to claim 1, characterized in that, The energy consumption data mentioned in step S1 includes the power of the kiln main motor and the flow rate of pulverized coal; the carbon emission data is the CO2 emission amount; The key operating parameters include the firing zone temperature, kiln rotation speed, clinker f-CaO content, clinker output, and induced draft fan opening.
3. The method according to claim 1, characterized in that, The missing value handling in step S2 uses linear interpolation to fill in the missing values, as shown below: Where, x t x represents the missing value to be filled at time point t; prev x represents the valid observation value before the missing time point t; next For the valid observations at the next time point after the missing time point t; t prev For x prev The corresponding time point; t next For x next The corresponding time point; The outlier removal method, which is based on statistics, includes the following steps: First, take one hour of data from the initial dataset and calculate the mean μ and standard deviation σ of each parameter in the initial dataset; then remove data points that exceed the range of μ±3σ; finally, mark the missing positions after removal as missing values and fill the missing values using the missing value filling method.
4. The method according to claim 3, characterized in that, The dynamic weight adaptive normalization process is specifically formulated as follows: Where, x' i (t) is the parameter x i The normalized value of x is calculated at time t. i (t) is the value at time t before normalization; dynamic boundary L i (t) and U i (t) is defined as: Where τ is the sliding time window, with a default of 1 hour; min(x i ) [t-τ,t] For x i Minimum value within one hour; max(x) i ) [t-τ,t] For x i Maximum value in one hour; R i The maximum value is max(x) i ) [t-τ,t] and minimum value min(x) i ) [t-τ,t] The difference.
5. The method according to claim 1, characterized in that, The reinforcement learning described in step S3 includes the following steps: 1) State space s t Based on the standardized dataset, construct the state vector: s t =[P' kiln,t ,Q' coal,t ,C' CO2,t ,T' t ,,R' rpm ,f' CaO,t ,V' IDfan,t ] Where t is the time point; P' kiln,t Q' is the power of the kiln's main motor; coal,t C' is the pulverized coal flow rate; CO2,t T' represents the CO2 emissions. t R' is the firing temperature of the zone. rpm f' is the rotational speed of the kiln cylinder; CaO,t V' is the f-CaO content of the clinker; IDfan,t The opening degree of the induced draft fan; 2) Action space a t Action space includes specific control strategies. a t =[ΔQ coal ,ΔV IDfan ,ΔT set ] Where, ΔQ coal The pulverized coal flow rate adjustment amount; ΔV IDfan The adjustment amount of the induced draft fan opening t; ΔT set The adjustment amount for the setpoint of the firing zone temperature is within the following range: 3) Reward function r t : r t =0.6×η EE,t -0.4×ξ CE,t -0.1×||a t || 2 Among them, the energy efficiency index η EE,t It is expressed as follows: Among them, C Clinker The clinker production rate; carbon emission index ξ CE,t It is expressed as follows: Action penalty item || a t || 2 It is expressed as follows: ||a t || 2 =(ΔQ coal ) 2 +(ΔV IDfan ) 2 +(ΔT set ) 2 4) Policy function π θ (a t |s t The update formula for the network parameter θ in the policy function is expressed as follows: Where, θ t θ represents the parameter value of the policy network at time step t. t+1 This represents the updated policy network parameter values at time step t+1; α is the learning rate. Here is the gradient operator with respect to θ; π θ (a t |s t )and π θ (a t |s t ) is the current policy network in state s t Take action a t The probability of; It is the policy network θ before the update. old In state s t Take action a t The probability; clip is the clipping function, which will... The value of is restricted to the interval [1-∈, 1+∈]; ∈ is a hyperparameter used to limit the magnitude of policy updates, preventing excessively large policy updates from causing model instability; A t The dominance function is expressed as follows: Where T is the total length of the entire time series, t is the current time step; γ is the discount factor, ranging from [0,1], used to measure the importance of future rewards; λ is used to control the contribution weight of different time steps in the multi-step guided algorithm; δ t The calculation formula is δ t+k =r t+k +γV(s t+k+1 )-V(s t+k ), where r t+k Action a is executed at time step t+k. t+k The immediate reward obtained afterward; V(s) t+k ) is state S t+k The value function; V(s) t+k+1 ) is state s t+k+1 The value function, γV(s) t+k+1 This represents a discounted estimate of the value of the next state. 5) Industrial constraint handling: Setting safe ranges for pulverized coal flow and process limit thresholds for temperature. Among them, Q coal,t Let ΔQ be the pulverized coal flow rate at time point t. coal The pulverized coal flow rate adjustment amount; T set,t The value of the firing zone temperature at time point t; ΔT set The adjustment amount for the setpoint of the firing zone temperature.
6. The method according to claim 1, characterized in that, The control strategy mentioned in step S3 is the action in the reinforcement learning action space, including the coal powder flow rate adjustment amount, the induced draft fan opening adjustment amount, and the calcination zone temperature setpoint. The coal powder flow rate is controlled by adjusting the belt speed through the frequency converter of the coal feeder; The kiln tail induced draft fan controls the air volume by adjusting the output frequency of the frequency converter; The setpoint for the firing zone temperature is written into the firing zone temperature controller, and the new setpoint is tracked by a PID algorithm.
7. The method according to claim 6, characterized in that, The deviation in the control effect includes deviations in energy consumption performance and carbon emission performance, taking into account the deviation in energy consumption performance δ. η and carbon emission implementation deviation δ CE Construct a comprehensive deviation score, the formula is: Score=0.7δ η +0.3δ CE When Score > 0.15, the parameter update of the reinforcement learning optimization model is triggered, that is, steps S1, S2 and S3 are re-executed.
8. A reinforcement learning-based system for the coordinated optimization and control of industrial energy and carbon emissions, characterized in that, include: Data acquisition module; An initial dataset is constructed by collecting energy consumption data, carbon emission data, and key operating parameters in the industrial production process in real time using sensors. Data preprocessing module: preprocesses the initial dataset, including missing value imputation, outlier removal, and dynamic weight adaptive normalization, to generate a standardized dataset; Reinforcement learning optimization module: Based on the standardized dataset, a reinforcement learning optimization model is constructed, aiming at the weighted comprehensive optimality of energy efficiency indicators and carbon emission indicators, and dynamically generating control strategies for equipment operating parameters; Control execution and feedback module: The control strategy is sent to the production equipment for execution, and the changes in energy consumption and carbon emissions caused by the control strategy are monitored in real time. The control effect deviation is calculated. When the control effect deviation exceeds a preset threshold, the parameters of the reinforcement learning optimization model are updated.
Citation Information
Cited By
Multi-objective collaborative optimization control method of environmental protection equipment and related device
CN122151568A
A multi-objective collaborative optimization control method for an environmental protection device and a related device
CN122151568B