Barrel storage and coal blending integrated intelligent scheduling method based on reinforcement learning
By adopting a reinforcement learning-based intelligent scheduling method for integrated coal storage and distribution in silos, the problems of unstable spray flow and high energy consumption in coal-fired power plants have been solved. The optimization of spray flow and discharge cycle time has been achieved, thereby improving production efficiency and economy.
Patent Information
- Application Number
- CN202511745762.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-26
- Publication Date
- 2026-02-24
AI Technical Summary
Existing coal-fired power plants and coal preparation and storage systems suffer from problems such as unstable spray flow, high energy consumption, and increased costs in dust control and cycle scheduling. They cannot effectively manage the spray system and discharge cycle, leading to repeated adjustments of overspraying or conservative speed reduction, which affects production efficiency and economy.
An intelligent scheduling method for integrated coal storage and distribution in silos based on reinforcement learning is adopted. By constructing a dust concentration prediction model and a cost-based elasticity index, and combining the spray system parameters, reinforcement learning state vectors are integrated to optimize the spray flow rate and discharge cycle time. Variable frequency drive commands for the spray pump and belt conveyor are generated to achieve joint optimization of spray flow rate and discharge cycle time.
Under the constraint of dust limits, the system achieves online identification of minimum spray flow rate and incorporates pump power cost into decision-making, reduces overspray and reciprocating speed reduction adjustments, stabilizes material throughput efficiency, reduces pump electricity and water costs, and improves the intelligence and economy of silo coal storage and distribution scheduling.
Smart Images

Figure CN121553724A_ABST
Abstract
Description
Technical Field field
[0001] This invention relates to the field of reinforcement learning technology, and more specifically, to a reinforcement learning-based intelligent scheduling method for integrated coal storage and distribution in silos. Background Technology
[0002] Coal-fired power plants and coal preparation and storage systems commonly employ silo storage and multi-section belt conveyor technology. Transfer points are equipped with spray dust suppression devices and dust monitoring devices, and dust concentration control thresholds must be strictly met during production. Increasing the discharge cycle time leads to a rise in dust source intensity due to increased material drop and induced airflow. Conventional solutions include increasing spray flow rate or reducing belt speed. However, increasing spray volume raises the surface area and total moisture content of the coal, reducing bulk flowability and causing problems such as arching, rodent burrowing, and adhesion, further limiting the discharge cycle time and potentially requiring unconventional unblocking methods.
[0003] In actual operation, cross-belt or point-based monitoring suffers from time delays in the transmission of sampling and aggregation signals and actuator responses. The moisture content and particle size of different coal sources vary randomly, and nozzle wear, pump efficiency fluctuations, and pipeline pressure differentials also alter the effective intensity of a unit spray flow rate. These factors combined mean that the minimum spray flow rate required to maintain a certain discharge cycle within the control threshold is not a static function, but rather drifts rapidly with time and operating conditions. Sudden increases in water consumption and pump power occur in certain cycle intervals, leading managers to rely on experience to reserve a large safety margin, resulting in over-spraying or conservative speed reduction.
[0004] Existing scheduling methods primarily manage the spray belt and feeding operations separately using rules or segmented control. This fails to integrate the costs of pumping electricity and water due to dust limits constraining the spray system's pressure differential and efficiency, as well as the limitations on cycle time caused by the deterioration in flowability due to spraying, into a unified decision-making process. When the load increases or changes in electricity, water, and environmental assessments necessitate a rapid increase in the discharge cycle time, the inability to anticipate sudden increases in water consumption and pump power leads to repeated adjustments of acceleration followed by overspraying and then deceleration. This increases energy and water costs, reduces effective throughput and equipment availability, and creates significant technical problems. Summary of the Invention
[0005] This invention provides an integrated intelligent scheduling method for coal storage and distribution in silos based on reinforcement learning, which solves the technical problems mentioned in the background.
[0006] This invention provides an intelligent scheduling method for integrated coal storage and distribution in silos based on reinforcement learning, comprising the following steps: Step S101: Collect state parameters to construct a dust concentration prediction model, solve the minimum spray flow rate corresponding to the target discharge cycle with the control threshold as a constraint, calculate the marginal rate of change of the flow rate relative to the target discharge cycle to obtain the yield elasticity index, and combine the spray system parameters to convert it into a cost elasticity index. Step S102: Integrate the previous time step's running data, minimum spray flow rate, cost elasticity index, and demand load to construct a reinforcement learning state vector and input it into the reinforcement learning network; Step S103: Output candidate actions through the policy network, determine the feasible region based on the dust concentration prediction model and control threshold, and correct the illegal candidate actions into feasible actions through differentiable projection along the model gradient. Step S104: Integrate the costs of pumping electricity, water, conveying energy consumption and overspray penalty to form the periodic cost. Use the product of the cost elasticity index and the change in discharge cycle as the shaded price item, and sum and inverse the two to obtain the immediate reward. Step S105: Using the difference between the predicted dust concentration and the control threshold as a constraint, update the Lagrange multiplier, and construct the advantage function by combining the constraint, the cost elasticity index and the immediate reward, and update the strategy and value network in a coordinated manner. Step S106: Convert the feasible discharge cycle time and spray flow rate into frequency conversion drive commands for the spray pump and belt conveyor, respectively. In the case of multiple silos, the discharge cycle time is allocated according to the coal blending plan to generate the feed gate opening command.
[0007] Furthermore, a set of state parameters is collected, including: environmental dust baseline, material humidity, belt speed, loading rate, chute geometry, and upstream coal characteristics; a dust concentration prediction model is constructed, with the set of state parameters, target discharge cycle time, and spray flow rate as inputs, and the predicted dust concentration value as output; Under the constraint that the predicted dust concentration does not exceed the control threshold, the minimum spray flow rate corresponding to the target discharge cycle is obtained by solving the problem. When the predicted dust concentration is equal to the control threshold, the partial derivatives of the dust concentration prediction model with respect to the target discharge cycle and the spray flow rate are calculated respectively. The marginal rate of change of the minimum spray flow rate relative to the target discharge cycle is calculated based on these two partial derivative values as the yield elasticity index. The operating pressure difference and overall efficiency of the spray system are obtained. The yield elasticity index is linearly transformed by the operating pressure difference and overall efficiency to obtain the cost elasticity index.
[0008] Furthermore, the previous time-instance operational data includes: discharge cycle time, spray flow rate, and predicted dust concentration; the reinforcement learning state vector consists of the previous time-instance operational data, minimum spray flow rate, cost elasticity index, and demand load.
[0009] Furthermore, the policy network receives the reinforcement learning state vector and outputs the discharge cycle adjustment amount and spray flow rate adjustment amount; Add the discharge cycle adjustment amount to the discharge cycle of the previous moment to obtain the candidate discharge cycle; add the spray flow rate adjustment amount to the spray flow rate of the previous moment to obtain the candidate spray flow rate; Based on the dust concentration prediction model and the control threshold, the feasible region is determined, which is the set of all combinations of discharge cycle time and spray flow rate that ensure the dust concentration prediction value output by the dust concentration prediction model does not exceed the control threshold.
[0010] Furthermore, the dust concentration prediction model is used to calculate the predicted dust concentration values corresponding to the candidate discharge cycle and candidate spray flow rate. When the predicted dust concentration value exceeds the control threshold, the dust concentration prediction model is used to calculate the gradient of the candidate discharge cycle and candidate spray flow rate respectively. The candidate discharge cycle and candidate spray flow rate are corrected along the gradient direction using the projection step size constant to obtain the discharge cycle and spray flow rate of the action. When the predicted dust concentration does not exceed the control threshold, the candidate discharge cycle is directly used as the discharge cycle of the actionable action, and the candidate spray flow rate is directly used as the spray flow rate of the actionable action.
[0011] Furthermore, the period cost consists of four parts: pump electricity cost, water cost, transportation energy consumption cost, and over-spraying penalty cost; Pump electricity costs are calculated based on electricity prices, the operating pressure differential of the spray system, the overall efficiency of the spray system, and the spray flow rate of the movable action. Water fees are calculated based on water prices and the spray flow rate of the action. The energy consumption cost of conveying is calculated based on the energy consumption cost function of conveying with the discharge cycle time of the movable operation as the independent variable; The overspray penalty fee is calculated based on the electricity price, the working pressure difference of the spray system, the overall efficiency of the spray system, the spray flow rate of the movable action, and the minimum spray flow rate. The overspray penalty fee is only included when the spray flow rate of the movable action is greater than the minimum spray flow rate. Add up the pump electricity cost, water cost, transportation energy consumption cost, and over-spraying penalty cost to obtain the period cost.
[0012] Furthermore, the difference between the discharge cycle time of the action and the discharge cycle time of the previous moment is calculated to obtain the change in discharge cycle time. The shaded price term is obtained by multiplying the cost elasticity index by the absolute value of the change in the discharge cycle time. Add the periodic cost to the shaded price, take the negative value of the sum, and obtain an immediate reward.
[0013] Furthermore, the difference between the predicted dust concentration and the dust concentration control threshold is calculated to obtain the constraint function value; Set the initial Lagrange multipliers and the Lagrange multiplier update step size. Add the current Lagrange multipliers to the product of the Lagrange multiplier update step size and the constraint function value, and take the non-negative value of the result to obtain the updated Lagrange multipliers. By setting a discount factor, the value of the current state vector corresponding to the reinforcement learning state vector and the value of the next state vector corresponding to the next time step after executing an action are calculated through the value network. The temporal difference term is calculated based on the immediate reward, discount factor, current state value, and next state value. The elastic augmented advantage function is obtained by subtracting the product of the current Lagrange multiplier and the constraint function value from the time-series difference term.
[0014] Furthermore, the learning rate of the policy network is set, the conditional probability density of the policy network generating corresponding actions under the reinforcement learning state vector is obtained, the gradient of the logarithm of the conditional probability density with respect to the policy network parameters is calculated, and the policy network parameters are updated by the product of the policy network learning rate, the gradient and the elastic augmentation advantage function. Set the learning rate of the value network, calculate the temporal difference error by using the immediate reward, discount factor, current state value and next state value, calculate the gradient of the current state value with respect to the value network parameters, and update the value network parameters by multiplying the value network learning rate, temporal difference error and the gradient.
[0015] Furthermore, the flow calibration coefficient of the spray pump is obtained. The flow calibration coefficient is obtained through a one-time test. The ratio of the feasible spray flow rate to the flow calibration coefficient is calculated to obtain the spray pump speed. The first proportional coefficient from the speed to the frequency conversion drive command is obtained. The product of the first proportional coefficient and the spray pump speed is calculated to obtain the frequency conversion drive command of the spray pump. The material bulk density, effective cross-sectional area of the belt, and belt loading rate are obtained. The belt speed of the belt conveyor is obtained by the ratio of the feasible discharge cycle time to the product of the material bulk density, effective cross-sectional area of the belt, and belt loading rate. The second proportional coefficient from the belt speed to the frequency conversion drive command is obtained. The product of the second proportional coefficient and the belt speed of the belt conveyor is calculated to obtain the frequency conversion drive command of the belt conveyor. When blending coal in multiple silos, the allocation coefficients corresponding to the coal blending plan are obtained. Each allocation coefficient is a non-negative value and the sum is 1. The allocated discharge cycle of each silo is obtained by multiplying the feasible discharge cycle by the corresponding allocation coefficient. The opening conversion coefficient corresponding to each feed gate is obtained. The opening command of each feed gate is obtained by multiplying the allocated discharge cycle of the corresponding silo by the opening conversion coefficient of the feed gate.
[0016] The beneficial effects of this invention are as follows: Under dust limit constraints, this invention identifies the minimum spray flow rate required to achieve the target discharge cycle time online and incorporates the corresponding pump power cost into the reinforcement learning decision, realizing joint optimization of spray flow rate and discharge cycle time. This automatically avoids the risk of sudden increases in water and energy consumption and wet coal blockage caused by slight speed increases. Through the convergence of the real-time reward-driven strategy composed of periodic costs and elasticity terms, compared with traditional rule control, it significantly reduces the occurrence of overspraying and repeated speed reduction adjustments, stabilizing material throughput efficiency and effectively reducing pump electricity and water costs. Simultaneously, the optimization results output by reinforcement learning are uniquely transformed into frequency conversion drive commands for the spray pump and belt conveyor, as well as opening commands for the multi-silo feed gates, through explicit physical mapping. This reduces manual intervention, improves the fulfillment rate of coal blending plans, and comprehensively enhances the intelligence and economy of silo coal storage and blending scheduling. Attached Figure Description
[0017] Figure 1 This is a flowchart of an intelligent scheduling method for integrated coal storage and distribution in silos based on reinforcement learning, according to the present invention. Detailed Implementation
[0018] The subject matter described herein will now be discussed with reference to exemplary embodiments. It should be understood that these embodiments are discussed only to enable those skilled in the art to better understand and implement the subject matter described herein, and changes may be made to the function and arrangement of the elements discussed without departing from the scope of this specification. Various processes or components may be omitted, substituted, or added as needed in the examples. Furthermore, features described in some examples may be combined in other examples.
[0019] It should be noted that, unless otherwise defined, the technical or scientific terms used in one or more embodiments of the present invention should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. The terms "first," "second," and similar terms used in one or more embodiments of the present invention do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" indicate that the element or object preceding the term encompasses the elements or objects listed following the term and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0020] like Figure 1 As shown, a reinforcement learning-based intelligent scheduling method for integrated coal storage and distribution in silos includes the following steps: Step S101: Collect state parameters to construct a dust concentration prediction model, solve the minimum spray flow rate corresponding to the target discharge cycle with the control threshold as a constraint, calculate the marginal rate of change of the flow rate relative to the target discharge cycle to obtain the yield elasticity index, and combine the spray system parameters to convert it into a cost elasticity index. Step S102: Integrate the previous time step's running data, minimum spray flow rate, cost elasticity index, and demand load to construct a reinforcement learning state vector and input it into the reinforcement learning network; Step S103: Output candidate actions through the policy network, determine the feasible region based on the dust concentration prediction model and control threshold, and correct the illegal candidate actions into feasible actions through differentiable projection along the model gradient. Step S104: Integrate the costs of pumping electricity, water, conveying energy consumption and overspray penalty to form the periodic cost. Use the product of the cost elasticity index and the change in discharge cycle as the shaded price item, and sum and inverse the two to obtain the immediate reward. Step S105: Using the difference between the predicted dust concentration and the control threshold as a constraint, update the Lagrange multiplier, and construct the advantage function by combining the constraint, the cost elasticity index and the immediate reward, and update the strategy and value network in a coordinated manner. Step S106: Convert the feasible discharge cycle time and spray flow rate into frequency conversion drive commands for the spray pump and belt conveyor, respectively. In the case of multiple silos, the discharge cycle time is allocated according to the coal blending plan to generate the feed gate opening command.
[0021] In one embodiment of the present invention, a set of state parameters is collected, including: environmental dust baseline, material humidity, belt speed, loading rate, chute geometric markings, and upstream coal characteristics; a dust concentration prediction model is constructed, with the set of state parameters, target discharge cycle time, and spray flow rate as inputs, and the predicted dust concentration value as output; Under the constraint that the predicted dust concentration does not exceed the control threshold, the minimum spray flow rate corresponding to the target discharge cycle is obtained by solving the problem. When the predicted dust concentration is equal to the control threshold, the partial derivatives of the dust concentration prediction model with respect to the target discharge cycle and the spray flow rate are calculated respectively. The marginal rate of change of the minimum spray flow rate relative to the target discharge cycle is calculated based on these two partial derivative values as the yield elasticity index. The operating pressure difference and overall efficiency of the spray system are obtained. The yield elasticity index is linearly transformed by the operating pressure difference and overall efficiency to obtain the cost elasticity index.
[0022] It should be noted that: the environmental dust baseline represents the dust concentration in the work area when there is no material transfer, which is obtained in real time by dust sensors; the material humidity represents the moisture content of the transferred material, which is obtained by contact or non-contact humidity sensors; the belt speed represents the operating speed of the belt conveyor, which is obtained by belt speed sensors; the loading rate represents the ratio of the actual material load on the belt to the rated load; the chute geometry represents the structural parameters of the chute, such as the inclination angle and cross-sectional dimensions, which can be obtained from the equipment design ledger; the upstream coal characteristics represent the particle size and hardness of the coal transported upstream, which can be obtained by online particle size analyzers or coal quality test reports; the target discharge cycle time represents the material output rate to be achieved by the silo coal blending operation plan, in tons per hour, which is issued by the production scheduling system according to downstream demand (such as power generation load and boiler coal consumption); and the spray flow rate represents the volumetric flow rate of water mist sprayed by the spray system into the work area, in liters per minute, which is controlled by the spray pump speed and can be obtained in real time by flow sensors on the pipeline.
[0023] It should be noted that the dust concentration prediction model can be built based on a multilayer perceptron or a convolutional neural network model. Historical state parameter sets, target discharge cycle time, and spray flow rate are collected as training sample data, and corresponding measured dust concentration values are collected as sample labels for the corresponding training samples. The training samples can be divided into training set and validation set. During the training process of the dust concentration prediction model, the mean square error between the predicted dust concentration value and the measured dust concentration value is used as the loss function. Then, the parameters of the dust concentration prediction model are updated in reverse through a gradient optimizer (such as Adam, AdaGrad, etc.), thereby reducing the loss value calculated by the loss function, until the maximum number of iterations is reached or the rate of change of the loss value of the loss function is lower than the rate of change threshold (such as 0.01) for 3 to 5 consecutive iterations, then the training of the model is completed.
[0024] It should be noted that the dust concentration control threshold represents the maximum allowable dust concentration in the work area, measured in milligrams per cubic meter (mg / m³), and can be set according to relevant regulations or the company's internal safety standards. Under the constraint that the predicted dust concentration does not exceed the control threshold, the purpose of solving for the minimum spray flow rate is to minimize water consumption and reduce operating costs while adhering to compliance. The gradient descent method is used to solve this problem, specifically including the following steps: Step S201, set the initial spray flow rate, for example, 0 liters per minute; Step S202, calculate the predicted dust concentration corresponding to this flow rate. If the predicted value exceeds the control threshold, increase the spray flow rate by a fixed step size (e.g., 0.5 liters per minute); if the predicted value is lower than the control threshold, decrease the spray flow rate by a fixed step size; Step S203, repeat step S202 until a flow rate is found where the predicted value is less than or equal to the control threshold and further reducing the flow rate would cause the predicted value to exceed the control threshold—this is the minimum spray flow rate. Alternatively, Newton's method can be used for acceleration, adjusting the step size through the partial derivative of the dust concentration prediction model with respect to the spray flow rate; this will not be elaborated upon here.
[0025] It should be noted that the yield elasticity index represents the marginal change in minimum spray flow rate for each unit change in the target discharge cycle time when the predicted dust concentration equals the dust concentration control threshold. It reflects the degree of influence of yield changes on minimum water consumption. The larger the value, the more spray flow rate is required to increase the cycle time. The first calculation result is obtained by using the automatic differentiation function of the dust concentration prediction model to calculate the rate of change of the predicted dust concentration with respect to the target discharge cycle time (i.e., the increase in the predicted dust concentration for each ton / hour increase in the target discharge cycle time). The second calculation result is obtained by calculating the rate of change of the predicted dust concentration with respect to the spray flow rate (i.e., the decrease in the predicted dust concentration for each liter / minute increase in the spray flow rate). Finally, the first calculation result is divided by the second calculation result, and the negative number is taken to obtain the yield elasticity index. The automatic differentiation can be achieved through the automatic differentiation tools of deep learning frameworks (such as PyTorchAutograd, TensorFlowGradientTape, etc.). The two rates of change can be obtained by inputting the current state parameters, the target discharge cycle time, and the minimum spray flow rate, which will not be elaborated here.
[0026] Specifically, the formula for calculating the yield elasticity index is as follows: , where represents the rate of change of the predicted dust concentration with respect to the target discharge cycle time, represents the rate of change of the predicted dust concentration with respect to the spray flow rate, and represents the predicted dust concentration equal to the dust concentration control threshold.
[0027] It should be noted that the cost elasticity index represents an indicator that converts the yield elasticity index from the technical dimension of flow rate corresponding to cycle time to the economic dimension of cost corresponding to cycle time. It is used to reflect the marginal change in the energy consumption cost of the spray system for each unit change in the target output cycle time. Specifically, the formula for calculating the cost elasticity index is as follows: , where represents the operating pressure difference of the spray system, and represents the overall efficiency of the spray system.
[0028] It should be noted that the working pressure difference of the spray system represents the pressure difference between the outlet of the spray pump and the outlet of the nozzle in the spray pipeline network, with the unit being Pascals. It is used to reflect the pressure level of the spray system and can be obtained by collecting the difference value from pressure sensors installed at the pump outlet and nozzle inlet. The overall efficiency of the spray system is a dimensionless parameter used to measure the energy conversion efficiency of the spray system. It is equal to the product of the spray pump efficiency, motor efficiency, and drive device (such as frequency converter) efficiency. It is used to reflect the conversion ratio from electrical energy input to effective hydraulic energy output from the spray pump. Its value ranges between 0 and 1. The spray pump efficiency is obtained by referring to the pump performance curve based on the actual flow rate and head of the pump. The motor efficiency is converted from the energy efficiency level on the motor nameplate (e.g., the efficiency of a level 2 energy efficiency motor is not less than 0.89) or collected by an energy efficiency tester. The drive device efficiency can be found in the equipment technical manual (e.g., the efficiency of a frequency converter is not less than 0.95). All three values are calculated in decimal form.
[0029] In one embodiment of the present invention, the previous time-in operation data includes: discharge cycle time, spray flow rate, and dust concentration prediction value; the reinforcement learning state vector consists of the previous time-in operation data, minimum spray flow rate, cost elasticity index, and demand load.
[0030] It should be noted that the discharge cycle time at the previous moment represents the actual material output rate achieved by the silo coal distribution system in the previous scheduling cycle, in tons per hour, obtained by collecting the average material mass flow rate of the previous cycle through belt scales; the spray flow rate at the previous moment represents the actual volumetric flow rate of water mist sprayed into the work area by the spray system in the previous scheduling cycle, in liters per minute, obtained by collecting the average flow rate of the previous cycle through flow sensors on the spray system pipelines; the dust concentration prediction value at the previous moment represents the estimated dust concentration of the work area calculated by the dust concentration prediction model in the previous scheduling cycle; and the demand load represents the short-term material demand rate issued by the enterprise's business plan at the current moment, in tons per hour, which can be obtained through the enterprise's production scheduling system. Specifically, it is the hourly coal demand plan of downstream processes (such as power plant boilers) at the current moment. If there are multiple downstream processes, the sum of the demands of each process is taken as the final demand load.
[0031] In one embodiment of the present invention, the policy network receives a reinforcement learning state vector and outputs a discharge cycle adjustment amount and a spray flow rate adjustment amount. Add the discharge cycle adjustment amount to the discharge cycle of the previous moment to obtain the candidate discharge cycle; add the spray flow rate adjustment amount to the spray flow rate of the previous moment to obtain the candidate spray flow rate; Based on the dust concentration prediction model and the control threshold, the feasible region is determined, which is the set of all combinations of discharge cycle time and spray flow rate that ensure the dust concentration prediction value output by the dust concentration prediction model does not exceed the control threshold.
[0032] It should be noted that the discharge cycle time adjustment amount represents the increment or decrement of the discharge cycle time at the previous moment, as output by the strategy network, with the unit being tons per hour. A positive value indicates an increase in the cycle time, and a negative value indicates a decrease in the cycle time. Specifically, the reinforcement learning state vector is input into the strategy network, processed by a fully connected layer and the ReLU activation function, and then the discharge cycle time adjustment amount is output. The adjustment amount is limited to the negative difference between the previous discharge cycle time and the equipment's maximum cycle time and the cycle time at the previous moment, to avoid candidate cycles exceeding the equipment's physical limits. The spray flow rate adjustment amount represents the increment or decrement of the spray flow rate at the previous moment, as output by the strategy network, with the unit being liters per minute. A positive value indicates an increase in the spray flow rate, and a negative value indicates a decrease in the spray flow rate. Specifically, the reinforcement learning state vector is input into the strategy network, processed by a fully connected layer and the ReLU activation function, and then the adjustment amount is output. The adjustment amount is limited to the negative difference between the previous spray flow rate and the equipment's maximum spray flow rate and the flow rate at the previous moment, to avoid candidate flow rates exceeding the equipment's physical limits.
[0033] It should be noted that the feasible region represents the set of all possible combinations of discharge cycle time and spray flow rate that ensure the predicted dust concentration does not exceed the control threshold under the current operating conditions. The feasible region is determined by grid search. Specifically, within a reasonable range of discharge cycle time 0 to the maximum cycle time of the equipment and spray flow rate 0 to the maximum spray volume of the equipment, all possible combinations of discharge cycle time and spray flow rate are generated at a fixed step size (cycle time step size 0.5 tons per hour, flow rate step size 0.2 liters per minute). These combinations are then substituted into the dust concentration prediction model one by one, and the combinations whose predicted values do not exceed the control threshold are selected to form the feasible region.
[0034] In one embodiment of the present invention, a dust concentration prediction model is used to calculate the predicted dust concentration values corresponding to the candidate discharge cycle and the candidate spray flow rate. When the predicted dust concentration value exceeds the control threshold, the dust concentration prediction model is used to calculate the gradient of the candidate discharge cycle and the candidate spray flow rate respectively. The candidate discharge cycle and the candidate spray flow rate are corrected along the gradient direction using a projection step size constant to obtain the discharge cycle and spray flow rate of the actionable action. When the predicted dust concentration does not exceed the control threshold, the candidate discharge cycle is directly used as the discharge cycle of the actionable action, and the candidate spray flow rate is directly used as the spray flow rate of the actionable action.
[0035] It should be noted that the predicted dust concentration value corresponding to the candidate action represents the estimated dust concentration obtained by substituting the candidate discharge cycle time and candidate spray flow rate into the dust concentration prediction model; the gradient is calculated using an automatic differentiation tool. At the input positions of the candidate discharge cycle time and candidate spray flow rate, the partial derivatives of the predicted dust concentration value with respect to the candidate discharge cycle time and candidate spray flow rate are calculated respectively, and the two partial derivative values together constitute the gradient. The projection step size constant represents a positive number that controls the correction magnitude of the candidate action. It can be determined through offline debugging. Specifically, 10 groups of historical violation candidate actions are selected, and step sizes in the range of 0.01 to 0.1 are tested. The proportion of the corrected action falling into the feasible region at each step size is counted, and the step size with the highest proportion and the smallest difference between the corrected action and the candidate action (usually 0.05) is selected as the final value.
[0036] Specifically, the formula for calculating differentiable projection is as follows: , where and represent the corrected discharge cycle time and spray flow rate of the action, respectively; and represent the candidate discharge cycle time and candidate spray flow rate, respectively; represents the projection step size constant; represents the gradient operator; represents the predicted dust concentration values corresponding to the candidate discharge cycle time and candidate spray flow rate output by the dust concentration prediction model; represents the dust concentration control threshold; represents the non-negative operator (i.e., set to 0 when less than 0).
[0037] It should be noted that if the predicted dust concentration corresponding to a candidate action does not exceed the control threshold, the discharge cycle time of the feasible action is equal to the candidate discharge cycle time; if the predicted value exceeds the control threshold, the discharge cycle time of the feasible action is equal to the candidate discharge cycle time minus the projection step size constant multiplied by the partial derivative of the predicted dust concentration in the gradient with respect to the candidate discharge cycle time; if the predicted value still exceeds the limit after one correction, the correction is repeated (up to 3 times) until compliance is achieved. If the predicted dust concentration corresponding to a candidate action does not exceed the control threshold, the spray flow rate of the feasible action is equal to the candidate spray flow rate; if the predicted value exceeds the control threshold, the spray flow rate of the feasible action is equal to the candidate spray flow rate minus the projection step size constant multiplied by the partial derivative of the predicted dust concentration in the gradient with respect to the candidate spray flow rate; if the predicted value still exceeds the limit after one correction, the correction is repeated (up to 3 times) until compliance is achieved. If the action still violates the rule after 3 corrections, the maximum spray flow rate of the equipment is taken as the spray flow rate of the feasible action.
[0038] In one embodiment of the present invention, the periodic cost consists of four parts: pump electricity cost, water cost, transportation energy consumption cost, and over-spraying penalty cost; Pump electricity costs are calculated based on electricity prices, the operating pressure differential of the spray system, the overall efficiency of the spray system, and the spray flow rate of the movable action. Water fees are calculated based on water prices and the spray flow rate of the action. The energy consumption cost of conveying is calculated based on the energy consumption cost function of conveying with the discharge cycle time of the movable operation as the independent variable; The overspray penalty fee is calculated based on the electricity price, the working pressure difference of the spray system, the overall efficiency of the spray system, the spray flow rate of the movable action, and the minimum spray flow rate. The overspray penalty fee is only included when the spray flow rate of the movable action is greater than the minimum spray flow rate. Add up the pump electricity cost, water cost, transportation energy consumption cost, and over-spraying penalty cost to obtain the period cost.
[0039] It should be noted that the period cost represents the total operating cost corresponding to the current feasible operation; the pump electricity cost represents the cost of electricity consumed by the spray pump, which is equal to the electricity price multiplied by (the spray system's operating pressure difference divided by the spray system's overall efficiency) and then multiplied by the feasible spray flow rate; where the product of "the spray system's operating pressure difference divided by the spray system's overall efficiency" and "the feasible spray flow rate" is the equivalent power of the spray pump, which is then multiplied by the electricity price to convert power into cost. The water cost represents the cost of water resources consumed during the spraying process, which is equal to the water price multiplied by the feasible spray flow rate. The electricity price is determined by the electricity contract signed between the enterprise and the power supplier or by the local power grid's unified pricing. The water price is determined by the water contract signed between the enterprise and the water supplier or by the local water supply department's unified pricing.
[0040] Specifically, the formula for calculating pump electricity costs is as follows: , where represents the electricity price, represents the operating pressure difference of the spray system, represents the overall efficiency of the spray system, and represents the spray flow rate for the action.
[0041] It should be noted that the conveying energy consumption cost function is a function relating the movable operation's discharge cycle time to the conveying energy consumption cost. First, collect the measured conveying energy consumption values for different daily discharge cycles over the past three months. Multiply the measured energy consumption value by the corresponding electricity price to obtain the conveying energy consumption cost. Then, use linear regression to fit the discharge cycle time and conveying energy consumption cost data to obtain the function form (e.g., conveying energy consumption cost = regression coefficient A × movable operation discharge cycle time + regression constant B). The regression coefficient A and constant B are calculated using the least squares method (to minimize the mean square error between the fitted value and the measured value). The overspray penalty cost represents the penalty cost for additional spray consumption when the movable operation's spray flow rate exceeds the minimum spray flow rate. First, calculate the difference between the movable operation's spray flow rate and the minimum spray flow rate. If the difference is positive (i.e., overspray exists), the overspray penalty cost equals the electricity price multiplied by (spray system operating pressure difference divided by spray system overall efficiency) multiplied by the difference. If the difference is zero or negative (no overspray), the overspray penalty cost is zero.
[0042] Specifically, the formula for calculating the over-boost penalty fee is as follows: , where represents the minimum spray flow rate.
[0043] In one embodiment of the present invention, the difference between the discharge cycle of the movable action and the discharge cycle of the previous moment is calculated to obtain the change in discharge cycle. The shaded price term is obtained by multiplying the cost elasticity index by the absolute value of the change in the discharge cycle time. Add the periodic cost to the shaded price, take the negative value of the sum, and obtain an immediate reward.
[0044] It should be noted that the change in discharge cycle time represents the difference between the actionable discharge cycle time and the previous discharge cycle time; the shaded price term represents the implicit cost quantified by the change in discharge cycle time; the immediate reward represents the feedback signal guiding the optimization of the strategy in reinforcement learning. The larger the value, the more the current action is in line with the goals of low cost and compliance. The immediate reward is equal to a negative value (the cost plus the shaded price term). The negative value is because reinforcement learning defaults to maximizing the reward, while the actual goal is to minimize the sum of the cost and the implicit cost. The goal alignment is achieved through sign transformation.
[0045] Specifically, the formula for calculating instant rewards is as follows: , where represents periodic cost, represents the cost elasticity index, and represents the change in the discharge cycle time.
[0046] In one embodiment of the present invention, the difference between the predicted dust concentration and the dust concentration control threshold is calculated to obtain the constraint function value; Set the initial Lagrange multipliers and the Lagrange multiplier update step size. Add the current Lagrange multipliers to the product of the Lagrange multiplier update step size and the constraint function value, and take the non-negative value of the result to obtain the updated Lagrange multipliers. By setting a discount factor, the value of the current state vector corresponding to the reinforcement learning state vector and the value of the next state vector corresponding to the next time step after executing an action are calculated through the value network. The temporal difference term is calculated based on the immediate reward, discount factor, current state value, and next state value. The elastic augmented advantage function is obtained by subtracting the product of the current Lagrange multiplier and the constraint function value from the time-series difference term.
[0047] It should be noted that the constraint function value is used to quantify the degree of violation of dust concentration constraints. A positive value indicates that the current action violates the dust control threshold (dust exceeds the standard), while a non-positive value indicates that the action is compliant. The initial Lagrange multiplier is set based on the frequency of dust violations over the past 3 months. If the violation frequency is below 5%, it is set to 0.1; if the violation frequency is between 5% and 15%, it is set to 0.3; and if the violation frequency is above 15%, it is set to 0.5. The Lagrange multiplier update step size represents the positive number controlling the update magnitude of the Lagrange multiplier. The sub-update step size is a fixed value between 0.001 and 0.01, specifically determined through offline testing. Five different step sizes (0.001, 0.003, 0.005, 0.008, 0.01) are selected, and the number of iterations required for the multiplier to converge to a stable value under each step size is tested. The step size with the fewest convergences and no oscillations (usually 0.005) is selected. The updated Lagrange multiplier represents the dynamically adjusted constraint penalty weight; the larger the value, the stronger the penalty for subsequent violations, ensuring the policy prioritizes compliant actions. The discount factor represents a dimensionless parameter balancing short-term and long-term rewards in reinforcement learning. The closer the value is to 1, the more the policy emphasizes long-term returns; the closer the value is to 0, the more the policy focuses on immediate returns. The discount factor is set according to the scheduling period: 0.8 for a 10-minute (short-term) scheduling period, 0.9 for a 30-minute (medium-term) scheduling period, and 0.95 for a 1-hour (long-term) scheduling period, ensuring the factor matches the time scale of the scheduling objective.
[0048] It should be noted that the value network represents the neural network used in reinforcement learning to evaluate the long-term benefits of a state, with parameters being weights and biases; the current state value represents the long-term benefit evaluation of the value network for the current reinforcement learning state vector, reflecting the cumulative reward that all subsequent actions may bring in the current state. The current state value is equal to the output result calculated by the value network after receiving the current reinforcement learning state vector, through the input layer, hidden layer (using the ReLU activation function), and output layer (linear activation function); the next state value represents the long-term benefit evaluation of the value network for the reinforcement learning state vector entered in the next time step after performing an action. The next state value is equal to the output result calculated by the value network after receiving the next reinforcement learning state vector, through the input layer, hidden layer (using the ReLU activation function), and output layer (linear activation function); the temporal difference term is used to measure the deviation between the immediate benefit and the long-term benefit of the current action; the temporal difference term is equal to the immediate reward plus a discount factor multiplied by the next state value, and then minus the current state value; where the discount factor multiplied by the next state value is the long-term benefit discount, which is added to the immediate reward to obtain the target value, and then subtracted from the current state value to obtain the deviation. The elastic augmented advantage function represents a scalar signal that integrates constraint penalties and action advantages, used to guide policy network updates and enable the policy to avoid illegal and high-cost actions.
[0049] Specifically, the formula for calculating the elasticity augmentation advantage function is as follows: , where represents the immediate reward, represents the discount factor, represents the state value of the next time step output by the value network, represents the state value of the current time step output by the value network, represents the Lagrange multiplier, and represents the constraint function value.
[0050] In one embodiment of the present invention, a learning rate for the policy network is set, the conditional probability density of the policy network generating corresponding actions under the reinforcement learning state vector is obtained, the gradient of the logarithm of the conditional probability density with respect to the policy network parameters is calculated, and the policy network parameters are updated using the product of the policy network learning rate, the gradient and the elastic augmentation advantage function. Set the learning rate of the value network, calculate the temporal difference error by using the immediate reward, discount factor, current state value and next state value, calculate the gradient of the current state value with respect to the value network parameters, and update the value network parameters by multiplying the value network learning rate, temporal difference error and the gradient.
[0051] It should be noted that the policy network learning rate is a positive number that controls the magnitude of policy network parameter updates. The policy network learning rate is determined through trial and error, initially set to 0.0001. If the average reward improvement of the policy after each iteration is less than 1%, the learning rate is increased to 1.2 times the original value; if the reward fluctuation exceeds 10%, the learning rate is decreased to 0.8 times the original value, until a learning rate that stably improves the reward is found (usually between 0.0001 and 0.001). The conditional probability density represents the probability density of the corresponding action that the policy network can generate given the reinforcement learning state vector. The conditional probability density is calculated through the output layer of the policy network (using the Softmax activation function, suitable for discrete actions; using a Gaussian distribution output, suitable for continuous actions). Specifically, the reinforcement learning state vector is input into the policy network, and after the hidden layer operation, the probability distribution parameters of the output action (such as the mean and variance of the Gaussian distribution) are output, and then substituted into the probability density formula of the corresponding distribution.
[0052] Specifically, the calculation formula for policy network updates is as follows: , where represents the parameters of the policy network, represents the learning rate of the policy network, represents the gradient of the conditional probability density of the output action of the policy network in the state with respect to the parameters, represents the conditional probability density of the output action of the policy network in the state, and represents the elastic augmentation advantage function.
[0053] It should be noted that the gradient of the logarithm of the conditional probability density with respect to the policy network parameters is calculated using an automatic differentiation tool. The policy network parameters are equal to the parameters before the update plus the policy network learning rate multiplied by the gradient of the logarithm of the conditional probability density with respect to the policy network parameters, and then multiplied by the elastic augmentation dominance function. If the dominance function is positive, the parameter update direction increases the probability of that action; otherwise, it decreases it. The value network learning rate is a positive number that controls the magnitude of the value network parameter update. The value network learning rate is 10 times that of the policy network learning rate (because the value network evaluation accuracy requirement is lower than that of the policy network, and faster convergence is required). If the policy network learning rate is 0.0005, then the value network learning rate is 0.005. If the temporal difference error of the value network increases by more than 20% after the update, the value network learning rate is reduced to 0.8 times the original value. Temporal difference error measures the prediction bias of the value network. A smaller error indicates a more accurate assessment of the state by the value network. Temporal difference error equals the immediate reward plus a discount factor multiplied by the value of the state at the next time step, minus the value of the current state. This is calculated in the same way as the temporal difference term, which is used in the advantage function, while the temporal difference error is used for value network updates. The gradient of the current state value with respect to the value network parameters is calculated using an automatic differentiation tool. The value network parameters equal the parameters before the update plus the value network learning rate multiplied by the temporal difference error, and then multiplied by the gradient of the current state value with respect to the value network parameters. If the error is positive, the parameter update direction increases the predicted value of the current state; conversely, it decreases it.
[0054] Specifically, the calculation formula for value network updates includes: , where represents the parameters of the value network, represents the learning rate of the value network, and represents the gradient of the value network with respect to the parameters in relation to the current state.
[0055] , where represents the time-series difference error, represents the immediate reward, represents the discount factor, represents the state value of the next time step output by the value network, and represents the state value of the current time step output by the value network.
[0056] In one embodiment of the present invention, the flow calibration coefficient of the spray pump is obtained. The flow calibration coefficient is obtained through a one-time test. The ratio of the feasible spray flow rate to the flow calibration coefficient is calculated to obtain the spray pump speed. The first proportional coefficient from the speed to the frequency conversion drive command is obtained. The product of the first proportional coefficient and the spray pump speed is calculated to obtain the frequency conversion drive command of the spray pump. The material bulk density, effective cross-sectional area of the belt, and belt loading rate are obtained. The belt speed of the belt conveyor is obtained by the ratio of the feasible discharge cycle time to the product of the material bulk density, effective cross-sectional area of the belt, and belt loading rate. The second proportional coefficient from the belt speed to the frequency conversion drive command is obtained. The product of the second proportional coefficient and the belt speed of the belt conveyor is calculated to obtain the frequency conversion drive command of the belt conveyor. When blending coal in multiple silos, the allocation coefficients corresponding to the coal blending plan are obtained. Each allocation coefficient is a non-negative value and the sum is 1. The allocated discharge cycle of each silo is obtained by multiplying the feasible discharge cycle by the corresponding allocation coefficient. The opening conversion coefficient corresponding to each feed gate is obtained. The opening command of each feed gate is obtained by multiplying the allocated discharge cycle of the corresponding silo by the opening conversion coefficient of the feed gate.
[0057] It should be noted that the flow rate calibration coefficient of the spray pump reflects the spray flow rate corresponding to a unit rotational speed of the spray pump. It is used to convert the feasible spray flow rate into the spray pump rotational speed. This can be achieved by selecting five different fixed spray pump rotational speeds (covering 30% to 100% of the pump's rated speed, such as 200 rpm, 400 rpm, 600 rpm, 800 rpm, and 1000 rpm), recording the spray flow rate at each stable operating speed, calculating the ratio of spray flow rate to rotational speed for each speed, and taking the average of the five ratios as the flow rate calibration coefficient. The spray pump rotational speed represents the target rotational speed that the spray pump needs to achieve. It is an intermediate parameter connecting the feasible spray flow rate and the variable frequency drive command. The higher the rotational speed, the greater the spray flow rate. If the calculated result exceeds the pump's rated rotational speed, the rated rotational speed is taken as the final spray pump rotational speed.
[0058] It should be noted that the first proportional coefficient represents the proportionality factor for converting the spray pump speed into the variable frequency drive command. This can be achieved by selecting five different target spray pump speeds (consistent with the speeds used in the flow calibration coefficient test), recording the drive commands output by the frequency converter that enable the pump to reach the corresponding speeds, calculating the ratio of the variable frequency drive command to the speed for each speed, and taking the average of the five ratios. The spray pump variable frequency drive command represents the control signal that the spray pump frequency converter can directly execute. If the calculated result exceeds the rated output range of the frequency converter (e.g., 0 to 50 Hz), the nearest value within the range is taken (0 Hz if below the lower limit, 50 Hz if above the upper limit).
[0059] It should be noted that the bulk density of a material represents the mass of a unit volume of material, which can be obtained from the material technical ledger (e.g., the bulk density of anthracite is approximately 0.8 to 1.0 tons per cubic meter, and that of bituminous coal is approximately 0.7 to 0.9 tons per cubic meter). The effective cross-sectional area of the belt conveyor represents the effective area of the cross-section of the material on the belt conveyor, which is determined by the belt width and the average stacking height of the material. The effective cross-sectional area of the belt conveyor is equal to the belt width multiplied by the average stacking height of the material. The belt width is obtained from the equipment ledger (e.g., 1.2 meters, 1.4 meters, etc.), and the average stacking height of the material is the average measured value of the belt conveyor during operation over the past 3 months. This can be obtained by subtracting the belt thickness from the distance from the top of the material to the belt using a laser rangefinder. Belt loading rate represents the ratio of the actual material load of the belt conveyor to the rated full-load load, dimensionless (value from 0 to 1); it reflects the load level of the belt conveyor and prevents overload or underload operation. It is calculated by real-time data collection from the belt scale. The belt loading rate equals the actual material mass flow rate collected by the belt scale divided by the rated full-load mass flow rate of the belt conveyor; where the rated full-load mass flow rate equals the material bulk density multiplied by the effective cross-sectional area of the belt multiplied by the rated maximum belt speed of the belt conveyor (obtained from the equipment ledger). Belt speed represents the target operating speed of the belt conveyor belt and is an intermediate parameter connecting the feasible discharge cycle time and the belt conveyor frequency conversion command. The higher the belt speed, the larger the discharge cycle time. Belt speed equals the feasible discharge cycle time divided by (material bulk density multiplied by the effective cross-sectional area of the belt multiplied by the belt loading rate). If the calculated result exceeds the rated maximum belt speed, the rated maximum belt speed is used; if it is lower than the rated minimum belt speed (to avoid belt slippage), the rated minimum belt speed is used.
[0060] It should be noted that the second proportional coefficient represents the ratio for converting the belt speed of the conveyor belt to the variable frequency drive command. This can be achieved by selecting five different target belt speeds (covering 30% to 100% of the rated belt speed, such as 0.5 m / s, 1.0 m / s, 1.5 m / s, 2.0 m / s, and 2.5 m / s), recording the drive commands output by the frequency converter that enable the belt to reach the corresponding speed, calculating the ratio of the variable frequency drive command to the belt speed for each speed, and taking the average of the five ratios. The variable frequency drive command of the conveyor belt represents the control signal that the frequency converter can directly execute. If the calculated result exceeds the rated output range of the frequency converter (e.g., 0 to 50 Hz), the closest value within the range is taken (0 Hz if below the lower limit, 50 Hz if above the upper limit).
[0061] It should be noted that the allocation coefficient corresponding to the coal blending plan represents the proportion of the discharge cycle time of each silo to the total feasible discharge cycle time when blending coal in multiple silos. It is dimensionless (values from 0 to 1). The sum of all allocation coefficients is 1 and all are non-negative, ensuring that the total discharge cycle time matches the allocated cycle time of each silo. It is determined by the upstream coal blending plan (combining the coal type inventory of each silo and downstream demand). The allocation coefficient is calculated based on the product of the silo inventory weight and the downstream demand weight. First, the inventory weight of each silo is calculated (the current inventory of the silo divided by the total inventory of all participating coal blending silos). Then, the demand weight of the corresponding coal type for each silo is calculated (the hourly demand of the downstream for this coal type divided by the total hourly coal demand of the downstream). The allocation coefficient of each silo is equal to the inventory weight multiplied by the demand weight. Finally, all allocation coefficients are normalized (each coefficient divided by the sum of the coefficients) to ensure that the sum is 1.
[0062] It should be noted that the number of silos represents the total number of silos participating in the current coal blending operation. The allocated discharge cycle time for each silo represents the target discharge cycle time that a single silo needs to achieve when blending coal from multiple silos. The allocated discharge cycle time for each silo is equal to the total feasible discharge cycle time multiplied by the allocation coefficient corresponding to that silo. If the calculated result exceeds the maximum discharge capacity of the silo (obtained from the equipment ledger), the maximum discharge capacity is taken, and the allocation coefficients of other silos are readjusted (keeping the total cycle time unchanged). The opening conversion coefficient corresponding to each feed gate represents the proportional coefficient for converting the allocated discharge cycle time of the silo into the feed gate opening. For a single feed gate, five different opening values are selected (0%, 25%, 50%, 75%, 100%, 0% is fully closed, and 100% is fully open). The discharge cycle time under each opening is recorded during stable operation. The ratio of the opening value to the discharge cycle time is calculated for each opening, and the average of the five ratios is the opening conversion coefficient for that feed gate. The opening command of each feed gate represents the target opening signal that the feed gate needs to achieve. It converts the silo allocation cycle into the final command for the feed gate action. The larger the opening, the larger the discharge cycle. The opening command of each feed gate is equal to the allocated discharge cycle of the corresponding silo multiplied by the opening conversion coefficient of that feed gate. If the calculation result exceeds the range of 0% to 100%, the nearest value in the range is taken (0% for values below 0%, 100% for values above 100%).
[0063] It should be noted that when the allocated cycle time of a silo exceeds its maximum capacity, the cycle time of that silo should first be set to its maximum capacity. The remaining total cycle time (total feasible cycle time minus the maximum capacity of that silo) should be calculated. Then, the remaining cycle time should be redistributed to other silos according to the original allocation coefficient ratio. This process should be repeated until the cycle time of all silos is within the capacity range, thereby resolving the issue of allocated cycle time exceeding capacity and ensuring that each silo is executable. When the feed gate undergoes maintenance (such as replacing the motor or adjusting the mechanical structure), or when the deviation between the measured opening and the commanded opening exceeds ±10% for three consecutive times, the opening conversion coefficient needs to be recalibrated using the method of measuring the opening at 5 opening points. This ensures that the coefficient matches the current state of the equipment, thereby resolving the issue of when the coefficient becomes invalid and avoiding deviations in feed gate operation.
[0064] It should be noted that the interval and threshold sizes are set for ease of comparison. The size of the threshold depends on the amount of sample data and the base number set by those skilled in the art for each set of sample data, as long as it does not affect the proportional relationship between the parameter and the quantized value. Furthermore, the above formulas are all dimensionless calculations, and the formulas are derived from software simulations using a large amount of collected data to obtain the most recent real-world results. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0065] The embodiments of this example have been described above. However, this example is not limited to the specific implementation methods described above. The specific implementation methods described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms based on the guidance of this example, and all of them are within the protection scope of this example.
Claims
1. A smart scheduling method for integrated coal storage and distribution in silos based on reinforcement learning, characterized in that, Includes the following steps: Step S101: Collect state parameters to construct a dust concentration prediction model, solve the minimum spray flow rate corresponding to the target discharge cycle with the control threshold as a constraint, calculate the marginal rate of change of the flow rate relative to the target discharge cycle to obtain the yield elasticity index, and combine the spray system parameters to convert it into a cost elasticity index. Step S102: Integrate the previous time step's running data, minimum spray flow rate, cost elasticity index, and demand load to construct a reinforcement learning state vector and input it into the reinforcement learning network; Step S103: Output candidate actions through the policy network, determine the feasible region based on the dust concentration prediction model and control threshold, and correct the illegal candidate actions into feasible actions through differentiable projection along the model gradient. Step S104: Integrate the costs of pumping electricity, water, conveying energy consumption and overspray penalty to form the periodic cost. Use the product of the cost elasticity index and the change in discharge cycle as the shaded price item, and sum and inverse the two to obtain the immediate reward. Step S105: Using the difference between the predicted dust concentration and the control threshold as a constraint, update the Lagrange multiplier, and construct the advantage function by combining the constraint, the cost elasticity index and the immediate reward, and update the strategy and value network in a coordinated manner. Step S106: Convert the feasible discharge cycle time and spray flow rate into frequency conversion drive commands for the spray pump and belt conveyor, respectively. In the case of multiple silos, the discharge cycle time is allocated according to the coal blending plan to generate the feed gate opening command.
2. The intelligent scheduling method for integrated coal storage and distribution in silos based on reinforcement learning according to claim 1, characterized in that, Collect a set of state parameters, including: environmental dust baseline, material humidity, belt speed, loading rate, chute geometric markings, and upstream coal characteristics; construct a dust concentration prediction model, taking the set of state parameters, target discharge cycle time, and spray flow rate as inputs, and the predicted dust concentration value as output; Under the constraint that the predicted dust concentration does not exceed the control threshold, the minimum spray flow rate corresponding to the target discharge cycle is obtained by solving the problem. When the predicted dust concentration is equal to the control threshold, the partial derivatives of the dust concentration prediction model with respect to the target discharge cycle and the spray flow rate are calculated respectively. The marginal rate of change of the minimum spray flow rate relative to the target discharge cycle is calculated based on these two partial derivative values as the yield elasticity index. The operating pressure difference and overall efficiency of the spray system are obtained. The yield elasticity index is linearly transformed by the operating pressure difference and overall efficiency to obtain the cost elasticity index.
3. The intelligent scheduling method for integrated coal storage and distribution in silos based on reinforcement learning according to claim 1, characterized in that, The previous time-instance operational data includes: discharge cycle time, spray flow rate, and predicted dust concentration; the reinforcement learning state vector consists of the previous time-instance operational data, minimum spray flow rate, cost elasticity index, and demand load.
4. The intelligent scheduling method for integrated coal storage and distribution in silos based on reinforcement learning according to claim 1, characterized in that, The policy network receives the reinforcement learning state vector and outputs the discharge cycle adjustment amount and spray flow rate adjustment amount; Add the discharge cycle adjustment amount to the discharge cycle of the previous moment to obtain the candidate discharge cycle; add the spray flow rate adjustment amount to the spray flow rate of the previous moment to obtain the candidate spray flow rate; Based on the dust concentration prediction model and the control threshold, the feasible region is determined, which is the set of all combinations of discharge cycle time and spray flow rate that ensure the dust concentration prediction value output by the dust concentration prediction model does not exceed the control threshold.
5. The intelligent scheduling method for integrated coal storage and distribution in silos based on reinforcement learning according to claim 4, characterized in that, The dust concentration prediction model is used to calculate the predicted dust concentration values corresponding to the candidate discharge cycle and candidate spray flow rate. When the predicted dust concentration value exceeds the control threshold, the dust concentration prediction model is used to calculate the gradient of the candidate discharge cycle and candidate spray flow rate respectively. The candidate discharge cycle and candidate spray flow rate are corrected along the gradient direction using the projection step size constant to obtain the discharge cycle and spray flow rate of the actionable operation. When the predicted dust concentration does not exceed the control threshold, the candidate discharge cycle is directly used as the discharge cycle of the actionable action, and the candidate spray flow rate is directly used as the spray flow rate of the actionable action.
6. The intelligent scheduling method for integrated coal storage and distribution in silos based on reinforcement learning according to claim 1, characterized in that, The periodic cost consists of four parts: pump electricity cost, water cost, transportation energy consumption cost, and over-spraying penalty cost; Pump electricity costs are calculated based on electricity prices, the operating pressure differential of the spray system, the overall efficiency of the spray system, and the spray flow rate of the movable action. Water fees are calculated based on water prices and the spray flow rate of the action. The energy consumption cost of conveying is calculated based on the energy consumption cost function of conveying with the discharge cycle time of the movable operation as the independent variable; The overspray penalty fee is calculated based on the electricity price, the working pressure difference of the spray system, the overall efficiency of the spray system, the spray flow rate of the movable action, and the minimum spray flow rate. The overspray penalty fee is only included when the spray flow rate of the movable action is greater than the minimum spray flow rate. Add up the pump electricity cost, water cost, transportation energy consumption cost, and over-spraying penalty cost to obtain the period cost.
7. The intelligent scheduling method for integrated coal storage and distribution in silos based on reinforcement learning according to claim 6, characterized in that, Calculate the difference between the discharge cycle time of the action and the discharge cycle time of the previous moment to obtain the change in discharge cycle time; The shaded price term is obtained by multiplying the cost elasticity index by the absolute value of the change in the discharge cycle time. Add the periodic cost to the shaded price, take the negative value of the sum, and obtain an immediate reward.
8. The intelligent scheduling method for integrated coal storage and distribution in silos based on reinforcement learning according to claim 1, characterized in that, Calculate the difference between the predicted dust concentration and the dust concentration control threshold to obtain the constraint function value; Set the initial Lagrange multipliers and the Lagrange multiplier update step size. Add the current Lagrange multipliers to the product of the Lagrange multiplier update step size and the constraint function value, and take the non-negative value of the result to obtain the updated Lagrange multipliers. By setting a discount factor, the value of the current state vector corresponding to the reinforcement learning state vector and the value of the next state vector corresponding to the next time step after executing an action are calculated through the value network. The temporal difference term is calculated based on the immediate reward, discount factor, current state value, and next state value. The elastic augmented advantage function is obtained by subtracting the product of the current Lagrange multiplier and the constraint function value from the time-series difference term.
9. The intelligent scheduling method for integrated coal storage and distribution in silos based on reinforcement learning according to claim 8, characterized in that, Set the learning rate of the policy network, obtain the conditional probability density of the policy network to generate corresponding actions under the reinforcement learning state vector, calculate the gradient of the logarithm of the conditional probability density with respect to the policy network parameters, and update the policy network parameters with the product of the policy network learning rate, the gradient and the elastic augmentation advantage function. Set the learning rate of the value network, calculate the temporal difference error by using the immediate reward, discount factor, current state value and next state value, calculate the gradient of the current state value with respect to the value network parameters, and update the value network parameters by multiplying the value network learning rate, temporal difference error and the gradient.
10. The intelligent scheduling method for integrated coal storage and distribution in silos based on reinforcement learning according to claim 1, characterized in that, The flow rate calibration coefficient of the spray pump is obtained through a one-time test. The ratio of the feasible spray flow rate to the flow rate calibration coefficient is calculated to obtain the spray pump speed. The first proportional coefficient from the speed to the frequency conversion drive command is obtained. The product of the first proportional coefficient and the spray pump speed is calculated to obtain the frequency conversion drive command of the spray pump. The material bulk density, effective cross-sectional area of the belt, and belt loading rate are obtained. The belt speed of the belt conveyor is obtained by the ratio of the feasible discharge cycle time to the product of the material bulk density, effective cross-sectional area of the belt, and belt loading rate. The second proportional coefficient from the belt speed to the frequency conversion drive command is obtained. The product of the second proportional coefficient and the belt speed of the belt conveyor is calculated to obtain the frequency conversion drive command of the belt conveyor. When blending coal in multiple silos, the allocation coefficients corresponding to the coal blending plan are obtained. Each allocation coefficient is a non-negative value and the sum is 1. The allocated discharge cycle of each silo is obtained by multiplying the feasible discharge cycle by the corresponding allocation coefficient. The opening conversion coefficient corresponding to each feed gate is obtained. The opening command of each feed gate is obtained by multiplying the allocated discharge cycle of the corresponding silo by the opening conversion coefficient of the feed gate.