Ventilation resistance coefficient inversion method based on deep reinforcement learning
Patent Information
- Application Number
- CN202310833644.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-07
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-07-07
AI Technical Summary
通风阻力系数可以通过测量来获取,但是对于大型复杂通风系统,测量的巷道多达上百条,工作难度大
[0038] This invention presents a deep reinforcement learning-based method for inverting ventilation resistance coefficients, enabling the automatic acquisition of mine ventilation resistance coefficients. The mine ventilation resistance coefficient is a crucial parameter in real-time ventilation network calculations; an accurate coefficient should ensure consistency between the calculated airflow and the sensor readings of the monitoring system. For large and complex ventilation systems, measuring hundreds of roadways presents significant challenges. Furthermore, while ventilation resistance coefficients can be obtained using empirical formulas, these formulas cannot resolve the contradiction between universality and accuracy. These issues restrict the acquisition and accuracy of ventilation resistance coefficients, directly limiting the extent of digital twin creation for ventilation systems. Therefore, this invention proposes a deep reinforcement learning-based method for inverting ventilation resistance coefficients, which is of great significance for the safety management, digitalization, and intelligentization of ventilation systems.
Smart Images

Figure CN117332469B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of mine ventilation technology, specifically relating to a method for inverting ventilation resistance coefficients based on deep reinforcement learning. Background Technology
[0002] The distribution of mine ventilation resistance coefficient is not only closely related to the stability of the mine ventilation system, but also a core parameter in the digitalization and intelligentization of the ventilation system. Its accuracy directly affects the reliability of real-time ventilation network calculation. As a decisive parameter for real-time ventilation network calculation, the accurate ventilation resistance coefficient should ensure that the calculated air volume of the real-time mine ventilation network is consistent with the sensor air volume readings of the monitoring system. The ventilation resistance coefficient can be obtained through measurement, but for large and complex ventilation systems, hundreds of roadways need to be measured, making the work very difficult. In practice, the ventilation resistance coefficient is often measured only for a portion of the roadways, and the accuracy of the measurement is affected by underground production activities, and the instrument itself also has certain errors. Furthermore, the ventilation resistance coefficient can also be obtained using empirical formulas, but empirical formulas cannot resolve the contradiction between universality and accuracy. Deng Lijun proposed using traditional genetic algorithms to invert the wind resistance coefficient, but traditional heuristic algorithms are prone to getting trapped in local optima and are difficult to solve high-dimensional problems. These problems restrict the acquisition and accuracy of the ventilation resistance coefficient, directly limiting the extent of digital twinning of the ventilation system. Therefore, the use of deep reinforcement learning artificial intelligence algorithms to invert the ventilation drag coefficient is of great significance for the safety management, digitalization, and intelligentization of ventilation systems. Summary of the Invention
[0003] The purpose of this invention is to provide a method for inverting ventilation resistance coefficient based on reinforcement learning. Through continuous trial and error learning by an intelligent agent, the optimal ventilation resistance coefficient is automatically searched, providing technical support for real-time ventilation network calculation.
[0004] To achieve the above objectives, the present invention adopts the following technical solution:
[0005] A method for inverting ventilation drag coefficient based on reinforcement learning has been invented, and the specific steps are as follows:
[0006] S1: Define the agent's interaction environment env, whose environment Where MVSS is the ventilation simulation system, and R is the reward function. To monitor airflow, rew = R(S) t A t ) indicates that in environmental state S t Take action A t The reward feedback value of the time environment, S t The ventilation resistance coefficient r at time t-1 t-1 ={r(t-1)1 ,r (t-1)2 ,r (t-1)3 ,…r (t-1)n}, A t Let be the ventilation resistance coefficient output by the agent at time t;
[0007] S2: Define the agent, which includes a policy neural network (Actor-Net) with θ parameters and a value neural network (Critic-Net) with w parameters. The Actor-Net determines the current environment state based on S... t Output ventilation resistance coefficient distribution π θ A t Sampling at π θ Critic-Net determines the current state S. t Output state value V t ;
[0008] S3: Initialize agent parameters and learning training parameters. Agent parameters include network layer number (layers), number of neurons (sizes), and learning rate (lr). Training parameters include maximum number of epochs for agent updates (max-epochs), batch number of training epochs (batch-epoch), batch step size (steps), buffer size for experience retrieval, and discount factor γ.
[0009] S4: Collect trajectory data of the agent's interaction with the environment, and collect <S,A,rews,π> generated during the interaction. θ (S)>, where S is the set of environmental states, A is the set of searched ventilation resistance coefficient vectors (i.e., actions), rews is the set of reward values, and π θ (S) represents the distribution of the policy neural network mapping under state S;
[0010] S5: Sample and collect trajectory data, calculate the dominance function. Update θ and w to complete the max-epochs agent policy update and output the optimal drag coefficient.
[0011] In step S1, the agent's interaction environment env is defined. Where MVSS is the ventilation simulation system, and R is the reward function. To monitor air volume, R(S) t A t ) indicates that in environmental state S t Take action A t The reward feedback value of the time environment, S t The distribution of ventilation drag coefficient r at time t-1 t-1 ={r (t-1)1 ,r (t-1)2 ,r(t-1)3 ,…r (t-1)n}, A t Let be the ventilation resistance coefficient output by the agent at time t, and the reward function be defined as:
[0012]
[0013] In the formula, w i Weights are used for evaluation; metrics are evaluation indicators of the same dimension, used to evaluate... and The distance or error between them; For in A t The calculated air volume of the ventilation simulation system under the drag coefficient distribution corresponds to Air volume; α is the reward value scaling factor.
[0014] In step S2, an agent is defined, comprising a policy neural network Actor-Net with θ parameters and a value neural network Critic-Net with w parameters. Actor-Net determines the current environment state S based on the current environment state. t Output ventilation resistance coefficient distribution π θ A t Sampling at π θ Critic-Net determines the current state S. t Output state value V t The neural network structure described is a multilayer perceptron (MLP) model, including an input layer, hidden layers, activation layers, and an output layer; the agent adjusts the current environmental state S... t Output the corresponding A t Defined as:
[0015]
[0016] In the formula, A i Ventilation resistance coefficient.
[0017] In step S3, the agent parameters and learning training parameters are initialized. The agent parameters include the number of network layers (layers), the number of neurons (sizes), and the learning rate (lr). The training parameters include the maximum number of epochs for agent updates (max-epochs), the number of batch training epochs (batch-epoch), the step size of the batch training epoch (steps), the buffer size for experience retrieval (buffer-sizes), and the discount factor γ.
[0018] In step S4, trajectory data of the agent's interaction with the environment is collected, and <S,A,rews,π> generated during the interaction are collected.θ (S)>, where S is the set of environmental states, A is the set of searched ventilation resistance coefficients (i.e., actions), rews is the set of reward values, and π θ (S) represents the distribution of the policy neural network mapping under state S.
[0019] In step S5, trajectory data is sampled and collected, and the dominance function is calculated. Update θ and w to complete the max-epochs agent policy update and output the optimal drag coefficient;
[0020] Specifically, it is stated as follows:
[0021] S5.1: Sample interaction information:
[0022] D = {τ1,τ2,τ3,…,τ} batch-epoch} (3)
[0023]
[0024] In the formula, D is the i-th trajectory set; τ is an interactive trajectory segment; batch-epoch is the batch training episode; S t A represents the environmental state at time t; t The action at time t corresponds to the ventilation drag coefficient; rew t Let be the reward value at time t; Let be the policy distribution at time t;
[0025] S5.2: Calculate the dominance function
[0026]
[0027]
[0028] In the formula, γ is the discount factor; steps is the plot step size; λ is the decay factor; rew t V(S) represents the reward value at time t. t ) represents the mapping value of the Critic-Net network;
[0029] S5.3: Update Actor-Net network parameters θ:
[0030]
[0031]
[0032]
[0033] In the formula, loss is the loss function; ε is the shearing factor; and lr is the learning rate.
[0034] S5.4: Update Critic-Net network parameters w:
[0035]
[0036]
[0037] S5.5: Update the network by max-epochs times and output the optimal training result.
[0038] This invention presents a deep reinforcement learning-based method for inverting ventilation resistance coefficients, enabling the automatic acquisition of mine ventilation resistance coefficients. The mine ventilation resistance coefficient is a crucial parameter in real-time ventilation network calculations; an accurate coefficient should ensure consistency between the calculated airflow and the sensor readings of the monitoring system. For large and complex ventilation systems, measuring hundreds of roadways presents significant challenges. Furthermore, while ventilation resistance coefficients can be obtained using empirical formulas, these formulas cannot resolve the contradiction between universality and accuracy. These issues restrict the acquisition and accuracy of ventilation resistance coefficients, directly limiting the extent of digital twin creation for ventilation systems. Therefore, this invention proposes a deep reinforcement learning-based method for inverting ventilation resistance coefficients, which is of great significance for the safety management, digitalization, and intelligentization of ventilation systems. Attached Figure Description
[0039] Figure 1 This is a flowchart of a ventilation resistance coefficient inversion method based on deep reinforcement learning according to the present invention.
[0040] Figure 2 This is a diagram of a 17-branch ventilation network.
[0041] Figure 3 The graph shows the performance results of the deep reinforcement learning agent.
[0042] Figure 4 This is a graph showing the results of the inverted air volume. Detailed Implementation
[0043] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and are not intended to limit the scope of the invention.
[0044] like Figure 1 As shown, the method of this embodiment is as follows:
[0045] This invention provides a method for inverting ventilation drag coefficients based on deep reinforcement learning, comprising the following steps:
[0046] In step S1, the agent's interaction environment env is defined. Where MVSS is the ventilation simulation system, and R is the reward function. To monitor air volume, R(S) t A t ) indicates that in environmental state S t Take action A t The reward feedback value of the time environment, S t The distribution of ventilation drag coefficient r at time t-1 t-1 ={r (t-1)1 ,r (t-1)2 ,r (t-1)3 ,…r (t-1)n}, A t Let be the ventilation resistance coefficient output by the agent at time t, and the reward function be defined as:
[0047]
[0048] In the formula, x represents the number of monitored air volumes; For in A t The calculated air volume of the ventilation simulation system under the drag coefficient distribution corresponds to Air volume; α is the reward value scaling factor.
[0049] In step S2, an agent is defined, comprising a policy neural network Actor-Net with θ parameters and a value neural network Critic-Net with w parameters. Actor-Net determines the current environment state S based on the current environment state. t Output ventilation resistance coefficient distribution π θ A t Sampling at π θ Critic-Net determines the current state S. t Output state value V t The neural network structure described is a multilayer perceptron (MLP) model, including an input layer, hidden layers, activation layers, and an output layer; the agent adjusts the current environmental state S... t Output the corresponding A t Defined as:
[0050]
[0051] In the formula, A i Ventilation resistance coefficient.
[0052] In step S3, the agent parameters and learning training parameters are initialized. The agent parameters include the number of network layers (layers = 2), the number of neurons (sizes = [64, 32]), and the learning rate (lr = 10).-3 The training parameters include the maximum number of episodes updated by the agent (max-epochs) = 120, the number of episodes per batch (batch-epoch) = 30, the step size of the episodes per batch (steps) = 10, the buffer size of the experience retrieval (buffer-sizes) = 20000, and the discount factor γ = 0.95.
[0053] In step S4, trajectory data of the agent's interaction with the environment is collected, and <S,A,rews,π> generated during the interaction are collected. θ (S)>, where S is the set of environmental states, A is the set of searched ventilation resistance coefficients (i.e., actions), rews is the set of reward values, and π θ (S) represents the distribution of the policy neural network mapping under state S.
[0054] In step S5, trajectory data is sampled and collected, and the dominance function is calculated. Update θ and w to complete the max-epochs agent policy update and output the optimal drag coefficient;
[0055] Specifically, it is stated as follows:
[0056] S5.1: Sample interaction information:
[0057] D = {τ1,τ2,τ3,…,τ} batch-epoch} (3)
[0058]
[0059] In the formula, D is the i-th trajectory set; τ is an interactive trajectory segment; batch-epoch is the batch training episode; S t A represents the environmental state at time t; t The action at time t corresponds to the ventilation drag coefficient; rew t Let π be the reward value at time t; θk (A t |S t Let t be the policy distribution at time t;
[0060] S5.2: Calculate the dominance function
[0061]
[0062]
[0063] In the formula, γ is the discount factor; steps is the plot step size; λ is the decay factor; rew t V(S) represents the reward value at time t. t ) represents the mapping value of the Critic-Net network;
[0064] S5.3: Update Actor-Net network parameters θ:
[0065]
[0066]
[0067]
[0068] In the formula, loss is the loss function; ε is the shearing factor; and lr is the learning rate.
[0069] S5.4: Update Critic-Net network parameters w:
[0070]
[0071]
[0072] S5.5: Update the network by max-epochs times and output the optimal training result.
[0073] In this example, the ventilation drag coefficient inversion results are shown below.
[0074] Table 1. Results of Ventilation Drag Coefficient Inversion
[0075] <![CDATA[e1]]> 0.08 <![CDATA[e2]]> 0.4563 <![CDATA[e3]]> 2 <![CDATA[e4]]> 6.5 <![CDATA[e5]]> 2 <![CDATA[e6]]> 4.0412 <![CDATA[e7]]> 3.3703 <![CDATA[e8]]> 4.4670 <![CDATA[e9]]> 9.9098 <![CDATA[e 10 ]]> 3 <![CDATA[e 11 ]]> 2.4952 <![CDATA[e 12 ]]> 2.2457 <![CDATA[e 13 ]]> 3 <![CDATA[e 14 ]]> 2.7633 <![CDATA[e 15 ]]> 0.012 <![CDATA[e 16 ]]> 1.2188 <![CDATA[e 17 ]]> 0.013
[0076] The air volume at the monitoring location was calculated based on the inversion results of the ventilation resistance coefficient, and the results are shown below.
[0077] Table 2 Calculation error of air volume
[0078]
[0079] The results above show that the ventilation resistance coefficient derived by the intelligent agent can meet the requirements of real-time ventilation network calculation. The calculation error and the monitoring error are controlled within about 5%. This invention can provide technical support for the construction of intelligent ventilation.
[0080] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope defined by the claims of the present invention.
Claims
1. A method for inverting ventilation drag coefficient based on deep reinforcement learning, characterized in that, The specific steps are as follows: S1: Define the agent's interaction environment env, where env = {MVSS, R, Q} target } where MVSS is the ventilation simulation system, and R is the reward function. To monitor airflow, rew = R(S) t A t ) indicates that in environmental state S t Take action A t The reward feedback value of the environment at time t-1 is the ventilation resistance coefficient r. t-1 ={r (t-1)1 ,r (t-1)2 ,r (t-1)3 ,r (t-1)n }, A t Let be the ventilation resistance coefficient output by the agent at time t; S2: Define the agent, which includes a policy neural network (Actor-Net) parameterized by θ and a value neural network (Critic-Net) parameterized by w. The Actor-Net determines the current environment state S... t Output ventilation resistance coefficient distribution π θ A t Sampling at π θ Critic-Net determines the current state S based on the current state S. t Output state value V t ; S3: Initialize agent parameters and learning training parameters. Agent parameters include network layer number (layers), number of neurons (sizes), and learning rate (lr). Training parameters include maximum number of epochs for agent updates (max-epochs), batch number of training epochs (batch-epoch), batch step size (steps), buffer size for experience retrieval, and discount factor γ. S4: Collect trajectory data of the agent's interaction with the environment, and collect <S,A,rews,π> generated during the interaction. θ (S)>, where S is the set of environmental states, A is the set of searched ventilation resistance coefficient vectors (i.e., actions), rews is the set of reward values, and π θ (S) represents the distribution of the policy neural network mapping under state S; S5: Sample and collect trajectory data, calculate the dominance function. Update θ and w to complete the max-epochs agent policy update and output the optimal drag coefficient.
2. The ventilation drag coefficient inversion method based on deep reinforcement learning according to claim 1, characterized in that: In step S1, the agent's interaction environment env is defined. Where MVSS is the ventilation simulation system, and R is the reward function. To monitor air volume, R(S) t A t ) indicates that in environmental state S t Take action A t The reward feedback value of the time environment, S t The distribution of ventilation drag coefficient r at time t-1 t-1 ={r (t-1)1 ,r (t-1)2 ,r (t-1)3 ,r (t-1)n }, A t Let be the ventilation resistance coefficient output by the agent at time t, and the reward function be defined as: In the formula, w i Weights are used for evaluation; metrics are evaluation indicators of the same dimension, used to evaluate... and The distance or error between them; For in A t The calculated air volume of the ventilation simulation system under the drag coefficient distribution corresponds to Air volume; α is the reward value scaling factor.
3. The ventilation drag coefficient inversion method based on deep reinforcement learning according to claim 1, characterized in that: In step S2, an agent is defined, comprising a policy neural network Actor-Net with θ parameters and a value neural network Critic-Net with w parameters. Actor-Net determines the current environment state S based on the current environment state. t Output ventilation resistance coefficient distribution π θ A t Sampling at π θ Critic-Net determines the current state S based on the current state S. t Output state value V t The neural network structure described is a multilayer perceptron (MLP) model, including an input layer, hidden layers, activation layers, and an output layer; the agent adjusts the current environmental state S... t Output the corresponding A t Defined as: In the formula, A t Ventilation resistance coefficient.
4. The ventilation drag coefficient inversion method based on deep reinforcement learning according to claim 1, characterized in that: In step S3, the agent parameters and learning training parameters are initialized. The agent parameters include the number of network layers (layers), the number of neurons (sizes), and the learning rate (lr). The training parameters include the maximum number of epochs for agent updates (max-epochs), the number of batch training epochs (batch-epoch), the step size of the batch training epoch (steps), the buffer size for experience retrieval (buffer-sizes), and the discount factor γ.
5. The ventilation drag coefficient inversion method based on deep reinforcement learning according to claim 1, characterized in that: In step S4, trajectory data of the agent's interaction with the environment are collected, and <S,A,rews,π> generated during the interaction process are collected. θ (S)>, where S is the set of environmental states, A is the set of searched ventilation resistance coefficients (i.e., actions), rews is the set of reward values, and π θ (S) represents the distribution of the policy neural network mapping under state S.
6. The ventilation drag coefficient inversion method based on deep reinforcement learning according to claim 1, characterized in that: In step S5, trajectory data is sampled and collected, the advantage function A is calculated, θ and w are updated, the agent policy update is completed in max-epochs, and the optimal drag coefficient is output. Specifically, it is stated as follows: S5.1: Sample interaction information once: D={τ1,τ2,τ3,…,τ batch-epoch } (3) In the formula, D is the i-th trajectory set; τ is an interactive trajectory segment; batch-epoch is the batch training episode; S t A represents the environmental state at time t; t The action at time t corresponds to the ventilation drag coefficient; rew t Let be the reward value at time t; Let be the policy distribution at time t; S5.2: Calculate the dominance function In the formula, γ is the discount factor; steps is the plot step size; λ is the decay factor; rew t V(S) represents the reward value at time t. t ) represents the mapping value of the Critic-Net network; S5.3: Update Actor-Net network parameters θ: In the formula, loss is the loss function; ε is the shearing factor; and lr is the learning rate. S5.4: Update Critic-Net network parameters w: S5.5: Update the network by max-epochs times and output the optimal training result.
Citation Information
Patent Citations
HEV energy management method based on deep reinforcement learning A3C algorithm
CN111731303A
Near-end strategy optimization method based on graph convolutional neural network
CN115983373A