Reinforcement learning-based training method for self-adaptive control model of dust removal system of fully-mechanized excavation face

By applying the adaptive control model of the comprehensive excavation surface dust removal system based on reinforcement learning in the mine, the problems of uneven dust distribution and difficulty in adaptive control are solved, and efficient dust removal and intelligent management are achieved.

CN120103707AActive Publication Date: 2025-06-06ZAOZHUANG MINING GRP CO LTD

Patent Information

Application Number
CN202510255498.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-06
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

In mines, the dust distribution of comprehensive excavation surfaces is uneven, and it is difficult for the existing technology to achieve adaptive control, resulting in poor dust removal effect.

Method used

The training method of the adaptive control model of the comprehensive surface dust removal system based on reinforcement learning is adopted. By building a physical model and constructing a reinforcement learning model, the model is trained using dust sensor data to achieve dynamic adjustment of control strategies.

Benefits of technology

It effectively improves dust removal efficiency, reduces equipment power consumption and manual intervention, and promotes the intelligent development of coal mines.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120103707A_ABST
    Figure CN120103707A_ABST
Patent Text Reader

Abstract

The invention discloses a training method for a self-adaptive control model of a fully-mechanized excavation face dust removal system based on reinforcement learning. The training method comprises the following steps: building a fully-mechanized excavation face physical model, generating dust at an excavation point of the fully-mechanized excavation face physical model, capturing dust concentration of each position of a fully-mechanized excavation face by using a dust sensor, and obtaining dust distribution data; and a reinforcement learning model is constructed, the reinforcement learning model is trained according to the obtained dust distribution data, and the trained reinforcement learning model is used for adaptively controlling the ventilation system. A large amount of training data can be generated, the self-adaptive control ventilation dust removal model is obtained through training, the dust removal efficiency is effectively improved, equipment power consumption and manual intervention are effectively reduced, and intelligent development of a coal mine is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of mine ventilation and dust removal, and in particular relates to a training method for an adaptive control model of a comprehensive excavation face dust removal system based on reinforcement learning. Background Art

[0002] Reinforcement learning is a branch of machine learning. It refers to the process of gradually updating and optimizing the training of intelligent agents through continuous interaction with the environment, using actions, rewards and observations. In this process, the intelligent agent continuously learns from the interaction with the environment and adjusts its behavior strategy to achieve a higher level of intelligence. The current mine dust removal system can collect dust concentration data through dust concentration sensors, dynamically adjust the fan power and spray water volume to reduce dust concentration. However, in the comprehensive excavation face, the distribution of dust is extremely uneven, and the dust concentration captured by each sensor varies greatly, so the control strategy cannot be determined by a single parameter. Moreover, with the widespread application of the "long pressure and short extraction" ventilation and dust removal system in the excavation working face, there are multiple control variables that are coupled with each other, such as extraction and compression power and air outlet position, and the dust removal effect cannot be improved by simply increasing the extraction and compression air volume. Therefore, it is difficult to achieve adaptive control of the dust removal system of the excavation face. In view of the dynamically changing environment, the reinforcement learning model can continuously adjust the control strategy according to the real-time changes of the dust concentration in the comprehensive excavation face, and finally obtain an ideal model, which can reduce the dust concentration to the safe limit in the shortest time for various environmental conditions. However, the training of the reinforcement learning model requires a large amount of dynamic change data of the environment and equipment. At present, there is no device and method that can train the adaptive control model of the comprehensive excavation face dust removal system based on reinforcement learning. Summary of the invention

[0003] In order to solve the above technical problems, the present invention proposes a training method for an adaptive control model of a comprehensive excavation face dust removal system based on reinforcement learning, which can generate a large amount of training data. Through training, an adaptive control ventilation and dust removal model is obtained, which can effectively improve the dust removal efficiency, effectively reduce equipment power consumption and manual intervention, and promote the intelligent development of coal mines.

[0004] To achieve the above object, the present invention provides a training method for an adaptive control model of a comprehensive excavation face dust removal system based on reinforcement learning, comprising:

[0005] A physical model of a fully excavated face is constructed, dust is generated at the excavation point of the physical model of the fully excavated face, and dust concentrations at various positions of the fully excavated face are captured using dust sensors to obtain dust distribution data;

[0006] A reinforcement learning model is constructed, and the reinforcement learning model is trained according to the obtained dust distribution data. The trained reinforcement learning model is used for adaptively controlling the ventilation system.

[0007] Optionally, building a physical model of the comprehensive excavation face includes:

[0008] The physical model of the comprehensive excavation face is divided into three surfaces: the excavation face, the tunnel wall and the model opening. The excavation face is sealed and connected to the tunnel wall. The right side of the model opening is in an open state. The bottom side length of the tunnel is a preset length. The excavation face is composed of an upper arch and a lower rectangle. The physical model of the comprehensive excavation face includes motors, fans, dust generators, dust sensors and conventional suction and pressure blowers that are replaced in equal proportion. Motors, fans and dust sensors that support networking or are connected to the network through interface modules are selected.

[0009] Optionally, generating dust at the excavation point of the physical model of the comprehensive excavation face includes:

[0010] The outlet of the pressure-in type air duct is installed facing the excavation face, which is used to input fresh air into the excavation area and cover the dust source. The outlet of the exhaust type air duct is located on the right side of the model opening end, which is used to discharge the dust. The exhaust and pressure pipes are installed above the platform bracket and are connected to the exhaust and pressure fans respectively. The control end of the fan is connected to the host to control the ventilation equipment.

[0011] A number of holes are opened on the excavation face of the physical model of the comprehensive excavation face as dust inlets of the dust generator, dust generators are installed corresponding to the dust inlets, the output ports of the dust generators are connected to pipelines, and the pipelines are connected to the excavation face; dust is transmitted through the number of holes on the excavation face, and the dust generation is completed in combination with the control of the ventilation equipment.

[0012] Optionally, installing the dust sensors includes: evenly distributing the dust sensors above the brackets in an array according to the area and characteristics of the comprehensive excavation face.

[0013] Optionally, the dust concentration at each position of the fully mechanized excavation face captured by the dust sensor includes:

[0014]

[0015] Where, C is the overall dust concentration; C i is the dust concentration measured by the i-th sensor; W i is the weight of the i-th sensor.

[0016] Optionally, construct a reinforcement learning model including: an input layer, a hidden layer, and an output layer;

[0017] The input layer is used to receive the current state, which includes dust concentration, ambient temperature, humidity, wind speed and power data of each device;

[0018] The hidden layer is used to set a number of fully connected layers;

[0019] The output layer is used to output the expected reward of taking an action in the current state according to each action of the system, and the action includes adjusting the exhaust air volume, the compressed air volume and the air outlet position.

[0020] Optionally, before training the reinforcement learning model according to the obtained dust distribution data, it is necessary to design a strategy to select the optimal action and update the expected reward of taking the action in the current state according to the reinforcement learning model.

[0021] Optionally, design strategies to select optimal actions including:

[0022] a * =max a Q(s,a);

[0023] Where s is each state; a is the action selected by the agent with the maximum Q value; Q is the expected return of taking an action in the current state; a * is the optimal action selected under the current state s.

[0024] Optionally, updating the expected reward of taking an action in the current state according to the reinforcement learning model includes:

[0025] Q(s,a)=(1-λ t )[R+γmax a Q(s′, a′);

[0026] Where Q(s, a) is the total reward expected from taking action a in state s; s is each state; a is the action selected by the agent with the maximum Q value; λ t is the learning rate; R is the immediate reward; γ is the discount factor; max a Q(s′, a′) is the maximum Q value among all possible actions a′ in the next state s′; Q is the expected return of taking an action in the current state.

[0027] Optionally, training the reinforcement learning model according to the obtained dust distribution data includes:

[0028] Calculate the reward value of each data sample according to the set reward function, add the experience data with the reward value to the playback buffer for the early training of the model and generate a preliminary dust removal control strategy;

[0029] Using the experience replay mechanism according to the reinforcement learning model, the experience data with reward values ​​are stored in the replay buffer;

[0030] Calculate the Q value of each action in the current state and update the Q value by minimizing the loss function, where Q is the expected return of taking the action in the current state;

[0031] When the loss function reaches a predetermined convergence condition or a set threshold, the training is stopped and the trained reinforcement learning model is output for adaptively controlling the ventilation system.

[0032] Technical effect of the invention: The invention discloses a training method for an adaptive control model of a comprehensive excavation face dust removal system based on reinforcement learning, which can generate a large amount of training data. An adaptive control ventilation and dust removal model is obtained through training, which effectively improves the dust removal efficiency, effectively reduces equipment power consumption and manual intervention, and promotes the intelligent development of coal mines. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0034] Figure 1 A flow chart of a method for training an adaptive control model of a fully mechanized excavation face dust removal system based on reinforcement learning according to an embodiment of the present invention;

[0035] Figure 2 A schematic diagram of the structure of a physical model of a comprehensive excavation face constructed for an embodiment of the present invention;

[0036] Figure 3 This is a schematic diagram of the installation of the sensor bracket according to an embodiment of the present invention;

[0037] Figure 4 The figure is a schematic diagram of the topological structure of the network according to the embodiment of the present invention. DETAILED DESCRIPTION

[0038] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0039] It should be noted that the steps shown in the flowcharts of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and that, although a logical order is shown in the flowcharts, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0040] In terms of intelligent dust removal, the existing technologies include tunneling systems, long-pressure and short-extraction intelligent frequency conversion control systems, two-stage air curtain generation systems based on Laval-like nozzles, wet dust collector systems and intelligent control systems. A two-stage pneumatic air curtain system is used to isolate dust and a wet dust collector system is used to reduce dust, effectively preventing dust diffusion and solving the problem of poor dust prevention effect of a single technology. At the same time, intelligent control technology is used to achieve dynamic regulation of air volume and dust concentration, thereby improving production safety. For dynamically changing environments, the next control strategy is obtained by reconstructing the simulation model and performing calculations and analysis. This control method requires very high computing resources, and the simulation time is long, and the real-time performance is not strong. By building a training platform and training a reinforcement learning model, it is a control method based on data-driven rather than physical models. Once the model is trained, the system can generate adaptive strategies and actions in real time according to environmental changes, and can effectively control the dust concentration on the tunneling surface in a timely manner. The prior art also includes an image collector for collecting images of the excavation working face, an environmental monitoring module for real-time collection of environmental data of the excavation working face, an operation monitoring module for real-time monitoring of local ventilator operation data, and an operation optimization module. Time series prediction is performed based on environmental data and operation data to obtain a predicted value of required air volume. A convolutional neural network is constructed based on a combination of the image, environmental data and operation data of the excavation working face to obtain a required air volume auxiliary value. The required air volume control value is calculated by combining the required air volume prediction value and the required air volume auxiliary value. The controller is used to adjust the local ventilator according to the required air volume control value calculated by the operation optimization module. The air adjustment system has good scalability and flexibility, can adapt to the required air volume prediction needs of coal mines of different types and sizes, and has important application value and market prospects. The disclosure provides an operation optimization module, which performs time series prediction based on environmental data and operation data to obtain a predicted value of required air volume. A convolutional neural network is constructed based on a combination of the image, environmental data and operation data of the excavation working face to obtain a required air volume auxiliary value. The required air volume control value is calculated by combining the required air volume prediction value and the required air volume auxiliary value. The air volume control method provided in the disclosure is to predict the required air volume through deep learning and then control the air volume, which is different from the control principle of the present invention using reinforcement learning. In addition, the local ventilator controlled by the disclosure cannot control a multivariable system such as a "long pressure and short extraction" dust removal system.

[0041] like Figure 1 As shown, in this embodiment, a training method for an adaptive control model of a comprehensive excavation face dust removal system based on reinforcement learning is provided, comprising:

[0042] A physical model of a fully excavated face is constructed, dust is generated at the excavation point of the physical model of the fully excavated face, and dust concentrations at various positions of the fully excavated face are captured using dust sensors to obtain dust distribution data;

[0043] A reinforcement learning model is constructed, and the reinforcement learning model is trained according to the obtained dust distribution data. The trained reinforcement learning model is used for adaptively controlling the ventilation system.

[0044] Further, such as Figure 2 The physical model of the comprehensive excavation face shown includes:

[0045] The physical model of the comprehensive excavation face is divided into three surfaces: the excavation face, the tunnel wall and the model opening. The excavation face is sealed and connected to the tunnel wall. The right side of the model opening is in an open state, wherein the bottom side length of the tunnel is a preset length, and the excavation face is composed of an upper half arch and a lower half rectangle. The physical model of the comprehensive excavation face includes motors, fans, dust generators, dust sensors and proportional replacements for traditional suction and pressure air ducts, wherein motors, fans and dust sensors that support networking or are connected to the network through an interface module are selected; wherein A is the excavation face, B is the tunnel wall, and C is the model opening.

[0046] Furthermore, generating dust at the excavation point of the physical model of the comprehensive excavation face includes:

[0047] The outlet of the pressure-in type air duct is installed facing the excavation face, which is used to input fresh air into the excavation area and cover the dust source. The outlet of the exhaust type air duct is located on the right side of the model opening end, which is used to discharge the dust. The exhaust and pressure pipes are installed above the platform bracket and are connected to the exhaust and pressure fans respectively. The control end of the fan is connected to the host to control the ventilation equipment.

[0048] A number of holes are opened on the excavation face of the physical model of the comprehensive excavation face as dust inlets of the dust generator, dust generators are installed corresponding to the dust inlets, the output ports of the dust generators are connected to pipelines, and the pipelines are connected to the excavation face; dust is transmitted through the number of holes on the excavation face, and the dust generation is completed in combination with the control of the ventilation equipment.

[0049] Further, the dust sensors are installed corresponding to the dust generating inlets, including: according to the area and characteristics of the comprehensive excavation face, the dust sensors are evenly distributed above the bracket in an array, such as Figure 3 shown.

[0050] Specifically, according to the area and characteristics of the comprehensive excavation face, the sensors are evenly distributed above the bracket in an array to ensure coverage of high-risk areas and areas near dust sources.

[0051] Use a random function to generate dust generation rate values ​​to ensure that the dust generation rate fluctuates within the set range:

[0052] R(t)=R min +(R max -R min )·rand;

[0053] Where R(t) is the dust generation rate at the current time t; R min and R max are the minimum and maximum values ​​of the dust generation rate respectively; rand() generates a random number in the range of [0,1].

[0054] At the same time, the system is set to update R(t) every 1 minute to simulate a dynamic dust-generating environment. At regular intervals, the dust-generating port connection point of the dust generator is manually replaced to switch the pipeline distribution. This method is used to simulate the diversity of dust distribution in the excavation section, such as different high-concentration and low-concentration areas.

[0055] Start the ventilation and dust generating equipment, and transmit the data of dust concentration on the comprehensive excavation surface collected by the sensor to the host computer through wireless or wired network for subsequent processing and analysis.

[0056] Furthermore, the dust concentration at each location of the fully mechanized excavation face is captured using dust sensors, including:

[0057]

[0058] Where, C is the overall dust concentration; C i is the dust concentration measured by the i-th sensor; W i is the weight of the i-th sensor.

[0059] Specifically, according to the importance of the location of each sensor (the weight near the dust generator is large, and the weight far away from the dust generator is small), the dust concentration collected by each sensor is weighted averaged to obtain the dust concentration of the entire comprehensive excavation face.

[0060] Further, such as Figure 4 The reinforcement learning model shown includes: input layer, hidden layer and output layer;

[0061] The input layer is used to receive the current state, which includes dust concentration, ambient temperature, humidity, wind speed and power data of each device;

[0062] The hidden layer is used to set a number of fully connected layers;

[0063] The output layer is used to output the expected reward of taking an action in the current state according to each action of the system, and the action includes adjusting the exhaust air volume, the compressed air volume and the air outlet position.

[0064] Furthermore, before training the reinforcement learning model according to the obtained dust distribution data, it is necessary to design a strategy to select the optimal action and update the expected reward of taking the action in the current state according to the reinforcement learning model.

[0065] Furthermore, the design strategy for selecting the optimal action includes:

[0066] a * =max a Q(s,a);

[0067] Where s is each state; a is the action selected by the agent with the maximum Q value; Q is the expected return of taking an action in the current state; a * is the optimal action selected under the current state s.

[0068] Furthermore, updating the expected reward of taking an action in the current state according to the reinforcement learning model includes:

[0069] Q(s,a)=(1-λ t )[R+γmax a Q(s′, a′)];

[0070] Where Q(s,a) is the total reward expected from taking action a in state s; s is each state; a is the action selected by the agent with the maximum Q value; λ t is the learning rate; R is the immediate reward; γ is the discount factor; max a Q(s′, a′) is the maximum Q value among all possible actions a′ in the next state s′; Q is the expected return of taking an action in the current state.

[0071] The reward function is set based on the magnitude of the reduction in dust concentration. The faster the dust concentration decreases before reaching the safety threshold, the higher the reward.

[0072] Set the reward value:

[0073]

[0074] Where C is the current dust concentration; ΔC is the decrease in dust concentration per unit time, ΔC = C prev -C current ; W c is the weight of dust concentration. When the dust concentration reaches the safety threshold C safe Finally, the reward is fixed at zero to prevent over-adjustment.

[0075] Further, training the reinforcement learning model according to the obtained dust distribution data includes:

[0076] Calculate the reward value of each data sample according to the set reward function, add the experience data with the reward value to the playback buffer for the early training of the model and generate a preliminary dust removal control strategy;

[0077] Using the experience replay mechanism according to the reinforcement learning model, the experience data with reward values ​​are stored in the replay buffer;

[0078] Calculate the Q value of each action in the current state and update the Q value by minimizing the loss function, where Q is the expected return of taking the action in the current state;

[0079] Among them, the loss function is:

[0080] L(θ)=E[(R+γmax a∈A Q(s t+1 ,q t+1 ;θ′)-Q(s t , a t ;θ)) 2 ];

[0081] Among them, L(θ) is the loss function; E is the expected value; R is the immediate reward; γ is the discount factor; max a∈A Search for all possible actions a in the action space A; s is the action with the maximum Q value selected by the agent; Q is the expected return of taking the action in the current state; A is the action space (the set of all possible actions); Q(s t+1 , a t+1 ; θ′) is the state s of the target Q network for the next time step t+1 and action a t+1 Q value estimation of s t+1 To perform action a t The next state to be transferred to; a t+1 For state s t+1 The action selected by the agent under the condition; θ′ is the parameter of the target network; Q(s t , a t ;θ) 2 is the current Q network state s t and action a t Q value estimation of s t is the current state; a t For state a t The action selected next; θ is the parameter of the Q network.

[0082] When the loss function reaches a predetermined convergence condition or a set threshold, the training is stopped and the trained reinforcement learning model is output for adaptively controlling the ventilation system.

[0083] The present invention also discloses a training system for an adaptive control model of a comprehensive excavation face dust removal system based on reinforcement learning, including a dust generation module, a data acquisition module, a data transmission module, a reinforcement learning module and a comprehensive excavation face control module, wherein: the dust generation module is composed of a dust generator, a connecting pipe, and a dust generating port, and realizes dust generation in the area where the comprehensive excavation face is located by simulating the actual operation of the comprehensive excavation face, including dust of different concentrations and various dust distribution conditions. The data acquisition module is composed of a dust sensor, a host computer, and equipment data acquisition software, and is used to collect data in the area and transmit the data to the processing end. The reinforcement learning module, based on the DQN model network architecture, accepts the current state, and uses the reward function and the experience playback mechanism to perform reinforcement learning to realize the adaptive control of the dust of the comprehensive excavation face. The comprehensive excavation face control module includes a control program, a host computer, an exhaust fan, a compressed air fan, a wind tube, a motor, etc. The comprehensive excavation face control module makes real-time decisions according to the environmental state of the current physical model, including adjusting parameters such as exhaust power, compressed air power, and air outlet position, to realize intelligent dust removal control of the comprehensive excavation face.

[0084] The reinforcement learning control method adopted by the present invention dynamically controls the key equipment of the dust removal system through a reinforcement learning algorithm, including adjusting the power of each fan, the distance between the air outlet and the excavation face, and other operations. Adaptive regulation of ventilation and dust removal of underground comprehensive excavation faces is achieved. It is possible to achieve precise control of dust concentration under dynamic and non-steady-state conditions, effectively reduce dust concentration, and improve dust removal efficiency. The reinforcement learning training method adopted by the present invention inputs the real-time environmental status (dust concentration, temperature and humidity, wind speed, etc.) and equipment operating parameters into a reinforcement learning algorithm based on a DQN structure, and uses an experience replay mechanism to train the DQN model. The DQN reinforcement learning model outputs the optimal control action in the current state according to the current real-time environment. The model dynamically adjusts the ventilation equipment parameters, optimizes the dust control strategy, and effectively reduces the dust concentration in the excavation face area.

[0085] The above are only preferred specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

Claims

1. A training method for an adaptive control model of a fully mechanized excavation face dust removal system based on reinforcement learning, characterized in that: include: A physical model of a fully excavated face is constructed, dust is generated at the excavation point of the physical model of the fully excavated face, and dust concentrations at various positions of the fully excavated face are captured using dust sensors to obtain dust distribution data; A reinforcement learning model is constructed, and the reinforcement learning model is trained according to the obtained dust distribution data. The trained reinforcement learning model is used for adaptively controlling the ventilation system.

2. The training method of the adaptive control model of the comprehensive excavation face dust removal system based on reinforcement learning as claimed in claim 1 is characterized in that: Building a physical model of a comprehensive excavation face includes: The physical model of the comprehensive excavation face is divided into three surfaces: the excavation face, the tunnel wall and the model opening. The excavation face is sealed and connected to the tunnel wall. The right side of the model opening is in an open state. The bottom side length of the tunnel is a preset length. The excavation face is composed of an upper arch and a lower rectangle. The physical model of the comprehensive excavation face includes motors, fans, dust generators, dust sensors and conventional suction and pressure blowers that are replaced in equal proportion. Motors, fans and dust sensors that support networking or are connected to the network through interface modules are selected.

3. The training method of the adaptive control model of the comprehensive excavation face dust removal system based on reinforcement learning as claimed in claim 1 is characterized in that: The generation of dust at the excavation point of the physical model of the comprehensive excavation face includes: The outlet of the pressure-in type air duct is installed facing the excavation face, which is used to input fresh air into the excavation area and cover the dust source. The outlet of the exhaust type air duct is located on the right side of the model opening end, which is used to discharge the dust. The exhaust and pressure pipes are installed above the platform bracket and are connected to the exhaust and pressure fans respectively. The control end of the fan is connected to the host to control the ventilation equipment. A number of holes are opened on the excavation face of the physical model of the comprehensive excavation face as dust inlets of the dust generator, dust generators are installed corresponding to the dust inlets, the output ports of the dust generators are connected to pipelines, and the pipelines are connected to the excavation face; dust is transmitted through the number of holes on the excavation face, and the dust generation is completed in combination with the control of the ventilation equipment.

4. The training method of the adaptive control model of the comprehensive excavation face dust removal system based on reinforcement learning as claimed in claim 3 is characterized in that: The installation of dust sensors includes: according to the area and characteristics of the comprehensive excavation face, the dust sensors are evenly distributed above the bracket in an array.

5. The training method of the adaptive control model of the comprehensive excavation face dust removal system based on reinforcement learning as claimed in claim 1 is characterized in that: The dust concentration at each location of the fully mechanized excavation face is captured by dust sensors, including: Where, C is the overall dust concentration; C i is the dust concentration measured by the i-th sensor; W i is the weight of the i-th sensor.

6. The training method of the adaptive control model of the comprehensive excavation face dust removal system based on reinforcement learning as claimed in claim 1 is characterized in that: Building a reinforcement learning model includes: input layer, hidden layer and output layer; The input layer is used to receive the current state, which includes dust concentration, ambient temperature, humidity, wind speed and power data of each device; The hidden layer is used to set a number of fully connected layers; The output layer is used to output the expected reward of taking an action in the current state according to each action of the system, and the action includes adjusting the exhaust air volume, the compressed air volume and the air outlet position.

7. The training method of the adaptive control model of the comprehensive excavation face dust removal system based on reinforcement learning as claimed in claim 1 is characterized in that: Before training the reinforcement learning model according to the obtained dust distribution data, it is necessary to design a strategy to select the optimal action and update the expected return of taking the action in the current state according to the reinforcement learning model.

8. The training method of the adaptive control model of the comprehensive excavation face dust removal system based on reinforcement learning as claimed in claim 7 is characterized in that: Designing strategies to select optimal actions includes: a * =max a Q(s,a); Where s is each state; a is the action selected by the agent with the maximum Q value; Q is the expected return of taking an action in the current state; a * is the optimal action selected under the current state s.

9. The training method of the adaptive control model of the comprehensive excavation face dust removal system based on reinforcement learning as claimed in claim 7 is characterized in that: The expected return of taking an action in the current state according to the reinforcement learning model is updated as follows: Q(s,a)=(1-λ t )[R+γmax a Q(s′, a′)]; Where Q(s, a) is the total reward expected from taking action a in state s; s is each state; a is the action selected by the agent with the maximum Q value; λ t is the learning rate; R is the immediate reward; γ is the discount factor; max a Q(s′, a′) is the maximum Q value among all possible actions a′ in the next state s′; Q is the expected return of taking an action in the current state.

10. The training method of the adaptive control model of the fully mechanized excavation face dust removal system based on reinforcement learning as claimed in claim 7, characterized in that: Training the reinforcement learning model according to the obtained dust distribution data includes: Calculate the reward value of each data sample according to the set reward function, add the experience data with the reward value to the playback buffer for the early training of the model and generate a preliminary dust removal control strategy; Using the experience replay mechanism according to the reinforcement learning model, the experience data with reward values ​​are stored in the replay buffer; Calculate the Q value of each action in the current state and update the Q value by minimizing the loss function, where Q is the expected return of taking the action in the current state; When the loss function reaches a predetermined convergence condition or a set threshold, the training is stopped and the trained reinforcement learning model is output for adaptively controlling the ventilation system.

Citation Information

Patent Citations

  • Parameterization analogue simulation method for dust, gas and air flow on fully mechanized excavation face of coal mine

    CN111027252A

  • Dust cleaning system based on artificial intelligence

    CN113230801A

  • Ventilation and dust reduction method and device based on driving working face

    CN115342079A

  • Roadway fully-mechanized excavation face multi-stage collaborative dust fall simulation experiment method

    CN115749916A

  • Intelligent ventilation dust control and removal system for coal mine fully-mechanized excavation face based on optimal coordination mechanism

    CN119102722A

Cited By

  • Tunneling roadway ventilation-dust control cooperative adjustment method based on multi-mode intelligent control

    CN121785094A