Intelligent control method and system for gas extraction process
Through multi-source sensing data and deep reinforcement learning control combined with well-lane coupled physical model, intelligent control of the gas extraction process is realized, solving the problems of insufficient self-optimization of parameters and manual dependence in the existing technology, and improving extraction efficiency and system stability.
Patent Information
- Application Number
- CN202510737156.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-04
AI Technical Summary
The existing gas extraction technology lacks self-optimization of parameters during the long-term extraction stage, and the rough control cannot adapt to the rapid changes in geological and gas enrichment states, lack of prediction capabilities and high artificial dependence, resulting in low efficiency and safety hazards.
Multi-source sensing data is used to combine well-lane coupled physical model and deep reinforcement learning control to achieve real-time prediction and adaptive regulation of gas concentration and extraction efficiency, preventive maintenance through equipment health index, and maintenance work orders are generated.
It improves gas extraction efficiency, reduces energy consumption, and keeps the gas concentration within the safe range to ensure stable operation and safety of the system.
Smart Images

Figure CN120331861A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of gas drainage, and particularly to an intelligent control method and system for the gas extraction process. Background Art
[0002] The gas extraction (also known as gas drainage) in coal mines has developed from the early pure manual negative pressure extraction to the current semi-automatic mode of "threshold-trigger + manual intervention". A typical solution (such as the prior art CN 118167404B) collects data such as gas concentration, extraction rate, and pressure, sets intervention thresholds, determines whether manual intervention is required, and disposes of abnormalities in the "drilling" link.
[0003] Although the existing documents have introduced the "manual + intelligent" hybrid control, there are still the following problems:
[0004] (1) Limited scope - only covering the "drilling" process, lacking parameter self-optimization for the subsequent long-term extraction stage;
[0005] (2) Coarse control - the key control quantities (such as negative pressure and valve opening) are triggered by fixed thresholds and cannot adapt to the rapid changes in geological and gas enrichment states;
[0006] (3) Insufficient prediction ability - without a digital twin or machine learning model, unable to give early warnings of over-limit trends and equipment failures;
[0007] (4) High dependence on manual labor - complex terrain and multi-branch pipe networks still require personnel to make decisions, reducing efficiency and bringing potential safety hazards. Summary of the Invention
[0008] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide an intelligent control method and system for the gas extraction process, which realizes the intelligent monitoring and preventive maintenance of the equipment state, and effectively guarantees the safety and stability of the system operation.
[0009] To achieve the above object, the present invention provides the following solutions:
[0010] An intelligent control method for the gas extraction process, comprising:
[0011] Deploying a plurality of environment-working condition integrated sensing nodes in the target gas extraction area, and synchronously collecting current multi-source sensing data at a preset period; the multi-source sensing data includes the current gas concentration value, the current extraction flow rate value, the current negative pressure value, the pipe network temperature value, and the equipment vibration-electrical operation data;
[0012] Online correcting the multi-source sensing data according to the well roadway fluid-particle coupling physical model and historical extraction samples to obtain a predicted gas concentration value and a predicted extraction efficiency value;
[0013] Construct a working condition state vector by combining the predicted gas concentration value, the predicted drainage efficiency value, the current negative pressure value, and the current drainage flow rate value, input it into the deep reinforcement learning control model, obtain control instructions including the negative pressure adjustment amount, the valve opening adjustment amount, and the drainage flow rate adjustment amount, and continuously and adaptively optimize the deep reinforcement learning control model based on a comprehensive reward function including drainage efficiency, concentration deviation, and energy consumption factor;
[0014] Send the control instructions to the drainage execution unit, and collect new gas concentration values, drainage flow rate values, and negative pressure values through the drainage execution unit;
[0015] Write the new gas concentration value, drainage flow rate value, negative pressure value, and the working condition state vector into the deep reinforcement learning control model together to form a closed-loop self-learning;
[0016] Calculate the equipment health index based on the equipment vibration - electrical operation data, and automatically generate a maintenance work order when the health index is lower than a preset threshold.
[0017] Preferably, the drainage execution unit includes: a negative pressure regulating pump, a branch electric valve, and a variable frequency fan.
[0018] Preferably, online correction is performed on the multi-source sensing data according to the roadway fluid - particle coupling physical model and historical extraction samples to obtain the predicted gas concentration value and the predicted drainage efficiency value, including:
[0019] Construct a coupling simulation model that describes the gas flow, diffusion, and particle perturbation effects in the roadway space of the target gas extraction area;
[0020] Use the coupling simulation model to perform a preliminary prediction on the currently collected multi-source sensing data to obtain the physical model predicted concentration value and the predicted drainage efficiency value;
[0021] Construct a residual set based on the deviation between the physical model predicted value and the preset historical extraction sample data;
[0022] Use a pre-trained residual correction neural network model to perform a regression prediction on the residual set to obtain a residual correction value;
[0023] Perform weighted fusion on the physical model predicted value and the residual correction value to obtain the predicted gas concentration value.
[0024] Preferably, the expression of the roadway fluid - particle coupling physical model is:
[0025]
[0026] Where ρ gis the density of the gas; u g is the velocity vector of the gas; p is the gas pressure; μ g is the dynamic viscosity of the gas; β is the gas-solid momentum coupling coefficient; φ p is the particle volume fraction; u p is the particle velocity vector; ε is the porosity; C g is the volume fraction concentration of the gas; D eff is the equivalent diffusion coefficient; K d is the gas adsorption-desorption kinetic coefficient; C eq is the gas equilibrium concentration; β c is the concentration coupling coefficient; ‖·‖ represents the vector norm; η pred is the predicted drainage efficiency; Q calc is the drainage flow rate calculated by the model; Q leak is the leakage flow rate; x is the concentration attenuation coefficient; C target is the target control concentration; is the gradient operator; t is the time.
[0027] Preferably, the calculation formula of the residual correction value is:
[0028]
[0029] where r t is the residual vector at the current time step; w λ is the weight vector of the gating layer; b λ is the bias of the gating layer; σ(·) is the Sigmoid function, used to output the gating coefficient in the range of 0-1; W1 is the weight matrix of the first fully connected layer; b1 is the bias of the first layer; φ(·) is the non-linear activation function; W2 is the weight matrix of the second fully connected layer; b2 is the bias of the second layer; Δy corr (t) is the residual correction value at the current time step.
[0030] Preferably, the predicted gas concentration value, the predicted drainage efficiency value, the current negative pressure value and the current drainage flow rate value are jointly used to form a working condition state vector, which is input into the deep reinforcement learning control model to obtain control instructions including the negative pressure adjustment amount, the valve opening adjustment amount and the drainage flow rate adjustment amount, and based on a comprehensive reward function including drainage efficiency, concentration deviation and energy consumption factor, the deep reinforcement learning control model is continuously and adaptively optimized, including:
[0031] Perform zero-mean normalization on the predicted gas concentration value, the predicted drainage efficiency value, the current negative pressure value and the current drainage flow rate value, and splice them in a predetermined order to form the four-dimensional working condition state vector;
[0032] Input the working condition state vector into the deep reinforcement learning control model including a policy network and a value network to obtain a three-dimensional adjustment amount corresponding to negative pressure, valve opening, and drainage flow rate;
[0033] Encapsulate the three-dimensional adjustment amount into a control instruction according to the industrial fieldbus communication protocol;
[0034] After the execution of the control instruction, calculate the comprehensive reward value in real time according to the comprehensive reward function R = w e (η new -η old )-w c |C new -C target |-w p P elec ; where R is the comprehensive reward value; w e is the drainage efficiency weight; w c is the concentration deviation penalty weight; w p is the energy consumption penalty weight; η new and η old are the drainage efficiencies after and before adjustment respectively; C new is the adjusted gas concentration value; C target is the target control concentration; P elec is the electric power of the drainage equipment;
[0035] Use the temporal difference error Δ = R + dV(s t+1 ) - V(s t ) to perform backpropagation and Adam optimization on the weights of the policy network and the value network to achieve continuous adaptive adjustment of the deep reinforcement learning control model; where Δ is the temporal difference error; d is the discount factor; V(s t ) and V(s t+1 ) are the valuations of the value network for the current state and the next state respectively.
[0036] Preferably, the policy network is composed of two layers with 128 rectified linear units each.
[0037] Preferably, write the new gas concentration value, drainage flow rate value, negative pressure value, and the working condition state vector into the deep reinforcement learning control model together to form a closed-loop self-learning, including:
[0038] Collect the new gas concentration value, drainage flow rate value, and negative pressure value after the execution of the control instruction, and splice them together with the predicted drainage efficiency value in a linear normalization manner to form the next moment working condition state vector s t+1 ;
[0039] Write the current moment state vector s t , the executed three-dimensional adjustment amount at , the immediate comprehensive reward R t and the state vector s at the next moment t+1 are combined into a transition quadruple <s t , a t , R t , s t+1 >;
[0040] Write the quadruple into the cyclic experience replay pool, and eliminate the oldest data according to the first-in, first-out strategy, keeping the memory pool capacity at N;
[0041] Randomly select m quadruples from the cyclic experience replay pool in the priority sampling manner to form a training batch
[0042] For each sample in the training batch , calculate the temporal difference error according to the formula , and use as the loss function, and use the Adam algorithm with the learning rate α lr to update the policy network parameter θ π and the value network parameter θ v simultaneously to achieve online self-learning; where d is the discount factor, is the output of the value network, δ is the temporal difference error, L is the mean squared error loss, and m is the batch size for each training
[0043] Every K times of parameter updates, execute with the soft update coefficient τ∈(0,1) to improve the stability of the training process and prevent policy divergence; where is the target value network parameter, and K is the soft update interval.
[0044] Preferably, calculate the equipment health index based on the equipment vibration-electrical operation data, and when the health index is lower than the preset threshold, automatically generate a maintenance work order, including:
[0045] Perform a sliding window analysis on the equipment vibration-electrical operation data, and extract the effective value of the axial vibration velocity V rms , the root mean square value of the operating current I rms , and the time average value of the stator winding temperature T avg ;
[0046] Perform linear normalization on V rms , I rms and T avg according to their respective historical extreme value intervals to obtain dimensionless vibration characteristics, dimensionless current characteristics, and dimensionless temperature characteristics;
[0047] According to the formula Calculate the health index of the computing device; wherein, and are dimensionless vibration characteristics, dimensionless current characteristics, and dimensionless temperature characteristics respectively; w v , w i and w t are the weight coefficients of the dimensionless vibration characteristic, the dimensionless current characteristic, and the dimensionless temperature characteristic respectively, satisfying w v + w i + w t = 1;
[0048] Compare the calculated H idx with the preset health threshold H thr ; When H idx ≤ H thr , trigger the preset maintenance process and generate the maintenance work order.
[0049] An intelligent control system for the gas extraction process, comprising:
[0050] A multi-source sensing and acquisition module, used to deploy multiple environment-operation integrated sensing nodes in the target gas extraction area and synchronously collect the current multi-source sensing data according to a preset period; the multi-source sensing data includes the current gas concentration value, the current extraction flow value, the current negative pressure value, the pipeline network temperature value, and the equipment vibration-electrical operation data;
[0051] A coupling modeling and prediction module, used to perform online correction on the multi-source sensing data according to the wellbore fluid-particle coupling physical model and historical extraction samples to obtain the predicted gas concentration value and the predicted extraction efficiency value;
[0052] A working condition evaluation and control decision-making module, used to jointly form a working condition state vector with the predicted gas concentration value, the predicted extraction efficiency value, the current negative pressure value, and the current extraction flow value, input it into the deep reinforcement learning control model, obtain control instructions including the negative pressure adjustment amount, the valve opening adjustment amount, and the extraction flow adjustment amount, and continuously and adaptively optimize the deep reinforcement learning control model based on a comprehensive reward function including extraction efficiency, concentration deviation, and energy consumption factor;
[0053] A control instruction execution module, used to send the control instructions to the extraction execution unit and collect new gas concentration values, extraction flow values, and negative pressure values through the extraction execution unit;
[0054] A closed-loop self-learning update module, used to jointly write the new gas concentration value, extraction flow value, negative pressure value, and the working condition state vector into the deep reinforcement learning control model to form a closed-loop self-learning;
[0055] The device health assessment and maintenance linkage module is used to calculate the device health index based on the device vibration - electrical operation data, and automatically generate a maintenance work order when the health index is lower than a preset threshold.
[0056] The present invention discloses the following technical effects:
[0057] Based on the combination of multi - source sensing data acquisition, roadway coupling model correction, and deep reinforcement learning control, the present invention realizes the dynamic assessment, adaptive control, and closed - loop optimization of the gas extraction working conditions, can continuously adjust the control strategy according to the difference between the prediction and the actual working conditions, improve the extraction efficiency, reduce energy consumption, and keep the gas concentration within a safe range; at the same time, through the real - time calculation of the device health index and the automatic generation of maintenance work orders, the intelligent monitoring and preventive maintenance of the device state are realized, effectively ensuring the safety and stability of the system operation. Brief Description of the Drawings
[0058] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0059] Figure 1 It is the flowchart of the method provided by the embodiment of the present invention;
[0060] Figure 2 It is the schematic diagram of the system structure provided by the embodiment of the present invention. Detailed Embodiments
[0061] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0062] The purpose of the present invention is to provide an intelligent control method and system for the gas extraction process, which can continuously adjust the control strategy according to the difference between the prediction and the actual working conditions, improve the extraction efficiency, reduce energy consumption, and keep the gas concentration within a safe range, and realize the intelligent monitoring and preventive maintenance of the device state.
[0063] To make the above - mentioned objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0064] Figure 1The flowchart of the method provided by the embodiments of the present invention is as follows. Figure 1 As shown, the present invention provides an intelligent control method for the gas extraction process, including:
[0065] Step 100: Deploy multiple environment - working condition integrated sensing nodes in the target gas extraction area, and synchronously collect the current multi - source sensing data at a preset period; the multi - source sensing data includes the current gas concentration value, the current extraction flow rate value, the current negative pressure value, the pipeline network temperature value, and the equipment vibration - electrical operation data;
[0066] Step 200: Perform online correction on the multi - source sensing data according to the well - lane fluid - particle coupling physical model and historical extraction samples to obtain the predicted gas concentration value and the predicted extraction efficiency value;
[0067] Step 300: jointly form a working condition state vector with the predicted gas concentration value, the predicted extraction efficiency value, the current negative pressure value, and the current extraction flow rate value, input it into the deep reinforcement learning control model, obtain control instructions including the negative pressure adjustment amount, the valve opening adjustment amount, and the extraction flow rate adjustment amount, and continuously and adaptively optimize the deep reinforcement learning control model based on a comprehensive reward function including extraction efficiency, concentration deviation, and energy consumption factor;
[0068] Step 400: Send the control instructions to the extraction execution unit, and collect new gas concentration values, extraction flow rate values, and negative pressure values through the extraction execution unit;
[0069] Step 500: jointly write the new gas concentration value, extraction flow rate value, negative pressure value, and the working condition state vector into the deep reinforcement learning control model to form a closed - loop self - learning;
[0070] Step 600: Calculate the equipment health index based on the equipment vibration - electrical operation data, and automatically generate a maintenance work order when the health index is lower than the preset threshold.
[0071] Specifically, to achieve multi-dimensional state perception of the gas extraction process, a number of environment-operation integrated sensing nodes are arranged along the gas extraction pipeline, branch nodes and key equipment in the target gas extraction area. Each sensing node integrates a gas concentration sensor, a flow sensor, a negative pressure sensor, a temperature sensor and an electrical vibration monitoring module. The gas concentration sensor adopts the photo-acoustic composite principle and can measure the local methane concentration with an accuracy of no more than ±0.1 vol% CH4; the flow sensor is constructed by a thermal mass flowmeter or a differential pressure flowmeter and is suitable for the medium and low-speed coalbed methane extraction working conditions; the negative pressure sensor is a corrosion-resistant diffused silicon differential pressure module; the temperature sensor is arranged in the area where the pipeline is prone to condensation to collect the ambient temperature and the surface temperature of the equipment shell; the vibration-electrical monitoring module integrates a three-axis accelerometer, a current transformer and a winding thermistor to synchronously collect the axial vibration, current fluctuation and temperature change of the extraction equipment and other operating states. The above-mentioned sensing nodes are uniformly connected to the edge acquisition host through the RS485 or LoRa communication protocol to form a star or chain structure. The host synchronously schedules each node to complete a data acquisition according to the set sampling period (preferably 10 seconds to 60 seconds), and encodes the acquisition results into a multi-source sensing data packet, and uploads it to the control center in real time for subsequent analysis and modeling.
[0072] Preferably, the multi-source sensing data is corrected online according to the well roadway fluid-particle coupling physical model and historical extraction samples to obtain the predicted gas concentration value and the predicted extraction efficiency value, including:
[0073] Construct a coupling simulation model that describes the gas flow, diffusion and particle perturbation effects in the well roadway space of the target gas extraction area;
[0074] Use the coupling simulation model to make a preliminary prediction of the currently collected multi-source sensing data to obtain the physical model predicted concentration value and the predicted extraction efficiency value;
[0075] Construct a residual set based on the deviation between the physical model predicted value and the preset historical extraction sample data;
[0076] Use the pre-trained residual correction neural network model to perform regression prediction on the residual set to obtain the residual correction value;
[0077] Perform weighted fusion on the physical model predicted value and the residual correction value to obtain the predicted gas concentration value.
[0078] In this embodiment, a coupled simulation model is first constructed to characterize the gas flow, diffusion, and particle perturbation behaviors in the roadway space of the target gas extraction area. This model is based on the three-dimensional roadway structure, physical parameters of coal and rock masses, and gas occurrence conditions in the extraction area, combines the momentum equation and mass conservation equation in continuum fluid dynamics, and introduces a particle volume fraction perturbation term and a gas desorption kinetic factor to simulate the gas-particle two-phase coupled migration process in coal fractures. By inputting the currently collected gas concentration value, extraction flow value, negative pressure value, and pipeline network temperature value into this coupled model, preliminary prediction results in the physical sense can be obtained, namely the predicted concentration value of the physical model and the predicted extraction efficiency value. Among them, the predicted extraction efficiency value is calculated from the ratio between the simulated total extraction volume and the theoretical maximum gas supply capacity.
[0079] Furthermore, in this embodiment, a residual correction neural network model is used to online correct the above preliminary prediction results. Specifically, first, the predicted concentration value of the physical model is compared with the actual measured concentration in the historical extraction samples, the prediction error within the set time window is calculated, and a residual set is constructed. Subsequently, a trained two-layer feedforward neural network model is used to perform non-linear regression prediction on this residual set to obtain the corresponding residual correction value. The neural network in this embodiment uses the ReLU activation function, and the network weights are trained and optimized by minimizing the mean square error of the residuals. Finally, in this embodiment, the predicted concentration value of the physical model and the residual correction value are weighted and fused according to the empirically set fusion coefficient, and the predicted gas concentration value with adaptive correction ability is output as the key input for the subsequent reinforcement learning control link.
[0080] Preferably, the expression of the roadway fluid-particle coupled physical model is:
[0081]
[0082] where ρ g is the gas density; u g is the gas velocity vector; p is the gas pressure; μ g is the gas dynamic viscosity; β is the gas-solid momentum coupling coefficient; φ p is the particle volume fraction; u p is the particle velocity vector; ε is the porosity; C g is the gas volume fraction concentration; D eff is the equivalent diffusion coefficient; K d is the gas adsorption-desorption kinetic coefficient; C eq is the gas equilibrium concentration; β c is the concentration coupling coefficient; ‖·‖ represents the vector norm; η pred is the predicted extraction efficiency; Q calc is the extraction flow calculated by the model; Q leakis the leakage flow rate; x is the concentration attenuation coefficient; C target is the target control concentration; is the gradient operator; t is the time.
[0083] In this embodiment, each parameter involved in the roadway fluid-particle coupling physical model is obtained by means of on-site measurement, experimental calibration and literature data fusion. Among them, the basic thermal physical parameters such as the density, dynamic viscosity and volume fraction concentration of the gas are calculated by the gas state equation combined with the measured pressure, temperature and components; the gas pressure and velocity are collected in real time by the differential pressure sensor and flow sensor arranged in the drainage pipeline network; the particle volume fraction and particle velocity mainly come from the particle activity data of the coal-rock mass fracture structure scanning and microseismic monitoring system, and the representative disturbance factor is obtained by numerical fitting; the porosity is determined by the core porosity analysis experiment or inversely calculated from the formation stress and gas compression state; the equivalent diffusion coefficient and the gas adsorption-desorption kinetic coefficient are obtained by the gas desorption experiment or isothermal adsorption curve fitting; the concentration coupling coefficient and the concentration attenuation coefficient are adjusted empirically according to the existing simulation results of coal seam gas migration; the equilibrium concentration and the target control concentration are determined according to the safety production regulations and gas control standards; the drainage flow rate and the leakage flow rate are jointly inversely calculated by the online flowmeter and the negative pressure change trend; the drainage efficiency is dynamically estimated by the model according to the coupling relationship between the drainage and leakage ratio, concentration and time-varying pressure. Some of the above parameters are statically set, and some can be dynamically updated with on-site data to ensure the accuracy and adaptability of the model prediction.
[0084] Preferably, the calculation formula of the residual correction value is:
[0085]
[0086] where r t is the residual vector at the current time step; w λ is the weight vector of the gating layer; b λ is the bias of the gating layer; σ(·) is the Sigmoid function, which is used to output the gating coefficient in the range of 0-1; W1 is the weight matrix of the first fully connected layer; b1 is the bias of the first layer; φ(·) is the non-linear activation function; W2 is the weight matrix of the second fully connected layer; b2 is the bias of the second layer; Δy corr (t) is the residual correction value at the current time step.
[0087] In this embodiment, the calculation of the residual correction value depends on a pre-trained neural network structure, and all its parameters are obtained through the historical sampling data-driven method. First, the residual vector at the current time step is composed of the difference between the predicted concentration value and the actual measured concentration value of the roadway physical model, which can be directly calculated from the sensor-collected data and the model output; the weight vector and bias of the gating layer are obtained through optimization based on the cross-entropy loss function during the training process, and are used to adjust the channel intensity of the residual signal before entering the main network; the Sigmoid function is a fixed non-linear function, which is used to normalize the gating output between 0 and 1 without training; the weight matrices and bias values of the first and second fully-connected networks are all obtained through training with a large number of historical data samples in the offline stage. The Adam optimization algorithm is used in the training process, and the performance is evaluated in combination with the validation set to ensure that the network has good fitting ability; the non-linear activation function selects the ReLU function, and its parameters are not updated during the training process, and it is only used as a structural unit to introduce non-linear ability; the finally output residual correction value is the forward propagation result of the above network structure under the current input, and is used to perform error compensation and accuracy correction on the prediction result of the physical model. The structural parameters of the entire network can be kept fixed after deployment, or can be iteratively updated according to the on-site data to improve the adaptive ability of online correction.
[0088] Preferably, the predicted gas concentration value, the predicted drainage efficiency value, the current negative pressure value, and the current drainage flow value are jointly used to form a working condition state vector, which is input into the deep reinforcement learning control model to obtain control instructions including the negative pressure adjustment amount, the valve opening adjustment amount, and the drainage flow adjustment amount. And based on a comprehensive reward function including drainage efficiency, concentration deviation, and energy consumption factor, the deep reinforcement learning control model is continuously and adaptively optimized, including:
[0089] Perform zero-mean normalization processing on the predicted gas concentration value, the predicted drainage efficiency value, the current negative pressure value, and the current drainage flow value, and splice them in a predetermined order to form the four-dimensional working condition state vector;
[0090] Input the working condition state vector into the deep reinforcement learning control model including a policy network and a value network to obtain three-dimensional adjustment amounts corresponding to the negative pressure, valve opening, and drainage flow;
[0091] Package the three-dimensional adjustment amount into a control instruction according to the industrial fieldbus communication protocol;
[0092] After the execution of the control instruction is completed, according to the comprehensive reward function R = w e (η new -η old )-w c |C new -C target |-wp P elec Calculate the comprehensive reward value in real time; where R is the comprehensive reward value; w e is the extraction efficiency weight; w c is the concentration deviation penalty weight; w p is the energy consumption penalty weight; η new and η old are the extraction efficiencies after and before adjustment respectively; C new is the adjusted gas concentration value; C target is the target control concentration; P elec is the electric power of the extraction equipment;
[0093] Use the temporal difference error Δ = R + dV(s t+1 ) - V(s t ) to perform backpropagation and Adam optimization on the weights of the policy network and the value network to achieve continuous adaptive adjustment of the deep reinforcement learning control model; where Δ is the temporal difference error; d is the discount factor; V(s t ) and V(s t+1 ) are the valuations of the current state and the next state by the value network respectively.
[0094] In this embodiment, to implement the intelligent regulation strategy for gas extraction based on deep reinforcement learning, first perform zero-mean normalization on the predicted gas concentration value, predicted extraction efficiency value, current negative pressure value, and current extraction flow value to unify the dimensional differences of various physical quantities and ensure the processing accuracy of the neural network. Subsequently, the four normalized features are concatenated in a set order to form a four-dimensional working condition state vector, which is input into the reinforcement learning control model jointly composed of the policy network and the value network. Among them, the policy network consists of two fully connected layers, each containing 128 and 64 rectified linear units, and the output is a three-dimensional action vector including the negative pressure adjustment amount, valve opening adjustment amount, and extraction flow adjustment amount. The control model generates the optimal control instruction by collecting the current state vector and combining the historical interaction samples in the experience pool. This control instruction is packed into a standard message according to the industrial fieldbus communication protocol and sent to the extraction execution unit for parameter adjustment.
[0095] Furthermore, after the control instruction is executed, the present embodiment instantaneously collects the execution result to obtain the gas extraction efficiency values before and after adjustment, the gas concentration value after adjustment, and the electric power of the equipment. The system calculates the reward value of the current step according to the comprehensive reward function. Among them, the reward function consists of three items: one is the part of the improved gas extraction efficiency, which is given positive incentives; the second is the absolute difference between the gas concentration and the target concentration, which is used as a penalty item; the third is the energy consumption penalty item, which is calculated based on the real-time electric power. The weights of the three items are set according to the historical strategy performance and the optimization goal. Subsequently, a reinforcement learning optimization mechanism based on temporal difference is adopted to calculate the difference between the current reward value and the value network's estimation of the current state and the next state, forming a temporal difference error. Using this as the loss function, backpropagation training is performed on the parameters of the policy network and the value network. The optimization method uses the Adam adaptive gradient algorithm to ensure the convergence stability of the training process and the real-time nature of the policy update, thereby realizing the continuous adaptive update of the deep reinforcement learning model and the closed-loop evolution of the control strategy.
[0096] Preferably, the policy network consists of two layers, each with 128 rectified linear units.
[0097] In this embodiment, the policy network in the deep reinforcement learning control model preferably adopts a two-layer feedforward neural network structure, and each layer contains 128 rectified linear units, that is, ReLU (Rectified Linear Unit) activation function nodes. The neurons in the first layer receive a four-dimensional operating condition state vector as input, and after being weighted by matrix multiplication and bias, they are sent to the ReLU activation function to extract non-linear features; the neurons in the second layer use the output of the first layer as input, and the structure is the same as that of the first layer, and continue to perform high-order feature mapping and compression. The network weights and bias parameters are iteratively trained through a backpropagation algorithm with the minimization of the temporal difference error as the objective function, and the Adam optimizer is used for parameter update to ensure the convergence efficiency and learning stability. In this embodiment, the output layer of the policy network is directly connected to the action space, and three control quantities are output: namely, the negative pressure adjustment amount, the valve opening adjustment amount, and the gas extraction flow adjustment amount, for subsequent use by the control instruction encapsulation module. This structure takes into account both the expressive ability and the computational efficiency, and is suitable for real-time deployment and execution in edge computing nodes or industrial control hosts.
[0098] Preferably, the gas extraction execution unit includes: a negative pressure regulating pump, a branch electric valve, and a variable frequency fan.
[0099] In this embodiment, the control center encapsulates the negative pressure adjustment amount, valve opening adjustment amount, and gas extraction flow adjustment amount output by the deep reinforcement learning control model into standard control instructions conforming to the Modbus TCP or PROFIBUS protocol format, and sends them to the gas extraction execution unit through the industrial Ethernet or fieldbus network. To ensure real-time performance and redundant reliability, the control channel is equipped with a dual communication link and a CRC redundancy check mechanism. After receiving the control instructions, each execution device in the gas extraction execution unit automatically completes the target adjustment according to the adjustment amount: the negative pressure adjustment pump dynamically adjusts the local negative pressure intensity of the gas extraction system according to the negative pressure adjustment amount; the branch electric valve executes angle adjustment according to the valve opening adjustment amount to achieve flow balance in the multi-branch pipe network; the variable-frequency fan changes the speed according to the gas extraction flow adjustment amount, thereby controlling the total gas extraction volume and the steady-state air flow structure.
[0100] Further, after the adjustment action is completed, the gas extraction execution unit immediately triggers the multi-source sensing nodes associated with it to collect feedback data, and obtains the adjusted gas concentration value, gas extraction flow value, and negative pressure value as the input data for the next round of the control model. The collection process is managed by an embedded edge collection controller, which supports a multi-threaded concurrent scheduling mechanism to ensure the closed-loop timeliness between the adjustment response and the data feedback. After the collected data is preliminarily filtered locally and outliers are detected, it is transmitted to the upper computer system through the data interface. This closed-loop execution structure effectively ensures the real-time matching of control decisions and on-site conditions, provides dynamic environment support for the continuous learning and strategy optimization of the model, and helps to achieve intelligent and stable control of the gas extraction process.
[0101] Preferably, the new gas concentration value, gas extraction flow value, negative pressure value, and the working condition state vector are jointly written into the deep reinforcement learning control model to form a closed-loop self-learning, including:
[0102] Collect the new gas concentration value, gas extraction flow value, and negative pressure value after the execution of the control instructions, and splice them together with the predicted gas extraction efficiency value in a linear normalization manner to form the working condition state vector s at the next moment t+1 ;
[0103] Combine the current moment state vector s t , the executed three-dimensional adjustment amount a t , the immediate comprehensive reward R t , and the state vector s at the next moment t+1 into a transition quadruple <s t ,a t ,R t ,s t+1 >;
[0104] Write the quadruple into the cyclic experience replay pool, and eliminate the oldest data according to the first-in-first-out strategy to keep the memory pool capacity at N entries;
[0105] Randomly select m quadruples from the cyclic experience replay pool in the priority sampling manner to form a training batch
[0106] For each sample in the training batch calculate the temporal difference error according to the formula and use as the loss function, and use the Adam algorithm with the learning rate α lr to update the policy network parameters θ π and the value network parameters θ v simultaneously to achieve online self-learning; where d is the discount factor, is the output of the value network, δ is the temporal difference error, L is the mean squared error loss, and m is the batch size for each training
[0107] Every K times of parameter updates, perform with the soft update coefficient τ∈(0,1) to improve the stability of the training process and prevent policy divergence; where are the target value network parameters, and K is the soft update interval
[0108] In this embodiment, after the control instruction is executed, the adjusted gas concentration value, gas drainage flow value, and negative pressure value are collected by the sensor nodes linked by the gas drainage execution unit, and the above real-time collected data is standardized by using the linear normalization method. The three normalized environmental parameters are concatenated with the previously predicted gas drainage efficiency value to form the working condition state vector at the adjusted moment. Then, this embodiment combines the state vector with the working condition state vector before adjustment, the three-dimensional adjustment amount executed at the corresponding moment, and the instantaneously calculated comprehensive reward value to form a set of transition samples, that is, the state-action-reward-new state quadruple, and writes the quadruple into the cyclic experience replay pool with a pre-set capacity. To maintain the freshness and representativeness of the training data, this embodiment adopts the first-in-first-out strategy to automatically eliminate the oldest data samples, so that the replay pool always stores the latest interaction information
[0109] In the subsequent model training process, this embodiment adopts the priority sampling method to extract a number of interaction samples from the above-mentioned experience replay pool to form a training batch. For each sample, first calculate the learning deviation based on the value difference between the current state and the subsequent state, and generate an error signal required for policy update accordingly. This error signal is used to drive the joint optimization of the parameters of the policy network and the value network, and the gradient update method with an adaptive learning rate is adopted in the optimization process. After each completion of the preset number of parameter updates, this embodiment further adopts a soft update mechanism to transfer some parameters of the current value network to the target network in a weighted form, realizing the smooth update of the model structure and the stable control of the policy convergence process. Through the above mechanism, this embodiment realizes the continuous closed-loop self-learning and policy improvement of the reinforcement learning control model in the dynamic gas drainage scenario.
[0110] Preferably, calculate the equipment health index based on the equipment vibration-electrical operation data, and when the health index is lower than the preset threshold, automatically generate a maintenance work order, including:
[0111] Perform a sliding window analysis on the equipment vibration-electrical operation data, and respectively extract the effective value V of the axial vibration velocity rms 、the root mean square value I of the operating current rms 、the time average value T of the stator winding temperature avg ;
[0112] Perform linear normalization on V rms 、I rms and T avg according to their respective historical extreme value intervals to obtain dimensionless vibration characteristics, dimensionless current characteristics and dimensionless temperature characteristics;
[0113] Calculate the equipment health index according to the formula ; where, and are the dimensionless vibration characteristics, dimensionless current characteristics and dimensionless temperature characteristics respectively; w v 、w i and w t are the weight coefficients of the dimensionless vibration characteristics, the dimensionless current characteristics and the dimensionless temperature characteristics respectively, satisfying w v +w i +w t = 1;
[0114] Compare the calculated H idx with the preset health threshold H thr ; when H idx ≤ H thr , trigger the preset maintenance process and generate the maintenance work order.
[0115] In this embodiment, to achieve real-time assessment of the operating status of the extraction equipment, first, sliding window analysis is performed on the equipment vibration, electrical, and temperature data collected by the vibration acceleration sensor, current transformer, and thermistor module. Within a set time period, this embodiment extracts three key operating characteristics: one is the root mean square value of the axial vibration velocity, which is used to reflect the mechanical structure stability; the second is the root mean square value of the motor operating current, which is used to reflect the load change and drive stability; the third is the time average value of the stator winding temperature, which is used to evaluate the thermal load condition. These three characteristics are linearly normalized based on their historical operating data intervals respectively, and converted into dimensionless indicators to eliminate the influence of dimensional differences on the evaluation accuracy, and are respectively denoted as dimensionless vibration characteristics, dimensionless current characteristics, and dimensionless temperature characteristics.
[0116] After obtaining the above dimensionless characteristics, this embodiment calculates the health index of the equipment in the way of weighted summation. The weight coefficients of each characteristic can be obtained through three methods: one is the recommended configuration provided by the equipment manufacturer, which is the empirical parameter based on typical failure modes; the second is the optimal coefficient extracted through correlation analysis of historical operation and maintenance data, which reflects the sensitivity of each characteristic to the occurrence of faults; the third is the coefficient obtained by adaptive learning using Bayesian optimization or genetic algorithm under the set evaluation objective. The sum of the weight coefficients is a constant to ensure that the health index is within the standard range. This embodiment compares the calculated health index with the preset health threshold. When the health index is lower than the threshold, it is determined that the equipment has potential abnormalities, and the system automatically triggers the maintenance response process, generates a maintenance work order including equipment identification, predicted fault type, remaining available life assessment results, and recommended maintenance time window, and pushes it to the remote maintenance platform or on-site management terminal. At the same time, the system can automatically link to reduce the extraction load according to the equipment type to ensure the safe operation of the equipment until the maintenance is completed.
[0117] Corresponding to the above method, as Figure 2 shown, this embodiment also provides an intelligent control system for the gas extraction process, including:
[0118] A multi-source sensing and acquisition module, which is used to deploy multiple environment-condition integrated sensing nodes in the target gas extraction area and synchronously acquire the current multi-source sensing data according to a preset period; the multi-source sensing data includes the current gas concentration value, the current extraction flow value, the current negative pressure value, the pipeline network temperature value, and the equipment vibration-electrical operation data;
[0119] A coupling modeling and prediction module, which is used to perform online correction on the multi-source sensing data according to the roadway fluid-particle coupling physical model and historical extraction samples to obtain the predicted gas concentration value and the predicted extraction efficiency value;
[0120] The working condition evaluation and control decision-making module is used to jointly form a working condition state vector with the predicted gas concentration value, the predicted drainage efficiency value, the current negative pressure value, and the current drainage flow rate value, input it into the deep reinforcement learning control model, obtain control instructions including the negative pressure adjustment amount, the valve opening adjustment amount, and the drainage flow rate adjustment amount, and continuously and adaptively optimize the deep reinforcement learning control model based on a comprehensive reward function including drainage efficiency, concentration deviation, and energy consumption factor;
[0121] The control instruction execution module is used to send the control instructions to the drainage execution unit, and collect new gas concentration values, drainage flow rate values, and negative pressure values through the drainage execution unit;
[0122] The closed-loop self-learning update module is used to jointly write the new gas concentration value, drainage flow rate value, negative pressure value, and the working condition state vector into the deep reinforcement learning control model to form a closed-loop self-learning;
[0123] The equipment health assessment and maintenance linkage module is used to calculate the equipment health index based on the equipment vibration-electrical operation data, and automatically generate a maintenance work order when the health index is lower than a preset threshold.
[0124] The beneficial effects of the present invention are as follows:
[0125] (1) By constructing an environment-working condition integrated sensing network covering multi-dimensional information such as gas concentration, drainage flow rate, negative pressure, temperature, and equipment status, the present invention realizes high-timeliness and high-precision data collection for the whole process of gas extraction, providing a complete and reliable perception basis for subsequent modeling, prediction, and control.
[0126] (2) The present invention adopts the method of fusing the physical model of roadway fluid-particle coupling and the residual correction neural network to conduct online modeling of the gas migration process, and can output high-precision predicted concentration values and drainage efficiency estimation results in a complex mining pressure disturbance and gas fluctuation environment, significantly improving the accuracy and adaptability of the prediction model.
[0127] (3) The present invention introduces a deep reinforcement learning control model based on the working condition state vector, combines a multi-objective reward mechanism for improving drainage efficiency, controlling concentration deviation, and suppressing energy consumption, realizes the adaptive optimization and closed-loop update of the control strategy, has strong autonomous learning ability and dynamic adjustment ability, and can effectively improve the stability and intelligent level of the gas extraction process.
[0128] (4) The present invention establishes an equipment health index by fusing characteristic indicators such as vibration, current, and temperature, and automatically generates a maintenance work order in combination with threshold judgment, realizing predictive maintenance management for key drainage equipment, effectively reducing the risk of unplanned shutdown, and ensuring the continuous and safe operation of the drainage system.
[0129] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method section.
[0130] Specific examples are used in this article to elaborate on the principles and implementation manners of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea. At the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. An intelligent control method for the gas extraction process, characterized in that, Including: Deploy multiple environment - working condition integrated sensing nodes in the target gas extraction area, and synchronously collect current multi - source sensing data at a preset period; The multi - source sensing data includes the current gas concentration value, the current extraction flow value, the current negative pressure value, the pipeline network temperature value, and the equipment vibration - electrical operation data; Online correct the multi - source sensing data according to the roadway fluid - particle coupling physical model and historical extraction samples to obtain the predicted gas concentration value and the predicted extraction efficiency value; Jointly form a working condition state vector with the predicted gas concentration value, the predicted extraction efficiency value, the current negative pressure value, and the current extraction flow value, input it into the deep reinforcement learning control model, obtain control instructions including the negative pressure adjustment amount, the valve opening adjustment amount, and the extraction flow adjustment amount, and continuously and adaptively optimize the deep reinforcement learning control model based on a comprehensive reward function including extraction efficiency, concentration deviation, and energy consumption factor; Send the control instructions to the extraction execution unit, and collect new gas concentration values, extraction flow values, and negative pressure values through the extraction execution unit; Jointly write the new gas concentration value, extraction flow value, negative pressure value, and the working condition state vector into the deep reinforcement learning control model to form a closed - loop self - learning; Calculate the equipment health index based on the equipment vibration - electrical operation data, and automatically generate a maintenance work order when the health index is lower than the preset threshold.
2. The intelligent control method for the gas extraction process according to claim 1, characterized in that, The extraction execution unit includes: a negative pressure regulating pump, a branch electric valve, and a variable - frequency fan.
3. The intelligent control method for the gas extraction process according to claim 1, wherein Online correct the multi - source sensing data according to the roadway fluid - particle coupling physical model and historical extraction samples to obtain the predicted gas concentration value and the predicted extraction efficiency value, including: Construct a coupling simulation model that describes the gas flow, diffusion, and particle perturbation effects in the roadway space of the target gas extraction area; Use the coupling simulation model to perform a preliminary prediction on the currently collected multi - source sensing data to obtain the physical model predicted concentration value and the predicted extraction efficiency value; Construct a residual set based on the deviation between the physical model predicted value and the preset historical extraction sample data; Use the pre - trained residual correction neural network model to perform a regression prediction on the residual set to obtain the residual correction value; Perform weighted fusion on the physical model predicted value and the residual correction value to obtain the predicted gas concentration value.
4. The intelligent control method for the gas extraction process according to claim 3, characterized in that The expression of the roadway fluid - particle coupling physical model is: where ρ g is the density of the gas; u g is the velocity vector of the gas; p is the gas pressure; μ g is the dynamic viscosity of the gas; β is the gas-solid momentum coupling coefficient; φ p is the particle volume fraction; u p is the velocity vector of the particle; ε is the porosity; C g is the volume fraction concentration of the gas; D eff is the equivalent diffusion coefficient; K d is the gas adsorption-desorption kinetic coefficient; C eq is the equilibrium concentration of the gas; β c is the concentration coupling coefficient; ‖·‖ represents the vector norm; η pred is the predicted drainage efficiency; Q calc is the drainage flow calculated by the model; Q leak is the leakage flow; x is the concentration attenuation coefficient; C target is the target control concentration; is the gradient operator; t is the time.
5. The intelligent control method for the gas extraction process according to claim 3, characterized in that The calculation formula of the residual correction value is: where r t is the residual vector at the current time step; w λ is the weight vector of the gating layer; b λ is the bias of the gating layer; σ(·) is the Sigmoid function, which is used to output the gating coefficient in the range of 0-1; W1 is the weight matrix of the first fully connected layer; b1 is the bias of the first layer; φ(·) is the non-linear activation function; W2 is the weight matrix of the second fully connected layer; b2 is the bias of the second layer; Δy corr (t) is the residual correction value at the current time step.
6. The intelligent control method for the gas extraction process according to claim 1, wherein Jointly form a working condition state vector with the predicted gas concentration value, the predicted extraction efficiency value, the current negative pressure value, and the current extraction flow value, input it into the deep reinforcement learning control model, obtain control instructions including the negative pressure adjustment amount, the valve opening adjustment amount, and the extraction flow adjustment amount, and continuously and adaptively optimize the deep reinforcement learning control model based on a comprehensive reward function including extraction efficiency, concentration deviation, and energy consumption factor, including: Perform zero - mean normalization processing on the predicted gas concentration value, the predicted extraction efficiency value, the current negative pressure value, and the current extraction flow value, and splice them in a predetermined order to form the four - dimensional working condition state vector; Input the working condition state vector into the deep reinforcement learning control model including a policy network and a value network to obtain a three-dimensional adjustment quantity corresponding to negative pressure, valve opening, and extraction flow rate; Package the three-dimensional adjustment quantity into a control instruction according to the industrial fieldbus communication protocol; After the execution of the control instruction, calculate the comprehensive reward value in real time according to the comprehensive reward function R = w e (η new - η old ) - w c |C new - C target | - w p P elec ; where R is the comprehensive reward value; w e is the extraction efficiency weight; w c is the concentration deviation penalty weight; w p is the energy consumption penalty weight; η new and η old are the extraction efficiencies after and before adjustment respectively; C new is the adjusted gas concentration value; C target is the target control concentration; P elec is the electric power of the extraction equipment. Using the temporal difference error Δ = R + dV(s t+1 ) - V(s t ), perform backpropagation and Adam optimization on the weights of the policy network and the value network to achieve continuous adaptive adjustment of the deep reinforcement learning control model; where Δ is the temporal difference error; d is the discount factor; V(s t ) and V(s t+1 ) are the valuations of the current state and the next state by the value network respectively.
7. The intelligent control method for the gas extraction process according to claim 6, characterized in that The policy network consists of two layers with 128 rectified linear units each.
8. The intelligent control method for the gas extraction process according to claim 1, characterized in that, Write the new gas concentration value, extraction flow rate value, negative pressure value, and the working condition state vector into the deep reinforcement learning control model together to form closed-loop self-learning, including: The new gas concentration value, gas drainage flow value, and negative pressure value after the execution of the acquisition control instruction are spliced together with the predicted gas drainage efficiency value in a linear normalization manner to form the working condition state vector s at the next moment t+1 ; The state vector s at the current moment t , the three-dimensional adjustment amount a executed t , the immediate comprehensive reward R t and the state vector s at the next moment t+1 are combined into a transition quadruple <s t , a t , R t , s t+1 >; Write the quadruple into the cyclic experience replay pool, and eliminate the oldest data according to the first-in-first-out strategy to keep the memory pool capacity at N items; Randomly select m quadruples from the cyclic experience replay pool in the priority sampling manner to form a training batch For each sample in the training batch calculate the temporal difference error according to the formula and use as the loss function, and use the Adam algorithm with the learning rate α lr to update the policy network parameters θ π and the value network parameters θ v simultaneously to achieve online self-learning; where d is the discount factor, is the output of the value network, δ is the temporal difference error, L is the mean squared error loss, and m is the batch size for each training Every K parameter updates, perform with a soft update coefficient τ ∈ (0, 1) to improve the stability of the training process and prevent policy divergence; where are the target value network parameters, and K is the soft update interval.
9. The intelligent control method for the gas extraction process according to claim 1, characterized in that Calculate the equipment health index based on the equipment vibration-electrical operation data, and automatically generate a maintenance work order when the health index is lower than a preset threshold, including: Perform a sliding window analysis on the vibration-electrical operation data of the device, and respectively extract the effective value V of the axial vibration velocity rms , the root mean square value I of the operating current rms , and the time average value T of the stator winding temperature avg ; Normalize V rms , I rms and T avg by linear normalization according to their respective historical extreme value intervals to obtain dimensionless vibration characteristics, dimensionless current characteristics and dimensionless temperature characteristics; Calculate the device health index according to the formula ; where and are dimensionless vibration characteristics, dimensionless current characteristics, and dimensionless temperature characteristics respectively; w v , w i and w t are the weight coefficients of the dimensionless vibration characteristic, the dimensionless current characteristic, and the dimensionless temperature characteristic respectively, satisfying w v + w i + w t = 1; Compare the calculated H idx with a preset health threshold H thr ; when H idx ≤H thr , trigger a preset maintenance process and generate the maintenance work order.
10. An intelligent control system for the gas extraction process, characterized in that, Including: A multi-source sensing acquisition module for deploying multiple environment-condition integrated sensing nodes in the target gas extraction area and synchronously acquiring the current multi-source sensing data at a preset period; The multi-source sensing data includes the current gas concentration value, the current extraction flow rate value, the current negative pressure value, the pipeline network temperature value, and the equipment vibration-electrical operation data; A coupling modeling and prediction module for online correcting the multi-source sensing data according to the roadway fluid-particle coupling physical model and historical extraction samples to obtain a predicted gas concentration value and a predicted extraction efficiency value; A working condition evaluation and control decision module for jointly forming a working condition state vector with the predicted gas concentration value, the predicted extraction efficiency value, the current negative pressure value, and the current extraction flow rate value, inputting them into the deep reinforcement learning control model to obtain control instructions including a negative pressure adjustment quantity, a valve opening adjustment quantity, and an extraction flow rate adjustment quantity, and continuously and adaptively optimizing the deep reinforcement learning control model based on a comprehensive reward function including extraction efficiency, concentration deviation, and energy consumption factor; A control instruction execution module for sending the control instruction to the extraction execution unit and collecting new gas concentration values, extraction flow rate values, and negative pressure values through the extraction execution unit; A closed-loop self-learning update module for jointly writing the new gas concentration value, extraction flow rate value, negative pressure value, and the working condition state vector into the deep reinforcement learning control model to form closed-loop self-learning; An equipment health assessment and maintenance linkage module for calculating the equipment health index based on the equipment vibration-electrical operation data and automatically generating a maintenance work order when the health index is lower than a preset threshold.
Citation Information
Patent Citations
Intelligent control method for gas extraction drilling
CN118167404B
Coal and gas outburst early warning method based on field real-time data driving
CN115977736A
Intelligent decision-making regulation and control method and regulation and control platform for gas extraction
CN116050239A
Gas pre-extraction multi-objective optimization decision-making method based on multi-modal deep learning
CN120068597A
Tunnel multi-source fusion dynamic twin surrounding rock intelligent prediction and control method and system
CN120087772A