Underground equipment state intelligent control system and control method thereof

By deploying multimodal sensor arrays and multimodal fusion models in downhole equipment, combining optimized reinforcement learning strategies and deep deterministic strategy gradient algorithms, the shortcomings in downhole equipment status monitoring and control in the existing technology are solved, comprehensive monitoring and intelligent control of equipment status are achieved, and equipment operation efficiency and safety are improved.

CN120030445AActive Publication Date: 2025-05-23NUOWENKE BLOWER FAN BEIJING

Patent Information

Application Number
CN202510510358.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-23
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing technology is difficult to monitor the status of downhole equipment comprehensively and accurately, resulting in insufficient early warning capabilities for potential equipment failures, and traditional equipment control strategies cannot be dynamically adjusted according to real-time status and complex working conditions, resulting in increased energy waste and equipment wear.

Method used

The multi-dimensional state acquisition module is used to collect multi-dimensional state data through a multi-modal sensor array that integrates vibration, stress and temperature, and build a multi-modal fusion model through a deep feature processing module to extract multi-modal state features. Then, by optimizing reinforcement learning strategies and improving the deep deterministic strategy gradient algorithm, combining dynamic reward functions, the optimal scheduling strategy is generated and the device operation control instructions are output.

Benefits of technology

It realizes comprehensive and accurate monitoring and intelligent control of downhole equipment status, improves the early warning capability of equipment failure and equipment operation efficiency, reduces energy waste and equipment wear, and reduces maintenance costs and failure risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030445A_ABST
    Figure CN120030445A_ABST
Patent Text Reader

Abstract

The invention discloses an underground equipment state intelligent control system and a control method thereof, and the system comprises a multi-dimensional state obtaining module which is used for deploying a vibration, stress and temperature integrated multi-mode sensor array and collecting multi-dimensional state data through a timestamp synchronization mechanism; the depth feature processing module is used for constructing a multi-modal fusion model and extracting multi-modal state features from the multi-dimensional state data based on the multi-modal fusion model; and the state decision output module is used for constructing an optimized reinforcement learning strategy through the multi-modal state features, obtaining an optimal scheduling strategy from the multi-modal state features based on the optimized reinforcement learning strategy in combination with an improved depth deterministic strategy gradient algorithm, and outputting an equipment operation control instruction through the optimal scheduling strategy. Therefore, the system can better adapt to complex and changeable working conditions in the well, and more intelligent control instructions are provided for equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent control, and particularly relates to an intelligent control system for underground equipment status and its control method. Background Art

[0002] With the continuous expansion of the scale of underground operations and the improvement of equipment complexity, traditional manual experience judgment and simple automation control are difficult to meet the modern high-efficiency and safe production requirements. Therefore, it is urgent to develop a system and method that can monitor the status of underground equipment in real time and comprehensively, and make intelligent decisions and precise control based on multi-source data; In the existing technology, the methods for monitoring and controlling the status of underground equipment are relatively single. Usually, only a few types of sensors are relied on to obtain limited equipment operation information, which is difficult to comprehensively and accurately reflect the actual status of the equipment. For example, only monitoring the vibration of the equipment cannot timely understand other key parameters such as the stress distribution and temperature change of the equipment, resulting in insufficient early warning ability for potential equipment failures. At the same time, in terms of equipment control strategies, most adopt fixed control modes and fail to make dynamic adjustments according to the real-time status of the equipment and the complex and changeable underground working conditions. This control method not only cannot give full play to the performance of the equipment, but also will cause energy waste and increased equipment wear due to the equipment being in an unreasonable operating state for a long time, increasing the equipment maintenance cost and failure risk. Therefore, an intelligent control system for underground equipment status and its control method are proposed here. Summary of the Invention

[0003] In order to overcome the above-mentioned defects of the prior art and to achieve the above object, the present invention proposes the following technical solutions: An intelligent control system for underground equipment status, comprising: A multi-dimensional state acquisition module: Deploy an integrated multi-modal sensor array of vibration, stress, and temperature, and collect multi-dimensional state data through a timestamp synchronization mechanism; A deep feature processing module: Construct a multi-modal fusion model, and extract multi-modal state features from the multi-dimensional state data based on the multi-modal fusion model; A state decision output module: Construct an optimized reinforcement learning strategy through the multi-modal state features, obtain the optimal scheduling strategy from the multi-modal state features based on the optimized reinforcement learning strategy combined with the improved deep deterministic policy gradient algorithm, and output equipment operation control instructions through the optimal scheduling strategy; The multi-modal fusion model includes a temporal convolutional network model, a graph convolutional network model, and an improved GhostNet model. The improved GhostNet model dynamically adjusts the convolutional kernel configuration according to the spatial feature distribution of the temperature field image on the basis of the original Ghost model; The reward mechanism for optimizing the reinforcement learning strategy is a dynamic learning reward mechanism, the improved deep deterministic policy gradient algorithm is a dual-Critic network structure, and the capacity is set to The target policy network and the target critic network are optimized and updated through the target network soft update mechanism.

[0004] The multi-dimensional state data includes vibration time domain sequence, stress distribution matrix, and temperature field image data; The vibration time domain sequence is obtained by collecting the vibration signal of the equipment in real time through a micro-electromechanical accelerometer during operation, the stress distribution matrix is ​​obtained by collecting the stress value of each monitoring point of the equipment through a fiber grating stress sensor, and the temperature field image data is obtained by configuring an infrared thermal imaging device to capture the surface temperature distribution of the equipment in real time.

[0005] The multi-dimensional state data is obtained by simultaneously adding precise timestamps to the vibration time domain sequence, stress distribution matrix and temperature field image data and performing data fusion preprocessing, and finally performing time-space alignment based on the precise timestamps.

[0006] The multi-modal state characteristics include a vibration time series characteristic group, a stress space characteristic group and a temperature characteristic group; The vibration time series feature group is extracted from the vibration time domain sequence based on the time series convolutional network model, the stress space feature group is extracted from the stress distribution matrix based on the graph convolutional network model, and the temperature feature group is extracted from the temperature field image based on the improved GhostNet model.

[0007] The process of extracting the temperature feature group from the temperature field image based on the improved GhostNet model is as follows: The temperature feature group is divided into feature regions, and the convolution kernel configuration is dynamically adjusted to distinguish between low-frequency and high-frequency regions; The low-frequency region is convolved with a 5×5 convolution kernel to obtain a low-frequency temperature image; The high-frequency region is convolved with a 3×3 convolution kernel to obtain a high-frequency temperature image; The low-frequency temperature image and the high-frequency temperature image are spliced ​​together in the channel dimension to obtain a fused temperature field feature map; The temperature field feature maps obtained through multiple fusion processes are arranged according to the hierarchical order of the fusion process to form a temperature feature group.

[0008] The specific implementation process of the optimized reinforcement learning strategy is as follows: The intelligent agent for optimizing the reinforcement learning strategy includes a state space, an action space, and a dynamic reward function; Define the multimodal state features as a state space; Define the action space, which includes the equipment control instruction set, including the frequency conversion speed regulation parameters, hydraulic pressure adjustment gear and component start and stop combination; Construct a dynamic reward function, which includes safety reward, efficiency reward and energy consumption reward, and assign corresponding weights to different rewards; When the equipment stress is lower than the safety threshold, the safety bonus value approaches 1; when the ratio of the actual output of the equipment to the planned output approaches 1, the efficiency bonus value approaches 1; when the ratio of the actual energy consumption of the equipment to the standard energy consumption approaches 1, the bonus value approaches 1.

[0009] The process of combining the optimized reinforcement learning strategy with the improved deep deterministic policy gradient algorithm to obtain the optimal scheduling strategy is: Construct two Critic networks and use the MLP structure. Set two hidden layers. The activation function of each layer is the ReLU function. The output layer uses a linear activation function. The parameters of the two Critic networks are randomly initialized to evaluate the value of the state to the action. At the same time, the target policy network is initialized. Through dynamic experience playback, the current policy network interacts with the environment space, observes the device state at each time step, regulates the device according to the policy network output action, obtains rewards after executing the action, and transfers to the next state to form experience samples, which are stored in a During training, batches of data are randomly sampled from the buffer periodically, and the optimal scheduling strategy is obtained through the target network soft update mechanism.

[0010] The target network soft update mechanism implementation process is: For each sample obtained from the experience replay buffer, the target policy network is used to obtain the next action corresponding to the next state. Then, the corresponding value is calculated through the two target critic networks and the minimum value is taken. The Adam optimizer is used to update the policy network and the critic network. The gradient is calculated through the loss function to update the parameters of the critic network. The gradient approximation of the policy network parameters is calculated to optimize the policy network parameters. At the same time, the policy network is guided to adjust the parameters with the help of the Critic network evaluation so that the target network parameters slowly follow the current Critic network update. The policy network and the Critic network are repeatedly trained iteratively. When the output action of the policy network tends to be stable and the loss of the Critic network no longer decreases significantly, the training is stopped.

[0011] A method for intelligently controlling the state of underground equipment, the method steps are as follows: S1: Deploy a multi-modal sensor array integrating vibration, stress, and temperature and collect multi-dimensional state data through a timestamp synchronization mechanism; S2: Construct a multimodal fusion model and extract multimodal state features from multidimensional state data based on the multimodal fusion model; S3: Construct an optimized reinforcement learning strategy through multimodal state features, obtain the optimal scheduling strategy from the multimodal state features based on the optimized reinforcement learning strategy combined with the improved deep deterministic policy gradient algorithm, and output the device operation control instructions through the optimal scheduling strategy.

[0012] The present invention has the following beneficial effects: In the present invention, firstly, by deploying a multimodal sensor array integrating vibration, stress and temperature, and using a timestamp synchronization mechanism to collect multidimensional state data, the operating state information of downhole equipment can be obtained comprehensively and accurately, and a multimodal fusion model is constructed, including a temporal convolutional network model, a graph convolutional network model and an improved GhostNet model, which can extract multimodal state features from multidimensional state data in a targeted manner; Secondly, the improved GhostNet model can dynamically adjust the convolution kernel configuration according to the characteristics of the temperature field image and accurately extract the temperature features of different areas. These features provide rich and effective information for subsequent equipment status analysis and decision-making. Based on the optimized reinforcement learning strategy and the improved deep deterministic policy gradient algorithm, combined with the dynamic reward function, the system can automatically learn and generate the optimal scheduling strategy according to the real-time status of the equipment and the job requirements; Finally, the improved deep deterministic policy gradient algorithm adopts a dual critic network structure and a target network soft update mechanism, and sets a large-capacity experience replay buffer, so that the system can learn and optimize strategies more stably during training. The dual critic network structure can more accurately evaluate the value of the state to the action and reduce the error in the strategy optimization process. The target network soft update mechanism avoids excessive fluctuations during training, improves the convergence speed and stability of the policy network and the critic network, enables the system to better adapt to the complex and changeable working environment underground, and provide more intelligent control instructions for the equipment. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] Figure 1 This is a system block diagram of an intelligent control system for downhole equipment status and a control method thereof proposed by the present invention.

[0014] Figure 2 This is a method step diagram of an intelligent control system for downhole equipment status and a control method thereof proposed by the present invention. DETAILED DESCRIPTION

[0015] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0016] Embodiment 1: Figure 1 As shown, the present invention proposes an intelligent control system for downhole equipment status, comprising: Multi-dimensional state acquisition module: deploys a multi-modal sensor array integrating vibration, stress, and temperature and collects multi-dimensional state data through a timestamp synchronization mechanism. The multi-dimensional state data includes vibration time domain sequence, stress distribution matrix, and temperature field image data; Vibration time domain sequence acquisition process: Deploy micro-electromechanical accelerometers with a sampling rate of 10kHz and a measurement range of ±20g. They are fixed to vibration-sensitive parts of downhole equipment (motor bearing seats, transmission gearbox housings) using surface mount technology. The micro-electromechanical accelerometers collect vibration signals during equipment operation in real time to form a vibration time domain sequence. ; Specifically, the vibration time domain series Contains information about the acceleration amplitude of equipment vibration changing over time, which is used to characterize the dynamic vibration characteristics of equipment operation; Stress distribution matrix acquisition: Deploy fiber Bragg grating stress sensors with a demodulation accuracy of ±1pm. The sensors are installed in the key stress-bearing structures of the equipment (hydraulic support columns, conveying equipment racks) by embedding or surface bonding. A stress monitoring network is constructed. The position of each sensor is located by optical time domain reflection technology. The stress value of each monitoring point of the equipment is obtained based on the linear relationship between wavelength drift and stress. The process is expressed as follows:

[0017] in, is the wavelength shift, is the central wavelength, is the elastic-optical coefficient, It is the degree of deformation of the equipment structure under stress; Finally, the formula Get stress values ,in is the elastic modulus of the material, and finally generates the stress distribution matrix ,in Indicates the coordinate index of the monitoring point; Temperature field image data acquisition: Install infrared thermal imaging equipment above or on the side of the underground equipment to perform a panoramic scan of the equipment surface. The module uses the non-contact temperature measurement principle to capture the temperature distribution of the equipment surface in real time and generate temperature field image data. ; The timestamp synchronization mechanism integrates multidimensional data in the following process: A hardware synchronization clock is used to provide a unified clock reference for the multimodal sensor array. The error in the data acquisition time of vibration, stress, and temperature sensors is less than 10μs. At the same time, accurate timestamps are added to the vibration time domain series, stress distribution matrix, and temperature field image. Through data fusion preprocessing, multi-source heterogeneous data are time-space aligned based on accurate timestamps to synchronize the vibration time domain series. , stress distribution matrix , Temperature field image Integrate into a multidimensional state data set including the time dimension, expressed as:

[0018] Where T represents the timestamp, and the multidimensional state data set As the input of subsequent modules, it realizes comprehensive and synchronous characterization of the operating status of downhole equipment.

[0019] Deep feature processing module: build a multimodal fusion model and extract multimodal state features from multidimensional state data based on the multimodal fusion model; The multimodal fusion model includes the temporal convolutional network model, the graph convolutional network model and the improved GhostNet model; Based on the time series convolutional network model for vibration time domain sequence Extract vibration time series feature groups and analyze stress distribution matrix based on graph convolutional network model Extract stress spatial feature groups, and extract temperature feature groups from temperature field images based on the improved GhostNet model; The process of extracting time series feature groups is: Through the time series convolution network, the time series convolution network is good at processing time series data. Through the expansion convolution, the dependency of the vibration signal in the time dimension is captured. Then the time series convolution network stacks multiple convolution layers. The bottom convolution extracts local vibration details (such as instantaneous vibration peak) from the dependency, and the high-level convolution integrates the bottom features to extract the overall vibration trend. After multiple convolution layers are processed, the vibration time series feature group is output. ; Specifically, vibration signals are typical time series data, and the equipment operation status information (such as fault signs) contained in them depends on the change law of the time dimension. The time series convolutional network model adapts to the analysis needs of time series data through dilated convolution and multi-layer convolutional layer stacking, mines features from time dependencies, and uses the smaller receptive field through the bottom convolution layer to capture the instantaneous changes of vibration signals (such as peak values ​​and mutations). These details are the direct manifestation of equipment abnormalities. In addition, the high-level convolution layer uses a larger receptive field to integrate the bottom-level local features, identify the overall change trend of the vibration signal (such as periodicity and long-term deterioration trend), and integrate the overall trend to form a vibration time series feature group. The process of extracting stress space feature groups is: The stress transfer graph is constructed using the graph convolutional network model. The sensor locations are used as nodes, and the edge weights are determined according to the physical distance to form a stress transfer topology. Then, the stress spatial distribution characteristics are extracted through the multi-layer convolution operation of the graph convolutional network, and the stress spatial feature group is output. ; Specifically, the stress distribution of the equipment has spatial correlation (for example, stress on a certain part will affect the adjacent parts). The graph convolution network constructs a stress transfer graph with sensor locations as nodes and physical distances as edge weights, converts stress distribution into graph structure data, and adapts to spatial feature analysis. At the same time, through multi-layer graph convolution operations, it propagates and aggregates features on the graph structure, explores the spatial distribution law of stress on the equipment structure (such as stress concentration areas and transfer paths), and determines the weak points of the equipment structure; The process of extracting temperature feature groups is: An improved GhostNet model is constructed. Based on the original GhostNet model, the improved GhostNet model dynamically adjusts the convolution kernel configuration to obtain the temperature feature group according to the spatial feature distribution of the temperature field image. Specifically: The core of the original Ghost model is to generate a large number of feature maps through cheap operations (depth-wise separable convolution) to reduce the computational complexity of traditional convolution. A Ghost model consists of two parts. First, a small number of initial feature maps are generated using ordinary convolution. Then, additional feature maps are generated from these initial feature maps through a series of linear transformations (depth-wise separable convolution). Finally, the initial feature maps and the generated feature maps are concatenated to obtain the final output feature map. Improve the GhostNet model to dynamically adjust the convolution kernel configuration to divide the image into feature areas, distinguishing between low-frequency areas (large areas with uniform temperature) and high-frequency areas (temperature mutation points): Low-frequency area (large area with uniform temperature): In order to capture the macroscopic temperature distribution characteristics, a large-size convolution kernel (5×5) is used. The large-size convolution kernel can cover a larger area and better extract the overall characteristics of the low-frequency area. Suppose the input temperature feature map is , the convolution operation in the low-frequency region is expressed as:

[0020] in, represents a 5×5 convolution operation, Represents a low-frequency temperature image; High frequency area (temperature mutation point): In order to accurately extract detail features such as temperature mutation points, a 3×3 convolution kernel is used. The 3×3 convolution kernel has an advantage in extracting local detail features and can more keenly capture small changes in temperature. Suppose the temperature input feature map is , the convolution operation for the high-frequency region can be expressed as:

[0021] in, represents a 3×3 convolution operation, Represents a high-frequency temperature image; The temperature images obtained after the low-frequency area and the high-frequency area are processed by different convolution kernels are fused. and Splice them together in the channel dimension to get the fused temperature field feature map ; The entire improved GhostNet model includes multiple (5) such fusion processes. The feature maps output by each fusion process together constitute a temperature feature group. The feature maps are arranged in the hierarchical order of the fusion process to form a temperature feature group. ; Specifically, the original GhostNet model usually generates a large number of feature maps through deep separable convolution operations to reduce the amount of calculation. However, for data with special spatial feature distribution such as temperature field images, the original model is difficult to extract features in a targeted manner. For example, there are low-frequency areas with large areas of uniform temperature and high-frequency areas with temperature mutation points in the temperature field image. It is necessary to adjust the convolution kernel configuration according to the characteristics of these different areas to better extract temperature features. This is the starting point for improving the GhostNet model. Secondly, by distinguishing between low-frequency areas (large areas with uniform temperature) and high-frequency areas (temperature mutation points), the feature extraction requirements of different areas can be clarified. Because low-frequency areas pay more attention to the overall temperature distribution, and high-frequency areas focus on capturing subtle changes and mutation points in temperature, different processing methods can be used to extract features more accurately. Finally, the entire improved GhostNet model contains multiple such fusion processes. The feature maps output by each fusion process together constitute a temperature feature group. These feature maps are arranged in the hierarchical order of the fusion process to form a temperature feature group. ,The temperature feature group constructed in this way covers the temperature ,feature information of different scales and levels, and comprehensively describes the ,features of the temperature field image, providing a rich information basis for ,subsequent equipment status analysis and decision-making; Vibration time series feature group , stress space feature group and temperature feature group Composition of multimodal state features as input for subsequent models.

[0022] State decision output module: construct an optimized reinforcement learning strategy through multimodal state features, obtain the optimal scheduling strategy from the multimodal state features based on the optimized reinforcement learning strategy combined with the improved deep deterministic policy gradient algorithm, and output the equipment operation control instructions through the optimal scheduling strategy; The process of building an optimized reinforcement learning strategy is: Define the initial state space and integrate multimodal state features The vibration time series feature group in , stress space feature group and temperature feature group Construct the state space, and assume that the state space is expressed as , The state space index comprehensively represents the operating state of the equipment, including equipment vibration, stress distribution, temperature change and other information; Define the action space as , is the action space index, action space It is a set of equipment control instructions, including variable frequency speed control parameters (set to 1-7 gears), hydraulic pressure adjustment gear (set to 3 gears), component start and stop combinations (set to 6 types), covering the main control operations of underground equipment, specifically: Frequency conversion speed regulation parameters: set to 1-7 levels to control the running speed of the equipment; For example: For underground belt conveyors in coal mines, gear 1 is used for low-speed operation when the equipment is started to avoid impact on the equipment due to high load at the moment of starting; gear 7 is used for high-speed operation of the equipment under full load and normal operation to improve transportation efficiency. Each gear corresponds to a different motor frequency, gear 1 corresponds to 20Hz, gear 2 corresponds to 25Hz, and so on, gear 7 corresponds to 50Hz, and the speed of the equipment is controlled by adjusting the frequency; Hydraulic pressure adjustment gear: set to gear 3, used to adjust the pressure of the hydraulic system; For example, taking the coal mine hydraulic support as an example, gear 1 is the low-pressure gear, which is suitable for the fine-tuning operation of the support. For example, when the roof pressure is small, gear 1 can be used to slightly adjust the support position. Gear 2 is the normal gear, and gear 3 is the high-pressure gear, which is used to provide strong support for the support when the roof pressure is large, ensuring that the support can effectively support the roof and ensure the safety of underground operations; Component start-stop combination: 6 types are set, covering different start-stop combinations of various components of the equipment; For example, for a coal mine roadheader, the six combinations include starting the cutting head alone, starting the conveyor belt alone, starting the spray dust suppression device alone, starting the cutting head and conveyor belt at the same time, starting the cutting head and spray dust suppression device at the same time, and starting the conveyor belt alone and the spray dust suppression device at the same time. By properly selecting the start-stop combination of components, different operation requirements can be met and equipment energy consumption and wear can be reduced. For example, when performing underground tunnel cleaning operations, only the conveyor belt and the spray dust suppression device can be started, and the cutting head can be turned off to reduce unnecessary energy consumption. Construct a dynamic learning reward mechanism, which is: Building a dynamic reward function , dynamic reward function include (Safety Reward), (Efficiency Reward), (Energy Consumption Reward), , , The weights corresponding to different rewards; Safety Rewards The setting standard is: when the equipment stress is lower than the safety threshold, the reward value approaches 1; otherwise, it decreases, guiding the strategy to optimize the safety of equipment operation. For example, if the stress is within the limit, the safety reward Close to 1; Efficiency Rewards The setting standard is: when the ratio of the actual output of the equipment to the planned output approaches 1, the reward value also approaches 1. For example, if the actual output reaches the planned output, the efficiency reward is 1; Energy consumption reward The setting standard is: when the ratio of the actual energy consumption of the equipment to the standard energy consumption approaches 1, the reward value also approaches 1. For example, when the actual energy consumption is equal to the standard energy consumption, the efficiency reward is is 1; Build an improved deep deterministic policy gradient algorithm: The deep deterministic policy gradient algorithm includes a policy network With Critic Network ; Improve the deep deterministic policy gradient algorithm to a dual-Critic network structure and set the capacity to The experience replay buffer stores historical interaction experience and randomly samples data from the buffer during training. When the target policy network and target critic network parameters are updated, the target network parameters are updated slowly following the current critic network update through the target network soft update mechanism. The process of combining the reinforcement learning strategy with the improved deep deterministic policy gradient algorithm to obtain the optimal scheduling strategy is as follows: Constructing 2 Critic networks and , using MLP structure, setting two hidden layers, the number of neurons in the first hidden layer is 256, the number of neurons in the second hidden layer is 128, the activation function of each hidden layer is ReLU function, the output layer uses linear activation function, the parameters and of the two critic networks are randomly initialized, and the evaluation state is affected by the action The value of and initialization of the target policy network , policy network Using the multi-layer perceptron MLP structure, the input state , output action , the initial parameters are consistent with the current network; Through dynamic experience playback, according to the current strategy network Interact with the environment space and observe the device state at each time step , output actions according to the policy network Control the device and get rewards after performing actions , and transfer to the next state , forming an experience sample ( , , , ), store the experience samples in a During training, batches of data are randomly sampled from the buffer periodically (e.g., 32 samples are sampled each time) and optimized through the target network soft update mechanism. The target network soft update mechanism is optimized as follows: For each sample ( , , , ), through the target strategy network Get the next state The corresponding next action , and then pass through two target Critic networks respectively and , calculate the value and , take the minimum value of the two; Update the policy network using the Adam optimizer With Critic Network and , calculate the gradient through the loss function to update the Critic network and The parameters are expressed as follows:

[0023] in, is the target value, is the mathematical expectation; By calculating the policy network parameters The gradient approximation of , optimizing the policy network parameters, is expressed as:

[0024] in, represents the parameters of the policy network, is the mathematical expectation, represents the approximate values ​​of the policy network parameters, Represents the policy network with respect to parameters The gradient of Represents the action output by the current policy network Below, the gradient of the Critic network about action a, which reflects the change in the evaluation of the Critic network on the output action of the current strategy; At the same time, the critic network evaluation is used to guide the policy network to adjust parameters so that the target network parameters slowly follow the current critic network update, so that the policy output action can obtain higher value in the critic network evaluation, thereby optimizing the policy; Repeat the steps of data collection, dynamic experience playback, and network update, iteratively train the policy network and the critic network, and gradually learn a better way to select actions. When the policy network output action tends to be stable and the critic network loss no longer decreases significantly, stop iterative training and generate the optimal scheduling strategy. At this time, the policy network can output the best control instructions according to the current state of the equipment. Specifically, the optimal scheduling instruction is a set of operations output by the strategy network based on the current equipment status, which is used to control the operation of the equipment. Its representation depends on the definition of the action space, which includes variable frequency speed regulation parameters, hydraulic pressure adjustment gears, and component start and stop combinations. Example: In an underground coal mine operation environment, the system outputs the best scheduling command to control the underground equipment. When it is detected that the equipment is in the startup stage and the load is large, the best scheduling command output by the strategy network is: The variable frequency speed regulation parameter instruction is gear 1 (corresponding to the motor frequency of 20Hz), which enables the equipment to start at a low speed to avoid the impact of the high load at the moment of starting. When the equipment enters stable operation and is fully loaded, the best scheduling instruction may be gear 7 (corresponding to the motor frequency of 50Hz) to improve transportation efficiency. The hydraulic pressure adjustment gear position instruction is that the hydraulic pressure adjustment gear position is set to gear 3. If the top plate pressure is small, the optimal scheduling instruction output by the strategy network is that the hydraulic pressure adjustment gear position is set to gear 1 (low pressure gear position) for fine-tuning operation of the bracket; The component start and stop combination instructions are: when performing tunnel cleaning operations, the optimal scheduling instruction output by the strategy network is to only start the conveyor belt and the spray dust reduction device, turn off the cutting head, and reduce unnecessary energy consumption. When performing excavation operations, the optimal scheduling instruction is to start the cutting head, conveyor belt and spray dust reduction device at the same time to meet the operational requirements.

[0025] Embodiment 2: Figure 2 As shown, the present invention proposes an intelligent control method for downhole equipment status, and the method steps are as follows: S1: Deploy a multi-modal sensor array integrating vibration, stress, and temperature and collect multi-dimensional state data through a timestamp synchronization mechanism; S2: Construct a multimodal fusion model and extract multimodal state features from multidimensional state data based on the multimodal fusion model; S3: Construct an optimized reinforcement learning strategy through multimodal state features, obtain the optimal scheduling strategy from the multimodal state features based on the optimized reinforcement learning strategy combined with the improved deep deterministic policy gradient algorithm, and output the device operation control instructions through the optimal scheduling strategy.

[0026] In the application, several formulas involved are calculated by removing dimensions and taking their numerical values. The formulas are established by collecting a large amount of data and performing software simulation to obtain a formula for the most recent actual situation. Some coefficients or weights in the formulas are set by technicians in this field according to actual conditions, so they will not be elaborated here.

[0027] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented using software, the above embodiments may be implemented in whole or in part in the form of a computer program product. A person of ordinary skill in the art may appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein may be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution.

[0028] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent control system for downhole equipment status, characterized in that: include: Multi-dimensional state acquisition module: deploys a multi-modal sensor array that integrates vibration, stress, and temperature and collects multi-dimensional state data through a timestamp synchronization mechanism; Deep feature processing module: build a multimodal fusion model and extract multimodal state features from multidimensional state data based on the multimodal fusion model; State decision output module: construct an optimized reinforcement learning strategy through multimodal state features, obtain the optimal scheduling strategy from the multimodal state features based on the optimized reinforcement learning strategy combined with the improved deep deterministic policy gradient algorithm, and output the equipment operation control instructions through the optimal scheduling strategy; The multimodal fusion model includes a temporal convolutional network model, a graph convolutional network model and an improved GhostNet model. The improved GhostNet model dynamically adjusts the convolution kernel configuration based on the original Ghost model according to the spatial feature distribution of the temperature field image. The reward mechanism for optimizing the reinforcement learning strategy is a dynamic learning reward mechanism, the improved deep deterministic policy gradient algorithm is a dual-Critic network structure, and the capacity is set to The target policy network and the target critic network are optimized and updated through the target network soft update mechanism.

2. The intelligent control system for downhole equipment status according to claim 1, characterized in that: The multi-dimensional state data includes vibration time domain sequence, stress distribution matrix, and temperature field image data; The vibration time domain sequence is obtained by collecting the vibration signal of the equipment in real time through a micro-electromechanical accelerometer during operation, the stress distribution matrix is ​​obtained by collecting the stress value of each monitoring point of the equipment through a fiber grating stress sensor, and the temperature field image data is obtained by configuring an infrared thermal imaging device to capture the surface temperature distribution of the equipment in real time.

3. The intelligent control system for downhole equipment status according to claim 2, characterized in that: The multi-dimensional state data is obtained by simultaneously adding precise timestamps to the vibration time domain sequence, stress distribution matrix and temperature field image data and performing data fusion preprocessing, and finally performing time-space alignment based on the precise timestamps.

4. The intelligent control system for downhole equipment status according to claim 1, characterized in that: The multi-modal state characteristics include a vibration time series characteristic group, a stress space characteristic group and a temperature characteristic group; The vibration time series feature group is extracted from the vibration time domain sequence based on the time series convolutional network model, the stress space feature group is extracted from the stress distribution matrix based on the graph convolutional network model, and the temperature feature group is extracted from the temperature field image based on the improved GhostNet model.

5. The intelligent control system for downhole equipment status according to claim 4, characterized in that: The process of extracting the temperature feature group from the temperature field image based on the improved GhostNet model is as follows: The temperature feature group is divided into feature regions, and the convolution kernel configuration is dynamically adjusted to distinguish between low-frequency and high-frequency regions; The low-frequency region is convolved with a 5×5 convolution kernel to obtain a low-frequency temperature image; The high-frequency region is convolved with a 3×3 convolution kernel to obtain a high-frequency temperature image; The low-frequency temperature image and the high-frequency temperature image are spliced ​​together in the channel dimension to obtain a fused temperature field feature map; The temperature field feature maps obtained through multiple fusion processes are arranged according to the hierarchical order of the fusion process to form a temperature feature group.

6. The intelligent control system for downhole equipment status according to claim 1, characterized in that: The specific implementation process of the optimized reinforcement learning strategy is as follows: The intelligent agent for optimizing the reinforcement learning strategy includes a state space, an action space, and a dynamic reward function; Define the multimodal state features as a state space; Define the action space, which includes the equipment control instruction set, including the frequency conversion speed regulation parameters, hydraulic pressure adjustment gear and component start and stop combination; Construct a dynamic reward function, which includes safety reward, efficiency reward and energy consumption reward, and assign corresponding weights to different rewards; When the equipment stress is lower than the safety threshold, the safety bonus value approaches 1; when the ratio of the actual output of the equipment to the planned output approaches 1, the efficiency bonus value approaches 1; when the ratio of the actual energy consumption of the equipment to the standard energy consumption approaches 1, the bonus value approaches 1.

7. The intelligent control system for downhole equipment status according to claim 1, characterized in that: The process of combining the optimized reinforcement learning strategy with the improved deep deterministic policy gradient algorithm to obtain the optimal scheduling strategy is: Construct two Critic networks and use the MLP structure. Set two hidden layers. The activation function of each layer is the ReLU function. The output layer uses a linear activation function. The parameters of the two Critic networks are randomly initialized to evaluate the value of the state to the action. At the same time, the target policy network is initialized. Through dynamic experience playback, the current policy network interacts with the environment space, observes the device state at each time step, regulates the device according to the policy network output action, obtains rewards after executing the action, and transfers to the next state to form experience samples, which are stored in a During training, batches of data are randomly sampled from the buffer periodically, and the optimal scheduling strategy is obtained through the target network soft update mechanism.

8. The intelligent control system for downhole equipment status according to claim 1, characterized in that: The target network soft update mechanism implementation process is: For each sample obtained from the experience replay buffer, the target policy network is used to obtain the next action corresponding to the next state. Then, the corresponding value is calculated through the two target critic networks and the minimum value is taken. The Adam optimizer is used to update the policy network and the critic network. The gradient is calculated through the loss function to update the parameters of the critic network. The gradient approximation of the policy network parameters is calculated to optimize the policy network parameters. At the same time, the policy network is guided to adjust the parameters with the help of the Critic network evaluation so that the target network parameters slowly follow the current Critic network update. The policy network and the Critic network are repeatedly trained iteratively. When the output action of the policy network tends to be stable and the loss of the Critic network no longer decreases significantly, the training is stopped.

9. A method for intelligently controlling the state of downhole equipment, using the control system according to any one of claims 1 to 8, the method steps are as follows: S1: Deploy a multi-modal sensor array integrating vibration, stress, and temperature and collect multi-dimensional state data through a timestamp synchronization mechanism; S2: Construct a multimodal fusion model and extract multimodal state features from multidimensional state data based on the multimodal fusion model; S3: Construct an optimized reinforcement learning strategy through multimodal state features, obtain the optimal scheduling strategy from the multimodal state features based on the optimized reinforcement learning strategy combined with the improved deep deterministic policy gradient algorithm, and output the device operation control instructions through the optimal scheduling strategy.

Citation Information

Patent Citations

  • Autonomous controllable heterogeneous intelligent computing service platform and intelligent scene matching method

    CN115202868A

  • Deep groove ball bearing fault diagnosis method fusing frequency domain features and improved residual network

    CN118626822A

  • Desilting robot intelligent control method and system based on deep learning

    CN119392782A

  • Intelligent heterogeneous metadata scheduling optimization method and system based on reinforcement learning algorithm

    CN119719782A

  • Multi-intelligence federal reinforcement learning-based vehicle-road cooperative control system and method at complex intersection

    US11862016B1

Cited By

  • Intelligent prediction method and system for mine air demand

    CN120217905A

  • Mine fully-mechanized coal mining equipment state identification and Wi-Fi transmission method based on deep learning

    CN120687892A

  • Intelligent granary ventilation and energy consumption optimization decision-making method based on reinforcement learning

    CN120725247A

  • Intelligent warehouse ventilation and energy consumption optimization decision method based on reinforcement learning

    CN120725247B