An intelligent control system for downhole equipment status and its control method

By deploying multimodal sensor arrays and building multimodal fusion models, combined with optimization reinforcement learning strategies, the single problem of downhole equipment status monitoring and control is solved, comprehensive monitoring and intelligent control of equipment status is achieved, and failure risk and energy waste are reduced.

CN120030445BActive Publication Date: 2025-07-22NUOWENKE BLOWER FAN BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510510358.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-07-22
Estimated Expiration
2045-04-23

AI Technical Summary

Technical Problem

The existing underground equipment status monitoring and control methods are single, and it is difficult to fully and accurately reflect the equipment status, resulting in insufficient potential fault warning capabilities and the control strategy cannot be dynamically adjusted, resulting in energy waste and equipment wear.

Method used

Deploy a multimodal sensor array to collect multidimensional state data, build a multimodal fusion model and optimize reinforcement learning strategy, generate the optimal scheduling strategy through the improved GhostNet model and deep deterministic strategy gradient algorithm, and output device operation control instructions.

Benefits of technology

It realizes comprehensive and accurate monitoring and intelligent control of downhole equipment status, reduces the risk of equipment failure, optimizes energy utilization, and reduces maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030445B_ABST
    Figure CN120030445B_ABST
Patent Text Reader

Abstract

The present invention discloses an intelligent control system for downhole equipment status and its control method, including a multi-dimensional status acquisition module: deploying a multi-modal sensor array integrating vibration, stress, and temperature, and collecting multi-dimensional status data through a timestamp synchronization mechanism; a deep feature processing module: constructing a multi-modal fusion model, and extracting multi-modal status features from the multi-dimensional status data based on the multi-modal fusion model; a status decision output module: constructing an optimized reinforcement learning strategy through the multi-modal status features, obtaining an optimal scheduling strategy from the multi-modal status features based on the optimized reinforcement learning strategy combined with an improved deep deterministic policy gradient algorithm, and outputting equipment operation control instructions through the optimal scheduling strategy, so that the system can better adapt to the complex and changeable downhole working conditions environment and provide more intelligent control instructions for the equipment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent control, and particularly to an intelligent control system for downhole equipment status and its control method. Background Art

[0002] With the continuous expansion of downhole operation scale and the improvement of equipment complexity, traditional manual experience judgment and simple automation control are difficult to meet the modern high-efficiency and safe production requirements. Therefore, it is urgent to develop a system and method that can monitor the status of downhole equipment in real time and comprehensively, and make intelligent decisions and precise control based on multi-source data;

[0003] In the existing technology, the methods for monitoring and controlling downhole equipment status are relatively single. Usually, only a few types of sensors are relied on to obtain limited equipment operation information, which is difficult to comprehensively and accurately reflect the actual status of the equipment. For example, only monitoring the vibration of the equipment cannot timely understand other key parameters such as equipment stress distribution and temperature change, resulting in insufficient early warning ability for potential equipment failures. At the same time, in terms of equipment control strategies, most adopt fixed control modes and fail to make dynamic adjustments according to the real-time status of the equipment and the complex and changeable downhole working conditions. This control method not only cannot give full play to the performance of the equipment, but also will cause energy waste and increased equipment wear due to the equipment being in an unreasonable operation state for a long time, increasing the equipment maintenance cost and failure risk. Therefore, an intelligent control system for downhole equipment status and its control method are proposed here. Summary of the Invention

[0004] In order to overcome the above-mentioned defects of the prior art and achieve the above object, the present invention proposes the following technical solutions:

[0005] An intelligent control system for downhole equipment status, comprising:

[0006] A multi-dimensional state acquisition module: deploying an integrated multi-modal sensor array of vibration, stress, and temperature and through a timestamp synchronization mechanism, collecting multi-dimensional state data;

[0007] A deep feature processing module: constructing a multi-modal fusion model, and extracting multi-modal state features from the multi-dimensional state data based on the multi-modal fusion model;

[0008] A state decision output module: constructing an optimized reinforcement learning strategy through the multi-modal state features, obtaining an optimal scheduling strategy from the multi-modal state features based on the optimized reinforcement learning strategy combined with an improved deep deterministic policy gradient algorithm, and outputting equipment operation control instructions through the optimal scheduling strategy;

[0009] The multi-modal fusion model includes a temporal convolutional network model, a graph convolutional network model, and an improved GhostNet model. The improved GhostNet model dynamically adjusts the convolutional kernel configuration according to the spatial feature distribution of the temperature field image based on the original Ghost model.

[0010] The reward mechanism for optimizing the reinforcement learning strategy is a dynamic learning reward mechanism. The improved deep deterministic policy gradient algorithm has a dual Critic network structure and sets an experience replay buffer with a capacity of to optimize and update the target policy network and the target Critic network through the target network soft update mechanism.

[0011] The multi-dimensional state data includes vibration time-domain sequences, stress distribution matrices, and temperature field image data.

[0012] The vibration time-domain sequences are obtained by collecting vibration signals during the operation of the device in real time through a micro-electro-mechanical accelerometer. The stress distribution matrices are obtained by collecting stress values at each monitoring point of the device through fiber Bragg grating stress sensors. The temperature field image data is obtained by configuring an infrared thermal imaging device to capture the surface temperature distribution of the device in real time.

[0013] The multi-dimensional state data is obtained by adding accurate timestamps to the vibration time-domain sequences, stress distribution matrices, and temperature field image data simultaneously, performing data fusion preprocessing, and finally performing spatio-temporal alignment based on the accurate timestamps.

[0014] The multi-modal state features include vibration time-series feature groups, stress spatial feature groups, and temperature feature groups.

[0015] The vibration time-series feature groups are extracted from the vibration time-domain sequences based on the temporal convolutional network model. The stress spatial feature groups are extracted from the stress distribution matrices based on the graph convolutional network model. The temperature feature groups are extracted from the temperature field images based on the improved GhostNet model.

[0016] The process of extracting the temperature feature groups from the temperature field images based on the improved GhostNet model is as follows:

[0017] Perform feature region division on the temperature feature groups, and dynamically adjust the convolutional kernel configuration to distinguish the low-frequency region and the high-frequency region.

[0018] The low-frequency region uses a 5×5 size convolutional kernel for convolution operation to obtain a low-frequency temperature image.

[0019] The high-frequency region uses a 3×3 size convolutional kernel for convolution operation to obtain a high-frequency temperature image.

[0020] Stitch the low-frequency temperature image and the high-frequency temperature image together in the channel dimension to obtain a fused temperature field feature map.

[0021] The obtained temperature field feature maps are arranged in the hierarchical order of the fusion processes through multiple fusion processes to form a temperature feature group.

[0022] The specific implementation process of the optimized reinforcement learning strategy is as follows:

[0023] The agent of the optimized reinforcement learning strategy includes a state space, an action space, and a dynamic reward function;

[0024] Define the multi-modal state features as the state space;

[0025] Define the action space. The action space includes a device control instruction set, and the device control instruction set includes variable frequency speed regulation parameters, hydraulic pressure adjustment gears, and component start / stop combinations;

[0026] Construct a dynamic reward function. The dynamic reward function includes safety rewards, efficiency rewards, and energy consumption rewards, and corresponding weights are assigned to different rewards;

[0027] When the device stress is lower than the safety threshold, the safety reward value approaches 1. When the ratio of the actual output of the device to the planned output approaches 1, the efficiency reward value approaches 1. When the ratio of the actual energy consumption of the device to the standard energy consumption approaches 1, the reward value approaches 1.

[0028] The process of combining the optimized reinforcement learning strategy with the improved deep deterministic policy gradient algorithm to obtain the optimal scheduling strategy is as follows:

[0029] Construct 2 Critic networks and adopt the MLP structure. Set two hidden layers, with the activation function of each layer being the ReLU function, and the output layer using the linear activation function. The parameters of the two Critic networks are randomly initialized respectively to evaluate the value of the state to the action, and at the same time initialize the target policy network;

[0030] Through dynamic experience replay, interact with the environment space according to the current policy network, and observe the device state at each time step. Control the device according to the action output by the policy network, obtain the reward after executing the action, and transfer to the next state to form an experience sample. Store the experience sample in the experience replay buffer with a capacity of . During training, randomly sample batch data from the buffer regularly, and optimize through the target network soft update mechanism to obtain the optimal scheduling strategy.

[0031] The implementation process of the target network soft update mechanism is as follows:

[0032] For each sample sampled from the experience replay buffer, the next action corresponding to the next state is obtained through the target policy network. Then, the corresponding values are calculated through two target Critic networks and the minimum of the two is taken. The Adam optimizer is used to update the policy network and the Critic network, and the gradient is calculated through the loss function to update the parameters of the Critic network;

[0033] Calculate the gradient approximation value of the policy network parameters to optimize the policy network parameters. At the same time, with the help of the Critic network evaluation, guide the policy network to adjust the parameters so that the target network parameters slowly follow the update of the current Critic network. Repeat the iterative training of the policy network and the Critic network. When the actions output by the policy network tend to be stable and the loss of the Critic network no longer decreases significantly, stop the training.

[0034] An intelligent control method for downhole equipment state, and the method steps are as follows:

[0035] S1: Deploy a multi-modal sensor array integrating vibration, stress, and temperature, and through a timestamp synchronization mechanism, collect multi-dimensional state data;

[0036] S2: Construct a multi-modal fusion model, and extract multi-modal state features from the multi-dimensional state data based on the multi-modal fusion model;

[0037] S3: Construct an optimized reinforcement learning policy through the multi-modal state features, obtain the optimal scheduling policy from the multi-modal state features based on the optimized reinforcement learning policy combined with the improved deep deterministic policy gradient algorithm, and output device operation control instructions through the optimal scheduling policy.

[0038] The present invention has the following beneficial effects:

[0039] In the present invention, first, by deploying a multi-modal sensor array integrating vibration, stress, and temperature, and using a timestamp synchronization mechanism to collect multi-dimensional state data, the operation state information of downhole equipment can be comprehensively and accurately obtained, and a multi-modal fusion model is constructed, including a temporal convolutional network model, a graph convolutional network model, and an improved GhostNet model, which can specifically extract multi-modal state features from the multi-dimensional state data;

[0040] Secondly, the improved GhostNet model can dynamically adjust the convolutional kernel configuration according to the characteristics of the temperature field image, and accurately extract the temperature features of different regions. These features provide rich and effective information for subsequent equipment state analysis and decision-making. Based on the optimized reinforcement learning policy and the improved deep deterministic policy gradient algorithm, combined with a dynamic reward function, the system can automatically learn and generate the optimal scheduling policy according to the real-time state of the equipment and the operation requirements;

[0041] Finally, the improved Deep Deterministic Policy Gradient algorithm adopts a dual-Critic network structure and a target network soft update mechanism, and sets up a large-capacity experience replay buffer, enabling the system to learn and optimize policies more stably during training. The dual-Critic network structure can more accurately evaluate the value of states for actions, reducing errors in the policy optimization process. The target network soft update mechanism avoids excessive fluctuations during training, improving the convergence speed and stability of the policy network and the Critic network, enabling the system to better adapt to the complex and changeable working conditions underground and providing more intelligent control instructions for the equipment. Brief Description of the Drawings

[0042] Figure 1 It is a system block diagram of an intelligent control system and its control method for underground equipment states proposed by the present invention.

[0043] Figure 2 It is a method step diagram of an intelligent control system and its control method for underground equipment states proposed by the present invention. Detailed Embodiments

[0044] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0045] Embodiment 1: As Figure 1 shown, an intelligent control system for underground equipment states proposed by the present invention includes:

[0046] Multi-dimensional state acquisition module: Deploy an integrated multi-modal sensor array of vibration, stress, and temperature and through a timestamp synchronization mechanism, collect multi-dimensional state data, where the multi-dimensional state data includes vibration time-domain sequences, stress distribution matrices, and temperature field image data;

[0047] Process of obtaining vibration time-domain sequences:

[0048] Deploy micro-electromechanical accelerometers. The micro-electromechanical accelerometers have a sampling rate of 10 kHz and a measurement range of ±20 g, and are fixed to the vibration-sensitive parts of underground equipment (motor bearing seats, transmission gearbox housings) using surface mount technology. The micro-electromechanical accelerometers continuously collect vibration signals during equipment operation to form vibration time-domain sequences ;

[0049] Specifically, the vibration time-domain sequence contains information on the variation of the acceleration amplitude of equipment vibration over time and is used to characterize the dynamic vibration characteristics of equipment operation;

[0050] Obtaining the stress distribution matrix:

[0051] Deploy fiber Bragg grating stress sensors with a demodulation accuracy of ±1 pm. The sensors are installed in a buried or surface-mounted manner on the key stressed structures of the equipment (hydraulic support columns, conveyor equipment frames) to construct a stress monitoring network. Locate the positions of each sensor through optical time domain reflectometry technology, and obtain the stress values at each monitoring point of the equipment based on the linear relationship between wavelength drift and stress. The process is expressed as:

[0052]

[0053] Where, is the wavelength drift amount, is the central wavelength, is the elasto-optic coefficient, is the degree of deformation generated by the equipment structure under the action of stress;

[0054] Finally, obtain the stress value through the formula where , and is the material elastic modulus, and finally generate the stress distribution matrix where represents the coordinate index of the monitoring point;

[0055] Obtaining temperature field image data:

[0056] Configure an infrared thermal imaging device to be installed above or on the side of the underground equipment, and perform a panoramic scan of the equipment surface. The module captures the temperature distribution on the equipment surface in real time through the non-contact temperature measurement principle and generates temperature field image data ;

[0057] The process of integrating multi-dimensional data by the timestamp synchronization mechanism is as follows:

[0058] Use a hardware synchronous clock to provide a unified clock reference for the multi-modal sensor array. The error in the data acquisition time of the vibration, stress, and temperature sensors is less than 10 μs. At the same time, add accurate timestamps to the vibration time domain sequence, stress distribution matrix, and temperature field image respectively. Through data fusion preprocessing, perform spatio-temporal alignment on the multi-source heterogeneous data based on the accurate timestamps, and integrate the vibration time domain sequence , stress distribution matrix , and temperature field image into a multi-dimensional state dataset including the time dimension, expressed as:

[0059]

[0060] Where, T represents the timestamp, and use the multi-dimensional state dataset as the input of the subsequent module to achieve a comprehensive and synchronous characterization of the operating state of the underground equipment.

[0061] Deep feature processing module: Construct a multimodal fusion model, and extract multimodal state features from multi-dimensional state data based on the multimodal fusion model;

[0062] The multimodal fusion model includes a temporal convolutional network model, a graph convolutional network model, and an improved GhostNet model;

[0063] Based on the temporal convolutional network model for the vibration time-domain sequence Extract the vibration time-series feature group, and based on the graph convolutional network model for the stress distribution matrix Extract the stress space feature group, and based on the improved GhostNet model, extract the temperature feature group from the temperature field image;

[0064] The process of extracting the time-series feature group is as follows:

[0065] Through the temporal convolutional network, which is good at processing time-series data, through dilated convolution, capture the dependencies of the vibration signal in the time dimension, and then the temporal convolutional network stacks multiple convolutional layers. The bottom convolutional layer extracts local vibration details (such as instantaneous vibration peaks) from the dependencies, and the high-level convolutional layer integrates the bottom features to extract the overall vibration trend. After multiple convolutional layers are processed, the vibration time-series feature group is output ;

[0066] Specifically, the vibration signal is typical time-series data, and the equipment operating state information (such as fault symptoms) it contains depends on the variation law in the time dimension. The temporal convolutional network model adapts to the analysis requirements of time-series data through dilated convolution and multi-layer convolutional layer stacking, mines features from the time dependencies, and uses a smaller receptive field through the bottom convolutional layer to capture the instantaneous changes (such as peaks, mutations) of the vibration signal. These details are the direct manifestations of equipment abnormalities, and through the high-level convolutional layer using a larger receptive field, integrate the bottom local features, identify the overall change trend of the vibration signal (such as periodic patterns, long-term deterioration trends), and perform overall trend integration to form the vibration time-series feature group;

[0067] The process of extracting the stress space feature group is as follows:

[0068] Use the graph convolutional network model to construct a stress transfer graph, with the sensor positions as nodes, determine the edge weights according to the physical distance to form a stress transfer topological structure, and then through the multi-layer convolutional operations of the graph convolutional network, extract the stress spatial distribution features and output the stress space feature group ;

[0069] Specifically, the stress distribution of the equipment has spatial correlation (for example, stress on a certain part will affect the adjacent parts). The graph convolution network constructs a stress transfer graph with sensor locations as nodes and physical distances as edge weights, converts stress distribution into graph structure data, and adapts to spatial feature analysis. At the same time, through multi-layer graph convolution operations, it propagates and aggregates features on the graph structure, explores the spatial distribution law of stress on the equipment structure (such as stress concentration areas and transfer paths), and determines the weak points of the equipment structure;

[0070] The process of extracting temperature feature groups is:

[0071] An improved GhostNet model is constructed. Based on the original GhostNet model, the improved GhostNet model dynamically adjusts the convolution kernel configuration to obtain the temperature feature group according to the spatial feature distribution of the temperature field image. Specifically:

[0072] The core of the original Ghost model is to generate a large number of feature maps through cheap operations (depth-wise separable convolution) to reduce the computational complexity of traditional convolution. A Ghost model consists of two parts. First, a small number of initial feature maps are generated using ordinary convolution. Then, additional feature maps are generated from these initial feature maps through a series of linear transformations (depth-wise separable convolution). Finally, the initial feature maps and the generated feature maps are concatenated to obtain the final output feature map.

[0073] Improve the GhostNet model to dynamically adjust the convolution kernel configuration to divide the image into feature areas, distinguishing between low-frequency areas (large areas with uniform temperature) and high-frequency areas (temperature mutation points):

[0074] Low-frequency area (large area with uniform temperature): In order to capture the macroscopic temperature distribution characteristics, a large-size convolution kernel (5×5) is used. The large-size convolution kernel can cover a larger area and better extract the overall characteristics of the low-frequency area. Suppose the input temperature feature map is , the convolution operation in the low-frequency region is expressed as:

[0075]

[0076] in, represents a 5×5 convolution operation, Represents a low-frequency temperature image;

[0077] High frequency area (temperature mutation point): In order to accurately extract detail features such as temperature mutation points, a 3×3 convolution kernel is used. The 3×3 convolution kernel has an advantage in extracting local detail features and can more keenly capture small changes in temperature. Suppose the temperature input feature map is , the convolution operation for the high-frequency region can be expressed as:

[0078]

[0079] Among them, represents a 3×3 convolution operation, represents the high-frequency temperature image;

[0080] Fuse the temperature images obtained by processing the low-frequency region and the high-frequency region with different convolutional kernels, and and are concatenated in the channel dimension to obtain the fused temperature field feature map ;

[0081] There are multiple (5) such fusion processes in the entire improved GhostNet model. The feature maps output by each fusion process together constitute a temperature feature group. Arrange the feature maps in the hierarchical order of the fusion process to form a temperature feature group ;

[0082] Specifically, the original GhostNet model usually generates a large number of feature maps through depthwise separable convolution operations to reduce the computational amount. However, for data with special spatial feature distributions such as temperature field images, it is difficult for the original model to extract features specifically. For example, there are low-frequency regions with large areas of uniform temperature and high-frequency regions with temperature mutation points in the temperature field image. It is necessary to adjust the convolutional kernel configuration according to the characteristics of these different regions to better extract temperature features. This is the starting point of the improved GhostNet model;

[0083] Secondly, by distinguishing the low-frequency region (large area of uniform temperature region) and the high-frequency region (temperature mutation point), the feature extraction requirements of different regions can be clarified. Because the low-frequency region pays more attention to the overall temperature distribution, and the high-frequency region focuses on capturing the subtle changes and mutation points of temperature, so using different processing methods for them can extract features more accurately;

[0084] Finally, there are multiple such fusion processes in the entire improved GhostNet model. The feature maps output by each fusion process together constitute a temperature feature group. Arrange these feature maps in the hierarchical order of the fusion process to form a temperature feature group , and the temperature feature group constructed in this way covers temperature feature information of different scales and levels, comprehensively describes the features of the temperature field image, and provides a rich information basis for subsequent equipment status analysis and decision-making;

[0085] Vibration time series feature group , stress space feature group and temperature feature group constitute the multi-modal state features as the input of the subsequent model.

[0086] State decision output module: constructs an optimized reinforcement learning policy through multi-modal state features, obtains the optimal scheduling policy from the multi-modal state features based on the optimized reinforcement learning policy combined with the improved deep deterministic policy gradient algorithm, and outputs device operation control instructions through the optimal scheduling policy;

[0087] The process of constructing the optimized reinforcement learning policy is as follows:

[0088] Define the initial state space and integrate multi-modal state features the vibration time series feature group in and the stress space feature group and the temperature feature group constitute the state space, denoted as , where

[0089] is the state space index. This space comprehensively represents the operating state of the device, including information such as the vibration condition, stress distribution, and temperature change of the device;

[0089] Define the action space as , where is the action space index. The action space

[0090] is the device control instruction set, including variable frequency speed regulation parameters (set to 1-7 gears), hydraulic pressure adjustment gears (set to 3 gears), and component start-stop combinations (set to 6 types), covering the main control operations of underground equipment. Specifically:

[0091] Variable frequency speed regulation parameters: Set to 1-7 gears, used to control the operating speed of the device;

[0091] Example: For a belt conveyor in a coal mine underground, gear 1 is used for low-speed operation during equipment startup to avoid impact on the equipment caused by high load at startup; gear 7 is used for high-speed operation when the equipment is fully loaded and operating normally to improve transportation efficiency. Each gear corresponds to a different motor frequency. Gear 1 corresponds to 20Hz, gear 2 corresponds to 25Hz, and so on. Gear 7 corresponds to 50Hz. The speed of the equipment is controlled by adjusting the frequency;

[0092] Hydraulic pressure adjustment gears: Set to 3 gears, used to adjust the pressure of the hydraulic system;

[0093] Example: Taking a coal mine hydraulic support as an example, gear 1 is the low-pressure gear, suitable for fine adjustment operations of the support. For example, when the roof pressure is relatively small, gear 1 can be used for slight adjustment of the support position. Gear 2 is the normal gear, and gear 3 is the high-pressure gear, used to provide strong support for the support when the roof pressure is large, ensuring that the support can effectively support the roof and guaranteeing the safety of underground operations;

[0094] Component start-stop combinations: Set to 6 types, covering different start-stop combination methods of each component of the equipment;

[0095] Example: For a coal mine roadheader, the 6 combinations include the cutting head starting alone, the conveyor belt starting alone, the spray dust suppression device starting alone, the cutting head and the conveyor belt starting simultaneously, the cutting head and the spray dust suppression device starting simultaneously, and the conveyor belt starting alone and the spray dust suppression device starting simultaneously. By reasonably selecting the start-stop combinations of components, different operation requirements can be met, and the energy consumption and wear of the equipment can be reduced. For example, when cleaning the underground roadway, only the conveyor belt and the spray dust suppression device can be started, and the cutting head can be turned off to reduce unnecessary energy consumption;

[0096] Construct a dynamic learning reward mechanism, and the dynamic learning reward mechanism is as follows:

[0097] Construct a dynamic reward function , the dynamic reward function includes (safety reward), (efficiency reward), (energy consumption reward), , , are the weights corresponding to different rewards;

[0098] The safety reward is set as: when the stress of the equipment is lower than the safety threshold, the reward value approaches 1; otherwise, it decreases, guiding the strategy to optimize the operation safety of the equipment. For example, if the stress is not exceeded, the safety reward is close to 1;

[0099] The efficiency reward is set as: when the ratio of the actual output of the equipment to the planned output approaches 1, the reward value also approaches 1. For example, if the actual output reaches the planned output, the efficiency reward is 1;

[0100] The energy consumption reward is set as: when the ratio of the actual energy consumption of the equipment to the standard energy consumption approaches 1, the reward value also approaches 1. For example, when the actual energy consumption is equal to the standard energy consumption, the efficiency reward is 1;

[0101] Construct an improved deep deterministic policy gradient algorithm:

[0102] The deep deterministic policy gradient algorithm includes a policy network and a Critic network ;

[0103] The improved deep deterministic policy gradient algorithm is a double Critic network structure, and the capacity is set to The experience replay buffer stores historical interaction experiences and randomly samples data from the buffer during training. When optimizing and updating the parameters of the target policy network and the target Critic network through the target network soft update mechanism, the target network parameters slowly follow the updates of the current Critic network;

[0104] The process of combining the reinforcement learning policy with the improved deep deterministic policy gradient algorithm to obtain the optimal scheduling policy is as follows:

[0105] Construct two Critic networks and , using the MLP structure, setting two hidden layers. The number of neurons in the first hidden layer is 256, and the number of neurons in the second hidden layer is 128. The activation function for each hidden layer is the ReLU function, and the output layer uses a linear activation function. The parameters of the two Critic networks are randomly initialized respectively to evaluate the value of the state-action . At the same time, initialize the target policy network , and the policy network adopts the multi-layer perceptron MLP structure, takes the input state , and outputs the action . The initial parameters are the same as those of the current network;

[0106] Through dynamic experience replay, interact with the environment space according to the current policy network , and observe the device state at each time step. According to the action output by the policy network, regulate the device, obtain the reward after executing the action, and transfer to the next state , forming an experience sample ([[]] , , , ). Store the experience sample in the experience replay buffer with a capacity of . During training, regularly randomly sample a batch of data from the buffer (such as sampling 32 samples each time), and optimize through the target network soft update mechanism;

[0107] The target network soft update mechanism is optimized as follows:

[0108] For each sample ([[]] , , , ) sampled from the experience replay buffer, obtain the next action corresponding to the next state through the target policy network , and then respectively pass through the two target Critic networks and , calculate the value and , take the minimum of the two;

[0109] Update the policy network using the Adam optimizer with the Critic network and , calculate the gradient through the loss function to update the parameters of the Critic network and , and the formula is expressed as:

[0110]

[0111] where, is the target value, is the mathematical expectation;

[0112] By calculating the approximate value of the gradient of the policy network parameters , optimize the policy network parameters, which is expressed as:

[0113]

[0114] where, represents the parameters of the policy network, is the mathematical expectation, represents the approximate value of the policy network parameters, represents the gradient of the policy network with respect to the parameters , represents the action output by the current policy network under the condition, the gradient of the Critic network with respect to the action a, which reflects the evaluation change of the Critic network for the action output by the current policy;

[0115] At the same time, with the help of the evaluation of the Critic network, guide the policy network to adjust the parameters so that the target network parameters slowly follow the update of the current Critic network, so that the action output by the policy obtains a higher value in the evaluation of the Critic network, and then optimize the policy;

[0116] Repeat the steps of data collection, dynamic experience replay, and network update, iteratively train the policy network and the Critic network. The policy network gradually learns a better action selection method. When the action output by the policy network tends to be stable and the loss of the Critic network no longer decreases significantly, stop the iterative training to generate the optimal scheduling policy. At this time, the policy network can output the best control instruction according to the current state of the device;

[0117] Specifically, the optimal scheduling instruction is the set of operations output by the policy network based on the current device status for regulating the device operation. Its representation form depends on the definition of the action space, which includes variable frequency speed regulation parameters, hydraulic pressure adjustment gears, and component start / stop combinations.

[0118] For example:

[0119] In an underground coal mine operation environment, the system outputs the optimal scheduling instruction to control underground equipment. When it is detected that the equipment is in the startup phase and has a large load, the optimal scheduling instruction output by the policy network is:

[0120] The variable frequency speed regulation parameter instruction is at gear 1 (corresponding to a motor frequency of 20 Hz), enabling the equipment to start at a low speed and avoiding impact on the equipment caused by the high load during the startup moment. When the equipment enters stable operation and is fully loaded, the optimal scheduling instruction may be at gear 7 (corresponding to a motor frequency of 50 Hz) to improve the transportation efficiency.

[0121] The hydraulic pressure adjustment gear instruction is to set the hydraulic pressure adjustment gear to gear 3. If the roof pressure is small, the optimal scheduling instruction output by the policy network is that the hydraulic pressure adjustment gear is at gear 1 (low pressure gear) for fine-tuning operations of the support.

[0122] The component start / stop combination instruction is that during roadway cleaning operations, the optimal scheduling instruction output by the policy network only starts the conveyor belt and the spray dust suppression device and shuts down the cutting head to reduce unnecessary energy consumption. While during tunneling operations, the optimal scheduling instruction is to start the cutting head, conveyor belt, and spray dust suppression device simultaneously to meet the operation requirements.

[0123] Embodiment 2: As Figure 2 shown, an intelligent control method for underground equipment status proposed by the present invention has the following method steps:

[0124] S1: Deploy a multi-modal sensor array integrating vibration, stress, and temperature and, through a timestamp synchronization mechanism, collect multi-dimensional status data;

[0125] S2: Construct a multi-modal fusion model and extract multi-modal status features from the multi-dimensional status data based on the multi-modal fusion model;

[0126] S3: Construct an optimized reinforcement learning policy through the multi-modal status features, obtain the optimal scheduling policy from the multi-modal status features based on the optimized reinforcement learning policy combined with the improved deep deterministic policy gradient algorithm, and output device operation control instructions through the optimal scheduling policy.

[0127] In the application, several formulas involved are calculated by taking their numerical values after dimensionless treatment. The establishment of the formulas is obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. Some coefficients or weights in the formulas are set by those skilled in the art according to the actual situation, so no more elaboration will be made here.

[0128] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution.

[0129] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An intelligent control system for downhole equipment status, characterized in that, It includes: Multi-dimensional state acquisition module: Deploy an integrated multi-modal sensor array of vibration, stress, and temperature, and collect multi-dimensional state data through a timestamp synchronization mechanism. Deep feature processing module: Construct a multi-modal fusion model, and extract multi-modal state features from multi-dimensional state data based on the multi-modal fusion model. State decision output module: Construct an optimized reinforcement learning policy through multi-modal state features, obtain the optimal scheduling policy from multi-modal state features based on the optimized reinforcement learning policy combined with the improved deep deterministic policy gradient algorithm, and output device operation control instructions through the optimal scheduling policy. The multi-modal fusion model includes a temporal convolutional network model, a graph convolutional network model, and an improved GhostNet model. The improved GhostNet model dynamically adjusts the convolutional kernel configuration according to the spatial feature distribution of the temperature field image based on the original Ghost model. The reward mechanism for optimizing the reinforcement learning strategy is a dynamic learning reward mechanism. The improved deep deterministic policy gradient algorithm has a dual-Critic network structure and sets an experience replay buffer with a capacity of . The target policy network and the target Critic network are optimized and updated through the target network soft update mechanism.

2. The intelligent control system for the downhole equipment status according to claim 1, wherein, The multi-dimensional state data includes vibration time-domain sequences, stress distribution matrices, and temperature field image data. The vibration time-domain sequence is obtained by collecting vibration signals during the operation of the device in real time through a micro-electro-mechanical accelerometer. The stress distribution matrix is obtained by collecting stress values at each monitoring point of the device through a fiber Bragg grating stress sensor. The temperature field image data is obtained by configuring an infrared thermal imaging device to capture the surface temperature distribution of the device in real time.

3. The intelligent control system for the downhole equipment state according to claim 2, wherein The multi-dimensional state data is obtained by adding accurate timestamps to the vibration time-domain sequence, stress distribution matrix, and temperature field image data simultaneously, performing data fusion preprocessing, and finally performing spatio-temporal alignment based on the accurate timestamps.

4. An intelligent control system for downhole equipment status according to claim 1, characterized in that, The multi-modal state features include a vibration time-series feature group, a stress spatial feature group, and a temperature feature group. The vibration time-series feature group is extracted from the vibration time-domain sequence based on the temporal convolutional network model. The stress spatial feature group is extracted from the stress distribution matrix based on the graph convolutional network model. The temperature feature group is extracted from the temperature field image based on the improved GhostNet model.

5. The intelligent control system for downhole equipment status according to claim 4, characterized in that The process of extracting the temperature feature group from the temperature field image based on the improved GhostNet model is as follows: Perform feature region division on the temperature feature group, and dynamically adjust the convolutional kernel configuration to distinguish the low-frequency region and the high-frequency region. Perform convolution operations on the low-frequency region using a 5×5 size convolutional kernel to obtain a low-frequency temperature image. Perform convolution operations on the high-frequency region using a 3×3 size convolutional kernel to obtain a high-frequency temperature image. Stitch the low-frequency temperature image and the high-frequency temperature image together in the channel dimension to obtain a fused temperature field feature map. Arrange the obtained temperature field feature maps in the hierarchical order of the fusion process through multiple fusion processes to form a temperature feature group.

6. An intelligent control system for the state of downhole equipment according to claim 1, characterized in that The specific implementation process of the optimized reinforcement learning policy is as follows: The agent of the optimized reinforcement learning policy includes a state space, an action space, and a dynamic reward function. Define the multi-modal state features as the state space. Define the action space. The action space includes a device control instruction set, and the device control instruction set includes variable frequency speed regulation parameters, hydraulic pressure adjustment gears, and component start / stop combinations. Construct a dynamic reward function, which includes safety rewards, efficiency rewards, and energy consumption rewards, and assign corresponding weights to different rewards; When the device stress is lower than the safety threshold, the safety reward value approaches 1. When the ratio of the actual output of the device to the planned output approaches 1, the efficiency reward value approaches 1. When the ratio of the actual energy consumption of the device to the standard energy consumption approaches 1, the reward value approaches 1.

7. An intelligent control system for the state of downhole equipment according to claim 1, characterized in that, The process of combining the optimized reinforcement learning strategy with the improved deep deterministic policy gradient algorithm to obtain the optimal scheduling strategy is as follows: Construct 2 Critic networks with an MLP structure, set two hidden layers, with the ReLU function as the activation function for each layer, and the linear activation function for the output layer. The parameters of the two Critic networks are randomly initialized respectively to evaluate the value of the state to the action, and at the same time, initialize the target policy network; Through dynamic experience replay, interact with the environmental space according to the current policy network, observe the device state at each time step, regulate the device according to the action output by the policy network, obtain a reward after executing the action, and transfer to the next state to form an experience sample. Store the experience sample in an experience replay buffer with a capacity of . During training, randomly sample batch data from the buffer regularly, and optimize through the target network soft update mechanism to obtain the optimal scheduling strategy.

8. An intelligent control system for downhole equipment status according to claim 1, characterized in that, The implementation process of the target network soft update mechanism is as follows: For each sample sampled from the experience replay buffer, obtain the next action corresponding to the next state through the target policy network, then calculate the corresponding values through the two target Critic networks and take the minimum of the two, and use the Adam optimizer to update the policy network and the Critic network, and calculate the gradient through the loss function to update the parameters of the Critic network; Calculate the gradient approximation value of the policy network parameters to optimize the policy network parameters. At the same time, with the help of the Critic network evaluation, guide the policy network to adjust the parameters so that the target network parameters slowly follow the update of the current Critic network. Repeat the iterative training of the policy network and the Critic network. When the actions output by the policy network tend to be stable and the loss of the Critic network no longer decreases significantly, stop the training.

9. An intelligent control method for underground equipment status, using the control system described in any one of claims 1-8. The method steps are as follows: S1: Deploy a multi-modal sensor array integrating vibration, stress, and temperature, and collect multi-dimensional state data through a timestamp synchronization mechanism; S2: Construct a multi-modal fusion model, and extract multi-modal state features from the multi-dimensional state data based on the multi-modal fusion model; S3: Construct an optimized reinforcement learning strategy through the multi-modal state features, obtain the optimal scheduling strategy from the multi-modal state features based on the optimized reinforcement learning strategy combined with the improved deep deterministic policy gradient algorithm, and output the device operation control instruction through the optimal scheduling strategy.

Citation Information

Patent Citations

  • Autonomous controllable heterogeneous intelligent computing service platform and intelligent scene matching method

    CN115202868A

  • Deep groove ball bearing fault diagnosis method fusing frequency domain features and improved residual network

    CN118626822A