Pulse jet rock breaking perforation mechanism and technological parameter optimization method
By employing reinforcement learning algorithms, particularly deep Q-network algorithms, in high-pressure pulse jet rock breaking perforation, real-time optimization and intelligent control of process parameters were achieved. This solved the problems of weak theoretical foundation, inconsistent parameter selection, and poor adaptability to geological conditions in existing technologies, thereby improving perforation efficiency and production capacity.
Patent Information
- Application Number
- CN202410658313.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-27
- Publication Date
- 2025-11-28
AI Technical Summary
The existing high-pressure pulse jet rock breaking and perforation technology lacks a systematic theoretical foundation, lacks unified standards for process parameter selection, has incomplete data acquisition and analysis, poor adaptability to geological conditions, high cost, and cannot achieve real-time optimization and intelligent control of process parameters.
By employing reinforcement learning algorithms, particularly deep Q-network algorithms, and through interactive learning between real-time monitoring data of the perforation process and the environmental model, and utilizing sensors and real-time monitoring equipment, intelligent and real-time optimization of process parameters is achieved.
This technology enables real-time optimization and intelligent control of process parameters in high-pressure pulse jet rock-breaking perforation technology, improving perforation effect and production capacity. It also solves the problems of weak theoretical foundation, inconsistent parameter selection, incomplete data acquisition, and poor adaptability to geological conditions in existing technologies, thereby reducing costs.
Smart Images

Figure SMS_87 
Figure SMS_138
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of rock mechanics, and particularly relates to a method for optimizing the mechanism and process parameters of pulse jet rock breaking and perforation. BACKGROUND
[0002] Some shortcomings exist in the existing high-pressure pulse jet rock breaking and perforation mechanism and process parameter optimization technology. Lack of systematic theoretical basis: The mechanism of high-pressure pulse jet rock breaking and perforation is still not deep enough, and there is still a lack of theoretical support for the interaction mechanism between jet flow and rock.
[0003] Lack of unified standard for process parameter selection: At present, there is a lack of unified standard for the selection of process parameters for high-pressure pulse jet rock breaking and perforation, and there is a large difference in the parameter range used in different researches and applications, which makes it difficult to compare and verify different research results.
[0004] Inadequate data collection and analysis: In experiments and practical applications, there are still some problems in the data collection and analysis of high-pressure pulse jet rock breaking and perforation, including quantitative evaluation of perforation effect, measurement of energy distribution, etc.
[0005] Poor adaptability to geological conditions: The application range of high-pressure pulse jet rock breaking and perforation technology is relatively narrow, and it is not strong in adaptability to different geological conditions and reservoir characteristics, which needs to be adjusted and optimized according to specific conditions.
[0006] High cost: Compared with traditional perforation technology, the equipment and operation cost of high-pressure pulse jet rock breaking and perforation technology is higher, which leads to certain difficulties in popularization and popularization in practical application.
[0007] The Chinese patent document with publication number CN116879068A discloses an experimental method for breaking rock by shock wave in simulated formation environment, which comprises the following steps: S1, pretreating the rock sample to form a cubic structure; S2, drilling a reserved hole on the rock sample and installing a wellbore in the reserved hole; S3, placing the rock sample into a triaxial pulse shock wave induced rock cracking experimental system, loading stress in horizontal X-axis direction, horizontal Y-axis direction and vertical direction on the rock sample, and heating the rock sample; S4, calibrating crack propagation by crack propagation calibration method, and placing a shock wave generator into the wellbore; S5, cracking the rock sample by shock wave operation of the shock wave generator; S6, repeating step S5 until the crack breaks through the boundary of the rock sample and no new crack is generated; S7, recording the cracking radius and crack propagation morphology, drilling a sampling hole on the rock sample and taking a sample; S8, comparing and analyzing the sampling results to obtain the changes of the pore and permeability parameters of the rock sample and the changes of the mechanical parameters of the rock sample. The document cannot realize real-time optimization and intelligent control of process parameters in high-pressure pulse jet rock breaking and perforation technology. Summary of the Invention
[0008] The purpose of this invention is to provide a pulse jet rock-breaking perforation mechanism and process parameter optimization method to overcome the problem that the existing technology cannot achieve real-time optimization and intelligent control of process parameters in high-pressure pulse jet rock-breaking perforation technology.
[0009] Therefore, this invention provides a pulse jet rock-breaking perforation mechanism and process parameter optimization method, comprising the following steps: 1) Collect real-time perforation process data and preprocess the collected real-time perforation process data; 2) Build an environmental model using historical data and preprocessed data; 3) The process parameters in the perforation process are optimized using a reinforcement learning algorithm, including the following steps: real-time monitoring of perforation process data, interactive learning between the monitored perforation process data and the environment model, and autonomous learning and adjustment of perforation process parameters based on the feedback results of the real-time monitored perforation process data and the environment model. The above steps are repeated until the perforation requirements are met, thus completing the optimization of the process parameters in the perforation process.
[0010] Preferably, the reinforcement learning algorithm is a deep Q-network algorithm.
[0011] Preferably, the perforation process data is monitored in real time using a deep Q-network algorithm. The monitored perforation process data is then interactively learned from the environmental model. The reinforcement learning algorithm autonomously learns and adjusts the perforation process parameters based on the feedback results from the real-time monitored perforation process data and the environmental model, including the following steps: 3.1) The process parameters during the perforation process are considered as states; 3.2) Treat the process parameters that need to be optimized as actions; 3.3) Based on the optimization objective, define a reward function to measure the performance in each state; 3.4) Using neural networks to approximate A function, where the input is the state and the output is the state corresponding to each action. value; 3.5) Initialize the parameters of the neural network and prepare the memory playback buffer; 3.6) At each time step, based on the current state, use a neural network to calculate the corresponding action for each action. The algorithm selects an action based on a certain strategy, executes the action, observes the environmental feedback to obtain a reward and the next state, and stores the state, action, reward, and next state transitions in a memory replay buffer. A batch of transition samples is randomly selected from the memory replay buffer to train the neural network. Using the trained neural network, the algorithm selects the state with the maximum value based on the current state. values of the actions.
[0012] Preferably, the neural network is a multi-layer perceptron or a convolutional neural network.
[0013] Preferably, the batch of transition samples is randomly selected from the memory replay buffer for training the neural network by updating the network parameters by minimizing the mean squared error between the predicted values and the target values.
[0014] Preferably, the neural network is trained according to the update formula of the Q-learning algorithm to update the values of the current state-action pairs.
[0015] Preferably, the batch of transition samples is randomly selected from the memory replay buffer for training the neural network by updating the network parameters by minimizing the mean squared error between the predicted values and the target values.
[0016] Preferably, the network parameters are updated by minimizing the mean squared error between the predicted values and the target values. Preferably, a utility loss function is used to minimize the difference between the predicted values and the target values.
[0017] Preferably, the certain policy in step 3.6) comprises a greedy policy.
[0018] Preferably, the neural network is trained according to the update formula of the Q-learning algorithm to update the values of the current state-action pairs to where is the learning rate, is the discount factor, represents the action with the maximum value selected in the next state . is the current state, is the current action, ( s, a ) is the value of the current state-action pair.
[0019] Advantages of the present application: 1. The pulse jet rock breaking perforating mechanism and process parameter optimization method provided by the application, through the following steps: 1) collecting real-time perforating process data, and pre-processing the collected real-time perforating process data; 2) establishing an environment model by using historical data and the pre-processed data; 3) optimizing the process parameters in the perforating process by using a reinforcement learning algorithm, including the following steps: real-time monitoring of perforating process data, interactive learning of the monitored perforating process data and the environment model, autonomous learning and adjustment of the perforating process parameters by the reinforcement learning algorithm according to the feedback results of the real-time monitored perforating process data and the environment model; repeating the above steps until the perforating requirements are met, and completing the optimization of the process parameters in the perforating process. Real-time optimization and intelligent control of the process parameters in the high-pressure pulse jet rock breaking perforating technology are realized, and the perforating effect and productivity are improved.
[0020] 2. The pulse jet rock breaking perforating mechanism and process parameter optimization method provided by the application, by using the deep Q network algorithm, the continuous parameter problem can be better handled, the state space explosion can be avoided, the training sample scarcity and correlation problems can be solved, and the complex relationship of nonlinear function approximation can be handled; the best pressure pulse jet rock breaking perforating value can be more accurately found, so as to optimize the process parameters and improve the perforating effect.
[0021] 3. The pulse jet rock breaking perforating mechanism and process parameter optimization method provided by the application, the data in the perforating process can be collected in real time by using sensors and real-time monitoring equipment. DETAILED DESCRIPTION
[0022] The principles and characteristics of the application are described below, and the examples are only used to explain the application, and are not used to limit the scope of the application.
[0023] Example 1 A pulse jet rock breaking perforating mechanism and process parameter optimization method, including the following steps: 1) collecting real-time perforating process data, and pre-processing the collected real-time perforating process data; Specifically, the perforating process data includes jet flow rate, jet energy density, jet frequency, and characteristic data of the formation rock; the collected real-time perforating process data is pre-processed, including data cleaning, feature extraction, etc., so as to be used for establishing an environment model.
[0024] In operation, existing data acquisition and processing devices are used to collect and pre-process the real-time perforating process data.
[0025] Preferably, sensors and real-time monitoring equipment are used to collect the data in the perforating process in real time.
[0026] Preferably, all sensors and real-time monitoring devices include flow velocity sensors, flow meters, pressure sensors, vibration sensors, and acceleration sensors. The flow velocity of the jet is collected and monitored in real time by the flow velocity sensor or flow meter, the energy density of the jet is monitored in real time by the pressure sensor, and the vibration frequency of the perforation device is monitored in real time by the vibration sensor or acceleration sensor.
[0027] Specifically, jet velocity acquisition involves using a velocity sensor or a flow meter to monitor the jet velocity in real time. The velocity sensor is installed at the nozzle of the perforating device or in the perforation pipe, and obtains the jet velocity by measuring the speed of the liquid or gas. Jet energy density acquisition: The energy density of the jet is monitored in real time using a pressure sensor. The pressure sensor is installed at the nozzle of the perforation equipment or in the perforation pipe, and the energy density of the jet is calculated by measuring the pressure applied to the jet. Jet frequency acquisition: The vibration frequency of the perforation equipment is monitored in real time using vibration sensors or accelerometers. Vibration sensors are installed on the moving parts of the perforation equipment, and the jet frequency is obtained by detecting the frequency of the vibration signal. By using sensors and real-time monitoring equipment to collect process parameters during the perforation process, the real-time acquisition of these process parameters can be provided to intelligent algorithms for analysis and processing, thereby achieving optimization and control of the process parameters.
[0028] 2) Build an environmental model using historical data and preprocessed data; Specifically, environmental models can be built using different methods depending on the actual situation. A common approach is to use machine learning techniques, such as neural networks, to model the physical environment of the perforation process. This can be achieved by training with historical or simulated data. During training, the neural network can learn the relationship between input data and output results, thereby predicting the effects during the perforation process. Historical data includes historical environmental data, historical perforation process parameters, etc.; historical environmental data includes characteristics of the underlying rock, etc.
[0029] Preferably, an environment model is built using methods based on supervised learning or reinforcement learning.
[0030] Specifically, in the pulsed jet rock-breaking perforation process, the environmental model includes the physical properties of the formation rock, the parameters of the perforation equipment, and the dynamic changes during the perforation process. This information will help intelligent algorithms better understand the complexity of the perforation process.
[0031] The physical environment refers to the actual environment during the perforation process, including the characteristics of the formation rocks, the parameters and conditions of the perforation equipment, etc. When establishing an environmental model, it is necessary to consider the interaction between the perforation equipment and the formation rocks, as well as the influence of perforation parameters on the rock-breaking effect.
[0032] 3) The process parameters in the perforation process are optimized using a reinforcement learning algorithm, including the following steps: real-time monitoring of the perforation process data, interactive learning between the monitored perforation process data and the environment model, and autonomous learning and adjustment of the perforation process parameters based on the feedback results of the real-time monitored perforation process data and the environment model; repeat the above steps until the perforation requirements are met, thus completing the optimization of the process parameters in the perforation process. Specifically, the data acquisition and processing device transmits the acquired data to the central control system; the central control system is responsible for monitoring and controlling the perforation equipment. Reinforcement learning algorithms and environmental models are deployed in the central control system to process data in the perforation process in real time and optimize and adjust the perforation process parameters based on the current status and environmental feedback.
[0033] This invention enables real-time optimization and intelligent control of process parameters in high-pressure pulse jet rock-breaking perforation technology, thereby improving perforation efficiency and production capacity.
[0034] Example 2: Based on Example 1, high-pressure pulse jet technology was used for rock-breaking perforation in the fracturing stimulation of oil wells in an oilfield. Flow velocity sensors, pressure sensors, and vibration sensors were installed in the perforation equipment to monitor the process parameters during perforation in real time.
[0035] Jet velocity acquisition: By installing a velocity sensor at the nozzle of the perforation device, the velocity of the jet is monitored in real time. The velocity sensor obtains the velocity of the jet by measuring the speed of the liquid and transmits the data to the data acquisition and processing system. Jet energy density acquisition: A pressure sensor is installed at the nozzle of the perforation device to monitor the pressure applied by the jet in real time. The pressure sensor calculates the energy density of the jet by measuring the pressure of the jet and transmits the data to the data acquisition and processing system. Jet frequency acquisition: By installing vibration sensors on the moving parts of the perforation equipment, the vibration frequency of the perforation equipment is monitored in real time. The vibration sensors obtain the frequency of the jet by detecting the frequency of the vibration signal and transmit the data to the data acquisition and processing system. The data acquisition and processing system transmits the process parameters collected in real time during the perforation process to a cloud server or central control system for real-time data processing and analysis. By applying reinforcement learning algorithms (artificial intelligence algorithms and big data analysis technology), the system processes and analyzes the real-time acquired data to optimize process parameters, such as adjusting jet velocity, jet energy density, and jet frequency, to achieve the best rock-breaking perforation effect and improve perforation efficiency and production capacity.
[0036] Example 3: Based on Example 2, a method for optimizing the mechanism and process parameters of pulse jet rock breaking perforation is described below: 1) Collect real-time perforation process data and preprocess the collected real-time perforation process data; Specifically, real-time perforation process data is collected, including jet velocity, jet energy density, jet frequency, and formation rock characteristics. This data is then preprocessed, including data cleaning and feature extraction, to facilitate the creation of an environmental model.
[0037] 2) Build an environmental model using historical data and preprocessed data; Specifically, supervised learning or reinforcement learning methods are used to build an environment model. Taking supervised learning as an example, the model is trained using historical data, enabling it to predict perforation results based on input perforation process parameters and formation characteristics. In reinforcement learning, the environment model predicts the next state and reward based on the combination of the current state and actions.
[0038] The environment model is used to describe the physical environment and corresponding effects during the perforation process.
[0039] 3) The process parameters in the perforation process are optimized using a reinforcement learning algorithm, including the following steps: real-time monitoring of the perforation process data, interactive learning between the monitored perforation process data and the environment model, and autonomous learning and adjustment of the perforation process parameters based on the feedback results of the real-time monitored perforation process data and the environment model; repeat the above steps until the perforation requirements are met, thus completing the optimization of the process parameters in the perforation process.
[0040] The reinforcement learning algorithm is Q-learning, which is used to optimize the process parameters in the perforation process, specifically: In this algorithm, the perforating device is regarded as an intelligent agent that continuously learns and optimizes its behavioral strategies through interaction with the environment; Specifically, continuously learning and optimizing one's behavioral strategies includes strategy iteration and updates, real-time optimization and control; Policy iteration and update: In reinforcement learning algorithms, the agent selects a behavioral policy to execute based on the current state and environmental feedback. Based on the execution result, the agent receives a reward or penalty signal, and then updates its policy. Through multiple iterations and updates, the agent can gradually learn the optimal behavioral policy. In Q-learning, the agent selects an action based on the current state and updates its Q-value based on environmental feedback. For example, the agent selects an action in the current state, performs a piercing operation, receives a reward or penalty, and then updates its Q-value to adjust its next action selection.
[0041] Real-time optimization and control: The trained agent is deployed on a cloud server or central control system to monitor data during the perforation process in real time. Based on the current state and environmental feedback, the trained environmental model is used to predict and optimize process parameters. The reinforcement learning algorithm can autonomously adjust process parameters. Through real-time optimization and control, efficient high-pressure pulse jet rock-breaking perforation can be achieved.
[0042] In this example, reinforcement learning algorithms can autonomously learn and optimize process parameters based on real-time data and environmental feedback to improve the effectiveness and production efficiency of pulsed jet rock-breaking perforation. Through continuous iteration and optimization, the intelligent algorithm can gradually improve the performance of the perforation process and achieve more efficient perforation operations.
[0043] Furthermore, the Q-learning algorithm from reinforcement learning is used to optimize the process parameters in the high-pressure pulsed jet rock-breaking perforation process. Below is a simplified example, illustrated with formulas: Define State: The process parameters during the perforation process are considered as states. For example, jet velocity, jet energy density, and jet frequency can be represented as states, denoted as... .
[0044] Define Action: The process parameters that need optimization are defined as actions. For example, the adjustment values for jet velocity, jet energy density, and jet frequency can be represented as actions, denoted as... These adjustments can be small increments or decrements.
[0045] Define Reward: Based on the optimization objective, define a reward function to measure performance in each state. For example, the reward function can be defined as the improvement in rock-breaking effect. The reward function can be defined according to the actual situation, for example: If the rock-breaking effect is significantly improved, a positive reward will be given; If the rock-breaking effect does not improve or decreases, a negative reward will be given.
[0046] definition Q-value: Each state-action pair has a corresponding Q-value. The value represents the expected reward for taking a certain action in a given state. The values can be continuously updated through iterative learning. During initialization, all state-action pairs can be... Set the value to 0.
[0047] Iterative updates of reinforcement learning algorithms: At each time step, based on the current state Choose an action based on a certain strategy Commonly used strategies include -greedy strategy, that is, with The probability of choosing a random action, in order to The probability is used to select the action with the highest Q value at the moment.
[0048] Execute action Rewards are given for observing environmental feedback. and the next state .
[0049] Update the current state-action pair according to the update formula of the Q-learning algorithm. value: ,in, It's the learning rate. It is a discount factor. Indicates the next state The next selection has the largest Value action .
[0050] Through multiple iterations and updates, The value will gradually converge to the optimal value, that is, the best action is selected in each state. Value. Ultimately, by using the trained... The value table allows you to select the value with the highest value based on the current state. The value of the action, namely the optimized process parameters, is to achieve efficient high-pressure pulse jet rock breaking and perforation.
[0051] Examples of high-pressure pulse jet rock-breaking perforation mechanism and process parameters used as parameters for Q-learning are as follows: When liquid is instantaneously accelerated through a nozzle and sprayed onto a rock surface, a high-pressure pulse jet is generated, applying high pressure to the rock surface and damaging its structure. Process parameters, including jet velocity, jet energy density, and jet frequency, have a significant impact on the mechanism and effectiveness of the high-pressure pulse jet.
[0052] In this example, we use jet velocity and jet energy density as process parameters and optimize them using the Q-learning algorithm. The specific steps are as follows: Define the state: The jet velocity and jet energy density are considered as states. For example, the jet velocity can be denoted as... The jet energy density is denoted as .
[0053] Define Action: Designate the process parameters that need optimization as actions. For example, denote the adjustment value of the jet velocity as... The adjustment value for the jet energy density is denoted as .
[0054] Define the reward: Based on the optimization objective, define a reward function to measure performance at each state. For example, the reward function could be defined as the increase in the degree of rock fragmentation. The reward function can be defined according to the specific circumstances.
[0055] definition Q-value: Each state-action pair has a corresponding Q-value. The value represents the expected reward for taking an action in a given state. During initialization, all state-action pairs can be initialized. Set the value to 0.
[0056] Iterative updates of reinforcement learning algorithms: At each time step, based on the current state Choose an action based on a certain strategy .
[0057] Execute action Rewards are given for observing environmental feedback. and the next state .
[0058] Update the current state-action pair according to the update formula of the Q-learning algorithm. value: ,in, It's the learning rate. It is a discount factor. Indicates the next state The next selection has the largest Value action .
[0059] Through multiple iterations and updates, The value will gradually converge to the optimal value, that is, the best action is selected in each state. Value. Ultimately, by using the trained... The value table allows you to select the value with the highest value based on the current state. The value of the action, i.e., the optimization of process parameters.
[0060] For example, suppose the initial state is You can choose a suitable strategy, such as -greedy strategy, with a certain probability Choose a random action, to The probability of choosing the current option with the highest probability The action to determine the value. Based on the current state, assume the action to be chosen is... The reward observed after execution is To obtain the next state .
[0061] According to the update formula of Q-learning, it can be updated value:
[0062] Assuming learning rate Discount factor The updated calculation is performed according to the formula to obtain the updated result. value.
[0063] Through multiple iterations and updates, The value will gradually converge to the optimal value, that is, the best action is selected in each state. Value. Ultimately, by using the trained... The value table allows you to select the value with the highest value based on the current state. The value of the action, i.e., the optimization of process parameters.
[0064] Example 4: Based on Example 3, the reinforcement learning algorithm is the deep Q-network algorithm.
[0065] Specifically, the Deep Q-Network (DQN) algorithm can help find the optimal parameters for Q-learning when the rock-breaking perforation mechanism and process parameters of pressure pulse jet are used as learning parameters. value.
[0066] DQN is an algorithm that combines deep learning and reinforcement learning. It approximates... Functions can handle problems with high-dimensional state spaces. In pressure pulse jet rock-breaking perforation, both states and actions can be continuous parameters, not just discrete values, making traditional Q-learning algorithms difficult to handle in the examples above. Therefore, the DQN algorithm can better handle this case of continuous parameters.
[0067] Preferably, the perforation process data is monitored in real time, and the monitored perforation process data is interactively learned with the environmental model. The reinforcement learning algorithm learns and adjusts the perforation process parameters autonomously based on the feedback results of the real-time monitored perforation process data and the environmental model, including the following steps: 3.1) The process parameters during the perforation process are considered as states; Specifically, define the state: the jet velocity and jet energy density are considered as states. For example, the jet velocity is denoted as... The jet energy density is denoted as .
[0068] 3.2) Treat the process parameters that need to be optimized as actions; Specifically, define the action: designate the process parameters that need optimization as the action. For example, denote the adjustment value of the jet velocity as... The adjustment value for the jet energy density is denoted as .
[0069] 3.3) Based on the optimization objective, define a reward function to measure the performance in each state; Specifically, define the reward: Based on the optimization objective, define a reward function to measure performance in each state. For example, the reward function can be defined as the increase in the degree of rock fragmentation. The reward function can be defined according to the actual situation.
[0070] 3.4) Using neural networks to approximate A function, where the input is the state and the output is the state corresponding to each action. value; Specifically, the Deep Q-Network (DQN) algorithm is defined as follows: DQN uses neural networks to approximate... A function, where the input is the state and the output is the state corresponding to each action. The value. The neural network can be a multilayer perceptron (MLP) or a convolutional neural network (CNN), and the specific structure is designed according to the actual situation.
[0071] 3.5) Initialize the parameters of the neural network and prepare the memory playback buffer; Preferably, the step of randomly selecting a batch of transformed samples from the memory playback buffer for training the neural network, by minimizing the predicted... Values and Objectives The mean squared error between values is used to update the network parameters.
[0072] Specifically, initialize the DQN network and memory replay: initialize the parameters of the DQN network and prepare a memory replay buffer to store the previous state, action, reward and the transition to the next state.
[0073] 3.6) At each time step, based on the current state, use a neural network to calculate the corresponding action for each action. The algorithm selects an action based on a certain strategy, executes the action, observes the environmental feedback to obtain a reward and the next state, and stores the state, action, reward, and next state transitions in a memory replay buffer. A batch of transition samples is randomly selected from the memory replay buffer to train the neural network. Using the trained neural network, the algorithm selects the state with the maximum value based on the current state. The value is determined by the actions taken to obtain optimized process parameters.
[0074] Preferably, a certain strategy in step 3.6) includes -greedy strategy.
[0075] Specifically, reinforcement learning algorithms are iteratively updated: At each time step, based on the current state Use a DQN network to predict the corresponding action for each action. Value. That is, calculation. .
[0076] According to certain strategies (such as) -greedy strategy), select an action With a certain probability Choose a random action, to The probability of choosing the current option with the highest probability The action that determines the value.
[0077] Execute action Rewards are given for observing environmental feedback. and the next state .
[0078] Store the state, action, reward, and transition to the next state in the memory replay buffer.
[0079] A batch of transformed samples is randomly selected from the memory replay buffer to train the DQN network. This is achieved by minimizing the predicted... Values and Objectives The mean squared error between values is used to update the network parameters.
[0080] Preferably, the method for training the neural network updates the current state-action pair according to the update formula of the Q-learning algorithm. The values are used to train the neural network.
[0081] Specifically, based on the update formula of the Q-learning algorithm, the current state-action pair is updated. value: ,in, It's the learning rate. It is a discount factor. Indicates the next state The next selection has the largest Value action .
[0082] Through multiple iterations and updates, the parameters of the DQN network will gradually converge to the optimal value, that is, the optimal action is selected in each state. Ultimately, by using a trained DQN network, the value with the largest [value] can be selected based on the current state. The value of the action, i.e., the optimization of process parameters.
[0083] Preferably, the method of minimizing the predicted Values and Objectives When updating network parameters using the mean squared error between values, a loss function is used to minimize the predicted values. Values and Objectives The difference between the values.
[0084] Specifically, the loss function is:
[0085] in, This represents the parameters of the DQN network. This represents the parameters of the target network.
[0086] The parameters of the DQN network can be updated using the backpropagation algorithm, making the predictions more accurate. The value gradually approaches the target. Value. Furthermore, the use of memory replay and the target network can improve training stability and performance.
[0087] The DQN of this invention can handle high-dimensional state spaces: the states and parameters involved in pressure pulse jet rock-breaking perforation can be numerous, such as jet velocity, jet energy density, and jet frequency. Traditional Q-learning algorithms require discretization of continuous parameters, resulting in a very large state space. DQN, however, can directly handle continuous parameters, avoiding this problem.
[0088] DQN incorporates memory replay and a target network: In pressure pulse jet rock-breaking perforation, the scarcity of training samples and the correlation between samples may be challenges. DQN addresses these issues by using memory replay and a target network. Memory replay stores and reuses previous experience, reducing sample correlation. The target network reduces the correlation between samples. The fluctuations in value improve the stability of learning.
[0089] DQN can handle nonlinear function approximation: the mechanisms and effects of pressure pulse jet rock-breaking perforation often involve complex nonlinear functional relationships. Using traditional Q-learning algorithms, if linear function approximation is used... The given values may not be able to fit this nonlinear relationship well. DQN, on the other hand, uses neural networks as function approximators, which can handle more complex nonlinear relationships.
[0090] By using the DQN algorithm, problems involving continuous parameters can be better handled, avoiding state space explosion. It also addresses the issues of training sample scarcity and correlation, as well as the complex relationships involved in approximating nonlinear functions. This allows for a more accurate determination of the optimal perforation point for pressure pulse jet rock breaking. This allows for the optimization of process parameters and improvement of perforation performance.
[0091] The above examples are merely illustrative of the present invention and do not constitute a limitation on the scope of protection of the present invention. All designs that are the same as or similar to the present invention are within the scope of protection of the present invention.
Claims
1. A method for optimizing the mechanism and process parameters of pulse jet rock-breaking perforation, characterized in that: Includes the following steps: 1) Collect real-time perforation process data and preprocess the collected real-time perforation process data; 2) Build an environmental model using historical data and preprocessed data; 3) The reinforcement learning algorithm is used to optimize the process parameters in the perforation process, including the following steps: real-time monitoring of perforation process data, interactive learning between the monitored perforation process data and the environment model, and the reinforcement learning algorithm learns and adjusts the perforation process parameters autonomously based on the feedback results of the real-time monitored perforation process data and the environment model. Repeat the above steps until the perforation requirements are met, thus completing the optimization of the process parameters during the perforation process.
2. The pulse jet rock-breaking perforation mechanism and process parameter optimization method as described in claim 1, characterized in that: The reinforcement learning algorithm is the deep Q-network algorithm.
3. The pulse jet rock-breaking perforation mechanism and process parameter optimization method as described in claim 2, characterized in that: When the reinforcement learning algorithm is a deep Q-network algorithm, the perforation process data is monitored in real time, and the monitored perforation process data is interactively learned with the environment model. Based on the feedback results of the real-time monitored perforation process data and the environment model, the reinforcement learning algorithm autonomously learns and adjusts the perforation process parameters, including the following steps: 3.1) The process parameters during the perforation process are considered as states; 3.2) The process parameters to be optimized are used as actions; 3.3) Based on the optimization objective, define a reward function to measure the performance in each state; 3.4) Using neural networks to approximate A function, where the input is the state and the output is the state corresponding to each action. value; 3.5) Initialize the parameters of the neural network and prepare the memory playback buffer; 3.6) At each time step, based on the current state, use a neural network to calculate the corresponding action for each action. The algorithm selects an action based on a certain strategy, executes the action, observes the environmental feedback to obtain a reward and the next state, and stores the state, action, reward, and next state transitions in a memory replay buffer. A batch of transition samples is randomly selected from the memory replay buffer to train the neural network. Using the trained neural network, the algorithm selects the state with the maximum value based on the current state. The value is determined by the actions taken to obtain optimized process parameters.
4. The pulse jet rock-breaking perforation mechanism and process parameter optimization method as described in claim 3, characterized in that: The neural network is a multilayer perceptron or a convolutional neural network.
5. The pulse jet rock-breaking perforation mechanism and process parameter optimization method as described in claim 3, characterized in that: The process involves randomly selecting a batch of transformed samples from the memory playback buffer to train the neural network, minimizing the predicted... Values and Objectives The mean squared error between values is used to update the network parameters.
6. The pulse jet rock-breaking perforation mechanism and process parameter optimization method as described in claim 3, characterized in that: The method used to train the neural network updates the current state-action pair according to the update formula of the Q-learning algorithm. The values are used to train the neural network.
7. The pulse jet rock-breaking perforation mechanism and process parameter optimization method as described in claim 3, characterized in that: When randomly selecting a batch of transformed samples from the memory playback buffer for training the neural network, the prediction is minimized. Values and Objectives The mean squared error between values is used to update the network parameters.
8. The pulse jet rock-breaking perforation mechanism and process parameter optimization method as described in claim 7, characterized in that: The method of minimizing the prediction Values and Objectives When updating network parameters using the mean squared error between values, a loss function is used to minimize the predicted values. Values and Objectives The difference between the values.
9. The pulse jet rock-breaking perforation mechanism and process parameter optimization method as described in claim 3, characterized in that: The strategies in step 3.6) include -greedy strategy.
10. The pulse jet rock-breaking perforation mechanism and process parameter optimization method as described in claim 6, characterized in that: The method for training the neural network involves updating the current state-action correspondence using the update formula of the Q-learning algorithm. Value ,in, It's the learning rate. It is a discount factor. Indicates the next state The next selection has the largest Value action ; This is the current state. This is the current action. ( s, a () is the current state-action correspondence value.
Citation Information
Patent Citations
Shock wave rock breaking experiment method for simulating stratum environment
CN116879068A