A dynamic crack propagation regulation method and device based on reinforcement learning

By employing a dynamic fracture propagation control method based on deep reinforcement learning, and utilizing three-dimensional physical simulation and real-time data fusion, the problems of uncontrollable fracture propagation paths and imbalance in multi-objective optimization in traditional fracturing technology are solved, achieving high-precision, safe, and economical fracturing operations.

CN121744805BActive Publication Date: 2026-05-19KARAMAY BAIJIANTAN DISTRICT (KARAMAY HIGH TECH ZONE) PETROLEUM ENG FIELD (PILOT) LAB +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
KARAMAY BAIJIANTAN DISTRICT (KARAMAY HIGH TECH ZONE) PETROLEUM ENG FIELD (PILOT) LAB
Filing Date
2026-02-26
Publication Date
2026-05-19

AI Technical Summary

Technical Problem

Traditional hydraulic fracturing technology struggles to precisely control fracture propagation paths and proppant distribution in reservoirs with strong heterogeneity and complex geostress fields, resulting in insufficient conductivity and a lack of real-time sensing and dynamic control, making it difficult to balance multi-objective optimization.

Method used

A dynamic fracture propagation control method based on deep reinforcement learning is adopted. A training dataset is constructed through a three-dimensional fracture physical simulation system. A hybrid architecture combining deep Q-network and near-end policy optimization algorithm is used to collect microseismic and distributed fiber optic sensing data in real time. The system dynamically outputs control commands for fracturing fluid discharge, proppant concentration and viscosity, thereby achieving closed-loop adaptive control of fracture propagation path and proppant distribution.

Benefits of technology

It significantly improves the accuracy and adaptability of crack propagation prediction and control, realizes real-time perception-decision-execution closed-loop control, reduces construction risks, optimizes diversion efficiency and energy consumption, forms a self-learning and self-evolving intelligent control capability, and improves construction safety and economy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121744805B_ABST
    Figure CN121744805B_ABST
Patent Text Reader

Abstract

The application discloses a kind of dynamic fracture propagation regulation and control method and device based on reinforcement learning, wherein, method includes: through three-dimensional fracture physical simulation system, orthogonal physical model experiment is carried out, training data set is constructed, and the pre-training of deep reinforcement learning strategy network that depth Q network is combined with near-end strategy optimization algorithm is carried out using it;During fracturing operation, real-time acquisition microseismic and distributed optical fiber sensing data, after being handled by edge computing node, input pre-training network, dynamically output the regulation and control instruction of fracturing fluid discharge capacity, proppant concentration and viscosity, and be issued to pump injection system through Internet of Things control link and be executed;According to the change of crack conductivity capacity fed back by optical fiber sensing, the reward function weight coefficient of strategy network is dynamically iterated and optimized, and closed-loop adaptive control is realized.The application realizes real-time, intelligent, adaptive control to fracture propagation path and proppant distribution, significantly improves the conductivity efficiency and construction stability of complex reservoir fracturing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of petroleum engineering technology, specifically to a method and apparatus for dynamic crack propagation control based on reinforcement learning. Background Technology

[0002] Hydraulic fracturing is a core production enhancement technique for developing unconventional oil and gas resources (such as shale gas and tight oil). Its core objective is to create a complex network of fractures with high conductivity in underground reservoirs, thereby providing efficient seepage channels for oil and gas. However, in reservoirs with strong heterogeneity and complex geostress fields, the propagation path, geometry, and final distribution of proppant are difficult to control precisely. This often results in a simple fracture network with insufficient conductivity after fracturing, directly affecting single-well productivity and ultimate recovery rate.

[0003] Traditional fracturing design and construction control have the following three main limitations:

[0004] First, in the design phase, traditional methods heavily rely on static models and empirical formulas. Current fracture propagation prediction is mainly based on numerical simulation methods such as the finite element method (FEM) and the displacement discontinuity method (DDM). While these methods can simulate fracture morphology under specific preset parameters (such as homogeneous rock mechanical parameters and simplified geostress fields), their model parameters are often difficult to obtain accurately and cannot be updated in real time. Faced with the natural fractures, lithological variations, and heterogeneous geostress commonly found in actual reservoirs, the prediction results of static models often deviate significantly from the actual fracture propagation paths, making it difficult to achieve the pre-designed "ideal fracture network." For example, in reservoirs with significant differences in horizontal stress or in areas with well-developed high-dipping natural fractures, fractures are prone to unexpected deflections or even premature termination, leading to construction failure.

[0005] Secondly, during construction, there is a lack of an effective closed-loop system for real-time sensing and dynamic control. Traditional fracturing operations mainly rely on pre-set pumping procedures. Although adjustments can be made empirically based on parameters such as surface pump pressure and displacement, it is impossible to perceive the dynamic propagation behavior of underground fractures and the migration and distribution of proppant in real time and with precision. In recent years, technologies such as microseismic monitoring and distributed fiber optic sensing have been introduced to monitor fracture growth, but this monitoring data is mostly used for "post-event interpretation" and has not yet formed an efficient real-time linkage closed loop with the pumping control system.

[0006] Finally, it is difficult to balance the conflicting objectives in the optimization process. The ideal fracturing effect requires simultaneously maximizing fracture conductivity, increasing fracture network complexity, and minimizing construction costs (energy consumption). Traditional methods often employ single-objective optimization or empirical combinations based on fixed weights, which cannot dynamically adjust the focus of the optimization objectives according to real-time construction results.

[0007] In summary, existing technologies have inherent limitations in areas such as fracture propagation control, safety risk early warning, and multi-objective optimization. Therefore, there is an urgent need for an innovative method, particularly suitable for unconventional oil and gas characterized by high heterogeneity and complex operating conditions, to drive the transformation of resource extraction from "extensive experience-driven" to "intelligent and precise control." Its core advantages lie in real-time control of complex fracture network morphology, improved prediction accuracy, reduced sand plugging risk, and adaptability to extreme conditions such as high temperature and high pressure. This will promote the transformation of traditional resource extraction from experience-driven to intelligent closed-loop control, providing innovative solutions for efficient energy development and safe utilization of underground space.

[0008] In view of this, the present invention is hereby proposed. Summary of the Invention

[0009] This invention aims to solve the problems of uncontrollable fracture propagation path and low flow efficiency caused by uneven proppant distribution in traditional fracturing processes, and proposes a dynamic fracture propagation control method and device based on deep reinforcement learning.

[0010] Specifically, the following technical solution was adopted:

[0011] A dynamic crack propagation control method based on reinforcement learning includes the following steps:

[0012] Orthogonal physical model experiments were conducted using a three-dimensional crack physical simulation system to simulate the crack propagation process under different geostress conditions, and data on crack trajectory, proppant distribution, and dynamic changes in conductivity were collected to construct a training dataset.

[0013] Based on the training dataset, a deep reinforcement learning policy network model is pre-trained. The deep reinforcement learning policy network model adopts a hybrid architecture combining a deep Q-network and a proximal policy optimization algorithm. The deep reinforcement learning policy network model includes a parallel Critic network and an Actor network. The Critic network adopts a dual-delay deep deterministic policy gradient structure to evaluate action value. The Actor network is based on the proximal policy optimization algorithm and is used to output the action probability distribution. The input of the Critic network is a concatenation of a multidimensional feature vector and a multidimensional action parameter vector. It contains at least one hidden layer with multiple neural nodes and uses a linear rectified function as the activation function. The final output is the state-action value function Q value. A long short-term memory module is introduced into the Actor network to model the temporal characteristics of the input microseismic monitoring data and distributed fiber optic sensing data.

[0014] During the fracturing operation, microseismic monitoring data and distributed fiber optic sensing data are collected in real time. After noise filtering and feature extraction by edge computing nodes deployed at the well site, the data are input into the pre-trained deep reinforcement learning strategy network model.

[0015] The deep reinforcement learning strategy network model dynamically outputs control commands for fracturing fluid discharge, proppant concentration, and fracturing fluid viscosity based on the input real-time data.

[0016] The control commands are sent to the pump control system via the Internet of Things control link to dynamically adjust the construction parameters;

[0017] Based on the real-time changes in crack flow-guiding capacity obtained from the distributed optical fiber sensing data, the reward function weight coefficients of the deep reinforcement learning strategy network model are dynamically and iteratively optimized to achieve closed-loop adaptive control of crack propagation path and proppant distribution.

[0018] As an optional embodiment of the present invention, in the dynamic crack propagation control method based on reinforcement learning, an orthogonal physical model experiment is conducted using a three-dimensional crack physical simulation system to simulate the crack propagation process under different geostress conditions. Data on crack trajectory, proppant distribution, and dynamic changes in conductivity are collected to construct a training dataset, including:

[0019] A physical model was constructed, including rock samples, hydraulic devices, fiber optic sensors, and microseismic monitors.

[0020] Through orthogonal physical model experiments under different pressure conditions, underground pressure was simulated using a hydraulic device, and crack propagation data and proppant distribution data were collected in real time using the fiber optic sensor and microseismic monitor.

[0021] The collected fracture propagation data and proppant distribution data are transformed into a multi-dimensional feature vector containing rock Young's modulus, formation permeability, porosity, Poisson's ratio, geostress difference, fracture aperture, fracturing fluid discharge rate, proppant concentration, fracturing fluid viscosity, pumping pressure, natural fracture density, proppant settling rate, microseismic event density, and fracture propagation rate. The fracturing fluid discharge rate, proppant concentration, and fracturing fluid viscosity form the action parameter space of the deep reinforcement learning strategy network model, and each action parameter is discretized into multiple preset levels.

[0022] Data normalization and feature dimension optimization are performed for pre-training of the deep reinforcement learning policy network model.

[0023] As an optional embodiment of the present invention, in the dynamic crack propagation control method based on reinforcement learning described in the present invention, the deep reinforcement learning policy network model adopts a hybrid architecture combining a deep Q-network and a proximal policy optimization algorithm, including:

[0024] The deep reinforcement learning policy network model includes a parallel Critic network and an Actor network. The Critic network adopts a dual-delay deep deterministic policy gradient structure to evaluate the value of actions, and the Actor network is based on a proximal policy optimization algorithm to output the probability distribution of actions.

[0025] The input to the Critic network is a concatenation of a multidimensional feature vector and a multidimensional action parameter vector. It contains at least one hidden layer with multiple neural nodes and uses a linear rectified function as the activation function. The final output is the state-action value function Q-value.

[0026] The Actor network incorporates a long short-term memory module to model the temporal characteristics of the input microseismic monitoring data and distributed fiber optic sensing data.

[0027] As an optional embodiment of the present invention, in the dynamic crack propagation control method based on reinforcement learning described in the present invention, the reward function of the deep reinforcement learning policy network... Constructed as a normalized weighted sum of fracture conductivity gain, fracture branch number, and pumping energy consumption, its expression is:

[0028] w1* + w2* – w3* ;

[0029] in, For real-time traffic redirection capability, The number of crack branches, For pumping energy consumption, , , These are the preset target maximum value or theoretical maximum value of the corresponding parameter, respectively. w1, w2, and w3 are the weight coefficients of each item, and w1 + w2 + w3 = 1.

[0030] As an optional embodiment of the present invention, in the dynamic crack propagation control method based on reinforcement learning described in the present invention, the step of pre-training a deep reinforcement learning policy network model based on the training dataset includes:

[0031] The deep reinforcement learning policy network is trained offline using a crack propagation simulation dataset generated by a numerical model that couples the displacement discontinuity method and the finite element method.

[0032] The pre-training step employs a priority experience replay mechanism and performs weighted sampling on key event samples representing sand blockage and crack turning in the crack propagation simulation dataset.

[0033] As an optional embodiment of the present invention, the construction and simulation process of a numerical model using a coupling of the displacement discontinuity method and the finite element method in the dynamic crack propagation control method based on reinforcement learning described in the present invention includes:

[0034] A displacement discontinuity method model was constructed to simulate the displacement discontinuity on the crack surface and to calculate the Type I and Type II stress intensity factors at the crack tip to predict the crack branch propagation direction.

[0035] A finite element method model is constructed, and the stress field of the continuous medium is solved through finite element mesh, taking into account the coupling effect of rock elastic-plastic deformation and pore pressure.

[0036] At the crack boundary, data interaction between the displacement discontinuity method model and the finite element method model is achieved through nodal force transmission, and the global stress balance equation is solved iteratively to simulate the crack propagation process.

[0037] As an optional embodiment of the present invention, in the dynamic crack propagation control method based on reinforcement learning, the reward function weight coefficients of the deep reinforcement learning policy network model are dynamically and iteratively optimized according to the real-time changes in crack conduction capacity obtained from the distributed optical fiber sensing data, including:

[0038] Based on the distributed fiber optic sensing data, calculate the measured value of the current crack conduction capacity. With target value The deviation δ;

[0039] When the absolute value of the deviation δ exceeds a preset first threshold, the optimization mechanism is triggered;

[0040] Based on the direction and magnitude of the deviation δ, the value of the weight coefficient w1 is adjusted so that the optimization objective of the deep reinforcement learning policy network model shifts towards reducing the deviation δ.

[0041] As an optional embodiment of the present invention, in the dynamic fracture propagation control method based on reinforcement learning, the closed-loop adaptive control includes a periodic policy network incremental learning step: during fracturing operations, on-site operation data is collected at preset time intervals and stored in an experience playback buffer, and the deep reinforcement learning policy network model is fine-tuned and updated at a setting lower than the pre-training learning rate.

[0042] This invention also provides a dynamic crack propagation control device based on reinforcement learning, comprising:

[0043] A three-dimensional crack physical simulation system is used to conduct orthogonal physical model experiments to simulate the crack propagation process under different geostress conditions, and to collect dynamic data on crack trajectory, proppant distribution and conductivity to construct a training dataset.

[0044] The model pre-training module is used to pre-train the deep reinforcement learning policy network model based on the training dataset. The deep reinforcement learning policy network model adopts a hybrid architecture combining a deep Q-network and a proximal policy optimization algorithm. The deep reinforcement learning policy network model includes a parallel Critic network and an Actor network. The Critic network adopts a dual-delay deep deterministic policy gradient structure to evaluate action value. The Actor network is based on the proximal policy optimization algorithm and is used to output the action probability distribution. The input of the Critic network is a concatenation of a multidimensional feature vector and a multidimensional action parameter vector. It contains at least one hidden layer with multiple neural nodes and uses a linear rectified function as the activation function. The final output is the state-action value function Q value. The Actor network introduces a long short-term memory module to model the temporal characteristics of the input microseismic monitoring data and distributed fiber optic sensing data.

[0045] The data acquisition and processing module includes edge computing nodes deployed at the well site, which are used to acquire microseismic monitoring data and distributed fiber optic sensing data in real time during fracturing operations, and to perform noise filtering and feature extraction on the acquired data.

[0046] The intelligent decision-making and control module is communicatively connected to the data acquisition and processing unit. It is loaded with the pre-trained deep reinforcement learning strategy network model, which is used to receive processed real-time data and dynamically output control commands for fracturing fluid discharge, proppant concentration and fracturing fluid viscosity.

[0047] The pumping execution module is connected to the intelligent decision-making and control module via an Internet of Things control link, and is used to receive the control instructions and dynamically adjust the construction parameters;

[0048] The closed-loop adaptive optimization module is communicatively connected to the data acquisition and processing module and the intelligent decision-making and control module. It is used to dynamically and iteratively optimize the reward function weight coefficients of the deep reinforcement learning strategy network model based on the real-time changes in the crack conduction capacity obtained from the distributed optical fiber sensing data, so as to achieve closed-loop adaptive regulation of crack propagation path and proppant distribution.

[0049] As an optional embodiment of the present invention, the model pre-training module is further configured to: generate simulation data containing high-risk scenarios using a numerical model coupled with the displacement discontinuity method and the finite element method, and perform offline pre-training of the deep reinforcement learning strategy network model using a priority experience playback mechanism.

[0050] This invention proposes a dynamic fracture propagation control method based on reinforcement learning. By integrating multi-physics data in real-time and employing an adaptive decision-making closed loop, it overcomes the challenges of traditional fracturing technologies, such as reliance on static models, response lag, and imbalance in multi-objective optimization. It offers the following advantages:

[0051] 1. Significantly improves the accuracy and adaptability of crack propagation prediction and control.

[0052] This invention proposes a dynamic crack propagation control method based on reinforcement learning. It constructs a training dataset by "conducting orthogonal physical model experiments using a three-dimensional crack physical simulation system" and performs pre-training and real-time decision-making based on a "deep reinforcement learning policy network model (a hybrid architecture combining deep Q-network and proximal policy optimization algorithm)".

[0053] High-fidelity data covering complex geostress conditions is generated through physical simulation experiments. Combined with the powerful nonlinear fitting and sequential decision-making capabilities of deep reinforcement learning, the policy network can learn and master the deep-seated laws governing fracture propagation and proppant migration under heterogeneous and complex stress fields. Compared to methods relying on fixed empirical formulas or static numerical models, this invention can dynamically adapt to actual geological conditions, achieving more accurate prediction and more effective active control of fracture propagation paths. Especially in areas with naturally developed fractures or reservoirs with high stress differentials, it can effectively guide fracture deflection, forming complex fracture networks.

[0054] 2. Achieve true real-time perception-decision-execution closed-loop control and improve response speed.

[0055] During fracturing operations, "real-time acquisition of microseismic monitoring data and distributed fiber optic sensing data" is processed in milliseconds by "edge computing nodes deployed at the well site" and input into a pre-trained model to quickly generate control commands, which are then issued and executed through an "Internet of Things control link".

[0056] This system integrates and processes multi-source data, including microseismic data (reflecting crack dynamics) and distributed fiber optic sensing data (reflecting conductivity and proppant distribution), at the millisecond level. Real-time localized decision-making is achieved through edge computing, fundamentally changing the traditional lag-based "monitor-interpretation-readjustment" model. The system can respond to and intervene in risks such as crack propagation speed, abnormal direction, and proppant settlement at the second or even millisecond level, significantly reducing construction risks such as sand blockage and uncontrolled cracks, and improving the safety and stability of the construction process.

[0057] 3. Overcome the challenge of multi-objective dynamic collaborative optimization to maximize traffic diversion efficiency.

[0058] During fracturing operations, multi-parameter coordinated commands such as displacement, concentration, and viscosity are dynamically output based on real-time data, and the reward function weight coefficients of the deep reinforcement learning strategy network model are dynamically iteratively optimized based on the changes in fracture conductivity obtained in real time from distributed optical fiber sensing data.

[0059] The deep reinforcement learning strategy network model of this invention naturally integrates multiple objectives such as conductivity, fracture complexity, and construction energy consumption in its reward function. More importantly, through online monitoring of conductivity feedback, the system can dynamically adjust the weights of each objective in the reward function, achieving online adaptive optimization. This allows the system to not only consider multiple competing objectives simultaneously, but also intelligently adjust the optimization focus based on actual results (such as whether the conductivity meets the standard) during construction, thereby continuously seeking optimization under complex constraints and ultimately maximizing the combined reservoir conductivity (FCD) and effective support volume (ESV), thus improving single-well productivity.

[0060] 4. Develop self-learning and self-evolving intelligent control capabilities, reducing reliance on prior experience.

[0061] A complete technology chain has been built, from "physical simulation pre-training" to "real-time data closed-loop control" and then to "dynamic iterative optimization of reward function".

[0062] The system can not only learn general strategies through simulated data in the initial stage, but also continuously fine-tune the strategy network and its optimization objective (reward weights) using real feedback data (fiber optic sensing and flow guidance capabilities) in actual operations. This constitutes a dual learning loop: the inner layer is the parameter optimization of the deep reinforcement learning strategy network model, and the outer layer is the adaptive adjustment of the reward function weights. This enables the entire method to have the ability to continuously learn and self-improve, accumulate construction experience in different blocks and well types, continuously optimize strategies, gradually reduce reliance on historical experience or expert intervention in specific areas, and promote the transformation of fracturing construction towards standardization and intelligence.

[0063] 5. Improve construction safety and economy.

[0064] Real-time dynamic control prevents excessive crack propagation or ineffective proppant placement; closed-loop optimization seeks a balance between energy consumption and effectiveness.

[0065] Real-time intervention effectively prevents fractures from penetrating and contaminating adjacent strata, controls fracture height, and improves construction safety. Simultaneously, precise control of proppant placement within the fractures reduces proppant waste. Energy consumption considerations in multi-objective optimization also encourage the system to prioritize more energy-efficient pumping schemes while meeting geological objectives. This ensures increased production while reducing material and energy costs, thereby improving the overall economic efficiency of fracturing operations.

[0066] In summary, this invention constructs a fracture propagation control system characterized by "real-time perception, intelligent decision-making, precise execution, and adaptive objectives" by deeply integrating physical simulation, deep reinforcement learning, multi-source real-time sensing, and edge computing. Its technical effectiveness lies in fundamentally addressing the systemic shortcomings of traditional methods in terms of prediction accuracy, response speed, multi-objective balancing, and learning capabilities, providing an innovative technical solution for the efficient, safe, and economical development of complex reservoirs. Attached Figure Description

[0067] Figure 1 This is a flowchart of a dynamic crack propagation control method based on reinforcement learning according to the present invention;

[0068] Figure 2 This is a flowchart illustrating a specific implementation example of a dynamic crack propagation control method based on reinforcement learning according to the present invention. Detailed Implementation

[0069] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0070] Therefore, the following detailed description of embodiments of the present invention is not intended to limit the scope of the claimed invention, but merely illustrates some embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention.

[0071] It should be noted that, unless otherwise specified, the embodiments and features and technical solutions in the embodiments of the present invention can be combined with each other.

[0072] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0073] In the description of this invention, it should be noted that the terms "upper," "lower," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use, or the orientation or positional relationship commonly understood by those skilled in the art. These terms are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention. In addition, the terms "first," "second," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0074] like Figure 1 As shown, this embodiment of the invention provides a dynamic crack propagation control method based on reinforcement learning, comprising the following steps:

[0075] Orthogonal physical model experiments were conducted using a three-dimensional crack physical simulation system to simulate the crack propagation process under different geostress conditions, and data on crack trajectory, proppant distribution, and dynamic changes in conductivity were collected to construct a training dataset.

[0076] Based on the training dataset, a deep reinforcement learning policy network model (DRL) is pre-trained. The deep reinforcement learning policy network model adopts a hybrid architecture that combines a deep Q network (DQN) with a proximal policy optimization algorithm (PPO).

[0077] During the fracturing operation, microseismic monitoring data and distributed fiber optic sensing data are collected in real time. After noise filtering and feature extraction by edge computing nodes deployed at the well site, the data are input into the pre-trained deep reinforcement learning strategy network model.

[0078] The deep reinforcement learning strategy network model dynamically outputs control commands for fracturing fluid discharge, proppant concentration, and fracturing fluid viscosity based on the input real-time data.

[0079] The control commands are sent to the pump control system via the Internet of Things control link to dynamically adjust the construction parameters;

[0080] Based on the real-time changes in crack flow-guiding capacity obtained from the distributed optical fiber sensing data, the reward function weight coefficients of the deep reinforcement learning strategy network model are dynamically and iteratively optimized to achieve closed-loop adaptive control of crack propagation path and proppant distribution.

[0081] This invention proposes a dynamic fracture propagation control method based on reinforcement learning. By integrating multi-physics data in real-time and employing an adaptive decision-making closed loop, it overcomes the challenges of traditional fracturing technologies, such as reliance on static models, response lag, and imbalance in multi-objective optimization. It offers the following advantages:

[0082] 1. Significantly improves the accuracy and adaptability of crack propagation prediction and control.

[0083] The present invention proposes a dynamic crack propagation control method based on reinforcement learning. It constructs a training dataset by "conducting orthogonal physical model experiments using a three-dimensional crack physical simulation system" and performs pre-training and real-time decision-making based on "a deep reinforcement learning policy network model (a hybrid architecture combining a deep Q-network and a proximal policy optimization algorithm)".

[0084] High-fidelity data covering complex geostress conditions is generated through physical simulation experiments. Combined with the powerful nonlinear fitting and sequential decision-making capabilities of deep reinforcement learning, the policy network can learn and master the deep-seated laws governing fracture propagation and proppant migration under heterogeneous and complex stress fields. Compared to methods relying on fixed empirical formulas or static numerical models, this invention can dynamically adapt to actual geological conditions, achieving more accurate prediction and more effective active control of fracture propagation paths. Especially in areas with naturally developed fractures or reservoirs with high stress differentials, it can effectively guide fracture deflection, forming complex fracture networks.

[0085] 2. Achieve true real-time perception-decision-execution closed-loop control and improve response speed.

[0086] During fracturing operations, "real-time acquisition of microseismic monitoring data and distributed fiber optic sensing data" is processed in milliseconds by "edge computing nodes deployed at the well site" and input into a pre-trained model to quickly generate control commands, which are then issued and executed through an "Internet of Things control link".

[0087] This system integrates and processes multi-source data, including microseismic data (reflecting crack dynamics) and distributed fiber optic sensing data (reflecting conductivity and proppant distribution), at the millisecond level. Real-time localized decision-making is achieved through edge computing, fundamentally changing the traditional lag-based "monitor-interpretation-readjustment" model. The system can respond to and intervene in risks such as crack propagation speed, abnormal direction, and proppant settlement at the second or even millisecond level, significantly reducing construction risks such as sand blockage and uncontrolled cracks, and improving the safety and stability of the construction process.

[0088] 3. Overcome the challenge of multi-objective dynamic collaborative optimization to maximize traffic diversion efficiency.

[0089] During fracturing operations, multi-parameter coordinated commands such as displacement, concentration, and viscosity are dynamically output based on real-time data, and the reward function weight coefficients of the deep reinforcement learning strategy network model are dynamically iteratively optimized based on the changes in fracture conductivity obtained in real time from distributed optical fiber sensing data.

[0090] The deep reinforcement learning strategy network model of this invention naturally integrates multiple objectives such as conductivity, fracture complexity, and construction energy consumption in its reward function. More importantly, through online monitoring of conductivity feedback, the system can dynamically adjust the weights of each objective in the reward function, achieving online adaptive optimization. This allows the system to not only consider multiple competing objectives simultaneously, but also intelligently adjust the optimization focus based on actual results (such as whether the conductivity meets the standard) during construction, thereby continuously seeking optimization under complex constraints and ultimately maximizing the combined reservoir conductivity (FCD) and effective support volume (ESV), thus improving single-well productivity.

[0091] 4. Develop self-learning and self-evolving intelligent control capabilities, reducing reliance on prior experience.

[0092] A complete technology chain has been built, from "physical simulation pre-training" to "real-time data closed-loop control" and then to "dynamic iterative optimization of reward function".

[0093] The system can not only learn general strategies through simulated data in the initial stage, but also continuously fine-tune the strategy network and its optimization objective (reward weights) using real feedback data (fiber optic sensing and flow guidance capabilities) in actual operations. This constitutes a dual learning loop: the inner layer is the parameter optimization of the deep reinforcement learning strategy network model, and the outer layer is the adaptive adjustment of the reward function weights. This enables the entire method to have the ability to continuously learn and self-improve, accumulate construction experience in different blocks and well types, continuously optimize strategies, gradually reduce reliance on historical experience or expert intervention in specific areas, and promote the transformation of fracturing construction towards standardization and intelligence.

[0094] 5. Improve construction safety and economy.

[0095] Real-time dynamic control prevents excessive crack propagation or ineffective proppant placement; closed-loop optimization seeks a balance between energy consumption and effectiveness.

[0096] Real-time intervention effectively prevents fractures from penetrating and contaminating adjacent strata, controls fracture height, and improves construction safety. Simultaneously, precise control of proppant placement within the fractures reduces proppant waste. Energy consumption considerations in multi-objective optimization also encourage the system to prioritize more energy-efficient pumping schemes while meeting geological objectives. This ensures increased production while reducing material and energy costs, thereby improving the overall economic efficiency of fracturing operations.

[0097] In summary, this invention, through the deep integration of physical simulation, deep reinforcement learning, multi-source real-time sensing, and edge computing, constructs a fracture propagation control system characterized by "real-time perception, intelligent decision-making, precise execution, and adaptive objectives." Its technical effectiveness lies in fundamentally addressing the systemic shortcomings of traditional methods in terms of prediction accuracy, response speed, multi-objective balancing, and learning capabilities, providing an innovative technical solution for the efficient, safe, and economical development of complex reservoirs.

[0098] As an optional embodiment of the present invention, in a dynamic crack propagation control method based on reinforcement learning, an orthogonal physical model experiment is conducted using a three-dimensional crack physical simulation system to simulate the crack propagation process under different geostress conditions. Data on crack trajectory, proppant distribution, and dynamic changes in conductivity are collected to construct a training dataset, including:

[0099] A physical model was constructed, including rock samples, hydraulic devices, fiber optic sensors, and microseismic monitors.

[0100] Through orthogonal physical model experiments under different pressure conditions, underground pressure was simulated using a hydraulic device, and crack propagation data and proppant distribution data were collected in real time using the fiber optic sensor and microseismic monitor.

[0101] The collected fracture propagation data and proppant distribution data are transformed into a multi-dimensional feature vector (e.g., a 14-dimensional feature vector) containing rock Young's modulus, formation permeability, porosity, Poisson's ratio, geostress difference, fracture aperture, fracturing fluid discharge rate, proppant concentration, fracturing fluid viscosity, pumping pressure, natural fracture density, proppant settling rate, microseismic event density, and fracture propagation rate. The fracturing fluid discharge rate, proppant concentration, and fracturing fluid viscosity form the action parameter space of the deep reinforcement learning strategy network model, and each action parameter is discretized into multiple preset levels (e.g., divided into 5 levels of action space discretization).

[0102] Data normalization and feature dimension optimization are performed for pre-training of the deep reinforcement learning policy network model.

[0103] Optionally, the normalization process described in this embodiment adopts the Z-score standardization method; the feature dimension optimization process is achieved through Pearson correlation analysis to remove redundant features.

[0104] As an optional embodiment of the present invention, in a dynamic crack propagation control method based on reinforcement learning, the deep reinforcement learning policy network model adopts a hybrid architecture combining a deep Q-network and a proximal policy optimization algorithm, including:

[0105] The deep reinforcement learning policy network model includes a parallel Critic network and an Actor network. The Critic network adopts a dual-delay deep deterministic policy gradient structure to evaluate the value of actions, and the Actor network is based on a proximal policy optimization algorithm to output the probability distribution of actions.

[0106] The input to the Critic network is a concatenation of a multidimensional feature vector and a multidimensional action parameter vector. It contains at least one hidden layer with multiple neural nodes and uses a linear rectified function as the activation function. The final output is the state-action value function Q-value.

[0107] The Actor network incorporates a long short-term memory module to model the temporal characteristics of the input microseismic monitoring data and distributed fiber optic sensing data.

[0108] The reward function of the deep reinforcement learning policy network model described in this embodiment Constructed as a normalized weighted sum of fracture conductivity gain, fracture branch number, and pumping energy consumption, its expression is:

[0109] w1* + w2* – w3* ;

[0110] in, For real-time traffic redirection capability, The number of crack branches, For pumping energy consumption, , , These are the preset target maximum value or theoretical maximum value of the corresponding parameter, respectively. w1, w2, and w3 are the weight coefficients of each item, and w1 + w2 + w3 = 1.

[0111] Optionally, the weight coefficients of the reward function are initially set as follows: weight w1 = 0.6 for the flow capacity gain, weight w2 = 0.3 for the number of crack branches, and weight w3 = 0.1 for the pumping energy consumption.

[0112] The weighting coefficients w1, w2, and w3 described in this embodiment can be dynamically adjusted based on the changes in the crack's flow-guiding capacity obtained in real time from the distributed optical fiber sensing data.

[0113] As an optional embodiment of the present invention, in a dynamic crack propagation control method based on reinforcement learning, the step of pre-training a deep reinforcement learning policy network model based on the training dataset includes:

[0114] The deep reinforcement learning policy network model was trained offline using a crack propagation simulation dataset generated by a numerical model coupled with the displacement discontinuity method (DDM) and the finite element method (FEM).

[0115] The pre-training step employs a priority experience replay mechanism and performs weighted sampling on key event samples representing sand blockage and crack turning in the crack propagation simulation dataset.

[0116] Specifically, in the reinforcement learning-based dynamic crack propagation control method of this embodiment, the construction and simulation process of the numerical model using the coupling of the displacement discontinuity method and the finite element method includes:

[0117] A displacement discontinuity method model was constructed to simulate the displacement discontinuity on the crack surface and to calculate the Type I and Type II stress intensity factors at the crack tip to predict the crack branch propagation direction.

[0118] A finite element method model is constructed, and the stress field of the continuous medium is solved through finite element mesh, taking into account the coupling effect of rock elastic-plastic deformation and pore pressure.

[0119] At the crack boundary, data interaction between the displacement discontinuity method model and the finite element method model is achieved through nodal force transmission, and the global stress balance equation is solved iteratively to simulate the crack propagation process.

[0120] The finite element method model described in this embodiment uses more than 1 million finite element meshes.

[0121] The coupled numerical model described in this embodiment has a crack propagation path prediction error of less than 3% under complex natural crack network conditions.

[0122] The crack propagation simulation dataset described in this embodiment contains specific high-risk scenario samples generated by the coupled numerical model to characterize sand blockage events and crack turning events.

[0123] As an optional embodiment of the present invention, a dynamic fracture propagation control method based on reinforcement learning is provided. During hydraulic fracturing, microseismic monitoring data and distributed fiber optic sensing data are collected in real time. After noise filtering and feature extraction by edge computing nodes deployed at the well site, the data are input into the pre-trained deep reinforcement learning policy network model, including:

[0124] The edge computing node performs noise filtering on the microseismic monitoring data, adopts a noise reduction method based on wavelet transform, and optimizes the selection of wavelet basis functions according to the time-frequency characteristics of the fracturing signal.

[0125] Specifically, the edge computing node receives and processes the microseismic monitoring data in real time at a period of no less than 12 milliseconds.

[0126] The edge computing node described in this embodiment integrates a software development kit with a dedicated chip to build a high-speed data path for transmitting the microseismic monitoring data and distributed fiber optic sensing data to the deep reinforcement learning policy network.

[0127] The feature extraction steps described in this embodiment include: processing and fusing the microseismic monitoring data and distributed optical fiber sensing data into a dynamic state vector containing information on microseismic event density, crack propagation rate, proppant settlement rate, and real-time diversion capacity.

[0128] As an optional embodiment of the present invention, in a dynamic fracture propagation control method based on reinforcement learning, the deep reinforcement learning strategy network model dynamically outputs control commands for fracturing fluid discharge, proppant concentration, and fracturing fluid viscosity based on the input real-time data, including:

[0129] In the control commands output by the deep reinforcement learning strategy network model, each of the fracturing fluid discharge rate, proppant concentration, and fracturing fluid viscosity is selected from a preset set of discrete actions, wherein each action parameter is pre-divided into at least 5 levels.

[0130] After outputting the control command, it is determined whether the fracturing fluid discharge rate in the command exceeds the safe operating limit of the fracturing pump set. If so, the output value is corrected to the safe operating limit value.

[0131] Optionally, the accuracy of the fracturing fluid discharge control command is within ±0.5 m³ / min.

[0132] An embodiment of the present invention provides a dynamic fracture propagation control method based on reinforcement learning, which further includes a safety verification step: after outputting the control command, it is determined whether the fracturing fluid discharge rate in the command exceeds the safe operating limit of the fracturing pump group; if so, the output value is corrected to the safe operating limit value.

[0133] Furthermore, the deep reinforcement learning strategy network model integrates a time-series prediction module; when the time-series prediction module predicts that the wellhead pressure fluctuation rate exceeds a preset threshold based on the input time-series data, the control command output by the strategy network model includes a reduction in fracturing fluid discharge by a preset amount.

[0134] Furthermore, the control commands also include commands for starting / stopping the pulse injection mode and adjusting the frequency.

[0135] As an optional embodiment of the present invention, in a dynamic crack propagation control method based on reinforcement learning, the control command is sent to the pump injection control system via an Internet of Things control link to dynamically adjust the construction parameters, including:

[0136] The IoT control link simultaneously sends control commands to both the fracturing pump truck control system and the proppant truck control system to achieve synchronized adjustment of fracturing fluid discharge and proppant concentration.

[0137] The pumping control system includes a fracturing pump truck with a rated power of not less than 2250 horsepower. The adjustment range of the fracturing fluid discharge rate in the control command is between 16 m³ / min and 20 m³ / min, and the control accuracy is ±0.5 m³ / min.

[0138] The dynamic adjustment steps include: the pumping control system generates corresponding pumping pressure and flow rate setpoints according to the received control instructions, and drives the actuator to operate, with the entire instruction-execution closed-loop delay time controlled within 2 minutes.

[0139] Before the control command is issued, the method further includes: comparing and verifying the parameter value in the command with a preset safe operating range of the device, and only allowing the command to be issued and executed when the parameter value is within the safe operating range.

[0140] The pump control system uses a manifold system consisting of multiple high-pressure pipelines for fluid transport to reduce pipeline friction loss.

[0141] As an optional embodiment of the present invention, in a dynamic crack propagation control method based on reinforcement learning, the reward function weight coefficients of the deep reinforcement learning policy network model are dynamically and iteratively optimized according to the real-time changes in crack conduction capacity obtained from the distributed optical fiber sensing data, including:

[0142] Based on the distributed fiber optic sensing data, calculate the measured value of the current crack conduction capacity. With target value The deviation δ;

[0143] When the absolute value of the deviation δ exceeds a preset first threshold, the optimization mechanism is triggered;

[0144] Based on the direction and magnitude of the deviation δ, the value of the weight coefficient w1 is adjusted so that the optimization objective of the deep reinforcement learning policy network model shifts towards reducing the deviation δ.

[0145] Optionally, the first threshold is 15%.

[0146] As an optional embodiment of the present invention, in a dynamic fracture propagation control method based on reinforcement learning, the closed-loop adaptive control includes a periodic policy network incremental learning step: during fracturing operations, on-site operation data is collected at preset time intervals and stored in an experience playback buffer, and the deep reinforcement learning policy network model is fine-tuned and updated at a setting lower than the pre-training learning rate.

[0147] Optionally, the preset time interval is 8 to 12 hours of fracturing operations; the learning rate used for the fine-tuning update is on the order of 1×10⁻⁶. -5 .

[0148] like Figure 2 As shown, a specific example of a dynamic crack propagation control method based on reinforcement learning according to the present invention includes the following steps:

[0149] Step 1: Build a 3D rock model, create rock samples that resemble real strata, use a hydraulic device to simulate underground pressure, and install fiber optic sensors and microseismic monitors to capture crack changes in real time.

[0150] Step 2: Through orthogonal model experiments under different pressure conditions, record data such as crack shape and proppant distribution, and convert them into a computer-recognizable format (such as a 14-dimensional feature vector).

[0151] Step 3: Select 14-dimensional feature vectors of rock Young's modulus, formation permeability, porosity, Poisson's ratio, geostress difference, fracture aperture, fracturing fluid discharge rate, proppant concentration, fracturing fluid viscosity, pumping pressure, microseismic event density, and fracture propagation rate, and discretize the fracturing fluid discharge rate, proppant concentration, and proppant-carrying fluid viscosity into 5 levels of action space.

[0152] Step 4: Perform data normalization and optimize feature dimensions.

[0153] Step 5: Design the deep reinforcement learning policy network model framework, design the Critic network and Actor network, model the reward function, and set a comprehensive scoring rule of "the more complex the crack, the higher the flow capacity, and the lower the energy consumption".

[0154] Step 6: Offline pre-training strategy. Pre-training is performed using data generated by displacement discontinuity method (DDM) and finite element method (FEM). The Adam optimizer (learning rate 3e-4) is used, and the experience replay mechanism is prioritized to perform 5-fold weighted sampling of key events such as sand blockage and crack turning.

[0155] Step 7: Install a small computer near the fracturing truck, deploy edge nodes, optimize wavelet basis functions based on fracturing signal characteristics, integrate domestic chip manufacturer SDK to build a high-speed data path, receive microseismic data every 12ms, filter noise, and automatically adjust reward weights based on changes in flow guidance capacity fed back by fiber optic sensors (e.g., increase flow guidance capacity weight when the actual flow guidance effect deviation exceeds 15%), and dynamically optimize the strategy.

[0156] Step 8: After every 10 hours of fracturing operations, extract 2000 sets of new data from the playback buffer for incremental training, perform 2 rounds of fine-tuning training (learning rate reduced to 1e-5), and add a safety constraint layer;

[0157] The advantages of this invention, implemented through the above specific examples, are as follows:

[0158] This example proposes a dynamic fracture propagation control method based on deep reinforcement learning. By integrating multi-physics data in real time and using an adaptive decision-making closed loop, it overcomes the challenges of traditional fracturing technology, such as reliance on static models, response lag, and imbalance in multi-objective optimization.

[0159] A dual-Critic network is employed to evaluate action value in parallel, suppressing Q-value overestimation bias by minimizing the value. A dynamic state space (14-dimensional feature vector) is constructed by combining microseismic time-series data with fiber optic sensing feedback, including key parameters such as rock stress field differences, proppant settling rate, and fracture propagation rate, enabling real-time fracture path perception. New data is extracted every 10 minutes of fracturing operation cycle, incrementally updating the reward model parameters (learning rate decays to 1e-5). Multi-well historical data is aggregated to optimize the global weighting strategy, reducing conductivity prediction error and energy consumption in high-temperature reservoirs.

[0160] This invention also provides a dynamic crack propagation control device based on reinforcement learning, comprising:

[0161] A three-dimensional crack physical simulation system is used to conduct orthogonal physical model experiments to simulate the crack propagation process under different geostress conditions, and to collect dynamic data on crack trajectory, proppant distribution and conductivity to construct a training dataset.

[0162] The model pre-training module is used to pre-train the deep reinforcement learning policy network model based on the training dataset. The deep reinforcement learning policy network model adopts a hybrid architecture that combines a deep Q-network with a proximal policy optimization algorithm.

[0163] The data acquisition and processing module includes edge computing nodes deployed at the well site, which are used to acquire microseismic monitoring data and distributed fiber optic sensing data in real time during fracturing operations, and to perform noise filtering and feature extraction on the acquired data.

[0164] The intelligent decision-making and control module is communicatively connected to the data acquisition and processing unit. It is loaded with the pre-trained deep reinforcement learning strategy network model, which is used to receive processed real-time data and dynamically output control commands for fracturing fluid discharge, proppant concentration and fracturing fluid viscosity.

[0165] The pumping execution module is connected to the intelligent decision-making and control module via an Internet of Things control link, and is used to receive the control instructions and dynamically adjust the construction parameters;

[0166] The closed-loop adaptive optimization module is communicatively connected to the data acquisition and processing module and the intelligent decision-making and control module. It is used to dynamically and iteratively optimize the reward function weight coefficients of the deep reinforcement learning strategy network model based on the real-time changes in the crack conduction capacity obtained from the distributed optical fiber sensing data, so as to achieve closed-loop adaptive regulation of crack propagation path and proppant distribution.

[0167] Furthermore, the model pre-training module in this embodiment is also configured to: generate simulation data containing high-risk scenarios using a numerical model coupled with the displacement discontinuity method and the finite element method, and perform offline pre-training of the deep reinforcement learning strategy network model using a priority experience playback mechanism. Example 1

[0168] This embodiment provides a dynamic crack propagation control method based on reinforcement learning, including the following steps:

[0169] Step 1: Geological modeling and multi-scale fracture prediction. By combining pre-stack and post-stack seismic data, well logging interpretation, and natural fracture distribution, a refined reservoir model is constructed.

[0170] Step 2: Calibrate rock mechanics parameters. Obtain Young's modulus (14,000-42,000 MPa), Poisson's ratio (0.2-0.27), and geostress field (horizontal stress difference 5-30 MPa) through core experiments, which will be used as input parameters for the DRL model.

[0171] Step 3: Pre-train the deep reinforcement learning policy network model to generate a simulated dataset (including crack morphology, proppant distribution, and conductivity). Pre-train a model combining a hybrid deep Q-network and a proximal policy optimization algorithm. The state space includes 14-dimensional features such as microseismic event density and lithological brittleness index, while the action space includes parameters such as displacement (±2 m³ / min) and proppant concentration (±3%).

[0172] Step 4: Initial setting of reward weights: flow capacity 60%, crack complexity 30%, and energy consumption 10%. Fiber optic sensors provide real-time feedback on flow distribution deviations, triggering dynamic weight adjustments.

[0173] Step 5: Deployment of the continuous injection system, employing a linkage scheme between a 2250Hp fracturing pump truck (displacement 16-20 m³ / min) and a sand mixing truck to achieve displacement accuracy control of ±0.5 m³ / min. The high-pressure manifold system consists of four 4-inch inner diameter pipelines to reduce friction loss.

[0174] Step 6: Coordinated monitoring of microseismic and fiber optic systems. A microseismic array is deployed to monitor crack propagation direction, combined with distributed optical fiber sensing (DAS / DTS) to capture changes in proppant settling rate and conductivity. Data is filtered by edge nodes (NVIDIA Jetson AGX) and then input into the DRL model.

[0175] Step 7: The Long Short-Term Memory (LSTM) module predicts the pressure fluctuation trend and triggers automatic displacement reduction (displacement is reduced by 30% when ΔP>5 MPa / min) and boost pulse injection mode (optimizing proppant placement).

[0176] Step 8: Post-pressurization performance verification showed that the EUR (estimated final recovery rate) increased by 28%, the fracture fractal dimension increased from 1.3 to 1.7, and the effective proppant placement rate reached 86%. Example 2

[0177] This embodiment provides a dynamic crack propagation control method based on reinforcement learning, including the following steps:

[0178] Step 1: Multi-scale crack modeling, constructing a displacement discontinuity method (DDM) model to simulate the displacement discontinuity on the crack surface, calculating the stress intensity factor (KI, KII) at the crack tip, and predicting the crack branch propagation direction.

[0179] Step 2: Construct a finite element method (FEM) model and solve the stress field of the continuous medium using a finite element mesh (mesh count > 1 million), considering the coupling effect of rock elastoplastic deformation and pore pressure. Data exchange between the DDM and FEM is achieved at the fracture boundary through nodal force transfer, and the global stress balance equation is solved iteratively, with a fracture propagation accuracy error of <3%.

[0180] Step 3: Model validation and calibration, based on reservoir core experimental data (validating the DDM-FEM model's ability to predict fracture propagation paths in complex natural fracture networks (density 1.2 fractures / m²), compared with microseismic monitoring data).

[0181] Step 4: Training and optimization of the deep reinforcement learning policy network model. State input (14-dimensional vector): geostress field difference (σH_max - σh_min), stress intensity factor at the crack tip (KI, KII), and pore pressure gradient.

[0182] Natural fracture density and dip angle, real-time microseismic event density (10-second window), action output (4-dimensional continuous space), fracturing flow rate (±1.5 m³ / min), proppant concentration (±2%), fracturing fluid viscosity (±10 cP), pulse injection frequency (0.1-1 Hz), etc.

[0183] Step 5: Design the reward function and calculate the conductivity based on the crack width and proppant embedding depth output by DDM-FEM.

[0184] Step 6: Generate crack propagation scenarios (including high-risk events such as sand blockage and crack turning) using DDM-FEM, pre-train the TD3+PPO hybrid DRL model with an initial learning rate of 3e-4 and a batch size of 256.

[0185] Step 7: Employ a priority experience replay mechanism to assign a 5x sampling weight to sand blockage event samples, thereby accelerating the learning efficiency of the policy network for key scenarios.

[0186] Step 8: Real-time dynamic control and closed-loop verification. Deploy an NVIDIA A100 GPU edge server and integrate the DDM-FEM lightweight proxy model (computation time < 2 seconds / step) to predict fracture propagation paths in real time. The DRL policy network receives microseismic data (processed with wavelet denoising) every 12ms, generates control commands, and sends them to the fracturing pump truck.

[0187] The above embodiments are only used to illustrate the present invention and are not intended to limit the technical solutions described herein. Although the present invention has been described in detail with reference to the above embodiments, the present invention is not limited to the specific embodiments described above. Therefore, any modifications or equivalent substitutions to the present invention, as well as all technical solutions and improvements that do not depart from the spirit and scope of the invention, are covered within the scope of the claims of the present invention.

Claims

1. A dynamic crack propagation control method based on reinforcement learning, characterized in that, Includes the following steps: Orthogonal physical model experiments were conducted using a three-dimensional crack physical simulation system to simulate the crack propagation process under different geostress conditions, and data on crack trajectory, proppant distribution, and dynamic changes in conductivity were collected to construct a training dataset. Based on the training dataset, a deep reinforcement learning policy network model is pre-trained. The deep reinforcement learning policy network model adopts a hybrid architecture combining a deep Q-network and a proximal policy optimization algorithm. The deep reinforcement learning policy network model includes a parallel Critic network and an Actor network. The Critic network adopts a dual-delay deep deterministic policy gradient structure to evaluate action value. The Actor network is based on the proximal policy optimization algorithm and is used to output the action probability distribution. The input of the Critic network is a concatenation of a multidimensional feature vector and a multidimensional action parameter vector. It contains at least one hidden layer with multiple neural nodes and uses a linear rectified function as the activation function. The final output is the state-action value function Q value. A long short-term memory module is introduced into the Actor network to model the temporal characteristics of the input microseismic monitoring data and distributed fiber optic sensing data. During the fracturing operation, microseismic monitoring data and distributed fiber optic sensing data are collected in real time. After noise filtering and feature extraction by edge computing nodes deployed at the well site, the data are input into the pre-trained deep reinforcement learning strategy network model. The deep reinforcement learning strategy network model dynamically outputs control commands for fracturing fluid discharge, proppant concentration, and fracturing fluid viscosity based on the input real-time data. The control commands are sent to the pump control system via the Internet of Things control link to dynamically adjust the construction parameters; Based on the real-time changes in crack flow-guiding capacity obtained from the distributed optical fiber sensing data, the reward function weight coefficients of the deep reinforcement learning strategy network model are dynamically and iteratively optimized to achieve closed-loop adaptive control of crack propagation path and proppant distribution.

2. The dynamic crack propagation control method based on reinforcement learning according to claim 1, characterized in that, Orthogonal physical model experiments were conducted using a three-dimensional fracture physical simulation system to simulate the fracture propagation process under different geostress conditions. Data on fracture trajectory, proppant distribution, and dynamic changes in conductivity were collected to construct a training dataset, including: A physical model was constructed, including rock samples, hydraulic devices, fiber optic sensors, and microseismic monitors. Through orthogonal physical model experiments under different pressure conditions, underground pressure was simulated using a hydraulic device, and crack propagation data and proppant distribution data were collected in real time using the fiber optic sensor and microseismic monitor. The collected fracture propagation data and proppant distribution data are transformed into a multi-dimensional feature vector containing rock Young's modulus, formation permeability, porosity, Poisson's ratio, geostress difference, fracture aperture, fracturing fluid discharge rate, proppant concentration, fracturing fluid viscosity, pumping pressure, natural fracture density, proppant settling rate, microseismic event density, and fracture propagation rate. The fracturing fluid discharge rate, proppant concentration, and fracturing fluid viscosity form the action parameter space of the deep reinforcement learning strategy network model, and each action parameter is discretized into multiple preset levels. Data normalization and feature dimension optimization are performed for pre-training of the deep reinforcement learning policy network model.

3. The dynamic crack propagation control method based on reinforcement learning according to claim 2, characterized in that, The reward function of the deep reinforcement learning policy network model Constructed as a normalized weighted sum of fracture conductivity gain, fracture branch number, and pumping energy consumption, its expression is: w1* + w2* – w3* ; in, For real-time traffic redirection capability, The number of crack branches, For pumping energy consumption, , , These are the preset target maximum value or theoretical maximum value of the corresponding parameter, respectively. w1, w2, and w3 are the weight coefficients of each item, and w1 + w2 + w3 = 1.

4. The dynamic crack propagation control method based on reinforcement learning according to claim 3, characterized in that, The pre-training of the deep reinforcement learning policy network model based on the training dataset includes: The deep reinforcement learning strategy network model is trained offline using a crack propagation simulation dataset generated by a numerical model that couples the displacement discontinuity method and the finite element method. The pre-training step employs a priority experience replay mechanism and performs weighted sampling on key event samples representing sand blockage and crack turning in the crack propagation simulation dataset.

5. The dynamic crack propagation control method based on reinforcement learning according to claim 4, characterized in that, The process of constructing and simulating a numerical model using the coupling of the displacement discontinuity method and the finite element method includes: A displacement discontinuity method model was constructed to simulate the displacement discontinuity on the crack surface and to calculate the Type I and Type II stress intensity factors at the crack tip to predict the crack branch propagation direction. A finite element method model is constructed, and the stress field of the continuous medium is solved through finite element mesh, taking into account the coupling effect of rock elastic-plastic deformation and pore pressure. At the crack boundary, data interaction between the displacement discontinuity method model and the finite element method model is achieved through nodal force transmission, and the global stress balance equation is solved iteratively to simulate the crack propagation process.

6. The dynamic crack propagation control method based on reinforcement learning according to claim 3, characterized in that, Based on the real-time changes in crack conduction capacity obtained from the distributed optical fiber sensing data, the reward function weight coefficients of the deep reinforcement learning policy network model are dynamically and iteratively optimized, including: Based on the distributed fiber optic sensing data, calculate the measured value of the current crack conduction capacity. With target value The deviation δ; When the absolute value of the deviation δ exceeds a preset first threshold, the optimization mechanism is triggered; Based on the direction and magnitude of the deviation δ, the value of the weight coefficient w1 is adjusted so that the optimization objective of the deep reinforcement learning policy network model shifts towards reducing the deviation δ.

7. The dynamic crack propagation control method based on reinforcement learning according to claim 6, characterized in that, The closed-loop adaptive control includes a periodic incremental learning step for the strategy network: during fracturing operations, on-site operation data is collected at preset time intervals and stored in an experience playback buffer, and the deep reinforcement learning strategy network model is fine-tuned and updated with a learning rate lower than that of the pre-training model.

8. A dynamic crack propagation control device based on reinforcement learning, characterized in that, include: A three-dimensional crack physical simulation system is used to conduct orthogonal physical model experiments to simulate the crack propagation process under different geostress conditions, and to collect dynamic data on crack trajectory, proppant distribution and conductivity to construct a training dataset. The model pre-training module is used to pre-train the deep reinforcement learning policy network model based on the training dataset. The deep reinforcement learning policy network model adopts a hybrid architecture combining a deep Q-network and a proximal policy optimization algorithm. The deep reinforcement learning policy network model includes a parallel Critic network and an Actor network. The Critic network adopts a dual-delay deep deterministic policy gradient structure to evaluate action value. The Actor network is based on the proximal policy optimization algorithm and is used to output the action probability distribution. The input of the Critic network is a concatenation of a multidimensional feature vector and a multidimensional action parameter vector. It contains at least one hidden layer with multiple neural nodes and uses a linear rectified function as the activation function. The final output is the state-action value function Q value. The Actor network introduces a long short-term memory module to model the temporal characteristics of the input microseismic monitoring data and distributed fiber optic sensing data. The data acquisition and processing module includes edge computing nodes deployed at the well site, which are used to acquire microseismic monitoring data and distributed fiber optic sensing data in real time during fracturing operations, and to perform noise filtering and feature extraction on the acquired data. The intelligent decision-making and control module is communicatively connected to the data acquisition and processing unit. It is loaded with the pre-trained deep reinforcement learning strategy network model, which is used to receive processed real-time data and dynamically output control commands for fracturing fluid discharge, proppant concentration and fracturing fluid viscosity. The pumping execution module is connected to the intelligent decision-making and control module via an Internet of Things control link, and is used to receive the control instructions and dynamically adjust the construction parameters; The closed-loop adaptive optimization module is communicatively connected to the data acquisition and processing module and the intelligent decision-making and control module. It is used to dynamically and iteratively optimize the reward function weight coefficients of the deep reinforcement learning strategy network model based on the real-time changes in the crack conduction capacity obtained from the distributed optical fiber sensing data, so as to achieve closed-loop adaptive regulation of crack propagation path and proppant distribution.

9. A dynamic crack propagation control device based on reinforcement learning according to claim 8, characterized in that, The model pre-training module is also configured to: generate simulation data containing high-risk scenarios using a numerical model coupled with the displacement discontinuity method and the finite element method, and perform offline pre-training of the deep reinforcement learning strategy network model using a priority experience playback mechanism.