Comprehensive dessert-based deep coal rock gas fracturing construction scheme optimization method
By constructing a comprehensive sweet spot factor and a deep reinforcement learning model based on cluster-level data, the deep coal and rock gas fracturing construction scheme was optimized, which solved the problem of insufficient scientificity and reliability of the construction scheme design in traditional methods, and achieved a significant improvement in reservoir stimulation effect and gas production efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SOUTHWEST PETROLEUM UNIV
- Filing Date
- 2026-06-08
- Publication Date
- 2026-07-03
AI Technical Summary
Existing technologies are insufficient to achieve balanced and effective fracturing operations in deep coal and gas reservoirs. Traditional evaluation methods lack dynamic adjustment capabilities and cannot reflect the impact of changes in construction measures and reservoir responses, resulting in insufficient scientific rigor and reliability in construction scheme design. Furthermore, existing machine learning models lack global optimal decision-making capabilities and dynamic update capabilities in complex environments.
Based on cluster-level data, a dynamically updatable comprehensive sweet spot factor is constructed. By combining the LightGBM model, SHAP interpretation, entropy weight method and deep reinforcement learning, the fracturing operation scheme is optimized. Intelligent decision-making is carried out through multilayer perceptron and deep Q network to generate efficient fracturing operation parameter combinations.
It has significantly enhanced the overall potential of reservoirs, provided a more scientific, intelligent, and reliable fracturing construction design, and improved gas production efficiency and development economy.
Smart Images

Figure CN122332847A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of unconventional oil and gas development technology, specifically to an optimization method for deep coal and gas fracturing construction schemes based on a comprehensive sweet spot. Background Technology
[0002] Deep coal-rock gas reservoirs typically exhibit strong heterogeneity, complex pore structures, and significant differences in geostress distribution. This makes it difficult for traditional fracturing operations to achieve balanced and effective reservoir stimulation. Fracture propagation paths are highly uncertain, leading to substantial variations in reservoir stimulation effects, thus impacting gas production efficiency and overall development economics. Identification and evaluation of sweet spots are crucial for optimal fracturing design; however, current technologies often rely on single geological parameters or empirical judgments, failing to comprehensively reflect the reservoir's overall development potential. Especially under complex reservoir conditions, the accuracy of sweet spot identification remains limited.
[0003] Currently, commonly used methods for evaluating sweet spots mostly employ static weighted or empirical scoring models, assigning fixed weights to different evaluation parameters for comprehensive evaluation. These methods lack dynamic adjustment capabilities and struggle to reflect the impact of changes in construction measures and reservoir response on the evaluation results. Furthermore, when faced with multi-source geological, engineering, and construction data, traditional evaluation methods struggle to effectively handle dimensional differences and complex nonlinear relationships between different parameters, leading to discrepancies between predicted sweet spot grades and actual production capacity, thus affecting the scientific validity and reliability of construction scheme design.
[0004] A fundamental limitation of existing evaluation and optimization methods lies in the fact that their basic evaluation and decision-making units remain at the scale of "well" or "section." In reality, the stimulation and production contribution of fracturing occurs at the "cluster" level. Due to the strong heterogeneity of reservoirs, the response and contribution of different fracturing clusters within the same well section are extremely uneven. Designing and evaluating using averaged parameters cannot achieve truly precise stimulation. Therefore, to change the past empirical approach relying on averaged parameters and truly achieve refined design and evaluation based on individual clusters, it is necessary to obtain gas production data for individual clusters that reflects their stimulation effect and to establish a comprehensive sweet spot evaluation model that can dynamically respond to different fracturing operation parameters based on the geological conditions of individual clusters. Individual cluster gas production is the ultimate criterion for evaluating optimization effectiveness, and the cluster-level dynamic comprehensive sweet spot model is the core tool for achieving precise optimization decisions.
[0005] In terms of fracturing operation optimization, existing technologies mainly rely on manual experience or single-objective optimization methods to determine the combination of fracturing operation parameters. There is a lack of systematic exploration into the complexity of the fracturing operation parameter space. Key fracturing operation parameters such as displacement, total fluid volume, and total sand volume have a significant impact on reservoir stimulation and gas production, and there are complex coupling relationships between different parameters. Traditional optimization methods struggle to simultaneously consider gas production enhancement, operational risk control, and economic benefits, resulting in a lack of intelligent decision-making and adaptive adjustment capabilities in complex reservoir conditions.
[0006] In recent years, machine learning technology has been applied to reservoir evaluation and construction optimization. However, existing models still have certain limitations, such as high dependence on training data, insufficient model generalization ability, and high computational complexity. Furthermore, multi-objective optimization often requires manual setting of weight coefficients, making it difficult to obtain the globally optimal construction strategy in complex development environments, and lacking the ability to dynamically update and perform closed-loop optimization based on actual production data. Summary of the Invention
[0007] To address at least one of the aforementioned problems, this invention proposes an optimization method for deep coal and gas fracturing operations based on a comprehensive sweetness factor. This method constructs a dynamically updatable comprehensive sweetness factor based on cluster-level data and optimizes the fracturing operation scheme based on candidate scheme generation and deep reinforcement learning.
[0008] The technical solution of this invention to solve the above problems is as follows: an optimization method for deep coal and rock gas fracturing construction based on comprehensive sweet spot, comprising the following steps: S1. Obtain multi-source data of the deep coal-rock gas reservoir in the block where the target well is located, and preprocess it to obtain a standardized dataset. The multi-source data includes cluster-level geological parameters, fracturing operation parameters, and gas production. S2. Based on the standardized dataset, the LightGBM model is trained with gas production as the output, and the training results are interpreted using SHAP to obtain the contribution of each parameter to gas production. Based on the standardized dataset, the initial weights of each parameter are determined using the entropy weight method. The contribution and the initial weights are then fused to obtain the dynamic weights. S3. Based on geological parameters and their dynamic weights, an uncontrollable comprehensive sweetness factor is calculated; based on fracturing construction parameters and their dynamic weights, a controllable comprehensive sweetness factor is calculated; and based on both the uncontrollable and controllable comprehensive sweetness factors, a comprehensive sweetness factor is calculated: In the formula, F Indicates the overall sweetness factor. This represents the overall sweetness factor, which is uncontrollable. This represents the controllable overall sweetness factor. ohIndicates the weighting of the sweetness factor composition; S4. Using fracturing construction parameters as controllable parameters, geological parameters as uncontrollable parameters, and comprehensive sweetness factor as target parameter, establish a state transition regression model, and use a multilayer perceptron to train the aforementioned model to obtain a comprehensive sweetness factor prediction model. S5. Select historical effective fracturing operation parameters, generate multiple fracturing operation schemes using a diffusion model, and use a comprehensive sweetness factor prediction model to screen the aforementioned schemes. Repeat this step until a preset number of schemes are obtained, and use them as candidate schemes. S6. Establish a reward function that considers profit, comprehensive sweetness factor and risk penalty, with comprehensive sweetness factor and geological parameters as state input and fracturing operation parameters as action output. Train the function using a deep reinforcement learning model based on deep Q network. The trained model can then optimize the fracturing operation plan for the target well.
[0009] The beneficial effects of this invention are as follows: by optimizing fracturing construction parameters, this invention achieves a continuous and significant improvement in the comprehensive potential of reservoirs (comprehensive sweetness factor), thereby driving a leap in sweetness level in the evaluation system and providing more scientific, intelligent and reliable technical support for deep coal and gas fracturing construction design. Attached Figure Description
[0010] Figure 1 This is a flowchart of a method according to an embodiment of the present invention. Detailed Implementation
[0011] The specific embodiments of the present invention will be clearly and completely described below with reference to examples. Obviously, the described examples are only some embodiments of the present invention, and not all embodiments.
[0012] See Figure 1 An optimization method for deep coal and rock gas fracturing construction scheme based on comprehensive sweet spot includes the following steps: S1. Obtain multi-source data of the deep coal-rock gas reservoir in the block where the target well is located, and preprocess it to obtain a standardized dataset. The multi-source data includes cluster-level geological parameters, fracturing operation parameters, and gas production. In this step, the geological parameters include clay content, porosity, gas content, horizontal stress difference coefficient, brittleness coefficient, and cleavage development index. The fracturing construction parameters include construction flow rate, total fluid volume, total sand volume, average sand ratio, and pre-fracturing fluid percentage. Expandable parameters include, but are not limited to, elastic modulus, Poisson's ratio, maximum / minimum horizontal principal stress, etc. The gas production refers to the actual gas production of the fracturing cluster.
[0013] The above parameters are obtained through the following methods: geological parameters are mainly obtained through well logging interpretation (such as sonic logging, density logging, neutron logging, resistivity logging, etc.) and core experimental analysis; fracturing construction parameters are derived from construction instrument data and design reports; and the gas production of a single cluster needs to be estimated and obtained through tracer monitoring, production logging, or production capacity splitting models based on flowing pressure data.
[0014] Among them, porosity represents the size of the space in the reservoir for storing gas, and clay content represents the degree to which this space is blocked by impurities and the obstruction to gas flow. These two parameters, from the two orthogonal dimensions of "reservoir capacity" and "reservoir-permeability efficiency", provide the necessary and physically meaningful input features for the subsequent construction of a comprehensive sweet spot evaluation model and a production capacity prediction model to characterize the static endowment of the reservoir.
[0015] Gas content is a core parameter characterizing reservoir resource potential. For calculating gas content, a multivariate regression well logging interpretation model can be established based on the correspondence between measured gas content in other strata of the block and well logging parameters. In the formula, It is the gas content , It is the volume density. , It is the natural gamma value. The model coefficients are obtained through multivariate fitting regression. The fitted gas-bearing well logging interpretation model can achieve continuous prediction of gas content in different strata of this block.
[0016] The horizontal stress difference coefficient is used to characterize the control effect of stress conditions on fracture propagation. Engineering sweet spots are primarily evaluated based on reservoir fracturing capability and fracture network formation ability, assessing the ease of fracturing operations and the effectiveness of fracturing. Sweet spots have a smaller horizontal stress difference coefficient and are more likely to form complex fracture networks. The formula for calculating the horizontal stress difference coefficient is: In the formula, It is the coefficient of horizontal stress difference. It is the maximum principal stress , It is the minimum principal stress .
[0017] Rock brittleness directly affects fracture initiation and propagation behavior and is an important parameter reflecting reservoir fracturing capability. In this step, the Rickman brittleness index model is used to calculate the brittleness index based on Young's modulus and Poisson's ratio. In the formula, BI represents the brittleness index. Indicates Young's modulus , Represents Poisson's ratio. These represent the minimum and maximum values of Young's modulus, respectively. These represent the minimum and maximum values of Poisson's ratio, respectively.
[0018] To characterize the control effect of the cleavage system on the crack propagation path, a cleavage development index is introduced. As one of the evaluation parameters for engineered desserts: In the formula, This indicates the amplitude of fluctuations in the resistivity curve. This represents the change in the time difference of sound waves. Indicates the natural gamma anomaly response value. It is a weighting coefficient used to reflect the contribution of different logging parameters to the degree of cleavage development.
[0019] As those skilled in the art know, the fracturing operation parameters, including the displacement rate, total fluid volume, total sand volume, average sand ratio, and pre-fracturing fluid percentage, all have a significant impact on the final fracturing effect. Therefore, these parameters are used as fracturing operation parameters in this step.
[0020] In particular, due to the different geological parameters, the fracturing parameters of each fracturing segment vary during fracturing operations, especially multi-segment fracturing operations. Therefore, in order to study the fracturing operations of each cluster, the geological parameters mentioned above are the data collected for each cluster, and the fracturing parameters are the data collected during the single-segment fracturing operation. The cluster-average fracturing parameters will be used in subsequent studies.
[0021] For gas production, the contribution ratio of each cluster is obtained by inverting data from production logging, tracer monitoring, or distributed acoustic sensing during the production period. Multiplying this ratio by the actual gas production of this segment (or well) yields the cluster-level gas production label. If direct monitoring is unavailable, initial estimation can be made using a flow distribution model, combining geological sweet spots and fracturing parameters. This is a standard procedure in the field, and its specific operation will not be elaborated upon here.
[0022] Specifically, in actual fracturing processes, especially in staged fracturing, the fracturing parameters for each fracturing stage are the same. However, the actual geological conditions differ slightly between different clusters (usually with a certain interval between them, such as 10m or longer). Furthermore, considering actual horizontal wells, the depths between adjacent clusters and adjacent stages also differ. Therefore, in this step, the relevant parameters between different clusters within the same fracturing stage can be determined using the following method: First, determine the specific depth of the fracturing stage and cluster. Then, based on the logging information under these depth conditions, determine the relevant parameters. Considering the actual depth difference between clusters, a step size of 0.25m, 0.5m, etc., can be used for adjacent depths. Those skilled in the art can select an appropriate step size based on the actual situation.
[0023] After obtaining the above data, preprocessing is required. This step mainly involves three preprocessing methods: The obtained data may contain outliers and missing values; therefore, outliers and missing values must first be processed. For outliers, the quartile method is used for identification, and once identified, they are removed. After removing outliers, considering both the original missing values in the data and the missing values after outlier removal, data imputation is necessary. This step can use adjacent well interpolation or mean imputation to complete the missing values. The aforementioned methods are all conventional in this field, therefore their specific operations will not be elaborated upon here.
[0024] After completing the data, considering that the physical meanings of each parameter are different and there are significant differences in their dimensions and orders of magnitude, it is necessary to normalize each parameter in this step.
[0025] Specifically, in this step, the normalization methods used for positive parameters (i.e., parameters whose larger values lead to better target values) and negative parameters (i.e., parameters whose smaller values lead to better target values) are slightly different: the normalization formula for positive parameters is shown below: The normalization formula for the inverse parameter is shown below: In the formula, Indicates the first In the nth sample Normalized values of the parameters; It is the first In the nth sample The original values of the parameters, and They are the first The maximum and minimum values of each parameter across all cluster samples.
[0026] After the above operations, a standardized dataset is obtained: In the formula, This represents the standardized cluster-level multi-parameter data matrix; Indicates the first The first sample The normalized parameter values of each parameter, and i =1,2… n , j =1,2… m , n Indicates the total number of clusters. m Indicates the total number of parameters.
[0027] S2. Based on the standardized dataset, the LightGBM model is trained with gas production as the output, and the training results are interpreted using SHAP to obtain the contribution of geological parameters and fracturing construction parameters to gas production. Based on the standardized dataset, the initial weights of each parameter are determined using the entropy weight method. The contribution and initial weights are then fused to obtain the dynamic weights. The LightGBM model is composed of An ensemble of sequentially trained decision trees (CART trees) whose predictions are the sum of the outputs of all trees: in, Indicates the first Predicted gas production per sample; Indicates the first The input feature vector of each sample, which constitutes That is, the splicing of geological parameter vectors and construction parameter vectors; Indicates the first Tree pairs of input The predicted value; Indicates the first Geological parameter vectors for each sample; Indicates the first A vector of construction parameters for each sample; This represents the index of the tree; This represents the total number of decision trees, and its core structural characteristics and training mechanism include: Histogram Algorithm: When constructing each tree, LightGBM discretizes continuous floating-point feature values into an integer histogram and finds the optimal split point based on the histogram.
[0028] Depth-constrained leaf-based growth strategy: At each split, the algorithm selects the leaf with the highest gain from all current leaves for splitting, instead of growing layer by layer. This strategy achieves better accuracy with the same number of splits, while also... Parameters such as tree depth are used to control tree depth and prevent overfitting.
[0029] Gradient one-sided sampling: When calculating the gradient in each iteration, GOSS retains all samples with large gradients and only randomly samples samples with small gradients, thereby improving training efficiency while maintaining data distribution.
[0030] Mutually exclusive feature bundling: EFB technology bundles mutually exclusive features (i.e. features that rarely take non-zero values at the same time) into a new feature, thereby effectively reducing feature dimensionality and further improving training speed.
[0031] Model training: using complete data samples from historical wells Supervised training of the model is performed, where, This represents the corresponding actual gas production, and the training objective is to minimize the predicted value. and The loss function (such as mean squared error) between the decision trees is iteratively increased using the gradient boosting algorithm to gradually correct the model's prediction error.
[0032] SHAP is a common interpretive model in the field of machine learning. In this step, no changes are made to SHAP. Therefore, the specific operation can refer to the conventional SHAP in this field.
[0033] The formula for calculating the SHAP value in this step is as follows: In the formula, Indicates the first The SHAP value of each parameter, that is, the contribution of that parameter to the model output; Indicates that it does not include the first A subset of parameters of 1 parameter Representing a subset The number of parameters; To represent factorial; This represents the total number of all input parameters; This indicates that only a subset of parameters is considered. Output in the LightGBM model Indicates in subset Add the first The output of the model after each parameter; Indicates the first The marginal contribution brought about by the addition of each parameter; Represents the set of all input parameters, i.e. The complete set of indices for all parameters.
[0034] This process mainly includes the following sub-steps: S21. Use the geological parameters and fracturing operation parameters in the standardized dataset as input to the LightGBM model, and the gas production in the standardized dataset as output to train the LightGBM model. S22. Use SHAP to interpret the trained LightGBM model to obtain the SHAP values of geological parameters and fracturing operation parameters, and use these values as the contribution of the parameters. S23. The initial weights of each parameter are calculated using the entropy weight method; The entropy weight method used in this step is a prior art technique, and its operation is as follows: First, the normalized values of each parameter in different samples are converted into weights, representing the contribution of the parameter information to the sample: in: Indicates the first The parameters in the sample The contribution percentage in It is the total number of cluster samples. It is the first The normalized parameter values of each parameter across all samples. Indicates the first All parameters The sum of normalized values in each sample.
[0035] Then calculate the information entropy of each parameter: in: It is the first The entropy value of each parameter; The information utility values of all parameters are normalized to obtain the initial weight of each parameter. ,Right now , in: It is the first The initial weights of each parameter reflect their importance to the overall dessert evaluation. It is the first The entropy value of each parameter It represents the maximum number of geological parameters and fracturing construction parameters.
[0036] S24. The initial weights and contributions are combined using the following formula to obtain the dynamic weights: In the formula, Indicates the first j Dynamic weights of each parameter; l This represents the fusion coefficient, with a value ranging from 0 to 1; Indicates the first j The contribution of each geological parameter to the fracturing operation parameter; Indicates the first j The initial weights of each parameter. Specifically, for the fusion coefficients, the purpose is to control the balance between the initial weights of the entropy weighting method and the dynamic contribution of SHAP. l The closer the value is to 1, the more reliable the objective statistical weights are. l The closer a value is to 0, the more confident one is in the actual contribution of the parameter to production capacity. Generally, a moderate value between the two is acceptable, such as 0.4, 0.5, or 0.6.
[0037] S3. Based on geological parameters and their dynamic weights, an uncontrollable comprehensive sweetness factor is calculated; based on fracturing construction parameters and their dynamic weights, a controllable comprehensive sweetness factor is obtained; based on the uncontrollable and controllable comprehensive sweetness factors, the comprehensive sweetness factor is calculated as follows: In the formula, F Indicates the overall sweetness factor. This represents the overall sweetness factor, which is uncontrollable. This represents the controllable overall sweetness factor. oh Indicates the weighting of the sweetness factor composition; In this embodiment, the overall sweetness factor is used as a quantitative indicator for overall desserts.
[0038] The formula for calculating the uncontrollable overall sweetness factor is as follows: , In the formula, Indicates the first Within-group normalized dynamic weights of geological parameters; Indicates the first Normalized values of geological parameters; The maximum number representing the geological parameter. ; Indicates the first Dynamic weights of individual geological parameters; The formula for calculating the controllable overall sweetness factor is as follows: , In the formula, Indicates the first Within-group normalized dynamic weights of each fracturing operation parameter; Indicates the first Normalized values of each fracturing operation parameter; This indicates the maximum number of the fracturing operation parameters. ; Indicates the first Dynamic weights of each fracturing operation parameter.
[0039] In the above calculation formula, the specific symbolic representation is related to the representation of the dataset in this embodiment: In this embodiment, a total of m There are 10 parameters, of which the first 10 is the first 10. These parameters are geological parameters. arrive The parameters between these are the fracturing construction parameters, the first... m The parameter is the gas production rate. If those skilled in the art use other methods to number the parameters, they can appropriately modify the above formula.
[0040] Specifically, regarding the weighting of the sweetness factor synthesis... ohThe sweetness factor is used to balance the importance of the reservoir's "inherent endowment" (i.e., the reservoir's inherent exploitability) and "acquired modification" (i.e., fracturing-related operations) in the overall sweetness factor. Its value can be set according to different development stages: for example, in the early stages of production, fracturing operations emphasize the reservoir itself, therefore... oh Higher values can be set, such as 0.7; in the mid-to-late stages of fracturing operations, since the easily exploitable oil and gas in the reservoir has already been developed, the focus is more on the fracturing's effectiveness. oh It can be set to a relatively low value, such as 0.3.
[0041] Meanwhile, the overall sweetness factor of each cluster calculated in this step It is a continuous variable, and its range is usually within Between. In subsequent steps, It will be used as a core variable, along with geological parameters Together they constitute a continuous state vector describing the reservoir state. This is used to drive the optimization process of intelligent decision-making models. Based on the cluster-integrated sweetness factor... Reservoirs are classified into different sweet spot grades (e.g., Class I, Class II, Class III). These grades are typically determined by setting certain thresholds based on actual data and development experience (e.g., [missing information]). Class I, Class II (Class III).
[0042] S4. Using fracturing operation parameters as controllable parameters, geological parameters as uncontrollable parameters, and the comprehensive sweetness factor as the target parameter, establish a state transition regression model, and train the aforementioned model using a multilayer perceptron to obtain a comprehensive sweetness factor prediction model; this step includes the following sub-steps: In this step, the state transition regression model is first established: In the formula, h (·) represents the specific form of the function. This indicates the current overall sweetness factor. Geological parameters are uncontrollable parameters. This refers to the combination of controllable parameters used in fracturing operations. This represents the predicted overall sweetness factor after the action is performed.
[0043] This step uses a multilayer perceptron (MLP) as the function. h The specific implementation of (·) is as follows. MLP is a conventional model in this field, and its structure is given below: Input layer: Receives the concatenated composite feature vector. For example, in this embodiment It is 6-dimensional. If it is 5-dimensional, then the number of neurons in the input layer is 1 + 6 + 5 = 12; Hidden layers: consist of 2–3 fully connected layers, each including: a linear mapping and a non-linear activation function. A typical network structure example is shown below: in: This represents a linear fully connected layer, used to... dimensional input mapping is Dimensionality of the output; 12 indicates the dimension of the model input parameters, 64 and 32 represent the number of neurons in the first and second hidden layers, respectively; the ReLU activation function is used to enhance the non-linear expressive power of the model; This means that some neurons are randomly deactivated with a probability of 0.1 during training to reduce the risk of model overfitting and improve generalization ability.
[0044] Output layer: is a linear fully connected layer Output a single scalar: Indicates the execution of engineering actions The post-construction comprehensive sweetness factor was predicted.
[0045] The training samples for the aforementioned MLP are derived from historical data within this block. The training objective is to minimize the mean squared error between the predicted and actual values. The network parameters are iteratively optimized using the backpropagation algorithm. After training, the model can quickly predict changes in the overall sweetness factor.
[0046] S5. Select historical effective fracturing operation parameters, generate multiple fracturing operation schemes using a diffusion model, and use a comprehensive sweetness factor prediction model to screen the aforementioned schemes. Repeat this step until a preset number of schemes are obtained, and use them as candidate schemes. In this step, the diffusion model includes a diffusion structure and a denoising network. The diffusion structure is a Markov chain, and Gaussian noise is gradually added during the diffusion process. The denoising network adopts a U-Net encoder-decoder architecture, which includes an input layer, an encoder, an intermediate layer, a decoder, and an output layer.
[0047] Specifically, the diffusion structure is a Markov chain with progressively increasing noise, where the added noise is Gaussian noise, and each noise addition operation conforms to a pre-defined variance scheduling table. ,in, This indicates that the noise variance gradually increases with the number of diffusion steps, and always remains between 0 and 1; Indicates the first The variance of the noise; This indicates the current noise addition step number during the diffusion process; This represents the total number of noise-adding steps in the diffusion process; Indicates from step 1 to step 2. The noise variance or noise intensity of the step.
[0048] The parameters after adding noise can be expressed as: in, Indicates the first The parameters after adding noise. ; Indicates the first The parameters after the noise addition step are the same as the parameters of the previous noise addition step. This represents the retention coefficient of the parameters from the previous step; This represents the scaling factor for adding noise in the current step; Indicates the first Random noise obtained from step sampling; This indicates that the noise follows a standard normal distribution with a mean of 0 and a variance of 1. This represents a standard normal distribution.
[0049] As mentioned above, a denoising network consists of an input layer, an encoder, an intermediate layer, a decoder, and an output layer, with the specific structure as follows: Input layer: Receives a dimensional vector , representing the current noisy fracturing operation parameters, where, The dimension of the parameter vector (the number of parameters contained in a set of construction schemes). Indicates the first The parameters after adding noise.
[0050] Conditional information injection: Time step embedding: time step It can be converted into a fixed-dimensional vector through sinusoidal position encoding or a learnable embedding layer. ,in, Represents the time step embedding vector, which is composed of time steps. A fixed-dimensional vector obtained through sinusoidal position encoding or learnable embedding layer transformation is used to inject temporal information into the network.
[0051] State-condition embedding: Reservoir continuous state The vector is projected through a small multilayer perceptron (conditional encoder) to a conditional vector that matches the network feature dimensions. ,in, Indicates the first The corresponding state condition information (reservoir continuous state) for each step typically includes a comprehensive sweetness factor and engineering state parameters; Indicates the first The overall sweetness factor corresponding to each step is part of the state conditions; Indicates the first The reservoir state parameters corresponding to each step are part of the state conditions; This represents the conditional vector obtained after projection by the conditional encoder. Its dimension matches the feature dimension of the denoising network and is used for conditional generation.
[0052] Condition vector and time embedding The conditional information is injected into each residual block of U-Net through adaptive group normalization or feature concatenation. Specifically, in AdaGN, conditional information is used to dynamically generate scaling and translation parameters for the normalization layer, thereby achieving fine-grained conditional control of network behavior.
[0053] The encoder (downsampling path) consists of alternating residual blocks and downsampling layers. Each residual block typically contains group normalization, activation functions, convolutional layers, and conditional information. Downsampling layers (such as convolutions with a stride of 2) progressively reduce the spatial resolution of the feature maps (length in sequence data) and increase the number of channels to extract high-level abstract features.
[0054] The decoder (upsampling path) is symmetrical to the encoder and consists of multiple "upsampling blocks" concatenated. Each block typically contains upsampling operations (such as transposed convolution), skip connections to features from the corresponding encoder layer, group normalization, activation functions, convolutions, and conditional injection. Skip connections preserve detailed information from the encoder, which helps to accurately reconstruct the data.
[0055] Intermediate layer: Connects the encoder and decoder with the lowest resolution features, usually also consisting of residual blocks, used to process the highest level of abstract features.
[0056] Output layer: The last layer is usually a convolutional layer or a fully connected layer, whose output dimension is the same as the input dimension. Same, that is The layer outputs the predicted noise. ,in, The dimension of the parameter vector (the number of parameters contained in a set of construction schemes). This represents the noise predicted by the denoising network. This represents the learnable parameters in the denoising network; Indicates the first The state condition information corresponding to each step.
[0057] In the above denoising network model, the samples from the previous step are calculated according to the following formula: In the formula, The denoised result is the first... Step parameters; This represents the noisy input parameters; Indicates the first The candidate sample at the th The noisy parameters of the step; Indicate the first The candidate sample at the th The parameters after denoising step; Represents the first step in the diffusion sampling process. Each candidate sample number; This represents the cumulative retention factor, from step 1 to step 2. Step All Multiplication of: ; This indicates additional noise from the reverse sampling process; This indicates that the noise follows a standard normal distribution with a mean of 0 and a variance of 1. This represents a standard normal distribution.
[0058] The aforementioned denoising network yields K sets of schemes. However, these schemes still need to be screened based on the prediction results of the comprehensive sweetness factor prediction model. Only when the predicted value of any of these K schemes is greater than the original value (i.e., the original data) is the scheme considered to have a better reservoir stimulation effect than the original scheme, and thus the scheme is retained. Furthermore, the retained schemes require further evaluation: if the relevant parameters of the construction scheme significantly exceed the on-site construction level, such as exceeding the equipment capacity limit, these schemes also need to be eliminated. The final schemes obtained are high-quality, on-site executable candidate schemes.
[0059] In this step, historical data needs to be collected multiple times, and schemes are generated using a diffusion model and then filtered to obtain a sufficient number of samples as candidate schemes. In practice, the number of candidate schemes can be set to 50-200 groups, such as 50, 80, 100, 150, 180, 200, etc. S6. Establish a reward function that considers profit, comprehensive sweetness factor and risk penalty, with comprehensive sweetness factor and geological parameters as state input and fracturing operation parameters as action output. Train the function using a deep reinforcement learning model based on deep Q network. The trained model can then optimize the fracturing operation plan for the target well.
[0060] In this step, each fracturing operation parameter is first discretized according to the preset step size to obtain the discrete values of the fracturing operation parameter. The discrete values of all fracturing operation parameters are combined by Cartesian product to obtain the action space of the fracturing operation parameter. Since fracturing parameters can usually be selected within a range, such as the average sand ratio, which can usually be set in the range of 15~25 wt%, if they are set as point values, there will be too many point values, which will affect subsequent calculations.
[0061] Therefore, these fracturing operation parameters first need to be discretized. During discretization, appropriate step sizes are set for different fracturing operation parameters and different construction sites. For example, for the average sand ratio, the step size can be set to 2 wt%, 3 wt%, or 5 wt%. After setting the step size, the parameters are divided according to the conventional setting range, resulting in multiple discretized values. For example, for the average sand ratio, if a step size of 3 wt% is used, 12 discretized values can be obtained.
[0062] During the subsequent training of the deep Q-network, the corresponding parameters are selected from the action space mentioned above.
[0063] The reward function is as follows: In the formula, Indicates the first Momentary rewards; This indicates a reward for increasing sweetness. This indicates the change in instantaneous gas production. Indicate execution The change in net present value resulting from the construction activity; Indicates execution Net present value obtained after construction activities; Indicates the discount period index; This indicates the predicted gas production after performing an action in the current state; This indicates the gas production rate corresponding to the current state; This represents the comprehensive sweetness factor prediction model for S4; This indicates the current overall sweetness factor; Indicates construction risk indicators; This indicates the first prediction based on LightGBM in S2. Gas production during the period; This indicates the unit price of gas; Indicates the first Construction costs during the initial phase; Indicates operating and maintenance costs; Indicates the discount rate; Indicates the forecast period; This indicates the weight of each risk sub-item; it is determined comprehensively by combining the analysis of historical risk event data using the analytic hierarchy process; it is used to balance the relative importance of different risk types in the overall penalty. This indicates a deviation in crack prediction; This indicates a risk of construction pressure exceeding the standard; Indicates the risk of sand blockage; This represents the weighting coefficients of the reward function.
[0064] For deep Q-networks, which are also common models in this field, their architecture is as follows in this embodiment: Input layer: This layer receives the state feature vector. As input, where The dimension of the state vector is represented by the combined sweetness factor. and geological parameter vector The dimension determines the scope of this embodiment. .
[0065] Hidden layers: The core of the network consists of 2 to 3 fully connected layers, responsible for nonlinear transformations, feature extraction, and abstraction of the high-dimensional state features of the input. A typical and efficient structural configuration is shown below: First hidden layer: This layer projects the original state vector into a 256-dimensional high-dimensional space. It is then followed by a ReLU activation function, introducing nonlinearity and enabling the model to learn complex value function patterns. This is typically achieved by introducing... Regularization layers are used to randomly discard some neuron connections to prevent the model from overfitting to specific features during training and to enhance its generalization ability. Indicates a fully connected layer, from 1-dimensional mapping to 256-dimensional; This indicates a discard rate of 0.2%. dropout Layer; ReLU represents the activation function.
[0066] Second hidden layer: This layer further condenses and refines features, learning deeper, more decision-related abstract representations in the state space. The stacking of multiple hidden layers allows the network to approximate arbitrarily complex nonlinear Q-value surfaces. This represents a fully connected layer, which maps from 256 dimensions to 128 dimensions.
[0067] Output layer (Action Value Assessment): This is the key layer of the DQN architecture; it is a linear layer. ,in It is the refined fracturing construction parameter candidate set obtained after screening in step S3. The size of the output of this layer is [amount]. 3D vector Each component in the vector That is to say, given the current state and network parameters Under the condition of, choose the first Actions of candidate construction parameters The network's ultimate goal is to make these output values approximate the theoretically optimal Q-value, based on the estimated expected discounted future cumulative reward (Q-value). This indicates the size of the candidate set of refined fracturing operation parameters obtained after screening in step S3, i.e., the number of candidate actions; Represents the set of candidate actions. Include A vector of possible fracturing operation parameters; This represents a specific candidate action (each action is a vector of construction parameters). q Indicates the output of the output layer L dimensional vector, Each component corresponds to a candidate action. Q value.
[0068] In this embodiment, the training and optimization mechanism of the deep Q-network is as follows: State transition: Invoke the comprehensive sweetness factor change prediction model trained by S4. Calculate the predicted overall sweetness factor for the next time step. Thus, the next state is obtained. Reward calculation: based on And by calling the LightGBM gas production prediction model of S21 to calculate... and In conjunction with construction risks Calculate the instant reward based on the reward function formula of S6. This simulated environment allows intelligent agents to learn efficiently in a virtual world driven by historical data without having to conduct real, costly field experiments.
[0069] Continuously iterate and optimize network parameters The process. To improve training stability and sample efficiency, this invention integrates the following key technologies: Experience replay: To avoid training instability caused by strong correlations between sequence data, this invention establishes a fixed-capacity experience replay buffer. The transfer experience generated by the agent's interaction with the environment at each step, i.e., the four-tuple... The data will be stored in this buffer. During training, a small batch of independent historical experiences are uniformly and randomly sampled from the buffer for network updates. This effectively breaks the temporal correlation between data, greatly improving data utilization efficiency and training stability.
[0070] Target network: This is used to address the target value in temporal difference learning. Depending on network parameters To address the training oscillation problem caused by frequent changes, this invention introduces an independent target network with the same structure as the main network, but with different parameters. Updates are slow. This target network is used when calculating the temporal difference objective: in, This represents the target value, used as a monitoring signal when updating the Q-network; Represents the set of candidate actions Take the maximum value of all possible actions; Represent candidate action variables, iterate through Each action in; Indicates the state The set of all candidate actions that can be executed; These represent the parameters of the target Q-network; The discount factor represents how much importance the agent places on future rewards; Indicates the use of target network parameters The calculated state-action value function represents the state... Next action The expected cumulative return. Target network parameters. Regularly update the main network parameters via a soft update strategy. synchronous: This slow update method keeps the target Q value relatively stable in the short term, thereby guiding the main network toward robust convergence.
[0071] Loss Function and Parameter Update: The core of training is to minimize the error between the current network's predicted Q-value and the target Q-value. The temporal difference error loss function is defined as the mean squared error. in, Representing the loss function value, in terms of network parameters i The training objective for the variables is to minimize this loss; This represents one or more empirical samples randomly sampled from the empirical playback buffer. Calculate the mean; Buffer represents the experience replay buffer, used to store historical interaction experiences. ; Indicates the current Q Network parameters i Below, regarding the state S and actions A Output prediction Q value; S This represents the current state feature vector (the state in the sampled sample); A In state S The action to be executed (the action in the sampled sample); R Indicates the execution of an action A The immediate reward obtained afterward (the reward in the sampled sample); Indicate the next state (i.e.) ).
[0072] In each training iteration, the loss function is calculated. Regarding network parameters The gradient is calculated, and the backpropagation algorithm and optimizer (such as Adam) are used to update the gradient along the gradient descent direction. This allows the predicted value to continuously approach a better target value.
[0073] Exploration-Exploitation Strategy: In the early stages of training, the agent has limited understanding of the state-action space. Therefore, an exploration-exploitation strategy is employed. - Greedy strategy for selecting actions: based on probability Random from candidate set Choose one action (exploration) to discover potentially high-value areas; based on probability. Select the action with the highest current Q-value (exploitation). As training progresses, the exploration rate... The value decays linearly or exponentially from an initial value (e.g., 1.0) to a smaller value (e.g., 0.01), allowing the agent's behavior to gradually transition from extensive exploration to fully utilizing the learned knowledge.
[0074] After the above training, the optimal decision-making strategy is finally produced in this step. It is a deterministic function parameterized by a neural network. Its mathematical expression is a greedy strategy: In the formula, This indicates the optimal policy at time [time]. The output; Represents the current state vector; This represents the optimal network parameters after training. Indicates the action of candidate fracturing operation parameters; This represents the set of candidate fracturing operation parameters. This indicates taking the parameter that maximizes the function; Indicates the state Next action The predicted value of the future cumulative discount reward that can be obtained under the optimal strategy.
[0075] The final output is an optimal solution recommendation that includes the aforementioned fracturing construction parameters.
[0076] To help those skilled in the art understand the advantages of the embodiments of the present invention, specific examples are given below.
[0077] The inventors collected data from 23 wells, totaling approximately 20,000 records, a substantial amount. Therefore, in this example, statistical methods are used to represent the data. Historical development data for this block is shown in Table 1.
[0078] Table 1 Data Statistics Table This invention selects two wells in the block, quantifies the "comprehensive sweet spot" to identify reservoir quality, compares the production response of "high sweet spot + conventional construction" and "low sweet spot + enhanced construction", and verifies that improving construction parameters can effectively compensate for the poor physical properties of low sweet spot reservoirs, achieving a significant increase in production, and providing a scientific basis for differentiated fracturing design of unconventional oil and gas reservoirs.
[0079] Table 2 Comparison Results of Two Wells Although the reservoir conditions in well B were poor (the overall sweetness factor was only 0.46), through enhanced fracturing (increasing the sand addition intensity by 32% and the fracturing flow rate by 21%), its overall sweetness factor reached 0.71, which is 93% of that of the high sweetness well A (0.76), and is significantly better than the expected 0.46 before optimization. This shows that enhanced fracturing parameters can effectively promote the transition from low sweetness wells to high sweetness wells.
[0080] The present invention has been disclosed above with preferred embodiments. However, those skilled in the art should understand that these embodiments are only for describing the present invention and should not be construed as limiting the scope of the present invention. Further improvements can be made without departing from the principles of the present invention, and these improvements should also be considered within the scope of protection of the present invention.
Claims
1. A method for optimizing a deep coal rock gas fracturing operation plan based on a comprehensive dessert, characterized in that, Includes the following steps: S1. Obtain multi-source data of the deep coal-rock gas reservoir in the block where the target well is located, and preprocess it to obtain a standardized dataset. The multi-source data includes cluster-level geological parameters, fracturing operation parameters, and gas production. S2. Based on the standardized dataset, the LightGBM model is trained with gas production as the output, and the training results of the model are interpreted using SHAP to obtain the contribution of each parameter to gas production. Based on a standardized dataset and using the entropy weight method, the initial weights of each parameter are determined. The contribution level and the initial weight are combined to obtain the dynamic weight; S3. Based on geological parameters and their dynamic weights, the uncontrollable comprehensive sweetness factor is calculated. Based on the fracturing construction parameters and dynamic weights thereof, a controllable comprehensive sweetness factor is calculated; and based on the uncontrollable comprehensive sweetness factor and the controllable comprehensive sweetness factor, a comprehensive sweetness factor is calculated: , wherein, F represents the comprehensive sweetness factor, represents the uncontrollable comprehensive sweetness factor, represents the controllable comprehensive sweetness factor, ω represents a sweetness factor synthesis weight; S4. Using fracturing construction parameters as controllable parameters, geological parameters as uncontrollable parameters, and comprehensive sweetness factor as target parameter, establish a state transition regression model, and use a multilayer perceptron to train the aforementioned model to obtain a comprehensive sweetness factor prediction model. S5. Select historical effective fracturing operation parameters, generate multiple fracturing operation schemes using a diffusion model, and use a comprehensive sweetness factor prediction model to screen the aforementioned schemes. Repeat this step until a preset number of schemes are obtained, and use them as candidate schemes. S6. Establish a reward function that considers profit, comprehensive sweetness factor and risk penalty, with comprehensive sweetness factor and geological parameters as state input and fracturing operation parameters as action output. Train the function using a deep reinforcement learning model based on deep Q network. The trained model can then optimize the fracturing operation plan for the target well.
2. The method according to claim 1, wherein, In S1, the geological parameters include clay content, porosity, gas content, horizontal stress difference coefficient, brittleness coefficient, and cleavage development index. The fracturing construction parameters include construction discharge rate, total fluid volume, total sand volume, average sand ratio, and pre-fracturing fluid percentage. The gas production rate refers to the actual gas production rate of the fracturing cluster.
3. The method according to claim 1, wherein, In S1, the preprocessing includes: identifying and deleting outliers using the quartile method; filling in missing values using adjacent well interpolation or mean imputation; and normalizing the imputed data.
4. The method according to claim 1, wherein, In S2, the dynamic weights are calculated as follows: S21. Use the geological parameters and fracturing operation parameters in the standardized dataset as input to the LightGBM model, and the gas production in the standardized dataset as output to train the LightGBM model. S22. Use SHAP to interpret the trained LightGBM model to obtain the SHAP values of geological parameters and fracturing operation parameters, and use these values as the contribution of the parameters. S23. The initial weights of each parameter are calculated using the entropy weight method; S24, the initial weight and the contribution degree are fused by using the following formula to obtain a dynamic weight: , in the formula, indicates the dynamic weight of the first j parameter. λ This represents the fusion coefficient, with a value ranging from 0 to 1; Indicates the first j The contribution of each parameter , This indicates the maximum number of geological parameters and fracturing operation parameters; Indicates the first j The initial weights of the parameters.
5. The method for optimizing deep coal and rock gas fracturing construction scheme based on comprehensive sweet spot as described in claim 1, characterized in that, In S3, the formula for calculating the uncontrollable overall sweetness factor is as follows: , In the formula, Indicates the first Within-group normalized dynamic weights of geological parameters; Indicates the first Normalized values of geological parameters; The maximum number representing the geological parameter. ; Indicates the first Dynamic weights of individual geological parameters; The formula for calculating the controllable overall sweetness factor is as follows: , In the formula, Indicates the first Within-group normalized dynamic weights of each fracturing operation parameter; Indicates the first Normalized values of each fracturing operation parameter; This indicates the maximum number of the fracturing operation parameters. ; Indicates the first Dynamic weights of each fracturing operation parameter.
6. The method for optimizing deep coal and rock gas fracturing construction scheme based on comprehensive sweet spot as described in claim 1, characterized in that, In S5, the diffusion model includes a diffusion structure and a denoising network. The diffusion structure is a Markov chain, and Gaussian noise is gradually added during the diffusion process. The denoising network adopts a U-Net encoder-decoder architecture, which includes an input layer, an encoder, an intermediate layer, a decoder, and an output layer.
7. The method for optimizing deep coal and rock gas fracturing construction scheme based on comprehensive sweet spot as described in claim 1, characterized in that, In S5, when using the comprehensive sweetness factor prediction model to screen the aforementioned schemes, the following operations are included: after obtaining multiple schemes using the diffusion model, these schemes are substituted into the comprehensive sweetness factor prediction model for calculation, and schemes with calculated values greater than the comprehensive sweetness factor before diffusion are retained. Then, schemes that violate engineering constraints are removed from the retained schemes, and the final schemes are the candidate schemes.
8. The method for optimizing deep coal and rock gas fracturing construction scheme based on comprehensive sweet spot as described in claim 1, characterized in that, In S6, the reward function is as follows: In the formula, Indicates the first Momentary rewards; This indicates a reward for increasing sweetness. This indicates the change in instantaneous gas production. Indicate execution The change in net present value resulting from the construction activity; Indicates execution Net present value obtained after construction activities; Indicates the discount period index; This indicates the predicted gas production after performing an action in the current state; This indicates the gas production rate corresponding to the current state; This represents the comprehensive sweetness factor prediction model for S4; This indicates the current overall sweetness factor; Indicates construction risk indicators; This indicates the first prediction based on LightGBM in S2. Gas production during the period; Indicates the unit price of gas; Indicates the first Construction costs during the initial phase; Indicates operating and maintenance costs; Indicates the discount rate; Indicates the forecast period; This indicates the weight of each risk sub-item; it is determined comprehensively by combining the analysis of historical risk event data using the analytic hierarchy process; it is used to balance the relative importance of different risk types in the overall penalty. This indicates a deviation in crack prediction; This indicates a risk of construction pressure exceeding the standard; Indicates the risk of sand blockage; This represents the weighting coefficients of the reward function.