Multi-objective fracturing optimization design method for shale reservoirs based on reinforcement learning

Through a multi-objective optimization design method based on reinforcement learning, combined with numerical simulation and economic evaluation model, the fracturing design parameters of shale reservoir are optimized, which solves the problem that the existing technology cannot consider both production capacity and economic factors, and achieves efficient and coordinated optimization results.

CN119740494BActive Publication Date: 2025-05-23CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510246819.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-23
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The existing shale reservoir fracturing optimization design method cannot consider various factors such as production capacity and economy at the same time, resulting in poor optimization results.

Method used

A multi-objective optimization design method based on reinforcement learning is adopted to establish a capacity prediction model through numerical simulation and economic evaluation model, and combined with the Transformer architecture and dual-agent reinforcement learning algorithm, the fracturing design parameters are optimized to achieve coordinated optimization of production capacity and economy.

Benefits of technology

The prediction accuracy and efficiency of fracturing optimization design are improved, and the coordinated optimization of production capacity and economy is achieved, and the cost of fracturing is controlled and crude oil output is increased.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119740494B_ABST
    Figure CN119740494B_ABST
Patent Text Reader

Abstract

The present invention belongs to the technical field of shale reservoir fracturing optimization design, and specifically discloses a shale reservoir multi-objective fracturing optimization design method based on reinforcement learning. It effectively solves the problem that the current shale reservoir fracturing optimization design method cannot consider the coordinated optimization of multiple factors such as production capacity and economy. It includes: (1) simulating the production capacity calculation results through numerical simulation software, and establishing a training sample set of the production capacity prediction model based on the Monte Carlo sampling method in combination with the constructed economic evaluation model; (2) constructing a production capacity prediction model based on the Transformer architecture; (3) On the basis of the NSGA‑II multi-objective optimization algorithm, introducing a dual-agent reinforcement learning algorithm to establish a fracturing optimization agent model; (4) By optimizing the Pareto frontier of the multi-objective optimization of cumulative oil production and economic net present value, a specific fracturing parameter optimization scheme is selected in combination with actual needs. The present invention can effectively increase crude oil production while controlling fracturing costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of shale reservoir fracturing optimization design, and in particular relates to a shale reservoir multi-objective fracturing optimization design method based on reinforcement learning. Background Art

[0002] Shale reservoirs have the characteristics of rapid lithology changes, large differences in physical properties, and strong heterogeneity. Horizontal well staged fracturing construction technology is an important technology for developing shale reservoirs and increasing the controlled reserves of single wells, and hydraulic fracturing optimization design plays a key role in improving the post-fracturing production capacity of shale oil and gas horizontal wells. However, due to the complexity of shale reservoir fractures and reservoir characteristics and the fact that a large number of fracturing design variables have not been fully and reasonably optimized, the fracturing process in the mine reservoir transformation process fails or the fracture extension length is insufficient, and an effective seepage channel cannot be formed, which makes the research on hydraulic fracturing optimization design still face major challenges.

[0003] At present, the research on commonly used fracturing layout design and fracturing parameter optimization mainly focuses on numerical simulation and multivariate analysis methods. Among them, in terms of multivariate analysis methods: the Chinese invention patent with publication number CN113090258A discloses the calculation of the "geological-engineering" comprehensive evaluation index of the horizontal section of shale gas wells based on logging data, thereby dividing the compressibility classification, and optimizing the fracturing process methods for different horizontal well sections, but it does not consider the post-compression capacity corresponding to different fracturing parameter design schemes, so that the optimized fracturing process methods lack actual data support; in terms of numerical simulation methods: the Chinese invention patent with publication number CN117634166A discloses the determination of the optimal fracturing parameters of the target fracturing section by constructing a final expected recoverable reserves prediction model and using the unconventional natural gas fracture development and expansion model established by the GOHFER fracturing software, but the economic net present value is not taken into consideration during its optimization process.

[0004] In addition to the above methods, artificial intelligence technology is also increasingly used in fracturing design. Methods such as big data, machine learning, and intelligent optimization algorithms provide new ideas for the optimization design of fracture parameters. Among them, the Chinese invention patent with publication number CN109236258A discloses the optimization design of fracturing parameters for horizontal wells in unconventional oil reservoirs based on the construction of an adaptive agent model, but it does not achieve the simultaneous optimization of post-fracturing production capacity and economic net present value.

[0005] In summary, the current shale reservoir fracturing optimization design method has the problem that the single-index fracturing parameter optimization method cannot consider the coordinated optimization of multiple factors such as production capacity and economy. Summary of the invention

[0006] The purpose of the present invention is to provide a multi-objective fracturing optimization design method for shale reservoirs based on reinforcement learning, which effectively solves the problem that the current shale reservoir fracturing optimization design method has a single indicator fracturing parameter optimization method that cannot consider the coordinated optimization of multiple factors such as production capacity-economy.

[0007] In order to solve the above technical problems, the technical solution adopted by the present invention is: a multi-objective fracturing optimization design method for shale reservoirs based on reinforcement learning, comprising the following steps: S1, simulating the capacity calculation results through numerical simulation software, and establishing a training sample set of the capacity prediction model based on the Monte Carlo sampling method in combination with the constructed economic evaluation model; S2, constructing the capacity prediction model based on the Transformer architecture, the model input includes but is not limited to the discrete data of the segment cluster design scheme and the fracturing construction parameters and the continuous dynamic data generated by the numerical simulation software when it is running, including but not limited to the production data and the bottom hole flow pressure time series, and the numerical simulation software input Output dynamic production capacity prediction data, and calculate the economic net present value through the constructed economic evaluation model; S3, construct optimization objective functions for cumulative oil production and economic net present value respectively, and introduce a dual-agent reinforcement learning algorithm including a production capacity optimization agent and an economic optimization agent on the basis of the traditional NSGA-II multi-objective optimization algorithm. The production capacity optimization agent and the economic optimization agent jointly optimize the fracturing design parameters through competition and collaboration, and establish a fracturing optimization agent model; S4, through the Pareto frontier of multi-objective optimization of cumulative oil production and economic net present value, select the specific fracturing parameter optimization scheme in combination with actual production and development needs.

[0008] Furthermore, in step S1, the economic evaluation model is:

[0009] ;

[0010] in, represents the economic net present value of shale reservoir development, is the total number of time steps, For the The length of the time step, is the discount rate, and are the numbers of production wells and injection wells drilled in the reservoir, and are the sales prices of oil and gas, respectively. is the unit cost of produced water purification, , and Respectively The production horizontal well The average oil production, average gas production and average water production at each time step, For the The horizontal well injection The amount of water injected per time step, and represent drilling cost and completion cost respectively.

[0011] ;

[0012] in, Represents the length of the horizontal section of a horizontal well.

[0013] ;

[0014] in, Represents the fracture volume of each fracturing stage.

[0015] Furthermore, the Monte Carlo sampling method is to obtain a sample set that reflects the uncertainty of the system by constructing a probability distribution model and performing random sampling in a given parameter space. First, the probability distribution model of reservoir physical parameters, fracturing operation parameters and engineering parameters is determined; then, in the actual sampling process, the physical constraints of each parameter are considered, and after the samples are generated, the validity is verified, and the sample combinations that do not meet the physical constraints are eliminated, and finally a valid parameter combination sample is obtained.

[0016] Furthermore, the Transformer architecture consists of an encoder and a decoder. The encoder consists of several encoder blocks, each of which contains a multi-head self-attention mechanism and a position-aware feedforward neural network. The encoder input sequence is a collection of various data type features including discrete data and continuous dynamic data. Discrete data includes but is not limited to segment cluster design schemes and fracturing construction parameters. Continuous dynamic data includes but is not limited to daily oil production, daily water production, daily liquid production and bottom hole flow pressure time series data information generated by the numerical simulation model during operation. The decoder input sequence is divided into two parts, one of which is the result generated by running the reservoir numerical simulator during the training of the production capacity prediction model; the other part is the output of the previous time step of the Transformer in the prediction process as the decoder input.

[0017] Furthermore, the encoder contains six identical layers, and the processing process of each layer is: A1, multi-head self-attention calculation.

[0018] First, the input matrix of the current layer Generate queries through three different linear transformations ,key Sum :

[0019] , , ;

[0020] in, , and is a learnable parameter matrix; is the input matrix of the current encoder block, i.e., the output matrix of the previous encoder block.

[0021] Then, query ,key Sum Divide into multiple attention heads, and each attention head independently calculates the attention score matrix:

[0022] ;

[0023] in, Indicates Attention head, For the The query matrix of the attention heads, For the The key matrix of the attention heads, For the The attention score matrix of the attention heads, is the dimension of the key vector.

[0024] Next, through The function converts the attention score matrix into The attention weight matrix of the attention heads :

[0025] ;

[0026] Finally, the value matrix is ​​weighted and summed using the attention weight matrix:

[0027] ;

[0028] in, Indicates The output matrix of the attention head is , is the sequence length, is the dimension of the value vector, Indicates The value matrix of the attention heads.

[0029] Each attention head independently learns different attention modes, which enhances the model's ability to capture the correlation of features in different dimensions. The results of all attention heads are concatenated and linearly transformed to calculate the multi-head self-attention. :

[0030] ;

[0031] in, Represents a splicing operation, Indicates The output matrix of the attention head is is a learnable output projection matrix used to map the concatenated multi-head attention results back to the original feature dimension. , Indicates Attention head, is the hidden dimension of the encoder.

[0032] A2. Feedforward neural network processing. The output of multi-head self-attention passes through two layers of feedforward neural network:

[0033] ;

[0034] in, is the input feature, is the weight matrix of the first layer linear transformation, is the bias vector of the first layer, is the weight matrix of the second layer linear transformation, is the bias vector of the second layer; is the ReLU activation function, is the output of the feedforward neural network.

[0035] Each encoder block contains two sublayers: the first sublayer is a multi-head self-attention mechanism, and the second sublayer is a feed-forward neural network for position perception. The output of each sublayer is processed with residual connections and layer normalization:

[0036] ;

[0037] in, Represents the processing function of the current sublayer, Represents the layer normalization operation.

[0038] Furthermore, the decoder includes six identical processing layers, each of which includes three sublayers. The processing process of each sublayer is as follows: First sublayer: multi-head self-attention mechanism with masking. By constructing a masking matrix, it is ensured that the fracturing optimization proxy model can only use the information before the current position when predicting. The construction rule of the masking matrix is: when considering the position in the sequence and When: If , the mask value is 0; if , the masked value is negative infinity.

[0039] Second sublayer: Encoder-decoder attention mechanism. Use the output of the decoder as the query , using the encoder's output as the key Sum , calculate attention : .

[0040] The third sublayer: Position-aware feedforward neural network: The structure is the same as the feedforward neural network in the encoder. In the output layer, the output of the decoder passes through a linear layer and Layer, get the predicted output for each time step:

[0041] ;

[0042] in, represents the distribution of the model's prediction results for single time step production, represents the activation function, represents the output matrix of the last layer of the decoder, represents the weight matrix of the output layer used to map model features to the prediction target space, Represents the bias vector of the output layer used to provide a base offset for predictions.

[0043] Furthermore, in step S3, the nonlinear constrained optimization problem for cumulative oil production and economic net present value is defined as:

[0044] ;

[0045] in, is the standard minimization function, is the cumulative oil production, represents the decision variable vector, including segment cluster parameters and fracturing operation parameters, is the total number of optimizable parameters, represents the economic net present value, is the negative form of the economic net present value target, It is the negative value form of the cumulative oil production target.

[0046] Defining optimization variables :

[0047] ;

[0048] in, For single-stage liquid volume, For single stage displacement, is the cluster spacing, is the fracture conductivity.

[0049] Determine the range of each optimization variable based on the actual data distribution of the mine:

[0050] ;

[0051] ;

[0052] ;

[0053] ;

[0054] in, and are the minimum and maximum values ​​of the single-stage liquid volume, and are the minimum and maximum values ​​of the single-stage displacement, and are the minimum and maximum values ​​of cluster spacing, respectively. and are the minimum and maximum values ​​of fracture conductivity, respectively.

[0055] Further, in step S3, the traditional NSGA-II multi-objective optimization algorithm includes the following steps: C1, non-dominated sorting.

[0056] In multi-objective optimization, define the set of candidate solutions :

[0057] ;

[0058] Objective function vector :

[0059] ;

[0060] in, Indicates candidate solutions, including optimization variable combinations, .

[0061] If and only if:

[0062] ;

[0063] The solution Dominate , which forms the basis of non-dominated hierarchical sorting; among them, Index the candidate solutions. , and All are objective function indices.

[0064] C2. Crowding calculation. For solutions with the same non-dominated level, after sorting the two objectives according to the values ​​of each objective function, define The crowding distance is:

[0065] ;

[0066] in, Indicates the sorted The solution The objective function value, Indicates the sorted The solution The objective function value, To resolve the problem The distribution density in its neighborhood, express The first solution The maximum value of the objective function, express The first solution The minimum value of the objective function.

[0067] C3. Basic optimization process. First, the initial population is generated through initialization, and the current population is non-dominated sorted through iterative optimization, the individual crowding is calculated, and offspring are generated through selection, crossover and mutation. Finally, the parent and offspring are merged, and high-quality individuals are selected to form a new population.

[0068] Through the above methods, the traditional NSGA-II multi-objective optimization algorithm simultaneously optimizes the two objectives of cumulative oil production and economic net present value, obtains a series of evenly distributed optimal solutions in the objective space, and provides multiple optional trade-off solutions for the optimization of fracturing design schemes.

[0069] Furthermore, in step S3, the action space of the capacity optimization agent and the economic optimization agent is Both are defined as parameter adjustment schemes:

[0070] ;

[0071] in, Indicates increasing the parameter value. Indicates reducing the parameter value. Indicates remains unchanged.

[0072] Reward function for the capacity optimization agent and the reward function of the economic optimization agent They are:

[0073] ;

[0074] ;

[0075] in, Represents the output change value, Indicates the change in economic net present value and the synergistic optimization coefficient Control the sensitivity of capacity optimization to economic indicators, Control the sensitivity of economic optimization to yield indicators, and To achieve differentiated collaborative strategies.

[0076] Furthermore, in step S3, the dual-agent fusion process is as follows: in the dual-agent initialization stage, the Monte Carlo method is first used to sample multiple groups of fracturing design parameter combinations in the parameter space, the cumulative oil production and economic net present value of each group of parameters are evaluated, and the experience replay pools of the two agents are initialized.

[0077] After completing the parameter initialization, in each iteration of the traditional NSGA-II multi-objective optimization algorithm, the two agents select actions according to the current state, execute the selected actions, and obtain a new combination of fracturing parameters:

[0078] ;

[0079] in, Indicates the current state. Indicates optional actions, i.e. parameter adjustment schemes; Indicates optional actions Select The action with the largest value; The optimal parameter adjustment scheme selected for the current moment; express Function, indicating that in state Next select action Expected return rating.

[0080] After obtaining the new fracturing parameter combination, the cumulative oil production results calculated by the relevant numerical simulation software and the economic net present value calculated by the economic evaluation model in step S1 are used to calculate the reward value and update The value of the function:

[0081] ;

[0082] in, is the learning rate, is the discount factor, For instant rewards, is the expected return of the current state-action pair; is the optimal expected reward for the next state-action pair.

[0083] The optimized solutions of the capacity optimization agent and the economic optimization agent are added to the population, and the traditional NSGA-II multi-objective optimization algorithm is used for non-dominated sorting and congestion calculation, and high-quality individuals are selected to build a new population; the capacity optimization agent and the economic optimization agent independently optimize their goals through their respective reward functions, forming a benign competition in the process of parameter adjustment; through the collaborative optimization coefficient in the reward function, the two agents consider each other's goals while optimizing their own goals, thereby achieving collaborative optimization of capacity and economy.

[0084] Compared with the prior art, the beneficial technical effects of the present invention are: (1) The capacity forecasting model established by the present invention is based on the Transformer architecture, and combined with the capacity simulation results based on numerical simulation and the economic evaluation method based on economic net present value, it has higher prediction accuracy and shorter program running time than the traditional model. The efficient and accurate capacity forecasting method combined with the economic evaluation model can realize the dynamic production evaluation of capacity and economy in multiple dimensions, while saving a lot of numerical simulation program running time, and further improving the algorithm efficiency while ensuring the calculation accuracy.

[0085] (2) The multi-objective fracturing optimization design method based on multi-agent reinforcement learning established in the present invention can further optimize the original fracturing optimization design scheme. By integrating multi-agent reinforcement learning, a dual agent for maximizing production and optimizing economic net present value is constructed. The fracturing parameters are optimized through the interaction and competition between the production capacity optimization agent and the economic optimization agent. The proposed production capacity-economic fracturing optimization design scheme can achieve the coordinated optimization of post-fracturing production capacity and economic net present value, effectively increase crude oil production while controlling fracturing costs, and has broad prospects for large-scale application. BRIEF DESCRIPTION OF THE DRAWINGS

[0086] Figure 1 It is the residual distribution of the capacity forecasting model during the training iteration process.

[0087] Figure 2 It is the prediction result of the production capacity prediction model in Example 1 for a horizontal well in a certain study area in the previous three years.

[0088] Figure 3 It is the Pareto frontier and optimization results in the process of optimizing fracturing parameters based on cumulative oil production and economic net present value.

[0089] Figure 4 It is the designed liquid volume in each fracturing stage for different fracturing design schemes.

[0090] Figure 5 It is the designed sand amount in each fracturing section for different fracturing design schemes.

[0091] Figure 6 It is the designed displacement of each fracturing section for different fracturing design schemes.

[0092] Figure 7 It is the cumulative oil production of different fracturing design schemes.

[0093] Figure 8 is the economic net present value of different fracturing design options. DETAILED DESCRIPTION

[0094] Example 1: This example takes a horizontal well in a certain study area of ​​an oil field as an example to introduce in detail the multi-objective fracturing optimization design method for shale reservoirs based on reinforcement learning.

[0095] The method comprises the following steps: S1, simulating the capacity calculation results by numerical simulation software, and establishing a training sample set of the capacity prediction model based on the Monte Carlo sampling method in combination with the constructed economic evaluation model.

[0096] In the traditional fracturing parameter optimization design calculation process, a large number of numerical simulation schemes need to be run, which makes the calculation cost too high. In order to further improve the calculation efficiency under the condition of ensuring the calculation accuracy, it is necessary to build a capacity prediction model based on machine learning algorithm. The establishment of the capacity prediction model first needs to rely on the construction of the training data sample set. In the present invention, the training data samples are mainly derived from the production data and economic net present value calculation results generated by the numerical simulation software.

[0097] (1) Establish an economic evaluation model. The net present value (NPV) method is an important method for evaluating the economic benefits of an investment project. This method determines the feasibility of an investment by calculating the difference between the present value of the net cash flow expected to be generated by the project and the initial investment cost. If the net present value is positive, it means that the expected benefits of the project exceed the cost and the investment is feasible. The higher the net present value, the better the economic benefits. On the contrary, if the net present value is negative, it indicates that the investment is not advisable. In order to find the critical values ​​of the key parameters that affect the economic and effective development of shale reservoir fracturing, the calculation function of the economic net present value (and economic evaluation model) of shale reservoir development is:

[0098] ;

[0099] in, represents the economic net present value of shale reservoir development, is the total number of time steps, For the The length of the time step, is the discount rate, and are the numbers of production wells and injection wells drilled in the reservoir, and are the sales prices of oil and gas, respectively. is the unit cost of produced water purification, , and Respectively The production horizontal well The average oil production, average gas production and average water production at each time step, For the The horizontal well injection The amount of water injected per time step, and represent drilling cost and completion cost respectively.

[0100] The cost function is affected by many factors, including operating cost, material cost, equipment investment, personnel cost, etc. In order to simplify the optimization process, only the cost items related to the optimization variables are considered, and other cost items are regarded as constants and do not participate in the optimization process.

[0101] The drilling cost is defined as a regression function of the lateral length of the horizontal well:

[0102] ;

[0103] in, Represents the length of the horizontal section of a horizontal well.

[0104] The completion cost consists of the cost of purchasing raw materials (such as fracturing fluid or proppant) and the cost of fracturing construction. The completion cost of slickwater fracturing is defined as the fracture volume (fracture width Crack height The linear regression function of crack length is:

[0105] ;

[0106] in, Represents the fracture volume of each fracturing stage.

[0107] In the above formula, among other constant cost items, the discount rate is defined as 0.1, the unit cost of produced water purification is 5 US dollars per barrel, other economic parameters are assigned according to the actual application of the mine, and the crude oil price is the annual average of WTI crude oil price in 2023-2024, which is about 78 US dollars per barrel. Considering that one cubic meter of crude oil is approximately equal to 7 barrels, combined with the average exchange rate of the US dollar to the RMB in the past year of about 1:7.05, it can be inferred that the approximate market price of one cubic meter of crude oil is 3,850 yuan, and the approximate market price of one cubic meter of gas can be calculated by the same logic as 5.3 yuan. Other economic costs are calculated by substituting the numerical simulation results into the existing cost function.

[0108] (2) Construction of production capacity simulation and economic evaluation sample set. The oil, gas, water production and production capacity simulation results of horizontal wells in the above economic evaluation model are all derived from the simulation results of numerical simulation software. The modeling can be completed using conventional commercial numerical simulation software, and the production dynamic data can be output. The production capacity prediction model training sample set is constructed in combination with the calculation results of the economic evaluation model to provide a data basis for the construction of the production capacity prediction model. In order to systematically study the influence of various parameters on the development effect of shale oil reservoirs, the Monte Carlo sampling method is used to construct a parameter combination sample set. Monte Carlo sampling is a random sampling method based on probability statistics theory. Its core idea is to obtain a sample set that can reflect the uncertainty of the system by constructing a suitable probability distribution model and performing random sampling in a given parameter space. A total of 9 key influencing factors are selected, including reservoir physical parameters (matrix porosity, oil saturation, matrix permeability), fracturing construction parameters (single-stage sand volume, single-stage liquid volume, single-stage displacement) and engineering parameters (cluster spacing, hydraulic fracture conductivity, horizontal section length). First, the probability distribution model of each parameter needs to be determined.

[0109] Based on a large amount of field data analysis and geological statistical characteristics, reservoir physical property parameters (matrix porosity, oil saturation, matrix permeability) usually present a log-normal distribution. , its lognormal distribution is expressed as:

[0110] ;

[0111] in, is the logarithmic mean, is the logarithmic standard deviation, is a random variable that follows a standard normal distribution.

[0112] Fracturing operation parameters (Sand volume, liquid volume, displacement) are usually assumed to obey normal distribution based on engineering practice:

[0113] ;

[0114] in, is the mean, is the standard deviation, is a random variable that follows a standard normal distribution.

[0115] Engineering parameters such as cluster spacing and horizontal segment length It can be assumed that it follows a uniform distribution within its physical constraints:

[0116] ;

[0117] in, and They are engineering parameters The minimum and maximum values ​​of is a uniform random number in the interval [0,1].

[0118] In the actual sampling process, the physical constraints of the parameters also need to be considered. For example, the porosity value must be between 0 and 1, and the permeability must be greater than 0. Therefore, after the samples are generated, validity verification is required to eliminate sample combinations that do not meet the physical constraints, and finally obtain valid parameter combination samples. These samples not only maintain the statistical characteristics of the parameters, but also meet the physical constraints, providing a reliable input data basis for subsequent numerical simulations.

[0119] S2. Construct a capacity prediction model based on the Transformer architecture. The model input includes but is not limited to discrete data of segment cluster design scheme and fracturing construction parameters and continuous dynamic data including but not limited to production data and bottom hole pressure time series generated by the numerical simulation software during operation. The numerical simulation software outputs dynamic capacity prediction data and calculates the economic net present value through the constructed economic evaluation model.

[0120] The Transformer architecture mainly consists of encoder and decoder layers. The encoder is used to extract high-level features from the input data. It consists of several encoder blocks, each of which contains a multi-head attention mechanism and a position-aware feedforward neural network to facilitate the extraction of complex features. The encoder input sequence is a collection of various data type features including discrete and multivariate continuous variables. The discrete variables mainly include segment cluster design schemes and fracturing construction parameters, etc. The input continuous dynamic variables mainly include time series data information such as daily oil production, daily water production, daily liquid production, and bottom hole pressure generated by the numerical simulation model during operation. The decoder input sequence is divided into two parts, one of which is the result generated by running the reservoir numerical simulator during the model training process; the other part is the output of the previous time step of the Transformer during the prediction process as the decoder input. The decoder part will eventually generate an output sequence of daily oil production, and finally calculate the economic net present value result through the constructed economic measurement method.

[0121] For the input static features and dynamic time series data, standardization is first performed to normalize all parameters to the [0,1] interval to eliminate the dimension effect. The calculation formula is:

[0122] ;

[0123] in, is the original input data value, is the minimum value of the feature in the data set, is the maximum value of the feature in the data set, is the standardized data value, ranging from [0,1].

[0124] In order to enable the model to recognize the temporal characteristics of the data, Calculate the positional encoding:

[0125] ;

[0126] ;

[0127] in, Indicates the time step position in the sequence (e.g., day 1 is 1, day 2 is 2); Indicates the position of the feature dimension (from 0 to ); is the characteristic dimension of the model; Encode values ​​for positions in even dimensions; is the position encoding value on the odd dimension; 10000 is the scaling factor used to control the periodic changes in different dimensions.

[0128] The positional encoding is added directly to the normalized input data:

[0129] ;

[0130] in, A matrix describing the input to the model.

[0131] Specifically, the encoder contains six identical layers, and the processing process of each layer is: (1) multi-head self-attention calculation.

[0132] First, the input matrix of the current layer Generate queries through three different linear transformations ,key Sum :

[0133] , , ;

[0134] in, , and is a learnable parameter matrix; is the input matrix of the current encoder block, i.e., the output matrix of the previous encoder block.

[0135] Then, query ,key Sum Divide into multiple attention heads (take 8 heads as an example), and each attention head independently calculates the attention score matrix:

[0136] ;

[0137] in, Indicates Attention heads, ranging from 1 to 8; For the The query matrix of the attention heads, For the The key matrix of the attention heads, For the The attention score matrix of the attention heads, is the dimension of the key vector.

[0138] Next, through The function converts the attention score matrix into The attention weight matrix of each head :

[0139] ;

[0140] Finally, the value matrix is ​​weighted and summed using the attention weight matrix:

[0141] ;

[0142] in, Indicates The output matrix of the attention head is , is the sequence length, is the dimension of the value vector, Indicates The value matrix of the attention heads.

[0143] Each attention head independently learns different attention modes, which enhances the model's ability to capture the correlation of features in different dimensions. The results of all attention heads are concatenated and linearly transformed to calculate the multi-head self-attention. :

[0144] ;

[0145] in, Represents a splicing operation, Indicates The output matrix of the attention head is is a learnable output projection matrix used to map the concatenated multi-head attention results back to the original feature dimension. , is the hidden dimension of the encoder.

[0146] (2) Feedforward neural network processing. The output of multi-head self-attention passes through two layers of feedforward neural network:

[0147] ;

[0148] in, is the input feature, i.e. the output of multi-head attention; is the weight matrix of the first layer linear transformation, is the bias vector of the first layer, is the weight matrix of the second layer linear transformation, is the bias vector of the second layer; For the ReLU activation function, set negative values ​​to 0; is the output of the feedforward neural network.

[0149] Each encoder block contains two sublayers: the first sublayer is a multi-head self-attention mechanism, and the second sublayer is a feed-forward neural network for position perception. The output of each sublayer is processed with residual connections and layer normalization:

[0150] ;

[0151] in, Represents the processing function of the current sublayer, Represents the layer normalization operation.

[0152] The decoder also contains six identical processing layers, each of which contains three sub-layers. The processing process of each sub-layer is as follows.

[0153] First sublayer: multi-head self-attention mechanism with masking.

[0154] By constructing a masking matrix, we ensure that the capacity forecasting model can only use the information before the current position when forecasting. The construction rule of the masking matrix is: when considering the position in the sequence and When: If (Allows attention to the current location and historical information), the mask value is 0; if (shield future position information), the masking value is negative infinity.

[0155] Second sublayer: encoder-decoder attention mechanism.

[0156] Use the decoder's output as a query , using the encoder's output as the key Sum , calculate attention : .

[0157] The third sublayer: Position-aware feedforward neural network: has the same structure as the feedforward neural network in the encoder.

[0158] In the output layer processing, the output of the decoder passes through a linear layer and Layer, get the predicted output for each time step:

[0159] ;

[0160] in, represents the distribution of the model's prediction results for single time step production, represents the activation function, represents the output matrix of the last layer of the decoder, represents the weight matrix of the output layer used to map model features to the prediction target space, Represents the bias vector of the output layer used to provide a base offset for predictions.

[0161] In practical applications, a new static parameter combination is first input, and the model predicts step by step in an autoregressive manner, where the first time step uses the initial conditions, and each subsequent time step uses the previous prediction results as input, and finally outputs a complete production forecast sequence. Through this design, the model can capture the complex nonlinear relationship between parameters, learn the long-term dependence of production dynamics, consider the influence of historical information when predicting, and achieve rapid production forecasting. After completing the capacity forecasting task, the economic net present value can be calculated by the proposed economic evaluation model.

[0162] S3. Optimization objective functions are constructed for cumulative oil production and economic net present value respectively. Based on the traditional NSGA-II multi-objective optimization algorithm, a dual-agent reinforcement learning algorithm including a production capacity optimization agent and an economic optimization agent is introduced. The production capacity optimization agent and the economic optimization agent jointly optimize the fracturing design parameters through competition and collaboration, and a fracturing optimization agent model is established.

[0163] (1) Optimization objective function and constraints. The nonlinear constrained optimization problem for cumulative oil production and economic net present value is defined as:

[0164] ;

[0165] in, is the standard minimization function, is the cumulative oil production, represents the decision variable vector, including segment cluster parameters and fracturing operation parameters, is the total number of optimizable parameters, represents the economic net present value, is the negative form of the economic net present value target, It is the negative value form of the cumulative oil production target.

[0166] Defining optimization variables :

[0167] ;

[0168] in, For single-stage liquid volume, For single stage displacement, is the cluster spacing, is the fracture conductivity.

[0169] Determine the range of each optimization variable based on the actual data distribution of the mine:

[0170] ;

[0171] ;

[0172] ;

[0173] ;

[0174] in, and are the minimum and maximum values ​​of the single-stage liquid volume, and are the minimum and maximum values ​​of the single-stage displacement, and are the minimum and maximum values ​​of cluster spacing, respectively. and are the minimum and maximum values ​​of fracture conductivity, respectively.

[0175] (2) Construction of multi-objective optimization algorithm. Since cumulative oil production and economic net present value belong to different optimization objectives and there is a mutually restrictive relationship between them, the use of multi-objective optimization algorithm to optimize the fracturing design scheme is considered. According to the optimization variables and their constraints proposed above, NSGA-II (non-dominated sorting genetic algorithm II) improved based on multi-agent reinforcement learning is used to carry out intelligent optimization of fracturing design scheme. The traditional NSGA-II algorithm is a multi-objective optimization method based on genetic algorithm. It searches for the optimal solution by simulating natural selection and genetic mechanisms. It mainly includes the following steps: C1, non-dominated sorting.

[0176] In multi-objective optimization, define the set of candidate solutions :

[0177] ;

[0178] Objective function vector :

[0179] ;

[0180] in, Indicates candidate solutions, including optimization variable combinations, .

[0181] If and only if:

[0182] ;

[0183] The solution Dominate , which forms the basis of non-dominated hierarchical sorting; among them, Index the candidate solutions. , and All are objective function indices.

[0184] C2. Crowding calculation. For solutions with the same non-dominated level, after sorting the two objectives according to the values ​​of each objective function, define The crowding distance is:

[0185] ;

[0186] in, Indicates the sorted The solution The objective function value, Indicates the sorted The solution The objective function value, To resolve the problem The distribution density in its neighborhood, express The first solution The maximum value of the objective function, express The first solution The minimum value of the objective function.

[0187] C3. Basic optimization process. First, the initial population is generated through initialization, and the current population is non-dominated sorted through iterative optimization, the individual crowding is calculated, and offspring are generated through selection, crossover and mutation. Finally, the parent and offspring are merged, and high-quality individuals are selected to form a new population.

[0188] Through the above methods, the traditional NSGA-II multi-objective optimization algorithm can simultaneously optimize the two objectives of cumulative oil production and economic net present value, obtain a series of evenly distributed optimal solutions in the objective space, and provide multiple optional trade-off solutions for the optimization of fracturing design schemes.

[0189] Based on the traditional NSGA-II algorithm, this embodiment introduces a dual-agent reinforcement learning framework, which includes two agents: a production optimization agent and an economic optimization agent. The two agents jointly optimize the fracturing design parameters through competition and collaboration.

[0190] First, the action space of the capacity optimization agent and the economic optimization agent Both are defined as parameter adjustment schemes:

[0191] ;

[0192] in, Indicates increasing the parameter value. Indicates reducing the parameter value. Indicates remains unchanged.

[0193] Reward function for the capacity optimization agent and the reward function of the economic optimization agent They are:

[0194] ;

[0195] ;

[0196] in, Represents the output change value, Indicates the change in economic net present value and the synergistic optimization coefficient Control the sensitivity of capacity optimization to economic indicators, Control the sensitivity of economic optimization to yield indicators, and To achieve differentiated collaborative strategies.

[0197] Secondly, the dual-agent fusion process is as follows: in the dual-agent initialization stage, the Monte Carlo method is first used to sample multiple groups of fracturing design parameter combinations in the parameter space, the cumulative oil production and economic net present value of each group of parameters are evaluated, and the experience replay pools of the two agents are initialized.

[0198] After completing the parameter initialization, in each iteration of the traditional NSGA-II multi-objective optimization algorithm, the two agents select actions according to the current state, execute the selected actions, and obtain a new combination of fracturing parameters:

[0199] ;

[0200] in, Indicates the current state. Indicates optional actions, i.e. parameter adjustment schemes; Indicates optional actions Select The action with the largest value; The optimal parameter adjustment scheme selected for the current moment; express Function, indicating that in state Next select action Expected return rating.

[0201] After obtaining the new fracturing parameter combination, the cumulative oil production results calculated by the relevant numerical simulation software and the economic net present value calculated by the economic evaluation model in step S1 are used to calculate the reward value and update The value of the function:

[0202] ;

[0203] in, is the learning rate, is the discount factor, For instant rewards, is the expected return of the current state-action pair; is the optimal expected reward for the next state-action pair.

[0204] The optimized solutions of the capacity optimization agent and the economic optimization agent are added to the population, and the traditional NSGA-II multi-objective optimization algorithm is used for non-dominated sorting and congestion calculation, and high-quality individuals are selected to build a new population; the capacity optimization agent and the economic optimization agent independently optimize their goals through their respective reward functions, forming a benign competition in the process of parameter adjustment; through the collaborative optimization coefficient in the reward function, the two agents consider each other's goals while optimizing their own goals, thereby achieving collaborative optimization of capacity and economy.

[0205] S4. By optimizing the Pareto frontier of cumulative oil production and economic net present value multi-objective optimization, a specific fracturing parameter optimization scheme is selected in combination with actual production and development needs.

[0206] In this embodiment, 500 sets of parameter combinations are extracted according to the Monte Carlo sampling algorithm, and the statistical results of the sampling range of each parameter are shown in Table 1. The range of each parameter is determined according to the actual mine data distribution in the study area, and a total of 500 sets of experimental schemes are generated. The cumulative oil production changes and economic net present values ​​corresponding to each parameter scheme for 15 years are calculated based on the numerical simulation model. In addition, other time series data generated during the numerical simulation operation, such as daily water production, daily liquid production, and bottom hole pressure, are also retained and used in the production capacity prediction model training process.

[0207] Table 1 Statistics of sampling parameter ranges used in capacity forecasting models

[0208]

[0209] After constructing the training data set, the capacity forecasting model based on the Transformer algorithm can be constructed. In order to optimize the forecasting performance of the machine learning algorithm, the algorithm hyperparameters need to be adjusted according to the characteristics of the data set. Stratified five-fold cross validation is used to evaluate the generalization ability of the proxy model. The training data set is randomly divided into five sub-samples of equal number. In each training iteration, one sub-sample is selected as the validation set, and the remaining four sub-samples are used to train the model. The average accuracy of the five sub-samples is calculated to evaluate the model prediction performance. The grid search method is combined to determine the optimal hyperparameter range of the training model. The optimization range and optimal hyperparameter value of the hyperparameter are shown in Table 2.

[0210] Table 2 Hyperparameter optimization results

[0211]

[0212] A total of 600 iterations were performed during the training of the Transformer algorithm. The root mean square error (RMSE) was used as the criterion for evaluating the model performance. The distribution of the average residual of the five-fold cross-validation with the iteration process was as follows: Figure 1 The results show that as the number of training iterations increases, the prediction residual decreases significantly and eventually converges to a lower value, indicating that the prediction accuracy of the model is effectively improved during the training process.

[0213] Based on the Transformer architecture capacity prediction model, the program running time and the average root mean square error of the five-fold cross validation during the training process are statistically analyzed. The results are shown in Table 3. The results show that the capacity prediction model not only has high model accuracy, but also has a short program running time. This dual advantage highlights the superiority of the capacity prediction model in handling capacity prediction tasks.

[0214] Table 3 Training time and root mean square error results of capacity prediction model

[0215]

[0216] The trained capacity prediction model is used to predict the production of a horizontal well in a certain study area in the first three years. Figure 2 By observing the calculation results, it can be seen that the constructed capacity forecasting model can well describe the post-compression production dynamics, reduce the response delay to changes in output, increase the sensitivity to sudden changes in data, and more effectively handle the dynamics of output data forecasting, thereby improving the accuracy of capacity forecasting.

[0217] After completing the capacity forecasting task, the economic net present value can be calculated by the proposed economic evaluation model. Taking the existing fracturing optimization design scheme as the initial scheme, a fracturing optimization agent model is constructed to perform multi-objective optimization of production and economic net present value, and a series of optimization design schemes are obtained.

[0218] In the process of fracturing optimization design, it is necessary to first clarify the fracturing optimization parameters and their ranges. The statistical range of fracturing optimization design parameters is shown in Table 4, where the range of fracturing optimization design parameters is adjusted based on the existing fracturing design scheme.

[0219] Table 4 Parameter ranges for fracturing optimization design

[0220]

[0221] According to the existing fracturing design parameter range, 50 groups of fracturing construction parameter combinations were sampled using the Monte Carlo algorithm, and the corresponding cumulative oil production and economic net present value under 15-year depletion development conditions were calculated based on the fracturing optimization proxy model and the economic evaluation model. The results were substituted into the NSGA-II algorithm as input data. The parameters of the improved NSGA-II algorithm are as follows: the crossover coefficient is 0.65, the coefficient of variation is 0.08, the population size is 50, and the genetic generation is 10.

[0222] The optimization process of fracturing parameters based on cumulative oil production and economic net present value is as follows: Figure 3 As shown in the figure, the larger dots represent the final Pareto optimization frontier after 10 genetic iterations, and the smaller dots show all the solutions in the optimization process. The solutions on the Pareto optimization frontier are all the solutions with higher cumulative oil production and economic net present value in actual applications. In actual applications, the solutions with the best production capacity and the best economy are often selected as specific selection solutions. Therefore, in the optimization process, all the solutions on the Pareto optimization frontier corresponding to the highest cumulative oil production and the highest economic net present value are marked separately with the largest dots as the solutions with the best production capacity and the best economy.

[0223] Compared with the original fracturing layout design scheme with almost evenly distributed fracturing clusters, the productivity layout optimization design scheme designed denser fracturing clusters in some areas, thereby maximizing the post-fracturing production; while the economic layout optimization design scheme adopted a larger average cluster spacing at the measured depths of 2900-3300m and 3600-3700m, which saved the fracturing development cost to a certain extent.

[0224] The results of the single-stage liquid volume, sand volume and displacement of each fracturing stage designed by the original fracturing design scheme and the fracturing optimization design scheme are as follows: Figure 4 , Figure 5 and Figure 6As shown. The comparison results of fracturing construction parameters show that the fracturing optimization design scheme increases the design single-stage liquid volume, sand volume, and displacement in the corresponding high-density seam distribution area (the first to eighth sections), while reducing the corresponding fracturing construction parameter design in the area with lower seam density (the ninth to eleventh sections). The comparison results of the two fracturing optimization design schemes with the original fracturing design scheme show that the capacity optimization design scheme adopts a larger single-stage liquid volume, sand volume, and displacement, while the economic optimization design scheme has relatively lower design liquid volume, sand volume, and displacement, which minimizes the fracturing construction cost.

[0225] The corresponding cumulative oil production and economic net present value of the original fracturing design scheme, the production capacity layout optimization design scheme, and the economic layout optimization design scheme under the condition of 15 years of exhaustion development are calculated respectively. Figure 7 and Figure 8 As shown, according to the simulation results of different fracturing optimization design schemes, the capacity layout optimization design scheme can effectively improve the post-fracturing capacity within a longer production time to a certain extent due to the setting of a large number of fracturing segment clusters and large fracturing construction parameters; while the economic layout optimization design scheme has a slightly lower capacity than the capacity layout optimization design scheme, it can recover costs in a shorter time due to the fewer segment clusters and lower design single-stage liquid volume, sand volume, and displacement, and achieve higher economic net present value benefits within a longer production time. Compared with the original fracturing design scheme, the capacity layout optimization design scheme and the economic layout optimization design scheme have greatly improved in terms of capacity and economic net present value. In the actual production process, the corresponding layout optimization design scheme can be selected according to the needs of the site. The fracturing optimization design method proposed in the present invention can ensure high yield with fewer layout schemes, achieve accurate and effective fracturing, and is of great significance to reducing costs and increasing efficiency of shale oil.

[0226] Of course, the above description is not a limitation of the present invention, and the present invention is not limited to the above examples. Changes, modifications, additions or substitutions made by technicians in this technical field within the essential scope of the present invention should also fall within the protection scope of the present invention.

Claims

1. A multi-objective fracturing optimization design method for shale reservoirs based on reinforcement learning, characterized in that: The following steps are involved: S1. Use numerical simulation software to simulate the capacity calculation results, and combine the constructed economic evaluation model to establish a training sample set for the capacity prediction model based on the Monte Carlo sampling method; S2. Construct a capacity prediction model based on the Transformer architecture. The model input includes discrete data of segment cluster design scheme and fracturing operation parameters and continuous dynamic data including production data and bottom hole flow pressure time series generated when the numerical simulation software is running. The numerical simulation software outputs dynamic capacity prediction data and calculates the economic net present value through the constructed economic evaluation model. S3. Optimization objective functions are constructed for cumulative oil production and economic net present value respectively. Based on the traditional NSGA-II multi-objective optimization algorithm, a dual-agent reinforcement learning algorithm including a production capacity optimization agent and an economic optimization agent is introduced. The production capacity optimization agent and the economic optimization agent jointly optimize the fracturing design parameters through competition and collaboration, and a fracturing optimization agent model is established; In step S3, the dual-agent fusion process is as follows: in the dual-agent initialization stage, the Monte Carlo method is first used to sample multiple groups of fracturing design parameter combinations in the parameter space, the cumulative oil production and economic net present value of each group of parameters are evaluated, and the experience replay pools of the two agents are initialized; After completing the parameter initialization, in each iteration of the traditional NSGA-II multi-objective optimization algorithm, the two agents select actions according to the current state, execute the selected actions, and obtain a new combination of fracturing parameters: ; in, Indicates the current state. Indicates optional actions, i.e. parameter adjustment schemes; Indicates optional actions Select The action with the largest value; The optimal parameter adjustment scheme selected for the current moment; express Function, indicating that in state Next select action Expected return score; After obtaining the new fracturing parameter combination, the cumulative oil production results calculated by the relevant numerical simulation software and the economic net present value calculated by the economic evaluation model in step S1 are used to calculate the reward value and update The value of the function: ; in, is the learning rate, is the discount factor, For instant rewards, is the expected return of the current state-action pair; is the optimal expected return for the next state-action pair; The optimized solutions of the capacity optimization agent and the economic optimization agent are added to the population, and the traditional NSGA-II multi-objective optimization algorithm is used to perform non-dominated sorting and crowding calculation, and high-quality individuals are selected to build a new population; the capacity optimization agent and the economic optimization agent independently optimize their goals through their respective reward functions, forming a healthy competition in the process of parameter adjustment; through the collaborative optimization coefficient in the reward function, the two agents can consider each other's goals while optimizing their own goals, realizing the collaborative optimization of capacity and economy; S4. By optimizing the Pareto frontier of cumulative oil production and economic net present value multi-objective optimization, a specific fracturing parameter optimization scheme is selected in combination with actual production and development needs.

2. The multi-objective fracturing optimization design method for shale reservoirs based on reinforcement learning according to claim 1, characterized in that: In step S1, the economic evaluation model is: ; in, represents the economic net present value of shale reservoir development, is the total number of time steps, For the The length of the time step, is the discount rate, and are the numbers of production wells and injection wells drilled in the reservoir, and are the sales prices of oil and gas, respectively. is the unit cost of produced water purification, , and Respectively The production horizontal well The average oil production, average gas production and average water production at each time step, For the The horizontal well injection The amount of water injected per time step, and represent drilling cost and completion cost respectively.

3. The multi-objective fracturing optimization design method for shale reservoirs based on reinforcement learning according to claim 2, characterized in that: The Monte Carlo sampling method constructs a probability distribution model and performs random sampling in a given parameter space to obtain a sample set that reflects the uncertainty of the system; First, the probability distribution model of reservoir physical parameters, fracturing construction parameters and engineering parameters is determined; then, in the actual sampling process, the physical constraints of each parameter are considered. After the samples are generated, the validity is verified and the sample combinations that do not meet the physical constraints are eliminated, finally obtaining the valid parameter combination samples.

4. The multi-objective fracturing optimization design method for shale reservoirs based on reinforcement learning according to claim 3 is characterized in that: The Transformer architecture consists of an encoder and a decoder. The encoder consists of several encoder blocks, each of which contains a multi-head self-attention mechanism and a position-aware feedforward neural network. The encoder input sequence is a collection of various data type features including discrete data and continuous dynamic data. The discrete data includes segment cluster design scheme and fracturing operation parameters. The continuous dynamic data includes daily oil production, daily water production, daily liquid production and bottom hole flow pressure time series data information generated by the numerical simulation model during operation. The decoder input sequence is divided into two parts. One part is the result generated by running the reservoir numerical simulator during the production capacity prediction model training process; the other part is the output of the previous time step of the Transformer in the prediction process as the decoder input.

5. The multi-objective fracturing optimization design method for shale reservoirs based on reinforcement learning according to claim 4, characterized in that: The encoder consists of six identical layers, and the processing process of each layer is: A1, multi-head self-attention calculation; First, the input matrix of the current layer Generate queries through three different linear transformations ,key Sum : , , ; in, , and is a learnable parameter matrix; is the input matrix of the current encoder block, i.e., the output matrix of the previous encoder block; Then, query ,key Sum Divide into multiple attention heads, and each attention head independently calculates the attention score matrix: ; in, Indicates Attention head, For the The query matrix of the attention heads, For the The key matrix of the attention heads, For the The attention score matrix of the attention heads, is the dimension of the key vector; Next, through The function converts the attention score matrix into The attention weight matrix of the attention heads : ; Finally, the value matrix is ​​weighted and summed using the attention weight matrix: ; in, Indicates The output matrix of the attention head is , Represented as a row number , the number of columns is A real matrix of is the sequence length, is the dimension of the value vector, Indicates The value matrix of the attention heads; Each attention head independently learns different attention modes, which enhances the model's ability to capture the correlation of features in different dimensions. The results of all attention heads are concatenated and linearly transformed to calculate the multi-head self-attention. : ; in, Represents a splicing operation, Indicates The output matrix of the attention head is is a learnable output projection matrix used to map the concatenated multi-head attention results back to the original feature dimension. , For the Attention head, is the hidden dimension of the encoder; A2, feed-forward neural network processing; The output of the multi-head self-attention passes through a two-layer feedforward neural network: ; in, is the input feature, is the weight matrix of the first layer linear transformation, is the bias vector of the first layer, is the weight matrix of the second layer linear transformation, is the bias vector of the second layer; is the ReLU activation function, is the output of the feedforward neural network; Each encoder block contains two sublayers: the first sublayer is a multi-head self-attention mechanism, and the second sublayer is a feed-forward neural network for position perception. The output of each sublayer is processed with residual connections and layer normalization: ; in, Represents the processing function of the current sublayer, Represents the layer normalization operation.

6. The multi-objective fracturing optimization design method for shale reservoirs based on reinforcement learning according to claim 5, characterized in that: The decoder comprises six identical processing layers, each of which comprises three sub-layers, and the processing process of each sub-layer is as follows: The first sub-layer: multi-head self-attention mechanism with masking; By constructing a masking matrix, we ensure that the capacity forecasting model can only use the information before the current position when forecasting. The construction rule of the masking matrix is: when considering the position in the sequence and When: If , the mask value is 0; if , the masked value is negative infinity; Second sublayer: encoder-decoder attention mechanism; Use the decoder's output as a query , using the encoder's output as the key Sum , calculate attention : ; The third sublayer: Position-aware feedforward neural network: has the same structure as the feedforward neural network in the encoder; In the output layer processing, the output of the decoder passes through a linear layer and Layer, get the predicted output for each time step: ; in, represents the distribution of the model's prediction results for single time step production, represents the activation function, represents the output matrix of the last layer of the decoder, represents the weight matrix of the output layer used to map model features to the prediction target space, Represents the bias vector of the output layer used to provide a base offset for predictions.

7. The multi-objective fracturing optimization design method for shale reservoirs based on reinforcement learning according to claim 6, characterized in that: In step S3, the nonlinear constrained optimization problem for cumulative oil production and economic net present value is defined as: ; in, is the standard minimization function, is the cumulative oil production, represents the decision variable vector, including segment cluster parameters and fracturing operation parameters, is the total number of optimizable parameters, represents the economic net present value, is the negative form of the economic net present value target, It is the negative form of the cumulative oil production target; Defining optimization variables : ; in, For single-stage liquid volume, For single stage displacement, is the cluster spacing, is the fracture conductivity; Determine the range of each optimization variable based on the actual data distribution of the mine: ; ; ; ; in, and are the minimum and maximum values ​​of the single-stage liquid volume, and are the minimum and maximum values ​​of the single-stage displacement, and are the minimum and maximum values ​​of cluster spacing, respectively. and are the minimum and maximum values ​​of fracture conductivity, respectively.

8. The multi-objective fracturing optimization design method for shale reservoirs based on reinforcement learning according to claim 7, characterized in that: In step S3, the traditional NSGA-II multi-objective optimization algorithm includes the following steps: C1, non-dominated sorting; In multi-objective optimization, define the set of candidate solutions : ; Objective function vector : ; in, Indicates candidate solutions, including optimization variable combinations, ; If and only if: ; The solution Dominate , which forms the basis of non-dominated hierarchical sorting; among them, Index the candidate solutions. , and All are objective function indexes; C2, congestion calculation; For solutions with the same non-dominated level, after sorting the two objectives according to the values ​​of each objective function, we define The crowding distance is: ; in, Indicates the sorted The solution The objective function value, Indicates the sorted The solution The objective function value, To resolve the problem The distribution density in its neighborhood is express The first solution The maximum value of the objective function, express The first solution The minimum value of the objective function; C3, basic optimization process; First, the initial population is generated through initialization, and the current population is non-dominated sorted through iterative optimization, the individual crowding is calculated, and offspring are generated through selection, crossover and mutation. Finally, the parent and offspring are merged, and high-quality individuals are selected to form a new population. Through the above methods, the traditional NSGA-II multi-objective optimization algorithm simultaneously optimizes the two objectives of cumulative oil production and economic net present value, obtains a series of evenly distributed optimal solutions in the objective space, and provides multiple optional trade-off solutions for the optimization of fracturing design schemes.

9. The shale reservoir multi-objective fracturing optimization design method based on reinforcement learning according to claim 8, characterized in that: In step S3, the action space of the capacity optimization agent and the economic optimization agent Both are defined as parameter adjustment schemes: ; in, Indicates increasing the parameter value. Indicates reducing the parameter value. Indicates remaining unchanged; Reward function for the capacity optimization agent and the reward function of the economic optimization agent They are: ; ; in, Represents the output change value, Indicates the change in economic net present value and the synergistic optimization coefficient Control the sensitivity of capacity optimization to economic indicators, Control the sensitivity of economic optimization to yield indicators, and To achieve differentiated collaborative strategies.

Citation Information

Patent Citations

  • Dense oil reservoir fracturing horizontal well optimization method based on self-adaptive agent model

    CN109236258A

  • Deep shale gas horizontal well fracturing differentiation design method based on logging data

    CN113090258A

  • Method and device for optimizing fracturing parameters of fracturing section of unconventional natural gas well and storage medium

    CN117634166A

  • Shale reservoir well factory horizontal well fracturing parameter layering multistage optimization method

    CN117574755A

  • Shale gas well fracturing parameter and well control parameter optimization method

    CN118536373A