A shale gas well EUR prediction method and system based on multi-modal information fusion

CN122528110APending Publication Date: 2026-08-07CHONGQING UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHONGQING UNIVERSITY OF SCIENCE AND TECHNOLOGY
Filing Date
2026-06-29
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0005]本发明的目的在于提供一种基于多模态信息融合的页岩气井EUR预测方法及系统,解决了现有技术中因未投产页岩气井缺乏生产历史数据而难以直接预测,且现有数据驱动方法忽略压裂施工步骤顺序依赖、裂缝网络几何拓扑及储层空间非均质性,以及缺乏多模态信息有效耦合建模所导致的预测效率低、精度差的问题

Benefits of technology

[0016]This invention discloses a method and system for predicting the EUR (Earnings Regime) of shale gas wells based on multimodal information fusion. First, it acquires fracturing operation sequence data, fracture network data, and reservoir attribute point cloud data of the target shale gas well. Then, it performs targeted preprocessing on these three types of data: converting the fracturing operation sequence into a fracturing operation sequence input tensor, converting the fracture network into graph structure data containing node and edge features, and converting the three-dimensional reservoir attribute field into a point cloud input tensor with a fixed number of points. Next, it uses a sequence encoder to extract construction modal features reflecting the sequential dependencies of construction steps, a graph structure encoder to extract fracture network modal features reflecting fracture geometry and topological connectivity, and a point cloud encoder to extract reservoir attribute modal features reflecting reservoir spatial heterogeneity. Then, it inputs the three modal feature vectors into a cross-modal self-attention fusion network, and aggregates the modal information through learnable aggregated label vectors to obtain multimodal fusion features. Finally, it inputs the fusion features into a regression prediction network to output the predicted EUR value of the shale gas well. This invention addresses the challenge of lacking historical production data for non-produced shale gas wells, enabling EUR prediction without relying on dynamic production data. It also addresses the high cost and low efficiency of traditional numerical simulations by employing a deep learning surrogate model to significantly shorten prediction time, meeting the need for rapid selection across multiple wells and scenarios. Furthermore, it addresses the shortcomings of existing data-driven methods that simplify construction parameters to totals or averages, ignore irregular fracture network topologies, and reduce reservoir properties to scalars, thus losing spatial heterogeneity. Sequence encoding, graph encoding, and point cloud encoding are used to preserve the original structure and distribution information of each mode. Simultaneously, a cross-modal self-attention mechanism is utilized to explicitly learn the coupling relationship between fracturing operations, fracture networks, and reservoir properties, overcoming the difficulty of integrating heterogeneous information through simple feature splicing. Therefore, this invention comprehensively improves the accuracy, efficiency, and generalization stability of EUR prediction for non-produced shale gas wells, providing a reliable basis for fracturing scheme optimization and well location deployment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122528110A_ABST
    Figure CN122528110A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of oil and gas field development, and discloses a shale gas well ultimate recoverable gas (EUR) prediction method based on multi-modal information fusion, obtains fracturing operation sequence, fracture network and reservoir attribute point cloud data; respectively carries out sequence pretreatment, graph structure conversion and point cloud pretreatment; a sequence encoder extracts fracturing operation modal features, a graph structure encoder extracts fracture network modal features, and a point cloud encoder extracts reservoir attribute modal features; the three modal feature vectors are input into a cross-modal self-attention fusion network to obtain multi-modal fusion features; and a EUR prediction value is output through a regression prediction network. In the prediction stage, the EUR of an unproduced shale gas well can be quickly predicted without production history data of the well to be predicted, the time sequence characteristics of the fracturing operation sequence, the topological characteristics of the fracture network and the spatial distribution characteristics of the reservoir attribute point cloud are retained, the cross-modal self-attention mechanism is used to learn the coupling relationship of multi-source information, and the prediction accuracy and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of oil and gas field development technology, and in particular to a method and system for predicting EUR (Earnings Flow Rate) of shale gas wells based on multimodal information fusion. Background Technology

[0002] Shale gas, as an important unconventional natural gas resource, is characterized by tight reservoirs, strong heterogeneity, and reliance on volumetric fracturing for development. Ultimate recoverable gas (EUR) is a core indicator for evaluating the development potential of shale gas wells, the rationality of fracturing schemes, and the effectiveness of block production capacity construction. For shale gas wells that have been put into production and have a long production history, EUR prediction can be based on methods such as production decline curves, production dynamic analysis, or numerical simulation. However, for shale gas wells that have been drilled and completed or have fracturing designs but have not yet been put into production, traditional EUR prediction methods relying on production dynamics are difficult to apply directly due to the lack of continuous production history data. Existing EUR prediction for non-producing wells usually relies on numerical simulation methods, that is, by establishing geological models, fracturing fracture models, and shale gas reservoir production models, and conducting fracture propagation simulations and production numerical simulations to obtain cumulative gas production. This type of method can reflect physical processes such as fracturing, fracture diversion, and shale gas adsorption and desorption well. However, its modeling process is complex, it has many input parameters, and the simulation takes a long time. When it is necessary to quickly screen multiple un-produced wells or a large number of fracturing construction schemes, it is difficult to meet the efficiency requirements of engineering applications by carrying out complete numerical simulations one by one.

[0003] With the development of machine learning technology, data-driven methods for predicting shale gas well productivity are gradually being applied. Existing data-driven methods typically input scalar parameters such as total fluid volume, total sand volume, average displacement, reservoir porosity, and permeability into models such as random forests, support vector machines, or fully connected neural networks to establish a nonlinear mapping relationship between the input parameters and the EUR (Effective Reactor Parameter). However, such methods still have significant shortcomings: First, existing methods often treat fracturing parameters as totals or averages, making it difficult to retain the stage sequence information and engineering logic dependencies between different construction steps such as the pre-fracturing fluid stage, the proppant carrying stage, and the tailings stage during fracturing. Second, existing methods rarely directly utilize the geometric structure information of the fracture network formed after fracturing, making it difficult to reflect the control effect of irregular topological structures such as fracture length, fracture height, fracture orientation, fracture density, and fracture connectivity on EUR. Third, existing methods often simplify reservoir properties to scalar characteristics such as well-circumferential average porosity and average permeability, making it difficult to retain the spatial heterogeneity information of shale reservoirs, which has strong variations in both the planar and vertical directions. Fourth, there is a complex coupling relationship between the fracturing construction sequence, the fracture network structure, and the three-dimensional reservoir properties. Construction parameters determine the fracture propagation process, the fracture network changes the seepage channels, and reservoir properties constrain the gas supply capacity. All three together determine the EUR response, and existing simple feature splicing methods are unable to explicitly learn the interdependencies between different modes.

[0004] Therefore, how to construct a method that can rapidly and accurately predict EUR for unproduced shale gas wells by comprehensively utilizing multi-source heterogeneous information such as fracturing operation sequences, fracture networks, and reservoir attribute point clouds, while avoiding repeated complete numerical simulations, is an urgent problem to be solved in this field. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for predicting EUR of shale gas wells based on multimodal information fusion. This invention solves the problems of low prediction efficiency and poor accuracy caused by the lack of production history data for shale gas wells that have not yet been put into production, the neglect of the sequential dependence of fracturing construction steps, the geometric topology of fracture networks and the heterogeneity of reservoir space in existing data-driven methods, and the lack of effective coupling modeling of multimodal information.

[0006] To achieve the above objectives, this invention provides a shale gas well EUR prediction method based on multimodal information fusion, comprising the following steps: Acquire fracturing operation sequence data, fracture network data, and reservoir attribute point cloud data of shale gas wells; The fracturing operation sequence data is preprocessed to obtain the fracturing operation sequence input tensor; The crack network data is transformed into a graph structure to obtain crack network graph data; Point cloud preprocessing is performed on the reservoir attribute point cloud data to obtain the reservoir attribute point cloud input tensor; The fracturing operation sequence is input into a tensor input sequence encoder to obtain the operation mode feature vector; The crack network graph data is input into a graph structure encoder to obtain the crack network modal feature vector; The reservoir attribute point cloud is input into a tensor input point cloud encoder to obtain the reservoir attribute modal feature vector; The construction modal feature vector, the fracture network modal feature vector, and the reservoir attribute modal feature vector are input into a multimodal fusion network to obtain a multimodal fusion feature vector; The multimodal fusion feature vector is input into the regression prediction network to obtain the EUR prediction value for shale gas wells.

[0007] The fracturing operation sequence data includes multiple fracturing operation steps. Each fracturing operation step includes continuous parameters and discrete parameters. The continuous parameters include operation flow rate, injection fluid volume, and proppant concentration. The discrete parameters include fracturing fluid type and proppant type. The sequence preprocessing of the fracturing construction sequence data includes: standardizing the continuous parameters; class-indexing the discrete parameters and inputting the encoded discrete parameters into an embedding layer to obtain discrete parameter embedding vectors; concatenating the standardized continuous parameters with the discrete parameter embedding vectors and obtaining the step embedding vector for each construction step through a linear mapping layer; and performing padding processing on fracturing construction sequences of different lengths to generate sequence masks for identifying actual construction steps and filling steps.

[0008] The fracture network data includes spatial location data, geometric attribute data, and connectivity data of multiple fracture patches or fracture segments; the geometric attribute data includes at least one of fracture segment length, fracture height, fracture azimuth angle, fracture area, and fracture conductivity. The crack network data is transformed into a graph structure to obtain crack network graph data, including: converting the crack patches or crack segments into nodes in a graph structure; using the geometric properties of the crack patches or crack segments as node features; establishing edges in the graph structure based on the shared endpoint relationships, spatial proximity relationships, or flow connectivity relationships between the crack patches or crack segments; and using the distances, angles, height differences, or connectivity states between the crack patches or crack segments as edge features to obtain crack network graph data containing node features and edge features.

[0009] The reservoir attribute point cloud data includes multiple sampling points, each sampling point including spatial coordinates and reservoir attribute values; the reservoir attribute values ​​include at least one of permeability, porosity, gas saturation, geostress parameters, or rock mechanics parameters. Point cloud preprocessing is performed on the reservoir attribute point cloud data to obtain the reservoir attribute point cloud input tensor, including: sampling the farthest point of the reservoir attribute point cloud data to obtain a fixed number of point cloud sampling points; centering and scaling the spatial coordinates of the point cloud sampling points; standardizing the reservoir attribute values ​​of the point cloud sampling points; and concatenating the normalized spatial coordinates with the standardized reservoir attribute values ​​to obtain the reservoir attribute point cloud input tensor.

[0010] The sequence encoder includes a position encoding layer, a self-attention encoding layer, and a sequence pooling layer. The position encoding layer is used to superimpose the construction step sequence information into the construction step embedding vector. The self-attention encoding layer is used to learn the dependencies between different construction steps. The sequence pooling layer is used to convert the variable-length fracturing construction sequence into a fixed-dimensional construction modal feature vector.

[0011] The graph structure encoder is an edge-enhanced graph neural network encoder; the edge-enhanced graph neural network encoder aggregates node features and edge features through a message passing mechanism to obtain the crack network modal feature vector.

[0012] The point cloud encoder uses hierarchical sampling, local neighborhood construction, and local feature aggregation to extract local heterogeneous features and global spatial distribution features of reservoir attribute point clouds, thereby obtaining reservoir attribute modal feature vectors.

[0013] Specifically, the construction modal feature vector, the fracture network modal feature vector, and the reservoir attribute modal feature vector are input into a multimodal fusion network to obtain a multimodal fusion feature vector, which includes: Linear mapping is performed on the construction modal feature vector, the fracture network modal feature vector, and the reservoir attribute modal feature vector, respectively, so that the three modal feature vectors have the same feature dimension; Set a learnable aggregated tag vector; The learnable aggregated label vector, the linearly mapped construction modal feature vector, the linearly mapped fracture network modal feature vector, and the linearly mapped reservoir attribute modal feature vector are combined to form a modal feature sequence; The modal feature sequence is input into a cross-modal self-attention fusion network to obtain an updated learnable aggregated label vector; The updated learnable aggregated label vector is used as the multimodal fusion feature vector.

[0014] The step of inputting the multimodal fused feature vector into the regression prediction network to obtain the EUR prediction value for shale gas wells includes: The multimodal fusion feature vector is input into the regression prediction network to obtain the standardized EUR prediction value; the standardized EUR prediction value is then de-standardized to obtain the EUR prediction value of shale gas wells under the actual physical dimensions.

[0015] A shale gas well EUR prediction system based on multimodal information fusion includes: The data acquisition module is used to acquire fracturing operation sequence data, fracture network data, and reservoir attribute point cloud data of shale gas wells; The fracturing operation sequence processing module is connected to the data acquisition module and is used to receive the fracturing operation sequence data, perform sequence preprocessing, and output the fracturing operation sequence input tensor. A crack network processing module, connected to the data acquisition module, is used to receive the crack network data, perform graph structure conversion, and output crack network graph data. The point cloud processing module is connected to the data acquisition module and is used to receive the reservoir attribute point cloud data, perform point cloud preprocessing, and output the reservoir attribute point cloud input tensor. The construction feature extraction module is connected to the fracturing construction sequence processing module. It is used to receive the input tensor of the fracturing construction sequence and input it into the sequence encoder, and output the construction modal feature vector. A crack feature extraction module, connected to the crack network processing module, is used to receive the crack network graph data and input it into the graph structure encoder, and output the crack network modal feature vector. The attribute feature extraction module is connected to the point cloud processing module and is used to receive the reservoir attribute point cloud input tensor and input it into the point cloud encoder, and output the reservoir attribute modal feature vector. The multimodal fusion module is connected to the construction feature extraction module, the fracture feature extraction module, and the attribute feature extraction module, respectively, and is used to receive the construction modal feature vector, the fracture network modal feature vector, and the reservoir attribute modal feature vector, perform multimodal fusion, and output the multimodal fusion feature vector; The EUR prediction module, connected to the multimodal fusion module, is used to receive the multimodal fusion feature vector and input it into the regression prediction network, and output the EUR prediction value of shale gas wells.

[0016] This invention discloses a method and system for predicting the EUR (Earnings Regime) of shale gas wells based on multimodal information fusion. First, it acquires fracturing operation sequence data, fracture network data, and reservoir attribute point cloud data of the target shale gas well. Then, it performs targeted preprocessing on these three types of data: converting the fracturing operation sequence into a fracturing operation sequence input tensor, converting the fracture network into graph structure data containing node and edge features, and converting the three-dimensional reservoir attribute field into a point cloud input tensor with a fixed number of points. Next, it uses a sequence encoder to extract construction modal features reflecting the sequential dependencies of construction steps, a graph structure encoder to extract fracture network modal features reflecting fracture geometry and topological connectivity, and a point cloud encoder to extract reservoir attribute modal features reflecting reservoir spatial heterogeneity. Then, it inputs the three modal feature vectors into a cross-modal self-attention fusion network, and aggregates the modal information through learnable aggregated label vectors to obtain multimodal fusion features. Finally, it inputs the fusion features into a regression prediction network to output the predicted EUR value of the shale gas well. This invention addresses the challenge of lacking historical production data for non-produced shale gas wells, enabling EUR prediction without relying on dynamic production data. It also addresses the high cost and low efficiency of traditional numerical simulations by employing a deep learning surrogate model to significantly shorten prediction time, meeting the need for rapid selection across multiple wells and scenarios. Furthermore, it addresses the shortcomings of existing data-driven methods that simplify construction parameters to totals or averages, ignore irregular fracture network topologies, and reduce reservoir properties to scalars, thus losing spatial heterogeneity. Sequence encoding, graph encoding, and point cloud encoding are used to preserve the original structure and distribution information of each mode. Simultaneously, a cross-modal self-attention mechanism is utilized to explicitly learn the coupling relationship between fracturing operations, fracture networks, and reservoir properties, overcoming the difficulty of integrating heterogeneous information through simple feature splicing. Therefore, this invention comprehensively improves the accuracy, efficiency, and generalization stability of EUR prediction for non-produced shale gas wells, providing a reliable basis for fracturing scheme optimization and well location deployment. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0018] Figure 1 This is a flowchart of the steps of the EUR prediction method for shale gas wells based on multimodal information fusion according to the present invention.

[0019] Figure 2 This is a flowchart of the fracturing operation sequence feature extraction method of the present invention.

[0020] Figure 3 This is a flowchart of the crack network structure construction and feature extraction method of the present invention.

[0021] Figure 4 This is a flowchart of the reservoir attribute point cloud feature extraction method of the present invention.

[0022] Figure 5 This is a structural block diagram of the EUR prediction system for shale gas wells based on multimodal information fusion, as described in this invention.

[0023] In the figure: 1-Data acquisition module, 2-Fracturing construction sequence processing module, 3-Fracturing network processing module, 4-Point cloud processing module, 5-Construction feature extraction module, 6-Fracturing feature extraction module, 7-Attribute feature extraction module, 8-Multimodal fusion module, 9-EUR prediction module. Detailed Implementation

[0024] The embodiments of the present invention are described in detail below. Examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, but should not be construed as limiting the present invention.

[0025] First embodiment: Please refer to Figures 1 to 5 This invention provides a method for predicting the EUR of shale gas wells based on multimodal information fusion, comprising the following steps: S101: Acquire fracturing operation sequence data, fracture network data, and reservoir attribute point cloud data of shale gas wells; Specifically, multimodal data of the target shale gas well is acquired, including fracturing operation sequence data, fracture network data, and reservoir attribute point cloud data. For the model training phase, EUR tags corresponding to the training samples are also acquired. These EUR tags are used to supervise the training of the regression prediction network and are not used as input data for the EUR prediction phase of the well to be predicted.

[0026] In this embodiment, the target shale gas well can be a horizontal shale gas well that has completed drilling and completion design and has formed a fracturing construction plan but has not yet been put into production, or it can be a representative well or a historical simulation well used for model training. The fracturing construction sequence data is used to describe the pumping parameters of each construction step in the fracturing construction process, including at least the construction displacement, fracturing fluid type, injected fluid volume, proppant type, and proppant concentration. The fracture network data is used to describe the spatial morphology and connectivity of the artificial fractures formed after fracturing, including at least the fracture facet coordinates, fracture segment length, fracture height, fracture azimuth, fracture area, and fracture segment connection relationships. The reservoir attribute point cloud data is used to describe the spatial distribution of reservoir attributes around the target well, including at least the spatial coordinates and one or more of the following: permeability, porosity, gas saturation, geostress parameters, or rock mechanics parameters corresponding to the spatial coordinates.

[0027] It is understood that during the model training phase, the EUR tag can be obtained from field production data, numerical simulation results corrected for historical fitting, or reservoir engineering evaluation results. For samples of wells not yet in production, it is preferable to use a coupled approach of fracturing propagation simulation and shale gas reservoir production numerical simulation to obtain the cumulative gas production under a preset production regime and preset development years, and then determine the cumulative gas production as the EUR tag.

[0028] S102: Perform sequence preprocessing on the fracturing operation sequence data to obtain the fracturing operation sequence input tensor; Specifically, in this embodiment, the fracturing operation sequence data is a variable-length multivariate time series. Each operation step corresponds to a pumping stage, and there is a clear engineering sequence and stage dependency between different operation steps. The fracturing operation sequence data is preprocessed, including standardizing continuous parameters, class indexing and embedding mapping for discrete parameters, and padding and masking for fracturing operation sequences of different lengths to obtain a fracturing operation sequence input tensor that can be input into a neural network.

[0029] S103: Perform graph structure transformation on the crack network data to obtain crack network graph data; Specifically, in this embodiment, the crack network consists of multiple crack patches or crack segments, exhibiting irregular geometric structures and complex connectivity relationships. The crack patches or crack segments are abstracted as nodes in a graph structure, and the spatial connections, endpoint sharing, or flow-directing connectivity relationships between crack segments are abstracted as edges in a graph structure, thereby obtaining crack network graph data containing node and edge features.

[0030] S104: Perform point cloud preprocessing on the reservoir attribute point cloud data to obtain the reservoir attribute point cloud input tensor; Specifically, in this embodiment, the reservoir attribute point cloud data consists of multiple spatial sampling points, each of which includes spatial coordinates and reservoir attribute values. Point cloud preprocessing is performed on the reservoir attribute point cloud data, including point cloud sampling, coordinate centering, scale normalization, attribute value standardization, and point feature concatenation, to form a reservoir attribute point cloud input tensor with a fixed number of points and a fixed feature dimension.

[0031] S105: Input the fracturing operation sequence into a tensor input sequence encoder to obtain the construction mode feature vector; S106: Input the crack network graph data into a graph structure encoder to obtain the crack network modal feature vector; S107: Input the reservoir attribute point cloud into a tensor and input it into a point cloud encoder to obtain the reservoir attribute modal feature vector; Specifically, in this embodiment, the fracturing operation sequence is input into a tensor sequence encoder to obtain the operation mode feature vector; the fracture network diagram data is input into a graph structure encoder to obtain the fracture network mode feature vector; and the reservoir attribute point cloud is input into a tensor point cloud encoder to obtain the reservoir attribute mode feature vector. These three types of features respectively characterize the fracturing operation process, the artificial fracture conduction structure, and the reservoir spatial heterogeneity.

[0032] S108: Input the construction modal feature vector, the fracture network modal feature vector, and the reservoir attribute modal feature vector into the multimodal fusion network to obtain the multimodal fusion feature vector; S109: Input the multimodal fusion feature vector into the regression prediction network to obtain the EUR prediction value of shale gas wells.

[0033] Specifically, in this embodiment, the three types of modal features are mapped to a unified feature dimension and combined with learnable aggregated label vectors to form a modal feature sequence; a cross-modal self-attention fusion network is used to learn the coupling relationship between different modalities to obtain multimodal fusion features; then the multimodal fusion features are input into a regression prediction network to obtain the EUR prediction value of the target shale gas well.

[0034] Therefore, this embodiment uses three types of data—fracture operation sequence, fracture network, and reservoir attribute point cloud—to jointly characterize the fracturing development process of shale gas wells. It can achieve rapid EUR prediction using a multimodal deep learning surrogate model in the absence of historical production data of the target well. This helps reduce the cost of repeated calculations in numerical simulation and improves the efficiency of evaluating the development potential of shale gas wells that have not yet entered production and selecting the best fracturing scheme.

[0035] The following combination Figures 2 to 5 The method for constructing the multimodal sample set used to train the model and the specific feature extraction process for each encoder are further described. See [link to documentation]. Figure 1 As shown, this invention discloses a method for constructing a multimodal sample set, comprising: Step S201: Obtain the geological model, geomechanical model, and representative well trajectory data of the target shale gas block.

[0036] In this embodiment, the geological model includes parameters such as reservoir structure, stratigraphy, thickness, porosity, permeability, gas saturation, pressure, and temperature. The geomechanical model includes parameters such as Young's modulus, Poisson's ratio, rock density, maximum horizontal principal stress, minimum horizontal principal stress, natural fracture density, natural fracture aperture, and natural fracture orientation. The representative well trajectory data includes horizontal wellbore trajectory, horizontal segment length, segment location, cluster spacing, and perforation cluster parameters.

[0037] Step S202: Generate multiple sets of fracturing construction parameter combinations based on the on-site fracturing design rules.

[0038] In this embodiment, the boundaries of the fracturing parameters are first determined according to the on-site fracturing construction specifications for the target block, including the range of fracturing flow rate, the range of single-stage injection volume, the range of proppant concentration, the set of fracturing fluid types, and the set of proppant types. Then, under the condition of satisfying the on-site construction rules and engineering constraints, multiple sets of fracturing construction parameter combinations are generated using random sampling, Latin hypercube sampling, or orthogonal experimental design methods. Each set of fracturing construction parameter combinations includes multiple construction steps, and each construction step includes at least fracturing flow rate, fracturing fluid type, injection volume, proppant type, and proppant concentration.

[0039] It is understood that the combination of construction parameters can be derived from actual construction design or from simulated construction schemes constrained by on-site expert rules. To reduce input feature redundancy, proppant mass and construction time, which can be directly calculated from construction displacement, injection volume, and proppant concentration, can be saved as auxiliary parameters or not used as independent inputs to the model.

[0040] Step S203: Conduct fracture propagation simulation for each set of fracturing construction parameter combinations to obtain the corresponding fracture network data.

[0041] In this embodiment, each set of fracturing construction parameters is input into the fracturing fracture propagation simulation model. Under the joint constraints of geomechanical parameters, natural fracture parameters, and construction parameters, the artificial fracture propagation process is simulated to obtain the spatial distribution results of the fracture network. The fracture network can be composed of multiple fracture patches, each containing information such as endpoint coordinates, patch height, patch length, patch area, azimuth angle, and conductivity. The spatial coordinates of the fracture patch can be represented as x, y, and z coordinates in the global coordinate system, where z represents the vertical coordinate.

[0042] Step S204: Based on the fracture network data and reservoir attribute data, establish a numerical simulation model of the shale gas reservoir after fracturing to obtain reservoir attribute point cloud data.

[0043] In this embodiment, a fracturing network is embedded into a shale gas reservoir numerical simulation model to establish a post-fracturing shale gas reservoir model including artificial fractures, natural fractures, and a matrix system. Based on the changes in fracture conductivity and reservoir permeability enhancement after fracturing, a three-dimensional reservoir attribute field is extracted within the control range of the target well. The three-dimensional reservoir attribute field is preferably represented in point cloud form, where each point cloud sampling point includes spatial coordinates and a corresponding attribute value. Taking permeability as an example, the point cloud data can be represented as follows: in, , , Let i be the spatial coordinates of the i-th sampling point. Let i be the permeability attribute value corresponding to the i-th sampling point. This represents the number of sampling points.

[0044] Step S205: Conduct production numerical simulation under the preset production system, and determine the cumulative gas production within the preset development period as the EUR label.

[0045] In this embodiment, a shale gas reservoir numerical simulation model incorporating fracture networks and reservoir property fields is used to set production regime parameters and conduct production simulations. These parameters include minimum bottomhole flowing pressure, maximum daily gas production, production duration, abandonment pressure, and production constraints. After the simulation, daily gas production curves and cumulative gas production curves are obtained, and the cumulative gas production within the preset development duration is defined as the sample EUR label.

[0046] This leads to the construction of a multimodal sample set: in, This represents the fracturing operation sequence data for the i-th sample. For the crack network diagram data of the i-th sample, For the reservoir attribute point cloud data of the i-th sample, Let N be the EUR label corresponding to the i-th sample, and N be the number of samples.

[0047] Therefore, it can be seen that by constructing a multimodal sample set through the process of "generating fracturing construction parameters - simulating fracture propagation - extracting reservoir attribute point cloud - numerical simulation of production - determining EUR labels", the sample data can simultaneously contain engineering construction information, fracture diversion structure information and reservoir spatial heterogeneity information, thus providing a foundation for subsequent multimodal fusion model training.

[0048] The following description, in conjunction with the accompanying drawings, further illustrates the specific implementation methods of feature extraction for single-modal data by each encoder in this invention.

[0049] The fracturing operation sequence data is a variable-length, multivariable time series. For example... Figure 2 As shown, the sequence encoder of the present invention extracts construction modal features in the following manner: Step S301: Standardize the continuous parameters in the fracturing operation sequence.

[0050] In this embodiment, the continuous parameters in the fracturing operation sequence include the fracturing displacement, injected fluid volume, and proppant concentration. For any continuous parameter x, the training set statistics are used for standardization: Where x′ is the standardized continuous parameter, and x is the original continuous parameter. The mean of this parameter in the training set, This represents the standard deviation of the parameter in the training set. Standardization reduces the impact of different units and numerical ranges on the stability of model training.

[0051] Step S302: Perform category index encoding and embedding mapping on the discrete parameters in the fracturing operation sequence.

[0052] In this embodiment, the discrete parameters in the fracturing operation sequence include fracturing fluid type and proppant type. The fracturing fluid type may include one or more of acid, slickwater, low-viscosity fluid, and high-viscosity fluid; the proppant type may include one or more of proppant-free, 70 / 140 mesh proppant, and 30 / 50 mesh proppant. After class indexing the discrete parameters, the class index is input into the embedding layer to obtain a low-dimensional continuous embedding vector.

[0053] Step S303: Concatenate the continuous parameter and discrete parameter embedding vectors and map them to the step embedding vector.

[0054] In this embodiment, for the t-th construction step, the standardized continuous parameter vector and the discrete parameter embedding vector are concatenated along the feature dimension to obtain the comprehensive input vector for that construction step. Then, a linear mapping layer maps the comprehensive input vector to a unified embedding dimension to obtain the step embedding vector for the t-th construction step. in, Let t be the step embedding vector for the t-th construction step. Let be the continuous parameter vector for the t-th construction step. Let f(t) be the discrete parameter index for the t-th construction step, where Emb(·) represents the embedded mapping function and Linear(·) represents the linear mapping function.

[0055] Step S304: Encode the position of the embedded vector in step S304 and input it into the sequence encoder.

[0056] In this embodiment, since the fracturing construction steps have a clear sequence, in order for the model to identify the positional relationships of different construction stages such as the pre-flush fluid, the proppant-carrying section, and the tailings section, positional codes are superimposed on the step embedding vector to obtain a step vector containing construction sequence information: in, This is the step vector after overlaying position encoding. This is the location code for the t-th construction step.

[0057] The sequence encoder is preferably a Transformer encoder. The Transformer encoder includes a multi-head self-attention layer, a feedforward fully connected layer, a residual connected layer, and a normalization layer. The self-attention mechanism is used to learn the dependencies between different construction steps, and its calculation form is as follows: Where Q, K, and V are the query vector, key vector, and value vector, respectively. The dimension of the key vector.

[0058] Step S305: Perform sequence pooling based on the sequence mask to obtain the construction modal feature vector.

[0059] In this embodiment, the number of steps may vary for different fracturing operation samples. To ensure consistent input tensor size, the fracturing operation sequence is padded to a preset maximum length, and a sequence mask is constructed. The mask value corresponding to the actual operation step is 1, and the mask value corresponding to the padded step is 0. The sequence mask is used to shield the influence of the padded step during attention calculation and sequence pooling stages.

[0060] After the Transformer encoder outputs the hidden vectors for each construction step, attention-weighted pooling or average pooling is used to convert the variable-length sequence into a fixed-dimensional vector. Preferably, the construction modal feature vector is represented as: in, For construction modal feature vectors, This is the output code corresponding to the t-th construction step. Let t be the attention weight corresponding to the t-th construction step, where T is the actual number of construction steps.

[0061] Therefore, this embodiment transforms the fracturing construction parameter table from a regular two-dimensional table into a fracturing construction sequence input with sequential relationships. By using embedding mapping, position encoding, self-attention modeling, and sequence pooling to extract construction modal features, it can fully characterize the stage-based and cross-step dependencies in the fracturing construction process.

[0062] The crack network data has an irregular geometric shape and complex topological relationships. For example... Figure 3 As shown, the present invention extracts crack network modal features in the following manner, including: Step S401: Convert the crack patch or crack segment into a graph structure node.

[0063] In this embodiment, the fracture patches or fracture segments output from the fracturing simulation are used as nodes in the graph structure. For each node, the geometric center of the fracture patch or fracture segment is taken as the node location, and node features are extracted. The node features include one or more of the following: fracture segment length, fracture height, fracture azimuth angle, fracture area, fracture conductivity, and node spatial coordinates.

[0064] Step S402: Construct graph structure edges based on the spatial connectivity between crack segments.

[0065] In this embodiment, when two crack segments meet at least one of the following conditions: sharing an endpoint, spatial distance less than a preset distance threshold, spatial intersection, or existence of a flow-guiding connection, an edge is established between the corresponding nodes of the two crack segments. The edge is used to characterize the connectivity and spatial adjacency between the crack segments.

[0066] Step S403: Construct node features and edge features to obtain crack network graph data.

[0067] In this embodiment, the node features can be represented as: in, Let v be the node characteristics. The length of the crack segment. The height of the crack. The azimuth angle of the crack. The area of ​​the crack. This refers to the flow conductivity of the crack.

[0068] Edge features can be represented as: in, The edge features between node u and node v The spatial distance between nodes. For the height difference, This is the azimuth difference. This is a connectivity status identifier.

[0069] Step S404: Input the crack network graph data into the graph structure encoder to obtain the node embedding vector.

[0070] In this embodiment, the graph structure encoder preferably employs an edge-enhanced graph neural network encoder. Through a message passing mechanism, each node can aggregate features from its neighboring nodes and edges, thereby learning the local density, bifurcation structure, connectivity, and overall morphology of the fracture network. The node feature update process can be represented as follows: in, Let v be the feature vector of the l-th node. Let v be the set of neighboring nodes. The edge features between node u and node v (·) is the edge feature mapping function. For learnable weight matrix, For learnable parameters, (·) is a non-linear activation function, and ⊙ represents element-wise operation.

[0071] Step S405: Perform graph-level pooling on the node embedding vectors to obtain the modal feature vectors of the crack network.

[0072] In this embodiment, after multi-layer graph structure encoding, the node embedding vector of each crack node is obtained. Global mean pooling, global max pooling, or graph attention pooling are used to aggregate the node-level features into a graph-level feature vector. Preferably, the crack network modal feature vector is represented as: in, For the modal feature vectors of the crack network, V is the node feature vector output by the last layer graph structure encoder, where V is the set of nodes.

[0073] Therefore, this embodiment converts the artificial fracture network into graph structure data and extracts the geometric and topological connectivity information of the fractures through a graph neural network, which can effectively characterize the control effect of the artificial fracture system on the EUR of shale gas wells.

[0074] The reservoir attribute point cloud data has spatial heterogeneity and irregular sampling characteristics. For example... Figure 4 As shown, the present invention extracts reservoir attribute modal features in the following manner, including: Step S501: Obtain reservoir attribute point cloud data.

[0075] In this embodiment, the reservoir attribute point cloud data includes multiple spatial sampling points, each containing spatial coordinates and a reservoir attribute value. Taking permeability as an example, the i-th point cloud sampling point can be represented as: in, , , For spatial coordinates, This is the permeability attribute value.

[0076] Step S502: Perform point cloud sampling on the reservoir attribute point cloud data.

[0077] In this embodiment, since the original point cloud is large, directly inputting it into the point cloud encoder would increase computational overhead and reduce training efficiency. Preferably, a farthest-point sampling method is used to select a fixed number of sampling points from the original point cloud, ensuring that the sampling points cover the entire spatial range of the original point cloud as much as possible. It is understood that random sampling, uniform grid sampling, or layered sampling methods can also be used for point cloud sampling.

[0078] Step S503: Center and scale normalize the spatial coordinates of the sampled point cloud.

[0079] In this embodiment, the geometric center of the sampled point cloud is first calculated: Where c is the geometric center of the point cloud. This represents the number of sampling points. Then, the coordinates are centered. Considering the significant difference between the horizontal and vertical scales of shale reservoirs, it is preferable to normalize using both horizontal and vertical scale factors separately. The horizontal scale factor is: The vertical scaling factor is: Normalized coordinates are: Step S504: Standardize the reservoir attribute values ​​and construct point features.

[0080] In this embodiment, permeability, porosity, or other reservoir attribute values ​​are standardized using training set statistics. Taking permeability as an example, its standardized form is: in, This is the standardized penetration rate attribute value. This is the original permeability attribute value. and These represent the mean and standard deviation of the penetration rate attribute values ​​in the training set, respectively. Concatenating the normalized spatial coordinates with the standardized attribute values ​​yields the point features: Step S505: Input the point features into the point cloud encoder to obtain the reservoir attribute mode feature vector.

[0081] In this embodiment, the point cloud encoder preferably adopts the PointNet++ architecture. PointNet++ extracts the spatial distribution features of the point cloud through hierarchical sampling, local neighborhood construction, and local feature aggregation. For any center point, a local neighborhood is constructed using a fixed-radius spherical neighborhood: Where r is the neighborhood radius, This represents the Euclidean distance. If the number of points in the neighborhood is less than the preset number, repeated sampling is used to make up the difference; if the number of points in the neighborhood exceeds the preset number, a fixed number of points are sampled from the neighborhood points.

[0082] After passing through multiple layers of ensemble abstraction, the point cloud encoder outputs reservoir attribute modal feature vectors: in, This represents the reservoir property modal feature vector. (·) represents the point cloud encoder. Input tensors into the reservoir attribute point cloud.

[0083] Therefore, this embodiment expresses the spatial distribution of reservoir attributes in the form of point clouds, avoiding the information loss caused by forcibly interpolating irregular reservoir attribute fields into regular grids, and extracts the local heterogeneity and global spatial distribution features of the reservoir through a point cloud encoder.

[0084] After obtaining the construction mode feature vector, fracture network mode feature vector, and reservoir attribute mode feature vector respectively, this invention further fuses the three through a multimodal fusion network and finally outputs the EUR prediction result. Figure 5 As shown, the specific process of multimodal fusion and EUR prediction includes: Step S601: Map the three types of modal feature vectors to a unified feature dimension.

[0085] In this embodiment, since the feature dimensions output by the sequence encoder, graph structure encoder, and point cloud encoder may be different, a linear mapping layer is used to map the three types of modal features to a unified dimension d: in, , , These are the construction mode feature vector, fracture network mode feature vector, and reservoir attribute mode feature vector, respectively. , , These are the mapped modal feature vectors; , , For linear mapping weights; , , This is a bias term.

[0086] Step S602: Construct a modal feature sequence containing learnable aggregated tags.

[0087] In this embodiment, a learnable aggregated tag vector is set. And combine it with the three types of modal feature vectors to form a modal feature sequence: in, Used to aggregate interaction information between the three modalities Characterizes fracturing operation sequence information. Characterize the crack network structure information. Characterizes the spatial distribution information of reservoir properties.

[0088] Step S603: Input the modal feature sequence into the cross-modal self-attention fusion network to obtain multimodal fusion features.

[0089] In this embodiment, the cross-modal self-attention fusion network preferably adopts a Transformer encoder structure. The attention weights between different modal vectors are calculated through a self-attention mechanism, enabling the model to automatically learn the coupling relationship between fracturing operation sequences, fracture networks, and reservoir properties. After fusion, the updated learnable aggregated label vector is taken as the multimodal fusion feature. in, This is a multimodal fusion feature vector. This is the aggregated label vector updated after passing through a cross-modal self-attention fusion network.

[0090] Step S604: Input the multimodal fusion features into the regression prediction network to obtain the standardized EUR prediction value.

[0091] In this embodiment, the regression prediction network includes at least one fully connected layer, a nonlinear activation layer, and an output layer. After inputting the multimodal fusion feature vector into the regression prediction network, it outputs a standardized EUR prediction value. in, For standardized EUR forecasts, (·) represents the regression prediction network.

[0092] Step S605: Perform denormalization on the standardized EUR prediction values ​​to obtain the EUR prediction values ​​under the actual physical dimensions.

[0093] In this embodiment, to improve the stability of model training, the EUR labels are standardized during the training phase: in, The standardized EUR value. The original EUR value. For the EUR mean in the training set, The standard deviation of EUR in the training set.

[0094] During the prediction phase, the standardized EUR prediction values ​​output by the model are destandardized: in, This is the predicted value of EUR in actual physical dimensions.

[0095] Therefore, this embodiment realizes the information interaction between fracturing operation sequence, fracture network and reservoir attribute point cloud through cross-modal self-attention mechanism. Compared with the method of directly splicing multiple features and inputting them into the regression network, it can more fully capture the coupling relationship between multi-source heterogeneous information and improve the EUR prediction accuracy of non-production shale gas wells.

[0096] This invention can output the EUR of shale gas wells through a multimodal fusion prediction model.

[0097] In this embodiment, during model training, the multimodal sample set is divided into a training set, a validation set, and a test set, or a K-fold cross-validation method is used for model training and performance evaluation. The training set is used to update model parameters, the validation set is used to select model hyperparameters and prevent overfitting, and the test set is used to evaluate the model's generalization performance. The hyperparameters include the encoding dimension of the fracturing construction sequence, the number of attention heads, the number of graph neural network layers, the number of point cloud sampling points, the hidden dimension of the fusion layer, the learning rate, the batch size, the number of training epochs, and the dropout rate.

[0098] During model training, the standardized EUR value was used as the supervision signal, and the mean squared error was used as the loss function. Where L is the training loss and N is the number of training samples. Let be the standardized EUR prediction value for the i-th sample. Let be the standardized true EUR value of the i-th sample.

[0099] Preferably, the Adam optimizer is used to update model parameters, and a learning rate decay strategy is combined to improve model convergence stability. During training, Dropout layers and weight decay terms are set in sequence encoders, graph structure encoders, point cloud encoders, multimodal fusion networks, or regression prediction networks to reduce the risk of model overfitting.

[0100] After the model is trained, inverse standardization is performed on the test set samples for prediction, and the prediction performance is evaluated using the coefficient of determination R², mean absolute error (MAE), and root mean square error (RMSE). in, For the true EUR of the i-th sample, For the predicted EUR of the i-th sample, This represents the true EUR mean of the test sample.

[0101] To verify the effectiveness of multimodal information fusion, a single-modal model, a bimodal model, and a feature direct concatenation model can be used as comparative models. The single-modal model uses only one modal input from the fracturing operation sequence, fracture network, or reservoir attribute point cloud; the bimodal model uses any two modal inputs from the three modalities; the feature direct concatenation model directly concatenates the outputs of the three encoders and inputs them into the regression prediction network, without setting a cross-modal self-attention fusion network. The R-values ​​of the different models are compared. 2 MAE and RMSE can be used to evaluate the contribution of different modal information and cross-modal fusion mechanisms to EUR prediction performance.

[0102] Therefore, this embodiment achieves model training and verification through training loss function, inverse standardization prediction and multi-index evaluation, and verifies the complementary effect of three modes: fracturing operation sequence, fracture network and reservoir attribute point cloud through ablation experiment.

[0103] Once the above model training and validation are complete, the trained multimodal fusion prediction model can be applied to predict the EUR of the target unproduced shale gas wells. The specific process is as follows: For target shale gas wells that are not yet in production, the fracturing construction design scheme for the well is first obtained, and this scheme is then compiled into fracturing construction sequence data. Subsequently, based on the geomechanical model of the target well's location and the fracturing construction design parameters, fracture propagation simulation is conducted to obtain fracture network data for the target well. Furthermore, reservoir attribute point cloud data is extracted from the control area of ​​the target well.

[0104] The fracturing operation sequence data, fracture network data, and reservoir attribute point cloud data of the target well are processed using the same preprocessing methods as in the training phase to obtain the fracturing operation sequence input tensor, fracture network map data, and reservoir attribute point cloud input tensor. These three types of input data are then fed into the trained multimodal fusion prediction model to output the target well's EUR prediction value.

[0105] Understandably, when multiple fracturing construction design schemes exist for the target well, a corresponding fracturing construction sequence can be constructed for each scheme. The corresponding fracture network can then be obtained through fracture propagation simulation and input into a trained multimodal fusion prediction model to obtain EUR prediction results under different fracturing construction schemes. By comparing the EUR prediction values ​​of different schemes, a basis can be provided for optimizing fracturing construction parameters, selecting optimal well locations, and deploying development plans.

[0106] Therefore, this embodiment can obtain the EUR prediction result by performing only one forward prediction on the multimodal input of the target well after the model training is completed. Compared with carrying out full-process numerical simulation for each scheme, it can shorten the prediction time and improve the efficiency of evaluating the development potential of unproduced wells.

[0107] The invention will be further illustrated below with a specific application example. The implementation process of this example can be found by referring to [reference needed]. Figures 1 to 4 .

[0108] In this embodiment, a representative horizontal well within the target shale gas block is selected as the sample construction object, and a multimodal sample set is constructed based on the geological model and numerical simulation model of the target block calibrated with field data. First, the boundaries of the construction parameters are determined according to the field fracturing construction design specifications, including parameters such as the fracturing flow rate, injection volume, fracturing fluid type, proppant type, and proppant concentration; then, multiple sets of fracturing construction parameter combinations are generated, and fracture propagation simulation and production numerical simulation are carried out for each set of construction parameter combinations.

[0109] Each sample includes three types of input data and one EUR label. The first type of input is the fracturing operation sequence, with each step including the fracturing fluid displacement, fracturing fluid type, injected fluid volume, proppant type, and proppant concentration. The second type of input is fracture network data, including fracture patch coordinates, fracture length, fracture height, fracture azimuth, fracture area, and fracture segment connectivity. The third type of input is reservoir attribute point cloud data, with each point cloud sampling point including spatial coordinates and permeability attribute values. The output label is the cumulative gas production (EUR) within the preset development years under the preset production regime.

[0110] During model training, the fracturing operation sequence branch uses a Transformer encoder to extract operation modal features; the fracture network branch uses a graph neural network encoder to extract fracture topological and geometric features; and the reservoir attribute point cloud branch uses a PointNet++ encoder to extract reservoir spatial heterogeneity features. After the three types of modal features are uniformly mapped to the same feature dimension, they are combined with learnable aggregated label vectors to form a modal feature sequence, which is then input into a cross-modal Transformer fusion layer. The aggregated label vector output from the fusion layer is input into a regression prediction network to obtain standardized EUR prediction values, which are then de-standardized to recover the actual EUR prediction values.

[0111] To ensure the stability of the model evaluation results, this embodiment uses K-fold cross-validation for training and testing. In each fold, a subset of samples is selected as the test set, while the remaining samples serve as the training and validation sets. Model performance is calculated as the average of the test results from each fold. Evaluation metrics include R², MAE, and RMSE. By comparing the single-modal fracturing construction model, the multimodal direct splicing model, and the multimodal self-attention fusion model described in this invention, the effect of the cross-modal fusion mechanism on improving EUR prediction accuracy can be verified.

[0112] Therefore, this invention can transform complex multi-source information generated by numerical simulation into trainable multimodal samples, and establish a nonlinear mapping relationship between input information and EUR through a deep learning surrogate model, thereby realizing rapid prediction of EUR and screening of fracturing schemes for shale gas wells that have not yet entered production.

[0113] Second embodiment: Please refer to Figure 5 This invention provides a shale gas well EUR prediction system based on multimodal information fusion, comprising: Data acquisition module 1 is used to acquire fracturing operation sequence data, fracture network data, and reservoir attribute point cloud data of shale gas wells; The fracturing operation sequence processing module 2 is connected to the data acquisition module 1 and is used to receive the fracturing operation sequence data, perform sequence preprocessing, and output the fracturing operation sequence input tensor. The crack network processing module 3 is connected to the data acquisition module 1 and is used to receive the crack network data, perform graph structure conversion, and output crack network graph data. Point cloud processing module 4 is connected to the data acquisition module 1 and is used to receive the reservoir attribute point cloud data and perform point cloud preprocessing, and output the reservoir attribute point cloud input tensor. The construction feature extraction module 5 is connected to the fracturing construction sequence processing module 2. It is used to receive the input tensor of the fracturing construction sequence and input it into the sequence encoder, and output the construction modal feature vector. The crack feature extraction module 6 is connected to the crack network processing module 3 and is used to receive the crack network graph data and input it into the graph structure encoder, and output the crack network modal feature vector. The attribute feature extraction module 7 is connected to the point cloud processing module 4 and is used to receive the reservoir attribute point cloud input tensor and input it into the point cloud encoder, and output the reservoir attribute modal feature vector. The multimodal fusion module 8 is connected to the construction feature extraction module 5, the fracture feature extraction module 6, and the attribute feature extraction module 7, respectively, and is used to receive the construction modal feature vector, the fracture network modal feature vector, and the reservoir attribute modal feature vector, perform multimodal fusion, and output the multimodal fusion feature vector; The EUR prediction module 9, connected to the multimodal fusion module 8, is used to receive the multimodal fusion feature vector and input it into the regression prediction network, and output the EUR prediction value of the shale gas well.

[0114] Specifically, the specific processing procedures of the data acquisition module 1, fracturing construction sequence processing module 2, fracture network processing module 3, point cloud processing module 4, construction feature extraction module 5, fracture feature extraction module 6, attribute feature extraction module 7, multimodal fusion module 8, and EUR prediction module 9 can be referred to the corresponding contents disclosed in the aforementioned method embodiments, and will not be repeated here.

[0115] Therefore, the prediction system disclosed in this embodiment can realize multimodal data preprocessing, feature extraction, cross-modal fusion and EUR prediction of shale gas wells through multiple functional modules, and is suitable for deployment in software systems for evaluating shale gas well fracturing schemes, planning production capacity construction and rapidly screening development potential.

[0116] The above-disclosed embodiments are merely one or more preferred embodiments of this application and should not be construed as limiting the scope of this application. Those skilled in the art can understand that all or part of the processes for implementing the above embodiments and equivalent changes made in accordance with the claims of this application still fall within the scope of this application.

Claims

1. A shale gas well EUR prediction method based on multimodal information fusion, characterized in that, Includes the following steps: The fracturing operation sequence data, fracture network data, reservoir attribute point cloud data, and EUR data of shale gas wells are acquired. The fracturing operation sequence data is derived from the combination of simulated operation parameters generated according to the on-site fracturing operation design rules. The fracture network data is derived from the simulation results of fracturing fracture propagation based on the fracturing operation sequence data. The reservoir attribute point cloud data is derived from the spatial attribute data in geological models, geomechanical models, reservoir attribute models, or numerical simulation models of shale gas reservoirs after fracturing. The EUR data is derived from the production numerical simulation results, calibrated numerical simulation results, or gas reservoir engineering evaluation results carried out under a preset production regime, and is used to train the regression prediction network. The fracturing operation sequence data is preprocessed to obtain the fracturing operation sequence input tensor; The crack network data is transformed into a graph structure to obtain crack network graph data; Point cloud preprocessing is performed on the reservoir attribute point cloud data to obtain the reservoir attribute point cloud input tensor; The fracturing operation sequence is input into a tensor input sequence encoder to obtain the operation mode feature vector; The crack network graph data is input into a graph structure encoder to obtain the crack network modal feature vector; The reservoir attribute point cloud is input into a tensor input point cloud encoder to obtain the reservoir attribute modal feature vector; The construction modal feature vector, the fracture network modal feature vector, and the reservoir attribute modal feature vector are input into a multimodal fusion network to obtain a multimodal fusion feature vector; The multimodal fusion feature vector is input into the regression prediction network to obtain the EUR prediction value for shale gas wells.

2. The EUR prediction method for shale gas wells based on multimodal information fusion as described in claim 1, characterized in that, The fracturing operation sequence data includes multiple fracturing operation steps. Each fracturing operation step includes continuous parameters and discrete parameters. The continuous parameters include the operation displacement, the injected fluid volume, and the proppant concentration. The discrete parameters include the fracturing fluid type and the proppant type. Sequence preprocessing of the fracturing operation sequence data includes: standardizing the continuous parameters; The discrete parameters are categorically indexed and encoded, and the encoded discrete parameters are input into the embedding layer to obtain discrete parameter embedding vectors. The standardized continuous parameters are concatenated with the discrete parameter embedding vectors, and the step embedding vectors for each construction step are obtained through a linear mapping layer. The fracturing construction sequences of different lengths are padded, and sequence masks are generated to identify the actual construction steps and filling steps.

3. The EUR prediction method for shale gas wells based on multimodal information fusion as described in claim 1, characterized in that, The fracture network data includes spatial location data, geometric attribute data, and connectivity data of multiple fracture patches or fracture segments; the geometric attribute data includes at least one of fracture segment length, fracture height, fracture azimuth angle, fracture area, and fracture conductivity. The crack network data is transformed into a graph structure to obtain crack network graph data, including: converting the crack patches or crack segments into nodes in a graph structure; using the geometric properties of the crack patches or crack segments as node features; establishing edges in the graph structure based on the shared endpoint relationships, spatial proximity relationships, or flow connectivity relationships between the crack patches or crack segments; and using the distances, angles, height differences, or connectivity states between the crack patches or crack segments as edge features to obtain crack network graph data containing node features and edge features.

4. The EUR prediction method for shale gas wells based on multimodal information fusion as described in claim 1, characterized in that, The reservoir attribute point cloud data includes multiple sampling points, each sampling point including spatial coordinates and reservoir attribute values; the reservoir attribute values ​​include at least one of permeability, porosity, gas saturation, geostress parameters, or rock mechanics parameters; Point cloud preprocessing is performed on the reservoir attribute point cloud data to obtain the reservoir attribute point cloud input tensor, including: sampling the farthest point of the reservoir attribute point cloud data to obtain a fixed number of point cloud sampling points; centering and scaling the spatial coordinates of the point cloud sampling points; standardizing the reservoir attribute values ​​of the point cloud sampling points; and concatenating the normalized spatial coordinates with the standardized reservoir attribute values ​​to obtain the reservoir attribute point cloud input tensor.

5. The EUR prediction method for shale gas wells based on multimodal information fusion as described in claim 1, characterized in that, The sequence encoder includes a position encoding layer, a self-attention encoding layer, and a sequence pooling layer. The position encoding layer is used to superimpose the construction step sequence information into the construction step embedding vector. The self-attention encoding layer is used to learn the dependencies between different construction steps. The sequence pooling layer is used to convert the variable-length fracturing construction sequence into a fixed-dimensional construction modal feature vector.

6. The EUR prediction method for shale gas wells based on multimodal information fusion as described in claim 1, characterized in that, The graph structure encoder is an edge-enhanced graph neural network encoder; the edge-enhanced graph neural network encoder aggregates node features and edge features through a message passing mechanism to obtain the crack network modal feature vector.

7. The EUR prediction method for shale gas wells based on multimodal information fusion as described in claim 1, characterized in that, The point cloud encoder uses hierarchical sampling, local neighborhood construction, and local feature aggregation to extract local heterogeneous features and global spatial distribution features of reservoir attribute point clouds, thereby obtaining reservoir attribute modal feature vectors.

8. The EUR prediction method for shale gas wells based on multimodal information fusion as described in claim 1, characterized in that, The construction modal feature vector, the fracture network modal feature vector, and the reservoir attribute modal feature vector are input into a multimodal fusion network to obtain a multimodal fusion feature vector, specifically including: Linear mapping is performed on the construction modal feature vector, the fracture network modal feature vector, and the reservoir attribute modal feature vector, respectively, so that the three modal feature vectors have the same feature dimension; Set a learnable aggregated tag vector; The learnable aggregated label vector, the linearly mapped construction modal feature vector, the linearly mapped fracture network modal feature vector, and the linearly mapped reservoir attribute modal feature vector are combined to form a modal feature sequence; The modal feature sequence is input into a cross-modal self-attention fusion network to obtain an updated learnable aggregated label vector; The updated learnable aggregated label vector is used as the multimodal fusion feature vector.

9. The EUR prediction method for shale gas wells based on multimodal information fusion as described in claim 1, characterized in that, The step of inputting the multimodal fused feature vector into the regression prediction network to obtain the EUR prediction value for shale gas wells includes: The multimodal fusion feature vector is input into the regression prediction network to obtain the standardized EUR prediction value; the standardized EUR prediction value is then de-standardized to obtain the EUR prediction value of shale gas wells under the actual physical dimensions.

10. A shale gas well EUR prediction system based on multimodal information fusion, applied to the shale gas well EUR prediction method based on multimodal information fusion as described in claim 1, characterized in that, include: The data acquisition module is used to acquire fracturing operation sequence data, fracture network data, and reservoir attribute point cloud data of shale gas wells; The fracturing operation sequence processing module is connected to the data acquisition module and is used to receive the fracturing operation sequence data, perform sequence preprocessing, and output the fracturing operation sequence input tensor. A crack network processing module, connected to the data acquisition module, is used to receive the crack network data, perform graph structure conversion, and output crack network graph data. The point cloud processing module is connected to the data acquisition module and is used to receive the reservoir attribute point cloud data, perform point cloud preprocessing, and output the reservoir attribute point cloud input tensor. The construction feature extraction module is connected to the fracturing construction sequence processing module. It is used to receive the input tensor of the fracturing construction sequence and input it into the sequence encoder, and output the construction modal feature vector. A crack feature extraction module, connected to the crack network processing module, is used to receive the crack network graph data and input it into the graph structure encoder, and output the crack network modal feature vector. The attribute feature extraction module is connected to the point cloud processing module and is used to receive the reservoir attribute point cloud input tensor and input it into the point cloud encoder, and output the reservoir attribute modal feature vector. The multimodal fusion module is connected to the construction feature extraction module, the fracture feature extraction module, and the attribute feature extraction module, respectively, and is used to receive the construction modal feature vector, the fracture network modal feature vector, and the reservoir attribute modal feature vector, perform multimodal fusion, and output the multimodal fusion feature vector; The EUR prediction module, connected to the multimodal fusion module, is used to receive the multimodal fusion feature vector and input it into the regression prediction network, and output the EUR prediction value of shale gas wells.