Geological double sweet spot prediction method and device based on knowledge graph representation learning, equipment, medium and product
By constructing a geological knowledge graph and using graph attention networks for representation learning, combined with coupling constraint loss, we have achieved deep integration and collaborative optimization of geological and engineering attributes. This solves the problems of data silos and poor model interpretability in traditional sweet spot prediction methods, and provides accurate basis for exploration and development.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TONGJI UNIV
- Filing Date
- 2026-04-17
- Publication Date
- 2026-06-23
AI Technical Summary
Traditional dessert prediction methods suffer from problems such as data silos and missing correlations, simplistic definitions of desserts, and poor model interpretability, resulting in poor exploration success rates and development benefits.
A geological double-sweet spot prediction method based on knowledge graph representation learning is adopted. By constructing a geological knowledge graph of multi-source heterogeneous geological data, using graph attention network for representation learning, and combining coupling constraint loss, the deep integration and collaborative optimization of geological and engineering attributes are achieved.
It enables the synergistic prediction of geological and engineering sweet spots, providing accurate and reliable basis for exploration and development, improving the success rate of exploration and development benefits, solving the problems of data fragmentation and one-sided prediction, and enhancing the interpretability of the model.
Smart Images

Figure CN122066100B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of oil and gas exploration and development technology, and in particular to a geological double sweet spot prediction method, apparatus, equipment, medium and product based on knowledge graph representation learning. Background Technology
[0002] As global oil and gas exploration and development gradually shifts towards deep, unconventional, and complex structural areas, accurate prediction of high-quality reservoirs (i.e., "sweet spots") has become crucial in determining exploration success rates and development benefits. Traditional sweet spot prediction methods mainly rely on geophysical inversion, geostatistics, and empirical formulas, which suffer from three major bottlenecks: 1) Data silos and lack of correlation: Multi-source data such as well logging, seismic data, core data, and production dynamics are fragmented in traditional methods, making it difficult to form a unified knowledge system. The complex control relationships between geological entities such as reservoirs, caprocks, and faults cannot be quantitatively characterized, making it difficult for prediction models to unify them for systematic analysis; 2) Simplified definition of sweet spots: Traditional methods focus on "geological sweet spots," such as high-porosity and high-permeability reservoirs, while neglecting "engineering sweet spots" (such as reservoir fracturing). This leads to the prediction of favorable areas being difficult to develop effectively due to poor engineering conditions (such as high ground stress and poor brittleness), resulting in drilling failure or poor development results; 3) Poor model interpretability: Most existing machine learning-based sweet spot prediction models are "black boxes", and their prediction results lack clear geological mechanism explanations. Geologists find it difficult to understand and trust the decision-making basis of the model, which limits its application in key exploration decisions.
[0003] In recent years, artificial intelligence technology has provided new insights into geological prediction. However, current methods (such as convolutional neural networks for processing seismic data) fail to explicitly learn the topological relationships and spatial constraints between geological entities, resulting in a disconnect between their prediction process and the knowledge logic of geologists. The development of knowledge graphs and graph neural networks (GNNs) has offered a solution to these problems. GNNs are capable of deep representation learning of graph-structured data, making them naturally suitable for handling complex non-Euclidean spatial relationships between geological entities. However, directly applying general GNNs to geological prediction still faces core challenges, including how to effectively transform multi-source, multi-scale geological data into graph structures and how to enable the model to learn and follow geological laws.
[0004] Therefore, there is an urgent need in this field for a new method that can deeply integrate multi-source geological data, explicitly model geological relationship networks, and achieve collaborative intelligent prediction of both geological and engineering sweet spots. Summary of the Invention
[0005] The purpose of this application is to provide a geological double sweet spot prediction method, device, equipment, medium and product based on knowledge graph representation learning, which can realize geological double sweet spot prediction and provide accurate and reliable basis for efficient exploration and development of oil and gas reservoirs.
[0006] To achieve the above objectives, this application provides the following solution:
[0007] Firstly, this application provides a geological double sweet spot prediction method based on knowledge graph representation learning, including:
[0008] Acquire multi-source heterogeneous geological data of the target work area; the multi-source heterogeneous geological data includes geological and engineering data;
[0009] A geological knowledge graph is determined based on the multi-source heterogeneous geological data and the schema layer definition scheme. The geological knowledge graph includes entity nodes, attribute values, and relation edges. The schema layer definition scheme is determined based on the multi-source heterogeneous geological data. The schema layer definition scheme includes: an entity type list, an attribute list, and a relation type list.
[0010] A double-sweet spot prediction model is employed to determine the predicted outcome data based on the geological knowledge graph. This outcome data includes the spatial distribution information of the double-sweet spots, probability intensity, and a comprehensive assessment of fracturability. The double-sweet spot prediction model is trained based on the geological knowledge graph containing the known outcome data, with the objective of minimizing the total loss function. The total loss function includes classification loss, regression loss, and coupling constraint loss. The double-sweet spot prediction model uses a dual-prediction head model with a shared underlying GAT encoder. The GAT encoder employs a graph attention network. The graph attention network is used to learn geological knowledge representations from the geological knowledge graph and determine node embedding vectors. The graph attention network comprises multiple stacked graph attention layers, each of which aggregates node information from neighboring entity nodes through an attention mechanism to update node features and quantify the interaction strength between entity nodes. The dual prediction head is used to determine the probability intensity and the comprehensive assessment of fracturability based on the node embedding vectors, and to filter double-sweet spot regions based on a preset threshold to determine the spatial distribution information of the double-sweet spots.
[0011] Secondly, this application provides a geological double sweet spot prediction device based on knowledge graph representation learning, comprising:
[0012] The data acquisition module is used to acquire multi-source heterogeneous geological data of the target work area; the multi-source heterogeneous geological data includes geological and engineering data;
[0013] The geological knowledge graph determination module is used to determine a geological knowledge graph based on the multi-source heterogeneous geological data and the model layer definition scheme. The geological knowledge graph includes entity nodes, attribute values, and relation edges. The model layer definition scheme is determined based on the multi-source heterogeneous geological data. The model layer definition scheme includes: an entity type list, an attribute list, and a relation type list.
[0014] The prediction module is used to determine the predicted outcome data based on the geological knowledge graph using a double-sweet spot prediction model. The outcome data includes the spatial distribution information of the double-sweet spots, probability intensity, and a comprehensive assessment result of fracturability. The double-sweet spot prediction model is trained based on the geological knowledge graph of the known outcome data, with the objective of minimizing the total loss function of the model. The total loss function includes classification loss, regression loss, and coupling constraint loss. The double-sweet spot prediction model uses a dual-prediction head model with a shared underlying GAT encoder. The GAT encoder employs a graph attention network. The graph attention network is used to learn geological knowledge representations from the geological knowledge graph and determine node embedding vectors. The graph attention network includes multiple stacked graph attention layers, each of which aggregates node information from neighboring entity nodes through an attention mechanism to update node features, thereby quantifying the interaction strength between entity nodes. The dual prediction head is used to determine the probability intensity and the comprehensive assessment result of fracturability based on the node embedding vectors, and to filter double-sweet spot regions based on a preset threshold to determine the spatial distribution information of the double-sweet spots.
[0015] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-described geological double sweet spot prediction method based on knowledge graph representation learning.
[0016] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the aforementioned geological double sweet spot prediction method based on knowledge graph representation learning.
[0017] Fifthly, this application provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned geological double sweet spot prediction method based on knowledge graph representation learning.
[0018] According to the specific embodiments provided in this application, the following technical effects are disclosed:
[0019] This application provides a geological double sweet spot prediction method, device, equipment, medium, and product based on knowledge graph representation learning. It transforms multi-source heterogeneous geological data into a geological knowledge graph containing rich semantic relationships and machine-understandable characteristics, solving the problems of data silos and missing connections. Based on this, it employs a graph attention network for representation learning, dynamically quantifying the influence weights of different geological entities on the target reservoir. By constructing a geological knowledge graph and introducing a graph neural network with coupling constraints, it achieves a paradigm shift in geological sweet spot prediction from data-driven to knowledge-driven, and collaboratively optimized. Furthermore, it designs a dual prediction head and a coupling constraint loss to achieve deep integration and collaborative optimization of geological attributes (reservoir capacity) and engineering attributes (fracturability), ultimately realizing geological double sweet spot prediction and providing accurate and reliable data for the efficient exploration and development of oil and gas reservoirs. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart of a geological double sweet spot prediction method based on knowledge graph representation learning;
[0022] Figure 2 This is a flowchart illustrating the overall technical process of a geological double sweet spot prediction method based on knowledge graph representation learning.
[0023] Figure 3 A comprehensive columnar section showing the structural location and lithological stratigraphy of the Shahejie Formation, Sha-3 Member;
[0024] Figure 4 A visual diagram of a geological knowledge graph subnetwork;
[0025] Figure 5 This is a schematic diagram of a three-layer graph attention network encoder architecture.
[0026] Figure 6 This is a schematic diagram of a double dessert prediction model;
[0027] Figure 7 This is a planar distribution map of the probability of geological sweet spots.
[0028] Figure 8 A planar distribution of the probability of scoring engineering desserts;
[0029] Figure 9 A schematic diagram illustrating the impact of hyperparameters on the GAT architecture;
[0030] Figure 10This is a schematic diagram for verifying the effectiveness of the coupling constraint mechanism;
[0031] Figure 11 A schematic diagram illustrating the impact of data quality on map completeness;
[0032] Figure 12 This is a structural diagram of a geological double sweet spot prediction device based on knowledge graph representation learning.
[0033] Figure 13 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0034] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0035] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0036] In one exemplary embodiment, such as Figure 1 As shown, a geological double sweet spot prediction method based on knowledge graph representation learning is provided, including:
[0037] Step 100: Obtain multi-source heterogeneous geological data for the target work area. Multi-source heterogeneous geological data includes geological and engineering data;
[0038] Step 200: Determine the geological knowledge graph based on the multi-source heterogeneous geological data and the model layer definition scheme. The geological knowledge graph includes entity nodes, attribute values, and relation edges; the model layer definition scheme is determined based on the multi-source heterogeneous geological data; the model layer definition scheme includes: an entity type list, an attribute list, and a relation type list.
[0039] Among them, based on the multi-source heterogeneous geological data and model layer definition scheme, a geological knowledge graph is determined, specifically including:
[0040] Standardization preprocessing is performed on multi-source heterogeneous geological data to obtain a standardized multi-source dataset. Standardization preprocessing includes cleaning, correction, and coordinate alignment.
[0041] Based on the standardized multi-source dataset, entity types are determined according to geological genetic analysis and hydrocarbon accumulation models of the target work area, and an entity type list is determined based on the entity types; the attributes of each type of entity are determined according to the entity types, and an attribute list is determined based on the attributes; the relationship types are determined according to the semantic relationships between entities, and a relationship type list is determined based on the relationship types; and the pattern layer definition scheme is determined according to the entity type list, attribute list, and relationship type list.
[0042] Based on the schema layer definition scheme, attribute values are extracted from the standardized multi-source dataset, and each entity node is instantiated; based on spatial location, geological patterns, and attribute correlation, relational edges between entity nodes are instantiated to form entity-relation-entity triples; and the geological knowledge graph is determined based on the triples.
[0043] Step 300: Employ a double-sweet spot prediction model to determine the predicted outcome data based on the geological knowledge graph. The outcome data includes the spatial distribution information of the double sweet spots, probability intensity, and comprehensive fracturability assessment results. The double-sweet spot prediction model is trained based on the geological knowledge graph of the known outcome data, with the objective of minimizing the total loss function. The total loss function includes classification loss, regression loss, and coupling constraint loss. The double-sweet spot prediction model uses a dual-prediction head model with a shared underlying GAT encoder. The GAT encoder employs a Graph Attention Network (GAT), which is used to learn geological knowledge representations from the geological knowledge graph and determine node embedding vectors. The GAT network consists of multiple stacked GAT layers, each of which aggregates node information from neighboring entity nodes through an attention mechanism to update node features and quantify the interaction strength between entity nodes. The dual prediction heads determine the probability intensity and comprehensive fracturability assessment results based on the node embedding vectors and filter double sweet spot regions based on preset thresholds to determine the spatial distribution information of the double sweet spots.
[0044] The process of updating node features includes:
[0045] .
[0046] .
[0047] .
[0048] in, For the first The first layer of the attention layer in the layer diagram The updated node feature vector of each entity node; It is a non-linear activation function; For the first The set of neighboring nodes of a node; For the first The entity node for the first Attention coefficients of each neighboring node; For the first In a layered network, the first Feature vectors of each neighboring node; The original attention score; To Exponentiation; It is the transpose of a trainable parameter vector; For the first In a layered network, the first Feature vectors of each entity node; For splicing operations; These are trainable parameters; , , All are serial numbers.
[0049] The first prediction head in a dual prediction head uses... The function is a fully connected network with an activation function; the second predictor in the dual predictor heads uses a regression network.
[0050] Regression networks are used to identify target reservoir nodes based on their node embedding vectors. The embeddings of neighboring nodes are weighted and aggregated to output a comprehensive assessment result of fracturability. The formula for weighted aggregation is:
[0051] .
[0052] in, The target reservoir node is obtained after weighted aggregation calculation; and All are learnable aggregate weights; For the target reservoir node; To and The first one directly connected on the geological knowledge map Embedding vectors of neighboring nodes.
[0053] The expression for the total loss function is:
[0054] .
[0055] .
[0056] in, This is the total loss function; The loss function for the geological sweet spot prediction task; For probability intensity; The true label for geological desserts; The loss function for the engineering dessert prediction task; The predicted value for engineered desserts; This is the actual value of the engineered dessert; For constraint strength; This is the loss due to coupling constraints; The total number of samples in the batch; This is the tolerance threshold; For the first The predicted probability of geological sweetness for each sample; For the first Predicted values for engineered desserts for each sample; For serial numbers.
[0057] As an optional implementation, the geological double sweet spot prediction method based on knowledge graph representation learning further includes:
[0058] Monte Carlo The method involves multiple samplings, and the regional confidence level is calculated based on the variance of the predicted results data.
[0059] This application first standardizes and preprocesses multi-source heterogeneous geological data, including well logging, seismic, core, and production data. Based on a geological conceptual model, it clearly defines core entities such as reservoirs, faults, and source rocks, along with their attributes, and formally defines semantic relationships such as "control" and "influence," thereby instantiating and constructing a computable geological knowledge graph. Subsequently, it utilizes a graph attention network to perform deep representation learning on the geological knowledge graph, quantifying the interaction strength between geological entities through an attention mechanism to generate node embedding vectors containing geological semantics. Building upon this, a dual-predictor model with a shared encoder is constructed to predict the probability of geological sweet spots and the score of engineering sweet spots (i.e., the comprehensive assessment result of fracturing capability). Innovatively, a coupling constraint term (i.e., coupling constraint loss) is introduced into the loss function, forcing the dual-sweet spot prediction model to collaboratively optimize both objectives. The trained dual-sweet spot prediction model can predict the entire work area, outputting a comprehensive dual-sweet spot distribution map (including spatial distribution information, probability intensity, and comprehensive assessment result of fracturing capability), and providing node-level and region-level prediction confidence based on attention weights and Monte Carlo sampling. Finally, the entire process is deployed as a closed-loop, optimizable intelligent system that continuously iterates and evolves using new data. This application realizes a paradigm shift from "data-driven" to "knowledge-driven, collaborative optimization," solving the problems of fragmented data, biased predictions, and weak interpretability in current traditional methods. It provides a precise, reliable, and interpretable intelligent solution for the efficient exploration and development of complex oil and gas reservoirs.
[0060] In short, this application aims to solve the problems of data fragmentation, biased prediction, and weak interpretability of traditional methods by constructing a structured geological knowledge graph and using graph attention networks for representation learning, ultimately achieving collaborative prediction and optimization of geological and engineering sweet spots.
[0061] The overall technical process of this application is as follows: Figure 2 As shown, the specific steps include:
[0062] Step S1: Collection and standardization preprocessing of multi-source heterogeneous geological data.
[0063] The input for this step is all available raw geological and engineering data for the target work area. Specifically, it involves: systematically collecting well logging data from all drilled wells within the target work area, 3D seismic data volumes and interpretation results covering the area, core analysis data from key wells, and historical production dynamic data; performing standardized preprocessing on the collected multi-source heterogeneous geological data, including environmental correction and normalization of well logging data, denoising and scaling of seismic attributes, interpolating or correlating core data to continuous depth segments, and unifying all data to the same spatial reference coordinate system and depth system. The output of this step is a standardized multi-source dataset after cleaning, correction, and coordinate alignment. Standardization preprocessing may also include: depth matching, environmental correction, data normalization, and unifying the spatial coordinate benchmark.
[0064] Step S2: Define entities, attributes, and relationships in the geological knowledge graph.
[0065] The input for this step is the standardized multi-source dataset obtained in step S1. Specifically, based on geological genetic analysis and the hydrocarbon accumulation model of the study area, define the core entity types constituting the geological knowledge graph, including at least: reservoirs, source rocks, caprocks, faults, fracture zones, and tectonic units; define key attributes for each entity type, for example, reservoir entities define porosity, permeability, thickness, and oil saturation attributes; fracture zone entities define fracture density, aperture, and strike attributes; and fault entities define fault displacement, sealing index, and occurrence attributes; define semantic relationship types between entities, including at least: "control" (e.g., faults control reservoirs), "influence" (e.g., source rocks influence oil content), "caprock," "connectivity," and "containment." The output of this step is a complete geological knowledge graph model layer definition scheme, including a list of entity types, an attribute list, and a relationship type list.
[0066] Step S3: Instantiation and construction of the geological knowledge graph.
[0067] The inputs to this step are the standardized dataset from step S1 and the schema layer definition scheme from step S2. Specifically, based on the schema layer definition scheme in S2, attribute values for each specific geological object are extracted and calculated from the standardized multi-source dataset, and each entity node is instantiated. Relationship edges between entity nodes are instantiated based on spatial location, geological patterns, and attribute correlations, forming "entity-relationship-entity" triples. All triples are imported into a graph database (such as Neo4j) to construct a queryable and computable geological knowledge graph. During the construction process, graph quality is checked to ensure there are no isolated nodes and that key relationships are connected. The output of this step is an instantiated geological knowledge graph containing specific nodes, attribute values, and relationship edges.
[0068] Step S4: Geological knowledge representation learning based on graph attention network.
[0069] The input for this step is the geological knowledge graph constructed in step S3. Specifically, the following steps are performed: The attribute vector corresponding to the attributes of each entity node in the geological knowledge graph is used as the initial feature; a graph attention network (GNA) is used as the encoder, consisting of multiple stacked GNA layers. Each layer aggregates neighbor node information through an attention mechanism to update the feature representation of the central node; the hidden layer dimension of the GNA is adjusted (e.g., 256→128→64) to refine higher-order features layer by layer; node-level pre-training tasks are designed, such as classifying nodes as "sweet spot / non-sweet spot" using known well production labels, or randomly masking some relational edges for link prediction, to drive the model to learn deep association patterns in the graph. The output of this step is a trained GAT encoder model, which can map any geological entity node into a low-dimensional, dense, semantically rich vector representation (i.e., a node embedding vector). In other words, the GNA is used to aggregate and encode the attribute features of each entity node in the geological knowledge graph into a low-dimensional, dense vector representation; the encoder is pre-trained using node classification or link prediction tasks.
[0070] Step S5: Construction and co-training of the double sweet spot prediction model.
[0071] The input for this step is the node embedding vector obtained in step S4. Specifically, it involves constructing a dual-predictor model sharing a low-level GAT encoder. The first predictor is the geological sweet spot predictor, which is based on... The function is a fully connected network with an activation function. The input is the embedding vector of the target reservoir node, and the output is the probability value of it being a geological sweet spot (with high reservoir capacity), i.e., the probability strength. The second prediction head is the engineering sweet spot prediction head, which is a regression network. In addition to the embedding of the target reservoir node as input, it also aggregates the embedding information of its adjacent key engineering entities such as faults and fracture zones, and outputs a comprehensive assessment result characterizing its fracturing capability. This refers to the engineering dessert score. During training, a coupling constraint loss is introduced, which penalizes samples with a much higher probability of geological dessert scores than engineering dessert scores, forcing the model to learn features that satisfy both conditions simultaneously. The total loss of the model is the sum of the classification loss, regression loss, and coupling constraint loss. The output of this step is a trained end-to-end model with the ability to co-predict both dessert scores.
[0072] Step S6: Model prediction, result generation and confidence assessment.
[0073] The inputs for this step are the prediction model trained in step S5 (i.e., the double-sweet spot prediction model) and the complete knowledge graph (i.e., the geological knowledge graph) of the target work area. Specifically, the double-sweet spot prediction model is used to perform forward inference on all reservoir nodes in the entire work area, obtaining the data for each node in batches. and According to a preset threshold (e.g., set...); >0.7 and >0.65) to filter out "double sweet spot" regions; using the attention weight distribution of the GAT encoder, calculate the node-level prediction confidence to reveal the key neighbor entities affecting the prediction and their contribution; Monte Carlo simulation is then used. Uncertainty quantification methods are used to perform multiple samplings during inference, and the regional confidence level is calculated based on the variance of the prediction results. The output of this step is complete prediction data containing the spatial distribution of the two sweet spots, probability strength, and confidence assessment.
[0074] Step S7: Predictive system deployment and closed-loop optimization iteration.
[0075] The inputs for this step are the prediction data from step S6 and newly acquired drilling and production feedback data. Specifically, the entire process from steps S1 to S6 is automated and encapsulated into a software module or API that can provide services externally; the sweet spot probability volume and optimized target points are integrated into the geological interpretation and drilling design platform; and model performance monitoring indicators are set. When new drilling validation results are available or the performance of the double-sweet spot prediction model declines, closed-loop optimization is initiated: the new data is used as new samples, or the knowledge graph structure (step S3), feature engineering (step S1 / step S2), and model hyperparameters (step S4 / step S5) are adjusted based on error analysis, and the model is retrained to achieve continuous iteration and performance improvement of the system. This step ultimately outputs a software module that encompasses the entire process and can be provided externally.
[0076] This application achieves a paradigm shift in geological sweet spot prediction from data-driven to knowledge-driven and collaborative optimization by constructing a structured geological knowledge graph and introducing a graph neural network with coupling constraints. First, multi-source heterogeneous geological data (geological and engineering data) are transformed into a machine-understandable knowledge graph containing rich semantic relationships, fundamentally solving the problems of data silos and missing connections in traditional methods, and providing a solid interpretability foundation for model prediction. Based on this, a graph attention network is used for representation learning, which can dynamically quantify the influence weights of different geological entities on the target reservoir. Its decision-making process is transparent and traceable. For example, when the attention weight reaches 0.85, it can be clearly revealed that the enrichment of a certain sweet spot is mainly attributed to the lateral sealing effect of a specific fault, thereby greatly enhancing geologists' trust and acceptance of the intelligent prediction results. Crucially, this application innovatively designs a dual-prediction-head architecture and coupled constraint loss, forcing deep integration and synergistic optimization of geological attributes (reservoir capacity) and engineering attributes (fractureability) during the training phase. This ensures that the final sweet spot target area not only has great resource potential but also good developability, effectively avoiding drilling risks caused by incompatible engineering conditions and achieving precise positioning of "recoverable reserves" rather than just "geological reserves." Furthermore, the entire method is designed as a complete and standardized technical process encompassing "data-knowledge-model-application-optimization," with clear inputs and outputs and specific operations at each step, ensuring high reproducibility and industrialization potential. Finally, the system's built-in closed-loop optimization mechanism enables it to continuously absorb new drilling and production data, iteratively evolving the model and knowledge base, thus forming an intelligent exploration decision support system with adaptive and continuous learning capabilities. This significantly improves the accuracy, efficiency, and practical application value of sweet spot prediction in complex oil and gas reservoirs.
[0077] Example 1:
[0078] This embodiment takes a complex fault-block-lithological reservoir in a certain area as the target work area. The structural location and comprehensive stratigraphic columnar section of the Shahejie Formation, Member 3, in the target work area are shown below. Figure 3 As shown, this layer is a typical steep-slope deltaic deposit, with interbedded sandstone and mudstone, well-developed faults, and strong reservoir heterogeneity, resulting in low accuracy of traditional single-attribute prediction methods. This embodiment aims to comprehensively demonstrate how to use the method described in this application to accurately predict the double sweet spot region in this work area, which possesses both superior reservoir capacity and good fracturing capability, providing a direct basis for horizontal well trajectory optimization and fracturing design. The specific steps are as follows:
[0079] Step S1: Collection and standardization preprocessing of multi-source heterogeneous geological data.
[0080] The system collected conventional logging curves (GR (Gamma Ray, natural gamma), AC (Acoustic Log), DEN (Density Log, compensated density), CNL (Compensated Neutron Log), RD (Deep Resistivity), etc.) and microresistivity scanning imaging logging (Formation MicroScanner Image) from 42 wells within the work area. Data; collected 800km 2 The data includes 3D seismic data volumes and 12 derived attribute volumes such as coherence, curvature, and wave impedance; core experimental data totaling 35 meters from 15 key wells; and production data from nearly 10 years of stratified systems. Preprocessing includes environmental correction of logging curves (e.g., wellbore enlargement correction), using the following formula:
[0081] .
[0082] in, This is the corrected time difference value for sound waves; These are uncorrected sonic transit time logging values; For correction factors; Well diameter; This represents the drill bit diameter.
[0083] Z-score normalization is performed on all continuous data:
[0084] .
[0085] in, The data has been standardized using Z-score. The original geological data is to be standardized; Let X be the arithmetic mean of the data sequence to be standardized.
[0086] Core analysis data and well logging data were rigorously matched through depth relocation; finally, all data were unified to the WGS-84 coordinate system and elevation-depth datum. Key preprocessing parameters are shown in Table 1.
[0087] Table 1 Key parameters for data preprocessing
[0088]
[0089] Step S2: Define entities, attributes, and relationships in the geological knowledge graph.
[0090] Define 6 types of entities and their attributes: (porosity) Penetration rate ,thickness Oil saturation ), (thickness Overcoming pressure ), (Dislocation) Shale Gouge Ratio ), attitude, fracture zone (fracture density) Average opening ), constructing high points (overflow point elevation, closed area) ), Jingyuan rock (Total Organic Carbon), ).
[0091] Define the core relationship: →Control→Reservoir, fault→Cut→Mudstone interlayer, fracture zone→Developed at→Reservoir, structural high point→Trap→Reservoir.
[0092] Step S3: Instantiation and construction of the geological knowledge graph.
[0093] by Using a grid as the basic unit, the work area was discretized into 17,500 "reservoir" entity nodes, each assigned an attribute value based on its location. Eighty-six "fault" entities were instantiated from the seismic interpretation results. Finally, in A virtualized knowledge graph containing approximately 17,600 nodes and over 41,000 relational edges was constructed using the graph database. A visualization of a local subnetwork of the constructed geological knowledge graph is shown below. Figure 4 As shown, it clearly presents the topological relationships between entities such as faults and reservoirs. The weights of the relationship edges are assigned according to geological rules, such as the weight of the "control" relationship. Based on fault sealing ( Calculation of reservoir contact relationship:
[0094]
[0095] in, This is the contact type index, ranging from 0 to 1.
[0096] Step S4: Learning geological knowledge representation based on GAT.
[0097] The initial features of each node are 8-dimensional attribute vectors. A 3-layer GAT encoder is employed (the architecture and information flow principle of this 3-layer GAT encoder are as follows...). Figure 5 As shown), the hidden layer dimensions are 256-128-64. Node feature updates follow the core formula of GAT:
[0098] .
[0099] .
[0100] in, and All of these are trainable parameters. This indicates a splicing operation.
[0101] .
[0102] Using the Adam optimizer ( The model was trained for 200 epochs with a learning rate of 1e-4. After training, the model achieved a node classification accuracy of 88.7% on the validation set. The exponential decay rate estimated by the first moment, The exponential decay rate is the second moment estimate.
[0103] Step S5: Construction and co-training of the dual-sweetness model.
[0104] A dual prediction head is added to the pre-trained GAT encoder. A schematic diagram of the overall structure of this dual-sweetness prediction model is shown below. Figure 6 As shown, it includes a shared GAT encoder, two prediction heads (geological and engineering), and a coupling constraint module. The geological sweet spot head is a fully connected layer + Activation function, output The engineering sweet spot first targets the reservoir nodes. and its neighboring nodes (such as faults) Weighted aggregation is performed on the embedding of ) :
[0105] .
[0106] Then output .
[0107] The model's total loss function is:
[0108] .
[0109] Coupling constraint loss Defined as:
[0110] .
[0111] in, For constraint strength, This represents the tolerance threshold. See Table 2 for details on hyperparameters.
[0112] Table 2 Key Hyperparameter Settings for the Double Sweet Spot Prediction Model
[0113]
[0114] Step S6: Model prediction, result generation and confidence assessment.
[0115] The trained model is used to predict the Es3 layer of the entire work area. Figure 7 ( The planar distribution map shows that high-value areas (>0.75) are distributed in strips on the footwall of the main control fault, which coincides with the sedimentary facies zone. Figure 8 ( The planar distribution map shows that high-value areas (>0.70) are located in regions with high brittle mineral content and fracture development zones. Based on both factors, four "double sweet spot" target areas were selected. Post-drilling verification was performed on the new exploratory well PL19-3-XX in target area A. This well achieved a daily oil equivalent of 102 cubic meters at the predicted formation, far exceeding the average production of surrounding older wells (approximately 30 cubic meters / day).
[0116] Step S7: Predictive system deployment and closed-loop optimization iteration.
[0117] The above process is encapsulated as a RESTful API service for deployment. The system automatically monitors new drilling data, and automatically triggers an optimization iteration process when the accumulated new well data exceeds 5 wells or the model's accuracy on the latest validation set drops by more than 2%.
[0118] This embodiment fully verifies that the method provided in this application can systematically complete the entire process from multi-source data to double-sweet spot target area recommendation. The generated prediction results are not only highly accurate and geologically clear, but also have excellent engineering applicability, providing a brand-new technical means for the precise exploration and efficient development of complex oil and gas reservoirs.
[0119] Example 2: Hyperparameter sensitivity analysis and model robustness verification.
[0120] This embodiment aims to systematically evaluate the impact of key hyperparameters and input data quality on the performance of a double-sweet spot prediction model, verify the robustness of the method, and provide clear guidance on parameter configuration and data quality for practical applications. Experiments were conducted on a unified oilfield dataset from a specific region, with the baseline model being a "3-layer GAT + 8-head attention" model using coupling constraint loss.
[0121] (1) The influence of hyperparameters on graph neural network architecture.
[0122] The system tested the combined effects of different GAT layer numbers (2, 3, 4), attention head numbers (4, 8, 16), and hidden layer dimensions (baseline: 256-128-64, comparison: 128-64-32). Figure 9As shown, when the number of layers increases from 2 to 3, the F1 score on the test set significantly improves from 86.1% to 89.3%. This is because the 3-layer network can effectively capture multi-hop geological dependencies such as "source rock → migration path → reservoir". However, the 4-layer network causes the F1 score to drop to 87.9% and the training time to increase by 52%, indicating overfitting and reduced efficiency. In the attention head number experiment ( Figure 10 Eight attention heads strike the best balance between F1 score (89.3%) and training time; 16 attention heads only bring a 0.2% performance improvement, but increase the time by 40%, resulting in a low cost-benefit ratio. Reducing the hidden layer dimension reduces the number of model parameters by 65%, but also decreases the F1 score by 3.5%, indicating that the baseline dimension is necessary for fully representing complex geological features.
[0123] (2) Verification of the effectiveness of the coupling constraint mechanism.
[0124] To quantify the effect of coupling constraint loss, three training strategies were compared: A (no coupling constraint, i.e.) B (has coupling constraints) Strategy B (this application) achieved a "double sweet spot accuracy rate" (i.e., the proportion of predictions that are double sweet spots and are correctly verified by drilling) of 82.4%, significantly higher than A (64.7%) and C (71.2%). Strategy A's model produced a large number of "high-precision" predictions. Low The "false dessert" of strategy C, while strategy C suffers from a small and contradictory intersection region due to the lack of coordination between the two predictor heads. Strategy B's coupling constraints successfully guide the model towards... and The solution space with uniform height.
[0125] Table 3 Comparison of Model Performance under Different Training Strategies
[0126]
[0127] (3) The impact of data quality on the completeness of knowledge graphs.
[0128] Three data defect scenarios were simulated: I (randomly missing 20% of the "fault-control-reservoir" relationship edges), II (adding 20% Gaussian noise to node features), and III (missing "source rock" entities and all their relationships). Figure 11As shown, scenarios I and II resulted in a 4.1% and 5.8% decrease in the F1 score, respectively, indicating that missing relationships and feature noise directly impact model performance. However, the model did not completely fail, demonstrating the robustness of GAT to some extent. Scenario III, on the other hand, caused a sharp drop in the F1 score of 11.3%, highlighting how the absence of core geological entities severely impairs the semantic integrity of the knowledge graph, thereby significantly reducing its predictive ability. This experiment underscores the importance of accurately defining and constructing the knowledge graph in steps S2 and S3.
[0129] (4) Comparison with other machine learning methods.
[0130] This application was compared with three benchmark methods on the same test set: 1) Random Forest, ), taking the original attributes of the nodes as input; 2) Multilayer Perceptron, ), input the same 3) Standard Graph Convolutional Network (GCN) This application significantly outperforms other applications in both double-sweetness compliance (82.4%) and interpretability (which provides attention weighting). (58.8%) (61.5%) and (76.1%). Although it has some graph structure awareness, its equal neighbor aggregation mechanism cannot be like... Such dynamic focusing on key geological factors results in lower accuracy and interpretability compared to this application.
[0131] In summary, this embodiment demonstrates through detailed experimental analysis that the method proposed in this application achieves the recommended "3-layer" performance. Under the "head attention + coupling constraint" configuration, not only was optimal prediction performance achieved, but it also demonstrated robustness to common data defects. Furthermore, the research results rigorously demonstrate the crucial role of the coupling constraint mechanism and high-quality knowledge graph construction in achieving collaborative optimization prediction between geology and engineering, providing a solid and reliable basis for the practical application of the method.
[0132] In one exemplary embodiment, such as Figure 12 As shown, a geological double sweet spot prediction device based on knowledge graph representation learning is provided, including:
[0133] The data acquisition module is used to acquire multi-source heterogeneous geological data of the target work area; the multi-source heterogeneous geological data includes geological and engineering data.
[0134] The geological knowledge graph determination module is used to determine a geological knowledge graph based on the multi-source heterogeneous geological data and the model layer definition scheme. The geological knowledge graph includes entity nodes, attribute values, and relation edges. The model layer definition scheme is determined based on the multi-source heterogeneous geological data. The model layer definition scheme includes: an entity type list, an attribute list, and a relation type list.
[0135] The prediction module is used to determine the predicted outcome data based on the geological knowledge graph using a double-sweet spot prediction model. The outcome data includes the spatial distribution information of the double-sweet spots, probability intensity, and a comprehensive assessment result of fracturability. The double-sweet spot prediction model is trained based on the geological knowledge graph of the known outcome data, with the objective of minimizing the total loss function of the model. The total loss function includes classification loss, regression loss, and coupling constraint loss. The double-sweet spot prediction model uses a dual-prediction head model with a shared underlying GAT encoder. The GAT encoder employs a graph attention network. The graph attention network is used to learn geological knowledge representations from the geological knowledge graph and determine node embedding vectors. The graph attention network includes multiple stacked graph attention layers, each of which aggregates node information from neighboring entity nodes through an attention mechanism to update node features, thereby quantifying the interaction strength between entity nodes. The dual prediction head is used to determine the probability intensity and the comprehensive assessment result of fracturability based on the node embedding vectors, and to filter double-sweet spot regions based on a preset threshold to determine the spatial distribution information of the double-sweet spots.
[0136] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 13 As shown, this computer device includes a processor, memory, input / output (I / O) interfaces, and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores geological double-sweet spot prediction data based on knowledge graph representation learning. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the geological double-sweet spot prediction method based on knowledge graph representation learning.
[0137] Those skilled in the art will understand that Figure 13The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0138] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0139] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0140] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0141] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0142] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0143] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logic devices, etc., and are not limited to these.
[0144] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0145] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A geological double sweet spot prediction method based on knowledge graph representation learning, characterized in that, include: Acquire multi-source heterogeneous geological data of the target work area; The multi-source heterogeneous geological data includes geological and engineering data; Based on the aforementioned multi-source heterogeneous geological data and model layer definition scheme, a geological knowledge graph is determined; the geological knowledge graph includes entity nodes, attribute values, and relation edges. The model layer definition scheme is determined based on the multi-source heterogeneous geological data; The schema layer definition scheme includes: a list of entity types, a list of attributes, and a list of relationship types; A double-sweet spot prediction model is employed to determine the predicted outcome data based on the geological knowledge graph. This outcome data includes the spatial distribution information of the double-sweet spots, probability intensity, and a comprehensive assessment of fracturability. The double-sweet spot prediction model is trained based on the geological knowledge graph containing the known outcome data, with the objective of minimizing the total loss function. The total loss function includes classification loss, regression loss, and coupling constraint loss. The double-sweet spot prediction model uses a dual-prediction head model with a shared underlying GAT encoder. The GAT encoder employs a graph attention network. The graph attention network is used to learn geological knowledge representations from the geological knowledge graph and determine node embedding vectors. The graph attention network comprises multiple stacked graph attention layers, each of which aggregates node information from neighboring entity nodes through an attention mechanism to update node features and quantify the interaction strength between entity nodes. The dual prediction head is used to determine the probability intensity and the comprehensive assessment of fracturability based on the node embedding vectors, and to filter double-sweet spot regions based on a preset threshold to determine the spatial distribution information of the double-sweet spots.
2. The geological double sweet spot prediction method based on knowledge graph representation learning according to claim 1, characterized in that, Based on the aforementioned multi-source heterogeneous geological data and model layer definition scheme, a geological knowledge graph is determined, specifically including: The multi-source heterogeneous geological data are subjected to standardization preprocessing to obtain a standardized multi-source dataset; the standardization preprocessing includes: cleaning, correction and coordinate alignment. Based on the standardized multi-source dataset, entity types are determined according to geological genetic analysis and hydrocarbon accumulation models of the target work area, and an entity type list is determined based on the entity types. The attributes of each type of entity are determined based on the entity type, and an attribute list is determined based on the attributes. The relationship type is determined based on the semantic relationship between entities, and a list of relationship types is determined based on the relationship type; Determine the schema layer definition scheme based on the entity type list, attribute list, and relation type list; Based on the pattern layer definition scheme, attribute values are extracted according to the standardized multi-source dataset, and each entity node is instantiated. Based on spatial location, geological patterns, and attribute correlation, instantiate relational edges between entity nodes to form entity-relation-entity triples; A geological knowledge map is determined based on the aforementioned ternary set.
3. The geological double sweet spot prediction method based on knowledge graph representation learning according to claim 1, characterized in that, The first prediction head in a dual prediction head uses... The function is a fully connected network with an activation function; the second predictor in the dual predictor head uses a regression network; The regression network is used to target reservoir nodes based on the node embedding vector. The embeddings of neighboring nodes are weighted and aggregated to output a comprehensive assessment result of fracturability. The formula for weighted aggregation is: ; in, The target reservoir node is obtained after weighted aggregation calculation; and All are learnable aggregate weights; For the target reservoir node; To and The first one directly connected on the geological knowledge map Embedding vectors of neighboring nodes.
4. The geological double sweet spot prediction method based on knowledge graph representation learning according to claim 1, characterized in that, The expression for the total loss function is: ; ; in, This is the total loss function; The loss function for the geological sweet spot prediction task; For probability intensity; The actual label for the dessert; The loss function for the engineering dessert prediction task; The predicted value for engineered desserts; This is the actual value of the engineered dessert; For constraint strength; This is the loss due to coupling constraints; The total number of samples in the batch; This is the tolerance threshold; For the first The predicted probability of geological sweetness for each sample; For the first Predicted values for engineered desserts for each sample; For serial numbers.
5. The geological double sweet spot prediction method based on knowledge graph representation learning according to claim 1, characterized in that, The geological double sweet spot prediction method based on knowledge graph representation learning also includes: Monte Carlo The method involves multiple samplings, and the regional confidence level is calculated based on the variance of the predicted results data.
6. The geological double sweet spot prediction method based on knowledge graph representation learning according to claim 1, characterized in that, The process of updating node features includes: ; ; ; in, For the first The first layer of the attention layer in the layer diagram The updated node feature vector of each entity node; It is a non-linear activation function; For the first The set of neighboring nodes of a node; For the first The entity node for the first Attention coefficients of each neighboring node; For the first In a layered network, the first Feature vectors of each neighboring node; The original attention score; To Exponentiation; It is the transpose of a trainable parameter vector; For the first In a layered network, the first Feature vectors of each entity node; For splicing operations; These are trainable parameters; , , All are serial numbers.
7. A geological double sweet spot prediction device based on knowledge graph representation learning, characterized in that, include: The data acquisition module is used to acquire multi-source heterogeneous geological data of the target work area; The multi-source heterogeneous geological data includes geological and engineering data; The geological knowledge graph determination module is used to determine the geological knowledge graph based on the multi-source heterogeneous geological data and the model layer definition scheme; the geological knowledge graph includes entity nodes, attribute values, and relation edges; The model layer definition scheme is determined based on the multi-source heterogeneous geological data; The schema layer definition scheme includes: a list of entity types, a list of attributes, and a list of relationship types; The prediction module is used to determine the predicted outcome data based on the geological knowledge graph using a double-sweet spot prediction model. The outcome data includes the spatial distribution information of the double-sweet spots, probability intensity, and a comprehensive assessment result of fracturability. The double-sweet spot prediction model is trained based on the geological knowledge graph of the known outcome data, with the objective of minimizing the total loss function of the model. The total loss function includes classification loss, regression loss, and coupling constraint loss. The double-sweet spot prediction model uses a dual-prediction head model with a shared underlying GAT encoder. The GAT encoder employs a graph attention network. The graph attention network is used to learn geological knowledge representations from the geological knowledge graph and determine node embedding vectors. The graph attention network includes multiple stacked graph attention layers, each of which aggregates node information from neighboring entity nodes through an attention mechanism to update node features, thereby quantifying the interaction strength between entity nodes. The dual prediction head is used to determine the probability intensity and the comprehensive assessment result of fracturability based on the node embedding vectors, and to filter double-sweet spot regions based on a preset threshold to determine the spatial distribution information of the double-sweet spots.
8. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the geological double sweet spot prediction method based on knowledge graph representation learning as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the geological double sweet spot prediction method based on knowledge graph representation learning as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the geological double sweet spot prediction method based on knowledge graph representation learning as described in any one of claims 1-6.
Citation Information
Patent Citations
CN120338171A
CN121413818A