Reservoir prediction method and device based on mixed probability principal component analysis fusion attributes

By combining hybrid probabilistic principal component analysis and semi-supervised K-means clustering, the problem of insufficient quantification in seismic attribute fusion is solved, achieving efficient and accurate reservoir characteristic prediction and uncertainty quantification, adapting to complex geological conditions, and reducing computational complexity and technical threshold.

CN121679746APending Publication Date: 2026-03-17SHANGHAI BRANCH CHINA OILFIELD SERVICES
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202610016554.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-03-17

AI Technical Summary

Technical Problem

Existing seismic attribute fusion methods suffer from insufficient qualitative or quantitative reliability in reservoir prediction, lack systematic accuracy control and closed-loop feedback mechanisms, and have poor adaptability to complex geological conditions, resulting in multiple solutions and making it difficult to achieve efficient and accurate reservoir characteristic prediction.

Method used

Hybrid probabilistic principal component analysis (MPPCA) is used to reduce the dimensionality and fuse seismic data. A mapping model is established by combining semi-supervised K-means clustering. Through multi-Gaussian subspace probabilistic modeling and well point label guidance, the prediction of reservoir characteristics across the entire domain is achieved. A closed-loop process is formed through cross-validation and parameter correction.

Benefits of technology

It significantly improves the accuracy and stability of reservoir prediction, reduces computational complexity, is suitable for exploration areas with small amounts of data, provides uncertainty quantification capabilities, lowers the technical threshold, and adapts to complex geological conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121679746A_ABST
    Figure CN121679746A_ABST
Patent Text Reader

Abstract

The invention relates to a reservoir prediction method and device based on mixed probability principal component analysis fusion attributes, and relates to the technical field of oil and gas resource exploration and development, and the method comprises the steps: obtaining an original three-dimensional seismic data volume, well point logging data and known reservoir feature data of a target research area; preprocessing the original three-dimensional seismic data volume and then constructing an attribute matrix, and modeling the attribute matrix by using mixed probability principal component analysis to obtain a fusion attribute matrix; taking the fusion attribute matrix as an input feature, taking known reservoir feature data as a target variable, and establishing a mapping model by adopting semi-supervised K-means clustering in combination with well point label guidance classification; and calculating all data in the fusion attribute matrix of the research area according to the mapping model to obtain a global reservoir feature prediction result. The reservoir prediction method provided by the invention has the advantages of high reservoir prediction precision, high reliability and uncertainty quantification capability.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of oil and gas resource exploration and development, and particularly relates to a reservoir prediction method and device based on mixed probability principal component analysis and fused attributes. BACKGROUND

[0002] In the process of oil and gas resource exploration and development, reservoir prediction is one of the core links, and its accuracy directly affects the success rate of exploration and development benefits. Seismic exploration technology as an important means to obtain underground geological information, through extracting seismic attributes to invert reservoir characteristics has become the mainstream method in the industry. Seismic attributes can reflect the physical property differences of underground rock layers, such as amplitude attributes related to reservoir lithology and porosity, frequency attributes can indicate reservoir changes, and phase attributes can help identify stratigraphic interfaces.

[0003] However, the response of a single seismic attribute to reservoir characteristics has limitations, and is affected by the complexity of geological conditions, data acquisition noise and observation errors, and is prone to multiple solutions, making it difficult to accurately depict reservoir distribution. For this reason, seismic multi-attribute fusion technology has emerged, which integrates the complementary information of multiple attributes, reduces the multiple solutions, and improves the prediction reliability.

[0004] At present, common multi-attribute fusion methods include weighted fusion, RGB fusion, multiple linear regression fusion and machine learning-based fusion, but these methods have obvious shortcomings: weighted fusion relies on artificial experience to set weights, which is highly subjective, and unreasonable weight distribution can lead to biased fusion results; RGB fusion can only achieve qualitative or semi-quantitative analysis, and cannot meet the needs of quantitative prediction of reservoir characteristics; multiple linear regression fusion assumes that the attribute and reservoir characteristics are linearly related, which is difficult to adapt to nonlinear mapping under complex geological conditions; machine learning-based fusion (such as support vector machine, neural network) can handle nonlinear problems, but requires a large amount of sample data for training, and the model is complex and computationally expensive, with poor adaptability to small data research areas.

[0005] In the prior art, although probability principal component analysis (PPCA) or single subspace dimension reduction method is applied to seismic data processing, but for reservoir prediction scenarios, there is no complete standardized process formed - especially in the key links of multi-subspace weight distribution, probability threshold determination of principal component selection, uncertainty mapping of fused attributes and reservoir characteristics, and quantitative verification of prediction accuracy, lack of systematic solutions, resulting in the inability to fully exploit the ability to depict reservoir characteristics under complex geological conditions.

[0006] CN112346117A discloses a reservoir feature prediction method and device based on seismic attribute fusion. Instead of fusing the entire attribute map, the scheme first calculates the correlation coefficient of each sensitive attribute and the reservoir feature (such as porosity and thickness) at the well point, and selects "regional sensitive attributes" with high correlation in the specific reservoir feature value domain interval. Then normalize these regional attributes (unified dimension), and then convert to grid points for fusion, to avoid including areas insensitive to reservoir response in the fusion and improve accuracy.

[0007] CN113341480A discloses a sand hydrate reservoir prediction method based on frequency division RGB slicing and multi-attribute fusion. The scheme first performs frequency division processing (decomposition into low, medium and high frequency data bodies) on the original seismic data for specific geological targets (such as sand hydrate reservoirs), then calculates multiple attributes (such as maximum amplitude and coherence) in each frequency band, and finally uses the RGB three primary color principle to color fuse the same attribute or different attributes in different frequency bands, and delineate the target area through color anomalies, realizing the comprehensive use of frequency information and multiple attributes in complex geological backgrounds (such as wide continental shelves) to qualitatively identify high potential areas.

[0008] CN120233425A discloses a multi-attribute fusion thin reservoir prediction method, device, equipment and medium. The scheme first uses PCA to reduce the dimension of the extracted multiple attributes, then further optimizes through correlation analysis with well point reservoir thickness, and finally uses a pre-set empirical formula (A = peak amplitude / (trough amplitude x thickness)) to calculate the "reservoir development probability", providing a fast and intuitive prediction index (A value) for thin reservoirs and guiding drilling trajectory adjustment.

[0009] CN115508893A discloses a reservoir prediction method based on machine learning and multi-attribute fusion of frequency division, including: combining frequency division technology with machine learning (such as SVM support vector machine). First, frequency division is performed on seismic data, and the preferred attributes of each frequency band are extracted, then the well point sand body thickness is used as a label to train a nonlinear regression model, which can automatically learn the complex mapping relationship between multiple frequency division attributes and sand body thickness, and finally output a fusion attribute that can quantitatively represent sand body thickness, solving the linear limitation of RGB fusion and realizing quantitative prediction of sand body thickness, taking into account thin and thick layers.

[0010] In summary, although the existing seismic attribute fusion reservoir prediction technology is constantly evolving in terms of methods, it still generally has problems such as qualitative prediction results or insufficient quantitative reliability. For example, RGB color fusion relies on manual delineation of abnormal areas, which is highly subjective and difficult to output quantitative parameters that can be directly used for engineering decision-making. Some methods such as empirical formula fusion output numerical values, but lack geological or physical mechanism support, and even have logical contradictions, resulting in low result reliability. At the same time, most technologies pay insufficient attention to the explainability of the fusion process, and machine learning models often become "black boxes" that geologists cannot understand the prediction basis, while empirical formulas are too arbitrary and lack universality. In terms of process integrity, existing technologies generally lack systematic precision control and closed-loop feedback mechanisms, and the diagnosis and correction path of prediction errors is unclear, with only simple mention of "adjusting parameters" or "modifying well trajectories", which does not have the ability of low-cost and high-efficiency iterative optimization. In addition, insufficient attention is paid to the preprocessing of data quality, and the methods of outlier rejection and standardization are not rigorous enough, and the model selection is also relatively rigid, lacking flexible adaptation strategies according to different data characteristics. SUMMARY

[0011] In view of the problems in the prior art, the purpose of the present application is to provide a reservoir prediction method and device based on mixed probability principal component analysis fusion attribute, to realize efficient and accurate prediction and uncertainty quantification of reservoir characteristics.

[0012] To achieve this purpose, the present application adopts the following technical solutions:

[0013] In a first aspect, the present application provides a reservoir prediction method based on mixed probability principal component analysis fusion attribute, which comprises:

[0014] Obtaining the original three-dimensional seismic data body, well point logging data and known reservoir characteristic data of the target study area;

[0015] Constructing an attribute matrix after preprocessing the original three-dimensional seismic data body, and modeling the attribute matrix using mixed probability principal component analysis to obtain a fusion attribute matrix;

[0016] Taking the fusion attribute matrix as the input feature and the known reservoir characteristic data as the target variable, a mapping model is established by combining semi-supervised K-means clustering with well point label guided classification;

[0017] According to the mapping model, all data in the fusion attribute matrix of the study area are calculated to obtain the global reservoir characteristic prediction result.

[0018] The reservoir prediction method based on mixed probability principal component analysis fusion attribute provided by the application fully gives play to the advantages of MPPCA in processing nonlinear and heterogeneous data, improves the accuracy, reliability and uncertainty quantification capability of reservoir prediction, and provides more accurate technical support for oil and gas exploration and development.

[0019] As a preferred technical solution of the application, the original three-dimensional seismic data volume includes one or a combination of at least two of amplitude attributes, frequency attributes and elastic parameter attributes.

[0020] Preferably, the amplitude attribute includes a root mean square amplitude attribute.

[0021] Preferably, the frequency attribute includes a 60Hz single frequency energy attribute based on matching pursuit time-frequency spectrum decomposition.

[0022] Preferably, the elastic parameter attribute includes a pre-stack Vp / Vs attribute.

[0023] As a preferred technical solution of the application, the original three-dimensional seismic data volume satisfies: data signal-to-noise ratio ≥ 3, and main frequency error is controlled within ± 5Hz.

[0024] As a preferred technical solution of the application, the well point logging data includes natural gamma, acoustic time difference, density and resistivity logging curves, and the depth correction error of the resistivity logging curve is ≤ 0.1m.

[0025] As a preferred technical solution of the application, the preprocessing includes: identifying and removing abnormal data beyond a reasonable range by using the Grubbs criterion, and standardizing the seismic attributes after removing outliers, converting each attribute value to the interval [0, 1] or the interval [-1, 1].

[0026] As a preferred technical solution of the application, the modeling includes: establishing a mapping relationship between attribute variables and potential low-dimensional features by a probability statistical method of multiple Gaussian subspaces, calculating principal component eigenvalues and corresponding eigenvectors of each subspace, selecting principal component components with a sum of variance contribution rates ≥ a set threshold value according to the eigenvalue size and variance contribution rate of each principal component, and obtaining a fusion attribute matrix based on the selected principal component components and corresponding eigenvectors.

[0027] As a preferred technical solution of the application, in the mapping model establishment, the known reservoir feature data are taken as semi-supervised labels.

[0028] Preferably, in the mapping model establishment, the fusion attribute data in the fusion attribute matrix are divided into a well point label sample set and a label-free seismic sample set.

[0029] Preferably, the well point label sample set is initialized with a hard constraint.

[0030] As a preferred technical solution of the present application, in the mapping model establishment, the class clusters obtained by clustering are matched with the well point labels in the well point label sample set to obtain the mapping model.

[0031] Preferably, in the mapping model establishment, cross-validation is used, and the accuracy, precision, recall and F1 value of the validation set are used as indexes, and the optimal parameters are obtained through grid search screening.

[0032] As a preferred technical solution of the present application, in the calculation, for the blank area not covered by the model, a method based on neighborhood majority discrimination is used for spatial reconstruction.

[0033] Preferably, in the calculation, a validation well point not participating in model training is selected for validation, and the number of the validation well point is ≥20% of the total number of wells.

[0034] Preferably, in the calculation, if the classification accuracy index meets the requirements, the global reservoir feature prediction result is output.

[0035] Preferably, the classification accuracy index meeting the requirements includes: the accuracy is ≥0.85, the recall is ≥0.85, and the F1 is ≥0.80.

[0036] In a second aspect, the present application provides a reservoir prediction device based on mixed probability principal component analysis fusion attribute, the device comprising: one or at least two combinations of module devices, electronic devices or computer storage media;

[0037] The module device comprises: an acquisition module for acquiring original three-dimensional seismic data volume, well point logging data and known reservoir feature data of a target study area; a matrix module for constructing an attribute matrix after preprocessing the original three-dimensional seismic data volume, modeling the attribute matrix by using mixed probability principal component analysis to obtain a fusion attribute matrix; a mapping model module for using the fusion attribute matrix as input features and using known reservoir feature data as target variables to establish a mapping model by using semi-supervised K-means clustering combined with well point label guidance classification; and a calculation and prediction module for calculating all data in the fusion attribute matrix of the study area according to the mapping model to obtain a global reservoir feature prediction result.

[0038] The electronic device comprises: at least one processor; and a memory in communication connection with the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the reservoir prediction method based on mixed probability principal component analysis fusion attribute of the first aspect.

[0039] The computer storage medium stores computer executable instructions, and the computer executable instructions are executed by a processor to realize the reservoir prediction method based on mixed probability principal component analysis fusion attributes in the first aspect.

[0040] Compared with the prior art, the present application has the following beneficial effects:

[0041] (1) Significant dimension reduction and redundancy elimination effect: through MPPCA fusion, 3 kinds of seismic attributes can be reduced to 1 principal component component, while retaining more than 85% of the core information, eliminating the redundancy correlation between the attributes, reducing the computational complexity of subsequent model training, and improving the computational efficiency by 40-60%.

[0042] (2) High prediction accuracy and strong stability: through pre-processing optimization and model adaptation, the prediction accuracy is improved compared with the traditional weighted fusion method; at the same time, through multiple rounds of verification and parameter correction, the error fluctuation range of the prediction result is controlled within ±10%, and the stability is significantly better than the prior art.

[0043] (3) MPPCA and semi-supervised K-means cooperation: through the multi-Gaussian subspace probability modeling of MPPCA to realize multi-attribute dimension reduction and uncertainty quantification, combined with the "well point label hard constraint + global unsupervised clustering" of semi-supervised K-means to build a mapping model, taking into account the adaptability of complex geological data, the adaptability of small sample scenes and the geological interpretability of clustering results, solving the problems of linear hypothesis limitation, unsupervised clustering and disconnection with geological significance of traditional methods.

[0044] (4) Wide applicability and simple operation: no need for a large amount of sample data (minimum sample size only 30 well points), suitable for small data amount exploration initial research area; each step can be realized based on conventional seismic interpretation software (such as Petrel, Landmark) and data analysis tools (such as Matlab, Python), without the need for special hardware equipment, reducing the technical application threshold. BRIEF DESCRIPTION OF DRAWINGS

[0045] Figure 1 is a flowchart of the reservoir prediction method based on mixed probability principal component analysis fusion attributes provided by the embodiment of the present application;

[0046] Figure 2 is a schematic diagram of a module device in the reservoir prediction device based on mixed probability principal component analysis fusion attributes provided by the embodiment of the present application;

[0047] Figure 3 is a schematic diagram of an electronic device provided by the embodiment of the present application;

[0048] Figure 4 is a root mean square amplitude attribute plane in Example 1 of the present application;

[0049] Figure 5 This is a 60Hz single-frequency energy attribute plane diagram based on time-spectral decomposition of matched pursuit in Embodiment 1 of the present invention;

[0050] Figure 6 This is the pre-stack Vp / Vs property planar diagram in Embodiment 1 of the present invention;

[0051] Figure 7 This is the fusion attribute planar diagram in Embodiment 1 of the present invention;

[0052] Figure 8 This is the reservoir prediction map obtained in Embodiment 1 of the present invention.

[0053] In the diagram: 10-Electronic device, 11-Processor, 12-ROM, 13-RAM, 14-Bus, 15-I / O interface, 16-Input unit, 17-Output unit, 18-Storage unit, 19-Communication unit.

[0054] The present invention will now be described in further detail. However, the examples described below are merely simplified examples of the present invention and do not represent or limit the scope of protection of the present invention. The scope of protection of the present invention is determined by the claims. Detailed Implementation

[0055] To better illustrate the present invention and facilitate understanding of its technical solutions, typical but non-limiting embodiments of the present invention are as follows:

[0056] Currently, in oil and gas exploration, the response of a single seismic attribute to reservoir characteristics is limited and highly ambiguous, while traditional multi-attribute fusion methods have significant shortcomings: weighted fusion relies on manual experience to set weights, which is highly subjective; RGB fusion can only achieve qualitative or semi-quantitative analysis and cannot meet the needs of quantitative prediction; multiple linear regression assumes that the attributes have a linear relationship with the reservoir, which is difficult to adapt to nonlinear mapping under complex geological conditions; although machine learning-based fusion methods can handle nonlinear problems, the models are complex, the computational cost is high, and the adaptability to small sample data is poor. Furthermore, although principal component analysis (PCA) or single probabilistic principal component analysis (PPCA) has dimensionality reduction capabilities, the former cannot handle nonlinearity and data heterogeneity, while the latter can only fit a single distribution, making it difficult to adapt to the multimodal characteristics of seismic attributes. Moreover, existing technologies lack standardized procedures for reservoir prediction, and systematic solutions are lacking in key aspects such as multi-subspace parameter optimization, principal component probability screening, uncertainty quantification, and semi-supervised mapping modeling, limiting their adaptability to complex geological conditions. Therefore, this invention provides a reservoir prediction method based on hybrid probabilistic principal component analysis with attribute fusion. Through systematic design, it achieves efficient and accurate prediction of reservoir characteristics and uncertainty quantification. This method first acquires the original seismic data, well logging data, and known reservoir characteristics of the study area, extracts various relevant seismic attributes, and preprocesses them to eliminate dimensional differences and noise interference. Subsequently, dimensionality reduction and fusion of the preprocessed multi-attribute data were performed using Mixed Probabilistic Principal Component Analysis (MPPCA) to construct a multi-Gaussian subspace probabilistic model. The EM algorithm was used to estimate the weights of each subspace and the principal component parameters. Principal components with a cumulative joint contribution greater than or equal to a set threshold were selected to construct the fused attribute body, while retaining the uncertainty information of the attribute fusion process. Next, using the fused attributes at well points as input and known reservoir characteristics as labels, a semi-supervised K-means clustering was used to establish a mapping model. Finally, the trained model was used to calculate the prediction results for the fused attributes across the entire domain. The accuracy was evaluated by comparing the results with the actual data from the validation wells. If the error did not meet the requirements, the preprocessing parameters, the number of MPPCA subspaces, or the semi-supervised clustering constraint strength could be adjusted to form an iteratively optimized closed-loop process to overcome the inherent defects of existing seismic attribute fusion reservoir prediction methods, as detailed below:

[0057] I. This embodiment provides a reservoir prediction method based on hybrid probabilistic principal component analysis and attribute fusion, the process of which is as follows: Figure 1 As shown, the reservoir prediction method includes:

[0058] Acquire the original 3D seismic data volume, well logging data, and known reservoir characteristic data of the target study area;

[0059] After preprocessing the original 3D seismic data volume, an attribute matrix is ​​constructed. The attribute matrix is ​​then modeled using mixed probabilistic principal component analysis to obtain the fused attribute matrix.

[0060] Using the fused attribute matrix as input features and known reservoir feature data as target variables, a mapping model is established by combining semi-supervised K-means clustering with well point labels to guide classification.

[0061] Based on the mapping model, all data in the fusion attribute matrix of the study area are calculated to obtain the prediction results of the reservoir characteristics of the entire region.

[0062] In this invention, Mixed Probabilistic Principal Component Analysis (MPPCA) is an advanced dimensionality reduction and data fusion algorithm that integrates probabilistic modeling and multi-subspace analysis. It can not only fit the complex heterogeneous distribution of seismic attributes through multi-Gaussian subspaces, but also eliminate attribute redundancy and correlation while retaining core geological information, thus achieving low-dimensional effective representation of high-dimensional data.

[0063] The original three-dimensional seismic data volume includes one or a combination of at least two of the following: amplitude-type attributes, frequency-type attributes, and elastic parameter-type attributes.

[0064] Among them, the amplitude-type attributes include: root mean square amplitude attribute.

[0065] For example, the extraction process of the root mean square amplitude attribute is as follows: In the seismic record, a sliding time window of a certain length is selected along the time axis with the target layer as the center. The average value of the square of the amplitude of each sampling point within the time window is calculated, and then the square root is taken to obtain the RMS amplitude value of the time window. Its mathematical expression is:

[0066] ;

[0067] In the formula, A RMS (t) represents the root mean square amplitude with time t as the center; A i The instantaneous amplitude value of the i-th sampling point within the time window is represented; N represents the total number of sampling points within the time window.

[0068] The frequency-related attributes include: a 60Hz single-frequency energy attribute based on matched pursuit time-spectral decomposition.

[0069] For example, the process of extracting the 60Hz single-frequency energy attribute based on the time-spectrum decomposition of matching pursuit is as follows:

[0070] (1) Signal representation: The seismic record s(t) is represented as a sparse linear combination of a set of atomic functions with different time shifts, frequencies, and phases: ;

[0071] In the formula, g γi(t) is an atomic function, a Gabor function, and a i R is a coefficient. N (t) represents the residual;

[0072] (2) Matching pursuit decomposition, through an iterative approach, selects the atomic function most relevant to the residual at each step:

[0073] ;

[0074] The residuals are updated until the energy or iteration count conditions are met.

[0075] (3) Time-frequency spectrum construction: The time-frequency distribution of the signal is obtained from the selected atomic functions and their parameters, forming a high-resolution time-frequency spectrum;

[0076] (4) Attribute extraction: Extract the energy distribution of the target frequency f0=60Hz from the decomposition results.

[0077] ;

[0078] A 60Hz single-frequency energy property volume was obtained, highlighting the formation's response characteristics to this frequency wave.

[0079] The elastic parameter class attributes include: pre-stack Vp / Vs attributes.

[0080] For example, the pre-stack Vp / Vs attribute extraction process is as follows:

[0081] (1) Perform pre-stack AVO analysis and use the Zoeppritz approximation. The specific formula of the Aki-Richards equation is as follows:

[0082] ;

[0083] In the formula, , , These are the reflection coefficients for the corresponding longitudinal waves, transverse waves, and densities, respectively. To obtain the average value of the incident angle and the transmitted angle, , , These represent the mean values ​​of the longitudinal wave velocity, transverse wave velocity, and density in the two media, respectively; Δv p This represents the difference in longitudinal wave velocity, i.e., the difference in longitudinal and transverse wave velocities between the upper and lower media; Δv s Δρ is the difference in transverse wave velocity; Δρ is the difference in density; v p1 v is the longitudinal wave velocity of the upper medium. p2 v is the longitudinal wave velocity of the underlying medium. s1 v is the transverse wave velocity of the upper medium. s2ρ1 is the shear wave velocity of the lower medium; ρ2 is the density of the upper medium; ρ3 is the density of the lower medium.

[0084] (2) The longitudinal wave impedance Ip, transverse wave impedance Is and density ρ are obtained simultaneously by pre-stack inversion method;

[0085] (3) Calculate Vp and Vs from the longitudinal wave impedance Ip=ρVp and the transverse wave impedance Is=ρVs;

[0086] (4) Calculate the ratio attribute Vp / Vs: Vp / Vs = Vp ÷ Vs.

[0087] In this invention, the seismic attribute extraction process can be based on the seismic interpretation horizon of the target segment in the study area, and can adopt either layer-by-layer extraction or isochronous slice extraction methods.

[0088] For example, the root mean square amplitude attribute is calculated by extending the construction time window by 10ms both upwards and downwards from the target layer to enhance the signal-to-noise ratio and reduce energy interference from adjacent layers.

[0089] For example, based on the 60Hz single-frequency energy attribute obtained from the matched pursuit time-spectral decomposition: the dominant frequency is calculated using the Hilbert transform method, and the time window length is set to 1.5 times the two-way travel time of the seismic reflection wave group in the target layer to avoid frequency calculation distortion. By performing Fourier transform on the seismic signal within the time window, the dominant frequency corresponding to the peak power spectrum is extracted, and the single-frequency energy attribute is obtained by matching the spectral decomposition results.

[0090] For example, the pre-stack Vp / Vs attribute: The Vp / Vs attribute is obtained from the pre-stack inversion result. The mean is calculated by extending the time window by 10ms above and below the target layer. The Vp / Vs attribute along the layer is extracted to ensure stability and reduce random noise interference.

[0091] The original three-dimensional seismic data volume satisfies the following conditions: data signal-to-noise ratio ≥3, and main frequency error controlled within ±5Hz.

[0092] In this invention, the original three-dimensional seismic data volume can be processed according to conventional seismic data processing standards if it meets the specified requirements. For example, static correction, pre-stack noise suppression, stacking imaging, and migration processing can be performed, and the specific processing can be carried out according to conventional procedures in the field.

[0093] The well point logging data includes: natural gamma, sonic transit time, density and resistivity logging curves, and the depth correction error of the resistivity logging curve is ≤0.1m.

[0094] In this invention, the depth correction error of the resistivity logging curve is ≤0.1m to ensure the depth-time matching accuracy of well point data and seismic data.

[0095] In this invention, known reservoir characteristic data refers to the "reservoir" or "non-reservoir" data at well points obtained through logging and other means within the target study area, which serves as labels for semi-supervised learning.

[0096] The preprocessing includes: using the Grubbs criterion to identify and remove anomalous data that exceeds a reasonable range, standardizing the seismic attributes after removing outliers, and converting each attribute value to the interval [0,1] or the interval [-1,1].

[0097] In this invention, an exemplary preprocessing procedure is as follows:

[0098] (1) Outlier removal: The Grubbs criterion is used to detect outliers in the well point data for each seismic attribute;

[0099] (2) Standardization: Due to the significant differences in the dimensions and numerical ranges of different seismic attributes (e.g., amplitude range of 0-5000, frequency range of 10-60Hz), Z-score standardization is used to convert the attribute values ​​into standardized data with a mean of 0 and a standard deviation of 1. The formula is as follows:

[0100] ;

[0101] In the formula, μ and σ are the mean and standard deviation of the attribute values, respectively.

[0102] If the data distribution is not normal, Min-Max standardization is used to map the data to the [0,1] interval, as shown in the formula:

[0103] ;

[0104] In the formula, x' is the attribute value after normalization, x max and x min These represent the maximum and minimum values ​​of the attribute, respectively, and x is the target value.

[0105] In this invention, the process of constructing the attribute matrix is ​​as follows:

[0106] Suppose there are n sample points in the study area, and m seismic attributes are extracted. Construct an n×m dimensional attribute matrix X=[x ij ] n×m (x) ij (This refers to the standardized value of the j-th attribute of the i-th sample).

[0107] The modeling includes: establishing a mapping relationship between attribute variables and potential low-dimensional features using a probabilistic statistical method for multi-Gaussian subspaces; calculating the principal component eigenvalues ​​and corresponding eigenvectors of each subspace; selecting principal component components whose sum of variance contribution rates is greater than or equal to a set threshold based on the magnitude of the eigenvalues ​​and variance contribution rates of each principal component; and obtaining a fused attribute matrix based on the selected principal component components and their corresponding eigenvectors.

[0108] In this invention, an exemplary modeling process is as follows:

[0109] (1) Multi-Gaussian subspace modeling: The attribute matrix X is modeled using probabilistic principal component analysis (PPCA). It is assumed that the data is composed of K probabilistic principal component analysis (PPCA) sub-models, and the probability density of each sub-model k is:

[0110] ;

[0111] In the formula, z k As a potential low-dimensional feature, W k Let μ be the load matrix. k Let σ be the mean vector. 2 Let θ be the noise variance. k ={W k μ k , σ 2} represents the parameters of the sub-model, N represents the normal distribution, and the weights of each sub-model are π. k ,satisfy ;

[0112] (2) Parameter estimation using the EM algorithm: The parameters of all sub-models are estimated iteratively using the Expectation-Maximization (EM) algorithm. :

[0113] Step E: Calculate the posterior probability γ of a sample belonging to each sub-model. ik =p(k∣x i , {π k θ k});

[0114] M-step: Update the weights of each sub-model. And for each sub-model k, based on the weighted sample γ ik x i Re-estimation of sub-model parameters W k μ k , σ 2 ;

[0115] (3) Subspace principal component fusion and screening: For each sub-model k, its principal components (eigenvalues ​​λ) are extracted by PPCA. k1 ≥λ k2 ≥……), calculate the contribution rate of principal components within the sub-model. (m) k (The principal component dimension of sub-model k); combining the principal components of all sub-models, and calculating the joint probability contribution (sub-model weight π). k Sorting by contribution rate ηkk′ within the sub-model, principal component components with a cumulative joint contribution of ≥85% are selected.

[0116] (4) Constructing fusion attributes: Based on the selected principal component components and combined with the probability parameters of each sub-model, the fusion attribute matrix F is reconstructed. This matrix contains the key features and probability uncertainty information of each subspace, and finally forms a low-dimensional fusion attribute representation that can reflect the multimodal characteristics of the reservoir.

[0117] In the mapping model establishment, known reservoir feature data are used as semi-supervised labels.

[0118] In the process of establishing the mapping model, the fusion attribute data in the fusion attribute matrix is ​​divided into a well point labeled sample set and an unlabeled seismic sample set.

[0119] The well point label sample set is initialized using hard constraints.

[0120] In the process of establishing the mapping model, the clusters obtained by clustering are matched with the well point labels in the well point label sample set to obtain the mapping model.

[0121] In the establishment of the mapping model, cross-validation is used, with the accuracy, precision, recall and F1 score of the validation set as indicators, and the optimal parameters are obtained through grid search.

[0122] For example, the process of establishing a mapping model is as follows:

[0123] (1) Data preparation and label constraints: Extract the fusion attribute values ​​from the fusion attribute matrix F from the well point samples, and use the known reservoir feature data at the well point as semi-supervised labels (such as "reservoir" and "non-reservoir" category labels); divide all data in the fusion attribute matrix F into "well point label sample set" and "unlabeled seismic sample set", and divide the training dataset and the validation dataset.

[0124] (2) Semi-supervised K-means initialization:

[0125] For well point labeled samples, hard constraint initialization is adopted: the mean of the fusion attributes of the "reservoir" labeled samples is used as the initial value of the reservoir cluster center, and the mean of the fusion attributes of the "non-reservoir" labeled samples is used as the initial value of the non-reservoir cluster center; for unlabeled seismic samples, several samples are randomly selected to initialize the remaining cluster centers.

[0126] (3) Semi-supervised clustering iteration:

[0127] Allocation phase: For well point labeled samples, force them to be assigned to the cluster corresponding to their label; for unlabeled seismic samples, calculate their Euclidean distance to each cluster center and assign them to the cluster with the closest distance.

[0128] Update phase: Recalculate the center of each cluster; repeat the "assign-update" process until the change in cluster center is less than the set threshold (the threshold can be any value within ≤0.001) or the number of iterations reaches the upper limit, or the number of iterations reaches more than 100, whichever comes first, or select one of the parameters as the requirement;

[0129] (4) Mapping between clusters and reservoir categories: The clusters obtained from clustering are matched with well point labels, and finally a mapping relationship between fusion attributes and reservoir categories is established to realize the classification of reservoir categories across the entire domain;

[0130] (5) Model optimization and evaluation: Cross-validation, such as 5-fold cross-validation, is adopted. The optimal parameters are selected by grid search, using the validation set accuracy, precision, recall, and F1 score as indicators. The specific formula is as follows:

[0131] Accuracy: ;

[0132] Precision: ;

[0133] Recall: ;

[0134] F1 value: ;

[0135] Where TP represents true positives, TN represents true negatives, FP represents false positives, and FN represents false negatives.

[0136] In this invention, the training dataset and the validation dataset are obtained by extracting the fusion attribute values ​​that coincide with the well point location and the reservoir feature data ("reservoir" or "non-reservoir" data) of the well point location from all data in the fusion attribute matrix F as the "well point label sample set". The "well point label sample set" is divided into the training dataset and the validation dataset in a ratio of (7-9):(1-3).

[0137] In the calculation, for blank areas not covered by the model, a method based on neighborhood majority discrimination is used for spatial reconstruction. Specifically, the neighborhood weighted average method can be used to complete the region based on the reservoir distribution of the surrounding predicted units, and finally form a reservoir range prediction plan map to achieve complete coverage of the reservoir spatial distribution.

[0138] In the calculation, validation wells that did not participate in model training are selected for validation, and the number of validation wells is ≥ 20% of the total number of wells.

[0139] In this invention, the total number of wells refers to all relevant wells in the study area.

[0140] If the classification accuracy index meets the requirements in the calculation, the prediction result of the reservoir characteristics of the whole domain will be output.

[0141] For example, the specific calculation process is as follows:

[0142] (1) Spatial distribution reconstruction of global prediction: Input all data in the fusion attribute matrix F into the trained mapping model to obtain the reservoir / non-reservoir classification results of each seismic trace grid point. For the blank areas not covered by the model, spatial reconstruction is carried out using the neighborhood majority discrimination method to ensure the continuity and geological rationality of the reservoir range distribution.

[0143] (2) Accuracy verification and parameter correction: Select verification wells that did not participate in model training (number ≥ 20% of the total number of wells), extract the prediction results and compare them with the actual reservoir feature data, and calculate the classification accuracy index (Accuracy, Precision, Recall, F1, etc.).

[0144] If the requirements of Accuracy ≥ 0.85, Recall ≥ 0.85, and F1 ≥ 0.80 cannot be met simultaneously, then parameter correction is required; otherwise, output the result index.

[0145] If the error originates from data noise, the original 3D seismic data volume is preprocessed again, such as adjusting the significance level of the Grubbs criterion (e.g., α=0.01) or changing the standardization method.

[0146] If the error originates from the loss of principal component information, then perform mixed probability principal component analysis again, such as lowering the threshold for the cumulative contribution rate of principal components (e.g., adjusting it to 80%).

[0147] If the error stems from poor model fit, then the mapping model is re-established, such as by changing the mapping model or optimizing the hyperparameters.

[0148] After correction, subsequent steps need to be repeated until the prediction accuracy meets the requirements. Finally, a reservoir feature prediction planar map is generated, and the prediction accuracy range is marked (e.g., F1≥0.9 for high accuracy region, 0.8≤F1<0.9 for medium accuracy region).

[0149] II. This embodiment provides a module device in a reservoir prediction apparatus based on hybrid probabilistic principal component analysis fusion attributes, such as... Figure 2 As shown, the module device includes:

[0150] The acquisition module 100 is used to acquire the original three-dimensional seismic data volume, well logging data and known reservoir characteristic data of the target study area;

[0151] Matrix module 200 is used to construct an attribute matrix from the original 3D seismic data volume after preprocessing, and to model the attribute matrix using mixed probabilistic principal component analysis to obtain a fused attribute matrix;

[0152] The mapping model module 300 is used to establish a mapping model by taking the fused attribute matrix as input features, taking known reservoir feature data as target variables, and using semi-supervised K-means clustering combined with well point labels to guide classification.

[0153] The calculation and prediction module 400 is used to calculate all data in the fusion attribute matrix of the study area based on the mapping model to obtain the prediction results of the reservoir characteristics of the whole area.

[0154] Regarding the modular devices of the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0155] III. This embodiment provides an electronic device intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0156] like Figure 3 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded into the RAM 13 from storage unit 18. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An I / O interface 15 is also connected to the bus 14.

[0157] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0158] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, central processing unit (CPU), graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as reservoir prediction methods based on mixed probabilistic principal component analysis fusion attributes.

[0159] In some embodiments, the reservoir prediction method based on the fusion attributes of mixed probabilistic principal component analysis can be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program can be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the reservoir prediction method based on the fusion attributes of mixed probabilistic principal component analysis described above can be performed. Alternatively, in other embodiments, processor 11 can be configured to perform the reservoir prediction method based on the fusion attributes of mixed probabilistic principal component analysis by any other suitable means (e.g., by means of firmware).

[0160] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0161] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0162] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0163] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0164] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0165] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0166] The server provided in this embodiment includes: a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the program, it implements a reservoir prediction method based on the fusion attributes of mixed probabilistic principal component analysis.

[0167] Unless otherwise specifically stated, terms such as processing, calculation, operation, determination, display, etc., may refer to the actions and / or processes of one or more processing or computing systems or similar devices that represent the manipulation and conversion of data representing physical (e.g., electronic) quantities within the registers or memory of the processing system into other data similarly representing physical quantities within the memory, registers, or other such information storage, transmission, or display devices of the processing system. Information and signals can be represented using any of a variety of different techniques and methods. For example, data, instructions, commands, information, signals, bits, symbols, and chips mentioned throughout the above description can be represented by voltage, current, electromagnetic waves, magnetic fields or particles, light fields or particles, or any combination thereof.

[0168] Those skilled in the art will also understand that the various illustrative logic blocks, modules, circuits, and algorithm steps described in conjunction with embodiments of the present invention can all be implemented as electronic hardware, computer software, or a combination thereof. To clearly illustrate the interchangeability between hardware and software, the various illustrative components, blocks, modules, circuits, and steps described above are generally described in terms of their functionality. Whether this functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art can implement the described functionality in alternative ways for each specific application; however, such implementation decisions should not be construed as departing from the scope of protection of the present invention.

[0169] The steps of the methods or algorithms described in conjunction with the embodiments herein can be directly embodied in hardware, software modules executed by a processor, or a combination thereof. The software modules can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disks, removable disks, CD-ROMs, or any other form of storage medium well known in the art. An exemplary storage medium is connected to the processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and storage medium can reside in an ASIC. The ASIC can reside in a user terminal. Alternatively, the processor and storage medium can exist as discrete components in the user terminal.

[0170] For software implementation, the techniques described in this invention can be implemented using modules (e.g., procedures, functions, etc.) that perform the functions described in this application. This software code can be stored in memory units and executed by a processor. The memory units can be implemented within the processor or externally; in the latter case, they are communicatively coupled to the processor via various means, as is well known in the art.

[0171] IV. To illustrate the predictive performance achievable by the reservoir prediction method based on hybrid probabilistic principal component analysis fusion attributes provided by this invention, the following example is used for explanation:

[0172] Example 1

[0173] This embodiment provides a reservoir prediction method based on hybrid probabilistic principal component analysis and attribute fusion, as detailed below:

[0174] In this embodiment, the target study area is specifically the Pinghu Formation in the Xihu Depression;

[0175] Acquire the original 3D seismic data volume, well logging data, and known reservoir characteristic data of the target study area;

[0176] After preprocessing the original 3D seismic data volume, an attribute matrix is ​​constructed. The attribute matrix is ​​then modeled using mixed probabilistic principal component analysis to obtain the fused attribute matrix.

[0177] Using the fused attribute matrix as input features and known reservoir feature data as target variables, a mapping model is established by combining semi-supervised K-means clustering with well point labels to guide classification.

[0178] Based on the mapping model, all data in the fusion attribute matrix of the study area are calculated to obtain the prediction results of the reservoir characteristics of the entire region.

[0179] The original 3D seismic data volume includes: amplitude-type attributes, frequency-type attributes, and elastic parameter-type attributes; the amplitude-type attributes are root mean square amplitude attributes; the frequency-type attributes are 60Hz single-frequency energy attributes based on matched pursuit time-spectrum decomposition; and the elastic parameter-type attributes are pre-stack Vp / Vs attributes.

[0180] Root mean square amplitude attribute: Using the target layer as the center, a 10ms construction window is extended upwards and downwards to calculate the root mean square amplitude attribute, in order to enhance the signal-to-noise ratio and reduce energy interference from adjacent layers, such as... Figure 4 As shown.

[0181] The 60Hz single-frequency energy attribute based on matched pursuit time-spectral decomposition: The dominant frequency was calculated using the Hilbert transform method, with the time window length set to 1.5 times the two-way travel time of the seismic reflection wave group in the target section to avoid frequency calculation distortion. By performing a Fourier transform on the seismic signal within the time window, the dominant frequency corresponding to the power spectrum peak was extracted, and the single-frequency energy attribute was obtained using the matched pursuit time-spectral decomposition results, such as... Figure 5 As shown.

[0182] Pre-stack Vp / Vs attributes: The Vp / Vs attributes are obtained from the pre-stack inversion results. The mean values ​​are calculated by extending the time window by 10ms above and below the target layer, and the Vp / Vs attributes along the layer are extracted to ensure stability and reduce random noise interference. Figure 6 As shown.

[0183] The original three-dimensional seismic data volume satisfies the following conditions: data signal-to-noise ratio ≥3, and main frequency error controlled within ±5Hz.

[0184] The well point logging data includes: natural gamma, sonic transit time, density and resistivity logging curves, with a depth correction error of ≤0.1m for the resistivity logging curve.

[0185] Known reservoir characteristic data refers to the "reservoir" or "non-reservoir" data at well points within the target study area obtained through logging and other means, which serves as labels for semi-supervised learning.

[0186] The preprocessing includes: using the Grubbs criterion to identify and remove outlier data that exceeds a reasonable range, standardizing the seismic attributes after removing outliers, and converting each attribute value to the interval [0,1].

[0187] The process of constructing the attribute matrix is ​​as follows: There are n sample points in the study area, m seismic attributes are extracted, and an n×m dimensional attribute matrix X=[x ij ] n×m (x) ij The j-th attribute value of the i-th sample after standardization is as follows: There are a total of 105,360,300 (501×701×300) sample points in the study area, and three kinds of seismic attributes are extracted (root mean square amplitude attribute, 60Hz single-frequency energy attribute based on matching pursuit time-spectral decomposition, and pre-stack Vp / Vs attribute).

[0188] The modeling process is as follows:

[0189] (1) Multi-Gaussian subspace modeling: The attribute matrix X is modeled using probabilistic principal component analysis (PPCA). It is assumed that the data is composed of K probabilistic principal component analysis (PPCA) sub-models, and the probability density of each sub-model k is:

[0190]

[0191] In the formula, z k As a potential low-dimensional feature, W k Let μ be the load matrix. k Let σ be the mean vector. 2 Let θ be the noise variance. k ={W k μ k , σ 2} represents the parameters of the sub-model, N represents the normal distribution, and the weights of each sub-model are π. k ,satisfy ;

[0192] (2) Parameter estimation using the EM algorithm: The parameters of all sub-models are estimated iteratively using the Expectation-Maximization (EM) algorithm. :

[0193] Step E: Calculate the posterior probability γ of a sample belonging to each sub-model. ik =p(k∣x i , {π k θ k});

[0194] M-step: Update the weights of each sub-model. And for each sub-model k, based on the weighted sample γ ik x i Re-estimation of sub-model parameters W k μk , σ 2 ;

[0195] (3) Subspace principal component fusion and screening: For each sub-model k, its principal components (eigenvalues ​​λ) are extracted by PPCA. k1 ≥λ k2 ≥...), calculate the contribution rate of principal components within the sub-model. (m) k (The principal component dimension of sub-model k); combining the principal components of all sub-models, and calculating the joint probability contribution (sub-model weight π). k Sorting by contribution rate ηkk′ within the sub-model, principal component components with a cumulative joint contribution of ≥85% are selected.

[0196] (4) Constructing fusion attributes: Based on the selected principal component components and combined with the probability parameters of each sub-model, the fusion attribute matrix F is reconstructed. This matrix simultaneously contains the key features and probabilistic uncertainty information of each subspace, ultimately forming a low-dimensional fusion attribute representation that can reflect the multimodal characteristics of the reservoir, such as... Figure 7 .

[0197] The process of establishing the mapping model is as follows:

[0198] (1) Data preparation and label constraints: Extract all fusion attribute values ​​from the fusion attribute matrix F from the well point samples, and use the known reservoir feature data at the well point as semi-supervised labels (such as "reservoir" and "non-reservoir" category labels); divide all data in the fusion attribute matrix F into "well point label sample set" and "unlabeled seismic sample set"; and divide the training dataset and the validation dataset. The training dataset and the validation dataset are the fusion attribute values ​​that coincide with the well point location and the known reservoir feature data ("reservoir" or "non-reservoir" data) at the well point location extracted from all data in the fusion attribute matrix F as the "well point label sample set", and divide the "well point label sample set" into the training dataset and the validation dataset in an 8:2 ratio.

[0199] (2) Semi-supervised K-means initialization:

[0200] For well point labeled samples, hard constraint initialization is adopted: the mean of the fusion attributes of the "reservoir" labeled samples is used as the initial value of the reservoir cluster center, and the mean of the fusion attributes of the "non-reservoir" labeled samples is used as the initial value of the non-reservoir cluster center; for unlabeled seismic samples, several samples are randomly selected to initialize the remaining cluster centers.

[0201] (3) Semi-supervised clustering iteration:

[0202] Allocation phase: For well point labeled samples, force them to be assigned to the cluster corresponding to their label; for unlabeled seismic samples, calculate their Euclidean distance to each cluster center and assign them to the cluster with the closest distance.

[0203] Update phase: Recalculate the center of each cluster; repeat the "assign-update" process until the change in cluster center is less than the set threshold (0.001) or the number of iterations reaches the upper limit of 100, whichever comes first.

[0204] (4) Mapping between clusters and reservoir categories: The clusters obtained by clustering are matched with well point labels, and finally the mapping relationship between fusion attributes and reservoir categories is established to realize the division of reservoir categories across the entire domain.

[0205] (5) Model optimization and evaluation: Cross-validation, such as 5-fold cross-validation, is adopted. The optimal parameters are selected through grid search, using validation set accuracy, precision, recall, and F1 score as metrics.

[0206] The specific calculation process is as follows:

[0207] (1) Spatial distribution reconstruction of global prediction: All data in the fusion attribute matrix F are input into the trained mapping model to obtain the reservoir / non-reservoir classification results for each seismic trace grid point. For blank areas not covered by the model, spatial reconstruction is carried out using a neighborhood majority discrimination method to ensure the continuity and geological rationality of the reservoir range distribution.

[0208] (2) Accuracy verification and parameter correction: Select verification wells that did not participate in model training (the number of which is 20% of the total number of wells), extract the prediction results and compare them with the actual reservoir feature data, and calculate the classification accuracy index (Accuracy, Precision, Recall, F1, etc.).

[0209] If the requirements of Accuracy≥0.85, Recall≥0.85 and F1≥0.80 cannot be met simultaneously, parameter correction is required; otherwise, output data.

[0210] If the error originates from data noise, the original 3D seismic data volume should be preprocessed again, such as adjusting the significance level of the Grubbs criterion (e.g., α=0.01) or changing the standardization method.

[0211] If the error stems from the loss of principal component information, then perform mixed probability principal component analysis again, such as by lowering the threshold for the cumulative contribution rate of principal components (e.g., adjusting it to 80%).

[0212] If the error stems from poor model fit, then rebuild the mapping model, such as by changing the mapping model or optimizing the hyperparameters.

[0213] After correction, subsequent steps need to be repeated until the prediction accuracy meets the requirements. Finally, a reservoir feature prediction planar map is generated, and the prediction accuracy range is marked (e.g., F1≥0.9 for high accuracy area, 0.8≤F1<0.9 for medium accuracy area).

[0214] in, Figure 6 The pre-stack Vp / Vs attribute planar map shows reservoir response characteristics in the areas of wells x-1, x-2, and x-3. However, verification by combining the actual interpretation data of well logging at the well points reveals that the reservoir response at wells x-2 and x-3 is a false feature and there is no actual reservoir development. This indicates that using a single seismic attribute for reservoir prediction has obvious limitations, making it difficult to distinguish between real reservoirs and false responses, and easily leading to prediction bias. Figure 8 The reservoir prediction map obtained based on the MPPCA fusion + semi-supervised K-means model of this invention can accurately identify reservoir distribution. The x-1 well area is accurately identified as a reservoir development zone, while the x-2 and x-3 well areas are clearly identified as non-reservoir zones. The prediction results are highly consistent with the actual logging data, demonstrating high prediction accuracy and successfully avoiding the spurious reservoir interference problem common in single attribute prediction. This verifies the advantages of the multi-attribute fusion method of this invention in improving the reliability of reservoir prediction.

[0215] In summary, the reservoir prediction method provided by this invention achieves a multi-dimensional breakthrough through MPPCA fusion—adapting complex heterogeneous distributions of seismic attributes to multi-Gaussian subspaces, quantifying and fusing uncertainties through probabilistic modeling, and retaining more key reservoir information compared to traditional PCA dimensionality reduction, reducing data redundancy by more than 60%, while improving computational efficiency by 30%-50%. Secondly, by combining the "few label constraints + global unsupervised clustering" characteristics of semi-supervised K-means, it solves the problem of insufficient generalization ability of supervised learning in small sample scenarios and avoids the defect of unsupervised clustering being disconnected from geological significance, improving prediction accuracy to over 85% and exhibiting significantly better stability than a single model. Thirdly, the method is highly adaptable to data distribution, can handle nonlinear and multimodal seismic attributes, and the entire process requires no complex hardware support, can be implemented on conventional geological modeling platforms, and has a low technical application threshold. Most importantly, this invention constructs a standardized process of "multi-subspace fusion - semi-supervised mapping - closed-loop verification". Through uncertainty quantification and iterative optimization mechanisms, it provides technical support with accuracy, reliability and interpretability for reservoir exploration target selection and drilling deployment in complex geological backgrounds.

[0216] The present invention is described in detail through the above embodiments, but the present invention is not limited to the above detailed structural features, that is, it does not mean that the present invention must rely on the above detailed structural features to be implemented. Those skilled in the art should understand that any improvements to the present invention, equivalent substitutions for the components used in the present invention, additions of auxiliary components, and selection of specific methods, etc., all fall within the protection scope and disclosure scope of the present invention.

[0217] The preferred embodiments of the present invention have been described in detail above. However, the present invention is not limited to the specific details in the above embodiments. Within the scope of the technical concept of the present invention, various simple modifications can be made to the technical solution of the present invention, and these simple modifications all fall within the protection scope of the present invention.

[0218] It should also be noted that the various specific technical features described in the above specific embodiments can be combined in any suitable manner without contradiction. In order to avoid unnecessary repetition, the present invention will not describe the various possible combinations separately.

[0219] Furthermore, various different embodiments of the present invention can be combined in any way, as long as they do not violate the spirit of the present invention, they should also be regarded as the content disclosed by the present invention.

Claims

1. A reservoir prediction method based on a hybrid probability principal component analysis fused attribute, characterized in that, The reservoir prediction method comprises: obtaining original three-dimensional seismic data, well point logging data and known reservoir characteristic data of a target study area; constructing an attribute matrix after preprocessing the original three-dimensional seismic data, modeling the attribute matrix by using hybrid probability principal component analysis to obtain a fusion attribute matrix; taking the fusion attribute matrix as input features and the known reservoir characteristic data as target variables, establishing a mapping model by using semi-supervised K-means clustering combined with well point label guided classification; calculating all data in the fusion attribute matrix of the study area according to the mapping model to obtain a global reservoir characteristic prediction result.

2. The reservoir prediction method of claim 1, wherein, The original three-dimensional seismic data comprises one or a combination of at least two of amplitude attributes, frequency attributes and elastic parameter attributes; Preferably, the amplitude attributes comprise root mean square amplitude attributes. Preferably, the frequency attributes comprise 60Hz single frequency energy attributes based on matching pursuit time-frequency spectrum decomposition. Preferably, the elastic parameter attributes comprise prestack Vp / Vs attributes.

3. The reservoir prediction method of claim 1, wherein, The original three-dimensional seismic data satisfies that the data signal-to-noise ratio is greater than or equal to 3 and the main frequency error is controlled within ±5Hz.

4. The reservoir prediction method of claim 1, wherein, The well point logging data comprises natural gamma, acoustic time difference, density and resistivity logging curves, and the depth correction error of the resistivity logging curve is less than or equal to 0.1m.

5. The reservoir prediction method of claim 1, wherein, The preprocessing comprises identifying and removing abnormal data beyond a reasonable range by using Grubbs criterion, standardizing the seismic attributes after removing the abnormal values, and converting the attribute values to the interval [0, 1] or the interval [-1, 1].

6. The reservoir prediction method of claim 1, wherein, The modeling comprises establishing a mapping relationship between attribute variables and potential low-dimensional characteristics by a probability statistical method of multi-Gaussian subspace, calculating principal component eigenvalues and corresponding eigenvectors of each subspace, screening principal component components whose sum of variance contribution rates is greater than or equal to a set threshold value according to the eigenvalue size and variance contribution rate of each principal component, and obtaining a fusion attribute matrix based on the selected principal component components and corresponding eigenvectors.

7. The reservoir prediction method of claim 1, wherein, In the establishment of the mapping model, the known reservoir characteristic data are used as semi-supervised labels. Preferably, in the establishment of the mapping model, the fusion attribute data in the fusion attribute matrix are divided into a well point label sample set and a label-free seismic sample set. Preferably, the well point label sample set is initialized by using hard constraint.

8. The reservoir prediction method of claim 1, wherein, In the establishment of the mapping model, the class clusters obtained by clustering are matched with well point labels in the well point label sample set to obtain the mapping model. Preferably, in the establishment of the mapping model, cross-validation is used, the accuracy, precision, recall rate and F1 value of the validation set are used as indexes, and the optimal parameters are selected by grid search.

9. The reservoir prediction method of claim 1, wherein, In the calculation, for blank areas not covered by the model, a method based on neighborhood majority discrimination is used for spatial reconstruction. Preferably, in the calculation, a validation well point not participating in model training is selected for validation, and the number of the validation well point is greater than or equal to 20% of the total number of wells. Preferably, in the calculation, if a classification accuracy index meets the requirements, a global reservoir characteristic prediction result is output. Preferably, the classification accuracy index meeting the requirements comprises that the accuracy is greater than or equal to 0.85, the recall rate is greater than or equal to 0.85, and the F1 is greater than or equal to 0.

80. 10.A reservoir prediction device based on a hybrid probability principal component analysis fusion attribute, characterized by, The reservoir prediction device comprises one or a combination of at least two of a module device, an electronic device or a computer storage medium; The module device comprises: an acquisition module, configured to acquire original three-dimensional seismic data, well point logging data and known reservoir characteristic data of a target research area; a matrix module, configured to construct an attribute matrix after preprocessing the original three-dimensional seismic data, model the attribute matrix by using hybrid probability principal component analysis, and obtain a fusion attribute matrix; a mapping model module, configured to take the fusion attribute matrix as input features, take the known reservoir characteristic data as target variables, and establish a mapping model by using semi-supervised K-means clustering combined with well point label guidance classification; and a calculation and prediction module, configured to calculate all data in the fusion attribute matrix of the research area according to the mapping model, and obtain a global reservoir characteristic prediction result. The electronic device comprises: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the reservoir prediction method based on hybrid probability principal component analysis fusion attribute according to any one of claims 1-9. The computer storage medium stores computer executable instructions, and the computer executable instructions are executed by a processor to implement the reservoir prediction method based on hybrid probability principal component analysis fusion attribute according to any one of claims 1-9.

Citation Information

Patent Citations

  • Reservoir feature prediction method and device based on seismic attribute fusion

    CN112346117A

  • Sandy hydrate reservoir prediction method based on frequency-division RGB slicing and multi-attribute fusion

    CN113341480A

  • Reservoir prediction method based on machine learning and frequency division multi-attribute fusion

    CN115508893A

  • Multi-attribute fusion thin reservoir prediction method, device, equipment and medium

    CN120233425A