Ultra-shallow gas-bearing reservoir interval transit time prediction method combining mRMR and LGBM models

By combining the mRMR and LGBM models, selecting representative features and training the model, the problem of insufficient accuracy in acoustic time difference prediction in shallow undiagenetic gas reservoirs was solved, and efficient and accurate prediction of P-wave and S-wave time differences was achieved.

CN120847901APending Publication Date: 2025-10-28HAINAN BRANCH OF CHINA NATIONAL OFFSHORE OIL (CHINA) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510825938.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Existing technologies lack accuracy in predicting acoustic time differences in shallow, undiagenetic gas-bearing reservoirs, especially in the ultra-shallow, undiagenetic reservoirs of the South China Sea, where the P- and S-wave time differences are small. Existing methods are unable to effectively extract acoustic time differences, making fluid identification and engineering parameter calculation difficult.

Method used

The method of combining mRMR and LGBM models is adopted. By collecting well logging data, performing data cleaning and standardization, the mRMR algorithm is used to select representative features, and the LGBM model is combined for training and verification to predict the P-wave and S-wave time differences.

Benefits of technology

It improves the accuracy and efficiency of acoustic wave time difference prediction, can quickly build models on large data sets, reduce model complexity, improve prediction accuracy and stability, and is suitable for P-wave and S-wave time difference prediction in ultra-shallow gas reservoirs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120847901A_ABST
    Figure CN120847901A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of geophysical logging, and discloses an ultra-shallow gas-bearing reservoir interval transit time prediction method combining mRMR and LGBM models, comprising: collecting logging original data of a target layer section, the logging original data comprising a plurality of features; performing standardization processing on the features; performing feature selection on the interval transit time by using an mRMR algorithm, and screening representative features based on correlation between the interval transit time and each feature; the representative features are divided into a training set and a test set, the features in the training set serve as input, the longitudinal wave time difference curve serves as output, and an LGBM model is trained; applying the trained LGBM model to interval transit time prediction of a well section without interval transit time measurement to obtain a predicted longitudinal wave time difference curve; and S4, taking the original longitudinal wave time difference curve and the obtained predicted longitudinal wave time difference curve as characteristic input, taking the transverse wave time difference curve as output, repeating the steps S1 to S6, and finally obtaining a longitudinal and transverse wave curve of a well section without longitudinal and transverse wave time difference measurement. The method improves the prediction precision of the interval transit time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of geophysical logging technology, and more specifically, to a method for predicting the acoustic transit time of ultra-shallow gas-bearing reservoirs using a combination of mRMR and LGBM models. Background Technology

[0002] Well logging is an indispensable and crucial technology in petroleum geological exploration. Its goal is to determine reservoir geological parameters such as porosity, permeability, and saturation by measuring the acoustic, electrical, and nuclear properties of the reservoir. Among these, sonic logging is a key technique for understanding formation and fluid information. It obtains geological parameters such as formation lithology, porosity, fluid type, and geostress by utilizing formation elastic information (e.g., sound velocity, transit time, attenuation, and dispersion). Furthermore, sonic logging can also be used for detecting fractures and small geological structures, tracing formation interfaces, and geological steering. However, the physical properties of shallow, unformed marine reservoirs differ from those of medium-deep, consolidated rocks. When acoustic waves propagate through these loosely structured strata, the P-wave and S-wave transit times increase (velocity decreases) and waveform amplitude attenuates significantly. In sections with high gas saturation, the P-wave and S-wave transit times increase sharply (velocity is extremely low, such as P-wave velocity being less than fluid velocity). The strong waveform attenuation makes it impossible to acquire effective acoustic logging wave trains, and acoustic transit time extraction is extremely difficult. This greatly restricts fluid identification, engineering parameter calculation, and seismic profile stratigraphic calibration in such ultra-shallow marine reservoirs.

[0003] Currently, for acoustic transit time prediction, Dong Hongchao et al. used Gaussian process regression to predict the shear wave velocity in low-permeability offshore reservoirs of the Bohai M oilfield; Sun Lijing et al. used quadratic polynomial relationship fitting to predict the shear wave velocity of the Sa-Zero group in the Changyuan area of ​​the Daqing oilfield; Zhang Guanjie et al. used the XU-White model to predict the shear wave transit time of low-permeability sandstone reservoirs; Li Xiongyan et al. compared the effects of cluster analysis, rock physics models, and multiple linear fitting in predicting P-wave transit time, and finally selected the most accurate multiple linear fitting method for P-wave transit time prediction. Based on this, they also found that the shear wave prediction model established using the Greengerg-Castagna (GC) formula has the highest accuracy. Chinese patent application, publication number CN117669785A, discloses a method and device for predicting shear wave transit time, using a neural network built by combining CNN and LSTM after data grouping to predict shear wave transit time. Publication number CN117148473A discloses a method for predicting shear wave transit time in continental shale gas reservoirs, which calculates the shear wave transit time of shale gas reservoirs using clay mineral content and sonic transit time logging values. Publication number CN118465832A discloses an intelligent prediction method for shale shear wave velocity based on activation graph-like combined physical constraints, using a two-layer convolutional neural network to intelligently predict shale shear wave velocity. Publication number CN114488311A discloses a shear wave transit time prediction method based on the SSA-ELM algorithm, which selects logging curves with strong correlation to shear wave transit time and uses a reservoir shear wave transit time prediction model optimized based on the SSA-ELM algorithm to predict the shear wave transit time curve of shale reservoir sections.

[0004] The above methods have two characteristics: (1) They are only applicable to reservoir transit time prediction for consolidation and diagenesis in specific regions, and are not fully applicable to shallow, undiagenetic reservoirs such as those in the South China Sea; (2) The input data of the above machine learning fluid prediction models lacks the extraction of features sensitive to acoustic transit time, and the predicted acoustic transit time type is basically shear wave. In shallow, undiagenetic reservoirs, both longitudinal and shear wave transit times are small. Accurate prediction of longitudinal wave transit time is a prerequisite, followed by shear wave transit time. Therefore, the existing prediction methods are insufficient in terms of the accuracy of acoustic transit time prediction for shallow, undiagenetic gas-bearing reservoirs. Summary of the Invention

[0005] The purpose of this invention is to overcome the shortcomings of existing prediction methods in the accuracy of acoustic transit time prediction for shallow, unformed gas-bearing reservoirs, and to provide a method for predicting acoustic transit time of ultra-shallow gas-bearing reservoirs by combining mRMR and LGBM models. This method achieves accurate prediction of acoustic transit time of shallow, unformed gas-bearing reservoirs.

[0006] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is as follows: A method for predicting acoustic transit time in ultra-shallow gas-bearing reservoirs using a combined mRMR and LGBM model includes the following steps: S1. Collect raw logging data of the target layer segment, wherein the raw logging data includes multiple features; S2. Perform data cleaning on the features, remove outliers, and standardize all features using data standardization / normalization methods; S3. Use the mRMR algorithm to select features for acoustic time difference. Based on the correlation between acoustic time difference and each feature, select representative features. S4. Divide the representative features obtained in step S3 into a training set and a test set. Use the features in the training set as input and the longitudinal wave time difference curve as output to train the LGBM model using the training set data. S5. Verify and evaluate the accuracy of the LGBM model, and then use test set data to further evaluate the accuracy of the LGBM model; S6. Use the trained LGBM model to predict the acoustic transit time of the unmeasured acoustic transit time well section to obtain the predicted P-wave transit time curve. S7. Using the original P-wave time difference curve and the predicted P-wave time difference curve obtained in step S6 as feature inputs and the S-wave time difference curve as output, repeat steps S1 to S6 to finally obtain the P-wave and S-wave curves of the well section with unmeasured P-wave and S-wave time differences.

[0007] Preferably, the plurality of features include at least natural gamma, deep lateral resistivity, shallow lateral resistivity, shallow induced resistivity, compensated neutrons, compensated density, and wellbore curve.

[0008] Preferably, step S2 specifically includes the following steps: S21. For each feature, calculate the correlation between the feature and the acoustic time difference. The correlation index is mutual information. S22. For each pair of features, calculate the redundancy between features, measured by mutual information, where the formula for mutual information is: (1) in, It is a random variable and The joint probability distribution, It is a random variable and The marginal probability distribution; S23. Use the following objective function to select features: (2) in, It is the selected feature set. It is a feature With sound wave time difference Mutual information between them It is a feature With features Mutual information between them It is a set Size; S24. Continuously update the feature set through iteration. In each iteration, select the option that maximizes the objective function. The characteristic is that it continues until the stopping condition is met; S25. Based on the sorting results of the mutual information of each input variable, select 1 to m input variables in the sequence as representative features.

[0009] Preferably, in step S3, the outlier value is a value that clearly does not conform to the logging law. After deleting the outlier value, the new value is calculated using the linear interpolation method. If outlier points appear continuously, all feature curve values ​​corresponding to this depth are deleted. Then, one of the following methods is used to standardize all features: Z-score method, Min-Max normalization, RobustScaler, Log transform, and StandardScaler.

[0010] Preferably, in step S3, the Z-score method standardizes all features: (3) in, For the feature data to be processed, The mean of the total data; Let Z be the standard deviation of the overall data, and Z be the standardized value of the data.

[0011] Preferably, in step S5, one of the following methods is used to verify and evaluate the accuracy of the LGBM model: five-fold cross-validation, K-fold cross-validation, leave-one-out method, hierarchical K-fold cross-validation, time series cross-validation, bootstrap method, and hold-out validation method.

[0012] Preferably, in step S5, the five-fold cross-validation method specifically includes the following steps: S51. Divide the entire dataset into five equal subsets and perform five rounds of training and validation. In each round, select one subset as the validation set and use the remaining four subsets as the training set. Use the remaining subset to validate the model's performance and calculate the corresponding performance metrics. S52. Summarize the results of the five rounds of verification and calculate the average performance index of the five rounds; S53. Finally, the model performance was further evaluated using the test set data. The performance metrics used were mean squared error (MSE), root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²). 2The formulas are as follows: (4) (5) (6) (7) In the formula For the true value, For predicted values, It is the mean of the averages.

[0013] Preferably, in step S6, the trained model is used to predict the longitudinal wave time difference of the well section without measured acoustic time difference, and the first prediction is the transverse wave time difference of the well section without measured acoustic time difference.

[0014] Preferably, the method further includes step S8, which involves filtering the longitudinal and transverse wave curves predicted in step S7 using a suitable filtering method.

[0015] Preferably, the filtering method is the Hamming window filtering method or the median filtering method.

[0016] Compared with the prior art, the beneficial effects of the present invention are: (1) The method for predicting the sonic transit time of ultra-shallow gas-bearing reservoirs by combining mRMR and LGBM algorithms provided by the present invention utilizes the mRMR algorithm, which can perform feature selection based on the correlation and redundancy of data. It selects the feature subsets that are most correlated with the transit time of P-waves or S-waves and have the lowest redundancy among them from the original logging data, avoiding the conventional linear correlation analysis steps, which helps to reduce model complexity and improve prediction accuracy.

[0017] (2) The acoustic time difference prediction method for ultra-shallow gas-bearing reservoirs provided by the present invention, which combines mRMR and LGBM algorithms, uses the ensemble learning LGBM algorithm, adopts a tree-based learner as a weak learner, and has a regularization term to prevent overfitting, thus exhibiting good stability.

[0018] (3) The method for predicting the acoustic time difference of ultra-shallow gas-bearing reservoirs by combining mRMR and LGBM algorithms provided by this invention adopts a gradient boosting strategy based on histograms. Compared with other tree-based ensemble learning algorithms (such as random forest, gradient boosting tree, etc.), it has higher computational efficiency, can quickly build and train models on large datasets, improve the efficiency of data processing and analysis, has high prediction accuracy, is fast, and can accurately predict the time difference of P-waves and S-waves. Attached Figure Description

[0019] Figure 1 This is a flowchart illustrating the method for predicting the acoustic transit time of ultra-shallow, non-diagenetic layers using the mRMR algorithm and its ensemble learning in this invention. Figure 2When predicting the P-wave time difference after collecting data for the target area of ​​this invention, the importance score ranking is calculated using the mRMR algorithm. Figure 3 When predicting shear wave time difference after collecting data for the target area of ​​this invention, the importance score ranking is calculated using the mRMR algorithm. Figure 4 This is a graph showing the measured P-wave time difference versus the predicted P-wave time difference in the five-fold cross-validation of the LGBM model in this invention. Figure 5 This is a graph showing the measured shear wave time difference versus the predicted shear wave time difference results of the five-fold cross-validation of the LGBM model in this invention. Figure 6 This is a graph showing the measured P-wave time difference versus the predicted P-wave time difference in the LGBM model test set of this invention. Figure 7 This is a graph showing the measured shear wave time difference versus the predicted shear wave time difference in the LGBM model test set of this invention. Figure 8 This is a diagram showing the predicted P-wave and S-wave time difference of the unmeasured time difference well section in the target area according to the present invention. The above Figure 4 , Figure 5 , Figure 6 , Figure 7 The thick solid lines are all fitted straight lines, and the thin solid lines are the lines connecting the perfect prediction points. In the equation, y is the vertical coordinate of each graph, x is the horizontal coordinate, and R² is the coefficient of determination. The coefficient of determination for the predicted P-wave time difference on the test set is 0.85, which is inferior to the predicted S-wave time difference of 0.94. It is speculated that this is because the P-wave time difference varies greatly between the reservoir and non-reservoir sections, affecting the prediction effect. Detailed Implementation

[0020] The present invention will be further described below with reference to specific embodiments.

[0021] Example 1 like Figure 1 As shown, a method for predicting acoustic transit time in ultra-shallow gas-bearing reservoirs using a combined mRMR and LGBM model includes the following steps: S1. Collect raw logging data of the target layer, wherein the raw logging data includes multiple features.

[0022] S2. Perform data cleaning on the features, remove outliers, and standardize all features using data standardization / normalization methods.

[0023] S3. Use the mRMR algorithm to select features for acoustic time difference. Based on the correlation between acoustic time difference and each feature, select representative features.

[0024] In step S3, the mRMR (Minimum Redundancy Maximum Relevance) algorithm is a feature selection method designed to choose the most representative features from high-dimensional data while minimizing feature redundancy. It is commonly used in machine learning tasks such as classification and regression to improve model performance and computational efficiency. Figure 2 As shown, the extracted longitudinal wave features are the most representative. Figure 3 It is the extraction of the most representative features of transverse waves.

[0025] S4. Divide the representative features obtained in step S3 into a training set and a test set. Use the features in the training set as input and the P-wave time difference curve as output to train the LGBM model using the training set data.

[0026] In step S4, the obtained representative features can be divided into training and test sets in an 8:2 ratio, but this ratio is not always necessary (i.e., 80% training, 20% test). The ratio of training to test sets can be adjusted according to the size of the specific dataset, model requirements, and task objectives. Common ratios include: 7:3 (70% training, 30% test): suitable for medium-sized datasets, with a slightly larger test set for more accurate model performance evaluation; 9:1 (90% training, 10% test): suitable for smaller datasets, prioritizing sufficient training data; other ratios, such as 6:4 or custom ratios, depending on the specific task. A tree model is selected as the weak learner for LGBM, with 100 weak learners, a learning rate of 0.1, and a maximum tree depth of -1 (indicating an unrestricted maximum depth). The LGBM model is then trained.

[0027] S5. Verify and evaluate the accuracy of the LGBM model, and then use test set data to further evaluate the accuracy of the LGBM model; S6. Use the trained LGBM model to predict the acoustic transit time of the unmeasured acoustic transit time well section to obtain the predicted P-wave transit time curve. S7. Using the original P-wave time difference curve and the predicted P-wave time difference curve obtained in step S6 as feature inputs and the S-wave time difference curve as output, repeat steps S1 to S6 to finally obtain the P-wave and S-wave curves of the well section with unmeasured P-wave and S-wave time differences.

[0028] In step S7, after the cable acoustic remeasurement, the original P-wave time difference curves of many wells were obtained, but the original P-wave time difference curves of some wells were not obtained. Therefore, the prediction using the original P-wave time difference curves obtained in this part did not obtain the original P-wave time difference curves. Then, both of these original P-wave time difference curves were used as feature inputs to predict the shear waves of the well sections with unmeasured time differences.

[0029] Among these features, at least the following are included: natural gamma (GR, API), deep lateral resistivity (RLLD, Ω.m), shallow lateral resistivity (RLLS, Ω.m), shallow induced resistivity (MSFL, Ω.m), compensated neutrons (CNL, pu), and compensated density (DEN, g / cm³). 3 ), wellbore curve (CAL, in).

[0030] In addition, step S2 specifically includes the following steps: S21. For each feature, calculate the correlation between the feature and the acoustic time difference. The correlation index is mutual information. S22. For each pair of features, calculate the redundancy between features, measured by mutual information, where the formula for mutual information is: (1) in, It is a random variable and The joint probability distribution, It is a random variable and The marginal probability distribution; S23. Use the following objective function to select features: (2) in, It is the selected feature set. It is a feature With sound wave time difference Mutual information between them It is a feature With features Mutual information between them It is a set Size; S24. Continuously update the feature set through iteration. In each iteration, select the option that maximizes the objective function. The features are selected until a stopping condition is met (e.g., the number of selected features reaches a preset value). S25. Based on the sorting results of the mutual information of each input variable, select 1 to m input variables in the sequence as representative features. In specific examples, when performing P-wave prediction, the top 6 features by importance score are selected, including compensated density (DEN, g / cm3), natural gamma (GR, API), compensated neutron (CNL, pu), shallow lateral resistivity (RLLS, Ω.m), deep lateral resistivity (RLLD, Ω.m), shallow induced resistivity (MSFL, Ω.m), and compensated neutron (CNL, pu). In a specific example, when performing S-wave prediction, the top 7 features by importance score are selected, including P-wave transit time (DT, μs / ft), compensated density (DEN, g / cm3), natural gamma (GR, API), compensated neutron (CNL, pu), shallow lateral resistivity (RLLS, Ω.m), deep lateral resistivity (RLLD, Ω.m), shallow induced resistivity (MSFL, Ω.m), and compensated neutron (CNL, pu). The remaining features are not selected to reduce computational load and improve model running speed.

[0031] In step S3, outliers are values ​​that clearly do not conform to well logging patterns, such as -9999. After deleting outliers, a new value is calculated using the linear interpolation method. If outliers occur consecutively, all feature curve values ​​corresponding to this depth are deleted. Then, one of the following methods is used to standardize all features: Z-score, Min-Max normalization, RobustScaler, Log transform, and StandardScaler. It should be noted that Min-Max normalization scales the data to a fixed range (e.g., [0, 1]), suitable for models requiring uniform units or sensitive to data range. RobustScaler standardizes based on the median and interquartile range, suitable for processing data with many outliers.

[0032] Log transformation: Performs a logarithmic transformation on the data, suitable for cases with skewed data distribution. StandardScaler (a variant of the Z-score method): Similar to the Z-score, but the implementation may differ slightly depending on the tool used. The choice of method depends on the data distribution characteristics, the number of outliers, and the model's sensitivity to data scaling. LGBM is not sensitive to data scaling, but standardization helps improve model stability.

[0033] In addition, in step S3, the Z-score method standardizes all features: (3) in, For the feature data to be processed, The mean of the total data; Let Z be the standard deviation of the overall data, and Z be the standardized value of the data.

[0034] In step S5, one of the following methods is used to evaluate the accuracy of the LGBM model: five-fold cross-validation, K-fold cross-validation, leave-one-out cross-validation, stratified K-fold cross-validation, time series cross-validation, bootstrap method, and hold-out validation. It should be noted that: K-fold cross-validation (K-Fold Cross-Validation): Divides the data into K parts (K is usually 3, 5, 10, etc.), using K-1 parts for training and 1 part for testing each time, repeating K times. The value of K can be adjusted according to the amount of data. Leave-one-out cross-validation (LOOCV): Leaves one sample as the test set each time, using the remaining samples for training; suitable for small datasets, but computationally expensive. Stratified K-fold cross-validation (Stratified K-Fold Cross-Validation): In classification tasks, ensures that the class distributions of the training and test sets are consistent; suitable for imbalanced datasets. Time series cross-validation (Time Series Cross-Validation): For time series data (such as well logging data which may involve depth sequences), the data is divided in time / depth order to avoid future data leakage. Bootstrap method: Generates multiple training and test sets by sampling with replacement to evaluate model performance; suitable for small datasets. Hold-out validation: Directly splits the data into training, validation, and test sets (e.g., 8:1:1), but may be sensitive to the splitting. Five-fold cross-validation is a compromise, balancing computational cost and evaluation stability. The choice of specific method needs to consider the amount of data, data characteristics (e.g., time series nature), and computational resources.

[0035] Example 2 The difference from Example 1 is that, in this example, step S5 specifically includes the following steps: S51. Divide the entire dataset into five equal subsets and perform five rounds of training and validation. In each round, select one subset as the validation set and use the remaining four subsets as the training set. Use the remaining subset to validate the model's performance and calculate the corresponding performance metrics. S52. Summarize the results of the five rounds of verification and calculate the average performance index of the five rounds; S53. Finally, the model performance was further evaluated using the test set data. The performance metrics used were mean squared error (MSE), root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²). 2 The formulas are as follows: (4) (5) (6) (7) In the formula For the true value, For predicted values, It is the mean of the averages.

[0036] In step S6, the trained model is used to predict the longitudinal wave time difference of the well section without measured acoustic time difference, and the first prediction is the transverse wave time difference of the well section without measured acoustic time difference.

[0037] like Figure 4 As shown in the figure, the fitting comparison diagram of the predicted P-wave time difference and the measured P-wave time difference of the validation set in the five-fold cross-validation is as follows. Figure 6 The figure shown is a comparison of the predicted P-wave time difference and the measured P-wave time difference fitting for the test set.

[0038] like Figure 5 and Figure 7 These are, respectively, the fitting comparison diagrams of the predicted and measured shear wave time difference of the validation set and the fitting comparison diagram of the predicted and measured shear wave time difference of the test set in the five-fold cross-validation obtained in step S5 after repeating steps S1 to S6.

[0039] Example 3 The difference from Example 1 is that it also includes step S8, in which the longitudinal and transverse wave curves predicted in step S7 are filtered using an appropriate filtering method. Figure 8 This study focuses on a well in the Lingshui 36-1 gas field within the Lingshui Depression. The target layer is the Ledong Formation, a deep-water, shallow-water reservoir with high clay content. The neutron density curve exhibits a reverse envelope due to the "digger effect," and resistivity values ​​are relatively high in the reservoir section. Predictive P-wave and S-wave curves show that in the gas layer, resistivity is high and P-wave and S-wave transit times increase significantly. In the mixed layer, P-wave and S-wave transit times are lower than in the gas layer, and no neutron density envelope phenomenon is observed. Overall, the predicted P-wave and S-wave curves correspond well with other curves.

[0040] It should be noted that step S8 is an optional optimization step, and its necessity depends on the quality of the predicted curve and the requirements of the application scenario. Filtering in shallow, non-diagenetic reservoirs can improve the quality of the acoustic curve and reduce spikes. Its necessity depends on the following factors: Data noise: The predicted P-wave and S-wave time difference curves may contain irregular fluctuations due to model errors or input data noise. Filtering can smooth the curves, reduce the impact of noise, and improve the reliability of the curves in practical applications.

[0041] Geological application requirements: In shallow, unformed gas-bearing reservoirs, the smoothness of the sonic transit time curve can have a significant impact on subsequent geological interpretations (such as reservoir evaluation and seismic wave propagation analysis). If the application scenario requires high curve smoothness, filtering is necessary.

[0042] Model performance: If the LGBM model's predictions are already very accurate and have low noise, then filtering may not be necessary.

[0043] The filtering method is either Hamming window filtering or median filtering.

[0044] In the specific implementation of the above embodiments, the technical features can be combined in any non-contradictory way. For the sake of brevity, not all possible combinations of the above technical features are described. However, as long as the combination of these technical features is not contradictory, it should be considered to be within the scope of this specification.

[0045] Obviously, the above embodiments of the present invention are merely examples for the purpose of clearly illustrating the present invention, and are not intended to limit the embodiments of the present invention. Those skilled in the art will appreciate that other variations or modifications can be made based on the above description. It is not necessary and impossible to enumerate all embodiments here. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the claims of the present invention.

Claims

1. A method for predicting acoustic transit time in ultra-shallow gas-bearing reservoirs using a combined mRMR and LGBM model, characterized in that, Includes the following steps: S1. Collect raw logging data of the target layer segment, wherein the raw logging data includes multiple features; S2. Perform data cleaning on the features, remove outliers, and standardize all features using data standardization / normalization methods; S3. Use the mRMR algorithm to select features for acoustic time difference. Based on the correlation between acoustic time difference and each feature, select representative features. S4. Divide the representative features obtained in step S3 into a training set and a test set. Use the features in the training set as input and the longitudinal wave time difference curve as output to train the LGBM model using the training set data. S5. Verify and evaluate the accuracy of the LGBM model, and then use test set data to further evaluate the accuracy of the LGBM model; S6. Use the trained LGBM model to predict the acoustic transit time of the unmeasured acoustic transit time well section to obtain the predicted P-wave transit time curve. S7. Using the original P-wave time difference curve and the predicted P-wave time difference curve obtained in step S6 as feature inputs and the S-wave time difference curve as output, repeat steps S1 to S6 to finally obtain the P-wave and S-wave curves of the well section with unmeasured P-wave and S-wave time differences.

2. The method for predicting acoustic transit time in ultra-shallow gas-bearing reservoirs using a combined mRMR and LGBM model as described in claim 1, characterized in that, The multiple features include at least natural gamma, deep lateral resistivity, shallow lateral resistivity, shallow induced resistivity, compensated neutrons, compensated density, and wellbore curve.

3. The method for predicting acoustic transit time in ultra-shallow gas-bearing reservoirs using a combined mRMR and LGBM model according to claim 2, characterized in that, Step S2 specifically includes the following steps: S21. For each feature, calculate the correlation between the feature and the acoustic time difference. The correlation index is mutual information. S22. For each pair of features, calculate the redundancy between features, measured by mutual information, where the formula for mutual information is: (1) in, It is a random variable and The joint probability distribution, It is a random variable and The marginal probability distribution; S23. Use the following objective function to select features: (2) in, It is the selected feature set. It is a feature With sound wave time difference Mutual information between them It is a feature With features Mutual information between them It is a set Size; S24. Continuously update the feature set through iteration. In each iteration, select the option that maximizes the objective function. The characteristic is that it continues until the stopping condition is met; S25. Based on the sorting results of the mutual information of each input variable, select 1 to m input variables in the sequence as representative features.

4. The method for predicting acoustic transit time in ultra-shallow gas-bearing reservoirs using a combined mRMR and LGBM model according to claim 1, characterized in that, In step S3, the outlier values ​​are values ​​that clearly do not conform to the logging rules. After deleting the outlier values, the new values ​​are calculated using the linear interpolation method. If outlier points appear continuously, all feature curve values ​​corresponding to this depth are deleted. Then, one of the following methods is used to standardize all features: Z-score method, Min-Max normalization, RobustScaler, Log transform, and StandardScaler.

5. The method for predicting acoustic transit time in ultra-shallow gas-bearing reservoirs using a combined mRMR and LGBM model according to claim 4, characterized in that, In step S3, the Z-score method standardizes all features: (3) in, For the feature data to be processed, The mean of the total data; Let Z be the standard deviation of the overall data, and Z be the standardized value of the data.

6. The method for predicting acoustic transit time in ultra-shallow gas-bearing reservoirs using a combined mRMR and LGBM model according to claim 1, characterized in that, In step S5, the accuracy of the LGBM model is evaluated using one of the following methods: five-fold cross-validation, K-fold cross-validation, leave-one-out method, hierarchical K-fold cross-validation, time series cross-validation, bootstrap method, and hold-out validation method.

7. The method for predicting acoustic transit time in ultra-shallow gas-bearing reservoirs using a combined mRMR and LGBM model as described in claim 6, characterized in that, In step S5, the five-fold cross-validation method specifically includes the following steps: S51. Divide the entire dataset into five equal subsets and perform five rounds of training and validation. In each round, select one subset as the validation set and use the remaining four subsets as the training set. Use the remaining subset to validate the model's performance and calculate the corresponding performance metrics. S52. Summarize the results of the five rounds of verification and calculate the average performance index of the five rounds; S53. Finally, the model performance was further evaluated using the test set data. The performance metrics used were mean squared error (MSE), root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²). 2 The formulas are as follows: (4) (5) (6) (7) In the formula is the true value, For predicted values, It is the mean of the averages.

8. The method for predicting acoustic transit time in ultra-shallow gas-bearing reservoirs using a combined mRMR and LGBM model according to claim 1, characterized in that, In step S6, the trained model is used to predict the longitudinal wave time difference of the well section without measured acoustic time difference, and the first prediction is the transverse wave time difference of the well section without measured acoustic time difference.

9. The method for predicting acoustic transit time of ultra-shallow gas-bearing reservoirs using a combined mRMR and LGBM model according to any one of claims 1 to 8, characterized in that, It also includes step S8, which involves filtering the longitudinal and transverse wave curves predicted in step S7 using an appropriate filtering method.

10. The method for predicting acoustic transit time in ultra-shallow gas-bearing reservoirs using a combined mRMR and LGBM model according to claim 9, characterized in that, The filtering method is either Hamming window filtering or median filtering.

Citation Information

Patent Citations

  • Transverse wave time difference prediction method based on SSA-ELM algorithm

    CN114488311A

  • Continental facies shale gas reservoir logging shear wave time difference prediction method

    CN117148473A

  • Transverse wave time difference prediction method and device

    CN117669785A

  • Shale shear wave velocity intelligent prediction method based on class activation diagram combined physical constraint

    CN118465832A