Ground surface settlement prediction method in oil well exploitation process based on machine learning

Through machine learning, the complexity of surface settlement prediction in oil well mining is solved, and accurate prediction and dynamic update of surface settlement are achieved, providing effective support for oil field management.

CN120336839APending Publication Date: 2025-07-18XI'AN PETROLEUM UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510408436.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the oil well mining, the model input depends on a single data source, and ignores the spatiotemporal coupling characteristics of multi-source data. Traditional algorithms are sensitive to noise data and lack dynamic feedback capabilities, making it difficult to meet the needs of intelligent oilfield management.

Method used

Using machine learning methods, multi-source data is integrated, including settlement scale, settlement dynamic monitoring, geological mechanics parameters and environmental influencing factors, the model is trained using a random forest algorithm, and modified with geological dynamic data to generate a surface settlement change trend chart.

Benefits of technology

Accurate prediction of surface settlement is achieved, geological disaster prevention and control and land resource management are supported, and the accuracy and adaptability of prediction are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336839A_ABST
    Figure CN120336839A_ABST
Patent Text Reader

Abstract

The invention relates to a machine learning-based ground surface settlement prediction method in an oil well exploitation process, and the method comprises the steps: collecting historical ground surface settlement data in the oil well exploitation process, carrying out the missing value, abnormal value and normalization processing of the historical ground surface settlement data, and constructing a data set; feature extraction is carried out on the data set, dimension reduction processing is carried out on the extracted features, and a low-dimension feature set is constructed; based on the low-dimensional feature set, training a random forest algorithm, constructing a settlement prediction model, and performing performance evaluation and optimization on the settlement prediction model to obtain a final settlement prediction model; and obtaining target ground surface settlement data in an oil well exploitation process, inputting the target ground surface settlement data into the final settlement prediction model, predicting a ground surface settlement change trend, and outputting the ground surface settlement change trend as a ground surface settlement change trend chart. According to the method, accurate prediction of ground surface settlement is realized by fusing multi-source data and an advanced machine learning technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of machine learning and oil well exploitation, and particularly relates to a method for predicting surface subsidence during the oil well exploitation process based on machine learning. Background Art

[0002] During the development of oil and gas resources, long-term exploitation of oil wells will cause changes in underground reservoir pressure and deformation of the formation structure, resulting in surface subsidence. Such subsidence may cause structural damage to surrounding infrastructure (such as pipelines, roads, buildings), and at the same time trigger environmental risks such as groundwater pollution and surface collapse. Especially in densely developed areas or ecologically sensitive regions, accurately predicting the subsidence trend is of great significance for formulating protective measures, optimizing the exploitation plan, and reducing economic losses. Predicting surface subsidence during oil well exploitation is a complex technical problem, and currently mainly relies on physical models (such as the finite element method, elastoplastic theory) and statistical regression methods. Physical models need to be based on accurate geomechanical parameters (such as rock modulus, pore pressure), but in the actual exploitation environment, formation parameters often have uncertainties, resulting in limited prediction accuracy of the models. In addition, the multi-field coupling effect under complex geological conditions (such as faults, heterogeneous rock formations) is difficult to accurately describe through simplified equations, and the calculation is time-consuming and difficult to update dynamically. Although statistical methods can establish empirical relationships through historical data, they cannot adapt to complex non-linear laws and have insufficient generalization ability.

[0003] In recent years, machine learning has shown advantages in solving non-linear and high-dimensional engineering problems. Its data-driven characteristics can effectively mine the implicit relationships among exploitation parameters, geological conditions, and subsidence amounts. However, existing research still has significant defects. The model input mostly relies on a single data source, ignoring the spatio-temporal coupling characteristics of multi-source data; traditional algorithms are sensitive to noise data and lack the ability to model dynamic feedback during long-term exploitation; most models are mainly static predictions and cannot update the prediction results in real time by fusing exploitation dynamic data, making it difficult to meet the intelligent management requirements of oil fields. Therefore, in view of the above technical problems, the present invention proposes a method for predicting surface subsidence during the oil well exploitation process based on machine learning. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for predicting surface subsidence during the oil well exploitation process based on machine learning, which integrates multi-source data and advanced machine learning technologies to achieve accurate prediction of surface subsidence.

[0005] To achieve the above purpose, the present invention provides the following solution:

[0006] A method for predicting surface subsidence during the oil well exploitation process based on machine learning, comprising:

[0007] Collect historical surface settlement data during the oil well exploitation process, perform missing value, outlier, and normalization processing on the historical surface settlement data, and construct a dataset;

[0008] Extract features from the dataset and perform dimensionality reduction processing on the extracted features to construct a low-dimensional feature set;

[0009] Train a random forest algorithm based on the low-dimensional feature set, construct a settlement prediction model, and perform performance evaluation and optimization on the settlement prediction model to obtain a final settlement prediction model;

[0010] Obtain target surface settlement data during the oil well exploitation process, input the target surface settlement data into the final settlement prediction model, predict the changing trend of surface settlement, and output the surface settlement changing trend as a surface settlement changing trend diagram.

[0011] Preferably, both the historical surface settlement data and the target surface settlement data include: settlement scale and spatial distribution data, settlement dynamic monitoring data, geomechanical parameter data, surface deformation parameter data, and environmental impact factor data.

[0012] Preferably, the missing value processing of the historical surface settlement data includes:

[0013] Identify the missing values in the historical surface settlement data;

[0014] Use a normality test method to judge the distribution type of the historical surface settlement data. If it shows a normal distribution, calculate the mean and fill the missing values with the mean; if it shows a skewed distribution, calculate the median and fill the missing values with the median.

[0015] Preferably, the outlier processing of the historical surface settlement data includes:

[0016] Use a preset threshold to detect outliers in the historical surface settlement data after missing value processing, and judge the outliers in combination with the distribution characteristics of geological features. If the outliers conform to the distribution characteristics of geological features, they are real outliers and are retained; if the outliers do not conform to the distribution characteristics of geological features, they are interference outliers and are filled with the median.

[0017] Preferably, the normalization processing of the historical surface settlement data includes:

[0018] Perform feature analysis on the historical surface settlement data after outlier processing to obtain the feature dimension difference and the feature distribution balance situation;

[0019] For data with a characteristic dimension difference reaching a preset difference value, the Z-score standardization method is used for processing; for data with a characteristic distribution balance reaching a preset balance value, the Min-Max standardization method is used for processing.

[0020] Preferably, feature extraction is performed on the dataset, and dimensionality reduction processing is performed on the extracted features to construct a low-dimensional feature set including:

[0021] Geological features, mining activity features, and groundwater dynamic features are extracted based on the dataset. Among them, the geological features include formation lithology, structural features, and hydrogeological conditions, the mining activity features include the advancing speed of the coal mining face and the mining intensity, and the groundwater dynamic features include water level, water quality, and flow rate changes;

[0022] Calculate the correlation coefficient matrix between features, and perform correlation judgment between features based on the correlation coefficient matrix. If the correlation reaches a preset correlation value, the principal component analysis method is used to perform dimensionality reduction processing on the corresponding features to obtain a low-dimensional feature set.

[0023] Preferably, based on the low-dimensional feature set, a random forest algorithm is trained to construct a settlement prediction model, and the performance of the settlement prediction model is evaluated and optimized to obtain the final settlement prediction model including:

[0024] S1. Input the low-dimensional feature set into the random forest algorithm for training to obtain an initial settlement prediction model;

[0025] S2. Judge the overfitting situation of the initial settlement prediction model. If overfitting exists, adjust the number and depth of decision trees in the random forest algorithm and train again until overfitting does not exist to obtain the settlement prediction model;

[0026] S3. Use the root mean square error and mean absolute error to evaluate the performance of the settlement prediction model. If the error does not exceed the preset threshold, save the settlement prediction model; if the error exceeds the preset threshold, use the feature extraction method to adjust the low-dimensional feature set to obtain an optimized feature set;

[0027] S4. Input the optimized feature set into the random forest algorithm for training again, and repeat S1 - S3 until the final settlement prediction model is obtained.

[0028] Preferably, input the target surface settlement data into the final settlement prediction model to predict the surface settlement change trend including:

[0029] Input the target surface settlement data into the final settlement prediction model and output the surface settlement prediction value;

[0030] Perform uncertainty judgment and consistency judgment on the predicted surface settlement value, and correct the predicted surface settlement value based on the judgment result to obtain the change trend of surface settlement.

[0031] The beneficial effects of the present invention are as follows:

[0032] The present invention first preprocesses the data, including missing value filling, outlier detection, and data normalization; then extracts geological features, mining activity features, and groundwater dynamic features from the cleaned data, and uses the principal component analysis method for dimensionality reduction; then, trains a model using the random forest algorithm and optimizes the model performance by adjusting hyperparameters or adding regularization terms; finally, applies the optimized model to the actual scenario, combines geological dynamic data for correction, and generates a surface settlement change trend map. By integrating multi-source data and advanced machine learning techniques, the present invention realizes accurate prediction of surface settlement and provides strong support for geological disaster prevention and control and land resource management. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0034] Figure 1 It is a flowchart of a method for predicting surface settlement during oil well production based on machine learning according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the protection scope of the present invention.

[0036] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0037] This embodiment provides a method for predicting surface settlement during oil well production based on machine learning, as Figure 1 shown, including:

[0038] Collect historical surface settlement data during oil well production, perform missing value, outlier, and normalization processing on the historical surface settlement data, and construct a data set;

[0039] Extract features from the dataset, perform dimensionality reduction on the extracted features, and construct a low-dimensional feature set;

[0040] Train a random forest algorithm based on the low-dimensional feature set, construct a settlement prediction model, evaluate and optimize the performance of the settlement prediction model, and obtain the final settlement prediction model;

[0041] Obtain the target surface settlement data during the oil well exploitation process, input the target surface settlement data into the final settlement prediction model, predict the change trend of surface settlement, and output the surface settlement change trend as a surface settlement change trend map.

[0042] Specifically, the method proposed in this embodiment first preprocesses the data, including missing value filling, outlier detection, and data normalization; then extracts geological features, mining activity features, and groundwater dynamic features from the cleaned data, and performs dimensionality reduction using the principal component analysis method; next, trains the model using the random forest algorithm, and optimizes the model performance by adjusting hyperparameters or adding regularization terms; finally, applies the optimized model to the actual scenario, corrects it in combination with geological dynamic data, and generates a surface settlement change trend map. The method proposed in this embodiment realizes the accurate prediction of surface settlement by integrating multi-source data and advanced machine learning technologies, providing strong support for geological disaster prevention and control and land resource management.

[0043] Furthermore, both the historical surface settlement data and the target surface settlement data include: settlement scale and spatial distribution data, settlement dynamic monitoring data, geomechanical parameter data, surface deformation parameter data, and environmental impact factor data.

[0044] Specifically, the acquisition of surface settlement data in this embodiment specifically includes the following contents:

[0045] (1) The settlement scale and spatial distribution data include:

[0046] The location and scope of the settlement funnel: Continuous exploitation of oil wells will form a significant surface settlement funnel, and its spatial distribution is directly related to the exploitation intensity of underground reservoirs. For example, typical settlement funnels have been monitored in the Shuguang Oil Production Plant and Huanxiling Oil Production Plant of the Liaohe Oilfield.

[0047] The maximum subsidence value: It refers to the maximum vertical displacement amount within the surface settlement area, usually related to factors such as reservoir thickness and exploitation depth.

[0048] (2) The settlement dynamic monitoring data include:

[0049] Settlement rate: The annual or phased rate of surface settlement can be obtained through technologies such as InSAR (Interferometric Synthetic Aperture Radar). For example, C-band and L-band radar data can be used for monitoring different vegetation-covered areas respectively.

[0050] Cumulative settlement: It reflects the total amount of surface settlement during the long-term mining process and needs to be calculated by combining time-series monitoring data.

[0051] (3) Geomechanical parameter data includes:

[0052] Reservoir parameter inversion data: It includes formation compressibility, porosity change, etc. Through settlement data inversion, the mechanical properties of the reservoir and the impact of mining on the geological structure can be evaluated.

[0053] Depth-to-thickness ratio (H / m) and width-to-depth ratio (D / H): They are used to describe the relationship between the mining depth and the reservoir thickness, and the ratio of the lateral expansion to the depth of the settlement area. These parameters are widely used in coal mine research and may have reference value for the settlement analysis of oil wells.

[0054] (4) Surface deformation parameter data includes:

[0055] Horizontal movement and tilt deformation: Surface settlement is often accompanied by horizontal displacement and tilt deformation, and such data can reflect the non-uniformity of settlement.

[0056] Curvature and horizontal strain: They are used to quantify the degree of surface bending and tensile / compressive deformation, which is crucial for evaluating the risk of infrastructure damage.

[0057] (5) Environmental impact factor data includes:

[0058] Ratio of the thickness of the loose layer to the mining depth (Hs / H): The ratio of the thickness of the loose overburden layer to the mining depth affects the settlement transfer efficiency. This parameter has been proven to be related to the surface subsidence coefficient in coal mine research.

[0059] Furthermore, the processing of missing values in the historical surface settlement data includes:

[0060] Identifying the missing values in the historical surface settlement data;

[0061] Using the normality test method to judge the distribution type of the historical surface settlement data. If it shows a normal distribution, calculate the mean value and fill the missing values with the mean value; if it shows a skewed distribution, calculate the median and fill the missing values with the median.

[0062] The processing of outliers in the historical surface settlement data includes:

[0063] Using a preset threshold, detect outliers in the historical surface subsidence data after missing value processing, and judge the outliers in combination with the distribution characteristics of geological features. If the outliers conform to the distribution characteristics of geological features, they are real outliers and are retained; if the outliers do not conform to the distribution characteristics of geological features, they are interference outliers and are filled with the median.

[0064] The normalization process of the historical surface subsidence data includes:

[0065] Conduct feature analysis on the historical surface subsidence data after outlier processing to obtain the difference in feature dimensions and the balance of feature distribution;

[0066] For data with the difference in feature dimensions reaching the preset difference value, use the Z-score standardization method for processing; for data with the balance of feature distribution reaching the preset balance value, use the Min-Max standardization method for processing.

[0067] Specifically, in this embodiment, after obtaining the historical surface subsidence data, missing value processing is first performed. Identify the missing values in the historical surface subsidence data, and according to the data distribution characteristics, use the normality test method to judge the data distribution type. If the data is normally distributed, calculate the data mean and fill the missing values with the mean; if the data is skewed, calculate the data median and fill the missing values with the median.

[0068] Then, perform outlier processing on the historical surface subsidence data after missing value filling, conduct outlier detection, and use the preset threshold rule to judge whether the data points exceed the range. If the data points exceed the preset threshold, analyze the distribution characteristics of the data points in combination with geological features to judge whether they are real outliers. For the data points judged as real outliers, retain their original values; for the data points judged as interference outliers, use the median filling method for processing. In this embodiment, the standard deviation method is used to set the threshold rule for outlier detection. For example, in the groundwater level monitoring data, if the value of a certain monitoring point suddenly rises to three times the surrounding water level, exceeding the preset range of plus or minus three standard deviations, it is necessary to analyze in combination with geological conditions. If there are faults or fissure developments in this area, it may lead to enhanced hydraulic connection and cause a sudden increase in the water level. In this case, the outliers should be retained; if the geological conditions in this area are stable and there are no special structures, it may be an interference value caused by equipment failure and should be replaced with the median. In the analysis of the chemical characteristics of groundwater in the mining area, the concentration of sulfate ions at a certain monitoring point suddenly increases and exceeds the normal range. Through analysis, it is found that this point is near a sulfur-bearing formation and there is an active fault, so this outlier is retained; in a similar situation, if the monitoring point is far from the sulfur-bearing formation and the structure is stable, the outlier is regarded as an interference value for processing.

[0069] Finally, the historical surface subsidence data after outlier processing is normalized, and the differences in characteristic dimensions are analyzed. If the dimensional differences are significant, the Z-score standardization method is used to process the data; if the feature distribution is unbalanced, the Min-Max standardization method is used to normalize the data. For example, the dimensional differences in geological data characteristics are common in parameters such as depth, pressure, and porosity. For instance, in a certain area, the depth range is from one thousand to three thousand meters, the pressure range is from ten to thirty megapascals, and the porosity range is from five to twenty percent. For such data with significant dimensional differences, standardization processing can make different features comparable. In practical applications, after standardization, the depth data is distributed near zero, and the pressure data is also converted to the same scale, facilitating subsequent analysis. For unbalanced features, such as permeability ranging from zero point one millidarcy to one thousand millidarcy, with a very uneven distribution, normalization processing can map the data to the interval from zero to one. Through normalization, the data differences between the high-permeability section and the low-permeability section are reasonably compressed, avoiding the excessive influence of extreme values on the analysis results.

[0070] Furthermore, feature extraction is performed on the dataset, and dimensionality reduction is performed on the extracted features to construct a low-dimensional feature set, including:

[0071] Geological features, mining activity features, and groundwater dynamic features are extracted based on the dataset. Among them, the geological features include formation lithology, structural features, and hydrogeological conditions, the mining activity features include the advancing speed of the coal mining face and the mining intensity, and the groundwater dynamic features include water level, water quality, and flow rate changes;

[0072] The correlation coefficient matrix between features is calculated, and the correlation between features is judged based on the correlation coefficient matrix. If the correlation reaches the preset correlation value, the principal component analysis method is used to perform dimensionality reduction on the corresponding features to obtain a low-dimensional feature set.

[0073] Specifically, in this embodiment, the extraction of geological features mainly focuses on formation lithology, structural features, and hydrogeological conditions. Taking a certain mining area as an example, through the analysis of borehole records, it is found that this area is mainly composed of interbedded sandstone and shale, with local coal seams intercalated, and the structure is mainly normal faults. The characteristics of mining activities include indicators such as the advancing speed of the coal mining face and the mining intensity. The annual mining volume of this mining area is about one million tons, and the average advancing speed of the working face is three meters per day. The dynamic characteristics of groundwater include changes in water level, water quality, and flow rate. The annual change range of the groundwater level in this area is about ten meters. The correlation analysis shows that there is a significant negative correlation between the groundwater level and the mining intensity, and the correlation coefficient reaches more than 0.8, indicating that mining activities have an obvious impact on the groundwater system. There is also a strong correlation between the degree of structural development and the groundwater flow rate, and the correlation coefficient is about 0.7, indicating that the fracture zone may become an important water-conducting channel. Principal component analysis retains key information through dimensionality reduction, selects principal components with eigenvalues greater than one, and the cumulative variance contribution rate exceeds 85%. In this embodiment, the first three principal components contain 88% of the information in the original data. The first principal component mainly reflects the impact of mining activities on groundwater, the second principal component reflects the geological structure characteristics, and the third principal component is related to the lithology combination.

[0074] Further, based on the low-dimensional feature set, a random forest algorithm is trained to construct a settlement prediction model, and the performance of the settlement prediction model is evaluated and optimized to obtain the final settlement prediction model, including:

[0075] S1. Input the low-dimensional feature set into the random forest algorithm for training to obtain an initial settlement prediction model;

[0076] S2. Judge the overfitting situation of the initial settlement prediction model. If overfitting exists, adjust the number and depth of decision trees in the random forest algorithm and train again until overfitting does not exist to obtain the settlement prediction model;

[0077] S3. Evaluate the performance of the settlement prediction model using the root mean square error and the mean absolute error. If the error does not exceed the preset threshold, save the settlement prediction model; if the error exceeds the preset threshold, adjust the low-dimensional feature set using the feature extraction method to obtain an optimized feature set;

[0078] S4. Input the optimized feature set into the random forest algorithm for training again, and repeat S1 - S3 until the final settlement prediction model is obtained.

[0079] Specifically, the random forest algorithm trains the model by constructing multiple decision trees. Each tree selects training samples and features through random sampling, and obtains the final result by voting or averaging. When the model is overfitting, it can be optimized by adjusting hyperparameters such as the number and depth of the trees. For example, the maximum depth of the tree can be adjusted from unlimited to ten layers, or the sample weight penalty term can be increased to make the model more generalized. Taking the prediction of surface subsidence in a certain mining area as an example, by restricting the maximum depth of a single tree to eight layers, the mean square error of the model on the test set is reduced from 0.3 to 0.15.

[0080] The root mean square error reflects the deviation degree between the predicted value and the actual value, and the mean absolute error represents the average level of the absolute value of the prediction error. If the root mean square error exceeds the preset threshold, the principal component analysis method is used to optimize the low-dimensional feature set. The optimized feature set can better express the essential features of the data and improve the prediction accuracy of the model.

[0081] Furthermore, inputting the target surface subsidence data into the final subsidence prediction model, the prediction of the surface subsidence change trend includes:

[0082] Inputting the target surface subsidence data into the final subsidence prediction model, and outputting the surface subsidence prediction value;

[0083] Performing uncertainty judgment and consistency judgment on the surface subsidence prediction value, and correcting the surface subsidence prediction value based on the judgment result to obtain the surface subsidence change trend.

[0084] Specifically, applying the final subsidence prediction model to the actual scenario to judge the uncertainty and consistency of the prediction result. Extracting the geological change rule data from the geological database to obtain the geological change rule value. Comparing the prediction result with the geological change rule value to judge their consistency. If the consistency meets the preset conditions, a surface subsidence visualization result is generated. If the consistency does not meet the preset conditions or the uncertainty is relatively high, dynamic geological information is extracted from the geological dynamic database, and the dynamic geological information is combined with the prediction result to generate a corrected data set to correct the model. Applying the corrected model again to obtain the final prediction result.

[0085] The above embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. A method for predicting surface subsidence during oil well exploitation based on machine learning, characterized in that, Including: Collecting historical surface subsidence data during the oil well exploitation process, processing the missing values, outliers, and normalization of the historical surface subsidence data, and constructing a dataset; Performing feature extraction on the dataset and performing dimensionality reduction processing on the extracted features to construct a low-dimensional feature set; Training a random forest algorithm based on the low-dimensional feature set, constructing a subsidence prediction model, and performing performance evaluation and optimization on the subsidence prediction model to obtain a final subsidence prediction model; Obtaining target surface subsidence data during the oil well exploitation process, inputting the target surface subsidence data into the final subsidence prediction model, predicting the change trend of surface subsidence, and outputting the surface subsidence change trend as a surface subsidence change trend diagram.

2. The method for predicting surface subsidence during oil well exploitation based on machine learning according to claim 1, wherein Both the historical surface subsidence data and the target surface subsidence data include: subsidence scale and spatial distribution data, subsidence dynamic monitoring data, geomechanical parameter data, surface deformation parameter data, and environmental impact factor data.

3. The method for predicting surface subsidence during oil well exploitation based on machine learning according to claim 1, wherein The processing of missing values for the historical surface subsidence data includes: Identifying the missing values in the historical surface subsidence data; Using a normality test method to judge the distribution type of the historical surface subsidence data. If it is normally distributed, calculate the mean and fill the missing values with the mean. If it is skewed, calculate the median and fill the missing values with the median.

4. The method for predicting surface subsidence during oil well exploitation based on machine learning according to claim 3, characterized in that, The processing of outliers for the historical surface subsidence data includes: Using a preset threshold to detect outliers in the historical surface subsidence data after missing value processing, and judging the outliers in combination with the distribution characteristics of geological features. If the outliers conform to the distribution characteristics of geological features, they are real outliers and are retained. If the outliers do not conform to the distribution characteristics of geological features, they are interference outliers and are filled with the median.

5. The method for predicting surface subsidence during oil well exploitation based on machine learning according to claim 4, wherein, The normalization processing of the historical surface subsidence data includes: Performing feature analysis on the historical surface subsidence data after outlier processing to obtain the feature dimension difference and the feature distribution balance; For the data with the feature dimension difference reaching the preset difference value, using the Z-score standardization method for processing; for the data with the feature distribution balance reaching the preset balance value, using the Min-Max standardization method for processing.

6. The method for predicting surface subsidence during oil well exploitation based on machine learning according to claim 1, wherein Performing feature extraction on the dataset and performing dimensionality reduction processing on the extracted features to construct a low-dimensional feature set includes: Extracting geological features, mining activity features, and groundwater dynamic features based on the dataset. Among them, the geological features include formation lithology, structural features, and hydrogeological conditions, the mining activity features include the advancing speed of the coal mining face and the mining intensity, and the groundwater dynamic features include water level, water quality, and flow rate changes; Calculating the correlation coefficient matrix between features and judging the correlation between features based on the correlation coefficient matrix. If the correlation reaches the preset correlation value, using the principal component analysis method to perform dimensionality reduction processing on the corresponding features to obtain a low-dimensional feature set.

7. The method for predicting surface subsidence during oil well exploitation based on machine learning according to claim 1, wherein Training a random forest algorithm based on the low-dimensional feature set, constructing a subsidence prediction model, and performing performance evaluation and optimization on the subsidence prediction model to obtain a final subsidence prediction model includes: S1. Input the low-dimensional feature set into the random forest algorithm for training to obtain an initial settlement prediction model; S2. Judge the overfitting situation of the initial settlement prediction model. If overfitting exists, adjust the number and depth of decision trees in the random forest algorithm and train again until overfitting does not exist to obtain the settlement prediction model; S3. Evaluate the performance of the settlement prediction model using the root mean square error and mean absolute error. If the error does not exceed the preset threshold, save the settlement prediction model; if the error exceeds the preset threshold, adjust the low-dimensional feature set using the feature extraction method to obtain an optimized feature set; S4. Input the optimized feature set into the random forest algorithm for training again, and repeat S1 - S3 until the final settlement prediction model is obtained.

8. The method for predicting surface subsidence during oil well exploitation based on machine learning according to claim 1, wherein Input the target ground settlement data into the final settlement prediction model. Predicting the ground settlement change trend includes: Input the target ground settlement data into the final settlement prediction model and output the ground settlement prediction value; Conduct uncertainty judgment and consistency judgment on the ground settlement prediction value, and correct the ground settlement prediction value based on the judgment results to obtain the ground settlement change trend.