Shale fracturing deformation machine learning prediction method based on mechanism model
By employing a machine learning-based prediction method for casing deformation in shale fracturing based on a mechanistic model, the efficiency and accuracy issues of casing deformation prediction in shale oil wells have been resolved. This method enables efficient and intelligent casing deformation prediction, supporting the efficient development of shale oil.
Patent Information
- Application Number
- CN202411886075.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-12-20
AI Technical Summary
Existing technologies cannot efficiently and universally predict casing deformation in shale oil wells. The 3D modeling process is cumbersome, the solution time is long and the results are unstable. Large-scale parameter calculations are not possible, and the error rate of casing deformation prediction is high.
The machine learning prediction method for shale fracturing and casing deformation based on mechanism model obtains a target simulation dataset that meets the acquisition criteria, performs preprocessing and feature selection, and combines machine learning algorithms to optimize numerical simulation results to predict casing deformation risk.
It improves the accuracy and efficiency of shale oil transformation prediction, reduces fracturing failures and resource waste, enhances development efficiency, realizes the application of intelligent and automated data processing processes, provides effective applications, and supports the efficient development of shale oil.
Smart Images

Figure CN119514388B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of shale casing deformation prediction, in particular to a shale fracturing casing deformation machine learning prediction method based on a mechanism model. BACKGROUND
[0002] In the process of reservoir fracturing reconstruction of shale oil well horizontal section, the hydraulic fracture tip induced stress changes the original stress state around the weak structural plane, thereby activating the weak structural plane within a certain probability, causing the weak structural plane to slip and damage, and triggering shear casing deformation; after the casing deformation occurs, due to the fact that the fracturing pipe string cannot be lowered to the designated position, part of the well section has to give up fracturing reconstruction, so the casing deformation problem seriously affects the efficient development of shale oil;
[0003] At present, for the casing deformation phenomenon caused by the fracturing process of shale oil well, numerical simulation method is mostly used to solve the slip-induced casing deformation, in the modeling process, in order to achieve more accurate results, three-dimensional modeling is often used to realize the relative position relationship between the actual weak structural plane and the wellbore; but in the process of three-dimensional modeling, there are the following problems: ① the relative position of the wellbore and the weak structural plane is not fixed, the model needs to be re-established every time the casing deformation is calculated, the process is tedious; ② the solving process of three-dimensional numerical model is complex, and the solving time is long; ③ when the initial and boundary conditions for solving are changed, the model may not converge, resulting in failure to solve; ④ the geometric structure of the weak structural plane is complex, and the grid may not be constructed; in summary, although the use of numerical simulation to directly solve the casing deformation risk and deformation under different engineering and geological parameters can provide more accurate results, it cannot be used for large-scale parameter calculation, and the solving time is long, with a high failure rate; therefore, it is particularly important to realize an efficient and universal shale oil well casing deformation prediction method.
[0004] Chinese patent application publication number CN114878340A discloses a numerical simulation method for shale hydraulic fracturing based on PFC discrete element method, including the following steps: Step 1: Establish an external wall unit, generate particle units within the wall, construct a circular region at the center, delete the particle units, and establish an injection wellbore; calibrate the shale microstructure parameters by comparing and analyzing the stress-strain curves output from uniaxial compression test results, Brazilian fracturing test results, and uniaxial compression simulation results and Brazilian fracturing simulation results; assign basic physical property parameters to the particle units, and bond adjacent particle units through parallel bonding units to generate a discrete element numerical model of the shale reservoir; Step 2: Set adjacent particle... The area enclosed by the unit cells is the fluid domain, and the vertical contact point between the particle units is defined as the fluid conduit. A network flow model is established on the discrete element numerical model of the shale reservoir. Step 3: Apply confining pressure to the shale boundary wall, set the injection velocity of the wellbore, and perform numerical simulation calculations of seepage-displacement-stress coupling. Step 4: Monitor the bonding state between particle units in real time to determine whether the bonding unit has failed. When the stress borne by the bonding unit is greater than the normal or tangential strength of the bonding unit, the bonding unit is considered to be in a state of failure, which is regarded as the generation of microcracks. Calculate the number of cracks and the hydraulic fracture propagation morphology. After the calculation is completed, export and save the simulation results of the hydraulic fracture propagation morphology of the shale reservoir.
[0005] Therefore, current shale oil well casing deformation prediction technology cannot achieve efficient and universally applicable prediction of shale oil well casing deformation. Summary of the Invention
[0006] Therefore, the purpose of this invention is to provide a machine learning prediction method for casing deformation in shale fracturing based on a mechanistic model, in order to overcome the problem that current shale oil casing deformation prediction technologies cannot achieve efficient and universal prediction of casing deformation in shale oil wells.
[0007] To achieve the above objectives, this invention provides a machine learning prediction method for shale fracturing casing deformation based on a mechanistic model, comprising:
[0008] Based on the characteristics of shale oil reservoirs, a target simulation dataset of shale oil reservoirs that meets the acquisition criteria is obtained. The target simulation dataset includes several sets of simulation features to be processed.
[0009] The target simulation dataset to be processed is preprocessed to obtain preprocessing results. Based on the preprocessing results, each set of simulation features to be processed is corrected to obtain the target actual simulation dataset, which includes several sets of actual simulation features.
[0010] Initial numerical simulations are performed on each set of actual simulated features to obtain initial numerical simulation results. Based on the initial numerical simulation results, the target shale oil well reservoir slippage variables under different influencing factors are determined.
[0011] acquire a plurality of invisible simulation feature sets of the influencing factors which cannot be measured actually, and perform inversion based on each invisible simulation feature set to optimize the initial numerical simulation result and obtain an actual numerical simulation result;
[0012] based on the selected initial prediction standard, perform prediction training and prediction testing on the actual numerical simulation result in sequence to optimize the initial prediction standard and obtain an actual prediction standard.
[0013] Further, the process of acquiring the target shale oil reservoir which meets the acquisition standard includes:
[0014] based on the characteristics of the shale oil reservoir, acquire each initial simulation feature set of the target shale oil reservoir;
[0015] for any initial simulation feature set, acquire the actual acquisition condition of the initial simulation feature set;
[0016] perform data labeling according to the actual acquisition condition;
[0017] determine a single acquisition condition based on the data labeling result, and determine to start a whole analysis mode, or a re-acquisition mode, or a preprocessing mode according to the single acquisition condition.
[0018] Further, the process of determining the single acquisition condition based on the data labeling result includes:
[0019] identify the error data, the missing data and the abnormal data in the actual acquisition condition, and perform corresponding cleaning and processing on the initial simulation feature set based on the identification result to obtain an intermediate simulation feature set;
[0020] perform type labeling on the error data, the missing data and the abnormal data, and acquire the number of each type of labeling;
[0021] determine an abnormal evaluation level of the initial simulation feature set according to the type labeling result, and perform a primary judgment on whether the initial simulation feature set meets the acquisition standard based on the abnormal evaluation level.
[0022] Further, the process of determining the single acquisition condition based on the data labeling result further includes:
[0023] calculate an acquisition evaluation value based on the primary judgment result and the number of each type of labeling, and perform a secondary judgment on whether the intermediate simulation feature set meets the acquisition standard according to the acquisition evaluation value; if yes, determine that the intermediate simulation feature set is the simulation feature set to be processed;
[0024] The type label result includes an error type, a missing type, and an abnormal type.
[0025] Further, the process of determining to start the overall analysis mode, or the re-collection mode, or the preprocessing mode according to the single collection condition includes:
[0026] According to the secondary judgment result, determine the actual qualified set number of the target initial simulation data set that meets the collection standard.
[0027] According to the actual qualified set number and a preset standard qualified set number, determine to start the overall analysis mode, or the re-collection mode, or the preprocessing mode.
[0028] Each initial simulation feature set constitutes the target initial simulation data set.
[0029] Further, according to the actual qualified set number and the total set number, calculate a collection qualified rate.
[0030] According to the first difference absolute value and a preset first evaluation value, determine to start the overall analysis mode, or the re-collection mode, or the preprocessing mode.
[0031] Based on the first difference absolute value being less than or equal to the first evaluation value, it is determined to start the overall analysis mode; based on the first difference absolute value being greater than the first evaluation value and the collection qualified rate being greater than a preset standard qualified rate, it is determined to start the preprocessing mode; based on the first difference absolute value being greater than the first evaluation value and the collection qualified rate being less than the standard qualified rate, it is determined to start the re-collection mode.
[0032] The first difference absolute value is an absolute value of a difference between the collection qualified rate and the standard qualified rate.
[0033] Further, based on starting the overall analysis mode, integrate a qualified number corresponding to each to-be-processed simulation feature set in the target initial simulation data set that meets the collection standard and an unqualified number corresponding to each abnormal simulation feature set in the target initial simulation data set that does not meet the collection standard.
[0034] According to the integration result, determine an overall abnormal score, and according to the overall abnormal score, determine to start the re-collection mode or the preprocessing mode.
[0035] Further, the process of preprocessing the target to-be-processed simulation data set to obtain a preprocessing result includes:
[0036] analyzing a first actual correlation degree between each to-be-processed simulation feature set, performing feature screening on the target to-be-processed simulation data set according to the first actual correlation degree, to obtain a target to-be-simulated data set;
[0037] The target to-be-simulated data set includes a plurality of to-be-simulated feature sets.
[0038] For any to-be-simulated feature set, analyze a second actual correlation degree between feature parameters in the to-be-simulated feature set, and perform feature screening on the to-be-simulated feature set according to the second actual correlation degree, to obtain the actual simulation feature set.
[0039] Further, the process of determining the target shale oil well reservoir slip variable under different influence factors based on the initial numerical simulation result includes:
[0040] determining actual geometric parameters and induced stress field distribution in the hydraulic fracturing process based on each actual simulation feature set;
[0041] determining the numerical range of the influence factor causing the weak structural surface to slip and the weak structural surface slip amount under different influence factors based on the actual geometric parameters and the induced stress field distribution;
[0042] determining the target shale oil well reservoir slip variable based on the numerical range of the influence factor causing the weak structural surface to slip and the weak structural surface slip amount under different influence factors.
[0043] Further, the process of performing inversion based on each invisible simulation feature set to optimize the initial numerical simulation result to obtain an actual numerical simulation result includes:
[0044] performing inversion on each invisible simulation feature set based on history matching to obtain the actual numerical simulation result;
[0045] comparing the actual numerical simulation result with the obtained actual numerical calculation result;
[0046] determining to optimize the initial prediction standard according to the comparison result, or performing history matching again.
[0047] Compared with existing technologies, the beneficial effects of this invention are as follows: By acquiring data based on shale oil reservoir characteristics, it ensures the high quality and relevance of the target simulation dataset, laying a solid foundation for subsequent simulation and prediction; it improves the accuracy and consistency of the data, further enhancing the precision of subsequent numerical simulations; it considers the shale oil well reservoir slip and casing variables under different influencing factors, enabling a comprehensive simulation of various situations that may occur during actual fracturing, helping to understand the complexity and diversity of casing deformation phenomena; by optimizing the initial numerical simulation results, it solves the problem of influencing factors that cannot be directly measured in actual engineering, making the simulation results closer to reality; through repeated optimization of the actual numerical simulation results, a reliable prediction standard is formed; it not only improves the accuracy and efficiency of prediction but also realizes the intelligence and automation of the prediction process; by accurately predicting the casing deformation risks and variables that may occur during shale oil well fracturing, it helps to formulate prevention and control measures in advance, reducing fracturing failures and well abandonment caused by casing deformation, thereby significantly reducing development costs and risks; through efficient numerical simulation and intelligent prediction optimization, it provides strong support for the efficient development of shale oil, helping to improve the overall efficiency and competitiveness of shale oil extraction.
[0048] By meticulously classifying influencing factors based on shale oil reservoir characteristics and acquiring each initial simulation feature set separately, the comprehensiveness and accuracy of the simulation data can be ensured. The refined data collection method helps to more realistically reflect the complexity and diversity of shale oil reservoirs, providing a reliable foundation for subsequent analysis. By labeling the actual acquisition of the initial simulation feature sets and flexibly selecting between overall analysis mode, re-acquisition mode, or preprocessing mode based on the labeling results, the data processing workflow becomes more flexible and efficient. The adaptive data processing mechanism can effectively address data quality issues under different conditions, improving the efficiency and accuracy of data processing. By pre-determining the influencing factors requiring numerical simulation and selectively acquiring each initial simulation feature set, the comprehensiveness and accuracy of the simulation data can be ensured. This approach avoids blindly collecting and processing irrelevant data, thereby reducing data collection and processing costs; it helps reduce repetitive work and resource waste caused by data quality issues; based on accurate and comprehensive simulation datasets, it enables the establishment of more precise and reliable shale fracturing casing deformation prediction models; through the application of machine learning algorithms such as random forests, it can deeply explore the potential patterns and correlations in the data, improving the model's prediction accuracy and generalization ability for shale fracturing casing deformation phenomena; through the data acquisition and processing methods described in this embodiment, it can more accurately predict the casing deformation risk during shale fracturing, providing a scientific basis for fracturing scheme design in actual production; it helps reduce fracturing failures and cost waste caused by casing deformation, improving the development efficiency and economic benefits of shale oil resources.
[0049] By identifying and processing erroneous, missing, and outlier data, the accuracy and integrity of the data are ensured; this helps reduce analytical biases or erroneous decisions caused by data quality issues; the introduction of anomaly evaluation levels and the calculation of collected evaluation values provide a more refined assessment of the dataset quality; the tiered evaluation mechanism can provide more accurate data quality judgments based on different error types and degrees, thereby guiding subsequent data processing or decision-making processes; adjustments based on the total amount of data and collection accuracy make the method highly flexible and adaptable, automatically identifying and processing errors and anomalies in the data, reducing the need for manual intervention, and improving the efficiency and accuracy of data processing; the automated evaluation process also reduces the subjectivity and inconsistency of human judgment; based on the initial judgment, a secondary judgment is made by calculating collected evaluation values, further improving the accuracy of data quality assessment; the multi-level, multi-step evaluation process helps ensure that the final dataset meets strict collection standards; after rigorous quality assessment and screening, the final dataset has higher credibility, providing a solid data foundation for subsequent data analysis and decision-making; through a refined data quality assessment process and automated processing mechanisms, the accuracy and credibility of the dataset are effectively improved, providing strong support for data analysis and decision-making.
[0050] By calculating the data collection pass rate and comparing it with a preset standard pass rate, the most suitable subsequent processing mode can be intelligently selected. This avoids unnecessary duplicate data collection, thereby optimizing the utilization of data processing resources. For datasets that meet the collection standards, the preprocessing mode can be directly activated, allowing for rapid entry into the preprocessing stage and shortening the data processing cycle. For datasets requiring overall analysis, holistic processing is performed, improving the targeting and efficiency of data processing. For datasets with a data collection pass rate lower than the standard, a re-collection mode is selected to ensure that the data used in subsequent analysis has high quality and reliability. This helps reduce analytical bias or erroneous decisions caused by data quality issues. It automates the selection of data processing modes, reducing the subjectivity and inconsistency of human decision-making and improving the accuracy and efficiency of decision-making. The automated processing flow reduces the need for human intervention and lowers the risk of data processing errors or omissions caused by improper human operation. Targeted processing of data according to different data processing modes can improve the accuracy and effectiveness of data processing, providing more reliable data support for subsequent data analysis and decision-making.
[0051] By assigning detailed numbers and scores to qualified and unqualified simulated feature sets, the overall anomaly level of the data can be assessed more comprehensively. Scoring methods based on detailed data reflect the true state of the data better than a single acquisition pass rate, thus improving the accuracy of subsequent processing mode selection. Calculating the overall anomaly score to intelligently select subsequent data processing modes can further optimize the data processing workflow. Datasets with low overall anomaly levels can be directly preprocessed for targeted processing; for datasets with high overall anomaly levels, a re-acquisition mode is selected to ensure data quality. This helps reduce unnecessary processing steps and improve data processing efficiency. For datasets with high overall anomaly scores, the re-acquisition mode ensures that the data used in subsequent analysis has high quality and reliability. This helps reduce analytical biases or erroneous decisions caused by data quality issues, improving the accuracy and reliability of data analysis. Numbering qualified and unqualified simulated feature sets helps to quickly locate and track problems in subsequent data processing or analysis; detailed records also help analyze potential problems in the data acquisition process, providing a basis for optimizing subsequent acquisition processes.
[0052] By calculating primary and secondary correlation coefficients and performing feature filtering based on these correlation coefficients, redundant and irrelevant features can be effectively removed, retaining only those features most influential on the prediction target or research question. This helps reduce data noise and improve data quality and purity. In machine learning or statistical modeling, using high-quality datasets for training and prediction can significantly improve model performance. Removing redundant and correlated features through preprocessing allows for a greater focus on features that truly influence the prediction results, thereby improving model accuracy and generalization ability. Feature filtering reduces the dimensionality of the dataset, lowering computational complexity and storage requirements, making data processing and model training more efficient. In data-driven decision support systems, preprocessed data can provide decision-makers with more accurate and useful information. Removing redundant and correlated features makes the decision-making process more focused and efficient, improving the quality and speed of decision-making. This significantly improves data quality while reducing computational costs and enhancing decision-making efficiency.
[0053] By employing machine learning methods, this approach addresses the challenges of traditional shale oil well casing deformation prediction, such as limited actual data, high modeling difficulty, complex and time-consuming calculation processes, and the significant impact of modeling precision on the calculation results. It achieves the goal of predicting casing variables quickly, conveniently, accurately, and with high universality, providing a new method for predicting casing deformation in shale oil well fracturing. Attached Figure Description
[0054] Figure 1This is a flowchart of a machine learning prediction method for shale fracturing casing deformation based on a mechanistic model, according to an embodiment of the present invention.
[0055] Figure 2 This is a flowchart illustrating the acquisition of the target simulation dataset in the shale fracturing and casing deformation machine learning prediction method based on a mechanism model, as described in this embodiment of the invention.
[0056] Figure 3 This is a flowchart illustrating the conformity determination process in the shale fracturing casing deformation machine learning prediction method based on a mechanism model, as described in this embodiment of the invention.
[0057] Figure 4 This is a flowchart of the preprocessing process in the shale fracturing and casing deformation machine learning prediction method based on the mechanism model in an embodiment of the present invention. Detailed Implementation
[0058] To make the objectives and advantages of the present invention clearer, the present invention will be further described below with reference to embodiments; it should be understood that the specific embodiments described herein are merely for explaining the present invention and are not intended to limit the present invention.
[0059] Preferred embodiments of the present invention will now be described with reference to the accompanying drawings. Those skilled in the art should understand that these embodiments are merely illustrative of the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0060] It should be noted that in the description of this invention, the terms "upper", "lower", "left", "right", "inner", "outer", etc., which indicate directions or positional relationships, are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and is not intended to indicate or imply that the device or element must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation of this invention.
[0061] See Figure 1 As shown, this invention provides a machine learning prediction method for shale fracturing casing deformation based on a mechanistic model, comprising:
[0062] Step S100: Based on the characteristics of shale oil reservoirs, obtain the target simulation dataset of the target shale oil reservoir that meets the acquisition criteria. The target simulation dataset includes several sets of simulation features to be processed.
[0063] Step S200: Preprocess the target simulation dataset to be processed to obtain preprocessing results, and correct each set of simulation features to be processed according to the preprocessing results to obtain the target actual simulation dataset, which includes several sets of actual simulation features.
[0064] Step S300: Perform initial numerical simulation on each actual simulation feature set to obtain initial numerical simulation results, and determine the target shale oil well reservoir slip sleeve variables under different influencing factors based on the initial numerical simulation results.
[0065] Step S400: Obtain influencing factors that cannot be measured in reality to obtain several sets of hidden simulation features; perform inversion based on each set of hidden simulation features to optimize the initial numerical simulation results and obtain the actual numerical simulation results.
[0066] Step S500: Based on the selected initial prediction standard, the actual numerical simulation results are sequentially subjected to prediction training and prediction testing to optimize the initial prediction standard and obtain the actual prediction standard.
[0067] Specifically, this invention, through data acquisition based on shale oil reservoir characteristics, ensures the high quality and relevance of the target simulation dataset, laying a solid foundation for subsequent simulation and prediction; it improves the accuracy and consistency of the data, further enhancing the precision of subsequent numerical simulations; it considers shale oil well reservoir slip and casing variables under different influencing factors, enabling a comprehensive simulation of various situations that may occur during actual fracturing, helping to understand the complexity and diversity of casing deformation phenomena; by optimizing the initial numerical simulation results, it solves the problem of influencing factors that cannot be directly measured in actual engineering, making the simulation results closer to reality; through repeated optimization of the actual numerical simulation results, a reliable prediction standard is formed; it not only improves the accuracy and efficiency of prediction but also realizes the intelligence and automation of the prediction process; by accurately predicting the casing deformation risks and casing variables that may occur during shale oil well fracturing, it helps to formulate prevention and control measures in advance, reducing fracturing failures and well abandonment caused by casing deformation, thereby significantly reducing development costs and risks; through efficient numerical simulation and intelligent prediction optimization, it provides strong support for the efficient development of shale oil, helping to improve the overall efficiency and competitiveness of shale oil extraction.
[0068] See Figure 2 The diagram shows a flowchart for obtaining the target simulation dataset to be processed; wherein, the process of obtaining the target simulation dataset of the target shale oil reservoir that meets the acquisition criteria includes:
[0069] Step S110: Based on the characteristics of shale oil reservoirs, obtain the initial simulation feature set of the target shale oil reservoir;
[0070] Step S120: For any initial simulated feature set, obtain the actual acquisition status of the initial simulated feature set;
[0071] Step S130: Mark the data according to the actual acquisition situation;
[0072] Step S140: Determine a single acquisition situation based on the data labeling results, and determine whether to enable the overall analysis mode, re-acquisition mode, or preprocessing mode based on the single acquisition situation.
[0073] Specifically, in this embodiment, based on the characteristics of shale oil reservoirs, the following influencing factors requiring numerical simulation were identified; these factors are categorized into eight types: ① fracturing operation parameters; ② fracturing fluid and proppant parameters; ③ formation stress environment parameters; ④ formation physical property parameters; ⑤ weak structural surface property parameters; ⑥ casing parameters; ⑦ cement sheath parameters; ⑧ wellbore structure parameters. Therefore, the initial simulation feature sets of the target shale oil reservoir are obtained as follows: ① initial simulated fracturing operation set; ② initial simulated fracturing fluid and proppant set; ③ initial simulated formation stress environment set; ④ initial simulated formation physical property set; ⑤ initial simulated weak structural surface property set; ⑥ initial simulated casing set; ⑦ initial simulated cement sheath set; ⑧ initial simulated wellbore structure set.
[0074] Specifically, this invention, by meticulously classifying influencing factors based on shale oil reservoir characteristics and acquiring each initial simulation feature set, ensures the comprehensiveness and accuracy of the simulation data. The refined data collection method helps to more realistically reflect the complexity and diversity of shale oil reservoirs, providing a reliable foundation for subsequent analysis. By marking the actual acquisition of the initial simulation feature sets and flexibly selecting between overall analysis mode, re-acquisition mode, or preprocessing mode based on the marking results, the data processing workflow becomes more flexible and efficient. The adaptive data processing mechanism effectively addresses data quality issues under different conditions, improving the efficiency and accuracy of data processing. By pre-determining the influencing factors requiring numerical simulation and selectively acquiring each initial simulation feature, the invention ensures the comprehensiveness and accuracy of the simulation data. Data aggregation can avoid blindly collecting and processing irrelevant data, thereby reducing the cost of data collection and processing; it helps reduce duplication of work and waste of resources caused by data quality issues; based on accurate and comprehensive simulation datasets, more accurate and reliable shale fracturing casing deformation prediction models can be established; through the application of machine learning algorithms such as random forests, the potential patterns and correlations in the data can be deeply explored, improving the model's prediction accuracy and generalization ability for shale fracturing casing deformation phenomena; through the data acquisition and processing methods described in this embodiment, the casing deformation risk in the shale fracturing process can be predicted more accurately, providing a scientific basis for fracturing scheme design in actual production; it helps reduce fracturing failure and cost waste caused by casing deformation, and improves the development efficiency and economic benefits of shale oil resources.
[0075] Specifically, the process of determining a single collection scenario based on data labeling results in this embodiment includes:
[0076] Step S1411: Identify erroneous data, missing data, and abnormal data in the actual acquisition situation, and clean and process the initial simulated feature set accordingly based on the identification results to obtain an intermediate simulated feature set;
[0077] Step S1412: Type-label the erroneous data, the missing data, and the abnormal data, and obtain the number of times each type is labeled;
[0078] Step S1413: Determine the anomaly evaluation level of the initial simulated feature set based on the type labeling result, and make an initial judgment on whether the initial simulated feature set meets the acquisition criteria based on the anomaly evaluation level.
[0079] Specifically, in this embodiment, the process of determining the anomaly evaluation level of the initial simulated feature set based on the identification result, and making an initial judgment on whether the initial simulated feature set meets the collection criteria based on the anomaly evaluation level, is as follows:
[0080] Analyze the number and type of items that meet a single judgment condition in the identification results;
[0081] Determine whether the project type that meets the single determination condition includes an anomaly type. If it does, determine whether the number of projects is 1. If it does, determine that the anomaly evaluation level is Level 1, and determine the initial judgment result as: the initial simulation feature set meets the collection standard. If it does not, determine that the anomaly evaluation level is Level 2, and determine the initial judgment result as: enable secondary judgment.
[0082] If the project type that meets the single judgment condition does not include the abnormal type, then determine whether the number of projects is equal to 1. If yes, then determine the abnormality evaluation level as level two, and determine the initial judgment result as: start the second judgment. If no, then determine the abnormality evaluation level as level three, and determine the initial judgment result as: the initial simulation feature set does not meet the collection standard.
[0083] Specifically, for any type of labeling result, if the absolute value of the labeling difference corresponding to any type of labeling result is greater than the preset labeling evaluation value and the corresponding number of labeling of the type is greater than the preset standard number of labeling, then any type of labeling result is determined to meet the single determination condition; the absolute value of the labeling difference is the absolute value of the difference between the number of labeling of the type and the number of standard labeling.
[0084] Specifically, in this embodiment, for any type of marking result, the absolute value S1 of the marking difference is calculated based on the type marking frequency X and the standard marking frequency X0 of the type marking result;
[0085] S1 = |X - X0|;
[0086] If S1≤S10, then the result of this type of labeling does not meet the single determination condition;
[0087] If S1 > S10 and X < X0, then the result of this type of labeling does not meet the single determination condition.
[0088] If S1 > S10 and X > X0, then the result of this type of labeling is determined to meet the single determination condition.
[0089] Wherein, S10 is the preset mark evaluation value;
[0090] In the specific implementation process, the standard marking frequency is set to 8; the standard evaluation value is set to 2; the specific settings are also affected by the total amount of data in the initial simulation feature set and the acquisition accuracy.
[0091] Specifically, the process of determining a single collection case based on the data labeling results described in this embodiment further includes:
[0092] Based on the initial judgment result and the number of times each type of marker is used, the collection evaluation value is calculated, and a second judgment is made based on the collection evaluation value to determine whether the intermediate simulated feature set meets the collection standard; if it meets the standard, the intermediate simulated feature set is determined to be the simulated feature set to be processed; if it does not meet the standard, the intermediate simulated feature set is determined to be an abnormal simulated feature set.
[0093] The type labeling results include: the error type, the missing type, and the exception type.
[0094] See Figure 3 As shown, this is a flowchart for the conformity determination process. The process of calculating the collection evaluation value based on the initial judgment result and the number of times each type of marker is used, and then determining whether the intermediate simulated feature set conforms to the collection criteria based on the collection evaluation value for a secondary judgment, is as follows:
[0095] When the anomaly evaluation level is level two, the intermediate simulated feature set is determined to meet the collection standard based on the absolute value of the second difference and the preset second evaluation value.
[0096] Determine whether the absolute value of the second difference is less than or equal to the second evaluation value. If it is less than or equal to, then the intermediate simulated feature set is determined to meet the acquisition standard. If it is greater than, then the acquisition evaluation value is determined to be less than the preset acquisition standard value. If it is less than, then the intermediate simulated feature set is determined to meet the acquisition standard. If it is greater than, then the intermediate simulated feature set is determined to not meet the acquisition standard.
[0097] Wherein, the absolute value of the second difference is the absolute value of the difference between the collected evaluation value and the collected standard value.
[0098] Specifically, in this embodiment, the absolute value of the second difference S2 is calculated based on the collected evaluation value H1 and the preset collected standard value H0;
[0099] S2 = |H1 - H0|; where S20 is the preset second evaluation value;
[0100] When the initial judgment result indicates that a second judgment is enabled, the collected evaluation value H1 is calculated based on the number of first type markers A1 for the error type, the number of second type markers A2 for the missing type, and the number of third type markers A3 for the anomaly type.
[0101] H1 = A1 × b1 + A2 × b2 + A3 × b3;
[0102] Wherein, b1 is the first calculation compensation parameter for the first type of marking count A1 on the collected evaluation value H1; b2 is the second calculation compensation parameter for the second type of marking count A2 on the collected evaluation value H1; and b3 is the third calculation compensation parameter for the third type of marking count A3 on the collected evaluation value H1.
[0103] In the specific implementation process, the standard value H0 is set to 9.6; the second evaluation value is set to 1.7; the first calculation compensation parameter b1 is set to 0.6; the second calculation compensation parameter b2 is set to 0.6; and the third calculation compensation parameter b3 is set to 0.8. The specific settings are also affected by the total number of data in the initial simulation feature set and the acquisition accuracy.
[0104] Specifically, this invention ensures data accuracy and integrity by identifying and processing erroneous, missing, and abnormal data; it helps reduce analytical biases or erroneous decisions caused by data quality issues; by introducing anomaly evaluation levels and calculating collection evaluation values, it provides a more refined assessment of the quality of the dataset; the tiered evaluation mechanism can provide more accurate data quality judgments based on different error types and degrees, thereby guiding subsequent data processing or decision-making processes; and it can be adjusted according to the specific total amount of data and collection accuracy, making the method highly flexible and adaptable, automatically identifying and processing errors and anomalies in the data, and reducing the need for manual intervention. This approach improves the efficiency and accuracy of data processing; the automated evaluation process reduces the subjectivity and inconsistency of human judgment; based on the initial judgment, a secondary judgment is made by calculating the collected evaluation values, further improving the accuracy of data quality assessment; the multi-level, multi-step evaluation process helps ensure that the final dataset meets strict collection standards; after rigorous quality assessment and screening, the final dataset has higher credibility, providing a solid data foundation for subsequent data analysis and decision-making; through a refined data quality assessment process and automated processing mechanism, the accuracy and credibility of the dataset are effectively improved, providing strong support for data analysis and decision-making.
[0105] Specifically, in this embodiment, the process of determining whether to activate the overall analysis mode, the re-acquisition mode, or the preprocessing mode based on the single acquisition situation includes:
[0106] Step S141: Determine the number of actual qualified sets in the target initial simulation dataset that meet the acquisition criteria based on the secondary judgment results;
[0107] Step S142: Based on the actual number of qualified sets combined with the total number of sets, determine whether to enable the overall analysis mode, or enable the re-acquisition mode, or enable the preprocessing mode.
[0108] The initial simulation feature sets constitute the target initial simulation dataset.
[0109] Specifically, in this embodiment, the collection pass rate is calculated based on the actual number of qualified sets combined with the total number of sets;
[0110] The collection pass rate = actual number of pass sets / total number of sets;
[0111] The overall analysis mode is activated, or the re-acquisition mode is activated, or the preprocessing mode is activated, based on the absolute value of the first difference and the preset first evaluation value.
[0112] If the absolute value of the first difference is less than or equal to the first evaluation value, then the overall analysis mode is activated; if the absolute value of the first difference is greater than the first evaluation value and the collection pass rate is greater than the preset standard pass rate, then the preprocessing mode is activated; if the absolute value of the first difference is greater than the first evaluation value and the collection pass rate is less than the standard pass rate, then the re-collection mode is activated.
[0113] Wherein, the absolute value of the first difference is the absolute value of the difference between the collection pass rate and the standard pass rate; the absolute value of the first difference = |collection pass rate - standard pass rate|;
[0114] In this embodiment, the standard pass rate is set to 80%-85%; the first evaluation value is set to 3%-5%; the specific setting is also affected by the total amount of data in the initial simulated feature set and the accuracy of data collection.
[0115] Specifically, this invention, by calculating the data collection pass rate and comparing it with a preset standard pass rate, can intelligently select the most suitable subsequent processing mode; it avoids unnecessary repeated data collection, thereby optimizing the utilization of data processing resources; for datasets that meet the collection standards, the preprocessing mode is directly activated, allowing for rapid entry into the preprocessing stage and shortening the data processing cycle; for datasets requiring overall analysis, holistic processing is performed, improving the targeting and efficiency of data processing; for datasets with a data collection pass rate lower than the standard, a re-collection mode is selected to ensure that the data used in subsequent analysis has high quality and reliability; it helps reduce analytical bias or erroneous decisions caused by data quality issues; it achieves automated selection of data processing modes, reducing the subjectivity and inconsistency of human decision-making and improving the accuracy and efficiency of decision-making; the automated processing flow reduces the need for human intervention, lowering the risk of data processing errors or omissions caused by improper human operation; and targeted processing of data according to different data processing modes can improve the accuracy and effectiveness of data processing, providing more reliable data support for subsequent data analysis and decision-making.
[0116] Specifically, in this embodiment, based on enabling the overall analysis mode, the number of qualified sets of each simulated feature set to be processed that meet the acquisition criteria in the target initial simulation dataset and the number of unqualified sets of each abnormal simulation feature set that do not meet the acquisition criteria in the target initial simulation dataset are integrated.
[0117] Based on the integration results, an overall anomaly score is determined, and based on the overall anomaly score, it is determined whether to enable the re-collection mode or the preprocessing mode.
[0118] Specifically, the initial target simulation dataset is integrated to obtain a set of simulation features to be processed that meet the acquisition criteria, and a pass / fail score is calculated. Conversely, a set of abnormal simulation features that do not meet the acquisition criteria is obtained, and a fail / fail score is calculated.
[0119] 1) Number each qualified set of simulated features to be processed, denoted as the first qualified set W11, the second qualified set W12, ..., the m-th qualified set W1m.
[0120] 2) Number the sets of non-compliant simulated features, denoted as the first set of anomalies W21, the second set of anomalies W22, ..., the nth set of anomalies W2n.
[0121] Based on the integration results, an overall anomaly score K is determined, and a setting is made.
[0122] K= 1-
[0123] Where m is the number of qualified items; n is the number of unqualified items; E1 is the qualified compensation parameter of the number of qualified items m to the overall anomaly score K; E2 is the unqualified compensation parameter of the number of unqualified items n to the overall anomaly score K.
[0124] If K ≥ K0, then the preprocessing mode is enabled.
[0125] If K < K0, then the re-acquisition mode is activated.
[0126] Wherein, K0 is the standard evaluation value for calculation.
[0127] In this embodiment, the qualified compensation parameter E1 is set to 0.7; the unqualified compensation parameter E2 is set to 0.8; the specific setting is also affected by the total amount of data in the initial simulation feature set and the acquisition accuracy.
[0128] Specifically, this invention, by assigning detailed numbers and scores to qualified and unqualified simulated feature sets, can more comprehensively assess the overall anomaly level of the data. The scoring method based on detailed data reflects the true state of the data better than a single acquisition pass rate, thereby improving the accuracy of subsequent processing mode selection. Intelligent selection of subsequent data processing modes by calculating the overall anomaly score can further optimize the data processing workflow. Datasets with low overall anomaly levels are directly preprocessed for targeted processing; for datasets with high overall anomaly levels, a re-acquisition mode is selected to ensure data quality. This helps reduce unnecessary processing steps and improve data processing efficiency. For datasets with high overall anomaly scores, the re-acquisition mode ensures that the data used in subsequent analysis has high quality and reliability. This helps reduce analytical biases or erroneous decisions caused by data quality issues, improving the accuracy and reliability of data analysis. Numbering qualified and unqualified simulated feature sets helps to quickly locate and track problems in subsequent data processing or analysis. Detailed records also help analyze potential problems in the data acquisition process, providing a basis for optimizing subsequent acquisition processes.
[0129] See Figure 4 As shown, this is a preprocessing flowchart; specifically, the process of preprocessing the target simulated dataset to obtain preprocessing results in this embodiment includes:
[0130] Step S210: Analyze the first actual correlation degree between each set of simulated features to be processed, and perform feature filtering on the target simulated dataset to be processed based on the first actual correlation degree to obtain the target simulated dataset;
[0131] The target dataset to be simulated includes several sets of features to be simulated;
[0132] Step S220: For any set of features to be simulated, analyze the secondary actual correlation degree between each feature parameter in the set of features to be simulated, and perform feature filtering on the set of features to be simulated based on the secondary actual correlation degree to obtain the actual simulated feature set.
[0133] Specifically, in this embodiment, the actual correlation degree refers to the degree of correlation or mutual influence between different sets of simulated features to be processed. By analyzing the actual correlation degree, a set of features that can fully characterize the predicted target is selected from the initial simulated dataset of the target, forming the target simulated dataset.
[0134] In this embodiment, the Pearson correlation coefficient method is used to calculate the actual correlation degree between each set of simulated features to be processed;
[0135] The Pearson correlation coefficient is used to measure the degree of linear correlation between two sets of simulated features to be processed, and its value ranges from -1 to 1.
[0136] If the actual correlation degree between any two sets of simulated features to be processed is close to 1, it indicates that the linear relationship between the two sets of simulated features to be processed is stronger; the closer it is to 0, the weaker the linear relationship between the two sets of simulated features to be processed is.
[0137] The actual correlation degree is calculated for each pair of simulated feature sets to be processed, so as to filter out each simulated feature set with weak linear relationship, so as to obtain the target simulated dataset;
[0138] For any set of features to be simulated, the Pearson correlation coefficient method is used to calculate the secondary actual correlation degree between each feature parameter in the set of features to be simulated.
[0139] The actual correlation degree of each feature parameter is calculated twice in pairs to filter out the feature parameters with weak linear relationships, so as to obtain the actual simulated feature set.
[0140] In this embodiment, the target actual simulation dataset includes several actual simulation feature sets, and each actual simulation feature set includes several feature parameters;
[0141] ① Actual simulated fracturing operation set: 1.1 Injection rate, m 3 / min; 1.2 total injection volume, m 3 1.3 Construction pressure, MPa; 1.4 Construction temperature, °C; 1.5 Cluster spacing, m; 1.6 Number of clusters.
[0142] ② Actual simulation of fracturing fluid and proppant combination: 2.1 Fracturing fluid viscosity, mPa·s; 2.2 Propane concentration, kg / m³ 3 2.3 Proppant particle size, m; 2.4 Proppant dosage, m 3 .
[0143] ③ Actual simulated formation stress environment set: 3.1 Maximum horizontal principal stress, MPa; 3.2 Minimum horizontal principal stress, MPa; 3.3 Vertical stress, MPa; 3.4 Formation pore pressure, MPa.
[0144] ④ Set of actual simulated formation physical properties: 4.1 Elastic modulus of reservoir rock, GPa; 4.2 Poisson's ratio of reservoir rock; 4.3 Permeability of reservoir rock, mD; 4.4 Porosity of reservoir rock; 4.5 Density of reservoir rock, kg / m³ 3 ; 4.6 Energy release rate of reservoir rock fracture; 4.7 Formation temperature, °C.
[0145] ⑤ Actual simulated weak structural surface characteristics set: 5.1 Weak structural surface dip angle, °; 5.2 Weak structural surface length, m; 5.3 Curvature at the bend of the weak structural surface, 1 / m; 5.4 Position of the weak structural surface relative to the horizontal well, m; 5.5 Distance between the weak structural surface and the fracturing point, m; 5.6 Friction coefficient of the weak structural surface; 5.7 Fracture energy release rate of the weak structural surface.
[0146] ⑥ Actual simulated casing assembly: 6.1 Casing diameter, m; 6.2 Casing wall thickness, m; 6.3 Casing elastic modulus, GPa; 6.4 Casing Poisson's ratio.
[0147] ⑦ Actual simulated cement ring assembly: 7.1 Cement ring thickness, m; 7.2 Cement ring elastic modulus, GPa; 7.3 Cement ring Poisson's ratio.
[0148] ⑧ Actual simulated wellbore structure set: 8.1 Well diameter, m; 8.2 Well inclination angle, °; 8.3 Azimuth angle, °.
[0149] Specifically, this invention, by calculating primary and secondary actual correlation degrees and performing feature filtering based on these correlation degrees, can effectively remove redundant and irrelevant features, retaining only those features most influential on the prediction target or research question. This helps reduce data noise and improve data quality and purity. In machine learning or statistical modeling, using high-quality datasets for training and prediction typically significantly improves model performance. Removing redundant and related features through preprocessing allows for a greater focus on features that truly influence the prediction results, thereby improving model accuracy and generalization ability. Feature filtering reduces the dimensionality of the dataset, lowering computational complexity and storage requirements, making data processing and model training more efficient. In data-driven decision support systems, preprocessed data can provide decision-makers with more accurate and useful information. Removing redundant and related features makes the decision-making process more focused and efficient, improving both the quality and speed of decision-making. This significantly improves data quality while reducing computational costs and enhancing decision-making efficiency.
[0150] Specifically, in this embodiment, the process of determining the target shale oil well reservoir slip casing variables under different influencing factors based on the initial numerical simulation results includes:
[0151] Step S310: Determine the actual geometric parameters and the distribution of the induced stress field during the hydraulic fracturing process based on each actual simulation feature set;
[0152] Step S320: Based on the actual geometric parameters and the distribution of the induced stress field, determine the numerical range of the influencing factors that cause slippage of the weak structural surface and the amount of slippage of the weak structural surface under different influencing factors.
[0153] Step S330: Determine the target shale oil well reservoir slip variable based on the numerical range of influencing factors of slippage of weak structural surfaces and the slippage amount of weak structural surfaces under different influencing factors.
[0154] The specific process includes:
[0155] (1) Establishing a numerical model for fracture propagation in shale reservoirs: calculating the geometry, size, and orientation of the fracturing body under different influencing factors; wherein, the actual geometric parameters include: geometry, size, and orientation; simultaneously calculating the distribution of the induced stress field during hydraulic fracturing to provide numerical basis for calculating the slip of weak structural surfaces during fracturing; the process of establishing the numerical model for fracture propagation in shale reservoirs includes:
[0156] 1) Establish a three-dimensional geometric model, with the model size controlled within the well depth +2000m.
[0157] 2) Based on the coordinates of the well, construct the wellbore in the geometric model, and design the fracturing point location according to the cluster spacing and number of clusters.
[0158] 3) Construct a grid and assign permeability and porosity properties to the grid cells.
[0159] 4) Add material property parameters. These include: ① fracturing fluid viscosity, mPa·s; ② proppant particle size, m; ③ proppant concentration, kg / m³. 3 ; ④ Elastic modulus of reservoir rock, GPa; ⑤ Poisson's ratio of reservoir rock; ⑥ Permeability of reservoir rock, mD; ⑦ Porosity of reservoir rock; ⑧ Density of reservoir rock, kg / m³ 3 ; ⑨ Energy release rate of reservoir rock fracture.
[0160] 5) Add boundary conditions and initial values. Boundary conditions include: ① Injection rate, m 3 / min; ② Construction pressure, MPa; ③ Construction temperature, °C; ④ Proppant dosage, m 3 The initial values include: ① Formation temperature, °C; ② Maximum horizontal principal stress, MPa; ③ Minimum horizontal principal stress, MPa; ④ Vertical stress, MPa; ⑤ Formation pore pressure, MPa.
[0161] 6) Set the solution time step, based on the ratio of total injection volume to injection rate.
[0162] 7) Solve the model.
[0163] (2) Establish a numerical model for the slip of weak structural surfaces during fracturing: calculate the slip of weak structural surfaces during fracturing of shale reservoirs under different influencing factors; use the Mohr-Coulomb criterion to identify the slip of weak structural surfaces, determine the slip of weak structural surfaces under the induced stress field of the shale reservoir calculated in step (1), and obtain the numerical range of influencing factors that cause slip of weak structural surfaces and the slip of weak structural surfaces under different influencing factors; the process of establishing a numerical model for the slip of weak structural surfaces during fracturing includes:
[0164] 1) Based on step (1), a parametric surface characterizing the weak structural surface in the shale reservoir is established according to the geometric coordinates of the weak structural surface. The design parameters include: ① dip angle of the weak structural surface, °; ② length of the weak structural surface, m; ③ curvature at the bend of the weak structural surface, 1 / m; ④ position of the weak structural surface relative to the horizontal well, m; ⑤ distance between the weak structural surface and the fracturing point, m.
[0165] 2) Assign material properties to the weak structural surface, including: ① friction coefficient of the weak structural surface; ② fracture energy release rate of the weak structural surface.
[0166] 3) Set the model boundary conditions and initial values, and ensure that the design parameters are consistent with the actual simulation dataset of the target.
[0167] 4) The model is solved in steady state, and the slip of the weak structure surface under the induced stress field obtained in step (1) is calculated.
[0168] (3) Establish a numerical model of shale oil well reservoir slip and casing deformation, and calculate the casing deformation under influencing factors. Under the condition of slip of the weak structural plane of the shale reservoir during the fracturing process calculated in step (2), calculate the shear stress on the casing. When the shear stress exceeds the shear strength of the casing, the casing undergoes shear deformation, and the casing deformation caused when the shear stress exceeds the shear strength is obtained. The process of establishing a numerical model of shale oil well reservoir slip and casing deformation includes:
[0169] 1) Establish a cuboid with dimensions of 5m × 1m × 1m as the geometric model of the casing-cement sheath-formation assembly. Tubular geometric bodies are added to the model based on the geometric parameters of the casing and cement sheath. The parameters involved include: ① casing diameter, m; ② casing wall thickness, m; ③ cement sheath thickness, m.
[0170] 2) Construct a mesh and assign solid mechanical properties to the mesh cells.
[0171] 3) Add material property parameters, including: ① Casing elastic modulus, GPa; ② Casing Poisson's ratio; ③ Cement sheath elastic modulus, GPa; ④ Cement sheath Poisson's ratio; ⑤ Reservoir rock elastic modulus, GPa; ⑥ Reservoir rock Poisson's ratio; ⑦ Reservoir rock density, kg / m³ 3 .
[0172] 4) The slip of the weak structural surface calculated in step (2) is used as the displacement boundary condition of the numerical model of the slip sleeve of the shale oil well reservoir.
[0173] 5) Solve the model to obtain the shale oil well reservoir slippage variables under different influencing factors.
[0174] Specifically, in this embodiment, the process of optimizing the initial numerical simulation results by performing inversion based on each set of latent simulation features to obtain the actual numerical simulation results includes:
[0175] Step S410: Based on historical fitting, invert each set of latent simulation features to obtain the actual numerical simulation results;
[0176] Step S420: Compare the actual numerical simulation results with the obtained actual numerical calculation results;
[0177] Step S430: Determine whether to optimize the initial prediction standard based on the comparison results, or to perform historical fitting again.
[0178] The specific process includes:
[0179] (4) Invert influencing factors that cannot be measured in practice through historical fitting to optimize the numerical model. The parameters that need to be inverted include: ① reservoir rock fracture energy release rate; ② friction coefficient of weak structural surfaces; ③ fracture energy release rate of weak structural surfaces. If there are other parameters without measurement results in the influencing factors in the actual simulation dataset of the target in actual engineering, they also need to be inverted through historical fitting to optimize the numerical model.
[0180] (5) The random forest algorithm is used as the machine learning method. 80% of the different influencing factors obtained in step (3) are used as the training set to establish a shale fracturing and deformation prediction model.
[0181] (6) Use 20% of the set variables obtained in step (3) as the test set to optimize the machine learning algorithm.
[0182] (7) After optimization in step (6), a machine learning prediction method for shale fracturing jacket deformation based on mechanism model is finally formed.
[0183] This invention addresses the problems of limited actual data, high modeling difficulty, complex and time-consuming calculation processes, and significant impact of modeling precision on the prediction of casing deformation in traditional shale oil wells by employing machine learning methods. It achieves the goal of predicting casing deformation quickly, conveniently, accurately, and with high universality, providing a new method for predicting casing deformation in fracturing shale oil wells. Example 1:
[0184] The tectonic location of this region is at the junction of the Sichuan Basin and the Yunnan-Guizhou Plateau, between the low-steep tectonic zone of the southern Sichuan ancient depression and the Loushan fold belt. It is influenced by the westward extension of the eastern Sichuan fold-thrust belt to the north and controlled by the evolution of the Loushan fold belt to the south, forming a tectonic complex integrating the characteristics of both. In this block, the casing deformation of platform H is relatively severe. The strata along the horizontal extension of this platform dip down by approximately 5 degrees; the actual drilled horizontal section of the lower half has a burial depth of 2430~2600m, a formation pressure coefficient of approximately 1.8, and a dip angle of approximately 3.0~5.0°. Within platform H, the casing shear deformation in the horizontal sections of wells X1, X2, and X3 is particularly severe. A machine learning prediction method based on a mechanistic model for shale fracturing casing deformation is used to predict casing variables. The steps are as follows:
[0185] (1) Establish a numerical model for fracture propagation in shale reservoirs. Solve for the geometry, size and orientation of the modified bodies generated by fracturing the three horizontal wells X1, X2 and X3 in platform H, and calculate the distribution of the induced stress field during hydraulic fracturing.
[0186] A geometric model with dimensions of 5000m × 5000m × 3000m (length × width × height) was established. After the geometric modeling of wells X1, X2, and X3 was completed, material properties and boundary conditions required for numerical solution were assigned based on actual fracturing data. The parameter ranges used in the solution are shown in Table 1.
[0187] Table 1. Parameter Selection Table for Numerical Model of Fracture Propagation in Shale Reservoirs
[0188]
[0189] Based on the parameters in Table 1, the geometry, size, and orientation of the hydraulic fracture modification body during the fracturing process, as well as the distribution of the induced stress field in the shale reservoir after fracturing, are determined.
[0190] (2) Establish a numerical model for the slip of weak structural planes during the fracturing process. Based on the numerical model of fracture propagation in shale reservoirs, add the geometric distribution of weak structural planes and the distribution of their characteristic parameters, where the characteristic parameters of weak structural planes are shown in Table 2.
[0191] Table 2 Selection of Characteristic Parameters for Weak Structural Surfaces
[0192]
[0193] Based on the model including weak structural surfaces, and combined with the long-range distribution of the induced stress field, the stress distribution at the weak structural surfaces is solved. Slip is determined according to the Mohr-Coulomb criterion, and the numerical range of influencing factors causing slip at the weak structural surfaces and the slip amount under different influencing factors are obtained. The calculation results of the slip amount are shown in Table 3.
[0194] Table 3. Slip of weak structural surfaces under different influencing factors
[0195]
[0196] (3) Establish a numerical model of shale oil well reservoir slippage and casing deformation. First, a cuboid with dimensions of 5m × 1m × 1m is established as the geometric model of the casing-cement sheath-formation assembly. Tubular geometries are added to the model based on the geometric parameters of the casing and cement sheath. The simulation parameters for the casing and cement sheath are shown in Table 4.
[0197] Table 4 Simulated operating parameters for casing and cement sheath
[0198]
[0199] Table 5. Variables of casing under different casing and cement sheath parameters
[0200]
[0201] (4) Through historical data fitting, the following parameters were inverted from the casing deformation data of three wells (X1, X2, and X3) on platform H: ① reservoir rock fracture energy release rate; ② weak structural surface friction coefficient; ③ weak structural surface fracture energy release rate. After inversion, the reservoir rock fracture energy release rate was 0.02, the weak structural surface friction coefficient was 0.42, and the weak structural surface fracture energy release rate was 0.003. After correcting the model through historical data fitting, the accuracy of the model was verified. The accuracy verification results are shown in Table 6.
[0202] Table 6. Numerical Model Accuracy Validation
[0203]
[0204] (5) The random forest algorithm was used as the machine learning method. 80% of the variables under different influencing factors in Table (5) were used as the training set to establish a shale fracturing and deformation prediction model.
[0205] (6) 20% of the nested variables obtained in Table (5) are used as the test set to optimize the machine learning algorithm, and finally a machine learning prediction method for shale fracturing nested variables based on the mechanism model is formed.
[0206] The calculation compensation parameters and calculation adjustment parameters described in this invention serve two purposes: first, to balance the left and right dimensions of the formula; and second, to adjust the numerical results. In this embodiment, no specific values are assigned. Furthermore, in this embodiment, each calculation formula is used to intuitively reflect the adjustment relationship between the values, such as positive correlation or negative correlation. Unless otherwise specified, the values of parameters that are not specifically limited are all taken as positive.
[0207] The technical solution of the present invention has been described above with reference to the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will all fall within the scope of protection of the present invention.
[0208] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A machine learning prediction method for shale fracturing casing deformation based on a mechanistic model, comprising: Based on the characteristics of shale oil reservoirs, a target simulation dataset of shale oil reservoirs that meets the acquisition criteria is obtained. The target simulation dataset includes several sets of simulation features to be processed. The target simulation dataset to be processed is preprocessed to obtain preprocessing results. Based on the preprocessing results, each set of simulation features to be processed is corrected to obtain the target actual simulation dataset, which includes several sets of actual simulation features. Initial numerical simulations are performed on each set of actual simulated features to obtain initial numerical simulation results. Based on the initial numerical simulation results, the target shale oil well reservoir slippage variables under different influencing factors are determined. The influencing factors that cannot be measured in reality are obtained to obtain several sets of implicit simulation features. Based on each set of implicit simulation features, inversion is performed to optimize the initial numerical simulation results in order to obtain the actual numerical simulation results. Based on the preset initial prediction standard, the actual numerical simulation results are sequentially subjected to prediction training and prediction testing to optimize the initial prediction standard in order to obtain the actual prediction standard. The process of preprocessing the target simulated dataset to obtain preprocessing results includes: Analyze the first actual correlation degree between each set of simulated features to be processed, and perform feature filtering on the target simulated dataset to be processed based on the first actual correlation degree to obtain the target simulated dataset; The target dataset to be simulated includes several sets of features to be simulated; For any set of features to be simulated, the secondary actual correlation degree between each feature parameter in the set of features to be simulated is analyzed, and the feature set to be simulated is filtered according to the secondary actual correlation degree to obtain the actual simulated feature set. The process of determining the target shale oil well reservoir slip casing variables under different influencing factors based on the initial numerical simulation results includes: The actual geometric parameters and the distribution of induced stress field during hydraulic fracturing are determined based on the actual simulation feature sets. Based on the actual geometric parameters and the distribution of the induced stress field, the numerical range of the influencing factors that cause slippage of the weak structural surface and the amount of slippage of the weak structural surface under different influencing factors are determined. The target shale oil well reservoir slip set variables are determined based on the numerical range of influencing factors of slippage of weak structural surfaces and the slippage amount of weak structural surfaces under different influencing factors.
2. The shale fracturing casing deformation machine learning prediction method based on a mechanism model according to claim 1, characterized in that, The process of obtaining the target simulation dataset of the target shale oil reservoir that meets the acquisition criteria includes: Based on the characteristics of shale oil reservoirs, the initial simulated feature sets of the target shale oil reservoir are obtained; For any initial simulated feature set, obtain the actual acquisition status of the initial simulated feature set; Data is labeled based on the actual acquisition situation; Based on the data labeling results, a single acquisition scenario is determined, and based on the single acquisition scenario, the overall analysis mode, re-acquisition mode, or preprocessing mode is activated.
3. The shale fracturing casing deformation machine learning prediction method based on a mechanism model according to claim 2, characterized in that, The process of determining a single acquisition scenario based on data labeling results includes: Identify erroneous, missing, and abnormal data in the actual acquisition situation, and clean and process the initial simulated feature set accordingly based on the identification results to obtain an intermediate simulated feature set; The erroneous data, the missing data, and the abnormal data are tagged with types, and the number of times each type is tagged is obtained; The anomaly evaluation level of the initial simulated feature set is determined based on the type labeling results, and an initial judgment is made on whether the initial simulated feature set meets the acquisition criteria based on the anomaly evaluation level.
4. The shale fracturing casing deformation machine learning prediction method based on a mechanism model according to claim 3, characterized in that, The process of determining a single acquisition scenario based on data labeling results also includes: Based on the initial judgment result and the number of times each type of label is used, the collection evaluation value is calculated, and a second judgment is made based on the collection evaluation value to determine whether the intermediate simulated feature set meets the collection standard; if it meets the standard, the intermediate simulated feature set is determined to be the simulated feature set to be processed. The type labeling results include: error type, missing type, and exception type.
5. The shale fracturing casing deformation machine learning prediction method based on a mechanism model according to claim 4, characterized in that, The process of determining whether to activate the overall analysis mode, the re-acquisition mode, or the preprocessing mode based on the single acquisition situation includes: Based on the results of the second judgment, determine the number of actual qualified sets in the target initial simulation dataset that meet the collection criteria; The overall analysis mode, re-acquisition mode, or preprocessing mode is activated based on the actual number of qualified sets combined with the preset number of qualified sets. The initial simulation feature sets constitute the target initial simulation dataset.
6. The shale fracturing casing deformation machine learning prediction method based on a mechanistic model according to claim 5, characterized in that, The collection pass rate is calculated based on the actual number of qualified sets combined with the total number of sets. The overall analysis mode is activated, or the re-acquisition mode is activated, or the preprocessing mode is activated, based on the absolute value of the first difference and the preset first evaluation value. If the absolute value of the first difference is less than or equal to the first evaluation value, then the overall analysis mode is activated. If the absolute value of the first difference is greater than the first evaluation value and the collection pass rate is greater than the preset standard pass rate, then it is determined that the preprocessing mode is enabled. If the absolute value of the first difference is greater than the first evaluation value and the collection pass rate is less than the standard pass rate, then it is determined that the re-collection mode is activated. Wherein, the absolute value of the first difference is the absolute value of the difference between the sampling pass rate and the standard pass rate.
7. The shale fracturing casing deformation machine learning prediction method based on a mechanistic model according to claim 6, characterized in that, Based on enabling the overall analysis mode, the number of qualified simulation feature sets that meet the acquisition criteria in the target initial simulation dataset and the number of unqualified simulation feature sets that do not meet the acquisition criteria in the target initial simulation dataset are integrated. Based on the integration results, an overall anomaly score is determined, and based on the overall anomaly score, it is determined whether to enable the re-collection mode or the preprocessing mode.
8. The shale fracturing casing deformation machine learning prediction method based on a mechanism model according to claim 7, characterized in that, The process of optimizing the initial numerical simulation results by performing inversion based on each set of latent simulation features to obtain the actual numerical simulation results includes: Based on historical fitting, the sets of hidden simulation features are inverted to obtain the actual numerical simulation results; The actual numerical simulation results are compared with the actual numerical calculation results obtained. Based on the comparison results, determine whether to optimize the initial prediction criteria, or perform historical fitting again.
Citation Information
Patent Citations
Shale hydraulic fracturing numerical simulation method based on PFC discrete element
CN114878340A
Shale oil three-dimensional well pattern fracturing multivariate parameter optimization design method
CN117454755A
Wellbore complex well geological potential assessment method, electronic equipment and storage medium
CN118898218A