A method and system for optimizing the main controlling factors of shale oil production based on fuzzy classification

By classifying the factors affecting shale oil production using fuzzy classification, and using membership functions and parabolic membership degrees to identify highly correlated characteristic parameters, the hidden correlation problem in shale oil production prediction is solved, and the recovery rate is improved.

CN119693180BActive Publication Date: 2025-10-31CHINA NAT PETROLEUM CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311233126.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-22
Publication Date
2025-10-31
Estimated Expiration
2043-09-22

AI Technical Summary

Technical Problem

Existing technologies struggle to uncover hidden correlations among factors influencing shale oil production under complex geological conditions, resulting in low recovery rates.

Method used

The fuzzy classification method is used to classify the factors affecting shale oil production. By establishing membership functions and parabolic membership degree calculations, highly correlated characteristic parameters are identified as the main controlling factors.

Benefits of technology

By using fuzzy classification methods, clear trends can be extracted from chaotic data, and characteristic parameters that are highly correlated with shale oil production can be identified, thereby improving the accuracy of shale oil production prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119693180B_ABST
    Figure CN119693180B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for optimizing the main controlling factors of shale oil production based on fuzzy classification. The method includes: collecting and cleaning multiple basic sample data, and establishing a sample set using the cleaned sample data; wherein, the basic sample data includes: characteristic parameters that affect shale oil production and target parameters corresponding to the characteristic parameters; classifying each sample data in the sample set into categories using a fuzzy classification method, and plotting the mean change trend of each characteristic parameter in each category of sample data under different numbers of categories, so as to analyze the correlation between each characteristic parameter and the category under different numbers of categories; plotting the relationship curve between each characteristic parameter and the corresponding target parameter in each category of sample data, determining the correlation between each characteristic parameter and the corresponding target parameter and sorting them, and selecting the characteristic parameters with high correlation as the main controlling factors affecting shale oil production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of oil and gas extraction technology, and in particular to a method and system for optimizing the main controlling factors of shale oil production based on fuzzy classification. Background Technology

[0002] Shale oil is a very important unconventional oil and gas resource. China has abundant shale oil resources, ranking among the world's top, and has a very broad development prospect. However, shale oil reservoirs have diverse pore types and complex structures, and shale oil has diverse occurrence states and complex flow mechanisms, resulting in often low recovery rates. Many factors influence shale oil production, including geological factors, rock mechanics and geostress factors, and engineering factors.

[0003] For data with simple influencing factors, strong correlations, and explicit correlations, existing technologies typically involve directly analyzing sample point data to obtain the correlation between feature parameters and target parameters. Then, feature parameters are sorted according to the magnitude of the correlation, and feature parameters with strong correlations are selected as the main control parameters.

[0004] However, in situations with highly complex geological conditions and numerous engineering factors, existing technologies, when analyzing data based on sample points, often fail to uncover subtle correlations, which are hidden within large amounts of data. Summary of the Invention

[0005] The purpose of this invention is to provide a method and system for optimizing the main controlling factors of shale oil production based on fuzzy classification. This invention classifies sample data, taking one type of sample data as a sample point. By analyzing these new sample points, hidden correlations in the data can be discovered.

[0006] To achieve the above objectives, this invention provides a method for optimizing the main controlling factors of shale oil production based on fuzzy classification, comprising:

[0007] Multiple basic sample data are collected and cleaned, and a sample set is established using the cleaned sample data; wherein, the basic sample data includes: characteristic parameters that affect shale oil production and the target parameters corresponding to the characteristic parameters;

[0008] A fuzzy classification method is used to classify each sample data in the sample set, and the mean change trend of each feature parameter in each class of sample data is plotted under different numbers of classifications, so as to analyze the correlation between each feature parameter and the class under different numbers of classifications.

[0009] Plot the relationship curve between each feature parameter and the corresponding target parameter in each type of sample data, determine the correlation between each feature parameter and the corresponding target parameter and sort them, and select the feature parameters with high correlation as the main controlling factors affecting shale oil production.

[0010] Furthermore, multiple basic sample data are collected and cleaned, and a sample set is established using the cleaned sample data, including:

[0011] Collect multiple sets of the aforementioned feature parameters and the target parameters corresponding to each set of the aforementioned feature parameters;

[0012] The basic sample data in which the feature parameter or the target parameter is zero, as well as the basic sample data in which the feature parameter or the target parameter is abnormal, are cleaned and removed to obtain cleaned sample data;

[0013] A sample set is built using the cleaned sample data;

[0014] Each basic sample data includes: a set of the aforementioned feature parameters and their corresponding target parameters;

[0015] The set of characteristic parameters includes: net thickness, porosity, permeability, oil saturation, Poisson's ratio, cohesion, compressive strength, sand strength, main fracture half-length, cluster number, fracturing fluid volume, fracturing section length, number of fracturing sections, and horizontal section length;

[0016] The target parameter is the cumulative output within a predetermined number of days.

[0017] Furthermore, a fuzzy classification method is used to classify each sample data in the sample set into categories, including:

[0018] Based on the fuzzy classification method, by establishing the membership function of each category, each sample data in the sample set is classified into different categories;

[0019] The target parameter is used as the independent variable. Select the maximum value of the target parameter in the sample set. Minimum value The maximum and minimum values ​​are used as the horizontal axis; and the membership function value of each category is used as the vertical axis y, with the minimum value of the membership function being 0 and the maximum value being 1; the samples in the sample set are divided into n categories according to the requirements;

[0020] After dividing the samples in the sample set into n categories, the parabolic membership degree of each category corresponding to each sample data is calculated based on the membership function, and the category with the largest parabolic membership degree is selected as the classification result of the corresponding sample data.

[0021] Furthermore, based on the membership function, the parabolic membership degree of each category corresponding to each sample data is calculated, including:

[0022] The parabolic membership degree of each sample data point in the first class is given by the following formula:

[0023]

[0024] In the formula, and All represent the parabolic membership degree of the first type of sample data, d represents the mean, and Δd represents the deviation. This represents the independent variable, i.e., the target parameter in the sample set; This represents the minimum value of the target parameter in the sample set;

[0025] The parabolic membership degree of the i-th class corresponding to each sample data is given by the following formula:

[0026]

[0027] In the formula, and All indicate the first Parabolic membership degree of class sample data Indicates the first Categories , Indicates the total number of categories;

[0028] The parabolic membership degree of the nth class corresponding to each sample data is given by the following formula:

[0029]

[0030] In the formula, and Both represent the parabolic membership degrees of the nth class of sample data. This represents the maximum value of the target parameter in the sample set;

[0031] The average quantity is calculated using the following formula. for:

[0032] .

[0033] Furthermore, the mean change trend of each feature parameter in each class of sample data is plotted for different numbers of classifications, in order to analyze the correlation between each feature parameter and the class under different numbers of classifications, including:

[0034] Calculate the mean of each feature parameter in each class of sample data;

[0035] Plot the trend of the mean change of each feature parameter in the sample data of each class under different numbers of classifications;

[0036] Based on the mean change trend graph of each feature parameter, observe the trend of each feature parameter changing with the number of categories, so as to analyze the correlation between each feature parameter and the category.

[0037] Furthermore, the mean of each feature parameter is calculated using the following formula:

[0038]

[0039] In the formula, Let represent the mean of the j-th feature parameter in the i-th class of sample data. This represents the total number of samples in the i-th category. This represents the k-th value of the j-th feature parameter in the i-th class of sample data, where k represents the sample number.

[0040] Furthermore, the relationship curves between each feature parameter and its corresponding target parameter in each type of sample data are plotted. The correlation between each feature parameter and its corresponding target parameter is determined and ranked. Feature parameters with high correlation are selected as the main controlling factors affecting shale oil production, including:

[0041] Calculate the mean of the target parameter in each class of sample data;

[0042] Normalize the mean of each feature parameter and the mean of the corresponding target parameter respectively to obtain the normalized value of each feature parameter and the normalized value of the corresponding target parameter of the sample data of the corresponding category;

[0043] Plot the relationship curve between the normalized value of each feature parameter and the normalized value of the corresponding target parameter;

[0044] Based on the inflection point of the relationship curve corresponding to each feature parameter, determine the straight line segment of the relationship curve, and calculate the slope of the straight line segment corresponding to each feature parameter.

[0045] The absolute values ​​of the slopes of the straight line segments of each obtained feature parameter are sorted, and the feature parameter with the largest slope is selected as the main controlling factor affecting shale oil production.

[0046] Furthermore, the mean of the target parameter in each class of sample data is calculated using the following formula:

[0047]

[0048] In the formula, Let represent the mean of the target parameter in the i-th class of sample data. Let k be the target parameter of the i-th class of sample data, where k represents the sample number; This represents the total number of samples in the i-th class of sample data;

[0049] The mean of each feature parameter and the mean of the corresponding target parameter are normalized using the following formula:

[0050]

[0051] In the formula, This represents the normalized value of the kth total parameter of the i-th class of sample data, that is, the normalized value of the kth feature parameter of the i-th class of sample data and the normalized value of the corresponding target parameter. Let represent the mean of the k-th total parameter of the i-th class of sample data. This represents the minimum mean of the k-th total parameter. This represents the maximum mean of the k-th total parameter, where k represents the sample number.

[0052] Based on the same inventive concept, this invention also provides a fuzzy classification-based optimization system for key factors controlling shale oil production, comprising:

[0053] The cleaning unit is used to collect and clean multiple basic sample data, and to build a sample set using the cleaned sample data;

[0054] The analysis unit is used to classify each sample data in the sample set using a fuzzy classification method, and to plot the mean change trend of each feature parameter in each category of sample data under different numbers of categories, so as to analyze the correlation between each feature parameter and the category under different numbers of categories.

[0055] A selection unit is used to plot the relationship curve between each feature parameter and the corresponding target parameter in each type of sample data, determine the correlation between each feature parameter and the corresponding target parameter and sort them, and select the feature parameter with high correlation as the main controlling factor affecting shale oil production.

[0056] The basic sample data includes: characteristic parameters that affect shale oil production and the target parameters corresponding to those characteristic parameters.

[0057] Based on the same inventive concept, embodiments of the present invention also provide an electronic device, including: a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, it implements the aforementioned method for optimizing the main controlling factors of shale oil production based on fuzzy classification.

[0058] Based on the same inventive concept, embodiments of the present invention also provide a computer storage medium storing computer-executable instructions, which, when executed, implement the aforementioned method for optimizing the main controlling factors of shale oil production based on fuzzy classification.

[0059] The technical effects and advantages of this invention are as follows: By classifying sample data and treating each class of sample data as a sample point, this invention can extract clear trends from seemingly chaotic data and obtain the correlation between the classified feature parameters and the target parameters. In other words, this invention transforms the hidden correlation in the samples into an explicit correlation, which has profound significance for subsequent shale oil production classification and prediction research.

[0060] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description

[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0062] Figure 1 This is a flowchart illustrating a method for optimizing the main controlling factors of shale oil production based on fuzzy classification, according to an embodiment of the present invention.

[0063] Figure 2 This is a schematic diagram illustrating the classification results of sample data when the sample set is divided into 3 categories in an embodiment of the present invention;

[0064] Figure 3 This is a schematic diagram illustrating the classification results of sample data when the sample set is divided into 4 categories in an embodiment of the present invention;

[0065] Figure 4 This is a schematic diagram of the classification results of sample data when the sample set is divided into 5 categories in an embodiment of the present invention;

[0066] Figure 5 This is a schematic diagram of the data distribution between the horizontal segment length and the cumulative output over 30 days in an embodiment of the present invention;

[0067] Figure 6 This is a schematic diagram showing the data distribution between sand addition intensity and cumulative output over 30 days in an embodiment of the present invention;

[0068] Figure 7 This is a schematic diagram illustrating the data distribution between Poisson's ratio and 30-day cumulative output in an embodiment of the present invention;

[0069] Figure 8 This is a schematic diagram illustrating the trend of the mean length of the horizontal segment of each category of sample data when the sample data in the sample set is divided into 3 categories in an embodiment of the present invention.

[0070] Figure 9 This is a schematic diagram illustrating the trend of the mean length of the horizontal segment of each category of sample data when the sample data in the sample set is divided into 4 categories in an embodiment of the present invention.

[0071] Figure 10 This is a schematic diagram illustrating the trend of the mean length of the horizontal segment of each category of sample data when the sample data in the sample set is divided into 5 categories in an embodiment of the present invention.

[0072] Figure 11 This is a schematic diagram illustrating the trend of the mean sand addition intensity of each category of sample data when the sample data in the sample set is divided into three categories in an embodiment of the present invention.

[0073] Figure 12 This is a schematic diagram illustrating the trend of the mean sand addition intensity of each category of sample data when the sample data in the sample set is divided into four categories in an embodiment of the present invention.

[0074] Figure 13 This is a schematic diagram illustrating the trend of the mean sand addition intensity of each category of sample data when the sample data in the sample set is divided into 5 categories in an embodiment of the present invention.

[0075] Figure 14 This is a schematic diagram illustrating the trend of the mean Poisson's ratio of each category when the sample data in the sample set is divided into three categories in an embodiment of the present invention.

[0076] Figure 15 This is a schematic diagram illustrating the trend of the mean Poisson's ratio of each category when the sample data in the sample set is divided into four categories in an embodiment of the present invention.

[0077] Figure 16 This is a schematic diagram illustrating the trend of the mean Poisson's ratio of each category when the sample data in the sample set is divided into 5 categories in an embodiment of the present invention.

[0078] Figure 17 This is a schematic diagram showing the relationship between the normalized value of the horizontal segment length and the normalized value of the cumulative output over 30 days in an embodiment of the present invention.

[0079] Figure 18 This is a schematic diagram showing the relationship between the normalized value of sand addition intensity and the normalized value of cumulative output over 30 days in an embodiment of the present invention.

[0080] Figure 19This is a schematic diagram showing the relationship between the normalized Poisson's ratio and the normalized cumulative yield over 30 days in an embodiment of the present invention.

[0081] Figure 20 This is a schematic diagram showing the sorting results of the absolute values ​​of the slopes of the line segments of all feature parameters in this embodiment of the invention.

[0082] Figure 21 This is a schematic diagram of a fuzzy classification-based optimization system for the main controlling factors of shale oil production according to an embodiment of the present invention.

[0083] Figure 22 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0084] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0085] To address the shortcomings of existing technologies, this invention discloses a method for optimizing the main controlling factors of shale oil production based on fuzzy classification, such as... Figure 1 As shown, it includes:

[0086] Step S1: Collect and clean multiple basic sample data, and establish a sample set using the cleaned sample data; specifically including:

[0087] Based on geological research data and fracturing operation reports, multiple sets of characteristic parameters that affect shale oil production and the target parameters corresponding to each set of characteristic parameters were collected and organized. Among them, each basic sample data includes: a set of characteristic parameters that affect shale oil production and the target parameters corresponding to this set of characteristic parameters.

[0088] The target parameter is the cumulative output over 30 days;

[0089] Each set of characteristic parameters mainly includes: geological parameters, rock mechanics and geostress parameters, and engineering parameters; geological parameters include: net thickness (m), porosity (%), permeability (mD), oil saturation (%), etc.; rock mechanics and geostress parameters include: Poisson's ratio, cohesion (MPa), compressive strength (MPa), etc.; engineering parameters include: sand addition strength (t / m), main fracture half-length (m), cluster number, fracturing fluid volume (m³ / s). 3 The parameters include, for example, the length of the fracturing section (m), the number of fracturing sections, and the length of the horizontal section (m); therefore, each set of characteristic parameters includes a total of 14 parameters.

[0090] The basic sample data with zero feature parameters or target parameters, as well as the basic sample data with abnormal feature parameters or target parameters, are cleaned and removed to obtain cleaned sample data; a sample set is then established using the cleaned sample data.

[0091] The cleaned sample data is shown in Table 1:

[0092] Table 1

[0093]

[0094] Step S2: Use a fuzzy classification method to classify each sample data in the sample set, and plot the mean change trend of each feature parameter in each class of sample data under different numbers of classifications, to analyze the correlation between each feature parameter and the class under different numbers of classifications; specifically including:

[0095] Step S201: Classify each sample data in the sample set using a fuzzy classification method. The specific steps are as follows:

[0096] Based on the fuzzy classification method, each sample data in the sample set is classified into different categories by establishing a membership function for each category.

[0097] The target parameter (cumulative output over 30 days) is used as the independent variable. Select the maximum value of the target parameter in the sample set. (6250t), minimum value (694t) is used as the maximum and minimum values ​​of the horizontal axis; and the membership function value of each category is used as the vertical axis y, with the minimum value of the membership function being 0 and the maximum value being 1; the samples in the sample set are divided into n categories according to the requirements.

[0098] After dividing the samples in the sample set into n categories, the parabolic membership degree of each category corresponding to each sample data is calculated based on the membership function, and the category with the largest parabolic membership degree is selected as the classification result of the corresponding sample data.

[0099] The formula for calculating the parabolic membership degree of each category for each sample data includes:

[0100] The parabolic membership degree of each sample data point in the first class is given by the following formula:

[0101]

[0102] In the formula, and All represent the parabolic membership degree of the first type of sample data, d represents the mean, and Δd represents the deviation. This represents the independent variable, i.e., the target parameter in the sample data; This represents the minimum value of the target parameter in the sample set.

[0103] The parabolic membership degree of the i-th class corresponding to each sample data is given by the following formula:

[0104]

[0105] In the formula, and All indicate the first Parabolic membership degree of class sample data Indicates the first Categories , d represents the total number of categories; d represents the average value; Δd ​​represents the deviation. This represents the independent variable, i.e., the target parameter in the sample set; This represents the minimum value of the target parameter in the sample set. This represents the maximum value of the target parameter in the sample set.

[0106] The parabolic membership degree of the nth class corresponding to each sample is given by the following formula:

[0107]

[0108] In the formula, and ... This represents the independent variable, i.e., the target parameter in the sample set; This represents the minimum value of the target parameter in the sample set. This represents the maximum value of the target parameter in the sample set.

[0109] The average quantity is calculated using the following formula. for:

[0110]

[0111] In the formula, Indicates the total number of categories. This represents the minimum value of the target parameter in the sample set. This represents the maximum value of the target parameter in the sample set.

[0112] Therefore, (1) when the samples in the sample set are divided into 3 categories in the embodiment:

[0113] The parabolic membership degree corresponding to Class 1 for each sample data is calculated using the following formula:

[0114] In the formula, and Both represent the parabolic membership degrees of the first type of sample data. This represents the independent variable, i.e., the target parameter in the sample data;

[0115] The parabolic membership degree for each sample data corresponding to Class 2 is calculated using the following formula:

[0116] In the formula, and All indicate the first Parabolic membership degree of class sample data Indicates the first Each category, at this time , This represents the independent variable, i.e., the target parameter in the sample set;

[0117] The parabolic membership degree for each sample data corresponding to the 3rd class is calculated using the following formula:

[0118] In the formula, and Both represent the parabolic membership degrees of the nth class of sample data. This represents the total number of categories. ; This represents the independent variable, i.e., the target parameter in the sample set.

[0119] When dividing the samples in the sample set into three categories, the parabolic membership degree for each sample data corresponding to each category is calculated according to the membership degree calculation formula mentioned above. The category with the largest parabolic membership degree is selected as the classification result for that sample data, and the classification result is as follows: Figure 2 As shown, 51.8% of the sample data in the sample set were classified as Class III wells, 37.2% were classified as Class II wells, and 11% were classified as Class I wells.

[0120] (2) In the embodiment, when the samples in the sample set are divided into 4 categories:

[0121] The parabolic membership degree corresponding to Class 1 for each sample data is calculated using the following formula:

[0122] In the formula, and Both represent the parabolic membership degrees of the first type of sample data. This represents the independent variable, i.e., the target parameter in the sample set;

[0123] The parabolic membership degree for each sample data corresponding to Class 2 is calculated using the following formula:

[0124] In the formula, and All indicate the first Parabolic membership degree of class sample data Indicates the first Each category, at this time ; This represents the independent variable, i.e., the target parameter in the sample set;

[0125] The parabolic membership degree for each sample data corresponding to the 3rd class is calculated using the following formula:

[0126] In the formula and All indicate the first Parabolic membership degree of class sample data Indicates the first Each category, at this time ; This represents the independent variable, i.e., the target parameter in the sample set;

[0127] The parabolic membership degree for each sample data corresponding to the 4th class is calculated using the following formula:

[0128] In the formula, and Both represent the parabolic membership degrees of the nth class of sample data. This represents the total number of categories. ; This represents the independent variable, i.e., the target parameter in the sample set.

[0129] When the samples in the sample set are divided into 4 categories, the classification results are as follows: Figure 3 As shown, 36.4% of the sample data in the sample set were classified as Class IV wells, 35.9% as Class III wells, 21.9% as Class II wells, and 5.8% as Class I wells.

[0130] (3) In the embodiment, when the samples in the sample set are divided into 5 categories:

[0131] The parabolic membership degree corresponding to Class 1 for each sample data is calculated using the following formula:

[0132] In the formula, and Both represent the parabolic membership degrees of the first type of sample data. This represents the independent variable, i.e., the target parameter in the sample data;

[0133] The parabolic membership degree for each sample data corresponding to Class 2 is calculated using the following formula:

[0134] In the formula, and All indicate the first Parabolic membership degree of class sample data Indicates the first Each category, at this time ; This represents the independent variable, i.e., the target parameter in the sample set;

[0135] The parabolic membership degree for each sample data corresponding to the 3rd class is calculated using the following formula:

[0136] In the formula and All indicate the first Parabolic membership degree of class sample data Indicates the first Each category, at this time ; This represents the independent variable, i.e., the target parameter in the sample set;

[0137] The parabolic membership degree for each sample data corresponding to the 4th class is calculated using the following formula:

[0138] In the formula and All indicate the first Parabolic membership degree of class sample data Indicates the first Each category, at this time ; This represents the independent variable, i.e., the target parameter in the sample set;

[0139] The parabolic membership degree for each sample data point corresponding to the 5th class is calculated using the following formula:

[0140] In the formula, and Both represent the parabolic membership degrees of the nth class of sample data. This represents the total number of categories. ; This represents the independent variable, i.e., the target parameter in the sample set.

[0141] When the samples in the sample set are divided into 5 categories, the classification results are as follows: Figure 4As shown, 25.8% of the sample data in the sample set were classified as Class V wells, 31.6% as Class IV wells, 23.8% as Class III wells, 15.3% as Class II wells, and 3.5% as Class I wells.

[0142] Step S202: Plot the mean change trend of each feature parameter in the sample data of each class under different numbers of classifications, so as to analyze the correlation between each feature parameter and the class under different numbers of classifications; the specific steps are as follows:

[0143] like Figures 5-7 As shown, it can be observed that there is no obvious correlation between the data distribution of horizontal section length, sand addition intensity, Poisson's ratio and 30-day cumulative output; that is, no obvious correlation can be found between the data of characteristic parameters and target parameters (30-day cumulative output).

[0144] Therefore, the mean of each feature parameter in each class of sample data is calculated using the following formula:

[0145]

[0146] In the formula, Let represent the mean of the j-th feature parameter in the i-th class of sample data. This represents the total number of samples in the i-th category. This represents the k-th value of the j-th feature parameter in the i-th sample data (i.e., the k-th feature parameter value of the j-th feature parameter in the i-th sample data), where k represents the sample number.

[0147] Plot the trend of the mean change of each feature parameter in the sample data of each class under different numbers of categories; based on the trend of the mean change of each feature parameter, observe the trend of each feature parameter with the number of categories, so as to analyze the correlation between each feature parameter and the category.

[0148] like Figures 8-16 As shown: Taking the horizontal segment length, sand addition intensity, and Poisson's ratio as examples, according to the trend chart of the mean change of the horizontal segment length, sand addition intensity, and Poisson's ratio, it can be found that no matter whether the samples are divided into 3, 4, or 5 categories, the mean change trend of the horizontal segment length, sand addition intensity, and Poisson's ratio with the category is the same. This shows that the horizontal segment length, sand addition intensity, and Poisson's ratio have a good correlation with the category.

[0149] Step S3: Plot the relationship curve between each feature parameter and the corresponding target parameter in each type of sample data, determine the correlation between each feature parameter and the corresponding target parameter and sort them, and select the feature parameters with high correlation as the main controlling factors affecting shale oil production; including:

[0150] Set the number of categories to 50, and calculate the mean of the target parameter for each category of sample data:

[0151]

[0152] In the formula, Let represent the mean of the target parameter in the i-th class of sample data. Let k be the target parameter of the i-th class of sample data, where k represents the sample number; This represents the total number of samples in the i-th class of sample data.

[0153] Normalize the mean of each feature parameter and the mean of the corresponding target parameter to obtain the normalized value of each feature parameter and the normalized value of the corresponding target parameter for the corresponding category of sample data:

[0154]

[0155] In the formula, This represents the normalized value of the k-th total parameter of the i-th class of sample data. Let represent the mean of the k-th total parameter of the i-th class of sample data. This represents the minimum mean of the k-th total parameter. This represents the maximum mean of the k-th total parameter, where k represents the sample number;

[0156] Since the normalization formulas for the feature parameters and the target parameters are the same, the feature parameters and the target parameters are combined in the formula, that is, the total parameters include both feature parameters and target parameters.

[0157] Plot the relationship curves between the normalized value of each feature parameter and the corresponding normalized value of the target parameter; (each point represents a category, the horizontal axis represents the normalized value of the feature parameter of this category of sample data, and the vertical axis represents the normalized value of the target parameter of this category of sample data). Figure 17-19 As shown.

[0158] Based on the inflection point of the relationship curve corresponding to each feature parameter, the straight line segment of the relationship curve is determined, and the slope of the straight line segment corresponding to each feature parameter is calculated. In the embodiment, the slopes of the straight line segments of the relationship curves for horizontal segment length, sand addition intensity, and Poisson's ratio are 5.6537, 3.4878, and -1.2458, respectively.

[0159] The slopes of the linear segments of the relationship curves for all characteristic parameters are shown in Table 2.

[0160] Table 2

[0161]

[0162] The absolute values ​​of the slopes of the line segments for each of the obtained feature parameters are sorted, and the sorting results are as follows: Figure 20 As shown, 11 parameters (characteristic parameters with large slopes) were selected as the main controlling factors affecting shale oil production, including horizontal section length, number of fracturing sections, fracturing fluid volume, net thickness, half length of main fracture, oil saturation, sand addition intensity, permeability, cluster number, fracturing section length, and porosity.

[0163] Based on the same inventive concept, this invention also provides a fuzzy classification-based optimization system for key factors controlling shale oil production, such as... Figure 21 As shown, it includes:

[0164] The cleaning unit is used to collect and clean multiple basic sample data, and to build a sample set using the cleaned sample data;

[0165] The analysis unit is used to classify each sample data in the sample set using a fuzzy classification method, and to plot the mean change trend of each feature parameter in each category of sample data under different numbers of categories, so as to analyze the correlation between each feature parameter and the category under different numbers of categories.

[0166] A selection unit is used to plot the relationship curve between each feature parameter and the corresponding target parameter in each type of sample data, determine the correlation between each feature parameter and the corresponding target parameter and sort them, and select the feature parameter with high correlation as the main controlling factor affecting shale oil production.

[0167] The basic sample data includes: characteristic parameters that affect shale oil production and the target parameters corresponding to those characteristic parameters.

[0168] Regarding the system in the above embodiments, the specific manner in which each unit module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0169] Based on the same inventive concept, embodiments of the present invention also provide an electronic device, the structure of which is as follows: Figure 22 As shown, it includes: a memory and a processor, wherein the processor is used to read and execute the computer program stored in the memory to implement the aforementioned method for optimizing the main controlling factors of shale oil production based on fuzzy classification.

[0170] Based on the same inventive concept, embodiments of the present invention also provide a computer storage medium storing computer-executable instructions, which, when executed, implement the aforementioned method for optimizing the main controlling factors of shale oil production based on fuzzy classification.

[0171] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A method for optimizing the main controlling factors of shale oil production based on fuzzy classification, characterized in that, include: Multiple basic sample data are collected and cleaned, and a sample set is established using the cleaned sample data; wherein, the basic sample data includes: characteristic parameters that affect shale oil production and the target parameters corresponding to the characteristic parameters; A fuzzy classification method is used to classify each sample data in the sample set, and the mean change trend of each feature parameter in each class of sample data is plotted under different numbers of classifications, so as to analyze the correlation between each feature parameter and the class under different numbers of classifications. Plot the relationship curve between each feature parameter and the corresponding target parameter in each type of sample data, determine the correlation between each feature parameter and the corresponding target parameter and sort them, and select the feature parameters with high correlation as the main controlling factors affecting shale oil production; in, The fuzzy classification method is used to classify each sample data in the sample set into categories, including: Based on the fuzzy classification method, a membership function is established for each category to classify each sample data in the sample set into different categories; the target parameter is used as the independent variable. Select the maximum value of the target parameter in the sample set. Minimum value The maximum and minimum values ​​are used as the horizontal axis; and the membership function value of each category is used as the vertical axis y, with the minimum value of the membership function being 0 and the maximum value being 1; the samples in the sample set are divided into n categories according to the requirements; after dividing the samples in the sample set into n categories, the parabolic membership degree of each category corresponding to each sample data is calculated based on the membership function, and the category with the largest parabolic membership degree is selected as the classification result of the corresponding sample data; Furthermore, we plotted the mean change trend of each feature parameter in each class of sample data under different numbers of classifications, in order to analyze the correlation between each feature parameter and the class under different numbers of classifications, including: Calculate the mean of each feature parameter in each class of sample data; plot the trend of the mean of each feature parameter in each class of sample data under different numbers of categories; based on the trend of the mean of each feature parameter, observe the trend of each feature parameter with the number of categories, so as to analyze the correlation between each feature parameter and the category.

2. The method for optimizing the main controlling factors of shale oil production based on fuzzy classification according to claim 1, characterized in that, Multiple basic sample data were collected and cleaned. A sample set was then established using the cleaned sample data, including: Collect multiple sets of the aforementioned feature parameters and the target parameters corresponding to each set of the aforementioned feature parameters; The basic sample data in which the feature parameter or the target parameter is zero, as well as the basic sample data in which the feature parameter or the target parameter is abnormal, are cleaned and removed to obtain cleaned sample data; A sample set is built using the cleaned sample data; Each basic sample data includes: a set of the aforementioned feature parameters and their corresponding target parameters; The set of characteristic parameters includes: net thickness, porosity, permeability, oil saturation, Poisson's ratio, cohesion, compressive strength, sand strength, main fracture half-length, cluster number, fracturing fluid volume, fracturing section length, number of fracturing sections, and horizontal section length; The target parameter is the cumulative output within a predetermined number of days.

3. The method for optimizing the main controlling factors of shale oil production based on fuzzy classification according to claim 1, characterized in that, The parabolic membership degree for each category corresponding to each sample data is calculated based on the membership function, including: The parabolic membership degree of each sample data point in the first class is given by the following formula: In the formula, and All represent the parabolic membership degree of the first type of sample data, d represents the mean, and Δd represents the deviation. This represents the independent variable, i.e., the target parameter in the sample set; This represents the minimum value of the target parameter in the sample set; The parabolic membership degree of the i-th class corresponding to each sample data is given by the following formula: In the formula, and All indicate the first Parabolic membership degree of class sample data Indicates the first Categories , Indicates the total number of categories; The parabolic membership degree of the nth class corresponding to each sample data is given by the following formula: In the formula, and Both represent the parabolic membership degrees of the nth class of sample data. This represents the maximum value of the target parameter in the sample set; The average quantity is calculated using the following formula. for: 。 4. The method for optimizing the main controlling factors of shale oil production based on fuzzy classification according to claim 1, characterized in that, The mean of each feature parameter is calculated using the following formula: In the formula, Let represent the mean of the j-th feature parameter in the i-th class of sample data. This represents the total number of samples in the i-th category. This represents the k-th value of the j-th feature parameter in the i-th class of sample data, where k represents the sample number.

5. The method for optimizing the main controlling factors of shale oil production based on fuzzy classification according to claim 1, characterized in that, Plot the relationship curve between each feature parameter and its corresponding target parameter in each type of sample data, determine the correlation between each feature parameter and its corresponding target parameter and sort them, and select the feature parameters with high correlation as the main controlling factors affecting shale oil production, including: Calculate the mean of the target parameter in each class of sample data; Normalize the mean of each feature parameter and the mean of the corresponding target parameter respectively to obtain the normalized value of each feature parameter and the normalized value of the corresponding target parameter of the sample data of the corresponding category; Plot the relationship curve between the normalized value of each feature parameter and the normalized value of the corresponding target parameter; Based on the inflection point of the relationship curve corresponding to each feature parameter, determine the straight line segment of the relationship curve, and calculate the slope of the straight line segment corresponding to each feature parameter. The absolute values ​​of the slopes of the straight line segments of each obtained feature parameter are sorted, and the feature parameter with the largest slope is selected as the main controlling factor affecting shale oil production.

6. The method for optimizing the main controlling factors of shale oil production based on fuzzy classification according to claim 5, characterized in that, The mean of the target parameter in each class of sample data is calculated using the following formula: In the formula, Let represent the mean of the target parameter in the i-th class of sample data. Let k be the target parameter of the i-th class of sample data, where k represents the sample number; This represents the total number of samples in the i-th class of sample data; The mean of each feature parameter and the mean of the corresponding target parameter are normalized using the following formula: In the formula, This represents the normalized value of the kth total parameter of the i-th class of sample data, that is, the normalized value of the kth feature parameter of the i-th class of sample data and the normalized value of the corresponding target parameter; Let represent the mean of the k-th total parameter of the i-th class of sample data. This represents the minimum mean of the k-th total parameter. This represents the maximum mean of the k-th total parameter, where k represents the sample number.

7. A fuzzy classification-based optimization system for key factors controlling shale oil production, characterized in that, include: The cleaning unit is used to collect and clean multiple basic sample data, and to build a sample set using the cleaned sample data; The analysis unit is used to classify each sample data in the sample set using a fuzzy classification method, and to plot the mean change trend of each feature parameter in each category of sample data under different numbers of categories, so as to analyze the correlation between each feature parameter and the category under different numbers of categories. A selection unit is used to plot the relationship curve between each feature parameter and the corresponding target parameter in each type of sample data, determine the correlation between each feature parameter and the corresponding target parameter and sort them, and select the feature parameter with high correlation as the main controlling factor affecting shale oil production. in, The basic sample data includes: characteristic parameters that affect shale oil production and the target parameters corresponding to those characteristic parameters; The fuzzy classification method is used to classify each sample data in the sample set into categories, including: Based on the fuzzy classification method, a membership function is established for each category to classify each sample data in the sample set into different categories; the target parameter is used as the independent variable. Select the maximum value of the target parameter in the sample set. Minimum value The maximum and minimum values ​​are used as the horizontal axis; and the membership function value of each category is used as the vertical axis y, with the minimum value of the membership function being 0 and the maximum value being 1; the samples in the sample set are divided into n categories according to the requirements; after dividing the samples in the sample set into n categories, the parabolic membership degree of each category corresponding to each sample data is calculated based on the membership function, and the category with the largest parabolic membership degree is selected as the classification result of the corresponding sample data; Furthermore, we plotted the mean change trend of each feature parameter in each class of sample data under different numbers of classifications, in order to analyze the correlation between each feature parameter and the class under different numbers of classifications, including: Calculate the mean of each feature parameter in each class of sample data; plot the trend of the mean of each feature parameter in each class of sample data under different numbers of categories; based on the trend of the mean of each feature parameter, observe the trend of each feature parameter with the number of categories, so as to analyze the correlation between each feature parameter and the category.

8. An electronic device, characterized in that, include: Memory, processor; The processor is used to read and execute the computer program stored in the memory to implement the method for optimizing the main controlling factors of shale oil production based on fuzzy classification as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed, implement the method for optimizing the main controlling factors of shale oil production based on fuzzy classification as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Wind power plant power prediction method based on fuzzy clustering and deep reinforcement learning

    CN112288157A

  • Power load data weighted incremental clustering method for adaptively determining clustering number

    CN116451097A