A cost-sensitive learning-based fluid identification method for water-rich tight sandstone
By constructing a cost-sensitive gradient booster method, combining fluid-sensitive logging composite parameters and optimized weight updates, and using a sparrow search algorithm for optimization, the accuracy problem in fluid identification of tight sandstone reservoirs was solved, and the accuracy of identification was improved.
Patent Information
- Application Number
- CN202311124905.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-02
- Publication Date
- 2026-01-02
- Estimated Expiration
- 2043-09-02
AI Technical Summary
Existing technologies have low accuracy in fluid identification in tight sandstone reservoirs, especially in water-rich tight sandstone. This is mainly due to the complexity caused by the small pore throat and the difficulty in identification caused by the similarity of data features, which leads to technical problems that are difficult to solve with existing technologies.
A cost-sensitive gradient booster method was adopted, which improved the accuracy of fluid identification by constructing fluid-sensitive logging composite parameters and optimizing the weight update of the gradient booster decision tree, combined with the sparrow search algorithm for optimization.
It improves the accuracy of fluid identification, especially the accuracy of identifying gas layers, gas-water layers, and water layers, and solves the technical problems existing in the prior art.
Smart Images

Figure CN117150390B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of well logging fluid identification and machine learning, and particularly relates to a cost-sensitive integrated learning well logging fluid type intelligent discrimination method. BACKGROUND
[0002] Tight sandstone gas is widely distributed in major oil and gas basins in the world, and is one of the important unconventional gas types. With the depletion of structural oil and gas resources, tight sandstone reservoirs have gradually become the focus of attention of petroleum and natural gas engineers. Such reservoirs have strong heterogeneity, and have the characteristics of low porosity, low permeability and low saturation due to small pore throats. This makes the rock skeleton the primary influencing factor of logging response, reduces the contribution of different fluids in the reservoir to the logging response, and brings great challenges to well logging fluid interpretation work. At present, the fluid identification work of tight sandstone gas reservoirs begins to be combined with artificial intelligence.
[0003] Guo Y (Guo Y. Well logging interpretation and productivity prediction of tight sandstone reservoir based on seepage and electrical conductivity characteristics[D]. Jilin University, 2017.) used PSO-SVM algorithm and random forest algorithm for gas layer identification and gas-bearing property evaluation, with an accuracy of 76%. Han Y (Han Y. Intelligent identification of reservoir fluid in Daniudi gas field based on AdaBoost machine learning algorithm[J]. Oil Drilling Technology, 2022, 50(1): 112-118.) used Adaboost machine learning algorithm based on Bosting idea, and constructed a fluid type classifier based on decision tree as the base learner, with an average accuracy of 86.5%. This to some extent proves the potential of intelligent algorithms in improving reservoir evaluation efficiency and interpretation accuracy. However, due to the lack of targeted algorithms for tight sandstone reservoir fluid identification, the final accuracy of these algorithms is not high. Therefore, some algorithms for strong heterogeneity of tight sandstone reservoirs have been developed. Most of the ideas for solving strong heterogeneity are parallel integration, which learns the characteristics of different areas in strong heterogeneous reservoirs by establishing multiple expert classifiers, so that the whole model can perfectly master the reservoir fluid characteristics. Wang F et al. (Wang F, Yang X M, Zhang Y H, et al. Application of multi-algorithm collaborative classification in tight sandstone fluid identification[J]. Progress in Geophysics, 2015, 30(6): 2785-2792.) combined Bayesian algorithm, K-nearest neighbor algorithm, generalized neural network algorithm, principal component analysis algorithm, and support vector machine algorithm into a multi-algorithm collaborative classifier for fluid identification. Tan M et al. (Tan M, Bai Y, Zhang H, et al. Fluid typing in tight sandstone from wireline logs using classification committee machine[J]. Fuel, 2020, 271: 117601.) and Bai Y et al. (Bai Y, Tan M, Xiao C W, et al. Dynamic classification committee machine for fluid identification in tight sandstone gas reservoirs using wireline logging data[J]. Chinese Journal of Geophysics, 2021, 64(5): 1745-1758.) successively constructed static and dynamic committee classification machines. The latter provided more ordered training samples for each expert classifier through clustering algorithm to improve the specificity of each expert classifier. This approach weakens the data bias problem caused by strong heterogeneity to some extent by building weak classifiers in parallel and independently, and reduces the generalization error of strong classifiers. However, its performance is highly dependent on the performance of the weak classifiers integrated. If the weak classifiers cannot produce effective identification results, it is impossible to simply integrate a high-precision strong classifier. Therefore, under the condition of sufficient sample size, neural network is a powerful tool for fitting complex nonlinear mapping relationships.Yi et al. (Yi J, Zheng Z J, Wen H C, et al. A method for fluid identification in tight sandstone reservoirs based on deep learning 2022-02-01.) extracted high-dimensional nonlinear logging features through a convolutional neural network, used a bidirectional long short-term memory neural network to further learn the multi-scale features of the logging data, and improved the efficiency of fluid identification. However, in actual work areas, the logging data used for fluid type identification is often limited and unbalanced. In order to deal with this imbalance, fluid identification methods combined with sampling algorithms have been proposed. Yi et al. (Yi J, Zhong W X, Zheng Z J, et al. Integrated model for fluid identification in tight sandstone reservoirs based on improved ADASYN data enhancement [J]. Progress in Geophysics, : 1-15.) further proposed an improved ADASYN data sampling method to eliminate the unbalanced characteristics of logging samples, and then obtained the fluid identification results through an adaptive weight integrated model. Mei He et al. (He M, Gu H, Wan H. Log interpretation for lithology and fluid identification using deep neural network combined with MAHAKIL in a tight sandstone reservoir [J]. Journal of Petroleum Science and Engineering, 2020, 194: 107498.) combined the MAHAKIL oversampling method with the DNN network to solve the data imbalance classification problem. Undoubtedly, oversampling the minority class can strengthen the characteristics of the minority class and reduce the classification difficulty. However, for water-rich tight sandstone, its tightness and high water content lead to high similarity between different classes of samples; oversampling will strengthen these highly overlapping and similar features, making it more difficult to identify easily confused samples. Therefore, the key to improving the accuracy of fluid identification in water-rich tight sandstone is to improve the classification accuracy of single models for different classes of fluids with highly similar features.
[0004] The present application aims at complex logging response characteristics, and on the basis of logging data, the fluid identification problem of water-rich tight sandstone is analogous to the class imbalance classification problem. According to the idea of solving the class imbalance classification problem, a cost-sensitive gradient boosting decision tree (PC-SC-GBDT) based on parameter construction is proposed. This method first constructs logging composite parameters at the data level, and uses logging curves, composite parameters and physical parameters as input features of the model. Secondly, at the algorithm level, the sample weight update method in the gradient boosting decision tree is optimized, and the recall rate index is introduced into the weight update formula, and different weight update methods are given to training samples of different categories. Finally, through various optimization algorithms, the hyperparameters of the model are optimized to obtain the best performance model under the optimal hyperparameters. SUMMARY
[0005] The present application mainly overcomes the deficiencies in the prior art, and provides a high water cut tight sandstone reservoir fluid type discrimination method based on a cost-sensitive gradient boosting machine.
[0006] To achieve the above technical purposes, the present application adopts the following technical scheme:
[0007] 1. A high water cut tight sandstone reservoir fluid type discrimination method based on a cost-sensitive gradient boosting machine, characterized by comprising the following steps:
[0008] Step 1, establishing a logging data set X suitable for artificial intelligence model training:
[0009] (1) selecting a wells from a target layer of a certain block, a is a positive integer, each well contains 7 logging curves of natural gamma, natural potential, compensated neutron, compensated density, acoustic time difference, formation resistivity, and flushed zone resistivity, and 4 physical parameters of permeability, porosity, shale content, and water saturation;
[0010] (2) based on the logging curves and the physical parameters, 13 composite parameters are constructed, and the logging curves, the physical parameters and the composite parameters are collectively used as input characteristic parameters;
[0011] Table 1 Composite parameter table
[0012]
[0013] Wherein, ρ b is the compensated density logging value, g / cm 3 ; ρ ma is the rock skeleton density, ρ ma = 2.65 (g / cm 3 ), g / cm 3 ; H is the reservoir effective thickness, m; PERM is the reservoir permeability calculation value, mD; Δt is the acoustic time difference logging value, μm / s; S g is the reservoir gas saturation, %; Δt ma is the acoustic time difference value of the rock skeleton, μm / s, taking the acoustic time difference value when the reservoir lithology is pure; Δt sh is the acoustic time difference value of the shale skeleton, μm / s, taking the acoustic time difference value when the shale content is extremely high; Δt f is the acoustic time difference value of the fluid, μm / s; is the hydrogen index of the fluid, % / %; ρ f is the fluid density, g / cm 3 ; RT is the formation resistivity, Q·m; RXO is the flushed zone resistivity, Q·m; V shThe volume percentage content of shale in the formation, %, is calculated as follows:
[0014] V sh =(2 2·ΔGR -1) / (2 2 -1)
[0015] △GR=(GR-GR min ) / (GR max -GR min )
[0016] Wherein, GR is the natural gamma ray logging value, API, GR min is the natural gamma ray value of pure sandstone, GR max is the natural gamma ray value of pure mudstone; POR is the calculated value of reservoir porosity, %, and the calculation formula is as follows:
[0017]
[0018] V ma is the volume percentage content of the skeleton mineral (pure sandstone), %, and the calculation formula is as follows:
[0019] V ma =1-POR-V sh
[0020] is the shale-free content of the neutron porosity, % obtained according to the compensated neutron logging value, and the calculation formula is as follows:
[0021]
[0022] CNL is the compensated neutron logging value, v / v; is the hydrogen index of the shale skeleton; is the shale-free content of the density porosity, % obtained according to the compensated density logging value, and the calculation formula is as follows:
[0023]
[0024] ρ sh is the shale skeleton density, g / cm 3 , and ρ sh =2.6 (g / cm 3 ) is taken according to the geological conditions of the work area;
[0025] (3) The Z-score standardization method is adopted to make it conform to the normal distribution, wherein the SP curve adopts local standardization, and other curves adopt global standardization;
[0026] (4) Remove the corresponding non-reservoir section, mudstone interlayer, reservoir section top and bottom interface and data missing section in the selected curve;
[0027] (5) Sampling each well section according to a fixed sampling point number Q, Q is a positive integer, so that different thickness reservoir well sections have different resolutions, as the original curve data set X;
[0028] Step 2, construct a cost-sensitive gradient boosting machine, the specific steps are as follows:
[0029] (1) Set the training data set Define the number of basic learners as T∈(1, +∞), T is a positive integer, t∈[1, T] is the iteration number, N is the sample number, and L is the loss function;
[0030] (2) Let t=1, initialize the sample weight Initialize the basic classifier h t (D t (i)x;α t ), at this time the optimization objective function of the model is
[0031] H t (x)=H0(x)+h t (D t (i)x;α t )
[0032]
[0033] Where, α t is the model parameter of the basic classifier h t (D t (i)x;α t ), H0(x) is the total output of the model at this time, and ρ is the weight parameter of the basic classifier;
[0034] (3) When t≤T, calculate the residual of this round, update the sample weight, and construct the next round of basic classifier to fit the residual of this round:
[0035] a) Calculate the residual g t of H t (x);
[0036]
[0037] b) Divide the input sample U train of the current classifier into three categories, U correct is the correctly classified sample, U wrong is the incorrectly classified sample, and U pw is the minimum recall rate category of the incorrectly classified sample, which satisfies U train =U correct ∪U wrong ;
[0038] c) Weighting the samples belonging to U correct , U wrong and U pw respectively:
[0039]
[0040]
[0041]
[0042] where recall is the current global recall, TP is the number of true positives, representing the case that positive class is correctly classified, FN is the number of false negatives, representing the case that positive class is classified as negative class, b, c are parameters of the weighting function, card() represents the number of elements in a set;
[0043] d) Using the following formula to ensure that the sample weight is valid, does not produce zero value, null value and infinity.
[0044]
[0045] e) Generating a new base classifier h t (D t (i); a t ) to fit the residual g t of the last classifier,
[0046] where,
[0047]
[0048] b t is the weight of the classifier h t (D t x; a t ) when fitting the residual g t , D t is the updated sample weight;
[0049] f) Updating the model:
[0050] H t (x) = H t-1 (x) + p t h t (D t (i) x; a t )
[0051]
[0052] where p t is the new base classifier h t (D t(i) x; a t ) of the weight;
[0053] g) when t = T, output the model
[0054] Step 3, input the training set into the prediction model constructed in step 2 for training, and use the sparrow search algorithm for hyperparameter optimization, select the model with the highest accuracy in the training optimization process as the final model, and use the cross entropy loss function as the loss function L, whose calculation formula is as follows:
[0055]
[0056] Wherein, Z is the number of categories; G is the number of observation samples; y ic is a symbol function, taking 0 or 1, if the real category of sample i is c, then y ic is 1, otherwise y ic is 0; p ic is the probability that observation sample i belongs to category c; the specific steps of sparrow search algorithm are as follows:
[0057] (1) let matrix be the position matrix of sparrow, be the position of the i-th sparrow on the j-th hyperparameter, be the fitness function under the current position, t ∈ T is the number of iterations, T is a positive integer, and Q is the best fitness queue;
[0058] (2) update the position of the producer:
[0059]
[0060] Wherein, ti is the number of iterations, Q is a random number subject to normal distribution, a ∈ 0, 1 is a random number, is a d-dimensional matrix with elements being 1, W2 ∈ [0, 1] is an alarm value, and ST ∈ [0.5, 1] is a safety threshold; when the alarm value is greater than the safety threshold, it means that the predator is found, and all sparrows need to update the position and fly to a safer area; when the alarm value is less than the safety threshold, the producer continues to search for food;
[0061] (3) update the position of the explorer:
[0062]
[0063] Wherein, is the best position occupied by the producer at t+1 iteration, is the global worst position at the t-th iteration, A 1×d is a d-dimensional matrix composed of elements 1 and -1, A+ = A T (AA T ) -1 ; indicates that the i-th explorer sparrow is in a hungry state and needs to fly farther to find food, while the other explorers choose to be close to the producer or to seize the position of the producer;
[0064] (4) the global optimal position of the t-th iteration is taken as the center of the sparrow population, and the global optimal fitness of the t-th iteration corresponding thereto is f best When the fitness of the i-th sparrow at the t-th iteration is greater than f best , it means that the sparrow i is at the edge of the sparrow population, and they need to move closer to the sparrow occupying the global optimal position; When f wa , it means that the sparrow at the population center feels danger and needs to move closer to other sparrows, which is represented by the following formula:
[0065]
[0066] Wherein, beta is a step control parameter, and is a random number normal distribution with a mean of 0 and a variance of 1;K is a random number in [-1, 1];And epsilon is the smallest constant to avoid zero division error;
[0067] (5) When the variance of the elements in the queue Q is less than the threshold value omega, output the global optimal position corresponding to the best fitness in the queue Q;
[0068] (6) Substitute the obtained optimal hyperparameter combination into the model to obtain the final model;
[0069] Step 4: input the test set into the trained model to identify the reservoir fluid type and obtain the prediction result.
[0070] The innovation of the present application is as follows:
[0071] (1) The fluid identification problem of water-rich dense sandstone reservoir is analogous to the data imbalance classification problem, which provides a new idea for fluid identification under complex gas-water distribution conditions.
[0072] (2) Thirteen fluid-sensitive logging composite parameters are constructed as prior information of the model, which amplifies the fluid logging response characteristics and optimizes the dimension characteristics of the input data.
[0073] (3) The cost-sensitive gradient boosting model solves the problem of high dispersion and high similarity of logging characteristics of gas layers, gas-water layers and water layers in fluid identification work.
[0074] Beneficial effects:
[0075] Compared with the prior art, the present application has the following beneficial effects:
[0076] (1) The fluid-sensitive logging composite parameter is constructed according to experience, prior knowledge of the model is strengthened, and the purpose of mining and amplifying the fluid logging response is achieved while increasing the input feature dimension.
[0077] (2) The cost-sensitive gradient boosting decision tree is constructed, the recall rate index is introduced into the weight updating formula of the gradient boosting decision tree, different weight updating rules are given to the correctly classified samples, the misclassified samples and the misclassified confused class samples according to the classification results of each round. The sample weight of the correctly classified sample is a value less than 1 and greater than 0, which reduces the attention of the model to the correctly classified sample. The misclassified sample is weighted according to the overall recall rate of the model, and the misclassified confused class sample is weighted again according to the class recall rate of the class, which increases the attention of the model to the inter-class distinction of the confused sample.
[0078] (3) By comparing the performance of different optimization algorithms in the above model, the sparrow search algorithm is finally used to find the optimal hyperparameters of the constructed model, which can obtain the best performance of the classification model. BRIEF DESCRIPTION OF DRAWINGS
[0079] Figure 1 It is a structural flowchart of the fluid type discrimination method for high water cut tight sandstone reservoir based on the cost-sensitive gradient boosting machine, which represents the model structure and the process of fluid discrimination using the method. The left half is a schematic diagram of the model optimization algorithm, the upper right part is a schematic diagram of data preprocessing, and the lower right part is a schematic diagram of modeling;
[0080] Figure 2 It is a flowchart of the optimization algorithm, which shows the principle and process of the optimization part of the model;
[0081] Figure 3 It is a schematic diagram of fluid discrimination of the well section of a blind well in the study area, and the prediction result of the fluid identification model is consistent with the gas logging result. DETAILED DESCRIPTION
[0082] In order to make the purpose, technical scheme and advantages of the present application clearer and more apparent, the present application will be further described in detail below with examples. It should be understood that the specific examples described herein are only used to explain the present application and do not limit the present application.
[0083] Example:
[0084] A high water cut tight sandstone reservoir productivity intelligent prediction method based on ensemble learning, characterized by comprising the following steps:
[0085] 1. A cost-sensitive gradient boosting machine-based fluid type discrimination method for high water cut tight sandstone reservoirs, characterized by comprising the following steps:
[0086] Step 1: Establishing a logging data set X suitable for artificial intelligence model training:
[0087] (1) Selecting a wells from the target layer of a certain block, a = 183, each well containing 7 logging curves of natural gamma, spontaneous potential, compensated neutron, compensated density, acoustic time difference, formation resistivity, and flushed zone resistivity, and 4 physical parameters of permeability, porosity, shale content, and water saturation;
[0088] (2) Based on the logging curves and physical parameters, composite parameters X1, X2, X3, X4, X5, X6, X7, X8, X9, X10, AK, DT, and S are constructed wa The logging curves, physical parameters, reservoir effective thickness, and composite parameters are collectively used as input characteristic parameters;
[0089] (3) Using Z-score standardization method to make it conform to normal distribution, wherein the SP curve adopts local standardization and other curves adopt global standardization;
[0090] (4) Removing the corresponding non-reservoir section, mudstone interlayer, reservoir section top and bottom interface, and data missing section in the selected curve;
[0091] (5) Sampling each well section according to a fixed sampling point number Q, Q = 30 is a positive integer, so that different thickness of reservoir well sections have different resolutions, as the original curve data set X, the dimension of X is (183, 25, 30);
[0092] Step 2: Constructing a cost-sensitive gradient boosting machine, the specific steps are as follows:
[0093] (1) Set the training data set Define the number of basic learners T = 50, t ∈ [1, T] is the iteration number, the sample number N = 183, L is the loss function,
[0094]
[0095] Where M = 4, indicating that there are 4 types of fluid, N is the number of observation samples; y ic is a sign function, taking value 0 or 1, if the true class of sample i is c, then y ic takes 1, otherwise y ic takes 0; p ic is the probability that observation sample i belongs to class c;
[0096] (2) Let t = 1, initialize the sample weight Initialize the base classifier h t (D t (i) x; a t ), the optimization objective function of the model is
[0097] H t (x) = H0(x) + h t (D t (i) x; a t )
[0098]
[0099] where a t is the model parameter of the base classifier h t (D t (i) x; a t ), and H0(x) is the total output of the model at this time;
[0100] (3) When t≤T, calculate the residual of this round, update the sample weight, and construct the base classifier of the next round to fit the residual of this round:
[0101] a) Calculate the residual g t of H t (x);
[0102] b) Divide the input sample U train of the current classifier into three categories, U correct , U wrong , and U pw , which satisfy U train = U correct ∪U wrong ;
[0103] c) Weight the samples belonging to U correct , U wrong , and U pw , respectively:
[0104] d) Use the following formula to ensure that the sample weight is valid, and does not produce zero, null, and infinity.
[0105]
[0106] e) Generate a new base classifier h t (D t (i); a t ) to fit the residual g t of the previous classifier,
[0107] f) Update the model;
[0108] g) When t = T, output the model
[0109] Step 3, input the training set into the prediction model constructed in step 2 for training, and use the sparrow search algorithm for hyperparameter optimization, and select the model with the highest accuracy in the training optimization process as the final model; the specific steps of the sparrow search algorithm are as follows:
[0110] (1) Let the matrix be the position matrix of the sparrow, be the position of the i-th sparrow on the j-th hyperparameter, be the fitness function under the current position, t e T is the number of algorithm iterations, T is a positive integer, T = 50;
[0111] (2) Update the position of the producer:
[0112] (3) Update the position of the explorer:
[0113] (4) Take the global optimal position of the t-th iteration as the center of the sparrow population, and the global optimal fitness of the t-th iteration corresponding thereto is f best , and the position of each sparrow is updated using the following formula:
[0114]
[0115] Where, β is a step control parameter, and is a random number normal distribution with a mean of 0 and a variance of 1; K e [-1, 1] is a random number. ε is the smallest constant to avoid zero division error
[0116] (5) Threshold ω = 0.01, when the variance of the elements in the queue Q is less than ω, output the global optimal position corresponding to the best fitness in the queue Q;
[0117] (6) Substitute the optimal hyperparameter combination obtained into the model to obtain the final model;
[0118] Step 4: input the test set into the trained model for reservoir fluid type identification to obtain the prediction result;
[0119] Step 5, compare the fluid discrimination effects of different algorithms through the confusion matrix and global accuracy, including gradient boosting tree, CatBoost, static classification committee machine, random forest, decision tree, support vector machine, adaptive boosting machine and the method proposed in the application, and the comparison results are shown in Table 2.
[0120] Table 2 Comparison of classification performance of each model on the fluid discrimination problem of water-rich dense sandstone
[0121]
[0122]
[0123] The above description is not intended to limit the present application in any form, although the present application has been disclosed by the above examples, however, not to define the present application, any skilled person in the art, without departing from the technical solution of the present application, can utilize the disclosed technical content to make changes or modifications as equivalent examples of equivalent changes, but any simple modification, equivalent change and modification made to the above examples according to the technical essence of the present application, without departing from the technical solution of the present application, still belongs to the scope of the technical solution of the present application.
Claims
1. A cost-sensitive learning based method for fluid discrimination in water-rich tight sand reservoirs, characterized in that Comprise the following steps: Step 1, establish a set of well logging data suitable for artificial intelligence model training X: (1) select a wells from a target layer of a certain block, a is a positive integer, each well contains 7 logging curves of natural gamma, spontaneous potential, compensated neutron, compensated density, acoustic time difference, formation resistivity, flush zone resistivity, 4 physical parameters of permeability, porosity, shale content, water saturation; (2) based on the logging curves and physical parameters, 13 composite parameters are constructed, and the logging curves, physical parameters and composite parameters are used as input feature parameters; wherein p b is the density log value, g / cm 3 ; p ma is the rock matrix density, p ma = 2.65 (g / cm 3 ), g / cm 3 ; H is the reservoir net thickness, m; PERM is the reservoir permeability calculated value, mD; At is the acoustic traveltime log value, pm / s; S g is the reservoir gas saturation, %; At ma is the acoustic traveltime value of the rock matrix, pm / s, taking the acoustic traveltime value when the reservoir lithology is pure; At sh is the acoustic traveltime value of the shale matrix, pm / s, taking the acoustic traveltime value when the shale content is extremely high; At f is the acoustic traveltime value of the fluid, pm / s; is the hydrogen index of the fluid, % / %; p f is the fluid density, g / cm 3 ; RT is the formation resistivity, W-m; RXO is the flushed zone resistivity, W-m; V sh is the volume percent content of formation shale, %, and the calculation formula is as follows: V sh =(2 2·△GR -1) / (2 2 -1) AGR = (GR - GR min ) / (GR max -GR min ) Where, GR is natural gamma ray logging value, API, GR min is pure sandstone natural gamma ray value, GR max is pure mudstone natural gamma ray value, according to the geological conditions of the work area, take GR min = 28 (API), GR max = 140 (API); POR is the calculated value of reservoir porosity, %, and the calculation formula is as follows: V ma The volume percentage content, %, of the matrix mineral (pure sandstone) is calculated as follows: V ma = 1 - POR - V sh The neutron porosity, % free of shale content, is obtained from the compensated neutron log value and is calculated as follows: CNL is compensated neutron log value, v / v; is hydrogen index of shale matrix; is density porosity, % without shale content, calculated from compensated density log value, according to the following formula: ρ sh ρ 3 , g / cm sh = 2.6 (g / cm 3 ) (3) Z-score standardization method is adopted to make it meet the normal distribution, wherein SP curve adopts local standardization, and other curves adopt global standardization; (4) remove the corresponding non reservoir section, mudstone interlayer, reservoir section top and bottom interface and data missing section in the selected curve; (5) sample each well section according to fixed sampling point number Q, Q is a positive integer, so that different thickness of reservoir well section has different resolution, as the original curve data set X; Step 2, build a cost sensitive gradient boosting machine, the specific steps are as follows: (1) Set training data set U: (x i ,y i )|i∈[1,N], Define the number of base learners as T∈(1,+∞), T is a positive integer, t∈[1,T] is the iteration number, N is the sample number, and L is the loss function. (2) Let t = 1, initialize sample weights Initialize base classifier h t (D t (i) x; a t ), The optimization objective function of the model at this time is H t (x) = H0(x) + h t (D t (i) x; a t ) where α t is a model parameter of the base classifier h t (D t (i) x; a t is the total output of the model at this time, and p is a weight parameter of the base classifier. (3) when t ≤ T, calculate the residual of this round, update the sample weight, and build the next round of basic classifier to fit the residual of this round: a) calculating H t the residual g of (x) t ; b) Input sample U of the current classifier train Divided into three categories, U Correct For correctly classified samples, U wrong For misclassified samples, U pw Samples misclassified for the minimum recall category satisfy U train =U correct ∪U wrong ; c) weighting the samples belonging to U correct , U wrong and U pw , respectively: wherein recall is the current global recall rate, TP is the number of true positives, representing the case of correct classification of positive classes, FN is the number of false negatives, representing the case of classification of positive classes as negative classes, b and c are parameters of the weighting function, and card(*) represents the number of elements in a set; d) use the following formula to ensure that the sample weight is effective, and does not produce zero value, null value and infinity, e) generating a new base classifier h t (D t (i) ; a t ) fitting the residual g of the previous classifier t wherein, β t is the fitting residual g t is the classifier h t (D t x; a t ) of the weight; D t is the updated sample weight; f) update the model: H t (x) = H t-1 (x) + p t h t (D t (i) x; a t ) wherein p t is a new base classifier h t (D t (i) x; a t ) is the weight of the base classifier g) when t = T, output model Step 3, input the training set into the prediction model constructed in step 2 for training, and use sparrow search algorithm for hyperparameter optimization, select the model with the highest accuracy in the training optimization process as the final model, and the loss function L uses cross entropy loss function, and its calculation formula is as follows: where Z is the number of classes; G is the number of observation samples; y ic is the indicator function, taking value 0 or 1, y ic = 1 if the true class of sample i is c, and y ic = 0 otherwise; p ic is the probability that observation sample i belongs to class c; the specific steps of the sparrow search algorithm are as follows: (1) Set matrix is the position matrix of the sparrow, is the position of the ith sparrow on the jth hyperparameter, is the fitness function under the current position, t ∈ T is the number of algorithm iterations, T is a positive integer, and Q is the best fitness queue; (2) update the position of the producer: Wherein, t is the iteration number, Q is a random number subject to normal distribution, and a is a random number in 0, 1, is a d-dimensional matrix with elements being 1, W2 is an alarm value in [0, 1], and ST is a safety threshold in [0.5, 1]. (3) update the position of the explorer: wherein, is the best position occupied by the producer at iteration t+1, is the global worst position at iteration t, A 1×d is a d-dimensional matrix consisting of elements 1 and -1, A + = A T (AA T ) -1 ; (4) The global optimal position at the t-th iteration is The corresponding global optimal fitness at the t-th iteration is f. best ,Will with f best Add to queue Q; the fitness of the i-th sparrow in the t-th iteration. Greater than f best When the i-th sparrow moves closer to the sparrow that occupies the globally optimal position; When the i-th sparrow needs to move closer to other sparrows, it is represented by the following formula: Wherein, β as step length control parameter, is a random number normal distribution with mean value of 0 and variance of 1; K ∈ [-1, 1] is a random number, and ε is the smallest constant to avoid zero division error; (5) when the variance of the elements in the queue Q is less than the threshold ω, output the global optimal position corresponding to the best fitness in the queue Q; (6) combine the obtained optimal hyperparameters into the model to obtain the final model; Step 4: input the test set into the trained model to identify the reservoir fluid type, and obtain the prediction result.