Intelligent oil well paraffin precipitation prediction method based on LSTM (Long Short Term Memory)

Through the intelligent prediction method of oil well tungsten wax based on LSTM, combined with the Gray Wolf algorithm and SHAP value analysis, the accuracy and robustness of oil well tungsten wax prediction are solved, and efficient oil well management and prevention of wax are achieved.

CN120020767APending Publication Date: 2025-05-20CHINA PETROLEUM & CHEMICAL CORP +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202311538026.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

The prior art is difficult to accurately predict the trend of oil well wax tungsten, resulting in lower or shutdown of oil well production, and traditional RNNs are prone to gradient disappearance or explosion when dealing with long sequence problems.

Method used

The LSTM-based intelligent prediction method of oil well tungsten wax is adopted to obtain the oil well timing data sequence, perform data preprocessing and feature selection, build and optimize the LSTM network, optimize the model hyperparameters using the Gray Wolf algorithm, and interpret the model results through SHAP value analysis.

Benefits of technology

High accuracy and robust oil well wax prediction is achieved, and the traditional method overcomes the shortcomings of the relationship and timing of oil well wax and multi-characteristics, and provides a method with high prediction accuracy and interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020767A_ABST
    Figure CN120020767A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent oil well paraffin precipitation prediction method based on LSTM. The method comprises the steps that 1, an oil well time sequence data sequence used for training a model is acquired; 2, performing data preprocessing operation to obtain a training data set; 3, carrying out feature selection on the preprocessed data, and selecting features which have great influence on the area of an oil well indicator diagram; 4, constructing and optimizing an oil well paraffin precipitation prediction model based on the LSTM network; and step 5, performing prediction by using the optimized LSTM model, and performing anti-normalization operation on the predicted data to obtain predicted real data. The intelligent oil well paraffin precipitation prediction method based on the LSTM overcomes the problem that correlation between oil well paraffin precipitation and many features and the time sequence of the oil well paraffin precipitation are not considered in a traditional method, meanwhile, the LSTM model is optimized through the grey wolf algorithm, the oil well paraffin precipitation prediction method is high in prediction precision and robustness, and the method is suitable for being used for oil well paraffin precipitation prediction. The defects in the field in the prior art are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of deep learning and petroleum engineering, and particularly to an intelligent prediction method for wax deposition in oil wells based on LSTM. Background Art

[0002] Wax deposition in oil wells brings a large amount of work to daily management, increasing the possibility and probability of downhole accidents. After wax deposition in oil wells, the inner diameter of the oil outlet channel gradually decreases, increasing the oil flow resistance, reducing the productivity of the oil well, and even blocking the oil flow channel, resulting in a reduction or shutdown of the oil well production. Therefore, to address this problem and prevent the serious impact caused by wax deposition in oil wells, we hope to predict the wax deposition trend before wax deposition in oil wells and take measures to remove wax to avoid wax deposition.

[0003] The wax deposition data of oil wells is a time series data sequence. With the development of the field of artificial intelligence, it has become quite common to use machine learning methods to handle such problems. However, compared with traditional machine learning, deep learning is more intelligent and more accurate in dealing with big data problems. In deep learning, the RNN method is often used for time series prediction. However, traditional RNN may encounter gradient disappearance or explosion in dealing with long sequence problems, while the gate structure of LSTM can effectively alleviate this problem.

[0004] LSTM (Long Short-Term Memory) is a time-recurrent neural network, suitable for processing and predicting important events with relatively long intervals and delays in time series, and applicable to the analysis and fitting of time series. The LSTM algorithm has been applied in various fields of science and technology. And the prediction of wax deposition in oil wells is exactly a time series prediction problem, so the present invention uses it as a tool for predicting wax deposition in oil wells.

[0005] In the Chinese patent application with the application number: CN201811140212.4, it relates to a method for predicting the dynamic wax removal cycle of oil wells in offshore oilfields. The steps are as follows: According to the liquid production volume, string characteristics, and electric pump unit parameters during the normal production of the given target well, and initially assuming the wax removal cycle time, iteratively calculate the temperature profile of the wellbore along the way from the wellhead downwards, the wax deposition amount along the wellbore, and the fluid pressure of the wellbore at the pump discharge outlet within this cycle, and the fluid pressure of the wellbore from the bottom to the pump suction inlet within this cycle. Combining with the electric pump characteristic curve, the liquid production volume flowing through the electric pump can be inversely deduced until the initial liquid production volume under the condition of wax deposition influence is close to the liquid production volume inversely deduced according to the electric pump characteristic curve. Finally, iteratively calculate the ratio of the current effective liquid production volume to the initial liquid production volume. If it meets the initial set critical interference production percentage target, the calculation ends. The method of this invention solves the problem of accurately predicting the wax removal cycle of wax deposition wells with electric pumps in offshore oilfields and can guide the formulation of the production system for wax deposition wells in the oilfield site.

[0006] In the Chinese patent application with the application number: CN201910874803.2, it involves a method for planning the hot washing times and dates of oil wells, including the following steps: Step 1) Predict the hot washing dates of multiple wax-depositing oil wells on the last day of the next year and the hot washing cycles of multiple wax-depositing oil wells; Step 2) Obtain the total wax removal times of all wax-depositing oil wells based on the hot washing dates of multiple wax-depositing oil wells on the last day of the next year, the last hot washing dates of multiple wax-depositing oil wells, and the hot washing cycles of multiple wax-depositing oil wells; Step 3) Obtain the hot washing times of all oil wells in the next year according to the total wax removal times of all wax-depositing oil wells obtained in Step 2 and the wax removal times during pump inspection. Combining the hot washing cycles of wax-depositing oil wells and the last hot washing dates of each wax-depositing oil well, the specific hot washing dates of each wax-depositing oil well can be obtained, and the planning of the hot washing times and dates of all wax-depositing oil wells is completed.

[0007] In the Chinese patent application with the application number: CN201911043103.5, it involves a method for optimizing the hot washing wax removal method and parameters of oil wells, including: Step 1, Determine the structural parameters of the hot washing well and calculate the total heat transfer coefficient; Step 2, Based on the total heat transfer coefficient, study the distribution law of the hot washing temperature field and establish a mathematical model of temperature distribution; Step 3, Through experimental tests, analyze the characteristics of wax precipitation with water in different regions of the oil well, and calculate the wax precipitation model of the oil well according to the oil layer temperature distribution and the physical property parameters of the oil sample; Step 4, Substitute the wax deposition parameters into the heat transfer model and calculate the heat required for hot washing; Step 5, According to different hot washing structures, select the most optimized hot washing plan by calculating the conditions of the required hot washing fluid. This method for optimizing the hot washing wax removal method and parameters of oil wells effectively integrates the wax deposition degree, wax removal technology, and operation management, improves the treatment effect of wax-depositing wells, effectively improves the operation rate of wax-depositing wells, can predict the cleaning cycle, and is convenient and simple to use and easy to promote.

[0008] The above existing technologies are all quite different from the present invention and fail to solve the technical problems we want to solve. Therefore, we have invented a new intelligent prediction method for oil well wax deposition based on LSTM. Summary of the Invention

[0009] The object of the present invention is to provide an intelligent prediction method for oil well wax deposition based on LSTM with high reliability, good accuracy, and better prediction effect.

[0010] The object of the present invention can be achieved by the following technical measures: An intelligent prediction method for oil well wax deposition based on LSTM, which includes:

[0011] Step 1, Obtain the time series data sequence of oil wells for training the model;

[0012] Step 2, Perform data preprocessing operations to obtain the training data set;

[0013] Step 3, perform feature selection on the pre-processed data, and select features that have a greater impact on the area of ​​the oil well indicator diagram;

[0014] Step 4: Build and optimize the oil well wax deposition prediction model based on LSTM network;

[0015] Step 5: Use the optimized LSTM model to make predictions, and perform denormalization on the predicted data to obtain the predicted real data.

[0016] The purpose of the present invention can also be achieved through the following technical measures:

[0017] In step 1, obtain oil well data including oil well dynamometer area and historical production data with various characteristics from the oil field database; write the acquired historical production data into a spreadsheet and save it locally.

[0018] Step 2 includes:

[0019] Step 21. Read the spreadsheet in step 1 to obtain the time series data sequence;

[0020] Step 22. For the time series data obtained in step 21, delete missing data and irrelevant data, remove outliers, and perform min-max standardization to convert the data values ​​to the [0,1] interval;

[0021] Step 23. For the processed data, use most of the data as the training set and a small part of the data as the test set, and ensure that the training set is normal data.

[0022] In step 22, the formula for min-max normalization is as follows:

[0023]

[0024] Where min(x) is the minimum value in the data, and max(x) is the maximum value in the data.

[0025] In step 23, for the processed data, 80% of the data is used as a training set and 20% of the data is used as a test set.

[0026] Step 3 includes:

[0027] Step 31. Build an interpretable machine learning model;

[0028] Step 32. Define the various parameters in the machine learning model constructed in step 31, and use the data obtained in step 2 to train the model, that is, use the features other than the oil well indicator diagram data as input, and use the oil well indicator diagram data as labels to train the model to obtain a trained machine learning model;

[0029] Step 33. For the machine learning model trained in Step 32, use the SHAP tool to interpret the model prediction results.

[0030] In Step 33, the SHAP value formula is as follows:

[0031] Y i = y base + f(x i1 ) + f(x i2 ) + … + f(x ij )

[0032] where Y i is the predicted value of sample i, y base is the baseline value, and f(x ij ) is the contribution value of the j features of sample i.

[0033] Step 33 specifically includes:

[0034] Step 331. Calculate the baseline value. The prediction result of the machine learning model with all features except the oil well dynamometer card set to the average value is used as the baseline, that is, the average prediction result of the model on the training data is used as the baseline value;

[0035] Step 332. Randomly combine the features except the oil well dynamometer card. For each combination, calculate the contribution value of each feature to the prediction result, and aggregate the contribution values to obtain the SHAP value;

[0036] Step 333. Visualize the SHAP values obtained in Step 332, and select the features that have a greater impact on the prediction result, that is, on the dynamometer card data, according to the high and low of the SHAP values.

[0037] Step 4 includes:

[0038] Step 41. According to the features selected in Step 3 that have a greater impact on the area of the oil well dynamometer card, construct an LSTM network with the selected features as the input and the area of the oil well dynamometer card as the label;

[0039] Step 42. For the oil well paraffin deposition prediction model constructed in Step 41, use the Grey Wolf Optimization (GWO) algorithm to optimize the model hyperparameters, specifically, optimize the number of neurons in the first and second layers of the model neural network and the number of iterations.

[0040] In Step 41, when constructing the LSTM network, the loss function is selected as MSE, and its formula is as follows:

[0041]

[0042] where Yi is the true value of sample i, is its corresponding predicted value.

[0043] In step 42, the principle of the Grey Wolf Optimizer (GWO) is as follows:

[0044] In the GWO, the wolves, which represent the combination of hyperparameters, update their positions according to the three wolves with the best positions (α, β, γ). The formula for updating the position D of a wolf from iteration t to t + 1 is as follows:

[0045] D = CX p (t) - X(t), X(t + 1) = X p (t) - AD

[0046] where Xp is the position of the prey, i.e., the positions of the α, β, and γ wolves, X is the position of the grey wolf, D is the distance between the grey wolf and the prey, and the calculation formulas for A and C are as follows:

[0047] A = 2ar 1 -a, C = 2r 2

[0048] where r1 and r2 are random numbers between 0 and 1, and a linearly decreases from 2 to 0 with the number of iterations;

[0049] Calculate D α , D β , D γ , and X(t + 1) α , X(t + 1) β , X(t + 1) γ Then, the position X(t + 1) that the wolf individual needs to adjust is:

[0050]

[0051] Step 42 specifically includes:

[0052] Step 421. Define the number of iterations, the number of wolves, and the setting range of the GWO. One wolf represents a combination of hyperparameters including the number of neurons in the first and second layers of a neural network and the number of iterations;

[0053] Step 422. Randomly initialize the wolf pack and calculate the fitness value of each wolf. The fitness value is the value of the loss function in step 41. The positions of the three grey wolves with the top fitness values are respectively used as the positions of the α, β, and γ grey wolves;

[0054] Step 423.3. Update the positions of the remaining grey wolves according to the positions of the α, β, and γ grey wolves obtained in step 421, and ensure that all grey wolves are within the set range;

[0055] Step 423.4. Recalculate the new α, β, and γ grey wolves based on the wolf pack updated in step 423.3, and repeat step 423.3;

[0056] Step 423.5. Determine whether the termination condition is reached. If so, end the algorithm and output the optimal grey wolf, i.e., the optimal model hyperparameter combination. If not, repeat Step 423.4.

[0057] In Step 5, transfer the optimal model hyperparameter combination obtained in Step 423.5, including the optimal number of neurons and the number of iterations, to the model constructed in Step 41, and use the optimized LSTM model for prediction. Perform inverse normalization on the predicted data, compare it with the true value, and visualize it.

[0058] Compared with the existing technologies, the intelligent wax deposition prediction method for oil wells based on LSTM in the present invention uses feature importance analysis to overcome the problems that the traditional methods do not consider the association between oil well wax deposition and many features and its temporality. At the same time, the grey wolf algorithm is used to optimize the LSTM model, providing an oil well wax deposition prediction method with higher prediction accuracy and strong robustness, and making up for the deficiencies of the existing technologies in this field.

[0059] The present invention combines SHAP value analysis and GWO algorithm. SHAP value analysis is used to reveal the importance of features to the prediction results, and then the GWO algorithm is used to optimize the LSTM model. This combination makes the model more interpretable and optimized. Through SHAP value analysis, not only can the influence degree of features on the prediction results be understood, but also the prediction process of the model can be explained, improving the interpretability of the model and making the prediction results more persuasive. Traditional optimization algorithms are prone to falling into local optimal solutions when facing complex problems, while the GWO algorithm can effectively search for the optimal solution, improving the optimization performance of the model and making the model more adaptable to complex oil well wax deposition prediction problems. The LSTM model is used to model time series data, considering the time dependence, enabling the model to better capture the trends and changing rules of oil well wax deposition. Description of the Drawings

[0060] Figure 1 It is a flowchart of a specific embodiment of the intelligent wax deposition prediction method for oil wells based on LSTM of the present invention;

[0061] Figure 2 It is a bar chart for explaining the feature importance in a specific embodiment of the present invention;

[0062] Figure 3 It is a comparison chart of the LSTM prediction result and the true value in Example 1 of a specific embodiment of the present invention;

[0063] Figure 4 It is a comparison chart of the prediction result after optimizing the LSTM in Example 1 and the true value in a specific embodiment of the present invention;

[0064] Figure 5In a specific embodiment of the present invention, it is a comparison graph between the predicted result after the optimization of Instance 2 LSTM and the true value. Detailed implementation manners

[0065] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0066] It should be noted that the terms used herein are only for describing specific implementation manners and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, and / or combinations thereof.

[0067] As Figure 1 shown, Figure 1 It is a flowchart of the intelligent wax deposition prediction method for oil wells based on LSTM according to the present invention. The intelligent wax deposition prediction method for oil wells based on LSTM includes:

[0068] Step 1, obtain the time series data sequence of the oil well for training the model. The specific steps are as follows:

[0069] Step 11. Obtain the oil well data from the oilfield database, including the area of the oil well dynamometer card and historical production data with various different characteristics;

[0070] Step 12. Write the data obtained in Step 11 into a spreadsheet and save it locally.

[0071] Step 2, perform data preprocessing operations to obtain a training data set, including the following steps:

[0072] Step 21. Read the table in Step 12 to obtain the time series data sequence;

[0073] Step 22. For the time series data sequence obtained in Step 21, delete the missing data and irrelevant data, eliminate the outliers, and perform min-max normalization processing, so as to convert the data values into the interval [0,1];

[0074] Step 23. For the processed data, use most of the data as the training set and a small part of the data as the test set, and ensure that the training set is all normal data.

[0075] In Step 22, the formula for min-max normalization processing is as follows:

[0076]

[0077] Where min(x) is the minimum value in the data and max(x) is the maximum value in the data.

[0078] In step 23, for the processed data, 80% of the data is used as the training set and 20% of the data is used as the test set.

[0079] Step 3: Use the SHAP tool to select features from the preprocessed data, and select the features that have a greater impact on the area of the dynamometer card of the oil well, including the following steps:

[0080] Step 31. Construct an interpretable machine learning model;

[0081] Step 32. Define the parameters in the machine learning model constructed in step 31, and use the data obtained in step 2 to train the model, that is, use the features other than the dynamometer card data of the oil well as the input and the dynamometer card data of the oil well as the label to train the model, and obtain a trained machine learning model;

[0082] Step 33. For the machine learning model trained in step 32, use the SHAP tool to interpret the model prediction results. The SHAP value formula is as follows:

[0083] Y i = y base + f(x i1 ) + f(x i2 ) + … + f(x ij )

[0084] Where Y i is the predicted value of sample i, y base is the baseline value, and f(x ij ) is the contribution value of the jth feature of sample i.

[0085] The specific process is as follows:

[0086] Step 331. Calculate the baseline value. In the present invention, the prediction result of the machine learning model when all features except the dynamometer card of the oil well are set to the average value is used as the baseline, that is, the average prediction result of the model on the training data is used as the baseline value.

[0087] Step 332. Randomly combine the features except the dynamometer card of the oil well. For each combination, calculate the contribution value of each feature to the prediction result, and aggregate the contribution values to obtain the SHAP value.

[0088] Step 333. Visualize the SHAP values obtained in step 332, and select the features that have a greater impact on the prediction result (i.e., on the dynamometer card data) according to the high and low of the SHAP values.

[0089] Step 4: Construct and optimize the wax deposition prediction model for oil wells. In this invention, the LSTM network is used as the wax deposition prediction model for oil wells, and the specific process is as follows:

[0090] Step 41. Based on the features selected in Step 3 that have a greater impact on the area of the oil well indicator diagram, use the selected features as inputs and the area of the oil well indicator diagram as the label to construct an LSTM network. In this invention, the loss function is selected as MSE, and its formula is as follows;

[0091]

[0092] where Yi is the true value of sample i, is its corresponding predicted value.

[0093] Step 42. For the wax deposition prediction model for oil wells constructed in Step 41, use the Grey Wolf Optimization (GWO) algorithm to optimize the hyperparameters of the model. Specifically, optimize the number of neurons in the first and second layers of the model neural network and the number of iterations. The principle of the Grey Wolf Optimization algorithm is as follows:

[0094] In the Grey Wolf Optimization (GWO) algorithm, the population of wolves (hyperparameter combinations) updates their positions according to the three wolves (α, β, γ) with the optimal positions. The formula for updating the position D of the wolves from iteration t to t + 1 is as follows:

[0095] D = CX p (t) - X(t), X(t + 1) = X p (t) - AD

[0096] where Xp is the position of the prey, i.e., the positions of the α, β, γ wolves, X is the position of the grey wolf, D is the distance between the grey wolf and the prey, and the calculation formulas for A and C are as follows:

[0097] A = 2ar 1 -a, C = 2r 2

[0098] where r1 and r2 are random numbers from 0 to 1, and a linearly decreases from 2 to 0 with the number of iterations.

[0099] Calculate D α , D β , D γ , and X(t + 1) α , X(t + 1) β , X(t + 1) γ , then the position X(t + 1) that the wolf individual needs to adjust is:

[0100]

[0101] The specific optimization process is as follows:

[0102] Step 421. Define the number of iterations, the number of wolves, and the setting range of the Grey Wolf Optimization algorithm, where one wolf represents a combination of hyperparameters of the number of neurons in the first and second layers of a neural network and the number of iterations.

[0103] Step 422. Randomly initialize the wolf pack, calculate the fitness value of each wolf (the fitness value is the loss function value in Step 41), and take the positions of the top three wolves with the highest fitness values as the positions of the α, β, and γ wolves respectively.

[0104] Step 423.3 Update the positions of the remaining wolves according to the positions of the α, β, and γ wolves obtained in Step 421, and ensure that all wolves are within the set range.

[0105] Step 423.4 Recalculate the new α, β, and γ wolves based on the wolf pack updated in Step 423.3, and repeat Step 423.3.

[0106] Step 423.5 Determine whether the termination condition is reached. If so, end the algorithm and output the optimal wolf (the optimal model hyperparameter combination); otherwise, repeat Step 423.4.

[0107] Step 5. Transmit the optimal number of neurons and the number of iterations obtained in Step 423.5 to the model constructed in Step 41, perform prediction using the optimized LSTM model, perform inverse normalization on the predicted data, compare it with the true value, and visualize it.

[0108] The following are several specific embodiments of applying the present invention

[0109] Example 1: In Specific Example 1 of applying the present invention, the intelligent prediction method for wax deposition in oil wells based on LSTM includes:

[0110] Step 1. Obtain the time series data sequence of the oil well for training the model. The specific steps are as follows:

[0111] Step 11. In this example, obtain the historical production data of the oil well with well number D16 from 0:29:00 on January 21, 2020 to 19:01:00 on February 19, 2020 from the oil field database, including the area of the dynamometer card of the oil well and various different characteristics.

[0112] Step 12. Write the data obtained in Step 11 into a spreadsheet and save it locally.

[0113] Step 2. Perform data preprocessing operations to obtain the training data set, including the following steps:

[0114] Step 21. Read the spreadsheet in Step 12 to obtain the time series data sequence.

[0115] Step 22. For the time series data sequence obtained in Step 21, delete the missing data and irrelevant data, eliminate the outliers, and perform min-max normalization processing to convert the data values into the interval [0, 1].

[0116] Step 23. For the processed data, use most of the data as the training set and a small part of the data as the test set, and ensure that the training set consists of normal data only.

[0117] In Step 22, the formula for min-max normalization processing is as follows:

[0118]

[0119] where min(x) is the minimum value in the data and max(x) is the maximum value in the data.

[0120] In Step 23, for the processed data, use 80% of the data as the training set and 20% of the data as the test set. In this example, there are a total of 1400 data points, that is, there are 1120 data points in the training set and 280 data points in the test set.

[0121] Step 3. Use the SHAP tool to select features from the preprocessed data, and select the features that have a greater impact on the area of the dynamometer card of the oil well, including the following steps:

[0122] Step 31. In this example, construct an XGBoost as the machine learning model to be explained;

[0123] Step 32. Define the parameters (learning rate, maximum tree depth, number of trees) and regression metrics in the XGBoost model constructed in Step 31, and use the data obtained in Step 2 to train the model, that is, use the features other than the dynamometer card data of the oil well as the input and the dynamometer card data of the oil well as the label to train the model, and obtain the trained XGBoost model;

[0124] Step 33. For the XGBoost model trained in Step 32, use the SHAP tool to explain the prediction results of the model. The SHAP value formula is as follows:

[0125] Y i = y base + f(x i1 ) + f(x i2 ) + … + f(x ij )

[0126] where Y i is the predicted value of sample i, y base is the baseline value, and f(x ij ) is the contribution value of the jth feature of sample i.

[0127] The specific process is as follows:

[0128] Step 331. Calculate the benchmark value. In this example, the prediction result of XGBoost with all features except the dynamometer card of the oil well set to the average value is used as the benchmark, that is, the average prediction result of the model on the training data is used as the benchmark value.

[0129] Step 332. Randomly combine the features except the dynamometer card of the oil well. For each combination, calculate the contribution value of each feature to the prediction result, and aggregate the contribution values to obtain the SHAP value.

[0130] Step 333. Visualize the SHAP value obtained in Step 332. According to the level of the SHAP value, select the features that have a greater impact on the prediction result (that is, on the dynamometer card data). The bar chart of feature importance explanation in this example is as Figure 2 shown. Select the top eight features with the greatest impact, namely the features RYNL1_TJ_V_2, ZXZH_GTCJ, XXGL, RCYL, ZDZH_GTCJ, RYNL1_V_2, CMXS, and SXGL.

[0131] Step 4. Build and optimize the LSTM model. The specific process is as follows:

[0132] Step 41. According to the features selected in Step 3 that have a greater impact on the area of the oil well dynamometer card, use the selected features as the input and the area of the oil well dynamometer card as the label to build an LSTM network. In the present invention, the loss function is selected as MSE, and its formula is as follows;

[0133]

[0134] where Yi is the true value of sample i, is its corresponding predicted value.

[0135] Step 42. For the neural network built in Step 41, use the Grey Wolf Optimization (GWO) algorithm to optimize the hyperparameters of the model, specifically optimize the number of neurons in the first and second layers of the model neural network and the number of iterations. The principle of the Grey Wolf Optimization (GWO) algorithm is as follows:

[0136] In the Grey Wolf Optimization (GWO) algorithm, the wolf (hyperparameter combination) group updates its position according to the three wolves (α, β, γ) with the optimal position. The formula for updating the position D of the wolf from iteration t to t + 1 is as follows:

[0137] D = CX p (t) - X(t), X(t + 1) = X p (t) - AD

[0138] Among them, Xp is the position of the prey, which is the positions of wolves α, β, and γ, X is the position of the gray wolf, D is the distance between the gray wolf and the prey, and the calculation formulas for A and C are as follows:

[0139] A = 2ar 1 -a, C = 2r 2

[0140] Among them, r1 and r2 are random numbers from 0 to 1, and a linearly decreases from 2 to 0 with the number of iterations.

[0141] Calculate D respectively α , D β , D γ , and X(t + 1) α , X(t + 1) β , X(t + 1) γ Then the position X(t + 1) that the wolf individual needs to adjust is:

[0142]

[0143] The specific optimization process is as follows:

[0144] Step 421. In this example, the number of iterations of the gray wolf algorithm is defined as 200, the number of wolves is 50, and the search range is set. One wolf is a combination of model hyperparameters.

[0145] Step 422. Randomly initialize the wolf pack, calculate the fitness value of each wolf (the fitness value is the loss function value in Step 41), and take the positions of the gray wolves with the top three fitness values as the positions of gray wolves α, β, and γ respectively.

[0146] Step 423.3 Update the positions of the remaining gray wolves according to the positions of gray wolves α, β, and γ obtained in Step 421, and ensure that all gray wolves are within the set range.

[0147] Step 423.4 Recalculate the new gray wolves α, β, and γ according to the wolf pack updated in Step 423.3, and repeat Step 423.3.

[0148] Step 423.5 Judge whether the termination condition is reached. If it is reached, end the algorithm and output the optimal gray wolf (the optimal model hyperparameter combination). If not, repeat Step 423.4. The hyperparameter combination finally output in this example is that the number of neurons in the first layer is 200, the number of neurons in the second layer is 150, and the number of iterations is 600.

[0149] Step 5. Pass the optimal hyperparameter combination obtained in Step 423.5 to the model constructed in Step 41, and use the optimized LSTM model to make predictions on the test set. Perform an inverse normalization operation on the predicted data, compare it with the true value and visualize it. In this example, the last 100 data are visualized, and the results are asFigure 4 As shown, in order to compare the prediction results before and after the optimization of the Grey Wolf Optimizer algorithm, the LSTM model before optimization was also used for prediction, and the results are as Figure 3 shown, and the effect is not as good as that of the optimized model.

[0150] Example 2. The intelligent prediction method for paraffin deposition in oil wells based on LSTM includes:

[0151] Step 1. Obtain the time series data sequence of the oil well for model training. The specific steps are as follows:

[0152] Step 11. In this example, the historical production data of the oil well with well number N8 from 2016 / 11 / 3 12:49:18 to 2018 / 2 / 6 2:27:04, including the area of the dynamometer card of the oil well and various different features, was obtained from the oilfield database;

[0153] Step 12. Write the data obtained in Step 11 into a spreadsheet and save it locally.

[0154] Step 2. Perform data preprocessing operations to obtain the training dataset, including the following steps:

[0155] Step 21. Read the table in Step 12 to obtain the time series data sequence;

[0156] Step 22. For the time series data sequence obtained in Step 21, delete the missing data and irrelevant data, eliminate the outliers, and perform min-max normalization processing to convert the data values to the interval [0,1];

[0157] Step 23. For the processed data, most of the data is used as the training set, and a small part of the data is used as the test set, and it is ensured that the training set is all normal data.

[0158] In Step 22, the formula for min-max normalization processing is as follows:

[0159]

[0160] where min(x) is the minimum value in the data, and max(x) is the maximum value in the data.

[0161] In Step 23, for the processed data, 80% of the data is used as the training set, and 20% of the data is used as the test set. In this example, there are a total of 10,000 pieces of data, that is, there are 8,000 pieces of data in the training set and 2,000 pieces of data in the test set.

[0162] Step 3. Use the SHAP tool to select features from the preprocessed data, and select the features that have a greater impact on the area of the dynamometer card of the oil well, including the following steps:

[0163] Step 31. Construct an XGBoost as the machine learning model to be explained;

[0164] Step 32. Define the parameters (learning rate, maximum tree depth, number of trees) and regression metrics in the XGBoost model constructed in Step 31, and use the data obtained in Step 2 to train the model. That is, use the features other than the dynamometer card data of the oil well as the input and the dynamometer card data of the oil well as the label to train the model, and obtain a trained XGBoost model;

[0165] Step 33. For the trained XGBoost model in Step 32, use the SHAP tool to interpret the model prediction results. The SHAP value formula is as follows:

[0166] Y i = y base + f(x i1 ) + f(x i2 ) + … + f(x ij )

[0167] where Y i is the predicted value of sample i, y base is the baseline value, and f(x ij ) is the contribution value of the j features of sample i.

[0168] The specific process is as follows:

[0169] Step 331. Calculate the baseline value. In this example, the prediction result of XGBoost when all features except the dynamometer card of the oil well are set to the average value is used as the baseline, that is, the average prediction result of the model on the training data is used as the baseline value.

[0170] Step 332. Randomly combine the features except the dynamometer card of the oil well. For each combination, calculate the contribution value of each feature to the prediction result, and aggregate the contribution values to obtain the SHAP value.

[0171] Step 333. Visualize the SHAP values obtained in Step 332, and select the features that have a greater impact on the prediction result (that is, on the dynamometer card data) according to the high and low of the SHAP values.

[0172] Step 4. Construct and optimize the LSTM model. The specific process is as follows:

[0173] Step 41. According to the features that have a greater impact on the area of the dynamometer card of the oil well selected in Step 3, use the selected features as the input and the area of the dynamometer card of the oil well as the label to construct an LSTM network. In the present invention, the loss function is selected as MSE, and its formula is as follows;

[0174]

[0175] Where Yi is the true value of sample i, and its corresponding predicted value.

[0176] Step 42. For the neural network constructed in Step 41, use the Grey Wolf Optimizer (GWO) to optimize the hyperparameters of the neural network. Specifically, optimize the number of neurons in the first and second layers of the neural network and the number of iterations. The principle of the Grey Wolf Optimizer is as follows:

[0177] In the Grey Wolf Optimizer (GWO), the population of wolves (hyperparameter combinations) updates their positions according to the three wolves with the best positions (α, β, γ). The formula for updating the position D of the wolves from iteration t to t + 1 is as follows:

[0178] D = CX p (t) - X(t), X(t + 1) = X p (t) - AD

[0179] where Xp is the position of the prey, i.e., the positions of the α, β, γ wolves, X is the position of the grey wolf, D is the distance between the grey wolf and the prey, and the calculation formulas for A and C are as follows:

[0180] A = 2ar 1 -a, C = 2r 2

[0181] where r1 and r2 are random numbers from 0 to 1, and a linearly decreases from 2 to 0 with the number of iterations.

[0182] Calculate D α , D β , D γ , and X(t + 1) α , X(t + 1) β , X(t + 1) γ respectively. Then the position X(t + 1) that the wolf individual needs to adjust is:

[0183]

[0184] The specific optimization process is as follows:

[0185] Step 421. In this example, define the number of iterations of the Grey Wolf Optimizer as 200, the number of wolves as 50, and set the search range. One wolf represents a hyperparameter combination of the neural network.

[0186] Step 422. Randomly initialize the wolf population, calculate the fitness value of each wolf (the fitness value is the loss function value in Step 41), and take the positions of the three grey wolves with the top three fitness values as the positions of the α, β, γ grey wolves respectively.

[0187] Step 423.3 Update the positions of the remaining gray wolves according to the positions of the α, β, and γ gray wolves obtained in Step 421, and ensure that all gray wolves are within the set range.

[0188] Step 423.4 Recalculate the new α, β, and γ gray wolves based on the wolf pack updated in Step 423.3, and repeat Step 423.3.

[0189] Step 423.5 Determine whether the termination condition is reached. If so, end the algorithm and output the optimal gray wolf (optimal model hyperparameter combination); otherwise, repeat Step 423.4. The hyperparameter combination finally output in this example is that the number of neurons in the first layer is 100, the number of neurons in the second layer is 50, and the number of iterations is 500.

[0190] Step 5 Transfer the optimal hyperparameter combination obtained in Step 423.5 to the model constructed in Step 41, perform predictions on the test set using the optimized LSTM model, perform inverse normalization on the predicted data, compare it with the true value, and visualize it. In this example, the last 100 data are visualized, and the results are as Figure 5 shown.

[0191] Example 3: In Specific Example 3 of applying the present invention, the intelligent wax scaling prediction method based on LSTM includes:

[0192] Step 1 Obtain the time series data sequence of the oil well for training the model. The specific steps are as follows:

[0193] Step 11 In this example, obtain the historical production data including the area of the dynamometer card of the oil well with well number S72 from 2016 / 6 / 10:01 to 2019 / 4 / 18 10:11 from the oilfield database, including various different characteristics.

[0194] Step 12 Write the data obtained in Step 11 into a spreadsheet and save it locally.

[0195] Step 2 Perform data preprocessing operations to obtain a training data set, including the following steps:

[0196] Step 21 Read the table in Step 12 to obtain the time series data sequence.

[0197] Step 22 For the time series data sequence obtained in Step 21, delete the missing data and irrelevant data, eliminate the outliers, and perform min-max normalization processing to convert the data values to the interval [0, 1].

[0198] Step 23 For the processed data, use most of the data as the training set and a small part of the data as the test set, and ensure that the training set consists of normal data only.

[0199] In step 22, the formula for min-max normalization is as follows:

[0200]

[0201] where min(x) is the minimum value in the data and max(x) is the maximum value in the data.

[0202] In step 23, for the processed data, 80% of the data is used as the training set and 20% of the data is used as the test set. In this example, there are a total of 10,000 data, that is, there are 8,000 data in the training set and 2,000 data in the test set.

[0203] Step 3: Use the SHAP tool to select features from the preprocessed data and select the features that have a greater impact on the area of the dynamometer card of the oil well, including the following steps:

[0204] Step 31. Construct an XGBoost as the machine learning model to be explained;

[0205] Step 32. Define the parameters (learning rate, maximum tree depth, number of trees) and regression metrics in the XGBoost model constructed in step 31, and use the data obtained in step 2 to train the model, that is, use the features other than the dynamometer card data of the oil well as the input and the dynamometer card data of the oil well as the label to train the model, and obtain the trained XGBoost model;

[0206] Step 33. For the trained XGBoost model in step 32, use the SHAP tool to explain the model prediction results. The SHAP value formula is as follows:

[0207] Y i = y base + f(x i1 ) + f(x i2 ) + … + f(x ij )

[0208] where Y i is the predicted value of sample i, y base is the baseline value, and f(x ij ) is the contribution value of the jth feature of sample i.

[0209] The specific process is as follows:

[0210] Step 331. Calculate the baseline value. In this example, the prediction result of XGBoost when all features except the dynamometer card of the oil well are set to the average value is used as the baseline, that is, the average prediction result of the model on the training data is used as the baseline value.

[0211] Step 332. Randomly combine the features other than the oil well dynamometer card, calculate the contribution value of each feature to the prediction result for each combination, and aggregate the contribution values to obtain the SHAP value.

[0212] Step 333. Visualize the SHAP values obtained in Step 332, and select the features that have a greater impact on the prediction result (i.e., on the dynamometer card data) based on the magnitude of the SHAP values.

[0213] Step 4. Construct and optimize the LSTM model, and the specific process is as follows:

[0214] Step 41. Based on the features selected in Step 3 that have a greater impact on the area of the oil well dynamometer card, construct an LSTM network with the selected features as the input and the area of the oil well dynamometer card as the label. In the present invention, the loss function is selected as MSE, and its formula is as follows;

[0215]

[0216] where Yi is the true value of sample i, is its corresponding predicted value.

[0217] Step 42. For the neural network constructed in Step 41, use the Grey Wolf Optimization (GWO) algorithm to optimize the hyperparameters of the neural network, specifically optimize the number of neurons in the first and second layers of the neural network and the number of iterations. The principle of the Grey Wolf Optimization (GWO) algorithm is as follows:

[0218] In the Grey Wolf Optimization (GWO) algorithm, the group of wolves (hyperparameter combinations) updates their positions according to the three wolves (α, β, γ) with the optimal positions. The formula for updating the position D of the wolves from iteration t to t + 1 is as follows:

[0219] D = CX p (t) - X(t), X(t + 1) = X p (t) - AD

[0220] where, Xp is the position of the prey, i.e., the positions of the α, β, γ wolves, X is the position of the grey wolf, D is the distance between the grey wolf and the prey, and the calculation formulas for A and C are as follows:

[0221] A = 2ar 1 - a, C = 2r 2

[0222] where, r1, r2 are random numbers from 0 to 1, and a linearly decreases from 2 to 0 with the number of iterations.

[0223] Calculate D α , D β , D γ , and X(t + 1) α , X(t + 1) β , X(t + 1)γ , the position X(t+1) that the wolf individual needs to adjust is as follows:

[0224]

[0225] The specific optimization process is as follows:

[0226] Step 421. In this example, the iteration number of the grey wolf algorithm is defined as 300, the number of wolves is 60, and the search range is set. One wolf represents a combination of neural network hyperparameters.

[0227] Step 422. Randomly initialize the wolf pack, calculate the fitness value of each wolf (the fitness value is the loss function value in Step 41), and take the positions of the top three grey wolves with the best fitness values as the positions of the α, β, and γ grey wolves respectively.

[0228] Step 423.3 Update the positions of the remaining grey wolves according to the positions of the α, β, and γ grey wolves obtained in Step 421, and ensure that all grey wolves are within the set range.

[0229] Step 423.4 Recalculate the new α, β, and γ grey wolves based on the wolf pack updated in Step 423.3, and repeat Step 423.3.

[0230] Step 423.5 Determine whether the termination condition is reached. If so, end the algorithm and output the optimal grey wolf (the optimal model hyperparameter combination). If not, repeat Step 423.4. The hyperparameter combination finally output in this example is that the number of neurons in the first layer is 128, the number of neurons in the second layer is 64, and the iteration number is 200.

[0231] Step 5. Transfer the optimal hyperparameter combination obtained in Step 423.5 to the model constructed in Step 41, use the optimized LSTM model to make predictions on the test set, perform an inverse normalization operation on the predicted data, and compare it with the true value.

[0232] In summary, through several examples with different data and sample sizes, the present invention has shown excellent performance. It proves that the method of using the LSTM model to predict oil well wax deposition data, using the SHAP tool for feature selection, and optimizing it with the grey wolf algorithm has high accuracy and robustness.

[0233] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

[0234] Except for the technical features described in the specification, the rest are well-known technologies to those skilled in the art.

Claims

1. An intelligent prediction method for oil well wax deposition based on LSTM, characterized in that: The LSTM-based intelligent prediction method for oil well wax deposition includes: Step 1, obtaining the oil well time series data sequence for training the model; Step 2: perform data preprocessing to obtain a training data set; Step 3, performing feature selection on the preprocessed data, and selecting features that have a greater impact on the area of ​​the oil well indicator diagram; Step 4, construct and optimize the oil well wax deposition prediction model based on LSTM network; Step 5: Use the optimized LSTM model to make predictions, and denormalize the predicted data to obtain the predicted real data.

2. The LSTM-based intelligent prediction method for oil well wax deposition according to claim 1, characterized in that: In step 1, oil well data including oil well dynamometer area and historical production data of various characteristics are obtained from the oil field database; the obtained historical production data is written into a spreadsheet and saved locally.

3. The intelligent prediction method for oil well wax deposition based on LSTM according to claim 2, characterized in that: Step 2 includes: Step 21. Read the spreadsheet in step 1 to obtain the time series data sequence; Step 22. For the time series data obtained in step 21, delete missing data and irrelevant data, remove outliers, and perform min-max standardization to convert the data values ​​to the [0,1] interval; Step 23. For the processed data, use most of the data as the training set and a small part of the data as the test set, and ensure that the training set is normal data.

4. The intelligent prediction method for oil well wax deposition based on LSTM according to claim 3, characterized in that: In step 22, the formula for the min-max normalization process is as follows: Where min(x) is the minimum value in the data, and max(x) is the maximum value in the data.

5. The intelligent prediction method for oil well wax deposition based on LSTM according to claim 3, characterized in that: In step 23, for the processed data, 80% of the data is used as a training set and 20% of the data is used as a test set.

6. The intelligent prediction method for oil well wax deposition based on LSTM according to claim 1, characterized in that: Step 3 includes: Step 31. Build an interpretable machine learning model; Step 32. Define various parameters in the machine learning model constructed in step 31, and use the data obtained in step 2 to train the model, that is, use the features other than the oil well indicator diagram data as input, and use the oil well indicator diagram data as labels to train the model, so as to obtain a trained machine learning model; Step 33. For the machine learning model trained in step 32, use the SHAP tool to interpret the model prediction results.

7. The LSTM-based intelligent prediction method for oil well wax deposition according to claim 6, characterized in that: In step 33, the SHAP value formula is as follows: Y i =y base +f(x i1 )+f(x i2 )+…+f(x ij ) where Y i is the predicted value of sample i, y base is the reference value, f(x ij ) is the contribution value of the j features of sample i.

8. The LSTM-based intelligent prediction method for oil well wax deposition according to claim 7, characterized in that: Step 33 specifically includes: Step 331. Calculate the benchmark value, and take the prediction result of the machine learning model when all features except the oil well indicator diagram are set to the average value as the benchmark, that is, the average prediction result of the model on the training data as the benchmark value; Step 332: Randomly combine the features except the oil well performance diagram, calculate the contribution value of each feature to the prediction result for each combination, and aggregate the contribution values ​​to obtain the SHAP value; Step 333. Visualize the SHAP value obtained in step 332, and select features that have a greater impact on the prediction result, i.e., the dynamometer data, based on the SHAP value.

9. The intelligent prediction method for oil well wax deposition based on LSTM according to claim 1, characterized in that: Step 4 includes: Step 41. According to the features that have a greater impact on the area of ​​the oil well indicator diagram selected in step 3, an LSTM network is constructed with the selected features as input and the area of ​​the oil well indicator diagram as a label; Step 42. For the oil well wax deposition prediction model constructed in step 41, the Grey Wolf Algorithm (GWO) is used to optimize the model hyperparameters, specifically, the number of neurons in the first and second layers of the model neural network and the number of iterations are optimized.

10. The intelligent prediction method for oil well wax deposition based on LSTM according to claim 9, characterized in that: In step 41, when constructing the LSTM network, the loss function is selected as MSE, and its formula is as follows: Where Yi is the true value of sample i, is the corresponding predicted value.

11. The LSTM-based intelligent prediction method for oil well wax deposition according to claim 10, characterized in that: In step 42, the grey wolf algorithm works as follows: The wolves in the Grey Wolf Algorithm GWO, i.e., the hyperparameter combination group, update their positions according to the three wolves (α, β, γ) with the best positions. The formula for updating the wolf's position D from iteration number t to t+1 is as follows: D=CX p (t)-X(t),X(t+1)=X p (t)-AD Among them, Xp is the position of the prey, that is, the position of the α, β, and γ wolves, X is the position of the gray wolf, and D is the distance between the gray wolf and the prey. The calculation formulas of A and C are as follows: A=2ar1-a,C=2r2 Among them, r1, r2 are random numbers between 0 and 1, and a decreases linearly from 2 to 0 with the number of iterations; Calculate D α , D β ,D γ , and X(t+1) α ,X(t+1) β ,X(t+1) γ , then the position X(t+1) that the wolf needs to adjust is:

12. The LSTM-based intelligent prediction method for oil well wax deposition according to claim 11, characterized in that: Step 42 specifically includes: Step 421. Define the number of iterations of the gray wolf algorithm, the number of wolves, and the setting range, where a wolf is a hyperparameter combination of the number of neurons in the first and second layers of a neural network and the number of iterations; Step 422. Randomly initialize the wolf pack, calculate the fitness value of each wolf, the fitness value is the loss function value in step 41, and use the positions of the top three gray wolves in terms of fitness value as the α, β, and γ gray wolf positions respectively; Step 423.

3. Update the positions of the remaining gray wolves according to the positions of the α, β, and γ gray wolves obtained in step 421, and ensure that all gray wolves are within the set range; Step 423.

4. Recalculate new α, β, and γ gray wolves based on the wolf pack updated in step 423.3, and repeat step 423.3; Step 423.

5. Determine whether the termination condition is met. If so, terminate the algorithm and output the optimal gray wolf, i.e., the optimal model hyperparameter combination. If not, repeat step 423.

4.

13. The LSTM-based intelligent prediction method for oil well wax deposition according to claim 12, characterized in that: In step 5, the optimal model hyperparameter combination obtained in step 423.5, including the optimal number of neurons and the number of iterations, is passed to the model constructed in step 41, and the optimized LSTM model is used for prediction. The predicted data is denormalized, compared with the true value and visualized.

Citation Information

Patent Citations

  • Prediction method of dynamic paraffin removal period of oil well in offshore oilfield

    CN109339765A

  • Oil well thermal washing frequency and date planning method

    CN110807166A

  • Optimization of Oil Well Hot Washing and Dewaxing Methods and Parameters

    CN112796704B