Oil well yield prediction method and device, storage medium and electronic equipment
By combining the time-frequency Transformer and Bi-LSTM methods, the accuracy problem of oil well production prediction is solved. By utilizing time-frequency information fusion and long-term dependency capture, the prediction accuracy and model stability are improved, which is suitable for oil well production prediction in the field of oil extraction.
Patent Information
- Application Number
- CN202511249198.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-03
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-03
AI Technical Summary
Existing technologies are unable to accurately predict the oil production of target oil wells. They face challenges such as scarcity of training samples, significant production fluctuations, and complexity of sample labeling, which affect model prediction results and data processing efficiency.
The method of combining time-frequency Transformer and Bi-LSTM is adopted to capture the long-term dependency of time series through the fusion of time domain and frequency domain information and forward and backward LSTM. Combined with the multi-head attention mechanism and iterative training, the optimal parameter model is constructed to predict oil production.
It achieves accurate prediction of the oil production of target oil wells, improves the generalization ability and prediction accuracy of the model, and can accurately capture the long-range dependencies and periodic characteristics of complex time series data.
Smart Images

Figure CN120744874A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence selection in oil production, and in particular relates to a method, device, storage medium and electronic equipment for predicting oil well production. Background Art
[0002] The Junggar Basin is extremely rich in oil and gas resources, with total estimated oil resources of 8.6 billion tons and natural gas resources of 2.1 trillion cubic meters. Currently, the proven oil rate is only 21.4%, and the proven natural gas rate is less than 3.64%, demonstrating broad exploration prospects and enormous development potential. The provenance of the study area primarily originates in the northwest, with thickness gradually decreasing from northeast to southwest. The reservoirs in this area have a porosity of 16.5% and a permeability of 7.5 mD, exhibiting strong heterogeneity, relatively poor physical properties, and a complex pore structure.
[0003] In recent years, artificial intelligence has been widely used in the oil and gas industry. Many researchers have applied AI techniques to production forecasting based on dynamic data from oilfield development. Network structures such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), and transformers can effectively process time series data and are crucial for analyzing cyclical changes in oil production and production trends. Chen et al. designed a model based on a backpropagation neural network to efficiently predict the dynamic development of low-permeability oilfields. Su et al. used an LSTM network to learn production trends over two years and accurately predict production capacity for the next four months. Liang et al. proposed a BiLSTM-RF-MPA model designed to handle the complex nonlinear and nonstationary characteristics of shale gas production time series. It effectively captures inherent correlations in the data and reduces model uncertainty. Zha et al. designed a hybrid model combining a convolutional neural network (CNN) and an LSTM. This model has complementary advantages in feature extraction and learning sequence dependencies, enabling the CNN-LSTM model to accurately describe changing trends in the production process. Wang et al. used simplified static formation data and dynamic data collected during daily production to develop a production forecasting method based on the Transformer architecture. This method takes into account human interference factors such as well shut-ins that may occur during production. These studies demonstrate that combining advanced deep learning techniques can significantly improve the accuracy and reliability of oilfield production forecasts.
[0004] The key to building a yield prediction model lies in fully leveraging historical oilfield production data to establish an accurate prediction mechanism. However, in practical applications, this approach faces a series of challenges, such as the scarcity of training samples, the significance of yield fluctuations, and the complexity of sample annotation. These factors not only affect the model's predictive effectiveness but also make data processing time-consuming and labor-intensive. Therefore, to address these issues, this paper first designs an effective data processing strategy and then proposes a method for oil well yield prediction that combines a time-frequency Transformer with a Bi-LSTM. By fusing information from the time domain (self-attention path) and the frequency domain (Fourier path), the model can capture global contextual information while extracting and utilizing periodic features. Furthermore, through the use of forward LSTM and backward LSTM, the model captures long-term dependencies in time series data, improving the modeling capabilities for complex time series data.
[0005] Therefore, how to provide a prediction method that can accurately predict the oil production of a target oil well is a technical problem to be solved. Summary of the Invention
[0006] Based on this, it is necessary to provide a method, device, storage medium and electronic device for predicting oil well production to address the defect that existing technologies cannot accurately predict the oil production of target oil wells.
[0007] In a first aspect, an embodiment of the present invention provides a method for predicting oil well production, the method comprising: Get the training set and test set; inputting the training set data in the training set into the bidirectional long short-term memory network branch and the time-frequency converter branch in sequence for processing, and outputting a first prediction value and a second prediction value, wherein the training set data includes multivariate time series data of historical oil production of a target oil well; Calculating a Q matrix based on the first predicted value, and calculating a K matrix and a V matrix based on the second predicted value; The Q matrix, the K matrix, and the V matrix are fused through a multi-head attention mechanism to obtain a fused result; After transforming the fused result through a linear layer, an optimal parameter model is obtained through iterative training; The oil production of the target oil well is predicted by using the optimal parameter model to predict the oil production of the target oil well.
[0008] Optionally, after obtaining the optimal parameter model, the method further includes: The test set data in the test set is input into the optimal parameter model, and the model performance of the optimal parameter model is evaluated to obtain an evaluation result.
[0009] Optionally, the step of sequentially inputting the training set data in the training set into the bidirectional long short-term memory network branch and the time-frequency converter branch for processing includes: Inputting the training set data in the training set into the bidirectional long short-term memory network branch, processing the data through the forward long short-term memory network and the reverse long short-term memory network in the bidirectional long short-term memory network branch to capture the long-term dependency of the time series data and obtain an output state; and The training set data in the training set is input into the time-frequency converter branch, and the time domain features and frequency domain features of the sequence data are captured by combining self-attention and Fourier transform.
[0010] Optionally, the step of inputting the training set data in the training set into the bidirectional long short-term memory network branch, processing the data through the forward long short-term memory network and the reverse long short-term memory network in the bidirectional long short-term memory network branch to capture the long-term dependency of the time series data and obtain the output state includes: Processing the training set data from front to back to obtain a forward hidden state; Processing the training set data from back to front to obtain a reverse hidden state; The forward hidden state and the reverse hidden state are concatenated to capture the long-term dependency of the time series data, thereby obtaining the output state.
[0011] Optionally, processing the training set data from front to back to obtain a forward hidden state includes: Perform calculation based on the input, the weight matrix of the input gate, the first bias term, and the first activation function to obtain the input gate; Performing calculation based on the weight matrix of the forget gate, the input, the second bias term, and the first activation function to obtain a forget gate; Perform calculation based on the weight matrix of the candidate memory unit, the input, the third bias term, and the second activation function to obtain a first memory unit; Performing calculation based on a weight matrix of an output gate, the input, the first activation function, and a fourth bias term to obtain an output gate; Calculation is performed based on the output gate, the first memory unit, and the second activation function to obtain the forward hidden state.
[0012] Optionally, processing the training set data from back to front to obtain a reverse hidden state includes: The reverse hidden state is obtained by performing calculation based on the output gate corresponding to the back-to-forward processing data, the memory unit corresponding to the back-to-forward processing data, and the third activation function.
[0013] Optionally, inputting the training set data in the training set into the time-frequency converter branch, and combining self-attention and Fourier transform to capture the time domain features and frequency domain features of the sequence data, includes: Acquire the training set data, where the training set data is preprocessed time series data; For the preprocessed time series data, the relevance weight of each time step in the input sequence is calculated through the self-attention mechanism to generate a time series context representation; The pre-processed time series data is calculated by fast Fourier transform to obtain the input frequency domain features; Extracting the first k frequency components with the largest amplitudes from the frequency domain features, so that the first k frequency components with the largest amplitudes represent the period of the input sequence; For the periodic results of each period, continuous convolution operations with different convolution kernel sizes are used to extract periodic features and enhance the input time-frequency information. The mean amplitude of each cycle is used as the weight to perform weighted summation on all cycle results to aggregate the cycle characteristics; Based on the weight parameters, the time domain output and the frequency domain output are weightedly fused to obtain the comprehensive features; Adding the fusion result and the input sequence through a residual connection to obtain an addition result; and performing layer normalization processing on the addition result to obtain a normalized output; The normalized output is sequentially processed through a feedforward neural network, a residual connection process, and a layer normalization process to capture the time domain features and the frequency domain features of the sequence data.
[0014] Optionally, calculating a Q matrix based on the first predicted value, and calculating a K matrix and a V matrix based on the second predicted value, includes: Performing calculation based on the first predicted value, the queried linear transformation weight matrix, and a fifth bias term to obtain the Q matrix; Performing calculation based on the linear transformation weight matrix of the key, the second prediction value, and the sixth bias term to obtain the K matrix; The V matrix is obtained by performing calculation based on the linear transformation weight matrix of the value, the second prediction value and the seventh bias term.
[0015] Optionally, the fusing of the Q matrix, the K matrix, and the V matrix through a multi-head attention mechanism to obtain a fused result includes: Through the multi-head attention mechanism, fusion processing is performed based on the Q matrix, the K matrix, the V matrix, the dimension of the key vector and the transposition operation of K to obtain the fused result.
[0016] Optionally, before obtaining the training set and the test set, the method further includes: Acquire a target data set, wherein the data in the target data set is oil production data of a target oil well within a preset time period; Preprocessing each data in the target data set to obtain preprocessed data, and forming a preprocessed data set based on the preprocessed data; The data in the preprocessed data set is divided according to a preset ratio to obtain a training set and a test set.
[0017] Optionally, obtaining an optimal parameter model through iterative training includes: Define optimization goals; Smooth L1 loss is used as the loss function and Adam is used as the optimizer. The optimal parameter model is obtained through iterative training.
[0018] In a second aspect, an embodiment of the present invention provides a device for predicting oil well production, the device comprising: Acquisition module, used to obtain training set and test set; a processing module, configured to sequentially input training set data in the training set into a bidirectional long short-term memory network branch and a time-frequency converter branch for processing, and output a first predicted value and a second predicted value, wherein the training set data includes multivariate time series data of historical oil production of a target oil well; a calculation module, configured to calculate a Q matrix based on the first predicted value, and calculate a K matrix and a V matrix based on the second predicted value; A fusion module is used to fuse the Q matrix, the K matrix, and the V matrix through a multi-head attention mechanism to obtain a fused result; An iterative training module, configured to obtain an optimal parameter model through iterative training after converting the fused result through a linear layer; The prediction module is used to predict the oil production of the target oil well by using the optimal parameter model to predict the oil production of the target oil well.
[0019] In a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method of the first aspect.
[0020] In a fourth aspect, an electronic device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method of the first aspect is implemented.
[0021] In an embodiment of the present invention, a training set and a test set are obtained; the training set data in the training set are sequentially input into a bidirectional long short-term memory network branch and a time-frequency converter branch for processing, and a first predicted value and a second predicted value are output, wherein the training set data includes multivariate time series data of the historical oil production of a target oil well; a Q matrix is calculated based on the first predicted value, and a K matrix and a V matrix are calculated based on the second predicted value; the Q matrix, the K matrix, and the V matrix are fused through a multi-head attention mechanism to obtain a fused result; after transforming the fused result through a linear layer, an optimal parameter model is obtained through iterative training; and the oil production of the target oil well is predicted using the optimal parameter model to predict the oil production of the target oil well. The prediction method provided by the embodiment of the present invention can accurately predict the oil production of the target oil well. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] A more complete understanding of the exemplary embodiments of the present invention may be obtained by referring to the following drawings. The drawings are provided to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and are not intended to limit the present invention. In the drawings, the same reference numerals generally represent the same components or steps.
[0023] Figure 1 A flowchart of a method for predicting oil well production according to an exemplary embodiment of the present invention; Figure 2 This is a network model diagram for a specific application scenario; Figure 3 A diagram showing the fitting results of the training set used in the prediction method provided in an embodiment of the present invention; Figure 4 A fitting result diagram of the test set used by the prediction method provided in an embodiment of the present invention; Figure 5 The flowchart of the method for predicting oil well production in a specific application scenario is as follows; Figure 6 A schematic structural diagram of an oil well production prediction device provided according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0024] Exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments described herein. Rather, these embodiments are provided to enable a more thorough understanding of the present invention and to fully convey the scope of the present invention to those skilled in the art.
[0025] It should be noted that, unless otherwise specified, the technical or scientific terms used in the present invention should have the common meanings understood by those skilled in the art to which the present invention belongs.
[0026] In addition, the terms "first" and "second" are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, apparatus, product, or device comprising a series of steps or elements is not limited to the listed steps or elements, but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or device.
[0027] Embodiments of the present invention provide a method and apparatus for predicting oil well production, a computer-readable medium, and an electronic device, which are described below with reference to the accompanying drawings.
[0028] Please refer to Figure 1 , which shows a flow chart of a method for predicting oil well production provided by some embodiments of the present invention, such as Figure 1 As shown, the method for predicting oil well production may include the following steps: Step S101: Obtain a training set and a test set.
[0029] In one example, before obtaining the training set and the test set, the oil well production prediction method provided by the embodiment of the present invention may further include the following steps: Obtaining a target data set, where the data in the target data set is oil production data of a target oil well within a preset time period; Preprocessing each data in the target data set to obtain preprocessed data, and forming a preprocessed data set based on the preprocessed data; According to the preset ratio, the data in the preprocessed data set is divided into a training set and a test set.
[0030] In the actual application scenario, the oil production data of the target oil wells in the target oil field over the past few years were collected, and the three basic parameters in the oil well development process, oil pressure, casing pressure, and back pressure, were selected to construct the basic data set of the neural network.
[0031] The oil production dataset in the above steps includes: oil pressure, casing pressure, back pressure, and historical monthly average oil production. The oil pressure is the tubing pressure in MPa; the casing pressure is the casing pressure (casing annulus pressure) in MPa; and the back pressure is the pressure in the surface oil pipeline in MPa.
[0032] In practical applications, the data in the dataset was preprocessed, employing piecewise cubic Hermite interpolation polynomials for interpolation and filling to address zero values. Furthermore, to reduce noise in the raw production data, a three-point moving average filter was introduced to smooth the data and improve its quality. Furthermore, z-score normalization was used to transform the data into a distribution with a mean of 0 and a standard deviation of 1, effectively mitigating any non-stationarity in the input data.
[0033] In actual application scenarios, the data set is divided and the sliding window technology is used to reorganize the data. Specifically, the data of the previous m days (m is the length of the sliding window) is selected from the historical production data to predict the production value of the next time step, and the actual production value of the time step is used as the label for training. At the same time, the production sequence data is divided into training set and test set in a ratio of 8:2.
[0034] It should be noted that an appropriate sliding window length m is selected. For example, if m=7, it means using the data of the previous 7 days to predict the production value of the 8th day.
[0035] In the prediction method provided in the embodiment of the present invention, the above prediction ratio is not specifically limited and can be adjusted according to the requirements of different application scenarios, which will not be described in detail here.
[0036] Step S102: input the training set data in the training set into the bidirectional long short-term memory network branch and the time-frequency converter branch in sequence for processing, and output a first prediction value and a second prediction value. The training set data includes multivariate time series data of the historical oil production of the target oil well.
[0037] like Figure 2 As shown in Figure 1, it is a network model diagram in a specific application scenario. Figure 2 As shown, the data is first preprocessed and divided into training and test sets. The training set is then fed into a time-frequency Transformer module and a Bi-LSTM. The time-frequency Transformer combines self-attention and Fourier transforms to simultaneously capture both time-domain and frequency-domain features of sequence data, enhancing its modeling capabilities. The outputs of the two branches are then fused using a multi-head attention mechanism to achieve adaptive allocation of attention weights. After training, the optimal parameter model is obtained, and oil production prediction is performed using the test set.
[0038]
[0039] Table 1 Oil production prediction results using the prediction method provided by the embodiment of the present invention As can be seen from Table 1, the prediction method provided by the embodiment of the present invention performs well in the three error indicators of MAE, RMSE and MAPE. Compared with models such as SVM, LSTM and Transformer, the prediction results are closer to the true value, and the generalization ability of the model is significantly enhanced. The value reaches 0.95, far higher than other methods, indicating that the model has stronger explanatory power and higher prediction accuracy. This shows that the effective application of the time-frequency transformer and Bi-LSTM enables the model to accurately capture the complex patterns and long-range dependencies in time series, improving the accuracy of predictions.
[0040] In one example, the training set data in the training set are sequentially input into the bidirectional long short-term memory network branch and the time-frequency transformer branch for processing, including the following steps: Inputting the training set data in the training set into the bidirectional long short-term memory network branch, processing it through the forward long short-term memory network and the reverse long short-term memory network in the bidirectional long short-term memory network branch to capture the long-term dependency of the time series data and obtain the output state; and The training set data in the training set is input into the time-frequency converter branch, and the self-attention and Fourier transform are combined to capture the time domain features and frequency domain features of the sequence data.
[0041] In one example, the training set data in the training set is input into the bidirectional long short-term memory network branch, and processed by the forward long short-term memory network and the reverse long short-term memory network in the bidirectional long short-term memory network branch to capture the long-term dependency of the time series data and obtain the output state, including the following steps: Process the training set data from front to back to obtain the forward hidden state; Process the training set data from back to front to obtain the reverse hidden state; The forward hidden state and the reverse hidden state are concatenated to capture the long-term dependencies of the time series data and obtain the output state.
[0042] In one example, processing the training set data from front to back to obtain the forward hidden state includes the following steps: Perform calculation based on the input, the weight matrix of the input gate, the first bias term, and the first activation function to obtain the input gate; Calculate based on the weight matrix of the forget gate, the input, the second bias term and the first activation function to obtain the forget gate; Calculate based on the weight matrix, input, third bias term and second activation function of the candidate memory unit to obtain a first memory unit; Calculate based on the weight matrix of the output gate, the input, the first activation function and the fourth bias term to obtain the output gate; Calculation is performed based on the output gate, the first memory unit, and the second activation function to obtain the forward hidden state.
[0043] In the above steps, the first activation function is the sigmoid activation function, and the second activation function is the hyperbolic tangent activation function.
[0044] It should be noted that the data is processed from front to back to obtain the forward hidden state The specific process is as follows: Assume the input is ,First, calculate the input gate : Formula (1); in, Represents the weight matrix of the input gate; represents the first bias term; is the first activation function, the first activation function is the sigmoid activation function, is the previous hidden state.
[0045] Next, calculate the forget gate : Formula (2); in, Represents the weight matrix of the forget gate; represents the second bias term; Represents the sigmoid activation function.
[0046] Next, calculate the first memory unit
[0047] Formula (3); Formula (4); in, For the Gate of Forgetfulness; The weight matrix representing the candidate memory unit; represents the third bias term, is the hyperbolic tangent activation function, Represents element-wise multiplication.
[0048] Next, calculate the output gate
[0049] Formula (5); in, Represents the weight matrix of the output gate; represents the fourth bias term; Represents the sigmoid activation function.
[0050] Finally, calculate the hidden state , and get the forward hidden state.
[0051] Formula (6); in, is the output gate, represents element-wise multiplication, is hidden state, Is the second activation function, the second activation function is the hyperbolic tangent activation function.
[0052] In one example, processing the training set data from back to front to obtain the reverse hidden state includes the following steps: The reverse hidden state is obtained based on the output gate corresponding to the data processed from back to front, the memory unit corresponding to the data processed from back to front, and the third activation function.
[0053] In a specific application scenario, the training set data is processed from back to front to obtain the reverse hidden state , as follows: Formula (7); The calculation steps are the same as those for forward processing data. Represents the output gate in processing data from back to front; Represents a memory unit, tanh is the second activation function, which is the hyperbolic tangent activation function. Represents element-wise multiplication.
[0054] In actual application scenarios, after obtaining the above forward hidden state and reverse hidden state, the forward hidden state and reverse hidden state are merged to obtain the output state
[0055] Formula (8); in, is the output state, is the forward hidden state, is the reverse hidden state, It means concatenating the forward hidden state and the reverse hidden state along the last dimension.
[0056] In one example, the training set data in the training set is input into the time-frequency converter branch, and the self-attention and Fourier transform are combined to capture the time domain features and frequency domain features of the sequence data, including the following steps: Get the training set data, which is the preprocessed time series data; For the preprocessed time series data, the self-attention mechanism is used to calculate the correlation weight of each time step in the input sequence to generate a temporal context representation; For the preprocessed time series data, the fast Fourier transform is used to calculate and obtain the frequency domain features of the input; Extract the first k frequency components with the largest amplitudes from the frequency domain features, so that the first k frequency components with the largest amplitudes represent the period of the input sequence; For the periodic results of each period, continuous convolution operations with different convolution kernel sizes are used to extract periodic features and enhance the input time-frequency information. The mean amplitude of each cycle is used as the weight to perform weighted summation on all cycle results to aggregate the cycle characteristics; Based on the weight parameters, the time domain output and the frequency domain output are weightedly fused to obtain the comprehensive features; The fusion result and the input sequence are added through residual connection to obtain the addition result; and the addition result is layer-normalized to obtain the normalized output; The normalized output is processed sequentially through a feedforward neural network, residual connection processing, and layer normalization processing to capture the time domain features and frequency domain features of the sequence data.
[0057] In a specific application scenario, the training set data in the training set is input into the time-frequency transformer branch, and the self-attention and Fourier transform are combined to capture the time domain features and frequency domain features of the sequence data, as described below: Step a1: For the preprocessed time series data, the self-attention mechanism is used to calculate the relevance weight of each time step in the input sequence and generate a temporal context representation. The calculation formula is: Formula (9); in, Represents query (Query), key (Key) and value (Value), respectively. Indicates the dimension of the key vector, used for scaling, Express Perform a transpose operation.
[0058] Step a2: Similarly, for the preprocessed time series data, the frequency domain features of the input are calculated by fast Fourier transform. The calculation formula is Formula (10); in, Indicates the number of features.
[0059] Step a3: Extract the first k frequency components with the largest amplitudes from the frequency domain features to represent the main period of the input sequence. The calculation formula is: Formula (11); in, represents the time step, represents the amplitude sequence of frequencies, Indicates a period.
[0060] Step a4: Use continuous convolution operations with different convolution kernel sizes for each period result to extract period features and enhance the time-frequency information of the input.
[0061] Step a5: Using the mean amplitude of each period as the weight, perform weighted summation on all period results to achieve aggregation of period features.
[0062] Step a6: Perform weighted fusion of the time domain output and the frequency domain output according to the weight parameter to form a comprehensive feature. The calculation formula is: Formula (12); in, represents the fusion result, represents the frequency domain output, represents the time domain output, Represents the frequency domain weighting coefficient.
[0063] Step a7: The fusion result is added to the input sequence through residual connection, and then layer normalization is performed.
[0064] Step a8: Normalized output It passes through a feed-forward neural network and again through residual connections and layer normalization to ensure the stability and performance of the network.
[0065] like Figure 3 As shown in FIG, it is the fitting result diagram of the training set used by the prediction method provided by the embodiment of the present invention. Among them, the solid line represents the predicted value and the dotted line represents the actual value. Figure 3 It can be seen that the model can track the average daily oil production trend well throughout the entire time period, especially in the early and middle stages, showing high accuracy. It can well capture the peaks, valleys and trend changes in the data, proving its effective application in long-term oil production forecasting.
[0066] Step S103: Calculate the Q matrix based on the first prediction value, and calculate the K matrix and the V matrix based on the second prediction value.
[0067] In one example, calculating a Q matrix based on the first predicted value, and calculating a K matrix and a V matrix based on the second predicted value, includes the following steps: Performing calculation based on the first predicted value, the queried linear transformation weight matrix, and the fifth bias term to obtain a Q matrix; Calculate based on the linear transformation weight matrix of the key, the second prediction value and the sixth bias term to obtain a K matrix; The linear transformation weight matrix of the value, the second prediction value and the seventh bias term are calculated to obtain a V matrix.
[0068] In a specific application scenario, calculating the Q matrix based on the first predicted value, and calculating the K matrix and the V matrix based on the second predicted value, specifically includes the following steps: Step b1: Use the first prediction value output by the Bi-LSTM branch Calculate the Q matrix: Formula (13); in, The linear transformation weight matrix representing the query; represents the fifth bias term.
[0069] Step b2: Use the second prediction value output by the time-frequency Transformer branch Calculate K matrix and V matrix: Formula (14); Formula (15); in, The linear transformation weight matrix representing the key; The linear transformation weight matrix representing the value; represents the sixth bias term and Represents the seventh bias term.
[0070] Step b3: Calculate the attention score and normalize it, and perform weighted summation: Formula (16); Step S104: Through the multi-head attention mechanism, the Q matrix, K matrix and V matrix are fused to obtain the fused result.
[0071] In one example, the Q matrix, the K matrix, and the V matrix are fused through a multi-head attention mechanism to obtain a fused result, including the following steps: Through the multi-head attention mechanism, the fusion processing is performed based on the dimensions of the Q matrix, K matrix, V matrix, key vector and the transpose operation of K to obtain the fused result.
[0072] Step S105: After the fused result is converted through the linear layer, the optimal parameter model is obtained through iterative training.
[0073] In one example, obtaining an optimal parameter model through iterative training includes the following steps: Define optimization goals; Use smooth L1 loss as the loss function and Adam as the optimizer to obtain the optimal parameter model through iterative training.
[0074] Specifically, the optimal parameter model is obtained through iterative training, which includes the following steps: Define optimization goals; Using smoothed L1 loss as the loss function and Adam as the optimizer, we iteratively train the model using early stopping as a regularization strategy to prevent overfitting. Specifically, when the loss function converges and the validation set accuracy stops improving for a certain number of rounds, training is stopped, resulting in the optimal parameter model. The optimal parameters are the set of network weights and biases that demonstrate the strongest generalization ability on an independent validation set. These parameters optimize the network's predictions for the mean absolute error (MAE) and root mean squared error (RMSE).
[0075] In a specific application scenario, the calculation formula of the smooth L1 loss function is Formula (17); in, Represents the error between the predicted value and the true value.
[0076] Step S106: predicting the oil production of the target oil well using the optimal parameter model to predict the oil production of the target oil well.
[0077] In one example, after obtaining the optimal parameter model, the oil well production prediction method provided by the embodiment of the present invention may further include the following steps: The test set data in the test set is input into the optimal parameter model, the model performance of the optimal parameter model is evaluated, and the evaluation results are obtained.
[0078] like Figure 4 As shown in FIG, it is the fitting result diagram of the test set used by the prediction method provided by the embodiment of the present invention. Figure 4As can be seen from the figure, there is a certain deviation between the predicted and actual values in the first few months (months 0-5). However, as the months progress, the predicted values gradually approach the actual values and remain essentially consistent after 10 months. Overall, the model effectively captures the average daily oil production trends of a well in the Karamay Oilfield, demonstrating good stability and predictive ability.
[0079] In actual application scenarios, the solutions recommended by the machine are submitted to experts for review to verify the accuracy of the model, and the model is continuously adjusted and optimized based on the actual application results.
[0080] like Figure 5 As shown in FIG, it is a flowchart of a method for predicting oil well production in a specific application scenario. Figure 5 As shown, the method for predicting oil well production may include the following steps: Step S501: Collect historical oil production data of the target oil well.
[0081] Step S502: pre-process the data and divide the data set.
[0082] In actual application scenarios, the methods for preprocessing data specifically include: using piecewise cubic Hermite interpolation polynomials for interpolation and filling to solve the zero value problem in the data, introducing a three-point moving average filter to reduce the noise in the original production data, and using z-score normalization to enhance the stability of the data.
[0083] Step S503: input the training set data in the training set into the joint time-frequency Transformer branch and the Bi-LSTM network branch for processing.
[0084] In this step S503, the training set is input into the network, in which an adaptive fusion network of the joint time-frequency Transformer and BI-LSTM is designed. The time-frequency Transformer combines self-attention and Fourier transform to simultaneously capture the time domain and frequency domain features of the sequence data, enhancing the modeling ability of time series data. The Bi-LSTM captures the long-term dependencies of the time series data through forward LSTM and backward LSTM.
[0085] Step S504: Through the multi-head attention mechanism, the output features of the two branches are fused to obtain the fused result.
[0086] Step S505: After the fused result is converted through the linear layer, the optimal parameter model is obtained through iterative training.
[0087] Step S506: Evaluate the performance of the optimal parameter model using the test data in the test set. The accuracy and fitted curves in the experiment demonstrate that the model can effectively capture the average daily oil production trend of a well in the target oil field, demonstrating good stability and predictive ability, and outperforming other comparison methods.
[0088] The prediction method provided by the embodiment of the present invention can accurately predict the oil production of a target oil well. The prediction method provided by the embodiment of the present invention can perform long-term dynamic production forecasts for target oil wells in a target oil field, providing a decision-making basis for reservoir engineers and assisting in optimizing oil field production and improving quality and efficiency. Furthermore, the prediction method provided by the embodiment of the present invention not only fuses information from the time domain (self-attention path) and the frequency domain (Fourier path), enabling the model to capture global contextual information and extract and utilize periodic features; it also uses forward LSTM and reverse LSTM to capture long-term dependencies in time series data, improving the modeling capabilities for complex time series data.
[0089] In the above-mentioned embodiment, a method for predicting oil well production is provided. Accordingly, the present invention also provides an oil well production prediction device. The oil well production prediction device provided in the embodiment of the present invention can implement the above-mentioned oil well production prediction method. The oil well production prediction device can be implemented using software, hardware, or a combination of software and hardware. For example, the oil well production prediction device can include integrated or separate functional modules or units to perform the corresponding steps in each of the above-mentioned methods.
[0090] Please refer to Figure 6 , which shows a schematic diagram of an oil well production prediction device provided by some embodiments of the present invention. Since the device embodiments are substantially similar to the method embodiments, the description is relatively brief. For relevant details, please refer to the description of the method embodiments. The device embodiments described below are merely illustrative.
[0091] like Figure 6 As shown, the prediction device for oil well production may include: Acquisition module 601, used to obtain training set and test set; Processing module 602, configured to sequentially input training set data in the training set into the bidirectional long short-term memory network branch and the time-frequency converter branch for processing, and output a first predicted value and a second predicted value, wherein the training set data includes multivariate time series data of historical oil production of a target oil well; A calculation module 603 is configured to calculate a Q matrix based on the first prediction value, and calculate a K matrix and a V matrix based on the second prediction value; A fusion module 604 is used to fuse the Q matrix, the K matrix, and the V matrix through a multi-head attention mechanism to obtain a fused result; An iterative training module 605 is used to obtain an optimal parameter model through iterative training after converting the fused result through a linear layer; The prediction module 606 is used to predict the oil production of the target oil well by using the optimal parameter model to predict the oil production of the target oil well.
[0092] In some implementations of the embodiments of the present invention, the oil well production prediction device may include: Evaluation Module (in Figure 6 ), which is used to: after obtaining the optimal parameter model, input the test set data in the test set into the optimal parameter model, evaluate the model performance of the optimal parameter model, and obtain an evaluation result.
[0093] In some implementations of the embodiments of the present invention, the processing module 602 is specifically configured to: Inputting the training set data in the training set into the bidirectional long short-term memory network branch, processing it through the forward long short-term memory network and the reverse long short-term memory network in the bidirectional long short-term memory network branch to capture the long-term dependency of the time series data and obtain the output state; and The training set data in the training set is input into the time-frequency converter branch, and the self-attention and Fourier transform are combined to capture the time domain features and frequency domain features of the sequence data.
[0094] In some implementations of the embodiments of the present invention, the processing module 602 is specifically configured to: Process the training set data from front to back to obtain the forward hidden state; Process the training set data from back to front to obtain the reverse hidden state; The forward hidden state and the reverse hidden state are concatenated to capture the long-term dependencies of the time series data and obtain the output state.
[0095] In some implementations of the embodiments of the present invention, the processing module 602 is specifically configured to: Perform calculation based on the input, the weight matrix of the input gate, the first bias term, and the first activation function to obtain the input gate; Calculate based on the weight matrix of the forget gate, the input, the second bias term and the first activation function to obtain the forget gate; Calculate based on the weight matrix, input, third bias term and second activation function of the candidate memory unit to obtain a first memory unit; Calculate based on the weight matrix of the output gate, the input, the first activation function and the fourth bias term to obtain the output gate; Calculation is performed based on the output gate, the first memory unit, and the second activation function to obtain the forward hidden state.
[0096] In some implementations of the embodiments of the present invention, the processing module 602 is specifically configured to: The reverse hidden state is obtained based on the output gate corresponding to the data processed from back to front, the memory unit corresponding to the data processed from back to front, and the third activation function.
[0097] In some implementations of the embodiments of the present invention, the processing module 602 is specifically configured to: Get the training set data, which is the preprocessed time series data; For the preprocessed time series data, the self-attention mechanism is used to calculate the correlation weight of each time step in the input sequence to generate a temporal context representation; For the preprocessed time series data, the fast Fourier transform is used to calculate and obtain the frequency domain features of the input; Extract the first k frequency components with the largest amplitudes from the frequency domain features, so that the first k frequency components with the largest amplitudes represent the period of the input sequence; For the periodic results of each period, continuous convolution operations with different convolution kernel sizes are used to extract periodic features and enhance the input time-frequency information. The mean amplitude of each cycle is used as the weight to perform weighted summation on all cycle results to aggregate the cycle characteristics; Based on the weight parameters, the time domain output and the frequency domain output are weightedly fused to obtain the comprehensive features; The fusion result and the input sequence are added through residual connection to obtain the addition result; and the addition result is layer-normalized to obtain the normalized output; The normalized output is processed sequentially through a feedforward neural network, residual connection processing, and layer normalization processing to capture the time domain features and frequency domain features of the sequence data.
[0098] In some implementations of the embodiments of the present invention, the calculation module 603 is specifically configured to: Performing calculation based on the first predicted value, the queried linear transformation weight matrix, and the fifth bias term to obtain a Q matrix; Calculate based on the linear transformation weight matrix of the key, the second prediction value and the sixth bias term to obtain a K matrix; The linear transformation weight matrix of the value, the second prediction value and the seventh bias term are calculated to obtain a V matrix.
[0099] In some implementations of the embodiments of the present invention, the fusion module 604 is specifically configured to: Through the multi-head attention mechanism, the fusion processing is performed based on the dimensions of the Q matrix, K matrix, V matrix, key vector and the transpose operation of K to obtain the fused result.
[0100] In some implementations of the embodiments of the present invention, The acquisition module 601 may also be used to: before acquiring the training set and the test set, acquire a target data set, where the data in the target data set is the oil production data of the target oil well within a preset time period; The oil well production prediction device may further include: Preprocessing module (in Figure 6 (not shown), used to: preprocess each data in the target data set to obtain preprocessed data, and to form a preprocessed data set based on the preprocessed data; Divide the module (in Figure 6 ), which is used to divide the data in the preprocessed data set according to a preset ratio to obtain a training set and a test set.
[0101] In some implementations of the embodiments of the present invention, the iterative training module 605 is specifically configured to: Define optimization goals; Use smooth L1 loss as the loss function and Adam as the optimizer to obtain the optimal parameter model through iterative training.
[0102] In some implementations of the embodiments of the present invention, the oil well production prediction device provided by the embodiments of the present invention is based on the same inventive concept as the oil well production prediction method provided by the aforementioned embodiments of the present invention and has the same beneficial effects.
[0103] According to another embodiment, there is also provided a computer readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute a combination of Figure 1 The method described.
[0104] According to another embodiment, an electronic device is provided, comprising a memory and a processor, wherein the memory stores an executable code, and when the processor executes the executable code, the system realizes the combination of Figure 1 The method described.
[0105] Those skilled in the art will appreciate that, in one or more of the above examples, the functions described herein may be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0106] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solution of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for predicting oil well production, characterized in that: The method comprises: Get the training set and test set; inputting the training set data in the training set into the bidirectional long short-term memory network branch and the time-frequency converter branch in sequence for processing, and outputting a first prediction value and a second prediction value, wherein the training set data includes multivariate time series data of historical oil production of a target oil well; Calculating a Q matrix based on the first predicted value, and calculating a K matrix and a V matrix based on the second predicted value; The Q matrix, the K matrix, and the V matrix are fused through a multi-head attention mechanism to obtain a fused result; After transforming the fused result through a linear layer, an optimal parameter model is obtained through iterative training; The oil production of the target oil well is predicted by using the optimal parameter model to predict the oil production of the target oil well.
2. The prediction method according to claim 1, characterized in that After obtaining the optimal parameter model, the method further includes: The test set data in the test set is input into the optimal parameter model, and the model performance of the optimal parameter model is evaluated to obtain an evaluation result.
3. The prediction method according to claim 1, wherein: The step of sequentially inputting the training set data in the training set into the bidirectional long short-term memory network branch and the time-frequency converter branch for processing includes: Inputting the training set data in the training set into the bidirectional long short-term memory network branch, processing the data through the forward long short-term memory network and the reverse long short-term memory network in the bidirectional long short-term memory network branch to capture the long-term dependency of the time series data and obtain an output state; and The training set data in the training set is input into the time-frequency converter branch, and the time domain features and frequency domain features of the sequence data are captured by combining self-attention and Fourier transform.
4. The prediction method according to claim 3, characterized in that The training set data in the training set is input into the bidirectional long short-term memory network branch, and processed by the forward long short-term memory network and the reverse long short-term memory network in the bidirectional long short-term memory network branch to capture the long-term dependency of the time series data and obtain the output state, including: Processing the training set data from front to back to obtain a forward hidden state; Processing the training set data from back to front to obtain a reverse hidden state; The forward hidden state and the reverse hidden state are concatenated to capture the long-term dependency of the time series data, thereby obtaining the output state.
5. The prediction method according to claim 4, characterized in that The processing of the training set data from front to back to obtain a forward hidden state includes: Perform calculation based on the input, the weight matrix of the input gate, the first bias term, and the first activation function to obtain the input gate; Performing calculation based on the weight matrix of the forget gate, the input, the second bias term, and the first activation function to obtain a forget gate; Perform calculation based on the weight matrix of the candidate memory unit, the input, the third bias term, and the second activation function to obtain a first memory unit; Performing calculation based on a weight matrix of an output gate, the input, the first activation function, and a fourth bias term to obtain an output gate; Calculation is performed based on the output gate, the first memory unit, and the second activation function to obtain the forward hidden state.
6. The prediction method according to claim 4, characterized in that The processing of the training set data from back to front to obtain a reverse hidden state includes: The reverse hidden state is obtained by performing calculation based on the output gate corresponding to the back-to-forward processing data, the memory unit corresponding to the back-to-forward processing data, and the third activation function.
7. The prediction method according to claim 3, characterized in that Inputting the training set data in the training set into the time-frequency converter branch, and combining self-attention and Fourier transform to capture the time domain features and frequency domain features of the sequence data, including: Acquire the training set data, where the training set data is preprocessed time series data; For the preprocessed time series data, the relevance weight of each time step in the input sequence is calculated through the self-attention mechanism to generate a time series context representation; The pre-processed time series data is calculated by fast Fourier transform to obtain the input frequency domain features; Extracting the first k frequency components with the largest amplitudes from the frequency domain features, so that the first k frequency components with the largest amplitudes represent the period of the input sequence; For the periodic results of each period, continuous convolution operations with different convolution kernel sizes are used to extract periodic features and enhance the input time-frequency information. The mean amplitude of each cycle is used as the weight to perform weighted summation on all cycle results to aggregate the cycle characteristics; Based on the weight parameters, the time domain output and the frequency domain output are weightedly fused to obtain the comprehensive features; Adding the fusion result and the input sequence through a residual connection to obtain an addition result; and performing layer normalization processing on the addition result to obtain a normalized output; The normalized output is sequentially processed through a feedforward neural network, a residual connection process, and a layer normalization process to capture the time domain features and the frequency domain features of the sequence data.
8. The prediction method according to claim 1, wherein: The calculating of the Q matrix based on the first predicted value, and the calculating of the K matrix and the V matrix based on the second predicted value, include: Performing calculation based on the first predicted value, the queried linear transformation weight matrix, and a fifth bias term to obtain the Q matrix; Performing calculation based on the linear transformation weight matrix of the key, the second prediction value, and the sixth bias term to obtain the K matrix; The V matrix is obtained by performing calculation based on the linear transformation weight matrix of the value, the second prediction value and the seventh bias term.
9. The prediction method according to claim 1, characterized in that The Q matrix, the K matrix, and the V matrix are fused through the multi-head attention mechanism to obtain a fused result, including: Through the multi-head attention mechanism, fusion processing is performed based on the Q matrix, the K matrix, the V matrix, the dimension of the key vector and the transposition operation of K to obtain the fused result.
10. The prediction method according to claim 1, wherein: Before obtaining the training set and the test set, the method further includes: Acquire a target data set, wherein the data in the target data set is oil production data of a target oil well within a preset time period; Preprocessing each data in the target data set to obtain preprocessed data, and forming a preprocessed data set based on the preprocessed data; The data in the preprocessed data set is divided according to a preset ratio to obtain a training set and a test set.
11. The prediction method according to claim 1, wherein: The optimal parameter model is obtained through iterative training, including: Define optimization goals; Smooth L1 loss is used as the loss function and Adam is used as the optimizer. The optimal parameter model is obtained through iterative training.
12. A device for predicting oil well production, characterized in that: The device comprises: Acquisition module, used to obtain training set and test set; a processing module, configured to sequentially input training set data in the training set into a bidirectional long short-term memory network branch and a time-frequency converter branch for processing, and output a first predicted value and a second predicted value, wherein the training set data includes multivariate time series data of historical oil production of a target oil well; a calculation module, configured to calculate a Q matrix based on the first predicted value, and calculate a K matrix and a V matrix based on the second predicted value; A fusion module is used to fuse the Q matrix, the K matrix, and the V matrix through a multi-head attention mechanism to obtain a fused result; An iterative training module, configured to obtain an optimal parameter model through iterative training after converting the fused result through a linear layer; The prediction module is used to predict the oil production of the target oil well by using the optimal parameter model to predict the oil production of the target oil well.
13. A computer-readable storage medium, characterized in that A computer program is stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 11.
14. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Industrial internet prediction method and system based on frequency domain and long and short term feature fusion
CN115907154A
Radar echo space-time sequence depth prediction method based on Fourier neural operator
CN119047511A
TCN-LSTM wind power generation capacity prediction method combining time-frequency analysis and attention mechanism
CN119382103A
Hybrid data- and model-driven method for predicting remaining useful life of mechanical component
US20240289610A1