Oil reservoir data intelligent reconstruction method based on feedforward neural network
Through the intelligent reconstruction method of reservoir data based on feedforward neural network, the inaccuracy and lack of reservoir data caused by equipment failures during the development process is solved, and efficient intelligent filling and accuracy of data are achieved.
Patent Information
- Application Number
- CN202510064532.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-06-13
Smart Images

Figure CN120145802A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to an intelligent reservoir data reconstruction method based on a feedforward neural network. Background Art
[0002] During the process of reservoir development, due to various human operations or other factors, data inaccuracy and missing are inevitable. These factors may include equipment failures, measurement errors, environmental changes, etc., resulting in the incompleteness of dynamic and static monitoring data. Traditional methods for processing missing data, such as mean interpolation, mode interpolation, linear interpolation, etc., are inefficient in processing non-linear data and it is difficult to ensure the accuracy of the filled data. Therefore, developing an efficient reservoir data preprocessing and intelligent filling method is an important requirement in the field of petroleum engineering. Summary of the Invention
[0003] To solve the above technical problems existing in the prior art, the present invention proposes an intelligent reservoir data reconstruction method based on a feedforward neural network. By using the non-linear fitting ability of the neural network, the complex relationships between reservoir parameters are learned through training data, and then the missing or inaccurate data is intelligently filled, which can effectively improve the integrity and accuracy of reservoir data and provide an efficient data processing tool for the field of petroleum engineering.
[0004] To achieve the above object, the present invention provides an intelligent reservoir data reconstruction method based on a feedforward neural network, including:
[0005] Obtain the reservoir data to be processed, input the reservoir data to be processed into a feedforward neural network model for processing, output a prediction result, and fill the prediction result to obtain the reconstructed data;
[0006] Wherein, the feedforward neural network model is obtained through training with a training set and verification with a verification set. The training set is a reservoir data set; the feedforward neural network model includes an input layer, a hidden layer, and an output layer connected in sequence.
[0007] Preferably, constructing the training set and the verification set includes:
[0008] Preprocess reservoir data sets of different scales, and divide the preprocessed data sets according to a preset ratio to obtain the training set and the verification set; the reservoir data sets of different scales include a complete reservoir data set and an element missing data set, and the scale of the complete reservoir data set is smaller than that of the element missing data set.
[0009] Preferably, performing the preprocessing includes:
[0010] Feature selection: used to select specific input features as model inputs, with the target labels being the specific dynamic and static data of oil, water, and gas reservoirs that need to be predicted and completed;
[0011] Categorical feature processing: used to convert categorical features into numerical features using one-hot encoding;
[0012] Feature alignment: used to align the feature columns of different datasets to ensure consistent input dimensions;
[0013] Feature standardization: used to perform standardization processing on input features.
[0014] Preferably, constructing the training set further includes:
[0015] Using PyTorch to create a custom dataset class to encapsulate training samples and using DataLoader for data loading;
[0016] Among them, the custom dataset class is a custom dataset class defined using PyTorch, which encapsulates input features and target labels as training samples;
[0017] Performing the data loading includes loading data through PyTorch's DataLoader, which supports batch processing and randomized sampling.
[0018] Preferably, training the initial feedforward neural network model based on the training set includes:
[0019] Receiving the standardized data features through the input layer, processing the input data by the hidden layer, and obtaining the training result after passing through the output layer;
[0020] Among them, each hidden layer includes a ReLU activation function. The feedforward neural network model calculates the predicted value through forward propagation, updates the weights according to the error through backward propagation, and after several iterations, continuously optimizes the model until a trained feedforward neural network model is obtained, that is, the feedforward neural network model.
[0021] Preferably, the feedforward neural network model uses an Adam optimizer for adaptive learning rate adjustment and uses the mean squared error MSE as the loss function.
[0022] Preferably, the prediction result is a non-negative number.
[0023] Compared with the prior art, the present invention has the following advantages and technical effects:
[0024] (1) Improving data processing efficiency: Through automated feature extraction and selection, the workload of data preprocessing is reduced, improving efficiency;
[0025] (2)Enhance the robustness of the model: Multiple data imputation methods can adapt to different data characteristics and scenarios, enhancing the model's robustness to outliers and uncertainties;
[0026] (3)Improve prediction accuracy: Model-based imputation methods can consider the relationships between features and provide more accurate predictions;
[0027] (4)Reduce computational complexity: By selecting appropriate imputation methods, the computational complexity can be reduced while ensuring the quality of the results. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] The drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments and descriptions of this application are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0029] Figure 1 It is a flowchart of an intelligent reservoir data reconstruction method based on a feedforward neural network according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail this application.
[0031] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0032] The main problems faced in current reservoir data missing processing include the complexity of data preprocessing, which involves time-consuming and complex steps such as cleaning, handling missing values, detecting outliers, and data standardization. In addition, it is also a challenge to perform feature selection from a large amount of reservoir historical data to screen out features that have a significant impact on solving problems. At the same time, the process of extracting potential meaningful features from these data itself is also very complex and requires the use of machine learning algorithms and techniques.
[0033] This embodiment studies a set of reservoir data processing and intelligent filling methods based on a feedforward neural network for the problem of data missing in reservoir development. This method utilizes the non-linear fitting ability of the neural network to learn the complex relationships between reservoir parameters through training data, and then intelligently fills in the missing or inaccurate data. Through this method, the integrity and accuracy of reservoir data can be effectively improved, providing an efficient data processing tool for the petroleum engineering field.
[0034] The structure of a feedforward neural network (FNN) consists of an input layer, one or more hidden layers, and an output layer. Information flows in only one direction, from the input layer through the hidden layers and finally to the output layer, without feedback loops. This structure enables the neural network to learn the non-linear relationship between input data and output data and to handle complex pattern recognition and prediction tasks. Compared with linear fitting, the non-linear fitting of artificial neural networks has obvious advantages in dealing with non-linear relationships, being able to provide a lower mean squared error (MSE) and a higher coefficient of determination (R 2 ), thus achieving more accurate data fitting.
[0035] Based on the above technical status quo, this embodiment proposes an intelligent reservoir data reconstruction method based on a feedforward neural network, as Figure 1 follows:
[0036] Obtain the reservoir data to be processed, input the reservoir data to be processed into a feedforward neural network model for processing, and output the reconstructed data set;
[0037] Among them, the feedforward neural network model is obtained through training with a training set and verification with a verification set. The training set is a reservoir data set; the feedforward neural network model includes an input layer, a hidden layer, and an output layer connected in sequence.
[0038] This embodiment uses a feedforward neural network model to extract the high-order relationships between input features through the hidden layer. Through repeated iterative optimization, the predicted value can approximate the true value to the greatest extent. The multi-layer structure of the feedforward neural network model enables it to gradually learn the feature relationships from simple to complex, while the ReLU activation function ensures the training stability and efficiency of the model. During the optimization process, the Adam algorithm avoids the oscillation and slow convergence problems that may occur in traditional gradient descent methods by dynamically adjusting the learning rate.
[0039] Furthermore, constructing the training set and the verification set includes:
[0040] Preprocess reservoir data sets of different scales, and divide the preprocessed data sets according to a preset ratio to obtain a training set and a verification set; the data sets of different scales include a complete reservoir data set and a data set with missing elements. The scale of the complete reservoir data set is smaller than that of the data set with missing elements;
[0041] Specifically, integrate two data sets of different scales through feature selection, processing of categorical features, feature alignment and standardization, and data set division. The goal is to use the small-scale complete data set to predict and complete the missing dynamic and static data of oil, water, and gas in the large-scale data set.
[0042] Dataset Integration: First, read two datasets, a small-scale complete dataset and a larger but more missing dataset;
[0043] Feature Selection: Select specific input features (such as thickness, porosity, permeability, etc.) as model inputs, and the target labels are the specific dynamic and static data of oil, water, and gas reservoirs that need to be predicted and completed;
[0044] Categorical Feature Processing: Use one-hot encoding to convert categorical features (such as formation type) into numerical features for model processing;
[0045] Feature Alignment: Align the feature columns of the two datasets to ensure consistent input dimensions;
[0046] Feature Standardization: Standardize the input features so that they are distributed within a unified numerical range, which helps to accelerate model training and improve prediction accuracy;
[0047] Dataset Partitioning: Partition the training set and validation set on the small dataset to provide data support for subsequent neural network training and evaluation.
[0048] Furthermore, constructing the training set also includes:
[0049] Create a custom dataset class using PyTorch to encapsulate training samples and use DataLoader for data loading;
[0050] Among them, the custom dataset class is a custom dataset class defined using PyTorch, which encapsulates input features and target labels as training samples;
[0051] Performing the said data loading includes loading data through PyTorch's DataLoader, which supports batch processing and random sampling.
[0052] Furthermore, training an initial feedforward neural network model based on the training set includes:
[0053] Receive the standardized data features through the input layer, process the input data by the hidden layer, and obtain the training result after passing through the output layer;
[0054] Among them, each hidden layer is followed by a ReLU activation function. The feedforward neural network model calculates the predicted value through forward propagation, updates the weights according to the error through backward propagation, and after several iterations, continuously optimizes the model until a trained feedforward neural network model is obtained, that is, the feedforward neural network model.
[0055] Specifically, data is passed from the input layer to the neurons in the hidden layer. Each input value is multiplied by the corresponding connection weight, and then these weighted input values are passed to the neuron. After each hidden layer neuron receives the weighted input, an activation function is applied to introduce a non-linear transformation.
[0056] Suppose there is an input layer with an input vector x = [x 1 , x 2 , …, x n , a hidden layer containing m neurons, and an output layer. There is a weight w ij between each hidden layer neuron j and each neuron i in the input layer, and each hidden layer neuron has a bias b j .
[0057] For each neuron j in the hidden layer, its output h can be calculated through the following steps:
[0058] (a) Weighted sum: Calculate the dot product of the input vector x and the weight vector wj = [w 1j , w 2j , …, w nj , and then add the bias b j :
[0059]
[0060] In the formula, z j is the result of the weighted sum;
[0061] (b) Apply the activation function: Pass the result z j of the weighted sum through a non-linear activation function f to obtain the output h j of the hidden layer neuron j:
[0062]
[0063] The activation function in this embodiment is ReLU (Rectified Linear Unit): f(z) = max(0, z).
[0064] Therefore, the output vector h = [h 1 , h 2 , …, h m of the hidden layer can be expressed as: h = f(Wx + b);
[0065] where W is an m×n weight matrix and b is an m-dimensional bias vector.
[0066] This process can be generalized to the case of multiple hidden layers and continues until the output of the last hidden layer, which will be used as the input to the output layer for further processing.
[0067] Furthermore, the mean squared error is used as the loss function to measure the prediction error, and the Adam optimizer is adopted for adaptive learning rate adjustment to achieve fast convergence. Through iterative training of 100 epochs and performance evaluation on the validation set, the learning effect of the model is monitored and overfitting is prevented.
[0068] Loss function: The mean squared error (MSE) is used as the loss function in the training phase to measure the gap between the prediction result and the actual value. The calculation formula of MSE is as follows:
[0069]
[0070] where: n is the number of samples, y i is the actual observed value of the i-th sample, is the predicted value of the i-th sample.
[0071] Optimizer selection: The Adam algorithm is selected as the optimizer, which is a gradient descent method with adaptive learning rate adjustment and fast convergence.
[0072] Iterative training: The model is trained through multiple rounds of iteration (100 epochs), and the training data is traversed multiple times in each round.
[0073] Performance evaluation: The performance of the model is evaluated on the validation set. By comparing the training loss and the validation loss, the learning effect of the model is monitored and overfitting is prevented.
[0074] After the model training is completed, it is used to predict the missing values in a larger dataset. The prediction results are restricted to non-negative numbers and filled back into the original dataset to generate a complete imputed dataset:
[0075] Prediction application: After training, the model is used to predict the missing values in a larger dataset.
[0076] Non-negative restriction: The predicted values are restricted to non-negative numbers and filled into the original dataset to generate a complete imputed dataset.
[0077] Finally, the larger dataset after prediction and imputation is saved as a new Excel file, ensuring the scientificity and accuracy of the filling process.
[0078] The technical solution of this embodiment improves the data processing efficiency: Through automated feature extraction and selection, the workload of data preprocessing is reduced and the efficiency is improved. Multiple data imputation methods can adapt to different data characteristics and scenarios, enhancing the robustness of the model to outliers and uncertainties. The model-based imputation method can consider the relationships between features and provide more accurate predictions. By selecting appropriate imputation methods, the computational complexity can be reduced while ensuring the quality of the results.
[0079] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for intelligent reconstruction of reservoir data based on feedforward neural network, characterized in that: include: Acquire the reservoir data to be processed, input the reservoir data to be processed into a feedforward neural network model for processing, output a prediction result, fill the prediction result, and obtain reconstructed data; The feedforward neural network model is obtained by training with a training set and verifying with a verification set, wherein the training set is a reservoir data set; the feedforward neural network model comprises an input layer, a hidden layer and an output layer connected in sequence.
2. The method for intelligent reconstruction of reservoir data based on feedforward neural network according to claim 1, characterized in that: Constructing the training set and the validation set includes: Reservoir data sets of different sizes are preprocessed, and the preprocessed data sets are divided according to a preset ratio to obtain the training set and the validation set; the reservoir data sets of different sizes include a complete reservoir data set and an element missing data set, and the scale of the complete reservoir data set is smaller than the scale of the element missing data set.
3. The method for intelligent reconstruction of reservoir data based on feedforward neural network according to claim 2 is characterized in that: The pretreatment comprises: Feature selection: used to select specific input features as model inputs. The target labels are the dynamic and static data of specific reservoirs of oil, water, and gas that need to be predicted and completed. Categorical feature processing: used to convert categorical features into numerical features using one-hot encoding; Feature alignment: used to align feature columns of different data sets to ensure consistent input dimensions; Feature standardization: used to standardize input features.
4. The method for intelligent reconstruction of reservoir data based on feedforward neural network according to claim 2, characterized in that: Constructing the training set further includes: Use PyTorch to create a custom dataset class to encapsulate training samples and use DataLoader to load data; Among them, the custom dataset class is defined in PyTorch and put into the custom dataset class, encapsulating the input features and target labels as training samples; The data loading includes loading data through PyTorch's DataLoader, which supports batch processing and random sampling.
5. The method for intelligent reconstruction of reservoir data based on feedforward neural network according to claim 1, characterized in that: Training an initial feedforward neural network model based on the training set includes: The input layer receives the standardized data features, the hidden layer processes the input data, and obtains the training results after passing through the output layer; Among them, a ReLU activation function is included after each hidden layer. The feedforward neural network model calculates the prediction value through forward propagation, updates the weight according to the error through backward propagation, and continuously optimizes the model after several iterations until a trained feedforward neural network model is obtained, that is, the feedforward neural network model.
6. The method for intelligent reconstruction of reservoir data based on feedforward neural network according to claim 2, characterized in that: The feedforward neural network model uses the Adam optimizer to perform adaptive learning rate adjustment and uses the mean square error (MSE) as the loss function.
7. The method for intelligent reconstruction of reservoir data based on feedforward neural network according to claim 1, characterized in that: The prediction result is a non-negative number.