Hierarchical personalized federal learning method for oil field production prediction
By adopting a hierarchical personalized federal learning method in oil field production forecasting, privacy-protected oil well clustering and dual-model federal optimization are solved, and high-precision, stable and personalized production forecasts are achieved.
Patent Information
- Application Number
- CN202510644833.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-05-20
AI Technical Summary
The prior art is difficult to effectively deal with the differences in geological conditions and data heterogeneity between different oil wells in oil field production prediction, resulting in low model prediction accuracy and difficulty in taking into account global generalization performance and local personalization needs.
A hierarchical personalized federated learning method is proposed to improve the accuracy and stability of production prediction through privacy-protected oil well clustering, dual-model federal optimization and cluster adaptive prediction, combining hierarchical clustering and personalized model collaborative training.
While protecting data privacy, this method improves the convergence speed and accuracy of the prediction model, can effectively balance global generalization with local personalization, and adapt to the diversified production characteristics of heterogeneous oil wells.
Smart Images

Figure CN120163479A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of the integrated application of petroleum engineering and artificial intelligence, and particularly to a hierarchical personalized federated learning method for oilfield production prediction. Background Art
[0002] The prediction of oilfield production is a core link in the development and management of oil and gas fields, and is crucial for optimizing development strategies and improving oil production efficiency. At present, deep learning methods based on data-driven have been widely used in the field of production prediction, but they rely on large-scale training data, and it is easy to lead to the problem of low prediction accuracy of the model when the production data of a single well is limited. The traditional centralized training method requires collecting the original data of each production unit, which not only faces high data transmission and storage costs, but also has the risk of data privacy leakage. To achieve collaborative modeling among oil wells without centrally sharing the original data, federated learning (FL) is introduced as an effective solution. In a typical federated learning framework (such as FedAvg), a global model is constructed by weighted averaging the local model parameters of each node.
[0003] However, in the scenario of oilfield production data, due to significant differences in geological conditions, operation modes, and production stages of different oil wells, the data of each node presents non-independent and identically distributed (Non-IID) characteristics, which will significantly affect the convergence speed and prediction accuracy of standard algorithms such as FedAvg. To alleviate the performance degradation problem caused by Non-IID data, improved algorithms based on proximal terms such as FedProx are proposed. By adding a regularization term to the local objective function, the deviation of the local model update from the global model is restricted, thereby improving the training stability. However, such methods still only rely on a single global model, and it is difficult to simultaneously take into account the global generalization performance and the local personalized prediction requirements of each oil well. Moreover, they mainly focus on data heterogeneity in the statistical sense and fail to connect data heterogeneity with geological differences, such as differences in reservoir reserves and structural characteristics, which actually dominate the production characteristics of different oil wells.
[0004] There is still a lack of a systematic solution in the prior art for combining hierarchical clustering and personalized model collaborative training in the scenario of oilfield production prediction. Therefore, there is a need for a federated learning method that can not only protect data privacy but also perform hierarchical clustering and personalized model training for the Non-IID characteristics of production data among oil wells to improve the accuracy and stability of production prediction. Summary of the Invention
[0005] In order to overcome the above problems existing in the prior art, the present invention proposes a hierarchical personalized federated learning method for oilfield production prediction.
[0006] The technical solution adopted by the present invention to solve its technical problems is: a hierarchical personalized federated learning method for oilfield production prediction, specifically including: Step 1, privacy-preserving well clustering: Each well extracts production structure features from its private production data and encrypts the production structure features; the encrypted production structure features are uploaded to the server, and the server uses a hierarchical clustering algorithm to group wells with similar encrypted features. Step 2, dual-model federated optimization: The server aggregates a global model for capturing cross-cluster common knowledge and multiple cluster proxy models at the same time; each well updates the local version of the global model by minimizing a regularized loss function. Step 3, cluster adaptive prediction: For a new well to be predicted, it is assigned to the corresponding cluster according to its production structure features, and then the corresponding cluster proxy model is used to generate the production prediction result of the well.
[0007] In the above-mentioned hierarchical personalized federated learning method for oilfield production prediction, the Mann-Kendall test and Theil-Sen estimation method are used in Step 1 to extract the production pattern , and the production structure features of well k are expressed as: ; where is the cumulative production time, is the average production, is the standard deviation, is the coefficient of variation, is the skewness, is the kurtosis.
[0008] In the above-mentioned hierarchical personalized federated learning method for oilfield production prediction, differential privacy technology is used to encrypt the production structure features in Step 1, specifically: adding Laplace noise to the production structure features: ; where represents Laplace noise with a scale parameter of , is the sensitivity, is the privacy budget; the trade-off between privacy protection and prediction efficiency is achieved by adjusting the privacy budget value.
[0009] In the above-mentioned hierarchical personalized federated learning method for oilfield production prediction, Step 2 specifically includes: The global optimization objective is: ; where is the aggregation weight, satisfying ; Each oil well updates a local version of the global model by minimizing based on its local dataset ; For the learned model on the local dataset the expected loss; denotes the loss function; After rounds of local updates, each client sends its updated local global model to the server for aggregation: ; where N represents the total number of oil wells; k is the index of the oil well; The cluster proxy model is represented as: ; where K is the number of clusters; After aggregation, the server distributes the global model and the corresponding cluster proxy models to each client for local training; a regularization term is introduced into the optimization objective of each oil well to penalize the deviation from the corresponding cluster proxy model.
[0010] In the above hierarchical personalized federated learning method for oilfield production prediction, the introduction of the regularization term is specifically: each oil well minimizes the modified objective function : ; where is the weight coefficient, which determines the weight proportion of the regularization term in the objective function.
[0011] In the above hierarchical personalized federated learning method for oilfield production prediction, step 3 specifically includes: the new oil well to be predicted extracts its own production structure features based on the existing production data, encrypts them using the Laplace differential privacy mechanism and uploads them to the server. The server calculates the distance between its production structure features and the production structure features of the existing cluster center points, selects the one with the minimum distance as the cluster to which the new production well belongs, and uses the corresponding cluster proxy model to predict the production of the new production well.
[0012] The beneficial effects of the present invention are as follows. Compared with traditional data-centered training, it avoids data collection costs. Since the data remains local, the present invention does not disclose data privacy information. Compared with benchmark federated learning algorithms such as FedAvg and FedProx, it improves the convergence speed of the prediction model training; it provides an accurate, privacy-protected, and distributed training platform for various oil well production prediction models that is not affected by data heterogeneity. The personalized prediction strategy not only improves the prediction accuracy at the single oil well level but also ensures the consistency of the prediction performance for different oil wells, promotes model prediction fairness, and can effectively balance the relationship between global generalization and local personalization, thereby capturing and adapting to the diverse production characteristics inherent in heterogeneous oil wells. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Figure 1 is a schematic diagram of the overall framework of hierarchical personalized federated learning of the present invention; Figure 2 is a schematic diagram of the comparison of the convergence speed between the present invention and existing federated learning algorithms; Figure 3 are the verification results of different RNN models under the present invention and existing learning algorithms; among them, (a) is the verification result of LSTM under various federated learning algorithms; (b) is the verification result of GRU under various federated learning algorithms; Figure 4 is a schematic diagram of the prediction results of different federated learning algorithms for test wells in the embodiment of the present invention; among them, (a) is the average value of the evaluation index on the data of 6 test wells, and (b) is the standard deviation of the evaluation index on the data of 6 test wells. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] To enable those skilled in the art to better understand the technical solution of the present invention, the present invention will be described in detail below in conjunction with the drawings and specific embodiments.
[0015] As Figure 1 shown, this embodiment discloses a hierarchical personalized federated learning method for oilfield production prediction. The federated learning algorithm in this embodiment is represented by HierPFL and adopts a client-server architecture, where each oil well serves as a client. Each client independently calculates the encrypted production structure feature (PSF) and updates its local model through an optimization objective with regularization. The central server then clusters the clients based on the encrypted PSF and coordinates a dual aggregation process: including a global model that fuses cross-cluster knowledge and a cluster proxy model that captures the personalized features within each cluster.
[0016] The technical solution of HierPFL is divided into three steps: Step 1, Privacy-Preserving Well Clustering: Each client extracts production structure features (PSF) from its private production data, encrypts these features using the Laplace differential privacy mechanism, and then uploads the encrypted PSF to the server. The server uses a hierarchical clustering algorithm to group wells with similar encrypted features.
[0017] Step 2, Dual-Model Federated Optimization: The server aggregates a global model for capturing cross-cluster common knowledge and multiple cluster proxy models simultaneously. Each client updates the local version of the global model by minimizing a regularized loss function, where the regularization term limits the deviation between the local model and the corresponding cluster proxy model.
[0018] Step 3, Cluster-Adaptive Prediction: For a new well to be predicted, it is first assigned to the corresponding cluster based on its PSF, and then the corresponding cluster proxy model is used to generate the production prediction result for this well.
[0019] I. Privacy-Preserving Well Clustering Due to considerations of data security and collection costs, it is difficult to directly use the original production data for clustering. Therefore, production structure features (PSF) are first extracted from the production data of each well to ensure that these features can reflect both the statistical characteristics and production patterns of the data. Among them, the statistical features include: cumulative production time , average production , standard deviation , coefficient of variation , skewness and kurtosis etc. Considering the sparsity of well production data and the interference of data noise, the present invention uses the Mann-Kendall test and Theil-Sen estimation method to extract the production pattern . Then the production structure features of well k can be expressed as: ; Since the PSF still contains the production data distribution information of each well, the client needs to encrypt the PSF locally before uploading it to the server. The present invention uses differential privacy technology to achieve a balance between privacy protection and computational efficiency, and the specific implementation method is as follows: The client adds Laplace noise to the PSF: ; Among them, represents Laplace noise with a scale parameter of , is the sensitivity, is the privacy budget. A smaller means stronger privacy protection requirements, thus introducing greater noise, which may reduce the accuracy of the results. The present invention can achieve a trade-off between privacy protection and prediction performance by adjusting the value.
[0020] After receiving the encrypted PSF, the server groups the oil wells using a hierarchical clustering method. This process aims to identify the similarities between oil wells, thereby alleviating the impact of data heterogeneity on federated learning. The number of clusters needs to be consistent with the inherent characteristics of the production data and the training requirements of federated learning. The determination of the number of clusters in the present invention follows the following principles: 1. Basic clustering principle: The similarity between oil wells within each cluster should be as large as possible; at the same time, there should be significant differences between oil wells in different clusters.
[0021] 2. Sample balance principle: Each cluster should contain an equal number of oil wells. This principle can prevent a serious skew in the number of cluster members, thereby reducing the risk of model training in federated learning being biased towards certain specific clusters.
[0022] 3. Expert knowledge principle: Based on previous research on reservoir production patterns, the production rate can usually be divided into types such as violent decline, rapid decline, slow decline, stepwise decline, and fluctuation. This classification helps to align the clustering process with existing domain knowledge and improve the model's ability to capture relevant production behaviors.
[0023] II. Hierarchical personalized federated learning The global optimization objective of federated learning (FL) is: ; where is the aggregation weight, satisfying . Each client updates a local version of the global model by minimizing . is the expected loss of the learned model on the local dataset : ; where represents the loss function.
[0024] After rounds of local updates, each client sends its updated local global model to the server for aggregation: .
[0025] Based on After clustering the oil wells, it can be reasonably assumed that the production data of the oil wells within the same cluster approximately satisfy the independent and identically distributed (IID) assumption. Therefore, aggregating the cluster proxy models for the oil wells within the same cluster can effectively solve the problem of prediction accuracy differences caused by data heterogeneity.
[0026] The difference between this embodiment and the prior art is that: in addition to aggregating the global model, the server also aggregates multiple cluster proxy models: ; where is the number of clusters.
[0027] After the aggregation is completed, the server distributes the global model and the corresponding cluster proxy models to each client for local training. However, the non-independent and identically distributed (non-IID) characteristics of the production data will exacerbate the differences after the local model updates, thereby hindering model convergence and further degrading the overall performance. To prevent this divergence phenomenon, the present invention introduces a regularization term into the optimization objective of each client to penalize the deviation from the corresponding cluster proxy model.
[0028] Specifically, each client no longer only minimizes the standard local loss , but approximately minimizes the modified objective function : .
[0029] III. Adaptive Prediction for New Wells The newly produced well extracts its own production structure feature PSF based on the existing production data, encrypts it using the Laplace differential privacy mechanism and uploads it to the server. The server calculates the distance between its PSF and the PSF of the existing cluster center points, and selects the one with the minimum distance as the cluster to which the newly produced well belongs. Then, the corresponding cluster proxy model is used to predict the production of the newly produced well.
[0030] Based on the above method, this embodiment collected the production records of 172 oil wells distributed in 7 reservoir blocks (as shown in Table 1) as the dataset. To evaluate the performance of HierPFL, 6 wells with the shortest cumulative production time were designated as test wells. The remaining 166 wells have complete production history data for model training and effect evaluation.
[0031] Table 1 Summary of Production Data
[0032] The training configuration is as follows: the learning rate is set to 0.01, the batch size is 128, 5 local training epochs are performed in each round of communication, and a total of 100 rounds of global models are trained. The Adam optimizer is used to promote the stable and efficient convergence of the model.
[0033] To simulate the deployment of the model on newly developed oil wells, the oil well with the shortest cumulative production time was designated as the test set. The historical production data of the remaining oil wells was divided into a training set and a validation set at a ratio of 8:2. Four widely used error metrics were adopted to evaluate the prediction performance of the model: mean squared error (MSE), root mean squared error (RMSE), mean absolute error (MAE), and coefficient of determination ( ).
[0034] To evaluate the effectiveness of the proposed HierPFL framework, comparative experiments were conducted with two widely adopted federated learning benchmark methods: FedAvg and FedProx. In addition, to evaluate the stability of the framework, several recurrent neural network (RNN) architectures for time series prediction were implemented, including long short-term memory networks (LSTM) and gated recurrent units (GRU). These experiments helped to comprehensively compare the performance advantages and disadvantages of the HierPFL framework relative to traditional federated learning methods and different time series prediction models. The evaluation results are as Figure 3 shown. As can be seen from Figure 3 , HierPFL consistently outperformed FedAvg and FedProx in all model types. Especially when using the GRU model, the performance improvement of HierPFL was the most significant. Compared with FedProx, the RMSE decreased by 7% and the MAE decreased by 20%, and the coefficient of determination increased by 4%.
[0035] The convergence effects of the federated learning algorithm of the present invention, FedAvg, and FedProx were evaluated. The evaluation results are as Figure 2 shown. Both HierPFL and FedProx showed a faster convergence rate than FedAvg and reached convergence at around the 20th round of training. This acceleration can be attributed to the regularization term in the local objective, which effectively reduced the local update bias caused by data heterogeneity. In addition, the convergence rate of HierPFL was slightly faster than that of FedProx, indicating that its clustering strategy further alleviated the negative impact brought by non-independent and identically distributed (Non-IID) data, thus improving the overall training efficiency.
[0036] In addition, the oil wells in the test set were predicted, and the obtained prediction results are as Figure 4 shown. As can be seen from Figure 4It can be seen that HierPFL consistently outperforms FedAvg and FedProx in almost all evaluation metrics. Notably, HierPFL achieves the lowest standard deviation among all methods, highlighting its superior robustness and stability. These results indicate that the personalized prediction strategy of HierPFL not only improves the prediction accuracy at the individual well level but also promotes fairness by ensuring consistent performance across different wells. Additionally, the research findings validate that HierPFL can effectively balance the relationship between global generalization and local personalization, thereby capturing and adapting to the diverse production characteristics inherent in heterogeneous wells.
[0037] The above embodiments are only exemplary embodiments of the present invention and are not used to limit the present invention. Those skilled in the art can make various modifications or equivalent replacements to the present invention within the essence and protection scope of the present invention, and such modifications or equivalent replacements should also be regarded as falling within the protection scope of the present invention.
Claims
1. A hierarchical personalized federated learning method for oilfield production prediction, characterized in that: Specifically include: Step 1: Privacy-preserving oil well clustering: Each oil well extracts production structure features from its private production data and encrypts the production structure features; the encrypted production structure features are uploaded to the server, and the server uses a hierarchical clustering algorithm to group oil wells with similar encrypted features; Step 2, dual-model federated optimization: The server simultaneously aggregates a global model for capturing common knowledge across clusters and multiple cluster proxy models; each oil well updates the local version of the global model by minimizing a regularized loss function; Step 3, cluster adaptive prediction: For new oil wells that need to be predicted, they are assigned to the corresponding clusters according to their production structure characteristics, and then the corresponding cluster agent model is used to generate the production prediction results of the oil wells.
2. A hierarchical personalized federated learning method for oilfield production prediction according to claim 1, characterized in that: In step 1, the Mann-Kendall test and Theil-Sen estimation method are used to extract the production model , the production structure characteristics of oil well k are expressed as: ; in, is the cumulative production time, is the average output, is the standard deviation, is the coefficient of variation, is the skewness, is the kurtosis.
3. A hierarchical personalized federated learning method for oilfield production prediction according to claim 1, characterized in that: In step 1, differential privacy technology is used to encrypt the production structure features, specifically: adding Laplace noise to the production structure features: ; in, The scale parameter is The Laplace noise, It's sensitivity. is the privacy budget; the trade-off between privacy protection and prediction performance is achieved by adjusting the privacy budget value.
4. A hierarchical personalized federated learning method for oilfield production prediction according to claim 1, characterized in that: The step 2 specifically includes: The global optimization goal is: ; in, is the aggregation weight, satisfying ; Each oil well is based on its local data set by minimizing To update a local version of the global model ; For the learned model In local dataset Expected loss on represents the loss function; In progress After a round of local updates, each client will update its local global model Send to server for aggregation: ; Where N represents the total number of oil wells; k is the oil well index; The cluster agent model is expressed as: ; Where K is the number of clusters; After the aggregation is completed, the server distributes the global model and the corresponding cluster proxy model to each client for local training; a regularization term is introduced into the optimization objective of each oil well to penalize the deviation between it and the corresponding cluster proxy model.
5. A hierarchical personalized federated learning method for oilfield production prediction according to claim 4, characterized in that: The specific introduction of the regularization term is: each oil well minimizes the modified objective function : ; in, is the weight coefficient, which determines the weight of the regularization term in the objective function.
6. A hierarchical personalized federated learning method for oilfield production prediction according to claim 1, characterized in that: The step 3 specifically includes: extracting the production structure characteristics of the new oil well that needs to be predicted based on the existing production data, encrypting it using the Laplace differential privacy mechanism and uploading it to the server, the server calculating the distance between its production structure characteristics and the production structure characteristics of the existing cluster center point, selecting the cluster with the smallest distance as the cluster to which the new production well belongs, and using the corresponding cluster proxy model to predict the production of the new production well.
Citation Information
Patent Citations
Wireless traffic prediction method and system based on personalized grouping federated learning
CN115860153A
Personalized federal learning and recognition method and system based on differential privacy
CN115952533A
Layered personalized federal learning method and device in edge computing network and medium
CN116579417A
Clustering and knowledge distillation-based credible personalized federal learning method and device
CN116862024A
Personalized federal learning method and device with privacy protection and medium
CN117094382A