A pumping unit well multi-source semi-supervised class incremental working condition recognition method and system

By applying attention mechanisms and multi-source data distillation learning methods to identify the operating conditions of pumping wells, the limitations of single information sources and multi-source data fusion are overcome, achieving efficient and robust operating condition identification in incremental-like scenarios and improving the accuracy and efficiency of oilfield production.

CN120180382BActive Publication Date: 2025-11-28SHANDONG UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510256300.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-11-28
Estimated Expiration
2045-03-05

AI Technical Summary

Technical Problem

Existing methods for identifying the operating conditions of pumping wells suffer from limitations in single-source identification, issues with the accuracy of feature parameters, limitations in multi-source data fusion technology, difficulties in obtaining labeled samples, and poor identification performance in incremental learning scenarios, making it difficult to achieve efficient and robust operating condition identification.

Method used

Employing attention-based multi-source fusion technology, incremental learning, and semi-supervised learning, this method dynamically fuses measured ground dynamometer and measured electrical dynamometer features through a Squeeze-and-Excitation attention mechanism. Combined with multi-source data distillation learning using Kullback-Leibler divergence and a logistic regression classifier, it achieves efficient and robust identification of multi-source data.

Benefits of technology

It has improved the accuracy and robustness of pumping unit well condition identification, enhanced adaptability to complex and variable operating conditions, optimized oilfield production scheduling and fault prevention capabilities, and promoted the construction and development of smart oilfields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120180382B_ABST
    Figure CN120180382B_ABST
Patent Text Reader

Abstract

The application discloses a pumping unit well multi-source semi-supervised class incremental working condition recognition method and system and relates to the technical field of pumping unit well working condition recognition.The application is characterized in that the method comprises the following steps: step 1, respectively storing sample libraries of measured ground dynamometer cards and measured electric dynamometer cards; step 2, respectively constructing graph neural network teacher models for the two data sources of the measured ground dynamometer cards and the measured electric dynamometer cards; step 3, dynamically fusing prediction probabilities of the teacher models; step 4, performing multi-source data distillation learning; and step 5, performing semi-supervised working condition recognition by using a label propagation algorithm improved by a logistic regression classifier.Through the pumping unit well multi-source semi-supervised class incremental working condition recognition method and system, the multi-source fusion technology based on an attention mechanism, class incremental learning and semi-supervised learning are simultaneously applied to pumping unit well class incremental working condition recognition, a small amount of multi-source labeled working condition samples are fully utilized, and a large amount of multi-source unknown working condition samples are combined to realize more efficient, robust and practical pumping unit well working condition recognition.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of pumping unit well working condition recognition, and particularly relates to a pumping unit well multi-source semi-supervised class incremental working condition recognition method and system. BACKGROUND

[0002] In the field of oilfield operations, real-time monitoring and identification of pumping unit well working conditions are crucial for preventing equipment failures, optimizing production scheduling, and improving resource utilization. Existing pumping unit well working condition recognition methods mainly include the following:

[0003] 1. Recognition based on dynamometer card. The recognition method based on dynamometer card mainly relies on pump dynamometer card (model calculated dynamometer card) or measured ground dynamometer card, and combines artificial intelligence method for working condition recognition. For example, the technical solution recorded in the Chinese invention patent with the application number 202411505865.3 and the application date of October 28, 2024, entitled "Pumping unit working condition intelligent analysis method and system", and the technical solution recorded in the Chinese invention patent with the application number 202310347518.1 and the application date of April 3, 2023, entitled "Training method of working condition recognition model, working condition diagnosis method and device".

[0004] 2. Recognition based on electrical parameters. Working condition recognition based on electrical dynamometer card belongs to the category of working condition recognition based on electrical parameters, and working condition recognition based on electrical dynamometer card includes the following two types: (1) Model conversion electrical dynamometer card. (2) Measured electrical dynamometer card, such as the technical solution recorded in the Chinese invention patent with the application number 202311441260.8 and the application date of December 1, 2023, entitled "Pumping unit well measured electrical dynamometer card working condition recognition method and system".

[0005] 3. Recognition based on multi-source information. Working condition recognition based on multi-source information mainly combines dynamometer card, oil well production information (such as well condition data, production, pumping parameters, etc.) and other multiple information sources for working condition recognition. For example, the technical solution recorded in the Chinese invention patent with the application number 202311284299.3 and the application date of September 30, 2023, entitled "Fault diagnosis method, computing device and readable storage medium".

[0006] The above-mentioned working condition recognition method has the following limitations: (1) single information source recognition limitation: in a complex nonlinear system of mechanical, electrical and hydraulic coupling, relying on a single information source to judge the working condition of the pumping unit well is prone to false positives; (2) feature parameter precision problem: the calculation of pump indicator diagram and electrical diagram may be subject to damping coefficient and "zero division" problem, which affects the precision calculation of feature parameter value; (3) multi-source data fusion technology limitation: due to the technical limitations of traditional multi-feature connection recognition method, the complexity and variability of well conditions, and the uncertainty of artificial statistical production data, the existing working condition recognition method based on multi-source data needs to be improved in recognition effect and model robustness; (4) difficulty in obtaining labeled samples: most working condition recognition methods rely on a large number of labeled training samples, but in actual engineering, obtaining labeled samples is difficult and expensive, and the method of training without labeled samples often has poor recognition accuracy; (5) as the oilfield operation continues, the working condition categories of pumping unit wells continue to expand, gradually evolving into a typical class-incremental learning scenario. In this scenario, the emergence of new working condition categories requires the model not only to adapt to the introduction of new categories, but also to maintain the recognition ability of old working condition categories. This poses new challenges to traditional working condition recognition methods.

[0007] Under the background of big data oil production, the pumping unit well production system can obtain a large amount of working condition information containing various measured information sources such as ground indicator diagram and electrical parameters, and also obtains a large amount of multi-source unknown working condition information. In the face of the class-incremental actual production running scenario caused by the complexity and variability of working condition categories, how to effectively utilize the various measured information sources of pumping unit wells, and realize more efficient, robust and practical working condition recognition by fully combining a large amount of multi-source unlabeled working condition samples under a small amount of multi-source labeled working condition samples, so as to provide scientific, timely and accurate guidance for oil production decision-making, has become a key problem to be solved in the intelligent oilfield production construction and development. Therefore, developing a pumping unit well multi-source semi-supervised class-incremental working condition recognition method and system has important theoretical and application value for effectively solving the above key problems and promoting the intelligent oilfield production construction and development. SUMMARY

[0008] The technical problem to be solved by the present application is to overcome the shortcomings of the prior art and provide a pumping unit well multi-source semi-supervised class-incremental working condition recognition method and system. The multi-source fusion technology based on attention mechanism, class-incremental learning and semi-supervised learning are applied to pumping unit well class-incremental working condition recognition at the same time, and a small amount of multi-source labeled working condition samples are fully utilized in combination with a large amount of multi-source unknown working condition samples to realize more efficient, robust and practical pumping unit well working condition recognition.

[0009] The technical scheme adopted by the present application to solve the technical problem is that the pumping unit well multi-source semi-supervised class-incremental working condition recognition method comprises the following steps:

[0010] Step 1, respectively, store the sample library containing the corresponding measured ground dynamometer card, measured electrical dynamometer card of the pumping well working condition with and without label;

[0011] Step 2, respectively, construct the graph neural network teacher model for the measured ground dynamometer card and the measured electrical dynamometer card two data sources;

[0012] Step 3, dynamically fuse the prediction probability of each teacher model by using Squeeze-and-Excitation attention mechanism;

[0013] Step 4, use Kullback-Leibler divergence for multi-source data distillation learning;

[0014] Step 5, use the improved label propagation algorithm of the logistic regression classifier for semi-supervised working condition recognition.

[0015] Preferably, step 3 specifically includes the following steps:

[0016] Step 3-1, refer to Squeeze-and-Excitation attention mechanism to realize multi-source feature representation fusion module, which is composed of full connection layer- ReLU activation layer-full connection layer-Sigmoid activation layer, and the prediction probability output by the vth data source branch teacher model is represented as The corresponding attention score ω v is represented as:

[0017]

[0018] Wherein, sigmoid(·) is Sigmoid activation function, relu(·) is ReLU activation function, W1 and W2 are the weights of two full connection layers respectively;

[0019] Step 3-2, weight the obtained attention weight to the prediction result output by the multi-source teacher model, and obtain the fused multi-source teacher model prediction result p t is represented as:

[0020]

[0021] Wherein, p t is the fused multi-source teacher model prediction result, which contains the labeled data prediction result p t (L t ) and the unlabeled data prediction result p t (U t ).

[0022] Preferably, step 4 specifically includes the following steps:

[0023] Step 4-1, during learning task t, the model G t is the model G t-1 of task t-1 as a student model t , the distillation loss L KL between the output prediction probabilities of the teacher model and the student model is calculated by KL divergence:

[0024]

[0025] where T t-1 is the number of learned classes after task t-1, represents the prediction probability output by the multi-source teacher model of task t-1, represents the prediction probability output by the student model of task t.

[0026] Step 4-2, the total loss function L t of task t is composed of the cross-entropy loss function L c and the multi-source data distillation loss L KL , the cross-entropy loss function L c is represented as:

[0027]

[0028] where c t represents the number of sample classes of task t, p t (L t ) represents the model prediction result after fusion of the feature representation of the labeled sample in task t, represents the true label of the i-th labeled sample in task t.

[0029] Step 4-3, the total loss function L t is calculated as:

[0030] L t = L c + λL KL

[0031] where L c represents the cross-entropy loss function, L KL is the distillation loss between the output prediction probabilities of the teacher model and the student model, and λ is the weight of the multi-source data distillation loss function.

[0032] Preferably, step 5 specifically includes the following steps:

[0033] Step 5-1, the features output by the graph convolutional networks of different data sources are mean fused, and the feature matrix X t of task t after fusion is calculated as:

[0034]

[0035] wherein, X t contains the labeled data global feature matrix X L and the unlabeled data global feature matrix X U wherein: X t = [X L , X U ];

[0036] Step 5-2, the multi-source semi-supervised graph classification method based on enhanced label propagation is represented as:

[0037]

[0038] wherein, is the adjacent matrix constructed using the fused features, contains the adjacent matrix constructed from the labeled data and the adjacent matrix constructed from the unlabeled data The calculation formula is:

[0039]

[0040] wherein, represents the probability that the kth sample belongs to class c, The calculation formula is:

[0041]

[0042] wherein, y k represents the label of the kth sample;

[0043] Step 5-3, an initial prediction probability distribution is generated for the unlabeled data using a logistic regression classifier, which provides a more accurate initial label vector for label propagation. After completing the label propagation process, multiple logistic regression classifiers are used to calculate P LR Classify the label embedding vectors and calculate the average of these classifier results to obtain the final classification label. The final label y t The calculation formula is:

[0044]

[0045] wherein, is the label vector of task t after label propagation.

[0046] Preferably, step 2 specifically comprises the following steps:

[0047] Step 2-1, for each sample in a single data source, calculate the distance between it and all other samples in its data set, and select the two samples with the smallest distance as its nearest neighbor samples, construct the initial adjacency matrix of the graph convolution network, and the calculation formula of the graph convolution network of the t-th task is:

[0048]

[0049] wherein, is the adjacency matrix constructed for the v-th data source in task t, is the data of the v-th data source in task t, including the labeled data and unlabeled data of the v-th data source;

[0050] Step 2-2, use the convolution layer of the graph convolution network to extract node features:

[0051]

[0052] wherein, I K is the identity matrix, relu(·) is the ReLU activation function, and W is the trainable parameter matrix, is a diagonal matrix, and the elements on the diagonal are: is the feature output by the l-th layer of the graph convolution network;

[0053] Step 2-3, the prediction probability output by the v-th data source branch teacher model

[0054]

[0055] wherein, FC is a fully connected layer for classifying the features output by the graph convolution network, is the v-th data source in task t,

[0056] Preferably, in step 1, the measured ground indicator diagram is a binary image composed of a polished rod displacement and a polished rod load, and the measured electrical indicator diagram is a binary image composed of a polished rod displacement and a motor active power.

[0057] A multi-source semi-supervised class incremental working condition recognition system of a pumping unit well, characterized in that it comprises a working condition sample library storage module, a working condition recognition model construction module and a working condition recognition module, wherein the working condition sample library storage module further comprises a measured ground indicator diagram sample library and a measured electrical indicator diagram sample library.

[0058] The measured ground dynamometer card sample library and the measured electric dynamometer card sample library respectively store measured ground dynamometer cards and measured electric dynamometer cards, and are used to collect and manage the working condition samples of the pumping unit well in the system for training and identification; the measured ground dynamometer card sample library and the measured electric dynamometer card sample library are respectively connected with the working condition identification model construction module;

[0059] The working condition identification model construction module comprises a pumping unit well multi-source semi-supervised class incremental working condition identification model, and an output end of the working condition identification model is connected with an input end of the working condition identification module.

[0060] The working condition identification module is used for classifying and identifying the working condition of the pumping unit well to be identified through the working condition identification model.

[0061] Compared with the prior art, the pumping unit well multi-source semi-supervised class incremental working condition identification method and system have the beneficial effects that:

[0062] The pumping unit well multi-source semi-supervised class incremental working condition identification method and system simultaneously apply the multi-source fusion technology based on the attention mechanism, the class incremental learning and the semi-supervised learning to the class incremental working condition identification of the pumping unit well, fully utilize a small amount of multi-source labeled working condition samples, and combine a large amount of multi-source unknown working condition samples to realize more efficient, robust and practical working condition identification of the pumping unit well.

[0063] In the prior art, the Squeeze-and-Excitation attention mechanism is initially proposed for feature map fusion in a model. It compresses the feature map through a global pooling operation, and then learns the importance weight of each channel through a fully connected layer, so as to realize dynamic weighting between the channels in the feature map. The Squeeze-and-Excitation attention mechanism is mainly used to improve the feature representation ability in a single deep learning model, thereby improving the classification accuracy. In the field of multi-source data fusion, the common method is to fuse the features from different data sources through simple weighted summation, splicing or fixed weight distribution. Although these methods can combine multi-source data, they cannot dynamically adjust the fusion weight according to the importance and correlation of the data sources, and ignore the importance difference of different modal data in different scenes. In the pumping unit well multi-source semi-supervised class incremental working condition identification method and system, the Squeeze-and-Excitation attention mechanism is applied to multi-source feature fusion by using two data sources of the measured ground dynamometer card and the measured electric dynamometer card. The SE mechanism is used to dynamically weight the features of the measured ground dynamometer card and the measured electric dynamometer card, and the fusion proportion of the features of different data sources is dynamically adjusted according to the weight. This dynamic weight distribution method can better capture the complementary information between different data sources and enhance the adaptability of the model to multi-source data.

[0064] The oil pumping well multi-source semi-supervised class incremental working condition recognition method and system uses a multi-source data distillation method based on Kullback-Leibler divergence to measure the difference between the output logic of the teacher and student models, helps the student model to remember the old task while learning a new task, reduces the forgetting of old working condition categories in the class incremental learning process, and improves the class incremental working condition recognition accuracy.

[0065] In the prior art, most of the patents of graph neural networks, distillation learning and attention mechanisms apply distillation learning to ordinary classification tasks. The main purpose is to improve the classification accuracy of the student model through distillation learning, while reducing the model parameters to realize more efficient and lighter model deployment. The working condition categories are fixed and unchanged during model training. In this application, it is aimed at the special scenario of class incremental classification task. In the class incremental learning process, as new working condition categories appear, the model needs to learn new categories while maintaining the recognition ability of old categories. The present application introduces a multi-source data distillation loss based on Kullback-Leibler divergence to measure the difference between the output logic of the teacher and student models, helping the student model to remember the old task while learning a new task, reducing the forgetting of old working condition categories in the class incremental learning process. The application mode of this distillation learning is essentially different from the distillation learning of ordinary classification tasks in the prior art. It focuses more on solving the knowledge forgetting problem in class incremental learning, rather than simply improving the classification accuracy or reducing the model parameters.

[0066] The oil pumping well multi-source semi-supervised class incremental working condition recognition method and system can further improve the working condition recognition effect by breaking through the single information source misreporting limitation and the traditional multi-source fusion technology bottleneck, and enhance the robustness and engineering practicability of the working condition recognition model in the scene of complex and variable working condition categories and less training data. It has important theoretical, practical application value and economic significance for improving oil well fault prevention ability, optimizing oil production scheduling, improving oil well recovery rate, promoting the construction and development of intelligent oilfields. BRIEF DESCRIPTION OF DRAWINGS

[0067] Fig. 1 The oil pumping well multi-source semi-supervised class incremental working condition recognition method flowchart.

[0068] Fig. 2 The oil pumping well multi-source semi-supervised class incremental working condition recognition method flowchart.

[0069] Fig. 3 The oil pumping well multi-source semi-supervised class incremental working condition recognition system principle block diagram. DETAILED DESCRIPTION

[0070] Figs. 1-3 is the best embodiment of the present application, and the following will be combined with the accompanying drawings Figs. 1-3Further illustrate the present application.

[0071] As shown in the figure, a pumping unit well multi-source semi-supervised class incremental working condition recognition method comprises the following steps: Figs. 1-2

[0072] Step 1, respectively store the sample library of the measured ground dynamometer card and the measured electric dynamometer card corresponding to the labeled and unlabeled pumping unit well working condition;

[0073] The measured ground dynamometer card and the measured electric dynamometer card corresponding to the pumping unit well working condition are respectively stored in the corresponding sample library. Each sample in the sample library is composed of data points actually collected by the production site under different working conditions, ensuring the authenticity and representativeness of the sample. These data points record the running state of the pumping unit well at a specific time in detail. The samples in each sample library are divided into labeled and unlabeled pumping unit well working condition samples. The labeled samples are strictly selected according to the operation records of the oil well, ensuring the accuracy and reliability of the samples.

[0074] The measured ground dynamometer card is a binary image composed of polished rod displacement and polished rod load, and the measured electric dynamometer card is a binary image composed of polished rod displacement and motor active power. These sample libraries serve as the basis for data storage and processing, providing necessary data support for subsequent working condition recognition.

[0075] Step 2, respectively construct a graph neural network teacher model for the measured ground dynamometer card and the measured electric dynamometer card;

[0076] For each information source raw data, the local texture feature information of the data is extracted by using the histogram of oriented gradients, and then the nearest neighbor method is used to construct a graph structure for each data source. Specifically, it includes the following steps:

[0077] Step 2-1, for each sample in a single data source, calculate the distance between it and all other samples in its data set, and select the two samples with the smallest distance as its nearest neighbor samples, construct the initial adjacency matrix of the graph convolution network, and the calculation formula of the graph convolution network of the tth task is:

[0078]

[0079] Wherein, is the adjacency matrix constructed for the vth data source in task t, is the data of the vth data source in task t, including labeled data and unlabeled data of the vth data source.

[0080] Step 2-2, use the convolution layer of the graph convolution network to extract node features, and the calculation process is as follows:

[0081]

[0082] wherein, I K is an identity matrix, relu(·) is a ReLU activation function, W is a trainable parameter matrix, is a diagonal matrix, and the elements on the diagonal are: is the feature output by the l-th layer of the graph convolutional network.

[0083] Step 2-3, the prediction probability output by the v-th data source branch teacher model The calculation process is:

[0084]

[0085] wherein, FC is a fully connected layer, used for classifying the features output by the graph convolutional network, is the v-th data source data in task t, is the feature output by the l+1-th layer of the graph convolutional network.

[0086] Step 3, dynamically fuse the prediction probabilities of each teacher model using the Squeeze-and-Excitation attention mechanism;

[0087] The attention mechanism is used to dynamically learn the weights of the prediction probabilities of different data source teacher models, and by assigning different weights to the prediction results of different data source teacher models, the effective fusion of multi-source teacher models is realized. Specifically, the following steps are included:

[0088] Step 3-1, the multi-source feature representation fusion module is realized by referring to the Squeeze-and-Excitation attention mechanism, which is composed of a fully connected layer- ReLU activation layer- fully connected layer- Sigmoid activation layer. The prediction probability output by the v-th data source branch teacher model is represented as p t v The corresponding attention score ω v is represented as:

[0089]

[0090] wherein, sigmoid(·) is a Sigmoid activation function, relu(·) is a ReLU activation function, W1 and W2 are the weights of the two fully connected layers respectively.

[0091] Step 3-2, the obtained attention weight is weighted to the prediction result output by the multi-source teacher model, and the fused multi-source teacher model prediction result p t is represented as:

[0092]

[0093] where p t represents the prediction result of the fused multi-source teacher model, including the prediction result p t (L t ) of the labeled data and the prediction result p t (U t ) of the unlabeled data.

[0094] Step 4, multi-source data distillation learning using Kullback-Leibler divergence;

[0095] The specific process of the multi-source data distillation strategy is as follows:

[0096] Step 4-1, when learning task t, the model G t is the student model, the model G t-1 of task t-1 is the teacher model of the model G t . The distillation loss L KL between the output prediction probabilities of the teacher model and the student model is calculated by KL divergence:

[0097]

[0098] where T t-1 is the number of learned classes after task t-1, represents the prediction probability output by the multi-source teacher model of the t-1th task, represents the prediction probability output by the student model of the tth task.

[0099] Step 4-2, the total loss function L t of task t is composed of the cross-entropy loss function L c and the multi-source data distillation loss L KL , and the cross-entropy loss function L c is represented as:

[0100]

[0101] where c t represents the number of sample classes of task t, p t (L t ) represents the prediction result of the fused model of the labeled sample features in task t, represents the true label of the i-th labeled sample in task t.

[0102] Step 4-3, the total loss function L t is calculated as:

[0103] L t =L c +λL KL

[0104] wherein L c represents the cross-entropy loss function, L KL is the distillation loss between the output prediction probabilities of the teacher model and the student model, and λ is the weight of the multi-source data distillation loss function.

[0105] Step 5, using a logistic regression classifier to improve the label propagation algorithm for semi-supervised working condition recognition;

[0106] By fusing multi-source data features and using a logistic regression classifier to improve the label propagation algorithm, the limited labeled data and a large amount of unlabeled data are fully utilized, and the accuracy of the pumping well working condition recognition is improved. Specifically, the steps include:

[0107] Step 5-1, fusing the features output by the graph convolutional network of different data sources to obtain the feature matrix X t of the task t after fusion.

[0108]

[0109] wherein, X t contains the global feature matrix X L of the labeled data and the global feature matrix X U of the unlabeled data. t wherein: X L = [X U , X ].

[0110] Step 5-2, the multi-source semi-supervised graph classification method based on enhanced label propagation is represented as:

[0111]

[0112] wherein, is the adjacency matrix constructed using the fused features, contains the adjacency matrix constructed by the labeled data and the adjacency matrix constructed by the unlabeled data The calculation formula is:

[0113]

[0114] wherein, represents the probability that the kth sample belongs to class c. The calculation formula is:

[0115]

[0116] wherein, y k represents the label of the kth sample.

[0117] Step 5-3, generate initial prediction probability distribution for unlabeled data using logistic regression classifier, provide more accurate initial label vector for label propagation.

[0118] After completing the label propagation process, use multiple logistic regression classifiers to make P LR Classify the label embedding vectors and calculate the average of these classifier results to obtain the final classification label. The final label of task t is y t The calculation formula is:

[0119]

[0120] wherein, is the label vector of task t after label propagation.

[0121] To implement the above-mentioned pumping well multi-source semi-supervised class incremental working condition recognition method, in the present application, an identification system is also included, such as Fig. 3 As shown in the figure, the identification system includes a working condition sample library storage module, a working condition recognition model construction module, and a working condition recognition module, wherein the working condition sample library storage module further includes a measured ground dynamometer card sample library and a measured electric dynamometer card sample library.

[0122] The measured ground dynamometer card sample library and the measured electric dynamometer card sample library respectively store measured ground dynamometer cards and measured electric dynamometer cards, aiming to collect and manage the pumping well working condition samples in the system for training and to be identified. The measured ground dynamometer card sample library and the measured electric dynamometer card sample library are respectively connected with the working condition recognition model construction module. The working condition recognition model construction module contains a pumping well multi-source semi-supervised class incremental working condition recognition model, and the output end thereof is connected with the input end of the working condition recognition module. The working condition recognition module is used for classifying and identifying the pumping well working condition to be identified through the working condition recognition model.

[0123] Next, an example is taken to verify the actual recognition effect of the above-mentioned pumping well multi-source semi-supervised class incremental working condition recognition method:

[0124] Firstly, a data set containing 1650 samples is constructed, covering 11 typical working conditions, and each working condition contains 150 samples. In the class incremental learning setting, the initial task contains 2 categories, and then each new task increases 1 or 2 categories. This design simulates the scenario of the gradual appearance of new working condition categories in oilfield operations. Specifically, the 11 typical working conditions in the data set include: normal, insufficient liquid supply, continuous pumping and spraying, waxing, pump leakage, tubing leakage, fixed valve leakage, mobile valve leakage, pumping rod breakage, pump card, and pump mobile valve failure.

[0125] Then, following the actual operation of the oilfield pumping well, the initial working condition is set as two working condition categories commonly seen in oilfield production, i.e., "normal" and "insufficient liquid supply", and the semi-supervised learning proportion is set as 10%, 30%, 50%, 70% and 100% respectively.

[0126] Subsequently, the actual measured ground indicator diagram and the actual measured electrical indicator diagram are used for working condition recognition. When the semi-supervised learning proportion is set as 10%, the number of sample categories of the training set is set as 2, 3, 5, 7, 9 and 11 respectively, and the average recognition rate is 99.33%, 99.12%, 98.98%, 97.89%, 98.09% and 96.69% respectively; when the semi-supervised learning proportion is set as 30%, the number of sample categories of the training set is set as 2, 3, 5, 7, 9 and 11 respectively, and the average recognition rate is 99.47%, 99.30%, 99.15%, 98.35%, 98.98% and 98.29% respectively; when the semi-supervised learning proportion is set as 50%, the number of sample categories of the training set is set as 2, 3, 5, 7, 9 and 11 respectively, and the average recognition rate is 99.51%, 99.34%, 99.28%, 98.85%, 99.17% and 98.72% respectively; when the semi-supervised learning proportion is set as 70%, the number of sample categories of the training set is set as 2, 3, 5, 7, 9 and 11 respectively, and the average recognition rate is 99.61%, 99.45%, 99.39%, 98.98%, 99.19% and 98.84% respectively; when the semi-supervised learning proportion is set as 100%, the number of sample categories of the training set is set as 2, 3, 5, 7, 9 and 11 respectively, and the average recognition rate is 99.78%, 99.57%, 99.41%, 99.03%, 99.33% and 98.99% respectively.

[0127] As can be seen from the above, the oil pumping well multi-source semi-supervised incremental working condition recognition method exhibits good robustness under most marked data proportions and different category numbers. In the case of relatively scarce marked data (semi-supervised proportion 10% and category number 11), the accuracy of the method proposed by the present application reaches 96.69%. In the case of scarce data, the oil pumping well multi-source semi-supervised incremental working condition recognition method can effectively utilize limited marked data, and at the same time, fully utilize unmarked data through semi-supervised learning to improve the recognition performance. With the gradual increase of the proportion of marked data, the performance of the method of the present application is further improved. When the semi-supervised proportion is 100% and the category number is 11, the accuracy of the oil pumping well multi-source semi-supervised incremental working condition recognition method reaches 98.99%.

[0128] The above merely describes preferred embodiments of the present application, but is not intended to limit the present application to other forms, and any person skilled in the art can make changes or modifications to the above disclosed technical contents into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiments without departing from the technical solution content of the present application and according to the technical essence of the present application still belongs to the protection scope of the technical solution of the present application.

Claims

1. A method for identifying multi-source semi-supervised incremental operating conditions in oil pumping wells, characterized in that: Includes the following steps: Step 1: Store sample libraries containing measured surface dynamometer diagrams and measured electrical dynamometer diagrams corresponding to the operating conditions of marked and unmarked pumping wells, respectively. Step 2: Construct a graph neural network teacher model for the two data sources, the measured ground dynamometer diagram and the measured electrical power diagram; Step 3: Dynamically fuse the prediction probabilities of each teacher model using the Squeeze-and-Excitation attention mechanism; Step 4: Use Kullback-Leibler divergence to perform multi-source data distillation learning; Step 5: Use a label propagation algorithm improved by logistic regression classifier for semi-supervised work condition identification; Step 4 specifically includes the following steps: Step 4-1, when learning task t, model G is usually... t As the student model, the model G for task t-1 t-1 As model G t The distillation loss L between the output predicted probabilities of the teacher model and the student model. KL Calculated using KL divergence: Among them, T t-1 The number of categories learned after task t-1. This represents the predicted probability output by the multi-source teacher model for the (t-1)th task. This represents the predicted probability output by the student model for the t-th task; Step 4-2, Total loss function L for task t t By cross-entropy loss function L c and multi-source data distillation loss L KL Composition, cross-entropy loss function L c Represented as: Among them, c t p represents the number of sample categories in task t. t (L t The ) indicates that there are labeled samples in task t, and the features represent the predicted results of the fused model. This represents the true label of the i-th labeled sample in task t; Step 4-3, Total Loss Function L t The calculation formula is: THE t =L c +λL KL Among them, L c L represents the cross-entropy loss function. KL Let λ be the distillation loss between the output prediction probabilities of the teacher model and the student model, and let λ be the weight of the multi-source data distillation loss function.

2. The method for identifying multi-source semi-supervised incremental operating conditions of pumping wells according to claim 1, characterized in that: Step 3 specifically includes the following steps: Step 3-1: Implement the multi-source feature representation fusion module using the Squeeze-and-Excitation attention mechanism. This module consists of a fully connected layer, a ReLU activation layer, another fully connected layer, and a Sigmoid activation layer. The predicted probability output by the teacher model for the v-th data source branch is represented as... Its corresponding attention score ω v Represented as: Where sigmoid(·) is the Sigmoid activation function, relu(·) is the ReLU activation function, and W1 and W2 are the weights of the two fully connected layers, respectively. Step 3-2: The obtained attention weights are weighted onto the prediction results output by the multi-source teacher model to obtain the fused multi-source teacher model prediction result p. t Represented as: Where, p t This represents the prediction results of the fused multi-source teacher model, including the prediction results of the labeled data (p). t (L t ) and unlabeled data prediction results p t (U t ).

3. The method for identifying multi-source semi-supervised incremental operating conditions of pumping wells according to claim 1, characterized in that: Step 5 specifically includes the following steps: Step 5-1: Combine the features output by graph convolutional networks from different data sources. Mean fusion is performed, and the feature matrix X of task t after fusion is obtained. t The calculation formula is: in, X t The global feature matrix X containing labeled data L and the global feature matrix X of unlabeled data U , where: X t =[X L ,X U ]; Step 5-2, the multi-source semi-supervised graph classification method based on enhanced label propagation is represented as: in, The adjacency matrix is ​​constructed using the fused features. Adjacency matrix containing labeled data Adjacency matrix constructed from unlabeled data The calculation formula is: in, Let represent the probability that the k-th sample belongs to class c. The calculation formula is: Among them, y k The label represents the k-th sample; Step 5-3: Use a logistic regression classifier to generate an initial predicted probability distribution for the unlabeled data, providing a more accurate initial label vector for label propagation. After completing the label propagation process, use multiple logistic regression classifiers to perform P... LR The label embedding vectors are classified, and the average of these classifier results is calculated to obtain the final classification label; task t final label y t The calculation formula is: in, Let t be the label vector after label propagation.

4. The method for identifying multi-source semi-supervised incremental operating conditions of pumping wells according to claim 1, characterized in that: Step 2 specifically includes the following steps: Step 2-1: For each sample in a single data source, calculate its distance to all other samples in its dataset, and select the two samples with the smallest distance as its nearest neighbors. Construct the initial adjacency matrix of the graph convolutional network. The formula for calculating the graph convolutional network for the t-th task is: in, The adjacency matrix constructed for the v-th data source in task t. For the v-th data source in task t, there are labeled and unlabeled data from the v-th data source. Step 2-2: Extract node features using the convolutional layers of a graph convolutional network. in, I K Let W be the identity matrix, relu(·) be the ReLU activation function, and W be the trainable parameter matrix. It is a diagonal matrix, and the elements on the diagonal are: The features output by the l-th layer graph convolutional network; Steps 2-3: The predicted probability output by the teacher model in the v-th data source branch. In this context, FC stands for fully connected layer, used to classify the features output by the graph convolutional network. For the v-th data source in task t, These are the features output by the (l+1)th layer graph convolutional network.

5. The method for identifying multi-source semi-supervised incremental operating conditions of pumping wells according to claim 1, characterized in that: In step 1, the measured ground dynamometer diagram is a binary image composed of the rod displacement and rod load, and the measured electrical dynamometer diagram is a binary image composed of the rod displacement and motor active power.

6. An identification system for implementing the multi-source semi-supervised incremental operating condition identification method for pumping wells according to any one of claims 1 to 5, characterized in that: It includes a working condition sample library storage module, a working condition identification model construction module, and a working condition identification module. The working condition sample library storage module further includes a measured ground dynamometer diagram sample library and a measured electrical dynamometer diagram sample library. The measured surface dynamometer diagram sample library and the measured electrical dynamometer diagram sample library store measured surface dynamometer diagrams and measured electrical dynamometer diagrams respectively, aiming to collect and manage oil well operating condition samples used for training and identification in the system; the measured surface dynamometer diagram sample library and the measured electrical dynamometer diagram sample library are respectively connected to the operating condition identification model construction module; The working condition identification model construction module includes a multi-source semi-supervised incremental working condition identification model for pumping wells, and its output is connected to the input of the working condition identification module. The operating condition identification module is used to classify and identify the operating conditions of the pumping unit well to be identified through the operating condition identification model.

Citation Information

Patent Citations

  • Fault diagnosis method, computing device and readable storage medium

    CN117349726A

  • Working condition identification model training method, working condition diagnosis method and device

    CN118781443A

  • Pumping unit working condition intelligent analysis method and system

    CN119025832A

  • Method and system for identifying actual measurement electric power diagram working condition of rod-pumped well

    CN117152548A

  • Land utilization classification method using incomplete multi-view image

    CN118097219A