Pump failure prediction method and device

By combining supervised learning and unsupervised learning methods, linear discriminant analysis and principal component analysis are used to predict abnormal states of electric submersible oil pumps, solving the problem of failure prediction of electric submersible oil pumps and improving production efficiency and economic benefits.

CN120273911APending Publication Date: 2025-07-08SK INNOVATION CO LTD +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411401599.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-01-08
Filing Date
2024-10-09
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art is difficult to effectively predict the failure of electric submersible pumps, resulting in reduced productivity and economic losses.

Method used

Using a combination of supervised learning and unsupervised learning, the abnormal state of the pump is predicted through data acquisition, training and anomaly detection modules, using linear discriminant analysis, principal component analysis and Mahalanobis distance.

Benefits of technology

Early prediction of electric submersible pump failures is achieved, unnecessary inspection and maintenance work is reduced, and production efficiency and economic benefits are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120273911A_ABST
    Figure CN120273911A_ABST
Patent Text Reader

Abstract

The present disclosure relates to a pump failure prediction method and apparatus, the pump failure prediction apparatus including at least one processor, a storage device communicatively connected with the processor and storing program code running in the processor, and a communicator communicatively connected with the processor, wherein the program code comprises a data acquisition module and an exception detection module, and the exception detection module detects three exceptions. And inputting the real-time data into a first model in which the historical data has been supervised and learned before a fault occurs, so as to obtain a first anomaly occurring before the fault occurs. A second anomaly occurring outside the normal operating range of the pump by inputting the real-time data into a second model having unsupervised learning of the normal operating data. A third anomaly occurring outside the initial normal operating range of the pump by inputting the real-time data into a third model that has unsupervised learning of the initial operating data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to methods and devices for predicting pump failures. Background Art

[0002] An electric submersible pump is a device for pumping oil from an oil well. It is widely used due to its operational applicability and ability in various environments. The pump is an important device in oil production and consists of various mechanical devices such as motors, pumps, and valves. Any mechanical device may fail due to mechanical wear caused by sand flowing into the drilling hole or corrosion caused by chemical reactions during the production process of hydrocarbons. If some of the devices constituting the electric submersible pump fail, the pumping function of the pump may not be fully exerted. As a result, the productivity decreases or the operation stops, and time, cost, and effort required for inspection / repair may be needed, leading to economic losses. Summary of the Invention

[0003] The present disclosure aims to propose methods and devices for predicting pump failures using both supervised learning and unsupervised learning.

[0004] In an embodiment of the present disclosure, there is provided a pump failure prediction device, including: at least one processor; a storage device communicatively connected to the processor and storing program code running in the processor; and a communicator communicatively connected to the processor, where the program code includes: a data acquisition module that acquires real-time data related to the state of a pump installed in an oil well and operating; and an anomaly detection module that detects three types of anomalies - a first anomaly that appears before a failure by inputting the real-time data into a first model that has performed supervised learning on historical data, a second anomaly outside the normal operating range of the pump by inputting the real-time data into a second model that has performed unsupervised learning on normal operating data, and a third anomaly outside the initial normal operating range of the pump by inputting the real-time data into a third model that has performed unsupervised learning on initial operating data.

[0005] In an embodiment of the present disclosure, the historical data may include two types of operating data: operating data marked as normal and operating data marked as abnormal. The first model may classify the operating data included in the historical data into a normal category or an abnormal category using linear discriminant analysis, and the anomaly detection module may determine that the first anomaly occurs when the real-time data is classified into the abnormal category according to the linear discriminant criterion generated by the first model.

[0006] In an embodiment of the present disclosure, the normal operation data may be operation data obtained when the pump operates in a normal state. The second model may be generated and trained by obtaining the average value of multiple normal section errors between the initial operation data (normal operation data) and the reconstructed operation data (reconstructed data) calculated after applying principal component analysis. In the second model, principal component analysis is applied to the initial operation data to obtain eigenvectors. Based on the eigenvectors, the second model reconstructs the operation data; and the process of obtaining the normal section error between the initial operation data and the reconstructed operation data is repeated at each time step of the operation data to obtain multiple normal section errors. In the case where real-time eigenvectors are obtained by applying principal component analysis to real-time data, when the error level of the real-time data increases over time, the anomaly detection module may determine that a second anomaly has occurred. Based on the real-time eigenvectors, real-time reconstructed data is obtained; a real-time error between the real-time data and the real-time reconstructed data is obtained; the error level is obtained by dividing the real-time error by the average value of the multiple normal section errors generated by the second model; and the error level is recorded over time.

[0007] In an embodiment of the present disclosure, the initial operation data may be operation data obtained within a predetermined time period starting from the time when it is assumed that the pump operates normally without any problems in a stable state. The third model may be trained by applying principal component analysis to the initial operation data. After applying principal component analysis to the data, an eigenvector composed of at least two components is obtained. Then, the center point of the eigenvector is obtained based on the distribution of the vector in the multi-dimensional space. Based on the distribution of the eigenvector and the center point, the third model calculates the Mahalanobis distance from the center point to each eigenvector of the data. According to the range of the Mahalanobis distance, a reference distance is statistically determined. Then, principal component analysis is applied to the real-time operation data, and the third model calculates the Mahalanobis distance from the center point to each eigenvector of the real-time data, and determines whether the state of the pump is abnormal when the Mahalanobis distance is greater than the reference distance.

[0008] In an embodiment of the present disclosure, the threshold distance may be determined as a multiple of the average value of the multiple Mahalanobis distances obtained from the initial operation data.

[0009] In an embodiment of the present disclosure, the program code may further include a reporting module that provides a notification of detecting an anomaly in the pump based on the output of the anomaly detection module.

[0010] In an embodiment of the present disclosure, even before the pump operates normally, a second model can be generated and trained by using the normal operation data of another pump operating in a similar environment, and the second model can be regenerated by updating the normal operation data with the operation data of the pump itself in the normal state.

[0011] In an embodiment of the present disclosure, the program code may further include a training data generation module that generates training data for training the first model, the second model, and the third model by using the preprocessed operation data in a manner of generating operation data by recording real-time data over time and preprocessing the operation data. The training data generation module can generate historical data, thereby generating the training data of the first model in a manner of normalizing time series data in the operation data. The operation data corresponding to the previously recorded fault state is marked as abnormal, while the data of the normal operation part is marked as normal. The training data generation module can generate normal operation data, thereby generating the training data of the second model in a manner of applying principal component analysis to the operation data at each time step to extract feature vectors. Reconstructed data is obtained by reconstructing the feature vectors. The error between the reconstructed data and the operation data is calculated. When the error is less than or equal to the reference error, the operation data is included in the normal operation data; when the error is greater than the reference error, the operation data is excluded from the normal operation data. The training data generation module can generate the operation data collected within a predetermined period starting from the time when the pump operates in the normal state after the pump is installed, as the initial operation data for generating the training data of the third model.

[0012] In an embodiment of the present disclosure, a pump fault prediction method includes: collecting real-time data related to the state of a pump installed in an oil well; generating training data by preprocessing the operation data collected by recording real-time data over time; training the first model, the second model, and the third model by using the training data; detecting a first anomaly that occurs before a fault occurs by inputting the real-time data of the pump into the first model that has performed supervised learning on historical data, detecting a second anomaly outside the normal operation range of the pump by inputting the real-time data of the pump into the second model that has performed unsupervised learning on normal operation data, and detecting a third anomaly outside the initial normal operation range of the pump by inputting the real-time data of the pump into the third model that has performed unsupervised learning on initial operation data; and providing a notification of an anomaly detected in the pump based on the output of the anomaly detection.

[0013] In an embodiment of the present disclosure, when detecting an anomaly, when real-time data is classified as an abnormal category based on a linear discrimination criterion generated by a first model according to the first model, a first anomaly can be determined. The first model is a linear discriminant analysis model trained to classify historical data including operation data marked as normal and operation data marked as abnormal into a normal category and an abnormal category.

[0014] In an embodiment of the present disclosure, when detecting an anomaly, in the case of applying principal component analysis to normal operation data including operation data obtained when the pump is operating in a normal state to obtain eigenvectors, a second anomaly can be determined when the error level of the real-time data increases over time. Reconstructed data is obtained by reconstructing the eigenvectors; the process of obtaining the normal partial error between the normal operation data (normal operation data) and the reconstructed data (reconstructed data) calculated after applying principal component analysis is repeated at each time step of the operation data to obtain a plurality of normal partial errors. Based on a second model generated by training by obtaining the average value of the plurality of normal partial errors, principal component analysis is applied to the real-time data to obtain real-time eigenvectors. Based on the real-time eigenvectors, real-time reconstructed data is obtained; the real-time error between the real-time data and the real-time reconstructed data is obtained; the real-time error is divided by the average value of the plurality of normal partial errors generated by the second model to obtain an error level; and the error level is recorded over time.

[0015] In an embodiment of the present disclosure, when detecting an anomaly, in the case of applying principal component analysis to initial operation data, a third anomaly can be determined when the Mahalanobis distance is greater than a reference distance. The initial operation data is operation data obtained within a predetermined period starting from a time when it is assumed that the pump is operating normally in a stable state without any problems. The third model can be trained by applying principal component analysis to the initial operation data. After applying principal component analysis to the data, an eigenvector composed of at least two components is obtained. Then, based on the distribution of the vectors in the multi-dimensional space, the center point of the eigenvector is obtained. Based on the distribution and the center point of the eigenvector, the third model calculates the Mahalanobis distance from the center point to each eigenvector of the data. According to the range of the Mahalanobis distance, the reference distance is statistically determined. Then, principal component analysis is applied to the real-time operation data, and the third model calculates the Mahalanobis distance from the center point to each eigenvector of the real-time data, and determines whether the state of the pump is abnormal when the Mahalanobis distance is greater than the reference distance.

[0016] In an embodiment of the present disclosure, when generating training data, in order to generate training data for the first model, historical data can be generated by normalizing the time series data in the operation data. Mark the operation data corresponding to the previously recorded fault status as abnormal, and mark the data of the normal operation part as normal, so as to generate training data for the second model. Normal operation data can be generated by applying principal component analysis to the operation data at each specific time to extract feature vectors. Reconstructed data is obtained from the feature vectors. Calculate the error between the reconstructed data and the operation data; when the error is less than or equal to the reference error, the operation data is included in the normal operation data; when the error is greater than the reference error, the operation data is excluded from the normal operation data. In order to generate training data for the third model, the operation data collected within a predetermined period starting from the time when the pump operates in a normal state after the pump is installed can be generated as initial operation data.

[0017] In an embodiment of the present disclosure, when training the model, the second model can be generated by using the normal operation data of another pump operating in a similar environment, and the second model can be regenerated by updating the normal operation data with the operation data of the pump itself in a normal state.

[0018] Through the following detailed description based on the drawings, the features and advantages of the present disclosure will become more obvious.

[0019] Prior to this, based on the fact that the inventor can appropriately define the concept of terms in order to best explain the principles of the present disclosure, the terms or words used in this specification and claims should not be construed as conventional and dictionary meanings, but should be construed as meanings and concepts consistent with the technical idea of the present disclosure.

[0020] In some embodiments of the present disclosure, it is possible to predict the occurrence of a new type of fault while accurately predicting the faults that have occurred in the past. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] When the following detailed description is combined with the accompanying drawings, the above and other objects, features, and other advantages of the present disclosure will be more clearly understood, where:

[0022] Figure 1 is a view showing the operating environment of a pump based on an embodiment;

[0023] Figure 2 is a diagram showing a pump fault prediction device based on an embodiment;

[0024] Figure 3 is a diagram showing a module based on an embodiment;

[0025] Figure 4 is a diagram showing the data collected by the data acquisition module based on an embodiment;

[0026] Figure 5 is a flowchart showing a method for predicting a pump failure based on an embodiment;

[0027] Figure 6 is a diagram for explaining anomaly detection using a first model based on an embodiment;

[0028] Figure 7 is a diagram for explaining a first model trained to have three classes based on an embodiment;

[0029] Figure 8 is a diagram showing results of anomaly detection obtained over time by using a first model based on an embodiment;

[0030] Figure 9 is a diagram for explaining anomaly detection using a second model based on an embodiment;

[0031] Figure 10 is a diagram showing results of anomaly detection obtained over time by using a second model based on an embodiment;

[0032] Figure 11 is a diagram for explaining anomaly detection using a third model based on an embodiment;

[0033] Figure 12 is a diagram showing results of anomaly detection obtained over time by using a third model based on an embodiment; and

[0034] Figure 13 is a diagram showing results of anomaly detection obtained over time by using a first model, a second model, and a third model and accumulated based on an embodiment. DETAILED DESCRIPTION

[0035] The objects, advantages, and features of the present disclosure will become more apparent through the following detailed description and exemplary embodiments in conjunction with the accompanying drawings, but the present disclosure is not necessarily limited thereto. Additionally, when explaining the present disclosure, when it is determined that a detailed description of related known technologies may unnecessarily obscure the gist of the present disclosure, the detailed description thereof will be omitted.

[0036] When assigning reference numerals to components in the drawings, it should be noted that even if components are shown in different drawings, the same components are given the same reference numerals as much as possible, and similar components are given similar reference numerals.

[0037] The terms used to describe the embodiments of the present disclosure are not intended to limit the present disclosure. It should be noted that unless the context otherwise requires, singular expressions include plural expressions.

[0038] In this document, expressions such as "having", "may have", "including", or "may include" indicate the presence of related features (e.g., numerical values, functions, operations, or components such as parts) without precluding the presence of additional features.

[0039] Terms such as "one", "another", "the other", "first", "second", etc. are used to distinguish one component from other components, and the components are not limited by these terms.

[0040] The embodiments described in this document and the drawings are not intended to limit the present disclosure to specific embodiments. The present disclosure should be understood to cover various modifications, equivalents, and / or alternatives of the embodiments.

[0041] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. Herein, "pump 2" may refer to "electric submersible pump 2".

[0042] Figure 1 is a view showing an operating environment of pump 2 according to an embodiment.

[0043] According to the method and apparatus 10 for pump failure prediction of the present disclosure, the failure of pump 2 installed in oil well 1 and used for oil production can be predicted in advance by predicting the failure of pump 2. Oil well 1 is a borehole formed in a formation for pumping oil buried underground. The pump 2 installed in oil well 1 can be of various types. Pump 2 may include an electric submersible pump 2. The electric submersible pump 2 may include a motor, a gas separator, a pump, a cable, a production pipe, and various other devices. The devices constituting the electric submersible pump 2 may fail for various reasons. According to the method and apparatus for predicting the failure of pump 2 of the present disclosure, the failure of pump 2 can be predicted by detecting an abnormality occurring in the electric submersible pump 2 before pump 2 fails. The pump failure prediction apparatus 10 is capable of reporting a notification to administrator 4 when an abnormality is detected.

[0044] The oil production unit may include oil well 1, pump 2, pipelines, oil tanks, junction boxes, pump controller 3, and various other devices. The pump failure prediction apparatus 10 may collect data from various sensors installed in the oil production unit. The pump failure prediction apparatus 10 may collect data from pump controller 3 that controls the electric submersible pump 2. The pump failure prediction apparatus 10 may use data input by administrator 4. The pump failure prediction apparatus 10 may be configured as a computer device. As a non-limiting example, the pump failure prediction apparatus 10 may be configured as a PC, a server, a tablet PC, a remote operating system, and a device that performs other information processing functions.

[0045] Figure 2 is a diagram showing the pump failure prediction apparatus 10 based on an embodiment.

[0046] The pump fault prediction device 10 may include at least one processor 11, a storage device 12 communicatively connected to the processor 11 and storing program code that runs in the processor 11, and a communicator 13 communicatively connected to the processor 11. The pump fault prediction device 10 may further include an input / output interface 14 communicatively connected to the processor 11, which receives commands or data from the administrator 4, displays the status of the pump 2 to the administrator 4, and provides data or anomaly detection notifications. The storage device 12 may store the trained first model M1, second model M2, and third model M3. The storage device 12 may store program code written to execute the method for predicting the fault of the pump 2, and the program code may be written in units of modules.

[0047] Figure 3 is a diagram showing a module based on an embodiment. Refer to Figure 3 and Figure 2 .

[0048] Based on the embodiment, the module may be a component combining hardware and software. The module may operate in the processor 11. The module may include a data acquisition module 21, a training data generation module 22, a training module 23, an anomaly detection module 24, and a reporting module 25. The data acquisition module 21, the training data generation module 22, the training module 23, the anomaly detection module 24, and the reporting module 25 may cooperate to execute the method for predicting the fault of the pump 2. The pump fault prediction device 10 may operate the module in a manner that the program code, which is software stored in the storage device 12, is run by the processor 11. The data acquisition module 21, the training data generation module 22, the training module 23, the anomaly detection module 24, and the reporting module 25 may be stored in the storage device 12 by being written into the program code and may be run by the processor 11.

[0049] The program code may include: a data acquisition module 21 that acquires real-time data related to the status of the pump 2 installed in and operating in the oil well 1; a training data generation module 22 that generates operating data by recording the acquired real-time data over time and generates training data by preprocessing the operating data; and a training module 23 that trains the first model M1, the second model M2, and the third model M3 by using the training data.

[0050] The data acquisition module 21 may acquire real-time data related to the status of the pump 2 installed in and operating in the oil well 1 from the oil production unit. The data acquisition module 21 may generate operating data by recording the real-time data over time. The data acquisition module 21 may store the real-time data and the operating data in the storage device 12.

[0051] Figure 4It is a diagram showing the data collected by the data acquisition module 21 based on the embodiments.

[0052] The data acquisition module 21 can collect well data WD. The well data WD is data related to the oil production unit. The well data WD can include monitoring data MD and static data SD. The static data SD is data that does not change after the pump 2 is installed in the oil well 1. The static data SD can include drilling hole diagrams, well inclination measurements, installation reports, workover reports, and various other data. The static data SD may change when the administrator 4 updates the data or when the specifications of the pump 2 change. The data acquisition module 21 can receive the static data SD from the administrator 4.

[0053] The monitoring data MD is data measured when the pump 2 is running. The monitoring data MD can include first change data CD1 and second change data CD2. The first change data CD1 can be a parameter measured daily. The first change data CD1 can include daily production data (e.g., oil, gas, and water) and daily wellhead pressure data (e.g., tubing pressure and casing pressure). The second change data CD2 can be a parameter measured every 15 minutes or at a predetermined time interval. The second change data CD2 can include alarm data and sensor data. The sensor data can include bus voltage, fluid temperature, motor current, motor frequency, motor temperature, and various other elements.

[0054] In this embodiment, the real-time data collected by the data acquisition module 21 can be a part of the monitoring data MD. For example, the real-time data can include a part of the second change data CD2 measured every 15 minutes. The elements of the real-time data or operating data based on the embodiments can include bus voltage, fluid temperature, motor current, motor frequency, and motor temperature among the second change data CD2. The elements of the real-time data or operating data based on the embodiments can include motor frequency, inflow rate into the pump 2, motor temperature, fluid temperature, motor voltage, bus voltage, motor current, and temperature difference between the fluid and the pump 2. The real-time data or operating data can have values of multiple elements at each measurement time.

[0055] Returning to the reference Figure 4 。

[0056] The training data generation module 22 can generate operating data by recording real-time data over time, and can generate training data for training the first model M1, the second model M2, and the third model M3 by using the preprocessed operating data after preprocessing the operating data.

[0057] The training data generation module 22 can generate historical data for training the first model M1. The historical data can include operating data marked as normal and operating data marked as abnormal.

[0058] The training data generation module 22 can generate normal operation data for training the second model M2. The normal operation data can include operation data obtained when the pump 2 operates in a normal state.

[0059] The training data generation module 22 can generate initial operation data for training the third model M3. The initial operation data can include operation data obtained within a predetermined time period starting from a time when it is assumed that the pump 2 operates normally without problems in a steady state.

[0060] The training module 23 can train the first model M1, the second model M2, and the third model M3 by using the training data generated by the training data generation module 22, and can store the trained first model M1, second model M2, and third model M3 in the storage device 12. The training module 23 can independently execute the process of training the first model M1, the second model M2, or the third model M3.

[0061] The program code can include an anomaly detection module 24 that detects three types of anomalies - a first anomaly that occurs before a failure by inputting real-time data into the first model M1 that has undergone supervised learning of historical data, a second anomaly outside the normal operation range of the pump 2 by inputting real-time data into the second model M2 that has undergone unsupervised learning of normal operation data, and a third anomaly outside the initial normal operation range of the pump 2 by inputting real-time data into the third model M3 that has undergone unsupervised learning of initial operation data.

[0062] The anomaly detection module 24 can predict whether a failure will occur by using the trained first model M1. When the anomaly detection module 24 detects the first anomaly by using the first model M1, the anomaly detection module 24 can predict that a failure with an existing occurrence record will occur in the near future. Since the first model M1 has undergone supervised learning of historical data, the first model M1 can determine the state of the pump 2 based on past normal or abnormal operation data and can predict a failure when an anomaly is detected. However, when the first model M1 is applied to a pump 2 that does not have a past operation history, the first model M1 may have limitations.

[0063] Even if the pump has no past failure history, the anomaly detection module 24 can detect the abnormal state (i.e., anomaly) of the pump by using the trained second model M2 and third model M3. The second anomaly detected by the anomaly detection module 24 by using the second model M2 or the third anomaly detected by the anomaly detection module 24 by using the third model M3 is not determined based on past failures. The second anomaly or the third anomaly means that the pump 2 is outside the normal state. The second model M2 and the third model M3 can detect anomalies based on the difference between real-time data and normal operation data, and thus can even predict new types of failures.

[0064] The anomaly detection module 24 can predict various failures that may occur in the pump 2 by using the first model M1, the second model M2, and the third model M3 together.

[0065] The program code may further include a reporting module 25 that provides a notification that an anomaly has been detected in the pump 2 based on the output of the anomaly detection module 24. When an anomaly is detected in each of the first model M1, the second model M2, and the third model M3, the reporting module 25 can provide an anomaly detection notification.

[0066] The reporting module 25 can determine whether an anomaly has occurred by comprehensively judging the outputs of the first model M1, the second model M2, and the third model M3. When an anomaly is detected in at least two of the first model M1, the second model M2, and the third model M3, the reporting module 25 can provide a comprehensive anomaly detection notification. While providing the comprehensive anomaly detection notification, the reporting module 25 can provide the individual models in which the anomalies are detected.

[0067] The anomaly detection notification can indicate the first anomaly detected by using the first model M1, the second anomaly detected by using the second model M2, and the third anomaly detected by using the third model M3. The anomaly detection notification can indicate a comprehensive anomaly. The anomaly detection notification can be provided to the administrator 4 visually or auditorily through the input / output interface 14. The anomaly detection notification can be provided to the administrator 4 by phone, text message, email, etc. The input / output interface 14 can display the first anomaly, the second anomaly, and the third anomaly by using a display or an LED.

[0068] Figure 5 is a flowchart showing a method for predicting the failure of the pump 2 based on an embodiment. See Figure 5 and Figure 3 .

[0069] A method for predicting the failure of pump 2 according to an embodiment of the present disclosure may include: at S10, collecting real-time data related to the state of pump 2 that is assumed to operate normally without any problems in a steady state; at S20, generating training data by preprocessing the operation data formed by recording the real-time data over time; at S30, training a first model M1, a second model M2, and a third model M3 by using the training data; at S40, detecting a first anomaly that occurs before a failure by inputting the real-time data of pump 2 into the first model M1 that has performed supervised learning on historical data, detecting a second anomaly outside the normal operation range of pump 2 by inputting the real-time data of pump 2 into the second model M2 that has performed unsupervised learning on normal operation data, and detecting a third anomaly outside the initial normal operation range of pump 2 by inputting the real-time data of pump 2 into the third model M3 that has performed unsupervised learning on initial operation data; and based on the output of the anomaly detection at S40, reporting a notification of detecting an anomaly in pump 2 at S50.

[0070] The data collection at S10 can be performed in the data collection module 21. In the data collection at S10, real-time data is collected from an oil rig and stored in the storage device 12 to generate operation data. The data collection at S10 can be repeated in real time. The data collection at S10 can always be performed when the oil production unit is operating. In the data collection at S10, the administrator 4 can store the static data SD in the storage device 12. In the data collection at S10, the data collection module 21 can collect the monitoring data MD from sensors installed in the oil production unit, the pump controller 3, or other electronic devices at every predetermined time.

[0071] The generation of the training data at S20 can be performed in the training data generation module 22. The generation of the training data at S20 can further include a process of preprocessing the operation data. In the generation of the training data at S20, historical data for training the first model M1, normal operation data for training the second model M2, and initial operation data for training the third model M3 can be generated independently.

[0072] The training data generation module 22 can generate historical data by performing the generation of the training data at S20, and this historical data is the training data for generating the first model M1. In order to generate the training data of the first model M1 in the generation of the training data at S20, the time series data in the operation data is normalized, and the operation data corresponding to the past recorded failure states is marked as abnormal, and the data of the normal operation part is marked as normal, so that historical data can be generated.

[0073] In the generation of training data at S20, in order to generate the training data of the second model M2, principal component analysis is applied to the operation data at each specific time to extract the feature vectors. The reconstructed data is obtained from the feature vectors, and the error between the reconstructed data and the operation data is calculated. When the error is less than or equal to the reference error, the operation data is included in the normal operation data, and when the error is greater than the reference error, the operation data is excluded from the normal operation data, so that the normal operation data can be generated.

[0074] In the generation of training data at S20, in order to generate the training data of the third model M3, the operation data collected within a predetermined period starting from the time when the pump 2 has been operating in a normal state after the installation of the pump 2 can be generated as the initial operation data.

[0075] The generation of training data at S20 is executed when the model is first generated or when the model needs to be updated, so that the training data can be updated.

[0076] The model training at S30 can be executed in the training module 23. The model training at S30 is a process of making the model learn the training data. In the model training at S30, the first model M1, the second model M2, and the third model M3 can be trained independently of each other. The model training at S30 can be executed when the model is first generated or when the model needs to be updated.

[0077] The training module 23 can generate the first model M1, which classifies the historical data into a normal category or an abnormal category by performing the model training at S30. In the model training at S30, the operation data included in the historical data can be classified into a normal category or an abnormal category by using linear discriminant analysis (LDA).

[0078] The training module 23 can generate the second model M2, which can learn the normal range in the normal operation data by performing the model training at S30. In the model training at S30, the feature vectors are obtained by applying principal component analysis to the normal operation data, and the reconstructed data is obtained based on the feature vectors, and the process of obtaining the normal partial error between the normal operation data and the reconstructed data is repeated at each time step of the operation data to obtain a plurality of normal partial errors, and the second model M2 can be generated by learning the average value of the plurality of normal partial errors between the initial operation data (normal operation data) and the reconstructed operation data (reconstructed data) calculated after applying principal component analysis.

[0079] The training module 23 can generate a third model M3, and the third model M3 can learn the normal range in the initial operation data by performing model training at S30. In the training of the model, at each time step of the initial operation data, the process of obtaining at least two eigenvectors by applying principal component analysis to the initial operation data in the order of mainly considering the characteristics of the initial operation data is repeated to obtain the distribution and center point of the eigenvectors, so that the third model M3 can be generated.

[0080] The anomaly detection at S40 can be performed in the anomaly detection module 24. In the anomaly detection at S40, the real-time data can be analyzed by using each of the first model M1, the second model M2, and the third model M3. In the anomaly detection at S40, detecting whether the first anomaly occurs by using the first module, detecting whether the second anomaly occurs by using the second module, and detecting whether the third anomaly occurs by using the third module can be performed independently of each other.

[0081] In the anomaly detection at S40, the anomaly that may occur before the occurrence of a fault can be detected by inputting the real-time data into the first model M1. In the anomaly detection at S40, when the real-time data is classified into the anomaly category based on the linear discriminant criterion generated by the first model M1 according to the first model M1, it can be determined that the first anomaly occurs. The first model M1 is a linear discriminant analysis model trained to classify historical data including operation data marked as normal and operation data marked as abnormal into the normal category and the anomaly category.

[0082] In the anomaly detection at S40, the anomaly outside the normal operation range of the pump can be detected by inputting the real-time data into the second model M2. In the anomaly detection at S40, principal component analysis is applied to the normal operation data including the operation data obtained when the pump 2 operates in the normal state to obtain eigenvectors. Reconstructed data is obtained based on the eigenvectors, and the process of obtaining the normal partial error between the normal operation data and the reconstructed data is repeated at each time step of the operation data to obtain a plurality of normal partial errors. Based on the second model M2 generated by training the average value of the plurality of normal partial errors calculated between the initial operation data (normal operation data) and the reconstructed operation data (reconstructed data) after applying principal component analysis, principal component analysis is applied to the real-time data to obtain real-time eigenvectors. Real-time reconstructed data is obtained based on the real-time eigenvectors, and the real-time error between the real-time data and the real-time reconstructed data is obtained. The real-time error is divided by the average value of the plurality of normal partial errors generated by the second model M2 to obtain an error level, and the error level is recorded over time. When the error level increases over time, it can be determined that the second anomaly occurs.

[0083] In the detection of the third anomaly at S40, by inputting real-time data into the third model M3, anomalies outside the initial state assuming that pump 2 is operating normally without any problems in the steady state can be detected. In the detection of the third anomaly at S40, principal component analysis is applied to the initial operation data, which is operation data obtained within a predetermined time period starting from the time when it is assumed that pump 2 is operating normally without any problems in the steady state.

[0084] The third model M3 can be trained by applying principal component analysis to the initial operation data. After applying principal component analysis to the data, a feature vector consisting of at least two components is obtained. Then, the center point of the feature vector is obtained based on the distribution of the vectors in the multi-dimensional space. Based on the distribution and center point of the feature vectors, the third model calculates the Mahalanobis distance from the center point to each feature vector of the data. According to the range of the Mahalanobis distance, a reference distance is statistically determined. Then, principal component analysis is applied to the real-time operation data, and the third model calculates the Mahalanobis distance from the center point to each feature vector of the real-time data, and determines whether the state of the pump is abnormal when the Mahalanobis distance is greater than the reference distance.

[0085] The reporting of the notification at S50 can be executed in the reporting module 25. In the reporting of the notification at S50, based on the output of the anomaly detection at S40, a notification of the occurrence of an anomaly in pump 2 can be reported to the administrator 4.

[0086] In the reporting of the notification at S50, when the first anomaly is detected in the anomaly detection at S40 by outputting the anomaly category of the real-time data, an anomaly detection notification that the first model M1 has detected an anomaly can be provided to the administrator 4.

[0087] In the reporting of the notification at S50, when the second anomaly is detected by outputting the result that the error level of the real-time data analyzed in the anomaly detection at S40 is increasing, an anomaly detection notification that the second model M2 has detected an anomaly is provided.

[0088] In the reporting of the notification at S50, when the third anomaly is detected by outputting the result that the Mahalanobis distance of the real-time data analyzed in the anomaly detection at S40 is greater than the reference distance, an anomaly detection notification that the third model M3 has detected an anomaly is provided.

[0089] In the reporting of the notification at S50, when at least two of the first anomaly, the second anomaly, and the third anomaly are detected in the anomaly detection at S40, an anomaly detection notification that at least two models have detected an anomaly can be provided.

[0090] By checking the anomaly detection notification, the administrator 4 can identify a predicted failure of the pump 2. The administrator 4 can check the status of the pump 2 and has the opportunity to check the pump 2 before realizing that the pump 2 has failed. The pump failure prediction device 10 according to the present disclosure can reliably predict the failure of a pump having a failure history and can reliably identify a state other than the normal operation of the pump, so that unnecessary notifications are not generated. Therefore, the work required to check error notifications and the pump 2 can be reduced.

[0091] Figure 6 It is a diagram for explaining anomaly detection using the first model M1 based on an embodiment.

[0092] will be referred to Figure 6 to describe the generation of training data for generating the first model M1, the training of the first model M1, and the anomaly detection of the first model M1.

[0093] The first model M1 uses a trained linear discriminant analysis model to classify historical data into a normal category or an abnormal category. When the first model M1 receives real-time data of the pump 2, the first model M1 is trained to classify the data into either the normal category or the abnormal category. The historical data for training the first model M1 may include operating data marked as normal and operating data marked as abnormal.

[0094] In the generation of training data at S20, in order to generate the training data of the first model M1, the training data generation module 22 may standardize the time series data in the operating data. The process of standardizing the time series data in the operating data can increase the generality of the first model M1.

[0095] For the historical data, among the operating data formed by storing the real-time data collected by the data collection module 21 in the storage device 12, the operating data measured under the normal state is marked as normal, and the operating data measured under the failure state according to the failure record is marked as abnormal. Here, the failure state may include a state in which the pump 2 stops due to a failure or a state before the pump 2 stops due to a failure. Therefore, the first model M1 can learn the operating data of the state before the failure occurs and can detect the state before the failure occurs as the first anomaly. The operating data marked as abnormal is the data of the state before the failure occurs. Therefore, when similar operating data is detected, it can be predicted that a failure will occur in the near future.

[0096] The historical data may include the motor frequency, the inflow rate into the pump 2, the motor temperature, the fluid temperature, the motor voltage, the bus voltage, the motor current, and the temperature difference between the fluid and the pump 2 among the elements of the operating data, and the values of the elements may be measured at each predetermined time. The operating data can be marked as normal or abnormal each time the operation is performed. At Figure 6In the exemplary historical data LD1 shown, the number (No.) represents the number of the data, the time represents the time of the time interval for measuring the real-time data according to the data acquisition module 21, and multiple elements (e.g., element 1, element 2, ……) include a part of the second change data (e.g., Figure 4 of CD2). Data (e.g., AAA, BBB, ……) is measured for each element and is marked as normal when the operating state at the relevant time is in the normal state and marked as abnormal when the operating state at the relevant time is in the fault state.

[0097] In the model training at S30, the training module 23 can generate a first model M1 using a linear discriminant analysis model by obtaining a decision boundary that classifies the operating data included in the historical data into a normal category or an abnormal category by using linear discriminant analysis (LDA).

[0098] The training module 23 can search for a decision boundary that can classify the data at each time in the historical data into two categories of normal category and abnormal category. A decision boundary can be obtained to maximize the distance between the normal category and the abnormal category and minimize the distribution within the category. The trained first model M1 can minimize the distribution of the operating data included in the category and can have a decision boundary that maximizes the distribution of the operating data between the categories. When the first model M1 classifies the historical data into a normal category and an abnormal category, the training of the first model M1 can stop. The training module 23 can store the trained first model M1 in the storage device 12.

[0099] In the anomaly detection at S40, the anomaly detection module 24 can determine which category the real-time data belongs to based on the decision boundary of the first model M1. When the real-time data is classified into the abnormal category according to the linear discriminant criterion generated by the first model M1, the anomaly detection module 24 can determine that the first anomaly has occurred. The linear discriminant criterion can include the decision boundary. In the anomaly detection at S40, the real-time data determined by using the first model M1 can have the same elements as the historical data and can have no label.

[0100] The anomaly detection module 24 inputs the real-time data into the trained first model M1, where the first model M1 can be described as outputting the input real-time data belonging to one of the normal category and the abnormal category. When the data output by the first model M1 is classified as abnormal, the anomaly detection module 24 can determine that the first anomaly has occurred. The anomaly detection module 24 can provide the result of detecting the first anomaly or the result of determining the real-time data as the abnormal category to the reporting module 25.

[0101] In the report of the notification at S50, when the first model M1 outputs real-time data as an abnormal category, the reporting module 25 can provide an anomaly detection notification. When the data output by the first model M1 is classified as a normal category, the reporting module 25 does not determine that an anomaly has occurred. When the reporting module 25 receives an abnormal category from the anomaly detection module 24, the reporting module 25 can determine that a first anomaly has occurred and can provide an anomaly detection notification. For example, the reporting module 25 can display the presence of the first anomaly on the display of the input / output interface 14 and can automatically send a call or text to the contact of the administrator 4. Through the anomaly detection notification, the administrator 4 can identify a failure that will occur with a history of occurrence.

[0102] An update of the first model M1 will be described. When the pump 2 fails, the training data generation module 22 can update the historical data, and the training module 23 can retrain the first model M1 by using the updated historical data, and the anomaly detection module 24 can detect anomalies by using the retrained first model M1. The update of the first model M1 can be performed in the event of a failure, and the determination of the occurrence of the failure is made by the administrator 4. Therefore, the update of the first model M1 can be performed irregularly according to the instructions of the administrator 4.

[0103] Figure 7 is a diagram for explaining the training of the first model M1 with three categories based on the embodiment.

[0104] The first model M1 can be trained to have three categories, such as a normal category, an abnormal category, and an unstable category. In the case where the first model M1 is trained to have three categories and the real-time data is determined to be one of the three categories, the user can more detailedly identify the state of the pump 2.

[0105] In the generation of training data at S20, the training data generation module 22 can generate historical data, which includes operation data marked as normal, operation data marked as abnormal, and operation data marked as unstable. Here, the operation data marked as normal can be operation data obtained in a normal state, the operation data marked as abnormal can be operation data measured during a failure or within a predetermined period before the occurrence of a failure, and the operation data marked as unstable can be operation data in a normal or non-failure state. The operation data marked as unstable is operation data measured between the normal state and the failure state.

[0106] In the model training at S30, the training module 23 can generate the first model M1 by classifying the operation data included in the historical data into a normal category, an abnormal category, or an unstable category by using a linear discriminant analysis model. The trained first model M1 can classify the historical data into the same categories as Figure 7Regions corresponding to the three categories shown. The normal category may correspond to the first region Z1, the unstable category may correspond to the second region Z2, and the abnormal category may correspond to the third region Z3. As Figure 7 shown, the first region Z1, the second region Z2, and the third region Z3 do not overlap with each other, so the probability that the first model M1 misclassifies real-time data is very low.

[0107] In the anomaly detection at S40, when the real-time data is classified as an abnormal category or an unstable category according to the linear discriminant criterion generated by the first model M1, the anomaly detection module 24 can determine that a first anomaly has occurred. When the real-time data is classified as an unstable category in the anomaly detection at S40, it can be preset that the administrator 4 determines that there is no first anomaly.

[0108] In the reporting of the notification at S50, when the real-time data is classified as an unstable category or an abnormal category, the reporting module 25 can provide an anomaly detection notification to include the relevant category into which the real-time data is classified. When it is determined that a first anomaly has occurred because the real-time data is classified as an unstable category, or when it is determined that there is no first anomaly because the real-time data is classified as an unstable category, the administrator 4 can receive the anomaly detection notification.

[0109] Figure 8 is a diagram showing the result of displaying the anomaly detection result obtained over time by using the first model M1 based on the embodiment.

[0110] Figure 8 The horizontal axis of is the measurement time of the real-time data. As the analysis result of the real-time data, Figure 8 the vertical axis of indicates the normal state (NS) in a first color when no first anomaly is detected, and indicates the abnormal state (ANS) in a second color when a first anomaly is detected.

[0111] The reporting module 25 can display the results output by the anomaly detection module 24 over time to form Figure 8 a chart of. For illustration, Figure 8 shows the moment when the pump 2 stops due to a fault. It can be seen that the multiple times when the first anomaly is detected are distributed within three weeks before the moment when the pump 2 stops due to a fault. The administrator 4 can check the chart of the first anomaly displayed by the reporting module 25 over time and identify that a fault is likely to occur in the future.

[0112] Figure 9 is a diagram for explaining the anomaly detection using the second model M2 based on the embodiment.

[0113] will be described with reference to Figure 9 the generation of the training data for generating the second model M2, the training of the second model M2, and the anomaly detection by using the second model M2.

[0114] The second model M2 learns the normal operating range of the pump 2 by learning the normal operating data, which is used to determine whether the real-time data is outside the normal range.

[0115] In the generation of the training data at S20, by preprocessing all the operating data stored in the storage device 12, the data to be included in the training data and the data to be excluded from the training data can be separated. The normal operating data needs to include the operating data collected when the oil production unit and the pump 2 are operating normally.

[0116] To this end, in order to generate the training data of the second model M2, the training module 23 extracts the feature vectors by applying principal component analysis to the operating data at each specific time of the operating data, obtains the reconstructed data from the feature vectors, and calculates the error between the reconstructed data and the operating data. When the error is greater than the reference error, the operating data can be excluded from the normal operating data. When the error is less than or equal to the reference error, the operating data can be included in the normal operating data. When the pump 2 is operating normally, the error deviation between the operating data and the reconstructed data is not large. The operating data when the error is greater than the reference error can indicate that the pump 2 is outside the normal operating state.

[0117] Optionally, the training module 23 extracts the feature vectors by applying principal component analysis to the operating data at each specific time of the operating data, obtains the reconstructed data from the feature vectors, and calculates the error between the reconstructed data and the operating data. The errors are listed over time, and the operating data in the part where the errors continuously increase can be excluded from the normal operating data. The operating data in the part where the errors remain within a predetermined range can be included in the normal operating data.

[0118] The normal operating data can include the motor frequency, the inflow rate into the pump 2, the motor temperature, the fluid temperature, the motor voltage, the bus voltage, the motor current, and the temperature difference between the fluid and the pump 2 among the elements of the operating data, and the values of the elements can be measured at each predetermined time. The normal operating data has no labels. In Figure 9 In the exemplary normal operating data (LD2) shown, the number represents the number of the data, the time represents the time according to the time interval for measuring the real-time data by the data acquisition module 21, and multiple elements (for example, element 1, element 2,...) include a part of the second change data (for example, Figure 4 of CD2). The data (for example, AAA, BBB,...) can be measured for each element.

[0119] In the model training at S30, the second model M2 can be generated by learning normal operation data with a principal component analysis model. In the model training at S30, when applying principal component analysis to normal operation data to obtain feature vectors, the second model M2 can be generated by training to obtain the average of multiple normal partial errors. Based on the feature vectors, reconstructed data is obtained; the process of obtaining the normal partial error between the initial operation data (normal operation data) and the reconstructed operation data (reconstructed data) is repeated at each time step of the operation data to obtain multiple normal partial errors. The trained second model M2 can be stored in the storage device 12.

[0120] The principal component analysis model can reduce the dimension of data while preserving as much of the multi-dimensional data distribution as possible. By applying principal component analysis, multiple feature vectors considering the data characteristics can be obtained. When reconstructing the original data using multiple feature vectors, errors may occur. In order to compare the error obtained from normal operation data with the error obtained from real-time data, first apply principal component analysis to the normal operation data to be reconstructed to obtain the error, which is the process of generating the second model M2.

[0121] The process of generating the second model M2 will be described. For example, when applying principal component analysis to normal operation data including eight elements, multiple feature vectors can be obtained. Among the multiple feature vectors, two feature vectors can be selected in the order of mainly considering the characteristics of the normal operation data. Three or more feature vectors can be selected. Based on the two selected feature vectors, the eight elements can be reconstructed to obtain reconstructed data. For these eight elements, the difference between the operation data and the reconstructed data is obtained, the difference is squared, and the squared differences are added together to obtain the normal partial error. For example, the normal partial error can be the square of the difference between the operation data and the reconstructed data of the first element + the square of the difference between the operation data and the reconstructed data of the second element +... + the square of the difference between the operation data and the reconstructed data of the eighth element. The normal partial error can be obtained each time, so multiple normal partial errors can be obtained. For example, when the normal operation data appears at 15-minute intervals and appears 1000 times, one thousand normal partial errors can be obtained. Finally, the second model M2 can be generated by obtaining the average of multiple normal partial errors between the initial operation data (normal operation data) and the reconstructed operation data (reconstructed data) calculated after applying principal component analysis.

[0122] In the anomaly detection at S40, in the case where a real-time feature vector is obtained by applying principal component analysis to real-time data, when the error level of the real-time data increases over time, a second anomaly can be determined. Based on the real-time feature vector, real-time reconstructed data is obtained. A real-time error between the real-time data and the real-time reconstructed data is obtained; the error level is obtained by dividing the real-time error by the average of a plurality of normal partial errors generated by the second model M2; and the error level is recorded over time. The real-time data input to the second model M2 can have the same type as the operating data included in the normal operating data.

[0123] In the anomaly detection at S40, the anomaly detection module 24 can obtain a real-time feature vector by applying principal component analysis to a plurality of elements of the real-time data. The anomaly detection module 24 can obtain real-time reconstructed data by using the real-time feature vector. The difference between the real-time reconstructed data and the real-time data is obtained and squared, and all the squared differences are added together to obtain a real-time error. Dividing the real-time error by the average of the normal partial errors gives the error level. The error level is a value indicating how large the real-time error is relative to the average of the normal partial errors.

[0124] Figure 10 is a diagram showing the result of anomaly detection obtained over time using the second model M2 based on an embodiment.

[0125] When the error level is recorded over time, a Figure 10 graph as shown in can be obtained. A real-time error of a similar magnitude to each normal partial error can be formed near the error level 1. When the real-time error is large, the error level may also be large. When the real-time error gradually increases, the error level may also gradually increase. It can be determined that an increase in the error level indicates an increase in the degree to which the current operating state deviates from the normal range. Figure 10 shows the part where the error level gradually increases (pre-detection) and the part where the pump 2 stops due to a failure (shut down). In the part where the error level gradually increases (pre-detection), the error level significantly exceeds 1. The part where the error level gradually increases (pre-detection) can be observed before the part where the pump 2 stops (shut down). Therefore, when the part where the error level gradually increases is observed, the administrator 4 can identify that the pump 2 is operating outside the normal range.

[0126] In the report of the notification at S50, when the second model M2 detects a second anomaly, the reporting module 25 can provide an anomaly detection notification.

[0127] When the report module 25 receives an input for which the evaluation result of the output of the second model M2 from the anomaly detection module 24 is abnormal, the report module 25 can determine that a second anomaly has occurred and can provide an anomaly detection notification. For example, the report module 25 can display the presence of the second anomaly on the display of the input / output interface 14 and can automatically send a call or text to the contacts of the administrator 4. Through the anomaly detection notification, the administrator 4 can identify that a new type of failure may have occurred rather than a failure with a history of occurrence.

[0128] The process of updating the second model M2 will be described.

[0129] In the model training at S30, before the pump 2 is installed in the oil well 1 and assuming that the pump 2 operates normally in a steady state without any problems, the second model M2 can be generated by using the normal operation data of another pump with similar specifications to the pump 2. After the pump 2 is installed in the oil well 1 and assuming that the pump 2 operates normally in a steady state without any problems, the normal operation data of the pump 2 can be updated every preset time period to regenerate the second model M2.

[0130] Even if the same type of pump 2 is used, the operation data of the pump 2 installed in different oil wells 1 has different values. This is because the characteristics of the oil well 1 vary according to the location of the geological formation. Each of the pumps 2 may have different operation data due to various reasons. Therefore, the normal operation data of a specific pump 2 can only be obtained after the specific pump 2 is actually installed in the oil well 1 and operates. Therefore, the second model M2 is not immediately used after the pump 2 is installed in the oil well 1. To solve this problem, after the pump 2 is installed in the oil well 1, the second model M2 can be immediately generated by using the normal operation data obtained from the pumps 2 installed in different oil wells 1. This is because when the characteristics of different oil wells 1 are similar, the types of pumps 2 are similar, and the operation modes of the pumps 2 are similar, even the pumps 2 installed in different oil wells 1 can have similar normal operation data.

[0131] When the normal operation data of a specific pump 2 is sufficiently obtained while detecting anomalies by using the second model M2, the second model M2, which is generated by using the normal operation data of the pumps 2 installed in different oil wells 1, can be updated. When the specific pump 2 starts to operate normally, the second model M2 can be updated regularly every preset time period. When the pump 2 is operating, the states of each of the oil well 1 and the pump 2 are constantly changing, so the second model M2 also needs to be updated regularly.

[0132] Figure 11 FIG. is a diagram for explaining anomaly detection using the third model M3 based on the embodiment.

[0133] will be referred to Figure 11Describe the generation of training data for generating a third model M3, the training of the third model M3, and the anomaly detection of the third model M3.

[0134] The third model M3 learns the normal range of the initial part where the pump 2 is installed and operating normally by learning the initial operating data, and this is intended to determine whether the real-time data is outside the normal range of the initial part.

[0135] The third model M3 is a model that compares the current state of the pump 2 with the state of the pump 2 after the pump 2 is installed in the oil well 1 and assuming that the pump 2 is operating normally under stable conditions without any problems. In the initial stage when the pump 2 is installed in the oil well 1, it is less likely to fail due to other reasons without aging. Therefore, the third model M3 can learn the operating data collected in the initial stage when the pump 2 is installed and operating.

[0136] The operating data that can be included in the initial operating data is the data included in a predetermined period after the pump 2 starts to operate normally. For example, the initial operating data can include the data within seven days after the pump 2 starts to operate normally. In this embodiment, the initial operating data can include the operating data collected in a time period of about two to seven days after the pump 2 is installed. The time period for collecting the initial operating data may vary. However, preferably, the data on the day when the pump 2 is installed is excluded, and the operating data collected within one week after the pump 2 is installed is difficult to be classified as initial data, so it is preferably excluded.

[0137] When the oil production unit is operating and the pump 2 is used for oil production in the oil well 1, the state of the oil well 1 changes and the pump 2 ages. Therefore, the initial operating data obtained when the oil well 1 and the pump 2 are in a normal state without aging is preset as the standard for determining anomalies.

[0138] The initial operating data can include the motor frequency, the inflow rate into the pump 2, the motor temperature, the fluid temperature, the motor voltage, the bus voltage, the motor current, and the temperature difference between the fluid and the pump 2 among the elements of the operating data, and the values of the elements can be measured at each predetermined time. The initial operating data is not marked. In Figure 11 the exemplary initial operating data (LD3) shown, the number represents the number of the data, the time represents the time according to the time interval for measuring the real-time data by the data acquisition module 21, and multiple elements (e.g., element 1, element 2,...) include a part of the second change data (e.g., Figure 4 of CD2). The data (e.g., AAA, BBB,...) can be measured for each element.

[0139] In the model training at S30, the third model M3 can be generated by learning the normal operation data with a principal component analysis model. In the model training at S30, principal component analysis is applied to the initial operation data, and the process of obtaining at least two eigenvectors in the order mainly considering the characteristics of the initial operation data is repeated at each time step of the initial operation data to obtain the distribution and center point of the eigenvectors, thereby generating the third model M3. The trained third model M3 can be stored in the storage device 12.

[0140] The process of generating the third model M3 will be described. For example, when principal component analysis is applied to the initial operation data including eight elements, multiple eigenvectors can be obtained. Among the multiple eigenvectors, a predetermined number of eigenvectors can be selected in the order mainly considering the characteristics of the initial operation data. In this embodiment, two eigenvectors can be selected. By applying principal component analysis to the initial operation data at each time of the initial operation data, two eigenvectors can be obtained. The eigenvectors obtained each time can be arranged in a space with a predetermined dimension. In this embodiment, two eigenvectors are selected, so they can be arranged in a two-dimensional space. If three eigenvectors are selected, the eigenvectors can be arranged in a three-dimensional space. The multiple feature points shown in the dimensional space can be distributed in a specific form according to the characteristics of the initial operation data. The third model M3 can be generated by obtaining the center point of the eigenvectors based on the distribution of the eigenvectors. As Figure 11 shown, the eigenvectors arranged in the two-dimensional space can be distributed in a specific form. Based on the arranged positions of the eigenvectors, the center point (CP) of the eigenvectors can be obtained.

[0141] In the anomaly detection at S40, based on the distribution of the eigenvectors and the center point (CP), the third model M3 calculates the Mahalanobis distance from the center point (CP) to each eigenvector of the data. According to the range of the Mahalanobis distance, a reference distance is statistically determined. Then, principal component analysis is applied to the real-time operation data, and the third model M3 calculates the Mahalanobis distance from the center point (CP) to each eigenvector of the real-time data, and determines whether the state of the pump is abnormal when the Mahalanobis distance is greater than the reference distance.

[0142] In the anomaly detection at S40, the anomaly detection module 24 can obtain real-time eigenvectors by applying principal component analysis to multiple elements of the real-time data. The anomaly detection module 24 can arrange the real-time eigenvectors in the dimensional space of the third model M3. The anomaly detection module 24 can obtain the Mahalanobis distance Md2 between the real-time eigenvector RP arranged in the dimensional space of the third model M3 and the center point CP of the eigenvectors generated by the third model M3. When the Mahalanobis distance Md2 is greater than the reference distance, it can be determined that the third anomaly has occurred.

[0143] After obtaining multiple Mahalanobis distances Md1 by repeating, at each time step, the process of obtaining the Mahalanobis distance Md1 between the center point CP of the feature vectors obtained from the initial run data by the third model M3 and the real-time feature vectors, the reference distance can be determined as a multiple of the average value of the multiple Mahalanobis distances Md1. In this embodiment, the reference distance can be determined as a value three times the average value of the multiple Mahalanobis distances Md1 obtained by the third model M3 from the initial run data.

[0144] Figure 12 is a diagram showing the result of an anomaly detection result obtained over time by using the third model M3 based on an embodiment. Also refer to Figure 12 and Figure 11 .

[0145] In Figure 12 , the horizontal axis is time. In Figure 12 , the vertical axis is the Mahalanobis distance. Figure 12 is a graph recording the Mahalanobis distance Md2 between the real-time feature vector RP and the center point CP of the feature vectors trained by the third model M3 over time. Figure 12 The reference distance of Figure 12 is represented as a value three times the average value of the multiple Mahalanobis distances Md1 obtained by the third model M3 from the initial run data based on the embodiment. In

[0146] In the report of the notification at S50, when the third model M3 detects a third anomaly, the reporting module 25 can provide an anomaly detection notification.

[0147] When the reporting module 25 receives an abnormal evaluation result of the output of the third model M3 from the anomaly detection module 24, the reporting module 25 can determine that a third anomaly has occurred and can provide an anomaly detection notification. For example, the reporting module 25 can display the presence of the third anomaly on the display of the input / output interface 14 and can automatically send a call or text to the contact of the administrator 4. Through the anomaly detection notification, the administrator 4 can identify that a new type of failure may have occurred rather than a failure with a history of occurrence.

[0148] The third model M3 uses the operation data in the initial stage of the installation pump 2, and the initial operation data is only selected once and not updated. When repairing the pump 2, it is possible to additionally execute the generation of the initial operation data and the training of the third model M3 according to the command of the administrator 4.

[0149] Figure 13 It is a diagram showing the results of cumulative anomaly detection obtained over time by using the first model M1, the second model M2, and the third model M3 based on the embodiments.

[0150] In Figure 13 , the upper diagram shows the results of detecting the first anomaly by using the first model M1, the middle diagram shows the results of detecting the second anomaly by using the second model M2, and the lower diagram shows the results of detecting the third anomaly by using the third model M3. In Figure 13 , when viewing the part marked as the abnormal state (ANS), it can be checked that multiple first anomalies detected by using the first model M1 are marked, it can be checked that the error level detected by using the second model M2 gradually increases and maintains the increasing state of the error level, and it can be checked that the Mahalanobis distance detected by using the third model M3 suddenly increases and maintains the increasing state of the Mahalanobis distance.

[0151] The reporting module 25 can provide an anomaly detection notification to the administrator 4 by including the Figure 13 shown chart in the anomaly detection notification, and the administrator 4 can visually inspect the chart and can easily identify the abnormal operation state of the pump 2. The reporting module 25 can display an alarm in color for the model determined to have an anomaly.

[0152] The administrator 4 can receive the anomaly detection notification and can stop the operation of the pump 2. The reporting module 25 can stop the operation of the pump 2 when an anomaly is detected. If the pump (2) stops running, any anomaly occurring in the pump (2) can be prevented from further escalating.

[0153] The above has detailed the present disclosure through specific embodiments. These embodiments are intended to specifically illustrate the present disclosure, but the present disclosure is not limited thereto. Obviously, those skilled in the art can modify or improve the embodiments within the technical spirit scope of the present disclosure.

[0154] Simple modifications or changes to the present invention all fall within the scope of the present disclosure, and the specific protection scope of the present disclosure will be clarified by the appended claims.

Claims

1. A pump fault prediction device, comprising: At least one processor; A storage device communicatively connected to the processor and storing program code that runs in the processor; And A communicator communicatively connected to the processor, Wherein the program code includes: A data acquisition module that acquires real-time data related to the state of a pump installed in and operating in an oil well; And An anomaly detection module that detects three types of anomalies, including a first anomaly that occurs before a fault by inputting the real-time data into a first model that has performed supervised learning on historical data, a second anomaly outside the normal operating range of the pump by inputting the real-time data into a second model that has performed unsupervised learning on normal operating data, and a third anomaly outside the initial normal operating range of the pump by inputting the real-time data into a third model that has performed unsupervised learning on initial operating data.

2. The device according to claim 1, wherein the historical data includes two types of operating data: operating data marked as normal and operating data marked as abnormal, The first model uses linear discriminant analysis to classify the operating data included in the historical data into a normal category or an abnormal category, and When the real-time data is classified into the abnormal category according to the linear discriminant criterion generated by the first model, the anomaly detection module determines that the first anomaly has occurred.

3. The device according to claim 1, wherein the normal operating data is the operating data obtained when the pump operates in a normal state, The second model is generated by obtaining the average value of multiple normal partial errors between the initial operating data, i.e., the normal operating data, and the reconstructed operating data, i.e., the reconstructed data, after applying principal component analysis, The principal component analysis is applied to the initial operating data to obtain eigenvectors, Based on the eigenvectors, the second model reconstructs the reconstructed data, and the process of obtaining the normal partial errors between the initial operating data and the reconstructed operating data is repeated at each time step of the operating data to obtain multiple normal partial errors, and In the case where real-time eigenvectors are obtained by applying the principal component analysis to the real-time data, when the error level of the real-time data increases over time, the anomaly detection module determines that the second anomaly has occurred, and Based on the real-time eigenvectors, real-time reconstructed data is obtained; the real-time error between the real-time data and the real-time reconstructed data is obtained; the error level is obtained by dividing the real-time error by the average value of the multiple normal partial errors generated by the second model; and the error level is recorded over time.

4. The device according to claim 1, wherein the initial operating data is the operating data obtained within a predetermined period starting from a time when it is assumed that the pump operates normally in a stable state without any problems, The third model is trained by applying principal component analysis to the initial operating data to obtain an eigenvector composed of at least two components, Based on the distribution of the feature vectors in the multi-dimensional space, the center point of the feature vectors is obtained, and Based on the distribution and the center point of the feature vectors, the third model calculates the Mahalanobis distance from the center point to each feature vector of the data, wherein a reference distance is statistically determined according to the range of the Mahalanobis distance, wherein the principal component analysis is applied to the real-time operation data, and the third model calculates the Mahalanobis distance from the center point to each feature vector of the real-time data, and determines whether the state of the pump is abnormal when the Mahalanobis distance is greater than the reference distance.

5. The device according to claim 4, wherein the reference distance is determined as a multiple of the average value of a plurality of Mahalanobis distances obtained from the initial operation data.

6. The device according to claim 1, wherein the program code further includes a reporting module, and the reporting module provides a notification of detecting an abnormality in the pump based on the output of the abnormality detection module.

7. The device according to claim 3, wherein the second model is generated and trained by using the normal operation data of another pump operating in a similar environment before the pump operates normally, and is regenerated by updating the normal operation data with the operation data of the pump itself in the normal state.

8. The apparatus according to claim 1, wherein the program code further comprises: A training data generation module, which generates training data for training the first model, the second model, and the third model by generating operation data by recording the real-time data over time and preprocessing the operation data, and using the preprocessed operation data The training data generation module generates historical data in a manner of normalizing the time series data in the operation data to generate the training data of the first model, wherein the operation data corresponding to the previously recorded fault state is marked as abnormal, and the data of the normal operation part is marked as normal; The training data generation module further generates the normal operation data in a manner of applying the principal component analysis to the operation data at each time step to extract feature vectors to generate the training data of the second model; the reconstructed data is obtained by reconstructing the feature vectors; Calculate the error between the reconstructed data and the operation data; When the error is less than or equal to the reference error, the operation data is included in the normal operation data; and when the error is greater than the reference error, the operation data is excluded from the normal operation data; and The training data generation module further generates the operation data collected within a predetermined period starting from the time when the pump operates in the normal state after the pump is installed, as the initial operation data for generating the training data of the third model.

9. A pump fault prediction method, comprising: Collecting real-time data related to the state of a pump installed in an oil well; Generating training data by preprocessing the operation data collected by recording the real-time data over time; Training a first model, a second model, and a third model by using the training data A first anomaly that appears before a fault occurs is detected by inputting real-time data of the pump into a first model that has performed supervised learning on historical data, a second anomaly outside the normal operating range of the pump is detected by inputting the real-time data of the pump into a second model that has performed unsupervised learning on normal operating data, and a third anomaly outside the initial normal operating range of the pump is detected by inputting the real-time data of the pump into a third model that has performed unsupervised learning on initial operating data. Based on the output of the anomaly detection, a notification that an anomaly has been detected in the pump is provided.

10. The method according to claim 9, wherein in the step of further detecting the first anomaly, when the real-time data is classified into an anomaly category based on a linear discriminant criterion generated by the first model according to the first model, it is determined that the first anomaly has occurred, and the first model is a linear discriminant analysis model trained to classify historical data including operating data marked as normal and operating data marked as abnormal into a normal category and an anomaly category.

11. The method according to claim 9, wherein in the step of further detecting the second anomaly, in the case of applying principal component analysis to normal operation data including operation data obtained when the pump operates in a normal state to obtain eigenvectors, when the error level of the real-time data increases over time, it is determined that the second anomaly occurs, wherein the reconstructed data is obtained by reconstructing the eigenvectors; The process of obtaining the normal partial error between the initial operating data, i.e., the normal operating data, and the reconstructed operating data, i.e., the reconstructed data, calculated after applying principal component analysis, is repeated at each time step of the operating data to obtain a plurality of normal partial errors; based on a second model generated by training by obtaining the average value of the plurality of normal partial errors, principal component analysis is applied to the real-time data to obtain a real-time feature vector; based on the real-time feature vector, real-time reconstructed data is obtained; the real-time error between the real-time data and the real-time reconstructed data is obtained; the real-time error is divided by the average value of the plurality of normal partial errors generated by the second model to obtain the error level; And the error level is recorded over time.

12. The method according to claim 9, wherein in the step of further detecting the third anomaly, in the case of applying principal component analysis to the initial operating data, when the Mahalanobis distance is greater than a reference distance, it is determined that the third anomaly has occurred, and the initial operating data is the operating data obtained within a predetermined period starting from when the pump is installed in the oil well and assuming that the pump operates normally without any problems in a stable state. Based on the distribution and center point of the feature vectors, the third model calculates the Mahalanobis distance from the center point to each feature vector of the data, statistically determines the reference distance according to the range of the Mahalanobis distance, applies the principal component analysis to the real-time operating data, and the third model calculates the Mahalanobis distance from the center point to each feature vector of the real-time data, and determines whether the state of the pump is abnormal when the Mahalanobis distance is greater than the reference distance.

13. The method according to claim 9, wherein in the step of generating the training data, To generate the training data for the first model, historical data is generated in a manner that normalizes the time series data in the operating data. The operating data corresponding to the previously recorded fault states is marked as abnormal, and the data for the normal operating part is marked as normal. To generate the training data for the second model, the normal operating data is generated by applying principal component analysis to the operating data at specific time intervals to extract feature vectors; the reconstructed data is obtained from the feature vectors; the error between the reconstructed data and the operating data is calculated; when the error is less than or equal to the reference error, the operating data is included in the normal operating data; and when the error is greater than the reference error, the operating data is excluded from the normal operating data, and To generate the training data for the third model, the operating data collected within a predetermined period starting from the time when the pump starts operating in a normal state after the pump is installed is generated as the initial operating data.

14. The apparatus according to claim 9, wherein in the step of training the second model, the second model is generated by using the normal operating data of another pump operating in a similar environment before the pump operates normally, and the second model is regenerated by updating the normal operating data with the operating data of the pump itself in a normal state.

Citation Information

Cited By

  • Pump shaft anomaly detection method

    CN121382665A

  • A method for detecting pump shaft abnormalities

    CN121382665B