A method, apparatus and storage medium for identifying abnormal daily production data
By employing a two-stage screening method combining isolated forest and cluster analysis with pre-defined domain knowledge conditions, the problem of cumbersome process and missed detection in detecting abnormal daily production data is solved, enabling rapid and accurate detection of abnormal daily production data.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-07
- Publication Date
- 2026-03-06
AI Technical Summary
Existing technologies are cumbersome and time-consuming in detecting abnormal production data, and are also subject to strong subjectivity and the risk of missed detections.
An isolated forest-cluster analysis combined method was used to initially screen daily production data, and then verified by combining it with preset domain knowledge conditions. Abnormal daily production data were identified through two-level screening.
It significantly shortens the detection time, improves detection accuracy, and avoids missed detections.
Smart Images

Figure CN115840918B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of oil and gas extraction technology, specifically to a method, apparatus, and storage medium for determining abnormal daily production data. Background Technology
[0002] Daily production data from production wells is entered automatically by the system or manually. However, if the metering equipment malfunctions or personnel are negligent, the entered daily production data may differ from the actual daily production data. This can lead to biased decisions when daily production data is used as a crucial basis for implementing various production measures. Therefore, it is necessary to detect any abnormal data in the daily production data to ensure its accuracy.
[0003] The relevant technologies mostly use manual experience to detect abnormal daily production data. That is, during the well opening phase, the daily production data is checked one by one using the production decline pattern to find the daily production data that obviously does not conform to the pattern and identify the daily production data as abnormal daily production data. During the well shut-in phase, daily production data where the daily output or daily production time is not zero is identified as abnormal daily production data.
[0004] However, the manual experience method requires verifying all daily production data one by one based on the production history and production decline pattern of each well. This verification process is tedious and time-consuming. Furthermore, this method is highly subjective and carries a certain risk of missing data. Summary of the Invention
[0005] This invention provides a method for determining abnormal day production data, addressing the problems of cumbersome and time-consuming processes in determining abnormal day production data using related technologies, and the risk of missed detection due to strong subjectivity. The technical solution is as follows:
[0006] Firstly, a method for determining abnormal daily production data is provided, the method comprising:
[0007] Obtain M daily production data points for the target production well, where the target production well refers to the production well whose abnormal daily production data is to be studied. The daily production data includes daily output, daily nozzle size, and daily production time, and M is a positive integer greater than or equal to 30.
[0008] Based on the M daily production data, the initial abnormal daily production data are determined using a combined isolated forest-cluster analysis method.
[0009] Based on the initial abnormal daily production data, the M daily production data, and the preset domain knowledge conditions, the abnormal daily production data of the target production well are determined. The preset domain knowledge conditions refer to the rule conditions applicable to the screening of abnormal daily production data, which are formulated based on oil and gas extraction theory and practice.
[0010] Optionally, the step of determining initial abnormal daily production data based on the M daily production data using a combined isolated forest-cluster analysis method includes:
[0011] Based on the M daily production data, N production data sequences are determined, where N is a positive integer greater than or equal to 1 and less than M;
[0012] Based on the N production data sequences, the anomaly score corresponding to the daily production data of each of the N production data sequences is determined by the Isolation Forest algorithm, which is an outlier detection method based on the degree of sorting.
[0013] Based on the anomaly score corresponding to the daily production data of each of the N production data sequences, the initial abnormal daily production data is determined by a clustering analysis algorithm.
[0014] Optionally, determining N production data sequences based on the M daily production data includes:
[0015] Iterate through the M daily production data;
[0016] The daily production data with a daily output or daily production time of 0 among the M daily production data are divided into a non-production stage production data sequence.
[0017] Based on the nozzle size of the day, the daily production data that are not divided into non-production stage production data sequences in the M daily production data are divided into N-1 production stage production data sequences, wherein the daily production data of each production stage production data sequence in the N-1 production stage production data sequences have the same nozzle size of the day.
[0018] Based on the production data sequence of the non-production stage and the production data sequences of the N-1 production stages, the N production data sequences are determined.
[0019] Optionally, the step of determining the initial abnormal daily production data based on the anomaly score corresponding to the daily production data of each of the N production data sequences, using a clustering analysis algorithm, includes:
[0020] Arrange all the anomaly scores corresponding to each of the N production data sequences in descending order to obtain N anomaly score datasets;
[0021] Let r = 1, and determine the abnormal scores in the r-th abnormal score dataset that are located at the 95th percentile as the outlier representatives, and determine the abnormal scores that are located at the 5th percentile as the normal values representatives;
[0022] The outlier representative and the normal value representative are determined as the initial center;
[0023] Based on the initial centers and the abnormal scores in the r-th abnormal score dataset excluding the outlier representative and the normal value representative, the k-medoids algorithm is used to determine the abnormal score cluster set of the r-th abnormal score dataset. The k-medoids algorithm is a clustering analysis algorithm based on the centroid distance algorithm.
[0024] Determine whether r is equal to N;
[0025] If r is not equal to N, then let r = r + 1, and return to the step of determining the abnormal scores in the r-th abnormal score dataset that are located at the 95th percentile as the abnormal value representative and the abnormal scores that are located at the 5th percentile as the normal value representative.
[0026] If r equals N, then the daily production data corresponding to the abnormal scores in the N clusters of abnormal scores are determined as the initial abnormal daily production data.
[0027] Optionally, determining the abnormal daily production data of the target production well based on the initial abnormal daily production data, the M daily production data, and preset domain knowledge conditions includes:
[0028] Based on the initial abnormal day production data and the M days production data, determine the control day production data that is closest to the initial abnormal day production data upstream on the time axis;
[0029] Based on the initial abnormal day's production data and the control day's production data, the flow rate of the initial abnormal day's production data and the flow rate of the control day's production data are determined using the following formulas.
[0030] v=4q / (πCh 2 Pd)
[0031] In the formula: v is the flow rate, q is the daily output, Pd is the daily production time, and Ch is the nozzle size for the day;
[0032] The production data of the initial abnormal day and the flow rate of the initial abnormal day are determined as the production data of the intermediate abnormal day.
[0033] The production data of the control day and the flow rate of the production data of the control day are determined as the production data of the intermediate control day;
[0034] Based on the intermediate abnormal production data, the intermediate control daily production data, and the preset domain knowledge conditions, the abnormal daily production data of the target production well is determined.
[0035] Optionally, determining the abnormal daily production data of the target production well based on the intermediate abnormal production data, the intermediate control daily production data, and the preset domain knowledge conditions includes:
[0036] Based on the intermediate abnormal production data and the intermediate control daily production data, and using the following preset domain knowledge conditions, the intermediate abnormal production data is marked as normal daily production data.
[0037] If both the daily production time and daily output in the intermediate abnormal production data are 0, then the intermediate abnormal production data will be marked as normal daily production data; or,
[0038] If the ratio of the nozzle size of the intermediate control day production data to the nozzle size of the intermediate abnormal production data is not equal to 1, then the intermediate abnormal production data is marked as normal day production data; or,
[0039] If the flow rate in the intermediate abnormal production data is greater than 0.5 times the flow rate of the intermediate control day production data, but less than 2 times the flow rate of the intermediate control day production data, then the intermediate abnormal production data is marked as normal day production data.
[0040] If the intermediate abnormal production data is not marked as normal daily production data, then the daily production data corresponding to the intermediate abnormal production data is determined as the abnormal daily production data of the target production well.
[0041] Secondly, an apparatus for determining abnormal daily production data is provided, the apparatus comprising:
[0042] The acquisition module is used to acquire M daily production data of the target production well. The target production well refers to the production well whose abnormal daily production data is to be studied. The daily production data includes daily output, nozzle size on the day and daily production time. M is a positive integer greater than or equal to 30.
[0043] The first determining module is used to determine the initial abnormal daily production data based on the M daily production data using a combined method of isolated forest and cluster analysis.
[0044] The second determining module is used to determine the abnormal daily production data of the target production well based on the initial abnormal daily production data, the M daily production data, and preset domain knowledge conditions. The preset domain knowledge conditions refer to the rule conditions applicable to the screening of abnormal daily production data, which are formulated based on oil and gas extraction theory and practice.
[0045] Optionally, the first determining module includes:
[0046] The first determining unit is used to determine N production data sequences based on the M daily production data, where N is a positive integer greater than or equal to 1 and less than M;
[0047] The second determining unit is used to determine the anomaly score corresponding to the daily production data of each of the N production data sequences based on the N production data sequences and through the isolated forest algorithm. The isolated forest algorithm is an outlier detection method based on the degree of sorting.
[0048] The third determining unit is used to determine the initial abnormal daily production data based on the anomaly score corresponding to the daily production data of each of the N production data sequences, using a clustering analysis algorithm.
[0049] Optionally, the first determining unit specifically includes:
[0050] Iterate through the M daily production data;
[0051] The daily production data with a daily output or daily production time of 0 among the M daily production data are divided into a non-production stage production data sequence.
[0052] Based on the nozzle size of the day, the daily production data that are not divided into non-production stage production data sequences in the M daily production data are divided into N-1 production stage production data sequences, wherein the daily production data of each production stage production data sequence in the N-1 production stage production data sequences have the same nozzle size of the day.
[0053] Based on the production data sequence of the non-production stage and the production data sequences of the N-1 production stages, the N production data sequences are determined.
[0054] Optionally, the third determining unit specifically includes:
[0055] Arrange all the anomaly scores corresponding to each of the N production data sequences in descending order to obtain N anomaly score datasets;
[0056] Let r = 1, and determine the abnormal scores in the r-th abnormal score dataset that are located at the 95th percentile as the outlier representatives, and determine the abnormal scores that are located at the 5th percentile as the normal values representatives;
[0057] The outlier representative and the normal value representative are determined as the initial center;
[0058] Based on the initial centers and the abnormal scores in the r-th abnormal score dataset excluding the outlier representative and the normal value representative, the k-medoids algorithm is used to determine the abnormal score cluster set of the r-th abnormal score dataset. The k-medoids algorithm is a clustering analysis algorithm based on the centroid distance algorithm.
[0059] Determine whether r is equal to N;
[0060] If r is not equal to N, then let r = r + 1, and return to the step of determining the abnormal scores in the r-th abnormal score dataset that are located at the 95th percentile as the abnormal value representative and the abnormal scores that are located at the 5th percentile as the normal value representative.
[0061] If r equals N, then the daily production data corresponding to the abnormal scores in the N clusters of abnormal scores are determined as the initial abnormal daily production data.
[0062] Optionally, the second determining module includes:
[0063] The fourth determining unit is used to determine, based on the initial abnormal day production data and the M daily production data, the control day production data that is closest to the initial abnormal day production data upstream on the time axis.
[0064] The fifth determining unit is used to determine the flow rate of the initial abnormal day's production data and the flow rate of the control day's production data based on the initial abnormal day's production data and the control day's production data using the following formula:
[0065] v=4q / (πCh 2 Pd)
[0066] In the formula: v is the flow rate, q is the daily output, Pd is the daily production time, and Ch is the nozzle size for the day;
[0067] The sixth determining unit is used to determine the initial abnormal day production data and the flow rate of the initial abnormal day production data as intermediate abnormal day production data;
[0068] The seventh determining unit is used to determine the production data of the control day and the flow rate of the production data of the control day as the production data of the intermediate control day;
[0069] The eighth determining unit is used to determine the abnormal daily production data of the target production well based on the intermediate abnormal production data, the intermediate control daily production data, and the preset domain knowledge conditions.
[0070] Optionally, the eighth determining unit specifically includes:
[0071] Based on the intermediate abnormal production data and the intermediate control daily production data, and using the following preset domain knowledge conditions, the intermediate abnormal production data is marked as normal daily production data.
[0072] If both the daily production time and daily output in the intermediate abnormal production data are 0, then the intermediate abnormal production data will be marked as normal daily production data; or,
[0073] If the ratio of the nozzle size of the intermediate control day production data to the nozzle size of the intermediate abnormal production data is not equal to 1, then the intermediate abnormal production data is marked as normal day production data; or,
[0074] If the flow rate in the intermediate abnormal production data is greater than 0.5 times the flow rate of the intermediate control day production data, but less than 2 times the flow rate of the intermediate control day production data, then the intermediate abnormal production data is marked as normal day production data.
[0075] If the intermediate abnormal production data is not marked as normal daily production data, then the daily production data corresponding to the intermediate abnormal production data is determined as the abnormal daily production data of the target production well.
[0076] Thirdly, an apparatus for determining abnormal daily production data is provided, the apparatus comprising:
[0077] Processor and memory for storing processor-executable instructions;
[0078] The processor is configured to perform any of the methods described in the first aspect above.
[0079] Fourthly, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when executed by a processor, the computer program implements any of the methods provided in the first aspect above.
[0080] The technical solutions provided in this application can bring at least the following beneficial effects:
[0081] In this embodiment of the invention, M daily production data points of a target production well can be obtained. The target production well refers to the production well whose abnormal daily production data is to be studied. The daily production data includes daily output, nozzle size, and daily production time, where M is a positive integer greater than or equal to 30. Based on the M daily production data points, initial abnormal daily production data are determined using a combined isolated forest and cluster analysis method. Based on the initial abnormal daily production data, the M daily production data points, and preset domain knowledge conditions, the abnormal daily production data of the target production well are determined. The preset domain knowledge conditions refer to rules and conditions applicable to screening production anomalies, formulated based on oil and gas extraction theory and practice. In other words, the basic idea of this invention is to first perform a preliminary screening of the daily production data using a combined isolated forest and cluster analysis method when detecting abnormal daily production data. Then, the daily production data determined as initial abnormal daily production data after the preliminary screening are verified according to the preset domain knowledge conditions. This two-stage screening simplifies the process in related technologies that requires checking all daily production data one by one based on the production history and production decline pattern of the production well, significantly shortening the detection time. Furthermore, thanks to the improved clustering analysis algorithm, the detection accuracy has been enhanced, avoiding missed detections. Attached Figure Description
[0082] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0083] Figure 1 This is a flowchart illustrating a method for determining abnormal daily production data provided in an embodiment of the present invention;
[0084] Figure 2 This is a flowchart illustrating another method for determining abnormal daily production data provided in an embodiment of the present invention;
[0085] Figure 3 This is a schematic diagram of the structure of an abnormal day production data determination device provided in an embodiment of the present invention;
[0086] Figure 4 This is a schematic diagram of the structure of a terminal 400 provided in an embodiment of the present invention. Detailed Implementation
[0087] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0088] Before providing a detailed explanation of the embodiments of the present invention, the terms, application scenarios, and system architectures involved in the embodiments of the present invention will be explained.
[0089] First, the terms used in the embodiments of this invention will be introduced.
[0090] Daily output
[0091] Daily production refers to the total amount of oil or natural gas extracted by a target production well in a single day. When oil is used as the statistical indicator, the unit of daily production is t / d; when natural gas is used as the statistical indicator, the unit of daily production is m³. 3 / d. Daily production is an important basis for reflecting the oil-water variation pattern in the reservoir, verifying production management measures, analyzing dynamic changes, and formulating the next adjustment measures.
[0092] Nozzle size for the day
[0093] A nozzle is a throttle device used to control and regulate the production of a flowing well. The daily nozzle size refers to the orifice size of the nozzle used on the day the target production well's daily production data is recorded.
[0094] Daily production time
[0095] Production time refers to the actual operating time of the pumping unit when the target production well is normally pumping fluid on the day the daily production data is recorded. The unit of daily production time is hours.
[0096] Isolation Forest-Cluster Analysis Joint Method
[0097] The Isolation Forest-Cluster Analysis Joint Method refers to an algorithm that combines the Isolation Forest algorithm and the Cluster Analysis algorithm to achieve automated and accurate classification of data samples.
[0098] Secondly, the application scenarios involved in the embodiments of the present invention will be introduced.
[0099] When the production output of a well decreases or the water content increases, it is necessary to adjust the production system or take various adjustment measures to stabilize the production. Daily production data from the well is required to scientifically and effectively formulate production systems or adjustment measures. However, directly using currently recorded daily production data may lead to risks in the formulated production systems or adjustment measures due to outliers. Therefore, before using the daily production data, it is necessary to detect and eliminate abnormal data. In this case, the abnormal daily production data determination method provided in this embodiment of the invention can achieve automated, rapid, and accurate screening of daily production data through a two-level screening mode, providing accurate data for subsequent work.
[0100] Finally, the system architecture involved in the embodiments of the present invention will be described.
[0101] The method for determining abnormal daily production data provided in this embodiment of the invention can be applied to a terminal with data processing capabilities. Specifically, the terminal can be a smartphone, tablet computer, laptop computer, desktop computer, or other terminal capable of data processing.
[0102] Figure 1 This is a flowchart illustrating a method for determining abnormal daily production data according to an embodiment of the present invention. See also... Figure 1 The method includes the following steps:
[0103] Step 101: Obtain M daily production data for the target production well. The target production well refers to the production well whose abnormal daily production data is to be studied. The daily production data includes daily output, nozzle size on the day, and daily production time. M is a positive integer greater than or equal to 30.
[0104] Step 102: Based on the production data of M days, determine the initial abnormal daily production data through the combined method of isolated forest and cluster analysis.
[0105] Step 103: Based on the initial abnormal daily production data, M daily production data, and preset domain knowledge conditions, determine the abnormal daily production data of the target production well. The preset domain knowledge conditions refer to the rule conditions applicable to the screening of abnormal daily production data based on oil and gas extraction theory and practice.
[0106] In this embodiment of the invention, M daily production data points of a target production well can be obtained. The target production well refers to the production well whose abnormal daily production data is to be studied. The daily production data includes daily output, nozzle size, and daily production time, where M is a positive integer greater than or equal to 30. Based on the M daily production data points, initial abnormal daily production data are determined using a combined isolated forest and cluster analysis method. Based on the initial abnormal daily production data, the M daily production data points, and preset domain knowledge conditions, the abnormal daily production data of the target production well are determined. The preset domain knowledge conditions refer to rules and conditions applicable to screening production anomalies, formulated based on oil and gas extraction theory and practice. In other words, the basic idea of this invention is to first perform a preliminary screening of the daily production data using a combined isolated forest and cluster analysis method when detecting abnormal daily production data. Then, the daily production data determined as initial abnormal daily production data after the preliminary screening are verified according to the preset domain knowledge conditions. This two-stage screening simplifies the process in related technologies that requires checking all daily production data one by one based on the production history and production decline pattern of the production well, significantly shortening the detection time. Furthermore, thanks to the improved clustering analysis algorithm, the detection accuracy has been enhanced, avoiding missed detections.
[0107] Optionally, based on M days of production data, the initial abnormal daily production data are determined using a combined method of isolated forest and cluster analysis, including:
[0108] Based on M daily production data, determine N production data sequences, where N is a positive integer greater than or equal to 1 and less than M;
[0109] Based on N production data sequences, the anomaly score corresponding to the daily production data of each production data sequence in the N production data sequences is determined by the Isolation Forest algorithm. The Isolation Forest algorithm is an outlier detection method based on the degree of sorting.
[0110] Based on the anomaly score corresponding to the daily production data of each of the N production data sequences, the initial abnormal daily production data is determined through cluster analysis algorithm.
[0111] Optionally, based on M daily production data, N production data sequences are determined, including:
[0112] Iterate through the production data of M days;
[0113] Divide the daily production data with a daily output or daily production time of 0 from the M daily production data into a non-production stage production data sequence;
[0114] Based on the nozzle size of the day, the daily production data that are not divided into non-production stage production data sequences in the M daily production data are divided into N-1 production stage production data sequences. Among them, the daily production data of each production stage production data sequence in the N-1 production stage production data sequences have the same nozzle size of the day.
[0115] Based on the production data sequence of the non-production stage and the production data sequence of N-1 production stages, N production data sequences are determined.
[0116] Optionally, based on the anomaly score corresponding to the daily production data of each of the N production data sequences, a clustering analysis algorithm is used to determine the initial abnormal daily production data, including:
[0117] Arrange all the anomaly scores corresponding to each of the N production data sequences in descending order to obtain N anomaly score datasets.
[0118] Let r = 1, and determine the abnormal scores in the r-th abnormal score dataset that are located at the 95th percentile as the outlier representatives, and determine the abnormal scores that are located at the 5th percentile as the normal values representatives;
[0119] The outlier and normal values were used as the initial centers.
[0120] Based on the initial centers and the abnormal scores in the r-th abnormal score dataset excluding the outlier and normal value representatives, the k-medoids algorithm is used to determine the abnormal score cluster set of the r-th abnormal score dataset. The k-medoids algorithm is a clustering analysis algorithm based on the centroid distance algorithm.
[0121] Determine if r equals N;
[0122] If r is not equal to N, let r = r + 1, and return the steps of identifying the outlier scores in the r-th outlier score dataset that are at the 95th percentile as outlier representatives and the outlier scores that are at the 5th percentile as normal representatives.
[0123] If r equals N, then the daily production data corresponding to the abnormal scores in the N clusters of abnormal scores will be determined as the initial abnormal daily production data.
[0124] Optionally, based on the initial abnormal daily production data, M daily production data, and preset domain knowledge conditions, the abnormal daily production data of the target production well is determined, including:
[0125] Based on the initial abnormal day production data and M days of production data, determine the control day production data that is closest to the initial abnormal day production data upstream on the time axis;
[0126] Based on the production data from the initial abnormal day and the production data from the control day, the flow rate of the production data from the initial abnormal day and the flow rate of the production data from the control day are determined using the following formula.
[0127] v=4q / (πCh 2 Pd)
[0128] In the formula: v is the flow rate, q is the daily output, Pd is the daily production time, and Ch is the nozzle size for the day;
[0129] The production data of the initial abnormal day and the flow rate of the initial abnormal day production data are determined as the production data of the intermediate abnormal day.
[0130] The production data of the control day and the flow rate of the control day production data are determined as the intermediate control day production data;
[0131] Based on intermediate abnormal production data, intermediate control daily production data, and preset domain knowledge conditions, the abnormal daily production data of the target production well is determined.
[0132] Optionally, based on intermediate abnormal production data, intermediate control daily production data, and preset domain knowledge conditions, the abnormal daily production data of the target production well is determined, including:
[0133] Based on intermediate abnormal production data and intermediate control daily production data, and using the following pre-defined domain knowledge conditions, intermediate abnormal production data is marked as normal daily production data.
[0134] If both the daily production time and daily output in the intermediate abnormal production data are 0, then the intermediate abnormal production data will be marked as normal daily production data; or,
[0135] If the ratio of the nozzle size of the intermediate comparison day's production data to the nozzle size of the intermediate abnormal production data is not equal to 1, then the intermediate abnormal production data will be marked as normal day's production data; or,
[0136] If the flow rate in the intermediate abnormal production data is greater than 0.5 times the flow rate of the intermediate control day production data, but less than 2 times the flow rate of the intermediate control day production data, then the intermediate abnormal production data will be marked as normal day production data.
[0137] If intermediate abnormal production data is not marked as normal daily production data, then the daily production data corresponding to the intermediate abnormal production data will be identified as the abnormal daily production data of the target production well.
[0138] All the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of the present invention, and the embodiments of the present invention will not be described in detail one by one.
[0139] Figure 2 This is a flowchart illustrating another method for determining abnormal daily production data provided in an embodiment of the present invention. See also... Figure 2 The method includes the following steps:
[0140] Step 201: Obtain M daily production data for the target production well. The target production well refers to the production well whose abnormal daily production data is to be studied. The daily production data includes daily output, nozzle size on the day, and daily production time. M is a positive integer greater than or equal to 30.
[0141] Daily production output and daily production time can be input by the user, sent by other devices, or recorded by a field flow recorder. For example, a field flow recorder records and saves the daily production output and daily production time of a production well using a flow sensor. The daily nozzle size can be input by the user, sent by other devices, or retrieved from the production well's production system. For instance, when a production well is in production, the engineer manually inputs the daily nozzle size into the production well's production system. When detecting abnormal daily production data, the engineer retrieves the daily nozzle size from the production well's production system.
[0142] It should be noted that the method for determining abnormal daily production data provided in this embodiment of the invention is based on a two-stage method of initial screening using improved cluster analysis and manual review. Therefore, a sufficient amount of sample data is required to ensure the accuracy of the determination results. However, the production decline curve of a typical production well within one month of well opening is not significant, and there is generally no need to adjust the production system or measures. Therefore, when applying this method over time, M should be greater than 30 to ensure accurate results.
[0143] Furthermore, the daily production data can be in the form of a sequence or a set. The daily production data is stored in the form of a set (daily output, daily nozzle size, daily production time). For example, to obtain one daily production data, the daily production data is (50t / d, 5mm, 24h). The embodiments of the present invention do not specifically limit the storage format and data of the daily production data.
[0144] Step 202: Based on M daily production data, determine N production data sequences, where N is a positive integer greater than or equal to 1 and less than M.
[0145] It should be noted that the method for determining abnormal daily production data provided in this embodiment of the invention can directly perform initial screening of each day's production data, or it can first classify M days' production data to determine N production data sequences, so as to improve the accuracy of the initial screening.
[0146] Specifically, N production data sequences can be determined through the following steps 2021-2024.
[0147] Step 2021: Iterate through the M daily production data.
[0148] It should be noted that, since all M daily production data need to be divided, the M daily production data need to be traversed to determine whether the daily output, nozzle size, and production time of each daily production data meet the target conditions. Furthermore, although the larger M is, the longer the traversal time, this embodiment of the invention only determines whether the daily production data meets the target conditions during traversal, without performing any calculations. Therefore, the time consumption is relatively short, and the system resource consumption is low.
[0149] Step 2022: Divide the daily production data with a daily output or daily production time of 0 from the M daily production data into a non-production stage production data sequence.
[0150] It should be noted that during the non-production phase, the production wells are either shut down or under maintenance. In this case, the daily production data is characterized by a daily output or daily production time of 0. Therefore, based on this characteristic, the daily production data with a daily output or daily production time of 0 among the M determined daily production data are identified as the daily production data of the non-production phase, and all the daily production data of the non-production phase are divided into a non-production phase production data sequence.
[0151] Step 2023: Based on the nozzle size of the day, divide the daily production data that are not divided into non-production stage production data sequences from the M daily production data into N-1 production stage production data sequences, wherein the daily production data of each production stage production data sequence in the N-1 production stage production data sequences have the same nozzle size of the day.
[0152] It should be noted that the production decline curves differ depending on the nozzle size used in the production well. Therefore, to improve the accuracy of identifying abnormal daily production data, the daily production data with the same nozzle size among the M daily production data that have not yet been classified into non-production stage production data sequences can be grouped into one production stage production data sequence. This ensures that each production stage production data sequence corresponds to a production decline curve, thereby improving the accuracy of the determination results. For example, if there are still 1000 daily production data that have not been classified into non-production stage production data sequences, and step 2021 determines that there are 5 different nozzle sizes among these 1000 daily production data, then these 1000 daily production data are divided into 5 production stage production data sequences.
[0153] Step 2024: Based on the production data sequence of the non-production stage and the production data sequences of N-1 production stages, determine N production data sequences.
[0154] After dividing the production data sequence into non-production stage sequences and N-1 production stage sequences, all production data sequences are combined to form N production data sequences. For example, the dataset Z = {z1, z2, z3, ..., z...} M}, where z M This represents daily production data, where M is the number of daily production data points. Daily production data z M The data format is z M ={q M Ch M Pd M}, where q M Ch represents the daily output in the production data for the Mth day. M Let Pd be the nozzle size for the day in the production data for the Mth day. M This represents the daily production time in the Mth day's production data.
[0155] Iterate through the dataset Z to determine z in Z. mIf the daily output or daily production data is 0, then the next daily production data is determined; otherwise, the nozzle size for that day is determined. After traversal, a total of 5 nozzle sizes are applied to the daily production data where the daily output or daily production data is not 0. Therefore, the dataset Z is divided into 1 non-production stage production data sequence and 5 production stage production data sequences, that is, the dataset Z is divided into 6 production data sequences.
[0156] Step 203: Based on N production data sequences, use the Isolation Forest algorithm to determine the anomaly score corresponding to the daily production data of each of the N production data sequences. The Isolation Forest algorithm is an outlier detection method based on the degree of sorting.
[0157] It should be noted that, in the process of initially screening abnormal daily production data, the method for determining abnormal daily production data provided in this embodiment of the invention first uses the isolated forest algorithm to calculate the abnormal score corresponding to the daily production data of each of the N production data sequences, and then performs cluster analysis on the abnormal scores, simplifying the three-dimensional data into one-dimensional data for calculation, which further improves the calculation efficiency, saves system overhead, and shortens the calculation time.
[0158] It should be noted that the Isolation Forest algorithm is an outlier detection method based on the degree of computational sorting. The algorithm can be manually input, downloaded from a server, and processed locally, or daily production data can be uploaded to the server for server-side computation. Specifically, the basic steps of the Isolation Forest algorithm are as follows:
[0159] (1) Construct a binary tree to store the data from one production data sequence. Each day's production data is used as a sample and placed in the root node of the tree.
[0160] (2) Randomly specify a feature from the daily output, daily nozzle size and daily production time in the daily production data, and randomly generate a cutting point p, the value of p is between the maximum and minimum values corresponding to the feature.
[0161] (3) For a given feature, samples less than p are placed to the left of the current node, and samples greater than or equal to p are placed to the right of the current node.
[0162] (4) Repeat steps (2) to (3) to construct new child nodes until there is only one sample in the child node or the child node has reached the limit height.
[0163] (5) Repeat steps (1) to (4) to build 100 binary trees. For a sample, after traversing each binary tree, calculate the level of each tree where the sample finally falls, i.e., h(x), and then determine the average height of the sample in each tree, E(h(x)).
[0164] (6) Finally, calculate the anomaly score for each sample. The corresponding calculation formula is as follows:
[0165]
[0166] In the formula, Let n' be the number of sample daily production data in a production data sequence, n′ be the total number of daily production data in a production data sequence, and H be the harmonic series. for The average path length of an optimal binary tree is constructed from each sample. This is an abnormal score.
[0167] For example, for a production data sequence Z1 = {z1, z2, z3, ..., z...} 50} Through the above steps (1)-(6), the abnormal score corresponding to each daily production data in the production data sequence is determined as Z1={0.5,0.7,0.5,……,0.9}, thereby realizing the process of simplifying three-dimensional data into one-dimensional data.
[0168] Step 204: Based on the anomaly score corresponding to the daily production data of each of the N production data sequences, determine the initial abnormal daily production data through cluster analysis algorithm.
[0169] It should be noted that after determining the anomaly score corresponding to the daily production data of each production data sequence, the level of the anomaly score can preliminarily characterize the anomaly probability of the daily production data. Using a clustering analysis algorithm, the anomaly scores of the daily production data of a production data sequence can be clustered and divided into normal and anomaly categories. The daily production data in the normal category are normal daily production data, while the daily production data in the anomaly category are the preliminarily determined abnormal daily production data.
[0170] Furthermore, there are many clustering analysis algorithms, as long as they can cluster outlier scores, and this embodiment of the invention does not impose specific limitations on them. In one possible implementation, the k-medoids algorithm is used to perform more accurate clustering of outlier scores. The k-medoids algorithm can be manually input, obtained from a server and downloaded to the local machine, and can be calculated locally, or daily production data can be uploaded to the server and calculated on the server side.
[0171] Specifically, the initial abnormal day production data can be determined through the following steps (1)-(7):
[0172] (1) Arrange all the anomaly scores corresponding to each of the N production data sequences in descending order to obtain N anomaly score datasets.
[0173] (2) Let r = 1, and determine the abnormal scores in the rth abnormal score dataset that are located at the 95th percentile as the abnormal value representatives, and determine the abnormal scores that are located at the 5th percentile as the normal value representatives.
[0174] (3) Determine the outlier and normal values as the initial centers.
[0175] (4) Based on the initial center and the abnormal scores in the r-th abnormal score dataset excluding the outlier and normal values, the abnormal score cluster set of the r-th abnormal score dataset is determined by the k-medoids algorithm. The k-medoids algorithm is a clustering analysis algorithm based on the center distance algorithm.
[0176] (5) Determine whether r is equal to N.
[0177] (6) If r is not equal to N, let r = r + 1 and return the steps of identifying the abnormal scores in the r-th abnormal score dataset that are at the 95th percentile as the abnormal value representative and the abnormal scores that are at the 5th percentile as the normal value representative.
[0178] (7) If r equals N, then the daily production data corresponding to the abnormal scores in the N clusters of abnormal scores will be determined as the initial abnormal daily production data.
[0179] It should be noted that the conventional operation of the k-medoids algorithm does not specify initial center representatives. However, the embodiments of this invention are improved algorithms based on the k-medoids algorithm, which use the isolated forest algorithm to transform three-dimensional data into one-dimensional data, and then specify the 95th and 5th percentiles of the outlier scores as initial center representatives. Furthermore, this algorithm is based on the characteristics of daily production data, which can achieve rapid initial screening of daily production data with high accuracy, fast processing speed, and short processing time. The specific process of the k-medoids algorithm is a well-known technology, and the embodiments of this invention will not elaborate on it.
[0180] Step 205: Based on the initial abnormal daily production data, M daily production data, and preset domain knowledge conditions, determine the abnormal daily production data of the target production well. The preset domain knowledge conditions refer to the rule conditions applicable to the screening of abnormal daily production data based on oil and gas extraction theory and practice.
[0181] It should be noted that after determining the initial abnormal day production data through step 204, there is a certain probability that the initial abnormal day production data will be misjudged. Therefore, the initial abnormal day production data can be reviewed based on preset domain knowledge conditions. If it meets the preset domain knowledge conditions, the initial abnormal day production data is marked as normal day production data; if it does not meet the preset domain knowledge conditions, the initial abnormal day production data is determined as daily production data. Specifically, the abnormal day production data of the target production well can be determined through the following steps:
[0182] (1) Based on the initial abnormal day production data and M days of production data, determine the control day production data that is closest to the initial abnormal day production data in the upstream of the time axis.
[0183] It should be noted that the production data of the control day closest upstream of the initial abnormal day's production data on the timeline refers to the daily production data collected corresponding to the production time closest to the initial abnormal day's production data in the upstream direction on the timeline. For example, if the production time corresponding to the initial abnormal day's production data is December 3, 2022, then the closest control day's production data upstream on the timeline is the daily production data collected on December 2, 2022.
[0184] It should also be noted that the reference daily production data is not the production data of the initial abnormal day. If the daily production data collected corresponding to the closest production time upstream of the initial abnormal day's production data timeline is also the production data of the initial abnormal day, then that production time is skipped, and the daily production data collected at the second closest time is determined as the reference daily production data. For example, if the production time corresponding to the initial abnormal day's production data is December 3, 2022, and the daily production data collected on December 2, 2022 is also determined to be the production data of the initial abnormal day, then the reference daily production data for both is the daily production data collected on December 1, 2022.
[0185] (2) Based on the production data of the initial abnormal day and the production data of the control day, the flow rate of the initial abnormal day production data and the flow rate of the control day production data are determined by the following formula.
[0186] v=4q / (πCh 2 Pd)
[0187] In the formula: v is the flow rate, q is the daily output, Pd is the daily production time, and Ch is the nozzle size for the day.
[0188] It should be noted that since the drastic changes in flow rate are related to nozzle adjustment and well opening / closing, the flow rate is calculated using the above formula to comprehensively consider the effects of nozzle size and production time, so as to further improve the screening accuracy based on the new parameters.
[0189] (3) The production data of the initial abnormal day and the flow rate of the initial abnormal day production data are determined as the production data of the intermediate abnormal day.
[0190] It should be noted that by adding a flow velocity parameter to the initial three-parameter relationship of abnormal day production data, the intermediate abnormal day production data becomes a four-parameter relationship, thus fully integrating oil and gas extraction theory with production well production practice. The format of intermediate abnormal day production data is (daily output, daily nozzle size, daily production time, flow velocity), for example, intermediate abnormal day production data: (50t / d, 5mm, 24h, 16m / s).
[0191] (4) The flow rate of the control day production data and the control day production data is determined as the intermediate control day production data.
[0192] It should be noted that, based on the three-parameter relationship of the reference daily production data, a flow rate parameter is added, making the reference daily production data a four-parameter relationship. The format of the reference daily production data is (daily output, daily nozzle size, daily production time, flow rate), for example, reference daily production data: (45t / d, 5mm, 23h, 17m / s).
[0193] (5) Based on intermediate abnormal production data, intermediate control daily production data and preset domain knowledge conditions, determine the abnormal daily production data of the target production well.
[0194] It should be noted that the following principles are universal and cannot be violated in the theory and practice of oil and gas extraction:
[0195] (1) If the production time is 0, then the output is 0;
[0196] (2) Output fluctuates throughout the entire production decline cycle;
[0197] (3) If production suddenly increases, the production system may have changed, such as the nozzle becoming larger, the nozzle being removed, the well being opened after being shut down for a period of time, or the production time being extended.
[0198] (4) If production suddenly drops, the production system may have changed, such as the nozzle becoming smaller or the production time being shortened.
[0199] (5) In addition to changes in the production system, factors such as water accumulation can also cause fluctuations in output.
[0200] Therefore, the relationship between intermediate abnormal production data and intermediate control daily production data also needs to adhere to the above practical experience. Based on this, the relationship between intermediate abnormal production data and intermediate control daily production data can be judged using relevant patterns from practical experience. If the relationship between the two conforms to practical experience, it indicates that the intermediate abnormal production data is normal data; if the relationship does not conform to practical experience, it indicates that the intermediate abnormal production data is abnormal data. In this case, the daily production data corresponding to the intermediate abnormal production data can be determined as abnormal production data. Specifically, intermediate abnormal production data can be marked as normal daily production data using the following preset domain knowledge conditions.
[0201] (1) If both the daily production time and daily output in the intermediate abnormal production data are 0, then mark the intermediate abnormal production data as normal daily production data; or,
[0202] (2) If the ratio of the nozzle size of the intermediate control day production data to the nozzle size of the intermediate abnormal production data is not equal to 1, then the intermediate abnormal production data is marked as normal day production data; or,
[0203] (3) If the flow rate in the intermediate abnormal production data is greater than 0.5 times the flow rate of the intermediate control day production data, but less than 2 times the flow rate of the intermediate control day production data, then the intermediate abnormal production data will be marked as normal day production data.
[0204] If intermediate abnormal production data is not marked as normal daily production data, then the daily production data corresponding to the intermediate abnormal production data will be identified as the abnormal daily production data of the target production well.
[0205] In this embodiment of the invention, M daily production data points of a target production well can be obtained. The target production well refers to the production well whose abnormal daily production data is to be studied. The daily production data includes daily output, nozzle size, and daily production time, where M is a positive integer greater than or equal to 30. Based on the M daily production data points, initial abnormal daily production data are determined using a combined isolated forest and cluster analysis method. Based on the initial abnormal daily production data, the M daily production data points, and preset domain knowledge conditions, the abnormal daily production data of the target production well are determined. The preset domain knowledge conditions refer to rules and conditions applicable to screening production anomalies, formulated based on oil and gas extraction theory and practice. In other words, the basic idea of this invention is to first perform a preliminary screening of the daily production data using a combined isolated forest and cluster analysis method when detecting abnormal daily production data. Then, the daily production data determined as initial abnormal daily production data after the preliminary screening are verified according to the preset domain knowledge conditions. This two-stage screening simplifies the process in related technologies that requires checking all daily production data one by one based on the production history and production decline pattern of the production well, significantly shortening the detection time. Furthermore, thanks to the improved clustering analysis algorithm, the detection accuracy has been enhanced, avoiding missed detections.
[0206] All the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of the present invention, and the embodiments of the present invention will not be described in detail one by one.
[0207] Figure 3 This is a schematic diagram of a device for determining abnormal daily production data according to an embodiment of the present invention. See also... Figure 3 The device may include:
[0208] The acquisition module 301 is used to acquire M daily production data of the target production well. The target production well refers to the production well whose abnormal daily production data is to be studied. The daily production data includes daily output, daily nozzle size and daily production time. M is a positive integer greater than or equal to 30.
[0209] The first determination module 302 is used to determine the initial abnormal daily production data based on M daily production data through a combined method of isolated forest and cluster analysis.
[0210] The second determination module 303 is used to determine the abnormal daily production data of the target production well based on the initial abnormal daily production data, M daily production data and preset domain knowledge conditions. The preset domain knowledge conditions refer to the rule conditions applicable to the screening of abnormal daily production data based on oil and gas extraction theory and practice.
[0211] Optionally, the first determining module includes:
[0212] The first determining unit is used to determine N production data sequences based on M daily production data, where N is a positive integer greater than or equal to 1 and less than M;
[0213] The second determining unit is used to determine the anomaly score corresponding to the daily production data of each of the N production data sequences based on the N production data sequences and using the Isolation Forest algorithm. The Isolation Forest algorithm is an outlier detection method based on the degree of sorting.
[0214] The third determining unit is used to determine the initial abnormal daily production data based on the anomaly score corresponding to the daily production data of each of the N production data sequences, through a cluster analysis algorithm.
[0215] Optionally, the first determining unit specifically includes:
[0216] Iterate through the production data of M days;
[0217] Divide the daily production data with a daily output or daily production time of 0 from the M daily production data into a non-production stage production data sequence;
[0218] Based on the nozzle size of the day, the daily production data that are not divided into non-production stage production data sequences in the M daily production data are divided into N-1 production stage production data sequences. Among them, the daily production data of each production stage production data sequence in the N-1 production stage production data sequences have the same nozzle size of the day.
[0219] Based on the production data sequence of the non-production stage and the production data sequence of N-1 production stages, N production data sequences are determined.
[0220] Optionally, the third determining unit specifically includes:
[0221] Arrange all the anomaly scores corresponding to each of the N production data sequences in descending order to obtain N anomaly score datasets.
[0222] Let r = 1, and determine the abnormal scores in the r-th abnormal score dataset that are located at the 95th percentile as the outlier representatives, and determine the abnormal scores that are located at the 5th percentile as the normal values representatives;
[0223] The outlier and normal values were used as the initial centers.
[0224] Based on the initial centers and the abnormal scores in the r-th abnormal score dataset excluding the outlier and normal value representatives, the k-medoids algorithm is used to determine the abnormal score cluster set of the r-th abnormal score dataset. The k-medoids algorithm is a clustering analysis algorithm based on the centroid distance algorithm.
[0225] Determine if r equals N;
[0226] If r is not equal to N, let r = r + 1, and return the steps of identifying the outlier scores in the r-th outlier score dataset that are at the 95th percentile as outlier representatives and the outlier scores that are at the 5th percentile as normal representatives.
[0227] If r equals N, then the daily production data corresponding to the abnormal scores in the N clusters of abnormal scores will be determined as the initial abnormal daily production data.
[0228] Optionally, the second determining module includes:
[0229] The fourth determination unit is used to determine the control day production data that is closest to the initial abnormal day production data in the upstream of the time axis based on the initial abnormal day production data and M days of production data.
[0230] The fifth determining unit is used to determine the flow rate of the initial abnormal day's production data and the flow rate of the control day's production data based on the initial abnormal day's production data and the control day's production data using the following formula:
[0231] v=4q / (πCh 2 Pd)
[0232] In the formula: v is the flow rate, q is the daily output, Pd is the daily production time, and Ch is the nozzle size for the day;
[0233] The sixth determining unit is used to determine the production data of the initial abnormal day and the flow rate of the initial abnormal day as the production data of the intermediate abnormal day.
[0234] The seventh determining unit is used to determine the production data of the control day and the flow rate of the control day production data as the intermediate control day production data;
[0235] The eighth determination unit is used to determine the abnormal daily production data of the target production well based on intermediate abnormal production data, intermediate control daily production data, and preset domain knowledge conditions.
[0236] Optionally, the eighth determining unit specifically includes:
[0237] Based on intermediate abnormal production data and intermediate control daily production data, and using the following pre-defined domain knowledge conditions, intermediate abnormal production data is marked as normal daily production data.
[0238] If both the daily production time and daily output in the intermediate abnormal production data are 0, then the intermediate abnormal production data will be marked as normal daily production data; or,
[0239] If the ratio of the nozzle size of the intermediate comparison day's production data to the nozzle size of the intermediate abnormal production data is not equal to 1, then the intermediate abnormal production data will be marked as normal day's production data; or,
[0240] If the flow rate in the intermediate abnormal production data is greater than 0.5 times the flow rate of the intermediate control day production data, but less than 2 times the flow rate of the intermediate control day production data, then the intermediate abnormal production data will be marked as normal day production data.
[0241] If intermediate abnormal production data is not marked as normal daily production data, then the daily production data corresponding to the intermediate abnormal production data will be identified as the abnormal daily production data of the target production well.
[0242] In this embodiment of the invention, M daily production data points of a target production well can be obtained. The target production well refers to the production well whose abnormal daily production data is to be studied. The daily production data includes daily output, nozzle size, and daily production time, where M is a positive integer greater than or equal to 30. Based on the M daily production data points, initial abnormal daily production data are determined using a combined isolated forest and cluster analysis method. Based on the initial abnormal daily production data, the M daily production data points, and preset domain knowledge conditions, the abnormal daily production data of the target production well are determined. The preset domain knowledge conditions refer to rules and conditions applicable to screening production anomalies, formulated based on oil and gas extraction theory and practice. In other words, the basic idea of this invention is to first perform a preliminary screening of the daily production data using a combined isolated forest and cluster analysis method when detecting abnormal daily production data. Then, the daily production data determined as initial abnormal daily production data after the preliminary screening are verified according to the preset domain knowledge conditions. This two-stage screening simplifies the process in related technologies that requires checking all daily production data one by one based on the production history and production decline pattern of the production well, significantly shortening the detection time. Furthermore, thanks to the improved clustering analysis algorithm, the detection accuracy has been enhanced, avoiding missed detections.
[0243] It should be noted that the above-described embodiment of the abnormal day production data determination device is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the embodiments of the abnormal day production data determination device and the abnormal day production data determination method provided above belong to the same concept, and their specific implementation process is detailed in the method embodiment, which will not be repeated here.
[0244] Figure 4 This is a schematic diagram of the structure of a terminal 400 provided in an embodiment of the present invention. The terminal 400 can be: a smartphone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, or a desktop computer. The terminal 400 may also be referred to as a user device, a portable terminal, a laptop terminal, a desktop terminal, or other names.
[0245] Typically, terminal 400 includes a processor 401 and a memory 402.
[0246] Processor 401 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 401 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 401 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 401 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 401 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0247] Memory 402 may include one or more computer-readable storage media, which may be non-transitory. Memory 402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in memory 402 is used to store at least one instruction, which is executed by processor 401 to implement the abnormal daily production data determination method provided in the method embodiments of this application.
[0248] In some embodiments, the terminal 400 may also optionally include a peripheral device interface 403 and at least one peripheral device. The processor 401, memory 402, and peripheral device interface 403 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 403 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 404, a display screen 405, a camera 406, an audio circuit 407, a positioning component 408, and a power supply 409.
[0249] Peripheral device interface 403 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 401 and memory 402. In some embodiments, processor 401, memory 402 and peripheral device interface 403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 401, memory 402 and peripheral device interface 403 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0250] The radio frequency (RF) circuit 404 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 404 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 404 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 404 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 404 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 404 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0251] Display screen 405 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 405 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 401 for processing. In this case, display screen 405 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 405, which serves as the front panel of terminal 400; in other embodiments, there may be at least two display screens 405, respectively disposed on different surfaces of terminal 400 or in a folded design; in still other embodiments, display screen 405 may be a flexible display screen, disposed on a curved or folded surface of terminal 400. Furthermore, display screen 405 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 405 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).
[0252] The camera assembly 406 is used to acquire images or videos. Optionally, the camera assembly 406 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 406 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm light flash and a cool light flash, which can be used for light compensation at different color temperatures.
[0253] The audio circuit 407 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting them into electrical signals that are input to the processor 401 for processing, or to the radio frequency circuit 404 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal 400. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 401 or the radio frequency circuit 404 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 407 may also include a headphone jack.
[0254] The positioning component 408 is used to locate the current geographical location of the terminal 400 in order to enable navigation or LBS (Location Based Service). The positioning component 408 can be a positioning component based on the US GPS (Global Positioning System), China's BeiDou system, Russia's Granas system, or the EU's Galileo system.
[0255] Power supply 409 is used to power the various components in terminal 400. Power supply 409 can be AC power, DC power, a disposable battery, or a rechargeable battery. When power supply 409 includes a rechargeable battery, the rechargeable battery can support wired charging or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0256] In some embodiments, the terminal 400 further includes one or more sensors 410. The one or more sensors 410 include, but are not limited to: an accelerometer 411, a gyroscope 412, a pressure sensor 413, a fingerprint sensor 414, an optical sensor 415, and a proximity sensor 416.
[0257] Accelerometer 411 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by terminal 400. For example, accelerometer 411 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 401 can control display screen 405 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 411. Accelerometer 411 can also be used for games or for acquiring user motion data.
[0258] The gyroscope sensor 412 can detect the orientation and rotation angle of the terminal 400. The gyroscope sensor 412, in conjunction with the accelerometer sensor 411, can collect 3D motion data from the user on the terminal 400. Based on the data collected by the gyroscope sensor 412, the processor 401 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0259] The pressure sensor 413 can be disposed on the side bezel of the terminal 400 and / or on the lower layer of the display screen 405. When the pressure sensor 413 is disposed on the side bezel of the terminal 400, it can detect the user's grip signal on the terminal 400, and the processor 401 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 413. When the pressure sensor 413 is disposed on the lower layer of the display screen 405, the processor 401 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 405. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0260] The fingerprint sensor 414 is used to collect the user's fingerprint. The processor 401 identifies the user's identity based on the fingerprint collected by the fingerprint sensor 414, or the fingerprint sensor 414 identifies the user's identity based on the collected fingerprint. When the user's identity is identified as trusted, the processor 401 authorizes the user to perform relevant sensitive operations, including unlocking the screen, viewing encrypted information, downloading software, making payments, and changing settings. The fingerprint sensor 414 can be located on the front, back, or side of the terminal 400. When the terminal 400 has physical buttons or a manufacturer's logo, the fingerprint sensor 414 can be integrated with the physical buttons or manufacturer's logo.
[0261] An optical sensor 415 is used to collect ambient light intensity. In one embodiment, the processor 401 can control the display brightness of the display screen 405 based on the ambient light intensity collected by the optical sensor 415. Specifically, when the ambient light intensity is high, the display brightness of the display screen 405 is increased; when the ambient light intensity is low, the display brightness of the display screen 405 is decreased. In another embodiment, the processor 401 can also dynamically adjust the shooting parameters of the camera assembly 406 based on the ambient light intensity collected by the optical sensor 415.
[0262] The proximity sensor 416, also known as a distance sensor, is typically located on the front panel of the terminal 400. The proximity sensor 416 is used to detect the distance between the user and the front of the terminal 400. In one embodiment, when the proximity sensor 416 detects that the distance between the user and the front of the terminal 400 is gradually decreasing, the processor 401 controls the display screen 405 to switch from a screen-on state to a screen-off state; when the proximity sensor 416 detects that the distance between the user and the front of the terminal 400 is gradually increasing, the processor 401 controls the display screen 405 to switch from a screen-off state to a screen-on state.
[0263] That is, embodiments of the present invention not only provide a terminal, including a processor and a memory for storing processor-executable instructions, wherein the processor is configured to execute... Figure 1 or Figure 2 Furthermore, embodiments of the present invention also provide a computer-readable storage medium storing a computer program that, when executed by a processor, can implement... Figure 1 or Figure 2 The method for determining abnormal daily production data in the illustrated embodiment.
[0264] Those skilled in the art will understand that Figure 4 The structure shown does not constitute a limitation on terminal 400 and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0265] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.
[0266] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. An abnormal daily production data determination method characterized by comprising: The abnormal daily production data determination method comprises the following steps: M daily production data of a target production well are acquired, the target production well refers to a production well whose abnormal daily production data are to be researched, the daily production data comprise daily yield, daily choke size and daily production time, and M is a positive integer greater than or equal to 30; N production data sequences are determined based on the M daily production data, and N is a positive integer greater than or equal to 1 and less than M; An abnormal score corresponding to daily production data of each production data sequence in the N production data sequences is determined by using an isolation forest algorithm based on the N production data sequences, the isolation forest algorithm is an abnormal value detection method based on calculation of combing degree; All abnormal scores corresponding to each production data sequence in the N production data sequences are arranged in ascending order to obtain N abnormal score data sets; An abnormal score located at a 95% quantile in an rth abnormal score data set is determined as an abnormal value representative, and an abnormal score located at a 5% quantile is determined as a normal value representative, wherein r = 1; The abnormal value representative and the normal value representative are determined as initial centers; An abnormal score clustering set of the rth abnormal score data set is determined by using a k-medoids algorithm based on the initial centers and abnormal scores in the rth abnormal score data set except the abnormal value representative and the normal value representative, the k-medoids algorithm is a clustering analysis algorithm based on a center point distance algorithm; It is determined whether r is equal to N; If r is not equal to N, r = r + 1, and the step of determining the abnormal value representative and the normal value representative is returned; If r is equal to N, daily production data corresponding to abnormal scores in the N abnormal score clustering sets are determined as initial abnormal daily production data; Abnormal daily production data of the target production well are determined based on the initial abnormal daily production data, the M daily production data and a preset field knowledge condition, the preset field knowledge condition refers to a rule condition suitable for screening of abnormal daily production data based on oil and gas exploitation theory and practice.
2. The abnormal day production data determination method of claim 1, wherein, The determination of the N production data sequences based on the M daily production data comprises the following steps: The M daily production data are traversed; Daily production data in which daily yield or daily production time is 0 in the M daily production data are divided into a non-production stage production data sequence; Daily production data in the M daily production data which are not divided into the non-production stage production data sequence are divided into N-1 production stage production data sequences based on daily choke size, wherein daily production data in each production stage production data sequence in the N-1 production stage production data sequences have the same daily choke size; The N production data sequences are determined based on the non-production stage production data sequence and the N-1 production stage production data sequences.
3. The abnormal day production data determination method of claim 1, wherein, The initial abnormal daily production data is determined based on the M daily production data and preset domain knowledge conditions. The control daily production data closest to the initial abnormal daily production data in a time axis is determined based on the initial abnormal daily production data and the M daily production data. The flow rate of the initial abnormal daily production data is determined based on the initial abnormal daily production data and the control daily production data by the following formula, v =4 q / (π Ch 2 Pd ) wherein: said v is flow rate, said q is daily production, said Pd is daily production time, said Ch is the day's choke size; The initial abnormal daily production data and the flow rate of the initial abnormal daily production data are determined as intermediate abnormal daily production data. The control daily production data and the flow rate of the control daily production data are determined as intermediate control daily production data. The abnormal daily production data of the target production well is determined based on the intermediate abnormal daily production data, the intermediate control daily production data and the preset domain knowledge conditions.
4. The abnormal day production data determination method of claim 3, wherein, The abnormal daily production data of the target production well is determined based on the intermediate abnormal daily production data, the intermediate control daily production data and the preset domain knowledge conditions. The intermediate abnormal daily production data is marked as normal daily production data based on the intermediate abnormal daily production data, the intermediate control daily production data and the following preset domain knowledge conditions, if the daily production time and the daily production volume in the intermediate abnormal daily production data are both 0, the intermediate abnormal daily production data is marked as normal daily production data; or, if the ratio of the daily choke size of the intermediate control daily production data to the daily choke size of the intermediate abnormal daily production data is not equal to 1, the intermediate abnormal daily production data is marked as normal daily production data; or, if the flow rate of the intermediate abnormal daily production data is greater than 0.5 times the flow rate of the intermediate control daily production data and less than 2 times the flow rate of the intermediate control daily production data, the intermediate abnormal daily production data is marked as normal daily production data; if the intermediate abnormal daily production data is not marked as normal daily production data, the daily production data corresponding to the intermediate abnormal daily production data is determined as the abnormal daily production data of the target production well.
5. An abnormal day production data determination device characterized by comprising: The abnormal daily production data determination device comprises: An acquisition module is configured to acquire M daily production data of a target production well, wherein the target production well refers to a production well for which abnormal daily production data is to be researched, the daily production data comprises daily production volume, daily choke size and daily production time, and M is a positive integer greater than or equal to 30. A first determination module comprises, a first determination unit configured to determine N production data sequences based on the M daily production data, wherein N is a positive integer greater than or equal to 1 and less than M; a second determination unit configured to determine an abnormal score corresponding to daily production data in each production data sequence in the N production data sequences by an Isolation Forest algorithm based on the N production data sequences, wherein the Isolation Forest algorithm is an outlier detection method based on calculation of combing degree. The third determining unit is configured to arrange all abnormal scores corresponding to each production data sequence in the N production data sequences in ascending order to obtain N abnormal score data sets, set r=1, determine an abnormal score located at a 95% quantile in an rth abnormal score data set as an abnormal value representative, determine an abnormal score located at a 5% quantile as a normal value representative, determine the abnormal value representative and the normal value representative as initial centers, determine an abnormal score clustering set of the rth abnormal score data set based on the initial centers and the abnormal scores in the rth abnormal score data set except the abnormal value representative and the normal value representative by using a k-medoids algorithm, and determine initial abnormal daily production data of the N production data sequences based on the abnormal scores in the N abnormal score clustering sets. The second determining module is configured to determine abnormal daily production data of the target production well based on the initial abnormal daily production data, the M daily production data and a preset domain knowledge condition, wherein the preset domain knowledge condition refers to a rule condition suitable for abnormal daily production data screening and established based on oil and gas exploitation theory and practice.
6. An abnormal day production data determination device characterized by comprising: The apparatus includes: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to perform the method of any one of claims 1-4.
7. A computer readable storage medium characterized in that, The storage medium has stored therein a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1-4.
Citation Information
Patent Citations
Transaction data anomaly detection method and device, program product and storage medium
CN119128766A
Method and system for anomaly detection
US20230252568A1
A method for detection of anomalies in an image and anomaly detection system
WO2025180648A1