Natural gas abnormal sales volume data detection method and device, electronic equipment and medium
Through the method of non-coin sliding calculation and window subsequence distance calculation, abnormal data and time periods in natural gas sales data are identified, which solves the problem of inaccurate abnormal identification in the prior art and improves the accuracy of gas load prediction.
Patent Information
- Application Number
- CN202311783033.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-06-24
AI Technical Summary
There are abnormal data points in the existing natural gas sales data, which affect the prediction accuracy of gas loads, and the existing abnormal identification methods are inaccurate in long-term data identification.
By obtaining the natural gas sales time sequence data of the target user, the difference value sequence is obtained by using no overlapping sliding calculation, and then gradually sliding with the preset time window size to obtain each window subsequence, calculate the distance between each window subsequence and all subsequences, and generate a maximum value sequence to identify abnormal data and time periods.
It improves the accuracy of identifying abnormal data in natural gas sales timing data, and reduces the probability of identifying local abnormal but globally normal data points as abnormal.
Smart Images

Figure CN120198145A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of data analysis technology and new energy technology, and in particular, to a method, device, electronic device and medium for detecting abnormal natural gas sales data. Background Art
[0002] With the increasingly wide use of natural gas, ensuring stable gas supply has become the demand of consumers. For gas suppliers, accurately predicting the gas consumption load of natural gas is also beneficial to ensuring gas supply demand, and at the same time is of great significance for ensuring the safe operation of the natural gas pipeline network and optimizing the dispatching of the pipeline network. Generally speaking, the user load can be predicted based on the sales data of natural gas.
[0003] In the process of implementing the concept of the present disclosure, the inventors found that there are at least the following technical problems in the related art: there may be abnormal data points in the natural gas sales data, and these abnormal data points have an impact on the prediction accuracy of the gas consumption load. However, the existing abnormal identification methods often only consider local differences and are inaccurate in identifying long-term data. Summary of the Invention
[0004] In order to solve the above technical problems or at least partially solve the above technical problems, embodiments of the present disclosure provide a method, device, electronic device and medium for detecting abnormal natural gas sales data.
[0005] In a first aspect, embodiments of the present disclosure provide a method for detecting abnormal natural gas sales data. The detection method includes:
[0006] Obtain the time series data of the natural gas sales volume of the target user within the statistical period;
[0007] Perform non-overlapping sliding calculation on the time series data of the natural gas sales volume according to a preset window size to obtain a difference value sequence composed of the differences between the data means in adjacent windows, where the non-overlapping sliding means that there is no overlap between the data covered by adjacent windows;
[0008] For the difference value sequence, obtain each window subsequence by sliding step by step according to a preset window size, and calculate the distances between each window subsequence and all subsequences respectively to obtain a subsequence distance sequence;
[0009] Find the maximum value of the distances between each window subsequence and all subsequences in the obtained subsequence distance sequence to generate a maximum value sequence;
[0010] Detect the data distribution state of the maximum value sequence to obtain abnormal data and abnormal periods.
[0011] Further, for the difference value sequence, each window subsequence is obtained by gradually sliding according to a single data with a preset time window size, and the distances between each window subsequence and all subsequences are calculated respectively to obtain a subsequence distance sequence, including:
[0012] For the difference value sequence, the first window subsequence is first determined;
[0013] According to the determined first window subsequence, each window subsequence is obtained by gradually sliding according to a single data with a preset time window size;
[0014] The distances between the first window subsequence and all subsequences are calculated respectively to obtain the first window subsequence distance sequence;
[0015] Then, the distances between the remaining window subsequences and all subsequences are calculated respectively to obtain the subsequence distance sequences of each window subsequence.
[0016] Further, the data distribution state of the maximum value sequence is detected to obtain abnormal data and abnormal time periods, including:
[0017] The data in the maximum value sequence are sorted according to the value size;
[0018] According to the increasing trend and relative size of the sorted data, the data with values exceeding the set threshold are determined as abnormal evaluation data;
[0019] Determine the window subsequences corresponding to the above abnormal evaluation data;
[0020] According to the obtained window subsequences, find the natural gas sales time series data that generates the window subsequences, and analyze whether there is abnormal data or abnormal time periods in the natural gas data corresponding to the natural gas sales time series data.
[0021] Further, the expression of the set threshold is as follows:
[0022] Set threshold = Q3 + 1.5 * (Q3 - Q1),
[0023] where Q1 represents the value of the lower quartile in the sorted data; Q3 represents the value of the upper quartile in the sorted data.
[0024] Further, the distances between each window subsequence and all subsequences are calculated respectively, including: calculating the Euclidean distances between each window subsequence and all subsequences.
[0025] Further, according to the preset time window size, non-overlapping sliding calculations are performed on the natural gas sales time series data, and the difference value sequence formed by the differences in the data means within adjacent time windows includes:
[0026] For the time series data of natural gas sales volume, data is selected without overlap by sliding according to the preset window size to obtain two sets of data within adjacent windows;
[0027] For the two sets of data within adjacent windows, the mean values are calculated respectively and then the difference operation is performed to obtain the difference value;
[0028] The difference values obtained from all adjacent windows form a difference value sequence.
[0029] In a second aspect, an embodiment of the present disclosure provides a detection device for abnormal natural gas sales volume data. The device includes:
[0030] A data acquisition module, configured to acquire the time series data of the natural gas sales volume of a target user within a statistical period;
[0031] A first calculation module, configured to perform non-overlapping sliding calculation on the time series data of the natural gas sales volume according to the preset window size to obtain a difference value sequence composed of the differences between the data means within adjacent windows; the non-overlapping sliding means that there is no overlap between the data covered by adjacent windows;
[0032] A second calculation module, configured to, for the difference value sequence, slide step by step according to a single data with the preset window size to obtain each window subsequence, and calculate the distances between each window subsequence and all subsequences respectively to obtain a subsequence distance sequence;
[0033] An abnormal evaluation sequence generation module, configured to find the maximum value of the distances between each window subsequence and all subsequences in the obtained subsequence distance sequence to generate a maximum value sequence;
[0034] An abnormal detection module, configured to detect the data distribution state of the maximum value sequence to obtain abnormal data and abnormal time periods.
[0035] The above technical solutions provided by the embodiments of the present disclosure have at least some or all of the following advantages:
[0036] By obtaining the time series data of the natural gas sales volume of the target user within the statistical duration; performing non-overlapping sliding calculation on the above-mentioned time series data of the natural gas sales volume according to the preset window size to obtain a difference value sequence composed of the differences of the data means within adjacent windows; the above-mentioned difference value sequence can reflect the distribution change trend of the window data over time; for the said difference value sequence, each window subsequence is gradually slid step by step according to the preset window size to calculate the distances between each window subsequence and all subsequences, obtaining a subsequence distance sequence; finding the maximum value of the distances between each window subsequence and all subsequences in the obtained subsequence distance sequence to generate a maximum value sequence; the maximum value sequence can reflect, from a global perspective, the situations with relatively large relative change differences in the distribution change trend of the window data, and based on this, abnormal data and abnormal periods can be identified. The present invention can reduce the probability of misidentifying data points that are locally abnormal but globally normal as abnormal for the time series data of natural gas sales volume, and improve the identification accuracy and applicability of abnormal data in the time series data of natural gas sales volume. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure;
[0038] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or related technologies. Obviously, for those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts; Figure 1 Schematically shows a flowchart of a method for detecting abnormal natural gas sales volume data provided by an embodiment of the present disclosure;
[0039] Figure 2 Schematically shows (a) a schematic diagram of time series data of natural gas sales volume and (b) a schematic diagram of a difference value sequence composed of the differences of the data means within adjacent windows obtained by performing non-overlapping sliding calculation on the above-mentioned time series data of natural gas sales volume according to the preset window size in an embodiment of the present disclosure;
[0040] Figure 3 Schematically shows (a) a schematic diagram of determining the distance between the first window subsequence selected in the above-mentioned difference value sequence and all window subsequences as the degree of difference; (b) a schematic diagram of calculating the distance between the first window subsequence selected and the penultimate window subsequence using the Euclidean distance; (c) a schematic diagram of determining the distance between the second window subsequence selected in the above-mentioned difference value sequence and each window subsequence as the degree of difference; (d) a schematic diagram of representing the distance vectors corresponding to each window subsequence using a matrix structure (matrix profile);
[0041] Figure 4 Schematically shows the generated maximum value sequence, i.e., the anomaly evaluation sequence, of the embodiments of the present disclosure;
[0042] Figure 5 Schematically shows the specific flowchart of step S150 in the embodiments of the present disclosure;
[0043] Figure 6 Schematically shows the structural block diagram of the detection device for abnormal gas sales data according to the embodiments of the present disclosure;
[0044] Figure 7 Schematically shows the structural block diagram of the electronic device provided by the embodiments of the present disclosure. Detailed implementation manners
[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0046] The first exemplary embodiment of the present disclosure provides a method for detecting abnormal gas sales data.
[0047] During the process of users using natural gas, after uncontrollable factors or sudden events occur, the gas consumption of users will be impacted, which is reflected as gas consumption fluctuations in the natural gas sales data. To reduce the interference of the above gas consumption fluctuations on the subsequent prediction of gas consumption, a processing logic is to first identify the abnormal data in the natural gas sales data, analyze and process it, and then perform the subsequent prediction of gas consumption.
[0048] Figure 1 Schematically shows the flowchart of the method for detecting abnormal gas sales data provided by an embodiment of the present disclosure.
[0049] Refer to Figure 1 As shown, the method for detecting abnormal gas sales data provided by the embodiments of the present disclosure includes the following steps: S110 to S150.
[0050] Step S110, obtain the time series data of the natural gas sales volume of the target user within the statistical period.
[0051] The above target user can be various users faced by the gas provider. For example, it can be a user of a collective organization type using natural gas, or a residential user using natural gas, etc.
[0052] The above statistical duration is the time period corresponding to the time series data of natural gas sales volume. The time granularity within the above statistical duration can be counted in terms of weeks (which can be one week or multiple weeks), days (which can be one day or multiple days), hours (which can be one hour or multiple hours), etc. For example, the statistical duration can be counted in terms of consecutive years (such as 3 - 5 years), consecutive quarters (such as several quarters), or consecutive months, etc.
[0053] In some embodiments, the above time series data of natural gas sales volume can be a long - period sequence. Specifically, a sequence in which the ratio of the time period to the time granularity is greater than a set ratio value is called a long - period sequence.
[0054] Figure 2 Schematically shows (a) a schematic diagram of the time series data of natural gas sales volume in an embodiment of the present disclosure. Referring to Figure 2 as shown in (a) above, in the above time series data of natural gas sales volume, each sales volume value corresponds to a time. Arrange each sales volume value in chronological order. For the sake of simplification in the operation process, a corresponding relationship can be established between the time corresponding to each sales volume value and the subscript of the sales volume value. Subsequently, only each sales volume value can be used for simplified representation. For example, the data sequence corresponding to the time series data of natural gas sales volume can be expressed as [(t1, a1), (t2, a2), …, (t n , a n )], where n represents the total number of data in the time series data of natural gas sales volume. The abscissa t1~t n represents time, and the ordinate a1~a n represents the specific natural gas sales volume value. By establishing a corresponding relationship between the subscripts 1~n of the natural gas sales volume values and the corresponding times t1~t n , the time series data of natural gas sales volume can be simplified and represented in the following sequence form: [a1, a2, …, a n . For example, t1 is January 3, 2016, t2 is February 21, 2016, etc.
[0055] Step S120: Perform non - overlapping sliding calculation on the above time series data of natural gas sales volume according to a preset window size to obtain a difference value sequence composed of the differences between the data means within adjacent windows.
[0056] Figure 2 Schematically shows (b) a schematic diagram of performing non - overlapping sliding calculation on the above time series data of natural gas sales volume according to a preset window size to obtain a difference value sequence composed of the differences between the data means within adjacent windows in an embodiment of the present disclosure.
[0057] The above non - overlapping sliding means that there is no overlap between the data covered by adjacent windows, that is, when selecting adjacent windows, a whole - translation - type sliding is performed according to the preset window size. Taking the preset window size as w as an example.
[0058] In some embodiments, in the above step S120, according to the preset time window size, non-overlapping sliding calculation is performed on the above-mentioned time series data of natural gas sales volume, and a difference value sequence composed of the differences between the data means in adjacent time windows is obtained, including steps S121 - S123:
[0059] Step S121, for the time series data of natural gas sales volume A = [a1, a2,..., a n , where n represents the total number of data in the time series data of natural gas sales volume, non-overlapping sliding selection of data is performed according to the preset time window size w, and two groups of data in adjacent time windows are obtained, and these two groups of data are respectively expressed as: [a i , a i+1 , a i+2 ,..., a i+w -1], [a i+w , a i+w+1 , a i+w+2 ,..., a i+2w-1 , and the value range of i is 1 to s, where s = n - 2w + 1.
[0060] Step S122, for the two groups of data in adjacent time windows, calculate their respective means and then perform a subtraction operation to obtain the difference value x j .
[0061] Step S123, according to the above difference value x j , generate a difference value sequence X = [x1, x2, x3,..., x j ,..., x m , where m = s - 1.
[0062] In some embodiments, the above difference value x j = mean(a i + a i+1 + a i+2 +... + a i+w-1 ) - mean(a i+w + a i+w+1 + a i+w+2 +... + a i+2w-1 ), or:
[0063] x j = mean(a i+w + a i+w+1 + a i+w+2 +... + a i+2w-1 ) - mean(a i + a i+1 + a i+2 +... + a i+w-1 ), and the value range of j is 1 to s - 1;
[0064] For each group of adjacent time windows, the same method for calculating the difference value is used. That is, if the algorithm of subtracting the latter window from the former window is adopted, then all adjacent time windows adopt the algorithm of subtracting the latter window from the former window; if the algorithm of subtracting the former window from the latter window is adopted, then all adjacent time windows adopt the algorithm of subtracting the former window from the latter window. For example, referring to Figure 2 as shown in (b) of
[0065] , the algorithm of subtracting the former window from the latter window is used to calculate the difference value of all adjacent time windows. 10 Taking w = 5 and n = 100 as an example, the data corresponding to the first group of adjacent time windows selected are: [a1, a2, a3, a4, a5], [a6, a7, a8, a9, a 10 . The difference value corresponding to the first group of adjacent time windows is expressed as: x1 = mean(a1 + a2 + a3 + a4 + a5) - mean(a6 + a7 + a8 + a9 + a 10 , a 11 ); the data corresponding to the second group of adjacent time windows selected are: [a2, a3, a4, a5, a6], [a7, a8, a9, a 10 + a 11 . The difference value corresponding to the second group of adjacent time windows is expressed as: x2 = mean(a2 + a3 + a4 + a5 + a6) - mean(a7 + a8 + a9 + a 10 + a 11 ); and so on. For the s-th group, s = n - 2w + 1 = 100 - 10 + 1 = 91, the data corresponding to the adjacent time windows are: [a 91 , a 92 , a 93 , a 94 , a 95 , [a 96 , a 97 , a 98 , a 99 , a 100 . The difference value corresponding to the 91st group of adjacent time windows is expressed as: x 91 = mean(a 91 + a 92 + a 93 + a 94 + a 95 ) - mean(a 96 + a 97 + a 98 + a 99 + a 100 ).
[0066] It can be understood that the abscissa of the difference value also corresponds to time. For the sake of simplifying the operation process, the corresponding relationship between the subscript of the difference value and the corresponding time can also be established.
[0067] In some embodiments, the time corresponding to the middle sequence position of a group of adjacent time windows can be used as the time corresponding to the group of difference values. For example, the time t5 or t6 corresponding to a5 or a6 can be used as the time corresponding to the difference value x1. In other embodiments, according to the same selection algorithm (such as selecting the head position, middle position, second position, end position, or any intermediate calculation position, etc.), the time corresponding to a certain sequence position in a group of adjacent time windows can be used as the time corresponding to the group of difference values.
[0068] Step S130: For the difference value sequence X, slide step by step according to a single data with a preset time window size to obtain each window subsequence, and calculate the distances between each window subsequence and all subsequences respectively to obtain a subsequence distance sequence.
[0069] In some embodiments, this step at least includes sub-steps S131 - S132:
[0070] Step S131: For the difference value sequence X, slide step by step according to a single data with a preset time window size to obtain each window subsequence.
[0071] For the difference sequence X = [x1, x2, x3... x m , the first window subsequence [x r , x r+1 , x r+2 , …, x r+w-1 is obtained according to the preset time window size. r represents the serial number of the selected first sub-window data sequence, and the value range of r is 1 to m - w + 1. Referring to Figure 3 As shown in the first row of (a) in
[0072] Taking the preset window length w = 4 as an example, the first selected window subsequence is [0, 1, 3, 2]; the second selected window subsequence is [1, 3, 2, 9], and so on. The penultimate window subsequence is [1, 2, 2, 10], and the last window subsequence is [2, 2, 10, 7].
[0073] Step S132: Calculate the distances between each window subsequence and all subsequences respectively.
[0074] That is, calculate the distances between the first window subsequence and the first window subsequence, the second window subsequence, the third window subsequence until the last window subsequence respectively; then calculate the distances between the second window subsequence and the first window subsequence, the second window subsequence, the third window subsequence until the last window subsequence respectively; and so on, calculate the distances between the last window subsequence and the first window subsequence, the second window subsequence, the third window subsequence until the last window subsequence respectively. rk= distance([x r , x r+1 , x r+2 , …, x r+w-1 , [x k , x k+1 , x k+2 , …, x k+w-1 ); where distance() represents a function that calculates the difference between two vectors, and then uses the obtained distances to form a distance vector.
[0075] In practical applications, the following formula can be used to obtain the distance between each subsequence and all subsequences (using the Euclidean distance):
[0076]
[0077] E1 = [e 11 , e 12 , e 13 ... e 1j
[0078] where [x1, x2, x3, x4] is the first window subsequence, [x i , x i+1 , x i+2 , x i+3 represents all subsequences, j is the number of subsequences included in the difference sequence, e 1j represents the distance between the first window subsequence and all subsequences, where e 11 is the distance between the first window subsequence and the first window subsequence, e 12 is the distance between the first window subsequence and the second window subsequence, e 1j is the distance between the first window subsequence and the j-th window subsequence. E1 is the distance vector between the first subsequence and all subsequences.
[0079] Similarly, the distance between the second window subsequence and all subsequences can be obtained, and a distance vector E2 can be generated.
[0080]
[0081] E2 = [e 21 , e 22 , e 23 ... e 2j
[0082] Similarly, the distance between the j-th window and all subsequences can be obtained, and a distance set E j .
[0083] For ease of understanding, combined with Figure 3 As shown, the differential sequence X is now copied into two rows. When the first row is a determined window subsequence, the second row slides step by step according to the preset time window size to obtain each window subsequence, calculates the distance between one subsequence and all subsequences, and fills the distance values into the third row.
[0084] For example Figure 3 In (a), the first determined window subsequence is [0, 1, 3, 2]. The second row is used to obtain each window subsequence: [0, 1, 3, 2], [1, 3, 2, 9], …, [1, 2, 2, 10] and [2, 2, 10, 7]. The third row is used to record the distances between the subsequence [0, 1, 3, 2] and each window subsequence, which are E1 = [0, 7.4, 6.9, 14.7……].
[0085] Figure 3 (b) shows the method of calculating the distance between the first window subsequence [0, 1, 3, 2] and the ninth window subsequence [1, 2, 2, 9].
[0086] In Figure 3 (c), the second window subsequence obtained in the first row is [1, 3, 2, 9]. The second row is used to obtain each window subsequence: [0, 1, 3, 2], [1, 3, 2, 9], …, [1, 2, 2, 10] and [2, 2, 10, 7]. The third row is used to record the distances between the subsequence [1, 3, 2, 9] and each window subsequence, which are E2 = [7.4, 0, 10.9, 7.9……]
[0087] Step S140: Find the maximum value of the distances between each window subsequence and all subsequences in the obtained subsequence distance sequence to generate a maximum value sequence.
[0088] Use S to represent the maximum value sequence. In the maximum value sequence, the larger the data value, the higher the abnormality of the corresponding subsequence. Therefore, this sequence can be used as an anomaly evaluation sequence: S = [max(E1), max(E2), max(E3)... max(E j )]
[0089] Referring to Figure 3 as shown in (d), all distance vectors are represented in a matrix structure. And to highlight the maximum value in the above-mentioned difference degree, other elements in the matrix structure are hidden and only the value of the maximum value and the corresponding anomaly evaluation sequence are shown (as Figure 4 shown).
[0090] In step S150, detect the data distribution state of the maximum value sequence to obtain abnormal data and abnormal time periods.
[0091] In some embodiments, it specifically includes steps S151 - S154:
[0092] S151, sort the data in the above maximum value sequence according to the value size.
[0093] S152, determine the data whose value exceeds the set threshold as abnormal evaluation data according to the increasing trend and relative size of the sorted data.
[0094] The above set threshold is related to the data distribution state of the sorted evaluation data. In some embodiments, the expression of the above set threshold is as follows:
[0095] Set threshold = Q3 + 1.5 * (Q3 - Q1),
[0096] where, Q1 represents the value of the lower quartile in the sorted evaluation data; Q3 represents the value of the upper quartile in the sorted evaluation data. The above lower quartile refers to the data at the 1 / 4 position of the sequence. For example, if the total number of data in the sequence is p, if (p + 1) / 4 is an integer, then Q1 corresponds to the data at the (p + 1) / 4 position; if (p + 1) / 4 is not an integer, after rounding to the nearest integer to get the integer - form position number, the data corresponding to this position number is used as Q1. The upper quartile refers to the data at the 3 / 4 position of the sequence. For example, if the total number of data in the sequence is p, if 3(p + 1) / 4 is an integer, then Q1 corresponds to the data at the 3(p + 1) / 4 position; if 3(p + 1) / 4 is not an integer, after rounding to the nearest integer to get the integer - form position number, the data corresponding to this position number is used as Q3.
[0097] The statistical method of this quartile can be simply expressed by the following formula:
[0098] For the abnormal evaluation sequence S, those greater than Q3 + 1.5 * (Q3 - Q1) are abnormal. Q1 is the lower quartile, that is, the number at the 25% position after sorting from small to large; Q3 is the upper quartile, that is, the number at the 75% position after sorting from small to large.
[0099] The new vector generated after sorting the abnormal evaluation sequence S is Y = [y1, y2, y3... y p
[0100] Q1 = the value of the (p + 1) / 4 - th data. If (p + 1) / 4 is not an integer, round to the nearest integer
[0101] Q3 = the value of the (p + 1) / 4 * 3 - th data. If (p + 1) / 4 * 3 is not an integer, round to the nearest integer
[0102] When Q3 + 1.5 * (Q3 - Q1) > (p + 1), there is no outlier for this user.
[0103] When Q3 + 1.5 * (Q3 - Q1) <= (p + 1), y t >= Q3 + 1.5 * (Q3 - Q1) are all outliers, outlier = (y t , y t+1 ,,,, y p ).
[0104] S153. Determine the window subsequence corresponding to the above outlier evaluation data.
[0105] S154. According to the obtained window subsequence, find the natural gas sales time series data that generates the window subsequence, and analyze whether there are outlier data or outlier periods in the natural gas data corresponding to the natural gas sales time series data.
[0106] The method provided by the embodiments of the present disclosure calculates local outlier information by combining a sliding window, and then calculates outlier data using the idea of global similarity. The above maximum value sequence can reflect, from a global perspective, the situation with a relatively large degree of relative change difference in the distribution change trend of the time window data, and based on this, outlier data and outlier periods can be identified. This solution can reduce the probability of misidentifying data points that are locally outlier but globally normal as outliers for the natural gas sales time series data, and improve the recognition accuracy and applicability of outlier data in the natural gas sales time series data.
[0107] The second exemplary embodiment of the present disclosure provides a detection device for abnormal natural gas sales data.
[0108] Figure 6 Schematically shows a structural block diagram of a detection device for abnormal natural gas sales data according to an embodiment of the present disclosure.
[0109] Refer to Figure 6 As shown, the detection device 600 for abnormal natural gas sales data provided by the embodiments of the present disclosure includes: a data acquisition module 601, a first calculation module 602, a second calculation module 603, an outlier evaluation sequence generation module 604, and an outlier detection module 605.
[0110] The above data acquisition module 601 is used to acquire the natural gas sales time series data of the target user within the statistical duration.
[0111] The above statistical duration is the time period corresponding to the time series data of natural gas sales volume. The time granularity within the above statistical duration can be counted by week (which can be one week or multiple weeks), day (which can be one day or multiple days), hour (which can be one hour or multiple hours), etc. For example, the statistical duration can be counted by consecutive years (such as 3 - 5 years), consecutive quarters (such as several quarters), or consecutive months, etc.
[0112] The above first calculation module 602 is used to perform non - overlapping sliding calculation on the above time series data of natural gas sales volume according to a preset time window size, and obtain a difference value sequence composed of the differences between the data means within adjacent time windows; the non - overlapping sliding means that there is no overlap between the data covered by adjacent time windows.
[0113] Specifically, the first calculation module 602 is used for the time series data of natural gas sales volume A = [a1, a2, …, a n , where n represents the total number of data in the time series data of natural gas sales volume. Select data for non - overlapping sliding according to the preset time window size w, and obtain two groups of data within adjacent time windows, which are respectively expressed as: [a i , a i+1 , a i+2 , …, a i+w-1 , [a i+w , a i+w+1 , a i+w+2 , …, a i+2w - 1]. The value of i ranges from 1 to s, and s = n - 2w + 1.
[0114] It is also used to calculate the means of the two groups of data within adjacent time windows respectively, and then perform a subtraction operation to obtain the difference value x j . And according to the above difference value x j , generate a difference value sequence X = [x1, x2, x3, …, x j , …, x m , where m = s - 1.
[0115] The second calculation module 603 is used to, for the difference value sequence, obtain each window subsequence by gradually sliding according to a single data with a preset time window size, and calculate the distances between each window subsequence and all subsequences respectively, to obtain a subsequence distance sequence.
[0116] Specifically, the second calculation module 603 is used to, for the difference value sequence, first determine the first window subsequence; according to the determined first window subsequence, obtain each window subsequence by gradually sliding according to a single data with a preset time window size; calculate the distances between the first window subsequence and all subsequences respectively, to obtain the first window subsequence distance sequence; then calculate the distances between the remaining window subsequences and all subsequences respectively, to obtain the subsequence distance sequences of each window subsequence.
[0117] An abnormal evaluation sequence generation module 604 is configured to find the maximum value of the distances between each window subsequence and all subsequences in the obtained subsequence distance sequence, and generate a maximum value sequence.
[0118] In the maximum value sequence, the larger the data value, the higher the abnormality of the corresponding subsequence. Therefore, this sequence can be used as an abnormal evaluation sequence.
[0119] An abnormality detection module 605 is configured to detect the data distribution state of the maximum value sequence to obtain abnormal data and abnormal time periods.
[0120] Specifically, the abnormality detection module 605 is configured to sort the data in the above maximum value sequence according to the value size. According to the increasing trend and relative size of the sorted data, determine the data whose value exceeds the set threshold as abnormal evaluation data. Determine the window subsequence corresponding to the above abnormal evaluation data. Find the natural gas sales time series data that generates the window subsequence according to the obtained window subsequence, and analyze whether there is abnormal data or abnormal time periods in the natural gas data corresponding to the natural gas sales time series data.
[0121] The detection device for abnormal natural gas sales data of the present invention obtains the natural gas sales time series data of the target user within the statistical time period; performs non-overlapping sliding calculation on the natural gas sales time series data according to the preset window size to obtain a difference value sequence composed of the differences between the data means in adjacent time windows; the difference value sequence can reflect the distribution change trend of the time window data over time; for the difference value sequence, each window subsequence is obtained by sliding step by step according to the preset window size for each single data, and the distances between each window subsequence and all subsequences are calculated respectively to obtain a subsequence distance sequence; find the maximum value of the distances between each window subsequence and all subsequences in the obtained subsequence distance sequence, and generate a maximum value sequence; the maximum value sequence can reflect the situation with a relatively large change difference in the distribution change trend of the time window data from a global perspective, and based on this, abnormal data and abnormal time periods can be identified. For the natural gas sales time series data, the present invention can reduce the probability of identifying data points that are locally abnormal but globally normal as abnormal, and improve the identification accuracy of abnormal data in the natural gas sales time series data.
[0122] It can be understood that more details and beneficial effects of this embodiment can be referred to the description of the first embodiment, which will not be elaborated here.
[0123] Any number of the functional modules included in the above abnormal detection device 600 can be combined and implemented in one module, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. At least one of the functional modules included in the abnormal detection device 600 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), programmable logic array (PLA), system on chip, system on substrate, system on package, application specific integrated circuit (ASIC), or can be implemented by any other reasonable means such as hardware or firmware through circuit integration or packaging, or can be implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the functional modules included in the abnormal detection device 600 can be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions can be executed.
[0124] The third exemplary embodiment of the present disclosure provides an electronic device.
[0125] Figure 7 The structural block diagram of the electronic device provided by the embodiment of the present disclosure is schematically shown.
[0126] Refer to Figure 7 As shown, the electronic device 700 provided by the embodiment of the present disclosure includes a processor 701, a communication interface 702, a memory 703, and a communication bus 704. Among them, the processor 701, the communication interface 702, and the memory 703 complete mutual communication through the communication bus 704; the memory 703 is used to store a computer program; when the processor 701 executes the program stored on the memory, the detection method of the abnormal sales volume data of natural gas as described above is implemented.
[0127] The fourth exemplary embodiment of the present disclosure further provides a computer-readable storage medium. A computer program is stored on the above computer-readable storage medium, and when the computer program is executed by a processor, the detection method of the abnormal sales volume data of natural gas as described above is implemented.
[0128] The above computer-readable storage medium can be included in the device or apparatus described in the above embodiment; or it can exist alone and is not assembled into the device or apparatus. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiment of the present disclosure is implemented.
[0129] According to an embodiment of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or component.
[0130] It should be noted that in the technical solutions provided by the embodiments of the present disclosure, in aspects such as the collection, acquisition, update, analysis, processing, use, transmission, and storage of the user's personal information, they all comply with the provisions of relevant laws and regulations, are used for legal purposes, and do not violate public order and good customs. Necessary measures are taken for the user's personal information to prevent illegal access to the user's personal information data, and to safeguard the security of the user's personal information, network security, and national security.
[0131] It should be noted that in this document, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device including a series of elements not only includes those elements but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device including the said element.
[0132] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. A detection method for abnormal sales volume data of natural gas, characterized in that, Including: Obtain the time series data of the natural gas sales volume of the target user within the statistical duration; Perform non-overlapping sliding calculation on the time series data of the natural gas sales volume according to the preset time window size, and obtain a difference value sequence composed of the differences between the data means in adjacent time windows. The non-overlapping sliding means that there is no overlap between the data covered by adjacent time windows; For the difference value sequence, obtain each window subsequence by gradually sliding according to the preset time window size with a single data, and calculate the distances between each window subsequence and all subsequences respectively to obtain a subsequence distance sequence; Find the maximum value of the distances between each window subsequence and all subsequences in the obtained subsequence distance sequence, and generate a maximum value sequence; Detect the data distribution state of the maximum value sequence to obtain abnormal data and abnormal time periods.
2. The anomaly detection method according to claim 1, wherein For the difference value sequence, obtain each window subsequence by gradually sliding according to the preset time window size with a single data, and calculate the distances between each window subsequence and all subsequences respectively to obtain a subsequence distance sequence, including: For the difference value sequence, first determine the first window subsequence; According to the determined first window subsequence, obtain each window subsequence by gradually sliding according to the preset time window size with a single data; Calculate the distances between the first window subsequence and all subsequences respectively to obtain a first window subsequence distance sequence; Then calculate the distances between the remaining window subsequences and all subsequences respectively to obtain the subsequence distance sequences of each window subsequence.
3. The anomaly detection method according to claim 1, wherein Detect the data distribution state of the maximum value sequence to obtain abnormal data and abnormal time periods, including: Sort the data in the maximum value sequence according to the value size; According to the increasing trend and relative size of the sorted data, determine the data with a value exceeding the set threshold as abnormal evaluation data; Determine the window subsequence corresponding to the above abnormal evaluation data; Find the time series data of the natural gas sales volume that generates the window subsequence according to the obtained window subsequence, and analyze whether there is abnormal data or abnormal time periods in the natural gas data corresponding to the time series data of the natural gas sales volume.
4. The anomaly detection method according to claim 3, wherein The expression of the set threshold is as follows: Set threshold = Q3 + 1.5 * (Q3 - Q1), where Q1 represents the value of the lower quartile in the sorted data; Q3 represents the value of the upper quartile in the sorted data.
5. The anomaly detection method according to claim 1, wherein Calculate the distances between each window subsequence and all subsequences respectively, including: calculating the Euclidean distances between each window subsequence and all subsequences respectively.
6. The anomaly detection method according to claim 1, wherein Performing non-overlapping sliding calculation on the time series data of the natural gas sales volume according to the preset time window size to obtain a difference value sequence composed of the differences between the data means in adjacent time windows includes: For the time series data of the natural gas sales volume, perform non-overlapping sliding selection of data according to the preset time window size to obtain two groups of data in adjacent time windows; For the two groups of data in adjacent time windows, calculate their respective means and then perform a subtraction operation to obtain a difference value; Form a difference value sequence with the difference values obtained from all adjacent time windows.
7. A detection device for abnormal sales volume data of natural gas, characterized in that, Including: A data acquisition module for obtaining the time series data of the natural gas sales volume of the target user within the statistical duration; A first calculation module, configured to perform non-overlapping sliding calculation on the time series data of the natural gas sales volume according to a preset time window size, and obtain a difference value sequence composed of the differences between the data means in adjacent time windows; the non-overlapping sliding means that there is no overlap between the data covered by adjacent time windows; A second calculation module, configured to, for the difference value sequence, gradually slide according to a single data with a preset time window size to obtain each window subsequence, and calculate the distances between each window subsequence and all subsequences respectively, to obtain a subsequence distance sequence; An abnormal evaluation sequence generation module, configured to find the maximum value of the distances between each window subsequence and all subsequences in the obtained subsequence distance sequence, and generate a maximum value sequence; An abnormal detection module, configured to detect the data distribution state of the maximum value sequence to obtain abnormal data and abnormal time periods.
8. The detection device according to claim 7, characterized in that The abnormal detection module is specifically configured to: Sort the data in the maximum value sequence according to the value size; According to the increasing trend and relative size of the sorted data, determine the data with a value exceeding a set threshold as abnormal evaluation data; Determine the window subsequence corresponding to the above abnormal evaluation data; Find the time series data of the natural gas sales volume that generates the window subsequence according to the obtained window subsequence, and analyze whether there is abnormal data or abnormal time periods in the natural gas data corresponding to the time series data of the natural gas sales volume.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store a computer program; The processor, when executing the program stored on the memory, implements the method described in any one of claims 1-6.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method described in any one of claims 1-6.