Intelligent gas meter data acquisition system based on machine learning
By calculating the activity and historical evaluation index of gas data, dynamically setting the parameters of the DBSCAN algorithm, the problem of inaccurate clustering results is solved, and adaptive adjustment of gas data sampling frequency and optimization of system load is realized.
Patent Information
- Application Number
- CN202510607673.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-13
AI Technical Summary
In the prior art, the settings of clustering radius R and minimum number of points MinPts in the DBSCAN algorithm are susceptible to artificial intervention, resulting in inaccurate clustering results, which in turn affects the adaptive adjustment of the sampling frequency of gas meter data, resulting in redundant data and increasing system processing load.
By obtaining the gas data sequence of the target user, calculate the gas usage activity and historical activity evaluation index of each gas data, dynamically set the clustering radius and minimum number of points in the clustering algorithm, and then cluster and adaptively adjust the sampling frequency of the gas data.
It improves the accuracy of clustering results, realizes dynamic adjustment of the gas data sampling frequency, reduces redundant data and system processing load, and retains effective gas usage information.
Smart Images

Figure CN120180165A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural gas, and particularly to an intelligent gas meter data acquisition system based on machine learning. Background Art
[0002] Existing Internet of Things gas meters can automatically collect gas meter data and perform functions such as data transmission through the Internet, greatly increasing the convenience of user use. However, when traditional gas meter data is collected, it is often collected at a unified frequency through a fixed time interval (i.e., a fixed sampling frequency). This sampling method generally has a relatively high frequency, resulting in high storage overhead and large processing load.
[0003] In the prior art, a clustering algorithm (DBSCAN algorithm) is used to analyze the relationship between user behavior habits and time for the historical usage records (gas meter data) of users, so as to divide the historical usage records (gas meter data) of users into multiple clustering clusters. One clustering cluster represents a certain degree of gas usage habit, and then the sampling frequency of the gas meter data of users is adaptively adjusted according to the clustering clusters to obtain an adaptive sampling frequency. However, the clustering radius R and the minimum number of points MinPts in the DBSCAN algorithm are set manually and are easily affected by human intervention, which may lead to inaccurate clustering results, resulting in a large deviation in the adjustment of the adaptive sampling frequency, and then redundant data appears, increasing the system processing load.
[0004] Therefore, how to set the clustering radius and the minimum number of points in the DBSCAN algorithm to improve the accuracy of adjusting the sampling frequency of user gas meter data has become an urgent problem to be solved. Summary of the Invention
[0005] In view of this, an embodiment of the present invention provides an intelligent gas meter data acquisition system based on machine learning to solve the problem of how to set the neighborhood radius and the minimum number of points in the DBSCAN algorithm to improve the accuracy of adjusting the sampling frequency of user gas meter data.
[0006] An intelligent gas meter data acquisition system based on machine learning provided in an embodiment of the present invention includes a memory, a processor, and a computer program stored in the memory and running on the processor. The processor, when executing the computer program, implements the following steps: Obtain gas data at each sampling moment of a target user in the current period to obtain a gas data sequence; For any gas data in the gas data sequence, obtain the gas usage activity of the any gas data according to the difference between the any gas data and the neighboring gas data within its local range, and obtain the historical activity evaluation index of the any gas data according to the gas usage activity of the historical gas data at the same sampling moment as the any gas data within the historical period; Set the optimal clustering radius and the minimum number of points in the clustering algorithm according to the gas usage activity and the historical activity evaluation index of each gas data in the gas data sequence, and cluster the gas data sequence according to the optimal clustering radius and the minimum number of points to obtain multiple clustering clusters; Adaptively obtain the gas data sampling frequency of the target user according to the gas usage activity of the gas data in each clustering cluster, for collecting the gas data of the target user in the next period after the current period.
[0007] The beneficial effects of the embodiments of the present invention compared with the prior art are: The present invention analyzes the gas data sequence of the target user in the current period to adaptively set the clustering radius and the minimum number of points during clustering, so that the parameters during clustering can better fit the gas usage situation of the target user in the current period, thereby obtaining a more accurate clustering result. Further, adaptively adjust the sampling frequency of the gas data of the target user in the future period according to the clustering result, realize the dynamic acquisition adjustment of the gas data of the target user, so that the target user increases the sampling frequency when using gas frequently and decreases the sampling frequency when using gas infrequently, so as to avoid the appearance of redundant data, reduce the system load processing while retaining the effective gas usage information of the target user. Description of the Drawings
[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.
[0009] Figure 1 It is a method flow chart of a smart gas meter data acquisition method based on machine learning provided by Embodiment 1 of the present invention; Figure 2 It is an example diagram of a scatter plot provided by the embodiments of the present invention. Detailed Embodiments
[0010] Embodiments of the present disclosure will be described in detail below, and examples of the embodiments are shown in the accompanying drawings. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, and should not be construed as limiting the present disclosure.
[0011] It should be noted that the terms "first", "second", etc. in the specification of the present disclosure and the above-mentioned accompanying drawings are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure.
[0012] In order to illustrate the technical solution of the present invention, specific embodiments will be used for illustration below.
[0013] An embodiment of the present invention provides a smart gas meter data acquisition system based on machine learning, including a processor and a memory. The processor executes the computer program stored in the memory to implement a smart gas meter data acquisition method based on machine learning, as Figure 1 shown. The smart gas meter data acquisition method based on machine learning includes the following steps: Step S101, obtain gas data at each sampling moment of a target user in the current period to obtain a gas data sequence.
[0014] Take one of the users using an Internet of Things gas meter as the target user. At the same time, take one day as a period, and take the i-th day as the current period. In order to analyze the gas usage habits of the target user, set the sampling frequency to 1 time per minute, and obtain the gas data of the target user at each sampling moment on the i-th day, corresponding to the gas data sequence in time series, for subsequent adaptive setting of the clustering radius and the minimum number of points in the DBSCAN clustering algorithm based on the data feature distribution of the new gas data sequence, and then adjust the sampling frequency of the gas data of the target user according to the clustering result, reduce redundant data, and reduce the system processing load.
[0015] Step S102, for any gas data in the gas data sequence, obtain the gas usage activity of the any gas data according to the difference between the any gas data and the neighboring gas data within its local range, and obtain the historical activity evaluation index of the any gas data according to the gas usage activity of the historical gas data at the same moment as the sampling moment of the any gas data in the historical period.
[0016] Since the changes in gas data are closely related to users' usage habits and behavioral activities, the gas data may exhibit different change states at different time periods. Secondly, the change in gas data of the gas meter generally occurs due to the use of gas. Therefore, the gas usage amount is directly reflected in the data change. Thus, the gas data in the embodiments of the present invention includes the gas usage amount and the remaining gas amount, that is, two-dimensional monitoring data of the gas meter are collected at a sampling moment. Without limitation here, parameters such as the cumulative gas usage amount and unit usage amount of the gas meter can also be collected.
[0017] Since the clustering data in the DBSCAN clustering algorithm is at least two-dimensional data, in the embodiments of the present invention, each gas data in the gas data sequence is used as a clustering sample, and two attribute parameters of each clustering sample are obtained according to the change in the gas data in the gas data sequence for constructing the two-dimensional clustering space of the gas data sequence.
[0018] Taking any gas data in the gas data sequence as an example, for the first attribute parameter of any gas data, the gas usage change fluctuation corresponding to the sampling moment of any gas data can be obtained according to the gas usage situation of any gas data on the i-th day, that is, the gas usage activity of any gas data. The specific process is as follows: Taking the any gas data as the center in the gas data sequence, a gas data subsequence with a preset length is obtained, the standard deviation of all gas usage amounts in the gas data subsequence is calculated, and the fluctuation stability degree of the any gas data is obtained according to the reciprocal of the sum of the standard deviation and the constant 1; Obtaining the previous gas data of the any gas data in the gas data sequence, calculating the absolute value of the difference in the remaining gas amount between the any gas data and the previous gas data, and obtaining the change stability degree of the any gas data according to the reciprocal of the sum of the absolute value of the difference and the constant 1; Calculating the average value between the fluctuation stability degree and the change stability degree, and obtaining the gas usage activity of the any gas data according to the difference between the constant 1 and the average value.
[0019] In an embodiment, in order to prevent the data change of any gas data from being reduced due to the too large length of the gas data subsequence, or to avoid the inability to reflect the data change of any gas data due to the too small length of the gas data subsequence, taking any gas data as the center, the 5 gas data adjacent to it on the left and right are used to form the gas data subsequence. Among them, the calculation formula for the gas usage activity of any gas data is:
[0020] Wherein, represents the gas usage activity of any gas data x, and 1 represents a constant. represents the standard deviation of all gas usage amounts in the gas data subsequence of any gas data x. represents the absolute value of the difference in the remaining gas amount between any gas data x and any gas data x - 1.
[0021] It should be noted that The larger the value of, it indicates that there are large fluctuations in the gas usage amount in the gas data subsequence, characterizing that the gas usage of the target user is relatively frequent, that is, the change stability degree of any gas data x is smaller, and the gas usage activity corresponding to any gas data x is larger; The larger the value of, it indicates that there is a large difference in the gas usage state between any gas data x and any gas data x - 1, and the gas usage is in an active usage stage, that is, the change stability degree of any gas data x is smaller, and the gas usage activity corresponding to any gas data x is larger.
[0022] After obtaining the gas usage activity of any gas data, for the second attribute parameter of any gas data, considering that the user behavior habits are regular, the gas usage activities of some clustering samples on historical daily basis will remain the same or similar. However, due to the uncontrollability of user behavior, there will be some accidental usage phenomena, and the longer the time, the higher the possibility that the user's behavior habits change or vary. Therefore, in the embodiments of the present invention, the gas data of the recent historical days are used as the main analysis object to analyze the gas usage activity of any gas data at the sampling moment in the past daily, and then used as the historical activity evaluation index of any gas data. The specific process is as follows: Taking the past 5 days as the historical period, and according to the sampling moment corresponding to any gas data, obtaining the gas data at the same sampling moment in the past daily as the historical gas data. Referring to the method for obtaining the gas usage activity of any gas data as described above, obtaining the gas usage activity of each historical gas data; According to the interval period between each historical period and the current period, obtaining the weight corresponding to each historical period, the smaller the interval period, the larger the corresponding weight; according to the weight corresponding to each historical period, performing weighted summation on the gas usage activities of all historical gas data to obtain the historical activity evaluation index of any gas data.
[0023] In an embodiment, the calculation formula for the weight corresponding to the t - th historical period is:
[0024] Wherein, represents the weight corresponding to the t - th historical period, represents the normalization function. Denote the i-th day, which is also the current period, Denote the t-th historical period, Denote the interval period between the t-th historical period and the current period, and 1 represents a constant.
[0025] It should be noted that the longer the interval period, the lower the reference value of the historical gas data of the corresponding historical period and the smaller the corresponding weight.
[0026] The calculation expression of the historical activity evaluation index of any gas data x is:
[0027] Among them, Denote the historical activity evaluation index of any gas data x, N represents the number of historical periods, Denote the weight corresponding to the t-th historical period, Denote the gas usage activity of the historical gas data of the t-th historical period.
[0028] It should be noted that The larger the value of, the higher the active state of the target user in the same moment of multiple historical days when using gas.
[0029] So far, the gas usage activity and the historical activity evaluation index of any gas data have been obtained. Similarly, the gas usage activity and the historical activity evaluation index of each gas data in the gas data sequence are obtained.
[0030] Step S103, according to the gas usage activity and the historical activity evaluation index of each gas data in the gas data sequence, set the optimal clustering radius and the minimum number of points in the clustering algorithm, and cluster the gas data sequence according to the optimal clustering radius and the minimum number of points to obtain multiple clustering clusters.
[0031] After obtaining the gas usage activity and the historical activity evaluation index of each gas data in the gas data sequence, take the gas usage activity as the vertical axis and the historical activity evaluation index as the horizontal axis to construct a scatter plot, and map each gas data in the gas data sequence to obtain the corresponding scatter points in the scatter plot according to the gas usage activity and the historical activity evaluation index of each gas data in the gas data sequence.
[0032] Since the gas usage status of users is different in different time periods, the gas data will show phased characteristics, which are manifested as regional aggregation in the scatter plot. Therefore, the clustering radius in the DBSCAN clustering algorithm can be set according to the influence degree and distribution characteristics between the scatter points in the scatter plot, that is, the optimal clustering radius that conforms to the target user. The method for obtaining the optimal clustering radius is: (1) For any scatter point in the scatter plot, based on the Euclidean distance, the difference in gas usage activity, and the difference in historical activity evaluation index between the any scatter point and each other scatter point, the relative neighborhood radius of the any scatter point is obtained.
[0033] Specifically, each scatter point in the scatter plot except the any scatter point is used as an other scatter point. For any other scatter point, the absolute value of the difference between the gas usage activity and the historical activity evaluation index of the any scatter point is calculated, and the reciprocal of the absolute value of the difference is used as the first probability that the any scatter point is not an isolated point; the absolute value of the difference between the gas usage activity and the historical activity evaluation index of the any other scatter point is calculated, and the reciprocal of the absolute value of the difference is used as the second probability that the any other scatter point is not an isolated point; The reciprocal of the absolute value of the difference between the first probability and the second probability is used as the state similarity between the any scatter point and the any other scatter point. The Euclidean distance between the any scatter point and the any other scatter point is obtained. Based on the mean between the reciprocal of the Euclidean distance and the state similarity, the correlation coefficient between the any scatter point and the any other scatter point is obtained; Based on all the correlation coefficients corresponding to the any scatter point, the correlation coefficient threshold corresponding to the any scatter point is obtained. The other scatter points corresponding to the correlation coefficient greater than or equal to the correlation coefficient threshold are used as the reference scatter points of the any scatter point. The maximum Euclidean distance among the Euclidean distances between the any scatter point and each of the reference scatter points is obtained as the relative neighborhood radius of the any scatter point.
[0034] In an embodiment, taking the j-th scatter point as an example, there are some scatter points with relatively close distances and some scatter points with relatively far distances from the j-th scatter point in the scatter plot. The scatter points with relatively close distances may have an associated relationship with the j-th scatter point. Then, the calculation expression of the correlation coefficient between the j-th scatter point and each other scatter point is:
[0035] where, represents the correlation coefficient between the j-th scatter point and the b-th other scatter point, represents the Euclidean distance between the j-th scatter point and the b-th other scatter point, represents the historical activity evaluation index of the j-th scatter point, represents the gas usage activity of the j-th scatter point, represents the historical activity evaluation index of the b-th scatter point, represents the gas usage activity of the b-th scatter point, and | | represents the absolute value symbol.
[0036] It should be noted that, The larger the value is, the greater the change difference of the j-th scatter point compared with the historical gas data, and the more likely it is to be an outlier. Therefore, using represents the first probability that the j-th scatter point is not an outlier. Similarly, is used to represent the second probability that the b-th scatter point is not an outlier. The greater the difference between them, the less similar the gas usage states between the two scatter points, and the smaller the corresponding correlation coefficient between them. At the same time, the smaller the Euclidean distance between them, the greater the probability that the two scatter points belong to the same cluster class, and the greater the corresponding correlation coefficient.
[0037] Similarly, obtain the correlation coefficient between the j-th scatter point and each other scatter point, and construct a statistical histogram of all correlation coefficients according to the correlation coefficient between the j-th scatter point and each other scatter point. The horizontal axis of the statistical histogram is the correlation coefficient, and the vertical axis is the number of correlation coefficients. According to the clustering distribution characteristics, the members of the same cluster class will be concentrated, while the non-cluster class members will be discrete. Therefore, the correlation coefficients between the j-th scatter point and other scatter points will show a two-level performance, with some correlation coefficients being larger and some being smaller, and there may be a gap in the correlation coefficients in the middle. Therefore, count the correlation coefficient intervals where the number of correlation coefficients is continuously 0 in the statistical histogram to obtain the maximum correlation coefficient interval, and take the median value in the maximum correlation coefficient interval as the correlation coefficient threshold corresponding to the j-th scatter point.
[0038] After obtaining the correlation coefficient threshold of the j-th scatter point, compare the correlation coefficient between the j-th scatter point and each other scatter point with the correlation coefficient threshold respectively, and obtain the other scatter points whose correlation coefficient is greater than or equal to the correlation coefficient threshold as the reference scatter points of the j-th scatter point. These reference scatter points form the neighborhood of the j-th scatter point, and then obtain the maximum Euclidean distance among the Euclidean distances between the j-th scatter point and each reference scatter point as the relative neighborhood radius of the j-th scatter point.
[0039] Similarly, refer to the method for obtaining the relative neighborhood radius of the j-th scatter point to obtain the relative neighborhood radius of each scatter point in the scatter plot.
[0040] (2)Obtain the relative neighborhood radius of each scatter point in the scatter plot, and screen the best relative neighborhood radius among all relative neighborhood radii as the best clustering radius in the clustering algorithm.
[0041] Specifically, obtain the radius interval composed of all relative neighborhood radii, divide the radius interval into a preset number of sub-intervals, respectively count the number of scatter points corresponding to the relative neighborhood radii included in each sub-interval, and take the proportion of the number of scatter points corresponding to each sub-interval in the total number of scatter points as the frequency of each relative neighborhood radius included in the corresponding sub-interval; For any relative neighborhood radius, respectively use the scatter points corresponding to each other relative neighborhood radius as the centers, and the other relative neighborhood radii as the circle radii to construct circular regions; count the number of circular regions that contain the scatter points corresponding to the any relative neighborhood radius, and take it as the occurrence times of the scatter points corresponding to the any relative neighborhood radius; obtain the preferred evaluation value of the any relative neighborhood radius according to the product between the occurrence times of the scatter points corresponding to the any relative neighborhood radius and the frequency. Obtain the preferred evaluation value of each of the relative neighborhood radii, and take the relative neighborhood radius corresponding to the maximum preferred evaluation value as the optimal relative neighborhood radius.
[0042] In an embodiment, with 10 as the basic number of interval divisions, divide the radius interval composed of all relative neighborhood radii into sub-intervals according to the difference between the relative neighborhood radii. For example: the minimum value of the relative neighborhood radius is 10.6, and the maximum value of the relative neighborhood radius is 16.3, then set the size of the sub-interval as If the difference is too large, the multiple of 10 can be appropriately increased.
[0043] Since the clustering center is the easiest to find, the scatter points can be analyzed to find the scatter points with the densest distribution and the most likely to become the clustering center. The relative neighborhood radius of this scatter point is most likely to be the appropriate clustering radius. Therefore, take the scatter points corresponding to the relative neighborhood radii included in each sub-interval as the quantity value of the sub-interval, and take the proportion of the quantity value corresponding to each sub-interval in the total number of all scatter points as the frequency of each relative neighborhood radius included in the corresponding sub-interval.
[0044] At the same time, calculate the number of times each scatter point appears in the relative neighborhood regions of the remaining scatter points, that is, when constructing circular regions with the remaining scatter points as the centers and the relative neighborhood radii of the remaining scatter points as the circle radii, the number of times the selected scatter point (target scatter point) appears in these circular regions is taken as the occurrence times of the scatter points corresponding to the corresponding relative neighborhood radius.
[0045] Therefore, taking the k-th relative neighborhood radius as an example, the calculation expression for the preferred evaluation value of the k-th relative neighborhood radius as the optimal relative neighborhood radius is:
[0046] Wherein, represents the preferred evaluation value of the k-th relative neighborhood radius, represents the frequency of the k-th relative neighborhood radius, represents the occurrence times of the scatter points corresponding to the k-th relative neighborhood radius.
[0047] It should be noted that if the frequency of the relative neighborhood radius is relatively large and the number of occurrences of the corresponding scatter points is large, the preferred evaluation value of the relative neighborhood radius is relatively large. The larger the preferred evaluation value, the more likely the relative neighborhood radius is the clustering radius of the clustering center.
[0048] Similarly, obtain the preferred evaluation value of each relative neighborhood radius, and take the relative neighborhood radius corresponding to the largest preferred evaluation value as the optimal relative neighborhood radius. Furthermore, take the optimal relative neighborhood radius as the optimal clustering radius when clustering all the scatter points in the scatter plot.
[0049] Furthermore, after obtaining the optimal clustering radius, set the minimum number of points MinPts in the DBSCAN clustering algorithm according to the scatter point density distribution within the neighborhood range corresponding to the optimal clustering radius. Due to the differences in user habits, the form of the scatter plot may not necessarily be in the form of multiple obvious clusters, and it may also exhibit other characteristics. For example, when the user's gas usage habits remain consistent daily, at this time, the gas usage activity levels of most scatter points are consistent with the historical activity evaluation index, showing Figure 2 the scatter point distribution shown.
[0050] Since the scatter plot may contain discrete points, discrete points will affect the judgment results when analyzing user behavior habits. Secondly, when clustering, a standard cluster should have the characteristics of dense distribution of cluster members and no obvious discrete points within the cluster. Therefore, if the minimum number of points in the clustering is selected too small, it is easy for discrete points to become pseudo-clusters and affect the judgment results. However, for Figure 2 the scatter point distribution situation, if the minimum number of points is too large, it is easy to cause real clusters to be divided into noise. Therefore, in the embodiments of the present invention, according to the above-obtained optimal clustering radius and the scatter point distribution situation in the scatter plot, obtain the minimum number of points MinPts in the DBSCAN clustering algorithm.
[0051] First, judge whether the scatter point distribution in the scatter plot conforms to Figure 2 the distribution in, that is, judge whether the user's gas usage habits remain consistent daily. Then, adaptively obtain the minimum number of points for DBSCAN clustering of the scatter plot according to the judgment result. The specific process is as follows: Take the scatter points corresponding to the optimal clustering radius as target scatter points, and construct a sliding window with the target scatter points as the center and the optimal clustering radius as the side length. Use the sliding window to divide the scatter plot into regions, obtaining multiple scatter point regions containing scatter points. The size of the scatter point regions is the same as the size of the sliding window, and at the same time, ignore the regions without scatter points. Then analyze each scatter point region. For any scatter point region, obtain the region center of the any scatter point region, obtain the distance between each scatter point in the any scatter point region and the region center, and take the scatter point corresponding to the minimum distance as the central scatter point of the any scatter point region.
[0052] Similarly, obtain the central scatter points of each scatter point region, perform linear fitting on all the central scatter points in the scatter plot to obtain the fitting line and the corresponding line slope, obtain the fitting value of the gas usage activity of each central scatter point according to the fitting line, calculate the difference between the gas usage activity of each central scatter point and the corresponding fitting value, obtain the average difference value, and take the reciprocal of the average difference value as the distribution uniformity index of the central scatter points. It should be noted that the least squares method is used for linear fitting, and the least squares method belongs to the prior art and will not be elaborated here.
[0053] According to the experimental statistics, set the distribution uniformity index threshold to 0.8. If the distribution uniformity index of the central scatter points is greater than or equal to the preset distribution uniformity index threshold, it indicates that the scatter points in the scatter plot conform to Figure 2 the scatter point distribution in, which reflects that the gas usage habits of the target users remain consistent every day, and the probability of corresponding isolated points appearing is relatively low. Therefore, count the number of scatter points in each scatter point region to obtain the maximum number of scatter points, and use the data dimension corresponding to the scatter plot as the scatter point number threshold, that is, the scatter point number threshold is 2, and delete the scatter point regions with the number of scatter points less than or equal to 2.
[0054] In order to prevent real clusters from being classified as noise, obtain the minimum number of points according to the scatter point distribution uniformity of the remaining scatter point regions. The obtaining method is as follows: based on the number of scatter points in each remaining scatter point region, perform standard deviation analysis on the number of scatter points greater than 2 to obtain the standard deviation of the number of scatter points; calculate the absolute value of the difference between the reciprocal of the standard deviation of the number of scatter points and the constant 1, obtain the product of the absolute value of the difference and the maximum number of scatter points, and use the product as the point adjustment amount. Perform a floor operation on the difference between the maximum number of scatter points and the point adjustment amount to obtain the minimum number of points in the clustering algorithm.
[0055] In one embodiment, the calculation formula for the minimum number of points is:
[0056] where, represents the minimum number of points, represents the maximum number of scatter points, represents the standard deviation of the number of scatter points of all scatter point regions after excluding the scatter point regions with the number of scatter points less than or equal to the scatter point number threshold, and 1 represents a constant.
[0057] It should be noted that Indicates the uniformity of the scatter point distribution of each scatter point area. If the scatter point distribution of each scatter point area is relatively uniform, the minimum number of points can be close to the number of scatter points in the scatter point area with the most scatter points. If the scatter point distribution of each scatter point area is uneven, in order to ensure that the real clusters are not classified as noise, the minimum number of points needs to be set smaller. Therefore, the minimum number of points can be significantly different from the number of scatter points in the scatter point area with the largest number of scatter points.
[0058] If the distribution uniformity index of the central scatter point is less than the preset distribution uniformity index threshold, it means that the distribution of the scatter points in the scatter plot does not meet the requirements. Figure 2 The scatter point distribution in the scatter plot needs to obtain the minimum number of points according to the distribution of isolated points in the scatter plot. The specific acquisition method is: obtain the probability that each scattered point in the scatter plot is an isolated point, screen and obtain at least one suspected isolated point, obtain the circular area of each suspected isolated point, count the number of suspected isolated points in the circular area of each suspected isolated point, and obtain the maximum number of suspected isolated points and the mean number of suspected isolated points; calculate the difference between the constant 1 and the reciprocal of the mean number of suspected isolated points, obtain the product between the difference and the total number of suspected isolated points in the scatter plot, use the product as the point adjustment amount, round up the sum of the maximum number of suspected isolated points and the point adjustment amount, and obtain the minimum number of points in the clustering algorithm.
[0059] In one embodiment, the probability of each scattered point being an isolated point is obtained according to the historical activity evaluation index and gas usage activity corresponding to each scattered point in the scatter plot, and the calculation expression is: Since isolated points are accidental phenomena, a histogram is constructed for the probability of all scattered points being isolated points. The horizontal axis of the histogram is probability, and the vertical axis is the number of probabilities. Since there will be obvious low-frequency probabilities in the histogram, through experimental calculation and statistics, it can be obtained that the probability of obvious low frequency is in the range of 0.2. Therefore, the probability threshold is set to 0.2, and the scattered points with an isolated point probability lower than 0.2 are regarded as suspected isolated points. At this point, at least one suspected isolated point in the scatter plot is obtained.
[0060] According to the above circular area construction method, the circular area of each suspected isolated point is obtained, and the minimum number of points is obtained according to the distribution of the suspected isolated points in each circular area. The calculation formula for the minimum number of points is:
[0061] in, Indicates the minimum number of points, represents the maximum number of suspected isolated points, H represents the number of suspected isolated points, Represents the number of suspected isolated points in the circular area of the vth suspected isolated point, and 1 represents a constant.
[0062] It should be noted that characterizes the average number of suspected isolated points within the region composed of all suspected isolated points. The larger this average value, the denser the distribution of suspected isolated points, corresponding to a larger minimum number of points required. Conversely, the smaller this average value, the more discrete the distribution of suspected isolated points, corresponding to a smaller minimum number of points required.
[0063] Thus, the optimal clustering radius and the minimum number of points for performing DBSCAN clustering on all scatter points in the scatter plot are obtained. Then, based on the optimal clustering radius and the minimum number of points, the gas data sequence (i.e., all scatter points in the scatter plot) is clustered to obtain multiple clustering clusters. Among them, DBSCAN clustering belongs to the prior art and will not be elaborated here.
[0064] Step S104: Adaptively obtain the gas data sampling frequency of the target user according to the gas usage activity of the gas data in each clustering cluster, for collecting the gas data of the target user in the next cycle after the current cycle.
[0065] After clustering all scatter points in the scatter plot to obtain multiple clustering clusters, considering that each clustering cluster represents the gas usage situation of the target user at different time periods on the i-th day, and each scatter point in the clustering cluster corresponds to a gas data in the gas data sequence. Therefore, adaptively obtain the gas data sampling frequency of the target user according to the gas usage activity of the gas data in each clustering cluster, so as to collect the gas data on the (i + 1)-th day using the gas data sampling frequency.
[0066] Among them, adaptively obtaining the gas data sampling frequency of the target user according to the gas usage activity of the gas data in each clustering cluster includes: (1) Obtain the clustering cluster with the most gas data as the target clustering cluster. According to the gas usage activity corresponding to each gas data in the target clustering cluster, obtain the average value of gas usage activity. According to the difference between the constant 1 and the reciprocal of the number of clustering clusters, obtain the diversity index. Take the sum of the average value of gas usage activity and the diversity index as the representation value of the gas usage active state of the target user in the current cycle.
[0067] In one embodiment, the calculation expression of the representation value of the gas usage active state of the target user on the i-th day is:
[0068] Wherein, represents the representation value of the gas usage active state of the target user on the i-th day, K represents the number of gas data in the target clustering cluster, It represents the gas usage activity corresponding to the k-th gas data in the target clustering cluster, 1 represents a constant, and CN represents the number of clustering clusters.
[0069] It should be noted that The larger the value of, the more the target users use gas, and the greater the gas usage activity when using gas. is the diversity index. The larger its value, the more the number of clustering clusters, indicating that the target users may have various gas usage habits at different levels, and the larger the characterization value of the gas usage active state of the target users on the i-th day.
[0070] (2) According to the gas usage activity of each gas data in the gas data sequence, the difference between the number of gas data with gas usage activity of 0 and the number of gas data with non-zero gas usage activity is used as the sampling frequency modification value.
[0071] (3) Similarly, according to the method for obtaining the characterization value of the gas usage active state of the target users on the i-th day as described above, obtain the historical gas usage active state characterization value of the target users in the previous cycle (that is, the (i - 1)-th day). If the gas usage active state characterization value is greater than or equal to the historical gas usage active state characterization value, calculate the difference between the constant 1 and the average gas usage activity, calculate the average value between the difference and the reciprocal of the number of clustering clusters, obtain the product value between the average value and the sampling frequency modification value, and add the product value to the sampling frequency in the current cycle to obtain the gas data sampling frequency of the target users.
[0072] In an embodiment, the calculation expression of the gas data sampling frequency of the target users is:
[0073] Among them, represents the gas data sampling frequency of the target users, that is, the gas data sampling frequency on the (i + 1)-th day. represents the sampling frequency in the current cycle, that is, the gas data sampling frequency on the i-th day, and G represents the sampling frequency modification value.
[0074] It should be noted that if the characterization value of the gas usage active state of the target users on the i-th day is greater than the characterization value of the gas usage active state on the (i - 1)-th day, it indicates that the gas usage of the target users shows an increasing trend, and the sampling frequency of the gas data on the (i + 1)-th day should be increased to more efficiently detect the gas usage habits of the target users on the (i + 1)-th day.
[0075] (4) If the characterization value of the active state of gas usage is less than the historical characterization value of the active state of gas usage, then calculate the difference between 1 and the average value of the gas usage activity, calculate the average value between the difference and the reciprocal of the number of clustering clusters, obtain the product value between the average value and the sampling frequency modification value, and subtract the product value from the sampling frequency within the current period to obtain the gas data sampling frequency of the target user.
[0076] In an embodiment, the calculation expression of the gas data sampling frequency of the target user is:
[0077] Wherein, represents the gas data sampling frequency of the target user, that is, the gas data sampling frequency on the (i + 1)-th day, represents the sampling frequency within the current period, that is, the gas data sampling frequency on the i-th day, and G represents the sampling frequency modification value.
[0078] So far, based on the gas data sequence collected from the target user on the i-th day, the gas data sampling frequency of the target user on the (i + 1)-th day is adaptively obtained. Then, the gas data of the target user can be collected and stored using the gas data sampling frequency of the target user on the (i + 1)-th day, avoiding redundancy of stored data and reducing the system processing load. Similarly, for each user using an IoT gas meter, the gas data sampling frequency of each user every day can be adaptively obtained for gas data collection.
[0079] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. The gas meter data intelligent collection system based on machine learning is characterized by: The method comprises a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the following steps when executing the computer program: Obtain the gas data of the target user at each sampling time in the current cycle to obtain a gas data sequence; For any gas data in the gas data sequence, the gas usage activity of the any gas data is obtained according to the difference between the any gas data and the neighboring gas data in a local range, and the historical activity evaluation index of the any gas data is obtained according to the gas usage activity of the historical gas data at the same time as the sampling time of the any gas data in the historical period; According to the gas usage activity and the historical activity evaluation index of each gas data in the gas data sequence, setting the optimal clustering radius and the minimum number of points in the clustering algorithm, clustering the gas data sequence according to the optimal clustering radius and the minimum number of points to obtain a plurality of cluster clusters; According to the gas usage activity of the gas data in each of the clusters, the gas data sampling frequency of the target user is adaptively acquired to collect the gas data of the target user in a next cycle after the current cycle.
2. The gas meter data intelligent collection system based on machine learning according to claim 1 is characterized in that: The gas data includes gas usage and gas remaining. The gas usage activity of any gas data is obtained according to the difference between any gas data and neighboring gas data in a local range thereof, including: In the gas data sequence, taking any gas data as the center, obtaining a gas data subsequence of a preset length, calculating the standard deviation of all gas usage in the gas data subsequence, and obtaining the fluctuation stability of any gas data according to the reciprocal of the sum of the standard deviation and a constant 1; Obtaining the previous gas data of any gas data in the gas data sequence, calculating the absolute value of the difference in the remaining gas amount between any gas data and the previous gas data, and obtaining the degree of stability of change of any gas data according to the inverse of the sum of the absolute value of the difference and a constant 1; An average value between the fluctuation stability and the change stability is calculated, and the gas usage activity of any gas data is obtained according to a difference between a constant 1 and the average value.
3. The gas meter data intelligent collection system based on machine learning according to claim 1 is characterized in that: The obtaining of the historical activity evaluation index of any gas data according to the gas usage activity of the historical gas data at the same time as the sampling time of any gas data in the historical period includes: According to the interval period between each historical period and the current period, the weight corresponding to each historical period is obtained, and the smaller the interval period, the larger the corresponding weight; according to the weight corresponding to each historical period, the gas usage activity of all historical gas data is weightedly summed to obtain the historical activity evaluation index of any gas data.
4. The gas meter data intelligent collection system based on machine learning according to claim 1 is characterized in that: The step of setting the optimal clustering radius and the minimum number of points in the clustering algorithm according to the gas usage activity and the historical activity evaluation index of each gas data in the gas data sequence includes: A scatter plot is constructed with the gas usage activity as the vertical axis and the historical activity evaluation index as the horizontal axis, and a scatter point corresponding to each gas data in the gas data sequence is mapped in the scatter plot according to the gas usage activity and the historical activity evaluation index of each gas data in the gas data sequence; For any scatter point in the scatter plot, the relative neighborhood radius of any scatter point is obtained according to the Euclidean distance between the scatter point and each other scatter point, the difference in gas usage activity, and the difference in historical activity evaluation index; the relative neighborhood radius of each scatter point in the scatter plot is obtained, and the best relative neighborhood radius is selected from all relative neighborhood radii as the best clustering radius in the clustering algorithm; The scattered point corresponding to the optimal clustering radius is used as the target scattered point, and a sliding window is constructed with the target scattered point as the center and the optimal clustering radius as the side length, and the scatter plot is divided into regions using the sliding window to obtain a plurality of scattered point regions containing scattered points, and the size of the scatter point region is the same as the size of the sliding window; For any scattered point area, obtain the area center of any scattered point area, obtain the distance between each scattered point in any scattered point area and the area center, take the scattered point corresponding to the minimum distance as the central scattered point of any scattered point area, and obtain the minimum number of points in the clustering algorithm according to the distribution characteristics of the central scattered points of all scattered point areas.
5. The gas meter data intelligent collection system based on machine learning according to claim 4 is characterized in that: The relative neighborhood radius of any scattered point is obtained according to the Euclidean distance, gas usage activity difference and historical activity evaluation index difference between any scattered point and each other scattered point, including: Taking each scatter point in the scatter plot except the any scatter point as another scatter point, for any other scatter point, calculating the absolute value of the difference between the gas usage activity of any other scatter point and the historical activity evaluation index, and taking the reciprocal of the absolute value of the difference as a first probability that the any scatter point is not an isolated point; calculating the absolute value of the difference between the gas usage activity of any other scatter point and the historical activity evaluation index, and taking the reciprocal of the absolute value of the difference as a second probability that the any other scatter point is not an isolated point; Taking the reciprocal of the absolute value of the difference between the first probability and the second probability as the state similarity between the any scatter point and the any other less-taken point, obtaining the Euclidean distance between the any scatter point and the any other scatter point, and obtaining the correlation coefficient between the any scatter point and the any other scatter point according to the mean value between the reciprocal of the Euclidean distance and the state similarity; According to all the correlation coefficients corresponding to any one of the scatter points, a correlation coefficient threshold corresponding to any one of the scatter points is obtained, other scatter points corresponding to correlation coefficients greater than or equal to the correlation coefficient threshold are used as reference scatter points for any one of the scatter points, and the maximum Euclidean distance between any one of the scatter points and each of the reference scatter points is obtained as the relative neighborhood radius of any one of the scatter points.
6. The gas meter data intelligent collection system based on machine learning according to claim 5 is characterized in that: The acquiring, according to all the correlation coefficients corresponding to the any scatter point, a correlation coefficient threshold corresponding to the any scatter point, comprises: A statistical histogram of all correlation coefficients is constructed, wherein the horizontal axis of the statistical histogram is the correlation coefficient, and the vertical axis is the number of correlation coefficients. A maximum correlation coefficient interval is obtained in the correlation coefficient interval in which the number of statistical correlation coefficients in the statistical histogram is continuously 0, and an intermediate value in the maximum correlation coefficient interval is used as the correlation coefficient threshold corresponding to any scatter point.
7. The gas meter data intelligent collection system based on machine learning according to claim 4 is characterized in that: The step of selecting the best relative neighborhood radius from among all relative neighborhood radiuses as the best clustering radius in the clustering algorithm includes: Obtain a radius interval consisting of all relative neighborhood radii, divide the radius interval into a preset number of sub-intervals, count the number of scattered points corresponding to the relative neighborhood radius contained in each sub-interval, and use the proportion of the number of scattered points corresponding to each sub-interval in the total number of scattered points as the frequency of each relative neighborhood radius contained in the corresponding sub-interval; For any relative neighborhood radius, a circular area is constructed with the scattered points corresponding to each other relative neighborhood radius as the center of the circle and the other relative neighborhood radius as the radius of the circle; the number of circular areas containing the scattered points corresponding to any relative neighborhood radius is counted as the number of occurrences of the scattered points of any relative neighborhood radius; and the preferred evaluation value of any relative neighborhood radius is obtained according to the product between the number of occurrences of the scattered points of any relative neighborhood radius and the frequency; The preferred evaluation value of each relative neighborhood radius is obtained, and the relative neighborhood radius corresponding to the largest preferred evaluation value is taken as the optimal relative neighborhood radius.
8. The gas meter data intelligent collection system based on machine learning according to claim 4 is characterized in that: The step of obtaining the minimum number of points in the clustering algorithm according to the distribution characteristics of the central scattered points of all scattered point areas includes: Performing straight line fitting on all central scatter points in the scatter plot to obtain a fitted straight line and a corresponding straight line slope, obtaining a fitted value of the gas usage activity of each central scatter point according to the fitted straight line, calculating the difference between the gas usage activity of each central scatter point and the corresponding fitting value to obtain a mean of the differences, and using the inverse of the mean of the differences as a distribution uniformity indicator of the central scatter points; If the distribution uniformity index of the central scatter point is greater than or equal to the preset distribution uniformity index threshold, the number of scatter points in each of the scatter point areas is counted to obtain the maximum number of scatter points, and the standard deviation analysis is performed on the number of scatter points greater than the preset scatter point number threshold to obtain the standard deviation of the number of scatter points; The absolute value of the difference between the reciprocal of the standard deviation of the number of scattered points and the constant 1 is calculated to obtain the product of the absolute value of the difference and the maximum number of scattered points, and the product is used as the point adjustment amount. The difference between the maximum number of scattered points and the point adjustment amount is rounded down to obtain the minimum number of points in the clustering algorithm.
9. The gas meter data intelligent collection system based on machine learning according to claim 8 is characterized in that: The step of obtaining the minimum number of points in the clustering algorithm according to the distribution characteristics of the central scattered points of all scattered point areas further includes: If the distribution uniformity index of the central scatter point is less than the preset distribution uniformity index threshold, the probability of each scatter point in the scatter plot being an isolated point is obtained, at least one suspected isolated point is screened, the circular area of each suspected isolated point is obtained, the number of suspected isolated points in the circular area of each suspected isolated point is counted, and the maximum number of suspected isolated points and the mean number of suspected isolated points are obtained; Calculate the difference between the constant 1 and the reciprocal of the mean of the number of suspected isolated points, obtain the product of the difference and the total number of suspected isolated points in the scatter plot, use the product as the point adjustment amount, round up the sum of the maximum number of suspected isolated points and the point adjustment amount, and obtain the minimum number of points in the clustering algorithm.
10. The gas meter data intelligent collection system based on machine learning according to claim 4 is characterized in that: The adaptively acquiring the gas data sampling frequency of the target user according to the gas usage activity of the gas data in each of the clusters includes: The cluster containing the most gas data is obtained as the target cluster, and the mean gas usage activity is obtained according to the gas usage activity corresponding to each gas data in the target cluster. The diversity index is obtained according to the difference between a constant 1 and the reciprocal of the number of clusters, and the sum of the mean gas usage activity and the diversity index is used as the gas usage activity state representation value of the target user in the current period; Obtain a historical gas usage activity status representation value of the target user in the previous cycle, and according to the gas usage activity of each gas data in the gas data sequence, use the difference between the number of gas data with a gas usage activity of 0 and the number of gas data with a gas usage activity of not 0 as a sampling frequency modification value; If the gas usage active state characterization value is greater than or equal to the historical gas usage active state characterization value, the difference between the constant 1 and the mean value of the gas usage activity is calculated, the average value between the difference and the inverse of the number of clusters is calculated, the multiplication value between the average value and the sampling frequency modification value is obtained, and the sampling frequency in the current period is added to the multiplication value to obtain the gas data sampling frequency of the target user; If the gas usage active state characterization value is less than the historical gas usage active state characterization value, then the difference between the constant 1 and the mean of the gas usage activity is used to calculate the average value between the difference and the inverse of the number of clusters, and the multiplication value between the average value and the sampling frequency modification value is obtained. The sampling frequency in the current period is subtracted from the multiplication value to obtain the gas data sampling frequency of the target user.
Citation Information
Patent Citations
Gas pipe network preset management method for intelligent gas and Internet of Things system
CN115456315A
Intelligent gas safety management method based on user activeness and Internet of Things system
CN117350680A
Building energy consumption monitoring method and system
CN117421618A
Artificial intelligence energy-saving management method and system based on big data
CN117493921A
Multi-level treatment method for real-time data of new energy station power generation equipment
CN118171051A