A method and system for cleaning power data based on big data
By constructing a KD tree and analyzing the fluctuation and correlation of power data, screening out the dimensions to be processed, solving the problem of dimensional disasters in power data cleaning, and achieving efficient power data cleaning and processing.
Patent Information
- Application Number
- CN202510134340.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-07
AI Technical Summary
The prior art has dimensional disasters when processing power data, and the filtering efficiency is low, resulting in low accuracy in cleaning power data.
By constructing a KD tree, the correlation between the fluctuations and dimensions of data points in each dimension are analyzed, the noise level and importance are calculated, the dimensions to be processed are selected, and the power data cleaning is achieved through denoising and interpolation methods.
It effectively improves the efficiency of power data cleaning, reduces noise interference, and improves the accuracy and completeness of data processing.
Smart Images

Figure CN119577350B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power data processing, and in particular to a power data cleaning method and system based on big data. Background Art
[0002] The power system is the core of energy supply in modern society. It converts various energy sources into electrical energy to provide the required power for various equipment and facilities, thereby ensuring the normal operation of society. Power data is an important basis for monitoring the operation status of the power grid. By analyzing power data, energy distribution can be optimized, energy waste can be reduced, and energy utilization efficiency can be improved.
[0003] Power data is crucial for monitoring the status of the power grid, preventing failures and quickly responding to emergencies. There are many applications of power data in the existing technology. For example, the patent application document with publication number CN118798728A discloses a business indicator analysis method and application system based on power data. The application obtains the power data of the target enterprise within a preset analysis period through a power distribution map, and constructs a power distribution cluster including multiple business areas of the enterprise; after building a data sandbox according to the business area of the enterprise, a digital business report of the enterprise is generated through the power distribution cluster, and the economic status of the enterprise is judged according to the power data of the enterprise.
[0004] However, the amount of power data generated by the power system is very large, with characteristics of high dimensionality, time series, and noise. Traditional filtering methods usually have the problem of dimensionality curse when processing power data, and the filtering efficiency is low, resulting in the enterprise economic status obtained in the above-mentioned prior art being interfered by noise and having low accuracy.
[0005] Based on this, how to effectively improve the efficiency of power data cleaning is an urgent problem to be solved by technical personnel in this field. Summary of the invention
[0006] In order to solve the technical problem of how to effectively improve the efficiency of power data cleaning, the present invention provides a power data cleaning method and system based on big data.
[0007] In a first aspect, the present invention provides a power data cleaning method based on big data, which adopts the following technical solution:
[0008] A power data cleaning method based on big data, comprising the steps of:
[0009] Obtain data points in the power data sequence of each dimension, and obtain the data fluctuation degree of each dimension through the numerical difference between the data points in each dimension and the adjacent data points on the left and right sides; obtain the difference in data fluctuation degree between a dimension and other dimensions to form the fluctuation vector of the dimension, and record the modulus of the fluctuation vector of the dimension as the noise degree of the dimension; record the cumulative sum of the deviations of the normalized values of the data between the dimension and another dimension at each acquisition moment as the difference between the dimension and the other dimension, obtain the number of data values with a covariance greater than 1 between the dimension and the other dimension as the first index of the dimension, record the variance of the difference between the dimension and the other dimension as the second index, and determine the importance of the dimension by the ratio of the first index to the second index; calculate the possibility of each dimension to be processed: ; In the formula, Indicates The possibilities to be processed in dimensions, Indicates The noise level in each dimension, Indicates The importance of the dimension, represents an exponential function with e as the base; the dimension to be processed is obtained by comparing the possibility to be processed of each dimension with a preset threshold, and the dimension to be processed is denoised to achieve power data cleaning.
[0010] The present invention can realize the cleaning of multi-dimensional power data by constructing a KD tree. In this process, the present invention takes into account that the noise performance and importance of different dimensions are different, and the necessity of processing the dimensions with higher noise performance and importance is higher. Based on this, the present invention accurately obtains the noise level of each dimension by analyzing the fluctuation degree of data points in each dimension and the correlation of the fluctuation degree between dimensions, and then accurately obtains the importance of each dimension through the difference and joint change degree of data between dimensions. The combination of noise level and importance can accurately screen out the dimensions to be cleaned, effectively improving the efficiency of power data cleaning.
[0011] According to a power data cleaning method based on big data provided by the present invention, the data points in the power data sequence of each dimension are obtained, including: presetting a single collection time, taking the data value of each dimension obtained at each collection moment as a data point, and forming the data points obtained within the single collection time into a power data sequence of each dimension.
[0012] According to a power data cleaning method based on big data provided by the present invention, the data fluctuation degree of each dimension is obtained, including: respectively obtaining the absolute value of the difference between a data point in the power data sequence under each dimension and the adjacent data points on the left and right sides, recorded as the first difference and the second difference, and recording the ratio of the sum of the first difference and the second difference to the data value of the data point as the relative fluctuation degree of the data point; recording the ratio of the average absolute deviation of the relative fluctuation degree of each data point in the power data sequence under the dimension to the extreme value as the data fluctuation degree of the dimension.
[0013] The present invention can accurately obtain the relative fluctuation degree of each data point by obtaining the deviation between the data point in each dimension and the adjacent data on the left and right sides, and then accurately obtain the data fluctuation situation of each dimension by analyzing the change of the relative fluctuation degree of the data point in each dimension.
[0014] According to a power data cleaning method based on big data provided by the present invention, the difference between the dimension and another dimension satisfies the relationship:
[0015] ;
[0016] In the formula, Indicates The difference between the jth dimension and the jth dimension, Indicates The number of data points in the power data series under the dimension, Indicates Dimension The data value of the data point, Indicates the jth dimension The data value of the data point, represents the linear normalization function.
[0017] According to a power data cleaning method based on big data provided by the present invention, the dimension to be processed is obtained by comparing the possibility of being processed of each dimension with a preset threshold, including: a preset threshold; if the possibility of being processed of a dimension is greater than the preset threshold, then the dimension is the dimension to be processed; if the possibility of being processed of the dimension is not greater than the preset threshold, then the dimension is the dimension not to be processed.
[0018] The present invention can accurately screen out dimensions to be processed from all dimensions by acquiring the possibility of processing in each dimension, thus preparing for subsequent data cleaning.
[0019] According to a power data cleaning method based on big data provided by the present invention, the dimension to be processed is denoised to achieve power data cleaning, including: forming a data set with data points in all the dimensions to be processed; and constructing a KD tree based on the data points in the data set; obtaining the distance between each data point and the nearest adjacent data point in the KD tree, and removing the data points whose distance is greater than a preset distance threshold as noise points to achieve power data cleaning.
[0020] The present invention can effectively improve the efficiency of dimensional data acquisition and cleaning by constructing a KD tree to remove noise points, thereby effectively improving the overall quality of power data.
[0021] According to a power data cleaning method based on big data provided by the present invention, the power data cleaning is implemented, and then further includes: interpolating at the position of the noise point by an interpolation method.
[0022] The present invention takes into account that the missing position of the noise point after removal will affect the subsequent data processing, and therefore interpolation is used to supplement it, thereby effectively improving the integrity of the power data.
[0023] In a second aspect, the present invention provides a power data cleaning system based on big data, which adopts the following technical solution:
[0024] A power data cleaning system based on big data comprises: a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the power data cleaning method based on big data is implemented.
[0025] By adopting the above technical solution, the above-mentioned power data cleaning method based on big data is generated into a computer program and stored in a memory to be loaded and executed by a processor, so that a terminal device is made according to the memory and the processor for easy use.
[0026] The present invention has the following technical effects:
[0027] Based on the above technical scheme, when the present invention realizes the cleaning of power data, it can realize the cleaning of multi-dimensional power data by constructing a KD tree. In this process, the present invention accurately obtains the noise degree of each dimension by analyzing the fluctuation degree of data points in each dimension and the correlation of the fluctuation degree between dimensions, and then accurately obtains the importance of each dimension through the difference and joint change degree of data between dimensions. Combining the noise degree and the importance degree, the dimensions to be cleaned can be accurately screened out, avoiding the processing of dimensional data with low noise degree and importance, and effectively improving the efficiency of power data cleaning. BRIEF DESCRIPTION OF THE DRAWINGS
[0028] By reading the detailed description below with reference to the accompanying drawings, the above and other purposes, features and advantages of the exemplary embodiments of the present invention will become readily understood. In the accompanying drawings, several embodiments of the present invention are shown in an exemplary and non-restrictive manner, and the same or corresponding reference numerals represent the same or corresponding parts.
[0029] Figure 1 A schematic flow chart of a method for cleaning power data based on big data provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0030] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0031] It should be understood that when the terms "first", "second", etc. are used in the claims, descriptions, and drawings of the present invention, they are only used to distinguish different objects, rather than to describe a specific order. The terms "include" and "comprise" used in the description and claims of the present invention indicate the presence of the described features, wholes, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their collections.
[0032] It should be noted that the amount of power data generated by the power system is very large, with characteristics such as high dimensionality, time series, and noise. Traditional filtering methods usually have the problem of dimensionality disaster when processing power data, and the filtering efficiency is low, resulting in noise interference when the power data is used later, affecting the accuracy of data processing. Therefore, it is necessary to clean the power data.
[0033] K-Dimensional Tree (KD tree) is a tree structure used to store and retrieve multidimensional data. In the process of power data cleaning, K-dimensional tree can effectively organize multidimensional data and support efficient neighborhood query, so as to quickly identify noise points in power data that are significantly different from normal data distribution.
[0034] However, power data is multidimensional data. The K-dimensional tree needs to continuously traverse dimensions when processing power data. The more dimensions there are, the longer the time required. In addition, the noise performance and source of each dimension may be different, and the importance of each dimension is also different. If the power data of all dimensions are cleaned in the same way, the highly variable dimensional data may be misprocessed, resulting in low data cleaning efficiency.
[0035] Based on this, the embodiment of the present invention discloses a power data cleaning method based on big data, which selects the dimensions to be processed by obtaining the noise level and importance of the power data in each dimension, constructs a K-dimensional tree for the dimensions to be processed and identifies noise points to achieve power data cleaning. Figure 1 As shown, Figure 1 A flow chart of a method for cleaning power data based on big data is provided in accordance with an embodiment of the present invention. The method comprises the following steps S1 to S4.
[0036] S1: Obtain data points in the power data series of each dimension.
[0037] For example, in an embodiment of the present invention, data points in a power data sequence of each dimension are obtained, including: presetting a single collection time, taking the data value of each dimension obtained at each collection moment as a data point, and forming the data points obtained within the single collection time into a power data sequence of each dimension.
[0038] Among them, the single collection time can be preset to 30 minutes; the single collection time can be set according to actual needs, and the embodiment of the present invention does not impose too many restrictions on this.
[0039] For example, the categories of dimensions may be voltage, current, power, load, etc., and may be specifically set according to actual needs.
[0040] Specifically, by presetting a single collection time, data points of each dimension can be obtained through fixed frequency collection within the single collection time. At the same collection moment, the data points in the power data sequence under each dimension correspond one by one, and the number of data points in the power data sequence under each dimension finally obtained is the same.
[0041] After obtaining the power data sequences of each dimension based on the above steps, continue to perform the following steps.
[0042] S2: The data fluctuation degree of each dimension is obtained through the numerical difference between the data points in each dimension and the adjacent data points on the left and right sides; the difference in the data fluctuation degree between a dimension and other dimensions is obtained to form the fluctuation vector of the dimension, and the modulus of the fluctuation vector of the dimension is recorded as the noise degree of the dimension.
[0043] It should be noted that when processing multidimensional power data, traditional filtering methods suffer from the dimensionality curse, and the sparsity and complexity of the data cause the effect of the filtering algorithm to be significantly reduced. K-dimensional trees can be used to store and retrieve tree structures of multidimensional data, supporting efficient neighborhood queries. However, when cleaning power data, the source and changes of noise in each dimension may be different. If data of all dimensions are cleaned, the data processing efficiency will be low.
[0044] Based on this, the embodiment of the present invention obtains the noise degree of each dimension by analyzing the data fluctuation in each dimension. The higher the fluctuation degree, the greater the possibility that the dimension is affected by noise, so that the dimensional data with a higher noise degree can be screened out for data cleaning.
[0045] By way of example, in an embodiment of the present invention, the degree of data fluctuation in each dimension is obtained, including: respectively obtaining the absolute value of the difference between a data point in the power data sequence in each dimension and the adjacent data points on the left and right sides, recorded as the first difference and the second difference, and recording the ratio of the sum of the first difference and the second difference to the data value of the data point as the relative fluctuation degree of the data point; recording the ratio of the average absolute deviation of the relative fluctuation degree of each data point in the power data sequence in the dimension to the extreme value as the data fluctuation degree of the dimension.
[0046] For example, the relative fluctuation degree of the data points in each dimension is determined, and the specific relationship can be seen in the following formula:
[0047] ;
[0048] In the formula, Indicates The relative fluctuation degree of the i-th data point in the dimension, Indicates The data value of the i-th data point in the dimension, Indicates In the dimension The data value of the data point, Indicates In the dimension The data value of the data point, Represents the absolute value symbol.
[0049] In the above formula, represents the first difference, Represents the second difference. The first difference and the second difference reflect the difference between the current data point and the adjacent data points on the left and right sides. The larger the first difference and the second difference, the greater the difference between the current data point and the adjacent data points on the left and right sides. The larger the ratio of the difference to the data value of the current data point, the greater the relative fluctuation of the corresponding current data point.
[0050] After the relative fluctuation degree of each data point in the dimension is obtained based on the above steps, the data fluctuation degree of each dimension can be obtained based on the relative fluctuation degree of the data points in each dimension.
[0051] For example, to determine the degree of data fluctuation in each dimension, please refer to the following relationship:
[0052] ;
[0053] In the formula, Indicates The degree of data fluctuation in each dimension, Indicates The number of data points in the power data series under the dimension, Indicates The relative fluctuation degree of the i-th data point in the dimension, Indicates The mean relative fluctuation of all data points in the dimension, Indicates The maximum data value in the dimension, Indicates The minimum data value in the dimension, Represents the absolute value symbol.
[0054] In the above formula, Indicates The average absolute deviation of the relative fluctuation degree of each data point in the power data series under the dimension. The larger the value, the The greater the data fluctuation in the dimension, the smaller the value is. The smaller the data fluctuation in each dimension.
[0055] Indicates The range of the data values in a dimension, that is, the range of data variation in that dimension. The smaller the range, the smaller the overall difference in the data in that dimension, the smaller the impact of noise on the data, and the more stable and reliable the data fluctuation will be.
[0056] It should be further explained that, based on the above steps, the degree of data fluctuation of each dimension can be obtained. However, the noise performance of the data may not only be manifested in the data fluctuation itself, but also in the data of other dimensions. If the degree of fluctuation of one dimension is highly correlated with the degree of fluctuation of other dimensions, it means that the data of this dimension is relatively stable and the noise performance is smaller; conversely, if the degree of fluctuation of one dimension is less correlated with the degree of fluctuation of other dimensions, it means that the data of this dimension is unstable and the noise performance is greater.
[0057] Based on this, the embodiment of the present invention can also obtain the correlation between the data fluctuation degree of the current dimension and other dimensions, and correct the data fluctuation degree of the current dimension through the correlation.
[0058] For example, when correcting the data fluctuation degree of the current dimension through the correlation between the data fluctuation degree of the current dimension and other dimensions, the difference in the data fluctuation degree between a dimension and other dimensions can also be obtained to obtain the correlation between the dimension and other dimensions; the data fluctuation degree of the dimension is corrected through the modulus of the vector composed of the correlations between the dimension and all other dimensions to obtain the noise degree of the dimension.
[0059] Among them, the larger the difference in data fluctuation degree, the smaller the correlation between the fluctuation degree of this dimension and other dimensions, and the higher the corresponding noise degree of this dimension; conversely, the smaller the difference in data fluctuation degree, the greater the correlation between the fluctuation degree of this dimension and other dimensions, and the lower the corresponding noise degree of this dimension.
[0060] Based on the above steps, the correlation between the current dimension and other dimensions is obtained to correct the data fluctuation degree of the current dimension, so as to accurately obtain the noise degree of the current dimension and continue to execute the following steps.
[0061] S3: The cumulative sum of the deviations of the normalized values of the data between the dimension and another dimension at each collection moment is recorded as the difference between the dimension and the other dimension, the number of data values whose covariance between the dimension and the other dimension is greater than 1 is obtained and recorded as the first indicator of the dimension, the variance of the difference between the dimension and the other dimension is recorded as the second indicator, and the importance of the dimension is determined by the ratio of the first indicator to the second indicator.
[0062] It should be noted that the importance of power data in different dimensions is different. When obtaining dimensional data that needs to be cleaned, it is also necessary to determine the importance of each dimension.
[0063] It should be further explained that if the data changes of one dimension are relatively synchronized with those of other dimensions, its importance is higher; conversely, if there is a large difference between the data changes of one dimension and those of other dimensions, the impact of this dimension on the overall power data is more complex or unstable, the impact is smaller, and the importance is lower.
[0064] Based on this, the embodiment of the present invention obtains the importance of the dimension by acquiring the difference and joint change degree between the dimensional data.
[0065] For example, in an embodiment of the present invention, the difference between a dimension and another dimension is determined, and specifically, see the following relationship:
[0066] ;
[0067] In the formula, Indicates The difference between the jth dimension and the jth dimension, Indicates The number of data points in the power data series under the dimension, Indicates Dimension The data value of the data point, Indicates the jth dimension The data value of the data point, represents the linear normalization function.
[0068] In the above formula, Indicates The jth dimension is in the The deviation of the data value at the data point, the larger the value, the The jth dimension is in the The greater the difference between the data points.
[0069] By analyzing the data differences at each data point based on the above formula, we can get the overall difference between the data of the current dimension and the data of another dimension. Combined with the degree of joint change between the current dimension and other dimensions, we can accurately get the importance of the current dimension.
[0070] For example, the importance of each dimension can be determined by referring to the following relationship:
[0071] ;
[0072] In the formula, Indicates The importance of the dimension, Indicates The number of data values whose covariance between the dimension and other dimensions is greater than 1, Indicates the number of dimensions, Indicates The difference between the jth dimension and the jth dimension, Represents the mean of the differences between all dimensions.
[0073] In the above formula, covariance can be used to describe the The larger the covariance, the greater the correlation between the data in the first dimension and the jth dimension during the change process. The larger the covariance, the greater the correlation between the data in the two dimensions. When the data in this dimension changes, the possibility of causing changes in the data in the jth dimension is higher. Indicates the first indicator. The larger the value, the closer it is to the first indicator. The more dimensions that change jointly, the greater the number of dimensions.
[0074] Indicates The variance of the difference between the first dimension and other dimensions is the second indicator. The larger the value, the stronger the The greater the deviation between a dimension and other dimensions, the lower its importance.
[0075] After obtaining the noise level and importance of each dimension based on the above steps, continue to perform the following steps.
[0076] S4: Calculate the possibility of each dimension to be processed, obtain the dimension to be processed by comparing the possibility of each dimension to be processed with a preset threshold, and denoise the dimension to be processed to achieve power data cleaning.
[0077] It should be noted that based on the above steps, the noise degree and importance of each dimension in the power data can be obtained. The higher the noise degree of each dimension, the higher the necessity of cleaning the data of this dimension, and the higher the possibility that this dimension is the dimension to be processed; the higher the importance of each dimension, the higher the necessity of ensuring the accuracy of the power data of this dimension, and the higher the possibility that this dimension is the dimension to be processed.
[0078] Based on this, the embodiment of the present invention can accurately screen out the dimensions that need to be cleaned by combining the noise level and importance of each dimension.
[0079] For example, in the embodiment of the present invention, the possibility to be processed in each dimension is determined, and the details can be referred to the following relationship:
[0080] ;
[0081] In the formula, Indicates The possibilities to be processed in dimensions, Indicates The noise level in each dimension, Indicates The importance of the dimension, Represents an exponential function with base e.
[0082] After the possibility of processing of each dimension is obtained based on the above formula, the dimension to be processed can be obtained based on the comparison result of the possibility of processing of each dimension and the preset threshold.
[0083] By way of example, in an embodiment of the present invention, the dimension to be processed is obtained by comparing the possibility of each dimension to be processed with a preset threshold, including: a preset threshold; if the possibility of the dimension to be processed is greater than the preset threshold, the dimension is the dimension to be processed; if the possibility of the dimension to be processed is not greater than the preset threshold, the dimension is an unprocessed dimension.
[0084] Among them, the preset threshold can be the average value of the possibilities to be processed in all dimensions; the preset threshold can be set according to actual needs, and the embodiment of the present invention does not impose too many restrictions on this.
[0085] It can be understood that the higher the possibility of pending processing in each dimension, the higher the necessity for data cleaning; conversely, the lower the possibility of pending processing in each dimension, the lower the necessity for data cleaning.
[0086] After obtaining the dimension to be processed based on the above steps, the data points in the dimension to be processed can be cleaned and denoised by constructing a KD tree.
[0087] By way of example, in an embodiment of the present invention, the dimensions to be processed are denoised to achieve power data cleaning, including: forming a data set with data points in all the dimensions to be processed; and constructing a KD tree based on the data points in the data set; obtaining the distance between each data point and the nearest adjacent data point in the KD tree, and removing data points whose distance is greater than a preset distance threshold as noise points to achieve power data cleaning.
[0088] The distance threshold may be preset to 0.8, and may be specifically set according to actual needs.
[0089] Specifically, when constructing a KD tree based on data points in a data set, the dimension with the largest variance can be selected as the dimension for segmentation, the median of all data points on the selected dimension can be found, and the median point can be selected as the current segmentation point; the data set can be segmented into two parts on the selected dimension using the segmentation point, and all data points on the left side of the segmentation point will be located in the left subtree, and data points on the right side will be located in the right subtree. Similarly, for the left subtree and the right subtree, a segmentation dimension and a segmentation point are determined in the same way, and subtrees are recursively constructed. In response to the termination condition, a KD tree is finally constructed.
[0090] The termination condition may be set according to actual needs, and the embodiment of the present invention does not impose too many limitations on this.
[0091] For example, the distance between a data point and its nearest adjacent data point may be Euclidean distance, Mahalanobis distance, etc., which may be set according to actual needs, and the embodiment of the present invention does not impose too many limitations thereon.
[0092] Specifically, by taking the median of the data points in the split dimension as the split point, the data set can be first divided into a left subset and a right subset, wherein the data points in the data set that are smaller than the split point are divided into the left subset, and the data points in the data set that are not smaller than the split point are divided into the right subset; repeat the above steps to continue to split the left subset and the right subset respectively, and finally construct a KD tree.
[0093] It is understandable that if the distance between a data point and its nearest adjacent data point is greater than a preset distance threshold, it means that the data point is a noise point, and power data cleaning can be achieved by removing the noise point.
[0094] For example, in an embodiment of the present invention, power data cleaning is implemented, and then the method further includes: interpolating at the position where the noise point is located by using an interpolation method.
[0095] In this way, the embodiment of the present invention performs interpolation at the position of the removed noise point through the interpolation method, which can ensure the integrity of the power data and facilitate subsequent processing.
[0096] It can be seen that in an embodiment of the present invention, when realizing power data cleaning, data points in the power data sequence of each dimension can be obtained, and the data fluctuation degree of each dimension can be obtained by the numerical difference between the data points in each dimension and the adjacent data points on the left and right sides; the difference in the degree of data fluctuation between one dimension and other dimensions is obtained to form the fluctuation vector of the dimension, and the modulus of the fluctuation vector of the dimension is recorded as the noise degree of the dimension; the cumulative sum of the deviations of the normalized values of the data between the dimension and another dimension at each acquisition moment is recorded as the difference between the dimension and the other dimension, the number of data values with a covariance greater than 1 between the dimension and the other dimension is obtained and recorded as the first indicator of the dimension, the variance of the difference between the dimension and the other dimension is recorded as the second indicator, and the importance of the dimension is determined by the ratio of the first indicator to the second indicator; the possibility of processing of each dimension is calculated; the dimension to be processed is obtained by comparing the possibility of processing of each dimension with the preset threshold, and the dimension to be processed is denoised to realize power data cleaning.
[0097] In this way, the embodiment of the present invention can realize the cleaning of multi-dimensional power data by constructing a KD tree. In this process, the embodiment of the present invention takes into account that the noise performance and importance of data in different dimensions are different, and the dimensions with higher noise performance and importance need to be processed more frequently. Based on this, the embodiment of the present invention accurately obtains the noise level of each dimension by analyzing the fluctuation level of data points in each dimension and the correlation between the fluctuation levels of the dimensions, and then accurately obtains the importance of each dimension through the difference and joint change degree of the data between the dimensions. The combination of the noise level and the importance can accurately screen out the dimensions to be cleaned, effectively improving the efficiency of power data cleaning.
[0098] An embodiment of the present invention also discloses a power data cleaning system based on big data, including a processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a power data cleaning method based on big data provided by the present invention is implemented.
[0099] The above system also includes other components well known to those skilled in the art, such as a communication bus and a communication interface, and their configuration and functions are known in the art, so they will not be described in detail here.
[0100] In the present invention, the aforementioned memory may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium may be any suitable magnetic storage medium or magneto-optical storage medium, such as a resistive random access memory RRAM, a dynamic random access memory DRAM, a static random access memory SRAM, an enhanced dynamic random access memory EDRAM, a high bandwidth memory HBM, a hybrid memory cube HMC, etc., or any other medium that can be used to store the required information and can be accessed by an application, a module, or both. Any such computer storage medium may be part of a device or accessible or connectable to a device.
[0101] Although this specification has shown and described a number of embodiments of the present invention, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art will conceive of many modifications, changes and alternatives without departing from the ideas and spirit of the present invention. It should be understood that in the practice of the present invention, various alternatives to the embodiments of the present invention described herein may be employed.
[0102] The above are all preferred embodiments of the present invention, and are not intended to limit the protection scope of the present invention. Therefore, any equivalent changes made based on the structure, shape, and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A power data cleaning method based on big data, characterized in that: include: Obtain data points in the power data sequence of each dimension, and obtain the data fluctuation degree of each dimension through the numerical difference between the data point in each dimension and the adjacent data points on the left and right sides; Obtain the data fluctuation degree difference between one dimension and other dimensions to form the fluctuation vector of the dimension, and record the modulus of the fluctuation vector of the dimension as the noise degree of the dimension; The cumulative sum of the deviations of the normalized values of the data between the dimension and another dimension at each acquisition moment is recorded as the difference between the dimension and the other dimension. The number of data values between the dimension and the other dimension whose covariance is greater than 1 is recorded as the first index of the dimension. The variance of the difference between the dimension and the other dimension is recorded as the second index. The importance of the dimension is determined by the ratio of the first index to the second index. Calculate the likelihood of being processed in each dimension: ; In the formula, Indicates The possibilities to be processed in dimensions, Indicates The noise level in each dimension, Indicates The importance of the dimension, represents an exponential function with base e; The dimensions to be processed are obtained by comparing the possibilities to be processed of each dimension with the preset threshold, and the dimensions to be processed are denoised to achieve power data cleaning.
2. The method for cleaning power data based on big data according to claim 1, characterized in that: The step of obtaining data points in the power data sequence of each dimension includes: A single collection time is preset, and the data value of each dimension obtained at each collection moment is taken as a data point. The data points obtained within the single collection time are combined into a power data sequence of each dimension.
3. The method for cleaning power data based on big data according to claim 1, characterized in that: The data fluctuation degree of each dimension is obtained, including: The absolute values of the differences between a data point in the power data sequence under each dimension and the adjacent data points on the left and right sides are obtained respectively, recorded as the first difference and the second difference, and the ratio of the sum of the first difference and the second difference to the data value of the data point is recorded as the relative fluctuation degree of the data point; the ratio of the average absolute deviation of the relative fluctuation degree of each data point in the power data sequence under the dimension to the extreme value is recorded as the data fluctuation degree of the dimension.
4. The method for cleaning power data based on big data according to claim 1, characterized in that: The difference between the dimension and another dimension satisfies the relationship: ; In the formula, Indicates The difference between the jth dimension and the jth dimension, Indicates The number of data points in the power data series under the dimension, Indicates Dimension The data value of the data point, Indicates the jth dimension The data value of the data point, represents the linear normalization function.
5. The method for cleaning power data based on big data according to claim 1, characterized in that: The step of obtaining the dimension to be processed by comparing the possibility of processing of each dimension with a preset threshold value includes: A preset threshold; if the possibility of a dimension to be processed is greater than the preset threshold, the dimension is a dimension to be processed; If the possibility of a dimension to be processed is not greater than a preset threshold, the dimension is not processed.
6. The method for cleaning power data based on big data according to claim 5, characterized in that: Denoising the dimension to be processed to achieve power data cleaning includes: The data points in all dimensions to be processed are combined into a data set; and a KD tree is constructed based on the data points in the data set; The distance between each data point and its nearest adjacent data point is obtained in the KD tree, and data points with a distance greater than a preset distance threshold are removed as noise points to achieve power data cleaning.
7. The method for cleaning power data based on big data according to claim 6, characterized in that: The power data cleaning is then implemented, and further includes: Interpolation is performed at the location of the noise point using the interpolation method.
8. A power data cleaning system based on big data, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, a power data cleaning method based on big data according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Enterprise operation index analysis method and application system based on electric power data
CN118798728A
Electrical equipment cleaning method, device and system
CN116020819A
Environmental data monitoring method based on multi-dimensional data analysis
CN117216484A