Data cleaning method and system and data center platform

By analyzing the parameter correlation and relative difference degree of IoT data, combined with clustering and parameter matrix construction, the problem that traditional data cleaning methods are difficult to accurately process abnormal data is solved, and higher quality data cleaning and analysis are achieved.

CN120234536AActive Publication Date: 2025-07-01QINGDAO KLEIMA IOT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510429204.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-01
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

Due to the wide range of sources, complex formats and high real-time requirements of IoT data, traditional data cleaning methods are difficult to accurately identify and process abnormal data, affecting the accuracy and reliability of data analysis.

Method used

By obtaining the information data sets of different users, analyzing the local fluctuation correlation between different parameters corresponding to each category, calculating the parameter correlation degree, determining the relative difference between users, clustering, building a parameter matrix, calculating the abnormality, and cleaning the data.

Benefits of technology

This method can more accurately identify and process abnormal data, improve the quality of data cleaning, and improve the accuracy of data analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234536A_ABST
    Figure CN120234536A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a data cleaning method and system and a data center platform, and the method comprises the steps: obtaining information data sets of different users; calculating the parameter correlation degree of each category in the information data set; determining the relative difference degree of any two users; all the users are clustered; constructing a parameter matrix corresponding to various parameters in each cluster; and determining the anomaly degree of each column in the parameter matrix, and cleaning data in the parameter matrix corresponding to various parameters in all the clustering clusters to obtain a cleaned information data set of different users. The abnormal data can be recognized more accurately, the data cleaning quality is improved, and the accuracy of data mining analysis is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a data cleaning method, system, and data center platform. Background Art

[0002] In the context of the rapid development of Internet of Things technology, advanced Internet of Things sensing devices are widely used in the collection of various types of data. These rich data resources provide a solid data foundation and strong support for optimizing product performance, improving service quality, and expanding business areas.

[0003] However, Internet of Things data has the characteristics of wide sources, complex formats, and high real-time requirements, making data processing work face many severe challenges. In the original data, a large number of data quality problems will affect the accuracy and reliability of data analysis results. Traditional data cleaning methods usually perform anomaly processing and analysis on the data of one parameter of a single user. This method ignores the anomalies between different users with similar data characteristics, making it difficult for data cleaning work to accurately identify and process abnormal data, resulting in low data cleaning quality and affecting in-depth data mining and analysis. Summary of the Invention

[0004] In order to solve the above technical problems, a data cleaning method, system, and data center platform are provided to solve the existing problems.

[0005] The solution of this application to solve the technical problem is to provide a data cleaning method, system, and data center platform, including the following steps: In the first aspect, an embodiment of this application provides a data cleaning method, which includes the following steps: Obtain information data sets of different users, and preprocess the information data sets, where the information data sets contain data of various parameters corresponding to each category at different times; Analyze the correlation of local fluctuations between different parameters corresponding to each category in the information data set, and calculate the parameter correlation degree of each category in the information data set; Determine the relative difference degree between any two users through the data difference change of the same parameter in the information data set between any two users, and combine the parameter correlation degree; cluster all users based on the relative difference degree; Based on the data of all users at different times under various parameters in each cluster, construct a parameter matrix corresponding to various parameters in each cluster; determine the abnormality degree of each column in the parameter matrix through the abnormality degree of each column element within each row of the parameter matrix and the data abnormality situation of the corresponding user in each row, and clean the data in the parameter matrix corresponding to various parameters in all clusters to obtain the cleaned information data set of different users.

[0006] Preferably, calculating the parameter correlation degree of each category in the information data set includes: Form the time parameter sequence with the data of various parameters in the information data set at all times; Adopt the moving standard deviation algorithm to obtain the moving standard deviation of each element in the time parameter sequence; Calculate the correlation degree of the moving standard deviations of all elements in the time parameter sequence between any two parameters corresponding to each category in the information data set, and take the average value of the correlation degrees of all any two parameters corresponding to each category as the parameter correlation degree of each category in the information data set.

[0007] Preferably, determining the relative difference degree between any two users includes: Analyze the difference situation of the parameter correlation degrees of all categories in the information data set between any two users to obtain the correlation difference between any two users; The th user and the th user The calculation formula of the relative difference degree is: where is the correlation difference between the th user and the th user; is the time parameter sequence corresponding to the th type of parameter in the information data set of the th user, is the time parameter sequence corresponding to the th type of parameter in the information data set of the th user,

[0008] Preferably, the correlation difference is the metric distance of the parameter correlation degrees of all categories in the information data set between any two users.

[0009] Preferably, the construction process of the parameter matrix is: The time parameter sequences corresponding to all users under various parameters in each cluster are combined into a parameter matrix, where each row element corresponds to a user.

[0010] Preferably, a further measurement method for the abnormality degree is: performing abnormality detection on all column elements in each row of the parameter matrix respectively to obtain an abnormality score of each column element in each row.

[0011] Preferably, determining the abnormality of each column in the parameter matrix includes: Calculate the weight corresponding to each row in the parameter matrix through multi-criteria decision analysis; The calculation formula of the abnormality of each column in the parameter matrix is: ,in, is the parameter matrix The abnormality of the column, is the parameter matrix Line anomaly score of column elements, is the parameter matrix The weight corresponding to the row, is the number of all rows in the parameter matrix.

[0012] Preferably, the cleaning of the data in the parameter matrix corresponding to various parameters in all clusters includes: A column of data corresponding to the abnormality degree greater than a preset threshold in the parameter matrix of various parameters in each cluster is taken as abnormal data, all abnormal data in the parameter matrix are removed, and missing values ​​are filled in the removed data.

[0013] In a second aspect, an embodiment of the present application further provides a data cleaning system, the system comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of any one of the above-mentioned data cleaning methods when executing the computer program.

[0014] In a third aspect, an embodiment of the present application further provides a data center platform, which performs visual analysis on the cleaned information data set.

[0015] This application has at least the following beneficial effects: This application analyzes the variation correlation between different parameters under each category in the information dataset of a single user, and calculates the parameter correlation degree of each category. The beneficial effect is that it takes into account the data fluctuation situation in different time periods to reflect the response situation of data changes between different parameters under each category. Secondly, by calculating the data difference change between any two users under the same parameter and the difference situation of the parameter correlation degree between any two users, the relative difference degree of the any two users is calculated, and all users are clustered. The beneficial effect is that it takes into account the data feature difference situation between two users, and then through clustering, users with similar data features are divided into one category, so as to perform abnormal analysis on the data changes of all users in the same category in the future. Analyze the data at different moments of all users under various parameters in each clustering cluster, construct a parameter matrix, analyze the abnormal change situation of the elements in each row and column of the parameter matrix, and calculate the abnormality degree of each column in the parameter matrix. The beneficial effect is that it takes into account the data abnormality situation of all users in the same clustering cluster under the same parameter, and avoids performing abnormal processing analysis on the data of a single parameter of a single user, resulting in errors in the identification and processing of abnormal data. Furthermore, through the abnormality degree, the abnormal data in the parameter matrix is removed and filled. The beneficial effect is that by considering the data change situation of multiple users in the same clustering cluster, the abnormal data of various parameters of each user is analyzed, which can more accurately identify abnormal data, improve the quality of data cleaning, and enhance the accuracy of data mining and analysis. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The following further elaborates on a data cleaning method of this application with reference to the accompanying drawings.

[0017] Figure 1 is a flowchart of the steps of a data cleaning method provided by an embodiment of this application; Figure 2 is a flowchart of the steps of a method for obtaining the parameter correlation degree of each category in the information dataset provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0018] In order to make the purpose, technical solution and advantages of this application clearer, the following further elaborates on a data cleaning method, system and data center platform proposed by this application with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs.

[0020] Please refer to Figure 1, which shows a flowchart of the steps of a data cleaning method provided by an embodiment of the present application. The method includes the following steps: Step 1, obtain the information data sets of different users and preprocess the information data sets.

[0021] Obtain the information data sets of different users; In this embodiment, by performing data cleaning on the provided energy usage data of different users, therefore, through smart meters, smart water meters, and smart gas meters, the power data, water resource data, and gas data of each user are respectively obtained. Among them, the power data includes data such as real-time power, cumulative power consumption, voltage, and current parameters; the water resource data includes instantaneous water flow and cumulative water consumption, etc., and the gas data includes instantaneous gas flow and cumulative gas consumption, etc. Finally, an information data set is formed, where the acquisition time interval is 15 minutes. As other implementation manners, the implementer can set it by himself according to the actual situation.

[0022] Therefore, the information data set of each user contains the data of various parameters corresponding to each category at different times. Among them, each category includes power data, water resource data, and gas data, and various parameters include real-time power, cumulative power consumption, voltage, current, instantaneous water flow, cumulative water consumption, instantaneous gas flow, and cumulative gas consumption.

[0023] Furthermore, perform preliminary preprocessing on the information data set, and perform missing value filling and data type conversion on the information data set; in this embodiment, the mean filling method is used for missing value filling. Among them, the mean filling method is a well-known technology and will not be elaborated here. As other implementation manners, the implementer can adopt other methods of existing technologies, such as the mode filling method, the KNN algorithm filling method, etc. This embodiment does not make special restrictions on this; secondly, for character-type data, convert it into numerical-type data, where the data type conversion is a well-known technology and will not be elaborated here.

[0024] So far, the information data sets of different users after preprocessing are obtained.

[0025] Step 2, analyze the correlation of local fluctuations between different parameters corresponding to each category in the information data set, and calculate the parameter correlation degree of each category in the information data set.

[0026] Secondly, due to the influence of the equipment itself and environmental interference during the acquisition process, the quality of the acquired data is poor, and due to the differences in the data acquired at different time periods, therefore, only performing anomaly analysis on the data acquired by the same user may have a large judgment error. By considering the similarity characteristics of the data between different users, comparing and analyzing users with similar data characteristics, and then accurately analyzing and identifying the abnormal data of different users.

[0027] First, analyze the response of the changes between different types of parameters of the same user, calculate the parameter correlation degree, and the step flowchart of the method for obtaining the parameter correlation degree of each category in the information dataset provided by the embodiments of the present application is as Figure 2 shown, specifically including: Form a time parameter sequence with the data of various parameters in the information dataset at all times; In this embodiment, the real-time power, cumulative power consumption, voltage, current, instantaneous water flow rate, cumulative water consumption, instantaneous gas flow rate, and cumulative gas consumption at all times in the information dataset are respectively formed into time parameter sequences, so as to obtain the time parameter sequences corresponding to the real-time power, cumulative power consumption, voltage, current, instantaneous water flow rate, cumulative water consumption, instantaneous gas flow rate, and cumulative gas consumption respectively.

[0028] Adopt the moving standard deviation algorithm to obtain the moving standard deviation of each element in the time parameter sequence; It should be noted that the moving standard deviation algorithm is a well-known technology and will not be elaborated here; secondly, the moving standard deviation reflects the fluctuation of various parameters within a local time range.

[0029] Calculate the correlation degree of the moving standard deviation of all elements in the time parameter sequence between any two parameters corresponding to each category in the information dataset; take the average value of the correlation degrees of all any two parameters corresponding to each category as the parameter correlation degree of each category in the information dataset; In this embodiment, the correlation degree is measured by calculating the Pearson correlation coefficient of the moving standard deviation of all elements in the time parameter sequence between any two parameters under each category in the information dataset. Among them, the calculation of the Pearson correlation coefficient is a well-known technology and will not be elaborated here. As other implementation manners, implementers can adopt other methods of existing technologies, such as the Spearman correlation coefficient, etc. This embodiment does not make special restrictions on this. For example, calculate the Pearson correlation coefficient of the moving standard deviation of all elements in the time parameter sequence corresponding to any two parameters among the real-time power, cumulative power consumption, voltage, and current in the power data in the information dataset.

[0030] It should be noted that the parameter correlation degree reflects the change correlation between different types of parameters corresponding to each category. The larger the absolute value of the parameter correlation degree, the more certain the data change correlation between the various parameters corresponding to this category, and there is a response change between the data of different types of parameters.

[0031] Thus, the parameter correlation degree of each category in the information dataset is obtained.

[0032] Step 3: Determine the relative difference degree between any two users by combining the data difference changes of the same type of parameters in the information dataset between any two users and the parameter correlation degree; cluster all users based on the relative difference degree.

[0033] Furthermore, analyze the data difference situation between different users through the parameter correlation degree, and calculate the relative difference degree. Specifically: Calculate the difference in the parameter correlation degrees of all categories in the information dataset between any two users, denoted as the correlation difference. In this embodiment, calculate the Euclidean distance of the parameter correlation degrees of all categories in the information dataset between any two users, denoted as the correlation difference. The calculation of the Euclidean distance is a well-known technology and will not be elaborated here. As other implementation manners, implementers can adopt other methods of the existing technology, such as DTW distance, Manhattan distance, etc. This embodiment does not make special restrictions on this; for example, calculate the Euclidean distance between the parameter correlation degrees corresponding to the power data, water resource data, and gas data between any two users.

[0034] The calculation formula for the relative difference degree between any two users is:

[0035] where is the relative difference degree between the th user and the th user, is the th user and the th user's correlation difference; is the time parameter sequence corresponding to the th type of parameter in the information dataset of the th user, is the time parameter sequence corresponding to the th type of parameter in the information dataset of the th user, is the calculation distance. In this embodiment, by calculating and the DTW distance between them. The calculation of the DTW distance is a well-known technology and will not be elaborated here. is the number of all types of parameters corresponding to all categories in the information dataset.

[0036] It should be noted that the obtained relative difference degree reflects the data difference characteristics between different users by combining the differences between power data, water resource data, and gas data. The larger the relative difference degree, the greater the data characteristic difference between the two users.

[0037] Further, based on the relative difference degree, different users are divided as follows: Based on the relative difference degree, all users are clustered to obtain multiple clustering clusters; In this embodiment, the relative difference degree is used as the distance between different users, and the DBSCAN clustering algorithm is used to cluster all users. Among them, the DBSCAN clustering algorithm is a well-known technology and will not be elaborated here. As other implementation manners, implementers can adopt other methods of the existing technology, such as the hierarchical clustering algorithm, etc. This embodiment does not make special restrictions on this.

[0038] Thus, multiple clustering clusters are obtained.

[0039] Step 4, based on the data of all users at different times under various parameters in each clustering cluster, construct a parameter matrix corresponding to each parameter in each clustering cluster; determine the abnormality degree of each column in the parameter matrix through the abnormality degree of each column element in each row of the parameter matrix and the data abnormality situation of the corresponding user in each row; clean the data in the parameter matrix corresponding to each parameter in all clustering clusters to obtain a cleaned information data set of different users.

[0040] Further, compare the data changes of the same parameter of all users in each clustering cluster and analyze their abnormality situations, specifically: Form a parameter matrix with the time parameter sequences corresponding to the same parameter of all users in each clustering cluster; where each row element corresponds to a user; In this embodiment, taking the current parameter in power data as an example, form a parameter matrix corresponding to the current parameter in each clustering cluster with the time parameter sequences corresponding to the currents of all users in each clustering cluster; taking the cumulative gas consumption parameter in gas data as an example, form a parameter matrix corresponding to the cumulative gas consumption parameter in each clustering cluster with the time parameter sequences corresponding to the cumulative gas consumption parameters of all users in each clustering cluster.

[0041] Perform anomaly detection on all column elements in each row of the parameter matrix respectively to obtain the anomaly scores of each column element in each row; In this embodiment, the Local Outlier Factor (LOF) algorithm is used for anomaly detection. Among them, the LOF anomaly detection algorithm is a well-known technology and will not be elaborated here.

[0042] Calculate the weight corresponding to each row in the parameter matrix through multi-criteria decision analysis; In this embodiment, the multi-criteria decision analysis adopts the CRITIC weight method to calculate the weight corresponding to each row in the parameter matrix. Among them, the CRITIC weight method is a well-known technology and will not be elaborated here. As other implementation manners, implementers can adopt other methods of the prior art, such as the entropy weight method, the TOPSIS method, etc. This embodiment does not make special restrictions on this.

[0043] It should be noted that the greater the weight, the more significant the data difference characteristics of the same parameter among different users in the corresponding clustering cluster.

[0044] Furthermore, based on the anomaly score and the weight, calculate the anomaly degree corresponding to each column in the parameter matrix. Specifically: The calculation formula for the anomaly degree corresponding to each column in the parameter matrix of various parameters in each clustering cluster is:

[0045] Where, is the anomaly degree corresponding to the th column in the parameter matrix, is the anomaly score of the element in the th row and th column in the parameter matrix, is the weight corresponding to the th row in the parameter matrix, is the number of all rows in the parameter matrix.

[0046] It should be noted that the greater the anomaly degree, the greater the possibility that the data of all users corresponding to the th column in the parameter matrix is abnormal. By analyzing the data change situation of different users under the same parameter, calculate the anomaly degree to avoid deviation in the abnormal recognition of the data change of a single user, thereby reducing the accuracy of data cleaning.

[0047] Furthermore, based on the anomaly degree, perform anomaly elimination on the data in the parameter matrix corresponding to various parameters. Specifically: Take the column data corresponding to the anomaly degree greater than the preset threshold in the parameter matrix of various parameters in each clustering cluster as abnormal data, eliminate all the abnormal data in the parameter matrix, and fill in the eliminated data to obtain the cleaned information data set of each user; In this embodiment, the preset threshold value is taken as 1. As other implementation manners, the implementer can set it by himself according to the actual situation. Secondly, the mean filling method is used to fill the deleted data. Among them, the mean filling method is a well-known technology and will not be elaborated here. As other implementation manners, the implementer can adopt other methods of the existing technology. For example, the mode filling method, the KNN algorithm filling method, etc. This embodiment does not make special restrictions on this.

[0048] After the cleaned information data set is transmitted to the data center platform, it is stored. The data center platform performs visual analysis on the information data set through the PowerBI tool to perform data mining on the data.

[0049] It should be noted that the PowerBI tool is a well-known technology and will not be elaborated here. As other implementation manners, the implementer can adopt other visualization tools of the existing technology. For example, the Tableau tool, etc. This embodiment does not make special restrictions on this.

[0050] The embodiment of the present application also provides a data cleaning system, including a memory, a processor, and a computer program stored in the memory and running on the processor. When the processor executes the computer program, the steps of any one of the above-mentioned data cleaning methods are implemented.

[0051] Based on the same inventive concept as the above method, the embodiment of the present application also provides a data center platform, and the platform performs visual analysis on the cleaned information data set.

[0052] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover,

[0053] At least a part of the steps in

[0054] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation to the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several variations and improvements can still be made. Therefore, any simple modifications, equivalent changes, and decorations made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application all fall within the protection scope of the technical solution of the present application.

Claims

1. A data cleaning method, characterized in that: The method comprises the following steps: Acquire information data sets of different users and preprocess the information data sets, wherein the information data sets contain data of various parameters corresponding to each category at different times; Analyzing the correlation between the local fluctuations of different types of parameters corresponding to each category in the information data set, and calculating the parameter correlation degree of each category in the information data set; Determine the relative difference between any two users by comparing the data difference of the same parameter in the information data set between the two users and the parameter correlation; and cluster all users based on the relative difference; Based on the data of all users at different times under various parameters in each clustering cluster, a parameter matrix corresponding to various parameters in each clustering cluster is constructed; the abnormality degree of each column in the parameter matrix is ​​determined by the abnormality degree of the elements in each column in each row of the parameter matrix and the data abnormality of the user corresponding to each row, and the data in the parameter matrix corresponding to various parameters in all clustering clusters is cleaned to obtain the cleaned information data sets of different users.

2. A data cleaning method according to claim 1, characterized in that: The calculating the parameter correlation degree of each category in the information data set includes: The data of various parameters at all times in the information data set are combined into a time parameter sequence; Using a moving standard deviation algorithm, obtaining a moving standard deviation of each element in the time parameter sequence; The correlation degree of the moving standard deviation of all elements in the time parameter sequence between any two parameters corresponding to each category in the information data set is calculated, and the mean value of the correlation degree of all any two parameters corresponding to each category is taken as the parameter correlation degree of each category in the information data set.

3. A data cleaning method as claimed in claim 2, characterized in that: The determining the relative difference between the two arbitrary users includes: Analyze the difference in the parameter association degrees of all categories in the information data set between the arbitrary two users to obtain the association difference between the arbitrary two users; No. User and The relative difference between users The calculation formula is: ,in, For the User and The difference in said associations between the users; For the The information data of users The time parameter sequence corresponding to the parameters, For the The information data of users The time parameter sequence corresponding to the parameters, To calculate the distance, is the number of all types of parameters corresponding to all categories in the information data set.

4. A data cleaning method as claimed in claim 3, characterized in that: The association difference is a metric distance of the parameter associations of all categories in the information data set between any two users.

5. A data cleaning method as claimed in claim 2, characterized in that: The construction process of the parameter matrix is: The time parameter sequences corresponding to all users under various parameters in each cluster are combined into a parameter matrix, where each row element corresponds to a user.

6. A data cleaning method according to claim 1, characterized in that: A further measurement method of the abnormality degree is: performing abnormality detection on all column elements in each row of the parameter matrix respectively to obtain the abnormality score of each column element in each row.

7. A data cleaning method as claimed in claim 6, characterized in that: Determining the abnormality of each column in the parameter matrix includes: Calculate the weight corresponding to each row in the parameter matrix through multi-criteria decision analysis; The calculation formula of the abnormality of each column in the parameter matrix is: ,in, is the parameter matrix The abnormality of the column, is the parameter matrix Line anomaly score of column elements, is the parameter matrix The weight corresponding to the row, is the number of all rows in the parameter matrix.

8. A data cleaning method as claimed in claim 1, characterized in that: The cleaning of the data in the parameter matrix corresponding to various parameters in all clusters includes: A column of data corresponding to the abnormality degree greater than a preset threshold in the parameter matrix of various parameters in each cluster is taken as abnormal data, all abnormal data in the parameter matrix are removed, and missing values ​​are filled in the removed data.

9. A data cleaning system, characterized in that: The system includes a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of a data cleaning method as described in any one of claims 1 to 8 when executing the computer program.

10. A data center platform, applying a data cleaning method in claim 1, characterized in that: The platform performs visual analysis on the cleaned information data set.

Citation Information

Patent Citations

  • Distributed data cleaning system and method based on data analysis

    CN111858572A

  • Data anomaly detection method for user side intelligent load control system

    CN116628529A

  • Hierarchical protection method for bank multi-database operation and maintenance information

    CN117992809A

  • On-board data processing method and device, electronic device and storage medium

    EP4203442A1

  • Techniques for detecting anomalous data points in time series data

    US20250077534A1