Carbon emission monitoring key feature extraction method considering data exception and related device
By using principal component analysis (PCA) to reduce dimensionality, autoencoders to fill in missing data, and interquartile range (IQR) method to remove abnormal data, the problems of missing and abnormal data in urban carbon emission monitoring were solved, and the data quality and the accuracy of the monitoring model were improved.
Patent Information
- Application Number
- CN202510277769.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies have data omissions and anomalies when processing urban carbon emission monitoring data, which makes it difficult to retain key feature dimensions and cannot guarantee the original distribution of the data.
Principal component analysis (PCA) is used for dimensionality reduction, autoencoders are used to fill in missing data, and the interquartile range (IQR) method is used to detect and remove abnormal data to form a key feature dataset.
It effectively retains the key features of urban carbon emission data, reduces data dimensions, fills in missing data, removes abnormal data, ensures the original distribution of data, and improves the accuracy of the monitoring model.
Smart Images

Figure CN120687811A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of carbon emission monitoring, and in particular to a method for extracting key features of carbon emission monitoring taking data anomalies into consideration and a related device. Background Art
[0002] Against the backdrop of increasingly severe global climate change, carbon emissions monitoring and management in cities, as major sources of energy consumption and carbon emissions, have become a research hotspot. Implementing refined carbon emissions monitoring at the city level and providing support for energy optimization decisions will significantly advance the development of a low-carbon society. Therefore, to achieve refined carbon emissions monitoring at the city level, continued research into relevant carbon emissions monitoring methods is warranted.
[0003] Carbon emission monitoring research has invested heavily in technical methods, data acquisition, and model building. Artificial intelligence (AI) has shown particularly impressive results in this area. However, the application of AI technology relies on high-quality data, which is the foundation for its operation and optimization and directly impacts the performance, accuracy, and generalization of AI models. Faced with complex urban environments and diverse emission sources, collected data often suffer from omissions, anomalies, and redundant data features. Therefore, further research is needed to better address these issues and extract key features. However, collected datasets often suffer from omissions and anomalies. Existing key feature extraction methods, such as standardization and normalization, filtering, wrapping, embedding, and ensemble algorithms (such as XGBoost and LightGBM), fail to effectively preserve important data dimensions when dealing with omissions and anomalies. More importantly, they fail to guarantee the original data distribution. Summary of the Invention
[0004] The present invention provides a method and related devices for extracting key features of carbon emission monitoring taking data anomalies into consideration, which are used to solve the problem that the existing technology cannot well retain the important feature dimensions of the data and cannot guarantee the original distribution of the data when processing data omissions and anomalies.
[0005] In view of this, the first aspect of the present application provides a method for extracting key features of carbon emission monitoring taking into account data anomalies, the method comprising:
[0006] The carbon emission monitoring data set to be extracted is screened by a principal component analysis method to extract key features of the carbon emission monitoring data set to obtain a first key feature data set;
[0007] Filling missing data in the first key feature dataset using an autoencoder to obtain a second key feature dataset;
[0008] The abnormal data in the second key feature data set is detected and removed using the four-quarter distance method to obtain a third key feature data set.
[0009] Optionally, the method of screening the carbon emission monitoring dataset to be extracted by principal component analysis to extract key features of the carbon emission monitoring dataset to obtain a first key feature dataset includes:
[0010] Standardize the extracted carbon emission monitoring data set;
[0011] The correlation between different features in the carbon emission monitoring dataset after standardization is performed using a covariance matrix;
[0012] Obtaining characteristic values of the carbon emission monitoring data set according to the correlation;
[0013] Sorting the eigenvalues according to the cumulative variance contribution rate or variance maximization, and constructing the principal component space through the first K largest eigenvalues after sorting;
[0014] The data in the carbon emission monitoring data set to be extracted is projected into the principal component space to obtain a first key feature data set.
[0015] Optionally, obtaining characteristic values of the carbon emission monitoring dataset according to the correlation includes:
[0016] Based on the correlation, eigenvalue decomposition or singular value decomposition is used to obtain eigenvalues of the carbon emission monitoring data set.
[0017] Optionally, the detecting and removing abnormal data in the second key feature data set by using the four-quarter distance method to obtain a third key feature data set includes:
[0018] After sorting the data in the second key feature data set, calculating the IQR;
[0019] defining an upper bound and a lower bound of the second key feature data set;
[0020] Abnormal data in the second key feature data set is marked or removed according to the IQR, the upper bound, and the lower bound to obtain a third key feature data set.
[0021] A second aspect of the present application provides a carbon emission monitoring key feature extraction system considering data anomalies, the system comprising:
[0022] a dimensionality reduction unit, configured to screen the carbon emission monitoring data set to be extracted by using a principal component analysis method, extract key features of the carbon emission monitoring data set, and obtain a first key feature data set;
[0023] a filling unit, configured to fill missing data in the first key feature dataset using an autoencoder to obtain a second key feature dataset;
[0024] The removing unit is configured to detect and remove abnormal data in the second key feature data set by using a four-quarter distance method to obtain a third key feature data set.
[0025] Optionally, the dimension reduction unit is specifically configured to:
[0026] Standardize the extracted carbon emission monitoring data set;
[0027] The correlation between different features in the carbon emission monitoring dataset after standardization is performed using a covariance matrix;
[0028] Obtaining characteristic values of the carbon emission monitoring data set according to the correlation;
[0029] Sorting the eigenvalues according to the cumulative variance contribution rate or variance maximization, and constructing the principal component space through the first K largest eigenvalues after sorting;
[0030] The data in the carbon emission monitoring data set to be extracted is projected into the principal component space to obtain a first key feature data set.
[0031] Optionally, obtaining characteristic values of the carbon emission monitoring dataset according to the correlation includes:
[0032] Based on the correlation, eigenvalue decomposition or singular value decomposition is used to obtain eigenvalues of the carbon emission monitoring data set.
[0033] Optionally, the removing unit is specifically configured to:
[0034] After sorting the data in the second key feature data set, calculating the IQR;
[0035] defining an upper bound and a lower bound of the second key feature data set;
[0036] Abnormal data in the second key feature data set is marked or removed according to the IQR, the upper bound, and the lower bound to obtain a third key feature data set.
[0037] A third aspect of the present invention provides a device for extracting key features of carbon emission monitoring taking into account data anomalies, the device comprising a processor and a memory:
[0038] The memory is used to store program code and transmit the program code to the processor;
[0039] The processor is configured to execute the steps of the method for extracting key features of carbon emission monitoring considering data anomalies as described in the first aspect according to the instructions in the program code.
[0040] A fourth aspect of the present invention provides a computer-readable storage medium for storing program code, wherein the program code is used to execute the method for extracting key features of carbon emission monitoring considering data anomalies as described in the first aspect.
[0041] It can be seen from the above technical solutions that the present invention has the following advantages:
[0042] The present invention provides a method for extracting key features from carbon emissions monitoring that takes data anomalies into account. Principal component analysis (PCA) is first used to reduce data dimensionality and improve subsequent data processing efficiency. A deep learning autoencoder is then used to fill in missing data. Finally, anomaly data is processed using the interquartile range method (IQR). The present invention's principal component analysis (PCA) method is highly effective in extracting key features from urban carbon emissions datasets, eliminating redundant data features. The introduced autoencoder and IQR techniques, respectively, enable missing data filling and anomaly detection, laying the foundation for subsequent urban carbon emissions monitoring. This addresses the problem with existing technologies, which fail to preserve important data feature dimensions and maintain the original data distribution when processing missing and anomaly data. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0044] Figure 1 A flowchart of a method for extracting key features of carbon emission monitoring taking into account data anomalies provided by an embodiment of the present invention;
[0045] Figure 2 The dataset before feature extraction provided by the embodiment of the present invention;
[0046] Figure 3 The data set after key feature extraction provided by the embodiment of the present invention;
[0047] Figure 4 A structural diagram of a key feature extraction system for carbon emission monitoring taking data anomalies into consideration provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0048] In order to make the purpose, features, and advantages of the present invention more obvious and easy to understand, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described below are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0049] See also Figure 1 , a method for extracting key features of carbon emission monitoring considering data anomalies is provided in an embodiment of the present invention, comprising:
[0050] Step 101: Screen the carbon emission monitoring data set to be extracted by principal component analysis, extract key features of the carbon emission monitoring data set, and obtain a first key feature data set.
[0051] It should be noted that this step uses principal component analysis (PCA) to process the urban carbon emission monitoring dataset and retain the key features of the dataset.
[0052] Taking the urban carbon emissions dataset as an example, the relevant features of the urban carbon emissions dataset include more than ten types, such as urban GDP, related industrial energy consumption, and urban electricity consumption. This means that the collected data has a characteristic of excessively high feature dimensionality, which makes it difficult to identify important data features. Therefore, if the feature dimensionality of the dataset is too high, the construction results of the urban carbon emissions monitoring model may be very complex, ultimately leading to poor monitoring results of the urban carbon emissions monitoring model. Therefore, it is necessary to appropriately reduce the feature dimensionality of the data to identify key high-quality data features and achieve the purpose of simplifying the model during model construction. When retaining the key features of the data in the process of reducing the feature dimensionality of the data, the use of principal component analysis (PCA) can effectively achieve the retention of key data features. The specific steps of principal component analysis (PCA) are as follows: Steps 1011-1015:
[0053] In one embodiment, step 101 includes:
[0054] Step 1011: Standardize the extracted carbon emission monitoring data set.
[0055] It should be noted that variance reflects the volatility of data to a great extent and can be a good measure of the differences between data. PCA is based on the size of the data variance to select the appropriate principal components of the data set. Therefore, before using PCA to filter the features of the data, it is necessary to standardize the corresponding data. The standardization process will convert the mean of each feature to 0 and the variance to 1, thereby ensuring that the various features of the data set can be on the same calculation scale. Data standardization is calculated using the following formula:
[0056] ;
[0057] in, It is the original data, and represent the mean and standard deviation of the features respectively.
[0058] Step 1012: Correlation between different features in the standardized carbon emission monitoring dataset is analyzed using a covariance matrix.
[0059] It should be noted that the linear correlation between data features will be represented by an intermediate medium. In the present invention, the correlation between different features between data sets will be represented by the covariance matrix. PCA will first calculate the covariance matrix between each feature, and then use the calculation of the covariance matrix to capture the linear relationship between data features.
[0060]
[0061] in, is the covariance matrix of the dataset.
[0062] Step 1013: Obtain characteristic values of the carbon emission monitoring data set based on the correlation.
[0063] It should be noted that this step is based on correlation and uses eigenvalue decomposition or singular value decomposition to obtain the eigenvalues of the carbon emission monitoring data set. Eigenvalue decomposition is a means of expressing the essential properties of the data matrix. By finding the eigenvalues of the data to solve its corresponding eigenvectors, the purpose of simply analyzing the matrix is finally achieved. There are many ways to decompose the eigenvalues. PCA can obtain the eigenvalues by decomposing the covariance of the eigenvalues, and then the corresponding eigenvectors can be obtained. Singular value decomposition is an important matrix decomposition method widely used in signal processing, data compression, machine learning and other fields. PCA achieves data dimensionality reduction and noise elimination by retaining the largest singular value and its corresponding singular vector.
[0064] Step 1014: Sort the eigenvalues according to the cumulative variance contribution rate or the variance maximization, and construct the principal component space using the first K largest eigenvalues after sorting.
[0065] It should be noted that the eigenvalues are sorted according to the cumulative variance contribution rate or variance maximization, and the eigenvectors corresponding to the first K largest eigenvalues constitute the principal component space.
[0066] Step 1015: Project the data in the carbon emission monitoring dataset to be extracted into the principal component space to obtain a first key feature dataset.
[0067] It should be noted that the original data is projected into the principal component space constructed in the previous step to obtain the data representation after dimensionality reduction:
[0068] ;
[0069] in, is the data representation after dimensionality reduction, It is the feature vector corresponding to the first K principal components selected.
[0070] See also Figure 2-Figure 3 , Figure 2 is the dataset before feature extraction; Figure 3 It is a dataset after key features are extracted.
[0071] Step 102: Fill in the missing data in the first key feature data set through an autoencoder to obtain a second key feature data set.
[0072] It should be noted that the relevant data sets may have a certain degree of data omissions. Generally, when faced with data omissions, the method of deleting the corresponding data is often used to improve the quality of the data set. The premise for this method to achieve excellent results is that the data set is large enough. At this time, deleting the missing data has little effect on the construction of the subsequent model, and there are significant limitations. In order to solve the problem that the urban carbon emission monitoring data set is missing data and the amount of collected data may be small, the present invention introduces a deep learning-based autoencoder technology to fill the missing data, which is used to make up for the defects of the method of directly deleting the missing data.
[0073] The autoencoder consists of three parts: encoder, latent space and decoder. Its working process is as follows: a high-dimensional data sample is input, and the encoder maps the data sample to the latent space. Then, the reconstruction error is minimized by optimizing the network parameters. Finally, the decoder reconstructs the original data from the latent space to achieve the purpose of filling in the missing data.
[0074] Step 103: Use the Quartile Method to detect and remove abnormal data in the second key feature data set to obtain a third key feature data set.
[0075] It should be noted that in the case of data anomalies, this paper introduces an IQR-based method for monitoring abnormal urban carbon emissions data. This method defines the data distribution range based on the quartiles of the data and then identifies outliers that significantly deviate from this range. The specific implementation steps are as follows: Steps 1031-1033:
[0076] In one embodiment, step 103 includes:
[0077] Step 1031: After sorting the data in the second key feature data set, calculate the IQR.
[0078] It should be noted that the formula for calculating IQR is as follows:
[0079] ;
[0080] in, is the 25th percentile of the data sort, is the 75th percentile of the data sort
[0081] Step 1032: Define the upper bound and lower bound of the second key feature data set.
[0082] It should be noted that outliers are usually defined as values outside the upper and lower bounds. The upper and lower bounds are defined as follows:
[0083] ;
[0084] Where m is a multiplier, which can be adjusted according to specific needs to achieve the required effect.
[0085] Step 1033: Mark or remove abnormal data in the second key feature data set according to the IQR and the upper and lower bounds to obtain a third key feature data set.
[0086] It should be noted that after the IQR is calculated and the upper and lower limits of the outliers are determined through the above steps 1031 and 1032, the outliers that are out of range are marked or removed.
[0087] An embodiment of the present invention provides a method for extracting key features from carbon emissions monitoring that takes data anomalies into account. Principal component analysis (PCA) is first used to reduce data dimensionality and improve subsequent data processing efficiency. A deep learning autoencoder is then used to fill in missing data. Finally, anomaly data is processed using the interquartile range method (IQR). The principal component analysis (PCA) method of the present invention is highly effective in extracting key features from urban carbon emissions datasets, eliminating redundant data features. The introduced autoencoder and IQR techniques, respectively, implement missing data filling and anomaly detection, laying the foundation for subsequent urban carbon emissions monitoring. This solves the problem with existing technologies that, when processing missing and anomaly data, fail to preserve important data feature dimensions and maintain the original data distribution.
[0088] The above is a method for extracting key features of carbon emission monitoring taking into account data anomalies provided in an embodiment of the present invention. The following is a system for extracting key features of carbon emission monitoring taking into account data anomalies provided in an embodiment of the present invention.
[0089] See also Figure 4 , a carbon emission monitoring key feature extraction system considering data anomalies is provided in an embodiment of the present invention, including:
[0090] A dimensionality reduction unit 201 is configured to screen the carbon emission monitoring data set to be extracted by using a principal component analysis method, extract key features of the carbon emission monitoring data set, and obtain a first key feature data set;
[0091] A filling unit 202 is configured to fill missing data in the first key feature dataset using an autoencoder to obtain a second key feature dataset;
[0092] The removing unit 203 is configured to detect and remove abnormal data in the second key feature data set by using the four-quarter distance method to obtain a third key feature data set.
[0093] Furthermore, an embodiment of the present invention also provides a device for extracting key features of carbon emission monitoring that takes data anomalies into consideration. The device includes a processor and a memory:
[0094] The memory is used to store program code and transmit the program code to the processor;
[0095] The processor is configured to execute the steps of the method for extracting key features of carbon emission monitoring considering data anomalies as described in the above method embodiment according to the instructions in the program code.
[0096] Furthermore, an embodiment of the present invention also provides a computer-readable storage medium, which is used to store program code, and the program code is used to execute the key feature extraction method for carbon emission monitoring considering data anomalies described in the above method embodiment.
[0097] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0098] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0099] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0100] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0101] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the method of the present invention. The aforementioned storage medium includes various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks.
[0102] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A method for extracting key features of carbon emission monitoring considering data anomalies, characterized in that: include: The carbon emission monitoring data set to be extracted is screened by a principal component analysis method to extract key features of the carbon emission monitoring data set to obtain a first key feature data set; Filling missing data in the first key feature dataset using an autoencoder to obtain a second key feature dataset; The abnormal data in the second key feature data set is detected and removed using the four-quarter distance method to obtain a third key feature data set.
2. The method for extracting key features of carbon emission monitoring considering data anomalies according to claim 1 is characterized in that: The method of screening the carbon emission monitoring data set to be extracted by the principal component analysis method, extracting key features of the carbon emission monitoring data set, and obtaining a first key feature data set includes: Standardize the extracted carbon emission monitoring data set; The correlation between different features in the carbon emission monitoring dataset after standardization is performed using a covariance matrix; Obtaining characteristic values of the carbon emission monitoring data set according to the correlation; Sorting the eigenvalues according to the cumulative variance contribution rate or variance maximization, and constructing the principal component space through the first K largest eigenvalues after sorting; The data in the carbon emission monitoring data set to be extracted is projected into the principal component space to obtain a first key feature data set.
3. The method for extracting key features of carbon emission monitoring considering data anomalies according to claim 2 is characterized in that: The obtaining of characteristic values of the carbon emission monitoring data set according to the correlation includes: Based on the correlation, eigenvalue decomposition or singular value decomposition is used to obtain eigenvalues of the carbon emission monitoring data set.
4. The method for extracting key features of carbon emission monitoring considering data anomalies according to claim 1 is characterized in that: The method of detecting and removing abnormal data in the second key feature data set by using the four-quarter distance method to obtain a third key feature data set includes: After sorting the data in the second key feature data set, calculating the IQR; defining an upper bound and a lower bound of the second key feature data set; Abnormal data in the second key feature data set is marked or removed according to the IQR, the upper bound, and the lower bound to obtain a third key feature data set.
5. A carbon emission monitoring key feature extraction system considering data anomalies, characterized by: include: a dimensionality reduction unit, configured to screen the carbon emission monitoring data set to be extracted by using a principal component analysis method, extract key features of the carbon emission monitoring data set, and obtain a first key feature data set; a filling unit, configured to fill missing data in the first key feature dataset using an autoencoder to obtain a second key feature dataset; The removing unit is configured to detect and remove abnormal data in the second key feature data set by using a four-quarter distance method to obtain a third key feature data set.
6. The carbon emission monitoring key feature extraction system considering data anomalies according to claim 5 is characterized in that: The dimension reduction unit is specifically used to: Standardize the extracted carbon emission monitoring data set; The correlation between different features in the carbon emission monitoring dataset after standardization is performed using a covariance matrix; Obtaining characteristic values of the carbon emission monitoring data set according to the correlation; Sorting the eigenvalues according to the cumulative variance contribution rate or variance maximization, and constructing the principal component space through the first K largest eigenvalues after sorting; The data in the carbon emission monitoring data set to be extracted is projected into the principal component space to obtain a first key feature data set.
7. The carbon emission monitoring key feature extraction system considering data anomalies according to claim 6 is characterized in that: The obtaining of characteristic values of the carbon emission monitoring data set according to the correlation includes: Based on the correlation, eigenvalue decomposition or singular value decomposition is used to obtain eigenvalues of the carbon emission monitoring data set.
8. The carbon emission monitoring key feature extraction system considering data anomalies according to claim 5 is characterized in that: The removal unit is specifically used to: After sorting the data in the second key feature data set, calculating the IQR; defining an upper bound and a lower bound of the second key feature data set; Abnormal data in the second key feature data set is marked or removed according to the IQR, the upper bound, and the lower bound to obtain a third key feature data set.
9. A device for extracting key features of carbon emission monitoring considering data anomalies, characterized in that: The device includes a processor and a memory: The memory is used to store program code and transmit the program code to the processor; The processor is configured to execute the method for extracting key features of carbon emission monitoring considering data anomalies according to any one of claims 1 to 4 according to the instructions in the program code.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store program code, and the program code is used to execute the key feature extraction method for carbon emission monitoring considering data anomalies according to any one of claims 1 to 4.