An air quality data processing method based on neighborhood rough set attribute reduction
By using the neighborhood rough set attribute reduction method and the Shapley value to evaluate the importance of air attributes, a weighted neighborhood rough set is constructed, which solves the problem of redundant attributes in air quality data and improves the accuracy and efficiency of air quality classification.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHONGQING UNIV OF POSTS & TELECOMM
- Filing Date
- 2022-11-25
- Publication Date
- 2026-05-15
AI Technical Summary
In current air quality data processing, the redundant properties of high-dimensional data lead to the "curse of dimensionality," increasing the risk of overfitting, and the contribution of different air attributes is not effectively distinguished.
We employ a neighborhood rough set-based attribute reduction method, using Shapley values to assess attribute importance, constructing a weighted neighborhood rough set, defining weighted neighborhood similarity relationships, calculating the reduced attribute subset with weighted dependency lower bounds, and combining it with a kernel-based attribute reduction algorithm to find an optimized subset of empty attributes.
It effectively reflects the correlation between air attributes, reduces time complexity, finds a subset of air attributes that contribute significantly to decision-making, and improves air quality classification performance.
Smart Images

Figure CN115935160B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of air quality data processing, specifically relating to an air quality data processing method based on neighborhood rough set attribute reduction. Background Technology
[0002] With the development of sensor technology, people can more comprehensively represent things from multiple perspectives, resulting in a large number of complex high-dimensional datasets, of which air quality datasets are one type. The classification of air quality data also exists in real life; the ability to quickly identify air quality based on data characteristics across various dimensions has created a need for classifying high-dimensional data. However, due to the large amount of redundant information in high-dimensional data, air quality data classification suffers from the "curse of dimensionality" problem. Redundant attributes in the data can confuse learning algorithms, increasing the risk of overfitting. Therefore, addressing the problem of redundant attributes in air quality data is essential.
[0003] Air quality data holds considerable value in today's big data environment, primarily originating from ground-based monitoring and meteorological satellite data collection sites. Given the importance of air quality to people's lives, many scholars are utilizing machine learning methods to process air quality data, such as predicting trends in air pollutant concentrations and classifying and predicting air quality levels. Existing approaches to air quality data processing include: Tian et al. using VAE-based inline relationships for feature extraction and clustering to analyze the internal connections between data; Zhou et al. using traditional autoregressive moving average models to study air quality data, analyzing its stationarity, seasonality, and difference stabilization; and Li et al. using LSTM-based models for representation learning of air quality data to uncover its potential patterns of change.
[0004] However, the air quality data model mentioned above does not consider air quality data attributes such as temperature, humidity, and the weights of each pollutant. However, the contribution of each air attribute to the decision may be different. Therefore, air attributes need to be processed in different ways, that is, different weights need to be assigned to different air attributes. Summary of the Invention
[0005] To address the shortcomings of existing technologies, this invention utilizes the ideas of Shapley value and neighborhood rough set construction to create equivalence classes. It designs a method for attribute reduction based on neighborhood rough sets and applies this method to air quality datasets. This method uses the dependency of the neighborhood rough set and introduces the calculation of the Shapley value to calculate the weight of each attribute. Then, a weighted neighborhood rough set is constructed based on the attribute weights. The upper and lower approximation relationships of the weighted neighborhood are redefined, a new dependency is defined, and finally, an importance lower limit greater than the set importance lower limit is calculated based on the new dependency. The invention addresses the problem of consistent air attribute weights in existing air quality data processing. By using weighted calculations of neighborhood rough sets to obtain air attributes, it overcomes the shortcomings of neighborhood rough set models. This invention proposes an air quality data processing method based on neighborhood rough set attribute reduction, specifically including the following steps:
[0006] Obtain air quality sample data with multiple air attributes, arrange and combine each air attribute, and calculate the dependency of each subset of air attributes.
[0007] The importance of each air attribute is assessed using the Shapley value. A weighted rough set is constructed based on the importance of each air attribute, and the weighted neighborhood similarity relationship is determined.
[0008] We define a weighted neighborhood similarity class based on the weighted neighborhood similarity relationship, and obtain a lower approximate definition of the similarity class; based on the lower approximate definition, we obtain the positive domain for classification decision.
[0009] Based on the weighted neighborhood similarity relationship and the positive domain of the classification decision, the weighted dependency of the classification decision on the air attribute subset is obtained;
[0010] We use a kernel-inspired, forward approximation-based method to find a reduced subset of air properties and output optimized air quality data.
[0011] The method of the present invention has the following beneficial effects:
[0012] 1. Compared to traditional neighborhood rough sets that use the same air attribute weights, the method used in this invention takes into account the contribution of each conditional air attribute to the decision attribute and uses this as the weight of the conditional air attribute, thus better reflecting the correlation between each air attribute.
[0013] 2. This invention introduces the importance of conditional air attributes into neighborhood relations to define a weighted neighborhood rough set model, and uses the correlation of weighted neighborhood relations to measure the importance of the air attribute set.
[0014] 3. This invention utilizes a kernel-inspired attribute reduction algorithm based on positive approximation, which has low time complexity and can find a subset of air attributes that contribute significantly to decision-making. These processed subsets of air attributes can provide better air quality classification performance. Attached Figure Description
[0015] Figure 1 This is a flowchart of an air quality data processing method based on neighborhood rough set attribute reduction according to an embodiment of the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] Rough sets, proposed by Polish scholar Pawlak in 1982, are a mathematical tool for characterizing the classification of uncertain information. They are a powerful soft computation method for handling incomplete and uncertain knowledge. The basic idea of classical rough sets is to group samples in a set using equivalence relations, and then divide the set into three regions based on equivalence classes: the positive region, the boundary region, and the negative region. Based on these three regions, rule induction, knowledge mining, decision-making, and feature selection can be performed. However, classical rough sets are only suitable for processing discrete data. For continuous data, discretization is necessary, which often leads to information loss. Hu Qinghua et al. proposed the neighborhood rough set model, which can directly process continuous data without discretization, thus avoiding information loss during discretization.
[0018] In neighborhood rough sets, the magnitude of dependency reflects the importance of decision attributes to conditional attributes. This not only demonstrates the dependence of classification attributes on conditional attributes but also helps identify key air attributes that play a role in air quality classification. Therefore, based on the above analysis, this invention proposes an air quality data processing method based on neighborhood rough set attribute reduction, such as... Figure 1 As shown, the method specifically includes the following steps:
[0019] S1. Obtain air quality sample data with multiple air attributes, arrange and combine each air attribute, and calculate the dependency of each subset of air attributes.
[0020] In this embodiment of the invention, it is first necessary to obtain air quality sample data with multiple air properties, including... , Ground-level ozone, particulate matter, etc.; To more clearly illustrate the data processing procedure of this invention, this embodiment provides a decision information table, which is shown in Table 1 as IS=(U,C,D):
[0021] Table 1 Decision Information Table IS=(U,C,D)
[0022]
[0023] Among them, the air quality sample set Represents air conditions at different times; conditional air attribute set Indicates air , Ground-level ozone, particulate matter; decision attribute set To indicate whether the air quality is good or bad, you can set D=1 to indicate good air quality and D=2 to indicate bad air quality, or vice versa.
[0024] Based on the above decision information table, the conditional air attribute set By arranging and combining the various air properties in the data, we obtain: { }、{ }、{ }、{ }、{ }、{ }、{ }、{ }、{ }、{ }、{ }、{ }、{ }、{ }、{ Then, following the classic formula for calculating the dependency of neighborhood rough sets, the dependency of each air attribute subset is obtained, as follows:
[0025] =
[0026] in, This represents the dependence of the air attribute subset S on the decision attribute D, and is used to measure the ability of the air attribute subset S to approximate the classification decision attribute D. It is the neighborhood threshold; It is the positive domain of the classification decision attribute D relative to the subset S of air attributes; The cardinality of a set. This represents a set of air quality sample data.
[0027] in, , , , In this embodiment, Euclidean distance is used for measurement. In this case, the neighborhood threshold... .
[0028] Decision attribute D divides the universe of discourse U into two equivalence classes:
[0029] , ;
[0030] when { }hour: ;
[0031] when { }hour: ;
[0032] when { }hour: ;
[0033] when { }hour: ;
[0034] when { }hour: ;
[0035] when { }hour: ;
[0036] when { }hour: ;
[0037] when { }hour: ;
[0038] when { }hour: ;
[0039] when { }hour: ;
[0040] when { }hour: ;
[0041] when { }hour: ;
[0042] when { }hour: ;
[0043] when { }hour: ;
[0044] when { }hour: .
[0045] S2. The importance of each air attribute is evaluated using the Shapley value. A weighted rough set is constructed based on the importance of each air attribute, and the weighted neighborhood similarity relationship is determined.
[0046] In this embodiment of the invention, based on the dependency of each air attribute subset, and assuming that each air attribute subset is not empty, a Shapley value is introduced to calculate the importance of each air attribute. The calculation formula is as follows:
[0047]
[0048] in, Indicates air properties The weight, or importance, This represents the i-th type of air property. ; S represents the cardinality of the set of air attributes; S represents the set of attributes that do not contain air attributes. A subset of attributes that is not empty. The number of elements in the set S containing air attributes is represented by ; P represents the number of elements in the set S containing air attributes. The set formed by all subsets of air properties.
[0049] Therefore, we can conclude that:
[0050]
[0051] + + + +
[0052] + +
[0053]
[0054] 1.04
[0055]
[0056]
[0057]
[0058]
[0059]
[0060]
[0061] The above analysis shows that, under the above conditions, when judging whether air quality is good or bad, air properties... , It played a relatively large role, so it was given greater weight, i.e., importance. However, how to give specific weights to air attributes still needs to be determined in the following way.
[0062] S3. Define a weighted neighborhood similarity class based on the weighted neighborhood similarity relationship, and obtain a lower approximate definition of the similarity class; derive the positive domain for classification decision based on the lower approximate definition;
[0063] In this embodiment of the invention, a weighted rough set is constructed based on the importance of each air attribute, and a weighted neighborhood similarity relationship is determined. That is, if the weighted Euclidean distance between air quality sample x and air quality sample y is less than the neighborhood threshold, then air quality sample x and air quality sample y belong to a weighted neighborhood similarity relationship; otherwise, they do not belong to a weighted neighborhood similarity relationship.
[0064] In this model, different weights are used to calculate neighborhood similarity classes for different attributes, with the attribute weights measured by the Shapley value. This rough set model can also be called a weighted neighborhood rough set, and the weighted Euclidean distance, i.e., the weighted neighborhood similarity relation, can be expressed as:
[0065]
[0066] in, , It is an air property The weight, This indicates that air quality sample x represents air properties. The value below, Indicates the neighborhood threshold; when When calculating relationships, air properties Its importance will increase; when At that time, air properties Its importance will decrease; when At that time, air properties Its importance will remain unchanged.
[0067] S4. Based on the weighted neighborhood similarity relationship and the positive domain of the classification decision, obtain the weighted dependency of the classification decision on the air attribute subset;
[0068] In this embodiment of the invention, given the decision information table IS=(U,C,D) and the weighted neighborhood similarity class , , Compared to The upper and lower approximations are defined as follows:
[0069]
[0070]
[0071] Among them, when At that time, for Relationship It is precise; otherwise, for Relationship It is rough. Therefore, .
[0072] Given a decision information table IS=(U,C,D) and weighted neighborhood similarity relations For decision attributes D regarding weighted neighborhood similarity relations The upper and lower approximations are defined as follows:
[0073]
[0074] ;
[0075] Regarding weighted neighborhood relations The decision boundary region and decision positive region of decision attribute D are defined as follows:
[0076]
[0077] ;
[0078] in, Weighted neighborhood similarity is measured from both upper and lower approximation perspectives. Roughness; Weighted neighborhood similarity is measured from the perspective of lower approximation. The roughness.
[0079] For the computational purposes of this invention, this embodiment only needs to obtain the following approximate relationship to determine the corresponding positive region, specifically including the following:
[0080] Sample set Compared to weighted neighborhood similarity The lower approximation is defined as:
[0081]
[0082] For weighted neighborhood relations The positive region of decision attribute D is defined as:
[0083]
[0084] in, Indicates the threshold in the neighborhood The approximate definition of the similarity class of a subset X of air quality samples, where x represents an air quality sample. Indicates the threshold in the neighborhood The weighted neighbor set of air quality sample x, where X represents a subset of air quality samples. Represents a set of air quality samples. Represents the set of decision attributes. This represents a subset of samples categorized by decision attributes. It is the positive region of classification decision attribute D relative to the air attribute subset B under weighted neighborhood similarity relation, which measures the weighted neighborhood similarity relation from the perspective of lower approximation. Roughness; Indicates the threshold in the neighborhood Subset of samples for decision attribute classification The approximate definition of similar classes.
[0085] Given a decision information table IS=(U,C,D) and weighted neighborhood similarity relations Decision attribute D for weighted neighborhood similarity Dependency is defined as:
[0086]
[0087] in, This represents the weighted dependency of the air attribute subset B relative to the decision attribute D, used to measure the ability of the air attribute subset S to approximate the classification decision attribute D. It is the neighborhood threshold; It is the positive domain of the classification decision attribute D under the weighted neighborhood similarity relation relative to the air attribute subset B; The cardinality of a set. This represents a set of air quality sample data.
[0088] S5. Use a kernel-inspired, positive approximation-based method to find a reduced subset of air properties and output optimized air quality data.
[0089] In this embodiment of the invention, the definition of weighted dependency can be used to find a better subset of air attribute features for easier classification. The relevance importance of air attributes can be defined as:
[0090]
[0091] in, Indicates conditional air properties The importance of the air attribute subset B relative to the classification decision attribute D. ,if Then the properties of air It's superfluous. This represents the weighted dependency of the air attribute subset B relative to the decision attribute D. This indicates that in the subset B of air attributes, attributes are added. The weighted dependency of the decision attribute D.
[0092] As can be seen from the above definition, weighted neighborhood rough sets reflect the relationships between attributes and between attributes and decisions, which is more conducive to the accuracy of decisions. Furthermore, the dependency of weighted neighborhood rough sets increases with the increase of attributes. When reducing the decision table, the change in the relevance of the "air" attribute should be considered. If there is no change, it means that the "air" attribute is redundant.
[0093] This embodiment mainly illustrates the attribute reduction using a positive approximation-based method after establishing a model of a weighted neighborhood rough set.
[0094] A greedy search strategy is used to find a minimal subset of attributes that has the same ability to represent samples as the original attribute set. Given a decision information table IS=(U,C,D), the attribute subset... Neighborhood threshold and attributes If the following conditions are met, attribute subset B can be considered a reduction of the entire attribute set C relative to decision attribute D:
[0095] Sufficiency: ;
[0096] necessity: ;
[0097] Among them, the degree of dependence of the air attribute subset B on the decision attribute D is consistent with the degree of dependence of the air attribute universal set C on the decision attribute D; as can be seen from the necessity, all attributes in the air attribute subset B are necessary attributes; therefore, the air attribute subset B and the air attribute universal set C are the smallest attribute subsets with the same degree of dependence.
[0098] Conditional attributes The importance of the air attribute subset B relative to the decision attribute set D is defined as follows:
[0099] ;
[0100] in, ,if Then the properties of air It's unnecessary.
[0101] Based on the above analysis, the neighborhood rough set is weighted according to the weight of each air attribute, and the kernel of the conditional attribute set is calculated. Starting from the kernel, a minimum attribute subset is found according to the greedy search strategy, and the importance of each air attribute in the remaining air attributes after removing the kernel attribute set is calculated. The importance of air attributes is compared and the air attribute with the highest importance is added. Until the importance of the highest air attribute is 0, the air attribute most relevant to the air quality data can be determined.
[0102] Step 1: Calculate the neighborhood relationship matrix in Table 1 based on the weighted neighborhood similarity relationship:
[0103] in
[0104] when hour:
[0105] =0.3
[0106] when hour:
[0107] =0
[0108] when hour:
[0109] =0.1
[0110] when hour:
[0111] =0.3
[0112] when hour:
[0113] =0.2
[0114] in, When x and y belong to a weighted neighborhood similarity relationship, the neighborhood relationship matrix takes the value of 1; otherwise, it takes the value of 0.
[0115] Step 2:
[0116] For each ,if Then the properties of air Belongs to the kernel property set From the neighborhood relation matrix in step 1, we can see that the kernel attribute set... { }
[0117] Step 3:
[0118] Calculate the importance of each air attribute among the remaining attributes after removing the kernel attribute set, compare the importance of the attributes and select the attribute with the highest importance to add, until the highest importance attribute has an importance of 0, then stop.
[0119] From step 2, the remaining attribute set = { },and ;calculate 0, therefore air properties Not included in the kernel property set middle.
[0120] The final returned attribute reduction subset is { Therefore, this embodiment can utilize [the following method / mechanism] when judging air quality. , The air quality sample data is classified using the set of particulate matter attributes. Any existing classification model can be used to classify the air quality sample data, such as the SVM support vector machine model, the convolutional neural network model, and the deep learning model.
[0121] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include ROM, RAM, disk, or optical disk, etc.
[0122] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A method for processing air quality data based on neighborhood rough set attribute reduction, characterized in that, The method includes: Obtain air quality sample data with multiple air attributes, arrange and combine each air attribute, and calculate the dependency of each subset of air attributes. The importance of each air attribute is assessed using the Shapley value. A weighted rough set is constructed based on the importance of each air attribute, and the weighted neighborhood similarity relationship is determined. The use of Shapley values to assess the importance of various air properties includes: in, Indicates air properties The degree of importance; S represents the cardinality of the set of air attributes; S represents the set of attributes that do not contain air attributes. A subset of air properties that is not empty. The number of elements in the set S containing air attributes is represented by ; P represents the number of elements in the set S containing air attributes. The set formed by all subsets of air properties; This indicates that the subset S of air attributes is affected by the addition of air attributes. The degree of dependence on decision attribute D; This represents the dependence of the air attribute subset S on the decision attribute D; We define a weighted neighborhood similarity class based on the weighted neighborhood similarity relationship, and obtain a lower approximate definition of the similarity class; based on the lower approximate definition, we obtain the positive domain for classification decision. Based on the weighted neighborhood similarity relationship and the positive domain of the classification decision, the weighted dependency of the classification decision on the air attribute subset is obtained; We use a kernel-inspired, forward approximation-based method to find a reduced subset of air properties and output optimized air quality data.
2. The air quality data processing method based on neighborhood rough set attribute reduction according to claim 1, characterized in that, The air properties include any combination of PM2.5, PM10, SO2, CO, NO2, O3, and particulate matter.
3. The air quality data processing method based on neighborhood rough set attribute reduction according to claim 1, characterized in that, The process involves constructing a weighted rough set based on the importance of each air attribute and determining the weighted neighborhood similarity relationship. If the weighted Euclidean distance between air quality sample x and air quality sample y is less than a neighborhood threshold, then air quality sample x and air quality sample y belong to the weighted neighborhood similarity relationship; otherwise, they do not belong to the weighted neighborhood similarity relationship.
4. The air quality data processing method based on neighborhood rough set attribute reduction according to claim 3, characterized in that, The weighted Euclidean distance is expressed as: in, , It is an air property The weight, This indicates that air quality sample x represents air properties. The value below, Indicates the neighborhood threshold; when When calculating relationships, air properties Its importance will increase; when At that time, air properties Its importance will decrease; when At that time, air properties Its importance will remain unchanged.
5. The air quality data processing method based on neighborhood rough set attribute reduction according to claim 1, characterized in that, The step of defining a weighted neighborhood similarity class based on weighted neighborhood similarity relations yields a lower approximate definition of the similarity class; based on this lower approximate definition, the positive domain for classification decisions includes: in, Indicates the threshold in the neighborhood The approximate definition of the similarity class of a subset X of air quality samples, where x represents an air quality sample. Indicates the threshold in the neighborhood The weighted neighbor set of air quality sample x, where X represents a subset of air quality samples. Represents a set of air quality samples. Represents the set of decision attributes. This represents a subset of samples categorized by decision attributes. It is the positive region of classification decision attribute D relative to the air attribute subset B under weighted neighborhood similarity relation, which measures the weighted neighborhood similarity relation from the perspective of lower approximation. Roughness; Indicates the threshold in the neighborhood Subset of samples for decision attribute classification The approximate definition of similar classes.
6. The air quality data processing method based on neighborhood rough set attribute reduction according to claim 1, characterized in that, The weighted dependency of the classification decision on the air attribute subset, obtained based on the weighted neighborhood similarity relationship and the positive domain of the classification decision, includes: ; in, This represents the weighted dependency of the air attribute subset B relative to the decision attribute D, used to measure the ability of the air attribute subset S to approximate the classification decision attribute D. It is the neighborhood threshold; It is the positive domain of the classification decision attribute D under the weighted neighborhood similarity relation relative to the air attribute subset B; The cardinality of a set. This represents a set of air quality sample data.
7. The air quality data processing method based on neighborhood rough set attribute reduction according to claim 1, characterized in that, The method of finding a reduced subset of air attributes using a kernel-inspired positive approximation approach includes calculating the kernel of the conditional attribute set; starting from the kernel, finding a minimum subset of attributes using a greedy search strategy, and calculating the importance of each attribute among the remaining attributes after removing the kernel attribute set. Compare the importance of attributes and select the attribute with the highest importance to add; continue until the highest importance attribute has an importance of 0.
8. The air quality data processing method based on neighborhood rough set attribute reduction according to claim 7, characterized in that, The formula for calculating the importance of the air properties is as follows: in, Indicates conditional air properties The importance of the air attribute subset B relative to the classification decision attribute D. ,if Then the properties of air It's superfluous. This represents the weighted dependency of the air attribute subset B relative to the decision attribute D. This indicates that in the subset B of air attributes, attributes are added. The weighted dependency of the decision attribute D.