PM2.5 pollution prediction method and equipment based on adaptive sample expansion
Through the adaptive sample expansion method, high-quality PM2.5 data samples are generated using the adversarial network, which solves the small sample problem in the PM2.5 concentration prediction model and improves the prediction accuracy of PM2.5 pollution events and the model generalization ability.
Patent Information
- Application Number
- CN202510807392.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-06-17
AI Technical Summary
In the existing technology, there is a small sample problem in the training process of PM2.5 concentration prediction models. Traditional data expansion methods cannot generate high-quality samples that conform to the actual distribution and dynamic changes, resulting in limited model prediction ability and generalization performance for PM2.5 pollution events.
Through the adaptive sample expansion method, PM2.5 concentration data and related environmental data are obtained, preprocessed and feature extracted, a similarity matrix is constructed for cluster analysis, and new samples are generated using an adversarial network. Considering time dependence and sequence characteristics, high-quality PM2.5 data expansion samples are generated to improve the distribution balance of training data.
The prediction accuracy and generalization performance of PM2.5 pollution events have been significantly improved. The generated samples are consistent with the actual environmental characteristics, which makes up for the lack of training data and improves the prediction accuracy and reliability of the model.
Smart Images

Figure CN120354219B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of environmental monitoring technology, and in particular to a PM2.5 pollution prediction method and device based on adaptive sample expansion. Background Art
[0002] Surface PM2.5 concentration is a key indicator of air pollution severity. High-concentration pollution events often occur under specific meteorological conditions, such as calm winds, high humidity, or temperature inversions. High-concentration PM2.5 pollution events can damage environmental quality and threaten public health. However, due to the low frequency and low probability of high PM2.5 pollution, the corresponding sample data volume is relatively small, making it difficult for machine learning models to fully learn and grasp the pollution patterns in these special circumstances during training. This, in turn, limits the model's predictive power and generalization performance for PM2.5 pollution events.
[0003] To address the small sample size issue during model training, existing data augmentation techniques often use oversampling methods. These methods generate new synthetic samples by interpolating between existing minority class samples, thereby increasing the number of minority class samples and improving the balance of data distribution. For example, Chinese patent application CN118095050A discloses an artificial intelligence-based method for transient stability assessment of power systems that considers sample imbalance. This method uses a K-nearest neighbor algorithm to select suitable neighbors for minority class samples and then generates new samples through linear interpolation. This method achieves sample data augmentation by generating new samples through interpolation between minority class samples.
[0004] However, the above-mentioned oversampling-based data augmentation methods have significant limitations:
[0005] On the one hand, traditional oversampling-based data augmentation methods are usually based on the binary sample assumption, that is, samples need to be divided into majority and minority classes. The accuracy of this classification method is highly dependent on threshold selection. In the application scenario of PM2.5 air quality data, due to factors such as environment, regional characteristics and time changes, the classification standards are inconsistent, which in turn affects the accuracy of data augmentation. In particular, in scenarios with multiple classifications or blurred category boundaries, the accuracy of data augmentation will be significantly reduced.
[0006] On the other hand, the new samples generated by traditional oversampling-based data augmentation methods are only simple feature-based linear combinations, which cannot effectively capture the nonlinear relationships in complex feature spaces and have insufficient expansion capabilities when processing high-dimensional and complex feature data. PM2.5 data has time dependence and sequence characteristics. Traditional data augmentation methods are difficult to generate high-quality samples that conform to actual distribution and dynamic changes, and cannot provide sufficient data augmentation for the model training process.
[0007] In summary, the existing technology has a small sample size problem during the training process of PM2.5 concentration prediction models, and traditional data augmentation methods are not suitable for the expansion of PM2.5 concentration data. It is difficult to accurately generate high-quality PM2.5 sample data that conforms to the actual distribution and dynamic changes. As a result, there is a problem of insufficient training of PM2.5 high pollution concentration samples in machine learning, which limits the model's ability to predict PM2.5 pollution events and the generalization of the model. Summary of the Invention
[0008] The technical problem to be solved by the present invention is: In response to the technical problems existing in the prior art, the present invention provides a PM2.5 pollution prediction method and device based on adaptive sample expansion, which can dynamically cluster samples according to the data characteristics of PM2.5 concentration, while considering the time dependence and sequence characteristics of PM2.5 data, generating high-quality PM2.5 data expansion samples, improving the distribution balance of training sample data, and enhancing the accuracy and reliability of PM2.5 pollution prediction.
[0009] In order to solve the above technical problems, the technical solution proposed by the present invention is:
[0010] A PM2.5 pollution prediction method based on adaptive sample expansion includes the following steps:
[0011] Step S01: Acquire PM2.5 concentration data of a target prediction area at different times and an environmental data set corresponding to the PM2.5 concentration to form an initial sample data set, wherein the environmental data set includes air pollutant concentration data, meteorological data, and time information corresponding to the air pollutant concentration data and meteorological data;
[0012] Step S02: After preprocessing the environmental data set, relevant features that are correlated with PM2.5 concentration are extracted, and correlation coefficients between the extracted relevant features and PM2.5 concentration are calculated respectively. The relevant features are screened according to the calculated correlation coefficients to obtain a final set of relevant features;
[0013] Step S03: Calculating the similarity between samples in the initial sample data set using the relevant feature set as a constraint condition, constructing a PM2.5 concentration sample similarity matrix under environmental constraints, performing cluster analysis on the PM2.5 concentration sample similarity matrix to classify the sample data in the initial sample data set, obtaining sample categories of the initial sample data set, and calculating the number of samples in each sample category;
[0014] Step S04: Segmenting the initial sample data set to obtain multiple input samples, inputting each input sample into an adversarial network for model training, and obtaining a generative adversarial network model after training to generate PM2.5 concentration samples, wherein the PM2.5 concentration samples include PM2.5 concentration and the corresponding related feature set;
[0015] Step S05: Generate new samples for each of the sample categories using the trained generative adversarial network model, and determine the number of new samples to be generated based on the number of samples in each sample category, wherein the new samples include PM2.5 concentration data and corresponding related feature sets;
[0016] Step S06: Add the generated new samples to the initial sample data set to form an expanded sample set, and use the expanded sample set to train the machine learning model to obtain a PM2.5 concentration prediction model for realizing PM2.5 concentration prediction. During the model training process, the relevant feature set corresponding to the PM2.5 concentration is used as a prediction factor, and the PM2.5 concentration is used as the target variable.
[0017] Optionally, the relevant characteristics include air pollutant characteristics, meteorological characteristics, pollutant diffusion coefficients, and oxidant mass concentrations. The air pollutant characteristics include characteristics of any one or more pollutants among NO2, SO2, CO, and O3. The meteorological characteristics include any one or more of temperature, relative humidity, atmospheric visibility, solar radiation intensity, wind speed and direction, boundary layer height, and precipitation. The pollutant diffusion coefficient is calculated based on the boundary layer height and wind speed in the meteorological data. The oxidant mass concentration is calculated based on the NO in the atmosphere. x The concentration and O3 concentration are calculated.
[0018] Optionally, in step S03, the similarity between samples is calculated according to the following formula:
[0019]
[0020] in, Represents a sample and samples The similarity value of PM2.5 concentration between 、 Represents a sample and samples PM2.5 concentration, 、 Represents a sample and samples The corresponding set of relevant features, It is the scale parameter for adjusting the similarity in the Gaussian kernel function.
[0021] Optionally, step S03 further includes: after clustering is completed, calculating a sample distribution balance value based on the number of samples in each cluster, and determining whether to proceed to step S04 to trigger sample expansion based on the calculated sample distribution balance value. The sample distribution balance value is used to characterize the balance degree of the sample distribution, and the calculation expression of the sample distribution balance value is:
[0022]
[0023] in, B is the balance index of data distribution, is the i-th cluster The number of samples, is the average number of samples in all clusters, is the total number of sample clusters.
[0024] Optionally, in step S04, when each input sample is input into the adversarial network for model training, the total loss function of the generator is for:
[0025]
[0026] in, represents the generator loss function of WANG, represents a high concentration penalty term used to impose penalties on samples generated when the PM2.5 concentration exceeds the historical peak threshold. Represents a correlation constraint condition for making the PM2.5 concentration of the generated sample negatively correlated with the pollution diffusion coefficient and positively correlated with the oxidant concentration;
[0027] The loss function of the discriminator is:
[0028]
[0029] in, is the loss function of the discriminator, Indicates the number of multi-target output features, It is a feature The weight of Represents the discriminator pair consisting of conditional variables and random noise Features in the sample generated after input The judgment result of Represents the discriminator's response to the real sample Medium Features The judgment result of represents the distribution of generated sample data, Represents the distribution of real sample data.
[0030] Optionally, the calculation expression of WANG's generator loss function is:
[0031]
[0032] Where, is the expected value of the output target, Indicates the number of multi-target output features, It is a feature The weight of Represents the discriminator pair consisting of conditional variables and random noise Features in the sample generated after input The judgment result of Represents the distribution of generated sample data;
[0033] The calculation expression of the high concentration penalty term is:
[0034]
[0035] Where, is the penalty weight used to balance the strength of the adversarial loss and the physical constraint, To generate the PM2.5 concentration value of the sample, The historical peak threshold of PM2.5 concentration value for generating samples;
[0036] The calculation expression of the constraint condition of the correlation relationship is:
[0037]
[0038] Where, 、 are the diffusion coefficient and oxidant concentration, respectively, is the covariance, when >0 or <0, physical penalty items are applied , chemical penalty items , is a positive number.
[0039] Optionally, in step S05, the expansion weight of each category is calculated according to the number of samples in each sample category:
[0040]
[0041] in, represents the expansion weight of the i-th sample category, represents the number of samples of the i-th sample category, Indicates the number of samples of the category with the largest number of samples among all categories;
[0042] Use augmented weights for each category Calculate the number of new samples required to generate for each sample category:
[0043]
[0044] in, Indicates the number of new samples required to be generated for the i-th sample category, Indicates the preset initial number of sample expansions, Indicates the The expansion weight of the sample category.
[0045] Optionally, in step S05, generating new samples for each sample category using the trained generative adversarial network model includes:
[0046] Step S511: randomly selecting a corresponding number of time points from the current samples of each sample category according to the number of new samples required to be generated for each sample category, and using them as reference time points for generating new samples;
[0047] Step S512: At each selected time point, obtain historical environmental data and random noise at each selected time point, wherein the historical environmental data includes air quality and meteorological data, and input the historical environmental data and random noise into the trained generator model to generate a complete new sample data.
[0048] Optionally, in step S03, an affinity propagation algorithm model is used to perform cluster analysis on the similarity matrix of PM2.5 concentration samples; step S06 also includes obtaining model performance indicators corresponding to different expanded sample numbers through grid search, comparing the model performance indicators corresponding to different expanded sample numbers, and selecting the expanded sample number that best corresponds to the model performance indicator as the final expanded sample number.
[0049] The present invention also provides an electronic device, comprising a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute a PM2.5 pollution prediction method based on adaptive sample expansion.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] The present invention extracts relevant features correlated with PM2.5 concentration after preprocessing the environmental data set and constructs an initial sample data set. The relevant feature set is used as a constraint condition to calculate the similarity between each sample in the initial sample data set, constructs a PM2.5 concentration sample similarity matrix under environmental constraints, and performs cluster analysis on the similarity matrix to achieve dynamic adaptive classification of samples. New samples are generated by introducing an adversarial network, and the number of new samples required to be generated is determined based on the number of samples in each sample category. This method can not only capture the nonlinear relationship and time series dynamic law of complex feature space in PM2.5 pollution events, generate high-quality PM2.5 concentration samples, and make the generated sample division more consistent with the actual environmental characteristics of PM2.5 pollution, but also improve the balance of data distribution, effectively expand the training data set, make up for the defect of insufficient extreme pollution event samples in traditional machine learning model training, and significantly improve the prediction accuracy and generalization performance of extreme and low-frequency PM2.5 pollution events. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 This is a flow chart of a PM2.5 pollution prediction method based on adaptive sample expansion in an embodiment of the present invention. DETAILED DESCRIPTION
[0053] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0054] like Figure 1 As shown, the PM2.5 pollution prediction method based on adaptive sample expansion in this embodiment includes the following steps:
[0055] Step S01: Obtain PM2.5 concentration data of the target prediction area at different times and the corresponding environmental data set related to PM2.5 concentration to form an initial sample data set. The environmental data set includes air pollutant concentration data, meteorological data, and time information corresponding to the air pollutant concentration data and meteorological data.
[0056] In this embodiment, a daily historical data set related to PM2.5 concentration in the target prediction area is collected, including air pollutant concentration data, meteorological data, timestamp information, etc., and an initial sample data set is formed by the PM2.5 concentration data and related air pollutant concentration data, meteorological data and other components. The initial sample data set contains multiple PM2.5 concentration samples, and each PM2.5 concentration sample contains PM2.5 concentration data and related air pollutant concentration data, meteorological data, etc. obtained at the same time.
[0057] Specifically, the environmental data set is composed of PM2.5 concentration air quality data, air pollutant data related to PM2.5 concentration, and meteorological data collected on a daily basis through monitoring stations, APIs, or historical data platforms. The relevant air pollutant data include NO x , SO2, CO and O3 and other air pollutants related to the production of PM2.5; meteorological data include temperature, relative humidity, atmospheric visibility, solar radiation intensity, wind speed and direction, boundary layer height, precipitation and other relevant meteorological conditions data. x , SO2, CO and O3, etc.) and related meteorological data to extract relevant features as possible attribute features that affect PM2.5 concentration.
[0058] Step S02: After preprocessing the environmental data set, relevant features that are correlated with PM2.5 concentration are extracted, and the correlation coefficients between the extracted relevant features and PM2.5 concentration are calculated respectively. The relevant features are screened according to the calculated correlation coefficients to obtain the final relevant feature set.
[0059] This example first constructs domain feature engineering, combines the boundary layer height (PBLH) and wind speed (WS) in meteorological data, and calculates the PM2.5 diffusion coefficient DI to reflect the diffusion efficiency of pollutants. For example, the diffusion coefficient can be defined as follows:
[0060] (1)
[0061] Where PBLH is the boundary layer height and WS is the wind speed.
[0062] In addition, combined with atmospheric NO x Concentration, O3 concentration, calculate the atmospheric oxidant mass concentration RI to reflect the efficiency of atmospheric secondary aerosol formation. For example, the atmospheric oxidant mass concentration can be defined as follows:
[0063] (2)
[0064] in, Indicates NO x concentration, Indicates O3 concentration.
[0065] In this embodiment, the relevant characteristics include air pollutant characteristics, meteorological characteristics, pollutant diffusion coefficients, and oxidant mass concentrations. Air pollutant characteristics include characteristics of relevant pollutants such as NO2, SO2, CO, and O3. Meteorological characteristics include temperature, relative humidity, atmospheric visibility, solar radiation intensity, wind speed and direction, boundary layer height, precipitation, etc. The pollutant diffusion coefficient is calculated based on the boundary layer height and wind speed in the meteorological data. The calculation expression can be shown in formula (1). The oxidant mass concentration is calculated based on the NO in the atmosphere. x The concentration and O3 concentration are calculated, and the calculation expression can be shown in formula (2).
[0066] In a specific application embodiment, after obtaining the relevant feature set, pre-processing operations such as missing value checking and normalization processing are also included. For example, the missing values of PM2.5 and related features that are correlated with PM2.5 concentration are first checked. If the missing rate of a single feature is less than or equal to a preset threshold, time-series linear interpolation is used to fill the missing value; if the missing rate of a single feature is greater than the preset threshold, the attribute feature is directly eliminated; then Min-Max normalization processing is performed on the relevant graph features such as air quality pollutant data, meteorological data, diffusion coefficient, and oxidant concentration:
[0067] (3)
[0068] in, represents the feature normalization result, represents the features to be processed, represents the minimum value of the feature to be processed, Indicates the maximum value of the feature to be processed.
[0069] After obtaining the relevant feature set, feature selection techniques are further used to optimize the feature space and eliminate low-correlation or redundant features. Preferably, the Pearson correlation coefficient can be used for feature screening and optimization, calculating the correlation coefficient between relevant air pollutant characteristics (such as NO2, SO2, CO, and O3), meteorological characteristics (such as temperature, relative humidity, atmospheric visibility, solar radiation intensity, wind speed and direction, boundary layer height, precipitation, etc.), pollutant diffusion coefficients, oxidant concentrations, etc., and PM2.5 concentrations.
[0070] Specifically, the absolute value threshold of the correlation coefficient can be set to 0.3. If the calculated correlation coefficient of the relevant feature exceeds the threshold, it is considered to be a feature significantly correlated with PM2.5. Otherwise, it is judged to be a feature with low correlation. The features significantly correlated with PM2.5 are retained, and the features with low correlation are eliminated to reduce the interference of redundancy and noise on the analysis. At the same time, the diffusion coefficient is forced to be retained. , oxidant concentration Furthermore, the variance inflation factor (VIF) can be used to eliminate the impact of multicollinearity. That is, by calculating the variance inflation factor (VIF) of each relevant feature, features with a VIF greater than a preset VIF threshold are eliminated to eliminate multicollinearity.
[0071] Step S03: Calculate the similarity between samples in the initial sample data set using the relevant feature set as a constraint condition, construct a PM2.5 concentration sample similarity matrix under environmental constraints, perform cluster analysis on the PM2.5 concentration sample similarity matrix, obtain the sample categories of the initial sample data set, and calculate the number of samples in each sample category.
[0072] This embodiment comprehensively considers the correlation between the gaseous pollutant characteristics, meteorological characteristics, pollution diffusion coefficient, oxidant concentration and other related characteristics that affect PM2.5 concentration and PM2.5 concentration, calculates the similarity between each sample in the initial sample data set, and then performs cluster analysis based on the similarity. It can fully consider the characteristics of sample data at different times to dynamically cluster samples, thereby improving the accuracy and reliability of sample clustering.
[0073] In this embodiment, the environmental constraints are constructed using the relevant features after screening and optimization, and the similarity between PM2.5 concentration samples is calculated. The similarity between each PM2.5 concentration sample is used to construct a PM2.5 concentration sample similarity matrix under environmental constraints, so as to calculate the similarity between each PM2.5 concentration sample under the influence of the relevant features. For example, the similarity between each PM2.5 concentration sample can be calculated according to the following formula:
[0074] (4)
[0075] in, Represents a sample and samples The similarity value of PM2.5 concentration between 、 Represents a sample and samples PM2.5 concentration, 、 Represents a sample and samples The corresponding set of related features, such as , is the scale parameter for adjusting the similarity in the Gaussian kernel function, Represents the sample optimized by Gaussian kernel function ,sample The Euclidean distance of PM2.5 concentration, Indicates PM2.5 concentration samples ,sample Cosine similarity between related environmental constraints.
[0076] The constructed PM2.5 concentration sample similarity matrix is further clustered to obtain the sample categories of the initial sample data set and the number of samples in each sample category. In this embodiment, the PM2.5 concentration similarity matrix can be input into the affinity propagation algorithm model (Affinity Propagation), and the affinity propagation algorithm model is used to adaptively cluster the PM2.5 concentration sample similarity matrix, which can quickly and accurately analyze the number of samples in each cluster. The affinity propagation algorithm automatically determines the cluster center through the message passing mechanism, and uses the responsibility value (Responsibility, ) and availability value (Availability, ) These two kinds of information are updated alternately and iteratively until they converge.
[0077] Specifically, the steps for adaptive clustering the similarity matrix of PM2.5 concentration samples using the affinity propagation algorithm model are as follows:
[0078] First, initialize the responsibility value and availability value. All responsibility values and availability values are initialized to 0.
[0079] (5)
[0080] in, is the responsibility value, indicating that the sample For the sample The responsibility of being selected as a cluster center, Represents a sample Accept samples Availability as cluster centers.
[0081] Then, the responsibility value and availability value are calculated iteratively until the change of the responsibility value and the availability value is less than the predefined threshold value. The iteration ends and the cluster center is determined.
[0082] The responsibility value update formula is:
[0083] (6)
[0084] The formula for updating the availability value is:
[0085] (7)
[0086] At the end of the iteration, the changes in the responsibility value and availability value at all time points are less than the threshold When , the convergence condition is reached, the iteration stops, and the final clustering result is output. At this time, the cluster center is the time point with the highest responsibility value, which can be determined by the following formula:
[0087] (8)
[0088] in, is the final cluster center, for each time point , assign it to the cluster center with the largest responsibility value The cluster to which it belongs and outputs the time point The cluster number to which it belongs.
[0089] In this embodiment, after clustering is completed, the number of samples under each cluster number is calculated according to the cluster sample results, and the clustering results are set as , each cluster Include time point samples, where Indicates the cluster number.
[0090] (9)
[0091] in, It is clustering The number of samples, T is the total length of the time series, is the indicator function, when the time point PM2.5 concentration Belong to cluster , the function value is 1, otherwise it is 0.
[0092] After clustering is completed, this embodiment further includes calculating a sample distribution balance value based on the number of samples in each cluster. Based on the calculated sample distribution balance value, it is determined whether to proceed to step S04 to trigger sample expansion. The sample distribution balance value is used to characterize the degree of balance of the sample distribution, enabling dynamic identification of the sample distribution balance state and timely and dynamic triggering of sample expansion when the sample distribution is unbalanced. For example, sample expansion is triggered when the calculated sample distribution balance value is less than a preset balance threshold.
[0093] In a specific application embodiment, the calculation expression of the sample distribution balance value can be expressed as follows:
[0094] (10)
[0095] in, B is the balance index of data distribution, is the i-th cluster The number of samples, is the average number of samples in all clusters, is the total number of sample clusters. The closer B is to 1, the more balanced the data distribution is; the closer it is to 0, the more severely imbalanced the data sample distribution is. It is preferable to define sample expansion as triggering when B < 0.6.
[0096] Step S04: The initial sample data set is divided into multiple input samples, and each input sample is input into the adversarial network for model training. After the training is completed, a generative adversarial network model is obtained to generate PM2.5 concentration samples. The PM2.5 concentration samples include PM2.5 concentration and corresponding related feature sets.
[0097] In this embodiment, the initial sample data set is first split to obtain multiple input samples, and then each input sample is input into the adversarial network for model training. During the model training process, the loss function is optimized through physical and chemical rule constraints to establish a generative adversarial network model that can generate high-quality samples.
[0098] Specifically, all the initial sample data sets after data preprocessing and screening are segmented to convert the continuous time series data into multiple input samples of the adversarial network. Each input sample is a data segmentation window. Each input sample can be represented as ,in Indicates at a point in time The PM2.5 concentration at that time and the related gaseous pollutant characteristics, meteorological characteristics, pollution diffusion coefficient, oxidant concentration and other data.
[0099] To ensure that the temporal relationship is learned, the generator in this embodiment uses a long short-term memory network (LSTM) to capture the temporal characteristics of the time series. And random noise is input into the LSTM layer, and a fully connected layer is added after the LSTM layer to map the time series features extracted by LSTM to the output space to obtain the output , That is, the generator The predicted value at a time point is a multidimensional vector containing PM2.5 concentration and related gaseous pollutant characteristics, meteorological characteristics, pollution diffusion coefficient, and oxidant concentration. Its goal is to be closer to the real data. To prevent overfitting, the fully connected layer uses dropout regularization to avoid overfitting. In the construction of the discriminator, this embodiment uses a multi-layer perceptron network to process time series data. The discriminator has three hidden layers, each with 10 neurons. ReLU is used as the activation function between hidden layers, and a sigmoid function is used to output the discrimination result. The discriminator is used to distinguish whether the sample is from real data or pseudo data generated by the generator.
[0100] In this example, the adversarial network involves multiple output targets, including PM2.5 concentration and related gaseous pollutant characteristics, meteorological characteristics, pollution diffusion coefficient, oxidant concentration, and so on. The discriminator needs to be able to evaluate the authenticity of the generated samples for each target, with a primary focus on PM2.5 concentration. The remaining targets assist in PM2.5 generation. Therefore, this example introduces weight parameters to adjust the importance of each target. Based on the use of binary cross entropy as the generator and discriminator loss of the GAN, this example combines the Wasserstein distance in the WGAN to maximize the difference between the scores of real and generated samples, effectively reducing the problem of mode collapse and facilitating model training convergence.
[0101] Specifically, WANG's generator loss function is The design is as follows:
[0102] (11)
[0103] in, is the expected value of the output target, Indicates the number of multi-target output features, It is a feature Preferably, in order to highlight the greatest importance of generating target PM2.5, the importance of PM2.5 can be set to 0.5 (configurable), and the rest of the target features are . Represents the discriminator pair consisting of conditional variables and random noise Features in the sample generated after input The judgment result of Represents the distribution that generated the sample data.
[0104] Furthermore, in the training of the generative adversarial network, this embodiment introduces physical rule constraints to optimize the generator loss function. By designing a high-concentration penalty term, the PM2.5 concentration generated by the generator is limited to not exceed the historical peak range, avoiding the generation of unrealistic extremely high pollution samples. This can ensure that the generated PM2.5 pollution data not only meets the statistical distribution laws, but also conforms to the physical mechanisms in environmental science.
[0105] Specifically, the PM2.5 concentration value of the generated sample , if it exceeds the historical peak threshold (For example, 150μg / m³), a secondary penalty will be imposed on the excess, namely:
[0106] (12)
[0107] in, is the penalty weight (e.g. 0.1~0.3), which is used to balance the strength of the adversarial loss and the physical constraint. To generate the PM2.5 concentration value of the sample, The historical peak threshold of PM2.5 concentration value for generating samples.
[0108] Based on the above loss function, this embodiment further imposes relevant physical and chemical constraints, such as configurable correlation constraints:
[0109] (13)
[0110] in, To generate PM2.5 concentration data, 、 are the diffusion coefficient and oxidant concentration, respectively, is the covariance, when >0 or <0, physical penalty items are applied , chemical penalty items ,and 、 It can be 0.1~0.3, It is a small positive number that allows for slight fluctuations.
[0111] As shown in formula (13), by adopting the above constraints, this embodiment can ensure that the PM2.5 concentration of the generated sample is negatively correlated with the pollution diffusion coefficient, which conforms to the physical law that the larger the pollution diffusion coefficient, the stronger the pollutant diffusion ability and the lower the PM2.5 concentration; at the same time, it can ensure that the PM2.5 concentration of the generated sample is positively correlated with the oxidant concentration, that is, it conforms to the chemical law that the larger the oxidant concentration, the stronger the PM2.5 secondary generation ability and the higher the PM2.5 concentration.
[0112] In summary, in this embodiment, when each input sample is input into the adversarial network for model training, the total loss function of the generator is for:
[0113] (14)
[0114] in, represents the generator loss function of WANG, represents a high concentration penalty term used to impose penalties on samples generated when the PM2.5 concentration exceeds the historical peak threshold. The correlation constraint condition is expressed so that the PM2.5 concentration of the generated sample is negatively correlated with the pollution diffusion coefficient and the PM2.5 concentration of the generated sample is positively correlated with the oxidant concentration.
[0115] In this embodiment, in the framework of the Generative Adversarial Network (GAN), the core task of the discriminator is to distinguish between real data and generated data. The corresponding loss function can be configured to not directly include physical and chemical rule constraints. For example, the loss function of the discriminator can be designed as:
[0116] (15)
[0117] in, is the loss function of the discriminator, Indicates the number of multi-target output features, It is a feature The weight of the target feature is set to 0.5 (configurable) to highlight the importance of generating the target PM2.5. . Represents the discriminator pair consisting of conditional variables and random noise Features in the sample generated after input The judgment result of Represents the discriminator's response to the real sample Medium Features The judgment result of represents the distribution of generated sample data, Represents the distribution of real sample data.
[0118] When training the aforementioned adversarial network, the discriminator and generator iterate alternately. After each discriminator iteration, the generator iterates once, and this cycle repeats until the probability of the discriminator making a correct or incorrect judgment reaches a preset value (for example, 0.5). At this point, the optimal generative adversarial network model is obtained.
[0119] Step S05: Generate new samples for each sample category using the trained generative adversarial network model, and determine the number of new samples to be generated based on the number of samples in each sample category. The new samples include PM2.5 concentration data and corresponding related feature sets.
[0120] After cluster analysis, this embodiment obtains the number of samples in each category, that is, the clustering result is categories, each category Include To determine the expansion priority of each category, this embodiment uses an inverse proportional function to convert the number of samples in each category into an expansion weight. is the number of samples of the category with the largest number of samples among all categories, that is:
[0121] (16)
[0122] The expansion weight of each category is calculated according to the number of samples in each sample category:
[0123] (17)
[0124] in, represents the expansion weight of the i-th sample category, represents the number of samples of the i-th sample category, Indicates the number of samples of the category with the largest number of samples among all categories.
[0125] By determining the expansion weights in the above manner, it is possible to ensure that categories with fewer samples receive more expansion resources, while categories with more samples receive less expansion, thereby improving the quality of sample generation.
[0126] In this embodiment, an initial number of sample expansion is preset , that is, the total number of samples that need to be generated, and the amount of expansion is adjusted through the feedback mechanism in the future. Use the expansion weights of each category Calculate the number of new samples required to generate for each sample category:
[0127] (18)
[0128] in, Indicates the number of new samples required to be generated for the i-th sample category, Indicates the preset initial number of sample expansions, Indicates the The expansion weight of the sample category, is the sum of all the category expansion weights. According to the above method, the number of expansion samples of each category can be reasonably allocated according to the category expansion weight, thereby achieving a balance between categories.
[0129] In this embodiment, the steps of using the trained generative adversarial network model to generate new samples for each sample category include:
[0130] Step S511: Randomly select time points: randomly select a corresponding number of time points from the current samples of each sample category according to the number of new samples required to be generated for each sample category, and use them as reference time points for generating new samples.
[0131] Specifically, the random number generator function rand() can be used to generate the Randomly select time points from the existing samples in . Each category Contains several time series data samples ,in For time point, For the time point The collected multi-dimensional characteristic data includes PM2.5 concentration and related gaseous pollutant characteristics, meteorological characteristics, pollution diffusion coefficient, oxidant concentration, etc. Function, from the category Random selection time points . Each randomly selected time point As the benchmark time point for generating new samples. The random selection process can be expressed as:
[0132] (19)
[0133] in, For the Randomly selected time points, For category The number of expanded samples.
[0134] Step S512: Generate new samples based on randomly selected time points: At each selected time point, obtain the historical environmental data and random noise of each selected time point, the historical environmental data includes air quality and meteorological data, and input the historical environmental data and random noise into the trained generator model to generate a complete new sample data.
[0135] For each time from the category Randomly selected time points , using the historical data at that moment and random noise to generate a new complete data sample. Specifically, for the randomly selected time point , extract historical data before that moment, including the air quality and meteorological data at that moment, for example: ,in Indicates a time point Air quality and weather data, Divide the time series into windows. Previous historical data and random noise Input to the trained generator model The generator generates a complete sample data based on these inputs , which includes new air quality and meteorological data.
[0136] Specifically, the generated new sample can be expressed as:
[0137] (20)
[0138] Through the above steps, this embodiment, combined with the adversarial network, can generate high-quality virtual samples of high-quality PM2.5 sample data that conform to the time series rules, actual distribution and dynamic changes, effectively expand the training data set, solve the problem of insufficient training of PM2.5 high pollution concentration samples in traditional machine learning, and improve the model's prediction and generalization capabilities for low-frequency and extreme pollution events.
[0139] Step S06: Add the generated new samples to the initial sample data set to form an expanded sample set, and use the expanded sample set to train the machine learning model to obtain a PM2.5 concentration prediction model for PM2.5 concentration prediction. During the model training process, the relevant feature set corresponding to the PM2.5 concentration is used as a prediction factor, and the PM2.5 concentration is used as the target variable.
[0140] Specifically, for each randomly selected time point , the generator will generate a new complete sample based on historical data and random noise and add it to the category By doing this process for each category, we can obtain the required expanded samples for all categories. , repeat the above process times, generate new samples, and then compare all the generated new samples with the original category The data sets are merged to form an expanded data set, which can be expressed as:
[0141] (twenty one)
[0142] Finally, all the expanded datasets of different categories will be merged to form the final expanded dataset. , where m is the total number of categories.
[0143] To determine the optimal number of augmented samples, a grid search algorithm can be used to optimize different augmentation amounts. The grid search algorithm trains multiple machine learning models and evaluates their performance over a preset range of augmented sample sizes to determine the most suitable number of augmented samples.
[0144] For each preset number of augmented samples , based on the expanded dataset Train a machine learning model. The target variable is PM2.5 concentration, and the predictors are gaseous pollutant characteristics, meteorological characteristics, pollution diffusion coefficient, and oxidant concentration. The model uses these factors to predict PM2.5 concentration at future times. Machine learning models include but are not limited to support vector machines (SVMs), random forests, and gradient boosting.
[0145] For each trained model, its performance on the test set is further evaluated to verify the model's generalization ability. This evaluation consists of two aspects: first, a performance evaluation of the entire test set that was not used for training; second, a special evaluation of the model's predictive ability under specific circumstances, particularly for high-concentration PM2.5 pollution events (i.e., events where PM2.5 concentration exceeds a certain threshold). In this example, the PM2.5 concentration threshold of 75 µg / m3, the second-level air quality standard, is used. Daily average values exceeding 75 µg / m3 are considered pollution events:
[0146] (twenty two)
[0147] (twenty three)
[0148] in, is the root mean square error of all test set data, is the root mean square error of all daily mean concentrations exceeding 75 µg / m3. is the predicted value of the model, is the true value.
[0149] This embodiment obtains different numbers of expanded samples through a grid search process. Corresponding performance indicators 、 , compare these performance indicators with the corresponding number of expanded samples, and finally select the number of expanded samples with the best performance as the final optimization result, and then expand the training data set according to the optimal number of expanded samples to train the PM2.5 concentration prediction model and improve the model.
[0150] This embodiment also provides an electronic device, including a microprocessor and a memory connected to each other, wherein the microprocessor is programmed or configured to execute a PM2.5 pollution prediction method based on adaptive sample expansion.
[0151] Those skilled in the art will appreciate that the above-described embodiments of the present invention may be provided as methods, systems, or computer program products. Therefore, the present invention may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code. The present invention is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the functions described in the processes. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including the instruction device, which implements the function specified in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The above is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as above with preferred embodiments, it is not intended to limit the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the scope of protection of the technical solution of the present invention.
[0152] The above description is merely a preferred embodiment of the present invention and does not constitute any form of limitation to the present invention. Although the present invention has been disclosed above with reference to the preferred embodiment, it is not intended to limit the present invention. Therefore, any simple modifications, equivalent variations, and modifications to the above embodiment that do not depart from the technical solution of the present invention and are based on the technical essence of the present invention shall fall within the scope of protection of the technical solution of the present invention.
Claims
1. A PM2.5 pollution prediction method based on adaptive sample expansion, characterized in that: The following steps are involved: Step S01: Acquire PM2.5 concentration data of a target prediction area at different times and an environmental data set corresponding to the PM2.5 concentration to form an initial sample data set, wherein the environmental data set includes air pollutant concentration data, meteorological data, and time information corresponding to the air pollutant concentration data and meteorological data; Step S02: After preprocessing the environmental data set, relevant features that are correlated with PM2.5 concentration are extracted, and correlation coefficients between the extracted relevant features and PM2.5 concentration are calculated respectively. The relevant features are screened according to the calculated correlation coefficients to obtain a final set of relevant features; Step S03: Calculating the similarity between samples in the initial sample data set using the relevant feature set as a constraint condition, constructing a PM2.5 concentration sample similarity matrix under environmental constraints, performing cluster analysis on the PM2.5 concentration sample similarity matrix to classify the sample data in the initial sample data set, obtaining sample categories of the initial sample data set, and calculating the number of samples in each sample category; Step S04: Segmenting the initial sample data set to obtain multiple input samples, inputting each input sample into a generative adversarial network for model training, and obtaining a generative adversarial network model after the training is completed to generate PM2.5 concentration samples, wherein the PM2.5 concentration samples include PM2.5 concentration and the corresponding related feature set; Step S05: Generate new samples for each of the sample categories using the trained generative adversarial network model, and determine the number of new samples to be generated based on the number of samples in each sample category, wherein the new samples include PM2.5 concentration data and corresponding related feature sets; Step S06: Add the generated new samples to the initial sample data set to form an expanded sample set, and use the expanded sample set to train the machine learning model to obtain a PM2.5 concentration prediction model for realizing PM2.5 concentration prediction. During the model training process, the relevant feature set corresponding to the PM2.5 concentration is used as a prediction factor, and the PM2.5 concentration is used as the target variable.
2. The PM2.5 pollution prediction method based on adaptive sample expansion according to claim 1, characterized in that: The relevant characteristics include air pollutant characteristics, meteorological characteristics, pollutant diffusion coefficient and oxidant mass concentration. The air pollutant characteristics include characteristics of any one or more pollutants among NO2, SO2, CO and O3. The meteorological characteristics include any one or more of temperature, relative humidity, atmospheric visibility, solar radiation intensity, wind speed and direction, boundary layer height, and precipitation. The pollutant diffusion coefficient is calculated based on the boundary layer height and wind speed in the meteorological data. The oxidant mass concentration is calculated based on the NO in the atmosphere. x The concentration and O3 concentration are calculated.
3. The PM2.5 pollution prediction method based on adaptive sample expansion according to claim 1, characterized in that: In step S03, the similarity between samples is calculated according to the following formula: in, Represents a sample and samples The similarity value of PM2.5 concentration between 、 Represents a sample and samples PM2.5 concentration, 、 Represents a sample and samples The corresponding set of relevant features, It is the scale parameter for adjusting the similarity in the Gaussian kernel function.
4. The PM2.5 pollution prediction method based on adaptive sample expansion according to claim 1, characterized in that: Step S03 also includes: after clustering is completed, calculating the sample distribution balance value according to the number of samples in each cluster, and determining whether to proceed to step S04 to trigger sample expansion based on the calculated sample distribution balance value. The sample distribution balance value is used to characterize the balance degree of the sample distribution. The calculation expression of the sample distribution balance value is: in, B is the balance index of data distribution, is the number of samples of the i-th sample category, is the average number of samples in all clusters, is the total number of sample categories.
5. The PM2.5 pollution prediction method based on adaptive sample expansion according to claim 1, characterized in that: In step S04, when each input sample is input into the adversarial network for model training, the total loss function of the generator is for: in, represents the generator loss function of WGAN, represents a high concentration penalty term used to impose penalties on samples generated when the PM2.5 concentration exceeds the historical peak threshold. Represents a correlation constraint condition for making the PM2.5 concentration of the generated sample negatively correlated with the pollution diffusion coefficient and positively correlated with the oxidant concentration; The loss function of the discriminator is: in, is the loss function of the discriminator, Indicates the number of multi-target output features, It is a feature The weight of Represents the discriminator pair consisting of conditional variables and random noise Features in the sample generated after input The judgment result of Represents the discriminator's response to the real sample Medium Features The judgment result of represents the distribution of generated sample data, represents the distribution of real sample data, is the expected value of the output target.
6. The PM2.5 pollution prediction method based on adaptive sample expansion according to claim 5, characterized in that: The calculation expression of the generator loss function of WGAN is: Where, is the expected value of the output target, Indicates the number of multi-target output features, It is a feature The weight of Represents the distribution of generated sample data; The calculation expression of the high concentration penalty term is: Where, is the penalty weight used to balance the strength of the adversarial loss and the physical constraint, To generate the PM2.5 concentration value of the sample, The historical peak threshold of PM2.5 concentration value for generating samples; The calculation expression of the constraint condition of the correlation relationship is: Where, 、 are the diffusion coefficient and oxidant concentration, respectively, is the covariance, when >0 or <0, physical penalty items are applied , chemical penalty items , is a positive number.
7. The PM2.5 pollution prediction method based on adaptive sample expansion according to any one of claims 1 to 6, characterized in that: In step S05, the expansion weight of each category is calculated according to the number of samples in each sample category: in, represents the expansion weight of the i-th sample category, represents the number of samples of the i-th sample category, Indicates the number of samples of the category with the largest number of samples among all categories; Use augmented weights for each category Calculate the number of new samples required to generate for each sample category: in, Indicates the number of new samples required to be generated for the i-th sample category, Indicates the preset initial number of sample expansions, Indicates the The expansion weight of the sample category, is the total number of sample categories.
8. The PM2.5 pollution prediction method based on adaptive sample expansion according to any one of claims 1 to 6, characterized in that: In step S05, generating new samples for each sample category using the trained generative adversarial network model includes: Step S511: randomly selecting a corresponding number of time points from the current samples of each sample category according to the number of new samples required to be generated for each sample category, and using them as reference time points for generating new samples; Step S512: At each selected time point, obtain historical environmental data and random noise at each selected time point, wherein the historical environmental data includes air quality and meteorological data, and input the historical environmental data and random noise into the trained generator model to generate a complete new sample data.
9. The PM2.5 pollution prediction method based on adaptive sample expansion according to any one of claims 1 to 6, characterized in that: In step S03, an affinity propagation algorithm model is used to perform cluster analysis on the similarity matrix of PM2.5 concentration samples; step S06 also includes obtaining model performance indicators corresponding to different expanded sample numbers through grid search, comparing the model performance indicators corresponding to different expanded sample numbers, and selecting the expanded sample number that best corresponds to the model performance indicator as the final expanded sample number.
10. An electronic device comprising a microprocessor and a memory connected to each other, characterized in that: The microprocessor is programmed or configured to execute the PM2.5 pollution prediction method based on adaptive sample expansion as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Electric power system transient stability evaluation method considering sample imbalance based on artificial intelligence
CN118095050A
Ship radiation noise enhancement method and system based on generative adversarial network
CN114201995A
Power distribution network distribution robust optimization scheduling method based on conditional generative adversarial network
CN117291292A