Method and system for identifying high value of atmospheric monitoring standard station based on random forest
Through the adaptive partitioning model and threshold recognition method based on the random forest algorithm, the hysteresis and subjectivity of high-value recognition in atmospheric environmental monitoring are solved, and the automatic identification and real-time monitoring of high-values of standard stations are realized, which improves data accuracy and management efficiency.
Patent Information
- Application Number
- CN202510558501.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, high-value identification of atmospheric environment monitoring mainly relies on artificial methods, with subjectivity and lag, and real-time identification and screening cannot be achieved.
Using a method based on the random forest algorithm, a number of pollutant concentration data of the standard station are obtained, and an adaptive partition model is inputted after preprocessing. A random forest recognition model based on the threshold is established to automatically identify the high value of real-time pollutant concentrations.
It realizes automatic identification of high values of atmospheric monitoring standard stations, improves data accuracy and reliability, can quickly identify pollution events and pollution sources, and optimizes air quality management.
Smart Images

Figure CN120448992A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of environmental monitoring, and in particular relates to a method and system for identifying high values of atmospheric monitoring standard stations based on random forest. Background Art
[0002] The atmospheric environment monitoring standard station can monitor the concentration changes of conventional pollution factors such as SO2, nitrogen oxides (NOx, NO2, NO), O3, CO, particulate matter (PM2.5, PM10) in the ambient air in real time, and provide accurate and real-time data for air quality evaluation through continuous automatic sampling and measurement of atmospheric environmental quality. In atmospheric monitoring, some high or abnormal values of atmospheric pollutants are often monitored, and the real-time identification of high values can quickly identify pollution events or pollution sources, especially when the concentration rises sharply in a short period of time, which plays an important role in taking countermeasures. At the same time, high-value identification helps to detect and correct potential problems in monitoring data, such as sensor failure, data entry errors or other technical problems, thereby improving the accuracy and reliability of the data. By analyzing and identifying high-value data, pollution sources can be investigated and air quality management can be optimized.
[0003] Currently, high-value identification in atmospheric environmental monitoring is primarily done manually, relying on work experience and personal judgment to identify high-value sites. This is often subjective. Furthermore, manual identification often has a certain lag, making real-time identification and screening impossible. Summary of the Invention
[0004] In order to overcome the problems existing in the above-mentioned prior art, the present invention provides a method and system for finding high values of atmospheric monitoring pollutant concentrations based on a random forest algorithm, thereby realizing automatic identification of high values of standard stations.
[0005] In a first aspect, the present invention proposes a method for identifying high values of atmospheric monitoring standard stations based on random forest, which specifically includes the following steps:
[0006] Obtain concentration data based on the original data of multiple pollutants at standard stations;
[0007] The concentration data is pre-processed and then input into an adaptive partitioning model based on spatial constraints to generate high-value identification data;
[0008] Establishing a threshold-based recognition model based on a random forest model using the high-value recognition data;
[0009] The identification model is used to identify high values of pollutant concentrations at standard stations obtained in real time.
[0010] Furthermore, the step of obtaining concentration data based on the original data of multiple pollutants at the standard station includes obtaining the concentration data of multiple pollutants at different standard stations in real time through the data interface corresponding to the standard station.
[0011] Furthermore, the step of pre-processing the concentration data and inputting the pre-processed concentration data into an adaptive partitioning model based on spatial constraints to generate high-value identification data includes:
[0012] After splitting the concentration data, cleaning missing values and standardizing them, concentration data values corresponding to different pollutant types based on time series identification are generated;
[0013] The concentration data values are used to generate high-value identification data for different site areas based on an adaptive partitioning model with spatial constraints;
[0014] Wherein, the adaptive partition model is:
[0015]
[0016] Among them, (x i ,y i )、(x j ,y j ) are the latitude and longitude positions of sites i and j, where i is the latitude and longitude of the regional center site, C i 、C j is the average concentration of sites i and j, C' i , C' j is the normalized concentration data value, where μ c is the mean concentration, σ c is the standard deviation, α is the mixed distance weight, when d ij Those below the city-wide threshold are considered to be in the same area.
[0017] Furthermore, the step of establishing a threshold-based recognition model based on a random forest model using the high-value recognition data includes:
[0018] A feature matrix is created based on the high-value identification data, using different pollutant indices as features, and divided into training and test sets. Training and testing are performed based on the random forest algorithm. The cross-validation model used in the test is:
[0019]
[0020] Among them, L(θ) is the cross-validation accuracy, is the accuracy evaluation index, is the true label of the k-fold validation set, To predict the k-fold validation set under the parameter θ, we compare L(θ) for all θ and select the parameter with the highest L(θ) value.
[0021] Dynamic thresholds corresponding to different pollutant types are set in the random forest model that has completed training and testing, and a recognition model is established based on the dynamic thresholds.
[0022] Furthermore, the pollutants include PM 2.5 、PM 10 , NO2, SO2, CO and O3.
[0023] Furthermore, the steps of setting dynamic thresholds corresponding to different pollutant types and establishing a recognition model based on the dynamic thresholds include:
[0024] If the pollutant is NO2, SO2, CO, or O3, a first identification model is established. When the measured value of the pollutant is greater than 1.7 times the predicted value, the identification result is set to a high value.
[0025] If the pollutant is PM 2.5 、PM 10 , then a second recognition model is established, and the second recognition model outputs the recognition result based on the following formula:
[0026]
[0027] Among them, C obs Indicates the measured concentration value.
[0028] In a second aspect, the present invention further proposes an identification system for the aforementioned high-value identification method for atmospheric monitoring standard stations based on random forest, comprising:
[0029] An acquisition module is used to obtain concentration data based on the original data of multiple pollutants at the standard station;
[0030] A preprocessing module, configured to preprocess the concentration data and then input the preprocessed data into a spatial constraint-based adaptive partitioning model to generate high-value identification data;
[0031] A recognition model generation module, configured to establish a threshold-based recognition model based on a random forest model using the high-value recognition data;
[0032] The processing module is used to identify high values of the pollutant concentration of the standard station obtained in real time through the identification model.
[0033] In a third aspect, the present invention proposes an electronic device comprising: one or more processors; and a memory associated with the one or more processors, the memory being used to store program instructions, which, when read and executed by the one or more processors, execute the steps of the random forest-based high-value identification method for atmospheric monitoring standard stations described in the first aspect.
[0034] In a fourth aspect, the present invention further proposes a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the random forest-based high-value identification method for atmospheric monitoring standard stations described in the first aspect.
[0035] This paper proposes a random forest-based method and system for identifying high-value atmospheric concentrations at standard stations. This method obtains concentration data based on raw data from multiple pollutants at standard stations, preprocesses it, and then inputs it into an adaptive partitioning model based on spatial constraints to generate high-value identification data. A threshold-based recognition model is then established based on the random forest model. Finally, this recognition model identifies high-value concentrations at standard stations based on real-time pollutant concentrations. This method automatically identifies high-value concentrations at standard stations through the high-value data-based recognition model. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 This is a flow chart of a method for identifying high values of atmospheric monitoring standard stations based on random forests provided by an embodiment of the present invention.
[0037] Figure 2 This is a structural diagram of a high-value identification system for atmospheric monitoring standard stations based on random forests provided in an embodiment of the present invention.
[0038] Figure 3 Schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0040] In response to the problems existing in the prior art, the present invention provides a method and system for finding high values of atmospheric monitoring pollutant concentrations based on a random forest algorithm. The present invention is described in detail below with reference to the accompanying drawings.
[0041] like Figure 1 As shown, this embodiment provides a method for identifying high values of atmospheric monitoring standard stations based on random forest, the method comprising:
[0042] S101. Obtaining concentration data based on the original data of multiple pollutants at the standard station;
[0043] S102. Preprocessing the concentration data and inputting it into a spatially constrained adaptive partitioning model to generate high-value identification data;
[0044] S103. Establishing a threshold-based recognition model based on the random forest model using the high-value recognition data;
[0045] S104. Using the identification model, high-value identification is performed on the pollutant concentrations of the standard stations obtained in real time.
[0046] In this example, a random forest-based recognition model is established to identify high-value pollutants monitored in real time. This requires the construction of a random forest-based algorithm model, which is trained and tested using acquired data. This process requires the collection and preprocessing of raw data, which serves as input for model training and testing. In this example, the raw data consists of multiple pollutant concentration data from standard stations, imported in real time through a corresponding database interface.
[0047] In step S102, the data cleaning process is implemented by preprocessing the raw data through splitting and standardization, thereby screening out the data types that meet the requirements of random forest. The concentration data obtained through different database interfaces are of large types and data volumes, so these data need to be categorized.
[0048] By splitting and standardizing the raw data, cleaning missing values, and reading the concentrations of various pollutants, we can effectively address the problem of missing data during uncertain periods in standard station monitoring, thereby avoiding the inability to identify and program errors caused by missing data. The goal of the data processing process is to achieve data standardization, thereby facilitating data reading and subsequent random forest learning.
[0049] In this embodiment, the string "---" is usually used to represent missing data. However, during the processing, this string cannot be recognized. Therefore, the string "NaN" is used to represent it. The implementation code is as follows:
[0050] group_data[column]=group_data[column].replace('---',np.nan)
[0051] At the same time, a numerical conversion is performed, forcing the column data to a numeric type (pd.to_numeric). Unconvertible values are set to NaN. In addition, invalid rows are deleted. For each feature column, dropna is called to delete rows containing NaN to ensure that all features in each data set are valid values. Finally, the time column is converted to the datetime type and set as the index to facilitate time series identification.
[0052] In this embodiment, the main purpose of data splitting is to group the data to facilitate subsequent learning and high-value identification. Specifically, an adaptive partitioning model is constructed to generate the high-value identification data required by the random forest based on the concentration data of different site locations and different sites, so as to carry out training and testing.
[0053] Specifically, concentration data is grouped by pollutant type and site number, dividing site data into different regions to mitigate potential data discrepancies caused by different background environments. Loop statements are set as needed to traverse and clean the data for each group. This allows for grouping sites based on their geographical location and local differences in background concentration characteristics, allowing for regional screening of high values and avoiding misjudgments of long-term high values at some sites.
[0054] In this embodiment, the adaptive partition model is:
[0055]
[0056] Among them, (x i ,y i )、(x j ,y j ) are the latitude and longitude positions of sites i and j, where i is the latitude and longitude of the regional center site, C i 、C j is the average concentration of sites i and j, C' i , C' j is the normalized concentration data value, where μ c is the mean concentration, σ c is the standard deviation, α is the mixed distance weight, when d ij Those below the city-wide threshold are considered to be in the same area.
[0057] After obtaining high-value identification data, it is input into a random forest model for training and testing. In this example, a feature matrix X is created, using different pollutant indices as features, and y represents the input data for the current column, X = [pollutant1, pollutant2, …, station_id, timestamp]. The data is divided into training and test sets, with the test set accounting for 20%. The test data is randomly selected to enhance random classification diversity through random selection. The model is trained on the dataset through multiple learning cycles, optimizing internal parameters and generating model predictions. The model uses each pollutant, monitoring time, and station number as input features, grouping them by station number. Each grouped data is then divided into separate datasets for learning. Using the random forest algorithm, a decision tree model for predicting relevant pollutant data is established. Using a self-service sampling method, multiple subsets are randomly selected from the original dataset for training and learning, generating different predictions based on the different learning results. All prediction results are then integrated and averaged to obtain the final output. The generated results are evaluated using the mean squared error and coefficient of determination. The two key parameters in the random forest algorithm are the number of decision trees (n_estimators) and the random seed (random_state). The number of decision trees is selected by traversing (n_estimators = 10 to 200) and multiple cross-validations to select the n_estimators value that optimizes model performance. The random seed is selected by trying different integer values to select the appropriate parameter value.
[0058] The expression of the cross-validation model is:
[0059]
[0060] Among them, L(θ) is the cross-validation accuracy, is the accuracy evaluation index, is the true label of the k-fold validation set, To predict the k-fold validation set under the parameter θ, we compare L(θ) for all θ and select the parameter with the highest L(θ) value.
[0061] Dynamic thresholds corresponding to different pollutant types are set in the random forest model that has completed training and testing, and a recognition model is established based on the dynamic thresholds. It should be noted that the pollutants in this embodiment include: PM 2.5 、PM 10 , NO2, SO2, CO and O3, a total of six categories.
[0062] In this embodiment, the model predictions are compared with the measured data, with different thresholds set for different pollutant types and dynamic thresholds set for different concentration ranges of the same pollutant. The results are output as a Boolean array. In practice, since the measured concentrations of NO2, SO2, CO, and O3 are all relatively low, a first recognition model is constructed based on these three factors. The recognition process of this first recognition model is as follows: when the measured value is greater than 1.7 times the predicted value, it is identified as a high value. PM 2.5 、PM 10 The concentrations of the two pollutants often range from single digits to several hundred, with a relatively large range of variation. Therefore, thresholds are set according to the concentration range to avoid misjudging high values at low concentrations and missing high values at high concentrations.
[0063] Grid search and cross-validation are used to optimize the threshold parameters, with the goal of minimizing the false positive rate (FP) and the false negative rate (FN):
[0064]
[0065] Among them, weight: ω1+ω2=1, α j ∈[1.2, 2.0], α j Indicates the dynamic threshold of a certain pollutant;
[0066]
[0067] represents the adjusted dynamic threshold, represents the initial setting threshold. In this embodiment, α is dynamically adjusted according to the difference between FP and FN. j If FP>FN (too many false positives), it is possible to reduce α j To suppress FP; otherwise, increase α j , η is the step size.
[0068] Therefore, based on PM 2.5 、PM 10 Construct the second recognition model as follows:
[0069]
[0070] The specific content of the second recognition model is:
[0071] For PM 2.5 : When the measured concentration is less than 30 micrograms, the measured value is greater than 1.8 times the predicted value and is identified as a high value; when the measured concentration is between 30-60 micrograms, the measured value is greater than 1.4 times the predicted value and is identified as a high value; when the measured concentration is greater than 70 micrograms, the measured value is greater than 1.6 times the predicted value and is identified as a high value.
[0072] For PM 10: When the measured concentration is less than 50 micrograms, the measured value is greater than 1.8 times the predicted value and is identified as a high value; when the measured concentration is greater than 50 micrograms, the measured value is greater than 1.4 times the predicted value and is identified as a high value.
[0073] In this embodiment, a complete high-value identification model based on random forests is constructed through the above content, and accurate identification of high values is achieved through training and testing. Step S104 of this embodiment, that is, when the pollutant concentration data of a certain site or certain sites is received in real time, it is input into the constructed model for identification, thereby determining the high value. In this embodiment, a Boolean mask can be set to represent the output result. For example, if the Boolean mask is "1", it is determined to be a high value, corresponding to the "True" element in the table where the data is located, and the element marked as "False" indicates that the data is not a high value and can be excluded. The above content can accurately and automatically identify the high-value data of the standard station.
[0074] like Figure 2 As shown, the random forest-based high-value identification system 200 for atmospheric monitoring standard stations provided in an embodiment of the present invention includes:
[0075] An acquisition module 201 is used to acquire concentration data based on multiple pollutant raw data of a standard station;
[0076] A preprocessing module 202 is used to preprocess the concentration data and then input it into an adaptive partitioning model based on spatial constraints to generate high-value identification data;
[0077] A recognition model generation module 203 is configured to establish a threshold-based recognition model based on a random forest model using the high-value recognition data;
[0078] The processing module 204 is configured to identify high values of the pollutant concentrations of the standard stations obtained in real time by using the identification model.
[0079] Figure 3 An example of a physical structure diagram of an electronic device is shown below. Figure 3 As shown, the electronic device 300 may include: a processor 310, a communication interface 320, a memory 330, and a communication bus 340, wherein the processor 310, the communication interface 320, and the memory 330 communicate with each other via the communication bus 340. The processor 310 may call the logic instructions in the memory 330 to execute the above-mentioned random forest-based high value identification method for atmospheric monitoring standard stations.
[0080] In addition, the logic instructions in the above-mentioned memory 330 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0081] According to an embodiment disclosed in the present invention, the present invention also includes a computer-readable storage medium, which includes a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication interface 320, or installed from the memory 330, or installed from the ROM. When the computer program is executed, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.
[0082] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware portion can be implemented using dedicated logic; the software portion can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art will appreciate that the above-mentioned devices and methods can be implemented using computer-executable instructions and / or contained in processor control code, for example, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. The devices and modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, or can be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0083] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with this technical field within the technical scope disclosed by the present invention and within the spirit and principles of the present invention should be covered by the scope of protection of the present invention.
Claims
1. A high value identification method for atmospheric monitoring standard stations based on random forest is characterized by: The specific steps include: Obtain concentration data based on the original data of multiple pollutants at standard stations; The concentration data is pre-processed and then input into an adaptive partitioning model based on spatial constraints to generate high-value identification data; Establishing a threshold-based recognition model based on a random forest model using the high-value recognition data; The identification model is used to identify high values of pollutant concentrations at standard stations obtained in real time.
2. The method for identifying high values of atmospheric monitoring standard stations based on random forest according to claim 1 is characterized in that: The step of obtaining concentration data based on the original data of multiple pollutants at the standard station includes obtaining the concentration data of multiple pollutants at different standard stations in real time through the data interface corresponding to the standard station.
3. The method for identifying high values of atmospheric monitoring standard stations based on random forest according to claim 2 is characterized in that: The steps of pre-processing the concentration data and inputting the pre-processed concentration data into a spatially constrained adaptive partitioning model to generate high-value identification data include: After splitting the concentration data, cleaning missing values and standardizing them, concentration data values corresponding to different pollutant types based on time series identification are generated; The concentration data values are used to generate high-value identification data for different site areas based on an adaptive partitioning model with spatial constraints; Wherein, the adaptive partition model is: Among them, (x i ,y i )、(x j ,y j ) are the latitude and longitude positions of sites i and j, where i is the latitude and longitude of the regional center site, C i 、C j is the average concentration of sites i and j, C' i , C' j is the normalized concentration data value, where μ c is the mean concentration, σ c is the standard deviation, α is the mixed distance weight, when d ij Those below the city-wide threshold are considered to be in the same area.
4. The method for identifying high values of atmospheric monitoring standard stations based on random forest according to claim 3 is characterized in that: The steps of establishing a threshold-based recognition model based on a random forest model using the high-value recognition data include: A feature matrix is created based on the high-value identification data, using different pollutant indices as features, and divided into training and test sets. Training and testing are performed based on the random forest algorithm. The cross-validation model used in the test is: Among them, L(θ) is the cross-validation accuracy, is the accuracy evaluation indicator, is the true label of the k-fold validation set, To predict the k-fold validation set under the parameter θ, we compare L(θ) for all θ and select the parameter with the highest L(θ) value. Dynamic thresholds corresponding to different pollutant types are set in the random forest model that has completed training and testing, and a recognition model is established based on the dynamic thresholds.
5. The method for identifying high values of atmospheric monitoring standard stations based on random forest according to claim 4 is characterized in that: The pollutants include PM 2.5 、PM 10 , NO2, SO2, CO and O3.
6. The method for identifying high values of atmospheric monitoring standard stations based on random forest according to claim 5, characterized in that: The steps of setting dynamic thresholds corresponding to different pollutant types and establishing an identification model based on the dynamic thresholds include: If the pollutant is NO2, SO2, CO, or O3, a first identification model is established. When the measured value of the pollutant is greater than 1.7 times the predicted value, the identification result is set to a high value. If the pollutant is PM 2.5 、PM 10 , then a second recognition model is established, and the second recognition model outputs the recognition result based on the following formula: Among them, C obs Indicates the measured concentration value.
7. An identification system for the high-value identification method of atmospheric monitoring standard stations based on random forest according to any one of claims 1 to 6, characterized in that: include: An acquisition module is used to obtain concentration data based on the original data of multiple pollutants at the standard station; A preprocessing module, configured to preprocess the concentration data and then input the preprocessed data into a spatial constraint-based adaptive partitioning model to generate high-value identification data; A recognition model generation module, configured to establish a threshold-based recognition model based on a random forest model using the high-value recognition data; The processing module is used to identify high values of the pollutant concentration of the standard station obtained in real time through the identification model.
8. An electronic device, characterized in that: include: one or more processors; And a memory associated with the one or more processors, the memory being used to store program instructions, which, when read and executed by the one or more processors, execute the steps of the high-value identification method for atmospheric monitoring standard stations based on random forests as described in any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the high value identification method of the atmospheric monitoring standard station based on random forest are implemented as described in any one of claims 1 to 6.