A Method for Filtering Air Quality Multi-Model Forecast Results Based on Sliding Optimal Matching

CN122332974APending Publication Date: 2026-07-03CHINA NAT ENVIRONMENTAL MONITORING CENT
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA NAT ENVIRONMENTAL MONITORING CENT
Filing Date
2026-03-27
Publication Date
2026-07-03

Smart Images

  • Figure CN122332974A_ABST
    Figure CN122332974A_ABST
Patent Text Reader

Abstract

This invention provides a method for filtering air quality forecast results based on sliding optimal matching, relating to the fields of environmental air quality forecasting and data science technology. The method includes: acquiring multiple predicted air quality index data obtained from various air quality prediction models; acquiring multiple measured air quality index data at multiple times within a past first preset time period; setting a time window; filtering target air quality prediction models from among the multiple air quality prediction models; and obtaining air quality index forecast data. According to this invention, predicted air quality index data and measured air quality index data can be acquired, and error statistics can be performed based on the time window to filter target air quality prediction models. The sliding optimal matching algorithm enables real-time evaluation and automatic filtering of model performance, overcoming the technical bottlenecks of traditional manual operation and static integration, improving the real-time performance and efficiency of air quality prediction model filtering, and providing high-precision and high-efficiency forecast support.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of ambient air quality forecasting and data science technology, and in particular to a method for screening air quality multi-model forecast results based on sliding optimal matching. Background Technology

[0002] In recent years, multi-model integrated air quality forecasting technology has become a major means to improve forecast accuracy and precision. Numerical models mainly include CAMx, CMAQ, NAQPMS, and WRF-Chem, while statistical models mainly include XGBoost, LightGBM, DNN, and LSTM. When forecasting air quality, methods such as multi-model ensembles, statistical corrections, machine learning fusion, and manual correction are typically used to optimize forecast results. In these technologies, forecasters need to evaluate the forecast performance of multiple models and manually select the optimal air quality prediction model based on air quality forecast performance evaluation methods. The forecast results of this optimal model are then used as the reference model forecast results, and further manual corrections are made by integrating current pollution source emissions, real-time air quality monitoring results, and meteorological and climate forecast reference information. However, the above methods rely to some extent on manual correction, and ambient air quality prediction systems typically employ multiple air quality prediction models. Each model's forecasting performance (accuracy) varies under different conditions. For example, different air quality forecasting models exhibit spatiotemporal heterogeneity, meaning their forecasting performance differs across seasons and primary pollutants. The forecasting capabilities for pollutants such as PM2.5 and O3 vary significantly with season, region, and pollution stage. Therefore, there is a possibility of inaccurate air quality prediction model selection and inaccurate air quality predictions.

[0003] The information disclosed in the background section of this application is intended only to enhance the understanding of the general background of this application and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention

[0004] This invention provides a method for screening air quality multi-model forecast results based on sliding optimal matching, which can solve the technical problems of inaccurate air quality prediction model screening and inaccurate air quality prediction in related technologies.

[0005] According to a first aspect of the present invention, a method for filtering air quality multi-model forecast results based on sliding optimal matching is provided, comprising: Acquire multiple predicted air quality index data obtained from various air quality prediction models at multiple moments within a first preset time period in the past; Acquire multiple measured air quality data at various times within a first preset time period in the past; Set a time window, wherein the duration of the time window is a second preset duration, and the second preset duration is less than the first preset duration; During the process of the start time of the time window changing within the first preset time period, the target air quality prediction model is selected from multiple air quality prediction models based on the predicted air quality index data and the measured air quality index data within the time window. Based on the target air quality prediction model, air quality index data for multiple moments within a future third preset time period are predicted to obtain air quality index forecast data.

[0006] According to the present invention, the end time of the first preset time period is the current time; As the start time of the time window changes within the first preset time period, a target air quality prediction model is selected from multiple air quality prediction models based on the predicted air quality index data and the measured air quality index data within the time window. This includes: When the start time of the time window is aligned with each time point of the first preset time period, the predicted air quality index data of each air quality prediction model at multiple times within the time window, as well as the measured air quality index data, are obtained. Obtain the error index between the predicted air quality index data and the measured air quality index data at multiple times within the time window output by each air quality prediction model; Based on the error index corresponding to each sliding within the first preset time period of the time window, the forecast stability score of each air quality prediction model is determined. Using a trained reference selection model, a reference time period corresponding to the first preset time period is selected from the historical time periods; Based on the measured air quality index data within the reference time period, the predicted air quality index data of each air quality prediction model, and the forecast stability score, the target air quality prediction model is selected.

[0007] According to the present invention, the forecast stability score of each air quality prediction model is determined based on the error index corresponding to each sliding of the time window within a first preset time period, including: Based on the error index of the i-th air quality prediction model for the j-th air quality index after each sliding of the time window, determine the error variance of the i-th air quality prediction model for the j-th air quality index. Based on the error variance, determine the stability index of the i-th air quality prediction model for predicting the j-th air quality index; After each slide of the time window, obtain the ranking information of the error index of each air quality prediction model for the j-th air quality index. Based on the ranking information corresponding to multiple sliding of the time window, determine the stable and preferred index for the i-th air quality prediction model to predict the j-th air quality index. Based on the stability index and the stability optimization index, the forecast stability score of the i-th air quality prediction model for the j-th air quality index is determined.

[0008] According to the present invention, based on the error variance, a stability index is determined for the i-th air quality prediction model to predict the j-th air quality index, including: According to the formula Determine the stability index of the i-th air quality prediction model for predicting the j-th air quality index. ,in, Let be the error variance of the i-th air quality prediction model for the j-th air quality index, max be the maximum value function, min be the minimum value function, n be the number of air quality prediction models, i ≤ n, and i and n are both positive integers.

[0009] According to the present invention, based on the ranking information corresponding to multiple sliding of the time window, a stable and preferred index for predicting the j-th air quality index by the i-th air quality prediction model is determined, including: Based on the ranking information corresponding to multiple sliding of the time window, among multiple air quality prediction models, the optimal model in the interval with the smallest error index for predicting the j-th air quality index after each sliding of the time window is determined. Based on the interval optimal model corresponding to each slide, determine the maximum number of consecutive times that the i-th air quality prediction model is continuously used as the interval optimal model. According to the formula Determine the stable and optimal index for the i-th air quality prediction model to predict the j-th air quality index. ,in, To predict the j-th air quality index, the maximum number of consecutive times the i-th air quality prediction model is continuously used as the optimal model in the interval, where max is the maximum value function, min is the minimum value function, n is the number of air quality prediction models, i≤n, and i and n are both positive integers.

[0010] According to the present invention, selecting a reference time period corresponding to the first preset time period from a historical time period using a trained reference selection model includes: Divide the historical time period into multiple test time periods of the same length as the first preset time period; Based on the measured air quality data, weather data and seasonal data at each moment within the test time period, determine the input vector to be measured at each moment within the test time period, and obtain the reference input vector at each moment within the first preset time period based on the measured air quality data, weather data and seasonal data at each moment within the first preset time period. The first mapping network layer of the selected model is used to map the input vector to be tested to obtain the first input feature information to be tested, and the first reference input vector is also used to map the reference input vector to obtain the first reference input feature information. Multiple first test input feature information are combined into a first test input matrix, and multiple first control input feature information are combined into a first control input matrix; By referencing the 2D convolutional network layers of the selected model, the first test input matrix and the first control input matrix are processed respectively to obtain the test pollution feature vector and the control pollution feature vector respectively. Obtain the first similarity between the feature vector of the pollution to be tested and the feature vector of the control pollution, and determine the test time period with the highest first similarity as the reference time period for the first pollution status; The second mapping network layer of the reference selected model is used to map the input vector to be tested to obtain the second input feature information to be tested, and the control input vector is also mapped to obtain the second control input feature information. By referencing the self-attention mechanism of the selected model, the second test input feature information is processed to obtain the second test fusion feature information, and the test fusion feature information is determined based on the second test fusion feature information. By referencing the self-attention mechanism of the selected model, the second control input feature information is processed to obtain the second control fusion feature information, and the control fusion feature information is determined based on the second control fusion feature information. Obtain the second similarity between the fusion feature information to be tested and the fusion feature information of the control, and determine the test time period with the highest second similarity as the first pollution trend reference time period; Based on the weather data and seasonal data at each moment within the test period, a clean control input vector corresponding to the test period is constructed, and a clean feature vector corresponding to the clean control input vector is obtained. The third similarity between the pollution feature vector to be measured and the clean feature vector is determined, and the time period with the highest third similarity is determined as the second pollution status reference time period. Based on the weather data and seasonal data at each moment within the test period, a pollution control input vector corresponding to the test period is constructed, and the pollution feature vector corresponding to the pollution control input vector is obtained. Determine the fourth similarity between the pollution feature vector to be measured and the pollution feature vector, and determine the time period with the highest fourth similarity as the reference time period for the third pollution status; Based on the weather data and seasonal data at each moment within the test period, a pollution accumulation control input vector corresponding to the test period is constructed, and a pollution accumulation fusion vector corresponding to the pollution accumulation control input vector is obtained. The fifth similarity between the fusion feature information to be tested and the pollution accumulation fusion vector is obtained, and the time period with the highest fifth similarity is determined as the second pollution trend reference time period. Based on the weather data and seasonal data at each moment within the test period, a pollution reduction control input vector corresponding to the test period is constructed, and a pollution reduction fusion vector corresponding to the pollution reduction control input vector is obtained. Obtain the sixth similarity between the fusion feature information to be tested and the pollution reduction fusion vector, and determine the test time period with the highest sixth similarity as the third pollution trend reference time period; The first pollution status reference time period, the first pollution trend reference time period, the second pollution status reference time period, the third pollution status reference time period, the second pollution trend reference time period, and the third pollution trend reference time period are determined as reference time periods.

[0011] According to the present invention, a target air quality prediction model is selected based on measured air quality index data within a reference time period, predicted air quality index data of each air quality prediction model, and the forecast stability score, including: The air quality prediction model corresponding to the maximum forecast stability score is determined as the target stable model; Based on the measured air quality index data within the reference time period and the predicted air quality index data of each air quality prediction model, the following accurate indicators are obtained: the first accurate indicator corresponding to each sliding of the time window within the first pollution status reference time period, the second accurate indicator corresponding to each sliding of the time window within the second pollution status reference time period, the third accurate indicator corresponding to each sliding of the time window within the third pollution status reference time period, the fourth accurate indicator corresponding to each sliding of the time window within the first pollution trend reference time period, the fifth accurate indicator corresponding to each sliding of the time window within the second pollution trend reference time period, and the sixth accurate indicator corresponding to each sliding of the time window within the third pollution trend reference time period. Based on the first accuracy index, second accuracy index, third accuracy index, fourth accuracy index, fifth accuracy index, sixth accuracy index, and target stable model, a target air quality prediction model is selected.

[0012] According to the present invention, a target air quality prediction model is selected based on a first accuracy index, a second accuracy index, a third accuracy index, a fourth accuracy index, a fifth accuracy index, a sixth accuracy index, and a target stable model, including: After each slide of the time window within the first pollution status reference period, it is determined whether the difference between the maximum and second maximum values ​​of the first accurate indicator is greater than the average difference between the first accurate indicators of each air quality prediction model. If so, the air quality prediction model corresponding to the maximum value of the first accurate indicator is determined as the optimal model of the first interval corresponding to this slide of the time window; otherwise, the target stable model is taken as the optimal model of the first interval. The number of times each air quality prediction model is selected as the optimal model for the first interval is statistically analyzed. Based on the second accuracy index and the target stability model, each air quality prediction model is determined as the second degree of the optimal model in the second interval; Based on the third accuracy index and the target stability model, each air quality prediction model is determined as the third number of the optimal model in the third interval; Based on the fourth accuracy index and the target stability model, each air quality prediction model is determined as the fourth degree of the optimal model in the fourth interval; Based on the fifth accuracy index and the target stability model, the fifth degree of each air quality prediction model is determined as the optimal model for the fifth interval; Based on the sixth accuracy index and the target stability model, the sixth degree of each air quality prediction model is determined as the optimal model for the sixth interval. Based on the first, second, and third counts, determine the optimal pollution score for each air quality prediction model to predict the j-th air quality index. Based on the fourth, fifth, and sixth times, determine the optimal pollution trend score for each air quality prediction model to predict the j-th air quality index. Based on the optimal pollution status score and the optimal pollution trend score, a target air quality prediction model for predicting the j-th air quality index is selected.

[0013] According to the present invention, based on the first, second, and third data points, the optimal pollution score for each air quality prediction model to predict the j-th air quality index is determined, including: According to the formula Determine the optimal pollution score for the i-th air quality prediction model when predicting the j-th air quality index. ,in, The first similarity score corresponds to the reference time period of the first pollution condition. The third similarity is the reference time period for the second pollution condition. The fourth similarity is the reference time period for the third pollution status. For the first count, For the second number, This is the third time, and N is the total number of times the time window slides.

[0014] According to the present invention, based on the fourth, fifth, and sixth iterations, the optimal pollution trend score for each air quality prediction model to predict the j-th air quality index is determined, including: According to the formula Determine the optimal pollution score for the i-th air quality prediction model when predicting the j-th air quality index. ,in, This is the fourth time. This is the fifth time. This is the sixth time. The second similarity is the reference time period for the first pollution trend. The fifth similarity score corresponds to the second pollution trend reference time period. The sixth similarity is the reference time period for the third pollution trend, and N is the total number of sliding times of the time window.

[0015] By adopting the above technical solution, the present invention can achieve the following technical effects: According to this invention, a computer can automatically calculate the error index between the predicted air quality index data of an air quality prediction model and the actual measured air quality index data within a first preset time period. By statistically analyzing the prediction performance of each model at different positions within a sliding window, a target air quality prediction model suitable for predicting specific types of air quality indicators within a third preset time period can be automatically and dynamically selected. Prediction is then performed based on this model to obtain air quality index forecast data. This improves the real-time performance and efficiency of air quality prediction model selection. From error index calculation and target air quality prediction model selection to air quality index forecast data provision, the entire process requires no manual intervention, reducing reliance on human intervention. It overcomes the technical bottlenecks of traditional manual operation and static integration, adapting to the rapid evolution of air pollution and providing high-precision, high-efficiency forecast support for operational departments. Furthermore, the variance of the error index can be used to determine a stability index, and the ranking of the error index can determine a stable optimization index. Therefore, when calculating the prediction stability score, the accuracy and stability of the air quality prediction model in predicting air quality indicators are fully considered, providing a data foundation for selecting models with accurate and stable prediction performance that can adapt to complex and rapidly changing conditions. Furthermore, under constraints of weather conditions and seasonal data, the overall pollution status within the test time period can be obtained through 2D convolutional network layers, and the pollution change trend within the test time period can be obtained through a self-attention mechanism. This allows for the comprehensive selection of reference time periods based on both pollution status and pollution change trends, providing a richer and more comprehensive reference for selecting suitable air quality prediction models, thus enhancing the reference value and accuracy of air quality prediction model selection. Moreover, during the training phase of the reference selection model, auxiliary training layers can be used to improve the accuracy of the model's feature representation of pollution status and pollution trends, and to enhance its ability to represent weather and seasonal features. This helps to provide more accurate pollution status and pollution change trend features under weather and seasonal constraints, providing an accurate computational foundation for selecting reference time periods. Furthermore, the predictive performance of each air quality prediction model under different pollution conditions can be described by the first, second, and third times, and the predictive performance of each air quality prediction model under different pollution trends can be described by the fourth, fifth, and sixth times. In this way, the predictive performance of the models under multiple conditions can be comprehensively considered, and the air quality prediction models can be fully evaluated. This allows for the selection of the air quality prediction model with the strongest comprehensive performance under multiple conditions, thereby improving the accuracy and stability of subsequent predictions.

[0016] It should be understood that the foregoing general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Other features and aspects of the invention will become clearer from the following detailed description of exemplary embodiments with reference to the accompanying drawings. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other embodiments can be obtained based on these drawings without creative effort. Figure 1 An exemplary flowchart of a method for screening air quality multi-model forecast results based on sliding optimal matching according to an embodiment of the present invention is shown. Figure 2 A schematic diagram of a reference selection model according to an embodiment of the present invention is shown as an example. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] The technical solution of the present invention will be described in detail below with reference to specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0020] Figure 1 An exemplary flowchart illustrates a method for filtering air quality multi-model forecast results based on sliding optimal matching according to an embodiment of the present invention. The method includes: Step S1: Obtain multiple predicted air quality index data from multiple times within the past first preset time period, which are obtained from multiple air quality prediction models. Step S2: Obtain multiple measured air quality data at various times within the past first preset time period; Step S3: Set a time window, wherein the duration of the time window is a second preset duration, and the second preset duration is less than the first preset duration; Step S4: During the process of the start time of the time window changing within the first preset time period, the target air quality prediction model is selected from multiple air quality prediction models based on the predicted air quality index data and the measured air quality index data within the time window. Step S5: Based on the target air quality prediction model, predict the air quality index data for multiple moments within a future third preset time period to obtain air quality index forecast data.

[0021] According to an embodiment of the present invention, the air quality multi-model forecast result screening method based on sliding optimal matching can automatically calculate the error index between the predicted air quality index data of the air quality prediction model and the actual measured air quality index data during a first preset time period using a computer. By statistically analyzing the prediction performance of each model at various positions within a sliding window, a target air quality prediction model suitable for predicting specific types of air quality indicators during a third preset time period is automatically and dynamically selected. Prediction is then performed based on this model to obtain air quality index forecast data. This method improves the real-time performance and efficiency of air quality prediction model screening. From error index calculation and target air quality prediction model selection to the provision of air quality index forecast data, no manual intervention is required throughout the process, reducing reliance on human intervention. It overcomes the technical bottlenecks of traditional manual operation and static integration, adapts to the rapid evolution of air pollution, and provides high-precision and high-efficiency forecast support for operational departments.

[0022] Example 1: According to an embodiment of the present invention, in step S1, multiple predicted air quality index data obtained by multiple air quality prediction models at multiple moments within a first preset time period are acquired, wherein the end time of the first preset time period is the current time. Since air quality data changes rapidly over time, data from the most recent three weeks can represent the recent air quality changes, and the amount of data from three weeks is suitable for quickly obtaining calculation results; therefore, the first preset time period can be three weeks. At multiple moments within the first preset time period, multiple air quality prediction models (e.g., numerical models mainly include CAMx, CMAQ, NAQPMS, WRF-Chem, etc., and statistical models mainly include XGBoost, LightGBM, DNN, LSTM, etc.) are used to predict multiple air quality indicators (e.g., PM2.5). 2.5 PM 10 The invention predicts the concentrations of major air pollutants (such as ozone, sulfur dioxide, nitrogen dioxide, and carbon monoxide) to obtain various predicted air quality index data. The time interval between each moment within the first preset time period can be 1 hour, 3 hours, etc., and at each moment, multiple air quality prediction models can be used to predict various air quality indicators, resulting in a set of predicted air quality index data. For example, if the first preset time period is 3 weeks, and the time interval between each moment is 3 hours, one air quality prediction model can obtain 168 sets of predicted air quality index data. This invention does not limit information such as the time interval between adjacent moments, the length of the first preset time period, the type of model, or the type of air quality index data.

[0023] According to one embodiment of the present invention, in step S2, multiple measured air quality index data at multiple times within a past first preset time period are acquired. Within the past first preset time period, the content of multiple air quality indexes at multiple times is detected by multiple instruments to obtain the actual content of each air quality index, which is the multiple measured air quality indexes.

[0024] According to one embodiment of the present invention, in step S3, a time window is set, wherein the duration of the time window is a second preset duration, and the second preset duration is less than a first preset duration. Since ecological and environmental management mainly focuses on 72-hour forecast results, and these forecast results can effectively support air pollution control and heavy pollution emergency response, the duration of the time window (i.e., the second preset duration) can be 3 days (i.e., obtaining the error index between predicted air quality index data and measured air quality index data in a 3-day cycle). Since the time window needs to be slidable within the first preset duration (sliding by one moment each time) to statistically analyze the error index and select the optimal air quality prediction model after comparison, the second preset duration is less than the first preset duration.

[0025] Example 2: According to one embodiment of the present invention, in step S4, it can be determined whether various air quality prediction models are suitable for predicting air quality indicators in the future third preset time period based on the prediction performance of various air quality prediction models at multiple locations in the time window, thereby selecting air quality prediction models with good prediction performance to predict specific air quality indicators in the third time period.

[0026] According to one embodiment of the present invention, during the process of the start time of the time window changing within a first preset time period, a target air quality prediction model is selected from multiple air quality prediction models based on the predicted air quality index data and measured air quality index data within the time window. This includes: when the start time of the time window aligns with each moment of the first preset time period, acquiring the predicted air quality index data and measured air quality index data of each air quality prediction model at multiple moments within the time window; acquiring the error index between the predicted air quality index data and measured air quality index data output by each air quality prediction model at multiple moments within the time window; determining the forecast stability score of each air quality prediction model based on the error index corresponding to each sliding of the time window within the first preset time period; selecting a reference time period corresponding to the first preset time period from historical time periods using a trained reference selection model; and selecting a target air quality prediction model based on the measured air quality index data within the reference time period, the predicted air quality index data of each air quality prediction model, and the forecast stability score.

[0027] According to one embodiment of the present invention, when the start time of the time window aligns with the t-th time of the first preset time period, the predicted air quality index data and the measured air quality index data of each air quality prediction model at multiple times within the time window are acquired. The time window slides in one-time increments each time, and stops sliding when the end time of the time window aligns with the end time of the first preset time period. After each slide of the time window (including when the start time of the time window aligns with the 1st time within the first preset time period, which can be considered the 0th slide), the error index between the predicted air quality index data and the measured air quality index data at multiple times within the time window can be calculated. For example, the mean absolute error index (MAE). For instance, for the j-th air quality index, the absolute value of the difference between the predicted air quality index data and the measured air quality index data at each time within the time window output by the i-th air quality prediction model can be calculated, and then averaged to obtain the error index of the i-th air quality prediction model for the j-th air quality index after that slide. Similarly, the error index of various air quality prediction models for various air quality indices can be calculated after each slide of the time window.

[0028] Example 3: According to one embodiment of the present invention, after obtaining the above-mentioned error index, the performance of each air quality prediction model for predicting various air indicators can be determined based on the error index.

[0029] According to one embodiment of the present invention, the forecast stability score of each air quality prediction model is determined based on the error index corresponding to each sliding of the time window within a first preset time period. This includes: determining the error variance of the i-th air quality prediction model for the j-th air quality index based on the error index of the i-th air quality prediction model for each j-th air quality index after each sliding of the time window; determining the stability index of the i-th air quality prediction model for the j-th air quality index based on the error variance; obtaining ranking information of the error indices of each air quality prediction model for the j-th air quality index after each sliding of the time window; determining the optimal stability index of the i-th air quality prediction model for the j-th air quality index based on the ranking information corresponding to multiple slidings of the time window; and determining the forecast stability score of the i-th air quality prediction model for the j-th air quality index based on the stability index and the optimal stability index.

[0030] According to an embodiment of the present invention, based on the above calculation method, after each sliding of the time window, the error index of the i-th air quality prediction model for the j-th air quality index can be calculated. Therefore, after the time window slides multiple times within the first preset time period, multiple error indices of the i-th air quality prediction model for the j-th air quality index can be obtained. The variance of the multiple error indices, i.e., the error variance, can be solved to determine whether the ability of the i-th air quality prediction model to predict the j-th air quality index is stable, i.e., whether the error index fluctuates significantly.

[0031] According to an embodiment of the present invention, determining the stability index of the i-th air quality prediction model for predicting the j-th air quality index based on the error variance includes: determining the stability index of the i-th air quality prediction model for predicting the j-th air quality index according to formula (1). , (1) in, Let be the error variance of the i-th air quality prediction model for the j-th air quality index, max be the maximum value function, min be the minimum value function, n be the number of air quality prediction models, i ≤ n, and i and n are both positive integers.

[0032] According to one embodiment of the present invention, This represents the maximum variance of the prediction error for each air quality prediction model for the j-th air quality index. The minimum variance of the error in each air quality prediction model for the j-th air quality index is therefore... Let be the normalized error variance of the i-th air quality prediction model for the j-th air quality index. The error variance can be mapped to a value between 0 and 1, which facilitates the comparison and evaluation of the error variance of each model. After normalization, this term can represent the volatility of the error of the i-th air quality prediction model. By subtracting this term from 1, the stability index can be obtained. The higher the stability index, the more stable the performance of the i-th air quality prediction model, and the less the prediction error will fluctuate at multiple times.

[0033] According to an embodiment of the present invention, based on the ranking information corresponding to multiple sliding of the time window, a stable preferred index for the i-th air quality prediction model to predict the j-th air quality index is determined, including: based on the ranking information corresponding to multiple sliding of the time window, determining the interval optimal model with the smallest error index for predicting the j-th air quality index after each sliding of the time window among multiple air quality prediction models; determining the maximum number of consecutive times the i-th air quality prediction model is continuously used as the interval optimal model based on the interval optimal model corresponding to each sliding; and determining the stable preferred index for the i-th air quality prediction model to predict the j-th air quality index according to formula (2). , (2) in, To predict the j-th air quality index, the maximum number of consecutive times the i-th air quality prediction model is continuously used as the optimal model in the interval, where max is the maximum value function, min is the minimum value function, n is the number of air quality prediction models, i≤n, and i and n are both positive integers.

[0034] According to one embodiment of the present invention, the interval optimal model is the air quality prediction model with the best prediction performance and the smallest error index within the time interval of the time window. During the sliding of the time window, the interval optimal model may change. For example, in some time intervals, model A has the smallest error index, while in other time intervals, model B has the smallest error index. The i-th air quality prediction model continuously serving as the interval optimal model indicates that the i-th air quality prediction model consistently has the smallest error index and the best prediction performance during multiple consecutive sliding of the time window. The larger the maximum number of consecutive occurrences, the better the model's prediction performance and stability. In other words, the larger the maximum number of consecutive occurrences, the more stably and accurately the i-th air quality prediction model can predict, and the more stable and efficient it is in rapidly changing environments. It is suitable for predicting the j-th air quality index under various environments and can serve as an objective evaluation index for selecting stable, robust, adaptable, and accurate models. To normalize the maximum consecutive count of the i-th air quality prediction model, the maximum consecutive count of each air quality prediction model can be mapped to the interval between 0 and 1, which facilitates comparison and evaluation of the models. This term can be used as a stable and preferred index for the i-th air quality prediction model to predict the j-th air quality index, and is used to represent the accuracy and stability of the i-th air quality prediction model in predicting the j-th air quality index.

[0035] According to one embodiment of the present invention, the stability index and the stability optimization index can be weighted and summed to obtain the forecast stability score of the i-th air quality prediction model for the j-th air quality index. For example, the weight of the stability index is 0.4 and the weight of the stability optimization index is 0.6. After weighted summation, the forecast stability score can be obtained.

[0036] In this way, stability indicators can be determined by the variance of error indicators, and optimal stability indicators can be determined by the ranking of error indicators. Thus, when calculating the prediction stability score, the accuracy and stability of the air quality prediction model in predicting air indicators can be fully combined, providing a data foundation for screening models with accurate and stable prediction performance that can adapt to complex and rapid changes.

[0037] Example 4: According to one embodiment of the present invention, in addition to using the above-mentioned forecast stability score to evaluate the prediction performance of each air quality prediction model in the first preset time period, a reference time period with other data characteristics can also be selected in the historical time period to evaluate the prediction performance of each air quality prediction model in the reference time period, so as to select the model with the strongest comprehensive performance and most suitable for prediction in the future third prediction time period.

[0038] According to one embodiment of the present invention, a reference time period can be selected. In the example, a reference selection model can be used to automatically filter reference time periods with specific data characteristics, which can then be used to evaluate the predictive performance of each model under specific circumstances.

[0039] Figure 2 A schematic diagram of a reference selection model according to an embodiment of the present invention is shown as an example.

[0040] According to an embodiment of the present invention, a reference time period corresponding to the first preset time period is selected from a historical time period using a trained reference selection model. This includes: dividing the historical time period into multiple test time periods of equal length to the first preset time period; determining the test input vector for each moment within the test time period based on various measured air quality index data, weather data, and seasonal data at each moment; and obtaining the reference input vector for each moment within the first preset time period based on the various measured air quality index data, weather data, and seasonal data at each moment; and mapping the test input vector through the first mapping network layer of the reference selection model to obtain the first test time period. Input feature information and map the control input vector to obtain first control input feature information; combine multiple first test input feature information into a first test input matrix, and combine multiple first control input feature information into a first control input matrix; process the first test input matrix and the first control input matrix respectively through the 2D convolutional network layers of the reference selection model to obtain test pollution feature vector and control pollution feature vector respectively; obtain the first similarity between the test pollution feature vector and the control pollution feature vector, and determine the test time period with the highest first similarity as the first pollution status reference time period; map the test input vector through the second mapping network layers of the reference selection model. The process involves obtaining second test input feature information and mapping the control input vector to obtain second control input feature information; processing the second test input feature information using the self-attention mechanism of the reference selection model to obtain second test fusion feature information, and determining the test fusion feature information based on the second test fusion feature information; processing the second control input feature information using the self-attention mechanism of the reference selection model to obtain second control fusion feature information, and determining the control fusion feature information based on the second control fusion feature information; obtaining the second similarity between the test fusion feature information and the control fusion feature information, and determining the test time period with the highest second similarity as the first pollution trend reference time period; and so on. Based on the weather and seasonal data at each moment within the test period, a clean control input vector corresponding to the test period is constructed, and a clean feature vector corresponding to the clean control input vector is obtained; the third similarity between the pollution feature vector to be tested and the clean feature vector is determined, and the test period with the highest third similarity is determined as the second pollution status reference period; based on the weather and seasonal data at each moment within the test period, a pollution control input vector corresponding to the test period is constructed, and a pollution feature vector corresponding to the pollution control input vector is obtained; the fourth similarity between the pollution feature vector to be tested and the pollution feature vector is determined, and the test period with the highest fourth similarity is determined as the third pollution status reference period;Based on the weather and seasonal data at each moment within the test period, a pollution accumulation control input vector corresponding to the test period is constructed, and a pollution accumulation fusion vector corresponding to the pollution accumulation control input vector is obtained. The fifth similarity between the test fusion feature information and the pollution accumulation fusion vector is obtained, and the test period with the highest fifth similarity is determined as the second pollution trend reference period. Based on the weather and seasonal data at each moment within the test period, a pollution decrease control input vector corresponding to the test period is constructed, and a pollution decrease fusion vector corresponding to the pollution decrease control input vector is obtained. The sixth similarity between the test fusion feature information and the pollution decrease fusion vector is obtained, and the test period with the highest sixth similarity is determined as the third pollution trend reference period. The first pollution status reference period, the first pollution trend reference period, the second pollution status reference period, the third pollution status reference period, the second pollution trend reference period, and the third pollution trend reference period are determined as reference periods.

[0041] According to one embodiment of the present invention, the historical time period can be the past three years. The duration of the time period to be measured is equal to the duration of the first preset time period, for example, three weeks. In the example, the historical time period of the past three years can be divided into multiple time periods with a duration of three weeks, and each time period is the time period to be measured.

[0042] According to one embodiment of the present invention, multiple measured air quality index data, weather data, and seasonal data at each moment within the test time period can be combined. For example, weather data can be weather type data represented by numbers, such as 1 for sunny, 2 for rainy, 3 for snowy, and 4 for cloudy. Seasonal data can be seasonal type data represented by numbers, such as 1 for spring, 2 for summer, 3 for autumn, and 4 for winter. Multiple measured air quality index data, weather data, and seasonal data can be combined into a vector to obtain the test input vector for that moment. Similarly, the test input vector for each moment within the test time period can be obtained. On the other hand, a reference input vector for each moment within a first preset time period can be obtained in a similar manner. The present invention does not limit the specific representation of weather data and seasonal data.

[0043] According to one embodiment of the present invention, the first mapping network layer includes multiple fully connected layers and ReLU activation layers, which can perform mapping processing on the input vector to be tested to obtain first input feature information to be tested, thereby fusing measured air quality data, weather data, and seasonal data, and expressing the features of these data through high-dimensional vectors. Similarly, first reference input feature information can be obtained from the reference input vector.

[0044] According to one embodiment of the present invention, multiple first input feature information to be tested can be combined into a first input matrix to be tested. This first input matrix can then be processed through a 2D convolutional network layer. The 2D convolutional network layer can include multiple convolutional layers, pooling layers, and ReLU activation layers. Using the 2D convolutional network layer allows processing not only of data in the first input feature information at a single time step, but also, based on the characteristics of the 2D convolutional kernel, processing of data at adjacent time steps. This allows for the acquisition of overall data features across multiple time steps; in other words, it allows for the fusion of data features from multiple time steps, enabling a more accurate representation of the overall contamination status at multiple time steps. After processing by the 2D convolutional network layer, the contamination feature vector to be tested can be directly obtained. Alternatively, after processing by the 2D convolutional network layer, an intermediate result can be obtained, which is then passed through a multilayer perceptron network layer (including multiple fully connected layers and ReLU activation layers) to obtain the contamination feature vector to be tested. A control contamination feature vector can be obtained in a similar manner, and then the first similarity between the contamination feature vector to be tested and the control contamination feature vector can be calculated. The first similarity between the pollution feature vector to be tested and the control pollution feature vector for each test time period can be calculated in the same way, and the test time period with the highest first similarity is selected as the first pollution status reference time period. The first pollution status reference time period is the time period that is most similar to the overall pollution status or pollution level under certain weather and seasonal constraints. The performance of each model in this time period can be used as a reference under the condition of similar overall pollution level.

[0045] According to one embodiment of the present invention, the second mapping network layer may include multiple fully connected layers and ReLU activation layers, which can obtain second test input feature information of the test input vector through another mapping method, thereby expressing other aspects of data features, such as features for predicting the trend of data changes. Similarly, second reference input feature information of the reference input vector can be obtained.

[0046] According to one embodiment of the present invention, the self-attention mechanism enables the data features expressed by the second test input feature information at different times to be fused together, thereby allowing the second test fused feature information to carry data features from multiple times, thus expressing the changing trend of data features at multiple times. Further, the second test fused feature information corresponding to the last time step in the test time period can be directly determined as the test fused feature information. Alternatively, the second test fused feature information corresponding to the last time step can be input into a multilayer perceptron layer (including multiple fully connected layers and ReLU activation layers) to obtain the test fused feature information. That is, the second test fused feature information at the last time step is used to express the changing trend of data features throughout the entire test time period. In other words, the changing trend of data features in all previous times is observed from the perspective of the last time step to obtain the test fused feature information. Alternatively, the second test fused feature information corresponding to multiple times can be concatenated and input into a multilayer perceptron layer (including multiple fully connected layers and ReLU activation layers) to obtain the test fused feature information. That is, the global second test fused feature information is used to more comprehensively express the changing trend. Similarly, contrast fused feature information can be obtained. Furthermore, the second similarity between the fusion feature information to be tested and the control fusion feature information can be determined, and based on the same method, the second similarity between the fusion feature information to be tested and the control fusion feature information for each time period to be tested can be determined. Then, the time period with the highest second similarity is determined as the first pollution trend reference time period, that is, the time period with the pollution change trend most similar to the first preset time period under certain weather and seasonal constraints.

[0047] According to one embodiment of the present invention, in the clean control input vector, the weather data and seasonal data at each moment are consistent with the weather data and seasonal data at the corresponding moment within a first preset time period. Each air quality indicator is randomly generated within the interval corresponding to the lowest pollution level, thus obtaining the clean control input vector. A clean feature vector is then obtained in the same manner as the pollution feature vector to be measured. Further, a third similarity between the pollution feature vector to be measured and the clean feature vector can be determined. The time period with the highest third similarity is then determined as the second pollution status reference time period. This time period is the period when, under certain weather and seasonal constraints, the overall pollution situation is closest to the optimal air quality condition.

[0048] According to one embodiment of the present invention, in the pollution control input vector, the weather data and seasonal data at each moment are consistent with the weather data and seasonal data at the corresponding moment within a first preset time period. Each air quality indicator is randomly generated within an interval corresponding to at least one higher pollution level, thus obtaining the pollution control input vector. The pollution feature vector is then obtained in the same manner as the pollution feature vector to be tested. Further, a fourth similarity between the pollution feature vector to be tested and the clean feature vector can be determined. The time period with the highest fourth similarity is then determined as the third pollution status reference time period, which, under certain weather and seasonal constraints, represents the time period where the overall pollution situation is closest to a very poor air quality condition.

[0049] According to one embodiment of the present invention, in multiple pollution accumulation control input vectors, the weather data and seasonal data at each moment are consistent with the weather data and seasonal data at the corresponding moment within a first preset time period. Multiple air quality indicators gradually increase over time. For example, a value can be selected from the range of good air quality levels as the starting point, and a value can be selected from the range of worst air quality levels as the ending point. The air quality indicators at multiple moments between these two levels show an overall increasing trend. The pollution accumulation fusion vector corresponding to the pollution accumulation control input vector can be obtained in the same way as the fusion feature information to be measured. The fifth similarity between the fusion feature information to be measured and the pollution accumulation fusion vector is calculated, and the time period with the highest fifth similarity is determined as the second pollution trend reference time period, which, under certain weather and seasonal constraints, represents the time period where the overall pollution trend is closest to the trend of worsening air pollution.

[0050] According to one embodiment of the present invention, in multiple pollution reduction control input vectors, the weather data and seasonal data at each moment are consistent with the weather data and seasonal data at the corresponding moment within a first preset time period. Multiple air quality indicators gradually decrease over time. For example, a value can be selected from the range of the worst air quality level as the starting point, and a value can be selected from the range of good air quality levels as the ending point. The air quality indicators at multiple moments between these two levels show an overall trend. The pollution reduction fusion vector corresponding to the pollution reduction control input vector can be obtained in the same way as the fusion feature information to be measured. The sixth similarity between the fusion feature information to be measured and the pollution reduction fusion vector is calculated, and the time period with the highest sixth similarity is determined as the third pollution trend reference time period, which, under certain weather and seasonal constraints, represents the time period where the overall pollution trend is closest to the trend of air pollution reduction.

[0051] According to an embodiment of the present invention, the first pollution status reference time period, the first pollution trend reference time period, the second pollution status reference time period, the third pollution status reference time period, the second pollution trend reference time period, and the third pollution trend reference time period can all be used as reference time periods.

[0052] According to one embodiment of the present invention, the above-mentioned reference selection model can be trained before use, and auxiliary training layers can be added for auxiliary training. In the example, the measured air quality index data, weather data, and seasonal data at each moment in the training period can be used to form a training input vector. After the reference selection model is trained, a training pollution feature vector is output through a 2D convolutional network layer, and training fusion feature information is output through a self-attention mechanism. Subsequently, various auxiliary training layers can be added, for example, auxiliary training layers (e.g., multiple fully connected layers and softmax activation layers) can be added after the 2D convolutional neural network layer to obtain the probability of which pollution level each air quality index belongs to in the training period. The probability of the pollution level is then combined with the actual pollution level label (e.g., the actual pollution level is labeled as 1, and other pollution levels are labeled as 0) to form a cross-entropy loss function. For example, by adding auxiliary training layers (e.g., multiple fully connected layers and multiple sigmoid activation layers, each corresponding to a weather condition) after the 2D convolutional neural network layers, the probability of which weather conditions are included in the training time period can be obtained. A cross-entropy loss function can then be constructed using the actual weather condition labels (e.g., 1 for included weather conditions and 0 for excluded weather conditions) and this probability. Similarly, by adding auxiliary training layers (e.g., multiple fully connected layers and softmax activation layers) after the 2D convolutional neural network layers, the probability of which season the training time period belongs to can be obtained. A cross-entropy loss function can then be constructed using the actual season labels (e.g., 1 for the actual season and 0 for other seasons). These various loss functions not only improve the ability of the trained pollution feature vector to express pollution conditions but also enhance its ability to express various constraints such as season and weather conditions. This improves the overall feature representation ability of the trained pollution feature vector, providing an accurate data foundation for accurately representing the pollution feature vector to be tested and for comparing it with a control pollution feature vector to select the first pollution condition reference time period.

[0053] On the other hand, auxiliary training layers can be added after the self-attention mechanism. For example, auxiliary training layers (e.g., multiple fully connected layers and sigmoid activation layers) can be added after the self-attention mechanism to obtain the probability that the pollution trend of various air indicators is rising or falling during the training period. The cross-entropy loss function can be constructed by labeling the actual pollution trend (e.g., 1 for an actual pollution trend rising and 0 for an actual pollution trend falling) and the probability. The auxiliary training layers for seasonality and weather conditions are similar to those described above and will not be repeated here.

[0054] According to one embodiment of the present invention, a loss function for the overall model can be constructed based on the above-mentioned multiple cross-entropy loss functions. For example, the loss function of the overall model is obtained by weighted summation of multiple cross-entropy loss functions. During training, the values ​​of each cross-entropy loss function are reduced, and after multiple training iterations, training is completed to obtain the trained reference selection model. This improves the overall expressive power of the model's output information, enabling it to accurately express not only pollution status and trends but also weather and seasonal characteristics, thus providing accurate feature representations and a precise computational basis for selecting reference time periods.

[0055] In this approach, under the constraints of weather conditions and seasonal data, the overall pollution status within a given time period can be obtained through a 2D convolutional network hierarchy. Furthermore, a self-attention mechanism can be used to capture the pollution change trend within that time period. This allows for the comprehensive selection of reference time periods based on both pollution status and pollution change trends, providing a richer and more comprehensive reference for selecting suitable air quality prediction models, thus enhancing the reference value and accuracy of the selection process. Moreover, during the training phase of the reference selection model, auxiliary training layers can be used to improve the accuracy of the model's representation of pollution status and trends, and to enhance its ability to represent weather and seasonal features. This helps to provide more accurate pollution status and trend characteristics under weather and seasonal constraints, providing a precise computational foundation for selecting reference time periods.

[0056] Example 5: According to one embodiment of the present invention, after obtaining the above-mentioned multiple reference time periods, the performance of each air quality prediction model within the reference time periods can be evaluated, thereby comprehensively selecting the target air quality prediction model.

[0057] According to one embodiment of the present invention, selecting a target air quality prediction model based on measured air quality index data within a reference time period, predicted air quality index data of each air quality prediction model, and the forecast stability score includes: determining the air quality prediction model corresponding to the maximum forecast stability score as the target stable model; obtaining a first accurate index corresponding to each sliding of the time window within a first pollution state reference time period, a second accurate index corresponding to each sliding of the time window within a second pollution state reference time period, a third accurate index corresponding to each sliding of the time window within a third pollution state reference time period, a fourth accurate index corresponding to each sliding of the time window within a first pollution trend reference time period, a fifth accurate index corresponding to each sliding of the time window within a second pollution trend reference time period, and a sixth accurate index corresponding to each sliding of the time window within a third pollution trend reference time period; and selecting the target air quality prediction model based on the first accurate index, second accurate index, third accurate index, fourth accurate index, fifth accurate index, sixth accurate index, and the target stable model.

[0058] According to an embodiment of the present invention, in the previous steps, the forecast stability score of various air quality prediction models for the j-th air quality index was solved, and the maximum value of the forecast stability score can be used as the target stable model corresponding to the j-th air quality index.

[0059] According to one embodiment of the present invention, the prediction performance of various models (including the target stable model) within the aforementioned reference time period can be comprehensively analyzed to select a target air quality prediction model for predicting the j-th air quality index. In the example, for the j-th air quality index, the average absolute error of the measured air quality index data within the time window of the first pollution state reference time period and the predicted air quality index data of the i-th air quality prediction model can be calculated at multiple times. Based on the average absolute error of each air quality prediction model, the average absolute error of the i-th air quality prediction model is normalized to obtain an error score between 0 and 1. Then, by subtracting the error score from 1, the first accurate index of the i-th air quality prediction model is obtained. Similarly, the first accurate index of each air quality prediction model can be obtained. Further, the second, third, fourth, fifth, and sixth accurate indices of each air quality prediction model can be determined in a similar manner.

[0060] According to one embodiment of the present invention, the above-mentioned various accurate indicators can describe the performance of the air quality prediction model in time periods similar to the first preset time period, in time periods of severe pollution and mild pollution, and in time periods of gradual pollution accumulation and gradual pollution elimination. The target air quality prediction model can be selected by comprehensively referring to the above performance.

[0061] According to one embodiment of the present invention, a target air quality prediction model is selected based on a first accuracy index, a second accuracy index, a third accuracy index, a fourth accuracy index, a fifth accuracy index, a sixth accuracy index, and a target stable model. This includes: after each sliding of the time window within a first pollution state reference period, determining whether the difference between the maximum and second largest values ​​of the first accuracy index is greater than the average difference between the first accuracy indices of each air quality prediction model; if so, determining the air quality prediction model corresponding to the maximum value of the first accuracy index as the first interval optimal model corresponding to this sliding of the time window; otherwise, using the target stable model as the first interval optimal model; counting the first number of times each air quality prediction model is used as the first interval optimal model; determining the second number of times each air quality prediction model is used as the second interval optimal model based on the second accuracy index and the target stable model; and determining the second number of times each air quality prediction model is used as the second interval optimal model based on the third accuracy index and the target stable model. The following steps are taken: First, determine the third iteration of each air quality prediction model as the optimal model for the third interval; second, and third iterations are used to determine the fourth iteration of each air quality prediction model as the optimal model for the fourth interval; third, fifth, and sixth iterations are used to determine the optimal pollution status score for each air quality prediction model for the j-th air quality index; fourth, fifth, and sixth iterations are used to determine the optimal pollution trend score for each air quality prediction model for the j-th air quality index; finally, based on the optimal pollution status score and the optimal pollution trend score, a target air quality prediction model for predicting the j-th air quality index is selected.

[0062] According to one embodiment of the present invention, within the first pollution state reference time period, after each time window slide, a first accurate index of each air quality prediction model can be obtained. The air quality prediction models can be ranked based on the first accurate index. If the difference between the first accurate index of the first-ranked air quality prediction model and the first accurate index of the second-ranked air quality prediction model is greater than the average difference between the first accurate indices of all air quality prediction models in the matching sequence list (when calculating the average, the difference between the first accurate index of the first-ranked air quality prediction model and the first accurate index of the second-ranked air quality prediction model can be included or excluded; the present invention does not impose any restrictions on this), it indicates that the first-ranked air quality prediction model's prediction performance is significantly stronger than other air quality prediction models at that time window position. Therefore, the first-ranked air quality prediction model can be used as the optimal model for the first interval corresponding to this time window slide. Otherwise, the target stable model is used as the optimal model for the first interval. That is, when the difference between the various air quality prediction models is not significant, a more conservative approach is taken, and the target stable model is used as the optimal model for the first interval. The target stable model itself may also rank first, or even significantly outperform other models; the present invention does not impose any restrictions on this. Furthermore, the number of times each air quality prediction model is selected as the optimal model in the first interval can be counted. Air quality prediction models with higher first-time counts show better prediction performance in the reference time period of the first pollution condition.

[0063] According to one embodiment of the present invention, similarly to the above, the second number of times each air quality prediction model is the optimal model for the second interval can be determined based on the ranking of the second accuracy index. The third number of times each air quality prediction model is the optimal model for the third interval can be determined based on the ranking of the third accuracy index. The fourth number of times each air quality prediction model is the optimal model for the fourth interval can be determined based on the ranking of the fourth accuracy index. The fifth number of times each air quality prediction model is the optimal model for the fifth interval can be determined based on the ranking of the fifth accuracy index. The sixth number of times each air quality prediction model is the optimal model for the sixth interval can be determined based on the ranking of the sixth accuracy index.

[0064] According to one embodiment of the present invention, the optimal pollution score for each air quality prediction model to predict the j-th air quality index is determined based on the first, second, and third counts, including: determining the optimal pollution score for the i-th air quality prediction model to predict the j-th air quality index according to formula (3). , (3) in, The first similarity score corresponds to the reference time period of the first pollution condition. The third similarity is the reference time period for the second pollution condition. The fourth similarity is the reference time period for the third pollution status. For the first count, For the second number, This is the third time, and N is the total number of times the time window slides.

[0065] According to one embodiment of the present invention, the first number can be used to describe the performance of the air quality prediction model in a first pollution condition reference time period similar to the pollution condition of a first preset time period; the second number can be used to describe the performance of the air quality prediction model in a second pollution condition reference time period under clean conditions; and the third number can be used to describe the performance of the air quality prediction model in a second pollution condition reference time period under heavy pollution conditions. , and These are the normalized values ​​of the first, second, and third numbers, respectively. The normalized weight represents the similarity between the pollution conditions of the first pollution condition reference time period and the first preset time period. The normalized weights represent the similarity between the reference time period for the second pollution state and the clean time period. The similarity between the reference time period for the third pollution condition and the time period for heavy pollution is weighted after normalization. By weighting the first, second, and third normalized values ​​using the above weights, a comprehensive performance score describing the air quality prediction model under various pollution conditions can be obtained, i.e., the optimal pollution condition score.

[0066] According to one embodiment of the present invention, the optimal pollution trend score for each air quality prediction model to predict the j-th air quality index is determined based on the fourth, fifth, and sixth times, including: determining the optimal pollution status score for the i-th air quality prediction model to predict the j-th air quality index according to formula (4). , (4) in, This is the fourth time. This is the fifth time. This is the sixth time. The second similarity is the reference time period for the first pollution trend. The fifth similarity score corresponds to the second pollution trend reference time period. The sixth similarity is the reference time period for the third pollution trend, and N is the total number of sliding times of the time window.

[0067] According to one embodiment of the present invention, similar to formula (3), the optimal pollution status score is used to describe the comprehensive performance score of the air quality prediction model under multiple pollution trends.

[0068] According to one embodiment of the present invention, the optimal pollution status score and the optimal pollution trend score can be weighted and summed to obtain the comprehensive score of the i-th air quality prediction model for the j-th air quality index. That is, the score is made by combining the prediction performance of the i-th air quality prediction model under various pollution conditions and pollution trends. In another example, the optimal pollution status score, the optimal pollution trend score, and the forecast stability score can also be weighted and summed to obtain the comprehensive score of the i-th air quality prediction model for the j-th air quality index. This comprehensive score can be used to describe the overall prediction performance of the i-th air quality prediction model under various pollution conditions, various pollution trends, and within a first preset time period, and can be used to comprehensively evaluate the i-th air quality prediction model. The air quality prediction model with the highest comprehensive score can be used as the target air quality prediction model for predicting the j-th air quality index. Furthermore, target air quality prediction models for predicting other air quality indices can also be selected.

[0069] In this way, the predictive performance of each air quality prediction model under different pollution conditions can be described by the first, second, and third times, and the predictive performance under different pollution trends can be described by the fourth, fifth, and sixth times. This allows for a comprehensive evaluation of the air quality prediction models under multiple conditions, thereby selecting the air quality prediction model with the strongest overall performance under various conditions and improving the accuracy and stability of subsequent predictions.

[0070] Example 6: According to an embodiment of the present invention, in step S5, the target air quality prediction model can be used to predict air quality index data in the future for a third preset time period. For example, the third preset time period is 3 days long, and the target air quality prediction model can provide air quality index forecast data for the next three days.

[0071] According to an embodiment of the present invention, the air quality multi-model forecast result screening method based on sliding optimal matching can automatically calculate the error index between the predicted air quality index data of the air quality prediction model and the actual measured air quality index data in a first preset time period using a computer. By statistically analyzing the prediction performance of each model at various positions within a sliding window, a target air quality prediction model suitable for predicting specific types of air quality indicators in a third preset time period is automatically and dynamically selected. Prediction is then performed using this model as a benchmark to obtain air quality index forecast data. This improves the real-time performance and efficiency of air quality prediction model screening. From error index calculation and target air quality prediction model screening to air quality index forecast data provision, no manual intervention is required throughout the process, reducing reliance on human intervention. It overcomes the technical bottlenecks of traditional manual operation and static integration, and can adapt to the rapid evolution of air pollution, providing high-precision and high-efficiency forecast support for operational departments. Furthermore, the stability index can be determined through the variance of the error index, and the stability optimization index can be determined through the ranking of the error index. Therefore, when calculating the prediction stability score, the accuracy and stability of the air quality prediction model in predicting air quality indicators are fully considered, providing a data foundation for screening models with accurate and stable prediction performance that can adapt to complex and rapidly changing conditions. Furthermore, under constraints of weather conditions and seasonal data, the overall pollution status within the test time period can be obtained through 2D convolutional network layers, and the pollution change trend within the test time period can be obtained through a self-attention mechanism. This allows for the comprehensive selection of reference time periods based on both pollution status and pollution change trends, providing a richer and more comprehensive reference for selecting suitable air quality prediction models, thus enhancing the reference value and accuracy of air quality prediction model selection. Moreover, during the training phase of the reference selection model, auxiliary training layers can be used to improve the accuracy of the model's feature representation of pollution status and pollution trends, and to enhance its ability to represent weather and seasonal features. This helps to provide more accurate pollution status and pollution change trend features under weather and seasonal constraints, providing an accurate computational foundation for selecting reference time periods. Furthermore, the predictive performance of each air quality prediction model under different pollution conditions can be described by the first, second, and third times, and the predictive performance of each air quality prediction model under different pollution trends can be described by the fourth, fifth, and sixth times. In this way, the predictive performance of the models under multiple conditions can be comprehensively considered, and the air quality prediction models can be fully evaluated. This allows for the selection of the air quality prediction model with the strongest comprehensive performance under multiple conditions, thereby improving the accuracy and stability of subsequent predictions.

[0072] This invention can be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions loaded thereon for performing various aspects of the invention.

[0073] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are merely examples and do not limit the present invention. The objectives of the present invention have been fully and effectively achieved. The functions and structural principles of the present invention have been demonstrated and explained in the embodiments, and any variations or modifications may be made to the implementation of the present invention without departing from the stated principles.

[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for filtering air quality multi-model prediction results based on sliding optimal matching, characterized in that, include: Acquire multiple predicted air quality index data obtained from various air quality prediction models at multiple moments within a first preset time period in the past; Acquire multiple measured air quality data at various times within a first preset time period in the past; Set a time window, wherein the duration of the time window is a second preset duration, and the second preset duration is less than the first preset duration; During the process of the start time of the time window changing within the first preset time period, the target air quality prediction model is selected from multiple air quality prediction models based on the predicted air quality index data and the measured air quality index data within the time window. Based on the target air quality prediction model, air quality index data for multiple moments within a future third preset time period are predicted to obtain air quality index forecast data.

2. The air quality multi-model prediction result screening method based on sliding optimal matching according to claim 1, characterized in that, The end time of the first preset time period is the current time; As the start time of the time window changes within the first preset time period, a target air quality prediction model is selected from multiple air quality prediction models based on the predicted air quality index data and the measured air quality index data within the time window. This includes: When the start time of the time window is aligned with each time point of the first preset time period, the predicted air quality index data of each air quality prediction model at multiple times within the time window, as well as the measured air quality index data, are obtained. Obtain the error index between the predicted air quality index data and the measured air quality index data at multiple times within the time window output by each air quality prediction model; Based on the error index corresponding to each sliding within the first preset time period of the time window, the forecast stability score of each air quality prediction model is determined. Using a trained reference selection model, a reference time period corresponding to the first preset time period is selected from the historical time periods; Based on the measured air quality index data within the reference time period, the predicted air quality index data of each air quality prediction model, and the forecast stability score, the target air quality prediction model is selected.

3. The method for screening air quality multi-model forecast results based on sliding optimal matching according to claim 2, characterized in that, Based on the error index corresponding to each sliding within the first preset time period of the time window, the forecast stability score of each air quality prediction model is determined, including: Based on the error index of the i-th air quality prediction model for the j-th air quality index after each sliding of the time window, determine the error variance of the i-th air quality prediction model for the j-th air quality index. Based on the error variance, determine the stability index of the i-th air quality prediction model for predicting the j-th air quality index; After each slide of the time window, obtain the ranking information of the error index of each air quality prediction model for the j-th air quality index. Based on the ranking information corresponding to multiple sliding of the time window, determine the stable and preferred index for the i-th air quality prediction model to predict the j-th air quality index. Based on the stability index and the stability optimization index, the forecast stability score of the i-th air quality prediction model for the j-th air quality index is determined.

4. The method for screening air quality multi-model forecast results based on sliding optimal matching according to claim 3, characterized in that, Based on the error variance, the stability index of the i-th air quality prediction model for predicting the j-th air quality index is determined, including: According to the formula Determine the stability index of the i-th air quality prediction model for predicting the j-th air quality index. ,in, Let be the error variance of the i-th air quality prediction model for the j-th air quality index, max be the maximum value function, min be the minimum value function, n be the number of air quality prediction models, i ≤ n, and i and n are both positive integers.

5. The method for screening air quality multi-model forecast results based on sliding optimal matching according to claim 3, characterized in that, Based on the ranking information corresponding to multiple sliding of the time window, the stable and optimal indicators for predicting the j-th air quality index by the i-th air quality prediction model are determined, including: Based on the ranking information corresponding to multiple sliding of the time window, among multiple air quality prediction models, the optimal model in the interval with the smallest error index for predicting the j-th air quality index after each sliding of the time window is determined. Based on the interval optimal model corresponding to each slide, determine the maximum number of consecutive times that the i-th air quality prediction model is continuously used as the interval optimal model. According to the formula Determine the stable and optimal index for the i-th air quality prediction model to predict the j-th air quality index. ,in, To predict the j-th air quality index, the maximum number of consecutive times the i-th air quality prediction model is continuously used as the optimal model in the interval, where max is the maximum value function, min is the minimum value function, n is the number of air quality prediction models, i≤n, and i and n are both positive integers.

6. The method for screening air quality multi-model forecast results based on sliding optimal matching according to claim 2, characterized in that, Using a trained reference selection model, a reference time period corresponding to the first preset time period is selected from the historical time period, including: Divide the historical time period into multiple test time periods of the same length as the first preset time period; Based on the measured air quality data, weather data and seasonal data at each moment within the test time period, determine the input vector to be measured at each moment within the test time period, and obtain the reference input vector at each moment within the first preset time period based on the measured air quality data, weather data and seasonal data at each moment within the first preset time period. The first mapping network layer of the selected model is used to map the input vector to be tested to obtain the first input feature information to be tested, and the first reference input vector is also used to map the reference input vector to obtain the first reference input feature information. Multiple first test input feature information are combined into a first test input matrix, and multiple first control input feature information are combined into a first control input matrix; By referencing the 2D convolutional network layers of the selected model, the first test input matrix and the first control input matrix are processed respectively to obtain the test pollution feature vector and the control pollution feature vector respectively. Obtain the first similarity between the feature vector of the pollution to be tested and the feature vector of the control pollution, and determine the test time period with the highest first similarity as the reference time period for the first pollution status; The second mapping network layer of the reference selected model is used to map the input vector to be tested to obtain the second input feature information to be tested, and the control input vector is also mapped to obtain the second control input feature information. By referencing the self-attention mechanism of the selected model, the second test input feature information is processed to obtain the second test fusion feature information, and the test fusion feature information is determined based on the second test fusion feature information. By referencing the self-attention mechanism of the selected model, the second control input feature information is processed to obtain the second control fusion feature information, and the control fusion feature information is determined based on the second control fusion feature information. Obtain the second similarity between the fusion feature information to be tested and the fusion feature information of the control, and determine the test time period with the highest second similarity as the first pollution trend reference time period; Based on the weather data and seasonal data at each moment within the test period, a clean control input vector corresponding to the test period is constructed, and a clean feature vector corresponding to the clean control input vector is obtained. The third similarity between the pollution feature vector to be measured and the clean feature vector is determined, and the time period with the highest third similarity is determined as the second pollution status reference time period. Based on the weather data and seasonal data at each moment within the test period, a pollution control input vector corresponding to the test period is constructed, and the pollution feature vector corresponding to the pollution control input vector is obtained. Determine the fourth similarity between the pollution feature vector to be measured and the pollution feature vector, and determine the time period with the highest fourth similarity as the reference time period for the third pollution status; Based on the weather data and seasonal data at each moment within the test period, a pollution accumulation control input vector corresponding to the test period is constructed, and a pollution accumulation fusion vector corresponding to the pollution accumulation control input vector is obtained. The fifth similarity between the fusion feature information to be tested and the pollution accumulation fusion vector is obtained, and the time period with the highest fifth similarity is determined as the second pollution trend reference time period. Based on the weather data and seasonal data at each moment within the test period, a pollution reduction control input vector corresponding to the test period is constructed, and a pollution reduction fusion vector corresponding to the pollution reduction control input vector is obtained. Obtain the sixth similarity between the fusion feature information to be tested and the pollution reduction fusion vector, and determine the test time period with the highest sixth similarity as the third pollution trend reference time period; The first pollution status reference time period, the first pollution trend reference time period, the second pollution status reference time period, the third pollution status reference time period, the second pollution trend reference time period, and the third pollution trend reference time period are determined as reference time periods.

7. The method for screening air quality multi-model forecast results based on sliding optimal matching according to claim 6, characterized in that, Based on the measured air quality index data within the reference time period and the predicted air quality index data of each air quality prediction model, as well as the forecast stability score, target air quality prediction models are selected, including: The air quality prediction model corresponding to the maximum forecast stability score is determined as the target stable model; Based on the measured air quality index data within the reference time period and the predicted air quality index data of each air quality prediction model, the following accurate indicators are obtained: the first accurate indicator corresponding to each sliding of the time window within the first pollution status reference time period, the second accurate indicator corresponding to each sliding of the time window within the second pollution status reference time period, the third accurate indicator corresponding to each sliding of the time window within the third pollution status reference time period, the fourth accurate indicator corresponding to each sliding of the time window within the first pollution trend reference time period, the fifth accurate indicator corresponding to each sliding of the time window within the second pollution trend reference time period, and the sixth accurate indicator corresponding to each sliding of the time window within the third pollution trend reference time period. Based on the first accuracy index, second accuracy index, third accuracy index, fourth accuracy index, fifth accuracy index, sixth accuracy index, and target stable model, a target air quality prediction model is selected.

8. The method for filtering air quality multi-model forecast results based on sliding optimal matching according to claim 7, characterized in that, Based on the first accuracy index, second accuracy index, third accuracy index, fourth accuracy index, fifth accuracy index, sixth accuracy index, and target stable model, target air quality prediction models are selected, including: After each slide of the time window within the first pollution status reference period, it is determined whether the difference between the maximum and second maximum values ​​of the first accurate indicator is greater than the average difference between the first accurate indicators of each air quality prediction model. If so, the air quality prediction model corresponding to the maximum value of the first accurate indicator is determined as the optimal model of the first interval corresponding to this slide of the time window; otherwise, the target stable model is taken as the optimal model of the first interval. The number of times each air quality prediction model is selected as the optimal model for the first interval is statistically analyzed. Based on the second accuracy index and the target stability model, each air quality prediction model is determined as the second degree of the optimal model in the second interval; Based on the third accuracy index and the target stability model, each air quality prediction model is determined as the third number of the optimal model in the third interval; Based on the fourth accuracy index and the target stability model, each air quality prediction model is determined as the fourth degree of the optimal model in the fourth interval; Based on the fifth accuracy index and the target stability model, the fifth degree of each air quality prediction model is determined as the optimal model for the fifth interval; Based on the sixth accuracy index and the target stability model, the sixth degree of each air quality prediction model is determined as the optimal model for the sixth interval. Based on the first, second, and third counts, determine the optimal pollution score for each air quality prediction model to predict the j-th air quality index. Based on the fourth, fifth, and sixth times, determine the optimal pollution trend score for each air quality prediction model to predict the j-th air quality index. Based on the optimal pollution status score and the optimal pollution trend score, a target air quality prediction model for predicting the j-th air quality index is selected.

9. The method for screening air quality multi-model forecast results based on sliding optimal matching according to claim 8, characterized in that, Based on the first, second, and third data points, determine the optimal pollution score for each air quality prediction model to predict the j-th air quality index, including: According to the formula Determine the optimal pollution score for the i-th air quality prediction model when predicting the j-th air quality index. ,in, The first similarity score corresponds to the reference time period of the first pollution condition. The third similarity is the reference time period for the second pollution condition. The fourth similarity is the reference time period for the third pollution status. For the first count, For the second number, This is the third time, and N is the total number of times the time window slides.

10. The method for screening air quality multi-model forecast results based on sliding optimal matching according to claim 8, characterized in that, Based on the fourth, fifth, and sixth exponents, the optimal pollution trend score for each air quality prediction model to predict the j-th air quality index is determined, including: According to the formula Determine the optimal pollution score for the i-th air quality prediction model when predicting the j-th air quality index. ,in, This is the fourth time. This is the fifth time. This is the sixth time. The second similarity is the reference time period for the first pollution trend. The fifth similarity score corresponds to the second pollution trend reference time period. The sixth similarity is the reference time period for the third pollution trend, and N is the total number of sliding times of the time window.