An air pollution risk prediction and early warning method based on multi-task learning

By adopting a multi-task learning method in air pollution prediction and early warning, combined with multi-objective optimization and improved multi-head attention mechanism, the problems of insufficient utilization of multi-source data and low prediction accuracy in the existing technology are solved, and efficient and accurate prediction and early warning of air pollution risks are achieved.

CN119831104BActive Publication Date: 2025-06-20BEIJING YIQIYIBA TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510009006.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-06-20
Estimated Expiration
2045-01-03

AI Technical Summary

Technical Problem

Existing air pollution prediction methods are difficult to effectively utilize multi-source data, especially when dealing with multi-pollutants and dynamic changes in time and space, prediction accuracy and real-time performance are insufficient, and it is difficult to provide customized risk assessment and early warning information.

Method used

The air pollution risk prediction and early warning method based on multi-task learning is adopted, combined with multi-objective optimization, contrast learning and improved multi-head attention mechanism, dynamically capture the spatio-temporal characteristics of multiple pollutants, and improve prediction accuracy and early warning timeliness through dynamic early warning rules and a comprehensive risk assessment method combining time-space.

Benefits of technology

It significantly improves the accuracy and efficiency of air pollution concentration prediction and risk warning, can handle the concentration prediction of multiple pollutants at the same time, improves the robustness of data utilization and model, and enhances the timeliness and accuracy of risk assessment and early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119831104B_ABST
    Figure CN119831104B_ABST
Patent Text Reader

Abstract

The present invention discloses an air pollution risk prediction and early warning method based on multi-task learning, comprising the following steps: S1. Collect multi-source air pollution monitoring data and perform preprocessing to generate preprocessed data; S2. Construct a multi-task learning model, based on an improved multi-head attention mechanism, and generate a preliminary prediction result through the multi-task learning model; S3. Optimize the parameters of the shared feature extraction layer and the task-specific prediction layer; S4. Optimize the preliminary prediction result to generate an optimized prediction result; S5. Generate a risk level, and generate a risk assessment report by integrating risk information in the time and space dimensions; S6. Design a trigger condition according to the risk level, define a joint early warning rule, judge the trigger of the early warning through a dynamic threshold, and generate early warning information. The present invention combines multi-task learning with an improved multi-head attention mechanism to achieve multi-pollutant concentration prediction and risk early warning, and has the advantages of high timeliness, high prediction accuracy, and comprehensive risk assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of air pollution prediction and early warning, and particularly to an air pollution risk prediction and early warning method based on multi-task learning. Background Art

[0002] Traditional air pollution prediction methods mainly rely on statistical models and physical models. Statistical models, such as time series analysis models and regression models, use historical monitoring data for trend prediction. These methods perform well when the data is relatively stable, but they are insufficient in dealing with complex spatio-temporal correlations and the interaction capabilities between pollutants. Physical models rely on the principles of atmospheric chemistry and meteorological dynamics, and predict by simulating the transport and transformation processes of pollutants. Such models require accurate initial conditions and a large amount of computing resources, and have poor adaptability to multi-source data. It is difficult to meet the actual requirements in terms of prediction real-time performance and accuracy.

[0003] In recent years, with the rapid development of artificial intelligence technology, air pollution prediction methods based on machine learning and deep learning have gradually emerged. These methods can mine potential laws using large-scale multi-source data and achieve high-precision prediction by combining non-linear modeling capabilities. However, existing single-task learning-based models usually only focus on a specific pollutant or a single task, and fail to fully utilize the correlations between pollutants and the characteristics of spatio-temporal dynamic changes. This single-task prediction method easily leads to low data utilization rate, redundant model calculations, and bottlenecks in prediction accuracy. In addition, single-task models are difficult to handle multi-task objectives simultaneously and cannot effectively support regional-level comprehensive air quality risk assessment and early warning.

[0004] Although existing multi-task learning-based models have solved the above problems to a certain extent, they also have significant limitations. On the one hand, the shared feature extraction ability of multi-task learning models is limited, and it is difficult to accurately capture the complex spatio-temporal dynamic relationships of multiple pollutants when dealing with high-dimensional heterogeneous data. On the other hand, the feature extraction in existing methods usually adopts a standard attention mechanism and is not optimized for the air pollution scenario. For example, the deep associations between time series and spatial features are not fully considered. In addition, traditional multi-task learning methods often adopt a single loss optimization objective, lacking dynamic weight adjustment for task collaboration and effective optimization of feature discriminability, which results in poor performance of the model in the face of complex multi-task scenarios.

[0005] In the field of air pollution risk assessment and early warning, existing technologies mainly rely on fixed risk assessment criteria and simple rule-based triggers for early warning. This approach cannot dynamically adapt to the spatio-temporal variations and regional differences in pollutant concentrations, making it difficult to provide customized and accurate risk assessment reports and early warning information for different regions. In addition, existing visualization methods are mostly presented in the form of static charts and fail to dynamically display the evolution process of pollution risks by integrating the temporal and spatial dimensions, thus making it difficult to meet the actual needs of decision-making support.

[0006] Therefore, how to provide an air pollution risk prediction and early warning method based on multi-task learning is an urgent problem to be solved by those skilled in the art. Summary of the Invention

[0007] An object of the present invention is to propose an air pollution risk prediction and early warning method based on multi-task learning. The present invention innovatively combines multi-objective optimization, contrastive learning, and an improved multi-head attention mechanism to design an efficient prediction and early warning model capable of dynamically capturing the spatio-temporal characteristics of multiple pollutants. By introducing dynamic early warning rules and a time-space integrated comprehensive risk assessment method, the accuracy of air pollution prediction and the timeliness of early warning are improved.

[0008] An air pollution risk prediction and early warning method based on multi-task learning according to an embodiment of the present invention includes the following steps:

[0009] S1. Collect multi-source air pollution monitoring data, including pollutant concentration data, meteorological data, and geospatial data, and preprocess the multi-source air pollution monitoring data to generate preprocessed data;

[0010] S2. Construct a multi-task learning model, input the preprocessed data into the multi-task learning model, and based on the improved multi-head attention mechanism, generate a shared feature matrix through the shared feature extraction layer of the multi-task learning model, and input the shared feature matrix into the task-specific prediction layer of the multi-task learning model to generate a preliminary prediction result;

[0011] S3. Define the multi-objective optimization function of the multi-task learning model, construct a contrastive learning loss function according to the gradient of the shared feature matrix, and comprehensively optimize the parameters of the shared feature extraction layer and the task-specific prediction layer by combining the multi-objective optimization function and the contrastive learning loss function;

[0012] S4. Apply the optimized parameters of the shared feature extraction layer and the task-specific prediction layer to the multi-task learning model to optimize the preliminary prediction result and generate an optimized prediction result, where the optimized prediction result includes the concentration and spatio-temporal change trend of the pollutant;

[0013] S5. Map the predicted pollutant concentrations to the preset risk assessment criteria to generate risk levels for different regions and different times. Calculate the dynamic risk level based on the time series of the optimized prediction results, calculate the spatial risk distribution of the concentration prediction results of pollutants at different geographical locations, and generate a risk assessment report by integrating the risk information in the time and space dimensions;

[0014] S6. Design trigger conditions according to the risk levels, define joint warning rules, judge and trigger warnings through dynamic thresholds, and generate warning information.

[0015] Optionally, the S2 specifically includes:

[0016] S21. Construct a multi-task learning model, which includes a shared feature extraction layer and a task-specific prediction layer. The shared feature extraction layer extracts the spatio-temporal features of pollutants, and the task-specific prediction layer predicts the concentrations and spatio-temporal change trends of multiple pollutants respectively;

[0017] S22. Input the generated preprocessed data X into the multi-task learning model, perform multi-scale convolution operations on the preprocessed data X to capture the local relationships of the spatio-temporal features of pollutants, and splice the convolution results at different scales to generate a convolution feature matrix. The multi-scale convolution operations use convolution kernels with sizes of 3×3, 5×5, and 7×7:

[0018]

[0019] Among them, represents the convolution results at different scales, represents the weight matrix of the kth convolution kernel, represents the bias vector, σ represents the ReLU activation function, F conv represents the convolution feature matrix, Concat represents the splicing operation, and * represents the convolution operation;

[0020] S23. Add time embedding information to the convolution feature matrix F conv to generate a time-enhanced feature matrix F time :

[0021]

[0022] Among them, F time [i,j] represents the time-enhanced feature matrix, t i represents the timestamp corresponding to sample i, j represents the feature dimension index, and d represents the total number of features;

[0023] S24. For the time-enhanced feature matrix F timeApply the improved multi-head attention mechanism to deeply fuse time series and spatial correlation information. Based on the standard multi-head attention mechanism, add a time weight matrix to adjust features and calculate the query, key, and value matrices:

[0024] Q = W q ·(F time +δ q ), K = W k ·(F time +δ k ), V = W v ·(F time +δ v );

[0025] Among them, Q, K, and V respectively represent the query, key, and value matrices, W q , W k , and W v respectively represent the trainable weight matrices of the query, key, and value, δ q , δ k , and δ v represent the dynamic adjustment offsets;

[0026] S25. Introduce the time decay factor and spatial weight:

[0027]

[0028] Among them, Attention(Q, K, V) represents the multi-head attention mechanism formula, ⊙ represents element-wise multiplication, d k represents the dimension of the key vector, W s represents the spatial weight, b s represents the bias, T represents the transpose, α t represents the time decay factor, exp represents the natural exponential function, t i and t q represent the timestamps corresponding to sample i and q, λ represents the time decay coefficient, and softmax represents normalization;

[0029] S26. Add a hierarchical fusion module after the output of the multi-head attention mechanism to combine features at different time scales and generate a shared feature matrix:

[0030] F shared = β·Attention(Q, K, V)+(1 - β)·F time ;

[0031] β = sigmoid(W β ·F time +b β );

[0032] Among them, Fshared denotes the shared feature matrix, W β and b β denote the fusion weight matrix and bias, sigmoid denotes the activation function, and β denotes the weight of the fusion operation;

[0033] S27. Input the shared feature matrix F shared into the task-specific prediction layer to generate preliminary prediction results for each task:

[0034]

[0035] g t = sigmoid(W g ·z t + b g );

[0036] where denotes the preliminary prediction result, W t and b t denote the task-specific convolutional weight matrix and bias vector, g t denotes the gating factor, W g and b g denote the gating weight matrix and bias vector.

[0037] Optionally, the S3 specifically includes:

[0038] S31. Define the multi-objective optimization function of the multi-task learning model. The multi-objective optimization function guides the joint optimization of the shared feature extraction layer and the task-specific prediction layer. The optimization objectives of the multi-objective optimization function include minimizing the loss of each task and maximizing the collaborative effect between tasks:

[0039]

[0040] where denotes the multi-objective optimization function, denotes the loss function of task t, which is the cross-entropy loss function, γ t denotes the task weight, denotes the regularization coefficient, and respectively denote the shared feature matrices of task i and task j, denotes the square of the two-norm, and T denotes the total number of tasks;

[0041] S32. Dynamically adjust the task weight γ according to the change of the task loss function t :

[0042]

[0043] Among them, exp represents the natural exponential function;

[0044] S33. Calculate the gradient of the shared feature extraction layer and generate task-specific positive and negative sample features:

[0045]

[0046] Among them, represents the positive sample feature, represents the negative sample feature, L i represents the loss function of task i, L j represents the loss function of task j, ∈ represents the step size;

[0047] S34. Construct a contrastive learning loss function to further optimize the discriminability and collaboration of features between tasks:

[0048]

[0049] Among them, represents the contrastive learning loss function, sim represents the cosine similarity, F i represents the features extracted by the shared feature extraction layer;

[0050] S35. Combine the contrastive learning loss function and the multi-objective optimization function to optimize the parameters of the shared feature extraction layer and the task-specific prediction layer parameters:

[0051]

[0052]

[0053] Among them, represents the total loss function, λ contrast represents the weight of the contrastive learning loss function, θ s represents the parameters of the shared feature extraction layer, θ t represents the parameters of the task-specific prediction layer, η represents the learning rate, which is used to control the step size of parameter update.

[0054] Optionally, the S5 specifically includes:

[0055] S51. Map the pollutant concentration to the risk level according to the preset risk assessment criteria;

[0056] S52. Calculate the corresponding risk level using the predicted pollutant concentration:

[0057]

[0058] Among them, R(y) represents the risk level, y represents the predicted pollutant concentration, k represents the index of the risk level, K represents the number of risk levels, C kIndicates the risk level demarcation point;

[0059] S53. Calculate the overall risk level of the region according to the risk levels of each task and perform normalization;

[0060] S54. Calculate the dynamic risk level based on the time series of the optimized prediction results, and calculate the spatial risk distribution for the prediction results of pollutant concentrations at different geographical locations;

[0061] S55. Integrate the risk information in the time and space dimensions of all tasks to generate a risk assessment report, where the risk assessment report includes the risk levels of various pollutants, the regional comprehensive risk level, the time series risk change trend, and the spatial distribution map of risk levels.

[0062] Optionally, the specific content of S6 includes:

[0063] S61. Set the corresponding warning thresholds for risk levels:

[0064]

[0065] where T k represents the warning threshold for risk level k, k represents the index of the risk level, K represents the number of risk levels, C min and C max respectively represent the minimum and maximum values of pollutant concentrations;

[0066] S62. Set dynamic warning rules, where the dynamic warning rules include multi-dimensional trigger conditions of time, space, and pollutant concentration:

[0067]

[0068] where N t represents the time step, represents the risk level, (x j , y j ) represents the regional grid, represents the risk level of the regional grid, T time and T space respectively represent the warning thresholds for the risk levels in the time dimension and the space dimension;

[0069] S63. Perform a composite design on the trigger conditions of time, space, and pollutant concentration to define joint warning rules:

[0070] E alert = I(R overall > T overall and R time > T time and R space > T space);

[0071] Among them, E alert represents the combined warning value, I(·) represents the indicator function, and when the condition is satisfied, E alert is 1, otherwise it is 0, R overall 、R time and R space respectively represent the overall risk level, the risk level in the time dimension, and the risk level in the space dimension, T overall represents the warning threshold of the overall risk level, and and represents "AND";

[0072] S64. When the combined warning value is 1, warning information is generated, and the warning information includes the warning type, the affected area, the warning time, and the pollutant.

[0073] The beneficial effects of the present invention are as follows:

[0074] First of all, by introducing a multi-task learning model, the present invention significantly improves the accuracy and efficiency of air pollution concentration prediction and risk warning. Compared with the traditional single-task learning model, the present invention can simultaneously process the concentration prediction tasks of multiple pollutants and capture the complex correlations and spatio-temporal dynamic characteristics among pollutants. This collaborative modeling method not only improves the data utilization rate but also avoids the problems of data redundancy and waste of computing resources in the single-task model.

[0075] Secondly, through the improvement of the multi-head attention mechanism, the present invention deeply integrates time series information and spatial features, making the model perform more excellently when dealing with multi-source heterogeneous data. The introduction of the time decay factor and the spatial weight matrix enables the attention mechanism to adaptively adjust the feature weights according to the spatio-temporal dynamic changes of pollutants, ensuring that the extracted features are more spatio-temporally correlated and have predictive value. In addition, the design of the hierarchical fusion module enables the model to comprehensively integrate features at different time scales, improving the expression ability of the shared feature extraction layer and laying a foundation for the accurate prediction of pollutant concentration and risk level.

[0076] In addition, during the optimization process, the present invention combines multi-objective optimization and contrast learning techniques, not only minimizing the independent losses of each task but also optimizing the collaboration and feature distinguishability between tasks through the construction of specific positive and negative samples. The mechanism of dynamically adjusting the task weights ensures a reasonable allocation of attention to each task at different stages, thereby improving the overall optimization effect of the model. This multi-objective optimization method makes up for the deficiency of insufficient collaboration between tasks in traditional methods, making the prediction results more robust and reliable in complex scenarios.

[0077] Finally, in terms of air pollution risk assessment and early warning, the present invention proposes dynamic early warning rules and a joint early warning mechanism, which automatically adjusts the early warning threshold and trigger conditions according to the changes in pollutant concentrations in different regions and time periods. Compared with the traditional early warning method with fixed rules, the dynamic early warning mechanism of the present invention is more flexible, can identify high-risk pollution regions and time periods, and improves the timeliness and accuracy of risk early warning. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings:

[0079] Figure 1 is a flowchart of a method for predicting and warning air pollution risks based on multi-task learning proposed by the present invention;

[0080] Figure 2 is a flowchart of the working process of an improved multi-head attention mechanism for a method for predicting and warning air pollution risks based on multi-task learning proposed by the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0081] Now, the present invention will be further described in detail with reference to the drawings. These drawings are all simplified schematic diagrams, only showing the basic structure of the present invention in a schematic manner, so they only show the components related to the present invention.

[0082] Refer to Figure 1 and Figure 2 , a method for predicting and warning air pollution risks based on multi-task learning, includes the following steps:

[0083] S1. Collect multi-source air pollution monitoring data, including pollutant concentration data, meteorological data, and geospatial data, and preprocess the multi-source air pollution monitoring data to generate preprocessed data;

[0084] S2. Construct a multi-task learning model, input the preprocessed data into the multi-task learning model, based on the improved multi-head attention mechanism, generate a shared feature matrix through the shared feature extraction layer of the multi-task learning model, and input the shared feature matrix into the task-specific prediction layer of the multi-task learning model to generate a preliminary prediction result;

[0085] S3. Define the multi-objective optimization function of the multi-task learning model, construct a contrastive learning loss function according to the gradient of the shared feature matrix, and comprehensively optimize the parameters of the shared feature extraction layer and the task-specific prediction layer by combining the multi-objective optimization function and the contrastive learning loss function;

[0086] S4. Apply the optimized parameters of the shared feature extraction layer and task-specific prediction layer to the multi-task learning model to optimize the preliminary prediction results and generate optimized prediction results, where the optimized prediction results include the concentration and spatio-temporal change trend of pollutants.

[0087] S5. Map the predicted pollutant concentration to a preset risk assessment standard to generate risk levels for different regions and different times. Calculate the dynamic risk level based on the time series of the optimized prediction results, and calculate the spatial risk distribution for the concentration prediction results of pollutants at different geographical locations. Integrate the risk information in the time and space dimensions to generate a risk assessment report.

[0088] S6. Design trigger conditions according to the risk level, define joint warning rules, judge and trigger warnings through dynamic thresholds, and generate warning information.

[0089] In this embodiment, the S2 specifically includes:

[0090] S21. Construct a multi-task learning model, which includes a shared feature extraction layer and a task-specific prediction layer. The shared feature extraction layer extracts the spatio-temporal features of pollutants, and the task-specific prediction layer predicts the concentration and spatio-temporal change trend of multiple pollutants respectively.

[0091] S22. Input the generated preprocessed data X into the multi-task learning model, perform multi-scale convolution operations on the preprocessed data X to capture the local relationships of the spatio-temporal features of pollutants, and splice the convolution results of different scales to generate a convolution feature matrix. The multi-scale convolution operations use convolution kernels with sizes of 3×3, 5×5, and 7×7:

[0092]

[0093] Among them, represents the convolution results of different scales, represents the weight matrix of the k-th convolution kernel, represents the bias vector, σ represents the ReLU activation function, F conv represents the convolution feature matrix, Concat represents the splicing operation, and * represents the convolution operation;

[0094] S23. Add time embedding information to the convolution feature matrix F conv to generate a time-enhanced feature matrix F time :

[0095]

[0096] Among them, F time [i,j] represents the time-enhanced feature matrix, t iDenote the timestamp corresponding to sample i, j as the feature dimension index, and d as the total number of feature dimensions;

[0097] S24. Apply an improved multi-head attention mechanism to deeply fuse time series and spatial correlation information. Based on the standard multi-head attention mechanism, add a time weight matrix to adjust the features, and calculate the query, key, and value matrices: time Apply an improved multi-head attention mechanism to deeply fuse time series and spatial correlation information. Based on the standard multi-head attention mechanism, add a time weight matrix to adjust the features, and calculate the query, key, and value matrices:

[0098] Q = W q ·(F time + δ q ), K = W k ·(F time + δ k ), V = W v ·(F time + δ v );

[0099] Among them, Q, K, and V represent the query, key, and value matrices respectively, W q , W k and W v represent the trainable weight matrices of the query, key, and value respectively, δ q , δ k and δ v represent the dynamic adjustment offsets;

[0100] S25. Introduce a time decay factor and a spatial weight:

[0101]

[0102] Among them, Attention(Q, K, V) represents the formula of the multi-head attention mechanism, ⊙ represents element-wise multiplication, d k represents the dimension of the key vector, W s represents the spatial weight, b s represents the bias, T represents the transpose, α t represents the time decay factor, exp represents the natural exponential function, t i and t q represent the timestamps corresponding to sample i and q respectively, λ represents the time decay coefficient, and softmax represents normalization;

[0103] S26. Add a hierarchical fusion module after the output of the multi-head attention mechanism to combine features at different time scales and generate a shared feature matrix:

[0104] F shared = β·Attention(Q, K, V)+(1 - β)·F time ;

[0105] β = sigmoid(Wβ ·F time +b β );

[0106] Among them, F shared represents the shared feature matrix, W β and b β represent the fusion weight matrix and bias, sigmoid represents the activation function, and β represents the weight of the fusion operation;

[0107] S27. Input the shared feature matrix F shared into the task-specific prediction layer to generate preliminary prediction results for each task:

[0108]

[0109] g t = sigmoid(W g ·z t +b g );

[0110] Among them, represents the preliminary prediction result, W t and b t represent the task-specific convolution weight matrix and bias vector, g t represents the gating factor, W g and b g represent the gating weight matrix and bias vector.

[0111] In this embodiment, the S3 specifically includes:

[0112] S31. Define the multi-objective optimization function of the multi-task learning model. The multi-objective optimization function guides the joint optimization of the shared feature extraction layer and the task-specific prediction layer. The optimization objectives of the multi-objective optimization function include minimizing the loss of each task and maximizing the collaborative effect between tasks:

[0113]

[0114] Among them, represents the multi-objective optimization function, represents the loss function of task t, which is the cross-entropy loss function, γ t represents the task weight, represents the regularization coefficient, and respectively represent the shared feature matrices of task i and task j, represents the square of the two-norm, and T represents the total number of tasks;

[0115] S32. Dynamically adjust the task weight γ according to the change of the task loss function t :

[0116]

[0117] Among them, exp represents the natural exponential function;

[0118] S33. Calculate the gradient of the shared feature extraction layer and generate task-specific positive and negative sample features:

[0119]

[0120] Among them, represents the positive sample feature, represents the negative sample feature, L i represents the loss function of task i, L j represents the loss function of task j, ∈ represents the step size;

[0121] S34. Construct a contrastive learning loss function to further optimize the discriminability and cooperation of features between tasks:

[0122]

[0123] Among them, represents the contrastive learning loss function, sim represents the cosine similarity, F i represents the features extracted by the shared feature extraction layer;

[0124] S35. Combine the contrastive learning loss function and the multi-objective optimization function to optimize the parameters of the shared feature extraction layer and the task-specific prediction layer parameters:

[0125]

[0126]

[0127] Among them, represents the total loss function, λ contrast represents the weight of the contrastive learning loss function, θ s represents the parameters of the shared feature extraction layer, θ t represents the parameters of the task-specific prediction layer, η represents the learning rate, which is used to control the step size of parameter update.

[0128] In this embodiment, the S5 specifically includes:

[0129] S51. Map the pollutant concentration to the risk level according to the preset risk assessment criteria;

[0130] S52. Calculate the corresponding risk level using the predicted pollutant concentration:

[0131]

[0132] Among them, R(y) represents the risk level, y represents the predicted pollutant concentration, k represents the index of the risk level, K represents the number of risk levels, and C k represents the risk level demarcation point;

[0133] S53. Calculate the overall risk level of the region according to the risk levels of each task, and perform normalization;

[0134] S54. Calculate the dynamic risk level according to the time series of the optimized prediction results, and calculate the spatial risk distribution for the prediction results of pollutant concentrations at different geographical locations;

[0135] S55. Integrate the risk information in the time and space dimensions of all tasks to generate a risk assessment report, and the risk assessment report includes the risk levels of various pollutants, the comprehensive regional risk level, the risk change trend of the time series, and the spatial distribution map of the risk levels.

[0136] In this embodiment, the S6 specifically includes:

[0137] S61. Set the corresponding warning thresholds for the risk levels:

[0138]

[0139] Among them, T k represents the warning threshold for the risk level k, k represents the index of the risk level, K represents the number of risk levels, and C min and C max respectively represent the minimum value and the maximum value of the pollutant concentration;

[0140] S62. Set dynamic warning rules, and the dynamic warning rules include multi-dimensional trigger conditions of time, space, and pollutant concentration:

[0141]

[0142] Among them, N t represents the time step, represents the risk level, (x j , y j ) represents the regional grid, represents the risk level of the regional grid, T time and T space respectively represent the warning thresholds of the risk levels in the time dimension and the space dimension;

[0143] S63. Perform a composite design on the trigger conditions of time, space, and pollutant concentration, and define the combined warning rules:

[0144] E alert = I(R overall > Toverall and R time >T time and R space >T space );

[0145] Among them, E alert represents the joint warning value, I(·) represents the indicator function, and when the condition is met, E alert is 1, otherwise it is 0, R overall , R time and R space They represent the overall risk level, time dimension risk level and space dimension risk level respectively, T overall Indicates the warning threshold of the overall risk level, and means "and";

[0146] S64. When the joint warning value is 1, generate warning information, which includes warning type, involved area, warning time and pollutants.

[0147] Embodiment 1:

[0148] In the embodiments of the present invention, the air pollution prediction and early warning of a certain city is used as a research scenario to demonstrate the application effect based on the multi-task learning model. The city is a typical industrialized city, which is affected by PM2.5, PM10, NO2, and SO2 pollutants all year round. At the same time, it is accompanied by complex meteorological conditions and geographic spatial characteristics, and the prediction of air quality risks faces great challenges. The performance of traditional single-task learning models and fixed rule early warning methods in this scenario is subject to many limitations, and it is impossible to achieve high-precision, multi-dimensional pollution risk prediction and early warning.

[0149] The city's air quality monitoring data comes from 50 monitoring stations across the city, including pollutant concentration data (PM2.5, PM10, NO2, SO2), meteorological data (temperature, humidity, wind speed, precipitation) and geospatial data (station latitude and longitude, elevation, etc.) from January to December 2023. The monitoring data is collected once an hour, and a total of 438,000 records are obtained. After data cleaning, 426,780 valid records are retained.

[0150] In the present invention, the above-mentioned data is first cleaned and preprocessed, including handling missing values, outliers, and unifying the time step and spatial resolution. After preprocessing, the data is input into a multi-task learning model. The model extracts the spatio-temporal features of multiple pollutants through a shared feature extraction layer, and uses task-specific prediction layers to generate concentration prediction values of PM2.5, PM10, NO2, and SO2 respectively. The model adopts an improved multi-head attention mechanism, and through the introduction of a time decay factor and a spatial weight matrix, deeply fuses the time series and spatial features. At the same time, combined with a contrastive learning loss function and a multi-objective optimization function, the task weights are dynamically adjusted, improving the prediction accuracy and coordination of the model. According to the predicted pollutant concentration and the risk assessment standard, the risk level of each site is generated. On this basis, the overall regional risk level is calculated, and a risk assessment report is generated in combination with the time series and spatial distribution information. In addition, dynamic warning rules are used to trigger warnings, and the warning information includes the pollutant type, the affected area, and the warning duration.

[0151] Table 1 Comparison Table of Experimental Data

[0152] Index Traditional method Method of the present invention Improvement rate (%) Mean absolute error (PM2.5) 9.45 6.72 28.89 Mean absolute error (PM10) 12.31 8.59 30.22 <![CDATA[Mean Absolute Error (NO2)]]> 6.83 4.97 27.21 <![CDATA[Mean Absolute Error (SO2)]]> 4.27 3.15 26.22 Average early warning lead time (hours) 2.5 5.3 112.00 Early warning accuracy rate (%) 78.6 93.2 18.60 Risk assessment generation time (seconds / time) 8.2 3.6 56.10 Dynamic visualization generation time (seconds / time) 15.7 6.8 56.69 Risk report coverage rate (%) 82.3 98.1 19.23

[0153] It can be seen from the comparison table of experimental data that the present invention shows significant advantages over traditional methods in multiple core indicators of air pollution risk prediction and warning. First of all, in terms of prediction accuracy, the mean absolute error of the present invention for the four pollutants of PM2.5, PM10, NO2, and SO2 has been reduced by 28.89%, 30.22%, 27.21%, and 26.22% respectively, improving the accuracy of the model in predicting the concentrations of multiple pollutants. This improvement benefits from the multi-task learning design in the shared feature extraction layer of the present invention, and the ability to deeply fuse spatio-temporal features through the improved multi-head attention mechanism.

[0154] In terms of warning response performance, the dynamic warning rules of the present invention have increased the average warning lead time from 2.5 hours of the traditional method to 5.3 hours, with a growth rate of 112%. This result shows that the present invention can identify high-risk areas earlier and provide more timely warning information for relevant departments. In addition, the warning accuracy rate has been increased from 78.6% of the traditional method to 93.2%, an increase of 18.6 percentage points, fully demonstrating the reliability and effectiveness of the present invention in risk identification.

[0155] The efficiency of risk assessment has also been significantly improved. The time taken by the present invention to generate a complete risk assessment report has been shortened from 8.2 seconds by the traditional method to 3.6 seconds, with an efficiency improvement of 56.1%. The dynamic visualization generation time has also been significantly reduced, from 15.7 seconds to 6.8 seconds, with an improvement rate of 56.69%. This efficiency improvement benefits from the efficient multi-task learning model and contrastive learning optimization method in the present invention, making the utilization of computing resources more sufficient.

[0156] In terms of the coverage rate of the risk report, the present invention can cover 98.1% of the area, an increase of 19.23% compared with 82.3% of the traditional method. This improvement benefits from the enhanced risk analysis ability of the present invention in the time and space dimensions, especially the innovation in dynamic early warning rule design and multi-source data fusion. Generally speaking, the present invention not only greatly exceeds the traditional method in prediction accuracy and early warning response speed, but also shows excellent performance in the efficiency of risk assessment and dynamic visualization generation. These advantages fully reflect the technical value and practical significance of the present invention in solving the key problems of air pollution prediction and early warning.

[0157] The above is only the preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention, according to the technical solution and inventive concept of the present invention, makes equivalent substitutions or changes, and should be covered by the protection scope of the present invention.

Claims

1. An air pollution risk prediction and early warning method based on multi-task learning, characterized in that: The steps include: S1. Collect multi-source air pollution monitoring data, including pollutant concentration data, meteorological data and geographic space data, and pre-process the multi-source air pollution monitoring data to generate pre-processed data; S2. Build a multi-task learning model, input the preprocessed data into the multi-task learning model, generate a shared feature matrix through the shared feature extraction layer of the multi-task learning model based on the improved multi-head attention mechanism, and input the shared feature matrix into the task-specific prediction layer of the multi-task learning model to generate preliminary prediction results; S3. Define a multi-objective optimization function for the multi-task learning model, construct a contrastive learning loss function based on the gradient of the shared feature matrix, and optimize the parameters of the shared feature extraction layer and the task-specific prediction layer by combining the multi-objective optimization function and the contrastive learning loss function. S4, applying the optimized shared feature extraction layer and task-specific prediction layer parameters to the multi-task learning model, optimizing the preliminary prediction results, and generating optimized prediction results, wherein the optimized prediction results include the concentration and spatiotemporal variation trend of the pollutants; S5. Map the predicted pollutant concentrations with the preset risk assessment standards to generate risk levels for different regions and at different times. Calculate the dynamic risk level based on the time series of the optimized prediction results. Calculate the spatial risk distribution of the predicted concentration results of pollutants in different geographical locations. Generate a risk assessment report by integrating the risk information in time and space dimensions. S6. Design trigger conditions according to risk levels, define joint warning rules, trigger warnings through dynamic threshold judgments, and generate warning information.

2. The air pollution risk prediction and early warning method based on multi-task learning according to claim 1 is characterized in that: The S2 specifically includes: S21, constructing a multi-task learning model, wherein the multi-task learning model includes a shared feature extraction layer and a task-specific prediction layer, wherein the shared feature extraction layer extracts the spatiotemporal characteristics of pollutants, and the task-specific prediction layer predicts the concentrations and spatiotemporal variation trends of multiple pollutants respectively; S22. The generated preprocessed data X is input into the multi-task learning model, and a multi-scale convolution operation is performed on the preprocessed data X to capture the local relationship of the spatiotemporal characteristics of the pollutants, and the convolution results of different scales are spliced ​​to generate a convolution feature matrix. The multi-scale convolution operation uses convolution kernel sizes of 3×3, 5×5 and 7×7: in, Represents the convolution results of different scales, represents the weight matrix of the kth convolution kernel, represents the bias vector, σ represents the ReLU activation function, F conv Represents the convolution feature matrix, Concat represents the concatenation operation, and * represents the convolution operation; S23, convolution feature matrix F conv Add time embedding information to generate the time-enhanced feature matrix F time : Among them, F time [i,j] represents the time-enhanced feature matrix, t i represents the timestamp corresponding to sample i, j represents the feature dimension index, and d represents the total feature dimension; S24, time-enhanced feature matrix F time Apply the improved multi-head attention mechanism to deeply integrate time series and spatial correlation information. On the basis of the standard multi-head attention mechanism, add the time weight matrix to adjust the features and calculate the query, key and value matrices: Q=W q ·(F time +δ q ),K=W k ·(F time +δ k ),V=W v ·(F time +δ v ); Where Q, K, and V represent query, key, and value matrices, respectively, and W q , W k and W v Trainable weight matrices representing queries, keys, and values, δ q , δ k and δ v Indicates dynamic adjustment of offset; S25, introduce time decay factor and space weight: Among them, Attention(Q,K,V) represents the multi-head attention mechanism formula, ⊙ represents element-by-element multiplication, d k represents the dimension of the key vector, W s represents the spatial weight, b s represents bias, T represents transpose, α t represents the time decay factor, exp represents the natural exponential function, t i and t q represents the timestamp corresponding to samples i and q, λ represents the time decay coefficient, and softmax represents normalization; S26. Add a hierarchical fusion module after the output of the multi-head attention mechanism to combine features of different time scales and generate a shared feature matrix: F shared =β·Attention(Q,K,V)+(1-β)·F time ; β=sigmoid(W β ·F time +b β ); Among them, F shared represents the shared feature matrix, W β and b β represents the fusion weight matrix and bias, sigmoid represents the activation function, and β represents the weight of the fusion operation; S27, share the feature matrix F shared Enter the task-specific prediction layer to generate preliminary prediction results for each task: g t =sigmoid(W g ·z t +b g ); in, represents the preliminary prediction result, W t and b t represents the task-specific convolutional weight matrix and bias vector, g t represents the gating factor, W g and b g Represents the gating weight matrix and bias vector.

3. The air pollution risk prediction and early warning method based on multi-task learning according to claim 1 is characterized in that: The S3 specifically includes: S31. Define a multi-objective optimization function of a multi-task learning model, wherein the multi-objective optimization function guides the joint optimization of a shared feature extraction layer and a task-specific prediction layer, and the optimization objectives of the multi-objective optimization function include minimizing the loss of each task and maximizing the synergy between tasks: in, represents the multi-objective optimization function, represents the loss function of task t, which is the cross entropy loss function, γ t represents the task weight, represents the regularization coefficient, and Represent the shared feature matrices of task i and task j respectively, represents the square of the second norm, T represents the total number of tasks; S32. Dynamically adjust task weight γ using changes in task loss function t : Where, exp represents the natural exponential function; S33. Calculate the gradient of the shared feature extraction layer and generate task-specific positive and negative sample features: in, represents the positive sample feature, Represents the negative sample feature, L i represents the loss function of task i, L j represents the loss function of task j, ∈ represents the step size; S34. Construct a contrastive learning loss function to further optimize the distinguishability and synergy of features between tasks: in, represents the contrastive learning loss function, sim represents the cosine similarity, F i Represents the features extracted by the shared feature extraction layer; S35. Combine the contrastive learning loss function and the multi-objective optimization function to optimize the shared feature extraction layer parameters and the task-specific prediction layer parameters: in, represents the total loss function, λ contrast represents the weight of the contrastive learning loss function, θ s represents the shared feature extraction layer parameters, θ t represents the task-specific prediction layer parameters, and η represents the learning rate, which is used to control the step size of parameter update.

4. The air pollution risk prediction and early warning method based on multi-task learning according to claim 1 is characterized in that: The S5 specifically includes: S51. Map pollutant concentrations to risk levels according to preset risk assessment standards; S52. Calculate the corresponding risk level using the predicted pollutant concentration: Where R(y) represents the risk level, y represents the predicted pollutant concentration, k represents the index of the risk level, K represents the number of risk levels, and C k Indicates the risk level cut-off point; S53. Calculate the overall risk level of the region according to the risk level of each task and normalize it; S54, calculating the dynamic risk level based on the time series of the optimized prediction results, and calculating the spatial risk distribution of the pollutant concentration prediction results at different geographical locations; S55. Integrate the risk information of all tasks in time and space dimensions to generate a risk assessment report, which includes the risk level of each pollutant, the regional comprehensive risk level, the time series risk change trend and the risk level spatial distribution map.

5. The air pollution risk prediction and early warning method based on multi-task learning according to claim 1 is characterized in that: The S6 specifically includes: S61. Set the corresponding warning threshold for the risk level: Among them, T k represents the warning threshold of risk level k, k represents the index of risk level, K represents the number of risk levels, C min and C max Respectively represent the minimum and maximum values ​​of pollutant concentration; S62. Setting dynamic warning rules, wherein the dynamic warning rules include multi-dimensional triggering conditions of time, space and pollutant concentration: Among them, N t represents the time step, Indicates the risk level, (x j ,y j ) represents the regional grid, Indicates the risk level of the regional grid, T time and T space Respectively represent the warning thresholds of the risk level in the time dimension and the risk level in the space dimension; S63. Design trigger conditions of time, space and pollutant concentration in a composite way and define joint warning rules: E alert =I(R overall >T overall and R time >T time and R space >T space ); Among them, E alert represents the joint warning value, I(·) represents the indicator function, and when the condition is met, E alert is 1, otherwise it is 0, R overall , R time and R space They represent the overall risk level, time dimension risk level and space dimension risk level respectively, T overall Indicates the warning threshold of the overall risk level, and means "and"; S64. When the joint warning value is 1, generate warning information, which includes warning type, involved area, warning time and pollutants.

Citation Information

Patent Citations

  • Atmospheric pollutant concentration prediction method based on multi-task learning

    CN113159099A

  • Air quality prediction method based on traffic jam index and multi-source data fusion

    CN117494034A