Traffic pollutant prediction method and system

By combining BLSTM and DSTMA methods, the problem of insufficient accuracy in multi-pollutant prediction is solved, efficient processing and multi-task learning of complex time series data is achieved, the accuracy and robustness of pollutant prediction are improved, and a new method for identifying and monitoring pollutant sources is provided.

CN120373557APending Publication Date: 2025-07-25HEFEI INSTITUTE OF PHYSICAL SCIENCE CHINESE ACADEMY OF SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510500594.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-21
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing pollutant prediction methods have the problem that prediction accuracy cannot be fully optimized when dealing with multiple pollutants, especially in the processing of complex time series data and multi-task environments, and it is difficult to effectively capture long-term dependencies and multi-dimensional factors.

Method used

The bidirectional long and short-term memory network (BLSTM) is used to combine dynamic sharing and task-specific multi-head attention mechanism (DSTMA), and the forward and backward dependencies of the time series are captured through bidirectional information flow, and the commonality and personalized characteristics between different tasks are captured through sharing and task-specific attention mechanisms, the weight ratio is dynamically adjusted, and the hyperparameters are tuned in combination with Bayesian optimization.

Benefits of technology

It significantly improves the accuracy and robustness of multi-pollutant prediction, can adapt to key features, adapt to complex environmental data, provide sensitivity analysis to pollutant concentration, and reveal the relative impact of different environmental factors and traffic flows.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120373557A_ABST
    Figure CN120373557A_ABST
Patent Text Reader

Abstract

The invention discloses a traffic pollutant prediction method and system, and the method comprises the steps: firstly enabling the time sequence data of traffic pollutants to enter a bidirectional long-short-term memory network, capturing the forward and backward dependencies in a time sequence through a bidirectional information flow, and then carrying out the detection of the forward and backward dependencies in the time sequence; the dynamic sharing and task exclusive multi-head attention mechanism is executed on data passing through the two-way long-short-term memory network, common characteristics among different tasks are captured through shared attention, personalized modeling of each task is achieved through task exclusive attention, and it is ensured that the model is balanced between sharing and exclusive learning. The method not only can capture long-time dependence in a time sequence, but also adaptively focuses on the most critical characteristics for predicting different pollutants through dynamic sharing and a task-exclusive multi-head attention mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technology of pollutant prediction, in particular to a traffic pollutant prediction method. Background Art

[0002] With the acceleration of the global urbanization process, urban road traffic has become one of the main sources of air pollution. Pollutants such as nitrogen oxides (NO, NO2), carbon monoxide (CO), carbon dioxide (CO2), and methane (CH4) emitted by motor vehicles pose a serious threat to the health of residents and have a profound impact on global climate change. Therefore, accurately predicting the concentration of road pollutants is crucial for environmental monitoring and pollution control. However, the complexity between traffic emissions and environmental factors makes it quite challenging to conduct time-series prediction of multiple pollutants.

[0003] Traditional pollutant prediction methods, such as statistical analysis and linear regression models, although they can reveal the relationship between pollutant concentration and traffic flow, vehicle type, and environmental conditions, due to the assumption of a linear relationship between variables, they ignore the complex dynamic relationship between environmental conditions and traffic emissions. For example, Chen J and Gilbert N L found that vehicle type and traffic flow are the main factors affecting air pollution levels and characteristics. With the improvement of computing power, models such as computational fluid dynamics (CFD) and land use regression (LUR) have been widely used in the simulation of pollutant diffusion. Although these methods show high accuracy in simulating the flow and diffusion of pollutants in urban blocks, they are computationally complex and time-consuming, making it difficult to meet the requirements of real-time prediction. At the same time, machine learning methods such as support vector machine (SVM) and random forest (RF) have been introduced into the field of pollutant prediction, which has improved the prediction effect to a certain extent. However, traditional machine learning methods still have limitations in capturing long-term dependence relationships and multi-dimensional factors when dealing with time-series data.

[0004] In recent years, due to its superiority in processing time series data, the long short-term memory network (LSTM) has been widely used in the time series prediction of air pollutants. For example, Wen et al. proposed a spatio-temporal convolutional long short-term memory neural network extension (C-LSTME) model for air quality concentration prediction. By combining the convolutional neural network (CNN) and the long short-term memory neural network (LSTM), high-level spatio-temporal features are extracted, and meteorological data and aerosol data are fused to improve the model prediction performance. The model uses historical PM2.5 concentration data, meteorological data, and aerosol data with high relevance but often overlooked, achieving accurate and stable predictions. Xu et al. proposed an LSTM autoencoder multi-task learning model to predict the PM2.5 time series at multiple locations in the city. This model can automatically mine the internal correlations between different pollution monitoring stations, overcoming the limitations of traditional data-driven PM2.5 prediction methods. The autoencoding of meteorological auxiliary information also significantly enhances the prediction performance. Shengdong Du et al. designed a joint hybrid deep learning framework based on one-dimensional CNN and Bi-LSTM for shared representation feature learning of multivariate air quality-related time series data. CNN extracts local trends and spatial correlation features, while Bi-LSTM learns spatio-temporal dependencies, effectively capturing the local trends and interdependencies of multivariate air quality data (such as temperature, humidity, wind speed, SO2, PM10, PM2.5). Experiments show that this model has better performance compared to traditional shallow and other deep learning models, especially in dealing with non-linear correlations. Chen et al. further proposed an attention-based hybrid convolutional neural network and bidirectional LSTM model. By introducing the attention mechanism, this model can adaptively focus on the features most critical for predicting different pollutants, thereby improving the prediction accuracy and stability of the model. Through the analysis of pollutant and meteorological data, accurate predictions of major pollutants such as PM2.5 and PM10 are achieved. Although significant progress has been made in existing research, most models still have the problem of inflexible adjustment of feature weights when dealing with multiple pollutants, resulting in insufficient optimization of the prediction accuracy of different pollutants. Summary of the Invention

[0005] The technical problem to be solved by the present invention is how to improve the prediction accuracy of pollutants when dealing with multiple pollutants simultaneously.

[0006] The present invention realizes the solution of the above technical problems through the following technical means: A traffic pollutant prediction method. The time series data of traffic pollutants first enters a bidirectional long short-term memory network, and the bidirectional information flow is used to capture the forward and backward dependencies in the time series. Then, the data after passing through the bidirectional long short-term memory network performs a dynamic sharing and task-specific multi-head attention mechanism. The shared attention captures the common features between different tasks, and at the same time, the task-specific attention realizes the personalized modeling of each task.

[0007] As a further optimized technical solution, during the process of performing the dynamic sharing and task-specific multi-head attention mechanism, it also includes adjusting the ratio of dynamic sharing and task-specific multi-head attention flexibly according to the actual needs of the task through a dynamic weight mechanism.

[0008] As a further optimized technical solution, the bidirectional long short-term memory network structure includes several branches. Each branch includes two symmetric network branches, a forward LSTM and a backward LSTM. The two symmetric network branches can simultaneously process the observed input vector at the current time step, and respectively obtain the hidden state output of the forward LSTM at each time step and the hidden state output of the backward LSTM at each time step. Then, the bidirectional hidden state outputs are concatenated to obtain the final output result.

[0009] As a further optimized technical solution, performing the dynamic sharing and task-specific multi-head attention mechanism includes sequentially performing the following operations:

[0010] S11. Shared attention

[0011] After receiving multi-task inputs, the original features of each task are uniformly input. Assume that the given input X ∈ R L×d , L is the sequence length, d is the dimension of the input, and the shared attention calculates the attention weights for each attention head. The shared attention output is:

[0012]

[0013] where Q shared , K shared , V shared are respectively the query matrix, key matrix, and value matrix of the shared attention layer;

[0014] S12. Task-specific attention

[0015] The attention mechanism of each task works independently. The output of the i-th task-specific attention is expressed as:

[0016]

[0017] where Q task,i , K task,i , Vtask,i respectively represent the query, key, and value matrices of task i;

[0018] S13. Dynamic Fusion

[0019] The dynamic weights are calculated by applying the softmax function to the output of the shared attention:

[0020] W dynamic = softmax(W s ·A shared )

[0021] where W dynamic ∈ R N , N is the number of tasks, which is a vector or matrix representing the fusion weights of all tasks, and Ws is a learnable weight matrix;

[0022] The final output of task i is a weighted combination of the shared attention output and the task-specific attention output:

[0023]

[0024] where W dynamic,i represents the weight term corresponding to the i-th task in the dynamic weight vector.

[0025] As a further optimized technical solution, in the bidirectional long short-term memory network and the dynamic shared and task-specific multi-head attention mechanism, Bayesian optimization is used to tune the hyperparameters.

[0026] As a further optimized technical solution, the hyperparameters include the number of LSTM units in the bidirectional long short-term memory network, the number of attention heads, the dropout rate, and the learning rate.

[0027] As a further optimized technical solution, during the Bayesian optimization process, the Gaussian process is used to model the evaluation function of the hyperparameters. Assuming there are n known hyperparameter combinations θ1, θ2, …, θn and their corresponding validation losses f(θ1), f(θ2), …, f(θn), the Gaussian process model is updated by the following formula in Bayesian optimization:

[0028] p(f(θ)∣D) = N(μ(θ), σ 2 (θ))

[0029] where D is the dataset of known hyperparameters and validation losses,

[0030] Each time Bayesian optimization selects the next hyperparameter combination θ next by maximizing the acquisition function:

[0031] θ next = argmax θEI(θ)

[0032] As the new hyperparameter combinations are continuously evaluated, the model gradually converges to the optimal hyperparameters.

[0033] The present invention also provides a traffic pollutant prediction system corresponding to the above method, including:

[0034] A bidirectional long short-term memory network model. The time series data of traffic pollutants first enters the bidirectional long short-term memory network model, and the bidirectional information flow is used to capture the forward and backward dependencies in the time series.

[0035] A dynamic shared and task-specific multi-head attention mechanism model. The data after passing through the bidirectional long short-term memory network enters the dynamic shared and task-specific multi-head attention mechanism model. The shared attention is used to capture the common features between different tasks, and at the same time, the task-specific attention is used to perform personalized modeling for each task, ensuring that the model achieves a balance between shared and specific learning.

[0036] As a further optimized technical solution, the bidirectional long short-term memory network structure includes several branches. Each branch includes two symmetric network branches, a forward LSTM and a backward LSTM. The two symmetric network branches can simultaneously process the observed input vector at the current time step, respectively obtaining the hidden state output of the forward LSTM at each time step and the hidden state output of the backward LSTM at each time step. Then, the bidirectional hidden state outputs are concatenated to obtain the final output result.

[0037] As a further optimized technical solution, the dynamic shared and task-specific multi-head attention mechanism model includes:

[0038] A shared attention unit, which is used to uniformly input the original features of each task when receiving multi-task inputs. Assuming the given input X ∈ R T×d , T is the sequence length, d is the dimension of the input, and the shared attention calculates the attention weights for each attention head. The shared attention output is:

[0039]

[0040] where Q shared , K shared , V shared are the query matrix, key matrix, and value matrix of the shared attention layer, respectively;

[0041] A task-specific attention unit, which is used for the attention mechanism of each task to work independently. The output of the i-th task-specific attention is expressed as:

[0042]

[0043] where Q taski, K taski , V taski respectively represent the query, key, and value matrices of task i;

[0044] The dynamic fusion unit is used for dynamic fusion, and the dynamic weights are obtained by calculating the normalized exponential function of the shared attention output:

[0045] W dynamic = softmax(W s ·A shared )

[0046] where W dynamic ∈R T , T is the number of tasks, which is a vector or matrix representing the fusion weights of all tasks, and Ws is a learnable weight matrix;

[0047] The final output of task i is a weighted combination of the shared attention output and the task-specific attention output:

[0048]

[0049] where, W dynamic,i represents the weight term corresponding to the i-th task in the dynamic weight vector.

[0050] Other execution steps in this traffic pollutant prediction system are the same as those in the above traffic pollutant prediction method.

[0051] The advantages of the present invention are as follows: The present invention proposes a bidirectional long short-term memory network (BLSTM) method (DSTMA-BLSTM) combining dynamic sharing and task-specific multi-head attention (DSTMA) for temporal prediction and sensitivity analysis of road pollutants. This method can not only capture long-term dependencies in time series but also adaptively focus on the features most critical for predicting different pollutants through the dynamic sharing and task-specific multi-head attention mechanisms. Through the verification of actual monitoring data, the performance of the DSTMA-BLSTM method in predicting multiple pollutants such as NO, NO2, CO2, CO, and CH4 is evaluated. The research results show that the dynamic sharing and task-specific multi-head attention mechanisms significantly improve the prediction accuracy of the model, especially showing high flexibility and robustness in processing complex environmental data. In addition, the sensitivity analysis reveals the relative impacts of different environmental factors and traffic flows on pollutant concentrations, providing new insights for pollutant source identification and monitoring. This study not only demonstrates the effectiveness of the DSTMA-BLSTM model in multi-pollutant prediction but also provides new ideas for the joint prediction of future traffic and non-traffic source pollutants. Brief Description of the Drawings

[0052] Figure 1It is the overall architecture diagram of the traffic pollutant prediction method based on DSTMA-BLSTM in the embodiments of the present invention;

[0053] Figure 2 It is the diagram of on-site monitoring of meteorological data by monitoring equipment;

[0054] Figure 3 It is the model loss curve;

[0055] Figure 4 It is the characteristic sensitivity of different meteorological factors during the prediction of each pollutant. Among them: (a) The characteristic sensitivity of different meteorological factors during NO prediction; (b) The characteristic sensitivity of different meteorological factors during NO2 prediction; (c) The characteristic sensitivity of different meteorological factors during CO prediction; (d) The characteristic sensitivity of different meteorological factors during CO2 prediction; (e) The characteristic sensitivity of different meteorological factors during CH4 prediction.

[0056] Figure 5 It is the diagram of the sensitivity of different traffic flow characteristics during greenhouse gas prediction;

[0057] Figure 6 It is the diagram of the sensitivity of different traffic flow characteristics during NOx prediction;

[0058] Figure 7 It is the diagram of model average characteristic sensitivity analysis;

[0059] Figure 8 It is the prediction effect diagram of NO in the time dimension, where (a) is the comparison between the measured value and the prediction; (b) is the residual curve; (c) is the residual histogram;

[0060] Figure 9 It is the prediction effect of NO2 in the time dimension, where (a) is the comparison between the measured value and the prediction; (b) is the residual curve; (c) is the residual histogram;

[0061] Figure 10 It is the prediction effect of CO in the time dimension, where (a) is the comparison between the measured value and the prediction; (b) is the residual curve; (c) is the residual histogram;

[0062] Figure 11 It is the prediction effect of CO2 in the time dimension, where (a) is the comparison between the measured value and the prediction; (b) is the residual curve; (c) is the residual histogram;

[0063] Figure 12 It is the prediction effect of CH4 in the time dimension, where (a) is the comparison between the measured value and the prediction; (b) is the residual curve; (c) is the residual histogram. Specific implementation manner

[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0065] When ordinal numbers such as "first", "second", "third", etc. are used in the description and claims, they are used to modify the corresponding elements. They do not themselves imply any ordinal number for the element, nor do they represent the order of one element relative to another or the order in the manufacturing method. The use of these ordinal numbers is only to clearly distinguish one element with a certain name from another element with the same name.

[0066] The traffic pollutant prediction method based on DSTMA-BLSTM of the present invention combines a bidirectional long short-term memory network (BLSTM) with a dynamic shared and task-specific multi-head attention mechanism (DSTMA), aiming to process complex time series data and is particularly suitable for the pollutant concentration prediction scenario in multi-task prediction, including the following processes:

[0067] S1. Enter the bidirectional long short-term memory network to capture the forward and backward dependencies in the time series through bidirectional information flow.

[0068] In time series data, there are obvious long-term and short-term dependence problems, that is, the cumulative effect of past concentrations and future meteorological conditions will jointly affect the current pollutant concentration level. Therefore, how to effectively capture these dependencies becomes the key. To address this challenge, the model introduces BLSTM to capture the forward (historical information) and backward (future information) dependencies in the time series through bidirectional information flow. Traditional LSTM can only handle unidirectional dependencies, while BLSTM can simultaneously model the influence of the past and the future on the current state, significantly improving the accuracy of time series modeling.

[0069] S2. Execute the dynamic shared and task-specific multi-head attention mechanism

[0070] In addition, multi-task learning is of great significance in pollutant concentration prediction, and there are often complex associations between different pollutants. To effectively model these associations in a multi-task environment, the method designs a dynamic shared and task-specific multi-head attention mechanism (DSTMA). DSTMA captures the common features between different tasks through shared attention, while achieving personalized modeling for each task through task-specific attention, ensuring that the model strikes a balance between shared and specific learning. Through the dynamic weight mechanism, DSTMA can also flexibly adjust the proportion of dynamic shared and task-specific multi-head attention according to the actual needs of the tasks, further improving the performance of multi-task learning.

[0071] S3. In the bidirectional long short-term memory network and the dynamic shared and task-specific multi-head attention mechanism, Bayesian optimization is used to tune the hyperparameters.

[0072] To further improve the model performance, the method uses Bayesian optimization to tune the hyperparameters. Due to the complexity of the model, it is difficult to achieve the optimal effect by manual parameter tuning. Therefore, Bayesian optimization automatically searches the hyperparameter space by constructing a surrogate model and an acquisition function, thereby finding the best configuration and ensuring the robustness of the model in different task scenarios.

[0073] This method combines a bidirectional long short-term memory network (BLSTM) with a dynamic shared and task-specific multi-head attention mechanism (DSTMA) to process complex time series data. The main innovation of this method is the introduction of a dynamic shared and task-specific attention mechanism, which improves the model's ability to capture complex patterns in time series through shared learning and task customization, as Figure 1 shown.

[0074] Among them, in S1, entering the bidirectional long short-term memory network, the bidirectional long short-term memory network (BLSTM) structure is introduced to capture the forward and backward dependencies in the time series through bidirectional information flow. BLSTM (bidirectional long short-term memory network) is an extension of the traditional LSTM. In pollutant concentration prediction, LSTM can capture the cumulative effect of historical pollutant data. For example, the influence of past pollutant concentrations, traffic flows, meteorological conditions, etc. on the current pollutant state. However, LSTM is essentially unidirectional, that is, it can only predict the future from past time series information. BLSTM not only obtains information from past time steps but also combines information from future time steps for prediction.

[0075] BLSTM runs two independent LSTM networks in parallel: one propagates forward (from the past), and the other propagates backward (from the future). The bidirectional structure of BLSTM can capture the complex dependencies in time series data, and the formula is as follows:

[0076]

[0077] Among them, is the hidden state of the forward LSTM, is the hidden state of the backward LSTM, and h t is the concatenation result of the outputs of the bidirectional LSTM.

[0078] Figure 3 The following shows the schematic diagram of the BLSTM model structure adopted in the embodiment of the present invention. This structure includes several branches, and each branch includes a forward LSTM (LSTM + ) and a backward LSTM (LSTM-) which are two symmetric network branches. The two symmetric network branches can simultaneously process the observed input vector at the current time step, respectively obtaining the hidden state output of the forward LSTM at each time step and the hidden state output of the backward LSTM at each time step. Then, the bidirectional hidden state outputs are concatenated to obtain the final output result. Its core parameters and meanings are as follows:

[0079] w1, w2, w3: represent the observed input vector at the current time step;

[0080] fh1, fh2, fh3: respectively represent the hidden state outputs of the forward LSTM at each time step;

[0081] bh1, bh2, bh3: respectively represent the hidden state outputs of the backward LSTM at each time step;

[0082] h t-1 , h t , h t+1 represent the concatenated output results of the BLSTM at times t - 1, t, and t + 1, that is, [fh t ; bh t ;

[0083] LSTM + represents the forward LSTM cell;

[0084] LSTM - represents the backward LSTM cell.

[0085] Among them, S2, the execution of the dynamic shared and task-specific multi-head attention mechanism is designed to solve the multi-task correlation problem. In the roadside environment, there are often certain correlations among different pollutants (such as NO, NO2, CO, CO2, CH4). They may come from the same pollution source, such as vehicle exhaust emissions, or may affect each other. For example, NO2 and NO generated by combustion often appear simultaneously, and the concentrations of CO and CO2 are usually related to the combustion efficiency and emission level of vehicles. Therefore, it is very important to consider the relationships among these pollutants when modeling to improve the prediction accuracy. It can dynamically assign weights to each task, which can not only capture the shared information among different tasks (such as the common source of pollutants) but also flexibly adjust the task-specific attention (processing the personalized characteristics of pollutants). The design of the dynamic shared and task-specific multi-head attention mechanism enables the use of the information of NO when predicting NO2 and also fully utilizes the characteristics of CO2 when predicting CO, thereby capturing the complex pollutant interrelationships and common trends. The DSTMA mechanism is the core innovation of this method. It shares information among different tasks through the multi-task attention mechanism and at the same time allows for customized learning for each task. This mechanism sequentially performs the following operations:

[0086] S11, Shared Attention

[0087] The shared attention layer is used to capture the patterns shared among all tasks. After the system receives the multi-task input, it uniformly inputs the original features of each task. Assuming the given input X ∈ R L×d (where L is the sequence length and d is the dimension of the input), that is, the input X is the output sequence of the BLSTM, which is the output result of the above BLSTM...h t-1 , h t , h t+1 ...

[0088] The shared attention calculates the attention weights for each attention head. The output of the shared attention is:

[0089]

[0090] where Q shared , K shared , V shared are the query matrix, key matrix, and value matrix of the shared attention layer respectively, and T represents the transpose operation of the matrix.

[0091] This stage conducts joint modeling based on the overall information of all tasks, extracts the common semantic representations with global dependencies, and the output result of the shared attention serves as the common basis for all tasks and is then passed in parallel to the next stage.

[0092] S12, Task-Specific Attention

[0093] Each task has its own dedicated attention mechanism for capturing patterns unique to that task. The attention mechanisms for each task work independently, and the output of the dedicated attention for the i-th task is denoted as:

[0094]

[0095] where Q task,i , K task,i , V task,i represent the query, key, and value matrices for task i respectively. These task-specific attention heads enable the method to focus on task-specific parts of the input.

[0096] In this stage, an independent multi-head attention mechanism is configured for each task to extract personalized features that the task is concerned with. While each task receives the shared representation, it combines its own feature preferences to generate a task-specific representation result.

[0097] S13. Dynamic Fusion

[0098] A dynamic weight mechanism is introduced, which determines the weight of the dedicated attention output for each task in the final representation. The dynamic weight is calculated by applying the softmax function to the shared attention output:

[0099] W dynamic = softmax(W s ·A shared )

[0100] where W dynamic ∈ R N , N is the number of tasks, and it is a vector or matrix representing the fusion weights for all tasks; Ws is a learnable weight matrix.

[0101] The final output for task i is a weighted combination of the shared attention output and the dedicated attention output for the task:

[0102] O i = W dynamic,i ·A shared + (1 - W dynamic,i )·A task,i

[0103] where W dynamic,i represents the weight term corresponding to the i-th task in the dynamic weight vector, reflecting the degree of dependence of the task on the shared attention output in the final representation.

[0104] In this stage, based on the output results of the previous shared attention, the dependence degree of each task on the shared information in the final representation is calculated, and the combination ratio between the shared attention output and the task-specific attention output is determined accordingly. This process assigns differential attention fusion weights to different tasks, ensuring that the task outputs not only retain individual characteristics but also integrate global information.

[0105] Finally, the fused representations generated by each task are used as outputs and enter the downstream multi-task learning and prediction. This mechanism realizes the organic unity of shared modeling, individual expression, and fusion regulation among tasks, improving the representation ability and adaptability of multi-task modeling.

[0106] This mechanism enables the method to dynamically balance between the shared attention output and the task-specific attention output. For some tasks, the shared pattern may be more important, while for other tasks, more focus can be placed on task-specific patterns.

[0107] S3. In the bidirectional long short-term memory network and the dynamic shared and task-specific multi-head attention mechanism, Bayesian optimization is used to tune the hyperparameters, specifically including:

[0108] In the bidirectional long short-term memory network and the dynamic shared and task-specific multi-head attention mechanism, the selection of hyperparameters (such as the number of LSTM units, the number of attention heads, the dropout rate, the learning rate, etc.) has a significant impact on the performance of the model. Although direct brute-force search (such as grid search or random search) is simple, the computational cost is extremely high, especially when the training of the model itself takes a long time. Therefore, we introduce Bayesian optimization, which can intelligently and efficiently search the hyperparameter space and find a near-optimal combination of hyperparameters with fewer evaluation times.

[0109] The advantage of Bayesian optimization is that it can gradually construct a surrogate model using the known evaluation information and select the next evaluation point based on the trade-off between the current optimal result and uncertainty, thus more effectively finding the global optimum.

[0110] In the process of Bayesian optimization, the core is to model the evaluation function of hyperparameters through a Gaussian process. Assuming that there are n known combinations of hyperparameters θ1, θ2, …, θn and their corresponding validation losses f(θ1), f(θ2), …, f(θn), Bayesian optimization updates the Gaussian process model through the following formula:

[0111] p(f(θ)∣D) = N(μ(θ), σ 2 (θ))

[0112] where D is the dataset of known hyperparameters and validation losses.

[0113] Each time Bayesian optimization selects the next combination of hyperparameters θnext Selected by maximizing the acquisition function:

[0114] θ next = argmax θ EI(θ)

[0115] As new hyperparameter combinations are continuously evaluated, the model gradually converges to the optimal hyperparameters.

[0116] Bayesian optimization provides an efficient way to tune the hyperparameters of the BLSTM+DSTMA model, which can find near-optimal hyperparameter combinations with fewer evaluation times. Its advantage lies in constructing an approximate model of the objective function through Gaussian processes and using the acquisition function to intelligently balance exploration and exploitation, avoiding unnecessary computational overhead.

[0117] Through Bayesian optimization, the complex hyperparameter space of the BLSTM and DSTMA models can be effectively explored. Especially in the optimization process of multi-task learning and dynamic weight mechanisms, Bayesian optimization can intelligently adjust the exclusive and shared attention mechanisms for each task, thereby improving the accuracy and generalization ability of pollutant prediction.

[0118] The present invention deeply integrates the Bayesian optimization process. Specifically, the present invention applies Bayesian optimization to the adjustment of key hyperparameters of the BLSTM+DSTMA network structure. The adjustment dimensions include the number of attention heads, the number of LSTM units, the Dropout rate, and the learning rate, and realizes adaptive search through Gaussian process modeling and acquisition function strategies, constructing a structure adjustment framework that can perceive task differences. This optimization process is coupled with the dynamic weight mechanism, which can dynamically adjust the attention structure and modeling path according to the training feedback of each task, significantly improving the adaptive modeling ability and prediction performance of the model in complex multi-task scenarios, and possessing a higher level of innovation and systematicness.

[0119] The following content is the process of experimenting and validating the DSTMA-BLSTM model using the above method.

[0120] The instrument parameters and monitored substances are shown in Table 1. Meteorological data is obtained in real time through high-precision sensors installed at the monitoring points. As Figure 2 shown, these sensors can accurately monitor multiple key meteorological parameters such as temperature, humidity, wind speed, and wind direction, providing basic data for further analysis and prediction. At the same time, traffic flow data is collected in real time by the traffic management department through devices such as cameras, induction coils, and intelligent transportation systems on the roads. After preliminary processing, these data cover not only information such as the number and type of vehicles passing through.

[0121] Table 1 Instrument parameters and monitored substances

[0122]

[0123] The data collection time was one week, and the monitoring location was 33°33’57”N, 117°00’16”E. To ensure the consistency and comparability of the data, the data from different sources were standardized. Specifically, the traffic flow data were statistically analyzed, and the average value within 2 minutes was used to represent the vehicle passing situation. The meteorological information also reflected the changes in meteorological parameters such as temperature, humidity, and wind speed through the average value within 2 minutes. At the same time, the sensors arranged along the road also calculated the average value of the pollutant concentration within 2 minutes to provide accurate environmental pollution data. By standardizing the data through this unified time window, the traffic flow, meteorological information, and pollutant concentration data can be aligned in the time dimension, ensuring that they can be smoothly input into the model for further analysis and prediction, thereby improving the accuracy and reliability of the model.

[0124] For the DSTMA-BLSTM model constructed based on the actual monitoring data, the hyperparameters were systematically adjusted and optimized through the Bayesian optimization method. During the hyperparameter adjustment process of the model, the search ranges of multiple hyperparameters were set, including the number of LSTM units, dropout rate (Dropout), the number of heads in the multi-head attention mechanism, and the learning rate, etc. Specifically, the search range of the number of LSTM units was set from 32 to 1024, and the value of the number of LSTM units would be adjusted according to the previous results in each round of experiments. Finally, the number of LSTM units in the first layer and the second layer was determined to be 864 and 960 respectively. These two values can effectively learn complex time series data while avoiding overfitting of the model. During the Bayesian search process, the values of Dropout1 and Dropout2 were searched between 0.0 and 0.5 respectively, and the best values obtained were 0.3 and 0.1. Such a combination can effectively prevent overfitting during the model training process while retaining sufficient model capacity to learn complex features.

[0125] In the optimization of the multi-head attention mechanism, the Bayesian search determined the use of 5 attention heads (NumHeads = 5), and this choice makes the model more flexible in capturing complex time series relationships in the input sequence. The selection of the learning rate (LearningRate) was also automatically adjusted through the Bayesian search, and its search range was set from 1e-4 to 1e-2, and the optimal value finally obtained was 0.0012511, which achieved a good balance between the optimization speed and the model convergence. For the DSTMA-BLSTM model constructed based on the actual monitoring data, the best parameters obtained by the Bayesian search are shown in Table 2.

[0126] Table 2 Optimal Parameters of DSTMA-BLSTM

[0127]

[0128] These optimal parameters obtained through Bayesian optimization fully consider the balance between model complexity and training efficiency, enabling the DSTMA-BLSTM model to have strong generalization ability and good convergence when dealing with high-dimensional and complex time-series data.

[0129] As a further optimized technical solution, the traffic pollutant prediction method also includes an evaluation step. To ensure its robustness and performance, multiple evaluation metrics are adopted, including loss curves, residual analysis, and feature importance analysis. These metrics are not only used to monitor the training process in the method but also to evaluate the prediction effect of the method and its sensitivity to different features.

[0130] 1. Loss Curve

[0131] MSE, as the main loss function, optimizes the training process of the model. The loss curve helps monitor overfitting or underfitting problems during the training process. MSE is the main loss function used in the code, and it is used to measure the difference between the predicted value of the model and the true value. The mathematical definition of MSE is:

[0132]

[0133] In the formula, y i is the true value of the i-th sample, is the predicted value of the model, n is the number of samples, and the loss curve shows how the loss value of the model changes with the increase in the number of iterations during the training process. This can help identify whether the model is overfitting or underfitting.

[0134] 2. Feature Importance

[0135] Feature importance analysis reveals the importance of each input feature to the prediction result, which helps understand the decision-making process of the model. Feature importance evaluates the contribution of each feature to the model prediction by perturbing the feature values and observing the changes in model performance. Generally, permutation feature importance is used for calculation.

[0136] For each feature X j , calculate the change in model loss △L j after shuffling this feature:

[0137]

[0138] In the formula, is the predicted value after shuffling the j-th feature, is the original model loss.

[0139] 3. Residual Analysis

[0140] Residual analysis further validates the reliability of the model prediction, ensuring that there is no systematic bias in the residuals. The residual r i is the difference between the model prediction value and the true value. Residual analysis is used to check whether there is a systematic bias in the prediction error of the model.

[0141]

[0142] By plotting the distribution map of r i , it can be judged whether the residuals are evenly distributed near zero. If the residuals are evenly distributed and close to zero, it indicates that the model prediction error is random and the model has no systematic bias.

[0143] Figure 3 shows the loss curve of the model during the training process, including the change trends of the training loss and the validation loss. By analyzing this curve, the following conclusions can be drawn:

[0144] a. Fast convergence: As can be observed from Figure 3 , in the initial stage of model training (within the first 20 epochs), the training loss and the validation loss decrease rapidly. This indicates that the model can quickly learn the basic features in the data in the initial stage and effectively reduce the prediction error.

[0145] b. Good generalization ability: As the training progresses, there is no obvious separation between the training loss and the validation loss, indicating that the model maintains a low error on both the training set and the validation set and there is no overfitting phenomenon.

[0146] c. Robustness and convergence of the model: As can be seen from Figure 3 , in the later stage of training (after about 100 epochs), the loss curve tends to be stable, and both the training loss and the validation loss are close to 0, indicating that the model has reached an ideal convergence state.

[0147] The loss curve of the model indicates that the model has fast convergence, good generalization ability and a stable convergence state. The combination of BLSTM and DSTMA enables the model to efficiently learn the multi-task features in complex time series, and at the same time, through reasonable optimization strategies and regularization means, the problems of overfitting and underfitting are avoided.

[0148] Figures 4 to 7 shows the sensitivity feature analysis, including:

[0149] (1) Sensitivity analysis of environmental features

[0150] Figure 4(a) to (e) show the results of sensitivity analysis of different pollutants (NO, NO2, CO, CH4, CO2) to multiple meteorological characteristics (including wind speed, wind direction, air pressure, temperature, humidity). By using feature importance assessment methods (such as permutation feature importance), the contribution degree of each meteorological characteristic to pollutant concentration prediction can be measured.

[0151] Through the sensitivity analysis of different pollutants and meteorological characteristics, the following conclusions can be drawn:

[0152] Wind Speed is a key factor in the concentration change of almost all pollutants (NO, NO2, CO, CH4, CO2), indicating that the diffusion ability of pollutants in the atmosphere is mainly affected by wind speed;

[0153] Wind Directions are also very important for local pollutants such as NO and CH4, especially the distribution of these pollutants will change significantly under different wind direction conditions;

[0154] For some pollutants, such as NO2 and CO2, air pressure has a greater impact on concentration, reflecting the role of atmospheric pressure in the diffusion and accumulation of these gases;

[0155] Humidity has a certain impact on the concentration of some pollutants (such as CO, NO), but it is relatively less important compared to other meteorological conditions;

[0156] Temperature has a generally small impact on each pollutant, indicating its weak role in the change of pollutant concentration;

[0157] Meteorological conditions have a significant regulatory effect on the change of pollutant concentration. Especially under the combined influence of wind speed, wind direction and air pressure, the concentration of pollutants will show significant fluctuations.

[0158] (2) Sensitivity analysis of traffic flow characteristics

[0159] (2.1) Traffic flow characteristics and greenhouse gases

[0160] As can be seen from Figure 5 Greenhouse gases (such as CH4, CO and CO2) are important components of road traffic emissions, especially the emissions from diesel vehicles and gasoline vehicles, which have a significant impact on the concentration of greenhouse gases in the atmosphere. Through sensitivity analysis, the contribution of different types of vehicles to various greenhouse gases can be revealed:

[0161] The number of gasoline vehicles dominates the emissions of CO, CH4 and CO2. Especially in terms of CO and CO2, it shows that gasoline vehicles play a very important role in road pollution;

[0162] The number of diesel vehicles significantly contributes to CO2 emissions, has a certain contribution to CO, but has little impact on CH4. This reflects that diesel vehicles mainly emit carbon dioxide and a small amount of carbon monoxide, rather than methane;

[0163] The total traffic flow has a weak impact on these three gases, especially on CH4 and CO2, indicating that relying solely on traffic flow cannot effectively predict the emissions of these gases, and the specific vehicle types are the key factors.

[0164] (2.2) Traffic flow characteristics and NOx gases

[0165] From Figure 6 it can be seen that diesel vehicles dominate the emissions of NO and NO2, while gasoline vehicles also make a certain contribution but relatively small. The total traffic flow has a limited impact on these two nitrogen oxides, indicating that the emissions of nitrogen oxides more depend on the specific quantity and proportion of diesel vehicles and gasoline vehicles. Therefore, when controlling road NOx pollution, special attention should be paid to the emission management of diesel vehicles.

[0166] (3) Overall sensitivity analysis of the model

[0167] Overall impact analysis Figure 7 it can be seen the average importance of various factors on each characteristic of the model prediction results:

[0168] Wind speed is the most significant meteorological factor affecting all gases, indicating that meteorological conditions are crucial for the diffusion and dilution of pollutants;

[0169] The number of gasoline vehicles and the number of diesel vehicles are the main sources affecting roadside pollutants, and the emission source characteristics determine the direct impact of the emissions of these vehicles on pollutant concentrations;

[0170] The impacts of other meteorological factors such as wind direction, humidity, and temperature are relatively small, but still have a secondary effect on certain gases.

[0171] The prediction performance of the residual analysis in the time dimension is described as follows.

[0172] (1) NO prediction

[0173] From Figure 8 it can be seen that the model's overall prediction of NO concentration performs well, especially during the periods when the NO concentration is stable, and the model's prediction is almost the same as the true value. The distribution of the residual histogram is relatively symmetric, with the peak around -10, indicating that the model has a slight systematic underestimation at certain moments. Although a small number of residuals exceed the range of ±20 ppb, generally the distribution of residuals is close to a normal distribution, and most residuals are between -20 and 20.

[0174] (2) NO2 prediction

[0175] From Figure 9 it can be seen that the model can well capture the overall fluctuation trend of NO2. Especially when the NO2 concentration is relatively stable, the model predictions are highly consistent with the actual values. However, when the NO2 concentration changes sharply, the prediction performance decreases, especially at the peak positions where the model often underestimates. The residuals are mainly concentrated within the range of ±10 ppb, indicating that the overall model predictions are relatively accurate, but there are sometimes large errors, especially at the concentration peaks and troughs. The vast majority of the residuals are distributed between -10 and 5, being relatively concentrated, indicating that the model performs well in most NO2 predictions, but there is a certain negative bias, that is, the model tends to underestimate the NO2 concentration.

[0176] (3) CO prediction

[0177] From Figure 10 it can be seen that the model can generally well capture the concentration changes of CO. Especially when the CO concentration is relatively stable, the difference between the predicted value and the actual value is small. However, when the CO concentration fluctuates greatly, the prediction performance of the model decreases, especially at the peaks or troughs where the model's response is slightly lagged. The residuals fluctuate between ±0.1 ppm, and most of the residuals are concentrated in a relatively small range, indicating that the overall prediction error of the model is small. But the residuals are large when the concentration changes violently, showing that there is a certain instability in the model predictions in the high volatility regions. The residuals are generally concentrated around 0, but a few residuals exceed 0.1 ppm, indicating that in some extreme cases, the model has large errors.

[0178] (4) CO2 prediction

[0179] From Figure 11 it can be seen that, similar to CO, when the CO2 concentration changes violently, especially at the extreme peaks or troughs, there are certain differences between the predicted values of the model and the actual values. Most of the residuals are concentrated within ±10 ppm, indicating that the model can provide relatively accurate predictions most of the time. A small number of residuals are distributed outside ±10 ppm, indicating that there are errors in the model in extreme cases, but the occurrence frequency of these errors is low.

[0180] (5) CH4 prediction

[0181] From Figure 12It can be seen that the model can predict the overall fluctuation trend of CH4 well, but at some peak or trough moments, the prediction accuracy decreases. In particular, there is an obvious underestimation in the prediction of high-concentration CH4 concentration. Most of the residuals are concentrated within the range of ±0.2 ppm, but some residuals are around ±0.4 ppm, indicating that the model still has relatively large errors at some moments, especially when the CH4 concentration changes rapidly, and the prediction performance of the model is relatively poor. The distribution of the histogram is relatively symmetric, but there are a small number of larger residuals distributed at the tails, especially in the region close to ±0.4 ppm, which indicates that the model has relatively large errors in some extreme cases.

[0182] (6) Prediction evaluation in the time dimension

[0183] Table 3 shows the prediction evaluation indicators of the model for different pollutants, including root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R 2 ). Generally speaking, the model has an ideal prediction effect on most pollutants, especially outstanding in the gases closely related to traffic activities. The prediction performance for different pollutants can be further subdivided as follows:

[0184] The model has a very good prediction effect on roadside NO, NO2, and CO2, and has a strong correlation with traffic;

[0185] The prediction performance of roadside CO is also relatively satisfactory, and it has a relatively strong correlation with traffic;

[0186] The inaccurate prediction of CH4 is closely related to the low contribution of methane emissions from road traffic. The traffic-related input features may not be sufficient to effectively capture the changes in methane concentration because road traffic is not the main source of methane. Therefore, it is difficult for the model to provide accurate predictions without data on high-methane emission sources such as agriculture and industry.

[0187] Table 3 Model prediction evaluation indicators

[0188]

[0189] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A traffic pollutant prediction method, characterized in that: The time series data of traffic pollutants first enter the bidirectional long short-term memory network, which captures the forward and backward dependencies in the time series through bidirectional information flow. Then, the data after passing through the bidirectional long short-term memory network is subjected to a dynamic sharing and task-specific multi-head attention mechanism, which captures the common features between different tasks through shared attention and realizes personalized modeling for each task through task-specific attention.

2. The traffic pollutant prediction method according to claim 1, characterized in that: During the process of implementing the dynamic sharing and task-specific multi-head attention mechanism, it also includes flexibly adjusting the ratio of dynamic sharing and task-specific multi-head attention according to the actual needs of the task through a dynamic weight mechanism.

3. The traffic pollutant prediction method according to claim 1, wherein: The bidirectional long short-term memory network structure includes several branches. Each branch includes two symmetric network branches, a forward LSTM and a backward LSTM. The two symmetric network branches can simultaneously process the observed input vector at the current time step, respectively obtaining the hidden state output of the forward LSTM at each time step and the hidden state output of the backward LSTM at each time step. Then, the bidirectional hidden state outputs are concatenated to obtain the final output result.

4. The traffic pollutant prediction method according to claim 1, characterized in that: Implementing the dynamic sharing and task-specific multi-head attention mechanism includes sequentially performing the following operations: S11. Shared attention After receiving multi-task inputs, uniformly input the original features of each task. Assume that the given input is \(X\in\mathbb{R}\) L×d , \(L\) is the sequence length, \(d\) is the dimension of the input. The shared attention calculates the attention weights for each attention head, and the output of the shared attention is: where Q shared , K shared , V shared are the query matrix, key matrix, and value matrix of the shared attention layer, respectively; S12. Task-specific attention The attention mechanism for each task works independently. The output of the i-th task-specific attention is expressed as: where Q task,i , K task,i , V task,i represent the query, key, and value matrices of task i, respectively; S13. Dynamic fusion The dynamic weight is obtained by calculating the normalized exponential function of the shared attention output: W dynamic = softmax(W s ·A shared ) where W dynamic ∈R N , N is the number of tasks, which is a vector or matrix representing the fusion weights of all tasks, and Ws is a learnable weight matrix; The final output of task i is a weighted combination of the shared attention output and the task-specific attention output: Among them, W dynamic,i represents the weight term corresponding to the i-th task in the dynamic weight vector.

5. The traffic pollutant prediction method according to claim 1, characterized in that: In the bidirectional long short-term memory network and the dynamic sharing and task-specific multi-head attention mechanism, Bayesian optimization is used to tune the hyperparameters.

6. The traffic pollutant prediction method according to claim 5, characterized in that: The hyperparameters include the number of LSTM units in the bidirectional long short-term memory network, the number of attention heads, the dropout rate, and the learning rate.

7. The traffic pollutant prediction method according to claim 5, characterized in that: During the Bayesian optimization process, the Gaussian process is used to model the evaluation function of the hyperparameters. Assuming that there are n known hyperparameter combinations θ1, θ2, …, θn and their corresponding validation losses f(θ1), f(θ2), …, f(θn), Bayesian optimization updates the Gaussian process model through the following formula: p(f(θ)∣D) = N(μ(θ), σ 2 (θ)) where D is the dataset of known hyperparameters and validation losses, Each time Bayesian optimization selects the next hyperparameter combination θ next by maximizing an acquisition function: θ next = argmax θ EI(θ) As new hyperparameter combinations are continuously evaluated, the model gradually converges to the optimal hyperparameters.

8. A traffic pollutant prediction system, characterized in that: Including: The bidirectional long short-term memory network model. The time series data of traffic pollutants first enter the bidirectional long short-term memory network model, which captures the forward and backward dependencies in the time series through bidirectional information flow. The dynamic sharing and task-specific multi-head attention mechanism model. The data after passing through the bidirectional long short-term memory network enters the dynamic sharing and task-specific multi-head attention mechanism model, which captures the common features between different tasks through shared attention and realizes personalized modeling for each task through task-specific attention, ensuring that the model achieves a balance between shared and specific learning.

9. The traffic pollutant prediction system according to claim 8, characterized in that: The bidirectional long short-term memory network structure includes several branches. Each branch includes two symmetric network branches, namely, the forward LSTM and the backward LSTM. The two symmetric network branches can simultaneously process the observed input vector at the current time step, respectively obtaining the hidden state output of the forward LSTM at each time step and the hidden state output of the backward LSTM at each time step. Then, the bidirectional hidden state outputs are concatenated to obtain the final output result.

10. A traffic pollutant prediction system according to claim 8, characterized in that: The dynamic shared and task-specific multi-head attention mechanism model includes: A shared attention unit, which, after receiving multi-task inputs, uniformly inputs the original features of each task. Assume a given input X ∈ R L×d , where L is the sequence length and d is the dimension of the input. The shared attention calculates attention weights for each attention head, and the shared attention output is: where Q shared , K shared , V shared are the query matrix, key matrix, and value matrix of the shared attention layer, respectively; Task-specific attention units, where the attention mechanism for each task works independently. The output of the i-th task-specific attention is expressed as: where Q task,i , K task,i , V task,i represent the query, key, and value matrices of task i, respectively; A dynamic fusion unit for dynamic fusion. The dynamic weights are obtained by calculating the normalized exponential function of the shared attention output: W dynamic = softmax(W s ·A shared ) where W dynamic ∈R N , N is the number of tasks, which is a vector or matrix representing the fusion weights of all tasks, and Ws is a learnable weight matrix; The final output of task i is a weighted combination of the shared attention output and the task-specific attention output: Among them, W dynamic,i represents the weight term corresponding to the i-th task in the dynamic weight vector.