Virtual air station quality control data generation method based on space-time Bayesian neural field

By using a virtual air station quality control data generation method based on spatiotemporal Bayesian neural fields, the problem of calibrating mobile air quality observation data was solved, enabling accurate prediction of air quality and quantification of uncertainty, thereby improving the accuracy and reliability of air station quality control data.

CN121051702AActive Publication Date: 2025-12-02SUN YAT SEN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511584205.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-31
Publication Date
2025-12-02
Estimated Expiration
2045-10-31

AI Technical Summary

Technical Problem

Existing technologies are difficult to effectively calibrate mobile air quality observation data, especially when there are few convergence events, and lack the ability to quantify the uncertainty of the calibrated observations, thus failing to meet the requirements for data traceability, confidence, and compatibility with direct reporting.

Method used

A virtual air quality control data generation method based on spatiotemporal Bayesian neural fields is adopted. Through multi-layer spatiotemporal feature coding, feature enhancement and mean prediction modules, combined with Bayesian probability prediction modules, the conditional mean and variance set of target pollutant concentrations are generated, so as to achieve accurate prediction of air quality and quantification of uncertainty.

Benefits of technology

With discontinuous observation data, virtual station air quality results can be directly generated, and confidence intervals for corresponding spatiotemporal points can be provided, thereby improving the accuracy and reliability of air station quality control data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121051702A_ABST
    Figure CN121051702A_ABST
Patent Text Reader

Abstract

The invention discloses a virtual air station quality control data generation method based on a space-time Bayesian neural field, and relates to the technical field of data processing. A feature enhancement and mean prediction module is utilized to adaptively filter and enhance information features, a Bayesian probability prediction module is combined to jointly optimize data generation distribution and parameter uncertainty, long-range spatial-temporal dynamics in discontinuous observation data can be captured and enhanced, a virtual station air quality result is directly generated in multiple missing data modes, and the real-time performance of a virtual station is improved. The target pollutant concentration set is obtained, and the confidence interval of the corresponding time-space point is provided, so that the accuracy of air station quality control data can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a method for generating virtual air station quality control data based on spatiotemporal Bayesian neural fields. Background Technology

[0002] Air quality data is a fundamental information resource for environmental regulation, urban governance, public health early warning, and scientific research. To ensure the reliability of management decisions and public services, high spatiotemporal resolution and quality-controlled observation data are essential. Current ground-based observation systems primarily consist of fixed monitoring stations, supplemented by vehicle-mounted mobile observation methods using low-cost sensors. Fixed stations offer high data quality and traceability, but their spatial coverage is relatively sparse. Mobile observation can significantly improve spatial coverage density, filling gaps in road sections and microenvironmental observations that fixed stations cannot reach. However, the sensors in mobile observation equipment often face problems such as sensor drift, inconsistent calibration, disconnections, and discontinuous reporting during actual operation, rendering their raw observations unusable as compliant observation data.

[0003] To ensure the availability of mobile observation data, the industry commonly uses on-site calibration by comparing data obtained when mobile observation vehicles and fixed stations intersect, thereby adjusting for deviations in mobile measurements. This method is effective in correcting local biases when there are sufficient intersecting events, but its inherent limitations are also significant: intersecting events are relatively rare in time and space, and many mobile observation points are located outside the buffer zones of fixed stations, meaning that a large amount of mobile observation data cannot be directly used for calibration reference through intersecting. Furthermore, intersecting-based calibration typically only provides point-estimated correction values, lacking quantification of the uncertainty of the corrected observations, making it difficult to meet the industry's requirements for data traceability, confidence level descriptions, and compatibility with direct reporting. Summary of the Invention

[0004] In view of this, embodiments of this application provide a method, apparatus, electronic device, and storage medium for generating virtual air station quality control data based on spatiotemporal Bayesian neural fields, so as to improve the accuracy of air station quality control data.

[0005] One aspect of this application provides a method for generating virtual air station quality control data based on spatiotemporal Bayesian neural fields, the method comprising the following steps:

[0006] The raw data of the target pollutant concentration and the corresponding raw spatiotemporal features are brought into the data preprocessing module, and after sample filtering, time discretization and feature normalization, the normalized raw feature matrix is ​​obtained.

[0007] The original feature matrix is ​​fed into the multi-layer spatiotemporal feature encoding module, and after spatiotemporal interaction, time harmonic expansion, spatial harmonic expansion and spatial correlation, spatiotemporal interaction feature, spatial-spatial interaction feature, time seasonality feature, spatial Fourier feature and spatial aggregation feature are obtained. Then, the obtained features are concatenated with the original feature matrix to obtain a multi-layer spatiotemporal feature matrix.

[0008] The multi-level spatiotemporal feature matrix is ​​fed into the adaptive scaling and channel attention process in the multi-level spatiotemporal feature encoding module for processing, forming a high-dimensional feature matrix with weight redistribution.

[0009] The high-dimensional feature matrix is ​​fed into the feature enhancement and mean prediction module, and then passed through a three-layer channel-gated learning unit for feature enhancement. It is then connected to a fully connected layer network to generate the conditional mean set for the target pollutant concentration prediction.

[0010] The conditional mean set is input into the Bayesian probability prediction module. The original spatiotemporal features, the original data of the target pollutant concentration, and the conditional mean set are used to reconstruct the spatiotemporal distribution of the target pollutant. Based on the spatiotemporal distribution of the target pollutant, a target pollutant concentration set and a corresponding variance set are generated.

[0011] In some embodiments, the spatiotemporal interaction item features are obtained through the following steps:

[0012] The original feature matrix The set of timestamps in Respectively with the set of longitudes of the observation stations and the set of latitudes of observation stations Multiply to obtain the spatiotemporal interaction term features. :

[0013] ;

[0014] in, Represents element-wise product. This indicates the concatenation of vectors or matrices;

[0015] The empty-empty interaction item feature is obtained through the following steps:

[0016] The original feature matrix The longitude set of the observation stations mentioned in With the set of latitudes of the observation stations Multiply to obtain an empty-space interaction term. :

[0017] ;

[0018] The time-seasonal feature is obtained through the following steps:

[0019] The time harmonic expansion process uses orthogonal sine and cosine pairs as fixed basis functions, which are then smoothly integrated into gradient training to capture multiple cycles and form time-seasonal features. ;

[0020] The spatial Fourier term features are obtained through the following steps:

[0021] The longitude of the observation station is normalized by using the Fourier characteristics through the spatial harmonic expansion. and latitude of observation station The mapping is performed to form the spatial Fourier term features. ;

[0022] The spatial aggregation term features are obtained through the following steps:

[0023] By utilizing the aforementioned spatial associations, a single-layer multi-head graph attention network is used to learn domain-specific non-stationary spatial associations from the original feature matrix, forming the spatial aggregation term features. .

[0024] In some embodiments, the step of feeding the multi-level spatiotemporal feature matrix into the adaptive scaling and channel attention process of the multi-level spatiotemporal feature encoding module to form a high-dimensional feature matrix with weight redistribution includes the following steps:

[0025] The multi-level spatiotemporal feature matrix is ​​processed through the adaptive scaling and channel attention process. Automatic adjustment of the amplitude of each feature, amplification of key information, and suppression of redundancy are performed to form the high-dimensional feature matrix with weight redistribution. Specifically, it includes:

[0026] For the multi-level spatiotemporal feature matrix Use the following expression in each batch To scale each channel, automatic recalibration of the feature scale is achieved:

[0027] ;

[0028] in, A set of learnable scaling factors;

[0029] The scaled feature matrix is ​​expressed by the following expression. Perform global average pooling to extract channel-level statistics and eliminate spatial location interference:

[0030] ;

[0031] in, Index for the current batch; Total number of batches; This is the feature matrix after global average pooling;

[0032] The following expression uses a two-layer fully connected network to capture non-linear dependencies between channels and generates channel attention gating vectors. :

[0033] ;

[0034] in, and These are the weight matrices of the two fully connected network layers in the channel attention layer, respectively. and These are the bias vectors of the two fully connected network layers in the channel attention layer, respectively;

[0035] The channel attention gating vector is expressed by the following expression. With the scaled feature matrix Multiply to obtain the high-dimensional feature matrix with weight redistribution. :

[0036] .

[0037] In some embodiments, the high-dimensional feature matrix is ​​sequentially passed through the gated residual network, the learnable activation function, and the channel attention in each of the channel-gated learning units;

[0038] The gated residual network refers to the robust regulation of the information flow of the feature matrix through a gated residual network, specifically including:

[0039] First, set For the first The input feature matrix of the layer channel gated learning unit;

[0040] Then, the first is calculated by linear transformation according to the following expression. Intermediate feature matrix of layer gated unit :

[0041] ;

[0042] in, and They represent the first Linear transformation weight matrix of a layered network; and The first The bias vector of the linear projection of the layer network; This is a hierarchical index for channel-gated learning units;

[0043] Next, calculate the first according to the following expression. Layer gate vector Used for dynamically fusing original and enhanced features:

[0044] ;

[0045] in, and These represent the weight matrix and bias vector that generate the gate vector, respectively;

[0046] Finally, by utilizing the gating mechanism and the hybrid activation function in combination according to the following expression, the first... Layer-gated residual output vector :

[0047] ;

[0048] The learnable activation function refers to replacing the single activation function in the original gated residual network with a trainable Elu–Tanh convex hybrid activation function. Elu retains negative responses and accelerates convergence, while Tanh provides smooth and bounded outputs, thereby improving the fitting ability and uncertainty quantification capability of complex spatiotemporal patterns. Specifically, this includes:

[0049] The learnable activation function is defined by the following expression:

[0050] ;

[0051] Among them, the Learnable Hybrid Activation Functions of Layers Through the first Learnable activation blending coefficients of layers The ratio of the two activations is adaptively adjusted during training. is the input to the hybrid activation function, and Tanh is the activation function.

[0052] In some embodiments, the process of accessing the fully connected layer network to generate the conditional mean set of the target pollutant concentration prediction includes the following steps:

[0053] The conditional mean set is generated using a lightweight fully connected network via the following expression:

[0054] ;

[0055] in, Represents the set of conditional means. Indicates after the first The layer can learn the output features after processing by the activation function. and These represent the weight matrix and bias vector of the mean prediction network, respectively.

[0056] In some embodiments, the step of inputting the conditional mean set into the Bayesian probability prediction module and using the original spatiotemporal features, the original data of the target pollutant concentration, and the conditional mean set to jointly reconstruct the spatiotemporal distribution of the target pollutant includes the following steps:

[0057] The concentration distribution model of the target pollutant is defined as a Gaussian likelihood distribution, and a Logistic(0,1) prior distribution is independently set for each parameter in the parameter space.

[0058] The parameters of the Gaussian likelihood distribution are optimized by repeatedly using the maximum probability posterior method starting from random initialization points to obtain multiple local optimal solutions for the parameters;

[0059] The predicted distribution components are constructed using the local optimal solutions of each parameter, and all the predicted distribution components are mixed with equal weights to obtain the posterior predicted distribution as the spatiotemporal distribution of the target pollutant.

[0060] The Gaussian likelihood distribution is expressed as follows:

[0061] ;

[0062] in, The standard deviation of the observed noise, To collect all network parameters, It is the number of entries observed in the training loss. for The identity matrix;

[0063] The optimization objective of the maximum probability posterior method is expressed as:

[0064] ;

[0065] in, Represents the likelihood function; This is the posterior optimal solution for the parameters; This represents the estimated standard deviation of the observed noise in the output distribution during the prediction phase;

[0066] The posterior prediction distribution is represented as:

[0067] ;

[0068] in, Optimize the number of iterations to maximize the posterior probability; and The first The predicted mean and observation noise variance of each posterior mode; The likelihood model for the observed data is a Gaussian normal distribution; This represents the input feature vector of the prediction point; The model represents the input The predicted target output; This is the training dataset.

[0069] In some embodiments, generating a set of target pollutant concentrations and a corresponding set of variances based on the spatiotemporal distribution of the target pollutant includes the following steps:

[0070] The target pollutant concentration set and variance set are generated based on the spatiotemporal distribution of the target pollutant as follows:

[0071] ;

[0072] in, This represents the set of concentrations of the target pollutants. Let represent the variance set.

[0073] Another aspect of this application embodiment provides a virtual air station quality control data generation device based on spatiotemporal Bayesian neural fields, the device comprising:

[0074] The data preprocessing unit is used to input the raw data of the target pollutant concentration and the corresponding raw spatiotemporal features into the data preprocessing module, and obtain the normalized raw feature matrix after sample filtering, time discretization and feature normalization respectively.

[0075] The feature encoding unit is used to bring the original feature matrix into the multi-layer spatiotemporal feature encoding module, and obtain spatiotemporal interaction feature, spatial-spatial interaction feature, temporal-seasonal feature, spatial Fourier feature, and spatial aggregation feature through spatiotemporal interaction, temporal harmonic expansion, spatial harmonic expansion and spatial correlation, respectively. Then, the obtained features are concatenated with the original feature matrix to obtain a multi-layer spatiotemporal feature matrix.

[0076] The weight allocation unit is used to input the multi-level spatiotemporal feature matrix into the adaptive scaling and channel attention process in the multi-level spatiotemporal feature encoding module for processing, forming a high-dimensional feature matrix with weight redistribution.

[0077] The mean prediction unit is used to input the high-dimensional feature matrix into the feature enhancement and mean prediction module, and then perform feature enhancement through a three-layer channel-gated learning unit. Finally, it is connected to a fully connected layer network to generate the conditional mean set for the prediction of the target pollutant concentration.

[0078] The distribution restoration unit is used to input the conditional mean set into the Bayesian probability prediction module, and use the original spatiotemporal features, the original data of the target pollutant concentration, and the conditional mean set to restore the spatiotemporal distribution of the target pollutant, and generate a target pollutant concentration set and a corresponding variance set based on the spatiotemporal distribution of the target pollutant.

[0079] Another aspect of this application embodiment provides an electronic device, including a processor and a memory;

[0080] The memory is used to store programs;

[0081] The processor executes the program to implement any of the methods described above.

[0082] Another aspect of this application provides a computer-readable storage medium storing a program that is executed by a processor to implement the method described in any of the above embodiments.

[0083] This application includes at least the following beneficial effects:

[0084] This application captures multi-scale spatial dependencies and seasonal temporal dynamics through a multi-layer spatiotemporal feature coding module, adaptively filters and enhances information features using feature enhancement and mean prediction modules, and jointly optimizes data generation distribution and parameter uncertainty by combining a Bayesian probability prediction module. It can capture and enhance long-range spatiotemporal dynamics in discontinuous observation data, and directly generate virtual station air quality results, i.e., target pollutant concentration sets, under various missing data modes, and provide confidence intervals for corresponding spatiotemporal points, which can improve the accuracy of air station quality control data. Attached Figure Description

[0085] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0086] Figure 1 A flowchart illustrating the method for generating virtual air station quality control data based on spatiotemporal Bayesian neural fields, provided in this application embodiment;

[0087] Figure 2 An example flowchart of a virtual air station quality control data generation method based on spatiotemporal Bayesian neural fields provided in this application embodiment;

[0088] Figure 3 A technical framework diagram of a virtual air station quality control data generation method based on spatiotemporal Bayesian neural fields provided for embodiments of this application;

[0089] Figure 4 PM provided for embodiments of this application 10 Time series diagram of inference results;

[0090] Figure 5 A structural block diagram of a virtual air station quality control data generation device based on a spatiotemporal Bayesian neural field provided in an embodiment of this application. Detailed Implementation

[0091] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0092] Before providing a detailed description of the embodiments of this application, some related technologies involved in the embodiments of this application will be described first, as follows:

[0093] Given the current state and limitations of existing technologies, the concept of constructing "virtual ground air stations" at spatiotemporal locations of mobile observations has been proposed. This involves inferring the air quality value of a virtual fixed station at the spatiotemporal location where mobile observations are generated, and using this inferred value as a "pseudo-true value" to calibrate mobile observation data, particularly for a large number of observation points located outside the fixed station buffer zone. To realize this concept, there is an urgent need for a virtual ground air station quality control data generation scheme that can provide corresponding inferred air quality values ​​for any virtual station in scenarios with discontinuous air quality observation data, and quantify the confidence interval of the generated results.

[0094] To meet these needs, two-stage and end-to-end frameworks have been proposed. The two-stage framework first completes the missing data, then trains the generative model based on the complete dataset. However, this direct coupling has limitations: existing interpolation methods struggle to learn the fine spatiotemporal features required for high-precision air quality forecasting, and systematic research on how interpolation accuracy affects generation performance remains scarce. The end-to-end framework, on the other hand, simultaneously optimizes feature extraction, data completion, and target generation within a single framework. The Gaussian process method, the most widely used end-to-end method, combines the expressive power of deep networks with the closed-form interpolation and uncertainty quantification of probabilistic methods. However, applying this method to air quality generation faces two major challenges. First, posterior inference requires... The computational cost is enormous. Secondly, the selection of key parameters (such as the covariance kernel function and the mean function) is highly dependent on domain expert knowledge. Therefore, designing generation methods that can ensure high-precision, reliable uncertainty quantification while flexibly handling data with varying degrees of missing information, especially for discontinuous observation data, remains a crucial research challenge.

[0095] A search of existing technologies revealed that Chinese patent document CN116542551A (application date 2023-04-13) discloses an air quality prediction method based on incomplete data. This method addresses scenarios with incomplete observations, first categorizing historical monitoring data by "monitoring point..." × Time × Types of pollutants Organized into a three-dimensional tensor and using an observation index set Mark observable elements; for predicting the future Heaven, introducing the time sliding window dimension Expand the data into a four-dimensional tensor, and according to A corresponding weight tensor is constructed to handle the imbalance between missing and known values. Then, CP decomposition is performed on the four-dimensional tensor to establish an objective function with a regularization term. The method captures the correlations between spatial (site), temporal (including periodic sliding window), and categorical (multi-pollutant) factors by alternately minimizing the factor matrix; finally, it reconstructs the unknown elements of the tensor using the estimated factor matrix, and outputs the result. to This method provides predicted values ​​for each pollutant at each monitoring station. It addresses the problem of air quality forecasting based on incomplete data.

[0096] However, compared with this application, the method has the following technical limitations: (1) It only gives deterministic point inference values ​​and lacks the uncertainty characterization of the generated results. This method constructs a weight tensor and solves the factor matrix by alternating minimization, and directly gives the point estimate of "the predicted value of the m-th pollutant at the p-th monitoring point on day t". It does not provide quantitative results such as interval, variance or confidence level, and cannot assess the reliability and risk propagation of the generated results. In contrast, this application synchronously outputs the generated air quality values ​​of the points and their corresponding confidence intervals in the end-to-end framework, and supports the evaluation of the generated results by the average interval width and the relative average interval width index. (2) It lacks the adaptive ability of the model under different data missing modes. This method uses the observation index set The method generates a weight tensor and fits the factor matrix under the Frobenius norm, which is essentially focused on the "complete-fit" paradigm. When the missing data presents a structured missing data at the node level, timestamp level, or spatiotemporal block level, the method does not provide additional robustness or uncertainty characterization means. However, this application does not require prior imputation and directly learns and infers under any missing data mode (random missing data, node missing data, timestamp missing data, spatiotemporal block missing data). By using Bayesian parameters, noise modeling, and structural priors, the incompleteness of information caused by missing data is explicitly incorporated into uncertainty propagation. (3) The ability to express spatial dependence and multivariate correlation is limited. This method uses CP decomposition (CP rank R needs to be set) to capture the overall low-rank structure, but it does not explicitly model the spatiotemporal correlation between observation points, nor does it introduce learnable exogenous variables to assist in modeling, making it difficult to fully characterize the local heterogeneity and non-stationarity of the city. However, this application combines graph convolution module to model the adjacency relationship between observation network nodes and introduces the nonlinear effects of exogenous variables such as topography and other pollutant concentrations to improve the expression and robustness of local heterogeneity and sudden anomalies. (4) Time modeling is constrained by a fixed window, resulting in weak periodicity representation capabilities. This method incorporates the time dimension... Fixed at multiples of 7, and mapped and expanded accordingly, is equivalent to imposing a priori "weekly" structure on time dependence, which is not flexible enough for characterizing complex periodicity with multiple coexisting, drifting, or misaligned periodicity. This application captures superimposed periods at the hourly, daily, weekly, monthly, quarterly, and even annual scales using both harmonic and Fourier mechanisms, and adaptively reweights different time-scale components through gating and channel attention structures, breaking through the limitations of fixed-week windows and improving generalization ability under non-stationary and misaligned periodic conditions.

[0097] Reference Figure 1 This application provides a method for generating virtual air station quality control data based on spatiotemporal Bayesian neural fields, specifically including the following steps S100~S140:

[0098] S100: The raw data of the target pollutant concentration and the corresponding raw spatiotemporal features are brought into the data preprocessing module, and after sample filtering, time discretization and feature normalization, the normalized raw feature matrix is ​​obtained.

[0099] S110: The original feature matrix is ​​brought into the multi-layer spatiotemporal feature coding module, and after spatiotemporal interaction, time harmonic expansion, spatial harmonic expansion and spatial correlation, spatiotemporal interaction feature, spatial-spatial interaction feature, time seasonality feature, spatial Fourier feature and spatial aggregation feature are obtained. Then, the obtained features are concatenated with the original feature matrix to obtain a multi-layer spatiotemporal feature matrix.

[0100] S120: The multi-level spatiotemporal feature matrix is ​​fed into the adaptive scaling and channel attention process in the multi-level spatiotemporal feature coding module for processing, forming a high-dimensional feature matrix with weight redistribution.

[0101] S130: The high-dimensional feature matrix is ​​fed into the feature enhancement and mean prediction module, and then enhanced by a three-layer channel-gated learning unit. It is then connected to a fully connected layer network to generate the conditional mean set for the target pollutant concentration prediction.

[0102] S140: The conditional mean set is brought into the Bayesian probability prediction module. The original spatiotemporal features, the original data of the target pollutant concentration, and the conditional mean set are used to reconstruct the spatiotemporal distribution of the target pollutant. The target pollutant concentration set and the corresponding variance set are generated based on the spatiotemporal distribution of the target pollutant.

[0103] The following section will provide a detailed introduction and explanation of the solutions in the embodiments of this application, using specific application examples.

[0104] The purpose of this embodiment is to overcome the shortcomings of existing technologies and propose a method for generating virtual air quality control data based on a spatiotemporal Bayesian neural field. This method captures multi-scale spatial dependencies and seasonal temporal dynamics through multiple spatiotemporal feature coding layers, adaptively filters and enhances information features using feature enhancement and mean prediction layers, and jointly optimizes the data generation distribution and parameter uncertainty by combining a Bayesian probability prediction layer. This embodiment can capture and enhance long-range spatiotemporal dynamics in discontinuous observation data, directly generate virtual station air quality results under various missing data patterns, and provide confidence intervals for corresponding spatiotemporal points.

[0105] To achieve the above, this embodiment adopts the following quality control data generation process and specific technical solution, such as... Figure 2 and Figure 3 As shown. The method in this embodiment includes four core parts: (1) Data preprocessing module. This module performs time discretization and feature normalization on discontinuous data to form a standardized data input form. (2) Multi-layer spatiotemporal feature encoding module. This module integrates time harmonics, spatial Fourier embedding and graph attention mechanism techniques to capture seasonal cycles and cross-site dependencies in discontinuous spatiotemporal observation data. (3) Feature enhancement and mean prediction module. This module uses multi-layer channel gating units to adaptively filter and enhance spatiotemporal features and maps them to conditional means. (4) Bayesian probability prediction module. This module estimates the data distribution and generates virtual station quality control data generation results with confidence intervals through particle-based posterior maximum likelihood inference.

[0106] Specifically, this embodiment includes the following steps:

[0107] Step 1: Collect the raw data of the target pollutant concentration. and its corresponding original spatiotemporal features After being fed into the data preprocessing module, the original feature matrix after model normalization is obtained through sample filtering, time discretization, and feature normalization. The original spatiotemporal features include spatiotemporal information (time, station longitude, and station latitude) for all observation stations, as well as exogenous variable data. (Concentration data of other pollutants collected at the same time and location).

[0108] Step 2: The original feature matrix obtained in Step 1... After being fed into a multi-layer spatiotemporal feature encoding module and undergoing spatiotemporal interaction, temporal harmonic expansion, spatial harmonic expansion, and spatial correlation construction processes, the obtained spatiotemporal interaction item features are... Empty interaction item features Seasonal characteristics Spatial Fourier term features Spatial aggregation term characteristics The matrix is ​​concatenated with the original feature matrix to form a multi-level spatiotemporal feature matrix. Then, the acquired multi-level spatiotemporal feature matrix By incorporating adaptive scaling and channel attention processes, a high-dimensional feature matrix with weight redistribution is ultimately formed. .

[0109] Step 3: The high-dimensional feature matrix obtained in Step 2... The features are then fed into the feature enhancement and mean prediction modules, and enhanced through three layers of channel-gated learning units. Finally, they are connected to a fully connected network to generate a conditional mean set for pollutant concentration prediction. In this process, the feature matrix passes sequentially through three parts in each channel learning gating unit: a gated residual network, a learnable activation function, and channel attention, in order to maximize the enhancement of key feature information.

[0110] Step 4: Combine the set of conditional means predicted in Step 3. Inputting into the Bayesian probability prediction module, utilizing the original spatiotemporal features Raw data on target pollutant concentrations and the set of conditional means obtained from the prediction The spatiotemporal distribution trends of pollutants are jointly reconstructed, and the final pollutant concentration set is generated based on the reconstructed spatiotemporal distribution. and the corresponding variance set Specifically, the pollutant concentration distribution model is first defined as a Gaussian likelihood distribution, and a Logistic(0,1) prior distribution is independently set for each parameter in the parameter space. Then, the parameters of the defined conditional prediction probability model are optimized multiple times using the maximum probability posterior method starting from randomly initialized points, obtaining multiple local optima for the parameters. Next, prediction distribution components are constructed using each parameter solution, and all components are mixed with equal weights to obtain the posterior prediction distribution. Finally, the set of pollutant concentrations to be generated is obtained based on the obtained posterior prediction distribution. and its corresponding variance set .

[0111] Furthermore, the raw data of the target pollutant concentration shown in step 1 This refers to the concentration data of the target pollutant type collected from all observation stations within a certain historical time period. Station spatiotemporal information refers to the longitude and latitude coordinates of all observation stations and the time points corresponding to the pollutant concentration data collected at each station. Exogenous variable data. This refers to the concentration data of other pollutants collected at the same time and location as the target pollutant concentration at the observation station.

[0112] Furthermore, the sample filtering shown in step 1 refers to removing invalid sample records that lack the target value. Time discretization refers to calculating the time relative to the reference time. The offset will be any timestamp Mapped to integer index The minimum value is shifted to zero, as shown in Formula 1. Feature normalization refers to normalizing the feature columns other than time using the Z-score method. Here, `index()` represents the "discretized index mapping function," which maps time from the actual time axis to the integer axis.

[0113] (1)

[0114] Furthermore, the spatiotemporal interaction process shown in step 2 refers to the normalization of the original feature matrix. The set of timestamps in Respectively with the set of longitudes of the observation stations and the set of latitudes of observation stations Multiply to obtain the spatiotemporal interaction term. The formula is shown in Figure 2. Wherein, Represents element-wise product. This indicates the concatenation of vectors or matrices.

[0115] (2)

[0116] Furthermore, the empty-to-empty interaction process shown in step 2 refers to the normalized feature matrix... Longitude set in With latitude set Multiply to obtain empty-empty interaction terms. The formula is shown in Figure 3.

[0117] (3)

[0118] Furthermore, the time harmonic expansion process shown in step 2 refers to using orthogonal sine-cosine pairs as fixed basis functions and smoothly integrating them into gradient training to capture multiple periods (supporting minutes, hours, days, weeks, months, quarters, and years) to form time-seasonal features. Specifically, let's assume... Let be a base periodic set, for any base periodic set, the first... Time period Define its corresponding set of harmonic basis functions. (satisfy ).in, For the base periodic set The periodic elements in; For period The set of all harmonic basis functions generated below; For the first Time period The corresponding maximum harmonic order.

[0119] In remapped integer indices Next, regarding the first Time period The Time harmonic eigenvectors generated by first harmonic components This can be obtained through Formula 4. Subsequently, all harmonics of all fundamental periods are cascaded along their respective channels to form a seasonal characteristic term. .in, This is the harmonic order index.

[0120] (4)

[0121] Furthermore, the spatial harmonic expansion process shown in step 2 refers to using Fourier characteristics to analyze the normalized longitude. and latitude Mapping is performed to form spatial Fourier term features. Specifically, for spatial coordinate channel indexes , define the first Maximum harmonic order of the axis and the normalized first Axis coordinates Mapped to the Fourier eigenvectors of each spatial axis The formula is shown in Figure 5. Subsequently, all spatial mappings are concatenated along the channels to form spatial Fourier terms. .in, This is the spatial harmonic order index.

[0122] (5)

[0123] Furthermore, the spatial association construction process shown in step 2 refers to using a single-layer multi-head graph attention network to learn domain-specific non-stationary spatial associations from the data, forming spatial aggregation feature items. Specifically, for any observation station index Based on historical pollutant concentration data collected from the station, the average concentration, 25th percentile, and 75th percentile statistical indicators of the pollutant concentration at that station are obtained. These statistical indicators are then used to construct an initial node feature vector for a static station. .have The graph attention network of the size only follows the adjacency matrix Information propagates from non-zero edges, and its attention weights are modulated by Gaussian connection priors. This attention mechanism can be represented by Equation 6.

[0124] (6)

[0125] in, For the number of heads; and Indicates the index of any two adjacent sites; For the first Attention kernel vectors for each attention head; For the first The linear projection matrix of each attention head; The input feature vectors of adjacent nodes; For the first Unnormalized attention coefficients under each attention head; For nodes and nodes The Euclidean distance; To control the scale parameter of distance attenuation; The attention score is the weighted average. To normalize attention weights; For the set of observation stations; ReLU, softmax, and ELU are different activation functions.

[0126] (7)

[0127] The output of attention is connected and linearly projected according to Equation 7 to form nodes. Aggregated feature output vector .in, This represents the output projection matrix, and `concat` represents the concatenation of feature matrices. Finally, they are stacked along the time dimension. Ultimately, spatial aggregation features are obtained. .

[0128] Furthermore, the adaptive scaling and channel attention process shown in step 2 refers to the scaling of multi-level spatiotemporal feature matrices. Automatic adjustment of the amplitude of each feature, amplification of key information, and suppression of redundancy are performed to ultimately form a high-dimensional feature matrix with weight redistribution. Specifically, for multi-level spatiotemporal feature matrices First, following the method in Formula 8, by using in each batch This is used to scale each channel, enabling automatic recalibration of the feature scale. It is a set of learnable scaling factors.

[0129] (8)

[0130] Then, according to Formula 9, the scaled feature matrix is... Global average pooling is performed to extract channel-level statistics and eliminate spatial location interference. Index for the current batch; Total number of batches; This is the feature matrix after global average pooling.

[0131] (9)

[0132] Next, following Equation 10, a two-layer fully connected network is used to capture the non-linear dependencies between channels and generate channel attention gating vectors. .in, and These are the weight matrices of the two fully connected network layers in the channel attention layer, respectively. and These are the bias vectors of the two fully connected network layers in the channel attention layer, respectively.

[0133] (10)

[0134] Finally, according to Formula 11, the channel attention gating vector is... Compared with scaled features Multiply the matrices to obtain the high-dimensional feature matrix with weight redistribution. :

[0135] (11)

[0136] Furthermore, the gated residual network module shown in step 3 refers to the robust control of the information flow of the feature matrix through a gated residual network. Specifically, firstly, a... For the first The input feature matrix of the layer channel-gated learning unit. Then, according to Equation 12, the first layer is calculated through a linear transformation. Intermediate feature matrix of layer gated unit .in, and They represent the first Linear transformation weight matrix of a layered network; and The first The bias vector of the linear projection of the layer network; This is the hierarchical index for the channel gating unit.

[0137] (12)

[0138] Next, according to Formula 13, calculate the first... Layer gate vector This is used for dynamically fusing original and enhanced features. and These represent the weight matrix and bias vector used to generate the gated vector, respectively.

[0139] (13)

[0140] Finally, according to Formula 14, by utilizing the combined effect of gating mechanism and hybrid activation function, the first... Layer-gated residual output vector .

[0141] (14)

[0142] Furthermore, the learnable activation function shown in step 3 refers to replacing the single activation function in the original gated residual network with a trainable Elu–Tanh convex hybrid activation function. Elu retains the negative response and accelerates convergence, while Tanh provides a smooth and bounded output, thereby improving the fitting ability and uncertainty quantification capability of complex spatiotemporal patterns. Specifically, the learnable activation function can be defined by Equation 15. Wherein, the... Learnable Hybrid Activation Functions of Layers Through the first Learnable activation blending coefficients of layers The ratio of the two activations is adaptively adjusted during training. The input to this hybrid activation function module is Tanh, where Tanh is the activation function.

[0143] (15)

[0144] Furthermore, the conditional mean shown in step 3 is generated by a lightweight fully connected network. Specifically, This represents the set of concentration conditional means. Indicates after the first The layer can learn the output features after processing by the activation function. and These represent the weight matrix and bias vector of the mean prediction network, respectively.

[0145] (16)

[0146] Furthermore, the Gaussian likelihood distribution shown in step 4 can be expressed using Equation 17. Wherein, The standard deviation of the observed noise, To collect all network parameters (scaling factors, network weights, bias and activation combination coefficients), It is the number of entries observed in the training loss. for The identity matrix.

[0147] (17)

[0148] Furthermore, the optimization objective of the parameter estimation method based on maximum a posteriori probability shown in step 4 can be expressed by Equation 18. Wherein, Represents the likelihood function; This is the posterior optimal solution for the parameters; This represents the estimated standard deviation of the observed noise in the output distribution during the prediction phase.

[0149] (18)

[0150] Furthermore, the posterior prediction distribution shown in step 4 can be represented by Equation 19. Wherein, Optimize the number of iterations to maximize the posterior probability; and The first The predicted mean and observation noise variance of each posterior mode; The likelihood model for the observed data is a Gaussian normal distribution; This represents the input feature vector of the prediction point; The model represents the input The predicted target output; This is the training dataset.

[0151] (19)

[0152] Furthermore, the final generated output mean and variance shown in step 4 can be obtained using formula 20.

[0153] (20)

[0154] In summary, this embodiment includes the following key technical solutions:

[0155] 1. A method for generating virtual air station quality control data based on spatiotemporal Bayesian neural fields is provided, including the following four core steps: (1) Data preprocessing module. This module performs time discretization and feature normalization on discontinuous data to form a standardized data input format. (2) Multi-layer spatiotemporal feature encoding module. This module integrates time harmonics, spatial Fourier embedding and graph attention mechanism techniques to capture seasonal cycles and cross-site dependencies in discontinuous spatiotemporal observation data. (3) Feature enhancement and mean prediction module. This module uses multi-layer channel gating units to adaptively filter and enhance spatiotemporal features and maps them to conditional means. (4) Bayesian probability prediction module. This module estimates the data distribution and generates results with confidence intervals through particle-based posterior maximum likelihood inference.

[0156] 2. In step (1), the raw data of the target pollutant concentration collected. and its corresponding spatiotemporal characteristic information The data will be fed into the data preprocessing module, where it will undergo sample filtering, time discretization, and feature normalization to obtain the model-normalized feature matrix. The spatiotemporal feature information includes the spatiotemporal information (time, station longitude, and station latitude) of all observation stations, as well as exogenous variable data. (Concentrations of other types of pollutants collected during the other simultaneous observations).

[0157] 3. In step (2), the feature matrix obtained The resulting spatiotemporal feature will be fed into a multi-layer spatiotemporal feature encoding module, where it will undergo spatiotemporal interaction, temporal harmonic expansion, spatial harmonic expansion, and spatial correlation construction processes. The resulting spatiotemporal interaction feature will then be... Empty interaction item features Seasonal characteristics Spatial Fourier term features Spatial aggregation term characteristics The original feature matrix is ​​concatenated with the feature matrix to form a multi-level spatiotemporal feature matrix. .

[0158] 4. In step (2), the multi-level spatiotemporal feature matrix is ​​obtained. The input will be fed into an adaptive scaling and channel attention process to form a high-dimensional feature tensor with weight redistribution. .

[0159] 5. In step (3), the high-dimensional feature tensor obtained The data will be fed into the feature enhancement and mean prediction module, where it will undergo feature enhancement through a three-layer channel-gated learning unit. It will then be connected to a fully connected network to generate a conditional mean set for pollutant concentration prediction. .

[0160] 6. In step (3), each channel learning gating unit comprises three parts: a gated residual network, a learnable activation function, and channel attention. High-dimensional feature matrix The learning process sequentially passes through the three parts of the gating unit in each channel layer to maximize the enhancement of key feature information. The learnable activation function is a convex combination of two separate activation functions, Elu and Tanh.

[0161] 7. In step (4), the set of conditional means obtained by prediction It will be fed into the Bayesian probability prediction module, utilizing the original spatiotemporal features. Raw data on target pollutant concentrations and the set of conditional means obtained from the prediction The spatiotemporal distribution trends of pollutant concentrations are jointly reconstructed, and a final set of pollutant concentrations is generated based on the reconstructed spatiotemporal distribution. and the corresponding variance set .

[0162] 8. In step (4), the specific steps for restoring the spatiotemporal distribution trend of pollutant concentration are as follows: First, the pollutant concentration distribution model is defined as a Gaussian likelihood distribution, and a Logistic(0,1) prior distribution is independently set for each parameter in the parameter space. Then, the maximum probability posterior method starting from random initialization points is used multiple times to optimize the parameters of the defined conditional prediction probability model, obtaining multiple local optimal solutions for the parameters. Next, prediction distribution components are constructed using each parameter solution, and all components are mixed with equal weight to obtain the posterior prediction distribution. Finally, the pollutant concentration set to be generated is obtained based on the obtained posterior prediction distribution. and its corresponding variance set .

[0163] This embodiment provides a method for generating virtual air quality control data based on spatiotemporal Bayesian neural fields. It unifies feature extraction, target generation, and uncertainty quantification within a single system, enabling direct inference under four common data missing modes: random missing data, node missing data, timestamp missing data, and spatiotemporal block missing data. This avoids the additional steps of interpolation or padding required by traditional methods, thus improving processing efficiency and data generation reliability. Furthermore, this embodiment integrates graph attention mechanisms, temporal harmonic decomposition, and spatial Fourier embedding to construct a multi-scale spatiotemporal encoder, effectively capturing cross-scale spatiotemporal dependencies and correlations under irregular sampling conditions. Simultaneously, it introduces a channel-gated learning unit, combining channel attention mechanisms, gated residual structures, and learnable activation functions to dynamically filter and enhance key information channels, effectively suppressing noise and maintaining the stability of the training process. Experimental results show that this embodiment exhibits lower inference errors on real air quality datasets and generates inference results with narrower intervals, balancing inference accuracy and reliability. It provides reliable and efficient technical support for the generation of virtual ground air quality control data and related applications.

[0164] The following is sample data to verify this embodiment:

[0165] This embodiment selects a large-scale, discontinuous spatiotemporal dataset of air quality from a certain region for PM2.5 analysis. 10 Concentration generation experiment. The dataset includes collection time, station longitude, station latitude, and PM2.5 concentration. 10 Basic information on concentrations of SO2, NO2, O3, and PM2.5 2.5 Information on four types of exogenous covariates.

[0166] Table 1. Basic Information of the Dataset

[0167]

[0168] Table 2 Sample Data

[0169]

[0170] To evaluate the robustness of this embodiment under different levels of data missing conditions, four common missing data scenarios were simulated: (i) random missing data, i.e., intermittent loss of midpoints in the data stream (e.g., packet loss due to temporary degradation of communication quality); (ii) node missing data, i.e., all observations of a single node are missing for a relatively long period of time (e.g., continuous sensor failure); (iii) timestamp missing data, i.e., all nodes are simultaneously unable to acquire data at a specific time (e.g., partial power outage); and (iv) spatiotemporal block missing data, i.e., missing data forming continuous spatiotemporal blocks (e.g., persistent blind spots caused by a moving sensor traversing a tunnel). For each scenario, a classic missing rate of 30% was selected. Target values ​​were removed according to a specified missing pattern. In cross-validation, observation sites were randomly divided into five mutually exclusive subsets. In each fold, the records of the most recent month from the retained subset sites constituted the test set. The training set contained all remaining data: full-cycle observation data from other sites and non-test period data from the retained subset sites.

[0171] All experiments were conducted on a server equipped with an Intel Xeon Gold 6133 CPU and four NVIDIA GeForce RTX 4090 graphics cards. The generation method in this embodiment consists of three stacked channel-gated learning layers, each containing 512 hidden units and 16 particles. Training was performed using the AdamW optimizer with an initial learning rate of... The batch size was 512, and the maximum training epochs were 5000. Hyperparameters were tuned in the initial experiments, and the contribution of each module was evaluated through ablation studies. Under the same settings, the method in this embodiment and all baseline models were trained and validated on five non-overlapping split sets of each dataset, and average performance metrics were reported to achieve fair comparisons in scenarios with incomplete observations. Among them, four models were selected as baseline methods: Historical Average (HA), Random Forest (RF), Spatiotemporal Gradient Boosting Tree (STGBOOST), and Spatiotemporal Sparse Variational Gaussian Process (ST-SVGP). The data generation performance evaluation metrics were root mean square error (RMSE), mean absolute error (MAE), and coefficient of determination (R²). 2 There are six types: Symmetrical Mean Absolute Percentage Error (SMAPE), Average Interval Width (AIW), and Average Relative Interval Width (RIWM).

[0172] The method described in this embodiment is used to perform PM in scenarios with original data, random missing data, missing nodes, missing timestamps, and missing spatiotemporal blocks. 10 The concentration inference and evaluation indicators of the virtual station quality control data generation results are shown in Tables 3-7.

[0173] Table 3 PM in the original data scenario 10 Inference results

[0174]

[0175] Table 4 PM in random missing scenarios 10 Inference results

[0176]

[0177] Table 5 PM in Node Missing Scenarios 10 Inference results

[0178]

[0179] Table 6 PM in scenarios with missing timestamps 10 Inference results

[0180]

[0181] Table 7 PM in scenarios with missing spatiotemporal blocks 10 Inference results

[0182]

[0183] As can be seen from the data in the tables above, the generation method proposed in this embodiment is significantly superior to other air quality generation methods with uncertainty estimation capabilities in scenarios involving raw data, random missing data, missing nodes, missing timestamps, and missing spatiotemporal blocks. Furthermore, Figure 4 The results of generating a virtual station and its confidence interval for a certain time period under the original data scenario are shown.

[0184] Reference Figure 5 This application provides a virtual air station quality control data generation device based on spatiotemporal Bayesian neural fields, including:

[0185] The data preprocessing unit is used to input the raw data of the target pollutant concentration and the corresponding raw spatiotemporal features into the data preprocessing module, and obtain the normalized raw feature matrix after sample filtering, time discretization and feature normalization respectively.

[0186] The feature encoding unit is used to bring the original feature matrix into the multi-layer spatiotemporal feature encoding module, and obtain spatiotemporal interaction feature, spatial-spatial interaction feature, temporal-seasonal feature, spatial Fourier feature, and spatial aggregation feature through spatiotemporal interaction, temporal harmonic expansion, spatial harmonic expansion and spatial correlation, respectively. Then, the obtained features are concatenated with the original feature matrix to obtain a multi-layer spatiotemporal feature matrix.

[0187] The weight allocation unit is used to input the multi-level spatiotemporal feature matrix into the adaptive scaling and channel attention process in the multi-level spatiotemporal feature encoding module for processing, forming a high-dimensional feature matrix with weight redistribution.

[0188] The mean prediction unit is used to input the high-dimensional feature matrix into the feature enhancement and mean prediction module, and then perform feature enhancement through a three-layer channel-gated learning unit. Finally, it is connected to a fully connected layer network to generate the conditional mean set for the prediction of the target pollutant concentration.

[0189] The distribution restoration unit is used to input the conditional mean set into the Bayesian probability prediction module, and use the original spatiotemporal features, the original data of the target pollutant concentration, and the conditional mean set to restore the spatiotemporal distribution of the target pollutant, and generate a target pollutant concentration set and a corresponding variance set based on the spatiotemporal distribution of the target pollutant.

[0190] It is understood that the content of the above method embodiments is applicable to the present device embodiments. The specific functions implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.

[0191] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order shown in the operation diagrams. For example, depending on the functions / operations involved, two consecutively shown blocks may actually be executed substantially simultaneously, or the blocks may sometimes be executed in reverse order. Furthermore, the embodiments presented and described in the flowcharts of this application are provided by way of example to provide a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logic flows presented herein. Alternative embodiments are contemplated in which the order of various operations is changed and sub-operations described as part of a larger operation are executed independently.

[0192] Furthermore, although this application is described in the context of functional modules, it should be understood that, unless otherwise stated, one or more of the described functions and / or features may be integrated into a single physical device and / or software module, or one or more functions and / or features may be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding this application. Rather, given the properties, functions, and internal relationships of the various functional modules in the apparatus disclosed herein, the actual implementation of the module will be understood within the scope of conventional technology for an engineer. Therefore, those skilled in the art can implement the application set forth in the claims using ordinary techniques without excessive experimentation. It is also understood that the specific concepts disclosed are merely illustrative and not intended to limit the scope of this application, which is determined by the full scope of the appended claims and their equivalents.

[0193] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0194] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0195] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0196] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0197] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0198] Although embodiments of this application have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of this application, the scope of which is defined by the claims and their equivalents.

[0199] The above is a detailed description of the preferred embodiments of this application, but this application is not limited to the embodiments described. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of this application, and these equivalent modifications or substitutions are all included within the scope defined by the claims of this application.

Claims

1. A method for generating virtual air station quality control data based on spatiotemporal Bayesian neural fields, characterized in that, The method includes the following steps: The raw data of the target pollutant concentration and the corresponding raw spatiotemporal features are brought into the data preprocessing module, and after sample filtering, time discretization and feature normalization, the normalized raw feature matrix is ​​obtained. The original feature matrix is ​​fed into the multi-layer spatiotemporal feature encoding module, and after spatiotemporal interaction, time harmonic expansion, spatial harmonic expansion and spatial correlation, spatiotemporal interaction feature, spatial-spatial interaction feature, time seasonality feature, spatial Fourier feature and spatial aggregation feature are obtained. Then, the obtained features are concatenated with the original feature matrix to obtain a multi-layer spatiotemporal feature matrix. The multi-level spatiotemporal feature matrix is ​​fed into the adaptive scaling and channel attention process in the multi-level spatiotemporal feature encoding module for processing, forming a high-dimensional feature matrix with weight redistribution. The high-dimensional feature matrix is ​​fed into the feature enhancement and mean prediction module, and then passed through a three-layer channel-gated learning unit for feature enhancement. It is then connected to a fully connected layer network to generate the conditional mean set for the target pollutant concentration prediction. The conditional mean set is input into the Bayesian probability prediction module. The original spatiotemporal features, the original data of the target pollutant concentration, and the conditional mean set are used to reconstruct the spatiotemporal distribution of the target pollutant. Based on the spatiotemporal distribution of the target pollutant, a target pollutant concentration set and a corresponding variance set are generated.

2. The method for generating virtual air station quality control data based on spatiotemporal Bayesian neural fields according to claim 1, characterized in that, The spatiotemporal interaction term features are obtained through the following steps: The original feature matrix The set of timestamps in Respectively with the set of longitudes of the observation stations and the set of latitudes of observation stations Multiply to obtain the spatiotemporal interaction term features. : ; in, Represents element-wise product. This indicates the concatenation of vectors or matrices; The empty-empty interaction item feature is obtained through the following steps: The original feature matrix The longitude set of the observation stations mentioned in With the set of latitudes of the observation stations Multiply to obtain an empty-empty interaction term. : ; The time-seasonal feature is obtained through the following steps: The time harmonic expansion process uses orthogonal sine and cosine pairs as fixed basis functions, which are then smoothly integrated into gradient training to capture multiple cycles and form time-seasonal features. ; The spatial Fourier term features are obtained through the following steps: The longitude of the observation station is normalized by using the Fourier characteristics through the spatial harmonic expansion. and the latitude of the observation station The mapping is performed to form the spatial Fourier term features. ; The spatial aggregation term features are obtained through the following steps: By utilizing the aforementioned spatial associations, a single-layer multi-head graph attention network is used to learn domain-specific non-stationary spatial associations from the original feature matrix, forming the spatial aggregation term features. .

3. The method for generating virtual air station quality control data based on spatiotemporal Bayesian neural fields according to claim 1, characterized in that, The process of inputting the multi-level spatiotemporal feature matrix into the adaptive scaling and channel attention process of the multi-level spatiotemporal feature encoding module to form a high-dimensional feature matrix with weight redistribution includes the following steps: The multi-level spatiotemporal feature matrix is ​​processed through the adaptive scaling and channel attention process. Automatic adjustment of the amplitude of each feature, amplification of key information, and suppression of redundancy are performed to form the high-dimensional feature matrix with weight redistribution. Specifically, it includes: For the multi-level spatiotemporal feature matrix Use the following expression in each batch To scale each channel, automatic recalibration of the feature scale is achieved: ; in, A set of learnable scaling factors; The scaled feature matrix is ​​expressed by the following expression. Perform global average pooling to extract channel-level statistics and eliminate spatial location interference: ; in, Index for the current batch; Total number of batches; This is the feature matrix after global average pooling; The following expression uses a two-layer fully connected network to capture non-linear dependencies between channels and generates channel attention gating vectors. : ; in, and These are the weight matrices of the two fully connected network layers in the channel attention layer, respectively. and These are the bias vectors of the two fully connected network layers in the channel attention layer, respectively; The channel attention gating vector is expressed by the following expression. With the scaled feature matrix Multiply to obtain the high-dimensional feature matrix with weight redistribution. : 。 4. The method for generating virtual air station quality control data based on spatiotemporal Bayesian neural fields according to claim 1, characterized in that, The high-dimensional feature matrix is ​​sequentially passed through the gated residual network, the learnable activation function, and the channel attention in each of the channel-gated learning units. The gated residual network refers to the robust regulation of the information flow of the feature matrix through a gated residual network, specifically including: First, set For the first The input feature matrix of the layer channel gated learning unit; Then, the first is calculated by linear transformation according to the following expression. Intermediate feature matrix of layer gated unit : ; in, and They represent the first Linear transformation weight matrix of a layered network; and The first The bias vector of the linear projection of the layer network; This is a hierarchical index for channel-gated learning units; Next, calculate the first according to the following expression. Layer gate vector Used for dynamically fusing original and enhanced features: ; in, and These represent the weight matrix and bias vector that generate the gate vector, respectively; Finally, by utilizing the gating mechanism and the hybrid activation function in combination according to the following expression, the first... Layer-gated residual output vector : ; The learnable activation function refers to replacing the single activation function in the original gated residual network with a trainable Elu–Tanh convex hybrid activation function. Elu retains negative responses and accelerates convergence, while Tanh provides smooth and bounded outputs, thereby improving the fitting ability and uncertainty quantification capability of complex spatiotemporal patterns. Specifically, this includes: The learnable activation function is defined by the following expression: ; Among them, the Learnable Hybrid Activation Functions of Layers Through the first Learnable activation blending coefficients of layers The ratio of the two activations is adaptively adjusted during training. is the input to the hybrid activation function, and Tanh is the activation function.

5. The method for generating virtual air station quality control data based on spatiotemporal Bayesian neural fields according to claim 1, characterized in that, The process of connecting to the fully connected layer network to generate the conditional mean set for the predicted concentration of the target pollutant includes the following steps: The conditional mean set is generated using a lightweight fully connected network via the following expression: ; in, Represents the set of conditional means. Indicates after the first The layer can learn the output features after processing by the activation function. and These represent the weight matrix and bias vector of the mean prediction network, respectively.

6. The method for generating virtual air station quality control data based on spatiotemporal Bayesian neural fields according to claim 1, characterized in that, The step of inputting the conditional mean set into the Bayesian probability prediction module and using the original spatiotemporal features, the original data of the target pollutant concentration, and the conditional mean set to reconstruct the spatiotemporal distribution of the target pollutant includes the following steps: The concentration distribution model of the target pollutant is defined as a Gaussian likelihood distribution, and a Logistic(0,1) prior distribution is independently set for each parameter in the parameter space. The parameters of the Gaussian likelihood distribution are optimized by repeatedly using the maximum probability posterior method starting from random initialization points to obtain multiple local optimal solutions for the parameters; The predicted distribution components are constructed using the local optimal solutions of each parameter, and all the predicted distribution components are mixed with equal weights to obtain the posterior predicted distribution as the spatiotemporal distribution of the target pollutant. The Gaussian likelihood distribution is expressed as follows: ; in, The standard deviation of the observed noise To collect all network parameters, It is the number of entries observed in the training loss. for The identity matrix; The optimization objective of the maximum probability posterior method is expressed as: ; in, Represents the likelihood function; This is the posterior optimal solution for the parameters; This represents the estimated standard deviation of the observed noise in the output distribution during the prediction phase; The posterior prediction distribution is represented as: ; in, Optimize the number of iterations to maximize the posterior probability; and The first The predicted mean and observation noise variance of each posterior mode; The likelihood model for the observed data is a Gaussian normal distribution; This represents the input feature vector of the prediction point; The model represents the input The predicted target output; This is the training dataset.

7. The method for generating virtual air station quality control data based on spatiotemporal Bayesian neural fields according to claim 6, characterized in that, The step of generating a set of target pollutant concentrations and a corresponding set of variances based on the spatiotemporal distribution of the target pollutant includes the following steps: The target pollutant concentration set and variance set are generated based on the spatiotemporal distribution of the target pollutant as follows: ; in, This represents the set of concentrations of the target pollutants. Let represent the variance set.

8. A virtual air station quality control data generation device based on spatiotemporal Bayesian neural fields, characterized in that, The device includes: The data preprocessing unit is used to input the raw data of the target pollutant concentration and the corresponding raw spatiotemporal features into the data preprocessing module, and obtain the normalized raw feature matrix after sample filtering, time discretization and feature normalization respectively. The feature encoding unit is used to bring the original feature matrix into the multi-layer spatiotemporal feature encoding module, and obtain spatiotemporal interaction feature, spatial-spatial interaction feature, temporal-seasonal feature, spatial Fourier feature, and spatial aggregation feature through spatiotemporal interaction, temporal harmonic expansion, spatial harmonic expansion and spatial correlation, respectively. Then, the obtained features are concatenated with the original feature matrix to obtain a multi-layer spatiotemporal feature matrix. The weight allocation unit is used to input the multi-level spatiotemporal feature matrix into the adaptive scaling and channel attention process in the multi-level spatiotemporal feature encoding module for processing, forming a high-dimensional feature matrix with weight redistribution. The mean prediction unit is used to input the high-dimensional feature matrix into the feature enhancement and mean prediction module, and then perform feature enhancement through a three-layer channel-gated learning unit. Finally, it is connected to a fully connected layer network to generate the conditional mean set for the prediction of the target pollutant concentration. The distribution restoration unit is used to input the conditional mean set into the Bayesian probability prediction module, and use the original spatiotemporal features, the original data of the target pollutant concentration, and the conditional mean set to restore the spatiotemporal distribution of the target pollutant, and generate a target pollutant concentration set and a corresponding variance set based on the spatiotemporal distribution of the target pollutant.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory; The memory is used to store programs; The processor executes the program to implement the method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a program that is executed by a processor to implement the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Air quality prediction method and system based on incomplete data

    CN116542551A

  • Efficient and accurate pollution source monitoring quality control system and method

    CN119443922A

  • Source identification by non-negative matrix factorization combined with semi-supervised clustering

    US20180060758A1