Cross-region pm2.5 concentration prediction method based on improved spatial autoregressive model

CN122333347BActive Publication Date: 2026-09-08CHINA UNIV OF PETROLEUM (EAST CHINA)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610445632.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-04-07
Publication Date
2026-09-08
Estimated Expiration
2046-04-07

AI Technical Summary

Technical Problem

[0010]本发明的主要目的在于提供一种基于改进空间自回归模型的跨区域PM2.5浓度预测方法,以解决现有技术中跨区域PM2.5浓度预测中存在环境样本稀缺、传感器硬件测量误差的问题

Benefits of technology

1、步骤S1将传感器硬件测量误差剥离机制嵌入跨区域实际PM2.5浓度物理扩散预测架构。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122333347B_ABST
    Figure CN122333347B_ABST
Patent Text Reader

Abstract

The application provides a cross-region PM2.5 concentration prediction method based on an improved spatial autoregressive model, and specifically comprises the following steps: generating a mature regional information source domain set with real atmospheric pollution evolution reference value based on a residual error resampling technology; performing high-dimensional meteorological and emission feature preliminary screening to obtain a core meteorological and emission active feature set after removing sensor hardware measurement error interference; constructing a candidate model according to the pollution driving importance represented by the absolute value of the physical feature regression coefficient, and outputting the spatial parameter estimator matrix corresponding to each candidate model; and scoring the diffusion model by using the negative half number high-dimensional Bayesian information, and outputting the optimal weighted regional actual PM2.5 concentration spatial parameter estimator. The technical scheme of the application overcomes the problems of environmental sample scarcity and sensor hardware measurement error in the cross-region PM2.5 concentration prediction in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of PM2.5 concentration prediction, and specifically to a cross-regional PM2.5 concentration prediction method based on an improved spatial autoregressive model. Background Technology

[0002] With the deepening of the national action plan for continuous improvement of air quality, the cross-regional collaborative governance of fine particulate matter pollution has become a core issue in atmospheric environmental management. Accurately quantifying the cross-regional concentration evolution of air pollutants has practical engineering guiding value.

[0003] With the expansion of multi-source environmental sensing dimensions in modern IoT, the aforementioned air pollution diffusion analysis framework faces the dual challenges of the curse of dimensionality in high-dimensional features and measurement errors in underlying hardware data in practical engineering applications for emerging industrial areas. On the one hand, to assess the evolution mechanism of environmental pollution, the system needs to extract a large number of high-dimensional monitoring indicators, such as objectively existing but inaccurately observable meteorological and emission physical characteristics. However, effective high-precision monitoring samples are relatively scarce in emerging industrial areas that are in the early stages of construction. On the other hand, due to the widespread use of low-cost micro IoT sensors in large-scale monitoring networks, the observation feature matrix collected by the system, which includes hardware measurement errors, suffers from systematic hardware measurement errors due to the aging of physical sensors, weather interference, or signal transmission loss. The combination of these factors exposes four major technical defects in existing air pollution spatial prediction technologies.

[0004] The first technical drawback of existing technologies lies in their difficulty in handling inaccurate descriptions of environmental characteristics due to sensor aging or signal attenuation. This causes the calculated pollution-driven models to deviate from the actual physical diffusion patterns. In standard atmospheric pollution spatial physics model analysis, if basic maximum likelihood fitting calculations are performed directly based on the observation feature matrix, which includes hardware measurement errors, fluctuations in these errors introduce amplified biases into the underlying physical calculations. This bias causes the true regression coefficients of various core meteorological and emission characteristics to shrink towards zero, thereby systematically distorting the physical identification of the intensity of air pollution transmission between adjacent grids.

[0005] The second technical drawback of existing technologies lies in the estimation variance inflation caused by the limited historical high-precision monitoring data of emerging industrial areas. When faced with hundreds of micro-meteorological and industrial emission physical characteristics, the small effective environmental monitoring sample size of emerging industrial areas means that relying solely on the limited sensing data of the area itself for independent pollution prediction modeling leads to a large statistical variance in parameter estimation. This reduces the practical guiding significance of the final output of actual PM2.5 concentration prediction results for environmental protection planning and heavy pollution early warning.

[0006] The third technical drawback of existing technologies lies in the fact that sensor hardware measurement errors interfere with the extraction of core pollution-driving features and the robust fusion of cross-regional sensing data. To compensate for the lack of sensing data in emerging industrial areas, conventional solutions attempt to utilize mature areas with well-established surrounding monitoring networks for transfer learning. However, the underlying logic of existing spatial transfer learning algorithms is usually based on the assumption of pure hardware data. In environments containing systematic hardware measurement errors, traditional feature reduction and screening algorithms are prone to selecting false meteorological features with no real environmental driving significance into pollution prediction models. Furthermore, they struggle to effectively identify and eliminate monitoring data from mature surrounding areas that no longer have reference value due to local micro-topographical obstruction or changes in the underlying heavy industrial structure. This can lead to negative data contagion effects and affect the identification of the true evolution pattern of air pollution in the target area.

[0007] The fourth technical deficiency of existing technologies lies in the instability of the single optimal pollution prediction model selection strategy under noisy sensor hardware environments. Currently, mainstream air quality assessment systems generally employ a hard threshold model selection strategy, attempting to find an absolutely optimal physical diffusion prediction formula. However, the existence of sensor hardware measurement errors reduces the signal-to-noise ratio of the overall environmental sensing data, causing the optimal diffusion model evaluated based on a single information criterion to easily undergo structural jumps between different finite sample data slices. This prediction decision-making mechanism ignores the uncertainty of model selection in high-dimensional sparse environments, posing a high risk of misclassifying environmental pollution propagation mechanisms.

[0008] In summary, traditional hardware error correction techniques are difficult to effectively leverage historical sensing data from surrounding mature areas in emerging industrial zones with limited monitoring samples. Furthermore, existing spatial migration sensing frameworks have limitations in addressing systematic physical measurement biases and model selection fluctuations caused by low-cost IoT sensors.

[0009] Therefore, there is a need for a cross-regional PM2.5 concentration prediction method based on an improved spatial autoregressive model that can simultaneously overcome the scarcity of environmental samples, correct sensor hardware measurement errors, and output reliable statistical confidence. Summary of the Invention

[0010] The main objective of this invention is to provide a cross-regional PM2.5 concentration prediction method based on an improved spatial autoregressive model, in order to solve the problems of scarce environmental samples and sensor hardware measurement errors in the existing cross-regional PM2.5 concentration prediction technology.

[0011] To achieve the above objectives, this invention provides a method for predicting cross-regional PM2.5 concentrations based on an improved spatial autoregressive model, specifically including the following steps: S1, based on residual resampling technology, generates a set of mature regional information source domains with real reference value for the evolution of atmospheric pollution; S2, perform preliminary screening of high-dimensional meteorological and emission characteristics to obtain a set of core meteorological and emission activity characteristics after removing interference from sensor hardware measurement errors; S3: Based on the ranking of pollution driving importance represented by the absolute value of the regression coefficients of physical characteristics, construct candidate models and output the spatial parameter estimation matrix corresponding to each candidate model. S4 uses the negative half-dimensional high-dimensional Bayesian information criterion diffusion model for scoring, and outputs the optimal weighted regional actual PM2.5 concentration spatial parameter estimate to the business terminal of external environmental protection departments or public health institutions.

[0012] Furthermore, step S1 specifically includes the following steps: S1.1, receives the response variable vector of the actual PM2.5 concentration in the emerging industrial zone. Observation feature matrix including systematic sensor hardware measurement errors Physical adjacency matrix of the target area in the physical spatial topology of each monitoring grid And the sensor hardware measurement error covariance matrix that is pre-acquired and characterizes the correlation of multidimensional measurement errors. , as input data.

[0013] S1.2, In order to eliminate the bias of regression coefficients shrinking to zero due to hardware measurement errors from the data level, the negative modified log-likelihood function with absolute value norm penalty term is called for initial data processing: ; ; in, This is the initial high-dimensional physical feature regression coefficient vector; These are the initial spatial autoregressive coefficients; This is an estimate of the initial random disturbance variance of the atmospheric environment system; To find the set of independent variables that minimizes the objective function; The target domain is modified for the log-likelihood function; Initial adjustment parameters for controlling the sparsity of the high-dimensional feature space; This is the absolute value norm penalty term for the regression coefficient vector; This refers to the sample size for environmental monitoring in the target area. This is the spatial transformation matrix of the target region in the corresponding dimension; The variance of random disturbances in the atmospheric environment system; This represents the response variable vector of the actual PM2.5 concentration in the emerging industrial zone; This is the observation feature matrix that includes systematic sensor hardware measurement errors; A high-dimensional physical feature regression coefficient vector that measures the driving force of various features on actual PM2.5 concentration; The spatial autoregressive coefficient characterizing the intensity of interregional transport of actual PM2.5 concentration; This is the covariance matrix of the measurement error in the sensor hardware.

[0014] S1.3, utilizing and The residual extraction operator for observed structures is defined as follows: This involves performing mathematical operations to extract residuals from sensor hardware measurement errors. ; in, It is a residual sequence; It is the identity matrix. This is the physical spatial adjacency matrix of the target region.

[0015] S1.4, After obtaining the residual sequence, process the residual sequence... Centralized computation is performed to reconstruct the empirical distribution of the real PM2.5 concentration fluctuation residuals in memory for subsequent resampling algorithms. : ; in, The residual sequence after the centering operation. To take the mean multiplier.

[0016] Furthermore, step S1 also includes the following steps: S1.5, perform the full pseudo-data copy generation operation, while preserving the physical spatial adjacency matrix of the emerging industrial zone target domain. Under the premise of topological integrity, the inherent random fluctuations of atmospheric environmental sensing data are simulated, and multiple resampling loop logic is preset; in the first... In the independent cycles, from the empirical distribution Multiple completely independent random sampling with replacement is performed, thereby generating multiple groups of samples in the data memory, each with a sample size of [missing information]. Generate full residual vector The copy identifier .

[0017] S1.6, call Perform the generation and transformation of the pseudo-response variable based on the actual PM2.5 concentration: ; in, In the first A pseudo-response variable for PM2.5 concentration is generated during the resampling cycle; Based on this step, three complete pseudo-data copies of emerging industrial zones, retaining the original network dependency characteristics, were generated. , .

[0018] S1.7, for the first The loop of validating copies first obtains the baseline contamination parameter estimates on the two remaining target domain copies that do not contain mature region data. : ; in, Let be the set of spatial parameters to be estimated, and its specific expression is: , for Transpose of; For the first A copy of pseudo data from an emerging industrial zone; Subsequently, in the case of the first A mature regional source area Obtain cross-regional migration parameter estimates on the joint training set : ; in, For the first Sample size of a surrounding mature area The log-likelihood function is modified for mature regions.

[0019] S1.8, in The task of calculating the predicted loss is performed on the upper part of the system. Through multiple iterations and summaries, the predicted deterioration of the average actual PM2.5 concentration is obtained. Among them, the correction prediction loss for eliminating hardware measurement errors. The calculation formula is: ; in, In the first The corrected predicted loss, calculated on a verified copy, eliminates hardware measurement errors. The spatial parameter estimate to be verified; For the first A copy of pseudo data from an emerging industrial zone; This is the identity matrix for the corresponding dimension; These are the estimated spatial autoregressive coefficients; The generated PM2.5 concentration is a pseudo-response variable; This is the observation feature matrix that includes systematic sensor hardware measurement errors; This is the estimated high-dimensional regression coefficient vector; for The transpose of ; The squared second norm of a vector; This is the covariance matrix of the measurement error in the sensor hardware.

[0020] S1.9, Calculate the sample standard deviation of the average validation loss sequence across all iterations. An adaptive threshold screening criterion is constructed based on the sample standard deviation, and a set of effective information source domains for mature regions is defined. : ; in, The adjustment parameters are used to control the rigor of cross-regional data screening. This is the function for finding the maximum value.

[0021] Furthermore, step S2 specifically includes the following steps: S2.1, Find the high-dimensional physical feature regression coefficient vector that minimizes the joint modified likelihood function of the negative penalty. : ; in, This is an absolute value norm penalty term; For the first Regression coefficients of each physical characteristic; express The absolute value; The total dimension of the high-dimensional observation features.

[0022] S2.2, invoke the high-dimensional Bayesian information criterion operator to perform rigorous judgment of meteorological and emission characteristic dimensions in order to select the optimal penalty adjustment parameter. The high-dimensional Bayesian information criterion for: ; in, This represents the extreme value of the joint modified log-likelihood function under the corresponding penalty parameter. To apply a given penalty parameter Below, the estimated values ​​of the high-dimensional physical feature regression coefficient vector are obtained; To apply a given penalty parameter Below, the estimated values ​​of the spatial autoregressive coefficients are obtained; To apply a given penalty parameter Below, the estimated value of the random disturbance variance of the atmospheric environment system is obtained; To represent the total environmental monitoring sample size participating in the joint summation, It is a high-dimensional divergence factor. , This represents the number of feature variables in the candidate model whose estimated coefficients are not zero.

[0023] S2.3, Traverse the previously generated feature regression coefficient regularization path matrix, search and select ,based on The high-dimensional feature regression coefficient vector is truncated, and the indices of all feature variables whose estimated coefficients are not equal to zero are extracted, outputting the core meteorological and emission activity feature set. : ; in, For the index number of the physical characteristic variable, For the first A high-dimensional physical feature regression coefficient vector.

[0024] Furthermore, step S3 specifically includes the following steps: S3.1, based on the absolute values ​​of the characteristic regression coefficients obtained from the initial screening. Constructing a descending index sequence : ; Wherein, the feature index satisfies .

[0025] S3.2, based on descending index sequence Before truncation Several key features are used to generate a candidate model set, and a single candidate model. The mathematical expression is: .

[0026] S3.3, Based on concentration response variables Subtracting the mixed covariance matrix of sensor hardware measurement errors The systematic impact of this was investigated, and the baseline offset of the regression coefficients with analytical solutions was obtained. : ; in, For the joint observation feature matrix, for transpose, The actual PM2.5 concentration is used as the response variable. This is the joint sensor hardware measurement error covariance matrix corresponding to the global cross-regional hybrid dataset; By differentiating the error variance and setting it to zero, the corrected analytical solution for the variance of random disturbances in the atmospheric environment system is obtained: ; in, This represents the total number of samples in the global cross-regional mixed dataset.

[0027] Furthermore, step S3 also includes the following steps: S3.4, to prevent numerical compensation of spatial autoregressive coefficients due to parameter overfitting under the dual interference of small sample size and high-dimensional hardware measurement errors in emerging industrial areas, a hyperparameter is introduced into the log-likelihood objective function in the target domain correction set. The objective operator for optimizing the second-order norm penalty term for controlling the anchoring elasticity and the local difference term penalty is: ; in, This is a local residual correction term. The adjusted spatial autoregressive coefficients for the target emerging industrial zone Let Variance be the atmospheric perturbation variance in the target domain. This represents the difference vector of local regression coefficients corresponding to physical feature dimensions not included in the candidate model. This is the reference offset; Let be the square of the second norm of the local difference term vector.

[0028] S3.5, Calculation Spatial parameter estimator matrix : .

[0029] Further, step S4 includes the following steps: S4.1 Calculate the high-dimensional Bayesian information criterion value for each candidate model in the nested candidate model sequence. : ; in, It is a high-dimensional divergence factor. The maximum value of the modified log-likelihood of the emerging industrial zone in the target domain.

[0030] S4.2, calculate the negative half of the high-dimensional Bayesian information criterion values ​​as the penalized likelihood score for a specific candidate model. : .

[0031] S4.3, Extract To ensure the numerical stability of underlying floating-point operations, the smooth and continuous weights of each candidate model are calculated. : ; in, It is a natural exponential function. For the first The penalized likelihood score of each candidate model This is the summation index for iterating through all candidate models.

[0032] S4.4 uses a linear combination operator to perform a weighted summation operation on each single model parameter estimate in the candidate model pool and its corresponding smooth continuous weights: ; in, This is the spatial parameter estimate of the actual PM2.5 concentration in the optimal weighted region.

[0033] The present invention has the following beneficial effects: 1. Step S1 embeds the sensor hardware measurement error stripping mechanism into the cross-regional actual PM2.5 concentration physical diffusion prediction architecture.

[0034] To address the issues of pollution regression coefficient shrinkage to zero bias and impaired identification of cross-regional transport intensity in existing spatial autoregressive models when processing high-dimensional meteorological and emission characteristics containing sensor hardware measurement errors, this invention improves upon traditional parameter estimation techniques. By introducing a physical bias compensation term based on the product of multidimensional measurement error correlation and environmental monitoring sample size into the joint modified log-likelihood objective function of emerging industrial areas and surrounding mature areas, this mechanism logically neutralizes the residual expectation bias caused by variance inflation of hardware measurement errors, thus constructing an unbiased data foundation for dimensionality reduction of core meteorological and emission characteristics from multi-source sensing data.

[0035] 2. Step S2: Construct a residual resampling cross-regional source domain discrimination mechanism based on unbiased parameter initialization.

[0036] To address the shortcomings of traditional direct fusion of cross-regional multi-source sensing data, which is prone to negative data contagion, and traditional cross-validation, which easily disrupts the topological integrity of the physical adjacency matrix, this invention constructs a data screening pipeline. It utilizes a high-dimensional feature-penalized dimensionality reduction operator to extract unbiased observational structure residuals stripped of hardware measurement errors. While maintaining the integrity of the original spatial network dependencies of emerging industrial zones, it generates multiple full-scale pseudo-data copies of these emerging industrial zones using a spatial transformation inverse matrix operator. Coupled with an adaptive threshold segmentation logic built upon the predictive volatility of actual PM2.5 concentrations within these emerging industrial zones, this mechanism achieves robust screening and filtering of historical environmental sensing data streams from surrounding mature areas, even in physical environments characterized by limited high-precision monitoring samples and severe concurrent hardware measurement errors.

[0037] 3. Step S3 proposes a two-stage elastic re-estimation framework for candidate models to suppress compensatory distortion of spatial transmission parameters.

[0038] To address the challenges of spatial transmission intensity heterogeneity shifts and local feature failures caused by direct joint calculation of cross-regional multi-source sensing data, this invention designs a two-stage reassessment architecture that globally integrates surrounding mature regions and locally fine-tunes emerging industrial zones. In the first stage, global cross-regional joint calculations lock in the baseline shift of high-dimensional core meteorological and emission characteristic pollution regression coefficients. In the second stage, a second-order norm penalty term is added to the corrected log-likelihood objective function of the target emerging industrial zone. This allows for independent updates of local regression coefficient differences and the target region's spatial autoregressive coefficients using only limited monitoring samples from emerging industrial zones, thus preventing the risk of spatial transmission parameter variations caused by direct fitting of macro-level cross-regional data.

[0039] 4. Step S4: Construct an average inference strategy for a smooth diffusion prediction model based on model complexity penalty.

[0040] To address the shortcomings of single hard threshold selection mechanisms, such as physical diffusion structure jumps and model misspecifications when the signal-to-noise ratio of sensed data decreases, this invention constructs a continuous weight allocation process for high-dimensional sparse meteorological and emission feature systems. By incorporating a divergence factor related to the feature dimension into the high-dimensional Bayesian information criterion as a model complexity penalty mechanism, this overcomes the problem that traditional information criteria, due to insufficient penalty, easily select spurious noise features into the diffusion prediction model. This mechanism transforms the exponentially normalized penalized likelihood score into smooth continuous weights, concentrating the main posterior probabilities on a diffusion model set containing real atmospheric environmental physical driving signals, thus reducing the uncertainty of quantitative decision-making under hardware measurement error sensing environments. Attached Figure Description

[0041] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 A flowchart of a cross-regional PM2.5 concentration prediction method based on an improved spatial autoregressive model according to the present invention is shown.

[0042] Figure 2 A stacked comparison chart of the source domain discrimination accuracy of different methods under heavy measurement error conditions is shown.

[0043] Figure 3 The Monte Carlo convergence trajectory plots of the root mean square error of the regression coefficients for different methods as the number of independent experimental batches increases are shown.

[0044] Figure 4Box plots comparing the accuracy of root mean square error estimation of regression coefficients using different methods are shown.

[0045] Figure 5 Box plots comparing the accuracy of spatial autoregressive coefficient absolute deviation parameter estimation using different methods are shown.

[0046] Figure 6 The diagram shows a comparison of the Monte Carlo convergence trajectories of spatial autoregressive coefficient estimates from different methods as the number of independent experimental batches increases.

[0047] Figure 7 The diagram shows a comparison of Monte Carlo convergence trajectories for the absolute deviation density distribution of spatial parameters of different methods as the number of independent experimental batches increases.

[0048] Figure 8 Box plots comparing the net mean square prediction error performance of different methods on independent test sets.

[0049] Figure 9 The graph shows the mean and variance distribution of the average posterior selection probability of this invention across different model sizes.

[0050] Figure 10 This is a heatmap showing how the average weight of the smoothed model changes with the experimental batch and the number of variables included in the model. Detailed Implementation

[0051] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0052] like Figure 1 The method for predicting cross-regional PM2.5 concentration based on an improved spatial autoregressive model, as shown, specifically includes the following steps: S1, based on residual resampling technology, generates a set of mature regional information source domains with real reference value for the evolution of atmospheric pollution; S2, perform preliminary screening of high-dimensional meteorological and emission characteristics to obtain a set of core meteorological and emission activity characteristics after removing interference from sensor hardware measurement errors; S3: Based on the ranking of pollution driving importance represented by the absolute value of the regression coefficients of physical characteristics, construct candidate models and output the spatial parameter estimation matrix corresponding to each candidate model. S4 utilizes a negative half-dimensional high-dimensional Bayesian information criterion diffusion model for scoring, outputting the optimal weighted regional spatial parameter estimate of actual PM2.5 concentration to external environmental protection departments or public health institutions. This provides quantitative decision support for cross-regional physical diffusion prediction of actual PM2.5 concentration in emerging industrial areas.

[0053] This invention proposes a high-dimensional spatial autoregressive model transfer sensing method for predicting actual PM2.5 concentrations in cities, taking into account sensor hardware measurement errors. This method aims to overcome the dual underlying physical interference of systematic sensor hardware measurement errors prevalent in high-dimensional environmental sensing data and the limited high-precision monitoring samples in emerging industrial areas. The overall optimization objective of the system is to construct an end-to-end atmospheric pollution data processing pipeline to output a weighted estimate of the spatial parameters of actual PM2.5 concentrations in the region with robust generalization ability and asymptotic optimality to an external decision-making terminal.

[0054] Based on a spatial autoregressive model, the actual PM2.5 concentration response variable vector for each grid in the emerging industrial zone. for: ; in, It is a physical adjacency matrix. This is a matrix of meteorological and emission physical characteristics. This represents the random disturbance term of the atmospheric environment system.

[0055] Specifically, in the scenario of cross-regional physical diffusion prediction of PM2.5 concentration in emerging industrial zones, due to the limited historical high-precision monitoring records of these zones, relying solely on the target area's own sensing data to independently construct a physical diffusion prediction model leads to significant variance inflation in the prediction, rendering the prediction results meaningless for practical guidance in government environmental protection planning and heavy pollution early warning. Therefore, it is necessary to leverage mature surrounding areas with abundant historical sensing data as information sources for joint prediction.

[0056] However, blindly integrating monitoring data from mature areas with severe local topographical obstruction or altered underlying heavy industrial structures can trigger a negative data contagion effect. Existing traditional feature selection frameworks, when faced with physical feature data containing systematic hardware measurement errors such as aging low-cost sensors, suffer from compromised source domain separability assumptions, making it difficult to accurately identify noisy cities with mismatched airflow diffusion mechanisms. To address these technical shortcomings, this invention removes the zero-bound bias in pollution regression coefficients caused by sensor hardware measurement errors. While maintaining the integrity of the target area's physical spatial topology, it accurately identifies and retains mature areas consistent with the actual atmospheric pollution evolution patterns of emerging industrial zones, thus providing a high-quality cross-regional data foundation for subsequent joint model training.

[0057] Step S1 specifically includes the following steps: S1.1, receives the response variable vector of the actual PM2.5 concentration in the emerging industrial zone. Observation feature matrix including systematic sensor hardware measurement errors Physical adjacency matrix of the target area in the physical spatial topology of each monitoring grid And the sensor hardware measurement error covariance matrix that is pre-acquired and characterizes the correlation of multidimensional measurement errors. , as input data.

[0058] S1.2, In order to eliminate the bias of regression coefficients shrinking to zero due to hardware measurement errors from the data level, the negative modified log-likelihood function with absolute value norm penalty term is called for initial data processing: ; ; in, This is the initial high-dimensional physical feature regression coefficient vector; These are the initial spatial autoregressive coefficients; This is an estimate of the initial random disturbance variance of the atmospheric environment system; To find the set of independent variables that minimizes the objective function; The target domain is modified for the log-likelihood function; Initial adjustment parameters for controlling the sparsity of the high-dimensional feature space; This is the absolute value norm penalty term for the regression coefficient vector; This refers to the sample size for environmental monitoring in the target area. This is the spatial transformation matrix of the target region in the corresponding dimension; The variance of random disturbances in the atmospheric environment system; This represents the response variable vector of the actual PM2.5 concentration in the emerging industrial zone; This is the observation feature matrix that includes systematic sensor hardware measurement errors; A high-dimensional physical feature regression coefficient vector that measures the driving force of various features on actual PM2.5 concentration; The spatial autoregressive coefficient characterizing the intensity of interregional transport of actual PM2.5 concentration; This is the covariance matrix of the measurement error in the sensor hardware.

[0059] S1.3, utilizing and The residual extraction operator for observed structures is defined as follows: This involves performing mathematical operations to extract residuals from sensor hardware measurement errors. ; in, It is a residual sequence; It is an identity matrix, the dimensions of which are the same as the environmental monitoring sample size of the target area. Maintain consistency, i.e., dimension is The identity matrix, This is the physical spatial adjacency matrix of the target region.

[0060] S1.4, After obtaining the residual sequence, process the residual sequence... Centralized computation is performed to reconstruct the empirical distribution of the real PM2.5 concentration fluctuation residuals in memory for subsequent resampling algorithms. : ; in, The residual sequence after the centering operation. To take the mean multiplier.

[0061] Specifically, step S1 also includes the following steps: S1.5, perform the full pseudo-data copy generation operation, while preserving the physical spatial adjacency matrix of the emerging industrial zone target domain. Under the premise of topological integrity, the inherent random fluctuations of atmospheric environmental sensing data are simulated, and multiple resampling loop logic is preset; in the first... In the independent cycles, from the empirical distribution Multiple completely independent random sampling with replacement is performed, thereby generating multiple groups of samples in the data memory, each with a sample size of [missing information]. Generate full residual vector The copy identifier This embodiment specifically involves three sets of data.

[0062] S1.6, call Perform the generation and transformation of the pseudo-response variable based on the actual PM2.5 concentration: ; in, In the first In the resampling cycle, a pseudo-response variable for PM2.5 concentration is generated.

[0063] Three complete pseudo-data copies of emerging industrial zones, generated based on this operator, retain the original network dependency characteristics. Its underlying data structure is defined as .

[0064] S1.7, to avoid the extreme computational complexity and memory overflow risk caused by constructing a super-large cross-regional block diagonal spatial weight matrix, the likelihood summation data processing strategy is invoked. For the first... The loop of validating copies first obtains the baseline contamination parameter estimates on the two remaining target domain copies that do not contain mature region data. : ; in, Let be the set of spatial parameters to be estimated, and its specific expression is: , for Transpose of; For the first A copy of pseudo data from an emerging industrial zone; Subsequently, in the case of the first A mature regional source area Obtain cross-regional migration parameter estimates on the joint training set The joint optimization objective function of this estimator, based on the sum of the corrected log-likelihoods of the two target domain replicas, additionally includes the corrected log-likelihood function of the mature region and the natural weighting coefficient representing the ratio of sample sizes between the two locations. The product of . This operator is precisely equivalent in the underlying algorithmic logic to the physical operation of directly stitching together cross-regional environmental sensing data, and its formula expands to: ; in, For the first The sample size of each surrounding mature region (source domain), The log-likelihood function is modified for mature regions.

[0065] S1.8, in The task of calculating the predicted loss is performed on the upper part of the system. Through multiple iterations and summaries, the predicted deterioration of the average actual PM2.5 concentration is obtained. Among them, the correction prediction loss for eliminating hardware measurement errors. The calculation formula is: ; in, In the first The corrected predicted loss, calculated on a verified copy, eliminates hardware measurement errors. The spatial parameter estimate to be verified; For the first A copy of pseudo data from an emerging industrial zone; This is the identity matrix for the corresponding dimension; These are the estimated spatial autoregressive coefficients; The generated PM2.5 concentration is a pseudo-response variable; This is the observation feature matrix that includes systematic sensor hardware measurement errors; This is the estimated high-dimensional regression coefficient vector; for The transpose of ; The squared second norm of a vector; The sensor hardware measurement error covariance matrix; Based on this formula, the baseline prediction loss and migration prediction loss on the validation copy are output to memory respectively.

[0066] Through multiple loop aggregation calculations, the predicted deterioration of the average actual PM2.5 concentration for a specific candidate mature region across all resampling cycles and three validation copies is output to memory. That is, the global average arithmetic result of the difference between all migration prediction losses and the baseline prediction losses.

[0067] After obtaining two sets of high-dimensional spatial parameters—the baseline and the cross-regional migration parameters—it was mandated that the replicas be verified only in the retained independent emerging industrial zones. The task of calculating the predicted loss is performed on the upper part. It is strictly forbidden to introduce any mature regional source domain data for verification and evaluation in this stage, so as to prevent the verification results from being dominated by mature regions with large data volumes and thus obscuring the true performance of the physical diffusion prediction model in emerging industrial areas.

[0068] S1.9, Calculate the sample standard deviation of the average validation loss sequence across all iterations. This allows for the precise quantification of the inherent physical prediction volatility of environmental sensing data in emerging industrial zones under current sensor hardware measurement error levels. An adaptive threshold screening criterion is constructed based on the sample standard deviation to define the set of effective information source domains for mature areas. : ; in, To control the stringency of cross-regional data screening, an adjustment parameter of 0.01 was used as a physical safety net to prevent the standard deviation from reaching its limit due to the limited number of actual high-precision monitoring samples. After rigorous filtering by this comparison operator, the first-stage main module successfully eliminated noisy cities with variations in airflow diffusion mechanisms, ultimately generating a set of surrounding effective information source domains whose predicted deterioration level had not significantly increased. This dataset serves as high-quality physical input, which is then passed downstream to the main module for preliminary screening of high-dimensional meteorological and emission characteristics. The average actual PM2.5 concentration is used to predict the degree of deterioration. This is the function for finding the maximum value.

[0069] Specifically, step S2 includes the following steps: S2.1, in order to fully utilize cross-regional, multi-source, consistent historical environmental monitoring data for identifying core meteorological and emission physical characteristics, the cross-regional data physical splicing operation is abandoned. Instead, based on the statistical assumption of physical spatial independence, the modified log-likelihood function of the target domain of emerging industrial areas is used. With all mature regions' effective information source domains Modified log-likelihood function Perform a joint summation operation. Append a factor to the negative of the likelihood sum to reduce the spurious feature dimension. Absolute value norm penalty term Find the high-dimensional physical feature regression coefficient vector that minimizes the joint modified likelihood function of the negative penalty. : ; in, This is an absolute value norm penalty term; For the first Regression coefficients of each physical characteristic; express The absolute value; The total dimension of the high-dimensional observation features.

[0070] S2.2 utilizes high-dimensional information criteria for data filtering and parameter truncation of the data stream. After generating the regularized path matrix of physical feature regression coefficients, to avoid the huge computational overhead and cross-regional data reuse bias caused by traditional cross-validation in high-dimensional meteorological feature spaces, a high-dimensional Bayesian information criterion operator is invoked to perform rigorous truncation of meteorological and emission feature dimensions to select the optimal penalty adjustment parameter. The high-dimensional Bayesian information criterion for: ; in, This represents the extreme value of the joint modified log-likelihood function under the corresponding penalty parameter. To apply a given penalty parameter Below, the estimated values ​​of the high-dimensional physical feature regression coefficient vector are obtained; To apply a given penalty parameter Below, the estimated values ​​of the spatial autoregressive coefficients are obtained; To apply a given penalty parameter Below, the estimated value of the random disturbance variance of the atmospheric environment system is obtained; To represent the total environmental monitoring sample size participating in the joint summation, It is a high-dimensional divergence factor. , This represents the number of feature variables with non-zero estimated coefficients in the candidate model, used to quantify the complexity of the current model. The underlying physical meaning of the divergence factor is to impose an additional penalty on the model complexity for the perception data scenario of emerging industrial areas with extremely high dimensionality and including sensor hardware measurement errors, thus preventing the algorithmic defect of the standard information criterion from selecting false noise features with no real environmental driving significance into the physical diffusion prediction model due to insufficient penalty.

[0071] S2.3, Traverse the previously generated feature regression coefficient regularization path matrix, search and select ,based on The high-dimensional feature regression coefficient vector is truncated, and the indices of all feature variables whose estimated coefficients are not equal to zero are extracted, outputting the core meteorological and emission activity feature set. : ; in, For the index number of the physical characteristic variable, For the first A vector of high-dimensional physical feature regression coefficients; This core set of meteorological and emission activity characteristics As the key dimensionality reduction feature index that truly drives the actual PM2.5 concentration in the target area after eliminating the influence of sensor hardware measurement noise, it is uniformly packaged and directly passed downstream as a high-quality, pure physical input to the main module of nested candidate model construction and two-stage transfer re-estimation, thus completely completing the closed loop of high-dimensional feature dimensionality reduction data flow in stage two.

[0072] Specifically, step S3 includes the following steps: S3.1, to avoid the extreme computational complexity caused by traversing all feature subset combinations, the absolute values ​​of the feature regression coefficients obtained from the initial screening are used... Constructing a descending index sequence : ; Wherein, the feature index satisfies ; S3.2, based on descending index sequence Before truncation Several key features are used to generate a candidate model set, and a single candidate model. The mathematical expression is: ; The set of candidate model sequences is passed as structured index data to the subsequent dual-track re-estimation subunit.

[0073] S3.3 involves physically concatenating the observation feature matrix of the emerging industrial zone target domain, which includes hardware measurement errors, with the environmental feature data of the mature area's effective information source domain along the sample dimension to construct a joint observation feature matrix. Combined with the actual PM2.5 concentration response variable vector The physical space adjacency matrix of each region is constructed as a cross-region block diagonal physical space adjacency matrix. This would cut off the spatial spillover of man-made air pollution between the surrounding mature areas and the emerging industrial zones from the underlying physical structure, thus preventing any actual physical or meteorological significance.

[0074] Spatial autoregressive coefficients based on the actual inter-regional transport intensity of PM2.5 concentration Using the hybrid space transformation matrix Calculate the response variable of the transformed mixed actual PM2.5 concentration Based on concentration response variables Subtracting the mixed covariance matrix of sensor hardware measurement errors The systematic impact of this was investigated, and the baseline offset of the regression coefficients with analytical solutions was obtained. : ; in, For the joint observation feature matrix, for transpose, The actual PM2.5 concentration is used as the response variable. This is the joint sensor hardware measurement error covariance matrix corresponding to the global cross-regional hybrid dataset; Similarly, by differentiating the error variance and setting it to zero, we obtain the corrected analytical solution for the variance of random disturbances in the atmospheric environment system: ; in, This represents the total number of samples in the global cross-regional mixed dataset.

[0075] Specifically, step S3 also includes the following steps: S3.4, the analytical solution obtained from S3.3 and Substituting back into the joint modified log-likelihood function to eliminate the regression coefficients and variance variables yields only information about... Modified profile likelihood function Subsequently, an extreme point is searched within the one-dimensional grid to maximize the likelihood function of the modified profile. Finally, the baseline offset of the specific model regression coefficients, which incorporates the variance reduction advantage of large samples from mature regions, is extracted and output to memory. .

[0076] Receive the above reference offset This is then transformed into a fixed constant input, utilizing only the observation feature matrix of the emerging industrial zone, which is constrained by hardware measurement errors. The response variable vector of actual PM2.5 concentration Perform local parameter reestimation. To prevent numerical compensation of spatial autoregressive coefficients due to parameter overfitting under the dual interference of small sample size and high-dimensional hardware measurement errors in emerging industrial areas, a hyperparameter is introduced into the log-likelihood objective function in the target domain correction set. The objective operator for optimizing the second-order norm penalty term for controlling the anchoring elasticity and the local difference term penalty is: ; in, This is a local residual correction term. The adjusted spatial autoregressive coefficients for the target emerging industrial zone Let Variance be the atmospheric perturbation variance in the target domain. This represents the difference vector of local regression coefficients corresponding to physical feature dimensions not included in the candidate model, which is controlled to be 0 in the calculation. As the reference offset, Let be the square of the second norm of the local difference term vector.

[0077] Calling the alternating iteration operator, the autoregressive coefficients in the fixed space The local difference term is updated using the generalized least squares method with second-order norm penalty. Given this difference term, a one-dimensional bounded search is performed on the modified profile likelihood function to update the spatial autoregressive coefficients of the emerging industrial zone in the target domain. After convergence, this alternating iterative operator outputs a local residual correction term. Spatial autoregressive coefficients of the target emerging industrial zone after fine-tuning and the target domain atmospheric perturbation variance .

[0078] S3.5, adjust the reference offset. With local residual correction term Perform a high-dimensional vector addition operation and combine it with the re-estimated spatial autoregressive coefficients. and the target domain atmospheric perturbation variance Assemble together and calculate Spatial parameter estimator matrix : ; The entire candidate model sequence is automatically traversed, and the above basic estimation and local fine-tuning processes are performed iteratively. Finally, a complete set of spatial parameter estimates for all candidate models is output to the outside world. This complete set of parameter matrices will be directly passed to step S4 as high-quality physical input.

[0079] Specifically, step S4 includes the following steps: S4.1, Receive the output of step S3 , and As the core physical input, to overcome the risks of physical diffusion structure jumps and model misspecification that are easily generated by the traditional single hard threshold selection strategy under sensor hardware measurement error environments, a high-dimensional Bayesian information criterion is calculated for each candidate model in the nested candidate model sequence. : ; in, It is a high-dimensional divergence factor. The maximum modified log-likelihood of the emerging industrial zone in the target domain; S4.2, utilize the maximum modified log-likelihood of emerging industrial zones in the target domain. Evaluate the fit of the underlying environmental perception data and forcibly superimpose it using a high-dimensional divergence factor. The number of core meteorological and emission physical features included in the candidate models Natural logarithm of the target domain environmental monitoring sample size The model complexity penalty term is jointly constituted. The negative half of the high-dimensional Bayesian information criterion value is calculated as the penalized likelihood score for a specific candidate model. : ; After the calculation is completed, the performance score data vector corresponding to all candidate models is output.

[0080] S4.3, to achieve a smooth parameter transition among a group of optimal diffusion prediction models with similar performance across different feature dimensions, calls the exponential normalization function operator with a maximum value shift mechanism. Extraction To ensure the numerical stability of underlying floating-point operations, the smooth and continuous weights of each candidate model are calculated. : ; in, It is a natural exponential function. For the first The penalized likelihood score of each candidate model The summation index is used to iterate through all candidate models. After rigorous processing by this nonlinear transformation operator, a smooth, continuous weight vector with a value range between zero and one is output for all candidate models.

[0081] S4.4 uses a linear combination operator to perform a weighted summation operation on each single model parameter estimate in the candidate model pool and its corresponding smooth continuous weights: ; in, This is the spatial parameter estimate of the actual PM2.5 concentration in the optimal weighted region.

[0082] After calculation, the optimal weighted regional spatial parameter estimate of actual PM2.5 concentration is officially output to the operational terminals of external environmental protection departments or public health institutions. This estimate combines the stability of parameters from large samples in surrounding mature areas with the adaptability to the local characteristics of emerging industrial zones. This statistically robust estimator will serve as the underlying data foundation, directly supporting the core physical scenario of cross-regional physical diffusion prediction of actual PM2.5 concentrations in emerging industrial areas. Thus, the entire internal data flow of the core algorithm pipeline of this invention, overcoming the dual limitations of sensor hardware measurement errors and the limited high-precision monitoring samples, is declared a tightly closed loop.

[0083] like Figure 2 As shown, regarding the signal-to-noise ratio In a heavy sensor hardware measurement error environment, this invention performed a Monte Carlo simulation experiment. Experimental statistics show that the traditional target-domain-only naive Lasso method suffers from increased false positive rates due to spurious correlation interference caused by measurement errors. In contrast, the method of this invention precisely introduces the sensor hardware measurement error covariance matrix into the underlying joint modified log-likelihood function. The constructed compensation item maintains Under the premise of a high true positive rate, the interference of noise variables was effectively suppressed. Meanwhile, the mature regional information source identification module based on modified spatial residual resampling successfully eliminated noise source domains subjected to reverse random offsets in each independent experiment, confirming the statistical robustness of the adaptive threshold operator in blocking the spread of negative data.

[0084] like Figure 3 , Figure 4 and Figure 5 As shown, the root mean square error, which measures the accuracy of estimating the regression coefficient β of high-dimensional physical characteristics, is... In terms of metrics, the experiment revealed objective comparative results. Due to the dual interference of limited high-precision monitoring samples and sensor hardware measurement errors, the regression coefficient β of the naive Lasso method in the target domain only... Da The naive hard thresholding method without error correction also achieves... The method of this invention identifies and introduces a set of effective information source domains. After that, Achieve a significant decrease, converging to The experimental data demonstrate that the global fundamental estimation mechanism of this invention successfully captures the variance reduction benefit brought about by large cross-regional samples while eliminating the decay bias of the pollution regression coefficient.

[0085] like Figure 6 and Figure 7 As shown, the spatial autoregressive coefficient is a core indicator characterizing the intensity of cross-regional transport of actual PM2.5 concentration. In terms of estimation, the method of this invention exhibits stable parameter protection performance. Monte Carlo convergence trajectory and absolute deviation density distribution data show that when the traditional naive hard threshold migration method directly merges multi-source environmental sensing data, it is prone to compensatory shifts in spatial parameters of the target region due to the heterogeneity of airflow diffusion mechanisms between the source and target domains. In contrast, the method of this invention invokes a two-stage local elastic deviation correction mechanism. After locking the global regression coefficient baseline shift, it only uses data from the target emerging industrial area to adjust the spatial autoregressive coefficients. Independent reestimation is performed. Experiments confirm that this mechanism enables the spatial autoregressive coefficient estimation trajectory of the method of this invention to smoothly converge to the true parameters. Nearby, the density distribution of absolute deviation is concentrated, eliminating the risk of distortion of local spatial physical diffusion characteristics caused by direct splicing of cross-regional data.

[0086] like Figure 8 As shown, in the independent test set evaluation phase, the net mean square prediction error, used to measure the resilience to real-world weather and emission characteristics, was calculated. The average net mean square prediction error of the traditional target-domain-only naive Lasso method is... The naive hard thresholding method is The method of this invention, by leveraging the underlying high-dimensional Bayesian information criterion and the average weighting mechanism of the exponentially normalized smoothing model, significantly reduces the average net mean square prediction error to [missing value]. .

[0087] like Figure 9 and Figure 10 As shown, the heatmap of the average weights of the smoothed model and the distribution of the posterior selection probability reveal the underlying mechanism by which the prediction error of the method in this invention is significantly reduced. Because the initial high-dimensional variable screening module is extremely accurate and stable, the continuous weight allocation operator of this invention can precisely concentrate most of the posterior selection probabilities on the true model size that includes the correct number of variables. This weight allocation mechanism, dominated by the correct model size, not only preserves the true physical signal to the greatest extent but also successfully dilutes the overfitting bias caused by occasional noise variables.

[0088] In summary, the experimental data objectively demonstrates that in low signal-to-noise ratio environmental sensing data environments, traditional single hard threshold selection strategies are prone to structural shifts in diffusion mechanisms and model misspecifications between different samples. The method of this invention can effectively smooth out the uncertainty of model selection through smooth continuous weight allocation, thereby providing a quantitative decision-making foundation with statistical confidence in the task of predicting cross-regional physical diffusion of actual urban PM2.5 concentrations, where sensor hardware measurement errors and high-precision monitoring samples are limited.

[0089] Of course, the above description is not intended to limit the present invention, and the present invention is not limited to the examples given above. Any changes, modifications, additions or substitutions made by those skilled in the art within the scope of the present invention should also fall within the protection scope of the present invention.

Claims

1. A method for predicting cross-regional PM2.5 concentration based on an improved spatial autoregressive model, characterized in that, Specifically, the steps include the following: S1, based on residual resampling technology, generates a set of mature regional information source domains with real reference value for the evolution of atmospheric pollution; S2, perform preliminary screening of high-dimensional meteorological and emission characteristics to obtain a set of core meteorological and emission activity characteristics after removing interference from sensor hardware measurement errors; S3: Based on the ranking of pollution driving importance represented by the absolute value of the regression coefficients of physical characteristics, construct candidate models and output the spatial parameter estimation matrix corresponding to each candidate model. S4 uses the negative half-dimensional high-dimensional Bayesian information criterion to score candidate models and outputs the optimal weighted regional actual PM2.5 concentration spatial parameter estimate to the business terminal of external environmental protection departments or public health institutions; Step S1 specifically includes the following steps: S1.1, receives the response variable vector of the actual PM2.5 concentration in the emerging industrial zone. Observation feature matrix including systematic sensor hardware measurement errors Physical adjacency matrix of the target area in the physical spatial topology of each monitoring grid And the sensor hardware measurement error covariance matrix that is pre-acquired and characterizes the correlation of multidimensional measurement errors. , as input data; S1.2, In order to eliminate the bias of regression coefficients shrinking to zero due to hardware measurement errors from the data level, the negative modified log-likelihood function with absolute value norm penalty term is called for initial data processing: ; ; in, This is the initial high-dimensional physical feature regression coefficient vector; These are the initial spatial autoregressive coefficients; This is an estimate of the initial random disturbance variance of the atmospheric environment system; To find the set of independent variables that minimizes the objective function; The target domain is modified for the log-likelihood function; Initial adjustment parameters for controlling the sparsity of the high-dimensional feature space; This is the absolute value norm penalty term for the regression coefficient vector; This refers to the sample size for environmental monitoring in the target area. This is the spatial transformation matrix of the target region in the corresponding dimension; The variance of random disturbances in the atmospheric environment system; This represents the response variable vector of the actual PM2.5 concentration in the emerging industrial zone; This is the observation feature matrix that includes systematic sensor hardware measurement errors; A high-dimensional physical feature regression coefficient vector that measures the driving force of various features on actual PM2.5 concentration; The spatial autoregressive coefficient characterizing the intensity of interregional transport of actual PM2.5 concentration; The sensor hardware measurement error covariance matrix; S1.3, utilizing and The residual extraction operator for observed structures is defined as follows: This involves performing mathematical operations to extract residuals from sensor hardware measurement errors. ; in, It is a residual sequence; It is the identity matrix. The physical adjacency matrix of the target region; S1.4, After obtaining the residual sequence, process the residual sequence... Centralized computation is performed to reconstruct the empirical distribution of the real PM2.5 concentration fluctuation residuals in memory for subsequent resampling algorithms. : ; in, The residual sequence after the centering operation. To take the mean multiplier.

2. The method for predicting cross-regional PM2.5 concentration based on an improved spatial autoregressive model according to claim 1, characterized in that, Step S1 also includes the following steps: S1.5, perform the full pseudo-data copy generation operation, while preserving the physical spatial adjacency matrix of the emerging industrial zone target domain. Under the premise of topological integrity, the inherent random fluctuations of atmospheric environmental sensing data are simulated, and multiple resampling loop logic is preset; in the first... In the independent cycles, from the empirical distribution Multiple completely independent random sampling with replacement is performed, thereby generating multiple groups of samples in the data memory, each with a sample size of [missing information]. Generate full residual vector The copy identifier ; S1.6, call Perform the generation and transformation of the pseudo-response variable based on the actual PM2.5 concentration: ; in, In the first A pseudo-response variable for PM2.5 concentration is generated during the resampling cycle; Based on this step, three complete pseudo-data copies of emerging industrial zones, retaining the original network dependency characteristics, were generated. , ; S1.7, for the first The loop of validating copies first obtains the baseline contamination parameter estimates on the two remaining target domain copies that do not contain mature region data. : ; in, Let be the set of spatial parameters to be estimated, and its specific expression is: , for Transpose of; For the first A copy of pseudo data from an emerging industrial zone; Subsequently, in the case of the first A mature regional source area Obtain cross-regional migration parameter estimates on the joint training set : ; in, For the first Sample size of a surrounding mature area Adjust the log-likelihood function for the mature region; S1.8, in The task of calculating the predicted loss is performed on the upper part of the system. Through multiple iterations and summaries, the predicted deterioration of the average actual PM2.5 concentration is obtained. Among them, the correction prediction loss for eliminating hardware measurement errors. The calculation formula is: ; in, In the first The corrected predicted loss, calculated on a verified copy, eliminates hardware measurement errors. The spatial parameter estimate to be verified; For the first A copy of pseudo data from an emerging industrial zone; This is the identity matrix for the corresponding dimension; These are the estimated spatial autoregressive coefficients; The generated PM2.5 concentration is a pseudo-response variable; This is the observation feature matrix that includes systematic sensor hardware measurement errors; This is the estimated high-dimensional regression coefficient vector; for The transpose of ; The squared second norm of a vector; The sensor hardware measurement error covariance matrix; S1.9, Calculate the sample standard deviation of the average validation loss sequence across all iterations. An adaptive threshold screening criterion is constructed based on the sample standard deviation, and a set of effective information source domains for mature regions is defined. : ; in, The adjustment parameters are used to control the rigor of cross-regional data screening. This is the function for finding the maximum value.

3. The method for predicting cross-regional PM2.5 concentration based on an improved spatial autoregressive model according to claim 1, characterized in that, Step S2 specifically includes the following steps: S2.1, Find the high-dimensional physical feature regression coefficient vector that minimizes the joint modified likelihood function of the negative penalty. : ; in, This is an absolute value norm penalty term; For the first Regression coefficients of each physical characteristic; express The absolute value; The total dimension of the high-dimensional observation features; S2.2, invoke the high-dimensional Bayesian information criterion operator to perform rigorous judgment of meteorological and emission characteristic dimensions in order to select the optimal penalty adjustment parameter. The high-dimensional Bayesian information criterion for: ; in, This represents the extreme value of the joint modified log-likelihood function under the corresponding penalty parameter. To apply a given penalty parameter Below, the estimated values ​​of the high-dimensional physical feature regression coefficient vector are obtained; To apply a given penalty parameter Below, the estimated values ​​of the spatial autoregressive coefficients are obtained; To apply a given penalty parameter Below, the estimated value of the random disturbance variance of the atmospheric environment system is obtained; To represent the total environmental monitoring sample size participating in the joint summation, It is a high-dimensional divergence factor. , This represents the number of feature variables in the candidate model whose estimated coefficients are not zero. S2.3, Traverse the previously generated feature regression coefficient regularization path matrix, search and select ,based on The high-dimensional feature regression coefficient vector is truncated, and the indices of all feature variables whose estimated coefficients are not equal to zero are extracted, outputting the core meteorological and emission activity feature set. : ; in, For the index number of the physical characteristic variable, For the first A high-dimensional physical feature regression coefficient vector.

4. The method for predicting cross-regional PM2.5 concentration based on an improved spatial autoregressive model according to claim 1, characterized in that, Step S3 specifically includes the following steps: S3.1, based on the absolute values ​​of the characteristic regression coefficients obtained from the initial screening. Constructing a descending index sequence : ; Wherein, the feature index satisfies ; S3.2, based on descending index sequence Before truncation Several key features are used to generate a candidate model set, and a single candidate model. The mathematical expression is: ; S3.3, Based on concentration response variables Subtracting the mixed covariance matrix of sensor hardware measurement errors The systematic impact of this was investigated, and the baseline offset of the regression coefficients with analytical solutions was obtained. : ; in, For the joint observation feature matrix, for transpose, The actual PM2.5 concentration is used as the response variable. This is the joint sensor hardware measurement error covariance matrix corresponding to the global cross-regional hybrid dataset; By differentiating the error variance and setting it to zero, the corrected analytical solution for the variance of random disturbances in the atmospheric environment system is obtained: ; in, This represents the total number of samples in the global cross-regional mixed dataset.

5. The method for predicting cross-regional PM2.5 concentration based on an improved spatial autoregressive model according to claim 4, characterized in that, Step S3 also includes the following steps: S3.4, to prevent numerical compensation of spatial autoregressive coefficients due to parameter overfitting under the dual interference of small sample size and high-dimensional hardware measurement errors in emerging industrial areas, a hyperparameter is introduced into the log-likelihood objective function in the target domain correction set. The objective operator for optimizing the second-order norm penalty term for controlling the anchoring elasticity and the local difference term penalty is: ; in, This is a local residual correction term. The adjusted spatial autoregressive coefficients for the target emerging industrial zone. Let Variance be the atmospheric perturbation variance in the target domain. This represents the difference vector of local regression coefficients corresponding to physical feature dimensions not included in the candidate model. This is the reference offset; The square of the second norm of the local difference term vector; S3.5, Calculation Spatial parameter estimator matrix : 。 6. The method for predicting cross-regional PM2.5 concentration based on an improved spatial autoregressive model according to claim 1, characterized in that, Step S4 includes the following steps: S4.1 Calculate the high-dimensional Bayesian information criterion value for each candidate model in the nested candidate model sequence. : ; in, It is a high-dimensional divergence factor. The maximum modified log-likelihood of the emerging industrial zone in the target domain; S4.2, calculate the negative half of the high-dimensional Bayesian information criterion values ​​as the penalized likelihood score for a specific candidate model. : ; S4.3, Extract To ensure the numerical stability of underlying floating-point operations, the smooth and continuous weights of each candidate model are calculated. : ; in, It is a natural exponential function. For the first The penalized likelihood score of each candidate model The summation index for traversing all candidate models; S4.4 uses a linear combination operator to perform a weighted summation operation on each single model parameter estimate in the candidate model pool and its corresponding smooth continuous weights: ; in, This is the spatial parameter estimate of the actual PM2.5 concentration in the optimal weighted region.

Citation Information

Patent Citations

  • Aerosol prediction method and system based on artificial intelligence

    CN119964669A

  • Combination model pollutant concentration prediction method and system based on multiple strategies, storage medium and electronic equipment

    CN120564881A