Heterogeneous air quality data fusion method based on sparse matrix decomposition

Through the methods of sparse matrix decomposition and attention mechanism optimization, the problems of heterogeneity and missing data in air quality data fusion are solved, and high-precision and comprehensive air quality assessment is achieved.

CN120671084AInactive Publication Date: 2025-09-19HEFEI OUWO ENVIRONMENTAL PROTECTION TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510800713.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing air quality data fusion technology cannot fully explore the potential information of the data when processing large-scale, high-dimensional, and heterogeneous data, and lacks effective weighting of the differences in temporal and spatial characteristics between different data sources, resulting in errors and inconsistencies in the fusion results.

Method used

A method based on sparse matrix decomposition is adopted, combined with improved singular value decomposition and low-rank matrix completion technology. Heterogeneous data are processed through time synchronization correction and spatial alignment. The attention mechanism is introduced to optimize feature weighting, and multi-dimensional error measurement and consistency verification are performed.

Benefits of technology

It improves the accuracy and consistency of data fusion, effectively captures the spatiotemporal patterns of data, reduces errors, and improves the accuracy and credibility of air quality forecasts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671084A_ABST
    Figure CN120671084A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous air quality data fusion method based on sparse matrix decomposition. The method comprises the following steps: S1, collecting air quality data; s2, preprocessing the collected air quality data, and constructing a preliminary fusion matrix; s3, marking missing data in the preliminary fusion matrix, and constructing a sparse matrix; s4, carrying out matrix decomposition by adopting a mode of combining an improved singular value decomposition method and a matrix completion method based on low-rank representation, filling missing data, and carrying out feature weighted optimization on a matrix decomposition result by utilizing an attention mechanism; and S5, error measurement and consistency check are carried out. According to the method, sparse matrix decomposition and low-rank completion technologies are adopted, the fusion process of the air quality data is optimized in combination with an attention mechanism, and high data precision, high consistency and improvement of prediction accuracy are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data fusion, and in particular to a heterogeneous air quality data fusion method based on sparse matrix decomposition. Background Art

[0002] As global environmental problems become increasingly severe, air pollution has become a significant threat to human health and sustainable development. To address air pollution, many countries and regions have established multi-tiered air quality monitoring systems. These systems collect and acquire air quality data through a variety of means, including ground-based monitoring stations, satellite remote sensing, meteorological data, and social perception data. The data provided by these different data sources varies significantly in form and quality. Effectively integrating this heterogeneous data to build a high-precision, comprehensive air quality assessment system has become a major challenge in the field of environmental monitoring.

[0003] Existing air quality data fusion technologies primarily include methods based on statistical analysis, machine learning, and collaborative filtering using matrix factorization. Traditional statistical analysis-based fusion methods often rely on simplifying assumptions about the data, employing techniques such as weighted averaging or regression analysis to fuse data from different sources. While these methods work well in simple scenarios, they fail to fully exploit the potential information within the data when faced with complex spatiotemporal correlations, particularly when dealing with large-scale heterogeneous data. Consequently, their processing accuracy and generalization capabilities are poor, failing to meet the current high-precision requirements for air quality monitoring and prediction.

[0004] Machine learning-based fusion methods leverage the model's self-learning capabilities to automatically extract features and relationships from data. However, when faced with large-scale, high-dimensional, heterogeneous data, these methods often require large amounts of training data, long training times, and high data quality. While deep learning methods can effectively handle nonlinear and high-dimensional data, their complex training process and susceptibility to data noise often make accuracy and stability difficult to guarantee.

[0005] In terms of matrix decomposition methods, many studies have adopted collaborative filtering algorithms to deal with missing data. These methods approximate the relationship between user behavior and items by constructing low-rank matrices. However, such methods have obvious limitations in practical applications. First, traditional matrix decomposition methods generally assume that missing data is randomly distributed. Missing data in air quality data is often affected by factors such as spatiotemporal distribution and equipment failure, presenting a complex non-random distribution. Therefore, traditional matrix decomposition methods are less effective when processing such data. Second, existing methods often rely on a single decomposition technique, such as singular value decomposition, which is prone to overfitting or underestimating certain underlying patterns when processing data with high noise or high heterogeneity. In addition, existing matrix completion methods often fail to fully utilize the spatiotemporal correlation of data and the interactions between various data sources, resulting in consistency and accuracy issues in the fused results.

[0006] With the continuous development of deep learning technology, attention mechanisms have been widely used in various data fusion and modeling tasks. Attention mechanisms can automatically learn the weights of different parts of the data, improving the model's focus on key features. However, existing air quality data fusion methods do not fully integrate attention mechanisms with matrix factorization techniques, lacking effective weighting of differences in spatiotemporal features between different data sources. When processing high-dimensional, sparse, and heterogeneous air quality data, there is a lack of an effective mechanism to automatically optimize and adjust the data contribution, which can lead to errors and inconsistencies in the results of model fusion.

[0007] Therefore, existing technologies have several shortcomings in the air quality data fusion process: First, they cannot fully address the heterogeneity and missing data issues between different data sources. Traditional methods generally rely on static weights or simple interpolation techniques, which leads to low accuracy and consistency in the fusion results. Second, when processing large-scale sparse data, existing matrix decomposition techniques often cannot effectively capture the complex spatiotemporal correlations between data, resulting in poor generalization of prediction results. Third, there is a lack of optimization strategies that combine spatiotemporal characteristics and data quality, and it is unable to fully utilize the complementary information between different data sources, resulting in large errors after data fusion. Summary of the Invention

[0008] One objective of this invention is to propose a heterogeneous air quality data fusion method based on sparse matrix decomposition. By combining improved singular value decomposition with low-rank matrix completion techniques, this method effectively addresses the missing heterogeneous data problem and optimizes the feature weighting process of matrix decomposition by introducing an attention mechanism. Through temporal synchronization correction, spatial alignment, and normalization, an optimized data matrix adapted to air quality monitoring requirements is constructed. Multiple error metrics and consistency verification methods are employed to ensure the accuracy and consistency of data fusion.

[0009] A heterogeneous air quality data fusion method based on sparse matrix decomposition according to an embodiment of the present invention includes the following steps: S1. Collect air quality data from different data sources; S2. Preprocess the collected air quality data, and merge the preprocessed air quality data into a multidimensional matrix by combining time and space information to construct a preliminary fusion matrix; S3, marking the missing data in the preliminary fusion matrix and constructing a sparse matrix; S4. Based on the constructed sparse matrix, an improved singular value decomposition method is combined with a matrix completion method based on low-rank representation to perform matrix decomposition and fill in missing data. Sparse regularization is combined to optimize matrix completion, and the attention mechanism is used to perform feature weighted optimization on the matrix decomposition results to generate an optimized air quality data matrix. S5. Perform error measurement and consistency test on the optimized air quality data matrix, and apply the final optimized air quality data matrix to pollution trend prediction, environmental monitoring system or decision support system.

[0010] Optionally, the air quality data includes pollutant data and meteorological data.

[0011] Optionally, the preprocessing includes performing time synchronization correction on timestamps of different data sources, adjusting time resolution by using a resampling method, performing spatial alignment based on geographic coordinate information, and performing standardization processing by using normalization.

[0012] Optionally, the sparse matrix Defined as: ; in, represents a sparse matrix, Represents the elements of the sparse matrix. If the data is missing, then ,otherwise , represents the data in the preliminary fusion matrix, Represents the data source, represents the time step, Indicates the number of features of air quality data, represents the missing data label matrix, Represent elements in the missing data marker matrix: .

[0013] Optionally, the S4 specifically includes: S41. For sparse matrices Perform singular value decomposition: ; in: Represents the low-rank feature matrix of the data source, which contains the feature information of each data source; represents a diagonal matrix of singular values, where the singular values ​​represent the main modes of the data; The low-rank feature matrix representing the time step represents the features in the time dimension; Represents a transpose operation; represents the number of latent variables, satisfying , represents the low-rank feature dimension of the data; S42, take the front singular values ​​and initialize the low-rank matrix based on the corresponding singular vectors: ; ; in, Indicates the The data source is in The singular vectors on the features, Indicates the The time step is The singular vectors on the features, Indicates the singular values; S43. Use low-rank matrix to fill missing data: ; in, Indicates the air quality data value after filling, Represents a data source The corresponding low-rank eigenvector, Represents the time step The corresponding low-rank eigenvector; S44. Use the filled air quality data values ​​to replace the missing values ​​in the sparse matrix: ; S45. Define the objective function and use L2 regularization to optimize the objective function: ; in, represents the objective function, Represents the regularization coefficient, controls complexity, and prevents overfitting. represents the output layer bias, represents the low-rank feature matrix of the data source, represents the low-rank feature matrix of the time step, represents a sparse matrix, represents the Frobenius norm; S46. Use gradient descent method to iteratively update the low-rank matrix: ; ; in, represents the objective function, represents the learning rate, Indicates the The data source is in The singular vectors on the features, Indicates the The time step is The singular vectors on the features, Represents iteration The second The data source is in The singular vectors on the features, Represents iteration The second The time step is The singular vectors on the features, Represents iteration The second The data source is in The singular vectors on the features, Represents iteration The second The time step is Singular vectors on features; S47. Calculate attention weight based on attention mechanism: ; in, Represents a data source At time step The attention weight at represents the natural exponential function, Represents the weight parameter matrix of the attention mechanism, which is used to adjust the weights of different time steps. Indicates the total number of time dimensions, that is, the data source The total number of time steps sampled, Indicates the Data sources at time step The filled air quality data value; S48. Perform weighted summation on the padded data matrix using the attention weights to obtain the optimized air quality data matrix: ; in, represents the optimized air quality data matrix, represents the attention weight, represents the air quality data matrix reconstructed after singular value decomposition and low-rank completion, represents the low-rank feature matrix of the data source, represents the low-rank feature matrix of the time step, Represents a transpose operation.

[0014] Optionally, the S5 specifically includes: S51. Use mean square error and mean absolute error to measure the error of the optimized air quality data matrix: ; ; in, represents the mean square error, represents the mean absolute error, Indicates the number of data sources, Represents the total number of time dimensions, that is, the number of time steps, Represents a sparse matrix The data points in Represents the optimized air quality data matrix The data points in S52. Use the root mean square error to evaluate the data fusion accuracy and calculate the normalized root mean square error: ; ; in, represents the root mean square error, the normalized root mean square error, and Respectively represent the maximum and minimum values ​​of air quality data; S53. Calculate the change rate and average change rate between adjacent time step data. If the data change rate exceeds the set threshold, it is considered as a time consistency anomaly: ; ; in, represents the rate of change between adjacent time step data, represents the average rate of change; S54. Calculate the measurement value differences and average spatial differences of adjacent data sources. If the measurement value differences between the data sources exceed a set threshold, it is considered a spatial consistency anomaly. ; ; in, represents the difference in measurements between adjacent data sources, represents the average spatial difference; S55. The air quality data matrix after error measurement and consistency verification is stored in a unified data structure, and the finally optimized air quality data matrix is ​​applied to pollution trend prediction, environmental monitoring system or decision support system.

[0015] The beneficial effects of the present invention are: First, the present invention effectively fills missing data when processing large-scale sparse data by constructing a sparse matrix and combining it with an improved singular value decomposition and low-rank matrix completion method. Traditional matrix decomposition methods typically assume that data missingness is random and fail to fully account for the spatiotemporal correlations of the data. However, by optimizing the matrix completion process, the present invention effectively captures the potential spatiotemporal patterns and complex distribution of missing data in air quality data, thereby improving the accuracy of data completion.

[0016] Secondly, the attention mechanism introduced in this paper performs weighted optimization on the features of different data sources. By automatically learning the weights of data sources at different time steps, the contribution of each data source is dynamically adjusted based on the data quality and spatiotemporal correlation. This weighted optimization strategy effectively addresses the imbalanced contribution of different data sources in traditional methods. This is especially true when measurement errors between different time steps and data sources are large. The attention mechanism automatically weights the data, improving the accuracy of the fusion results.

[0017] Finally, the present invention uses a variety of error metrics, including mean square error, root mean square error, and mean absolute error, combined with consistency testing, to perform multi-dimensional tests on the optimized air quality data matrix. These evaluation methods not only ensure the accuracy of the model but also enable the timely identification and removal of outliers during the data fusion process, ensuring the high credibility of the final fusion results. By calculating the rate of change between data sources and time steps, the spatiotemporal consistency of the data is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 This is a flow chart of a heterogeneous air quality data fusion method based on sparse matrix decomposition proposed by the present invention; Figure 2 This is a flowchart of matrix decomposition based on singular value decomposition and low-rank completion method for a heterogeneous air quality data fusion method based on sparse matrix decomposition proposed by the present invention. DETAILED DESCRIPTION

[0019] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0020] refer to Figure 1 and Figure 2 , a heterogeneous air quality data fusion method based on sparse matrix decomposition, comprising the following steps: S1. Collect air quality data from different data sources; S2. Preprocess the collected air quality data, and merge the preprocessed air quality data into a multidimensional matrix by combining time and space information to construct a preliminary fusion matrix; S3, marking the missing data in the preliminary fusion matrix and constructing a sparse matrix; S4. Based on the constructed sparse matrix, an improved singular value decomposition method is combined with a matrix completion method based on low-rank representation to perform matrix decomposition and fill in missing data. Sparse regularization is combined to optimize matrix completion, and the attention mechanism is used to perform feature weighted optimization on the matrix decomposition results to generate an optimized air quality data matrix. S5. Perform error measurement and consistency test on the optimized air quality data matrix, and apply the final optimized air quality data matrix to pollution trend prediction, environmental monitoring system or decision support system.

[0021] In this embodiment, the air quality data includes pollutant data and meteorological data.

[0022] In this embodiment, the preprocessing includes time synchronization correction of timestamps from different data sources, adjusting time resolution using a resampling method, performing spatial alignment based on geographic coordinate information, and performing standardization using normalization.

[0023] In this embodiment, the sparse matrix Defined as: ; in, represents a sparse matrix, Represents the elements of the sparse matrix. If the data is missing, then ,otherwise , represents the data in the preliminary fusion matrix, Represents the data source, represents the time step, Indicates the number of features of air quality data, represents the missing data label matrix, Represent elements in the missing data marker matrix: .

[0024] In this embodiment, the S4 specifically includes: S41. For sparse matrices Perform singular value decomposition: ; in: Represents the low-rank feature matrix of the data source, which contains the feature information of each data source; represents a diagonal matrix of singular values, where the singular values ​​represent the main modes of the data; The low-rank feature matrix representing the time step represents the features in the time dimension; Represents a transpose operation; represents the number of latent variables, satisfying , represents the low-rank feature dimension of the data; S42, take the front singular values ​​and initialize the low-rank matrix based on the corresponding singular vectors: ; ; in, Indicates the The data source is in The singular vectors on the features, Indicates the The time step is The singular vectors on the features, Indicates the singular values; S43. Use low-rank matrix to fill missing data: ; in, Indicates the air quality data value after filling, Represents a data source The corresponding low-rank eigenvector, Represents the time step The corresponding low-rank eigenvector; S44. Use the filled air quality data values ​​to replace the missing values ​​in the sparse matrix: ; S45. Define the objective function and use L2 regularization to optimize the objective function: ; in, represents the objective function, Represents the regularization coefficient, controls complexity, and prevents overfitting. represents the output layer bias, represents the low-rank feature matrix of the data source, represents the low-rank feature matrix of the time step, represents a sparse matrix, represents the Frobenius norm; S46. Use gradient descent method to iteratively update the low-rank matrix: ; ; in, represents the objective function, represents the learning rate, Indicates the The data source is in The singular vectors on the features, Indicates the The time step is The singular vectors on the features, Represents iteration The second The data source is in The singular vectors on the features, Represents iteration The second The time step is The singular vectors on the features, Represents iteration The second The data source is in The singular vectors on the features, Represents iteration The second The time step is Singular vectors on features; S47. Calculate attention weight based on attention mechanism: ; in, Represents a data source At time step The attention weight at represents the natural exponential function, Represents the weight parameter matrix of the attention mechanism, which is used to adjust the weights of different time steps. Indicates the total number of time dimensions, that is, the data source The total number of time steps sampled, Indicates the Data sources at time step The filled air quality data value; S48. Perform weighted summation on the padded data matrix using the attention weights to obtain the optimized air quality data matrix: ; in, represents the optimized air quality data matrix, represents the attention weight, represents the air quality data matrix reconstructed after singular value decomposition and low-rank completion, represents the low-rank feature matrix of the data source, represents the low-rank feature matrix of the time step, Represents a transpose operation.

[0025] In this embodiment, the S5 specifically includes: S51. Use mean square error and mean absolute error to measure the error of the optimized air quality data matrix: ; ; in, represents the mean square error, represents the mean absolute error, Indicates the number of data sources, Represents the total number of time dimensions, that is, the number of time steps, Represents a sparse matrix The data points in Represents the optimized air quality data matrix The data points in S52. Use the root mean square error to evaluate the data fusion accuracy and calculate the normalized root mean square error: ; ; in, represents the root mean square error, the normalized root mean square error, and Respectively represent the maximum and minimum values ​​of air quality data; S53. Calculate the change rate and average change rate between adjacent time step data. If the data change rate exceeds the set threshold, it is considered as a time consistency anomaly: ; ; in, represents the rate of change between adjacent time step data, represents the average rate of change; S54. Calculate the measurement value differences and average spatial differences of adjacent data sources. If the measurement value differences between the data sources exceed a set threshold, it is considered a spatial consistency anomaly. ; ; in, represents the difference in measurements between adjacent data sources, represents the average spatial difference; S55. The air quality data matrix after error measurement and consistency verification is stored in a unified data structure, and the finally optimized air quality data matrix is ​​applied to pollution trend prediction, environmental monitoring system or decision support system.

[0026] Example 1: To verify the feasibility of the present invention, it was applied to the air quality monitoring and early warning system of a large city. The city's air quality data comes from multiple different monitoring stations, including ground-based monitoring stations, meteorological stations, remote sensing satellites, and social perception platforms. These data sources have inconsistent timestamps, different spatial coverage, and often contain missing data due to factors such as equipment failure and data transmission issues. How to effectively integrate these heterogeneous data sources, fill in missing data, and improve data integrity and accuracy is one of the key application scenarios of the present invention.

[0027] In this application scenario, the city's environmental monitoring center monitors the air quality in 10 major areas and collects different pollutants (such as PM2.5, PM10, ) and meteorological data. Due to inconsistent data collection times across various data sources and the potential for equipment failure, data is missing for certain periods. Therefore, the city's air quality monitoring system utilizes the heterogeneous data fusion method based on sparse matrix decomposition proposed in this paper to effectively fill in missing data and optimize data quality, ensuring data timeliness and accuracy.

[0028] The data used in the experiment include PM2.5, PM10, Hourly data on four major pollutants, including CO and CO. In addition to ground-based monitoring data, this also includes meteorological data such as temperature, humidity, and wind speed provided by weather stations. Because different data sources have different temporal and spatial resolutions, the data are aligned in time and space. All data are resampled to an hourly time step and spatially aligned based on geographic coordinates, mapping the data sources to the same geographic area to ensure consistent data in each region at every moment.

[0029] After data preprocessing and alignment, pollutant concentrations were combined with meteorological data to construct a multidimensional sparse matrix. The rows of this matrix represent different time steps (hourly), and the columns represent different monitoring stations and pollutant types. Missing data are marked as null values ​​and are explicitly represented in the sparse matrix.

[0030] For example, at a certain point in time, the PM2.5 data of monitoring station 1 may be lost due to equipment failure, and the PM2.5 data of monitoring station 3 may be lost due to equipment failure. Data is missing due to transmission problems, and wind speed data from weather station 2 may be missing because it cannot be collected. These missing data are marked with 0 in the sparse matrix and further processing is based on this matrix.

[0031] Singular value decomposition (SVD) is performed on the sparse matrix to extract low-rank features. The main goal of SVD is to decompose the original sparse matrix into the product of several matrices, thereby extracting the main patterns in the data. These patterns represent the distribution trends and temporal variations of pollutants. After SVD, the low-rank matrix will include the patterns of pollutant concentration changes over time, as well as the response of each monitoring station to these patterns. To improve the accuracy of the completion, a matrix completion method based on sparse regularization was introduced. By optimizing the objective function, the completed data is more consistent with actual physical laws and historical trends.

[0032] After matrix completion, an attention mechanism is used to further optimize the imputed results. Specifically, by calculating the attention weights for different pollutants and monitoring stations at each time step, missing values ​​are weighted and corrected, thereby generating a more accurate air quality data matrix. Error metrics and consistency checks are performed on the imputed data and compared with traditional data imputation methods. Conventional evaluation metrics such as mean squared error and root mean square error are used to verify the accuracy of the imputed results. Furthermore, a data consistency check is performed by calculating the data change rate between adjacent time steps and spaces to further ensure the temporal and spatial consistency of the imputed data.

[0033] Table 1 Experimental data comparison table ; From the above experimental data comparison table, it can be seen that the heterogeneous air quality data fusion method based on sparse matrix decomposition proposed in the present invention is superior to traditional methods in terms of data filling accuracy, consistency and air quality prediction accuracy. In the process of filling data, traditional methods (such as mean filling and interpolation) often cannot fully utilize the potential patterns of the data, resulting in poor performance of the filled data in error metrics. For example, in terms of mean square error, the error value of the traditional method after filling is 2.34, while the error value of the method of the present invention after filling is only 1.02, which is a decrease of 56.0% compared with the traditional method, indicating that the deviation between the filled data and the real data is significantly reduced. Similarly, in terms of root mean square error, the error of the method of the present invention is 0.97, while the error of the traditional method is 1.53, a decrease of 36.7%, further proving the improvement of the accuracy of data filling.

[0034] In terms of temporal consistency and spatial consistency, since air quality data usually has obvious spatiotemporal correlation, the traditional method lacks modeling of the complex correlations between data, resulting in poor consistency of the filled data in the temporal and spatial dimensions. For example, in terms of the temporal consistency anomaly rate, the anomaly rate after filling by the traditional method is 8%, while the anomaly rate after filling by the method of the present invention is only 5%, a reduction of 37.5%. In terms of the spatial consistency anomaly rate, the method of the present invention reduces it to 4%, compared with 6% of the traditional method, an improvement of 33.3%. This shows that the method of the present invention effectively maintains the temporal change trend and spatial distribution law of the data during the data fusion process, avoiding the instability that may be caused by traditional methods.

[0035] In the application scenario of air quality prediction, the quality of the filled data directly affects the accuracy of the prediction model. Experimental results show that the data filled by the method of the present invention can improve the accuracy of pollutants (such as PM2.5, The prediction accuracy of PM2.5 concentration increased from 82.7% of the traditional method to 88.2%, an increase of 6.7%. The prediction accuracy increased from 79.8% to 83.3%, an increase of 4.4%. This shows that the method of the present invention can effectively improve the performance of the air quality prediction model, making the prediction results closer to the actual pollution situation, and providing more reliable data support for air pollution control and early warning.

[0036] Overall, the proposed method outperforms traditional methods in terms of data filling, error measurement, data consistency, and air quality prediction accuracy. In particular, when processing large-scale sparse data, the method based on singular value decomposition and low-rank matrix completion can fully explore the data's potential patterns. Combined with the attention mechanism, it performs weighted optimization on the data, making the filled data more realistic and reducing unreasonable fluctuations.

[0037] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A heterogeneous air quality data fusion method based on sparse matrix decomposition, characterized in that: The steps include: S1. Collect air quality data from different data sources; S2. Preprocess the collected air quality data, and merge the preprocessed air quality data into a multidimensional matrix by combining time and space information to construct a preliminary fusion matrix; S3, marking the missing data in the preliminary fusion matrix and constructing a sparse matrix; S4. Based on the constructed sparse matrix, an improved singular value decomposition method is combined with a matrix completion method based on low-rank representation to perform matrix decomposition and fill in missing data. Sparse regularization is combined to optimize matrix completion, and the attention mechanism is used to perform feature weighted optimization on the matrix decomposition results to generate an optimized air quality data matrix. S5. Perform error measurement and consistency test on the optimized air quality data matrix, and apply the final optimized air quality data matrix to pollution trend prediction, environmental monitoring system or decision support system.

2. A heterogeneous air quality data fusion method based on sparse matrix decomposition according to claim 1, characterized in that: The air quality data includes pollutant data and meteorological data.

3. The heterogeneous air quality data fusion method based on sparse matrix decomposition according to claim 1 is characterized in that: The preprocessing includes time synchronization correction of time stamps of different data sources, adjustment of time resolution by using a resampling method, spatial alignment based on geographic coordinate information, and standardization by using normalization.

4. The heterogeneous air quality data fusion method based on sparse matrix decomposition according to claim 1 is characterized in that: The sparse matrix Defined as: ; in, represents a sparse matrix, Represents the elements of the sparse matrix. If the data is missing, then ,otherwise , represents the data in the preliminary fusion matrix, Represents the data source, represents the time step, Indicates the number of features of air quality data, represents the missing data label matrix, Represent elements in the missing data marker matrix: 。 5. The heterogeneous air quality data fusion method based on sparse matrix decomposition according to claim 1 is characterized in that: The S4 specifically includes: S41. For sparse matrices Perform singular value decomposition: ; in: Represents the low-rank feature matrix of the data source, which contains the feature information of each data source; represents a diagonal matrix of singular values, where the singular values ​​represent the main modes of the data; The low-rank feature matrix representing the time step represents the features in the time dimension; Represents a transpose operation; represents the number of latent variables, satisfying , represents the low-rank feature dimension of the data; S42, take the front singular values ​​and initialize the low-rank matrix based on the corresponding singular vectors: ; ; in, Indicates the The data source is in The singular vectors on the features, Indicates the The time step is The singular vectors on the features, Indicates the singular values; S43. Use low-rank matrix to fill missing data: ; in, Indicates the air quality data value after filling, Represents a data source The corresponding low-rank eigenvector, Represents the time step The corresponding low-rank eigenvector; S44. Use the filled air quality data values ​​to replace the missing values ​​in the sparse matrix: ; S45. Define the objective function and use L2 regularization to optimize the objective function: ; in, represents the objective function, Represents the regularization coefficient, controls complexity, and prevents overfitting. represents the output layer bias, represents the low-rank feature matrix of the data source, represents the low-rank feature matrix of the time step, represents a sparse matrix, represents the Frobenius norm; S46. Use gradient descent method to iteratively update the low-rank matrix: ; ; in, represents the objective function, represents the learning rate, Indicates the The data source is in The singular vectors on the features, Indicates the The time step is The singular vectors on the features, Represents iteration The second The data source is in The singular vectors on the features, Represents iteration The second The time step is The singular vectors on the features, Represents iteration The second The data source is in The singular vectors on the features, Represents iteration The second The time step is Singular vectors on features; S47. Calculate attention weight based on attention mechanism: ; in, Represents a data source At time step The attention weight at represents the natural exponential function, Represents the weight parameter matrix of the attention mechanism, which is used to adjust the weights of different time steps. Indicates the total number of time dimensions, that is, the data source The total number of time steps sampled, Indicates the Data sources at time step The filled air quality data value; S48. Perform weighted summation on the padded data matrix using the attention weights to obtain the optimized air quality data matrix: ; in, represents the optimized air quality data matrix, represents the attention weight, represents the air quality data matrix reconstructed after singular value decomposition and low-rank completion, represents the low-rank feature matrix of the data source, represents the low-rank feature matrix of the time step, Represents a transpose operation.

6. The heterogeneous air quality data fusion method based on sparse matrix decomposition according to claim 1 is characterized in that: The S5 specifically includes: S51. Use mean square error and mean absolute error to measure the error of the optimized air quality data matrix: ; ; in, represents the mean square error, represents the mean absolute error, Indicates the number of data sources, Represents the total number of time dimensions, that is, the number of time steps, Represents a sparse matrix The data points in Represents the optimized air quality data matrix The data points in S52. Use the root mean square error to evaluate the data fusion accuracy and calculate the normalized root mean square error: ; ; in, represents the root mean square error, the normalized root mean square error, and Respectively represent the maximum and minimum values ​​of air quality data; S53. Calculate the change rate and average change rate between adjacent time step data. If the data change rate exceeds the set threshold, it is considered as a time consistency anomaly: ; ; in, represents the rate of change between adjacent time step data, represents the average rate of change; S54. Calculate the measurement value differences and average spatial differences of adjacent data sources. If the measurement value differences between the data sources exceed a set threshold, it is considered a spatial consistency anomaly. ; ; in, represents the difference in measurements between adjacent data sources, represents the average spatial difference; S55. The air quality data matrix after error measurement and consistency verification is stored in a unified data structure, and the finally optimized air quality data matrix is ​​applied to pollution trend prediction, environmental monitoring system or decision support system.

Citation Information

Cited By

  • Dynamic weight matrix decomposition energy consumption prediction method and system

    CN121189573A