Atmospheric Pollutant Detection Method and System Based on Multi-Source Data Analysis

Through multi-source data analysis and multi-component decomposition methods, combined with atmospheric interference and diffusion models, the pollution hot spots and source locations are identified, and the problems of insufficient accuracy and coverage of traditional monitoring methods are solved, and the rapid and accurate monitoring of atmospheric pollutants is achieved.

CN119901687BActive Publication Date: 2025-06-24四川省遂宁生态环境监测中心站 +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510330090.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-20
Publication Date
2025-06-24
Estimated Expiration
2045-03-20

AI Technical Summary

Technical Problem

Traditional atmospheric pollutant monitoring methods rely on a single data source, which is difficult to meet the needs of complex pollution monitoring, and are disturbed by factors such as meteorological conditions and topographic characteristics, resulting in reduced monitoring accuracy and reliability.

Method used

The multi-source data analysis method is used to obtain hyperspectral data, environmental parameters and remote sensing images of pollutants in the atmosphere. The spectral characteristics of pollutants are extracted through the Gaussian distribution multi-component decomposition method, combined with the atmospheric interference model and pollutant diffusion model, identify the pollutant hot spots and the location of the pollution source, and optimize the diffusion model through time series analysis.

Benefits of technology

It realizes rapid and accurate monitoring of atmospheric pollutants in terms of spatial and temporal characteristics, improves the space-time coverage and accuracy of pollutant monitoring, and provides a scientific basis for pollution control and early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119901687B_ABST
    Figure CN119901687B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for detecting atmospheric pollutants based on multi-source data analysis. The method includes obtaining hyperspectral data, environmental parameters, and remote sensing images of pollutants; separating features in the spectral overlapping region by a multi-component decomposition method based on Gaussian distribution, matching a pre-stored pollution spectral library to determine the types and concentration ranges of pollutants; removing interference signals in the hyperspectral data through an atmospheric interference model to obtain pure data; inputting the environmental parameters, pollutant information, and pure data into a pollutant diffusion model, and performing layer superposition analysis in combination with the remote sensing image to identify pollution hotspots; identifying outliers in the hotspots based on the pure data to locate pollution sources; analyzing the time series of the spectral data of the pollution sources, and optimizing the diffusion model in combination with the environmental parameters; and finally outputting the types of pollutants, concentration ranges, hotspots, pollution source locations, and time characteristics to achieve rapid and accurate spatio-temporal monitoring of atmospheric pollutants.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of air pollutant monitoring, and particularly to an air pollutant detection method and system based on multi-source data analysis. Background Art

[0002] Traditional air pollutant monitoring methods mainly rely on a single data source, such as ground monitoring stations or satellite remote sensing data. Although these methods can provide basic information on pollutant concentrations to a certain extent, due to their inherent limitations, such as insufficient spatial coverage of ground monitoring stations and limited spatio-temporal resolution of satellite remote sensing data, it is difficult to meet the increasingly complex pollution monitoring requirements. In addition, the interference of factors such as meteorological conditions and terrain features further reduces the monitoring accuracy and reliability of traditional methods. With the rapid development of big data, artificial intelligence, and Internet of Things technologies, multi-source data fusion technology has shown significant advantages in the field of air pollution monitoring. By integrating multi-source information from ground monitoring stations, satellite remote sensing, unmanned aerial vehicle monitoring, meteorological data, and social media, this technology can effectively make up for the deficiencies of a single data source and provide more comprehensive, accurate, and real-time pollutant monitoring results. Multi-source data fusion technology can not only improve the spatio-temporal coverage ability and accuracy of pollutant monitoring, but also deeply explore the pollutant sources, diffusion laws, and their correlations with meteorological conditions through big data analysis and machine learning algorithms, thus providing a scientific basis for pollution control and early warning.

[0003] In the air pollutant monitoring method using a single data source of ground monitoring stations, first, several representative locations are selected in cities or industrial areas, and fixed air pollution monitoring stations are deployed. Each monitoring station is equipped with a variety of sensors for real-time monitoring of the concentrations of pollutants such as PM2.5, PM10, SO2, NO2, CO, and O3 in the air. The data collected by the sensors is transmitted to the data center through wired or wireless networks. The data center preprocesses the received raw data, including operations such as removing outliers, filling missing values, and data standardization. The processed data is integrated with the geographical location information of the monitoring station to generate a time series of air pollutant concentrations at specific stations. Then, the data center uses interpolation methods or statistical models, such as Kriging interpolation or inverse distance weighted interpolation, to expand the discrete station data into a spatial distribution map of pollutants in the entire monitoring area. Finally, the system regularly updates and publishes a large-scale air pollution status report for relevant departments and the public to query and reference.

[0004] Due to the limited spatial coverage of a single data source in ground monitoring stations, it is difficult to capture the changes in pollutants outside the monitoring area. Additionally, the sparse data in local areas caused by complex terrain, uneven distribution of monitoring stations, atmospheric background interference, and the influence of dynamic environmental parameters leads to significant deviations in the spatial distribution map of pollutant concentrations based on ground monitoring stations, restricting its application effect in large-scale fine monitoring and precise pollution source tracing. Therefore, the existing technology cannot avoid the influence of atmospheric background interference and dynamic environmental parameters, and cannot quickly and accurately monitor atmospheric pollutants in terms of spatio-temporal characteristics. Summary of the Invention

[0005] The present invention provides a method and system for detecting atmospheric pollutants based on multi-source data analysis to achieve rapid and accurate monitoring of atmospheric pollutants in terms of spatio-temporal characteristics by combining the influence of atmospheric background interference and dynamic environmental parameters.

[0006] In a first aspect, to solve the above technical problems, the present invention provides a method for detecting atmospheric pollutants based on multi-source data analysis, including:

[0007] Obtain hyperspectral data, environmental parameters, and remote sensing images of pollutants in the atmosphere, where the environmental parameters include aerosol concentration, water vapor concentration, wind speed, wind direction, temperature, and humidity;

[0008] According to the hyperspectral data, use a multi-component decomposition method based on Gaussian distribution to separate features in the spectral overlap region to obtain the spectral features of pollutants;

[0009] Match the spectral features with a pre-stored pollution spectral library to obtain the types of pollutants and the range of pollutant concentrations;

[0010] According to a pre-constructed atmospheric interference model and the environmental parameters, remove interference signals from the hyperspectral data to obtain pure data;

[0011] Input the environmental parameters, the types of pollutants, the range of pollutant concentrations, and the pure data into a pre-constructed pollutant diffusion model, and perform layer superposition analysis in combination with the remote sensing image to obtain pollution hotspots;

[0012] Identify abnormal points in the pollution hotspots based on the pure data to obtain the source locations of pollutants;

[0013] Perform time series analysis on the spectral data at the source locations of pollutants to obtain the time characteristics of pollutant emissions, and optimize the pollutant diffusion model in combination with the environmental parameters;

[0014] Output the types of pollutants, the range of pollutant concentrations, the pollution hotspots, the source locations of pollutants, and the time characteristics as the results of atmospheric pollutant detection.

[0015] In an alternative embodiment, based on the hyperspectral data, a multi-component decomposition method based on Gaussian distribution is used to separate the features of the spectral overlapping region to obtain the spectral features of the pollutant, including:

[0016] Convert the hyperspectral data into the form of a Gaussian function to obtain Gaussian data;

[0017] Perform Fourier transform on the Gaussian data to obtain frequency-domain data decomposed into different frequencies;

[0018] Perform inverse Fourier transform on each piece of the frequency-domain data to obtain wavelength-domain data of different wavelengths;

[0019] Calculate the mean square error between the decomposed spectral signal and the mixed signal according to the wavelength-domain data;

[0020] Calculate the spectral matching degree according to the wavelength-domain data and preset standard data, and compare the spectral matching degree with a preset spectral matching threshold. If the spectral matching degree is less than the preset spectral matching threshold, adjust the Gaussian function parameters with the goal of minimizing the mean square error;

[0021] If the spectral matching degree is greater than the preset spectral matching threshold, obtain the spectral features of the pollutant;

[0022] Among them, the hyperspectral data is converted into the form of a Gaussian function through the following formula:

[0023]

[0024] Among them, the mean square error is calculated through the following formula:

[0025]

[0026] Among them, the spectral matching degree is calculated through the following formula:

[0027]

[0028] Among them, represents the hyperspectral data, represents the amplitude of the th Gaussian function, represents the th central wavelength, represents the standard deviation of the hyperspectral data, represents the Gaussian function component, represents the mean square error, represents the wavelength-domain data at wavelength , represents the number of wavelength-domain data Indicates the wavelength domain data of a pollution signal at wavelength . Indicates the standard data of a pollution signal at wavelength . Indicates the spectral matching degree, where e is the natural constant.

[0029] In an alternative embodiment, the matching of the spectral characteristics with a pre-stored pollution spectral library to obtain the pollutant type and the pollutant concentration range includes:[[]]

[0030] Performing denoising processing on the spectral characteristics by using wavelet transform to obtain denoised data;

[0031] Extracting feature points and feature lines from the denoised data and matching them with a preset pollution spectral library to obtain the pollutant type;

[0032] Calculating the pollution concentration based on the denoised data and the pollution spectral library to obtain the pollutant concentration range;

[0033] Among them, the feature points and feature lines are calculated by the following formula:

[0034]

[0035]

[0036] Among them, the pollutant concentration range is calculated by the following formula:

[0037]

[0038] Among them, represents the feature point, represents the feature line, represents the wavelength, represents the denoised data, and represent the start wavelength and the end wavelength of the absorption peak, represents the pollutant concentration range, represents the reference data with wavelength in the pollution spectral library, represents the maximum wavelength, represents the preset molar extinction coefficient, represents the optical path length, represents the denoised data with wavelength , represents the constraint condition, taking the wavelength as the feature point . .

[0039] In an alternative embodiment, removing the interference signals from the hyperspectral data according to the pre-constructed atmospheric interference model and the environmental parameters to obtain pure data includes:

[0040] Calculating the interference amount of aerosol scattered light in the atmosphere according to the aerosol concentration to obtain the aerosol interference amount;

[0041] Calculating the influence amount of water vapor absorption light on the spectral value according to the water vapor concentration to obtain the water vapor influence amount;

[0042] Extracting the signal amount in the hyperspectral data according to the aerosol interference amount and the water vapor influence amount;

[0043] Correcting the signal amount through a linear regression model to obtain corrected data;

[0044] Extracting the pure value in the corrected data through a principal component analysis algorithm to obtain pure data;

[0045] The atmospheric interference model is as follows:

[0046]

[0047]

[0048]

[0049]

[0050]

[0051]

[0052] Wherein, represents the aerosol interference amount, represents the water vapor influence amount, represents the aerosol concentration, represents the wavelength of the scattering coefficient, represents the wavelength, represents the spectral value without absorption, represents the water vapor absorption coefficient, represents the water vapor concentration, represents the optical path length, represents the hyperspectral data, represents the signal amount, represents the corrected data, and represents the regression coefficient fitted by the least squares method, Represents a preset error term, Represents the feature vector matrix, Represents the eigenvalue matrix, Represents the transpose matrix of the feature vector matrix, Represents the first k matrix composed of feature vectors, and P represents the pure data.

[0053] In an alternative embodiment, the step of inputting the environmental parameters, the types of pollutants, the pollutant concentration range, and the pure data into a pre-constructed pollutant diffusion model, and performing layer overlay analysis in combination with the remote sensing image to obtain the pollution hotspot area includes:

[0054] Input the environmental parameters, the types of pollutants, the pollutant concentration range, and the pure data into a pre-constructed pollutant diffusion model to obtain predicted pollution data and generate a pollutant distribution map;

[0055] Perform layer overlay analysis on the pollutant distribution map and the remote sensing image to obtain the pollution hotspot area;

[0056] Among them, the pollutant diffusion model is as follows:

[0057]

[0058]

[0059]

[0060]

[0061] Among them, Represents the pollutant concentration, Represents the pollution type influence coefficient, Represents the wind speed vector, and Represents the diffusion coefficient, Represents the pollution emission height obtained through the remote sensing image, 、 、 and Represents the preset empirical coefficient, 、 and Represents the distances in the downwind direction, crosswind direction, and vertical direction from the pollutant.

[0062] In an alternative embodiment, the step of identifying outliers in the pollution hotspot area based on the pure data to obtain the pollution source location includes:

[0063] Use the K - value clustering algorithm to perform clustering analysis on the pure data corresponding to the pollution hot - spot area, obtain spectral categories, and take the spectral category containing the most pure data as the background spectral category;

[0064] Calculate the difference degree between each spectral category and the background spectral category, and compare the difference degree with a preset anomaly threshold. If the difference degree is less than the anomaly threshold, this spectral category is marked as a normal point;

[0065] If the difference degree is greater than the anomaly threshold, this spectral category is marked as an abnormal point;

[0066] Combine all the abnormal points for spatial distribution analysis to obtain the pollution source location;

[0067] Among them, the difference degree is calculated by the following formula:

[0068]

[0069] Among them, represents the difference degree, represents the spectral feature value of the spectral category, represents the spectral feature value of the background spectral category.

[0070] In an alternative embodiment, the time - series analysis of the spectral data at the pollution source location to obtain the time characteristics of pollutant emissions and optimize the pollutant diffusion model in combination with the environmental parameters includes:

[0071] Perform time - series analysis on the spectral data at the pollution source location and the pre - stored historical spectral data to obtain the time characteristics of pollutant emissions;

[0072] According to the time characteristics and the environmental parameters, optimize the parameters of the pollutant diffusion model by the gradient descent method.

[0073] In a second aspect, the present invention provides an atmospheric pollutant detection system based on multi - source data analysis, including:

[0074] A data acquisition module for acquiring hyperspectral data of pollutants in the atmosphere, environmental parameters, and remote sensing images, where the environmental parameters include aerosol concentration, water vapor concentration, wind speed, wind direction, temperature, and humidity;

[0075] A spectral feature analysis module for separating the features of the spectral overlapping region by using a multi - component decomposition method based on Gaussian distribution according to the hyperspectral data to obtain the spectral features of pollutants;

[0076] A pollutant matching module for matching the spectral features with a pre - stored pollution spectral library to obtain the pollutant types and pollutant concentration ranges;

[0077] A data optimization module, configured to remove interference signals from the hyperspectral data according to a pre-constructed atmospheric interference model and the environmental parameters, so as to obtain pure data;

[0078] A pollution area analysis module, configured to input the environmental parameters, the types of pollutants, the pollutant concentration range and the pure data into a pre-constructed pollutant diffusion model, and perform layer overlay analysis in combination with the remote sensing image to obtain a pollution hotspot area;

[0079] A pollution source analysis module, configured to identify abnormal points in the pollution hotspot area according to the pure data to obtain the location of the pollution source;

[0080] A diffusion model optimization module, configured to perform time series analysis on the spectral data at the location of the pollution source to obtain the time characteristics of pollutant emissions, and optimize the pollutant diffusion model in combination with the environmental parameters;

[0081] An output module, configured to output the types of pollutants, the pollutant concentration range, the pollution hotspot area, the location of the pollution source and the time characteristics as the results of atmospheric pollutant detection.

[0082] Compared with the prior art, the present invention has the following beneficial effects:

[0083] The present invention discloses an atmospheric pollutant detection method based on multi-source data analysis, which includes obtaining hyperspectral data, environmental parameters, and remote sensing images of pollutants in the atmosphere. Among them, the environmental parameters include aerosol concentration, water vapor concentration, wind speed, wind direction, temperature, and humidity; according to the hyperspectral data, a multi-component decomposition method based on Gaussian distribution is used to separate the characteristics of the spectral overlapping region to obtain the spectral characteristics of the pollutants; according to the spectral characteristics and the pre-stored pollution spectral library, the pollutant types and pollutant concentration ranges are obtained; according to the pre-constructed atmospheric interference model and the environmental parameters, interference signals are removed from the hyperspectral data to obtain pure data; the environmental parameters, the pollutant types, the pollutant concentration ranges, and the pure data are input into the pre-constructed pollutant diffusion model, and layer stacking analysis is performed in combination with the remote sensing image to obtain the pollution hot spot area; abnormal point recognition is performed on the pollution hot spot area according to the pure data to obtain the source location of the pollutant; time series analysis is performed on the spectral data of the source location to obtain the time characteristics of pollutant emissions, and the pollutant diffusion model is optimized in combination with the environmental parameters; the pollutant types, the pollutant concentration ranges, the pollution hot spot area, the source location of the pollutant, and the time characteristics are output as the results of atmospheric pollutant detection. The present invention realizes the rapid and accurate spatio-temporal monitoring of atmospheric pollutants by obtaining the hyperspectral data, environmental parameters, and remote sensing images of pollutants; separating the characteristics of the spectral overlapping region by a multi-component decomposition method based on Gaussian distribution, matching the pre-stored pollution spectral library, and determining the pollutant types and concentration ranges; removing the interference signals in the hyperspectral data through the atmospheric interference model to obtain pure data; inputting the environmental parameters, pollutant information, and pure data into the pollutant diffusion model, performing layer stacking analysis in combination with the remote sensing image, and identifying the pollution hot spot area; performing abnormal point recognition on the hot spot area based on the pure data to locate the source of the pollutant; analyzing the time series of the spectral data of the source of the pollutant, and optimizing the diffusion model in combination with the environmental parameters; finally outputting the pollutant types, concentration ranges, hot spot areas, source locations of the pollutants, and time characteristics. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 FIG. is a schematic flowchart of an atmospheric pollutant detection method based on multi-source data analysis provided by the first embodiment of the present invention;

[0085] Figure 2 FIG. is a schematic structural diagram of an atmospheric pollutant detection system based on multi-source data analysis provided by the second embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0086] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0087] Referring to Figure 1 , the first embodiment of the present invention provides an air pollutant detection method based on multi-source data analysis, including the following steps:

[0088] S11. Obtain hyperspectral data, environmental parameters, and remote sensing images of pollutants in the atmosphere, where the environmental parameters include aerosol concentration, water vapor concentration, wind speed, wind direction, temperature, and humidity;

[0089] S12. According to the hyperspectral data, use a multi-component decomposition method based on Gaussian distribution to separate features in the spectral overlap region to obtain the spectral features of the pollutants;

[0090] S13. Match the spectral features with a pre-stored pollution spectral library to obtain the types of pollutants and the range of pollutant concentrations;

[0091] S14. According to a pre-constructed atmospheric interference model and the environmental parameters, remove interference signals from the hyperspectral data to obtain pure data;

[0092] S15. Input the environmental parameters, the types of pollutants, the range of pollutant concentrations, and the pure data into a pre-constructed pollutant diffusion model, and perform layer stacking analysis in combination with the remote sensing image to obtain a pollution hotspot area;

[0093] S16. Identify abnormal points in the pollution hotspot area according to the pure data to obtain the source location of the pollution source;

[0094] S17. Perform time series analysis on the spectral data at the source location of the pollution source to obtain the time characteristics of pollutant emissions, and optimize the pollutant diffusion model in combination with the environmental parameters;

[0095] S18. Output the types of pollutants, the range of pollutant concentrations, the pollution hotspot area, the source location of the pollution source, and the time characteristics as the results of air pollutant detection.

[0096] In step S11, hyperspectral data, environmental parameters, and remote sensing images of pollutants in the atmosphere are obtained, where the environmental parameters include aerosol concentration, water vapor concentration, wind speed, wind direction, temperature, and humidity.

[0097] Hyperspectral data are collected by satellites, aircraft, or ground sensors. These data contain rich spectral information and can reflect the spectral characteristics of different substances in the atmosphere. The acquisition of environmental parameters mainly relies on weather stations, ground monitoring equipment, and satellite remote sensing technology. The aerosol concentration and water vapor concentration can be measured by devices such as lidar and sun photometers; the wind speed, wind direction, temperature, and humidity are obtained through weather stations or other meteorological sensors. Remote sensing images are taken using multispectral or hyperspectral satellite sensors, providing large-scale surface information and atmospheric conditions.

[0098] In step S12, according to the hyperspectral data, a multi-component decomposition method based on Gaussian distribution is used to separate the characteristics of the spectral overlapping region, and the spectral characteristics of the pollutants are obtained.

[0099] In a specific implementation manner, the step of using a multi-component decomposition method based on Gaussian distribution to separate the characteristics of the spectral overlapping region according to the hyperspectral data and obtain the spectral characteristics of the pollutants includes:

[0100] Convert the hyperspectral data into the form of a Gaussian function to obtain Gaussian data;

[0101] Perform Fourier transform on the Gaussian data to obtain frequency-domain data decomposed into different frequencies;

[0102] Perform inverse Fourier transform on each piece of the frequency-domain data respectively to obtain wavelength-domain data of different wavelengths;

[0103] Calculate the mean square error between the decomposed spectral signal and the mixed signal according to the wavelength-domain data;

[0104] Calculate the spectral matching degree according to the wavelength-domain data and preset standard data, and compare the spectral matching degree with a preset spectral matching threshold. If the spectral matching degree is less than the preset spectral matching threshold, take the minimization of the mean square error as the goal and adjust the Gaussian function parameters;

[0105] If the spectral matching degree is greater than the preset spectral matching threshold, obtain the spectral characteristics of the pollutants;

[0106] Among them, the hyperspectral data are converted into the form of a Gaussian function through the following formula:

[0107]

[0108] Among them, the mean square error is calculated through the following formula:

[0109]

[0110] Among them, the spectral matching degree is calculated through the following formula:

[0111]

[0112] Among them, represents the hyperspectral data, represents the amplitude of the th Gaussian function, represents the th central wavelength, represents the standard deviation of the hyperspectral data, represents the Gaussian function component, represents the mean square error, represents the wavelength domain data at wavelength represents the number of wavelength domain data, represents pollution signals at wavelength in the wavelength domain data, represents pollution signals at wavelength in the standard data, represents the spectral matching degree, where e is the natural constant.

[0113] Specifically, first, the hyperspectral data is converted into the form of Gaussian functions. Hyperspectral data usually contains the light intensity information of multiple wavelength points, and these data can be represented by the superposition of multiple Gaussian functions. The specific formula is:

[0114]

[0115] Among them, represents the hyperspectral data, represents the th amplitude of the Gaussian function, represents the th central wavelength, represents the standard deviation of the hyperspectral data, represents the Gaussian function component, where e is the natural constant. Through this formula, the hyperspectral data is decomposed into the superposition of multiple Gaussian functions, and each Gaussian function corresponds to a spectral component.

[0116] Next, the Gaussian data is subjected to Fourier transform to convert the time-domain data into frequency-domain data. The formula for Fourier transform is:

[0117]

[0118] The Fourier transform decomposes the Gaussian data into frequency-domain data of different frequencies, and each frequency component corresponds to a spectral feature. After Fourier transform, the inverse Fourier transform is performed on each frequency-domain data respectively to obtain the wavelength domain data of different wavelengths. The formula for inverse Fourier transform is:

[0119]

[0120] The inverse Fourier transform reconverts the frequency-domain data into wavelength-domain data, thereby obtaining the decomposed spectral signal. Based on the wavelength-domain data, the mean square error between the decomposed spectral signal and the original mixed signal is calculated. The formula for the mean square error is as follows:

[0121]

[0122] where represents the number of wavelength-domain data,[ represents the wavelength-domain data at wavelength , , and represent the amplitude, central wavelength, and standard deviation of the Gaussian function respectively, and e is the natural constant. The mean square error is used to measure the fitting degree between the decomposed spectral signal and the original signal, and the smaller the value, the better the fitting effect.

[0123] According to the wavelength-domain data and the preset standard data, the spectral matching degree is calculated. The formula for the spectral matching degree is as follows:

[0124]

[0125] where represents wavelength-domain data of pollution signals at wavelength represents wavelength-domain data of pollution signals at wavelength , the central wavelength and the standard deviation until the spectral matching degree is greater than the threshold, thereby obtaining the accurate spectral characteristics of the pollutants.

[0126] Through the above process, the spectral overlapping regions in the hyperspectral data are effectively decomposed, and the spectral characteristics of each pollutant are accurately extracted. The implementation of this method not only depends on the mathematical modeling of hyperspectral data, but also involves signal processing technologies such as Fourier transform, inverse Fourier transform, and calculation of spectral matching degree. During the decomposition process, the iterative optimization of the mean square error and spectral matching degree ensures the high precision of the analysis results.

[0127] In step S13, the spectral features are matched with a pre-stored pollution spectral library to obtain the types of pollutants and the ranges of pollutant concentrations.

[0128] In a specific embodiment, the matching of the spectral features with the pre-stored pollution spectral library to obtain the types of pollutants and the ranges of pollutant concentrations includes:

[0129] Perform denoising processing on the spectral features using wavelet transform to obtain denoised data;

[0130] Extract feature points and feature lines from the denoised data and match them with a preset pollution spectral library to obtain the types of pollutants;

[0131] Calculate the pollution concentration based on the denoised data and the pollution spectral library to obtain the ranges of pollutant concentrations;

[0132] Among them, the feature points and feature lines are calculated by the following formula:

[0133]

[0134]

[0135] Among them, the ranges of pollutant concentrations are calculated by the following formula:

[0136]

[0137] Among them, represents the feature point, represents the feature line, represents the wavelength, represents the denoised data, and represent the starting wavelength and the ending wavelength of the absorption peak, represents the ranges of pollutant concentrations, represents the reference data with the wavelength of in the pollution spectral library, represents the maximum wavelength, represents the preset molar absorption coefficient, represents the optical path length, represents the denoised data with the wavelength of , represents the constraint condition, taking the wavelength of as the feature point . .

[0138] Specifically, first, denoise the spectral features. The wavelet transform method is used to remove the noise in the spectral data. By decomposing the spectral signal into frequency components of different scales, the wavelet transform can effectively separate the noise and the useful signal. After removing the noise through the wavelet transform, the denoised data is obtained. .

[0139] Next, extract the feature points and feature lines from the denoised data. Feature points refer to the positions of extreme values or inflection points in the spectral curve, and feature lines refer to the integral of the spectral intensity within a specific wavelength range. The calculation formula for feature points is:

[0140]

[0141] where, represents the feature point, represents the derivative of the spectral curve at wavelength . The position where the derivative is zero is the feature point. The calculation formula for feature lines is:

[0142]

[0143] where, represents the feature line, and represent the starting wavelength and ending wavelength of the absorption peak respectively. By integrating the spectral intensity within a specific wavelength range, the feature line is obtained.

[0144] After extracting the feature points and feature lines, match them with the pre-stored contaminated spectral library. The contaminated spectral library contains standard spectral data of various pollutants. By comparing the similarity of the feature points and feature lines with the standard spectra, the types of pollutants can be determined. The matching process uses a spectral similarity algorithm, such as the Euclidean distance or the correlation coefficient, to calculate the similarity between the measured spectrum and the standard spectrum, and selects the type of pollutant corresponding to the standard spectrum with the highest similarity.

[0145] After determining the types of pollutants, further calculate the range of pollutant concentrations. According to the Lambert-Beer law, the pollutant concentration is proportional to the spectral absorption intensity. The calculation formula for pollutant concentration is:

[0146]

[0147] where, represents the range of pollutant concentrations, the intensity of the measured spectrum at the maximum wavelength , represents the reference intensity at wavelength in the contaminated spectral library, represents the preset molar absorption coefficient, represents the optical path length, represents the wavelength of Denoised data Indicates the constraint condition, taking Wavelength of As the characteristic point . Through this formula, the concentration range of pollutants in the sample to be measured can be calculated.

[0148] Through wavelet transform denoising, characteristic point and characteristic line extraction, spectral matching, and concentration calculation, accurate identification of pollutant types and concentration ranges is achieved. Wavelet transform denoising improves the signal-to-noise ratio of spectral data, characteristic point and characteristic line extraction enhances the recognition of spectral features, spectral matching realizes the accurate classification of pollutant types, and concentration calculation provides quantitative information on pollutant content, improving the efficiency and accuracy of pollution source identification and pollution diffusion analysis.

[0149] In step S14, according to the pre-constructed atmospheric interference model and the environmental parameters, interference signals are removed from the hyperspectral data to obtain pure data.

[0150] In a specific embodiment, the removing interference signals from the hyperspectral data according to the pre-constructed atmospheric interference model and the environmental parameters to obtain pure data includes:

[0151] Calculating the interference amount of aerosol scattering light in the atmosphere according to the aerosol concentration to obtain the aerosol interference amount;

[0152] Calculating the influence amount of water vapor absorption light on the spectral value according to the water vapor concentration to obtain the water vapor influence amount;

[0153] Extracting the signal amount in the hyperspectral data according to the aerosol interference amount and the water vapor influence amount;

[0154] Correcting the signal amount through a linear regression model to obtain corrected data;

[0155] Extracting the pure value in the corrected data through a principal component analysis algorithm to obtain pure data;

[0156] The atmospheric interference model is as follows:

[0157]

[0158]

[0159]

[0160]

[0161]

[0162]

[0163] Among them, represents the aerosol interference amount, represents the water vapor influence amount, represents the aerosol concentration, represents the wavelength of the scattering coefficient, represents the wavelength, represents the spectral value without absorption, represents the water vapor absorption coefficient, represents the water vapor concentration, represents the optical path length, represents the hyperspectral data, represents the signal amount, represents the calibration data, and represents the regression coefficient fitted by the least squares method, represents the preset error term, represents the eigenvector matrix, represents the eigenvalue matrix, represents the transpose matrix of the eigenvector matrix, represents the first k matrix composed of eigenvectors, and P represents the pure data.

[0164] Specifically, the interference amount of aerosol scattered light in the atmosphere is calculated according to the aerosol concentration. The interference amount of aerosol scattered light is related to the aerosol concentration and the wavelength, and the specific formula is:

[0165]

[0166] Among them, represents the aerosol interference amount, represents the aerosol concentration, represents the wavelength of the scattering coefficient. This formula shows that the interference amount of aerosol scattered light is proportional to the aerosol concentration and inversely proportional to the fourth power of the wavelength.

[0167] Next, the influence amount of water vapor absorption light on the spectral value is calculated according to the water vapor concentration. The influence amount of water vapor absorption light is calculated by the Lambert-Beer law, and the specific formula is:

[0168]

[0169] Among them, represents the water vapor influence amount, represents the spectral value without absorption, represents the water vapor absorption coefficient, represents the water vapor concentration, Indicates the optical path length. This formula shows that the influencing amount of water vapor absorbing light has an exponential decay relationship with the water vapor concentration and the optical path length.

[0170] According to the aerosol interference amount and the water vapor influence amount, the signal amount in the hyperspectral data is extracted. The calculation formula for the signal amount is:

[0171]

[0172] Among them, Indicates the signal amount, Indicates the hyperspectral data. Through this formula, the interference of aerosol scattered light and water vapor absorbed light is removed from the hyperspectral data to obtain the preliminary signal amount.

[0173] In order to further improve the accuracy of the data, the signal amount is corrected through a linear regression model. The form of the linear regression model is:

[0174]

[0175] Among them, Indicates the corrected data, And Indicates the regression coefficients fitted by the least squares method, Indicates the preset error term. The fitting process of the linear regression model is achieved by minimizing the sum of squared errors, thereby improving the linear relationship of the data.

[0176] Finally, the pure value in the corrected data is extracted through the principal component analysis algorithm. The principal component analysis projects the data into a low-dimensional space through eigenvalue decomposition, and the specific formula is:

[0177]

[0178] Among them, Indicates the eigenvector matrix, Indicates the eigenvalue matrix, Indicates the transpose matrix of the eigenvector matrix. By selecting the matrix k composed of the first eigenvectors, the pure data P is calculated:

[0179]

[0180] Among them, P represents the pure data. Through the principal component analysis, the noise and redundant information in the corrected data are removed, and the main feature information is retained, thereby obtaining high-precision pure data.

[0181] Through the calculation of aerosol interference and water vapor absorption, the extraction of signal quantities, linear regression correction, and principal component analysis, the effective elimination of interference signals in hyperspectral data is achieved. The interference models of aerosols and water vapor are based on physical laws to ensure the accuracy of interference quantity calculation; linear regression correction improves the linear relationship of the data; principal component analysis extracts the main features of the data, enhances the reliability of the analysis results, and improves the accuracy and efficiency of pollution source location and pollutant diffusion analysis.

[0182] In step S15, the environmental parameters, the types of pollutants, the pollutant concentration range, and the pure data are input into a pre-constructed pollutant diffusion model, and layer overlay analysis is performed in combination with the remote sensing image to obtain the pollution hotspot area.

[0183] In a specific implementation manner, the inputting the environmental parameters, the types of pollutants, the pollutant concentration range, and the pure data into a pre-constructed pollutant diffusion model, and performing layer overlay analysis in combination with the remote sensing image to obtain the pollution hotspot area includes:

[0184] Input the environmental parameters, the types of pollutants, the pollutant concentration range, and the pure data into a pre-constructed pollutant diffusion model to obtain predicted pollution data and generate a pollutant distribution map;

[0185] Perform layer overlay analysis on the pollutant distribution map and the remote sensing image to obtain the pollution hotspot area;

[0186] Among them, the pollutant diffusion model is as follows:

[0187]

[0188]

[0189]

[0190]

[0191] Among them, represents the pollutant concentration, represents the pollution type influence coefficient, represents the wind speed vector, and represent the diffusion coefficient, represents the pollution emission height obtained through the remote sensing image, 、 、 and represent the preset empirical coefficients, 、 and Indicates the distances in the downwind, crosswind, and vertical directions from the pollutant.

[0192] Specifically, first, based on environmental parameters, pollutant types, pollutant concentration ranges, and pure data, calculate the concentration distribution of pollutants at different spatial positions. The pollutant diffusion model is based on the Gaussian diffusion equation, and the specific formula is:

[0193]

[0194] Where, Indicates the pollutant concentration, Indicates the influence coefficient of the pollution type, Indicates the wind speed vector, And Indicates the diffusion coefficient, Indicates the pollution emission height obtained through remote sensing images, , And Indicates the distances in the downwind, crosswind, and vertical directions from the pollutant. The diffusion coefficients And Are calculated through empirical formulas:

[0195]

[0196]

[0197] Where, , , And Indicate preset empirical coefficients used to describe the variation law of the diffusion coefficient with distance. Through this model, predict the pollution data And generate a pollutant distribution map.

[0198] In the pollutant distribution map, the spatial distribution information of the pollutant concentration is visualized, and the high-concentration areas are marked as potential pollution hotspots. Then, perform a layer overlay analysis on the pollutant distribution map and the remote sensing image. The layer overlay analysis combines the pollutant distribution information with the geographical features (such as terrain, vegetation, buildings, etc.) in the remote sensing image by aligning the spatial coordinates. The specific steps of the overlay analysis include coordinate system unification, data interpolation, and layer fusion. Through the overlay analysis, the relationship between the spatial distribution of pollutants and the geographical environment is clarified, thereby further accurately locating the pollution hotspot areas.

[0199] The determination of the pollution hotspot areas is based on the pollutant concentration threshold , that is, when the pollutant concentration Exceeds the preset threshold, this area is marked as a pollution hotspot. This process is achieved through spatial queries and conditional filtering to form the final pollution hotspot area map.

[0200] Through the prediction of the pollutant diffusion model and the layer overlay analysis of remote sensing images, the precise positioning of pollution hotspots is achieved. The pollutant diffusion model is based on the Gaussian equation and empirical coefficients, accurately describing the diffusion behavior of pollutants in space; the layer overlay analysis combines the pollutant distribution with the geographical environment, enhancing the practicality and visualization effect of the analysis results. This method improves the accuracy and efficiency of pollution source identification and pollution diffusion analysis.

[0201] In step S16, according to the pure data, outlier identification is performed on the pollution hotspots to obtain the positions of pollution sources.

[0202] In a specific embodiment, the outlier identification of the pollution hotspots according to the pure data to obtain the positions of pollution sources includes:

[0203] Using the K-value clustering algorithm to perform clustering analysis on the pure data corresponding to the pollution hotspots, obtaining spectral categories and taking the spectral category containing the most pure data as the background spectral category;

[0204] Calculating the difference degree between each spectral category and the background spectral category, and comparing the difference degree with a preset outlier threshold. If the difference degree is less than the outlier threshold, this spectral category is marked as a normal point;

[0205] If the difference degree is greater than the outlier threshold, this spectral category is marked as an outlier;

[0206] Combining the spatial distribution analysis of all the outliers to obtain the positions of pollution sources;

[0207] Among them, the difference degree is calculated by the following formula:

[0208]

[0209] Among them, represents the difference degree, represents the spectral feature value of the spectral category, represents the spectral feature value of the background spectral category.

[0210] Specifically, first, use the K-value clustering algorithm to perform clustering analysis on the pure data corresponding to the pollution hotspots. The K-value clustering algorithm divides the data into K clusters, so that the data points within the same cluster have high similarity, while the data points between different clusters have low similarity. The goal of clustering analysis is to obtain spectral categories, and the spectral category represents the area with similar spectral characteristics. During the clustering process, by calculating the distance between data points, the pure data is divided into K clusters, and the distance calculation formula is the Euclidean distance:

[0211]

[0212] Among them, represents the distance between data points and ; and respectively represent the spectral feature values of data points and in the -th dimension of the spectrum; represents the dimension of the spectral feature. After clustering, the spectral class containing the most pure data is marked as the background spectral class.

[0213] Next, calculate the difference degree between each spectral class and the background spectral class. The difference degree is calculated by cosine similarity, and the specific formula is:

[0214]

[0215] Among them, represents the difference degree, represents the spectral feature value of the spectral class, represents the spectral feature value of the background spectral class, represents the dot product of vectors and ; and respectively represent the norms of vectors and . The difference degree reflects the similarity between the spectral class and the background spectral class.

[0216] Compare the difference degree with a preset anomaly threshold: if the difference degree is less than the anomaly threshold, the spectral class is marked as a normal point; if the difference degree is greater than the anomaly threshold, the spectral class is marked as an abnormal point. Abnormal points indicate areas where the spectral features are significantly different from the background spectral class and are potential locations of pollution sources.

[0217] Finally, combine all abnormal points for spatial distribution analysis to determine the location of the pollution source. Spatial distribution analysis identifies high-density areas as pollution source locations by calculating the spatial density and distribution pattern of abnormal points. Specific methods include kernel density estimation and spatial clustering. Kernel density estimation generates a density distribution map by calculating the density of abnormal points within a certain range, and the high-density area is the location of the pollution source. The kernel density estimation formula is:

[0218]

[0219] Among them, represents the density value at location ; represents the number of abnormal points, represents the bandwidth parameter, represents the kernel function, represents the position and the distance between outliers Through kernel density estimation, the high-density area is identified as the pollution source location.

[0220] The above steps achieve the accurate identification of the pollution source location through K-value clustering analysis, difference calculation, and spatial distribution analysis. K-value clustering divides the pure data into spectral categories, difference calculation identifies outliers significantly different from the background spectral category, and spatial distribution analysis locates the high-density area as the pollution source. This method effectively locates the pollution source quickly, improving the efficiency and accuracy of environmental pollution monitoring and management.

[0221] In step S17, time series analysis is performed on the spectral data of the pollution source location to obtain the time characteristics of pollutant emissions, and the pollutant diffusion model is optimized in combination with the environmental parameters.

[0222] In a specific implementation manner, the time series analysis of the spectral data of the pollution source location to obtain the time characteristics of pollutant emissions and optimize the pollutant diffusion model in combination with the environmental parameters includes:

[0223] Performing time series analysis on the spectral data of the pollution source location and the pre-stored historical spectral data to obtain the time characteristics of pollutant emissions;

[0224] Optimizing the parameters of the pollutant diffusion model by the gradient descent method according to the time characteristics and the environmental parameters.

[0225] In step S17, time series analysis is performed on the spectral data of the pollution source location to obtain the time characteristics of pollutant emissions and optimize the pollutant diffusion model in combination with the environmental parameters. The specific implementation process is as follows: First, extract the spectral data of the pollution source location and integrate it with the pre-stored historical spectral data to form a time series data set. Time series analysis models the time variation law of the spectral data to reveal the dynamic characteristics of pollutant emissions. The time series analysis method uses the autoregressive integrated moving average model (ARIMA), and the ARIMA model is expressed as:

[0226]

[0227] where, represents the time of the spectral data of the polluted area at time represents the autoregressive coefficient polynomial, represents the moving average coefficient polynomial, represents the lag operator, represents the difference order, represents the random error term. By fitting the time series data with the ARIMA model, the time characteristics of pollutant emissions are extracted, including periodic changes, trend patterns, and abnormal fluctuations.

[0228] Next, using the extracted time characteristics and environmental parameters, the parameters of the pollutant diffusion model are optimized by the gradient descent method. The parameters of the pollutant diffusion model include the influence coefficient of pollution types , the empirical coefficient of the diffusion coefficient , , and . The optimization goal is to minimize the error between the prediction results of the model and the measured data. The error function is defined as the mean squared error:

[0229]

[0230] where, represents the error function, represents the total number of data points, represents the pollutant concentration predicted by the model, represents the measured pollutant concentration. The model parameters are iteratively updated by the gradient descent method to minimize the error function. The parameter update formula of the gradient descent method is:

[0231]

[0232] where, represents the model parameter at the -th iteration, represents the learning rate, represents the gradient of the error function with respect to the parameter . Through multiple iterations, the model parameters gradually converge to the optimal values, thereby optimizing the pollutant diffusion model.

[0233] The optimized pollutant diffusion model can more accurately predict the spatial diffusion behavior of pollutants. Combining the time characteristics obtained from time series analysis further enhances the dynamic prediction ability of the model. The form of the optimized model is:

[0234]

[0235] where, represents the optimized pollutant concentration, represents the optimized influence coefficient of pollution types, and through the optimized empirical coefficient , , and Calculation

[0236] Through time series analysis and gradient descent optimization, the extraction of the time characteristics of pollutant emissions and the optimization of the pollutant diffusion model are realized. Time series analysis reveals the dynamic law of pollutant emissions, and the gradient descent method improves the prediction accuracy of the model. The optimized model provides a reliable tool for the real-time monitoring, dynamic prediction and precise treatment of environmental pollution.

[0237] In step S18, the pollutant type, the pollutant concentration range, the pollution hot spot area, the pollution source location, and the time characteristics are output as the results of atmospheric pollutant detection.

[0238] Specifically, first, the pollutant type, the pollutant concentration range, the pollution hot spot area, the pollution source location, and the time characteristics are integrated into a structured data set. The pollutant type is represented by a classification label The pollutant concentration range is represented by the minimum concentration and the maximum concentration The pollution hot spot area is represented by a set of spatial coordinates The pollution source location is represented by spatial coordinates The time characteristics are represented by extracting time characteristics through a time series model represented.

[0239] The structured data set of the output result is represented in matrix form as:

[0240]

[0241] To facilitate the understanding of the present invention, some preferred embodiments of the present invention will be further described below.

[0242] The working process of the present invention is described below by taking a relatively common scenario as an example. The specific implementation manner of the present invention includes the following steps:

[0243] In a certain industrial area, first, hyperspectral data of pollutants in the atmosphere, environmental parameters including aerosol concentration, water vapor concentration, wind speed, wind direction, temperature, and humidity, and multi-spectral remote sensing images are obtained through satellites, aircraft, and ground sensors.

[0244] After the hyperspectral data is processed by a multi-component decomposition method based on Gaussian distribution, the spectral overlapping regions are effectively separated, and the spectral characteristics of the pollutants are extracted. By matching with a pre-stored pollution spectral library, the main pollutant types in the industrial area are determined to be sulfur dioxide (SO2) and nitrogen oxides (NO X )), and the concentration range of SO2 is calculated to be 50 - 80 μg / m³, NO XThe concentration range is 60 - 100 μg / m³.

[0245] Subsequently, the interference signals in the hyperspectral data were removed using an atmospheric interference model to obtain pure data. The environmental parameters, pollutant information, and pure data were input into the pollutant diffusion model, and layer stacking analysis was performed in combination with remote sensing images to identify two main pollution hotspots. The first pollution hotspot is located in the northeast corner of the industrial area, and the second hotspot is located in the southwest corner.

[0246] By identifying the abnormal points in the pure data, the location of the pollution source was further located. The pollution source in the northeast corner was determined to be the discharge port of a chemical plant, while the pollution source in the southwest corner is the blast furnace of a steel plant. Subsequently, time series analysis was performed on the spectral data of the pollution source, and it was found that the emission peak of SO2 appears at 8 am and 6 pm, and the emission peak of NO X appears at 12 noon.

[0247] Combining the time characteristics of pollutant emissions and environmental parameters, the pollutant diffusion model was optimized using the gradient descent method. The optimized model shows that wind speed and wind direction have a significant impact on the diffusion of pollutants. Especially when the wind speed exceeds 5 m / s, the pollutants will diffuse towards the southeast direction, and the affected area expands to the residential areas around the industrial area.

[0248] Finally, the system outputs the pollutant types (SO2 and NO X ), concentration ranges (SO2: 50 - 80 μg / m³, NO X : 60 - 100 μg / m³), pollution hotspot areas (northeast corner and southwest corner), pollution source locations (chemical plant and steel plant), and the time characteristics of pollutant emissions (SO2 at 8 am and 6 pm, NO X at 12 noon). These results provide accurate monitoring data for the environmental protection department to support the formulation of effective pollution control measures.

[0249] In the practical application of the industrial area, the method effectively improves the accuracy and efficiency of pollution source identification and pollutant diffusion analysis.

[0250] Referring to Figure 2 , the second embodiment of the present invention provides an atmospheric pollutant detection system based on multi-source data analysis, including:

[0251] A data acquisition module for acquiring hyperspectral data, environmental parameters, and remote sensing images of pollutants in the atmosphere, where the environmental parameters include aerosol concentration, water vapor concentration, wind speed, wind direction, temperature, and humidity;

[0252] A spectral feature analysis module, used to perform feature separation on the spectral overlapping area according to the hyperspectral data by using a multi-component decomposition method based on Gaussian distribution to obtain the spectral features of the pollutants;

[0253] A pollutant matching module, used to match the spectral characteristics with a pre-stored pollution spectrum library to obtain the type of pollutants and the range of pollutant concentrations;

[0254] A data optimization module, used to remove interference signals from the hyperspectral data according to a pre-built atmospheric interference model and the environmental parameters to obtain pure data;

[0255] A pollution area analysis module is used to input the environmental parameters, the pollutant types, the pollutant concentration range and the pure data into a pre-built pollutant diffusion model, and perform layer overlay analysis in combination with the remote sensing image to obtain pollution hotspot areas;

[0256] A pollution source analysis module, used to identify abnormal points in the pollution hotspot area according to the clean data to obtain the location of the pollution source;

[0257] A diffusion model optimization module, used to perform time series analysis on the spectral data at the pollution source location to obtain the time characteristics of pollutant emissions, and optimize the pollutant diffusion model in combination with the environmental parameters;

[0258] The output module is used to output the pollutant type, the pollutant concentration range, the pollution hotspot area, the pollution source location and the time characteristics as the result of atmospheric pollutant detection.

[0259] It should be noted that the atmospheric pollutant detection device based on multi-source data analysis provided in an embodiment of the present invention is used to execute all the process steps of the atmospheric pollutant detection method based on multi-source data analysis in the above-mentioned embodiment. The working principles and beneficial effects of the two correspond one to one, and therefore will not be repeated here.

[0260] The specific embodiments described above further illustrate the purpose, technical solutions and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. It is particularly pointed out that for those skilled in the art, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for detecting air pollutants based on multi-source data analysis, characterized in that: include: Obtain hyperspectral data, environmental parameters and remote sensing images of pollutants in the atmosphere, where environmental parameters include aerosol concentration, water vapor concentration, wind speed, wind direction, temperature and humidity; According to the hyperspectral data, a multi-component decomposition method based on Gaussian distribution is used to separate the characteristics of the spectral overlapping area to obtain the spectral characteristics of the pollutants; According to the spectral characteristics, the spectral characteristics are matched with a pre-stored pollution spectrum library to obtain the type of pollutants and the range of pollutant concentrations; According to the pre-built atmospheric interference model and the environmental parameters, interference signals are removed from the hyperspectral data to obtain pure data; The environmental parameters, the pollutant types, the pollutant concentration range and the pure data are input into the pre-built pollutant diffusion model, and the remote sensing images are combined for layer overlay analysis to obtain the pollution hotspot area. The pollutant diffusion model is as follows: in, represents the pollutant concentration, represents the impact coefficient of pollution type, represents the wind speed vector, and represents the diffusion coefficient, represents the pollution emission height obtained through remote sensing images, , , and represents the preset empirical coefficient, , and Indicates the distance from the pollutant in the downwind direction, crosswind direction and vertical direction; Identify abnormal points in the pollution hotspot area according to the clean data to obtain the location of the pollution source; Performing time series analysis on the spectral data at the pollution source location to obtain the time characteristics of pollutant emissions, and optimizing the pollutant diffusion model in combination with the environmental parameters; The pollutant type, the pollutant concentration range, the pollution hotspot area, the pollution source location and the time characteristics are output as the results of atmospheric pollutant detection.

2. The method for detecting air pollutants based on multi-source data analysis according to claim 1, characterized in that: According to the hyperspectral data, a multi-component decomposition method based on Gaussian distribution is used to separate the characteristics of the spectral overlapping area to obtain the spectral characteristics of the pollutants, including: Converting the hyperspectral data into a Gaussian function form to obtain Gaussian data; Performing Fourier transformation on the Gaussian data to obtain frequency domain data decomposed into different frequencies; Performing inverse Fourier transform on each frequency domain data to obtain wavelength domain data of different wavelengths; Calculate the mean square error between the decomposed spectral signal and the mixed signal according to the wavelength domain data; Calculating the spectral matching degree according to the wavelength domain data and the preset standard data, and comparing the spectral matching degree with a preset spectral matching threshold value, if the spectral matching degree is less than the preset spectral matching threshold value, adjusting the Gaussian function parameters with the goal of minimizing the mean square error; If the spectral matching degree is greater than a preset spectral matching threshold, the spectral characteristics of the pollutant are obtained; Among them, the hyperspectral data is converted into Gaussian function form by the following formula: The mean square error is calculated by the following formula: The spectral matching degree is calculated by the following formula: in, represents hyperspectral data, Indicates The amplitude of a Gaussian function, Indicates The central wavelength, represents the standard deviation of the hyperspectral data, represents the Gaussian function component, represents the mean square error, The wavelength is The wavelength domain data of Indicates the number of wavelength domain data, express The contamination signal is at wavelength The wavelength domain data at express The contamination signal is at wavelength Standard data, represents the spectral matching degree, and e is a natural constant.

3. The method for detecting air pollutants based on multi-source data analysis according to claim 1, characterized in that: The method of matching the spectral features with a pre-stored pollution spectrum library to obtain the type of pollutants and the range of pollutant concentrations includes: Using wavelet transform to perform denoising on the spectral features to obtain denoised data; Extracting characteristic points and characteristic lines according to the denoised data, and matching them with a preset pollution spectrum library to obtain pollutant types; Calculate the pollution concentration according to the denoised data and the pollution spectrum library to obtain a pollutant concentration range; Among them, the characteristic points and characteristic lines are calculated by the following formula: Among them, the pollutant concentration range is calculated by the following formula: in, Represents feature points, represents the characteristic line, represents the wavelength, represents the denoised data, and Indicates the starting wavelength and ending wavelength of the absorption peak. Indicates the pollutant concentration range, The wavelength in the contamination spectrum library is Reference data, represents the maximum wavelength, Indicates the preset molar absorptivity, represents the optical path length, The wavelength is The denoised data, Represents the constraint condition, taking Wavelength As a feature point .

4. The method for detecting air pollutants based on multi-source data analysis according to claim 1, characterized in that: The step of removing interference signals from the hyperspectral data according to the pre-built atmospheric interference model and the environmental parameters to obtain pure data includes: Calculating the interference amount of aerosol scattered light in the atmosphere according to the aerosol concentration to obtain the aerosol interference amount; Calculate the influence of water vapor absorption light on the spectral value according to the water vapor concentration to obtain the water vapor influence; Extracting a signal amount in the hyperspectral data according to the aerosol interference amount and the water vapor influence amount; Correcting the signal quantity by a linear regression model to obtain correction data; Extracting pure values ​​from the correction data by a principal component analysis algorithm to obtain pure data; The atmospheric disturbance model is as follows: in, represents the amount of aerosol interference, represents the amount of water vapor influence, represents the aerosol concentration, Indicates wavelength The scattering coefficient, represents the wavelength, represents the spectrum value when there is no absorption, represents the water vapor absorption coefficient, represents the water vapor concentration, represents the optical path length, represents hyperspectral data, Indicates a semaphore. represents the calibration data, and represents the regression coefficients fitted by the least squares method, represents the preset error term, represents the eigenvector matrix, represents the eigenvalue matrix, represents the transposed matrix of the eigenvector matrix, Before k The matrix consists of eigenvectors, and P represents the pure data.

5. The method for detecting air pollutants based on multi-source data analysis according to claim 1, characterized in that: The environmental parameters, pollutant types, pollutant concentration ranges and pure data are input into a pre-built pollutant diffusion model, and layer overlay analysis is performed in combination with the remote sensing image to obtain pollution hotspots, including: Inputting the environmental parameters, the pollutant types, the pollutant concentration range and the pure data into a pre-built pollutant diffusion model to obtain predicted pollution data and generate a pollutant distribution map; The pollutant distribution map and the remote sensing image are subjected to layer overlay analysis to obtain pollution hotspot areas.

6. The method for detecting air pollutants based on multi-source data analysis according to claim 1, characterized in that: The step of identifying abnormal points in the pollution hotspot area according to the clean data to obtain the location of the pollution source includes: Using a K-value clustering algorithm to perform cluster analysis on the clean data corresponding to the pollution hotspot area, obtain a spectrum category and use the spectrum category containing the most clean data as the background spectrum category; Calculating the difference between each of the spectrum categories and the background spectrum category, and comparing the difference with a preset abnormal threshold, if the difference is less than the abnormal threshold, marking the spectrum category as a normal point; If the difference is greater than the abnormal threshold, the spectrum category is marked as an abnormal point; Combining all the abnormal points to perform spatial distribution analysis and obtain the location of the pollution source; The difference is calculated by the following formula: in, Indicates the difference, spectral feature values ​​representing spectral classes, The spectral feature value representing the background spectral category.

7. The method for detecting air pollutants based on multi-source data analysis according to claim 1, characterized in that: The time series analysis of the spectral data at the pollution source location is performed to obtain the time characteristics of pollutant emissions, and the pollutant diffusion model is optimized in combination with the environmental parameters, including: Performing time series analysis on the spectral data of the pollution source location and pre-stored historical spectral data to obtain the time characteristics of pollutant emissions; According to the time characteristics and the environmental parameters, the parameters of the pollutant diffusion model are optimized by using a gradient descent method.

8. An atmospheric pollutant detection system based on multi-source data analysis, characterized in that: include: A data acquisition module is used to obtain hyperspectral data, environmental parameters and remote sensing images of pollutants in the atmosphere, wherein the environmental parameters include aerosol concentration, water vapor concentration, wind speed, wind direction, temperature and humidity; A spectral feature analysis module, used to perform feature separation on the spectral overlapping area according to the hyperspectral data by using a multi-component decomposition method based on Gaussian distribution to obtain the spectral features of the pollutants; A pollutant matching module, used to match the spectral characteristics with a pre-stored pollution spectrum library to obtain the type of pollutants and the range of pollutant concentrations; A data optimization module, used to remove interference signals from the hyperspectral data according to a pre-built atmospheric interference model and the environmental parameters to obtain pure data; The pollution area analysis module is used to input the environmental parameters, the pollutant types, the pollutant concentration range and the pure data into the pre-built pollutant diffusion model, and combine the remote sensing image to perform layer overlay analysis to obtain the pollution hotspot area. The pollutant diffusion model is as follows: in, represents the pollutant concentration, represents the impact coefficient of pollution type, represents the wind speed vector, and represents the diffusion coefficient, represents the pollution emission height obtained through remote sensing images, , , and represents the preset empirical coefficient, , and Indicates the distance from the pollutant in the downwind direction, crosswind direction and vertical direction; A pollution source analysis module, used to identify abnormal points in the pollution hotspot area according to the clean data to obtain the location of the pollution source; A diffusion model optimization module, used to perform time series analysis on the spectral data at the pollution source location to obtain the time characteristics of pollutant emissions, and optimize the pollutant diffusion model in combination with the environmental parameters; The output module is used to output the pollutant type, the pollutant concentration range, the pollution hotspot area, the pollution source location and the time characteristics as the result of atmospheric pollutant detection.

Citation Information

Patent Citations

  • Multi-source data fusion and environmental pollution source and pollutant distribution analysis method

    CN110186820A

  • Volatile organic pollutant diffusion simulation and tracing method and system

    CN117610438A