Deep learning based hyperspectral-vocs combined with underway method for traceability
By combining deep learning and convolutional neural networks, the problems of missing spatiotemporal dynamic response and limited component identification accuracy in VOCs emission source tracing have been solved, achieving efficient source tracing of VOCs emission sources and improving identification accuracy and anti-interference ability.
Patent Information
- Application Number
- CN202510735655.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-06-04
AI Technical Summary
Existing technologies for VOCs emission source tracing suffer from problems such as lack of spatiotemporal dynamic response, deviation in pollution area contour determination, limited accuracy of component identification, and mis-correlation or missed detection of redundant items, making it difficult to cope with non-continuous and irregular emission behaviors.
A deep learning-based hyperspectral-VOCs mobile source tracing method is adopted. By constructing a concentration anomaly response task flow, the indicators of concentration temporal difference, spatial distribution standard deviation and local concentration fluctuation anomaly rate are obtained. Combined with the gradient change judgment of remote sensing images, VOCs components are identified by convolutional neural network. The posterior probability of emission source and component is calculated by Bayesian inversion method to generate a pollution source tracing list.
It enhances the ability to respond to sudden or covert emission events, improves the robustness of pollution source boundary identification and the accuracy of component identification, and ensures the data support and quantitative reliability of source tracing results.
Smart Images

Figure CN120635736B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of hyperspectral satellite remote sensing technology, and in particular to a hyperspectral-VOCs mobile source tracing method based on deep learning. Background Technology
[0002] The field of hyperspectral satellite remote sensing encompasses a technological system for acquiring and analyzing high-spectral-resolution data on the Earth's surface, atmosphere, and water bodies using hyperspectral imaging equipment. The core of this field lies in achieving precise identification and quantitative inversion of the composition, spatial distribution characteristics, and dynamic changes of various substances through detailed observation of continuous bands in the electromagnetic spectrum. The overall technological system covers multiple observation methods, including hyperspectral satellite remote sensing, ground-based horizontal remote sensing, and targeted imaging remote sensing. It enables quantitative monitoring and spatial source tracing of pollutants such as PM2.5, O3, NO2, and VOCs in the atmosphere at different spatial and temporal resolutions, and is widely used in urban environmental monitoring, air quality assurance for major events, pollution source identification, and regulatory assessment.
[0003] The hyperspectral-VOCs mobile monitoring combined source tracing method refers to a method that combines hyperspectral remote sensing imaging data with VOCs mobile monitoring data to identify high-value pollutant source areas and conduct emission source tracing analysis. Addressing the issue of volatile organic compound emission source identification, it covers technical aspects such as simultaneous data acquisition, hyperspectral feature extraction, screening of abnormal VOCs concentration points, and intelligent identification of high-value pollution source areas.
[0004] Current technologies for tracing VOC emissions generally rely on static data for planar matching and human experience, resulting in a lack of spatiotemporal dynamic response, and are particularly ill-suited for handling non-continuous and irregular emission behaviors. While hyperspectral remote sensing possesses area source coverage capabilities, it often suffers from deviations in identifying pollution area contours during source boundary identification due to complex pixel grayscale distributions and blurred gradient changes, easily misidentifying differences in illumination or ground features as pollution boundaries. VOC mobile monitoring data lacks systematic screening standards for anomaly concentration identification, exhibiting the problem of high-frequency fluctuations mixing with genuine anomalies, leading to accumulated errors in task flow generation. Furthermore, pollutant component identification methods rely on conventional feature comparisons, neglecting band structure correlations and reflectance differences, thus limiting the accuracy of species classification. During source tracing, the relationship between emission points and components is often handled through simple matching or distance constraints, failing to reflect the multifactorial relationship between component abundance and source characteristics, and easily leading to redundant terms being incorrectly associated or missed detections. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing a hyperspectral-VOCs mobile source tracing method based on deep learning.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: a hyperspectral-VOCs mobile source tracing method based on deep learning, comprising the following steps:
[0007] S1: Obtain the VOCs concentration response sequence of the monitoring area through a mobile monitoring vehicle, filter out abnormal VOCs concentrations, and construct a concentration anomaly response task flow;
[0008] S2: Obtain three sets of indicators from the concentration anomaly response task flow: concentration time-series difference parameters, spatial distribution standard deviation, and local concentration fluctuation anomaly rate, and perform standardization processing. Based on the three sets of indicators, determine the pollution phase transition point and obtain the pollution phase transition judgment label.
[0009] S3: Obtain the remote sensing image of the monitoring area corresponding to the pollution phase change determination label during the corresponding time period, select the boundary location of the suspected pollution source area from the remote sensing image and determine the gradient change of the location, and output the boundary enhancement response map.
[0010] S4: Obtain hyperspectral fixed-point pixels of the suspected pollution source area in the boundary enhancement response map, and input them into the convolutional neural network model to identify VOCs compound components and obtain the component identification results;
[0011] S5: Based on the component identification results, the correspondence between emission sources and VOCs components is calculated using the Bayesian inversion method, and the sources are sorted and filtered according to the probability values to obtain a pollution source tracing list.
[0012] As a further aspect of the present invention, the concentration anomaly response task flow includes the concentration anomaly time, the anomaly concentration amplitude, and the anomaly duration; the pollution phase change determination label includes the phase change occurrence time, the phase change spatial location, and the phase change type; the boundary enhancement response map specifically includes the boundary location coordinates, the gradient change direction, and the boundary continuity index; the component identification result includes the VOCs species type, the reflectance feature vector, and the relative abundance value; and the pollution source tracing list specifically refers to the emission source coordinates, the emission source type, and the posterior probability value.
[0013] As a further aspect of the present invention, the step of obtaining the concentration anomaly response task flow specifically includes:
[0014] S111: The VOCs concentration response sequence of the monitoring area is obtained by mobile monitoring vehicle, the continuous TVOCs concentration data stream of single mass spectrometry equipment at a frequency of one second is collected, and the continuously collected TVOCs concentration values are arranged in time order to generate a concentration time series.
[0015] S112: Based on the TVOCs value at each time point in the concentration time series, using the 1ppm concentration threshold as the judgment benchmark, record the time index of all matching judgment benchmarks, and obtain the over-threshold concentration index sequence.
[0016] S113: Based on the time points recorded in the above-threshold concentration index sequence, extract the corresponding TVOCs concentration values and occurrence times, and construct a concentration anomaly response task flow.
[0017] As a further aspect of the present invention, the step of obtaining the pollution phase change determination label specifically includes:
[0018] S211: Obtain VOCs concentration data for the time segment corresponding to the concentration anomaly response task flow, and calculate the concentration temporal difference, spatial distribution standard deviation, and local concentration fluctuation anomaly rate to obtain three change indicators.
[0019] S212: The three change indicators are standardized using the Z-score method to obtain a standardized indicator sequence;
[0020] S213: Based on the standardized index sequence, set a double sliding time window, identify the jump trend of the local concentration fluctuation anomaly rate, mark the location area where the jump occurs, and generate a pollution phase change judgment label.
[0021] As a further aspect of the present invention, the step of obtaining the boundary enhancement response map specifically includes:
[0022] S311: Based on the pollution phase change determination label, obtain remote sensing image data of the corresponding monitoring area and time period, extract the gray level change of each pixel in the image, identify the boundary line of continuous pixels with prominent gradient changes, screen the edge area location of suspected pollution sources, and obtain the boundary candidate location set.
[0023] S312: Extract the position coordinates, gradient values and gradient point density of the boundary candidate position set in three consecutive frames of images, calculate the position offset, gradient change rate and density increase / decrease of each point in the time dimension, determine whether there is a continuous upward trend, and obtain a continuously enhanced boundary point set.
[0024] S313: Input the continuous enhanced boundary point set into the attention mechanism module in the convolutional neural network structure, dynamically allocate spatial weights according to the gradient strength and change trend of each point, perform response enhancement processing on the boundary region, and output the boundary enhancement response map.
[0025] As a further aspect of the present invention, the step of obtaining the component identification result specifically includes:
[0026] S411: Obtain the location of hotspot regions in the boundary enhancement response map where the brightness enhancement exceeds the boundary enhancement response threshold, and select the corresponding hyperspectral fixed-point pixels and image spatial location index from the hotspot region locations to obtain the hotspot pixel location set;
[0027] S412: Extract the reflectance value of each pixel in multiple bands from the set of hot spot pixel locations, obtain the spectral features in the 3.3μm, 6.2μm and 9.6μm band channels, construct a continuous spectral reflectance sequence and encapsulate it as an input tensor to obtain the spectral feature input sequence;
[0028] S413: Input the spectral feature input sequence into the convolutional neural network classification model, call the convolutional extraction layer and the fully connected discriminant layer to identify and compare the spectral angle structure, output the corresponding VOCs compound name label, and obtain the component identification result.
[0029] As a further aspect of the present invention, the steps for obtaining the pollution source tracing list are specifically as follows:
[0030] S511: Based on the component identification results, extract the VOCs type and relative abundance information, and merge similar items according to the standard format of component names to obtain the VOCs component feature set;
[0031] S512: Based on the VOCs component feature set, obtain the spatial range marked in the pollution phase change label, and extract the species information and coordinate data of all emission points in the target area from the spatial emission database to construct a regional emission source list;
[0032] S513: Based on the regional emission source list and the VOCs component feature set, the Bayesian inversion method is called to calculate the posterior probability value of the component corresponding to each emission point, the probability results are sorted, and a pollution source tracing list is generated.
[0033] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0034] In this invention, a spatiotemporal linkage screening method is used to screen VOCs concentration exceeding the threshold by constructing a concentration anomaly response task flow, effectively forming highly sensitive markers for pollution responses. Based on the joint analysis of three sets of indicators—concentration temporal difference, spatial standard deviation, and local volatility—standardization processing and jump trend identification methods are used to achieve multidimensional dynamic discrimination of pollution phase transition phenomena, thereby enhancing the response capability to sudden or covert emission events. The boundaries of suspected pollution areas in remote sensing images are finely characterized by gray-level gradient change trends, and their stability is judged by combining temporal displacement and local gradient density changes, effectively eliminating artifact disturbance factors and strengthening the robustness of pollution source boundary identification. In hyperspectral pixel selection, hotspot areas are focused based on brightness enhancement thresholds to ensure that the acquired spectral data has typical characteristics. A spectral input sequence is constructed using multi-band reflectance, and VOCs types and abundances are identified by combining deep convolutional structures, effectively improving the accuracy of component identification and anti-interference capability. In the component-emission source association linking link, a spatially constrained emission database is introduced, and the emission source list is probabilistically ranked in the form of posterior probability evaluation, so that the source tracing results have significant data support and quantitative reliability. Attached Figure Description
[0035] Figure 1 This is a schematic diagram of the main steps of the present invention;
[0036] Figure 2 This is a flowchart of step S1 of the present invention;
[0037] Figure 3 This is a flowchart of step S2 of the present invention;
[0038] Figure 4 This is a flowchart of step S3 of the present invention;
[0039] Figure 5 This is a flowchart of step S4 of the present invention;
[0040] Figure 6 This is a flowchart of step S5 of the present invention. Detailed Implementation
[0041] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0042] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0043] Please see Figure 1 This invention provides a technical solution: a deep learning-based hyperspectral-VOCs mobile source tracing method, comprising the following steps:
[0044] S1: Obtain the VOCs concentration response sequence of the monitoring area through a mobile monitoring vehicle, filter out abnormal VOCs concentrations, and construct a concentration anomaly response task flow;
[0045] S2: Obtain three sets of indicators from the concentration anomaly response task flow: concentration temporal difference parameters, spatial distribution standard deviation, and local concentration fluctuation anomaly rate, and perform standardization processing. Based on the three sets of indicators, determine the pollution phase transition point and obtain the pollution phase transition judgment label.
[0046] S3: Obtain remote sensing images of the pollution phase change determination label monitoring area corresponding to the time period, select the boundary location of the suspected pollution source area from the remote sensing image and judge the gradient change of the location, and output the boundary enhancement response map.
[0047] S4: Obtain hyperspectral fixed-point pixels of suspected pollution source areas in the boundary enhancement response map, and input them into the convolutional neural network model to identify VOCs compound components and obtain component identification results;
[0048] S5: Based on the component identification results, the correspondence between emission sources and VOCs components is calculated using the Bayesian inversion method. The sources are then sorted and filtered according to the probability values to obtain a list of pollution sources.
[0049] The concentration anomaly response task flow includes the time of the concentration anomaly, the magnitude of the anomaly concentration, and the duration of the anomaly. The pollution phase change determination label includes the phase change occurrence time, the spatial location of the phase change, and the phase change type. The boundary enhancement response map specifically includes the boundary location coordinates, the gradient change direction, and the boundary continuity index. The component identification results include the VOCs species type, the reflectance feature vector, and the relative abundance value. The pollution source tracing list specifically refers to the emission source coordinates, the emission source type, and the posterior probability value.
[0050] Please see Figure 2The specific steps for obtaining the concentration anomaly response task flow are as follows:
[0051] S111: The VOCs concentration response sequence of the monitoring area is obtained by mobile monitoring vehicle, the continuous TVOCs concentration data stream of single mass spectrometry equipment at a frequency of one second is collected, and the continuously collected TVOCs concentration values are arranged in time order to generate a concentration time series.
[0052] When acquiring VOCs concentration response sequences in a monitoring area using a mobile monitoring vehicle, continuous monitoring along a fixed route is required. The onboard single-mass spectrometer collects the total VOCs concentration value in the air once per second, recording three items during sampling: sampling time, sampling location, and TVOCs value. The sampling device automatically samples on a second-by-second basis, generating a concentration data stream corresponding to the duration and frequency. After receiving the sampling information for each second, the device first arranges it in chronological order and stores it in a buffer, forming a sampling time axis, with each time point corresponding to a sampling concentration value. During continuous travel, for example, if a monitoring task is set to travel for 30 minutes (1800 seconds), then 1800 sets of concentration response data will be collected. This data structure is maintained using a linear index, with time point indices and concentration values recorded in pairs, such as 1 second corresponding to 1.03 ppm, 2 seconds corresponding to 0.96 ppm, etc. After the initial sampling is completed, the sequence is organized into a one-dimensional time series for subsequent screening operations, resulting in a time series that represents the mapping relationship between TVOCs concentration and time.
[0053] S112: Based on the TVOCs values at each time point in the concentration time series, using the 1ppm concentration threshold as the judgment benchmark, record the time index of all matching judgment benchmarks, and obtain the concentration index sequence exceeding the threshold.
[0054] After the concentration time series is established, the system will perform conditional filtering based on the set concentration judgment threshold. In this embodiment, the judgment benchmark is set to 1 ppm, meaning that if the concentration value per second is greater than 1 ppm, it is judged as an anomaly. The system traverses all data points in the concentration time series and performs individual comparisons one by one, using the logic of "if the current value > 1 ppm, then it is an anomaly" for judgment. Time points that meet the conditions will have their indices in the time series recorded, while retaining the corresponding original concentration values. During system execution, if the concentration values collected at positions such as 12 seconds, 30 seconds, and 85 seconds are 1.14 ppm, 1.23 ppm, and 1.07 ppm respectively, these positions will be marked as anomalies. All anomaly point indices will be summarized into a sequence structure and passed to the next stage of analysis. It should be noted that, to eliminate sampling error interference, the system requires at least three anomaly points to appear within a continuous time interval to constitute a valid segment; if individual anomaly points exist in isolation, they will be removed. The final record retains all sampling point indices that are higher than the judgment benchmark value and have the condition of continuity, forming an over-threshold concentration index sequence.
[0055] S113: Based on the time points recorded in the over-threshold concentration index sequence, extract the corresponding TVOCs concentration values and occurrence times, and construct a concentration anomaly response task flow;
[0056] The system extracts the concentration values and sampling times for all corresponding time points based on the selected excess-threshold concentration index sequence, and combines this with synchronously recorded location information from the vehicle's GPS to construct structured task units. Each task item includes a start time, end time, maximum concentration value, and the corresponding GPS location point on the sampling path. If consecutive points constitute an abnormal segment, the system records the start and end times of the segment and calculates the maximum concentration value within that segment as a representative value to identify the pollution intensity. To enhance the logical clarity of the task flow, the system assigns a unique number to each segment, such as Task 1, Task 2, etc., and the task information is organized as a key-value structure. For example, if a segment starts at 200 seconds and ends at 210 seconds, with a maximum concentration of 1.88 ppm, then the content of this task item would be: Start 200 seconds, End 210 seconds, Maximum value 1.88 ppm, Located at the corresponding location on the monitoring path. After all task items are summarized, a task data structure that can be called by downstream modules is formed, i.e., the concentration anomaly response task flow.
[0057] Please see Figure 3 The specific steps for obtaining the pollution phase change determination label are as follows:
[0058] S211: Obtain VOCs concentration data for the time segment corresponding to the concentration anomaly response task flow, and calculate the concentration temporal difference, spatial distribution standard deviation, and local concentration fluctuation anomaly rate to obtain three change indicators.
[0059] When processing the time segment corresponding to the concentration anomaly response task flow, three key parameters need to be calculated: concentration temporal difference, spatial distribution standard deviation, and local concentration fluctuation anomaly rate. First, the concentration temporal difference reflects the concentration change at a monitoring point between the current time and the previous time, and its calculation formula is as follows:
[0060] ΔC i (t)=C i (t)-C i (t-1);
[0061] Where: ΔC i (t): Represents the concentration difference (in ppm) at the i-th monitoring point at time t; C i (t): Represents the concentration value of the i-th monitoring point at the current time t; C i (t-1): represents the concentration value at the previous moment.
[0062] Taking actual sampling data as an example, the concentration value at t=121 seconds is [1.30, 1.15, 0.92, 1.00, 0.83] ppm, and the concentration at t=120 seconds is [1.20, 1.10, 0.95, 1.05, 0.88] ppm. The difference result is: ΔC(121) = [0.10, 0.05, -0.03, -0.05, -0.05] ppm. Secondly, the spatial distribution standard deviation measures the degree of concentration dispersion at different monitoring points within the same time slice. The calculation formula is as follows:
[0063]
[0064] Where: σ(t): standard deviation of spatial distribution at time t (unit: ppm); N: number of monitoring points; C i (t): The concentration value at the i-th monitoring point; The average concentration at all monitoring points at time t.
[0065] Taking the 121st second as an example, the concentration at the monitoring point was [1.30, 1.15, 0.92, 1.00, 0.83] ppm, with a mean of:
[0066]
[0067] Substituting into the formula, we get: σ(121)≈0.167ppm.
[0068] Third, the local concentration fluctuation anomaly rate is used to represent the proportion of monitoring points exceeding a set threshold at a certain moment, and the calculation formula is: Where: R abn (t): Local concentration fluctuation anomaly rate at time t; M tN: The number of monitoring points with a concentration value higher than the set threshold T at the current time; T: The TVOCs anomaly judgment threshold, which is set to 1.00 ppm here.
[0069] At 121 seconds, the concentration value is [1.30, 1.15, 0.92, 1.00, 0.83] ppm, which satisfies C. i There are 3 points in T, therefore: In summary, at 121 seconds, the concentration time series difference for this time slice is [0.10, 0.05, -0.03, -0.05, -0.05] ppm, with a spatial distribution standard deviation of 0.167 ppm and a local concentration fluctuation anomaly rate of 0.6. These three parameters constitute a complete three-dimensional indicator of pollution dynamics, used for subsequent standardization processing and trend change identification.
[0070] S212: The three change indicators are standardized using the Z-score method to obtain a standardized indicator sequence;
[0071] After obtaining the three change index sequences output by the preceding operations—namely, the concentration time-series difference, the spatial distribution standard deviation, and the local concentration fluctuation anomaly rate—the system performs standardization processing on each sequence to address the issues of inconsistent units and numerical scale differences among the different indicators, facilitating trend comparison and jump identification under a unified reference. The standardization process uses the Z-score method, which involves subtracting the mean of the sequence from the value at each time point in the current index sequence, and then dividing by the standard deviation of the sequence. This process maps the original data to a dimensionless numerical sequence centered at 0, making the indicators comparable. In this embodiment, the local concentration fluctuation anomaly rate has a known mean of 0.45 and a standard deviation of 0.12 in a set of samples. If the value is 0.60 at the 121st second, the standardization result is 1.25. Similarly, in the spatial distribution standard deviation sequence, if the value at a certain time is 0.19, the sequence mean is 0.14, and the standard deviation is 0.03, the standardization result is 1.67. Using the above method, the system obtains three standardized index sequences, whose structure remains consistent with the original time series. Each moment corresponds to a set of three standardized values. This structure will be used as the input for the next sliding window trend recognition module and will be uniformly named the standardized index sequence.
[0072] S213: Based on a standardized index sequence, a dual sliding time window is set to identify the jump trend of the local concentration fluctuation anomaly rate, mark the location area where the jump occurs, and generate a pollution phase change judgment label.
[0073] Based on the standardized indicator sequence described above, the system employs a sliding time window strategy to identify and classify trend changes in the sequence, thereby determining abrupt changes in pollution concentration over time and space. The system sets two window lengths: a fixed window for obtaining historical averages over a reference period, and a variable window for detecting short-term fluctuations in the indicators. For example, the fixed window length is set to 30 seconds, and the variable window length is set to 10 seconds. The two windows slide synchronously in steps of one second. At each sliding position, the standardized value of the local concentration fluctuation anomaly rate is averaged in both windows, and the difference between the means is compared. If this difference exceeds a set jump threshold, the current time point is marked as a suspected abrupt change. The jump threshold is determined based on experimental statistical analysis and is set to 1.5 in this example. If at second 310, the fixed window mean is 0.42 and the variable window mean is 2.10, the mean difference is 1.68, exceeding the set threshold of 1.5, thus indicating a jump at that moment. Simultaneously, the system uses the spatial distribution standard deviation trend at the corresponding time point for joint judgment. If the indicator increases continuously within the current time and the previous two seconds, it indicates that the pollution is spreading more spatially, further strengthening the confirmation of the abrupt change at that time point. After the judgment is confirmed, the system generates unique pollution phase change identification information based on the time and spatial location of the pollution point. The information includes fields such as abrupt change time, abrupt change coordinates, and indicator jump value, and outputs a pollution phase change judgment label.
[0074] Please see Figure 4 The specific steps for obtaining the boundary enhancement response map are as follows:
[0075] S311: Based on the pollution phase change determination label, obtain remote sensing image data of the corresponding monitoring area and time period, extract the gray level change of each pixel in the image, identify the boundary line of continuous pixels with prominent gradient changes, screen the edge area location of suspected pollution sources, and obtain the boundary candidate location set.
[0076] When acquiring remote sensing image data based on pollution phase change determination labels, the system first determines the monitoring area and corresponding time period covered by the label. It then retrieves the multispectral or hyperspectral image frame sequence of the corresponding area within that time period from the remote sensing database and arranges them chronologically to ensure alignment between the time axis and the phase change label. After image loading, the system performs grayscale gradient calculation on all pixels in each frame, calling the grayscale value difference between each pixel and its adjacent pixels (up, down, left, and right), extracting the grayscale change amplitude, and marking the direction of change to construct the gradient map of the current frame. Subsequently, the system searches for connected regions composed of pixels with continuously increasing grayscale change amplitudes in the same image, constructs boundary lines according to contour extraction logic, and retains closed structures with boundary contour lengths exceeding 10 pixels, identifying them as continuous boundary lines. After extracting all candidate boundary lines, the system determines whether the gradient amplitude of the boundary grayscale values exceeds a threshold setting (20 gray levels in this embodiment). If the average gradient value of any segment on the continuous boundary line exceeds this threshold, the area corresponding to that boundary is determined as a suspected pollution source edge area. The boundary coordinates of all eligible region pixels are aggregated to form a set of candidate boundary locations.
[0077] S312: Extract the position coordinates, gradient values and gradient point density of the boundary candidate location set in three consecutive frames of images, calculate the position offset, gradient change rate and density increase / decrease of each point in the time dimension, determine whether there is a continuous upward trend, and obtain the continuously enhanced boundary point set.
[0078] After extracting the candidate boundary locations, the system compares them sequentially in three adjacent frames according to their time index. The spatial position of each boundary point in the three consecutive frames is marked as a point coordinate sequence. The system calculates the position offset based on this sequence. If the coordinate position of a point at three time points shows a unidirectional movement trend and the cumulative offset exceeds two pixels, it is determined that the position change is continuous. At the same time, the system obtains the gradient value corresponding to the point in each frame and calculates the gradient change rate. Using the gradient value of the first frame as a benchmark, the system observes the direction of change and the increase ratio of the gradient value in the subsequent two time points. If the gradient value increases continuously and the increase is greater than 20%, it is marked as a point with strong gradient change. To further describe the boundary intensity structure, the system establishes a 3×3 pixel sliding window in the neighborhood of each boundary point, counts the number of pixels in the window marked as boundary points, calculates the local gradient density value, and analyzes the magnitude of the change of this density in the three frames. If the density increases continuously and the increment is not less than one pixel, it indicates that the local structure is in an enhanced state. For a boundary point that simultaneously satisfies the conditions of significant positional offset, positive gradient rate of change, and increasing density, the system determines it as a continuously enhanced boundary point, and the set of all such points constitutes the continuously enhanced boundary point set.
[0079] S313: Input the continuous enhanced boundary point set into the attention mechanism module in the convolutional neural network structure, dynamically allocate spatial weights according to the gradient strength and change trend of each point, perform response enhancement processing on the boundary region, and output the boundary enhancement response map;
[0080] The continuously enhanced boundary point set is used as the input feature set and fed into the attention mechanism module of the convolutional neural network structure. The module needs to assign spatial weights to each boundary point to adjust its response intensity during the convolution extraction process. The weight assignment is not directly based on a single gradient value, but is calculated by combining three dynamic factors: the gradient intensity G of the current frame. i Gradient rate of change R i Local gradient density increase D i All three are normalized before being introduced into the weight function, with the normalization interval unified to [0,1]. Among them, the gradient strength G... i The gradient rate R represents the proportion of the current frame pixel's grayscale gradient value to the maximum gradient value in that frame. i This indicates the degree of linear increase in density over three consecutive frames, where D represents the density increase. i This represents the ratio of the increase in the number of gradient points in the local window to the maximum value. The weighting coefficients α, β, and γ are set based on the contribution of these three indicators to boundary continuity and the stability of abrupt change detection. Empirical calculations and experimental statistics show that: gradient strength is directly related to structure localization and should have the highest weight, so α = 0.5; the gradient change rate reflects the boundary motion state and is less important for trend detection, so β = 0.3; density increase mainly characterizes the integrity of changes in the local area, and because of its high volatility, it has the lowest weight, so γ = 0.2. The weighting function is as follows:
[0081]
[0082] Taking a continuously enhanced boundary point as an example, its current frame gradient intensity is 28, the maximum gradient in the current image frame is 40, and the gradient over three frames is 20→28→35. The normalized gradient change rate is 0.75, the local density increases from 5 to 8, and the maximum density is 10. The normalized parameter is: G i =0.70, R i =0.75, D i =0.30. Substitute into the calculation: The convolutional attention weight obtained at this boundary point is 0.635. The system directly applies this weight to the activation value of the corresponding pixel in the input tensor of the convolutional layer as a factor. Pixels with higher weights express stronger responses in subsequent feature maps, while those with lower weights are suppressed. After performing the above process on all consecutive enhanced boundary points, the system completes the generation of the spatial attention map. This map is then convolved bitwise with the input image. The convolution output retains the response of the enhanced boundary and suppresses the response of the non-changed region, forming a boundary enhancement response map.
[0083] Please see Figure 5 The specific steps for obtaining the component identification results are as follows:
[0084] S411: Obtain the location of hotspot regions in the boundary enhancement response map where the brightness enhancement exceeds the boundary enhancement response threshold. Select the corresponding hyperspectral fixed-point pixels and image spatial location index from the hotspot region locations to obtain the set of hotspot pixel locations.
[0085] When identifying regions in the boundary enhancement response map where brightness enhancement exceeds a threshold, the system first scans each pixel in the map point-by-point based on its enhancement intensity value. The boundary enhancement response threshold is set to 0.6 (normalized dimensionless value), a weighted value derived from the median of the intensity distribution. Statistical experiments have confirmed its effectiveness in screening areas with strong pollution source responses. The system evaluates all pixels; if a pixel's value is greater than or equal to the threshold, its corresponding coordinates are marked as a hotspot. Between multiple hotspots, the system extracts adjacent hotspot pixels using a region clustering approach to construct hotspot regions. A minimum connected region of 3×3 pixels is set, and isolated points are excluded to determine the hotspot region boundaries. Multiple pixels are sampled at the region center and edges for spectral data extraction. Based on the pixel location of each hotspot region, the system obtains its spatial index in the image matrix and compiles it into a list structure. This set, named the hotspot pixel location set, contains the coordinate index information of all selected fixed-point pixels and their respective hotspot numbers.
[0086] S412: Extract the reflectance value of each pixel in multiple bands from the hot spot pixel location set, obtain the spectral features in the 3.3μm, 6.2μm and 9.6μm band channels, construct a continuous spectral reflectance sequence and encapsulate it as an input tensor to obtain the spectral feature input sequence;
[0087] Based on the obtained set of hotspot pixel locations, the system sequentially reads the complete band reflectance information of each pixel in the hyperspectral image and selects the reflectance values of three target band channels (3.3μm, 6.2μm, and 9.6μm) as feature dimensions. These bands are where volatile organic compounds have strong absorption characteristics in the mid-infrared region, as verified by standard band response experiments. In the actual extraction process, the system first obtains the spectral reflectance sequence of each pixel from the original hyperspectral cube image data, totaling N bands. The numerical unit of each band is the reflectance ratio (dimensionless, range 0–1). Then, the 45th, 88th, and 132nd bands are read from the N-dimensional bands, corresponding to center wavelengths of 3.3μm, 6.2μm, and 9.6μm, respectively, to construct a three-dimensional reflectance vector. The system encapsulates this vector and its spatial index position together as a single input sample. Finally, all samples form an input tensor with a data structure dimension of M×3, where M is the number of hotspot pixels, resulting in the spectral feature input sequence, which serves as the input content for subsequent model recognition.
[0088] S413: Input the spectral feature input sequence into the convolutional neural network classification model, call the convolutional extraction layer and the fully connected discriminant layer to identify and compare the spectral angle structure, output the corresponding VOCs compound name label, and obtain the component identification result;
[0089] The system inputs the spectral feature sequence into a convolutional neural network classification model. Each row of the input tensor represents the spectral reflectance vector of a hotspot pixel in a selected band, denoted as a three-dimensional vector x =
[0090] [x a ,x b ,x c ], where: x a : Indicates that the pixel is in band λ a Reflectance at (i.e., 3.3 μm); x b : Indicates that the pixel is in band λ b Reflectance at (i.e., 6.2 μm); x c : Indicates that the pixel is in band λ c The reflectance at (i.e., 9.6 μm); the first layer of the convolutional network is a one-dimensional convolutional layer, and its convolution kernel parameters are denoted as w = [w a ,w b ,w c ], corresponding to the weight parameters for each input dimension, where: w a : Convolution kernel for band λ a Response weights; w b : Convolution kernel for band λ b Response weights; w c : Convolution kernel for band λ c The response weights are set to b; the bias term is set to b, and the output of the convolutional layer is:
[0091] y = σ(w) a ·x a +w b ·x b +w c ·x c +b);
[0092] Where: y: is the activation value output by the convolutional layer; σ(·): is the activation function, using the ReLU function σ(z)=max(0,z).
[0093] Set the input value to x a =0.52, x b =0.65, x c =0.70, let the convolution kernel parameter be w a =0.4, w b =0.3, w c =0.2, bias term b = -0.1, substituting, we get:
[0094] y=σ(0.4·0.52+0.3·0.65+0.2·0.70-0.1)=σ(0.443)=0.443.
[0095] The system then uses this activation value as input to a fully connected discriminant layer for classification and recognition. Cosine similarity is used to compare the input vector with the standard reflectance vectors of known VOC components. The formula is as follows:
[0096]
[0097] Where: x = [x a ,x b ,x c ]: The reflectance vector of the input pixel; t = [t a ,t b ,t c ]: Reflectance vector of standard VOCs components; t a t b t c : These represent the standard species in band λ a , λ b , λ c The reflectance is given by S(x,t); S(x,t) is the cosine similarity between the input and the standard vector.
[0098] Let the input pixel vector be the aforementioned value, and the reflectance vector of the standard species ethylbenzene be t. a =0.50, t b =0.66, t c =0.68, then the calculation yields:
[0099]
[0100] The system sets a similarity threshold of θ = 0.92. When S(x,t) ≥ θ, the sample is determined to belong to the ethylbenzene component. For the convolution kernel weight w... a ,w b ,w c The system settings are based on the response capabilities of each band in the VOCs spectrum to component identification. Through statistical analysis of a large number of labeled samples, the following settings were determined: Band λ a (3.3 μm) is the main absorption region for most aromatic structures in VOCs, with stable and highly discriminative signals. Therefore, w is set as... a Maximum weight; band λ b (6.2μm) represents the reflectance characteristics of some ketones and alkanes, with moderate signal. Let w b Secondly; band λ c (9.6μm) is more sensitive to heterocyclic molecules, resulting in a lower signal-to-noise ratio. Let w c The weight is minimized. In summary, all input samples undergo convolution calculation and classification recognition according to the above process, ultimately forming VOCs compound component identification results corresponding to the hotspot pixel location set, which can be used for subsequent pollution source inversion.
[0101] Please see Figure 6 The specific steps for obtaining the pollution source tracing list are as follows:
[0102] S511: Based on the component identification results, extract the VOCs type and relative abundance information, and merge similar items according to the standard format of component names to obtain the VOCs component feature set;
[0103] Based on the component identification results, the types and relative abundance information of the contained VOCs are extracted. First, according to the real-time spectral data collected by the detection instrument, the characteristic absorption peaks of the atmospheric sample in the target area are analyzed. The analysis process includes spectral background correction, spectral line denoising, baseline flattening, and peak identification. Background correction is achieved by introducing an environmental reference signal and the original signal point by point. In the denoising stage, weak perturbation bands are removed according to a set amplitude threshold. The average value of the baseline stable band is set as the normalization benchmark. Then, the positions of the identified characteristic peaks are compared one by one with the characteristic absorption positions included in the standard VOCs database. During the comparison, the wavelength difference is not more than ±0.5 nm as the matching criterion. At the same time, the peak height is also adjusted. The relative difference is calculated, and a successful match is confirmed if the relative error is less than 10%. Target components such as benzene, toluene, and cumene are matched sequentially in a set of samples. Then, the relative abundance of each component is estimated proportionally. The abundance estimation is done by normalizing the total peak area to obtain the proportion value. The corresponding proportion of each component is calculated by dividing the integral intensity by the total integral intensity. Based on this, components with the same name or similar structure are grouped into a unified category. For example, toluene and p-methylbenzene are merged into the "toluene class", benzene is retained separately, and cumene is classified into the alkylbenzene class according to its structure. This completes the initial component merging. The compound index standard is used for unified naming, and finally, a VOCs component feature set is obtained, which includes the standardized component category name and its relative abundance value.
[0104] S512: Based on the VOCs component feature set, obtain the spatial range marked in the pollution phase change label, and extract the species information and coordinate data of all emission points in the target area from the spatial emission database to construct a list of regional emission sources;
[0105] Based on the VOCs component feature set, the spatial boundary coordinate information marked in the pollution phase change label is first analyzed. A target analysis rectangle is constructed based on the four boundary points obtained, and a spatial filtering logic is set accordingly. This logic is used as the filtering condition for emission point data in the spatial emission database. A two-way interval judgment method is adopted to determine whether the latitude and longitude of each emission point simultaneously fall within the set coordinate interval. If the condition is met, the emission point is considered a valid emission point within the target area. Subsequently, the emission species information corresponding to the emission point is extracted. The extraction process is based on the species name and emission concentration or intensity value recorded in the emission database. All registered emission component information for that point is recorded and summarized by category. During the summarization process, the logic of component merging in the feature set is maintained, i.e., classified according to "toluene", "benzene", "alkylbenzene", etc., and their original emission values and relative proportions are recorded. This operation is performed sequentially on all filtered emission points to construct a regional emission source list containing geographical location, emission species type, and emission amount, ensuring that each emission point in the emission source list has a traceable location and corresponding VOCs type and emission information.
[0106] S513: Based on the regional emission source list and VOCs component feature set, the Bayesian inversion method is called to calculate the posterior probability value of the component corresponding to each emission point, the probability results are sorted, and a pollution source tracing list is generated.
[0107] Based on the regional emission source list and VOCs component feature set, the emission information of each emission point and the feature set need to be mapped at the element level first. Specifically, the emission species list of each emission point is used as the basic dataset, and cross-referencing is performed with the standardized and merged component list in the feature set. For component items that appear in both sets, their emission intensity values in the emission point are extracted, and their relative proportions are calculated. Let the emission intensity of the i1th component in the emission point be denoted as . Total emission intensity is The standardized emission intensity of this component is then: Simultaneously extract the relative abundance of corresponding components in the feature set. This forms the core parameter pair involved in the posterior probability calculation. Then, a posterior probability value is calculated for each emission point. The posterior probability value represents the likelihood that the emission point is a pollution source. The calculation formula is as follows:
[0108]
[0109] in, The posterior probability value of emission point j1 represents the relative likelihood that this point is a pollution source; The emission intensity of the i1th component at emission point j1; The sum of the emission intensities of all components at emission point j1; The relative abundance of the i1th component in the feature set; Weighting coefficients are used to enhance the role of certain key components in the overall matching; Z: normalization constant, which makes the sum of the posterior probabilities of all emission points equal to 1.
[0110] Among them, the weighting coefficient The following criteria were used to set the weights: Based on pollution event data from key national monitoring stations over the past ten years, the frequency of occurrence of various VOCs components as the dominant species in major pollution events was statistically analyzed. Components with a frequency higher than 20% were assigned a weight of 1.2, those with a frequency between 10% and 20% were assigned a weight of 1.1, and those with a frequency lower than 10% were assigned a weight of 1.0. Toluene components were assigned a weight of 1.2 due to a frequency of 26.3%, benzene components were assigned a weight of 1.1 due to a frequency of 18.7%, and alkylbenzenes were assigned a weight of 1.0 due to a frequency of 9.1%. This setting ensures that the weights are supported by historical statistics and avoids subjective assumptions.
[0111] Let the relative abundances of toluene, benzene, and alkylbenzene compounds in the feature set be 0.324, 0.287, and 0.195, respectively, and their corresponding weights be 1.2, 1.1, and 1.0, respectively. Let the emission intensities of these three components at emission point j1 = B be 17.9 mg / s, 21.3 mg / s, and 14.2 mg / s, respectively, with a total emission intensity of 53.4 mg / s. Then:
[0112] Toluene-based standardized strength is The error is |0.335-0.324|=0.011, and the weighted error is 0.011×1.2=0.0132;
[0113] The normalized strength of benzene is The error is |0.399-0.287|=0.112, and the weighted error is 0.112×1.1=0.1232;
[0114] The normalized strength of alkylbenzenes is The error is |0.266-0.195|=0.071, and the weighted error is 0.071×1.0=0.071.
[0115] The sum of the exponent terms is:
[0116]
[0117] Substituting further into the formula, the unnormalized posterior probability of emission point j1 = B is:
[0118]
[0119] After completing the above operations at all emission points, for all Summing and normalizing, for example, given three emission points A, B, and C with unnormalized values of 0.813, 0.785, and 0.642 respectively, the normalization constant Z = 0.813 + 0.785 + 0.642 = 2.240. The final posterior probability of emission point B is:
[0120]
[0121] The posterior probability values calculated for all emission points are sorted in descending order to form a pollution source tracing list. This list is sorted in descending order of the posterior probability of each emission point, representing its probability level as the source of the target pollution event. This result is used in the key emission source identification stage of subsequent source tracing and emergency response strategies to complete the source tracing closed loop from identifying characteristic components to the corresponding emission points.
[0122] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A deep learning-based hyperspectral-VOCs mobile source tracing method, characterized in that, Includes the following steps: S1: Obtain the VOCs concentration response sequence of the monitoring area through a mobile monitoring vehicle, filter out abnormal VOCs concentrations, and construct a concentration anomaly response task flow. The concentration anomaly response task flow includes the concentration anomaly time, the abnormal concentration amplitude, and the anomaly duration. S2: Obtain three sets of indicators from the concentration anomaly response task flow: concentration time-series difference parameters, spatial distribution standard deviation, and local concentration fluctuation anomaly rate, and perform standardization processing. Based on the three sets of indicators, determine the pollution phase transition point and obtain the pollution phase transition judgment label. The pollution phase transition judgment label includes the phase transition occurrence time, phase transition spatial location, and phase transition type. S3: Obtain the remote sensing image of the monitoring area corresponding to the pollution phase change judgment label during the corresponding time period, select the boundary position of the suspected pollution source area from the remote sensing image and judge the gradient change of the position, and output the boundary enhancement response map. The boundary enhancement response map specifically includes the boundary position coordinates, gradient change direction, and boundary continuity index. S4: Obtain hyperspectral fixed-point pixels of the suspected pollution source area in the boundary enhancement response map, and input them into the convolutional neural network model to identify VOCs compound components and obtain the component identification results; S5: Based on the component identification results, the correspondence between emission sources and VOCs components is calculated using the Bayesian inversion method, and the sources are sorted and filtered according to the probability values to obtain a pollution source tracing list.
2. The deep learning-based hyperspectral-VOCs mobile source tracing method according to claim 1, characterized in that, The component identification results include VOC species type, reflectance feature vector, and relative abundance value, while the pollution source tracing list specifically refers to emission source coordinates, emission source type, and posterior probability value.
3. The deep learning-based hyperspectral-VOCs mobile source tracing method according to claim 1, characterized in that, The specific steps for obtaining the concentration anomaly response task flow are as follows: S111: The VOCs concentration response sequence of the monitoring area is obtained by mobile monitoring vehicle, the continuous TVOCs concentration data stream of single mass spectrometry equipment at a frequency of one second is collected, and the continuously collected TVOCs concentration values are arranged in time order to generate a concentration time series. S112: Based on the TVOCs value at each time point in the concentration time series, using the 1ppm concentration threshold as the judgment benchmark, record the time index of all matching judgment benchmarks, and obtain the over-threshold concentration index sequence. S113: Based on the time points recorded in the above-threshold concentration index sequence, extract the corresponding TVOCs concentration values and occurrence times, and construct a concentration anomaly response task flow.
4. The deep learning-based hyperspectral-VOCs mobile source tracing method according to claim 3, characterized in that, The specific steps for obtaining the pollution phase change determination label are as follows: S211: Obtain VOCs concentration data for the time segment corresponding to the concentration anomaly response task flow, and calculate the concentration temporal difference, spatial distribution standard deviation, and local concentration fluctuation anomaly rate to obtain three change indicators. S212: The three change indicators are standardized using the Z-score method to obtain a standardized indicator sequence; S213: Based on the standardized index sequence, set a double sliding time window, identify the jump trend of the local concentration fluctuation anomaly rate, mark the location area where the jump occurs, and generate a pollution phase change judgment label.
5. The deep learning-based hyperspectral-VOCs mobile source tracing method according to claim 4, characterized in that, The specific steps for obtaining the boundary enhancement response map are as follows: S311: Based on the pollution phase change determination label, obtain remote sensing image data of the corresponding monitoring area and time period, extract the gray level change of each pixel in the image, identify the boundary line of continuous pixels with prominent gradient changes, screen the edge area location of suspected pollution sources, and obtain the boundary candidate location set. S312: Extract the position coordinates, gradient values and gradient point density of the boundary candidate position set in three consecutive frames of images, calculate the position offset, gradient change rate and density increase / decrease of each point in the time dimension, determine whether there is a continuous upward trend, and obtain a continuously enhanced boundary point set. S313: Input the continuous enhanced boundary point set into the attention mechanism module in the convolutional neural network structure, dynamically allocate spatial weights according to the gradient strength and change trend of each point, perform response enhancement processing on the boundary region, and output the boundary enhancement response map.
6. The deep learning-based hyperspectral-VOCs mobile source tracing method according to claim 5, characterized in that, The specific steps for obtaining the component identification results are as follows: S411: Obtain the location of hotspot regions in the boundary enhancement response map where the brightness enhancement exceeds the boundary enhancement response threshold, and select the corresponding hyperspectral fixed-point pixels and image spatial location index from the hotspot region locations to obtain the hotspot pixel location set; S412: Extract the reflectance value of each pixel in multiple bands from the set of hot spot pixel locations, obtain the spectral features in the 3.3μm, 6.2μm and 9.6μm band channels, construct a continuous spectral reflectance sequence and encapsulate it as an input tensor to obtain the spectral feature input sequence; S413: Input the spectral feature input sequence into the convolutional neural network classification model, call the convolutional extraction layer and the fully connected discriminant layer to identify and compare the spectral angle structure, output the corresponding VOCs compound name label, and obtain the component identification result.
7. The deep learning-based hyperspectral-VOCs mobile source tracing method according to claim 6, characterized in that, The specific steps for obtaining the pollution source tracing list are as follows: S511: Based on the component identification results, extract the VOCs type and relative abundance information, and merge similar items according to the standard format of component names to obtain the VOCs component feature set; S512: Based on the VOCs component feature set, obtain the spatial range marked in the pollution phase change label, and extract the species information and coordinate data of all emission points in the target area from the spatial emission database to construct a regional emission source list; S513: Based on the regional emission source list and the VOCs component feature set, the Bayesian inversion method is called to calculate the posterior probability value of the component corresponding to each emission point, the probability results are sorted, and a pollution source tracing list is generated.
Citation Information
Patent Citations
Pollution source simulation and prediction method based on multistage verification
CN118709388A
Method for realizing VOCS inversion of satellite remote sensing column concentration by using machine learning
CN119903362A