Atmospheric pollutant transmission trajectory clustering analysis method and system

By calculating the toluene/benzene concentration ratio and spherical distance, a semantically connected set is constructed to generate a cluster of atmospheric pollutant transport trajectories. This solves the problem of misleading pollutant transport paths in existing technologies and achieves more accurate pollution source tracing.

CN121980291APending Publication Date: 2026-05-05BEIJING ZHONGYI ENVIRONMENTAL TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING ZHONGYI ENVIRONMENTAL TECHNOLOGY CO LTD
Filing Date
2026-01-27
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies for analyzing atmospheric pollutant transport rely on real-time concentration readings from ground monitoring stations and meteorological parameters for trajectory backtracking. This makes it difficult to accurately distinguish the mixed results of multiple pollution sources, and errors are easily generated under complex wind field conditions, leading to misleading pollutant transport paths.

Method used

By calculating the toluene/benzene concentration ratio and combining the spherical distance and ratio difference, a semantic connected set is constructed to generate a cluster of atmospheric pollutant transport trajectories. Cosine similarity is then used to match the source component spectrum to pinpoint the pollution source.

Benefits of technology

It improves the accuracy of tracing pollutant transport paths, eliminates interference points that are physically close but have different compositions, solves the problem of multi-source mixing based on the inherent composition law, and gets rid of dependence on meteorological parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121980291A_ABST
    Figure CN121980291A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of GIS (Geographic Information System), in particular to an atmospheric pollutant transmission trajectory clustering analysis method and system, which comprises the following steps of: extracting sampling longitude and latitude time and toluene concentration, calculating a ratio to construct trajectory points, selecting core points to screen spatial neighborhoods, calculating a ratio difference and performing threshold filtering to form semantic neighborhoods; according to the method, the chemical fingerprints are constructed by extracting the concentration of methylbenzene and the concentration of benzene and calculating the ratio, concentration fluctuation interference caused by atmospheric dilution is avoided, double constraints are implemented by combining the spherical distance and the ratio difference value, interference points which are physically adjacent but have different components are removed, and the pollution source is determined. The method comprises the following steps: constructing a semantic connected set according to a difference sequence to generate a track cluster, reconstructing a homologous transmission path, calculating an in-cluster mean value, matching a source spectrum by using cosine similarity, quantifying component similarity to lock a pollution source, solving the problem of multi-source mixing by using an internal component rule, getting rid of meteorological parameter dependence, and improving traceability accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of GIS technology, and in particular to a method and system for cluster analysis of atmospheric pollutant transport trajectories. Background Technology

[0002] The field of spatiotemporal trajectory data management and analysis in Geographic Information Systems (GIS) involves the storage, management, and analysis of continuous location data of moving objects changing over time within geographic space. This field collects spatiotemporal sequence data from Global Positioning System coordinates, remote sensing images, and ground monitoring stations, and uses spatiotemporal indexing and trajectory compression techniques to establish a database system that reflects the movement patterns and spatial distribution characteristics of objects. Traditional atmospheric pollutant transport trajectory clustering analysis methods utilize real-time concentration values ​​collected by environmental monitoring units combined with wind force and direction parameters obtained from meteorological monitoring equipment for correlation analysis. Addressing the issue of ozone and other pollutants being susceptible to airflow and drifting, leading to discrepancies between monitoring points and actual volatilization sources, existing technologies rely on readings from fixed ground monitoring stations, combined with atmospheric diffusion models, to infer the diffusion paths of pollutants under different meteorological conditions, and estimate the initial volatilization area of ​​ozone by comparing wind field data in historical meteorological databases.

[0003] Current technologies for analyzing atmospheric pollutant transport primarily rely on real-time concentration readings from ground-based monitoring stations combined with wind speed and direction parameters obtained from meteorological monitoring equipment. This over-reliance on external meteorological parameters and atmospheric diffusion models often leads to cumulative errors when dealing with pollutants like ozone that are susceptible to long-distance migration due to airflow limitations, such as spatiotemporal resolution limitations of wind field data or local microclimate disturbances. Simply relying on concentration gradients and wind direction for trajectory backtracking can easily lead to incorrect association of high-concentration points that are spatially close but not from the same source, ignoring the objective fact that different pollution sources converge and overlap in the same area. Without effective identification of the intrinsic chemical composition of pollutants, physical diffusion laws alone cannot distinguish whether the data from monitoring points originates from direct transport from a single emission source or is a secondary result of mixing multiple emission sources. When wind speed and direction change rapidly in the actual environment or when complex terrain obstructs the path, estimation models based on historical meteorological databases often fail to accurately reconstruct the true transport path, resulting in significant spatial offsets in the location of the initial volatile matter area. Using only spatial distribution and concentration values ​​as the basis for clustering is highly susceptible to interference from random high-concentration noise points, resulting in the generated transmission trajectories containing a large amount of pseudo-correlation data. This makes it difficult to truly reflect the actual movement patterns and spatial distribution characteristics of specific pollutants, thus misleading subsequent pollution control decisions. Summary of the Invention

[0004] To address the technical problems existing in the prior art, embodiments of the present invention provide a method for clustering analysis of atmospheric pollutant transport trajectories, comprising the following steps: S1: Obtain discrete sampling records, extract longitude values, latitude values, sampling timestamps, toluene concentration values, and benzene concentration values, calculate the toluene / benzene concentration ratio, and construct multidimensional trajectory data points; S2: Select a core capture object from the multidimensional trajectory data points, calculate the spatial physical distance between the core capture object and the data points, filter data points whose spatial physical distance is less than the neighborhood radius, and construct a spatial neighborhood candidate set; S3: Obtain the toluene / benzene concentration ratio of the spatial neighborhood candidate set elements and the core capture object, calculate the absolute difference between the toluene / benzene concentration ratios between the core capture object and the elements, and generate a feature ratio difference sequence. S4: Compare the values ​​in the feature ratio difference sequence with the preset feature difference threshold, remove candidate points in the spatial neighborhood candidate set whose values ​​are greater than the feature difference threshold, retain candidate points whose values ​​are less than or equal to the feature difference threshold, and construct a semantically connected neighbor set. S5: When the number of semantically connected neighbor sets meets the minimum number of included points, merge the core capture object and the semantically connected neighbor set to construct an atmospheric pollutant transport trajectory cluster, calculate the arithmetic mean of the toluene / benzene concentration ratio in the atmospheric pollutant transport trajectory cluster, compare the arithmetic mean with the source component spectrum based on cosine similarity, and output the pollution source classification result.

[0005] As a further aspect of the present invention, the multidimensional trajectory data points include longitude values, latitude values, sampling timestamps, and toluene / benzene concentration ratios; the spatial neighborhood candidate set includes core capture objects and neighboring data points whose spatial physical distance is less than the neighborhood radius; the feature ratio difference sequence includes the absolute difference in toluene / benzene concentration ratios between the core capture objects and the candidate points; the semantically connected neighbor set includes retained candidate points whose feature difference values ​​meet the feature difference threshold requirements; and the pollution source classification result includes pollution source categories that are consistent with the source component spectrum.

[0006] As a further aspect of the present invention, the core capture object is the currently processed data point selected sequentially in the sampling time order among the multi-dimensional trajectory data points; The spatial physical distance is a spherical distance calculated based on longitude and latitude values; The neighborhood radius is a pre-set spatial distance threshold used to limit data points that are geographically adjacent to the core capture object.

[0007] As a further aspect of the present invention, the characteristic difference threshold is a pre-set upper limit for the difference in the toluene / benzene concentration ratio; The feature difference threshold is determined based on the statistical distribution of the toluene / benzene concentration ratio in historical pollution source sample data. When the absolute difference between the toluene / benzene concentration ratio of the core capture object and the data points in the spatial neighborhood candidate set is less than or equal to the feature difference threshold, it is determined that the two have connectivity in the semantics of pollution sources.

[0008] As a further aspect of the present invention, the specific steps of S1 are as follows: S101: Collect discrete sampling records through atmospheric environmental monitoring stations, perform field parsing and data extraction operations on the records, extract longitude values, latitude values, sampling timestamps, toluene concentration values, and benzene concentration values, standardize and integrate the extracted data, and generate an environmental discrete sampling feature set. S102: Call the environmental discrete sampling feature set to obtain the toluene concentration value and the benzene concentration value, calculate the concentration ratio characterizing the chemical components of pollutants, and associate and anchor the calculated value with the original data record index to generate the toluene / benzene concentration ratio. S103: Based on the toluene / benzene concentration ratio, retrieve the corresponding longitude, latitude, and sampling timestamp within the environmental discrete sampling feature set, establish a mapping relationship between the concentration ratio and the spatiotemporal coordinate dimension, and restructure and encapsulate the data according to the mapping relationship to construct multidimensional trajectory data points.

[0009] As a further aspect of the present invention, the specific steps of S2 are as follows: S201: Traverse the set of multi-dimensional trajectory data points, select the data item to be processed as the core capture object, synchronously index the multi-dimensional trajectory data points in the set, extract the geographical coordinate values ​​of both parties and establish the corresponding relationship, and generate a set of coordinate pairings for the core capture object. S202: Call the coordinate pairing set of the core capture object, and perform spherical distance calculation operation on the core capture object and multi-dimensional trajectory data points according to the spatial distribution characteristics under the geographic coordinate system, quantify the spatial interval between data points, and generate a spatial physical distance measurement sequence. S203: Based on the spatial physical distance measurement sequence, obtain a preset neighborhood radius threshold, compare the distance measurement value with the neighborhood radius threshold, filter data items that meet the spatial range constraints, aggregate and reorganize the trajectory data points that meet the conditions, and construct a spatial neighborhood candidate set.

[0010] As a further aspect of the present invention, the specific steps of S3 are as follows: S301: Based on the index identifiers included in the spatial neighborhood candidate set, backtrack to retrieve the toluene / benzene concentration ratio, extract the concentration ratio data of the core capture object and the candidate point respectively, establish a correspondence between the ratio data of the core object and the ratio data of the candidate point, and generate a core candidate ratio pairing set; S302: Call the core candidate ratio pairing set, perform numerical difference measurement calculation on the pairing data, quantify the degree of deviation between the core capture object and the candidate point in the chemical pollutant composition, and generate a concentration ratio deviation measurement. S303: Based on the concentration ratio deviation metric, the deviation metric values ​​are serialized, recombined, and encapsulated according to the original arrangement order of data points within the spatial neighborhood candidate set to construct a feature ratio difference sequence.

[0011] As a further aspect of the present invention, the specific steps of S4 are as follows: S401: Call the feature ratio difference sequence, introduce a preset feature difference threshold as a screening benchmark, perform threshold discrimination operation on the difference values ​​in the sequence, define whether the value fluctuation is within the value range defined by the threshold, generate a conformity mark for the element position that meets the judgment condition, construct a screening algorithm based on dynamic weight, and generate a candidate point conformity evaluation table. S402: Based on the candidate point conformity evaluation table, perform index mapping and data cleaning on the spatial neighborhood candidate set, remove candidate points marked as invalid in the logical mask, lock and extract trajectory data items whose difference values ​​meet the threshold constraints, and queue up and reorganize the filtered data entities to generate a homogeneity retention point set. S403: Based on the homogeneity retention point set, perform attribute verification on the remaining data points that have been verified by both spatial and chemical attributes, structure and integrate the trajectory points with homogeneous characteristics, establish the association set of the core capture object under the target semantic rules, and construct the semantically connected neighbor set.

[0012] As a further aspect of the present invention, the specific steps of S5 are as follows: S501: Call the semantically connected neighbor set, count the total number of data items included in the set, introduce a preset minimum included point threshold as the reference boundary for density clustering determination, perform numerical comparison and verification between the count and the minimum included point threshold, after verifying that the count meets the density constraint, extract the core capture object data, and perform a set merging operation with the trajectory data items in the neighbor set to integrate spatially adjacent and attribute-similarity data entities to generate an atmospheric pollutant transmission trajectory cluster; S502: Based on the atmospheric pollutant transport trajectory cluster, traverse all trajectory data points encapsulated within the cluster, extract the toluene to benzene concentration ratio associated with the data points, construct a chemical feature numerical group to be processed, perform arithmetic mean operation on the numerical group, quantify the overall concentration trend of the cluster in chemical composition, establish a central metric index reflecting the chemical characteristics of the transport trajectory, and generate the average concentration ratio within the cluster. S503: Call the mean concentration ratio within the cluster to obtain the source component spectrum including standard feature vectors of multiple pollution sources, map the mean within the cluster and the standard vectors in the source component spectrum to the same feature space, perform numerical difference measurement, calculate the difference between the mean within the cluster and the standard vectors in the source component spectrum, determine the category of pollution source based on the difference value, match the corresponding pollution source label information, and generate pollution source classification results.

[0013] An atmospheric pollutant transport trajectory clustering analysis system, comprising: The trajectory point construction module is used to execute S1: obtain discrete sampling records containing atmospheric environmental monitoring data through atmospheric environmental monitoring stations, extract longitude values, latitude values, sampling timestamps, toluene concentration values, and benzene concentration values ​​from the discrete sampling records, divide the toluene concentration value by the benzene concentration value to calculate the toluene / benzene concentration ratio that characterizes the chemical composition of pollutants, establish the mapping relationship between the toluene / benzene concentration ratio and the longitude values, latitude values, and sampling timestamps, and construct multidimensional trajectory data points; The spatial neighborhood screening module is used to execute S2: traverse the set of multi-dimensional trajectory data points to select the core capture object to be processed, call the spherical distance calculation formula to calculate the spatial physical distance between the core capture object and the other multi-dimensional trajectory data points in the set in the geographic coordinate system, compare the spatial physical distance with the preset neighborhood radius, filter data points whose spatial physical distance is less than the neighborhood radius, and construct a spatial neighborhood candidate set. The feature difference calculation module is used to execute S3: obtain the toluene / benzene concentration ratio of the core capture object and the candidate points in the spatial neighborhood candidate set respectively, calculate the absolute difference of the toluene / benzene concentration ratio between the core capture object and each candidate point, and generate a feature ratio difference sequence corresponding to the spatial neighborhood candidate set; The semantic connectivity construction module is used to perform S4: compare each value in the feature ratio difference sequence with a preset feature difference threshold, remove candidate points whose corresponding values ​​are greater than the feature difference threshold from the spatial neighborhood candidate set, retain candidate points whose corresponding values ​​are less than or equal to the feature difference threshold, and construct a semantic connectivity neighborhood set. The trajectory cluster source resolution module is used to execute S5: count the total number of elements in the semantically connected neighbor set. When the total number of elements meets the preset minimum number of included points, merge the core capture object and the semantically connected neighbor set to generate an atmospheric pollutant transport trajectory cluster. Calculate the arithmetic mean of the toluene / benzene concentration ratio of all data points in the atmospheric pollutant transport trajectory cluster. Compare the arithmetic mean of the concentration ratios within the cluster with the standard feature values ​​in the source component spectrum to measure the numerical difference. Calculate the absolute difference and perform normalization to output the pollution source classification result.

[0014] Compared with the prior art, the advantages and positive effects of the present invention are as follows: In this invention, a chemical fingerprint is constructed by extracting the concentration ratio of toluene and benzene to avoid concentration fluctuation interference caused by atmospheric dilution. The invention also implements dual constraints by combining spherical distance and ratio difference to eliminate interference points that are physically close but have different compositions. Based on the difference sequence, a semantic connected set is constructed to generate trajectory clusters, reconstructing the homogeneous transmission path. The mean within the cluster is calculated and the source spectrum is matched using cosine similarity. The composition similarity is quantified to lock the pollution source. The invention solves the problem of multi-source mixing by utilizing the inherent composition law, gets rid of dependence on meteorological parameters, and improves the accuracy of source tracing. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0016] Figure 1 This is a schematic diagram of the steps of the present invention; Figure 2 This is a detailed schematic diagram of S1 of the present invention; Figure 3 This is a detailed schematic diagram of S2 of the present invention; Figure 4 This is a detailed schematic diagram of S3 of the present invention; Figure 5 This is a detailed schematic diagram of S4 of the present invention; Figure 6 This is a detailed schematic diagram of S5 of the present invention; Figure 7 This is a system module diagram of the present invention. Detailed Implementation

[0017] The technical solution of the present invention will now be described with reference to the accompanying drawings.

[0018] To make the technical problems, technical solutions and advantages of the present invention clearer, a detailed description will be given below in conjunction with the accompanying drawings and specific embodiments.

[0019] Please see Figure 1 This invention provides a method for cluster analysis of atmospheric pollutant transport trajectories, comprising the following steps: S1: Obtain discrete sampling records through atmospheric environmental monitoring stations, extract longitude, latitude, sampling timestamp, toluene concentration, and benzene concentration from the discrete sampling records, divide the toluene concentration by the benzene concentration to calculate the toluene / benzene concentration ratio, which characterizes the chemical composition of pollutants, establish the mapping relationship between the toluene / benzene concentration ratio and the longitude, latitude, and sampling timestamp, and construct multidimensional trajectory data points; S2: Traverse the set consisting of multiple multidimensional trajectory data points to select the core capture object to be processed, call the spherical distance calculation formula to calculate the spatial physical distance between the core capture object and the other multidimensional trajectory data points in the set in the geographic coordinate system, compare the spatial physical distance with the preset neighborhood radius (Eps), filter the data points whose spatial physical distance is less than the neighborhood radius (Eps), and construct a spatial neighborhood candidate set; S3: Obtain the toluene / benzene concentration ratio of the core capture object and the candidate points in the spatial neighborhood candidate set respectively, calculate the absolute difference between the toluene / benzene concentration ratio between the core capture object and the candidate points, and generate a feature ratio difference sequence corresponding to the spatial neighborhood candidate set. S4: Compare each value in the feature ratio difference sequence with the preset feature difference threshold, remove candidate points whose corresponding values ​​are greater than the feature difference threshold from the spatial neighborhood candidate set, and retain candidate points whose corresponding values ​​are less than or equal to the feature difference threshold to construct a semantically connected neighbor set. S5: Count the total number of elements in the semantically connected neighbor set. When the total number of elements meets the preset minimum number of included points (MinPts), merge the core capture object and the semantically connected neighbor set to generate an atmospheric pollutant transport trajectory cluster. Calculate the arithmetic mean of the toluene / benzene concentration ratio of all data points in the atmospheric pollutant transport trajectory cluster. Compare the arithmetic mean of the concentration ratios in the cluster with the preset source component spectrum to measure the difference. Compare the difference value between the cluster and the source component spectrum and output the pollution source classification result.

[0020] The multidimensional trajectory data points include longitude values, latitude values, sampling timestamps, and toluene / benzene concentration ratios. The spatial neighborhood candidate set includes the core capture object and neighboring data points whose spatial physical distance is less than the neighborhood radius. The feature ratio difference sequence includes the absolute difference in toluene / benzene concentration ratios between the core capture object and the candidate points. The semantically connected neighbor set includes retained candidate points whose feature difference values ​​meet the feature difference threshold requirements. The pollution source classification results include pollution source categories that are consistent with the source component spectrum.

[0021] Please see Figure 2 The specific steps of S1 are as follows: S101: Collect discrete sampling records through atmospheric environmental monitoring stations, perform field parsing and data extraction operations on the records, extract longitude values, latitude values, sampling timestamps, toluene concentration values, and benzene concentration values, standardize and integrate the extracted data, and generate an environmental discrete sampling feature set. Each record is written to the original record queue in arrival order, with each record retaining its index field. Then, the current record is parsed, sequentially locating the longitude, latitude, sampling timestamp, toluene concentration, and benzene concentration fields, extracting the field values ​​from the text as independent numerical items. Longitude and latitude are validated: longitude is limited to -180.000000 to 180.000000, and latitude is limited to -90.000000 to 90.000000. Records not meeting these criteria are added to an exception list and removed. The sampling timestamp is formatted uniformly, converting the station's upload time to Coordinated Universal Time (UTC) and retaining it to the second. Toluene and benzene concentrations are standardized in units, converting milligrams per cubic meter to micrograms per cubic meter by multiplying the milligram value by 1000, while removing null values ​​and non-numeric characters. In the example, record index 0001 extracts longitude 113.264500, latitude 23.129100, sampling timestamp 2026-01-14 08:00:00, toluene concentration 18.6 μg / m³, and benzene concentration 3.2 μg / m³. Record index 0002 extracts longitude 113.271200, latitude 23.130800, sampling timestamp 2026-01-14 08:10:00, toluene concentration 21.4 μg / m³, and benzene concentration 3.6 μg / m³. After extraction, each field is encapsulated into an environmental discrete sampling feature set in a unified order, with the index consistent with the original record index, and written into the verification status field.

[0022] S102: Call the environmental discrete sampling feature set to obtain the toluene concentration value and the benzene concentration value, calculate the concentration ratio characterizing the chemical components of pollutants, and associate and anchor the calculated value with the original data record index to generate the toluene / benzene concentration ratio. The toluene and benzene concentration values ​​are read sequentially according to the record index, and the concentration ratio is calculated for each record. To avoid abrupt changes in the ratio due to an excessively small divisor, a lower limit check is first performed on the benzene concentration. Records with benzene concentrations less than 0.10 micrograms per cubic meter are marked as uncalcifiable and removed. This lower limit is obtained by statistically analyzing the stable reading range near the instrument's detection limit. The statistical sample consists of data from stations transmitted over 30 consecutive days. The proportion of jumps in the period when the concentration is below 0.10 micrograms per cubic meter is higher than 20%, so 0.10 is set as the stable calculation lower limit. For calculable records, the toluene and benzene concentration values ​​are divided to obtain the concentration ratio, and the concentration ratio is linked and anchored to the original data record index to form an independent set of ratios. In the example, record index 0001 divides the toluene concentration of 18.6 by the benzene concentration of 3.2, resulting in a concentration ratio of 5.81; record index 0004 divides the toluene concentration of 33.8 by the benzene concentration of 5.1, resulting in a concentration ratio of 6.63; and record index 0008 divides the toluene concentration of 24.9 by the benzene concentration of 3.9, resulting in a concentration ratio of 6.38. Each record in the ratio set is simultaneously written to the toluene concentration snapshot and benzene concentration snapshot fields, ensuring that subsequent difference measurements can be retrospectively verified, and outputting the toluene / benzene concentration ratio.

[0023] S103: Based on the toluene / benzene concentration ratio, retrieve the corresponding longitude, latitude, and sampling timestamp within the discrete environmental sampling feature set, establish a mapping relationship between the concentration ratio and the spatiotemporal coordinate dimension, and restructure and encapsulate the data according to the mapping relationship to construct multidimensional trajectory data points; The system retrieves the discrete environmental sampling feature set record by record, backtracking from the record index to obtain the longitude, latitude, and sampling timestamp corresponding to each index. A mapping relationship is then established between the concentration ratio and these three spatiotemporal fields. This mapping is achieved using index-based equal-value matching. Upon successful matching, the longitude, latitude, sampling timestamp, and concentration ratio are encapsulated into a single data entity and written to the point sequence number field. Subsequently, the data entities are structurally reorganized based on the sampling timestamps. First, they are sorted in ascending order by sampling timestamp; if times are the same, they are sorted in ascending order by record index. The sorted results are then sequentially written into the multidimensional trajectory data point set. In the example, record index 0001 is encapsulated as point number 1, with longitude 113.264500, latitude 23.129100, sampling timestamp 2026-01-14 08:00:00, and concentration ratio 5.81; record index 0002 is encapsulated as point number 2, with longitude 113.271200, latitude 23.130800, sampling timestamp 2026-01-14 08:10:00, and concentration ratio 5.94. To maintain trajectory continuity, the sampling time difference between adjacent points is calculated synchronously and written into the segment marker field. A time difference greater than 3600 seconds is marked as a segment. In the example, point numbers 1 to 8 are all 600-second intervals, and the segment is marked as no segment. After encapsulation, a multi-dimensional trajectory data point set is output, providing input for subsequent spatial neighborhood filtering.

[0024] Please see Figure 3 The specific steps of S2 are as follows: S201: Traverse the set consisting of multiple multi-dimensional trajectory data points, select the data item to be processed as the core capture object, synchronously index the multi-dimensional trajectory data points in the set, extract the geographic coordinate values ​​of both parties and establish the corresponding relationship, and generate a set of coordinate pairings for the core capture object. The current data point is selected as the core capture object in ascending order of its point number, and its record index is locked as the core index. Then, the remaining data points in the index set are synchronized to generate a candidate list. The point number and record index of each candidate point are extracted from the candidate list. Longitude and latitude values ​​are extracted for both the core capture object and the candidate points. A one-to-one pairing record is established between the core longitude / latitude and the candidate longitude / latitude, forming a core capture object coordinate pairing set. Each pairing record in the coordinate pairing set is simultaneously written with the candidate point number and candidate record index to ensure that subsequent distance measurement results can locate the specific candidate point. In the example, when point number 1 is the core capture object, the core coordinates are longitude 113.264500 and latitude 23.129100; paired with candidate point number 2, the candidate coordinates are longitude 113.271200 and latitude 23.130800; paired with candidate point number 3, the candidate coordinates are longitude 113.279900 and latitude 23.127600. During the pairing set generation process, a valid coordinate marker is written. The marker is based on whether the latitude and longitude are within a valid range and are retained to 6 decimal places. If invalid, the pairing is skipped in subsequent distance measurements. After completion, the core capture object coordinate pairing set is output.

[0025] S202: Call the core capture object coordinate pairing set, and perform spherical distance calculation operation on the core capture object and multi-dimensional trajectory data points according to the spatial distribution characteristics under the geographic coordinate system, quantify the spatial interval between data points, and generate a spatial physical distance measurement sequence. Spherical distance calculation is performed for each pairing record. Before the calculation, the longitude and latitude differences are converted into meter-level distance components. The conversion is based on the calibration results of this embodiment. Around 23 degrees latitude, a longitude difference of 0.001000 degrees corresponds to 102 meters, and a latitude difference of 0.001000 degrees corresponds to 111 meters. This calibration is obtained by fitting 20 sets of known measured distances with longitude and latitude differences. For each pairing record, the longitude and latitude differences between the core and candidate are first calculated, and then converted into longitude and latitude distance components respectively. Subsequently, the two directional components are combined into a single spatial physical distance value according to the spherical distance calculation logic. In the example, the longitude difference between core point number 1 and candidate point number 2 is 0.006700 degrees, which translates to approximately 683 meters, and the latitude difference is 0.001700 degrees, which translates to approximately 189 meters. Substituting these values ​​into the synthesis logic, the resulting physical distance is approximately 709 meters. Similarly, the longitude difference between core point number 1 and candidate point number 3 is 0.015400 degrees, which translates to approximately 1571 meters, and the latitude difference is -0.001500 degrees, which translates to approximately 167 meters. Substituting these values ​​into the synthesis logic, the resulting physical distance is approximately 1580 meters. To suppress extremely small distance anomalies caused by jitter in station coordinate reporting, distances less than 5 meters are uniformly written as 5 meters. The calculation results are written into the physical distance metric sequence in ascending order of candidate point number, and this sequence is output.

[0026] S203: Based on the spatial physical distance measurement sequence, obtain the preset neighborhood radius threshold, compare the distance measurement value with the neighborhood radius threshold, filter the data items that meet the spatial range constraints, aggregate and reorganize the trajectory data points that meet the conditions, and construct a spatial neighborhood candidate set; The neighborhood radius threshold is read and filtering is performed. In this embodiment, the neighborhood radius threshold is set to 2000 meters. The threshold setting was determined through experimental verification. 240 trajectory points in the same area were selected for candidate point filtering at 1000 meters, 2000 meters, and 3000 meters respectively. The dispersion of the number of points entering the candidate set for each core object and the subsequent chemical differences was statistically analyzed. Under the condition of 2000 meters, the number of candidate points was concentrated between 4 and 12, and the dispersion of chemical differences was stable. Under the condition of 1000 meters, the number of candidate points was repeatedly lower than 4. Under the condition of 3000 meters, the dispersion of chemical differences increased significantly, so 2000 meters was fixed. The filtering action compares each distance value in the sequence with 2000 meters. If it is less than or equal to 2000 meters, it is marked as passed; if it is greater than 2000 meters, it is marked as failed. In the example, the distance between core point number 1 and candidate point number 2 is 709 meters, which is marked as passed, and the distance between core point number 1 and candidate point number 3 is 1580 meters, which is also marked as passed. All candidate points that have passed the marker are aggregated and reorganized in ascending order of candidate point number, and simultaneously written to the candidate point number list, candidate record index list, and distance list. If the number of passed points is 0, an empty set marker is written. After completion, the spatial neighborhood candidate set is output.

[0027] Please see Figure 4 The specific steps of S3 are as follows: S301: Based on the index identifiers included in the spatial neighborhood candidate set, backtrack to retrieve the toluene / benzene concentration ratio, extract the concentration ratio data of the core capture object and the candidate point respectively, establish a correspondence between the ratio data of the core object and the ratio data of the candidate point, and generate a core candidate ratio pairing set; The process reads the record indexes of core points and candidate record indexes, and backtracks to retrieve the toluene / benzene concentration ratio set. The backtracking operation uses index equality matching to extract the core concentration ratio and the concentration ratio of each candidate point. Then, a correspondence is established between the core concentration ratio and the candidate point concentration ratio, generating a core-candidate ratio pairing set. Each record in the pairing set contains the core point number, core record index, core concentration ratio, candidate point number, candidate record index, and candidate concentration ratio. In the example, core point number 1 corresponds to a concentration ratio of 5.81, candidate point number 2 corresponds to a concentration ratio of 5.94, forming one pairing record; candidate point number 3 corresponds to a concentration ratio of 5.43, forming another pairing record. To ensure the verifiability of subsequent difference calculations, the pairing set synchronously writes the ratio calculation status field. If a candidate point ratio is not calculable, an invalid mark is written to that pairing record, and it is skipped in subsequent iterations. The pairing set is stored in ascending order of candidate point number, maintaining consistency with the spatial neighborhood candidate set. Upon completion, the core-candidate ratio pairing set is output.

[0028] S302: Call the core candidate ratio pairing set, perform numerical difference measurement calculation on the pairing data, quantify the degree of deviation between the core capture object and the candidate point in the chemical pollutant composition, and generate a concentration ratio deviation measurement. Numerical difference metric calculation is performed on each paired record. The difference metric uses absolute difference logic, subtracting the core concentration ratio from the candidate concentration ratio and taking the absolute value. In the example, substituting the core concentration ratio of 5.81 and the candidate concentration ratio of 5.94 into the difference and taking the absolute value yields a deviation metric of 0.13; substituting the core concentration ratio of 5.81 and the candidate concentration ratio of 5.43 into the difference and taking the absolute value yields a deviation metric of 0.38. If the candidate concentration ratio is 7.17, the same logic is applied as with 5.81 to obtain a deviation metric of 1.36. No calculation is performed on invalid paired records, and the deviation metric is written as a negative 1. Each deviation metric is bound to the candidate point number and the candidate record index and written to the concentration ratio deviation metric set, which is sorted in ascending order of candidate point number. This provides consistent input for subsequent threshold filtering, and the deviation metric is retained to two decimal places. The concentration ratio deviation metric is then output upon completion.

[0029] S303: Based on the concentration ratio deviation measure, the deviation measure values ​​are serialized, recombined and encapsulated according to the original arrangement order of data points within the spatial neighborhood candidate set to construct a feature ratio difference sequence. The system reads the original arrangement order within the candidate set of spatial neighborhoods, sorted in ascending order by the candidate point indices. Then, it serializes and reassembles the deviation metrics according to this order, writing the candidate point indices, candidate record indices, and deviation metric values ​​into the feature ratio difference sequence one by one. In the example, if the candidate point indices are 2, 3, and 4, and the corresponding deviation metrics are 0.13, 0.38, and 0.82, then the first position in the difference sequence is written with candidate point indices 2 and 0.13, the second position with candidate point indices 3 and 0.38, and the third position with candidate point indices 4 and 0.82. For records with a deviation metric of -1, their positions are preserved in the difference sequence, and an invalid marker is written to prevent misalignment between the sequence position and the candidate point index. To facilitate subsequent logical mask generation, a discrimination marker reserved field is synchronously written into the difference sequence; this field is initially empty. After encapsulation, the feature ratio difference sequence is output.

[0030] Please see Figure 5 The specific steps of S4 are as follows: S401: Call the feature ratio difference sequence, introduce the preset feature difference threshold as the screening benchmark, perform threshold discrimination operation on the difference values ​​in the sequence, define whether the value fluctuation is within the value range defined by the threshold, generate conformity markers for the positions of elements that meet the judgment conditions, construct a screening algorithm based on dynamic weights, and generate a candidate point conformity evaluation table. A feature difference threshold is introduced to perform threshold discrimination. In this embodiment, the feature difference threshold is set to 0.60. The threshold setting was obtained through experimental verification. A total of 1680 trajectory points over 7 consecutive days were selected, and thresholds of 0.40, 0.60, and 0.80 were used for filtering. The dispersion of the number of retained points and the deviation metric of retained points was statistically analyzed. Under the condition of 0.40, the number of retained points repeatedly fell below the minimum included point threshold. Under the condition of 0.80, the dispersion of the deviation metric of retained points increased significantly. Under the condition of 0.60, both statistics remained stable simultaneously, so 0.60 was fixed. The discrimination action compares each deviation metric in the sequence with 0.60. If it is less than or equal to 0.60, a conformity flag of 1 is written; if it is greater than 0.60, a conformity flag of 0 is written; and a deviation metric of negative 1 is directly written as 0. In the example, deviation metrics of 0.13 and 0.38 are written as 1, and deviation metric of 0.82 is written as 0. Subsequently, a logical index table is generated based on the conformity markers. The logical index table is written with 1s or 0s sequentially according to the position of the difference sequence, forming a candidate point conformity evaluation table. The candidate point sequence number and candidate record index are simultaneously written to ensure consistency in subsequent index mappings. After completion, the candidate point conformity evaluation table is output.

[0031] S402: Based on the candidate point conformity evaluation table, perform index mapping and data cleaning on the spatial neighborhood candidate set, remove candidate points marked as invalid in the logical mask, lock and extract trajectory data items whose difference values ​​meet the threshold constraints, queue up and reorganize the filtered data entities to generate a homogeneous retention point set; The mapping process iterates through the logical mask positions, reading elements marked as 1. Candidate point numbers and record indices at the same position are added to the retention list, while elements marked as 0 are added to the removal list with a removal reason field. The reason field indicates a deviation metric exceeding 0.60 or an invalid deviation metric. Subsequently, the candidate point number list of the spatial neighborhood candidate set is filtered, retaining only candidate points in the retention list. Simultaneously, the distance list is filtered to ensure a one-to-one correspondence between distance and candidate points. In the example, core point number 1 has candidate points 2, 3, and 4, with a mask of 1, 1, and 0. After cleaning, points 2 and 3 are retained, and point 4 is added to the removal list. After cleaning, the filtered data entities are queued and reorganized. Retained points are written to the homogeneous retention point set in ascending order of candidate point number, with the queue number field incrementing from 1. If the number of retained points is 0, the homogeneous retention point set is marked as empty.

[0032] S403: Based on the homogeneity retention point set, attribute verification is performed on the remaining data points that have been verified by both spatial and chemical attributes. Trajectory points with homogeneous characteristics are structurally integrated and encapsulated to establish the association set of the core capture object under the target semantic rules and construct a semantically connected neighbor set. After retaining the homogeneous set of points, attribute verification is performed on the remaining data points. Attribute verification first involves time consistency checking, which involves backtracking the sampling timestamp of each retained point and calculating the time difference with the sampling timestamp of the core capture object. Points with a time difference greater than 7200 seconds are marked as inconsistent and removed. The 7200-second threshold is obtained by statistically analyzing the duration window of pollution plumes during stable wind periods in the region. The statistical sample is the similarity ratio duration of 10 consecutive transport events, with a median close to 6900 seconds; therefore, 7200 seconds is used as the fixed threshold. Secondly, duplicate points are removed. For points with identical longitude, latitude, and sampling timestamp values, only the first occurrence is retained. After verification, the core capture object and the verified retained points are merged and encapsulated into a semantically connected neighbor set. The encapsulated fields include the core point number, core record index, neighbor point number list, neighbor record index list, corresponding distance list, and corresponding deviation metric list. In the example, if core point number 1 has two reserved points, 2 and 3, and their time differences are both less than 7200 seconds, then the neighbor list will contain points 2 and 3, along with deviation metrics of 0.13 and 0.38. If a reserved point has a time difference of 9000 seconds, it will be removed and will not be added to the neighbor list.

[0033] Please see Figure 6 The specific steps of S5 are as follows: S501: Call the semantically connected neighbor set, count the total number of data items included in the set, introduce a preset minimum included point threshold as the reference boundary for density clustering, perform numerical comparison and verification between the count and the minimum included point threshold, after verifying that the count meets the density constraint, extract the core capture object data, and perform set merging operation with the trajectory data items in the neighbor set, integrate spatially adjacent and attribute-similarity data entities, and generate atmospheric pollutant transmission trajectory clusters; The total number of data items in the set is counted by adding the core point count (1) to the number of neighboring points, thus obtaining the number of connected points. A minimum included point threshold is introduced to perform density determination; in this embodiment, the minimum included point threshold is set to 4. The threshold setting is verified through experiments, with cluster output statistics performed at thresholds of 3, 4, and 5 respectively. When the threshold is 3, the number of trajectory clusters increases significantly and the boundaries between clusters overlap; when the threshold is 5, the number of trajectory clusters decreases and the proportion of noise points increases; when the threshold is 4, the number of trajectory clusters and the proportion of noise points remain stable, so it is fixed at 4. The determination action compares the number of connected points with 4. If it is greater than or equal to 4, a set merging is performed, and the core capture object and all trajectory points in the neighbor set are written into the same trajectory cluster container. In the example, core point number 1 has neighboring points 2, 3, and 5, and the number of connected points is 4, which meets the threshold of 4. Therefore, a trajectory cluster is generated and written to cluster identifier 1-1, with point numbers 1, 2, 3, and 5 within the cluster. If core point number 6 has only neighboring point 7 and the number of connected points is 2, which does not meet the threshold of 4, a trajectory cluster is not generated, and the core point is written to the noise list. After completing the traversal, output the cluster of atmospheric pollutant transport trajectories.

[0034] S502: Based on the atmospheric pollutant transport trajectory cluster, traverse all trajectory data points encapsulated within the cluster, extract the toluene to benzene concentration ratio associated with the data points, construct a set of chemical feature values ​​to be processed, perform arithmetic mean operation on the set of values, quantify the overall concentration trend of the cluster in terms of chemical composition, establish a central metric index reflecting the chemical characteristics of the transport trajectory, and generate the average concentration ratio within the cluster. The algorithm iterates through all trajectory data points within a cluster, extracting the toluene-to-benzene concentration ratio for each point and writing them into a chemical feature value group in ascending order of point number. Then, it performs an arithmetic mean operation on the value group, summing all ratios and dividing by the number of points to obtain the average concentration ratio within the cluster. To ensure the mean is not affected by extreme values ​​when the number of points is large, a truncation cleanup is performed first when the number of points is greater than or equal to 6, removing one maximum and one minimum value from the sorted values ​​before averaging the remaining values; no truncation is performed when the number of points is less than 6. In the example, the point numbers in cluster 1-1 are 1, 2, 3, and 5, with corresponding concentration ratios of 5.81, 5.94, 5.43, and 7.13. With 4 points, no truncation is performed. Substituting these values ​​into the averaging logic yields an average concentration ratio of 6.08 within the cluster. Another cluster has intra-cluster ratios of 2.10, 2.35, 2.28, and 2.05. Substituting these into the averaging logic, the intra-cluster mean is 2.20. The mean intra-cluster concentration ratio is then bound to the cluster identifier and the number of points, written into the cluster center metric set, and also written into the mean calculation status field. After completion, the mean intra-cluster concentration ratio is output.

[0035] S503: Call the mean concentration ratio within the cluster to obtain the source component spectrum including standard feature vectors of multiple pollution sources, map the mean within the cluster and the standard vectors in the source component spectrum to the same feature space, perform numerical difference measurement, calculate the difference between the mean within the cluster and the standard vectors in the source component spectrum, determine the category of pollution source based on the difference value, match the corresponding pollution source label information, and generate pollution source classification results. The standard feature vector records of multiple pollution sources in the source composition spectrum are read. Each record contains a pollution source label and a standard ratio vector value. The standard ratio vector values ​​are obtained by statistically analyzing 30 consecutive days of data from monitoring stations near known pollution sources. The statistical process first calculates the daily ratio of the daily average concentrations of toluene and benzene, and then averages these daily ratios over 30 days to obtain the standard ratio. In the example, the standard ratio for traffic exhaust sources is 6.00, for solvent-related sources it is 2.10, and for petrochemical process-related sources it is 9.20. The cluster mean and each standard ratio are mapped to the same feature space using a unified normalization logic. All ratios are divided by the maximum standard ratio value of this batch, 9.20, to obtain the normalized ratio. The cluster mean concentration ratio is represented as a vector and its cosine similarity is calculated with the standard feature vector in the source composition spectrum. The pollution source category is determined based on the cosine similarity value, which is limited to the range of 0 to 1. In the example, cluster identifier 1-1 has a mean of 6.08, which is normalized to approximately 0.66. Its absolute difference with traffic exhaust-related sources (0.65) is approximately 0.01, resulting in a similarity of 0.99 when substituted into the scoring logic. Its absolute difference with solvent usage-related sources (0.23) is approximately 0.43, yielding a similarity of 0.57. Its absolute difference with petrochemical process-related sources (1.00) is approximately 0.34, yielding a similarity of 0.66. The highest similarity is used to assign the corresponding pollution source label, generating the pollution source classification result as traffic exhaust-related source. The same comparison is performed on all clusters to output the pollution source classification result.

[0036] Please see Figure 7 An atmospheric pollutant transport trajectory clustering analysis system, comprising: The trajectory point construction module is used to execute S1: obtain discrete sampling records containing atmospheric environmental monitoring data through atmospheric environmental monitoring stations, extract longitude values, latitude values, sampling timestamps, toluene concentration values, and benzene concentration values ​​from the discrete sampling records, divide the toluene concentration value by the benzene concentration value to calculate the toluene / benzene concentration ratio that characterizes the chemical composition of pollutants, establish the mapping relationship between the toluene / benzene concentration ratio and the longitude values, latitude values, and sampling timestamps, and construct multidimensional trajectory data points; The spatial neighborhood screening module is used to execute S2: traverse the set consisting of multiple multidimensional trajectory data points to select the core capture object to be processed, call the spherical distance calculation formula to calculate the spatial physical distance between the core capture object and the other multidimensional trajectory data points in the set in the geographic coordinate system, compare the spatial physical distance with the preset neighborhood radius, filter data points whose spatial physical distance is less than the neighborhood radius, and construct a spatial neighborhood candidate set; The feature difference calculation module is used to execute S3: obtain the toluene / benzene concentration ratio of the core capture object and the candidate points in the spatial neighborhood candidate set respectively, calculate the absolute difference of the toluene / benzene concentration ratio between the core capture object and each candidate point, and generate the feature ratio difference sequence corresponding to the spatial neighborhood candidate set; The semantic connectivity construction module is used to execute S4: compare each value in the feature ratio difference sequence with a preset feature difference threshold, remove candidate points whose corresponding values ​​are greater than the feature difference threshold from the spatial neighborhood candidate set, and retain candidate points whose corresponding values ​​are less than or equal to the feature difference threshold to construct a semantic connectivity neighborhood set. The trajectory cluster source resolution module is used to execute S5: count the total number of elements in the semantically connected neighbor set. When the total number of elements meets the preset minimum number of included points, merge the core capture object and the semantically connected neighbor set to generate an atmospheric pollutant transport trajectory cluster. Calculate the arithmetic mean of the toluene / benzene concentration ratio of all data points in the atmospheric pollutant transport trajectory cluster. Compare the arithmetic mean of the concentration ratios within the cluster with the standard feature values ​​in the source component spectrum to measure the numerical difference. Calculate the absolute difference and perform normalization to output the pollution source classification result.

[0037] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of protection of the described technical solutions.

Claims

1. A method for cluster analysis of atmospheric pollutant transport trajectories, characterized in that, Includes the following steps: S1: Obtain discrete sampling records, extract longitude values, latitude values, sampling timestamps, toluene concentration values, and benzene concentration values, calculate the toluene / benzene concentration ratio, and construct multidimensional trajectory data points; S2: Select a core capture object from the multidimensional trajectory data points, calculate the spatial physical distance between the core capture object and the data points, filter data points whose spatial physical distance is less than the neighborhood radius, and construct a spatial neighborhood candidate set; S3: Obtain the toluene / benzene concentration ratio of the spatial neighborhood candidate set and the core capture object, calculate the absolute difference between the toluene / benzene concentration ratio of the core capture object and the element, and generate a feature ratio difference sequence. S4: Compare the values ​​in the feature ratio difference sequence with a preset feature difference threshold, remove candidate points in the spatial neighborhood candidate set whose values ​​are greater than the preset feature difference threshold, retain candidate points whose values ​​are less than or equal to the preset feature difference threshold, and construct a semantically connected neighbor set. S5: When the number of semantically connected neighbor sets meets the minimum number of included points, merge the core capture object and the semantically connected neighbor set to construct an atmospheric pollutant transport trajectory cluster, calculate the arithmetic mean of the toluene / benzene concentration ratio in the atmospheric pollutant transport trajectory cluster, compare the arithmetic mean with the source component spectrum based on cosine similarity, and output the pollution source classification result.

2. The atmospheric pollutant transport trajectory clustering analysis method according to claim 1, characterized in that, The multidimensional trajectory data points include longitude values, latitude values, sampling timestamps, and toluene / benzene concentration ratios. The spatial neighborhood candidate set includes the core capture object and neighboring data points whose spatial physical distance is less than the neighborhood radius. The feature ratio difference sequence includes the absolute difference in toluene / benzene concentration ratios between the core capture object and the candidate points. The semantically connected neighbor set includes retained candidate points whose feature difference values ​​meet the feature difference threshold requirements. The pollution source classification result includes pollution source categories that are consistent with the source component spectrum.

3. The atmospheric pollutant transport trajectory clustering analysis method according to claim 1, characterized in that, The core capture object is the currently processed data point selected sequentially from the multi-dimensional trajectory data points according to the sampling time order; The spatial physical distance is a spherical distance calculated based on longitude and latitude values; The neighborhood radius is a pre-set spatial distance threshold used to limit data points that are geographically adjacent to the core capture object.

4. The atmospheric pollutant transport trajectory clustering analysis method according to claim 1, characterized in that, The characteristic difference threshold is a pre-set upper limit for the difference in the toluene / benzene concentration ratio; The feature difference threshold is determined based on the statistical distribution of the toluene / benzene concentration ratio in historical pollution source sample data. When the absolute difference between the toluene / benzene concentration ratio of the core capture object and the data points in the spatial neighborhood candidate set is less than or equal to the feature difference threshold, it is determined that the two have connectivity in terms of pollution source semantics.

5. The atmospheric pollutant transport trajectory clustering analysis method according to claim 1, characterized in that, The specific steps of S1 are as follows: S101: Collect discrete sampling records through atmospheric environmental monitoring stations, perform field parsing and data extraction operations on the records, extract longitude values, latitude values, sampling timestamps, toluene concentration values, and benzene concentration values, standardize and integrate the extracted data, and generate an environmental discrete sampling feature set. S102: Call the environmental discrete sampling feature set to obtain the toluene concentration value and the benzene concentration value, calculate the concentration ratio characterizing the chemical components of pollutants, and associate and anchor the calculated value with the original data record index to generate the toluene / benzene concentration ratio. S103: Based on the toluene / benzene concentration ratio, retrieve the corresponding longitude, latitude, and sampling timestamp within the environmental discrete sampling feature set, establish a mapping relationship between the concentration ratio and the spatiotemporal coordinate dimension, and restructure and encapsulate the data according to the mapping relationship to construct multidimensional trajectory data points.

6. The atmospheric pollutant transport trajectory clustering analysis method according to claim 1, characterized in that, The specific steps of S2 are as follows: S201: Traverse the set of multi-dimensional trajectory data points, select the data item to be processed as the core capture object, synchronously index the multi-dimensional trajectory data points in the set, extract the geographical coordinate values ​​of both parties and establish the corresponding relationship, and generate a set of coordinate pairings for the core capture object. S202: Call the coordinate pairing set of the core capture object, and perform spherical distance calculation operation on the core capture object and multi-dimensional trajectory data points according to the spatial distribution characteristics under the geographic coordinate system, quantify the spatial interval between data points, and generate a spatial physical distance measurement sequence. S203: Based on the spatial physical distance measurement sequence, obtain a preset neighborhood radius threshold, compare the distance measurement value with the neighborhood radius threshold, filter data entries with distance values ​​less than the neighborhood radius threshold, aggregate and reorganize the trajectory data points that meet the conditions, and construct a spatial neighborhood candidate set.

7. The atmospheric pollutant transport trajectory clustering analysis method according to claim 1, characterized in that, The specific steps for S3 are as follows: S301: Based on the index identifiers included in the spatial neighborhood candidate set, backtrack to retrieve the toluene / benzene concentration ratio, extract the concentration ratio data of the core capture object and the candidate point respectively, establish a correspondence between the ratio data of the core object and the ratio data of the candidate point, and generate a core candidate ratio pairing set; S302: Call the core candidate ratio pairing set, perform numerical difference measurement calculation on the pairing data, quantify the degree of deviation between the core capture object and the candidate point in the chemical pollutant composition, and generate a concentration ratio deviation measurement. S303: Based on the concentration ratio deviation metric, the deviation metric values ​​are serialized, recombined, and encapsulated according to the original arrangement order of the data points within the spatial neighborhood candidate set to construct a feature ratio difference sequence.

8. The atmospheric pollutant transport trajectory clustering analysis method according to claim 1, characterized in that, The specific steps of S4 are as follows: S401: Call the feature ratio difference sequence, introduce a preset feature difference threshold as a screening benchmark, perform threshold discrimination operation on the difference values ​​in the sequence, define whether the value fluctuation is within the value range defined by the threshold, generate a conformity mark for the element position that meets the judgment condition, construct a screening algorithm based on dynamic weight, and generate a candidate point conformity evaluation table. S402: Based on the candidate point conformity evaluation table, perform index mapping and data cleaning on the spatial neighborhood candidate set, remove candidate points marked as invalid in the logical mask, lock and extract trajectory data items whose difference values ​​meet the threshold constraints, and queue up and reorganize the filtered data entities to generate a homogeneity retention point set. S403: Based on the homogeneity retention point set, perform attribute verification on the remaining data points that have been verified by both spatial and chemical attributes, structure and integrate the trajectory points with homogeneous characteristics, establish the association set of the core capture object under the target semantic rules, and construct the semantically connected neighbor set.

9. The atmospheric pollutant transport trajectory clustering analysis method according to claim 1, characterized in that, The specific steps of S5 are as follows: S501: Call the semantically connected neighbor set, count the total number of data items included in the set, introduce a preset minimum included point threshold as the reference boundary for density clustering determination, perform numerical comparison and verification between the count and the minimum included point threshold, after verifying that the count meets the density constraint, extract the core capture object data, and perform a set merging operation with the trajectory data items in the neighbor set to integrate spatially adjacent and attribute-similarity data entities to generate an atmospheric pollutant transmission trajectory cluster; S502: Based on the atmospheric pollutant transport trajectory cluster, traverse all trajectory data points encapsulated within the cluster, extract the toluene to benzene concentration ratio associated with the data points, construct a chemical feature numerical group to be processed, perform arithmetic mean operation on the numerical group, quantify the overall concentration trend of the cluster in chemical composition, establish a central metric index reflecting the chemical characteristics of the transport trajectory, and generate the average concentration ratio within the cluster. S503: Call the mean concentration ratio within the cluster to obtain the source component spectrum including standard feature vectors of multiple pollution sources, map the mean within the cluster and the standard vectors in the source component spectrum to the same feature space, calculate the cosine similarity, determine the category of pollution source based on the similarity value, match the corresponding pollution source label information, and generate pollution source classification results.

10. A clustering analysis system for atmospheric pollutant transport trajectories, characterized in that, The system is used to implement the atmospheric pollutant transport trajectory clustering analysis method according to any one of claims 1-9, the system comprising: The trajectory point construction module is used to execute S1: obtain discrete sampling records containing atmospheric environmental monitoring data through atmospheric environmental monitoring stations, extract longitude values, latitude values, sampling timestamps, toluene concentration values, and benzene concentration values ​​from the discrete sampling records, divide the toluene concentration value by the benzene concentration value to calculate the toluene / benzene concentration ratio that characterizes the chemical composition of pollutants, establish the mapping relationship between the toluene / benzene concentration ratio and the longitude values, latitude values, and sampling timestamps, and construct multidimensional trajectory data points; The spatial neighborhood screening module is used to execute S2: traverse the set of multi-dimensional trajectory data points to select the core capture object to be processed, call the spherical distance calculation formula to calculate the spatial physical distance between the core capture object and the other multi-dimensional trajectory data points in the set in the geographic coordinate system, compare the spatial physical distance with the preset neighborhood radius, filter data points whose spatial physical distance is less than the neighborhood radius, and construct a spatial neighborhood candidate set. The feature difference calculation module is used to execute S3: obtain the toluene / benzene concentration ratio of the candidate points in the spatial neighborhood candidate set and the core capture object respectively, calculate the absolute difference of the toluene / benzene concentration ratio between the core capture object and each candidate point, and generate a feature ratio difference sequence. The semantic connectivity construction module is used to perform S4: compare each value in the feature ratio difference sequence with a preset feature difference threshold, remove candidate points whose corresponding values ​​are greater than the feature difference threshold from the spatial neighborhood candidate set, retain candidate points whose corresponding values ​​are less than or equal to the feature difference threshold, and construct a semantic connectivity neighborhood set. The trajectory cluster source resolution module is used to execute S5: count the total number of elements in the semantically connected neighbor set. When the total number of elements meets the preset minimum number of included points, merge the core capture object and the semantically connected neighbor set to generate an atmospheric pollutant transport trajectory cluster. Calculate the arithmetic mean of the toluene / benzene concentration ratio of all data points in the atmospheric pollutant transport trajectory cluster. Compare the arithmetic mean of the concentration ratios within the cluster with the standard feature values ​​in the source component spectrum to measure the numerical difference. Calculate the absolute difference and perform normalization to output the pollution source classification result.