A satellite ground measurement data and communication signaling fusion air pollution exposure assessment method
By constructing a method for fusing satellite ground measurement data with communication signaling, the problem of separation between satellite-ground observation and signaling data in existing technologies has been solved. This enables high spatiotemporal resolution air pollution exposure assessment for multiple pollutants and at multiple time scales, improving the accuracy and reliability of the assessment and supporting environmental health risk assessment and management decisions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA UNIV OF MINING & TECH
- Filing Date
- 2026-04-07
- Publication Date
- 2026-06-23
AI Technical Summary
In existing technologies, there is a lack of an integrated framework for fusion of satellite and ground observation and operator signaling data, making it difficult to accurately characterize individual spatiotemporal activity patterns, support air pollution exposure assessments of multiple pollutants and multiple time scales simultaneously, and lack the ability to quantify uncertainty, resulting in insufficient accuracy and reliability of exposure assessment results.
A unified spatiotemporal reference framework is constructed, integrating satellite remote sensing and communication signaling data. A high spatiotemporal resolution pollutant concentration field is inverted through a gradient boosting tree model. Individual activity patterns are identified by combining communication signaling data, multi-pollutant exposure is calculated, and uncertainty is quantified through Monte Carlo simulation.
It achieves high spatiotemporal resolution for multi-pollutant exposure assessment, reduces spatiotemporal mismatch errors in exposure assessment, supports exposure analysis at multiple time scales, and provides confidence intervals for exposure results, thereby improving the reliability of assessment and the scientific nature of management decisions.
Smart Images

Figure CN122264646A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of environmental health and air pollution control, and in particular to an air pollution exposure assessment method that integrates satellite ground measurement data and communication signaling. Background Technology
[0002] Fine particulate matter (PM2.5), ground-level ozone (O3), and other air pollutants have been extensively confirmed by epidemiological studies to be closely related to respiratory diseases, cardiovascular diseases, and even premature death. To accurately assess the impact of air pollution on human health, it is necessary to precisely characterize individual pollution exposure levels at different spatiotemporal scales. Traditional exposure assessment methods typically allocate pollutant concentrations from environmental monitoring stations or model simulation results to the population by administrative division or residential address, using this as a proxy variable for exposure in epidemiological statistical analysis. While these methods are simple to implement, they often lead to "exposure misclassification" because they ignore the fine spatial heterogeneity of pollution fields and individuals' daily movement behaviors, thus reducing the accuracy of health effect assessments.
[0003] With the development of remote sensing technology and big data analysis methods, existing technologies have gradually formed high-resolution air quality estimation methods based on "satellite-ground observation fusion". By combining aerosol optical depth (AOD) retrieved from multi-source satellites with auxiliary variables such as ground monitoring station data, meteorological elements, and land use, and using machine learning models such as random forests and long short-term memory neural networks, researchers have been able to estimate the PM2.5 concentration field at a 1km spatial resolution on a national scale, forming a series of high-resolution datasets such as ChinaHighPM2.5 and CHAP. In addition, some studies have further utilized deep learning methods to integrate ground monitoring data, satellite products, and chemical transport model outputs to separate the spatiotemporal field of PM2.5 inorganic component concentration at a 1km resolution in China since 2000. Regarding ozone pollution, researchers have used spatiotemporal artificial intelligence models to couple multi-source satellite observations and ground station data, achieving 24-hour continuous surface ozone concentration retrieval at a 1km resolution in China. Datasets such as ChinaHighO3 provide the maximum 8-hour average ozone concentration data at a 1km resolution and daily scale in China since 2000. These high-resolution contaminated fields provide a more refined environmental data foundation for exposure assessment.
[0004] Meanwhile, the widespread adoption of mobile internet and smart terminals has made spatiotemporal activity characterization of populations based on mobile phone data or telecommunications operator signaling an important supplementary tool for exposure assessment. Among existing technologies, one type of research combines social media location service (LBS) data with high-quality PM2.5 concentration data retrieved from satellites to conduct high spatiotemporal resolution dynamic assessments of PM2.5 exposure and health risks in multiple cities in the Beijing-Tianjin-Hebei region. Another type of research directly extracts residents' spatiotemporal trajectories from mobile phone data or operator signaling data. For example, in Beijing, cellular network data is used to infer the hourly location distribution of the entire population, and ground-monitored PM2.5 concentration data is overlaid to assess the dynamic exposure characteristics of urban residents. In a big data analysis framework built in Shanghai, 6 billion call detail records (CDRs) were used to identify the dwell points and complete daily trajectories of nearly one million users, mapping them onto a 1km grid PM2.5 concentration field to quantify individual hourly PM2.5 exposure levels. Compared with exposure results based solely on residential concentrations, this reveals significant biases in traditional methods.
[0005] Building on this foundation, existing research has further utilized mobile phone data to explore exposure inequalities. For example, based on hourly PM2.5 concentrations derived from carrier data and machine learning, the study analyzed the exposure differences among residents at different housing price levels in megacities; by combining mobile phone population distribution with model-simulated NO2 concentrations, the study characterized the spatiotemporal pattern of NO2 exposure from regional road traffic emissions and the exposure differences among different population groups; and by using residential and employment density information from mobile phone data products, overlaid with a 1km resolution ozone dataset, the study assessed weekday surface ozone exposure, showing that commuting patterns have a significant impact on exposure.
[0006] However, existing technologies still have significant shortcomings: First, there is a lack of an integrated framework for combining satellite and ground observations with operator signaling data. Current studies typically construct high-resolution pollution fields first, and then overlay population distribution or mobile phone statistics in post-processing for exposure estimation. The two are relatively independent in terms of spatiotemporal grids, data update frequency, and error control. Second, existing methods do not make sufficient use of signaling data. Most methods only infer residential and workplace locations or estimate group location distribution based on base station connection records, and do not fully mine information such as handover events, dwell time, and multi-anchor travel chains in the original signaling data. This makes it difficult to accurately reconstruct complex travel behaviors (such as multiple transfers, temporary stops, weekend trips, etc.), leading to systematic biases in exposure assessments. Third, existing technologies cannot simultaneously support exposure assessments of multiple pollutants and multiple time scales. Most are limited to single pollutant PM2.5 or only attempted in short-term ozone exposure studies, and the time scales are mostly concentrated on daily averages or a few representative periods, making it difficult to meet the need to simultaneously study chronic exposure effects and acute short-term high-exposure events. Finally, existing methods generally lack a systematic quantification of the uncertainty of exposure assessment results, making it difficult to assess the reliability of the results and limiting their application value in environmental health management decision-making.
[0007] Therefore, there is an urgent need for an air pollution exposure assessment method that can deeply integrate satellite and ground observation data with communication operator signaling data, finely characterize individual spatiotemporal activity patterns, support multi-pollutant multi-timescale exposure assessment, and have the ability to quantify uncertainty, so as to overcome the limitations of existing technologies and provide a scientific basis for environmental health risk assessment and refined air quality management. Summary of the Invention
[0008] This invention addresses the shortcomings of existing technologies by providing a method for assessing air pollution exposure that integrates satellite ground measurement data with communication signaling.
[0009] To achieve the above-mentioned objectives, the technical solution adopted by the present invention is as follows:
[0010] An air pollution exposure assessment method that integrates satellite ground measurement data and communication signaling includes the following steps:
[0011] S1: Constructing a unified spatiotemporal reference framework:
[0012] Acquire spatial boundary information and basic geographic data of the target study area, divide the study area into regular grids to form a unified spatial grid system; set the grid resolution according to the spatial resolution of satellite remote sensing products and the positioning error of communication signaling data, and set a unified temporal resolution and time range according to application requirements, and map all subsequent data onto the spatiotemporal grid and time axis.
[0013] S2: Acquisition and Preprocessing of Multi-Source Satellite-Ground Observation Data:
[0014] Multi-source environmental data of the target area are collected, including pollutant concentration data and meteorological observation data released by ground ambient air quality monitoring stations, environmental parameter data and auxiliary meteorological data in satellite remote sensing products, and the above data are subjected to quality control, outlier removal and missing data filling. Through spatial resampling and temporal interpolation, data with different spatial and temporal resolutions are unified to the spatial grid and time step.
[0015] S3: Construction of Atmospheric Pollutant Concentration Field Based on Satellite-Ground Data Fusion:
[0016] Using the measured pollutant concentrations at ground-based ambient air quality monitoring stations as supervisory variables and the multi-source feature data obtained from S2 as input features, a satellite-ground data fusion concentration inversion model is constructed to predict the outdoor pollutant concentrations at all grid-time points within the study area, thereby obtaining a multi-pollutant high spatiotemporal resolution atmospheric environmental pollutant concentration field.
[0017] The aforementioned satellite-to-ground data fusion concentration inversion model refers to a calculation model that fuses satellite remote sensing data (satellite) with ground monitoring station data (ground) through a mathematical model to inversely calculate (invert) the pollutant concentration at various locations throughout the entire study area. This model leverages the wide coverage advantage of satellite data and the high precision of ground monitoring to overcome the limitations of a single data source, thereby obtaining a high spatiotemporal resolution pollutant concentration distribution.
[0018] The grid-time point refers to the basic spatiotemporal unit formed by dividing the study area into regular grids (e.g., 1km × 1km) and dividing time into fixed intervals (e.g., 1 hour). Each "grid-time point" represents a combination of specific spatial locations at a specific time, and is a unified spatiotemporal reference framework for multi-source heterogeneous data.
[0019] S4: Acquisition and preprocessing of signaling data by telecommunications operators:
[0020] Obtain raw signaling data generated by telecommunications operators within the target area. The signaling data includes anonymous user identifiers, location information, and timestamps. Clean the signaling data and map the location of each signaling record to the spatial grid in S1. Discretize the timestamps to a uniform time step to obtain discrete time-series data in the form of "user-time step-grid".
[0021] S5: Time-Activity Pattern Recognition Based on Signaling Data:
[0022] User dwell points are identified using preset time and space thresholds. Nighttime dwell features and weekday daytime dwell features for each user are constructed to identify residential grids and work grids. Dwell points are classified into different activity scenarios by combining dwell time period, dwell duration and spatial information, and path interpolation is performed between adjacent dwell points to generate travel chains containing dwell and movement information. The spatial grid where the user is located and its activity type are determined according to a unified time step to form a user-level time-activity pattern table.
[0023] The time-activity pattern refers to the types of activities and spatial location patterns of individuals or groups at different times. By analyzing communication signaling data, it is possible to identify users' stop points, travel chains, residences / workplaces, etc., and then construct the user's activity types (such as residence, work, commuting, other outdoor activities) and their corresponding spatial locations at each time step, reflecting the real exposure scenario.
[0024] S6: Star-ground concentration field and time-activity pattern fusion and single pollutant exposure calculation:
[0025] The multi-pollutant atmospheric environmental pollutant concentration field constructed in S3 is matched with the user time-activity pattern table obtained in S5 to obtain the pollutant environmental concentration of each user in the grid at each time step; for a given time window, the environmental concentration is integrated over time to obtain the cumulative outdoor exposure of each user in the time window.
[0026] S7: Construction and Statistical Analysis of Multi-Pollutant Exposure Indicators at Multiple Time Scales
[0027] Based on the user-level single pollutant exposure index, a weighting parameter is introduced to weight and synthesize the exposure amounts of multiple pollutants to obtain the multi-pollutant comprehensive outdoor exposure index; according to different time scales, the single pollutant exposure index and comprehensive exposure index of individuals and population subgroups are statistically analyzed.
[0028] S8: Uncertainty Modeling and Confidence Assessment of Exposure Results:
[0029] Based on the validation results of the star-ground concentration inversion model, a statistical model of the prediction error of pollutant concentration at different grid-time points is constructed; by adjusting the activity pattern recognition parameters, the stability of the user's grid and activity type is analyzed; by using error propagation or Monte Carlo simulation, the concentration field and time-activity pattern are perturbed to obtain the confidence intervals of user-level and population-level exposure indicators.
[0030] Furthermore, the resolution of the spatial grid is adaptively set based on the original spatial resolution of the satellite product and the positioning error of the communication signaling data; the time resolution is set to any value among 1 hour, 3 hours, or 24 hours.
[0031] Furthermore, in step S2:
[0032] Multi-source satellite-to-ground observation data should include at least:
[0033] 1) Hourly concentrations of PM2.5, PM10, O3, and NO2 pollutants and routine meteorological data released by ground-based ambient air quality monitoring stations;
[0034] 2) Data on aerosol optical depth (AOD), tropospheric NO2 column concentration, cloud cover, and surface albedo from satellite remote sensing products;
[0035] 3) Further analyze the meteorological field, boundary layer height, and planetary boundary layer structure parameters;
[0036] The data is unified to the spatial grid and time step using bilinear interpolation, nearest neighbor interpolation, linear interpolation, and spline interpolation methods.
[0037] Furthermore, in step S3:
[0038] The concentration inversion model of satellite-ground data fusion adopts the gradient boosting tree model XGBoost. Its input feature vector includes satellite observation variables, meteorological variables, land use characteristics and time variables, and the output is the pollutant concentration at the corresponding grid-time point. During the model training process, the measured concentration at the grid-time point where the monitoring station is located is used as the supervision sample, and cross-validation and / or station-based verification are used to evaluate the model performance.
[0039] Furthermore, in step S4:
[0040] Telecommunications operator signaling data includes records of various network events, such as location updates, call setup, SMS sending and receiving, and mobile data sessions.
[0041] When preprocessing the raw signaling data, the events are first sorted by user identifier, and records with missing key fields or abnormal timestamps are deleted; then each record is mapped to the corresponding spatial grid, and the timestamp is discretized to the nearest time step.
[0042] Furthermore, the time-activity pattern recognition in step S5 includes:
[0043] 1) Stop point identification: When a user is located in the same or adjacent spatial grid within a certain number of consecutive time steps, and the cumulative stop time exceeds the preset stop time threshold and the spatial displacement is less than the preset distance threshold, the time period is merged into a stop point;
[0044] 2) Residential / Workplace Identification: Based on the distribution of users' stops during the nighttime hours, the duration and frequency of nighttime stays are statistically analyzed, and the grid with the highest duration and frequency of stays is designated as the residential grid; based on the distribution of stops during the weekday daytime hours, the characteristics of stays during the workday hours are statistically analyzed, and the grid with the highest percentage of stays during the weekday daytime hours is designated as the work grid.
[0045] 3) Activity type identification and path interpolation: Based on the time period of the stay, the duration of the stay, the land use ratio within the grid, and the POI type information, the stay points are divided into different activity types; for the movement process between adjacent stay points, the user's location is interpolated on the time axis using the spatial grid adjacency matrix and road network accessibility index.
[0046] 4) Time-Activity Pattern Generation: Based on a uniform time step, the stationary and moving states are mapped to the activity type and spatial grid corresponding to each time step, and a user-level time-activity pattern table is constructed.
[0047] Furthermore, in step S6:
[0048] For time steps in commuting or other outdoor activity scenarios, the corresponding breathing rate is selected from preset breathing rate values based on the activity intensity level mapped by the activity scenario, and the environmental concentration is multiplied by the breathing rate to obtain a dose-type exposure index.
[0049] Within a given time window, the exposure concentration or dose-related indicators at each time step are weighted to obtain the cumulative exposure and time-weighted average exposure concentration within that time window.
[0050] Furthermore, the method for constructing the multi-pollutant comprehensive exposure index in step S7 is as follows:
[0051] For each pollutant i, the user-level exposure amount is calculated, and the weight is determined according to the pollutant's toxicity equivalent, health risk coefficient, or ambient air quality standard limit. The exposure amounts of each pollutant are then calculated using a weighted composite formula to obtain the multi-pollutant comprehensive exposure index.
[0052] Furthermore, the uncertainty modeling and exposure result evaluation in step S8 includes:
[0053] A grid-time point pollutant concentration error distribution was constructed based on the residuals of the star-ground concentration inversion model on the validation samples.
[0054] By perturbing the dwell time identification threshold and activity type discrimination threshold within a reasonable range, multiple sets of candidate time-activity patterns are generated.
[0055] At least 500 Monte Carlo simulations were performed to jointly perturb the pollutant concentration field and user trajectory, and user-level and population-level exposure indicators were repeatedly calculated to obtain the probability distribution of the exposure indicators, from which 95% confidence intervals were extracted.
[0056] Furthermore, the exposure assessment results are used as input to epidemiological exposure-response models, or for population exposure inequality analysis, regional health risk early warning, and the development of refined urban air quality management strategies.
[0057] The exposure-response model is a core concept in environmental health science and epidemiology, referring to a mathematical model that describes the quantitative relationship between the level of exposure to environmental pollutants and the health effects on organisms.
[0058] Compared with the prior art, the advantages of the present invention are as follows:
[0059] 1. Improve the spatiotemporal precision and coverage of pollutant concentration fields.
[0060] This invention integrates ground-based monitoring and satellite remote sensing data within a unified spatial grid and temporal resolution framework, taking into account both the high precision of ground-based observations and the wide coverage of satellite observations, to construct a high-resolution spatiotemporal concentration field with multiple pollutants, significantly improving the ability to characterize the spatiotemporal distribution of regional air pollution.
[0061] 2. Implement individualized exposure assessment based on time-activity trajectory.
[0062] This invention utilizes signaling data from telecommunications operators to reconstruct the temporal location and time-activity patterns of a large sample population. It breaks through the traditional practice of using the average concentration of a residence or administrative region as an exposure proxy, and can characterize the exposure process of individuals in real activity space at fine grid and fine temporal resolution, thereby reducing the spatiotemporal mismatch error in exposure assessment.
[0063] 3. By taking into account the actual exposure levels of individuals, the accuracy of exposure estimates can be improved.
[0064] Based on the fusion of star-ground concentration field and time-activity pattern, this invention introduces correction parameters such as inhalation dose coefficient to transform environmental concentration into a concentration or dose that is closer to the actual inhalation, thereby achieving a fine characterization of the individual's true exposure level under different scenarios and activity intensities.
[0065] 4. Support the construction of exposure indicators for multiple pollutants and multiple time scales.
[0066] This invention enables joint exposure assessment of multiple air pollutants within a unified framework, and allows for flexible integration and accumulation of exposure amounts at different time scales such as hours, days, weeks, and months, depending on research or management needs. It is applicable to both acute short-term effect studies and chronic long-term health risk assessments and policy effectiveness evaluations.
[0067] 5. Possesses the ability to quantify uncertainty, enhancing the interpretability of results and their value for decision-making.
[0068] This invention can quantify the uncertainty of exposure assessment results through parameter perturbation, Monte Carlo simulation, or sensitivity analysis, and provide confidence intervals or credibility characterizations for exposure estimates. This helps management departments and researchers to scientifically formulate environmental health policies based on an understanding of the reliability and robustness of the results. Attached Figure Description
[0069] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0070] Figure 1 This is an overall flowchart of an air pollution exposure assessment method that integrates satellite ground measurement data and communication signaling, as described in an embodiment of the present invention.
[0071] Figure 2 This is a flowchart of step S5 in an embodiment of the present invention;
[0072] Figure 3 This is a flowchart of steps S6 and S7 in an embodiment of the present invention. Detailed Implementation
[0073] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0074] like Figure 1As shown, the present invention provides an air pollution exposure assessment method based on the fusion of satellite ground measurement data and communication signaling. The method mainly includes steps such as satellite-ground observation data acquisition and preprocessing, satellite-ground observation data fusion and concentration field construction, communication operator signaling data acquisition and preprocessing, time-activity pattern identification, satellite-ground concentration field and time-activity pattern fusion, and multi-scale exposure index construction and uncertainty assessment. The steps are performed sequentially according to the order shown by the arrows.
[0075] This embodiment uses common air pollutants such as fine particulate matter (PM2.5) and ozone (O3) as examples for illustration, but the present invention is not limited to these and is also applicable to other gaseous and particulate pollutants supported by monitoring data and satellite inversion data.
[0076] Step S1: Construct a unified spatiotemporal reference framework
[0077] Acquire spatial boundary information and basic geographic data such as administrative divisions, road networks, and land use for the target study area. Based on a preset spatial resolution (the grid resolution can be adaptively set according to signaling positioning errors and satellite product resolution), and considering the impact of positioning errors in uncertainty analysis, divide the target area into regular grids to form a unified spatial grid system. According to application requirements and data timeliness, set a unified time resolution (e.g., 1h, 3h, or 24h) and time range, and construct a unified time axis using the format of "date + time period." Map all subsequent data types to this spatiotemporal grid and time axis to achieve unified encoding of time and space.
[0078] Step S2: Acquisition and Preprocessing of Multi-Source Satellite-Ground Observation Data
[0079] Based on the unified spatiotemporal grid constructed in step S1, multi-source air quality observation data of the target area are collected, including:
[0080] 1) Hourly concentrations of pollutants such as PM2.5, PM10, O3, and NO2 released by ground-based ambient air quality monitoring stations, as well as routine meteorological observation data (temperature, relative humidity, wind speed and direction, air pressure, etc.).
[0081] 2) Satellite remote sensing products, such as aerosol optical depth (AOD), tropospheric NO2 column concentration, cloud cover, and surface albedo;
[0082] 3) Other relevant auxiliary data, such as reanalysis of meteorological fields, boundary layer height, planetary boundary layer structure parameters, etc.
[0083] The above data undergoes quality control, outlier removal, and missing data imputation. Through spatial resampling and temporal interpolation, data with different spatial and temporal resolutions are unified to the spatial grid and time step defined in step S1, forming a multi-source feature dataset.
[0084] Step S3: Construction of Atmospheric Pollutant Concentration Field Based on Satellite-Ground Data Fusion
[0085] This step, based on the multi-source feature data from step S2, uses the measured concentrations from ground monitoring stations as supervisory variables and employs the XGBoost gradient boosting tree model to construct a satellite-ground data fusion model to estimate the outdoor pollutant concentrations at all grid points and time points in the study area. Specifically, it includes:
[0086] 1) Select grid-time points containing ground monitoring stations as training samples. Extract feature vectors for each training sample containing the following: satellite observation variables: e.g., AOD, NO2 column concentration, cloud cover, etc.; meteorological variables: surface temperature, relative humidity, boundary layer height, 10m wind speed and direction, etc.; geographical and land use characteristics: e.g., elevation, urban / suburban / rural classification, road density, industrial land ratio, green space ratio, etc.; time variables: e.g., hour, day of the week, whether it is a weekday, season, etc. The corresponding label is the ground monitoring concentration for that grid-time point, e.g., PM2.5. 2.5 Or O3 concentration.
[0087] 2) Based on the above features and labels, an XGBoost outdoor concentration inversion model is constructed. Using the squared error as the loss function and the ground-monitored concentration as the regression target, key hyperparameters are set. Optimal hyperparameter combinations are found on the training set using methods such as grid search, random search, or Bayesian optimization to balance fitting accuracy and model generalization ability.
[0088] 3) Evaluate the performance of the XGBoost model using methods such as cross-validation (e.g., K-fold cross-validation), standby testing, or time-based partitioning of the training and validation sets, and calculate the coefficient of determination R. 2 Metrics such as root mean square error (RMSE) and mean square bias (MB) can be used. If the model performance does not meet the preset threshold, the feature combination or hyperparameters can be adjusted and the model can be retrained.
[0089] 4) Once the XGBoost model performance meets the requirements, the trained model is applied to the feature data of all grid-time points to predict the outdoor concentration values of each pollutant at each grid-time point. For areas and times without ground monitoring station coverage, the model provides estimates based on satellite observations and meteorological fields, thereby constructing a high-resolution atmospheric pollutant concentration field covering the entire study area and study period.
[0090] Step S4: Acquisition and preprocessing of signaling data by telecommunications operators.
[0091] This step acquires communication operator signaling data from users within the study area and performs denoising, gridding, and time alignment to provide a foundation for subsequent time-activity pattern recognition. Specifically, it includes:
[0092] 1) Subject to compliance with laws, regulations, and data security requirements, obtain anonymized signaling data of users within the research area from telecommunications operators. The original signaling data should include at least: anonymous user identifier, base station identifier and its latitude and longitude (or user location latitude and longitude), event timestamp, etc.
[0093] 2) Preprocess the signaling data, clean up incomplete records, and remove records that are missing user identifiers or timestamps; remove duplicate records with obviously disordered time order or abnormally small time intervals; group the records by user identifier and sort them in ascending order by timestamp.
[0094] 3) Map the location (base station location or user latitude and longitude) of each signaling record to the spatial grid defined in step S1 to obtain an event sequence in the form of "user-time-grid".
[0095] 4) Based on a unified time step, merge or downsample records of the same user that are repeatedly connected to the same grid in multiple adjacent time steps. Within a short time window (e.g., 5 min) less than the time step Δt, merge duplicate records to reduce data redundancy and noise, and provide input data that is easy to process for subsequent stop point identification.
[0096] Step S5: Time-activity pattern identification based on signaling data.
[0097] like Figure 2 As shown, this step processes the gridded signaling data through multiple functional modules to identify the user's main activity types and their corresponding spatial locations at different times, thereby constructing a user-level time-activity pattern. Specifically, this includes:
[0098] 1) Static Spatial Data Preparation Module M1:
[0099] Auxiliary data for describing the spatial functions of the study area are acquired and preprocessed, including: spatial grid division results, point of interest (POI) data, land use type data, road network data, and spatial adjacency relationships. Each spatial grid is correlated with the distribution of POI types, the proportion of land use types (e.g., residential, commercial, industrial land), road density, and other information within its coverage area to generate functional attribute information corresponding to the spatial grid. A grid-to-grid adjacency matrix and a road network-based accessibility index are constructed to provide a basis for path interpolation and activity type determination.
[0100] 2) Signaling preprocessing and meshing module M2:
[0101] The user-level signaling event sequence output in step S4 is received, further denoised and compressed, and each record is mapped to the corresponding spatial grid and time step to obtain discrete time-series data in the form of "user-time step-grid".
[0102] 3) Stop Point Recognition Module M3:
[0103] Dwell points are identified based on gridded user time-series data. A threshold-based rule-based approach can be used: when a user stays in the same or adjacent grids for several consecutive time steps, and the total duration exceeds a preset dwell time threshold (e.g., 20–60 minutes), while the spatial displacement is less than a preset distance threshold (e.g., 500–1000 meters), this time period is classified as a dwell point. For each dwell point, its start time, end time, dwell time, and corresponding representative grid number are calculated to form a "user-dwelling sequence." The aforementioned time and spatial thresholds can be adjusted according to the city scale of the study area, signaling sampling frequency, and user movement characteristics.
[0104] 4) Residence / Workplace Identification Module M4:
[0105] The locations of each user's stays are statistically analyzed using spatial grids, constructing nighttime stay characteristics and weekday daytime stay characteristics separately. Typical rules include: grids with the longest cumulative nighttime stay (e.g., 22:00–06:00) and appearing on most study days (e.g., more than 50% of the days) are identified as the user's residential grids; grids with the longest cumulative daytime stay (e.g., 09:00–18:00) and appearing on most weekdays are identified as the user's work grids. Further calculations of the percentage of stay time and frequency of occurrence for residential and work grids can be used to evaluate the confidence level of the identification results.
[0106] 5) Other outdoor activity recognition modules M5:
[0107] For stops not marked as residential or work grids, the activity type is determined by combining the stay time period (weekday / non-weekday, daytime / nighttime), stay duration, and static space functional attributes.
[0108] When a stop point is located in a grid with high road network density and POIs mainly consisting of transportation facilities, it can be identified as a commuting transfer or travel activity; when a stop point is located in a grid with a high proportion of commercial, leisure, or green space, and the stop occurs outside of working hours, it can be identified as other outdoor activities; stop points that cannot be identified can be temporarily marked as unknown activities.
[0109] In this embodiment, to highlight scenarios related to outdoor exposure, the time-activity type is limited to living, working, commuting, and other outdoor activities, without further modeling indoor sub-scenarios.
[0110] 6) Trip Chain and Path Interpolation Module M6:
[0111] By connecting two adjacent stop points of the same user in chronological order, a "stop-move-stop" travel chain is constructed. Using grid-to-grid adjacency relationships and road network data, path interpolation is performed between two stop grids to identify the grid sequence traversed during movement and the approximate travel distance. Based on whether the departure and arrival grids are respectively residential and work grids, commuter travel can be distinguished from non-commuter travel.
[0112] 7) Time-Activity Pattern Generation Module M7:
[0113] The timeline is discretized using a uniform time step (e.g., 1 hour), and information on each user's stop points, travel chains, and activity types is integrated to generate a time-activity pattern table. Within each time step, the spatial grid of each user is determined. And the corresponding activity type (residential, working, commuting, other outdoor activities, or unknown). In this invention, the exposure in residential and working scenarios is estimated using the outdoor concentration of the grid in which it is located as the exposure proxy value, while the dose for commuting and other outdoor activities is directly estimated using the outdoor concentration at the corresponding spatial location and a higher respiratory rate.
[0114] Based on the duration of stay, activity type, and scenario (e.g., overnight stay, weekday work stay, commuting), each time step can be mapped to an activity intensity level (e.g., sedentary, light physical activity, moderate physical activity) to support integration with other health risk models when needed. This embodiment provides two types of indicators: exposure indicators based on environmental concentration and inhalation volume indicators based on respiratory rate mapped to activity intensity.
[0115] 8) Parameter and Quality Monitoring Module M8:
[0116] Key indicators such as the number of stop points, average stay duration, percentage of unknown activities, and commuting distance distribution are statistically analyzed and monitored to check the rationality of the identification results. If the percentage of unknown activities is too high, the system prompts adjustments to the stay threshold or activity identification rules. Based on the monitoring results, the parameters of each module are updated and returned to M2–M5, forming a closed-loop optimization mechanism.
[0117] 9) Group Travel Pattern Analysis Module M9:
[0118] Based on user-level time-activity pattern tables and travel chain information, feature vectors describing commuting distance, travel time, travel frequency, and activity time distribution are constructed. Typical travel patterns, such as "short-distance commuting mode" and "long-distance cross-regional commuting mode", are mined through cluster analysis to provide a basis for exposure assessment by different population groups.
[0119] Step S6: Fusion of star-ground concentration field with time-activity pattern and calculation of pollutant exposure ( Figure 3 ).
[0120] This step fuses the high-resolution atmospheric pollutant concentration field obtained in step S3 with the user time-activity pattern table obtained in step S5 to calculate individual outdoor exposure indices at different time scales. The study period is discretized into equal-length time steps Δt, and the time corresponding to the k-th time step is denoted as Δt. k=1,2,…,N. For the u-th user, the raster index of its position at each time step can be obtained in step S5. And the type of activity scenario, including commuting, outdoor recreation, etc. Based on the high-resolution concentration field obtained in step S3, the environmental concentration of pollutant p at grid i and time step k is denoted as... Then the outdoor environmental concentration at user u's location at time step k can be obtained:
[0121]
[0122] For a given time interval [T1,T2], let the set of time steps it covers be K(T1,T2), then the cumulative outdoor exposure of user u to pollutant p can be expressed as:
[0123]
[0124] The corresponding time-weighted average outdoor exposure concentration is:
[0125]
[0126] Depending on the application scenario, different time windows can be selected to construct outdoor exposure indicators. For example, daily exposure: take [T1,T2] as 00:00–24:00 on a single day to obtain the daily cumulative outdoor exposure amount and the daily average outdoor exposure concentration; weekday exposure: take 09:00–18:00 on all weekdays as the time window to calculate the average outdoor exposure level during the working hours on weekdays; commuting time exposure: include the time steps of 07:00–09:00 and 17:00–19:00 on each day in the exposure calculation to obtain the outdoor exposure indicators during commuting hours.
[0127] For time steps in outdoor settings (including commuting / other outdoor activities), the respiratory rate derived from the time-activity pattern mapping is introduced. respiratory rate It is related to an individual's activity state (such as rest, commuting, exercise, etc.), and therefore depends on the individual's activity type. , representing the air inhalation rate of user u under different activity scenarios at time step k. The outdoor inhalation amount of individual u within the time interval [T1, T2] is calculated, that is, the total amount of pollutant p inhaled by the individual during that time period. The calculation method is to calculate the pollutant concentration within each time step. With the individual's respiratory rate Multiply by the product, then multiply by the time step Δt, and finally sum over all time steps, as shown in the following formula:
[0128]
[0129] To facilitate alignment with concentration-based indicators, the average inhalation volume per unit time is calculated:
[0130]
[0131] Step S7: Construction and statistical analysis of exposure indicators for multiple pollutants and multiple time scales ( Figure 3 ).
[0132] Based on user-level outdoor exposure calculations, outdoor exposure indicators with multiple time scales and multiple pollutant combinations are constructed according to research needs, and statistical analyses are performed at individual, population subgroup, and regional scales. Specifically, this includes:
[0133] 1) Multi-timescale outdoor exposure indicators: Daily scale: Calculate daily cumulative exposure and daily average exposure concentration using each calendar day as a window; Weekly / monthly scale: Calculate weekly / monthly cumulative exposure and average exposure concentration using a continuous week or month as a window, used to assess medium- to long-term exposure levels; Specific time-period scale: Such as "school hours" and "outdoor activity hours" of concern to high-exposure-sensitive populations, through flexible definition. accomplish.
[0134] 2) Comprehensive exposure index of multiple pollutants:
[0135] For multiple pollutants {p1, p2, ..., A comprehensive exposure index can be constructed by introducing weights based on toxicity, ambient air quality standard limits, or health risk coefficients, combined with the time-weighted average outdoor exposure concentration from step S6 within a given time window.
[0136]
[0137] When weight When using the ratio of pollutant concentration to standard limit or based on relative weights of health risk, this composite index can reflect the relative exposure level under the combined effects of different pollutants.
[0138] 3) Individual and population subgroup exposure statistics:
[0139] Users are grouped according to preset population classification rules (such as by residential administrative division, workplace grid, commuting distance, travel mode type, etc.), and the outdoor exposure index distribution of different population subgroups is calculated, including mean, median, standard deviation and different quantiles.
[0140] 4) Regional-scale outdoor exposure space analysis:
[0141] By spatially aggregating user-level exposure indicators into residential grids, workplace grids, or administrative units, the average or high-quantile outdoor exposure level of the population within each spatial unit is calculated, forming regional-scale outdoor exposure distribution maps and classification maps. This identifies areas with high outdoor exposure population concentrations and areas where outdoor exposure peaks frequently occur during specific time periods, providing a basis for subsequent health risk assessments and policy interventions.
[0142] Step S8: Uncertainty modeling and confidence assessment of exposure results.
[0143] This step models and analyzes the main sources of uncertainty in the outdoor concentration inversion, time-activity pattern identification, and exposure calculation processes, and provides a confidence assessment of the outdoor exposure indicators. Specifically, it includes:
[0144] 1) Uncertainty modeling of atmospheric pollutant concentration field:
[0145] Based on the model cross-validation results or station-based test results from step S3, estimate the prediction error distribution of different pollutants at different grid-time points. For example, assume that the error follows a zero-mean normal distribution or a log-normal distribution, and estimate the variance based on the statistical characteristics of the observation-prediction residuals. If necessary, stratified error models can be used for different regions or different concentration levels.
[0146] 2) Time-Activity Pattern Uncertainty Analysis:
[0147] The analysis examines the uncertainties that may be introduced in the identification thresholds for dwell points, the rules for identifying residences / workplaces, and the classification of activity types. For example, it assesses the stability of the grid and activity type to which a user belongs at different times by adjusting spatial and temporal thresholds and using different rule combinations for sensitivity analysis.
[0148] 3) Error propagation and Monte Carlo simulation:
[0149] Based on the aforementioned error model, Monte Carlo simulations are used to perturb the atmospheric pollutant concentration field and time-activity patterns. Specifically, in each simulation, the concentration at each grid-time point is randomly sampled according to the error distribution. Simultaneously, a set of possible time-activity trajectories is generated based on different parameter settings. Then, the corresponding outdoor exposure indicators are calculated according to steps S6 and S7. After multiple simulations (e.g., 500 or 1000), the outdoor exposure indicator distribution for each user and population subgroup is obtained.
[0150] 4) Result confidence assessment and visualization:
[0151] Based on the simulation results, confidence intervals (e.g., 95% confidence intervals), coefficients of variation, or confidence levels for user-level and population-level outdoor exposure indicators are calculated. High-uncertainty areas are overlaid on the spatial exposure distribution map to indicate that the assessment results for these areas require comprehensive judgment in conjunction with more monitoring data or field investigations, thereby improving the interpretability and reliability of the outdoor exposure assessment results of this invention in practical applications.
[0152] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0153] In another embodiment, an air pollution exposure assessment system integrating satellite ground measurement data and communication signaling is provided. This system corresponds one-to-one with the air pollution exposure assessment methods described in the above embodiments. Detailed descriptions of each functional module are as follows:
[0154] The spatiotemporal reference frame construction module is used to acquire spatial boundary information and basic geographic data of the target study area, divide the study area into regular grids to form a unified spatial grid system, set a grid resolution that is compatible with the spatial resolution of satellite remote sensing products and the positioning error of communication signaling, as well as a unified temporal resolution and time range.
[0155] The satellite-ground observation data acquisition and preprocessing module is used to collect pollutant concentration data, meteorological observation data, environmental parameter data and auxiliary meteorological data released by ground ambient air quality monitoring stations, perform quality control, outlier removal and missing data filling on the data, and unify data with different spatial and temporal resolutions to the spatial grid and time step through spatial resampling and temporal interpolation.
[0156] The pollutant concentration field construction module is used to construct a satellite-ground data fusion concentration inversion model by taking the measured pollutant concentrations of ground monitoring stations as supervisory variables and the preprocessed multi-source feature data as input. It predicts the outdoor pollutant concentrations of all grid-time points in the study area and generates a multi-pollutant high spatiotemporal resolution atmospheric environmental pollutant concentration field.
[0157] The signaling data processing module is used to acquire the original signaling data of the communication operator under the premise of meeting the requirements of data security and privacy protection, clean the signaling data and remove abnormal records, map the position of each signaling record to the spatial grid, discretize the timestamp to a unified time step, and generate discrete time series data in the form of "user-time step-grid".
[0158] The time-activity pattern recognition module is used to identify user dwell points based on preset time and space thresholds. It identifies residential grids and work grids by analyzing nighttime and weekday daytime dwell characteristics. It classifies dwell points by activity type by combining dwell time period, dwell duration and spatial information. It generates travel chains by path interpolation between adjacent dwell points and determines the grid and activity type of the user at each time step according to a unified time step.
[0159] The exposure calculation module is used to match the atmospheric environmental pollutant concentration field with the user's time-activity pattern, obtain the pollutant environmental concentration of the grid where the user is located at each time step, perform time integration on the environmental concentration for a given time window, calculate the user's cumulative outdoor exposure, and introduce the breathing rate parameter to calculate the dose-type exposure index for the time step in the outdoor activity scenario.
[0160] The multi-pollutant comprehensive exposure analysis module is used to weight and synthesize the exposure of multiple pollutants based on toxicity weights, ambient air quality standard limits or health risk coefficients, generate a multi-pollutant comprehensive exposure index, and perform statistical analysis on the exposure indicators of individuals and population subgroups at daily, weekly / monthly and specific time scales.
[0161] The uncertainty assessment module is used to construct a statistical model of concentration prediction error based on the validation results of the concentration inversion model, analyze the trajectory stability by identifying parameters through perturbation activity patterns, and use error propagation or Monte Carlo simulation methods to perturb the concentration field and time-activity patterns multiple times to calculate the confidence intervals of user-level and population-level exposure indicators.
[0162] The modules are connected sequentially through data interfaces to form a complete processing flow from raw data input to exposure assessment results output, enabling accurate assessment of air pollution exposure levels of individuals and populations at multiple time scales.
[0163] Specific limitations regarding the air pollution exposure assessment system can be found in the limitations of air pollution exposure assessment methods described above, and will not be repeated here. Each module in the above system can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0164] In another embodiment of the present invention, a terminal device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions to achieve a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used for the operation of an air pollution exposure assessment method.
[0165] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory). This computer-readable storage medium is a memory device in a terminal device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the terminal device and extended storage media supported by the terminal device. The computer-readable storage medium provides storage space that stores the terminal's operating system. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device.
[0166] One or more instructions stored in a computer-readable storage medium can be loaded and executed by a processor to implement the corresponding steps of the air pollution exposure assessment method in the above embodiments; one or more instructions in the computer-readable storage medium are loaded and executed by a processor.
[0167] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0168] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0169] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for air pollution exposure assessment by fusing satellite ground measurement data and communication signaling, characterized in that, Includes the following steps: S1: Constructing a unified spatiotemporal reference framework: Acquire spatial boundary information and basic geographic data of the target study area, divide the study area into regular grids to form a unified spatial grid system; set the grid resolution according to the spatial resolution of satellite remote sensing products and the positioning error of communication signaling data, and set a unified temporal resolution and time range according to application requirements, and map all subsequent data onto the spatiotemporal grid and time axis. S2: Acquisition and Preprocessing of Multi-Source Satellite-Ground Observation Data: Multi-source environmental data of the target area are collected, including pollutant concentration data and meteorological observation data released by ground ambient air quality monitoring stations, environmental parameter data and auxiliary meteorological data in satellite remote sensing products, and the above data are subjected to quality control, outlier removal and missing data filling. Through spatial resampling and temporal interpolation, data with different spatial and temporal resolutions are unified to the spatial grid and time step. S3: Construction of Atmospheric Pollutant Concentration Field Based on Satellite-Ground Data Fusion: Using the measured pollutant concentrations at ground-based ambient air quality monitoring stations as supervisory variables and the multi-source feature data obtained from S2 as input features, a satellite-ground data fusion concentration inversion model is constructed to predict the outdoor pollutant concentrations at all grid-time points within the study area, thereby obtaining a multi-pollutant high spatiotemporal resolution atmospheric environmental pollutant concentration field. S4: Acquisition and preprocessing of signaling data by telecommunications operators: Acquire raw signaling data generated by telecommunications operators within the target area, the signaling data including anonymous user identifiers, location information, and timestamps; The signaling data is cleaned, and the location of each signaling record is mapped to the spatial grid in S1. The timestamps are discretized to a uniform time step to obtain discrete time series data in the form of "user-time step-grid". S5: Time-Activity Pattern Recognition Based on Signaling Data: User dwell points are identified using preset time and space thresholds. Nighttime dwell features and weekday daytime dwell features for each user are constructed to identify residential grids and work grids. Dwell points are classified into different activity scenarios by combining dwell time period, dwell duration and spatial information, and path interpolation is performed between adjacent dwell points to generate travel chains containing dwell and movement information. The spatial grid where the user is located and its activity type are determined according to a unified time step to form a user-level time-activity pattern table. S6: Star-ground concentration field and time-activity pattern fusion and single pollutant exposure calculation: The multi-pollutant atmospheric environmental pollutant concentration field constructed in S3 is matched with the user time-activity pattern table obtained in S5 to obtain the pollutant environmental concentration of each user in the grid at each time step; for a given time window, the environmental concentration is integrated over time to obtain the cumulative outdoor exposure of each user in the time window. S7: Construction and Statistical Analysis of Multi-Pollutant Exposure Indicators at Multiple Time Scales Based on the user-level single pollutant exposure index, a weighting parameter is introduced to weight and synthesize the exposure amounts of multiple pollutants to obtain the multi-pollutant comprehensive outdoor exposure index; according to different time scales, the single pollutant exposure index and comprehensive exposure index of individuals and population subgroups are statistically analyzed. S8: Uncertainty Modeling and Confidence Assessment of Exposure Results: Based on the validation results of the star-ground concentration inversion model, a statistical model of the pollutant concentration prediction error at different grid-time points is constructed; by adjusting the activity pattern recognition parameters, the stability of the user's grid and activity type is analyzed. By employing error propagation or Monte Carlo simulation, the concentration field and time-activity pattern are perturbed to obtain confidence intervals for user-level and population-level exposure indicators.
2. The method of claim 1, wherein, In step S1: The resolution of the spatial grid is adaptively set based on the original spatial resolution of the satellite product and the positioning error of the communication signaling data; the time resolution is set to any value among 1 hour, 3 hours or 24 hours.
3. The method of claim 1, wherein, In step S2: The multi-source satellite-to-ground observation data includes at least: 1) Hourly concentrations of PM2.5, PM10, O3, and NO2 pollutants and routine meteorological data released by ground-based ambient air quality monitoring stations; 2) Data on aerosol optical depth (AOD), tropospheric NO2 column concentration, cloud cover, and surface albedo from satellite remote sensing products; 3) Further analyze the meteorological field, boundary layer height, and planetary boundary layer structure parameters; The data is unified to the spatial grid and time step using bilinear interpolation, nearest neighbor interpolation, linear interpolation, and spline interpolation methods.
4. The method of claim 1, wherein, In step S3: The satellite-to-ground data fusion concentration inversion model adopts the gradient boosting tree model XGBoost. Its input feature vector includes satellite observation variables, meteorological variables, land use characteristics and time variables, and the output is the pollutant concentration at the corresponding grid-time point. During the model training process, the measured concentration at the grid-time point where the monitoring station is located is used as the supervision sample, and cross-validation and / or station-based verification are used to evaluate the model performance.
5. The method of claim 1, wherein, In step S4: The communication operator signaling data includes records of various network events, such as location updates, call setup, SMS sending and receiving, and mobile data sessions. When preprocessing the raw signaling data, the events are first sorted by user identifier, and records with missing key fields or abnormal timestamps are deleted; then each record is mapped to the corresponding spatial grid, and the timestamp is discretized to the nearest time step.
6. The method of claim 1, wherein, The time-activity pattern recognition in step S5 includes: 1) Stop point identification: When a user is located in the same or adjacent spatial grid within a certain number of consecutive time steps, and the cumulative stop time exceeds the preset stop time threshold and the spatial displacement is less than the preset distance threshold, the time period is merged into a stop point; 2) Residential / Workplace Identification: Based on the distribution of users' stops during the nighttime hours, the duration and frequency of nighttime stays are statistically analyzed, and the grid with the highest duration and frequency of stays is designated as the residential grid; based on the distribution of stops during the weekday daytime hours, the characteristics of stays during the workday hours are statistically analyzed, and the grid with the highest percentage of stays during the weekday daytime hours is designated as the work grid. 3) Activity type identification and path interpolation: Based on the time period of the stay, the duration of the stay, the land use ratio within the grid, and the POI type information, the stay points are divided into different activity types; for the movement process between adjacent stay points, the user's location is interpolated on the time axis using the spatial grid adjacency matrix and road network accessibility index. 4) Time-Activity Pattern Generation: Based on a uniform time step, the stationary and moving states are mapped to the activity type and spatial grid corresponding to each time step, and a user-level time-activity pattern table is constructed.
7. The method of claim 1, wherein, In step S6: For time steps in commuting or other outdoor activity scenarios, the corresponding breathing rate is selected from preset breathing rate values based on the activity intensity level mapped by the activity scenario, and the environmental concentration is multiplied by the breathing rate to obtain a dose-type exposure index. Within a given time window, the exposure concentration or dose-related indicators at each time step are weighted to obtain the cumulative exposure and time-weighted average exposure concentration within that time window.
8. The method of claim 1, wherein, The method for constructing the multi-pollutant comprehensive exposure index in step S7 is as follows: For each pollutant i, the user-level exposure amount is calculated, and the weight is determined according to the pollutant's toxicity equivalent, health risk coefficient, or ambient air quality standard limit. The exposure amounts of each pollutant are then calculated using a weighted composite formula to obtain the multi-pollutant comprehensive exposure index.
9. The method of claim 1, wherein, The uncertainty modeling and exposure result evaluation in step S8 include: A grid-time point pollutant concentration error distribution was constructed based on the residuals of the star-ground concentration inversion model on the validation samples. By perturbing the dwell time identification threshold and activity type discrimination threshold within a reasonable range, multiple sets of candidate time-activity patterns are generated. At least 500 Monte Carlo simulations were performed to jointly perturb the pollutant concentration field and user trajectory, and user-level and population-level exposure indicators were repeatedly calculated to obtain the probability distribution of the exposure indicators, from which 95% confidence intervals were extracted.
10. The method according to any one of claims 1 to 9, characterized in that, The exposure assessment results are used as input for epidemiological exposure-response models, or for population exposure inequality analysis, regional health risk early warning, and the development of refined urban air quality management strategies.