Dynamic report generation method and system based on environmental health data fusion

By constructing a population exposure transmission network and multi-factor collaborative interaction test, combined with time-series pattern mining and parallel clustering processing, a hierarchical health assessment report is generated. This solves the problems of inaccurate identification of susceptible populations and lack of specificity in reports in existing technologies, and improves the accuracy of risk quantification and the practicality of reports.

CN121862292APending Publication Date: 2026-04-14WUXI XUEYUN DATA TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WUXI XUEYUN DATA TECH CO LTD
Filing Date
2025-12-31
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing health risk assessment methods fail to effectively identify the spatial distribution patterns of susceptible populations and the lagged response characteristics of health effects, neglect the synergistic interactions between pollutants and between pollutants and meteorological factors, lack a complete transmission path tracing from pollution sources to health effects, resulting in risk quantification indicators deviating from actual exposure and reports lacking specificity.

Method used

By constructing a population exposure transmission network, conducting distributed lag correlation analysis and multi-factor collaborative interaction tests, and combining time-series pattern mining and parallel clustering, a hierarchical health assessment report is generated. This achieves a deep correlation between pollutant concentration data and disease occurrence data, identifies susceptible population clusters and reconstructs exposure response chains, and generates dynamic reports using evidence weight fusion and dynamic length allocation.

Benefits of technology

It enables precise identification of vulnerable populations, improves the accuracy and relevance of risk quantification, identifies the contribution rate of dominant pollution sources, provides detailed risk analysis and protection recommendations, and enhances the practicality and operability of the report.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121862292A_ABST
    Figure CN121862292A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic report generation method and system based on environmental health data fusion, and the method comprises the steps: carrying out the time-space alignment of pollutant concentration data and disease occurrence data, generating a health association data set, constructing a crowd exposure transmission network, recognizing a susceptible crowd accumulation region, and constructing a layered exposure map; carrying out multi-factor collaborative interaction inspection to obtain a preliminary risk quantitative index; time sequence mode mining is carried out to establish a pollutant cumulative effect relationship, a composite exposure scene is identified, and spatial correction is carried out on the preliminary risk quantitative index to form a risk assessment matrix; parallel clustering processing is carried out on the risk assessment matrix to generate a risk grading identifier, factor source analysis is carried out to determine a dominant pollution source contribution table, an exposure response chain is reconstructed, and a health effect factor set is established; and based on the risk assessment matrix and the health effect factor set, content priorities are determined according to crowd hierarchy, and a hierarchical health assessment report is output, so that a comprehensive, accurate and operable scientific basis is provided for pollution prevention and control decisions and public health protection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental health risk assessment technology, and in particular to a method and system for generating dynamic reports based on environmental health data fusion. Background Technology

[0002] The impact of air pollutants on human health has become a significant issue in the field of public health. Combined exposure to different pollutants in the environment can produce cumulative health effects on susceptible populations, and these effects exhibit significant differences across different regions and populations. Existing health risk assessment methods typically separate pollutant concentration monitoring data from disease incidence data, lacking in-depth analysis of the spatiotemporal correlation between the two, making it difficult to accurately identify the spatial distribution patterns of susceptible populations and the lagged response characteristics of health effects.

[0003] Traditional assessment methods, when dealing with multi-pollutant exposure scenarios, often overlook the synergistic interactions between pollutants and between pollutants and meteorological factors, leading to deviations in risk quantification indicators from actual exposure levels. Furthermore, existing methods lack the ability to systematically trace the complete transmission pathway from pollution sources to health effects, failing to effectively identify the contribution patterns and regional differences of dominant pollution sources. In the risk assessment report generation stage, traditional methods employ a standardized content organization approach, failing to provide targeted assessment information and protective recommendations based on the varying risk levels and susceptibility of different population groups. Therefore, a method is urgently needed to address at least one of the aforementioned problems. Summary of the Invention

[0004] This invention discloses a dynamic report generation method and system based on environmental health data fusion. It aims to achieve deep correlation between pollutant concentration data and disease occurrence data, integrating pollution transmission path tracing. The system constructs a population exposure transmission network from health-related datasets and performs distributed lag correlation analysis. It quantifies the compound exposure effect through multi-factor collaborative interaction testing, establishes pollutant cumulative effect relationships through time-series pattern mining, reconstructs the exposure response chain through parallel clustering and pollution source analysis, and establishes a set of health effect factors. Finally, it generates a hierarchical health assessment report based on evidence weight fusion and dynamic length allocation, providing a scientific basis for environmental health management and public health protection.

[0005] The first aspect of this invention proposes a dynamic report generation method based on environmental health data fusion, comprising the following steps: Acquire pollutant concentration data and disease occurrence data, and perform spatiotemporal alignment on the pollutant concentration data and disease occurrence data to generate a health-related dataset; Based on the aforementioned health-related dataset, a population exposure transmission network is constructed. Distributional lag association analysis is performed on this network to identify susceptible population clusters. Differential response features are extracted from these clusters to construct a hierarchical exposure map. The hierarchical exposure map is then used to conduct multi-factor collaborative interaction tests to obtain preliminary risk quantification indicators. Time-series pattern mining is performed on the aforementioned health-related dataset to establish pollutant cumulative effect relationships. Based on these relationships, health impact weights for different regions are determined. These weights are then used to identify compound exposure scenarios. Finally, the preliminary risk quantification indicators are spatially corrected by region according to the compound exposure scenarios to form a risk assessment matrix. Parallel clustering is performed on the risk assessment matrix to generate risk classification identifiers. Factor source analysis is then performed on the risk classification identifiers to determine the contribution table of dominant pollution sources. Key impact pathways are extracted from the contribution table of dominant pollution sources to reconstruct the exposure-response chain. Based on the exposure-response chain, regional difference characteristics are extracted to establish a set of health effect factors. Based on the risk assessment matrix and the set of health effect factors, content priorities are determined by population stratification, and stratified health assessment reports are output using the content priorities.

[0006] A second aspect of this invention proposes a dynamic report generation system based on environmental health data fusion, comprising: The data acquisition module is used to acquire pollutant concentration data and disease occurrence data, and to perform spatiotemporal alignment of the pollutant concentration data and the disease occurrence data to generate a health-related dataset; The exposure analysis module is used to construct a population exposure transmission network based on the health association dataset, perform distributional lag association analysis on the population exposure transmission network to identify susceptible population clusters, extract differential response features from the susceptible population clusters to construct a hierarchical exposure map, and use the hierarchical exposure map to conduct multi-factor collaborative interaction tests to obtain preliminary risk quantification indicators; The risk assessment module is used to conduct time-series pattern mining on the health-related dataset to establish pollutant cumulative effect relationships, determine health impact weights for different regions based on the pollutant cumulative effect relationships, identify compound exposure scenarios using the health impact weights, and perform regional spatial corrections on the preliminary risk quantification indicators according to the compound exposure scenarios to form a risk assessment matrix; The factor reconstruction module is used to perform parallel clustering processing on the risk assessment matrix to generate risk classification identifiers, perform factor source analysis on the risk classification identifiers to determine the contribution table of dominant pollution sources, extract key impact paths from the contribution table of dominant pollution sources to reconstruct the exposure response chain, and extract regional difference features based on the exposure response chain to establish a health effect factor set; The report generation module is used to determine the content priority by population stratification based on the risk assessment matrix and the health effect factor set, and to output a stratified health assessment report using the content priority.

[0007] The beneficial effects of this invention are reflected in the following aspects: First, by using spatiotemporal alignment technology to deeply correlate pollutant concentration data with disease occurrence data, it solves the problem of inaccurate identification of susceptible populations caused by the separation and processing of the two types of data in traditional methods, and achieves accurate positioning of susceptible population clusters. Second, by employing distributed lag correlation analysis to construct a multi-scale lag effect response model, it identifies the differentiated response characteristics of acute and chronic exposure paths, overcoming the risk assessment bias caused by neglecting lag effects. Third, by constructing a multi-factor interaction matrix and using stepwise regression screening, it identifies the synergistic enhancement factor combination of pollutants and meteorological factors, solving the problem that single-factor assessment cannot reflect the true risk of compound exposure, and improving the accuracy and relevance of risk quantification. Fourth, addressing the problem that traditional methods neglect the synergistic effect of multi-pollutant compound exposure, this invention establishes the cumulative effect relationship of pollutants through time-series pattern mining and determines the health impact weights of different regions. Based on the health impact weights, it generates a comprehensive exposure score, identifies the periods of multi-pollutant exceedances, extracts exceedance combination patterns, and summarizes them into a compound exposure scenario library. Based on the compound exposure scenarios, it corrects the risk enhancement coefficient of the preliminary risk quantification indicators, making the risk assessment results closer to the actual exposure situation, avoiding the risk underestimation problem caused by single-pollutant assessment, and improving the risk prediction capability under compound exposure scenarios. Finally, addressing the shortcomings of traditional methods in tracing pollution sources and producing monotonous reports, this invention employs a parallel clustering algorithm and a positive definite matrix factorization model to identify dominant pollution sources and quantify their contribution rates. By constructing a complete five-stage transmission path to reconstruct the exposure-response chain, it achieves full-chain tracing from pollution source to health effects, providing clear source information for pollution prevention and control. Furthermore, it establishes a multi-source data fusion mechanism based on evidence weights and adopts a dynamic length allocation algorithm to automatically allocate report content length and determine content priorities according to the risk levels and population sizes of different groups. This solves the problem that standardized reports cannot meet differentiated needs, enabling high-risk groups to obtain more detailed risk analysis and protection recommendations, thus improving the practicality and operability of the reports. Attached Figure Description

[0008] The accompanying drawings illustrate specific examples of the technical solutions described in this invention and, together with the detailed embodiments, form part of the specification, serving to explain the technical solutions, principles, and effects of this invention.

[0009] Unless otherwise specified, the same reference numerals in different figures represent the same or similar technical features, and different reference numerals may be used to represent the same or similar technical features.

[0010] Figure 1 This is a flowchart illustrating a dynamic report generation method based on environmental health data fusion according to the present invention.

[0011] Figure 2 This is a structural block diagram of a dynamic report generation system based on environmental health data fusion according to the present invention. Detailed Implementation

[0012] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0013] It should be understood that, when used in this application specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or a collection thereof.

[0014] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0015] The technical solutions of the embodiments of this application will be described below.

[0016] like Figure 1 As shown, this embodiment of the invention provides a dynamic report generation method based on environmental health data fusion, including the following steps S110-S150: Step S110: Obtain pollutant concentration data and disease occurrence data, and perform spatiotemporal alignment on the pollutant concentration data and disease occurrence data to generate a health-related dataset.

[0017] Specifically, pollutant concentration data and disease incidence data are acquired. Pollutant concentration data is obtained from environmental monitoring stations distributed throughout the study area, which record the concentrations of major atmospheric pollutants in real time. Pollutant concentration data includes hourly concentrations of six indicators: PM2.5, PM10, SO2, NO2, CO, and O3. Data is collected hourly, covering a continuous multi-year period. Each concentration data record includes the monitoring station number, geographical coordinates (longitude and latitude), monitoring timestamp, and the concentration value of each pollutant. The spatial coverage of the concentration data is determined based on the distribution of monitoring stations, typically one station is set up every 5-10 square kilometers within the urban built-up area. Disease incidence data is obtained from medical and health institutions, including general hospitals, specialized hospitals, and community health service centers. Disease incidence data records patient medical information, including the patient's residential address, date of visit, type of diagnosed disease, and severity of disease. Emphasis is placed on the incidence of respiratory and cardiovascular diseases. Respiratory diseases include asthma, chronic obstructive pulmonary disease (COPD), and pneumonia; cardiovascular diseases include hypertension, coronary heart disease, and stroke. Disease data is recorded with daily time accuracy, and spatial information is converted into geographic coordinates using patient residence addresses. Quality control is performed on pollutant concentration data to identify and remove outliers, including out-of-range data due to equipment malfunction and measurements significantly deviating from the normal range. Missing concentration data for different time periods is interpolated using the average concentration from adjacent stations or time-series interpolation. Disease occurrence data is anonymized, removing patient names, ID numbers, and other private information while retaining statistically relevant information such as disease type, onset time, and place of residence.

[0018] A health-related dataset is generated by spatiotemporally aligning pollutant concentration data and disease occurrence data. Pollutant concentration data and disease occurrence data are matched and aligned in both temporal and spatial dimensions to establish a correspondence between pollution exposure and disease occurrence. For temporal alignment, the lag effect of pollutant health effects is considered, as there is a time delay between pollutant exposure and disease onset. For acute health effects, a short lag is used, matching the onset date in the disease data with concentration data from 1-7 days prior. For chronic health effects, a long lag is used, matching the disease data with the average concentration data from 1-3 months prior. Different lag periods are set according to disease type; acute respiratory attacks typically use a 1-3 day lag, while cardiovascular disease attacks typically use a 3-7 day lag. For spatial alignment, the patient's residential coordinates in the disease data are matched with the concentration data from the nearest monitoring station. The distance between the patient's residence and each monitoring station is calculated, and the nearest station is selected as the source of pollution exposure for that patient. When the patient's residence is located between multiple monitoring stations, an inverse distance weighting method is used to spatially interpolate the concentration data from multiple stations, with the interpolation weight inversely proportional to the square of the distance. The aligned data forms a health-related dataset, in which each record contains the patient's disease information, onset time, residential coordinates, and pollutant concentration values ​​at the corresponding spatiotemporal location.

[0019] Step S120: Construct a population exposure transmission network based on the health association dataset, perform distributional lag association analysis on the population exposure transmission network to identify susceptible population clusters, extract differential response features from susceptible population clusters to construct a hierarchical exposure map, and use the hierarchical exposure map to conduct multi-factor collaborative interaction tests to obtain preliminary risk quantification indicators.

[0020] Specifically, a population exposure transmission network is constructed based on a health-related dataset. The residential coordinates and pollutant exposure concentrations of patients in the dataset are mapped to network nodes, and pollutant propagation paths are mapped to network edges. Spatial connections are established based on the geographical proximity of patients in the dataset; when the distance between the residences of two patients is less than 500 meters, a spatial connection edge is established between the corresponding nodes. Exposure connections are established based on the similarity of exposure concentrations among patients in the dataset. The correlation coefficient of pollutant exposure doses between patients is calculated; when the correlation coefficient is greater than 0.7, an exposure connection edge is established. In the network, node attributes include the patient's disease type, onset time, exposure concentration, and demographic characteristics; edge weights represent the spatial distance or exposure similarity between nodes. The spatiotemporal information of the health-related dataset is mapped to the nodes and edges of the network, forming a dynamically evolving population exposure transmission network. The temporal resolution of the network is set to daily, and the spatial resolution is determined by the location accuracy of the patients' residences.

[0021] In some embodiments, the step of performing distributional lag association analysis on the population exposure transmission network to identify susceptible population clusters includes: decomposing the population exposure transmission network into acute exposure paths and chronic exposure paths according to the exposure duration dimension; constructing multi-scale lag effect responses for the acute exposure paths and the chronic exposure paths to obtain spatiotemporal interactive lag response curves; identifying the stratified susceptibility intensity of the population based on the spatiotemporal interactive lag response curves to generate a stratified susceptibility feature table; and performing high susceptibility region extraction on the stratified susceptibility feature table to select high susceptibility clusters as susceptible population clusters.

[0022] The population exposure transmission network was decomposed into acute exposure pathways and chronic exposure pathways based on the exposure duration. Pathways were classified according to the duration of contaminant exposure for patients in the network, with exposure duration determined by the cumulative exposure time of contaminants before the onset of illness. For each node in the network, the number of days the corresponding patient was continuously exposed to high concentrations of contaminants before the onset of illness was counted. When the continuous exposure days were less than 7 days, the node and its connected edges were classified as acute exposure pathways, corresponding to health effects caused by short-term, high-intensity pollution events. When the continuous exposure days were greater than 30 days, the node and its connected edges were classified as chronic exposure pathways, corresponding to health effects caused by long-term, cumulative pollution exposure. For nodes with exposure durations between 7 and 30 days, the pathway type was determined based on the trend of contaminant concentration changes: those showing a rapid upward trend in concentration were classified as acute exposure pathways, while those maintaining a consistently high concentration were classified as chronic exposure pathways.

[0023] Multi-scale lag effect response models were constructed for both acute and chronic exposure pathways to obtain spatiotemporal interactive lag response curves. For patients in the acute exposure pathway, the short-term lag relationship between pollutant exposure and disease onset was analyzed. The onset time of patients in the acute pathway was paired with the exposure time, and the time delay from peak pollutant concentration to disease onset was calculated; this time delay is called the lag period. The relative risk of disease onset at different lag periods was calculated; the relative risk represents the multiple of the probability of disease occurrence at that lag period relative to the baseline level. The relationship curve between lag period and relative risk was plotted, reflecting the lag effect response characteristics of the acute pathway. For patients in the chronic exposure pathway, the lag relationship between long-term cumulative exposure and disease onset was analyzed. The cumulative exposure dose of patients in the chronic pathway was calculated within different time windows: 1 month, 3 months, and 6 months. The association strength between the cumulative exposure dose and disease onset risk in each time window was analyzed; the association strength was quantified using correlation coefficients. A multi-scale lag response model was constructed by combining the lag effects of the acute and chronic exposure pathways. The multi-scale model includes short-term and long-term lag dimensions. The short-term lag dimension covers 0-7 days, and the long-term lag dimension covers 1-6 months. Lag response characteristics are plotted on a two-dimensional spatiotemporal plane, with the horizontal axis representing lag time and the vertical axis representing spatial location. Color depth indicates relative risk intensity, forming a spatiotemporal interactive lag response curve. The curve illustrates the disease risk response patterns in different spatial regions at different lag times, identifying the spatiotemporal locations where risk peaks occur.

[0024] Based on spatiotemporal interactive lag response curves, a stratified susceptibility intensity identification method is used to generate a stratified susceptibility characteristic table. The spatiotemporal interactive lag response curves contain relative risk values ​​for each spatial location at different lag times. For each spatial location, the maximum relative risk value over all lag times is calculated; this maximum value serves as the susceptibility intensity index for that location. The susceptibility intensity index reflects the sensitivity of the population in that area to pollutant exposure; a higher value indicates a greater susceptibility to illness due to pollutant exposure. Based on the distribution of the susceptibility intensity index in the curve, the population is divided into three levels: high susceptibility, moderate susceptibility, and low susceptibility. Stratification thresholds are set: a susceptibility intensity index greater than 2.0 is classified as high susceptibility, between 1.2 and 2.0 as moderate susceptibility, and less than 1.2 as low susceptibility. High susceptibility typically includes children, the elderly, and people with underlying diseases; moderate susceptibility includes the general adult population; and low susceptibility includes healthy young and middle-aged adults. The susceptibility intensity index, susceptibility hierarchy classification, and demographic characteristics are compiled into a hierarchical susceptibility characteristic table, which includes fields such as spatial location, susceptibility intensity, population stratification, and characteristic attributes.

[0025] High-susceptibility regions were extracted from the hierarchical susceptibility feature table, and high-susceptibility clusters were selected as areas of concentration for susceptible populations. The hierarchical susceptibility feature table was filtered to obtain records with a high susceptibility level, corresponding to the spatial locations of high-susceptibility populations. The geographic coordinates of these high-susceptibility records were marked on a map as high-risk locations. The spatial distribution patterns of these high-susceptibility locations were analyzed to identify the clusters formed by these locations. A spatial clustering algorithm was used to cluster these high-susceptibility locations. The DBSCAN density clustering method was selected, as it can identify clusters of arbitrary shapes. Clustering parameters were set: a neighborhood radius of 1 km and a minimum number of points of 10, ensuring that the identified clusters were statistically significant. The clustering algorithm grouped spatially adjacent and sufficiently dense high-susceptibility locations into the same cluster, forming high-susceptibility clusters. Each cluster represents an area where a high-susceptibility population is concentrated, and the extent of the cluster is determined by the outer envelope of all points within the cluster. Calculate the mean susceptibility intensity within each cluster; a higher mean indicates a stronger overall susceptibility of the population in that area. Calculate the population size within each cluster, estimated using resident population statistics within the cluster area. Identified high-susceptibility clusters are marked as susceptible population gathering areas, which are key areas for pollutant health risk prevention and control.

[0026] Differential response characteristics were extracted from susceptible population clusters to construct stratified exposure maps. For each susceptible population cluster, pollutant exposure characteristics and population response characteristics were analyzed. Pollutant exposure characteristics included the average concentration, peak concentration, coefficient of variation, and cumulative exposure dose of each pollutant. Population response characteristics included disease incidence, relative risk value, lag distribution, and disease type composition. By comparing the exposure and response characteristics of different clusters, differentiated response patterns were identified. Some clusters showed significant responses to PM2.5 exposure, with a marked increase in respiratory disease incidence when PM2.5 concentrations rose, while the effects of other pollutants were relatively small. Some clusters showed significant responses to NO2 exposure, with a marked increase in cardiovascular disease incidence when NO2 concentrations rose. Clusters were stratified and classified according to the main responding pollutant type and response intensity. Areas with high PM2.5 response were classified as PM2.5 sensitive layers, areas with high NO2 response as NO2 sensitive layers, and areas responding to multiple pollutants as multi-pollutant sensitive layers. Clusters were colored on the map according to their sensitivity layer type, with different colors representing different sensitivity layers. By integrating the spatial distribution of clusters, the classification of sensitive layers, exposure characteristics, and response characteristics into a single map, a stratified exposure map is formed. This map visually illustrates the exposure response characteristics of different populations at different locations within the study area to different pollutants, providing spatial evidence for precise health risk management.

[0027] In some embodiments, the step of using the stratified exposure map to conduct a multi-factor synergistic interaction test to obtain a preliminary risk quantification index includes: extracting pollutant concentration variables and meteorological factor variables from the stratified exposure map to construct a multi-factor interaction matrix; performing interaction term regression analysis on the multi-factor interaction matrix to obtain an interaction effect coefficient table; screening the interaction effect coefficient table for significance levels to identify synergistic enhancement factor combinations; and weighting the synergistic enhancement factor combinations according to interaction intensity to form a preliminary risk quantification index.

[0028] A multi-factor interaction matrix was constructed by extracting pollutant concentration variables and meteorological factor variables from a stratified exposure map. Based on pollutant concentration data from each sensitive layer region of the stratified exposure map, daily average concentrations of PM2.5, PM10, SO2, NO2, CO, and O3 were obtained. Based on the spatiotemporal range marked on the stratified exposure map, corresponding meteorological factor variables, including daily average temperature, relative humidity, wind speed, air pressure, and precipitation, were obtained. The pollutant concentration variables and meteorological factor variables were aligned temporally and spatially to ensure that the data for each variable originated from the same spatiotemporal location. A multi-factor interaction matrix was constructed, where each row corresponds to a different observation sample, with each sample representing an observation record at a specific spatiotemporal location. The columns of the matrix correspond to various factor variables, including 6 pollutant concentration variables and 5 meteorological factor variables, totaling 11 variables. The element values ​​of the multi-factor interaction matrix are the measured values ​​of the corresponding samples for their respective variables, with units determined according to the variable type: pollutant concentration in μg / m³, temperature in °C, humidity in %, wind speed in m / s, and air pressure in hPa. The variables in the matrix are standardized to eliminate differences in dimensions. The standardization method is Z-score standardization, which converts each variable into a standard normal distribution with a mean of 0 and a standard deviation of 1.

[0029] For example, the step of performing interaction term regression analysis on the multi-factor interaction matrix to obtain an interaction effect coefficient table includes: extracting each factor variable in the multi-factor interaction matrix to construct interaction terms and generate an interaction term feature set; performing stepwise regression screening on the interaction term feature set to obtain a significant interaction term subset; using the significant interaction term subset to estimate effect coefficients and generate an effect coefficient vector; and organizing the effect coefficient vector according to the factor combination type to form an interaction effect coefficient table.

[0030] Interaction term feature sets are generated by extracting interaction terms from the factor variables in the multi-factor interaction matrix. Each factor variable in the multi-factor interaction matrix is ​​paired, including 6 pollutant concentration variables and 5 meteorological factor variables. Interaction term features are constructed by multiplication, where X_ij = X_i × X_j, and X_i and X_j are the i-th and j-th factor variables, respectively, and X_ij is the interaction term variable. For example, the interaction term between PM2.5 and temperature is the product of PM2.5 concentration and temperature, and the interaction term between NO2 and relative humidity is the product of NO2 concentration and relative humidity. Since the factor variables have been standardized, the interaction term features are dimensionless values ​​and can be directly used for subsequent analysis. Pairwise combinations of the 11 factor variables result in C(11,2) = 55 second-order interaction term features. The original 11 factor variables and 55 interaction term features are merged to form an interaction term feature set containing 66 features. Each column of the feature set corresponds to a feature variable, and each row corresponds to an observed sample. The feature set fully describes the interaction patterns among multiple factors.

[0031] A significant subset of interaction terms is obtained by performing stepwise regression screening on the interaction term feature set. Using the interaction term feature set as the independent variable and the disease incidence rate or relative risk value as the dependent variable, a multiple linear regression model is constructed. Stepwise regression is used for variable screening, which includes two processes: forward selection and backward elimination. In the forward selection process, interaction term features with the strongest explanatory power for the dependent variable are introduced one by one from the interaction term feature set. After introducing one feature at a time, the goodness of fit and significance level of the model are calculated. In the backward elimination process, interaction term features that do not meet the significance level requirement are removed from the model. The significance level threshold is set to α=0.05. After multiple rounds of iterative screening, interaction term features that pass the significance level test are retained, and these features constitute the significant subset of interaction terms. The number of interaction terms in the significant subset is usually much smaller than the original 55 interaction terms; typically, 10-20 significant interaction terms are retained. Each interaction term in the subset represents a significant synergistic effect of the corresponding factor combination on disease incidence, and the strength of the effect is quantified by regression coefficients.

[0032] Effect coefficient vectors are generated by estimating effect coefficients using a subset of significant interaction terms. A multiple linear regression model is then reconstructed based on this subset of significant interaction terms, with the model form Y = β0 + Σ(β... k ×X_k)+ε, where Y is the dependent variable (relative risk of disease), X_k is the k-th interaction term feature in the subset, and β k Let β0 be the corresponding regression coefficient, β0 be the intercept term, and ε be the error term. The least squares method is used to estimate the regression coefficients, minimizing the sum of squared residuals. The regression coefficients β0 for each interaction term are then calculated. kThe coefficient values ​​reflect the direction and strength of the interaction term's impact on disease risk. A positive coefficient indicates that the interaction term increases disease risk, while a negative coefficient indicates that the interaction term decreases disease risk. The larger the absolute value of the coefficient, the stronger the impact. Since the factor variables are standardized, the regression coefficients are comparable, allowing direct comparison of the relative importance of different interaction terms. The regression coefficient values ​​of all significant interaction terms are arranged in order of interaction terms to form an effect coefficient vector. The dimension of the vector equals the number of interaction terms contained in the subset of significant interaction terms, and each element in the vector corresponds to the effect coefficient of one interaction term. For example, if the subset contains 15 interaction terms, the effect coefficient vector will be a 15-dimensional vector.

[0033] The effect coefficient vectors are organized into an interaction effect coefficient table according to factor combination types. Each element in the effect coefficient vector is associated with its corresponding interaction term information. The factor combination involved in each interaction term is identified; for example, one interaction term might be a combination of PM2.5 and temperature, while another might be a combination of NO2 and relative humidity. Interaction terms are categorized according to factor combination type, grouping those involving the same factor type together. Pollutant-pollutant interactions are grouped into pollutant-pollutant interaction groups, and pollutant-meteorological factor interactions are grouped into pollutant-meteorological interaction groups. An interaction effect coefficient table is constructed, containing four fields: Factor 1, Factor 2, Interaction Term Type, and Effect Coefficient. The Factor 1 and Factor 2 fields record the names of the two factor variables involved in the interaction term, the Interaction Term Type field indicates the group to which the interaction term belongs, and the Effect Coefficient field records the corresponding coefficient value. The interaction effect coefficient table systematically displays the synergistic effect strength of each factor combination, providing data support for identifying synergistically enhancing factor combinations.

[0034] The interaction effect coefficient table is used to screen for synergistic enhancement factor combinations based on significance levels. Interactions with positive effect coefficients are selected, indicating a synergistic enhancement effect of the factor combination on disease risk. These positive coefficient interactions are then tested for significance, and the p-value of each interaction coefficient is calculated. Interactions with p-values ​​less than 0.05 are considered statistically significant. Interactions with positive effect coefficients and p-values ​​less than 0.05 are selected; the factor combinations corresponding to these interactions are the synergistic enhancement factor combinations. The compositional characteristics of these synergistic enhancement factor combinations are analyzed to identify which combinations of pollutants and meteorological factors have a significant synergistic amplification effect on health risk. The effect coefficient values ​​for each synergistic enhancement factor combination are obtained; these coefficient values ​​represent the strength of the combination's enhancement of disease risk, known as the interaction strength. The synergistic enhancement factor combination records include the types of factors involved, their corresponding standardized variable values, and the interaction strength coefficient.

[0035] The synergistic enhancement factor combinations are weighted according to their interaction strength to form a preliminary risk quantification index. Weighted calculations are performed on each combination within the synergistic enhancement factor combinations, recording the standardized variable values ​​and interaction strength coefficients of the two factors in each combination. The weighted interaction effect of each combination is calculated as E_i = β_i × X_i1 × X_i2, where E_i is the weighted interaction effect of the i-th combination, β_i is the interaction strength coefficient of that combination, and X_i1 and X_i2 are the standardized variable values ​​of the two factors in the combination, respectively. The weighted interaction effects of all synergistic enhancement factor combinations are summed to obtain the comprehensive interaction effect E_total = ΣE_i. The comprehensive interaction effect reflects the overall health risk intensity under the synergistic effect of multiple factors, and this value is the preliminary risk quantification index. The unit of the preliminary risk quantification index is a dimensionless relative risk value; a higher value indicates a higher risk of disease occurrence under the synergistic effect of multiple factors. The preliminary risk quantification index is standardized and mapped to a risk level scoring range of 0-100 for easy risk assessment and decision-making applications. The indicator integrates multi-dimensional information such as pollutant concentration, meteorological factors, and population susceptibility, laying a quantitative foundation for subsequent refined risk assessment.

[0036] Step S130: Conduct time-series pattern mining on the health-related dataset to establish the cumulative effect relationship of pollutants, determine the health impact weights of different regions based on the cumulative effect relationship of pollutants, use the health impact weights to identify compound exposure scenarios, and perform regional spatial correction on the preliminary risk quantification indicators according to the compound exposure scenarios to form a risk assessment matrix.

[0037] Specifically, time-series pattern mining was conducted on the health association dataset to establish the cumulative effect relationship of pollutants. Time-series data on pollutant concentrations and disease incidence were organized in the health association dataset and arranged on a daily basis. Cumulative effect analysis was performed on the pollutant concentration sequences in the dataset to calculate the cumulative exposure dose within different time windows. The cumulative exposure dose calculation formula is D(t,w)=Σ[C(ti)], where D(t,w) represents the cumulative value of pollutant concentration within a w-day window before time t, i ranging from 0 to w-1, and C(ti) is the pollutant concentration on day ti, in μg / m³. The cumulative dose unit is μg·day / m³. Multiple time window lengths were set, including 3 days, 7 days, 14 days, and 30 days, corresponding to short-term, medium-term, and long-term cumulative effects, respectively. Association analysis was performed on the cumulative exposure dose and disease incidence sequences for each time window, and a distributed lag nonlinear model was used to evaluate the dose-response relationship between cumulative exposure and disease risk. The dose-response relationship curve describes the changing trend of relative disease risk as the cumulative exposure dose increases, and the slope of the curve reflects the intensity of the cumulative effect of pollutants. The dose-response parameters for each pollutant at different time windows were obtained, including baseline risk, dose coefficient, and threshold concentration. The cumulative effect window, dose-response parameters, and correlation strength of each pollutant were integrated into a cumulative effect relationship for the pollutants.

[0038] The health impact weights for different regions are determined based on the cumulative effect relationships of pollutants. These relationships include the dose coefficients and correlation strength indices for each pollutant. The dose coefficient reflects the increase in relative disease risk per unit increase in pollutant concentration, while the correlation strength is quantified using the coefficient of determination (R²). The dose coefficients of each pollutant are standardized to a comparable dimensionless scale. The comprehensive impact intensity of each pollutant is calculated, combining the standardized dose coefficient and the correlation strength (R²). A higher R² value indicates a more significant health impact. Local dose-response parameters are obtained for different spatial locations within the study area. Due to differences in population structure, baseline health status, and environmental conditions across regions, the health impacts of the same pollutant may vary. The comprehensive impact intensity of pollutants is calculated for each region to identify the main health-affecting pollutants. In densely industrialized areas, SO2 and PM10 have higher comprehensive impact intensities, while in densely trafficked areas, NO2 and CO have higher comprehensive impact intensities. The overall impact intensity of each pollutant in each region is normalized so that the sum of the impact intensities of all pollutants in each region equals 1. The normalized overall impact intensity is the health impact weight of each pollutant in that region. A health impact weight matrix is ​​constructed, where rows correspond to different spatial regions, columns correspond to different pollutants, and matrix elements represent the health impact weight of the corresponding pollutant in the corresponding region.

[0039] In some embodiments, the step of using the health impact weights to identify composite exposure scenarios includes: extracting pollutant weight vectors from the health impact weights; performing weighted summation on the pollutant weight vectors according to concentration standardization to generate a comprehensive exposure score; identifying multiple pollutant exceedance periods and extracting exceedance combination patterns on the comprehensive exposure score; and summarizing the exceedance combination patterns into composite exposure scenarios.

[0040] Pollutant weight vectors are extracted from the health impact weights. The weight data for each region in the health impact weights have been normalized and weighted, with each region corresponding to a weight vector containing six elements, representing the health impact weights of PM2.5, PM10, SO2, NO2, CO, and O3, respectively. The sum of the elements in the weight vector equals 1 to ensure the normalization constraint of the weights. The distribution characteristics of the weight vectors in different regions are analyzed to reflect the differences in regional pollution characteristics. In the weight vectors of densely industrialized areas, the SO2 weight is typically 0.25-0.35, and the PM10 weight is 0.30-0.40, reflecting the dominant role of industrial emissions in the health impact of this region. In the weight vectors of densely trafficked areas, the NO2 weight is typically 0.30-0.45, and the PM2.5 weight is 0.25-0.35, reflecting the characteristics of traffic pollution. The weight vectors of residential areas show a balanced distribution of multiple pollutants, with the weights of each pollutant ranging from 0.15 to 0.20, and no clearly dominant pollutant. The weight vector for commercial areas falls between that of industrial and residential areas, with PM2.5 weighting at 0.25-0.30 and NO2 weighting at 0.20-0.28. The pollutant weight vector records the relative impact intensity of each pollutant in the area and is used for subsequent comprehensive exposure assessment. The vectors represent PM2.5, PM10, SO2, and NO2, respectively. The pollutant weight vector is standardized by concentration and then weighted and summed to generate a comprehensive exposure score. Pollutant concentration values ​​for each region at each time point are read from pollutant monitoring data, which includes the measured concentrations of six pollutants at each time point. Pollutant concentrations are standardized to eliminate differences in concentration magnitudes between different pollutants. The standardized concentration calculation formula is C_norm = C_actual / C_standard, where C_norm is the standardized concentration, C_actual is the measured concentration, and C_standard is the daily average standard limit for that pollutant. The daily average standard limits are 75 μg / m³ for PM2.5, 150 μg / m³ for PM10, 150 μg / m³ for SO2, 80 μg / m³ for NO2, 4 mg / m³ for CO, and 160 μg / m³ for O3. After standardization, C_norm = 1 indicates that the concentration meets the standard limit, and C_norm > 1 indicates that the concentration exceeds the limit. The standardized concentration vector is then weighted and summed with the pollutant weight vector to calculate the comprehensive exposure score. The weighted summation formula is E_i(t) = Σ[w_ij × C_normj(t)], where E_i(t) is the overall exposure score of the i-th region at time t, w_ij is the weight of the j-th pollutant in that region, and C_normj(t) is the standardized concentration of that pollutant at time t. The overall exposure score comprehensively reflects the exposure intensity under the combined effect of multiple pollutants; a higher score indicates a greater exposure risk. When the score is greater than 1, it indicates that the overall exposure level of multiple pollutants in that region exceeds the health risk benchmark.

[0041] This study identifies multiple pollutant exceedance periods and extracts exceedance combination patterns based on the comprehensive exposure score. Exceedance thresholds for the comprehensive exposure score are set: 1.0 for mild risk, 1.5 for moderate risk, and 2.0 for severe risk. Periods exceeding these thresholds are identified from the comprehensive exposure score time series; these periods are considered multiple pollutant exceedance periods. For each exceedance period, the concentration levels of each pollutant are analyzed to identify which pollutants are exceeding the limits. Pollutant exceedance is determined based on whether its standardized concentration (C_norm) is greater than 1; the exceeding pollutants constitute the pollutant combination for that period. The characteristics of pollutant combinations for each exceedance period are recorded, including the types of pollutants included in the combination, the exceedance multiple of each pollutant, and the overall risk intensity of the combination. Pattern recognition is performed on all pollutant combinations for exceedance periods within the study period, and the frequency of occurrence of various combination patterns is statistically analyzed. Common exceedance combination patterns include PM2.5 + PM10 dual particulate matter exceedance pattern, NO2 + PM2.5 traffic pollution exceedance pattern, O3 + NO2 photochemical pollution exceedance pattern, and multi-pollutant comprehensive exceedance pattern, etc.

[0042] Exceeding pollution levels were grouped into complex exposure scenarios. Multiple identified exceedance combinations were categorized based on the dominant pollutant type and risk level. Combinations with the same dominant pollutant and similar risk levels were grouped into the same category of complex exposure scenarios. Characteristic attributes of complex exposure scenarios were defined, including scenario name, dominant pollutant, secondary pollutant, risk level, typical concentration range, spatiotemporal conditions of susceptibility, and risk amplification coefficient. The risk amplification coefficient reflects the risk amplification factor of this scenario relative to exposure to a single pollutant, determined based on epidemiological evidence of synergistic effects of multiple pollutants. For example, in a complex particulate matter exposure scenario during the winter heating season, the dominant pollutants are PM2.5 and PM10, the secondary pollutant is SO2, the risk level is moderate to severe, the PM2.5 concentration range is 100-300 μg / m³, and it is prone to occur in northern heating areas from November to February of the following year, with a risk amplification coefficient of 1.3-1.8. In summer, the combined exposure scenario of photochemical pollution is characterized by dominant pollutants O3 and NO2, with VOCs as a secondary pollutant. The risk level is mild to moderate, with O3 concentrations ranging from 150 to 250 μg / m³. This scenario is more likely to occur during hot, sunny weather from May to September, with a risk enhancement coefficient of 1.2 to 1.5. In the combined exposure scenario of traffic pollution, the dominant pollutants are NO2 and PM2.5, with CO as a secondary pollutant. The risk level is moderate, with a risk enhancement coefficient of 1.4 to 1.7. In the combined exposure scenario of industrial emissions, the dominant pollutants are SO2 and PM10, with heavy metal particulate matter as a secondary pollutant. The risk level is severe, with a risk enhancement coefficient of 1.5 to 2.0. A combined exposure scenario library is constructed, containing all identified scenario types and their detailed characteristic descriptions.

[0043] A risk assessment matrix is ​​formed by spatially correcting the preliminary risk quantification indicators according to the combined exposure scenarios. The combined exposure scenario library contains risk enhancement coefficients for each scenario, which are used to correct the baseline risk value. The study area is divided into several spatial grid units, with grid size determined based on monitoring station density and population distribution; typical grid sizes are 1km×1km or 2km×2km. The combined exposure scenario type for each grid unit in the current time period is identified, based on the matching degree between the pollutant concentration combination of that unit and the scenario characteristics in the scenario library. The preliminary risk quantification indicators from step S120 are read to obtain the preliminary risk value for each grid unit. The preliminary risk value of each grid unit is corrected using the corresponding scenario's risk enhancement coefficient, with the correction formula being R_corrected = R_initial × K_scenario, where R_corrected is the corrected risk value, R_initial is the preliminary risk value, and K_scenario is the risk enhancement coefficient of the scenario to which the grid belongs. The corrected risk value considers the synergistic enhancement effect of multiple pollutant combined exposures. The corrected risk value is spatially smoothed to eliminate discontinuities at grid boundaries. Arrange the corrected risk values ​​of each grid cell according to their row and column positions to construct a two-dimensional risk assessment matrix. The rows of the matrix correspond to the latitude direction of the grid, the columns correspond to the longitude direction of the grid, and the matrix element values ​​are the corrected risk values ​​of the corresponding grid cells.

[0044] Step S140: Parallel clustering is performed on the risk assessment matrix to generate risk classification labels. Factor source analysis is performed on the risk classification labels to determine the contribution table of the dominant pollution sources. Key impact paths are extracted from the contribution table of the dominant pollution sources to reconstruct the exposure response chain. Based on the exposure response chain, regional difference characteristics are extracted to establish a health effect factor set.

[0045] Specifically, parallel clustering is performed on the risk assessment matrix to generate risk level labels. The corrected risk values ​​of each grid cell in the risk assessment matrix are read and used as input features for cluster analysis. Since the matrix contains tens of thousands of grid cells, a GPU-accelerated parallel clustering algorithm is used to improve computational efficiency. A GPU-optimized version of the K-means algorithm is chosen, which distributes distance calculation and cluster center update processes across multiple GPU cores for parallel execution. The number of clusters, K, is set to 5, corresponding to 5 risk levels: very low risk, low risk, medium risk, high risk, and very high risk. Five cluster centers are initialized, with the center positions selected from the quintiles of the risk values. For all grid cells in the matrix, the Euclidean distance between each cell and the 5 cluster centers is calculated in parallel using the GPU, and each cell is assigned to the category of the nearest cluster center. Each cluster center is updated to the mean of the risk values ​​of all grid cells within that category. The distance calculation and center update process is iteratively executed until the cluster centers no longer change significantly or the maximum number of iterations (100) is reached. After clustering convergence, each grid cell is assigned to its corresponding risk level category. A label is assigned to each risk level: 1 for extremely low risk, 2 for low risk, 3 for medium risk, 4 for high risk, and 5 for extremely high risk. The risk level labels for each grid cell are then labeled according to their spatial position in the matrix, forming a risk grading label. This risk grading label displays the spatial distribution of risk levels within the study area in a grid format, with different levels represented by different label values.

[0046] Factor source apportionment was performed on risk classification markers to determine the dominant pollution source contribution table. Risk classification markers were screened to obtain grid cells with marker values ​​of 4 and 5, corresponding to high-risk and very high-risk areas. Pollutant concentration and chemical composition data for the grid cells in high-risk areas were obtained, including inorganic ions, carbon components, and trace elements in particulate matter. Pollution source apportionment was performed on high-risk areas, using a positive definite matrix factorization model to identify the main source types of pollutants. The source apportionment model decomposes the pollutant concentration matrix into the product of a pollution source contribution matrix and a source composition spectrum matrix, solving for both matrices through iterative optimization. The number of pollution source categories was set to 6-8, with common pollution source types including coal combustion, vehicle exhaust, industrial emissions, dust, biomass combustion, and secondary emissions. After model convergence, the pollution source contribution matrix contained the contribution rate of each pollution source to the risk value of each grid cell. The contribution rate is calculated using the formula CR_ik = (C_ik / ΣC_ik) × 100%, where CR_ik is the contribution rate (%) of the k-th pollution source to the i-th grid cell, C_ik is the pollutant mass concentration contributed by the k-th pollution source to the i-th grid cell, and ΣC_ik is the total mass concentration contributed by all pollution sources to the i-th grid cell. The pollution source type with the highest contribution rate is identified for each grid cell and marked as the dominant pollution source for that cell. The dominance frequency and average contribution rate of each pollution source type in high-risk areas are statistically analyzed; pollution sources with high frequency and large contribution rates are identified as key pollution sources in the area. The dominant pollution source types and contribution rates of each grid cell are compiled into a dominant pollution source contribution table, which contains the contribution rate data for each pollution source.

[0047] In some embodiments, the step of extracting key impact paths and reconstructing the exposure response chain from the dominant pollution source contribution table includes: constructing a source contribution ranking map by performing source contribution ranking analysis through the dominant pollution source contribution table; obtaining a set of transmission paths from pollution sources to health effects by performing path tracing detection on the source contribution ranking map; obtaining key impact path identifiers by performing path integrity identification based on the transmission path set; and reconstructing the exposure response chain according to the transmission order of the key impact path identifiers.

[0048] A source contribution ranking map was constructed using a dominant pollution source contribution table. The dominant pollution source contribution table contains contribution rate data for each pollution source, and the sources are ranked in descending order of contribution rate. The frequency and contribution rate distribution characteristics of each pollution source type in high-risk areas were statistically analyzed. Frequency reflects the prevalence of the pollution source, and contribution rate reflects the intensity of its impact. A weighted importance score was calculated for each pollution source, taking into account both the mean contribution rate and frequency of occurrence. The formula is S_source = CR_mean × f^0.5, where S_source is the importance score of the pollution source, CR_mean is the average contribution rate of the pollution source across all grid cells, f is the percentage of grid cells in which the pollution source is the dominant pollution source, and the exponent 0.5 indicates that the frequency is weighted by the square root to avoid overemphasizing prevalence. Pollution sources were ranked according to their importance scores, with higher-scoring sources ranked higher. The pollution sources are arranged sequentially on the horizontal axis, and their contribution rates are marked on the vertical axis. A bar chart of the pollution source contribution rates is then drawn to form a source contribution ranking chart, which intuitively shows the relative importance of each pollution source to health risks.

[0049] Path tracing detection is performed on the source contribution ranking map to obtain the set of transmission paths from pollution sources to health effects. Based on the source contribution ranking map, the top-ranked major pollution sources are selected as the starting points for path tracing. For each major pollution source, the transmission and diffusion process of its emitted pollutants in the environment, the human exposure process, and the occurrence of health effects are analyzed to construct a complete transmission chain from pollution source to health effect. The transmission chain includes five key links: pollution source emission, atmospheric transmission, human exposure, dose response, and health effect. Each link is labeled with nodes and connections; nodes represent key states in the transmission process, and connections represent transition paths between states. In the atmospheric transmission link, the diffusion direction, transmission distance, and concentration decay characteristics of pollutants are labeled. In the human exposure link, the distribution of exposed population, exposure duration, and exposure concentration are labeled. In the dose response link, dose-response relationship parameters and health risk thresholds are labeled. In the health effect link, disease type, number of cases, and relative risk value are labeled. Path coding is performed on the transmission chain of each pollution source, and the path coding includes information such as the starting pollution source type, transmission characteristics, exposure characteristics, and effect characteristics. The transmission paths of all major pollution sources are summarized to form a transmission path set, which contains multiple transmission paths originating from different pollution sources and reaching health effects.

[0050] For example, the step of performing path integrity identification based on the transmission path set to obtain key impact path identifiers includes: performing causal chain continuity checks based on the transmission path set to obtain a chain break location map; performing missing link completion reasoning on the chain break location map to generate a completed transmission path set; extracting path integrity indicators from the completed transmission path set to generate a path integrity score vector; and filtering the path integrity score vector according to a integrity threshold to obtain key impact path identifiers.

[0051] A causal chain continuity test is performed based on the transmission path set to obtain a chain break location map. For each path in the transmission path set, it is checked whether all five key links exist and are interconnected. The causal relationship between each link is tested for validity, verified by correlation analysis and temporal sequence. When a link is missing or the causal relationship between two links is invalid, it is marked as a chain break location. The break locations and causes of each path are statistically analyzed, including missing data, insignificant causal relationships, and unreasonable temporal logic. A spatial distribution map of chain break locations is drawn, marking the links and grid positions where the break locations are located. The distribution patterns of break locations are analyzed to identify high-frequency break links and high-frequency break regions. High-frequency break links are usually those with weak data monitoring or insufficient understanding of the mechanism, while high-frequency break regions are usually areas with sparse monitoring stations or incomplete population statistics. The break location information of each path is compiled into a chain break location map, which identifies the break links of each path.

[0052] The missing link completion reasoning is performed on the chain break location map to generate a complete transmission path set. The chain break location map is identified to determine the completeable breaks, which are those with missing data but obtainable through reasoning or interpolation. For breaks in atmospheric transport, atmospheric diffusion models are used to calculate pollutant transport trajectories and concentration fields. For breaks in population exposure, population density data and activity pattern data are used to calculate the exposed population and exposure dose. For breaks in dose-response, dose-response relationship parameters from similar pollutants or populations are used as substitutes. For breaks in health effects, disease burden models are used to calculate the number of cases and health losses. The completed link data is then filled into the corresponding paths, updating the path integrity status. The causal chain continuity is re-checked on the completed paths to ensure the causal relationship between the completed link and adjacent links is valid. All completed paths are summarized to form a complete transmission path set, improving the integrity of each path in the complete path set.

[0053] A path completeness score vector is generated by extracting path completeness indicators from the completed path set. For each path in the completed path set, the completeness of its five key links is statistically analyzed, and the path completeness is obtained by dividing the number of complete links by the total number of links. The path completeness calculation formula is I_path = N_complete / N_total, where I_path is the path completeness, N_complete is the number of complete links in the path, and N_total is the total number of links in the path. A completeness score of 1 indicates that all five links of the path are complete and interconnected, while a completeness score less than 1 indicates that the path has broken or incomplete links. A credibility weight is assigned to each incomplete link: the weight of the original data link is 1.0, the weight of the inference incomplete link is 0.8, and the weight of the interpolation incomplete link is 0.6. The weighted completeness of each path is calculated, taking into account the credibility differences of the incomplete links. The weighted completeness is calculated as the sum of the weights of each link in the path divided by the total number of links; the closer the value is to 1, the higher the path quality. Obtain the weighted completeness scores for each path, arrange them in order of path number, and form a path completeness score vector. The dimension of the score vector is equal to the number of paths, and each element in the vector corresponds to the completeness score of one path.

[0054] Key impact pathway identifiers are obtained by filtering the pathway integrity score vectors according to integrity thresholds. A selection threshold for pathway integrity is set, determined based on the required research precision: a threshold of 0.9 or higher is required for high-precision studies, and 0.8 or higher for medium-precision studies. Pathways exceeding the integrity score threshold are selected, indicating higher data integrity and causal reliability. The importance of these high-integrity pathways is assessed, considering risk intensity, affected population, and spatial coverage. Risk intensity is quantified by the health effect at the pathway's endpoint; affected population is quantified by demographic data of the area covered by the pathway; and spatial coverage is quantified by the number of grid cells traversed by the pathway. A comprehensive importance score is calculated for each pathway; pathways with high scores have a critical impact on regional health risk. The top 20% of pathways by comprehensive importance score are selected and marked as key impact pathways. Each key impact pathway is assigned a unique identifier code containing information such as pollution source type and health effect type, forming a key impact pathway identifier.

[0055] The exposure response chain is reconstructed based on the transmission sequence of key impact pathway identifiers. Key impact pathway identifiers contain coded information for each pathway, analyzing the starting pollution source, transmission sequence, and endpoint health effect. Pathways are sequentially arranged and connected according to the chronological order of transmission, constructing a complete chain from pollution source emission to the occurrence of health effects. Multiple pathways originating from the same pollution source are merged to form a comprehensive exposure response chain for that pollution source. Exposure response chains from different pollution sources are ranked according to the importance of the pollution source, with chains from more important sources listed first. Key parameters and conversion coefficients for each stage are labeled on the chain; diffusion distance and concentration decay rate are labeled for transmission stages; the exposed population and exposure dose are labeled for exposure stages; and relative risk values ​​and the number of cases are labeled for response stages. The chain is visualized using a flowchart to show the causal and quantitative relationships between the pollution source, transmission stages, and health effects. The width of the chain represents the impact intensity, and the color represents the risk level. All key impact pathway chains are integrated into a single map to form a complete exposure response chain. The exposure response chain systematically demonstrates the complete pathway and mechanism by which major pollution sources ultimately lead to health effects through environmental transport and human exposure processes.

[0056] In some embodiments, the step of extracting regional difference features based on the exposure response chain to establish a set of health effect factors includes: performing spatial distribution gradient analysis on the exposure response chain to identify high-risk regional clusters; identifying risk attenuation regions at the boundaries of the high-risk regional clusters to generate attenuation markers; tracing low-risk transmission features step by step along the spatial gradient direction based on the attenuation markers to obtain regional difference candidates; and performing spatial correlation matching between the regional difference candidates and the high-risk regional clusters to establish a set of health effect factors.

[0057] Spatial distribution gradient analysis of exposure response chains is used to identify high-risk clusters. An exposure response chain encompasses the spatial coverage of each path, represented by the set of grid cells traversed by the path. The number of paths traversed and the sum of the risk intensities of each path are counted for each grid cell; cells with a large number of paths and high risk intensities are considered high-risk cells. Criteria for identifying high-risk cells are established: a cell is considered high-risk when the sum of the path risk intensities exceeds a threshold or the number of paths exceeds three. Connectivity analysis is performed on spatially adjacent high-risk cells, and a connected component labeling algorithm is used to identify connected high-risk areas. Connected high-risk areas form clusters, with cells within each cluster exhibiting similar high-risk characteristics and being spatially closely connected. The area, perimeter, and shape parameters of each cluster are calculated; the area reflects the spatial scale of the cluster, and the shape parameter reflects its spatial morphology. The distribution of pollution source types and the characteristics of the dominant exposure response chains within each cluster are analyzed to identify the risk sources and transmission mechanisms of the clusters. The population covered by each cluster and the number of disease cases are statistically analyzed to assess the health impact intensity of the clusters. The identified high-risk connected areas are marked as high-risk clusters, which are spatial areas where health risks are most concentrated and significant.

[0058] In high-risk area clustering zones, risk attenuation regions are identified at the boundary, generating attenuation markers. The spatial boundary of the high-risk area clustering zone is obtained, formed by the outline of grid cells surrounding the clustering zone. Risk level analysis is performed on adjacent grid cells outside the boundary, calculating the difference between the risk value of adjacent cells and the risk value inside the clustering zone. When the risk value of an adjacent cell is significantly lower than that inside the clustering zone, that adjacent cell is considered to be in a risk attenuation region. A risk attenuation criterion is set: when the risk value of an adjacent cell is lower than 70% of the average risk value of the clustering zone, it is considered an attenuation region. Extending outwards along the clustering zone boundary, attenuation regions with gradually decreasing risk values ​​are identified layer by layer, forming a buffer zone around the clustering zone. The risk attenuation rate of each attenuation region cell is calculated; the attenuation rate is the ratio of the cell's risk value to the average risk value of the clustering zone. The spatial gradient direction of risk attenuation is analyzed, determined by the direction of the fastest change in attenuation rate. The spatial location and attenuation gradient direction of the risk attenuation regions are marked on the map, using color intensity to represent the degree of attenuation. Attenuation markers containing spatial gradient direction information are assigned to each attenuation region cell.

[0059] Regional difference candidates are obtained by tracing low-risk transmission characteristics step by step along the spatial gradient direction using attenuation markers. The attenuation markers contain spatial gradient direction information, tracing outwards layer by layer from the high-risk cluster zone. During the tracing process, the risk level, pollution source type, exposure characteristics, and population characteristics of each grid unit are recorded. The differences between each grid unit and the high-risk cluster zone are analyzed, including differences in risk value, pollution source composition, exposure dose, and population susceptibility. When the differences in a certain grid unit exceed a set threshold, that layer is marked as a regionally significant difference layer. A regionally significant difference layer represents a spatial location where risk characteristics undergo a significant shift; this location is a transition zone between high-risk and low-risk areas. The spatial location, difference characteristic type, and difference intensity of the regionally significant difference layer are obtained. Simultaneously, environmental and socioeconomic characteristic factors of this layer are recorded, including pollution source type, meteorological conditions, topographic features, population density, and socioeconomic level. The tracing results of multiple cluster zones are summarized to identify all regionally significant difference locations within the study area. These significant difference locations and their characteristic factors are compiled into regional difference candidates, which include difference location and environmental and socioeconomic characteristic factors.

[0060] A set of health effect factors is established by spatially matching candidate regional differences with high-risk regional clusters. The spatial correlation between each candidate regional difference location and the high-risk regional cluster is calculated. Spatial correlation is comprehensively quantified by proximity, risk gradient continuity, and transmission path similarity. Proximity reflects the spatial distance between the difference location and the cluster; closer proximity indicates higher correlation. Risk gradient continuity reflects the smoothness and continuity of risk changes from the cluster to the difference location; gradient continuity indicates high correlation. Transmission path similarity reflects the similarity between the exposure response chain of the difference location and the exposure response chain of the cluster; path similarity indicates high correlation. A comprehensive spatial correlation score is calculated for each difference location; difference locations with high scores have a close spatial-risk correlation with the cluster. Difference locations with correlation scores exceeding a threshold are selected, as the difference characteristics of these locations have a regulatory or influencing effect on the health effects of the cluster. Characteristic factors of highly correlated difference locations are obtained, including pollution source type, meteorological conditions, topographic features, population density, and socioeconomic level. These characteristic factors are correlated with the health effect data of the high-risk regional cluster to identify key factors with significant impact on health effects. The identified key characteristic factors were organized into a health effect factor set, which includes information such as factor name, factor type, influence intensity and mechanism of action, revealing how regional differences affect the spatial distribution and intensity differences of health risks.

[0061] Step S150: Based on the risk assessment matrix and the set of health effect factors, determine the content priority by population stratification, and use the content priority to output a stratified health assessment report.

[0062] In some embodiments, determining content priority based on population stratification according to the risk assessment matrix and the health effect factor set includes: performing multi-source data consistency cross-validation on the risk assessment matrix and the health effect factor set to generate an evidence weight table; performing weighted fusion of population risks based on the evidence weight table to obtain corrected population risk values; dynamically allocating the corrected population risk values ​​according to risk levels using content length coefficients to generate a population content allocation table; and performing priority sorting on the population content allocation table to determine content priority.

[0063] A cross-validation of the risk assessment matrix and the health effect factor set was performed using multi-source data to generate an evidence weight table. The risk assessment matrix contains gridded risk value data calculated based on a multi-factor collaborative interaction model. The health effect factor set contains regional difference characteristic factor data obtained based on spatial analysis of the exposure-response chain. The two data sources were spatially aligned, with grid cells as the basic unit. For each grid cell, the consistency between the risk value provided by the matrix and the risk level derived from the factor set was compared. A rank consistency test was used for consistency assessment; if the risk levels determined by the two data sources were the same or differed by no more than one level, they were considered consistent. The consistency status of each grid cell was statistically analyzed; consistent cells were marked as high confidence, and inconsistent cells as low confidence. For inconsistent cells, the reasons for the differences between the two data sources were analyzed, including differences in data coverage, analytical methods, and time scale mismatches. Evidence weights were assigned to the two data sources based on their confidence and coverage completeness. The risk assessment matrix, based on a large amount of monitoring data and statistical models, has a wide data coverage and was assigned a weight of 0.6. The health effect factor set, based on mechanistic analysis and causal reasoning, has a strong ability to identify local features and is assigned a weight of 0.4. The weights of the two data sources are organized according to grid units and population levels to form an evidence weight table, with the sum of the weights equal to 1.

[0064] The corrected population risk value is obtained by weighted fusion of population risks based on an evidence weight table. The evidence weight table contains the weights of the two data sources corresponding to each grid cell and the population level. The risk assessment matrix contains the risk value of the population level corresponding to each grid cell, denoted as R_matrix. The risk assessment derived from the health effect factor set contains the corresponding risk value, denoted as R_factor. The two risk values ​​of the same grid cell and population level are weighted and fused using the fusion formula R_fused = w_matrix × R_matrix + w_factor × R_factor, where R_fused is the fused risk value, w_matrix is ​​the evidence weight of the risk assessment matrix, R_matrix is ​​the risk value provided by the matrix, w_factor is the evidence weight of the health effect factor set, and R_factor is the risk value derived from the factor set. The weight constraint is w_matrix + w_factor = 1. The fused risk value integrates information from both data sources, reducing the bias and uncertainty of a single data source. Quality control is performed on the fused risk value to identify outliers and extreme values. When the difference between the merged risk value and both original risk values ​​exceeds 30%, it is marked as an outlier, and the merged value is replaced by a simple average of the two original risk values. The average merged risk value for each population level within the study area is calculated. This average is obtained by weighting the merged risk values ​​of all activity grid units within that population level, with the weight being the population size of each unit. The average merged risk value for each population level is recorded as the corrected population risk value, which is a quantitative indicator of the overall health risk level faced by that population level.

[0065] The revised population risk values ​​are dynamically allocated according to risk levels, generating a population content allocation table. The revised population risk values ​​include risk values ​​for each population level, and populations are classified into risk levels based on these values. Thresholds for risk level classification are set: risk values ​​less than 1.0 are low-risk, 1.0-2.0 are medium-risk, 2.0-3.0 are high-risk, and greater than 3.0 are extremely high-risk. The number of population levels and population size included in each risk level are statistically analyzed, with population size obtained by aggregating demographic data for the corresponding regions. A content length coefficient is assigned to each risk level in the health assessment report, reflecting the proportion of content appropriate for that level in the report. The allocation of length coefficients follows the principle that higher risk levels require more content, while also considering the impact of population size. The content length coefficient is calculated using the formula C_level = (R_level / ΣR_level) × α + (P_level / ΣP_level) × (1-α), where C_level is the content length coefficient for a certain risk level, R_level is the average risk value for that level, ΣR_level is the sum of the average risk values ​​for all levels, P_level is the population size for that level, ΣP_level is the total population for all levels, α is the risk weight factor with a value of 0.7, and (1-α) is the population weight factor with a value of 0.3. The content length coefficient considers both the level of risk and the size of the affected population; α = 0.7 indicates a greater emphasis on high-risk levels. Specific content length values ​​are assigned to each population level, with the length value equal to the content length coefficient of that level multiplied by the total report length. The corrected risk values, population size, content length coefficients, and content length values ​​for each population level are compiled into a population content allocation table.

[0066] Content priority is determined by prioritizing the content allocation table for each population group. The table includes the content length coefficient, adjusted risk value, and population size for each group level, with the length coefficient serving as the primary basis for priority ranking. The population levels in the allocation table are sorted in descending order of length coefficient, with the level having the largest length coefficient ranking first and having the highest priority. When the length coefficients of two population levels are close, the adjusted risk value in the allocation table is used as a secondary ranking criterion, with the level having a higher risk value having a higher priority. When both the length coefficient and risk value are close, the population size in the allocation table is used as a tertiary ranking criterion, with the level having a larger population having a higher priority. Priority numbers are assigned to the sorted population levels, starting from 1 and increasing sequentially, with smaller numbers indicating higher priority. These priority numbers are then linked to the population level information in the allocation table to form the content priority ranking result for each population level.

[0067] A stratified health assessment report is generated based on content priority. The report content is organized in descending order of priority. The report adopts a hierarchical structure, with independent chapters or sections for each population level. A population content allocation table includes the content length values ​​for each population level, determining the specific page count and level of detail for each chapter. For the highest priority population level, a dedicated chapter provides a detailed explanation, including risk source analysis, exposure route identification, health effect assessment, and protective recommendations. For medium-priority population levels, sections provide moderate levels of detail. For low-priority population levels, simplified or combined descriptions are used to summarize the risk status and general recommendations. An overall overview chapter is included at the beginning of the report, outlining the overall health risk status of the study area, major pollution sources, and key impact pathways. Spatial distribution maps and data charts obtained from the aforementioned analysis are referenced in each population chapter. Visualization charts are included in the report, including risk distribution maps, population risk comparison bar charts, pollution source contribution pie charts, and exposure response chain flowcharts. The report concludes with sections on protective measures and policy recommendations, proposing differentiated health protection measures and pollution control strategies for different population groups and risk areas. The report content undergoes a quality review to ensure accurate data citations, rigorous analytical logic, and feasible recommendations. The approved report content is then formatted according to standard formats to generate the final stratified health assessment report.

[0068] To implement the above-described method embodiments, a dynamic report generation method based on environmental health data fusion is proposed to achieve the corresponding functionalities and technical effects. See also... Figure 2 , Figure 2 This diagram illustrates a structural block diagram of a dynamic report generation system 200 based on environmental health data fusion, as provided in an embodiment of this application. For ease of explanation, only the parts relevant to this embodiment are shown. The dynamic report generation system 200 based on environmental health data fusion provided in this embodiment includes: Data acquisition module 201 is used to acquire pollutant concentration data and disease occurrence data, and to perform spatiotemporal alignment on the pollutant concentration data and the disease occurrence data to generate a health-related dataset; Exposure analysis module 202 is used to construct a population exposure transmission network based on the health association dataset, perform distributional lag association analysis on the population exposure transmission network to identify susceptible population clusters, extract differential response features from the susceptible population clusters to construct a hierarchical exposure map, and use the hierarchical exposure map to conduct multi-factor collaborative interaction tests to obtain preliminary risk quantification indicators; Risk assessment module 203 is used to conduct time-series pattern mining to establish pollutant cumulative effect relationships in the health-related dataset, determine health impact weights for different regions based on the pollutant cumulative effect relationships, identify compound exposure scenarios using the health impact weights, and perform regional spatial corrections on the preliminary risk quantification indicators according to the compound exposure scenarios to form a risk assessment matrix; The factor reconstruction module 204 is used to perform parallel clustering processing on the risk assessment matrix to generate risk classification identifiers, perform factor source analysis on the risk classification identifiers to determine the contribution table of dominant pollution sources, extract key impact paths from the contribution table of dominant pollution sources to reconstruct the exposure response chain, and extract regional difference features based on the exposure response chain to establish a health effect factor set; The report generation module 205 is used to determine the content priority by population stratification based on the risk assessment matrix and the health effect factor set, and to output a stratified health assessment report using the content priority.

[0069] The aforementioned dynamic report generation system 200 based on environmental health data fusion can implement the dynamic report generation method based on environmental health data fusion described in the above method embodiments. The options in the above method embodiments are also applicable to this embodiment and will not be detailed here. The remaining content of this application embodiment can be referred to the content of the above method embodiments, and will not be repeated in this embodiment.

[0070] The purpose of the above embodiments is to reproduce and derive the technical solution of the present invention by way of example, and to fully describe the technical solution, purpose and effect of the present invention. The purpose is to enable the public to have a more thorough and comprehensive understanding of the disclosure of the present invention, and not to limit the scope of protection of the present invention.

[0071] The above embodiments are not an exhaustive list based on the present invention, and there may be many other embodiments not listed. Any substitutions and improvements made without departing from the concept of the present invention are within the protection scope of the present invention.

Claims

1. A method for generating dynamic reports based on environmental health data fusion, characterized in that, include: Acquire pollutant concentration data and disease occurrence data, and perform spatiotemporal alignment on the pollutant concentration data and disease occurrence data to generate a health-related dataset; Based on the aforementioned health-related dataset, a population exposure transmission network is constructed. Distributional lag association analysis is performed on this network to identify susceptible population clusters. Differential response features are extracted from these clusters to construct a hierarchical exposure map. The hierarchical exposure map is then used to conduct multi-factor collaborative interaction tests to obtain preliminary risk quantification indicators. Time-series pattern mining is performed on the aforementioned health-related dataset to establish pollutant cumulative effect relationships. Based on these relationships, health impact weights for different regions are determined. These weights are then used to identify compound exposure scenarios. Finally, the preliminary risk quantification indicators are spatially corrected by region according to the compound exposure scenarios to form a risk assessment matrix. Parallel clustering is performed on the risk assessment matrix to generate risk classification identifiers. Factor source analysis is then performed on the risk classification identifiers to determine the contribution table of dominant pollution sources. Key impact pathways are extracted from the contribution table of dominant pollution sources to reconstruct the exposure-response chain. Based on the exposure-response chain, regional difference characteristics are extracted to establish a set of health effect factors. Based on the risk assessment matrix and the set of health effect factors, content priorities are determined by population stratification, and stratified health assessment reports are output using the content priorities.

2. The method according to claim 1, characterized in that, The step of performing distributional lag association analysis on the population exposure transmission network to identify susceptible population clusters includes: The population exposure transmission network is decomposed into acute exposure pathways and chronic exposure pathways according to the exposure duration dimension; Multiscale hysteresis response curves were constructed for the acute exposure pathway and the chronic exposure pathway to obtain spatiotemporal interactive hysteresis response curves; Based on the spatiotemporal interaction hysteresis response curve, a stratified susceptibility intensity identification of the population is performed to generate a stratified susceptibility feature table. The high-susceptibility region is extracted from the hierarchical susceptibility feature table, and high-susceptibility clusters are selected as the susceptible population cluster areas.

3. The method according to claim 1, characterized in that, The preliminary risk quantification indicators obtained by conducting multi-factor collaborative interaction tests using the stratified exposure map include: A multi-factor interaction matrix was constructed by extracting pollutant concentration variables and meteorological factor variables from the stratified exposure map. Interaction term regression analysis was performed on the multi-factor interaction matrix to obtain the interaction effect coefficient table; The interaction effect coefficient table is used to screen for significance levels to identify combinations of synergistic enhancement factors; The synergistic enhancement factors are combined and weighted according to their interaction strength to form a preliminary risk quantification index.

4. The method according to claim 1, characterized in that, The method of identifying compound exposure scenarios using the health impact weights includes: Extract the pollutant weight vector from the health impact weights; The pollutant weight vector is standardized by concentration and then weighted and summed to generate a comprehensive exposure score. Based on the comprehensive exposure score, identify the periods of multi-pollutant exceedance and extract the combination patterns of exceedance; The aforementioned combinations of exceeding the standard are summarized into a composite exposure scenario.

5. The method according to claim 1, characterized in that, The step of extracting key impact pathways and reconstructing the exposure-response chain from the dominant pollution source contribution table includes: A source contribution ranking diagram is constructed by performing source contribution ranking analysis using the aforementioned dominant pollution source contribution table. Path tracing detection is performed on the source contribution ranking graph to obtain the set of transmission paths from pollution sources to health effects; Based on the aforementioned transmission path set, perform path integrity identification to obtain key impact path identifiers; The exposure response chain is reconstructed based on the transmission order of the key impact path identifiers.

6. The method according to claim 1, characterized in that, The step of establishing a set of health effect factors based on regional difference features extracted from the exposure response chain includes: Spatial distribution gradient analysis was performed on the exposure response chain to identify high-risk area clusters; Risk attenuation zones are identified at the boundaries of the high-risk cluster zones, and attenuation markers are generated. Based on the attenuation marker, low-risk transmission characteristics are traced step by step along the spatial gradient direction to obtain regional difference candidates; A set of health effect factors is established by spatially matching the regional difference candidates with the high-risk regional clusters.

7. The method according to claim 1, characterized in that, The method of determining content priority by population stratification based on the risk assessment matrix and the set of health effect factors includes: A weighted evidence table is generated by performing multi-source data consistency cross-validation on the risk assessment matrix and the health effect factor set. Based on the aforementioned evidence weighting table, the population risk is weighted and fused to obtain the corrected population risk value; The revised population risk values ​​are dynamically allocated according to risk levels using content length coefficients to generate a population content allocation table. The content allocation table for the aforementioned population is sorted by priority to determine the content priority.

8. The method according to claim 3, characterized in that, The step of performing interaction term regression analysis on the multi-factor interaction matrix to obtain the interaction effect coefficient table includes: Extract the factor variables from the multi-factor interaction matrix to construct interaction term feature sets; Perform stepwise regression filtering on the feature set of the interaction items to obtain a significant subset of interaction items; Effect coefficient vectors are generated by estimating effect coefficients using the subset of significant interaction terms; The effect coefficient vectors are organized into an interaction effect coefficient table according to the factor combination type.

9. The method according to claim 5, characterized in that, The step of performing path integrity identification based on the transmission path set to obtain key impact path identifiers includes: Based on the transmission path set, a causal chain continuity test is performed to obtain a chain break location map; Perform missing link completion reasoning on the chain break location map to generate a complete transmission path set; The path completeness index is extracted from the completed transmission path set to generate a path completeness score vector; The path integrity score vector is filtered according to the integrity threshold to obtain the key impact path identifiers.

10. A dynamic report generation system based on environmental health data fusion, characterized in that, include: The data acquisition module is used to acquire pollutant concentration data and disease occurrence data, and to perform spatiotemporal alignment of the pollutant concentration data and the disease occurrence data to generate a health-related dataset; The exposure analysis module is used to construct a population exposure transmission network based on the health association dataset, perform distributional lag association analysis on the population exposure transmission network to identify susceptible population clusters, extract differential response features from the susceptible population clusters to construct a hierarchical exposure map, and use the hierarchical exposure map to conduct multi-factor collaborative interaction tests to obtain preliminary risk quantification indicators; The risk assessment module is used to perform time-series pattern mining on the health-related dataset to establish pollutant cumulative effect relationships, determine health impact weights for different regions based on the pollutant cumulative effect relationships, identify compound exposure scenarios using the health impact weights, and perform regional spatial corrections on the preliminary risk quantification indicators according to the compound exposure scenarios to form a risk assessment matrix; The factor reconstruction module is used to perform parallel clustering processing on the risk assessment matrix to generate risk classification identifiers, perform factor source analysis on the risk classification identifiers to determine the contribution table of dominant pollution sources, extract key impact paths from the contribution table of dominant pollution sources to reconstruct the exposure response chain, and extract regional difference features based on the exposure response chain to establish a health effect factor set; The report generation module is used to determine the content priority by population stratification based on the risk assessment matrix and the health effect factor set, and to output a stratified health assessment report using the content priority.