A building individual scale dengue fever prevention and control priority classification method
By combining interpretable machine learning models with individual building attributes, the problem of insufficient dengue fever risk identification at the individual building scale in existing technologies has been solved. This enables the classification of prevention and control priority levels at the individual building scale, improving the accuracy of dengue fever prevention and control and the efficiency of resource allocation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-02
- Publication Date
- 2026-07-10
AI Technical Summary
Existing technologies are insufficient to identify dengue fever risk at the individual building scale, lack effective methods for prioritizing prevention and control, fail to reflect the correlation between urban morphology and dengue fever risk, and lack differentiated characterization of individual buildings.
By employing an interpretable machine learning model combined with individual building attributes, and constructing building weights, the priority levels for prevention and control at the individual building scale are determined. This includes identifying key driving factors and predicting risks, and comprehensively considering meteorological, urban morphology, and socio-economic factors.
It improved the accuracy and spatial characterization of dengue fever risk prediction, enabled the classification of prevention and control priorities at the individual building scale, and enhanced the precision of prevention and control measures and the efficiency of resource allocation.
Smart Images

Figure CN122370000A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of infectious disease risk assessment and spatial analysis technology, and in particular to a method for classifying dengue fever prevention and control priority levels at the individual building scale. Background Technology
[0002] Dengue fever is a mosquito-borne infectious disease caused by the dengue virus. Its occurrence and spread are closely related to factors such as climate conditions, urban spatial environment, population activity characteristics, and mosquito breeding conditions. For urban areas at risk of dengue fever importation and transmission, identifying the key factors influencing dengue fever occurrence and spatially determining high-risk areas are important technical issues for carrying out refined prevention and control.
[0003] In existing technologies, research on dengue fever risk primarily relies on risk assessment models built around factors such as meteorological conditions, land use types, landscape characteristics, and population distribution. These models are then combined with statistical analysis or machine learning methods to generate grid-scale spatial distribution results of the risk. Such methods can reflect the overall distribution of dengue fever risk at the city scale and can explain the relationship between some influencing factors and disease risk using techniques such as feature importance analysis and marginal effect analysis.
[0004] However, existing technologies still have the following shortcomings: On the one hand, the urban internal spatial environment is highly heterogeneous, and different urban morphology types have significant differences in building height, building density, and the combination relationship between buildings and surrounding green and blue spaces. Existing dengue fever risk prediction methods do not adequately consider urban morphology factors, making it difficult to effectively reveal the relationship between urban morphology type and dengue fever risk. On the other hand, existing technologies usually remain at the grid or area scale for risk identification, lacking technical means to further distinguish buildings in high-risk areas by combining the individual socio-economic functional attributes of buildings, thus making it difficult to support the classification of prevention and control priority levels at the individual building scale. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a method for classifying dengue fever prevention and control priority levels at the individual building scale, which can reveal the correlation between urban morphology and dengue fever risk and achieve fine identification of potential high-risk areas, thereby improving the accuracy of dengue fever risk prediction and spatial characterization capabilities.
[0006] To achieve the above objectives, the present invention provides the following solution:
[0007] A method for prioritizing dengue fever control at the individual building scale, comprising:
[0008] Using grid cells as the basic spatial units, an analysis grid is constructed within the study area, and urban morphology type, meteorological factors, and socio-economic factors are calculated for each grid cell to obtain influencing factor data.
[0009] Based on the addresses and latitude and longitude values of local dengue fever cases in the city, the number of dengue fever cases in each grid unit is calculated, and it is determined whether there are dengue fever cases in each grid unit, so as to obtain the classification target variable used to characterize the dengue fever occurrence status of the grid unit.
[0010] An interpretable machine learning model was constructed based on influencing factor data and categorical target variables. The relative contribution of each influencing factor to the prediction of dengue fever risk was quantified based on the interpretable machine learning model, and the response relationship between key driving factors and dengue fever risk was identified.
[0011] Based on an interpretable machine learning model, the risk of dengue fever occurrence is predicted for each grid cell, a dengue fever risk map is drawn, and potential high-risk areas are identified.
[0012] Obtain individual attribute data of buildings in potentially high-risk areas, and construct building weights by comprehensively considering building volume, social function attributes, and building aging effects.
[0013] Within the grid cells corresponding to potentially high-risk areas, the risk of each grid cell is allocated based on building weights to achieve risk conservation, thereby obtaining the relative dengue fever risk of each individual building. Based on the relative dengue fever risk of each individual building, the priority level of building prevention and control is determined.
[0014] Preferably, urban morphology type, meteorological factors, and socioeconomic factors are calculated for each grid cell to obtain influencing factor data, including:
[0015] Calculate meteorological factors for each grid cell; meteorological factors include monthly average temperature, monthly average total rainfall, and monthly average relative humidity.
[0016] Calculate the urban morphology type for each grid cell; the urban morphology type includes the proportion of dense high-rise buildings, dense mid-rise buildings, dense low-rise buildings, open high-rise buildings, open mid-rise buildings, open low-rise buildings, large low-rise buildings, sparse buildings, heavy industrial land, dense forests, sparse trees, shrubs, low vegetation, rock or paved land, bare soil or sand, and water surface.
[0017] Calculate socioeconomic factors for each grid cell; socioeconomic factors include population density, GDP per capita, and road density.
[0018] Logarithmic transformation of population density and GDP per capita yields data on influencing factors.
[0019] Preferably, the classification target variables used to characterize the dengue fever occurrence state of the grid cell include:
[0020] Based on the addresses and latitude and longitude values of local dengue fever cases in the city, the number of dengue fever cases in each grid unit is counted.
[0021] Determine whether there are dengue fever cases in each grid unit;
[0022] Grid cells with dengue fever cases are marked as 1, and grid cells without dengue fever cases are marked as 0.
[0023] The result, which is recorded as 1 or 0, is used as the classification target variable.
[0024] Preferably, an interpretable machine learning model is constructed based on influencing factor data and categorical target variables, including:
[0025] The research area was divided into spatially isolated training and validation areas according to administrative boundaries;
[0026] Based on influencing factor data and classification target variables, a random forest binary classification model is used to model the risk of dengue fever occurrence;
[0027] The random forest binary classification model was trained and validated using the 10-fold cross-validation method.
[0028] The parameters of the random forest binary classification model are optimized using a grid search method, including the number of decision trees and the maximum depth of a single decision tree.
[0029] The parameter combination with the highest average AUC value during cross-validation was selected as the target parameter combination. The model training process was repeated multiple times based on the target parameter combination to obtain multiple random forest binary classification models, which can be used as interpretable machine learning models.
[0030] Preferably, the relative contribution of each influencing factor to dengue risk prediction is quantified based on an interpretable machine learning model, and the response relationship between key driving factors and dengue risk is identified, including:
[0031] The SHAP technique is used to interpret the model prediction results of the interpretable machine learning model and quantify the relative contribution of each influencing factor to the prediction of dengue fever risk.
[0032] Identifying key driving factors based on relative contributions;
[0033] PDP technology was used to analyze the response relationship between key driving factors and dengue fever risk prediction results.
[0034] Preferably, a dengue fever risk prediction is performed on each grid cell based on an interpretable machine learning model, a dengue fever risk map is drawn, and potentially high-risk areas are identified, including:
[0035] Based on multiple random forest binary classification models, the entire study area is traversed to predict the risk of dengue fever occurrence in each grid unit;
[0036] The predicted values output by multiple random forest binary classification models are averaged to form a dengue fever risk map;
[0037] The risk values in the dengue fever risk map were reclassified using the quartile method, and the risk levels were divided according to the three quartiles.
[0038] Grid cells with risk values above the third or fourth quartile are identified as potentially high-risk areas.
[0039] Preferably, a building weight is constructed by comprehensively considering building volume, social function attributes, and the aging effect of the building, including:
[0040] Obtain individual attribute data for buildings within potentially high-risk areas;
[0041] Calculate the building volume based on the building footprint and building height;
[0042] Within each grid cell, the building volume is truncated at the 5th to 95th percentile, and then standardized using the median of the building volume within the grid to obtain the building volume.
[0043] Residential building indicator variables are set based on the results of building function determination, serving as social function attributes;
[0044] The building age is calculated based on the year the building was built, and the building aging effect is constructed based on the logistic function;
[0045] Building weights are constructed based on building volume, social function attributes, and the effects of building aging.
[0046] Preferably, building weights are constructed based on building volume, social function attributes, and building aging effects, including:
[0047] Let V denote the building volume, and let Res represent the social functional attributes as an indicator variable for residential buildings. A value of 1 for Res indicates a residential building, and a value of 0 for Res indicates a non-residential building.
[0048] The building age (Age) is calculated based on the year the building was built, and the building age (Age) is substituted into the logistic function to construct the building age aging effect (AgeFactor). The building age aging effect (AgeFactor) satisfies: AgeFactor=0.85+0.40·σ((Age-25) / 8), where σ(x)=1 / (1+exp(-x)), σ(x) is the logistic function, x represents the normalized deviation of the building age (Age) from 25 years, 25 is the age inflection point, and 8 is the smoothing scale;
[0049] The building weights w are constructed according to the weight formula w=V^α×(ε+Res)^β×AgeFactor; where w represents the building weight, α represents the influence coefficient of building volume V, β represents the influence coefficient of social functional attribute Res, and ε represents the smoothing parameter used to ensure that non-residential buildings still retain non-zero weights.
[0050] The parameter α is set to 0.7, the parameter β is set to 2.0, and the parameter ε is set to 0.05.
[0051] Preferably, risk conservation allocation of grid cells based on building weights includes:
[0052] Obtain the grid risk value R of the grid cell corresponding to the potential high-risk area;
[0053] Based on the grid risk value R and the building weight, normalized risk conservation allocation is performed within the grid cell;
[0054] The risk value assigned to the building scale is used as the relative dengue risk Rb for the individual building.
[0055] Preferably, the priority level of building prevention and control is classified according to the relative dengue fever risk of individual buildings, including:
[0056] The dengue fever case location data was matched and statistically analyzed with the spatial location of buildings to obtain the building-level observed case count Yb;
[0057] All buildings within the study area were ranked from highest to lowest according to their relative dengue fever risk (Rb), and the priority level for building prevention and control was determined based on the ranking results.
[0058] The top k% of buildings in the sorted list are denoted as Top(k);
[0059] The Top-k capture rate is defined as the ratio of the total number of cases covered by the buildings ranked in the top k% to the total number of cases matched by all buildings in the study area.
[0060] The Top-k capture rate was used to assess the ability of building-based prevention and control priority levels to locate cases.
[0061] The present invention discloses the following beneficial effects:
[0062] This invention, in the dengue fever risk identification stage, incorporates urban morphology type, meteorological factors, and socioeconomic factors into an interpretable machine learning framework. This not only enables spatial prediction of dengue fever risk but also identifies key urban morphology types that play a major role in risk prediction, thereby revealing the correlation between urban morphology type and dengue fever risk. Compared to existing technologies that rely solely on macro-level assessments based on factors such as meteorology, land use, or population distribution, this invention more comprehensively reflects the impact of differences in the urban spatial environment on dengue fever transmission risk, improving the targeting of potential high-risk area identification and the precision of risk map representation.
[0063] Based on the identification of potential high-risk areas, this invention further integrates the age, volume, and social functional attributes of individual buildings to differentiate between buildings within high-risk areas, and accordingly establishes a prevention and control priority level at the building-by-building scale. Compared to existing technologies that typically only provide risk warnings at the grid or area scale, this invention can further implement risk identification results down to specific building objects, achieving refined positioning and priority division of prevention and control targets, thereby improving the accuracy, hierarchy, and resource allocation efficiency of dengue fever prevention and control measures. Attached Figure Description
[0064] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0065] Figure 1 A flowchart of the method provided in an embodiment of the present invention;
[0066] Figure 2 A flowchart of data collection and processing provided for embodiments of the present invention;
[0067] Figure 3 This is a schematic diagram illustrating machine learning modeling and model interpretation provided in an embodiment of the present invention;
[0068] Figure 4 This is a schematic diagram of risk assessment at the individual scale of a building, provided as an embodiment of the present invention. Detailed Implementation
[0069] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0070] The purpose of this invention is to provide a method for classifying dengue fever prevention and control priorities at the individual building scale. Based on the identification of potential high-risk areas, this method can further classify prevention and control priorities at the individual building scale, thereby improving the accuracy of dengue fever prevention and control measures and the rationality of resource allocation.
[0071] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0072] Figure 1 The method flowchart provided in the embodiments of the present invention is as follows: Figure 1 As shown, this invention provides a method for prioritizing dengue fever control at the individual building scale, including:
[0073] Step 100: Using grid cells as the basic spatial units, construct an analysis grid within the study area, and calculate urban morphology type, meteorological factors, and socio-economic factors for each grid cell to obtain influencing factor data;
[0074] Step 200: Based on the addresses and latitude and longitude values of local dengue fever cases in the city, calculate the number of dengue fever cases in each grid unit, and determine whether there are dengue fever cases in each grid unit, so as to obtain the classification target variable used to characterize the dengue fever occurrence status of the grid unit.
[0075] Step 300: Construct an interpretable machine learning model based on influencing factor data and categorical target variables, and quantify the relative contribution of each influencing factor to dengue risk prediction based on the interpretable machine learning model, and identify the response relationship between key driving factors and dengue risk.
[0076] Step 400: Based on an interpretable machine learning model, predict the risk of dengue fever occurrence in each grid cell, draw a dengue fever risk map, and identify potential high-risk areas;
[0077] Step 500: Obtain individual attribute data of buildings in potential high-risk areas, and construct building weights by comprehensively considering building volume, social function attributes, and building aging effects;
[0078] Step 600: Within the grid cells corresponding to the potential high-risk areas, risk conservation allocation is performed on the grid cells based on building weights to obtain the relative dengue fever risk of individual buildings, and the priority level of building prevention and control is determined according to the relative dengue fever risk of individual buildings.
[0079] like Figure 2 , Figure 3 and Figure 4 As shown, the technical process of this embodiment is as follows:
[0080] Step 1: Draw detailed spatial analysis units
[0081] Using a 1-kilometer grid cell as the basic spatial unit, an analysis grid of uniform spatial scale was constructed within the study area based on ArcMap 10.6 software.
[0082] Step 2: Calculate urban morphology, weather, and socioeconomic factors.
[0083] Big data on geospatial conditions was collected, and meteorological factors (including monthly average temperature, monthly average total rainfall, and monthly average relative humidity), urban morphology (including the proportion of dense high-rise buildings, dense mid-rise buildings, dense low-rise buildings, open high-rise buildings, open mid-rise buildings, open low-rise buildings, large low-rise buildings, sparse buildings, heavy industrial land, dense forests, sparse trees, shrubs, low vegetation, rocky or paved land, bare soil or sandy land, and water surface) and socioeconomic factors (including population density, GDP per capita, and road density) were calculated for each grid cell in the study area (as shown in Table 1). Logarithmic transformation preprocessing was performed on population density and GDP per capita.
[0084] Table 1. Urban morphology, weather, and socioeconomic factors
[0085] Step 3: Calculate the risk variables for predicting the response.
[0086] Based on the addresses and latitude and longitude values of local dengue fever cases within the city, the number of dengue fever cases in each grid is calculated, and it is determined whether there are dengue fever cases in each grid (marked as 1 if present, 0 if absent). This is used as the classification target variable to characterize the dengue fever occurrence status of the grid unit.
[0087] Step 4: Build an interpretable machine learning model
[0088] The study area was divided into two spatially isolated parts according to administrative boundaries, serving as the training and testing areas for the model, respectively. First, districts and counties that had recorded local cases were selected as the research samples. Then, the central urban area of Chongqing was designated as the training area, and three typical districts and counties in northern Chongqing (Zhongxian County, Wanzhou District, and Fengjie County), spatially isolated from it, were selected as independent validation areas. The central urban area contains 5717 grids, of which 187 grids had reported cases, totaling 750 local cases; the three northern districts and counties contain 9849 grids, of which 66 grids had reported cases, totaling 313 local cases.
[0089] Based on the constructed kilometer-scale sample data, a random forest binary classification model was used to model the risk of dengue fever occurrence. The model was trained and validated using 10-fold cross-validation, and the model parameters were optimized using a grid search approach. The adjusted parameters included the number of decision trees in the random forest (n_estimators) and the maximum depth of a single decision tree (max_depth). The area under the curve (AUC) was used as the model performance evaluation metric. The parameter search space was set as: "max_depth":[2,4,6,8,5,10,15,20], "n_estimators":[2,4,6,8,40,80,120,160]. The parameter combination with the highest average AUC value during cross-validation and its corresponding model were selected as the optimal model configuration. This process was repeated multiple times to reduce the impact of randomness, ultimately obtaining a stable random forest model. After the model training is completed, the obtained optimal model is applied to an independent region (such as a county outside the study area) for risk prediction, and the corresponding AUC value is calculated to evaluate the model's prediction performance and generalization ability under spatial migration conditions.
[0090] In this exemplary embodiment, one grid corresponds to one sample, and the case labels are statistically analyzed by year, all being local dengue fever cases in Chongqing that occurred between August and November 2019. Meteorological data were selected from the monthly temperature, relative humidity, and total precipitation data of August to November 2019, and averaged to align the meteorological data with the case data over time. The positive and negative sample categories were processed using a random forest classification model at ratios of 1:1, 1:2, and 1:3, respectively, with the parameter search space as described above.
[0091] Step 5: Explain the association between influencing factors and dengue fever risk
[0092] After obtaining a stable random forest model, the Shapley Additive Explanation (SHAP) method is used to interpret the model's prediction results, quantifying the relative contribution of each influencing factor to dengue risk prediction and identifying key driving factors. Simultaneously, Partial Dependency Plot (PDP) is used to analyze the response relationship between key factors and dengue risk prediction results, thereby revealing potential nonlinear patterns affecting dengue spread.
[0093] Step 6: Create a dengue fever risk map and identify potential high-risk areas.
[0094] Based on the trained model, the entire study area was traversed, and the dengue fever risk was predicted for each grid cell. The average of all model predictions was calculated to form a dengue fever risk map for a 1-kilometer grid cell. Then, the risk values were reclassified using the quartile method: risk levels were divided according to three quartiles, and grid cells with risk values higher than the third quartile were identified as potential high-risk dengue fever areas.
[0095] Step 7: Quantify the comprehensive characteristics of individual buildings
[0096] We obtained publicly available individual building attribute data (as shown in Table 2), and used the building set within each grid as the object. We constructed building weights by comprehensively considering building volume, social function attributes, and building aging effects. We then performed risk conservation allocation within the grid according to the weights to obtain the relative dengue fever risk of individual buildings, and subsequently formulated building prevention and control priority levels.
[0097] (1) Building volume calculation and standardization. For each building, the building volume is calculated as: V = Area × Height, where Area is the building footprint and Height is the building height. To reduce the dominant effect of extremely large buildings on the allocation results, V is truncated to the 5th-95th percentile within each grid and standardized using the median of the building volume within the grid to obtain the standardized volume.
[0098] (2) Characterization of social function attributes. Based on the results of the building function determination, an indicator variable for residential buildings was set: Res=1 (residential buildings), Res=0 (non-residential buildings). This variable is used to characterize whether the building has continuous human residential exposure characteristics.
[0099] (3) Construction of Building Age Modulation Factor. The building age (Age) is calculated based on the year the building was constructed, and a logistic function is introduced to characterize the process of the risk of old buildings gradually increasing over time: AgeFactor=0.85+0.40·σ((Age-25) / 8), where σ(x)=1 / (1+exp(-x)) is the logistic function. The parameter 25 years is the age inflection point, reflecting that problems such as aging drainage systems, weakened environmental management, and container accumulation are more common after residential buildings are built 20-30 years ago; the parameter 8 years is the smoothing scale, used to describe the gradual increase of the aging effect over a wider time interval; the value range of AgeFactor is approximately 0.85-1.25, ensuring that the building age is only used as a risk modulation factor rather than the dominant factor.
[0100] (4) Construction of Building Comprehensive Weights. The individual building weights were constructed by integrating building volume, residential building attributes, and building age modulation effects: w = V^α × (ε + Res)^β × AgeFactor. The parameters in the weight formula were determined using a combination of empirical constraints and sensitivity analysis. First, the volume influence coefficient α was set to a sublinear parameter less than 1 to avoid excessive dominance of large-scale buildings in risk allocation within the grid; the residential building weighting coefficient β was set to an enhancing parameter greater than 1 to highlight the role of continuous residential exposure in dengue fever transmission; the smoothing constant ε was taken as a small positive value to ensure that non-residential buildings retained non-zero weights. Furthermore, different combinations of α, β, and ε were compared in the study area sample, and the fitting effect of building-level case counts was used as the evaluation criterion. Finally, α = 0.7, β = 2.0, and ε = 0.05 were selected as the stable parameter combination.
[0101] As an example, the parameters in the weighting formula were determined using a combination of empirical constraints and sensitivity analysis. First, the volumetric influence coefficient α was set as a sublinear parameter less than 1 to avoid excessive dominance of large-scale buildings in risk allocation within the grid. The residential building weighting coefficient β was set as an enhancing parameter greater than 1 to highlight the role of continuous residential exposure in dengue fever transmission. The smoothing constant ε was chosen as a small positive value to ensure that non-residential buildings retained non-zero weights. Furthermore, different combinations of α, β, and ε were compared in the study area sample, with the fitting effect of building-level case counts used as the evaluation criterion. Finally, α=0.7, β=2.0, and ε=0.05 were selected as the stable parameter combination.
[0102] For the social function attributes of buildings, the building function field in the CMAB multi-attribute building dataset is used first for determination. When this field is missing, attribute information related to building use or building name in the CBF dataset is used to assist in the determination. Buildings marked as residential, apartment, dormitory, apartment building, or residential community supporting housing in the building function field are uniformly identified as residential buildings and recorded as Res=1. Buildings marked as office, commercial, industrial, warehousing, public service, educational, medical, or transportation facilities are uniformly identified as non-residential buildings and recorded as Res=0. For buildings that cannot be directly determined, supplementary identification is performed by combining building name keywords and clustering characteristics of similar surrounding buildings.
[0103] Table 2 Individual Attribute Data of Buildings
[0104]
[0105] Step 8: Prioritize and verify the control measures at the building scale
[0106] Risk conservation and allocation within the grid. For each grid cell, the risk is normalized within the grid according to the building weight w and allocated to the building scale using the grid risk value R. Dengue fever case location data are matched and statistically analyzed with building spatial locations to obtain the number of observed cases at the building level (i.e., the number of cases falling into or matching a building), which is used for subsequent ranking and discriminant analysis. The Top-k capture rate is then used to evaluate the ability of building-level risk to locate cases. Specifically, after obtaining the relative risk value for each building... and the number of observed cases Then, all buildings within the study area were classified according to their relative risk values. Sort the buildings from highest to lowest, and denote the top k% of the buildings as Top(k). Define the Top-k capture rate as: .
[0107] in: This represents the total number of cases covered by buildings ranked in the top k% of risk.
[0108] This represents the total number of cases matched across all buildings within the study area.
[0109] The aforementioned indicator reflects the proportion of all cases that can be covered when prevention and control measures are prioritized only for the top k% of buildings with the highest risk. A higher Top-k capture rate indicates that the building-scale risk ranking is more effective in identifying buildings with clustered cases, and the stronger the localization capability of the plan.
[0110] Here, parameter k represents the proportion of buildings that the public health department can prioritize intervening in within a prevention and control cycle, and its value is determined based on the actual prevention and control resource capacity. The prevention and control resources include, but are not limited to: the number of available epidemic prevention personnel; the number of buildings that can be inspected door-to-door in a single day or week; the building coverage capacity for breeding ground cleanup and environmental remediation; and the maximum carrying capacity for pest control, public awareness campaigns, and key inspections.
[0111] Specifically, the maximum number of buildings that can be intervened in a certain period can be determined first based on the actual prevention and control capabilities, M, and then estimated based on the total number of buildings N in the study area, i.e.: k=M / N*100%.
[0112] As an example, in this embodiment, the case address is converted into latitude and longitude using Amap (Gaode Maps), and case points falling within a 100-meter radius circle centered on the building are counted. If a case point falls within the 100-meter radius of multiple buildings, the closest building is selected for counting.
[0113] In this embodiment, a random forest model is used as the dengue fever risk prediction model. Combined with the SHAP importance ranking method and PDPs partial dependency graph, the correlation between weather factors, urban morphology, socioeconomic factors, and dengue fever risk is analyzed, and a dengue fever risk map at a 1-kilometer grid scale is drawn accordingly. It should be noted that the random forest model is only one optional modeling method for implementing the technical solution of this invention. Without departing from the core idea of this invention, other machine learning models capable of dengue fever risk classification and prediction or spatial risk modeling can be used instead, such as the XGBoost extreme gradient boosting model and the SVM support vector machine model.
[0114] In this embodiment, a 1-kilometer grid is used as the spatial representation unit for dengue fever risk to perform gridded prediction of dengue fever occurrence risk within the study area, and to identify potentially high-risk areas based on this prediction. It should be noted that the 1-kilometer grid is only a preferred spatial scale in one embodiment. This invention does not limit the grid scale itself. In practical applications, other grid scales can be selected as spatial analysis units, such as 500m × 500m grids, 2km × 2km grids, etc., depending on the study area, data resolution, and risk identification accuracy requirements, to achieve spatial characterization of dengue fever risk and identification of potentially high-risk areas.
[0115] In this embodiment, the SHAP importance ranking method and PDP partial dependency graph are used to interpret the model prediction results, in order to identify key influencing factors and reveal the response relationship between each influencing factor and dengue fever risk. It should be noted that the above-described interpretable method is only one implementation of the technical solution of this invention. Without affecting the identification of key influencing factors and the interpretation of the relationships between factors, other model interpretation methods can be used instead, such as the Permutation Importance method, Local Interpretable Model-agnostic Explanations, LIME method, etc.
[0116] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0117] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for prioritizing dengue fever control at the individual building scale, characterized in that, include: Using grid cells as the basic spatial units, an analysis grid is constructed within the study area, and urban morphology type, meteorological factors, and socio-economic factors are calculated for each grid cell to obtain influencing factor data. Based on the addresses and latitude and longitude values of local dengue fever cases in the city, the number of dengue fever cases in each grid unit is calculated, and it is determined whether there are dengue fever cases in each grid unit, so as to obtain a classification target variable used to characterize the dengue fever occurrence status of the grid unit. An interpretable machine learning model is constructed based on the influencing factor data and the classification target variable. The relative contribution of each influencing factor to the prediction of dengue fever risk is quantified based on the interpretable machine learning model, and the response relationship between key driving factors and dengue fever risk is identified. Based on the interpretable machine learning model, the risk of dengue fever occurrence is predicted for each grid cell, a dengue fever risk map is drawn, and potential high-risk areas are identified. Obtain individual attribute data of buildings within the potential high-risk areas, and construct building weights by comprehensively considering building volume, social function attributes, and building aging effects. Within the grid cells corresponding to the potential high-risk areas, the risk of the grid cells is allocated based on the building weights to obtain the relative dengue fever risk of individual buildings, and the priority level of building prevention and control is determined according to the relative dengue fever risk of individual buildings.
2. The method for prioritizing dengue fever control at the individual building scale according to claim 1, characterized in that, For each grid cell, urban morphology type, meteorological factors, and socioeconomic factors are calculated to obtain influencing factor data, including: Meteorological factors are calculated for each of the grid cells; the meteorological factors include monthly average temperature, monthly average total rainfall, and monthly average relative humidity. Calculate the urban morphology type for each grid cell; the urban morphology type includes the proportion of dense high-rise buildings, dense mid-rise buildings, dense low-rise buildings, open high-rise buildings, open mid-rise buildings, open low-rise buildings, large low-rise buildings, sparse buildings, heavy industrial land, dense forests, sparse trees, shrubs, low vegetation, rock or paved land, bare soil or sand, and water surface. Calculate socioeconomic factors for each grid cell; these socioeconomic factors include population density, GDP per capita, and road density. Logarithmic transformation is performed on the population density and GDP per capita to obtain the influencing factor data.
3. The method for prioritizing dengue fever control at the individual building scale according to claim 1, characterized in that, The classification target variables used to characterize the dengue fever occurrence state of the grid cells are obtained, including: Based on the addresses and latitude and longitude values of local dengue fever cases within the city, the number of dengue fever cases in each grid unit is counted. Determine whether there are dengue fever cases in each of the grid units; The grid cell containing a dengue fever case is recorded as 1, and the grid cell containing no dengue fever case is recorded as 0; The result, which is recorded as 1 or 0, is used as the classification target variable.
4. The method for prioritizing dengue fever control at the individual building scale according to claim 1, characterized in that, Based on the influencing factor data and the classification target variable, an interpretable machine learning model is constructed, including: The research area was divided into spatially isolated training and validation areas according to administrative boundaries; Based on the aforementioned influencing factor data and the aforementioned classification target variable, a random forest binary classification model is used to model the risk of dengue fever occurrence. The random forest binary classification model was trained and validated using a 10-fold cross-validation method. The parameters of the random forest binary classification model are optimized using a grid search method, including the number of decision trees and the maximum depth of a single decision tree. The parameter combination with the highest average AUC value during cross-validation is selected as the target parameter combination, and the model training process is repeated multiple times based on the target parameter combination to obtain multiple random forest binary classification models, which serve as the interpretable machine learning model.
5. The method for prioritizing dengue fever control at the individual building scale according to claim 4, characterized in that, Based on the interpretable machine learning model, the relative contributions of each influencing factor to dengue risk prediction are quantified, and the response relationship between key driving factors and dengue risk is identified, including: The SHAP technique is used to interpret the model prediction results of the interpretable machine learning model and quantify the relative contribution of each of the influencing factors to the prediction of dengue fever risk. Key driving factors are identified based on the relative contributions; The response relationship between the key driving factors and the dengue fever risk prediction results was analyzed using PDP technology.
6. The method for prioritizing dengue fever control at the individual building scale according to claim 4, characterized in that, Based on the interpretable machine learning model, the risk of dengue fever occurrence is predicted for each grid cell, a dengue fever risk map is drawn, and potentially high-risk areas are identified, including: Based on the multiple random forest binary classification models, the entire study area is traversed, and the risk of dengue fever occurrence is predicted for each of the grid cells. The predicted values output by the multiple random forest binary classification models are averaged to form a dengue fever risk map; The risk values in the dengue fever risk map were reclassified using the quartile method, and the risk levels were divided according to the three quartiles. Grid cells with risk values above the third quartile are identified as potentially high-risk areas.
7. The method for prioritizing dengue fever control at the individual building scale according to claim 1, characterized in that, Building weights are constructed by comprehensively considering building volume, social function attributes, and the effects of building aging, including: Obtain individual attribute data of buildings within the potential high-risk area; Calculate the building volume based on the building footprint and building height; Within each grid cell, the building volume is truncated to the 5th-95th percentile and standardized using the median of the building volume within the grid to obtain the building volume. Based on the building function determination results, residential building indicator variables are set as the aforementioned social function attributes; The building age is calculated based on the year the building was built, and the building age aging effect is constructed based on the logistic function; The building weights are constructed based on the building volume, the social functional attributes, and the building aging effect.
8. The method for prioritizing dengue fever control at the individual building scale according to claim 7, characterized in that, The building weights are constructed based on the building volume, the social functional attributes, and the building aging effect, including: Let the building volume be denoted as V, and let the social functional attributes be represented by the residential building indicator variable Res, where a value of 1 for Res indicates a residential building, and a value of 0 for Res indicates a non-residential building; The building age Age is calculated based on the year the building was built, and the building age Age is substituted into the logistic function to construct the building age aging effect AgeFactor. The building age aging effect AgeFactor satisfies: AgeFactor=0.85+0.40·σ((Age-25) / 8), where σ(x)=1 / (1+exp(-x)), σ(x) is the logistic function, x represents the normalized deviation of the building age Age from 25 years, 25 is the age inflection point, and 8 is the smoothing scale; The building weight w is constructed according to the weight formula w=V^α×(ε+Res)^β×AgeFactor; where w represents the building weight, α represents the influence coefficient of building volume V, β represents the influence coefficient of social functional attribute Res, and ε represents the smoothing parameter used to ensure that non-residential buildings still retain non-zero weights. The parameter α is set to 0.7, the parameter β is set to 2.0, and the parameter ε is set to 0.
05.
9. The method for prioritizing dengue fever control at the individual building scale according to claim 7, characterized in that, Risk conservation allocation of the grid cells based on the building weights includes: Obtain the grid risk value R of the grid cell corresponding to the potential high-risk area; Based on the grid risk value R and the building weight, a normalized risk conservation allocation is performed within the grid cell; The risk value assigned to the building scale is used as the relative dengue risk Rb for the individual building.
10. The method for prioritizing dengue fever control at the individual building scale according to claim 9, characterized in that, Building prevention and control priorities are assigned based on the relative dengue fever risk of the individual buildings, including: By matching dengue fever case location data with building spatial locations, the number of observed cases at the building level, Yb, is obtained. All buildings within the study area were ranked from highest to lowest according to their relative dengue fever risk (Rb), and the priority level for building prevention and control was determined based on the ranking results. The top k% of buildings in the sorted list are denoted as Top(k); The Top-k capture rate is defined as the ratio of the total number of cases covered by the buildings ranked in the top k% to the total number of cases matched by all buildings in the study area. The Top-k capture rate was used to assess the ability of the building prevention and control priority level to locate cases.