Disease spatial distribution analysis method and device based on environmental risk factors
By constructing a methodological framework that integrates case and environmental data and adopting a geographically weighted regression model, the problem of assessing the spatiotemporal heterogeneity of disease spatial distribution was solved, providing data support for detailed analysis and strategy formulation.
Patent Information
- Application Number
- CN202510594681.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-09
- Publication Date
- 2025-09-05
AI Technical Summary
When evaluating the relationship between epidemics and environmental risk factors, existing technologies ignore the spatiotemporal heterogeneity of disease incidence and environmental risk factors, resulting in biased regression results and making it difficult to reflect the true spatial pattern and spatiotemporal evolution of the impact of environmental risk factors on diseases.
A methodological framework integrating case data, geographic data and environmental data was constructed, and a geographically weighted regression model was used to consider the spatiotemporal heterogeneity of environmental risk factors. The impact of environmental risk factors on the spatial distribution of diseases was evaluated through spatial autocorrelation analysis and multicollinearity test.
It provides spatial distribution analysis of diseases at a fine spatial scale, helps understand the causes of diseases and their influencing factors, and provides high-quality data support for the formulation of targeted environmental intervention measures and disease prevention and control strategies.
Smart Images

Figure CN120600295A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of environmental epidemiological analysis, and specifically to a method and device for analyzing the spatial distribution of diseases based on environmental risk factors. Background Art
[0002] Spatiotemporal heterogeneity, the spatial variation of point distributions or changes in qualitative or quantitative values of surface patterns, is a fundamental characteristic of spatial data. Previous assessments of the relationship between epidemics and environmental risk factors have typically relied on traditional regression models. However, these traditional statistical models often assume that environmental exposures and disease risks are spatially homogeneous within the study area, ignoring the spatiotemporal heterogeneity of disease incidence and environmental risk factors. Furthermore, due to the lack of disease and environmental risk factor data with fine spatial resolution, few methods have addressed the spatial heterogeneity of disease distribution and the environmental risk factors that drive it.
[0003] However, simplification of the relationship between epidemics and environmental risk factors can lead to biased regression results, making it difficult to reflect the true spatial pattern and spatiotemporal evolution of the impact of environmental risk factors on diseases. Therefore, a spatiotemporal analysis method that considers the spatiotemporal heterogeneity of disease and environmental risk factors is needed to fully understand the impact of environmental risk factors on the spatial distribution of diseases and provide sound data support for epidemic prevention and management. Summary of the Invention
[0004] This application provides a method and device for analyzing the spatial distribution of diseases based on environmental risk factors. By constructing a method framework that integrates case data, geographic data, and environmental data, it considers the impact of the spatiotemporal heterogeneity of environmental risk factors on the spatial distribution of diseases at a fine spatial scale, and uses a geographically weighted regression model to predict the relationship between environmental risk factors and the spatial distribution of diseases. It provides new ideas and methods for understanding the causes of the spatial distribution of diseases and their influencing factors, and provides high-quality data support for the formulation of targeted environmental intervention measures and disease prevention and control strategies.
[0005] In a first aspect, the present application provides a method for analyzing the spatial distribution of diseases based on environmental risk factors, the method comprising:
[0006] After determining the spatial distribution analysis task of a specific epidemic disease in the target area, obtain the initial environmental risk factor data and preprocess them to obtain the target environmental risk factor data of the corresponding study area granularity;
[0007] Obtain basic disease statistics and calculate the disease incidence rate at the study area level based on the basic disease statistics;
[0008] Based on the incidence of diseases at the regional level, spatial autocorrelation analysis was used to examine the spatial distribution of diseases;
[0009] Combined with the target environmental risk factor data and the spatial distribution of the disease, all environmental risk factors were subjected to multicollinearity tests, significance tests, and least squares model regression. Based on the test results, environmental risk factors that did not meet the requirements were eliminated to obtain the target environmental risk factors.
[0010] Taking the target environmental risk factors as independent variables and the disease incidence as the dependent variable, the geographically weighted regression method was used to evaluate the effect of the target environmental risk factors on the spatial distribution of the disease.
[0011] In a second aspect, the present application provides a device for analyzing the spatial distribution of diseases based on environmental risk factors, the device comprising:
[0012] The first acquisition unit is used to obtain initial environmental risk factor data and preprocess it after determining the disease spatial distribution analysis task of a specific epidemic directed to the target area, so as to obtain target environmental risk factor data of the granularity corresponding to the study area;
[0013] The second acquisition unit is used to obtain basic disease statistical data and calculate the disease incidence rate at the research area level based on the basic disease statistical data;
[0014] The test unit is used to test the spatial distribution of diseases through spatial autocorrelation analysis based on the incidence of diseases at the study area level;
[0015] Elimination unit is used to combine the target environmental risk factor data and the spatial distribution of the disease, conduct multicollinearity test, significance test and least squares model regression on all environmental risk factors, and eliminate environmental risk factors that do not meet the requirements based on the test results to obtain the target environmental risk factors;
[0016] The evaluation unit is used to use the target environmental risk factor as the independent variable and the disease incidence as the dependent variable, and to evaluate the effect of the target environmental risk factor on the spatial distribution of the disease based on the geographically weighted regression method.
[0017] In a third aspect, the present application provides a processing device comprising a processor and a memory, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the method provided in the first aspect of the present application or any possible implementation of the first aspect of the present application is executed.
[0018] In a fourth aspect, the present application provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for a processor to load to execute the method provided in the first aspect of the present application or any possible implementation of the first aspect of the present application.
[0019] From the above content, it can be concluded that this application has the following beneficial effects:
[0020] Aiming to evaluate and analyze the impact of environmental risk factors on the spatial distribution of diseases, this application constructs a methodological framework that integrates case data, geographic data, and environmental data. It considers the impact of the spatiotemporal heterogeneity of environmental risk factors on the spatial distribution of diseases at a fine spatial scale, and uses a geographically weighted regression model to predict the relationship between environmental risk factors and the spatial distribution of diseases. This provides new ideas and methods for understanding the causes of the spatial distribution of diseases and their influencing factors, and provides high-quality data support for the formulation of targeted environmental intervention measures and disease prevention and control strategies. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0022] Figure 1 A schematic diagram of a process for analyzing the spatial distribution of diseases based on environmental risk factors in this application;
[0023] Figure 2 This is a schematic diagram of the structure of a disease spatial distribution analysis device based on environmental risk factors in this application;
[0024] Figure 3 This is a structural diagram of the processing equipment for this application. DETAILED DESCRIPTION
[0025] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0026] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not necessarily limited to those steps or modules clearly listed, but may include other steps or modules that are not clearly listed or that are inherent to these processes, methods, products or devices. The naming or numbering of steps in this application does not mean that the steps in the method flow must be executed in the time / logical sequence indicated by the naming or numbering. The process steps that have been named or numbered can be changed in the execution order according to the technical purpose to be achieved, as long as the same or similar technical effects can be achieved.
[0027] The division of modules in this application is a logical division. In actual application, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection between modules can be electrical or other similar forms, which are not limited in this application. Moreover, the modules or submodules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed into multiple circuit modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application.
[0028] Before introducing the disease spatial distribution analysis method based on environmental risk factors provided by this application, the background content involved in this application is first introduced.
[0029] The disease spatial distribution analysis method, device and computer-readable storage medium based on environmental risk factors provided in this application can be applied to processing equipment. By constructing a method framework that integrates case data, geographic data and environmental data, it considers the impact of the spatiotemporal heterogeneity of environmental risk factors on the spatial distribution of diseases at a fine spatial scale, and uses a geographically weighted regression model to predict the relationship between environmental risk factors and the spatial distribution of diseases. It provides new ideas and methods for understanding the causes of the spatial distribution of diseases and their influencing factors, and provides high-quality data support for the formulation of targeted environmental intervention measures and disease prevention and control strategies.
[0030] The disease spatial distribution analysis method based on environmental risk factors mentioned in this application can be executed by a disease spatial distribution analysis device based on environmental risk factors, or a server, physical host or even user equipment (UE) and other different types of processing devices that integrate the disease spatial distribution analysis device based on environmental risk factors. Among them, the disease spatial distribution analysis device based on environmental risk factors can be implemented in hardware or software. The UE can specifically be a terminal device such as a smart phone, tablet computer, laptop computer, desktop computer or personal digital assistant (PDA), and the processing device can be set up in the form of a device cluster.
[0031] In specific applications, it can be understood that the solution of this application is mainly data processing based on existing data, and does not involve the direct collection of source data. Therefore, for the processing equipment that executes the disease spatial distribution analysis method based on environmental risk factors of this application or is equipped with the corresponding application service of the disease spatial distribution analysis method based on environmental risk factors of this application, it usually only needs to have the required data processing capabilities. The specific equipment type and equipment deployment form are relatively flexible and can be adaptively configured according to actual needs.
[0032] As an example, the processing device may be a physical host in a laboratory, or may be a server providing cloud computing services.
[0033] If it also involves work responsibilities such as the ability to directly collect source data or display results, it is easy to understand that further adaptive configuration of the processing equipment is required.
[0034] Taking result display as an example, the processing device needs to be equipped with a corresponding display screen (including a touch screen) itself, or the processing device can also meet the result display requirements through an external display screen device or other device with a display screen.
[0035] Next, we will introduce the disease spatial distribution analysis method based on environmental risk factors provided by this application.
[0036] First, see Figure 1 , Figure 1 A flow chart of the disease spatial distribution analysis method based on environmental risk factors of the present application is shown. The disease spatial distribution analysis method based on environmental risk factors provided by the present application may specifically include the following steps S101 to S105:
[0037] Step S101: After determining a disease spatial distribution analysis task for a specific epidemic in a target area, initial environmental risk factor data is obtained and preprocessed to obtain target environmental risk factor data of a granularity corresponding to the study area.
[0038] It can be understood that the spatial distribution analysis and processing of diseases based on environmental risk factors in this application is actually carried out in the form of tasks, and is carried out for specific regions and specific epidemics. Correspondingly, in the initial stage of obtaining input data, it is necessary to follow the processing scope.
[0039] For the convenience of explanation, the area targeted by the current disease spatial distribution analysis task is recorded as the target area, and the targeted epidemic is recorded as the specific epidemic.
[0040] As an example, the target area may be the entire country, or a provincial or prefectural administrative region, and the specific epidemic may be an epidemic of thyroid disease or the like.
[0041] The task information of the current disease spatial distribution analysis task may directly include the required data input or indirectly indicate the required data input, which can be flexibly adjusted according to actual conditions.
[0042] Similarly, the acquisition of the current disease spatial distribution analysis task can be either manually entered, autonomously generated according to the corresponding autonomous generation strategy, or actively / passively obtained from external devices.
[0043] The initial environmental risk factor data obtained here are data describing environmental factors / elements. In this application, these environmental factors / elements are related to the spatial distribution of diseases and are therefore called environmental risk factors.
[0044] After obtaining the initial environmental risk factor data related to the current situation, this application can pre-process it to obtain high-quality data for subsequent use.
[0045] Among them, the preprocessing that can be carried out is generally for enhancing data quality. Therefore, the specific preprocessing operations that can be applied can refer to the existing technology, and this application does not make too many detailed explanations in this regard.
[0046] As an exemplary embodiment herein, the environmental risk factors involved in the initial environmental risk factor data may specifically include PM 2.5 、PM 10, CO (carbon monoxide), O3 (ozone), NO2 (nitrogen dioxide), SO2 (sulfur dioxide), temperature, humidity, Normalized Difference Vegetation Index (NDVI), GDP (Gross Domestic Product) and industrial point of interest (POI) density.
[0047] It can be seen that the environmental risk factors that this application may specifically involve include air, ground, and economic aspects such as GDP and density of industrial points of interest. The latter is an extension and deepening of the environmental risk factors in this application. It is believed that it can also have a certain degree of impact on the spatial distribution of diseases and needs to be analyzed as an influencing factor.
[0048] The density of industrial points of interest refers specifically to the density of industry-related elements such as factory addresses, industrial waste gas emission points, wastewater emission points, and waste residue emission points.
[0049] At the same time, for preprocessing, this application also provides a set of implementation solutions from a deeper level that are highly consistent with the scenarios cited in this application and help to better promote subsequent data processing.
[0050] Specifically, as an exemplary embodiment herein, initial environmental risk factor data is obtained and preprocessed to obtain target environmental risk factor data of a granularity corresponding to the study area, which may specifically include:
[0051] 1.1. Obtain an environmental risk raster dataset with a resolution of 1 km × 1 km based on the integration of satellite remote sensing and ground monitoring;
[0052] It can be seen that this application uses a fusion data product of remote sensing technology and ground monitoring technology, thereby obtaining a high-resolution environmental risk raster dataset that is configured with raster or grid units for environmental risk factors.
[0053] Among them, for the environmental risk raster dataset, in actual situations, open source, that is, publicly available, related data products can be used, or it can be obtained by integrating open source satellite remote sensing data products and open source ground monitoring data products.
[0054] 1.2. Based on the base map of the primary study area, the initial environmental risk factor data configured in raster data format are clipped to obtain granular environmental risk raster data for the primary study area.
[0055] After obtaining the initial environmental risk factor data of the target area, the corresponding clipping tools or data segmentation algorithms can be used to perform fine-grained clipping and division according to the base map of the primary study area (that is, the map base map of different study areas divided by the target area) to form the granular environmental risk raster data of the primary study area.
[0056] Obviously, there are multiple first-level research areas, which constitute the overall target area.
[0057] 1.3. Continue to perform zoning statistics on the environmental risk raster data to obtain granular target environmental risk factor data for the secondary study area.
[0058] After the environmental risk raster data corresponding to the first-level study area has been preliminarily cropped and divided, this application can also correspond to the second-level study area that is smaller / smaller than the first-level study area, and continue to crop and divide, that is, zoning statistics, to obtain the granular target environmental risk factor data of the second-level study area.
[0059] It is understandable that the primary and secondary research areas involved here can usually be configured according to the corresponding administrative divisions. Of course, it is not ruled out that similar to the target area, flexible settings with independent planning can be adopted to meet the diverse application needs of the scheme in actual situations.
[0060] At this time, the granularity of the secondary research area of the target environmental risk factor data also corresponds to the granularity / unit level of subsequent data processing. In other words, the spatial distribution analysis and processing of diseases based on environmental risk factors to be implemented in this application is the granularity of the secondary research area.
[0061] Step S102, obtaining basic disease statistical data, and calculating the disease incidence rate at the study area level based on the basic disease statistical data;
[0062] On the other hand, this application also involves the statistics of disease incidence to provide data basis for subsequent analysis of disease spatial distribution.
[0063] To this end, this application can obtain basic disease statistics of specific epidemics in the target area of this mission, and on this basis, process them into study area-level disease incidence rates that align with the study area granularity of the previous target environmental risk factor data (such as the granularity of the previous exemplary secondary study area).
[0064] Among them, the basic disease statistical data obtained here are themselves open source data products, which can be obtained from local hospitals and other channels.
[0065] Corresponding to the granularity of the secondary study area in the previous exemplary embodiment, the disease incidence rate at the study area level is calculated based on the basic disease statistical data. As an exemplary embodiment here, it may specifically include:
[0066] 2.1. Obtain disease statistics and the geographic locations of cases, and use spatial join tools to obtain the number of disease cases in each secondary study area;
[0067] It is understandable that the corresponding cases of the disease statistical data obtained above may directly carry the corresponding geographical location of the cases, or may be identified or provided for finding the corresponding geographical location of the cases through indirect means.
[0068] In this way, it can be mapped to the spatial level and matched to the secondary research area to which it belongs, so that the number of disease cases in each research area can be obtained through the corresponding spatial connection tools.
[0069] 2.2. Calculate the incidence rate per 100,000 people in each secondary study area to obtain the disease incidence rate at the study area level. The corresponding calculation formula is:
[0070]
[0071] Among them, I i is the incidence rate per 100,000 people in the i-th secondary study area, N i is the total number of cases in the ith secondary study area, P i is the total population of the i-th secondary study area.
[0072] It can be understood that here, starting from the quantitative formula level, a calculation plan for how to calculate the incidence rate per 100,000 people in the secondary research area is given.
[0073] Step S103, based on the regional-level disease incidence, examine the spatial distribution of the disease through spatial autocorrelation analysis;
[0074] After the previous processing obtained the disease incidence rate at the study area level, for example, the previous exemplary embodiment obtained the incidence rate per 100,000 people in the secondary study area, it can be understood that this application can perform further data analysis to obtain the specific spatial distribution of the disease or the type of spatial distribution of the disease, to provide data support for subsequent research and evaluation of the role of target environmental risk factors in the spatial distribution of the disease, which may specifically involve data analysis and processing of spatial autocorrelation analysis.
[0075] As an exemplary embodiment herein, based on the study of regional-level disease incidence, spatial autocorrelation analysis is performed to examine the spatial distribution of the disease, which may specifically include:
[0076] Based on the regional disease incidence rate, global spatial autocorrelation analysis and local spatial autocorrelation analysis were conducted using the ArcGIS platform. The evaluation index selected for the global spatial autocorrelation analysis was the global Moran's index, and the evaluation index selected for the local spatial autocorrelation analysis was the local Moran's index.
[0077] It can be seen that in the embodiment of this application, the spatial autocorrelation analysis task is completed by global spatial autocorrelation analysis based on the global Moran's index indicator and local spatial autocorrelation analysis based on the local Moran's index indicator.
[0078] For this purpose, the existing ArcGIS platform / software can be used.
[0079] It is understandable that the Global Moran's I and Local Moran's I indicators are existing indicators. Their corresponding contents can be referred to as follows:
[0080] 3.1. Global Moran's I
[0081] When the p-value is less than 0.05, when the Global Moran's I index is greater than 0, it indicates that the disease presents a significant clustered distribution pattern in space; when the Global Moran's I index is less than 0, it presents a significant discrete distribution pattern; and when the Global Moran's I index is equal to 0, it indicates that the disease incidence rate is randomly distributed in space.
[0082] The Global Moran's I indicator is defined as follows:
[0083]
[0084] Among them, z i is the deviation of the disease incidence rate in the i-th secondary study area from the mean value, w i,j is the location weight between the i-th secondary study area and the j-th secondary study area, n is the total number of secondary study areas, and S0 represents the sum of all weight values.
[0085] 3.2. Local Moran's I
[0086] Local spatial autocorrelation analysis is used to test the degree of spatial clustering of incidence rates between adjacent regions. When the Local Moran's I index is selected, a positive Local Moran's I index indicates that high-incidence areas are adjacent to high-incidence areas or low-incidence areas are connected to low-incidence areas (HH cluster or LL cluster), and a negative Local Moran's I index indicates that high-incidence areas are surrounded by low-incidence areas as spatial hot spots or low-incidence areas are surrounded by high-incidence areas as spatial cold spots (HL cluster or LH cluster).
[0087] The Local Moran's I indicator is defined as follows:
[0088]
[0089] Among them, w ij is the row-normalized adjacency matrix, x i is the disease incidence rate in the i-th secondary study area, x j The disease incidence rate of the jth secondary study area adjacent to the i-th secondary study area, μ is the average incidence rate of each area.
[0090] Step S104: Combine the target environmental risk factor data and the spatial distribution of the disease, perform multicollinearity test, significance test and least squares model regression on all environmental risk factors, and eliminate environmental risk factors that do not meet the requirements based on the test results to obtain the target environmental risk factors;
[0091] At the same time, corresponding to subsequent research and evaluation of the role of target environmental risk factors in the spatial distribution of diseases, this application also involves the screening and simplification of environmental risk factors, which means that environmental risk factors that do not meet the requirements need to be eliminated in order to promote more accurate and simpler treatment effects.
[0092] In this regard, this application specifically conducts multicollinearity test, significance test and least squares model regression on all environmental risk factors involved at the beginning based on the target environmental risk factor data and disease spatial distribution obtained previously, and then eliminates environmental risk factors that do not meet the requirements based on the test results of these three aspects.
[0093] Specifically, as an exemplary embodiment, here, in combination with the target environmental risk factor data and the spatial distribution of the disease, all environmental risk factors are subjected to multicollinearity test, significance test and least squares model regression, and environmental risk factors that do not meet the requirements are eliminated based on the test results to obtain the target environmental risk factors, which may include:
[0094] 4.1. Perform a multicollinearity test on all environmental risk factors involved in the target environmental risk factor data. If the variance inflation factor (VIF) of the current environmental risk factor is less than 5, it is considered that there is no strong multicollinearity and meets the requirements. The formula for calculating the variance inflation factor is:
[0095]
[0096] in, is the square of the multiple correlation coefficient of the kth environmental risk factor to other factors;
[0097] 4.2. All environmental risk factors were subjected to correlation analysis and significance testing using the spatial least squares model on the ArcGIS platform. If the correlation coefficient was greater than 0.7, it was considered to have a strong positive correlation and meet the requirements. If the significance p-value was greater than 0.05, it was considered to have failed the significance test and did not meet the requirements.
[0098] It can be understood that in the embodiment here, based on specific formulas and thresholds, a specific filtering scheme for environmental risk factors is given, thereby comprehensively considering the variance inflation factor, the correlation coefficient between environmental risk factors and disease incidence, and the significance p-value to effectively eliminate some environmental risk factors, so that the variance inflation factors of the remaining environmental risk factors are all less than 5 and pass the significance test.
[0099] Step S105 , using the target environmental risk factor as an independent variable and the disease incidence rate as a dependent variable, the effect of the target environmental risk factor on the spatial distribution of the disease is evaluated based on the geographically weighted regression method.
[0100] It can be seen that after the previous series of processing, the simplified target environmental risk factors and the spatial distribution of diseases are obtained. The former can be used as the independent variable, and the disease incidence involved in the latter can be used as the dependent variable. The geographically weighted regression method can be specifically introduced to efficiently and accurately carry out the evaluation and analysis of the role of target environmental risk factors in the spatial distribution of diseases.
[0101] In layman's terms, the geographically weighted regression method is a regression analysis method based on geographical weighting technology. In traditional spatial regression analysis, it is assumed that the weights of all sample points are equal, that is, each sample point has the same impact on the regression analysis. However, in geographic space, there may be geographic spatial correlation between sample points, that is, adjacent sample points may have a greater impact on the regression results. Therefore, through the geographically weighted matrix, the weight of each sample point is combined with its geographic spatial correlation, which can more accurately reflect the regression relationship in geographic space. Furthermore, the geographic spatial relationship can be combined to more deeply and accurately complete the evaluation goal of the role of target environmental risk factors in the spatial distribution of diseases, and overcome the problems of spatial non-independence and spatial heterogeneity.
[0102] Specifically, as an example here, the formula involved in the geographically weighted regression model can be:
[0103]
[0104] Among them, y i is the disease incidence rate in the ith secondary study area, (u i ,v i ) is the geographical coordinate of the i-th secondary study area, β0(u i ,v i ) is the intercept of the i-th secondary study area, β k (u i ,v i ) is the local coefficient of the i-th secondary study area, p is the number of factors of the dependent variable, x ik is the independent variable at the i-th secondary research area, ε i is the random error term of the ith secondary study area, and the geographically weighted regression model uses R 2 The AICc index and local regression coefficient were used to evaluate the processing results of the geographically weighted regression model.
[0105] Among them, it is understandable that R 2 The three evaluation indicators, index, AICc index and local regression coefficient, are existing indicators themselves, so they are not explained in detail. The same applies to the indicators involved in the previous factor elimination link.
[0106] Specifically, in the present application, the higher R 2This shows that areas with high environmental risk factors can better explain the changes in local disease incidence; the AICc index of different models is compared, and a smaller AICc index indicates a better model regression effect; by visualizing the local regression coefficients of environmental risk factors and incidence in geographically weighted regression, the spatial non-stationarity of disease incidence under the influence of environmental risk factors can be verified, and the spatial differences in the strength of each environmental risk factor in different regions can be evaluated.
[0107] To keep the PM 2.5 Taking the two target environmental risk factors of thyroid disease and industrial point of interest density as examples, for the task of analyzing the spatial distribution of basic thyroid diseases in a certain area, the geographically weighted regression results can be as shown in the following example in Table 1:
[0108] Table 1 - Example of Geographically Weighted Regression Results
[0109] 2016 2017 2018 2019 2020 <![CDATA[R 2 ]]> 0.680 0.540 0.505 0.552 0.551 AICc 1661.7 1748.9 1767.7 1664.1 1711.6
[0110] At the same time, after completing the evaluation of the effect of target environmental risk factors on the spatial distribution of diseases through the geographically weighted regression method, or obtaining the evaluation results of the effect of target environmental risk factors on the spatial distribution of diseases, further data application links may be involved.
[0111] For example, the evaluation results of the effect of target environmental risk factors on the spatial distribution of diseases can be stored locally, stored remotely, pushed, output as a prompt to indicate the completion of the evaluation, displayed, or further analyzed and processed. These are all possible, and can be adaptively processed with pre-configured and real-time configured data application strategies to meet the flexible and changeable solution application needs in actual situations.
[0112] Finally, regarding the above program content, in general, for the goal of evaluating and analyzing the impact of environmental risk factors on the spatial distribution of diseases, this application constructs a method framework that integrates case data, geographic data, and environmental data, considers the impact of the spatiotemporal heterogeneity of environmental risk factors on the spatial distribution of diseases at a fine spatial scale, and uses a geographically weighted regression model to predict the relationship between environmental risk factors and the spatial distribution of diseases, providing new ideas and methods for understanding the causes of the spatial distribution of diseases and their influencing factors, and providing high-quality data support for the formulation of targeted environmental intervention measures and disease prevention and control strategies.
[0113] The above is an introduction to the disease spatial distribution analysis method based on environmental risk factors provided by this application. In order to facilitate better implementation of the disease spatial distribution analysis method based on environmental risk factors provided by this application, this application also provides a disease spatial distribution analysis device based on environmental risk factors from the perspective of functional modules.
[0114] See Figure 2 , Figure 2 This is a schematic diagram of the structure of a disease spatial distribution analysis device based on environmental risk factors in this application. In this application, the disease spatial distribution analysis device 200 based on environmental risk factors may specifically include the following structure:
[0115] The first acquisition unit 201 is used to acquire initial environmental risk factor data and perform preprocessing after determining a disease spatial distribution analysis task for a specific epidemic in a target area to obtain target environmental risk factor data of a granularity corresponding to the study area;
[0116] The second acquisition unit 202 is used to acquire basic disease statistical data and calculate the disease incidence rate at the research area level based on the basic disease statistical data;
[0117] A testing unit 203 is used to test the spatial distribution of diseases through spatial autocorrelation analysis based on the disease incidence rate at the research area level;
[0118] Elimination unit 204 is used to combine the target environmental risk factor data and the spatial distribution of the disease, perform multicollinearity test, significance test and least squares model regression on all environmental risk factors, and eliminate environmental risk factors that do not meet the requirements based on the test results to obtain the target environmental risk factors;
[0119] The evaluation unit 205 is configured to use the target environmental risk factor as an independent variable and the disease incidence rate as a dependent variable to evaluate the effect of the target environmental risk factor on the spatial distribution of the disease based on a geographically weighted regression method.
[0120] As an exemplary embodiment, the environmental risk factors involved in the initial environmental risk factor data include PM 2.5 、PM 10 , CO, O3, NO2, SO2, temperature, humidity, normalized difference vegetation index, GDP and density of industrial points of interest.
[0121] As another exemplary embodiment, the first acquiring unit 201 is specifically configured to:
[0122] Obtain an environmental risk raster dataset with a resolution of 1km×1km based on the fusion of satellite remote sensing and ground monitoring;
[0123] Combined with the base map of the primary study area, the initial environmental risk factor data configured in the form of raster data are clipped to obtain the granular environmental risk raster data of the primary study area;
[0124] The environmental risk raster data is further subjected to zoning statistics to obtain the granular target environmental risk factor data for the secondary study area.
[0125] As another exemplary embodiment, the second acquiring unit 202 is specifically configured to:
[0126] Obtain disease statistics and the geographic location of cases, and use spatial join tools to obtain the number of disease cases in each secondary study area;
[0127] Calculate the incidence rate of 100,000 people in each secondary study area to obtain the disease incidence rate at the study area level. The corresponding calculation formula is:
[0128]
[0129] Among them, I i is the incidence rate per 100,000 people in the i-th secondary study area, N i is the total number of cases in the ith secondary study area, P i is the total population of the i-th secondary study area.
[0130] As another exemplary embodiment, the inspection unit 203 is specifically configured to:
[0131] Based on the regional disease incidence rate, global spatial autocorrelation analysis and local spatial autocorrelation analysis were conducted using the ArcGIS platform. The evaluation index selected for the global spatial autocorrelation analysis was the global Moran's index, and the evaluation index selected for the local spatial autocorrelation analysis was the local Moran's index.
[0132] As another exemplary embodiment, the elimination unit 204 is specifically configured to:
[0133] Perform a multicollinearity test on all environmental risk factors involved in the target environmental risk factor data. If the variance inflation factor of the current environmental risk factor is less than 5, it is considered that there is no strong multicollinearity and meets the requirements. The calculation formula of the variance inflation factor involved is:
[0134]
[0135] in, is the square of the multiple correlation coefficient of the kth environmental risk factor to other factors;
[0136] The ArcGIS platform was used to conduct correlation analysis and significance test of the spatial least squares model for all environmental risk factors. If the correlation coefficient was greater than 0.7, it was considered to have a strong positive correlation and meet the requirements. If the significance p value was greater than 0.05, it was considered to have failed the significance test and did not meet the requirements.
[0137] As another exemplary embodiment, the formula involved in the geographically weighted regression model is:
[0138]
[0139] Among them, y i is the disease incidence rate in the ith secondary study area, (u i ,v i ) is the geographical coordinate of the i-th secondary study area, β0(u i ,v i ) is the intercept of the i-th secondary study area, β k (u i ,v i ) is the local coefficient of the i-th secondary study area, p is the number of factors of the dependent variable, x ik is the independent variable at the i-th secondary research area, ε i is the random error term of the ith secondary study area, and the geographically weighted regression model uses R 2 The AICc index and local regression coefficient were used to evaluate the processing results of the geographically weighted regression model.
[0140] This application also provides a processing device from the perspective of hardware structure, see Figure 3 , Figure 3 The schematic diagram of the structure of the processing device of the present application is shown. Specifically, the processing device of the present application may include a processor 301, a memory 302 and an input / output device 303. The processor 301 is used to execute the computer program stored in the memory 302 to implement the following Figure 1 Each step of the disease spatial distribution analysis method based on environmental risk factors in the corresponding embodiment; or, when the processor 301 is used to execute the computer program stored in the memory 302, the following is implemented Figure 2 The memory 302 is used to store the functions of each unit in the embodiment corresponding to the processor 301. Figure 1 The computer program required for the method for analyzing the spatial distribution of diseases based on environmental risk factors in the corresponding embodiment.
[0141] For example, the computer program may be divided into one or more modules / units, one or more of which are stored in the memory 302 and executed by the processor 301 to complete the present application. One or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in a computer device.
[0142] The processing device may include, but is not limited to, a processor 301, a memory 302, and an input / output device 303. Those skilled in the art will appreciate that the illustrations are merely examples of processing devices and do not limit the processing device. The processing device may include more or fewer components than shown, or a combination of certain components, or different components. For example, the processing device may also include a network access device, a bus, etc., and the processor 301, the memory 302, the input / output device 303, etc. are connected via a bus.
[0143] The processor 301 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor is the control center of the processing device and connects various parts of the entire device using various interfaces and lines.
[0144] The memory 302 can be used to store computer programs and / or modules. The processor 301 implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory 302 and accessing the data stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area. The program storage area may store an operating system, at least one application required for a function, etc.; the data storage area may store data created based on the use of the processing device, etc. In addition, the memory may include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0145] When the processor 301 is used to execute the computer program stored in the memory 302, it can specifically implement the following functions:
[0146] After determining the spatial distribution analysis task of a specific epidemic disease in the target area, obtain the initial environmental risk factor data and preprocess them to obtain the target environmental risk factor data of the corresponding study area granularity;
[0147] Obtain basic disease statistics and calculate the disease incidence rate at the study area level based on the basic disease statistics;
[0148] Based on the incidence of diseases at the regional level, spatial autocorrelation analysis was used to examine the spatial distribution of diseases;
[0149] Combined with the target environmental risk factor data and the spatial distribution of the disease, all environmental risk factors were subjected to multicollinearity tests, significance tests, and least squares model regression. Based on the test results, environmental risk factors that did not meet the requirements were eliminated to obtain the target environmental risk factors.
[0150] Taking the target environmental risk factors as independent variables and the disease incidence as the dependent variable, the geographically weighted regression method was used to evaluate the effect of the target environmental risk factors on the spatial distribution of the disease.
[0151] Those skilled in the art will clearly understand that for the convenience and brevity of description, the specific working process of the disease spatial distribution analysis device, processing equipment and corresponding units described above based on environmental risk factors can be referred to as follows: Figure 1 The description of the disease spatial distribution analysis method based on environmental risk factors in the corresponding embodiment will not be repeated here.
[0152] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be accomplished by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.
[0153] To this end, the present application provides a computer-readable storage medium, which stores a plurality of instructions, which can be loaded by a processor to execute the present application as follows: Figure 1 The steps of the disease spatial distribution analysis method based on environmental risk factors in the corresponding embodiment, the specific operations can be referred to as follows Figure 1 The description of the disease spatial distribution analysis method based on environmental risk factors in the corresponding embodiment will not be repeated here.
[0154] The computer-readable storage medium may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0155] Due to the instructions stored in the computer readable storage medium, the present application can be executed as follows: Figure 1 The steps of the disease spatial distribution analysis method based on environmental risk factors in the corresponding embodiment can thus realize the present application as follows Figure 1The beneficial effects that can be achieved by the disease spatial distribution analysis method based on environmental risk factors in the corresponding embodiment are detailed in the previous description and will not be repeated here.
[0156] The above is a detailed introduction to the disease spatial distribution analysis method, device, processing equipment and computer-readable storage medium based on environmental risk factors provided by this application. Specific examples are used in this article to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the core idea of this application; at the same time, for technical personnel in this field, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on this application.
Claims
1. A disease spatial distribution analysis method based on environmental risk factors, characterized in that: The method comprises: After determining the spatial distribution analysis task of a specific epidemic disease in the target area, obtain the initial environmental risk factor data and preprocess them to obtain the target environmental risk factor data of the corresponding study area granularity; Obtain basic disease statistical data and calculate the disease incidence rate at the study area level based on the basic disease statistical data; Based on the incidence of diseases at the study area level, spatial autocorrelation analysis was used to examine the spatial distribution of diseases; Combining the target environmental risk factor data and the spatial distribution of the disease, performing multicollinearity tests, significance tests, and least squares model regression on all environmental risk factors, and eliminating environmental risk factors that do not meet the requirements based on the test results to obtain target environmental risk factors; The target environmental risk factors were used as independent variables and the disease incidence rate was used as the dependent variable. The effect of the target environmental risk factors on the spatial distribution of the disease was evaluated based on the geographically weighted regression method.
2. The method according to claim 1, characterized in that The environmental risk factors involved in the initial environmental risk factor data include PM 2.5 、PM 10 , CO, O3, NO2, SO2, temperature, humidity, normalized difference vegetation index, GDP and density of industrial points of interest.
3. The method according to claim 1, characterized in that The initial environmental risk factor data is obtained and preprocessed to obtain target environmental risk factor data of the corresponding study area granularity, including: Obtain an environmental risk raster dataset with a resolution of 1km×1km based on the fusion of satellite remote sensing and ground monitoring; In combination with the base map of the primary study area, the initial environmental risk factor data configured in the form of raster data are clipped to obtain granular environmental risk raster data of the primary study area; The environmental risk grid data is further subjected to zoning statistics to obtain the target environmental risk factor data of the granularity of the secondary research area.
4. The method according to claim 3, characterized in that The disease incidence rate at the study area level is calculated based on the basic disease statistical data, including: Obtain disease statistics and the geographical locations of cases, and use spatial join tools to obtain the number of disease cases in each secondary study area; Calculate the incidence rate of 100,000 people in each secondary study area to obtain the disease incidence rate at the study area level. The corresponding calculation formula is: Among them, I i is the incidence rate per 100,000 people in the i-th secondary study area, N i is the total number of cases in the i-th secondary study area, P i is the total population of the i-th secondary study area.
5. The method according to claim 3, characterized in that Based on the disease incidence rate at the study area level, the spatial distribution of the disease was examined through spatial autocorrelation analysis, including: Based on the disease incidence rate at the study area level, global spatial autocorrelation analysis and local spatial autocorrelation analysis were performed using the ArcGIS platform. The evaluation index selected for the global spatial autocorrelation analysis was the global Moran's index index, and the evaluation index selected for the local spatial autocorrelation analysis was the local Moran's index index.
6. The method according to claim 3, characterized in that The target environmental risk factor data and the spatial distribution of the disease are combined to perform multicollinearity test, significance test and least squares model regression on all environmental risk factors, and environmental risk factors that do not meet the requirements are eliminated according to the test results to obtain target environmental risk factors, including: A multicollinearity test is performed on all the environmental risk factors involved in the target environmental risk factor data. If the variance inflation factor of the current environmental risk factor is less than 5, it is considered that there is no strong multicollinearity and the requirements are met. The variance inflation factor calculation formula involved is: in, is the square of the multiple correlation coefficient of the kth environmental risk factor to other factors; The ArcGIS platform was used to conduct correlation analysis and significance test of the spatial least squares model for all the environmental risk factors. If the correlation coefficient was greater than 0.7, it was considered that there was a strong positive correlation and the requirements were met. If the significance p value was greater than 0.05, it was considered that the significance test had not been passed and the requirements were not met.
7. The method according to claim 3, characterized in that The formula involved in the geographically weighted regression model is: Among them, y i is the disease incidence rate of the i-th secondary study area, (u i ,v i ) is the geographical coordinate of the i-th secondary study area, β0(u i ,v i ) is the intercept of the i-th secondary research area, β k (u i ,v i ) is the local coefficient of the i-th secondary study area, p is the number of factors of the dependent variable, x ik is the independent variable at the i-th secondary research area, ε i is the random error term of the ith secondary study area, and the geographically weighted regression model uses R 2 The AICc index and local regression coefficient were used to evaluate the processing results of the geographically weighted regression model.
8. A disease spatial distribution analysis device based on environmental risk factors, characterized in that: The device comprises: The first acquisition unit is used to obtain initial environmental risk factor data and preprocess it after determining the disease spatial distribution analysis task of a specific epidemic directed to the target area, so as to obtain target environmental risk factor data of the granularity corresponding to the study area; The second acquisition unit is used to obtain basic disease statistical data and calculate the disease incidence rate at the research area level based on the basic disease statistical data; A testing unit is used to test the spatial distribution of diseases through spatial autocorrelation analysis based on the disease incidence rate at the study area level; An elimination unit is used to combine the target environmental risk factor data and the spatial distribution of the disease, perform multicollinearity test, significance test and least squares model regression on all environmental risk factors, and eliminate environmental risk factors that do not meet the requirements according to the test results to obtain target environmental risk factors; An evaluation unit is used to use the target environmental risk factor as an independent variable and the disease incidence rate as a dependent variable to evaluate the effect of the target environmental risk factor on the spatial distribution of the disease based on a geographically weighted regression method.
9. A processing device, characterized in that The method comprises a processor and a memory, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the method according to any one of claims 1 to 7 is executed.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a plurality of instructions, and the instructions are suitable for being loaded by a processor to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
Disease distribution area intelligent analysis method and system
CN121638941A