Avian influenza virus overflow space risk assessment method based on interpretable machine learning and space weighted sampling

By integrating global multi-source data based on interpretable machine learning and spatially weighted sampling, generating environmental characteristic raster layers, screening key driving factors, and quantifying the risk of avian influenza virus spillover, the problem of inaccurate risk prediction in existing technologies is solved, and a refined assessment and factor interpretation of global epidemic risks are achieved.

CN120613152APending Publication Date: 2025-09-09INST OF GEOGRAPHICAL SCI & NATURAL RESOURCE RES CAS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510677353.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies make it difficult to accurately predict the spillover risk of the highly pathogenic avian influenza virus H5N1 on a global scale, especially when the dataset construction and the number and spatial distribution of case data are uneven. It is impossible to effectively draw a spatially continuous risk map, which affects the research and prevention and control of the epidemic pattern.

Method used

A method based on interpretable machine learning and spatially weighted sampling was used to integrate global multi-source and multi-temporal resolution data. Environmental feature raster layers were generated through time scale aggregation. A training set was constructed using spatially weighted random sampling to screen key driving factors. Feature importance analysis and partial dependence diagrams were combined to quantify threshold effects and interactions, generating a spatially continuous avian influenza risk distribution.

Benefits of technology

It has achieved refined spatial prediction of epidemic risks and interpretation of key factors at the global scale, significantly improved the scientificity and practicality of spatial risk assessment of zoonotic diseases, and accurately identified and quantified the multidimensional driving factors affecting avian influenza spillover.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120613152A_ABST
    Figure CN120613152A_ABST
Patent Text Reader

Abstract

The invention provides an avian influenza virus overflow space risk assessment method based on interpretable machine learning and space weighted sampling, and relates to the technical field of disease transmission prediction. According to the method, global multi-source heterogeneous data are integrated, annual scale environment feature grids are generated on a 25km grid through a spatial aggregation algorithm based on a geographic space big data cloud platform, samples are established by combining spatial weight resampling, an interpretable machine learning model is trained, key factors are screened by applying a recursive elimination method, and the method is applied to the cloud platform. And outputting spatial continuous global'wild bird-wild animal 'highly pathogenic avian influenza H5N1 overflow risk distribution through a partial dependency graph quantization threshold value and nonlinear influence. According to the method, scientificity and practicability of global pathogen overflow risk assessment and key mechanism analysis can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of disease transmission prediction, and in particular to a method for assessing the spatial risk of avian influenza virus spillover based on interpretable machine learning and spatial weighted sampling. Background Art

[0002] Currently, 60% of known human infectious diseases and 75% of emerging infectious diseases worldwide are zoonotic. Climate and land-use changes exacerbate the risk of zoonotic diseases. For example, driven by current climate and land-use changes, the highly pathogenic avian influenza virus H5N1 has demonstrated its ability to cross species barriers. Numerous cross-species spillovers of H5N1 from wild birds to wildlife have occurred globally, posing a serious threat to wildlife conservation and human health. There is an urgent need to characterize the drivers and risk patterns of cross-species spillovers of avian influenza using a One Health indicator system.

[0003] The spread of zoonotic pathogens is often influenced by multiple factors within the "One Health" ecosystem of humans, animals, and the environment, resulting in spatiotemporal clustering of zoonotic pathogen transmission. The current spillover of the highly pathogenic avian influenza virus H5N1 from wild birds to wildlife is influenced by factors such as the distribution and diversity of wild birds, the spatial distribution and diversity of wild mammals, climatic factors such as temperature, rainfall, humidity, and drought, and environmental factors such as farms and livestock farms that affect the interactions between birds and terrestrial mammals. Therefore, characterizing these factors and constructing a dataset is extremely challenging. Furthermore, building upon this dataset, developing a spatially continuous risk map based on the currently reported H5N1 cases in wildlife and their spatial distribution is crucial for future epidemic research and prevention.

[0004] Explainable machine learning refers to methods and techniques used to make machine learning models more transparent and understandable. Unlike traditional black-box models, it aims to provide insights into the model's decision-making process. The use of explainable machine learning methods can identify high-risk drivers that affect the occurrence and spread of zoonotic diseases. By establishing risk prediction models and simulating different scenarios, a detailed risk atlas can be drawn, high-risk points can be accurately predicted, risk factors can be identified, and ultimately the spread trends of zoonotic diseases in the context of climate change and land use can be clarified. However, the prediction of pathogen spillover risk requires the construction of an explainable machine learning model using currently reported case data as label data. This method is often affected by the number and spatial distribution (small number and spatial distribution offset) of existing case data. Summary of the Invention

[0005] In order to overcome the shortcomings of the existing technology, the purpose of the present invention is to provide a spatial risk assessment method for avian influenza virus spillover based on interpretable machine learning and spatially weighted sampling, which realizes the refined spatial prediction of epidemic risks and interpretation of key factors at the global scale, and significantly improves the scientificity and practicality of spatial risk assessment of zoonotic diseases.

[0006] To achieve the above object, the present invention provides the following solutions:

[0007] A spatial risk assessment method for avian influenza virus spillover based on interpretable machine learning and spatially weighted sampling, including:

[0008] Based on the geospatial big data cloud computing platform, we integrate multi-source and multi-temporal resolution data from around the world covering the research period to obtain the original data.

[0009] The raw data is temporally synthesized by a time scale aggregation algorithm to generate an annual scale environmental feature raster layer;

[0010] Using a 25km grid as the spatial unit, perform spatial aggregation on the annual scale environmental characteristic raster layer to generate a spatial distribution map of environmental characteristics for each driving factor;

[0011] Train the machine learning model based on a training set constructed by spatially weighted random sampling;

[0012] Based on the trained machine learning model, we use the feature importance analysis method to calculate the contribution of each environmental feature of each data point to the prediction result. Then, we use the recursive feature elimination method to gradually eliminate the features with the lowest average contribution and screen out the set of key driving factors.

[0013] Partial dependence plots were used to analyze the nonlinear impact trend of the key driver set on risk prediction to quantify threshold effects and interactions;

[0014] The optimized machine learning model is applied to global environmental characteristic raster data, outputting the predicted probability value for each raster cell and generating a spatially continuous H5N1 avian influenza risk distribution.

[0015] Preferably, the machine learning model is trained based on a training set constructed by spatially weighted random sampling, comprising:

[0016] Cover the entire world with a 25km*25km grid, and set the center point of each grid as the data point;

[0017] Based on the literature database, we searched for the number of academic papers related to avian influenza in various countries to calculate the attention and reporting capabilities of various countries to avian influenza events, and assigned the reporting capabilities as spatial weights to the data points of the corresponding countries.

[0018] A stratified 10-fold cross-validation method was used to divide the data points of the global regions into a training set and a test set. For the training set, data points with and without epidemic records were randomly sampled using the reporting capacity as the weight to generate a balanced positive and negative sample data set.

[0019] A machine learning model is used to train the balanced data set, and the hyperparameters of the machine learning model are optimized using a grid search method; the input features of the machine learning model are environmental features; and the output label of the machine learning model is a binary label indicating whether an epidemic exists or not.

[0020] Preferably, the original data includes: environmental data, animal distribution data, human activity data and H5N1 avian influenza outbreak site records.

[0021] Preferably, the environmental data include vegetation index, rainfall, average temperature, maximum temperature, minimum temperature, maximum maximum temperature, minimum minimum temperature, average wind speed, Palmer Drought Severity Index, altitude, water surface coverage, wetland coverage, bare area coverage, forest coverage, permanent snow and ice cover, shrub cover and tundra cover.

[0022] Preferably, the animal distribution data includes wild bird diversity and wild animal diversity data.

[0023] Preferably, the human activity data include spatial distribution of the number of farms worldwide, poultry density, and livestock density data.

[0024] Preferably, after applying the optimized machine learning model to the global environmental characteristic grid data, outputting the predicted probability value of each grid cell, and generating a spatially continuous H5N1 avian influenza risk distribution, the method further includes:

[0025] The AUC and TSS indicators of the test set were calculated through 10-fold cross validation to evaluate the generalization ability of the optimized machine learning model in unknown areas.

[0026] Preferably, the machine learning model is any one of an XGBoost model and a Random Forest model.

[0027] Preferably, the feature importance analysis method is any one of a permutation feature importance analysis method and a Shapley additivity feature interpretation method.

[0028] Preferably, the spatial distance is not less than 25 km.

[0029] The present invention discloses the following technical effects:

[0030] The present invention provides a spatial risk assessment method for avian influenza virus spillover based on interpretable machine learning and spatial weighted sampling, comprising: integrating global multi-source heterogeneous data based on a geospatial big data cloud computing platform to obtain raw data of a "One Health" indicator system that affects the spread of H5N1 virus from wild birds to wild mammals; synthesizing the raw data on an annual scale using a time scale aggregation algorithm to generate an annual scale environmental characteristic raster layer; performing spatial aggregation operations on the annual scale environmental characteristic raster layer based on a 25km grid to generate an environmental characteristic spatial map of each driving factor; and performing a training set constructed by spatial weighted random sampling to generate an environmental characteristic spatial map of each driving factor. The machine learning model is trained on the training set; based on the trained machine learning model, the contribution of each environmental feature of each data point to the prediction result is calculated using the feature importance analysis method, and the features with the lowest average contribution are gradually eliminated through the recursive feature elimination method to screen out a set of key driving factors; the nonlinear influence trend of the key driving factor set on risk prediction is analyzed using the partial dependence graph to quantify the threshold effect and interaction; the optimized machine learning model is applied to the global environmental feature grid data, and the predicted probability value of each grid unit is output to generate a spatially continuous "wild bird-wild animal" H5N1 virus spillover risk distribution map. By combining interpretable machine learning with spatial weighted sampling, the present invention can accurately identify and quantify the multidimensional driving factors that affect the "wild bird-mammal" H5N1 avian influenza spillover, realize the refined spatial prediction of epidemic risk and the interpretation of key factors at the global scale, and significantly improve the scientificity and practicality of the spatial risk assessment of zoonotic diseases. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0032] Figure 1 A flow chart of a method provided by an embodiment of the present invention;

[0033] Figure 2 A schematic diagram of a technical route provided by an embodiment of the present invention;

[0034] Figure 3 A schematic diagram showing the importance ranking of factors affecting the "wild bird-to-mammal" H5N1 spillover based on SHAP analysis of animal, environmental, and human factors provided by an embodiment of the present invention;

[0035] Figure 4A schematic diagram illustrating the mechanism of the impact of animal, environmental, and human factors on the "wild bird-to-mammal" H5N1 spillover based on PDPs provided in an embodiment of the present invention;

[0036] Figure 5 Schematic diagram of spatial risk prediction for H5N1 spillover from wild birds to mammals provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0037] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0038] The purpose of the present invention is to provide a spatial risk assessment method for avian influenza virus spillover based on interpretable machine learning and spatially weighted sampling, which can achieve refined spatial prediction of epidemic risks and interpretation of key factors at a global scale, significantly improving the scientificity and practicality of spatial risk assessment of zoonotic diseases.

[0039] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0040] Figure 1 A flow chart of the method provided in the embodiment of the present invention is shown in FIG. Figure 1 As shown, the present invention provides a spatial risk assessment method for avian influenza virus spillover based on interpretable machine learning and spatial weighted sampling, comprising:

[0041] Step 100: Integrate multi-source and multi-temporal resolution data from around the world covering the study period based on the geospatial big data cloud computing platform to obtain raw data;

[0042] Step 200: Temporally synthesize the original data using a time scale aggregation algorithm to generate an annual scale environmental feature raster layer;

[0043] Step 300: Using 25km grid as the spatial unit, perform spatial aggregation on the annual scale environmental characteristic raster layer to generate a spatial distribution map of environmental characteristics of each driving factor;

[0044] Step 400: Training a machine learning model based on a training set constructed by spatially weighted random sampling;

[0045] Step 500: Based on the trained machine learning model, use the feature importance analysis method to calculate the contribution of each environmental feature of each data point to the prediction result, and use the recursive feature elimination method to gradually eliminate the features with the lowest average contribution to screen out the set of key driving factors;

[0046] Step 600: Analyze the nonlinear impact trend of the key driver set on risk prediction using a partial dependence diagram to quantify threshold effects and interactions;

[0047] Step 700: Apply the optimized machine learning model to the global environmental characteristic grid data, output the predicted probability value of each grid cell, and generate a spatially continuous H5N1 avian influenza risk distribution.

[0048] Preferably, step 400 of this embodiment specifically includes:

[0049] Divide the world's countries into multiple regions according to administrative divisions, and generate evenly distributed data points with preset spatial spacing within each region;

[0050] Based on the literature database, we searched for the number of academic papers related to avian influenza in various countries to calculate the reporting capacity of each country on avian influenza events, and assigned the reporting capacity as a spatial weight to the data points of the corresponding country;

[0051] A stratified 10-fold cross-validation method was used to divide the data points of the global regions into a training set and a test set. For the training set, data points with and without epidemic records were randomly sampled using the reporting capacity as the weight to generate a balanced positive and negative sample data set.

[0052] A machine learning model is used to train the balanced data set, and the hyperparameters of the machine learning model are optimized using a grid search method; the input features of the machine learning model are environmental features; and the output label of the machine learning model is a binary label indicating whether an epidemic exists or not.

[0053] In an embodiment of the present invention, a method for analyzing infectious disease driving factors based on machine learning feature optimization is provided, comprising the steps of:

[0054] Step S1: Determine the time range of the research target and obtain the start time and end time of the research time period.

[0055] Step S2: Use Google Earth Engine (GEE) to obtain raster images of environmental factors (environmental characteristics, such as those shown in Table 1) such as average temperature, total rainfall, and vegetation index within the timeframe of the study. Based on the timeframe, perform data synthesis on an annual scale using the algorithm for these driving factors.

[0056] Step S3: Spatially aggregate the raster data synthesized using the above method within a 25km grid at a scale of 25km to calculate the features for each year and form a feature combination for each year. Combined with the spatial distribution data (labels) of wild bird-wildlife overflow cases for that year, an interpretable machine learning model is constructed.

[0057] Table 1 Environmental characteristics and data sources used in this patent technology

[0058]

[0059] In an embodiment of the present invention, steps for resampling infectious disease occurrence records are provided, specifically as follows:

[0060] Step S5: Divide the countries with larger areas, including Brazil, China, the United States, Russia, Australia, and Canada, into provinces, states, and other small regions; generate a 25 km*25 km grid globally to cover all countries / regions, extract the center point of each grid as a data point, and assign driving factor values ​​to all data points.

[0061] Step S6: Mark the data points in countries / regions with records of mammals infected with H5N1 as labeled points, and the remaining data points are unlabeled points.

[0062] Step S7: Use the search formula "((animal[Title / Abstract])AND(highlypathogenic avian influenza[Title / Abstract]))OR((animal[Title / Abstract])AND(H5N1[Title / Abstract]))OR((mammal[Title / Abstract])AND(H5N1[Title / Abstract]))OR((mammal[Title / Abstract])AND(highly pathogenic avian influenza[Title / Abstract]))" on PubMed to obtain a list of all relevant academic journal articles. Download the full name of the first author's affiliation and the country / region to which the first author's affiliation belongs for each article. Use the Python geotext library to identify the first affiliation of the i-th paper belonging to the c-th paper. i countries. For the jth country, the number of papers whose first unit is located in this country is:

[0063]

[0064] Where δ(x) is the Dirac Delta function, which is 1 when x=0 and 0 otherwise. Based on this, the reporting effort of each country is calculated by normalization:

[0065]

[0066] The above r j Assigned as value to all data points within that country.

[0067] Step S8: (Nested 10-fold begins here) Use the Stratified 10-fold method to divide the countries / regions into ten subsets. "Stratified" means that the ratio of labeled countries / regions to unlabeled countries / regions in each subset is close to the ratio of labeled countries / regions to unlabeled countries / regions globally. At the same time, if a subset contains only labeled countries / regions or only unlabeled countries / regions, the ten subsets are re-divided. For the above 10 subsets, when each subset is used as a test set, the remaining subsets are used together as training sets, thus forming 10 training / test set country / region combinations.

[0068] Step S9: For each training / test set region combination, 100 positive and negative sample data points are collected in the training set countries / regions and the test set countries / regions respectively. The rules for collecting positive and negative samples are as follows: when collecting positive samples, random sampling with replacement is performed among the labeled data points with (reporting effort) as the weight; when collecting negative samples, random sampling with replacement is performed among the unlabeled data points with (reporting effort) as the weight, and the number of samples is the same as that of positive samples. Thus, 10 groups of 100 training / test set data pairs are generated for the 10 training / test set region combinations. In each data pair, the training set data points and the test set data points belong to different countries / regions (that is, the training set data points and the test set data points will not appear in the same country / region at the same time), but the 100 test set data points belong to the same group of countries / regions, and the 100 training sets belong to the same group of countries / regions.

[0069] In an embodiment of the present invention, specific steps for constructing an interpretable machine learning model, optimizing features, analyzing driving mechanisms, and mapping risks are provided:

[0070] Step S10: (This is the inner layer repeated sampling cross-validation) In a set of 100 training / test set data pairs obtained in step S9, for each data pair, randomly select 30% of the data points from the training set data points as the validation set, use the random forest model to fit the remaining 70% of the training set data points, calculate the accuracy (including AUC, TSS) on the validation set, calculate the accuracy (including AUC, TSS) on the test set, calculate the SHAP value of each feature on each data point in the training set, and take the average of the absolute values ​​to obtain the SHAP feature importance of each feature. Finally, the accuracy on the 100 validation sets, the accuracy on the 100 test sets, and the feature importance on the 100 training sets are calculated.

[0071] Step S11: Based on the average AUC on the 100 validation sets, use the grid search method to adjust the hyperparameters of the random forest model. Repeat step S10 to obtain the random forest hyperparameter combination with the optimal average AUC on the validation set and the corresponding 100 random forest models. Use these models to calculate the accuracy of the 100 test sets and obtain the 100 feature importances.

[0072] Step S12: (Here is the outer 10-fold) Repeat steps S10 and S11 for the 10 sets of training / test set region combinations obtained in steps S8-S9 to obtain the average accuracy on 1000 validation sets, the average accuracy on 1000 test sets, 1000 feature importance values, and 1000 random forest models.

[0073] Step S13: Calculate the average feature importance obtained in step S12, remove the least important feature, and return to step S8 until only the last feature remains. This provides the order in which features are removed, the change in validation set accuracy during the removal process, the change in test set accuracy, and the feature importance of each feature combination.

[0074] Step S14: When step S12 is completed for the first time, all features are used, 1000 random forest models and the training sets used by each model are obtained, and these models are used to analyze the driving mechanism: for a certain feature, the numerical range of the feature is counted, and 30 values ​​are evenly taken from it. For this feature, for each random forest model and the training set used by it, all the values ​​of the feature in the training set are set to these 30 values, and the random forest model is used for prediction and the average value of the predicted probability is calculated, so that for the random forest model, a broken line with 30 points is obtained with the feature value and the average predicted probability as the horizontal and vertical coordinates respectively. Repeat the above process for 1000 random forest models and the training sets used by them to obtain 1000 broken lines, calculate the average value and confidence interval, and obtain the driving mechanism of the feature. Repeat the above process for each feature to obtain the driving mechanism of each feature.

[0075] For example, other machine learning models for tabular data, such as XGBoost, can replace the random forest (RF) model; permutation feature importance analysis can replace SHAP feature importance analysis; in addition to being used for feature importance analysis, SHAP values ​​can also be used to replace PDP analysis.

[0076] Step S15: Continue using the 1000 random forest binary classification models used in step S14, and use the raster image set of step S3 to output the predicted probability for each raster on the global land to form a risk image of zoonotic diseases.

[0077] Specifically, given that the occurrence records of the H5N1 avian influenza virus in mammals are unevenly distributed spatially, the information on the reporting locations of some infection records is not specific, and the availability of various indicators in the "One Health" dataset at a global scale, this embodiment spatially quantifies the reporting capabilities of different countries for such infection events based on literature big data, uses spatial weights as weighted sampling for samples, and resamples event records of the H5N1 avian influenza epidemic in mammals to form training and test sets.

[0078] Furthermore, this embodiment reveals the importance of human, environmental and animal factors by calculating the SHAP values ​​of 20 driving factors in the zoonotic disease risk prediction model. The three categories of features have obvious differences in their contributions to the model. Among the human-related driving factors, population density plays the most important role, and population density has a greater impact on the occurrence of zoonotic diseases. Among the environmental factors, the SHAP values ​​of factors such as average wind speed, minimum monthly minimum temperature, average annual precipitation, forest coverage and elevation are relatively high, indicating that climatic conditions and geographical characteristics play a key role in the spread of zoonotic diseases, especially climatic factors, which may play a key role in the spread of zoonotic diseases. Among the animal factors, the species diversity of mammals contributes the most to the model, followed by the number of livestock and poultry, further highlighting the impact of animal species and numbers on zoonotic diseases.

[0079] like Figure 2As shown, this embodiment couples the random forest model (RF), Shapley Additive Explanation (SHAP), partial dependency plots (PDPs), and nested cross-validation (NCV) to form a method for identifying and analyzing avian influenza driving factors and assessing spatial occurrence risk based on the "geospatial big data + cloud computing + artificial intelligence" scientific research paradigm. The performance of the random forest is evaluated by calculating the area under the receiver-operator characteristic curve (AUC) and the true skill statistics (TSS). In terms of analyzing the driving mechanism of driving factors, this embodiment uses the PDPs method to visualize the relationship between driving factor variables and occurrence risk, and combines the ecological and epidemiological knowledge related to avian influenza to analyze the relationship between the two. Given that the occurrence records of H5N1 avian influenza viruses in mammals are unevenly distributed spatially, the reporting locations of some infection records are not specific, and the availability of various indicators in the "One Health" dataset at a global scale, this example spatially quantifies the reporting capabilities of different countries for such infection events based on literature big data, uses spatial weights to perform weighted sampling, and resamples the event records of H5N1 avian influenza outbreaks in mammals to form training and test sets. The above method can screen out important factors affecting the spillover of H5N1 from "wild birds to mammals" from the potential factors listed in the "One Health" dataset ( Figure 3 and Figure 4 ) and predicted the probability of "wild bird-mammal spillover" events on a global scale ( Figure 5 ).

[0080] Depend on Figure 3 As can be seen, the predicted value increases significantly with increasing population density and average wind speed. Increases in the minimum monthly temperature and average annual precipitation cause fluctuations in the predicted risk value, initially increasing, then decreasing, and finally stabilizing. With increasing mammalian species diversity, the predicted risk initially increases and then stabilizes. Increases in altitude decrease the predicted risk value, which then stabilizes after reaching a certain value. Changes in the remaining factors have little impact on the predicted risk value, which remains stable overall.

[0081] Depend on Figure 4 It can be seen that the areas with higher risks are mainly concentrated in North America (central and eastern United States and southeastern Canada), Europe (France, Germany, the Netherlands, the United Kingdom, etc.), some countries or regions in East Asia and Central Asia, and some countries or regions in South America.

[0082] Figure 5 The risk data ranges from 0 to 1, with redder colors indicating a higher risk of wild bird-to-mammal H5N1 spillover and greener colors indicating a lower risk of wild bird-to-mammal H5N1 spillover.

[0083] The beneficial effects of the present invention are as follows:

[0084] This paper constructs a "One Health" system database of typical zoonotic pathogen spillover risks based on geospatial big data cloud computing. Furthermore, based on literature big data, the paper spatially quantifies the reporting capacity or level of attention paid to such infection events by different countries or regions. This data is used as spatial weights to weight samples, and event records of H5N1 avian influenza outbreaks in mammals are resampled to form training and test sets. An interpretable machine learning model is constructed and optimized to more realistically analyze the impact of different environmental, animal, and human factors on H5N1 spillover from "wild birds and wildlife," thereby enabling a spatial assessment of global risks.

[0085] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referenced to each other.

[0086] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The above examples are only intended to help understand the method and core concept of the present invention. At the same time, those skilled in the art will find that the specific implementation methods and application scopes may vary based on the concept of the present invention. In summary, the contents of this specification should not be construed as limiting the present invention.

Claims

1. A spatial risk assessment method for avian influenza virus spillover based on interpretable machine learning and spatially weighted sampling, characterized in that: include: Based on the geospatial big data cloud computing platform, we integrate multi-source and multi-temporal resolution data from around the world covering the research period to obtain the original data. The raw data is temporally synthesized by a time scale aggregation algorithm to generate an annual scale environmental feature raster layer; Using a 25km grid as the spatial unit, perform spatial aggregation on the annual scale environmental characteristic raster layer to generate a spatial distribution map of environmental characteristics for each driving factor; Train the machine learning model based on a training set constructed by spatially weighted random sampling; Based on the trained machine learning model, we use the feature importance analysis method to calculate the contribution of each environmental feature of each data point to the prediction result. Then, we use the recursive feature elimination method to gradually eliminate the features with the lowest average contribution and screen out the set of key driving factors. Partial dependence plots were used to analyze the nonlinear impact trend of the key driver set on risk prediction to quantify threshold effects and interactions; The optimized machine learning model is applied to global environmental characteristic raster data, outputting the predicted probability value for each raster cell and generating a spatially continuous H5N1 avian influenza risk distribution.

2. The method for spatial risk assessment of avian influenza virus spillover based on interpretable machine learning and spatial weighted sampling according to claim 1, characterized in that: The machine learning model is trained on a training set constructed by spatially weighted random sampling, including: Cover the entire world with a 25km*25km grid, and set the center point of each grid as the data point; Based on the literature database, we searched for the number of academic papers related to avian influenza in various countries to calculate the attention and reporting capabilities of various countries to avian influenza events, and assigned the reporting capabilities as spatial weights to the data points of the corresponding countries. A stratified 10-fold cross-validation method was used to divide the data points of the global regions into a training set and a test set. For the training set, data points with and without epidemic records were randomly sampled using the reporting capacity as the weight to generate a balanced positive and negative sample data set. A machine learning model is used to train the balanced data set, and the hyperparameters of the machine learning model are optimized using a grid search method; the input features of the machine learning model are environmental features; and the output label of the machine learning model is a binary label indicating whether an epidemic exists or not.

3. The method for spatial risk assessment of avian influenza virus spillover based on interpretable machine learning and spatial weighted sampling according to claim 1, characterized in that: The original data includes: environmental data, animal distribution data, human activity data and H5N1 avian influenza outbreak site records.

4. The method for spatial risk assessment of avian influenza virus spillover based on interpretable machine learning and spatial weighted sampling according to claim 3, characterized in that: The environmental data include vegetation index, rainfall, average temperature, maximum temperature, minimum temperature, maximum maximum temperature, minimum minimum temperature, average wind speed, Palmer Drought Severity Index, altitude, water surface cover, wetland cover, bare area cover, forest cover, permanent snow and ice cover, shrub cover and tundra cover.

5. The method for spatial risk assessment of avian influenza virus spillover based on interpretable machine learning and spatial weighted sampling according to claim 3, characterized in that: The animal distribution data includes wild bird diversity and wild animal diversity data.

6. The method for spatial risk assessment of avian influenza virus spillover based on interpretable machine learning and spatial weighted sampling according to claim 3, characterized in that: The human activity data include the spatial distribution of the number of farms worldwide, poultry density, and livestock density data.

7. The method for spatial risk assessment of avian influenza virus spillover based on interpretable machine learning and spatial weighted sampling according to claim 2, characterized in that: After applying the optimized machine learning model to global environmental characteristic raster data and outputting the predicted probability value for each raster cell to generate a spatially continuous H5N1 avian influenza risk distribution, it also includes: The AUC and TSS indicators of the test set were calculated through 10-fold cross validation to evaluate the generalization ability of the optimized machine learning model in unknown areas.

8. The method for spatial risk assessment of avian influenza virus spillover based on interpretable machine learning and spatial weighted sampling according to claim 1, characterized in that: The machine learning model is any one of an XGBoost model and a Random Forest model.

9. The method for spatial risk assessment of avian influenza virus spillover based on interpretable machine learning and spatial weighted sampling according to claim 1, characterized in that: The feature importance analysis method is any one of a permutation feature importance analysis method and a Shapley additivity feature interpretation method.

10. The method for spatial risk assessment of avian influenza virus spillover based on interpretable machine learning and spatial weighted sampling according to claim 2, characterized in that: The spatial distance is not less than 25km.

Citation Information

Patent Citations

  • Power failure processing method, device and equipment for electric power guarantee area and medium

    CN118504991A

  • Diabetes prediction and interpretability analysis method and computer program product

    CN119418954A