A method and device for evaluating the risk of spillover of pathogen epidemic ecology in animals
By constructing a pathogen spillover risk model and using step-by-step method and machine learning methods for automated evaluation, the problems of inefficient spillover risk assessment in the existing technology are solved, and the effect of rapid, accurate and comprehensive evaluation is achieved.
Patent Information
- Application Number
- CN202410925829.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-11
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-07-11
AI Technical Summary
Existing pathogen spillover risk assessment tools are inefficient, difficult to achieve early warning and high throughput assessment, and relying on experts to discuss leads to cumbersome processes and high evaluation costs, making it difficult to fully evaluate the risks of all known strains.
By obtaining pathogen index data, a step-by-step method and machine learning method are used to build a spillover risk model to realize automated assessment of pathogen spillover risk and get rid of the cumbersome process of expert consultation.
It achieves rapid and accurate assessment of pathogen spillover risk, improves the comprehensiveness and efficiency of assessment, and reduces assessment costs.
Smart Images

Figure CN118800476B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of pathogen spillover risk assessment, and in particular to a method and device for assessing the spillover risk of pathogen epidemic ecology in animals. Background Art
[0002] Spillover refers to the transmission of pathogens from animal hosts to humans, and spillover risk assessment has important public health significance. In 2020, the second edition of the Tool for Influenza Pandemic Risk Assessment (TIPRA) expanded the scope of assessment from viruses with human cases to animal influenza viruses with important public health significance. However, the criteria for determining important public health significance were not clearly defined, and most assessment reports were still for viruses with human cases. In addition, the assessment method still adopts the form of expert consultation and is carried out in the form of monthly risk assessment. Before the meeting, relevant information covering public health, animal health, virology and other aspects needs to be prepared. Although the reliability of the results is guaranteed, it is difficult to meet the timeliness requirements. Among them, the influenza risk assessment tool based on expert consultation, for example, the Influenza Risk Assessment Tool (IRAT), only evaluated 24 strains of viruses from 2011 to 2023. It can be seen that the current pandemic assessment tools are inefficient and difficult to achieve early warning and high-throughput assessment.
[0003] In addition, the application of the TIPRA tool requires the preparation of three types of materials, including virus characteristics, disease characteristics, and the ecology and prevalence of the virus in animals. The first two types require certain biochemical experiments and population surveys, which are difficult to collect in a short period of time. The third type includes indicators such as rapid geographical spread, increase in the number of host species, and spread in mammals, especially pigs, which need to be summarized and analyzed by expert consultation through gene sequences. Expert consultation only assesses the risk of a single strain. Due to the high cost of assessment, it is difficult to assess all known strains. It is easy to ignore some risk strains by only assessing strains that some experts believe are at higher risk. Summary of the invention
[0004] The purpose of this application is to provide a method and equipment for assessing the spillover risk of pathogen epidemic ecology in animals, which can get rid of the cumbersome process of expert consultation and realize the automated assessment of pathogen spillover risk, thereby improving the comprehensiveness of pathogen spillover risk assessment while providing assessment efficiency.
[0005] To achieve the above objectives, this application provides the following solutions:
[0006] In a first aspect, the present application provides a method for assessing the spillover risk of pathogen epidemic ecology in animals, comprising:
[0007] Obtain pathogen indicator data; pathogen indicators include pathogen spatial distribution indicators, infected host width indicators and epidemic intensity indicators;
[0008] Based on the pathogen indicator data, a first spillover risk model was formed by induction using a stepwise approach;
[0009] Based on the pathogen indicator data, a re-spillover risk model is constructed using a machine learning method;
[0010] Determine the indicator data of the pathogen to be evaluated in each subtype year; the subtype year is used to indicate the spillover status of different pathogens in different years, or to indicate the spillover status of different subtypes or lineages in different years;
[0011] Determining the type of pathogen to be evaluated, and selecting a spillover risk assessment model based on the type of pathogen to be evaluated; the spillover risk assessment model is a first spillover risk model or a second spillover risk model;
[0012] The selected spillover risk assessment model is used to complete the spillover risk assessment based on the indicator data of the pathogen to be assessed in each subtype year, and the early warning situation of each subtype year is obtained.
[0013] Optionally, obtain pathogen indicator data, including:
[0014] Obtain pathogen entry data according to the pathogen name, and generate an Excel file based on the pathogen entry data;
[0015] Standardizing the pathogen entry data in the Excel file to obtain a data set;
[0016] Based on the data set, n indicators of each pathogen subtype year are obtained to generate the pathogen indicator data.
[0017] Optionally, pathogen entry data is obtained according to the pathogen name, specifically including:
[0018] In the Protein database of the GenBank website, keyword searches were performed according to the pathogen names to obtain sequence data;
[0019] Regular expressions are used to extract sequence information of each sequence from the sequence data; the sequence information includes: version number, strain name, location, time, host, subtype, segment and original sequence;
[0020] The sequence information is preprocessed to obtain pathogen entry data.
[0021] Optionally, the pathogen entry data in the Excel file is standardized to obtain a data set, specifically including:
[0022] Taking the comparison table as the standard, matching based on the Excel file forms a host field; the host field includes a host type and a host species;
[0023] Using the comparison table as a standard, a host type field is formed based on the Excel file matching; the host type field includes a "whether it is a host in close contact with humans" field and a "whether it is a mammalian host" field;
[0024] Based on the comparison table, a location field is formed based on the Excel file matching; the location field includes a continent, a country, and a province;
[0025] Based on the comparison table, the year field is formed based on the Excel file matching;
[0026] The data set is generated based on the host field, the host type field, the location field, and the year field.
[0027] Optionally, based on the pathogen indicator data, a first spillover risk model is formed by adopting a stepwise method, specifically including:
[0028] According to the first spillover event of each subtype of pathogen occurring in the set time period, the subtype-years meeting the inclusion criteria in the set time period are divided into new-onset group and non-new-onset group;
[0029] Based on the pathogen index data, obtaining the pathogen index for each subtype year, and determining the minimum value of each pathogen index in the new outbreak group;
[0030] Taking the minimum value of each pathogen index in the new outbreak group as the standard, the subtype-years meeting the set standard are screened in the non-new outbreak group and counted;
[0031] The pathogen indicators of each subtype year are screened out, and the pathogen indicators that change the number of subtype years that meet the set standards are retained;
[0032] The retained pathogen index and the corresponding minimum value of this pathogen index in the new outbreak group are used as the first spillover risk standard;
[0033] Determine the early warning status of different pathogen subtypes in different years based on the first spillover risk criteria to form external verification results;
[0034] Based on the external validation results, the first spillover risk standard corresponding to the minimum total number of subtypes per year is obtained, and this first spillover risk standard is used as the final first spillover risk standard;
[0035] The first spillover risk model is constructed based on the final first spillover risk standard.
[0036] Optionally, based on the pathogen indicator data, a re-spillover risk model is constructed using a machine learning method, specifically including:
[0037] Based on the first spillover events of each pathogen subtype occurring in the set time period and the number of annual animal infection cases of the pathogen, the subtype-years that meet the inclusion criteria in the set time period are divided into recurrence group and non-recurrence group;
[0038] Generate year labels for each subtype in the relapse group and year labels for each subtype in the non-relapse group to obtain data labels;
[0039] The pathogen indicator data is used as input, and the data label is used as output. Model training is carried out based on different machine learning methods until the set training conditions are met, and the trained model is used as the re-spillover risk model.
[0040] In a second aspect, the present application provides a computer device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the above-mentioned methods for assessing the spillover risk of pathogen epidemic ecology in animals.
[0041] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0042] The present application provides a method and device for assessing the spillover risk of pathogen epidemic ecology in animals. By acquiring pathogen indicator data, the infection distribution dynamics and epidemic situation of pathogens in animals can be obtained to quickly characterize the process ecology of pathogens in animals. Based on pathogen indicator data, a stepwise method is used to summarize and form a first spillover risk model, and a machine learning method is used to construct a second spillover risk model. Different spillover risk assessment methods can be provided for different pathogen types to achieve rapid and accurate assessment of pathogen risk, get rid of the cumbersome process of expert consultation, and realize the automated assessment of pathogen spillover risk. While providing assessment efficiency, it can improve the comprehensiveness of pathogen spillover risk assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0044] Figure 1This is an application environment diagram of a method for assessing the epidemic ecology of a pathogen in animals in one embodiment of the present application;
[0045] Figure 2 A schematic flow chart of a method for assessing the risk of a pathogen spreading in an animal according to an embodiment of the present application;
[0046] Figure 3 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0047] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0048] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0049] The method for assessing the risk of pathogen epidemic ecology in animals provided in the embodiments of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the index data of the pathogen to be evaluated in each subtype year to the server 104. After the server 104 receives the index data of the pathogen to be evaluated in each subtype year, for the index data of the pathogen to be evaluated in each subtype year, the server 104 determines the type of the pathogen to be evaluated, and based on the type of the pathogen to be evaluated, chooses whether to use the first overflow risk model as the overflow risk assessment model or the second overflow risk model as the overflow risk assessment model. Using the selected overflow risk assessment model, the overflow risk assessment is completed based on the index data of the pathogen to be evaluated in each subtype year, and the early warning situation of each subtype year is obtained. The server 104 can feedback the obtained overflow risk assessment results and the early warning situation of each subtype year to the terminal 102. In addition, in some embodiments, the spillover risk assessment method of the epidemic ecology of pathogens in animals can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly perform a spillover risk assessment on the indicator data of the pathogens to be assessed in each subtype year, or the server 104 can obtain the indicator number of the pathogens to be assessed in each subtype year from the data storage system, and perform a spillover risk assessment on the indicator number of the pathogens to be assessed in each subtype year.
[0050] The terminal 102 may be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers, or may be a cloud server.
[0051] In an exemplary embodiment, Figure 2 As shown, a method for assessing the risk of spillover of pathogen epidemic ecology in animals is provided. The method is executed by a computer device, and can be executed by a computer device such as a terminal or a server alone, or by a terminal and a server together. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in the example is used as an example to illustrate, including the following steps 200 to 205. Among them:
[0052] Step 200: Obtain pathogen index data. Pathogen indexes include pathogen spatial distribution index, infected host width index and epidemic intensity index.
[0053] Step 201: Based on the pathogen indicator data, a first spillover risk model is formed by adopting a stepwise method.
[0054] Step 202: Based on the pathogen indicator data, a re-spillover risk model is constructed using a machine learning method.
[0055] Step 203: Determine the indicator data of the pathogen to be evaluated in each subtype year. When evaluating the spillover risk of different pathogens, the subtype year can represent the spillover status of different pathogens in different years; when evaluating the spillover risk of different subtypes or lineages of a specific pathogen, the subtype year can represent the spillover status of different subtypes or lineages in different years.
[0056] Step 204: Determine the type of pathogen to be evaluated, and select a spillover risk assessment model based on the type of pathogen to be evaluated. The spillover risk assessment model is a first spillover risk model or a second spillover risk model.
[0057] Step 205: Using the selected spillover risk assessment model, the spillover risk assessment is completed based on the indicator data of the pathogen to be assessed in each subtype year, and the warning status of each subtype year is obtained.
[0058] The implementation of the above steps 200 to 205 gets rid of the cumbersome process of expert consultation, realizes the automated assessment of pathogen spillover risk, and can improve the comprehensiveness of pathogen spillover risk assessment while providing assessment efficiency. In addition, the present application can also obtain the infection distribution dynamics and prevalence of pathogens in animals by obtaining pathogen indicator data, so as to quickly characterize the process ecology of pathogens in animals.
[0059] In another exemplary embodiment of the present application, in order to avoid the problem of evaluating only strains that some experts believe to be at higher risk and ignoring some risky strains, the above step 200 is replaced by the following steps (1) to (3):
[0060] (1) Obtain pathogen entry data according to the pathogen name, and generate an Excel file based on the pathogen entry data.
[0061] Taking influenza A virus as an example, there are two ways to generate Excel files, one of which is:
[0062] (1.1) Search the Protein database on the GenBank website with "InfluenzaAvirus" as the keyword and download all returned entries, saving them as GenPept files. The gene sequences shared by GenBank and GISAID contain information such as subtypes, time, space, and hosts, providing data support for a detailed description of the historical epidemic situation of each subtype.
[0063] (1.2) The version number, strain name, location, time, host, subtype, segment, and original sequence of each sequence in the GenPept file were extracted using regular expressions. Irrelevant information such as influenza B, protein subchains, reassortant strains, and sequences collected before January 1, 1981 were deleted. Duplicates were removed based on the strain name, location, time, host, subtype, and other fields, and the files were saved as Excel files.
[0064] Another way to generate an Excel file is to retrieve all Type A records from January 1, 1981 to June 18, 2022 from the EpiFlu database on the GISAID website and download all returned entries and save them as an Excel file.
[0065] (2) The pathogen entry data in the Excel file was standardized to obtain a data set.
[0066] Taking influenza A virus as an example, the standardization process can be described as:
[0067] (2.1) The host field content of the Excel file is matched with the reference table to form a host field containing fields such as host type and host species.
[0068] (2.2) The host type field content of the Excel file is matched with the reference table as the standard to form a host type field containing fields such as "whether it is a host in close contact with humans" and "whether it is a mammalian host".
[0069] (2.3) The location field of the Excel file is matched with the comparison table to form the location field of the continent, country, province and other fields.
[0070] (2.4) Automatically identify and obtain the year information through the content of the time field in the Excel file to form a year field.
[0071] (2.5) Generate a data set based on the host field, host type field, location field and year field.
[0072] Among them, the comparison tables involved in steps (2.1) to (2.4) can be obtained by conventional methods. For example, the comparison table of locations in step (2.3) can be obtained by associating with a standard map attribute table or by querying the Internet.
[0073] Based on the above description, an influenza A virus discovery information dataset based on GenBank or GISAID gene sequences can be generated. The dataset contains nine fields: subtype, year, continent, country, province, host type, host species, whether it is a host in close contact with humans, and whether it is a mammalian host.
[0074] (3) Based on the data set, n indicators of each pathogen subtype year are obtained to generate pathogen indicator data.
[0075] For example, the calculation method of various indicators in the pathogen indicator data is explained by taking the subtype year H7N9-2013 as an example. Specifically:
[0076] (3.1) Specific time frame meaning.
[0077] The background is from 1981 to the year before, which in this case is from 1981 to 2012. The year 1 is 2012, the year before 2 is 2011 to 2012, and so on, the year before 10 is 2003 to 2012.
[0078] (3.2) The meaning of the indicators of known spatial distribution range.
[0079] Taking the provinces in the previous year as an example, all records of H7N9 (subtype) in 2012 (year) are obtained, and the provinces are counted after deleting duplicates. Other specific time ranges or spatial levels are similar to the above example.
[0080] (3.3) Changes in the spatial distribution range.
[0081] Taking the provinces in the previous year as an example, ① filter and obtain all records of H7N9 (subtype) in 2012 (year), and remove duplicate provinces to form province set 1. ② Filter and obtain all records of H7N9 (subtype) from 1981 to 2011 (year), and remove duplicate provinces to form province set 2. ③ Take the provinces in province set 1 but not in province set 2, and count them to obtain the diffusion changes of the spatial distribution range (spatial level is province) in the previous year.
[0082] (3.4) Known host breadth of infection.
[0083] Taking the host type of close contact with humans in the previous year as an example, all records of H7N9 (subtype) close contact with humans (host range) in 2012 (year) were obtained, and the host types were counted after duplication. Other specific time ranges, host levels or host ranges are similar to the above examples.
[0084] (3.5) Changes in host distribution range.
[0085] Taking the host types of hosts in close contact with humans in the previous year as an example, ① all records of H7N9 (subtype) in close contact with humans (host range) in 2012 (year) were screened and obtained, and after deduplication of host types, host type set 1 was formed. ② All records of H7N9 (subtype) in close contact with humans (host range) from 1981 to 2011 (year) were screened and obtained, and host types were deduplicated to form host type set 2. ③ The host types in host type set 1 but not in host type set 2 were taken, and the changes in the host distribution range (host level is host type, host range is hosts in close contact with humans) in the previous year were obtained by counting.
[0086] (3.6) Number of sequences.
[0087] Taking the hosts that had close contact with humans in the previous year as an example, all records of H7N9 (subtype) in close contact with humans (host range) in 2012 (year) were obtained and counted. Other specific time ranges or host ranges are similar to the above example.
[0088] (3.7) Frequency.
[0089] Taking the hosts that had close contact with humans in the past 10 years as an example, all records of H7N9 (subtype) and hosts that had close contact with humans (host range) from 2003 to 2012 (year) were obtained, and the years were counted after duplication. Other specific time ranges or host ranges are similar to the above example.
[0090] Based on the above description, 255 indicators can be obtained for each pathogen subtype year, including 63 pathogen spatial distribution indicators, 126 infected host width indicators, and 66 epidemic intensity indicators. Among them, pathogen spatial distribution indicators include the historical epidemic background of a specific subtype year and the known spatial distribution range in the previous 1 to 10 years (33 indicators at 3 spatial levels and 11 time ranges) and the diffusion changes of the spatial distribution range in the previous 1 to 10 years (30 indicators at 3 spatial levels and 10 time ranges). Infected host width indicators include the historical background of a specific subtype year and the known infected host width in the previous 1 to 10 years (66 indicators at 2 host levels, 3 host ranges and 11 time ranges) and the changes in the host distribution range in the previous 1 to 10 years (60 indicators at 2 host levels, 3 host ranges and 10 time ranges). Epidemic intensity indicators include the historical epidemic background of a specific subtype and the number of sequences reported in different hosts in the previous 1 to 10 years (33 indicators in 3 host ranges and 11 time ranges) and frequency (i.e. the number of years of occurrence. 33 indicators in 3 host ranges and 11 time ranges).
[0091] A subtype-year is a research unit used to analyze the spillover risk of a subtype (a specific pathogen subtype or lineage) in a specific year. The time range includes background (1981 to the previous year, a total of 1 time range) and the previous 1 to 10 years (the previous year, the previous 2 years...the previous 10 years, a total of 10 time ranges). The spatial level includes continents, countries, and provinces (sub-national administrative units, including provinces, states, etc.). The infection host width includes the infection host species (the "species" in the zoological classification "kingdom, phylum, class, order, family, genus, species") and host types (including humans, pigs, companion animals, horses, other mammals, wild birds, poultry, domesticated birds, and the environment. Companion animals include cats and dogs). The host range includes all hosts, hosts in close contact with humans (including humans, pigs, companion animals, horses, poultry, and domesticated birds), and mammalian hosts (including humans, pigs, companion animals, horses, and other mammals).
[0092] In another exemplary embodiment of the present application, the construction of the first spillover risk model of avian influenza virus subtypes (influenza A virus subtypes excluding H1N1, H3N2 and H1N2) is taken as an example to illustrate the process of forming the first spillover risk model in the above step 201. In this embodiment, based on the infection distribution dynamics and prevalence of pathogens in animals and the first spillover events that occurred between 2001 and 2022, the first spillover warning standard is formed by induction through a stepwise method, specifically:
[0093] Through literature research or data sharing by authoritative institutions, the first spillover events of various subtypes of avian influenza virus that occurred between 2001 and 2022 are summarized as follows, a total of 11 subtype years: H7N2-2002, H7N3-2003, H10N7-2004, H6N1-2013, H7N9-2013, H10N8-2013, H5N6-2014, H7N4-2017, H5N8-2020, H10N3-2021 and H3N8-2022.
[0094] (1) Based on the first spillover event of each pathogen subtype occurring within a set time period, the subtype-years that meet the inclusion criteria within the set time period are divided into new-emergence group and non-new-emergence group.
[0095] For example, based on the first spillover events of each subtype of avian influenza virus that occurred between 2001 and 2020, the subtype years that met the inclusion criteria from 2001 to 2020 were divided into the emerging group and the non-emerging group. The inclusion criteria were: (1) subtype years that had not yet spilled over were the non-emerging group, and (2) subtype years that had first spillover were the emerging group.
[0096] The first H7N9 spillover event occurred in 2013, so 12 subtype years including H7N9-2001, H7N9-2002...H7N9-2012 were classified as non-emerging groups. One subtype year including H7N9-2013 was classified as emerging group. Seven subtype years including H7N9-2014, H7N9-2015...H7N9-2020 did not meet the inclusion criteria and were excluded. H3N8 did not have its first spillover event between 2001 and 2020, so 20 subtype years including H3N8-2001, H3N8-2002...H3N8-2020 were classified as non-emerging groups. So far, two sets of data and corresponding labels are obtained.
[0097] (2) Based on the pathogen index data, the pathogen index for each subtype year was obtained (i.e., 255 indicators were calculated for each subtype year), and the minimum value of each pathogen index in the emerging group was determined.
[0098] (3) Using the minimum value of each pathogen indicator in the emerging group as the standard, the subtype-years that meet (greater than or equal to) the set standards are screened in the non-emerging group and counted.
[0099] (4) The pathogen indicators for each subtype year are screened and the pathogen indicators that change the number of subtype years that meet the set criteria are retained.
[0100] For example, remove one of the 255 indicators in turn, repeat steps (2) and (3), and make the following judgments: 1) Whether the number of subtype-years that meet the criteria in the non-newly diagnosed group remains unchanged before and after the removal of a certain indicator. If it does not change, it means that the indicator does not have the ability to predict the first overflow, and the indicator is permanently deleted. If the number of subtype-years that meet the criteria changes, the indicator is retained.
[0101] (5) The retained pathogen index and the corresponding minimum value of this pathogen index in the new outbreak group are used as the first spillover risk standard.
[0102] For example, calculate the number of indicators remaining after step (4). Remove one of the remaining indicators in turn, repeat steps (2) and (3), and make a judgment on whether to keep or remove the indicator (the judgment criteria are the same as those of step (4)). Calculate the number of remaining indicators and make the following judgment: If it is the same as the number of indicators obtained in step (4), return the current remaining indicators and their minimum value in the new group, that is, obtain the first overflow risk standard. If it is different from the number of indicators obtained in step (4), remove one of the remaining indicators, repeat steps (2) and (3), and calculate the number of remaining indicators after repetition. If it is the same as the number of indicators before repetition, return the current remaining indicators and their minimum value in the new group. Otherwise, continue to repeat.
[0103] (6) Determine the early warning status of different pathogen subtypes in different years based on the first spillover risk criteria to form external verification results.
[0104] For example, based on the first spillover risk standard obtained in step (5), the warning status (whether it meets the standard) of 22 subtype years such as H3N8-2001, H3N8-2002...H3N8-2022 and 21 subtype years such as H10N3-2001, H10N3-2002...H10N3-2021 is calculated to form an external verification result.
[0105] (7) Based on the external validation results, the first spillover risk standard corresponding to the minimum total number of subtypes per year is obtained, and this first spillover risk standard is used as the final first spillover risk standard.
[0106] For example, remove the relevant indicators of one time range in turn, remove the order (first 10 years, first 9 years...first 2 years), and complete steps (2) to (6) (the total number of indicators changes from 255 to 221, 197...39 in turn), and form 10 groups (coded B+9, B+8...B+1 and B+10 without removing any time range) of first spillover risk standards and their external verification results together with the results of the relevant indicators of any time range. According to the external verification results of H3N8-2022 and H10N3-2021, select the first spillover risk standards that meet the requirements (the requirement is that 2 subtype years such as H3N8-2022 and H10N3-2021 must meet the warning standards), and eliminate the first spillover risk standards that do not meet the requirements (only 3 groups such as B+1, B+2 and B+3 remain). Calculate the total number of warning subtypes per year in the remaining groups of external validation results, and take the first spillover risk standard corresponding to the group with the least total number of warning subtypes per year as the final standard (B+3 has the least total number of warning subtypes per year).
[0107] (8) Based on the final first spillover risk standard, a first spillover risk model is constructed.
[0108] In another exemplary embodiment of the present application, still taking the avian influenza virus subtype as an example, based on the infection distribution dynamics and prevalence of pathogens in animals, and the re-spillover events that occurred between 2001 and 2022, a re-spillover risk model is trained by machine learning methods. Among them, the annual number of human infection cases of each subtype of avian influenza virus that occurred between 2001 and 2022 is summarized and summarized by literature research or data sharing by authoritative institutions as shown in Table 1. Based on this, the above step 202 of the present application can be replaced by the following steps (1) to (3). Among them:
[0109] (1) Based on the first spillover events of each pathogen subtype occurring in a set time period and the number of annual animal infection cases of the pathogen, the subtype-years that meet the inclusion criteria in the set time period are divided into a recurrence group and a non-recurrence group.
[0110] For example, based on the first spillover events of each subtype of avian influenza virus that occurred between 2001 and 2020 and the annual number of human infection cases, the subtype years that met the inclusion criteria from 2001 to 2020 were divided into recurring group and non-recurring group.
[0111] The inclusion criteria in step 1 are as follows: (1) A subtype has spilled over for the first time before a specific year and spilled over again in a specific year, i.e., human infection cases occurred. This subtype-year is classified as the recurrence group (a total of 53 subtype-years). (2) A subtype has spilled over for the first time before a specific year, but no human infection cases occurred in a specific year. This subtype-year is classified as the non-recurrence group (a total of 88 subtype-years).
[0112] If the first H7N9 spillover event occurred in 2013 and there were human cases from 2014 to 2019, then the six subtype years, H7N9-2014, H7N9-2015, ... H7N9-2019, were classified as the recurring group. One subtype year, H7N9-2020, was classified as the non-recurring group. Thirteen subtype years, H7N9-2001, H7N9-2002, ... H7N9-2013, did not meet the inclusion criteria and were excluded. If the first H5N1 spillover event occurred before 2001 and there were human cases from 2003 to 2017, 2019, and 2020, then three subtype years, H5N1-2001, H5N1-2002, and H5N1-2018, were classified as the non-recurring group. Fifteen subtype-years, including H5N1-2003, H5N1-2004…H5N1-2017, and two subtype-years, including H5N1-2019 and H5N1-2020, are classified as the recurrent group.
[0113] The infection distribution dynamics and prevalence of the pathogen in animals for each subtype year were further calculated, that is, 255 indicators for each subtype year were calculated.
[0114] (2) Generate year labels for each subtype in the recurrence group and year labels for each subtype in the non-recurrence group to obtain data labels. For example, the year labels for each subtype in the recurrence group are marked as P, and the year labels for each subtype in the non-recurrence group are marked as N.
[0115] (3) Using pathogen indicator data as input and data labels as output, conduct model training based on different machine learning methods until the set training conditions are met, and use the trained model as the re-spillover risk model. For example, use 255 indicators (independent variables) to conduct model training based on data labels (dependent variables. The subtype-year labels of the recurrence group are marked as P, and the subtype-year labels of the non-recurrence group are marked as N). The specific training process can be:
[0116] (3.1): Optimal parameter selection for machine learning models.
[0117] There are four machine learning methods, namely boosted regression trees (BRT), extreme gradient boosting trees (XGB), random forest (RF) and support vector machines (SVM). The machine learning methods are parallel, and any one of them can be used, or multiple methods can be used to compare the results (select the best method based on the accuracy of the external verification results), and other machine learning methods can also be used. In this solution, the best machine learning method is selected based on the comparison of the results.
[0118] (3.11) Establish a modeling data set by random sampling: According to the grouping in step (1) of this example, there are 53 positive data and 88 negative data. The modeling data set is established according to the positive-negative ratio of 1:1, that is, 53 items are randomly selected from the negative data and combined with all the positive data to form a modeling data set with a total of 106 items.
[0119] (3.12) The modeling data set is split into a training set and a test set by stratified sampling. The sample size ratio of the training set to the test set is 4:1, that is, the test set contains 53*0.2=10.6≈10 negative (positive) data, a total of 20 (negative + positive) data. The training set contains 106-20=86 data.
[0120] (3.13) The training set is used for model training. During the training process, 10 fold cross validation resampling is repeated 10 times: the training set is divided into 10 equal parts, and any 9 of them are used for model training, and the other one is used for validation, that is, 10 model results are generated. The above process is repeated 10 times, that is, this step generates a total of 100 models.
[0121] (3.14) Set candidate hyperparameters. According to the R language 'caret' package, there is only one adjustable hyperparameter for RF, mtry. Based on experience, mtry is set to 2 to 20, for a total of 19 candidate hyperparameters.
[0122] (3.15) Set the model evaluation criteria: Use ROC as the model evaluation criterion, that is, calculate the area under the ROC curve (AUC).
[0123] (3.16) Optimal hyperparameter selection: 19 candidate hyperparameters are used to train the model respectively, and the AUC of the generated 100 models is averaged as the evaluation result of the candidate hyperparameter. Finally, the candidate hyperparameter with the highest evaluation result is selected as the optimal hyperparameter.
[0124] For example, the optimal parameter selection is performed by using 10-fold cross validation resampling and receiver operating characteristic (ROC) evaluation method with 10 repetitions. The optimized parameters here are model hyperparameters, and the R language RF model related code is used as an example for specific explanation. The related code is:
[0125]
[0126]
[0127] (3.2): Use the optimal parameters obtained in step (3.1) of this example to train the model.
[0128] The computer randomly splits the modeling data set into 80% training set and 20% test set, and the random sampling method is the same as in step (3.1). Repeat the training 100 times, and obtain the prediction results of the model obtained after each training in the test set, that is, the possibility of each data being positive. Then, by comparing with the actual results, the cutoff value for obtaining the maximum Youden index is obtained, that is, setting any cutoff value can divide a group of numbers into two results (greater than this value is positive, otherwise it is negative). The cutoff value to be obtained here is the one that can make the predicted results and the actual results match best, and the standard is the highest Youden index. Use this cutoff value to obtain the binary classification prediction results of the test set.
[0129] This is not the final prediction result, but the prediction result of the test set. The prediction result will be compared with the actual label to get the AUC. The binary prediction result refers to positive (P) or negative (N). For example, the brief training code for model training using the optimal hyperparameters in step (3.1) is as follows:
[0130] set.seed(seed+i)
[0131] gbm.train<-train(
[0132] x=x.train,
[0133] y=y.train,
[0134] method="rf",
[0135] trControl=fitControl,
[0136] verbose=FALSE,
[0137] tuneGrid=bestTune,
[0138] metric="ROC" )
[0140] predict(gbm.train,
[0141] newdata=train_test, #train_test is the test set, containing 10 positive and 10 negative data
[0142] type = 'prob')
[0143] #Here we will get the prediction results of the test set, that is, the possibility of each data being positive. Then, by comparing with the actual results, we can get the cutoff value for obtaining the maximum Youden index, that is, setting any cutoff value can divide a group of numbers into two results (greater than this value is positive, otherwise it is negative). The cutoff value to be obtained here is the one that can make the prediction result and the actual result match best, and the standard is the highest Youden index.
[0144] (3.3): Use the test set to verify the model.
[0145] The consistency between the prediction results and labels in the test set was calculated, and internal validation was carried out using indicators such as the area under the ROC curve (AUC) and F1 score. The results showed that the model effects of the four machine learning models were all good (AUC range 0.887-0.970).
[0146] (3.4): Use the external validation set to validate the different machine learning models trained above (the data of the external validation set is different from the data set used in the modeling process).
[0147] (3.41) The computer marks the subtype years that meet the inclusion criteria from 2021 to 2022 based on the first spillover events of each subtype of avian influenza virus and the annual number of human infection cases between 2021 and 2022 (the standards are the same as step 1) to form an external validation set.
[0148] (3.42) The computer calculates the consistency between the prediction results and the labels in the external validation set, and uses the accuracy rate as an indicator for external validation.
[0149] For example, the external validation set contains a total of 25 subtype-years, including 13 subtypes other than H3N8 in 2022 (H3N8-2022 is the first overflow, not a second overflow) and 12 subtypes other than H3N8 and H10N3 in 2021 (H10N3-2021 is the first overflow. H3N8-2021 has not yet overflowed), of which 7 subtype-years have human infection cases and 18 subtype-years have no human infection cases.
[0150] (3.43): The computer selects the optimal machine learning model based on the external validation results (the machine learning method with the highest accuracy) (RF has the best external validation results, with an accuracy of 88% or 22 / 25).
[0151] (3.44): The computer extracted the relative contribution of each indicator in the best model. The results showed that the three indicators with higher relative contribution were the number of sequences of this subtype in hosts in close contact with humans in the previous two years (relative contribution of 3.33%, 95% confidence interval of 2.42% to 4.43%), the number of host species in close contact with humans where this subtype was found in the previous two years (relative contribution of 3.30%, 95% confidence interval of 2.00% to 4.32%), and the number of sequences of this subtype in hosts in close contact with humans in the previous three years (relative contribution of 3.08%, 95% confidence interval of 2.21% to 4.22%), all of which are variables related to hosts in close contact with humans.
[0152] Table 1 Number of human infection cases of avian influenza virus in 2001-2022
[0153]
[0154]
[0155]
[0156] In another exemplary embodiment of the present application, taking avian influenza virus as an example, the process of the above steps 203 to 205 is described in detail. Among them:
[0157] Step 1: The computer determines which risk assessment method to use based on the pathogen type and the initial spill status.
[0158] ① If the pathogen to be evaluated is an avian influenza virus subtype, the evaluation can be carried out directly by referring to the relevant results in the present invention:
[0159] a. If there have been no human cases of infection with this avian influenza virus subtype, the first spillover risk assessment criteria can be used for the assessment.
[0160] b. If human infection has occurred with this avian influenza virus subtype, the re-spillover risk assessment criteria can be used to conduct the assessment.
[0161] ② If the pathogen to be evaluated is a certain lineage of avian influenza virus, the 255 relevant indicators in the present invention can be directly used, and the evaluation criteria can be re-established by referring to the relevant methods in the present invention, and then the evaluation can be carried out.
[0162] ③ If the pathogen to be evaluated is not avian influenza virus, the relevant indicators and methods in the present invention can be referred to, the host in close contact with humans can be redefined and relevant indicators can be set, and then evaluation standards can be established to carry out evaluation.
[0163] Step 2: The computer calculates the infection distribution dynamics and prevalence of pathogens in animals for each subtype year.
[0164] Step 3: The computer determines the warning status of each subtype year based on the first spillover risk standard.
[0165] Taking the avian influenza virus subtype as an example, it meets (is greater than or equal to) the following requirements (the first spillover risk standard of Group B+3): 1) Found in 3 continents within the background time range. 2) 44 sequences were found within the background time range. 3) 23 sequences were found in hosts in close contact with humans within the background time range. 4) There were records of discovery of hosts in close contact with humans in the 4 years within the background time range. 5) 1 sequence was found in the previous year. 6) Found in 3 provinces in the previous two years. 7) Found in 1 new province in the previous two years. 8) 6 sequences were found in the previous two years. 9) Found in 6 provinces in the previous three years. 10) Found in 4 host species in the previous three years. 11) Found in 1 new host species in the previous three years. 12) 10 sequences were found in the previous three years. 13) 3 sequences were found in hosts in close contact with humans in the previous three years.
[0166] Step 4: The computer evaluates the warning status of each subtype year based on the re-spillover risk model.
[0167] Taking the avian influenza virus subtype as an example, all 255 indicator data of a specific subtype year were input, the output value was calculated based on the RF model, and it was divided into binary variables according to the cutoff value (if the output value was greater than the cutoff value, it was positive, otherwise it was negative).
[0168] In summary, this application is based on the gene sequences collected and / or independently sampled and monitored by GenBank, GISAID, etc., which can comprehensively analyze the infection distribution dynamics and prevalence of pathogens in animals, and establish a rapid assessment method for pathogen spillover risk, realize the rapid identification and threat assessment of high-risk pathogens, and make up for the shortcomings of existing risk assessment tools in terms of timeliness and automation. Among them, by analyzing the relevant indicators of the years of historical spillover, it is found that the important role of hosts in close contact with humans in assessing spillover risks is found; by constructing epidemic ecological indicators of pathogens in animals, its spillover potential is fully characterized, that is, whether different strains can infect mammals (mammalian adaptability), whether they are prevalent in hosts in close contact with humans (human-animal contact surface), whether there is a host expansion phenomenon (host adaptability change), and whether there is a spatial expansion phenomenon (increased infectivity). In addition, the background situation can reflect the infectivity and host adaptability of the virus, and the recent epidemic situation can reflect the changes in the animal population immunity level, the size of the human-animal interface, and the infectivity, host adaptability and other characteristics of the virus. In order to explore the most suitable near-term time range for characterizing the spillover of avian influenza virus, the above example sets 10 near-term time ranges such as the previous 1 to 10 years for screening. The above indicators are assigned by collecting information on gene sequences in public databases, thus getting rid of the cumbersome process of expert consultation and realizing the (computer) automated assessment of pathogen spillover risk.
[0169] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Figure 3 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store overflow risk assessment data. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a method for evaluating the overflow risk of a pathogen epidemic ecology in animals is implemented.
[0170] Those skilled in the art will understand that Figure 3The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0171] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0172] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0173] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0174] Those of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0175] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, etc., but is not limited thereto.
[0176] The technical features of the above embodiments may be arbitrarily combined. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0177] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A method for assessing the spillover risk of pathogen epidemic ecology in animals, characterized in that: The spillover risk assessment method for the epidemic ecology of the pathogen in animals includes: Obtain pathogen indicator data; pathogen indicators include pathogen spatial distribution indicators, infected host width indicators and epidemic intensity indicators; the pathogen spatial distribution indicators include known spatial distribution ranges and changes in the spread of spatial distribution ranges; infected host width indicators include known infected host widths and changes in host distribution ranges; epidemic intensity indicators include the number and frequency of sequences reported in different hosts; frequency refers to the number of years of occurrence; Based on the pathogen indicator data, a first spillover risk model was formed by induction using a stepwise approach; Based on the pathogen indicator data, a re-spillover risk model is constructed using a machine learning method; Determine the indicator data of the pathogen to be evaluated in each subtype year; the subtype year is used to indicate the spillover status of different pathogens in different years, or to indicate the spillover status of different subtypes or lineages in different years; Determining the type of pathogen to be evaluated, and selecting a spillover risk assessment model based on the type of pathogen to be evaluated; the spillover risk assessment model is a first spillover risk model or a second spillover risk model; The selected spillover risk assessment model is used to complete the spillover risk assessment based on the indicator data of the pathogen to be assessed in each subtype year, and the early warning situation of each subtype year is obtained; Based on the pathogen indicator data, a first spillover risk model was formed using a stepwise approach, specifically including: According to the first spillover event of each subtype of pathogen occurring in the set time period, the subtype-years meeting the inclusion criteria in the set time period are divided into new-onset group and non-new-onset group; Based on the pathogen index data, obtaining the pathogen index for each subtype year, and determining the minimum value of each pathogen index in the new outbreak group; Taking the minimum value of each pathogen index in the new outbreak group as the standard, the subtype-years meeting the set standard are screened in the non-new outbreak group and counted; The pathogen indicators of each subtype year are screened out, and the pathogen indicators that change the number of subtype years that meet the set standards are retained; The retained pathogen index and the corresponding minimum value of this pathogen index in the new outbreak group are used as the first spillover risk standard; Determine the early warning status of different pathogen subtypes in different years based on the first spillover risk criteria to form external verification results; Based on the external validation results, the first spillover risk standard corresponding to the minimum total number of subtypes per year is obtained, and this first spillover risk standard is used as the final first spillover risk standard; The first spillover risk model is constructed based on the final first spillover risk standard; Based on the pathogen indicator data, a re-spillover risk model is constructed using a machine learning method, specifically including: Based on the first spillover events of each pathogen subtype occurring in the set time period and the number of annual animal infection cases of the pathogen, the subtype-years that meet the inclusion criteria in the set time period are divided into recurrence group and non-recurrence group; Generate year labels for each subtype in the relapse group and year labels for each subtype in the non-relapse group to obtain data labels; The pathogen indicator data is used as input, and the data label is used as output. Model training is carried out based on different machine learning methods until the set training conditions are met, and the trained model is used as the re-spillover risk model.
2. The method for assessing the risk of spillover of pathogen epidemic ecology in animals according to claim 1, characterized in that: Obtain pathogen indicator data, including: Obtain pathogen entry data according to the pathogen name, and generate an Excel file based on the pathogen entry data; Standardizing the pathogen entry data in the Excel file to obtain a data set; Based on the data set, n indicators of each pathogen subtype year are obtained to generate the pathogen indicator data.
3. The method for assessing the risk of spillover of pathogen epidemic ecology in animals according to claim 2, characterized in that: Get pathogen entry data according to pathogen name, including: In the Protein database of the GenBank website, keyword searches were performed according to the pathogen names to obtain sequence data; Using regular expressions to extract sequence information of each sequence from the sequence data; the sequence information includes: version number, strain name, location, time, host, subtype, segment and original sequence; The sequence information is preprocessed to obtain pathogen entry data.
4. The method for assessing the risk of spillover of pathogen epidemic ecology in animals according to claim 2, characterized in that: The pathogen entry data in the Excel file is standardized to obtain a data set, which specifically includes: Taking the comparison table as the standard, matching based on the Excel file forms a host field; the host field includes a host type and a host species; Taking the comparison table as the standard, a host type field is formed based on the Excel file matching; the host type field includes: a "whether it is a host in close contact with humans" field and a "whether it is a mammalian host" field; Based on the comparison table, a location field is formed based on the Excel file matching; the location field includes a continent, a country, and a province; Based on the comparison table, the year field is formed based on the Excel file matching; The data set is generated based on the host field, the host type field, the location field, and the year field.
5. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the method for assessing the epidemic ecology of pathogens in animals as described in any one of claims 1 to 4.