A method for quantitatively assessing the diversity of pathogen evolutionary directions

Through the feature extraction and cluster evaluation algorithm of pathogen genotype data, the diversity of pathogen evolution directions is quantified, and the problem of insufficient analysis complexity and accuracy in the existing technology is solved, and rapid and accurate assessment of pathogen evolution directions and identification of transmission risks is achieved.

CN120015105BActive Publication Date: 2025-08-22ACADEMY OF MILITARY MEDICAL SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510075198.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-08-22
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The prior art has complex analysis process, difficult to calculate, small data volume, insufficient accuracy of results when evaluating the evolutionary direction of pathogens, making it difficult to quickly identify and quantify the diversity of evolutionary directions of pathogens.

Method used

By collecting genotype data of pathogens, using feature extraction and cluster evaluation methods, a feature matrix was generated, and a cluster evaluation algorithm was used to quantify the evolutionary direction diversity of pathogens, and combining intra-cluster density and inter-cluster separation were used to calculate the evolutionary direction diversity evaluation index.

Benefits of technology

A rapid, concise and accurate diversity assessment of the evolution direction of pathogens is achieved, which can effectively identify the risk of spread spillover, improve the analysis efficiency and stability of results, and reduce the deviation of manual judgment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015105B_ABST
    Figure CN120015105B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for quantitatively assessing pathogen evolutionary diversity, comprising the following steps: 1) collecting core gene sequence data of pathogens based on specified pathogen species and performing quality control on the collected data to obtain high-quality sequence data; 2) then extracting features from the collected sequence data using Shannon entropy to obtain a feature vector for each sequence; 3) constructing a feature matrix based on the extracted feature data, and calculating the pathogen evolutionary diversity score using a clustering evaluation calculation method. By extracting sequence features, the present invention can achieve rapid and efficient evolutionary diversity assessment and transmission spillover risk analysis without the need for evolutionary tree analysis or modeling training. This method can provide technical support and reference for subsequent pathogen genetic analysis, related epidemic detection and monitoring, and the development of specific drugs and antibodies, and has a wide range of related applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology, and in particular to a method for calculating the diversity of pathogen evolutionary directions based on a clustering quality evaluation method. Background Art

[0002] Pathogens are widespread and diverse in nature, posing a significant threat to human health and life. Different pathogens exhibit distinct evolutionary characteristics and patterns of transmission. Rapidly identifying and understanding the evolutionary trajectory of pathogens is a key challenge facing biosecurity and disease control.

[0003] The evolutionary direction of pathogens is often related to their habitats. Important influencing factors include the pathogen's geographic distribution and host distribution. For example, some pathogens can spread across multiple regions and even cross regions, suggesting that their evolutionary direction in terms of geographic distribution tends to be diversified and presents a risk of spillover, such as avian influenza. Other pathogens, such as Ebola and Lassa viruses, spread only within specific regions and rarely cross regions, suggesting that their evolutionary direction in terms of geographic distribution tends to be stable and presents a lower risk of spillover.

[0004] Quantitative assessment of the evolutionary direction of pathogens can not only help researchers and medical staff gain a deeper understanding of the inherent evolutionary laws of pathogens, but also predict the risk of future spread and spillover of pathogens, and provide theoretical basis and technical support for whether cities or countries around the epidemic area should upgrade or downgrade their defenses.

[0005] Current methods for detecting pathogen evolutionary trends typically include multiple steps, including genetic data collection, tagging, multiple sequence alignment, evolutionary analysis, cluster analysis, and geographic phylogenetic analysis. These analyses rely on subjective judgment, are complex, and computationally challenging, resulting in low efficiency, limited analyzable data volumes, and inaccurate results.

[0006] Assessing pathogen evolutionary diversity essentially involves examining whether the genetic evolutionary characteristics of pathogens differ significantly across regions (or hosts). If the genetic evolutionary characteristics of pathogens vary significantly, with each region (or host) displaying distinct evolutionary characteristics, this indicates that the pathogen's evolutionary direction is relatively stable and concentrated. Conversely, if the genetic evolutionary characteristics of pathogens vary little, this indicates that the pathogen's evolutionary direction is relatively diverse, posing a risk of spillover.

[0007] Using interdisciplinary approaches, the biological problem of studying pathogen evolutionary diversity can be transformed into the computational problem of estimating sample distributions. Pathogen sequence features serve as data samples, and pathogen distribution attributes (geographic or host information) serve as sample labels. The analysis of evolutionary diversity can be transformed into an assessment of the degree of disorder in the multi-category data distribution. A more disordered distribution indicates a greater diversity of evolutionary directions, while a less disordered distribution indicates a more stable evolutionary direction. Summary of the Invention

[0008] The purpose of the present invention is to provide a method for quantitatively evaluating the diversity of pathogen evolutionary directions.

[0009] The present invention provides a method for quantitatively evaluating the diversity of pathogen evolutionary directions. For a given pathogen, the genotype data of the pathogen is collected, and features are extracted from the sequence using a related method. A feature matrix is ​​then generated based on the extracted feature vectors and natural clustering results, and a clustering evaluation method is used to quantitatively calculate the diversity of pathogen evolutionary directions.

[0010] Further, the following steps are included:

[0011] 1) Collect core gene sequences of target pathogens;

[0012] 2) Perform feature extraction on each sequence obtained in step 1) to obtain the feature vector of the core gene sequence;

[0013] 3) Generate a feature matrix based on the feature vector obtained in step 2) and the natural clustering result, and use the clustering evaluation method to quantitatively calculate the diversity of the pathogen's evolutionary direction to obtain an evaluation index for the pathogen's evolutionary direction diversity.

[0014] Furthermore, the core gene sequence in step 1) is a specific gene sequence or a full-length genome sequence having core characteristics of the pathogen, and the sequence type includes a nucleic acid sequence or an amino acid sequence.

[0015] 4. The method for quantitatively assessing pathogen evolutionary diversity according to claim 2, wherein the method for extracting the feature vector comprises the following steps:

[0016] 21) Distance Measure based on k-tuple (DMk) is used to calculate the position and occurrence information of k-tuples in the sequence. When k=3, k-tuples can be regarded as codons. The position of each k-tuple is recorded as ,in Representative The position of the k-tuple that appears the first time, , is the number of times the k-tuple appears in the sequence;

[0017] 22) Calculate the intervals between the positions where k tuples appear, and combine all the intervals to be , its mathematical calculation formula is as follows:

[0018] ;

[0019] 23) Based on the interval information sequence of k tuples, calculate the interval sum and combine them into a sequence which can be recorded as , its mathematical calculation formula is as follows:

[0020]

[0021] in, It is determined by the number and position of the k-tuple. At the same time, the number and position of the k-tuple can also be obtained through intervals and sequences;

[0022] 24) Calculate the Shannon entropy of the sequence based on the interval and probability, and define a discrete probability distribution based on the interval and ,in The mathematical formula for this is as follows:

[0023]

[0024] Then the Shannon entropy can be calculated:

[0025]

[0026] Shannon entropy reflects the number and position information of a k-tuple, which is used as the feature of a k-tuple in the virus sequence;

[0027] 25) Repeat steps 21)-24) for each k-tuple to extract its features, and finally get a feature vector, which is recorded as , where , this vector is the feature vector extracted from the sequence.

[0028] Furthermore, the calculation method of the pathogen evolutionary direction diversity evaluation index in step 3) includes the following steps:

[0029] 31) The attribute labels of pathogens are regarded as natural clustering results. The sequence features of the same type of pathogens are combined into a class feature matrix. The intra-class variance is used to calculate the intra-cluster density of the feature matrix. The calculation formula is as follows:

[0030]

[0031] in, represents the number of categories, Indicates the All samples in a class, Indicates the The cluster center corresponding to each host, represents the dimension of the feature vector, Represents the feature vector No. bit vector value, Represents the feature vector No. bit vector value;

[0032] 32) Use the inter-class variance to calculate the inter-cluster separation of the feature matrix. The calculation formula is as follows:

[0033]

[0034] in, represents the number of categories, Indicates the The number of samples in a class, represents the number of all samples, represents the global center, represents the dimension of the feature vector, represents the global center vector No. bit vector value, Represents the center vector within the class No. bit vector value;

[0035] 33) Taking into account the influence of intra-cluster compactness, inter-cluster separation and the number of sample clusters, the calculation formula for the evolutionary direction diversity evaluation index is:

[0036]

[0037] in, It is an evaluation index of the diversity of the evolutionary direction of pathogens.

[0038] The application of the above-mentioned method for quantitatively evaluating the diversity of pathogen evolutionary directions in the preparation of products for evaluating the risk of pathogen transmission and spillover.

[0039] Furthermore, the standard for evaluating the risk of pathogen transmission spillover is that the larger the value of the pathogen's evolutionary direction diversity evaluation index, the more chaotic the category distribution, the higher the evolutionary direction diversity, and the higher the risk of pathogen transmission spillover; the smaller the value of the pathogen's evolutionary direction diversity evaluation index, the greater the difference between categories, the more concentrated the evolutionary directions, and the lower the risk of pathogen transmission spillover.

[0040] A product for evaluating the risk of pathogen transmission spillover, including a computer program, characterized in that when the computer program is executed by a processor, it implements the steps of the method for quantitatively evaluating the diversity of pathogen evolutionary directions, obtains an evaluation index of the diversity of pathogen evolutionary directions; and evaluates the risk of pathogen transmission spillover based on the size of the evaluation index of the diversity of pathogen evolutionary directions.

[0041] Furthermore, the standard for evaluating the risk of pathogen transmission spillover is that the larger the value of the pathogen's evolutionary direction diversity evaluation index, the more chaotic the category distribution, the higher the evolutionary direction diversity, and the higher the risk of pathogen transmission spillover; the smaller the value of the pathogen's evolutionary direction diversity evaluation index, the greater the difference between categories, the more concentrated the evolutionary directions, and the lower the risk of pathogen transmission spillover.

[0042] This paper uses an interdisciplinary approach to regard distribution attribute (geographic or host) labels as the result of "natural clustering" of pathogens, and uses a clustering evaluation algorithm to evaluate the clustering quality of "natural clustering" to determine the degree of chaos in the data distribution, thereby completing a quantitative assessment of the diversity of pathogen evolutionary directions.

[0043] Compared with the existing technology, the advantages of the present invention are: 1. The calculation is simple and intuitive, a large amount of data can be analyzed, the analysis efficiency is high, and the degree of diversity in the evolutionary direction of pathogens can be effectively tracked and quantified; 2. The clustering evaluation algorithm is used to evaluate the diversity of evolutionary directions, which is highly innovative. By utilizing the advantages of interdisciplinary technology, it can effectively avoid the result deviation caused by inconsistent pathogen sequence quality and the increased difficulty of calculation, and can cover more data information, effectively handle noise and missing data, and have higher analysis accuracy; 3. The method has the characteristics of full-process automated calculation, which avoids deviations and misleading caused by manual measurement, and the analysis results have strong stability and reliability; 4. According to this method, the diversity of pathogen evolutionary directions is quantitatively evaluated, which can effectively perceive the geographical (or host) spread and spillover situation of pathogens, provide technical support and reference for pathogen identification and epidemic prevention and control, and has strong practicality.

[0044] In summary, the present invention realizes the quantitative assessment of the diversity of pathogen evolutionary directions through the process methods of pathogen sequence collection quality control, feature extraction, cluster evaluation, etc. The assessment speed is fast, the method is simple and effective, and the spread and spillover situation and risk of pathogens can be systematically and comprehensively measured. This invention avoids the complexity problems existing in the classical methods, such as low computational efficiency, small amount of processable data, and difficulty in interpreting analysis results, and overcomes the influence of interference factors such as data noise and quantity scale. This method is mainly reflected in the rapid assessment of the diversity of pathogen evolutionary directions, and the analysis of pathogen transmission risks can be realized without evolutionary tree analysis or modeling training. Utilizing this technical advantage will provide technical methods and references for pathogen prevention and epidemic risk prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 Computational flowchart for quantitatively assessing the diversity of pathogen evolutionary directions.

[0046] Figure 2 This is a diagram showing the quantitative analysis of the diversity in the evolutionary direction of Ebola virus and the verification of the results. DETAILED DESCRIPTION SUMMARY OF THE INVENTION

[0048] The method provided by the present invention is to collect the gene sequence and annotation information of a given type of pathogen, use relevant methods to extract features from the sequence, and then use a clustering evaluation algorithm to calculate the extracted feature vectors to complete a quantitative assessment of the diversity of the pathogen's evolutionary direction. Specifically, the present invention first collects the gene sequence data of the pathogen based on the specified pathogen type, completes the attribute label information annotation, and performs quality control on the collected data to obtain a high-quality sequence data set; then, feature extraction is performed on the collected pathogen sequence data to obtain a feature vector for each sequence; finally, the sequence feature vectors are classified and combined according to the label information to obtain a feature matrix for each type of pathogen, and a clustering evaluation algorithm is used to quantitatively evaluate the degree of chaos in the pathogen sequence data, completing a quantitative assessment of the diversity of the evolutionary direction.

[0049] The above method comprises the following steps:

[0050] 1. Collect and download the core gene sequence of the target pathogen.

[0051] 1.1 Collection of pathogen core gene sequence data

[0052] For pathogens, the method of the present invention requires the collection and organization of as much core gene sequence data as possible to more accurately measure the evolutionary diversity of the target pathogen. Through database searches (such as NCBI GenBank, GISAID, and other databases), literature review, biological experiments, and other means, as comprehensive a collection of core gene sequence data as possible is necessary. These core genes are generally those related to the pathogen's host adaptability and can include a specific gene or the entire pathogen's genes. The collected core gene sequence data should generally be no fewer than 100 entries.

[0053] 1.2 Pathogen sequence data quality control and attribute label information annotation

[0054] To ensure the stability and effectiveness of the experimental results, it is necessary to perform quality control on the core gene sequence data of the pathogen and delete sequences that are insufficient in length, contain abnormal characters, or are repeated or redundant.

[0055] Determine the pathogen attribute label type, selecting geographic distribution, host distribution, or other attribute information as the attribute label, and label each sequence. Delete sequences with unclear labels to complete data cleaning and screening. The filtered dataset serves as the baseline dataset for subsequent analysis and processing.

[0056] 2. Perform feature extraction on each sequence obtained in step 1) as follows:

[0057] 2.1 Use distance measure based on k-tuple (DMk) to calculate the position and occurrence information of k-tuples in the sequence. k =3, the k-tuple can be regarded as a codon in the nucleotide sequence. Therefore, the value of k in DMk can be various, but in this method, it is generally taken as 3 because it involves nucleic acid sequences and codons. The position of each k-tuple is recorded as ,in Representative The position where the k-tuple appears for the first time. is the number of times the k-tuple occurs in the sequence. It should be noted that this method supports the study of both nucleic acid and amino acid sequence data. Depending on the sequence type, the k value can be adjusted to capture more accurate sequence features.

[0058] 2.2 Calculate the intervals between the positions where k-tuples appear, and combine all the intervals to be , its mathematical calculation formula is as follows:

[0059]

[0060] 2.3 Based on the interval information sequence of k-tuples, calculate the interval sum and combine them into a sequence which can be recorded as , its mathematical calculation formula is as follows:

[0061]

[0062] in, It is determined by the number and position of the k-tuple. At the same time, the number and position of the k-tuple can also be obtained through intervals and sequences.

[0063] 2.4 Calculate the Shannon entropy of the sequence based on the interval and probability. Define a discrete probability distribution based on the interval and ,in The mathematical formula for this is as follows:

[0064]

[0065] Then the Shannon entropy can be calculated:

[0066]

[0067] Shannon entropy reflects the number and position information of a k-tuple, which is used as the characteristic of a k-tuple in the pathogen sequence.

[0068] 2.5 Repeat steps 2.1-2.4 for each k-tuple to extract its features, and finally obtain a feature vector, which is recorded as ,in, ,when k =3, t = 64. This vector is the feature vector of the pathogen sequence.

[0069] 3. Use the clustering evaluation algorithm to calculate and evaluate the feature vectors of all sequences obtained in step 2. The specific steps are as follows:

[0070] 3.1 Sequences with the same attribute labels in the dataset are considered to be of the same type, and the feature vectors of sequences of the same type are combined into a matrix as the class feature matrix.

[0071] 3.2 Use the intra-class variance to calculate the intra-cluster closeness of the feature matrix (intra-cluster closeness), and the calculation formula is as follows:

[0072]

[0073] in, represents the number of categories, Indicates the All samples in a class, Indicates the The cluster center corresponding to each host, represents the dimension of the feature vector, Represents the feature vector No. bit vector value, Represents the feature vector No. A vector value of bits.

[0074] 3.3 Use the inter-class variance to calculate the inter-cluster separation of the feature matrix. The calculation formula is as follows:

[0075]

[0076] in, represents the number of categories, Indicates the The number of samples in a class, represents the number of all samples, represents the global center, represents the dimension of the feature vector, represents the global center vector No. bit vector value, Represents the center vector within the class No. A vector value of bits.

[0077] 3.4 Taking into account the influence of intra-cluster compactness, inter-cluster separation, and the number of sample clusters, the calculation formula for the evolutionary direction diversity evaluation index is:

[0078]

[0079] in, It is an evaluation index of the diversity of the evolutionary direction of pathogens. The larger the value, the more chaotic the category distribution, the higher the diversity of evolutionary directions, and the higher the risk of spillover. The smaller the value, the greater the difference between categories, the more concentrated the evolutionary direction, and the lower the risk of spillover.

[0080] Through the above three steps, a quantitative score of the evolutionary diversity of the pathogen to be tested can be obtained.

[0081] The present invention will be further described in detail below in conjunction with specific embodiments. The examples provided are only for illustrating the present invention and are not intended to limit the scope of the present invention. The examples provided below can serve as a guide for further improvements by those skilled in the art and are not intended to limit the present invention in any way.

[0082] Unless otherwise specified, the experimental methods in the following examples are conventional methods and were performed according to the techniques or conditions described in the literature in the field or according to the product instructions. The materials and reagents used in the following examples, unless otherwise specified, were all commercially available.

[0083] Unless otherwise specified, the quantitative tests in the following examples were performed three times, and the results were averaged.

[0084] Example 1: Quantitative evaluation and trend tracking analysis of Ebola virus evolutionary diversity

[0085] This embodiment establishes a method for quantitatively calculating the diversity of virus evolutionary directions based on a clustering evaluation algorithm, quantitatively calculates and tracks the changing trends of the diversity of the evolutionary directions of the Ebola virus (Ebola) in different time intervals, and completes the risk assessment of transmission spillover.

[0086] 1. Collection and preprocessing of viral genome data

[0087] 1.1 Collection of viral genome nucleic acid sequence data

[0088] For viruses, the method of the present invention requires the collection and organization of as much viral genome sequence data as possible to more accurately measure the virus's evolutionary diversity. In this example, the publicly available Ebola virus genome data was downloaded from the NCBI GenBank online database (https: / / www.ncbi.nlm.nih.gov). The GP protein gene was selected as the core Ebola virus gene for analysis. A total of 3,850 DNA sequence data entries were obtained, which served as the viral genome dataset for this case.

[0089] 1.2 Virus sequence data preprocessing

[0090] To ensure the stability and validity of the experimental results, this example performed quality control on the viral sequence data. The GP protein DNA sequence of the Ebola virus reference strain is 2031 characters long. Therefore, sequences less than 90% of the reference sequence length were deleted. Specifically, sequences less than 1818 characters were deleted from the viral genome dataset. Sequences containing a large number of abnormal characters or duplicates were also deleted. After quality control, the viral genome dataset contained a total of 730 sequences.

[0091] We selected geographic information from the Ebola virus as attribute tags for quantitative evaluation of evolutionary diversity. The geographic attribute tags in the viral genome dataset included 13 countries and regions: Zaire, Guinea, Uganda, Sierra Leone, Gabon, Liberia, Germany, Nigeria, USA, United Kingdom, Italy, Switzerland, and Mali. These 13 regions served as natural clustering tags in subsequent analyses.

[0092] 2. Extract virus feature vectors

[0093] Extract feature vectors from the gene sequences of all virus strains obtained in step 1 using the following method:

[0094] 2.1 Use k-tuple-based distance measurement to calculate the position and occurrence information of all k-tuples in each sequence. In this embodiment, let k=3 , treat k-tuples as codons to facilitate genomic genetic analysis. For each k-tuple, count the number of times it appears in the sequence and its position information. Recorded as ,in Representative The position of the k-tuple that appears the first time, , is the number of times the k-tuple appears in the sequence.

[0095] 2.2 For each k-tuple, calculate the interval between adjacent occurrence positions, recorded as , its mathematical calculation formula is as follows:

[0096]

[0097] 2.3 Based on the interval information of k-tuples, calculate the interval sum, which is recorded as , its mathematical calculation formula is as follows:

[0098]

[0099] in, Determined by the number and position of the k-tuple.

[0100] 2.4 Calculate the Shannon entropy of the sequence. Define a discrete probability distribution ,in The mathematical formula for this is as follows:

[0101]

[0102] Then the Shannon entropy can be calculated:

[0103]

[0104] Shannon entropy reflects the number and position information of a k-tuple, which is used as the characteristic of the k-tuple in the viral gene sequence.

[0105] 2.5 Repeat steps 2.1-2.4 for each k-tuple to extract its features, and finally get a feature vector, which is recorded as , where .because ,so , this vector is the characteristic vector of the virus sequence.

[0106] All sequence eigenvectors under the natural clustering label are combined together to obtain the eigenvector matrix of the label category.

[0107] 3. Quantitative evaluation and calculation of virus evolutionary diversity

[0108] Calculate the eigenvectors of all sequences obtained in step 2. The specific steps are as follows:

[0109] 3.1 Consider the attribute labels of the virus as natural clustering results, and use the intra-class variance to calculate the intra-cluster density of the feature matrix. The calculation formula is as follows:

[0110]

[0111] in, represents the number of categories, Indicates the All samples in a class, Indicates the The cluster center corresponding to each host, represents the dimension of the feature vector, Represents the feature vector No. bit vector value, Represents the feature vector No. A vector value of bits.

[0112] 3.2 Use the inter-class variance to calculate the inter-cluster separation of the feature matrix. The calculation formula is as follows:

[0113]

[0114] in, represents the number of categories, Indicates the The number of samples in a class, represents the number of all samples, represents the global center, represents the dimension of the feature vector, represents the global center vector No. bit vector value, Represents the center vector within the class No. A vector value of bits.

[0115] 3.3 Taking into account the influence of intra-cluster compactness, inter-cluster separation and the number of sample clusters, the calculation formula of the evolutionary direction diversity evaluation index is:

[0116]

[0117] in, It is an evaluation index of the diversity of the evolutionary direction of pathogens. The larger the value, the more chaotic the category distribution, the higher the diversity of evolutionary directions, and the higher the risk of spillover. The smaller the value, the greater the difference between categories, the more concentrated the evolutionary direction, and the lower the risk of spillover.

[0118] Through the above three steps, we can obtain the quantitative score of the evolutionary diversity of Ebola virus in terms of geographical attributes.

[0119] To track the evolutionary diversity of Ebola virus over time, the viral dataset can be divided into four datasets based on year: "2005 and earlier," "2010 and earlier," "2015 and earlier," and "2020 and earlier." Based on the above three steps, quantitative scores for evolutionary diversity were calculated for each dataset, with the results shown in Table 1. To assess only the current evolutionary diversity of Ebola virus, all current sequences can be used for calculations, without the need to partition the viral genome dataset.

[0120] Table 1 Quantitative assessment and tracking of Ebola virus evolutionary diversity

[0121]

[0122] The calculation results in this example show that the evolutionary diversity scores for the "2005 and earlier virus datasets" and "2010 and earlier virus datasets" are low, indicating that the risk of geographic spread of Ebola virus before 2010 was low and the evolutionary diversity was relatively simple. However, the evolutionary diversity scores for the "2015 and earlier virus datasets" and "2020 and earlier virus datasets" gradually increased, indicating that the risk of geographic spread of Ebola virus after 2015 increased and the evolutionary diversity became more complex. This is consistent with actual news reports and the spread of the epidemic.

[0123] To further verify the accuracy of the method of the present invention, this example used a geographic phylogenetics approach to experimentally analyze the virus. This approach can simulate historical virus transmission routes. A greater number of cross-regional transmission routes indicates a greater risk of virus spillover and a more complex evolutionary diversity. Conversely, a lower risk of virus spillover and a simpler evolutionary diversity indicate a lower risk of virus spillover. Geographic phylogenetics is currently one of the main analytical methods for determining virus transmission trends. Experimental results based on geographic phylogenetic data revealed that before 2010, Ebola virus had only one cross-regional transmission route, indicating a low risk of spillover. After 2015, Ebola virus developed five cross-regional transmission routes, indicating an increased risk of spillover, consistent with the quantitative assessment results of the method of the present invention. Furthermore, before 2010, Ebola virus had only two major geographic transmission areas, but by around 2020, the number had expanded to 13, indicating that Ebola virus spillover has indeed occurred, which is consistent with the analysis results of the method of the present invention. This demonstrates the effectiveness and accuracy of the method of the present invention.

[0124] The present invention has been described in detail above. It will be apparent to those skilled in the art that the present invention may be practiced over a wide range of parameters, concentrations, and conditions without departing from the spirit and scope of the present invention and without unnecessary experimentation. Although specific embodiments have been given herein, it should be understood that further modifications may be made to the present invention. In summary, this application is intended to encompass any variations, uses, or improvements to the present invention, including those made by conventional techniques known in the art that depart from the scope of the present invention. Applications of the essential features may be made within the scope of the following claims.

Claims

1. A method for quantitatively assessing the diversity of pathogen evolutionary directions, characterized in that: For a given pathogen, the genotype data of the pathogen is collected, and the sequence is feature extracted using relevant methods. A feature matrix is ​​generated based on the extracted feature vectors and the natural clustering results. The clustering evaluation method is used to quantitatively calculate the diversity of the pathogen's evolutionary direction. The process includes the following steps: 1) Collect core gene sequences of target pathogens; 2) Perform feature extraction on each sequence obtained in step 1) to obtain the feature vector of the core gene sequence; 3) Generate a feature matrix based on the feature vectors obtained in step 2) and the natural clustering results, and use the clustering evaluation method to quantitatively calculate the diversity of the pathogen's evolutionary direction to obtain an evaluation index for the pathogen's evolutionary direction diversity; The method for extracting the feature vector includes the following steps: 21) Distance Measure based on k-tuple (DMk) is used to calculate the position and number of occurrences of k-tuples in the sequence. When k=3, the k-tuple is regarded as a codon; the position of each k-tuple is recorded as ,in Representative The position of the k-tuple that appears the first time, , is the number of times the k-tuple appears in the sequence; 22) Calculate the intervals between the positions where k tuples appear, and combine all the intervals to be , its mathematical calculation formula is as follows: ; 23) Based on the interval information sequence of k tuples, calculate the interval sum and combine them into a sequence recorded as , its mathematical calculation formula is as follows: in, It is determined by the number and position of the k-tuple. At the same time, the number and position of the k-tuple are obtained by interval and sequence; 24) Calculate the Shannon entropy of the sequence based on the interval and probability, and define a discrete probability distribution based on the interval and ,in The mathematical formula for this is as follows: Then the Shannon entropy can be calculated: Shannon entropy reflects the number and position information of a k-tuple, which is used as the feature of a k-tuple in the virus sequence; 25) Repeat steps 21)-24) for each k-tuple to extract its features, and finally get a feature vector, which is recorded as , where , this vector is the feature vector extracted from the sequence; The calculation method of the pathogen evolutionary direction diversity evaluation index in step 3) includes the following steps: 31) The attribute labels of pathogens are regarded as natural clustering results. The sequence features of the same type of pathogens are combined into a class feature matrix. The intra-class variance is used to calculate the intra-cluster density of the feature matrix. The calculation formula is as follows: in, represents the number of categories, Indicates the All samples in a class, Indicates the The cluster center corresponding to each host, represents the dimension of the feature vector, Represents the feature vector No. bit vector value, Represents the feature vector No. bit vector value; 32) Use the inter-class variance to calculate the inter-cluster separation of the feature matrix. The calculation formula is as follows: in, represents the number of categories, Indicates the The number of samples in a class, represents the number of all samples, represents the global center, represents the dimension of the feature vector, represents the global center vector No. bit vector value, Represents the center vector within the class No. bit vector value; 33) Taking into account the influence of intra-cluster compactness, inter-cluster separation and the number of sample clusters, the calculation formula for the evolutionary direction diversity evaluation index is: in, It is an evaluation index of the diversity of the evolutionary direction of pathogens.

2. The method for quantitatively evaluating the diversity of pathogen evolutionary directions according to claim 1, characterized in that: The core gene sequence in step 1) is a specific gene sequence or a full-length genome sequence having core characteristics of the pathogen, and the sequence type includes a nucleic acid sequence or an amino acid sequence.

3. Use of the method for quantitatively assessing the diversity of pathogen evolutionary directions as described in any one of claims 1-2 in the preparation of products for evaluating the risk of pathogen transmission and spillover.

4. The use according to claim 3, characterized in that The standard for evaluating the risk of pathogen transmission spillover is that the larger the value of the pathogen's evolutionary direction diversity evaluation index, the more chaotic the category distribution, the higher the evolutionary direction diversity, and the higher the risk of pathogen transmission spillover; the smaller the value of the pathogen's evolutionary direction diversity evaluation index, the greater the difference between categories, the more concentrated the evolutionary directions, and the lower the risk of pathogen transmission spillover.

5. A device for evaluating the risk of pathogen spread and spillover, comprising a computer program, characterized in that: When the computer program is executed by a processor, the steps described in any one of claims 1 to 2 are implemented to obtain an evaluation index of the diversity of the evolutionary direction of the pathogen; and the risk of pathogen transmission spillover is evaluated based on the size of the evaluation index of the diversity of the evolutionary direction of the pathogen.

6. The device according to claim 5, characterized in that The standard for evaluating the risk of pathogen transmission spillover is that the larger the value of the pathogen's evolutionary direction diversity evaluation index, the more chaotic the category distribution, the higher the evolutionary direction diversity, and the higher the risk of pathogen transmission spillover; the smaller the value of the pathogen's evolutionary direction diversity evaluation index, the greater the difference between categories, the more concentrated the evolutionary directions, and the lower the risk of pathogen transmission spillover.

Citation Information

Patent Citations

  • Analysis method for diversity of tick-borne pathogens

    CN114058716A

  • Non-evolutionary tree-dependent segmented RNA virus reconfiguration method

    CN115910377A