Method for quantitatively evaluating diversity of pathogen evolution direction

By performing feature extraction and cluster evaluation of pathogen genotype data, quantitatively assessing its evolutionary direction diversity, the problems of low analysis efficiency and insufficient results accuracy in the existing technology are solved, and efficient and accurate assessment of pathogen transmission risk is achieved.

CN120015105AActive Publication Date: 2025-05-16ACADEMY OF MILITARY MEDICAL SCIENCES
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510075198.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

The prior art is difficult to effectively quantify the diversity of pathogen evolution directions, resulting in low analysis efficiency, small data volume, insufficient results accuracy, and high calculation difficulty.

Method used

By collecting genotype data of pathogens, using feature extraction methods to generate feature matrix, and using cluster evaluation methods to quantify the diversity of pathogen evolution directions to generate evolution direction diversity evaluation indicators.

Benefits of technology

A rapid, concise and intuitive diversity assessment of pathogen evolution directions is achieved, and the analysis efficiency and accuracy are improved, and the risk of spillover of pathogen transmission can be effectively tracked.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure SMS_9
    Figure SMS_9
Patent Text Reader

Abstract

The invention discloses a method for quantitatively evaluating pathogen evolution direction diversity, which comprises the following steps: 1) collecting core gene sequence data of pathogens based on specified pathogen types, and performing quality control on the collected data to obtain high-quality sequence data; 2) performing feature extraction on the collected sequence data by using Shannon entropy to obtain a feature vector of each sequence; and 3) based on the extracted feature data, constructing a feature matrix, and calculating a pathogen evolution direction diversity score by using a clustering evaluation calculation method. According to the method, by extracting the sequence features, evolution direction diversity evaluation and propagation overflow risk analysis can be rapidly and efficiently realized without evolutionary tree analysis or modeling training, and technical support and reference can be provided for subsequent pathogen genetic analysis, related epidemic situation detection and monitoring, and specific medicine and antibody development; and the related application is very wide.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biotechnology, and in particular to a method for calculating the diversity of pathogen evolution directions based on a clustering quality evaluation method. Background Art

[0002] Pathogens are widely present in nature and are diverse, and are one of the important factors threatening human life and health. Different pathogens have different evolutionary characteristics and transmission patterns. How to quickly identify and perceive the evolutionary direction of pathogens is an important problem facing current biosafety defense and disease control.

[0003] The evolutionary direction of pathogens is often related to their living and spreading environment. Among them, the more important influencing factors include the geographical distribution of pathogens and the distribution of hosts. Taking geographical distribution as an example, some pathogens can spread in multiple regions and spread across regions, which shows that the evolutionary direction of pathogens in terms of geographical distribution tends to be diversified and there is a risk of spillover, such as avian influenza virus. Some pathogens only spread in specific areas and rarely spread across regions, which shows that the evolutionary direction of pathogens in terms of geographical distribution tends to be stable and the risk of spillover is relatively small, such as Ebola virus and Lassa virus.

[0004] Quantitative assessment of the evolutionary direction of pathogens can not only help researchers and medical staff to gain a deeper understanding of the internal evolutionary laws of pathogens, but also predict the risk of future spread of pathogens, and provide theoretical basis and technical support for whether cities or countries around the epidemic area should upgrade or downgrade their defenses.

[0005] At present, the perception methods for pathogen evolution generally include gene data collection, label information annotation, multiple sequence alignment, evolutionary analysis, cluster analysis, geographic phylogenetic analysis, etc. The analysis process relies on manual subjective judgment, the analysis process is complex, the calculation is difficult, and there are many problems such as low analysis efficiency, small amount of analyzable data, and insufficient accuracy of results.

[0006] Assessing the diversity of pathogen evolutionary directions is essentially to detect whether the genetic evolutionary characteristics of pathogens in different regions (or hosts) are significantly different. If the genetic evolutionary characteristics of pathogens vary greatly, each region (or host) has its own evolutionary characteristics, indicating that the evolutionary direction of pathogens is relatively stable and concentrated. On the contrary, if the genetic evolutionary characteristics of pathogens vary little, it means that the evolutionary direction of pathogens is relatively diverse, and there is a risk of spillover.

[0007] Using interdisciplinary methods, the biological problem of studying the diversity of pathogen evolutionary directions can be transformed into a computational problem of estimating sample distribution. The pathogen sequence features are data samples, and the pathogen distribution attributes (geographic information or host information) are sample labels. The analysis of the diversity of evolutionary directions can be transformed into an assessment of the degree of chaos in the distribution of multi-category data. The more chaotic the distribution, the more diverse the evolutionary directions, and vice versa. Summary of the invention

[0008] The purpose of the present invention is to provide a method for quantitatively evaluating the diversity of pathogen evolutionary directions.

[0009] The present invention provides a method for quantitatively evaluating the diversity of pathogen evolutionary directions. For a given pathogen, the genotype data of the pathogen is collected, and features of the sequence are extracted using a related method. A feature matrix is ​​generated based on the extracted feature vectors and natural clustering results, and a clustering evaluation method is used to quantitatively calculate the diversity of pathogen evolutionary directions.

[0010] Further, the following steps are included: 1) Collect core gene sequences of target pathogens; 2) Perform feature extraction on each sequence obtained in step 1) to obtain the feature vector of the core gene sequence; 3) Generate a feature matrix based on the feature vector obtained in step 2) and the natural clustering result, use the clustering evaluation method to quantitatively calculate the diversity of pathogen evolutionary directions, and obtain an evaluation index for the diversity of pathogen evolutionary directions.

[0011] Furthermore, the core gene sequence in step 1) is a specific gene sequence or a full-length genome sequence having core characteristics of the pathogen, and the sequence type includes a nucleic acid sequence or an amino acid sequence.

[0012] 4. The method for quantitatively evaluating the diversity of pathogen evolutionary directions according to claim 2, characterized in that the method for extracting the feature vector comprises the following steps: 21) Distance Measure based on k-tuple (DMk) is used to calculate the position and occurrence information of k-tuples in the sequence. When k=3, k-tuples can be regarded as codons. The position of each k-tuple is recorded as ,in Representative The position of the k-tuple that appears for the first time, , is the number of times the k-tuple appears in the sequence; 22) Calculate the intervals between the positions where k-tuples appear, and combine all the intervals to be , its mathematical calculation formula is as follows: ; 23) According to the interval information sequence of k-tuples, calculate its interval sum and combine it into a sequence which can be recorded as , its mathematical calculation formula is as follows:

[0013] in, It is determined by the number and position of the k-tuple. At the same time, the number and position of the k-tuple can also be obtained by intervals and sequences; 24) Calculate the Shannon entropy of the sequence based on the interval and probability, and define a discrete probability distribution based on the interval and ,in The mathematical formula for this is as follows:

[0014] Then the Shannon entropy can be calculated:

[0015] Shannon entropy reflects the number and position information of a k-tuple, which is used as the feature of a k-tuple in the virus sequence; 25) Repeat steps 21)-24) for each k-tuple to extract its features, and finally get a feature vector, recorded as , where , this vector is the feature vector extracted from the sequence.

[0016] Furthermore, the calculation method of the pathogen evolutionary direction diversity evaluation index in step 3) includes the following steps: 31) The attribute labels of pathogens are regarded as natural clustering results, and the sequence features of the same type of pathogens are combined into a class feature matrix. The intra-class variance is used to calculate the intra-cluster density of the feature matrix. The calculation formula is as follows:

[0017] in, represents the number of categories, Indicates All samples in the class, Indicates The cluster center corresponding to each host, represents the dimension of the feature vector, Represents the feature vector No. A vector value of bits, Represents the feature vector No. A vector value of bits; 32) Use the inter-class variance to calculate the inter-cluster separation of the feature matrix. The calculation formula is as follows:

[0018] in, represents the number of categories, Indicates The number of samples in a class, represents the number of all samples, represents the global center, represents the dimension of the feature vector, represents the global center vector No. A vector value of bits, Represents the center vector within the class No. A vector value of bits; 33) Taking into account the influence of intra-cluster compactness, inter-cluster separation and the number of sample clusters, the calculation formula of the evolutionary direction diversity evaluation index is:

[0019] in, It is an evaluation index for the diversity of the evolutionary direction of pathogens.

[0020] The above-mentioned method for quantitatively evaluating the diversity of pathogen evolutionary directions is used in the preparation of products for evaluating the risk of pathogen spread and spillover.

[0021] Furthermore, the standard for evaluating the risk of pathogen transmission spillover is that the larger the value of the pathogen's evolutionary direction diversity evaluation index, the more chaotic the category distribution, the higher the evolutionary direction diversity, and the higher the risk of pathogen transmission spillover; the smaller the value of the pathogen's evolutionary direction diversity evaluation index, the greater the difference between categories, the more concentrated the evolutionary directions, and the lower the risk of pathogen transmission spillover.

[0022] A product for evaluating the risk of pathogen transmission spillover, including a computer program, characterized in that when the computer program is executed by a processor, the steps of the method for quantitatively evaluating the diversity of pathogen evolutionary directions are implemented to obtain an evaluation index for the diversity of pathogen evolutionary directions; and the risk of pathogen transmission spillover is evaluated based on the size of the evaluation index for the diversity of pathogen evolutionary directions.

[0023] Furthermore, the standard for evaluating the risk of pathogen transmission spillover is that the larger the value of the pathogen's evolutionary direction diversity evaluation index, the more chaotic the category distribution, the higher the evolutionary direction diversity, and the higher the risk of pathogen transmission spillover; the smaller the value of the pathogen's evolutionary direction diversity evaluation index, the greater the difference between categories, the more concentrated the evolutionary directions, and the lower the risk of pathogen transmission spillover.

[0024] The present invention uses an interdisciplinary approach to regard the distribution attribute (geography or host) labels as the result of "natural clustering" of pathogens, and uses a clustering evaluation algorithm to evaluate the clustering quality of "natural clustering" to determine the degree of chaos in the data distribution, thereby completing a quantitative evaluation of the diversity of pathogen evolutionary directions.

[0025] Compared with the prior art, the advantages of the present invention are: 1. The calculation is simple and intuitive, a large amount of data can be analyzed, the analysis efficiency is high, and the diversity of the evolutionary direction of pathogens can be effectively tracked and quantified; 2. The clustering evaluation algorithm is used to evaluate the diversity of evolutionary directions, which is highly innovative. By utilizing the advantages of interdisciplinary technology, it can effectively avoid the result deviation caused by inconsistent pathogen sequence quality and the increased difficulty of calculation, and can cover more data information, effectively handle noise and missing data, and have higher analysis accuracy; 3. The method has the characteristics of full-process automated calculation, which avoids deviations and misleadings caused by manual measurement, and the analysis results have strong stability and reliability; 4. According to the method, the diversity of pathogen evolutionary directions is quantitatively evaluated, which can effectively perceive the geographical (or host) spread and spillover situation of pathogens, provide technical support and reference for pathogen identification and epidemic prevention and control, and has strong practicality.

[0026] In summary, the present invention realizes the quantitative assessment of the diversity of pathogen evolutionary directions through the process methods of pathogen sequence collection quality control, feature extraction, cluster evaluation, etc. The assessment speed is fast, the method is simple and effective, and the spread and spillover trends and risks of pathogens can be systematically and comprehensively measured. The invention avoids the complexity problems existing in the classical methods, such as low computational efficiency, small amount of processable data, and difficulty in interpreting analysis results, and overcomes the influence of interference factors such as data noise and quantity scale. The method is mainly reflected in the rapid assessment of the diversity of pathogen evolutionary directions, and the risk analysis of pathogen transmission can be realized without evolutionary tree analysis or modeling training. Utilizing this technical advantage will provide technical methods and references for pathogen prevention and epidemic risk prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 Computational flow chart for quantitatively assessing the diversity of pathogen evolutionary directions.

[0028] Figure 2 This is a diagram of the quantitative analysis of the diversity of the evolutionary direction of the Ebola virus and the verification of the results. DETAILED DESCRIPTION SUMMARY OF THE INVENTION The method provided by the present invention is to collect the gene sequence and annotation information of the pathogen for a given type of pathogen, use relevant methods to extract features from the sequence, and then use a clustering evaluation algorithm to calculate the extracted feature vector to complete a quantitative evaluation of the diversity of the pathogen's evolutionary direction. Specifically, the present invention first collects the gene sequence data of the pathogen based on the specified pathogen type, completes the attribute label information annotation, and performs quality control on the collected data to obtain a high-quality sequence data set; then, the collected pathogen sequence data is feature extracted to obtain a feature vector for each sequence; finally, the sequence feature vectors are classified and combined according to the label information to obtain a feature matrix for each type of pathogen, and a clustering evaluation algorithm is used to quantitatively evaluate the degree of chaos of the pathogen sequence data to complete a quantitative evaluation of the diversity of the evolutionary direction.

[0030] The above method comprises the following steps: 1. Collect and download the core gene sequences of target pathogens.

[0031] 1.1 Collection of pathogen core gene sequence data For pathogens, the method of the present invention requires collecting and collating as much core gene sequence data of target pathogens as possible to more accurately measure the diversity of the evolutionary direction of target pathogens. Through database search (such as NCBI GenBank, GISAID and other databases), literature review, biological experiments and other means, collect and collate as comprehensive core gene sequence data of pathogens as possible. The core gene is generally a gene related to the host adaptability of the pathogen, which can be a certain gene of the pathogen or the full length of the pathogen gene. The collected pathogen core gene sequence data should generally be no less than 100.

[0032] 1.2 Pathogen sequence data quality control and attribute label information annotation In order to ensure the stability and effectiveness of the experimental results, it is necessary to perform quality control on the core gene sequence data of the pathogen and delete sequences that are insufficient in length, contain abnormal characters, or are repeated and redundant.

[0033] Determine the type of pathogen attribute label, and select geographic distribution, host distribution or other attribute information as the attribute label, and label each sequence. Delete the sequences with unclear labels to complete data cleaning and screening. The screened data set is used as the benchmark data set for subsequent analysis and processing.

[0034] 2 Extract features from each sequence obtained in step 1) as follows: 2.1 Distance Measure based on k-tuple (DMk) is used to calculate the position and occurrence information of k-tuples in the sequence.k =3, the k-tuple can be regarded as a codon in the nucleotide sequence. Therefore, the value of k in DMk can be various, but in this method, it is generally taken as 3 because it involves nucleic acid sequences and codons. The position of each k-tuple is recorded as ,in Representative The position where the k-tuple appears for the first time. is the number of times the k-tuple appears in the sequence. It should be noted that this method supports the study of both nucleic acid sequences and amino acid sequences. According to different sequence types, the k value can be adjusted to capture more accurate sequence features.

[0035] 2.2 Calculate the intervals between the positions where k-tuples appear, and combine all the intervals to , its mathematical calculation formula is as follows:

[0036] 2.3 Based on the interval information sequence of k-tuples, calculate the interval sum and combine them into a sequence which can be recorded as , its mathematical calculation formula is as follows:

[0037] in, It is determined by the number and position of the k-tuple. At the same time, the number and position of the k-tuple can also be obtained through intervals and sequences.

[0038] 2.4 Calculate the Shannon entropy of the sequence based on the interval and probability. Define a discrete probability distribution based on the interval and ,in The mathematical formula for this is as follows:

[0039] Then the Shannon entropy can be calculated:

[0040] Shannon entropy reflects the number and position information of a k-tuple, which is used as the characteristic of a k-tuple in the pathogen sequence.

[0041] 2.5 Repeat steps 2.1-2.4 for each k-tuple to extract its features, and finally get a feature vector, denoted as ,in, ,when k =3, t = 64. This vector is the feature vector of the pathogen sequence.

[0042] 3 Use the clustering evaluation algorithm to calculate and evaluate the feature vectors of all sequences obtained in step 2. The specific steps are as follows: 3.1 The sequences with the same attribute labels in the dataset are regarded as the same type, and the feature vectors of the sequences of the same type are combined into a matrix as the class feature matrix.

[0043] 3.2 Use the intra-class variance to calculate the intra-cluster closeness of the feature matrix (intra-cluster closeness), and the calculation formula is as follows:

[0044] in, represents the number of categories, Indicates All samples in the class, Indicates The cluster center corresponding to each host, represents the dimension of the feature vector, Represents the feature vector No. A vector value of bits, Represents the feature vector No. A vector value of bits.

[0045] 3.3 Use the inter-class variance to calculate the inter-cluster separation of the feature matrix. The calculation formula is as follows:

[0046] in, represents the number of categories, Indicates The number of samples in a class, represents the number of all samples, represents the global center, represents the dimension of the feature vector, represents the global center vector No. A vector value of bits, Represents the center vector within the class No. A vector value of bits.

[0047] 3.4 Taking into account the influence of intra-cluster compactness, inter-cluster separation and the number of sample clusters, the calculation formula of the evolutionary direction diversity evaluation index is:

[0048] in, It is an evaluation index for the diversity of the evolutionary direction of pathogens. The larger the value, the more chaotic the category distribution, the higher the diversity of evolutionary directions, and the higher the risk of spillover; The smaller the value, the greater the difference between categories, the more concentrated the evolutionary direction, and the lower the risk of spillover.

[0049] Through the above three steps, a quantitative score of the diversity of the evolutionary direction of the pathogen to be tested can be obtained.

[0050] The present invention is further described in detail below in conjunction with specific embodiments, and the examples provided are only for illustrating the present invention, rather than for limiting the scope of the present invention. The examples provided below can be used as a guide for further improvements by those of ordinary skill in the art, and do not constitute a limitation of the present invention in any way.

[0051] The experimental methods in the following examples, unless otherwise specified, are all conventional methods, and are performed according to the techniques or conditions described in the literature in the field or according to the product instructions. The materials, reagents, etc. used in the following examples, unless otherwise specified, can all be obtained from commercial channels.

[0052] Unless otherwise specified, the quantitative tests in the following examples were performed three times and the results were averaged.

[0053] Example 1: Quantitative evaluation and trend tracking analysis of the diversity of the evolutionary direction of Ebola virus This embodiment establishes a method for quantitatively calculating the diversity of virus evolutionary directions based on a clustering evaluation algorithm, quantitatively calculates and tracks the changing trends of the diversity of the evolutionary directions of the Ebola virus (Ebola) in different time intervals, and completes the risk assessment of transmission spillover.

[0054] 1. Collection and preprocessing of viral genome data 1.1 Collection of viral genome nucleic acid sequence data For viruses, the method of the present invention requires collecting and organizing as much genome sequence data of the virus as possible to more accurately measure the diversity of the evolutionary direction of the virus. In this embodiment, the public Ebola virus genome data was downloaded from the NCBI GenBank online database (https: / / www.ncbi.nlm.nih.gov). The GP protein gene was selected as the Ebola virus core gene for analysis and research. A total of 3850 DNA sequence data were obtained as the virus genome data set for this case.

[0055] 1.2 Virus sequence data preprocessing To ensure the stability and effectiveness of the experimental results, this embodiment performs quality control on the virus sequence data. The length of the GP protein DNA sequence of the Ebola virus reference strain is 2031. Therefore, sequences with a length less than 90% of the reference sequence length are deleted, that is, sequences with a sequence length less than 1818 are deleted in the virus genome data set. In addition, sequences with a large number of abnormal characters or repeated redundant sequences in the data set are also deleted. After quality control, the virus genome data set contains a total of 730 sequences.

[0056] The geographic information of Ebola virus was selected as the attribute label for quantitative evaluation of evolutionary diversity. The geographic attribute labels in the virus genome dataset include 13 countries and regions: "Zaire", "Guinea", "Uganda", "Sierra Leone", "Gabon", "Liberia", "Germany", "Nigeria", "USA", "United Kingdom", "Italy", "Switzerland", "Mali", etc. In the subsequent analysis, these 13 regions are natural clustering labels.

[0057] 2. Extract virus feature vector The gene sequences of all virus strains obtained in step 1 are used to extract feature vectors according to the following method: 2.1 Use k-tuple-based distance measurement to calculate the position and occurrence information of all k-tuples in each sequence. In this embodiment, let k=3 , treat k-tuples as codons to facilitate genomic genetic analysis. For each k-tuple, count the number of times it appears in the sequence and its position information. Recorded as ,in Representative The position of the k-tuple that appears for the first time, , is the number of times the k-tuple appears in the sequence.

[0058] 2.2 For each k-tuple, calculate the interval between adjacent occurrences, denoted as , its mathematical calculation formula is as follows:

[0059] 2.3 Based on the interval information of k-tuples, calculate the interval sum, recorded as , its mathematical calculation formula is as follows:

[0060] in, Determined by the number and position of the k-tuple. 2.4 Calculate the Shannon entropy of the sequence. Define a discrete probability distribution ,in The mathematical formula for this is as follows:

[0061] Then the Shannon entropy can be calculated:

[0062] Shannon entropy reflects the number and position information of a k-tuple, which is used as the characteristic of the k-tuple of the viral gene sequence.

[0063] 2.5 Repeat steps 2.1-2.4 for each k-tuple to extract its features, and finally get a feature vector, denoted as , where .because ,so , this vector is the feature vector of the virus sequence.

[0064] All sequence feature vectors under the natural clustering label are combined together to obtain the feature vector matrix of the label category.

[0065] 3. Quantitative evaluation and calculation of the diversity of viral evolutionary directions Calculate the feature vectors of all sequences obtained in step 2. The specific steps are as follows: 3.1 Consider the attribute labels of the virus as natural clustering results, and use the intra-class variance to calculate the intra-cluster density of the feature matrix. The calculation formula is as follows:

[0066] in, represents the number of categories, Indicates All samples in the class, Indicates The cluster center corresponding to each host, represents the dimension of the feature vector, Represents the feature vector No. A vector value of bits, Represents the feature vector No. A vector value of bits.

[0067] 3.2 Use the inter-class variance to calculate the inter-cluster separation of the feature matrix. The calculation formula is as follows:

[0068] in, represents the number of categories, Indicates The number of samples in a class, represents the number of all samples, represents the global center, represents the dimension of the feature vector, represents the global center vector No. A vector value of bits, Represents the center vector within the class No. A vector value of bits.

[0069] 3.3 Taking into account the influence of intra-cluster compactness, inter-cluster separation and the number of sample clusters, the calculation formula of the evolutionary direction diversity evaluation index is:

[0070] in, It is an evaluation index for the diversity of the evolutionary direction of pathogens. The larger the value, the more chaotic the category distribution, the higher the diversity of evolutionary directions, and the higher the risk of spillover; The smaller the value, the greater the difference between categories, the more concentrated the evolutionary direction, and the lower the risk of spillover.

[0071] Through the above three steps, we can obtain the quantitative score of the diversity of the evolutionary direction of Ebola virus in terms of geographical attributes.

[0072] In order to track the changing trend of the diversity of the evolutionary direction of Ebola virus over time, the virus dataset can be divided into four datasets according to the year information: "2005 and before virus dataset", "2010 and before virus dataset", "2015 and before virus dataset", and "2020 and before virus dataset". The quantitative scores of the diversity of evolutionary direction were performed according to the above three steps, and the results are shown in Table 1. If only the current diversity of the evolutionary direction of Ebola virus is evaluated, all current sequences can be used for calculation, and there is no need to divide the virus genome dataset.

[0073] Table 1 Quantitative assessment and tracking of the diversity of Ebola virus evolution

[0074] The calculation results of the embodiment show that the evolutionary diversity scores of the "virus datasets in 2005 and before" and "virus datasets in 2010 and before" are low, indicating that the risk of geographical spread of Ebola virus before 2010 was low and the diversity of evolutionary direction was relatively simple. However, the evolutionary diversity scores of the "virus datasets in 2015 and before" and "virus datasets in 2020 and before" gradually increase, indicating that the risk of geographical spread of Ebola virus has increased since 2015, and the diversity of evolutionary direction has become more complex, which is consistent with the actual news reports and the spread of the epidemic.

[0075] To further verify the accuracy of the method of the present invention, the present embodiment uses the geographic phylogenetic method to perform experimental analysis on the virus. The geographic phylogenetic method can simulate the historical transmission route of the virus, wherein if there are more cross-regional transmission routes, it means that the risk of virus transmission spillover is greater, and the diversity of evolutionary direction is more complex, otherwise it means that the risk of virus transmission spillover is smaller, and the diversity of evolutionary direction is simpler. The geographic phylogenetic method is one of the main analysis methods for determining the trend of virus transmission at present. According to the experimental results of the data of geographic phylogenetics, it is found that before 2010, Ebola virus only had one cross-regional transmission route, indicating that the risk of transmission spillover is low. After 2015, 5 cross-regional transmission routes appeared in Ebola virus, indicating that the risk of transmission spillover increased, which is consistent with the results of quantitative evaluation of the method of the present invention. At the same time, before 2010, there were only 2 main geographical transmission areas of Ebola virus, and around 2020, the main geographical transmission areas spread to 13, indicating that Ebola virus did have a transmission spillover phenomenon, which is also consistent with the analysis results of the method of the present invention. The effectiveness and accuracy of the method of the present invention are illustrated.

[0076] The present invention has been described in detail above. It will be apparent to those skilled in the art that the present invention may be implemented in a wide range under equivalent parameters, concentrations and conditions without departing from the spirit and scope of the present invention and without the need for unnecessary experimentation. Although the present invention provides specific embodiments, it should be understood that further improvements may be made to the present invention. In short, according to the principles of the present invention, this application intends to include any changes, uses or improvements to the present invention, including changes made by conventional techniques known in the art that depart from the scope disclosed in this application. Applications of some of the basic features may be made within the scope of the following appended claims.

Claims

1. A method for quantitatively evaluating the diversity of pathogen evolutionary directions, characterized in that: For a given pathogen, the genotype data of the pathogen is collected, and the sequence is feature extracted using relevant methods. Then, a feature matrix is ​​generated based on the extracted feature vectors and natural clustering results, and the clustering evaluation method is used to quantitatively calculate the diversity of the pathogen's evolutionary direction.

2. The method for quantitatively evaluating the diversity of pathogen evolutionary directions according to claim 1, characterized in that: The steps include: 1) Collect core gene sequences of target pathogens; 2) Extract features from each sequence obtained in step 1) to obtain the feature vector of the core gene sequence; 3) Generate a feature matrix based on the feature vector obtained in step 2) and the natural clustering result, use the clustering evaluation method to quantitatively calculate the diversity of pathogen evolutionary directions, and obtain an evaluation index for the diversity of pathogen evolutionary directions.

3. The method for quantitatively evaluating the diversity of pathogen evolutionary directions according to claim 2, characterized in that: The core gene sequence in step 1) is a specific gene sequence or a full-length genome sequence having core characteristics of the pathogen, and the sequence type includes a nucleic acid sequence or an amino acid sequence.

4. The method for quantitatively evaluating the diversity of pathogen evolutionary directions according to claim 2, characterized in that: The method for extracting the feature vector comprises the following steps: 21) Distance Measure based on k-tuple (DMk) is used to calculate the position and occurrence information of k-tuples in the sequence. When k=3, k-tuples can be regarded as codons; the position of each k-tuple is recorded as ,in Representative The position of the k-tuple that appears for the first time, , is the number of times the k-tuple appears in the sequence; 22) Calculate the intervals between the positions where k-tuples appear, and combine all the intervals to be , its mathematical calculation formula is as follows: ; 23) According to the interval information sequence of k-tuples, calculate its interval sum and combine it into a sequence which can be recorded as , its mathematical calculation formula is as follows: in, It is determined by the number and position of the k-tuple. At the same time, the number and position of the k-tuple can also be obtained by intervals and sequences; 24) Calculate the Shannon entropy of the sequence based on the interval and probability, and define a discrete probability distribution based on the interval and ,in The mathematical formula for this is as follows: Then the Shannon entropy can be calculated: Shannon entropy reflects the number and position information of a k-tuple, which is used as the feature of a k-tuple in the virus sequence; 25) Repeat steps 21)-24) for each k-tuple to extract its features, and finally get a feature vector, recorded as , where , this vector is the feature vector extracted from the sequence.

5. The method for quantitatively evaluating the diversity of pathogen evolutionary directions according to claim 2, characterized in that: The calculation method of the pathogen evolutionary direction diversity evaluation index in step 3) comprises the following steps: 31) The attribute labels of pathogens are regarded as natural clustering results, and the sequence features of the same type of pathogens are combined into a class feature matrix. The intra-class variance is used to calculate the intra-cluster density of the feature matrix. The calculation formula is as follows: in, represents the number of categories, Indicates All samples in the class, Indicates The cluster center corresponding to each host, represents the dimension of the feature vector, Represents the feature vector No. A vector value of bits, Represents the feature vector No. A vector value of bits; 32) Use the inter-class variance to calculate the inter-cluster separation of the feature matrix. The calculation formula is as follows: in, represents the number of categories, Indicates The number of samples in a class, represents the number of all samples, represents the global center, represents the dimension of the feature vector, represents the global center vector No. A vector value of bits, Represents the center vector within the class No. A vector value of bits; 33) Taking into account the influence of intra-cluster compactness, inter-cluster separation and the number of sample clusters, the calculation formula of the evolutionary direction diversity evaluation index is: in, It is an evaluation index for the diversity of the evolutionary direction of pathogens.

6. Application of the method for quantitatively evaluating the diversity of pathogen evolutionary directions as described in any one of claims 1-5 in the preparation of products for evaluating the risk of pathogen spread and spillover.

7. The use according to claim 6, characterized in that: The standard for evaluating the risk of pathogen transmission spillover is that the larger the value of the pathogen's evolutionary direction diversity evaluation index, the more chaotic the category distribution, the higher the evolutionary direction diversity, and the higher the risk of pathogen transmission spillover; the smaller the value of the pathogen's evolutionary direction diversity evaluation index, the greater the difference between categories, the more concentrated the evolutionary directions, and the lower the risk of pathogen transmission spillover.

8. A product for evaluating the risk of pathogen spread and spillover, comprising a computer program, characterized in that: When the computer program is executed by a processor, the steps described in any one of claims 1 to 5 are implemented to obtain an evaluation index of the diversity of the evolutionary direction of the pathogen; and the risk of pathogen transmission spillover is evaluated based on the size of the evaluation index of the diversity of the evolutionary direction of the pathogen.

9. The product according to claim 8, characterized in that The standard for evaluating the risk of pathogen transmission spillover is that the larger the value of the pathogen's evolutionary direction diversity evaluation index, the more chaotic the category distribution, the higher the evolutionary direction diversity, and the higher the risk of pathogen transmission spillover; the smaller the value of the pathogen's evolutionary direction diversity evaluation index, the greater the difference between categories, the more concentrated the evolutionary directions, and the lower the risk of pathogen transmission spillover.

Citation Information

Patent Citations

  • Analysis method for diversity of tick-borne pathogens

    CN114058716A

  • Non-evolutionary tree-dependent segmented RNA virus reconfiguration method

    CN115910377A

  • Methods for identifying sequence motifs, and applications thereof

    US20090208955A1