Prediction of mutations using within-host deep sequencing data

A computational framework analyzing within-host deep sequencing data predicts key mutations by comparing fitness levels, addressing the limitations of population-level virus evolution models and enabling early intervention strategies.

WO2025214227A1PCT designated stage Publication Date: 2025-10-16THE CHINESE UNIVERSITY OF HONG KONG
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/086809
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-10
Filing Date
2025-04-02
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing techniques for modeling virus evolution focus on population-level analysis, failing to identify key mutations that provide survival advantages within individual hosts, which can lead to immune-resistant strains and disease spread.

Method used

A computational framework that analyzes within-host deep sequencing data to predict mutations with high fitness potential by comparing within-host and between-host fitness, enabling early identification of 'key' mutations using next-generation sequencing techniques.

Benefits of technology

Predicts key mutations several months before they become dominant, allowing for timely development of vaccines or treatments targeting these strains.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025086809_16102025_PF_FP_ABST
    Figure CN2025086809_16102025_PF_FP_ABST
Patent Text Reader

Abstract

A statistical model can be used to evaluate the potential of specific mutations to become dominant in a population of hosts by leveraging a contrast between a measure of within-host fitness and between-host fitness for a given mutation. The model can based on deep sequencing data that provides information variant frequencies within a single host. Deep sequencing across multiple hosts enables a determination of between-host detection frequency. Based on discrepancies between the within-host variant frequency and the between-host detection frequency, mutations with high fitness potential can be identified (or predicted) as potential "key" mutations before they become dominant across a host population.
Need to check novelty before this filing date? Find Prior Art

Description

PREDICTION OF MUTATIONS USING WITHIN-HOST DEEP SEQUENCING DATACROSS-REFERENCE TO RELATED APPLICATION

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 632,442, filed April 10, 2024, the disclosure of which is incorporated herein by reference.BACKGROUND

[0002] Viruses and other pathogens that can spread from host to host cause a significant number of illnesses worldwide. When a host is infected with a pathogen, the pathogen can replicate rapidly within the host, generating a pathogen population that can escape and infect others and that can also cause disease symptoms in the host. Rapidly replicating pathogens such as viruses often have error-prone replication mechanisms that generate mutations in a stochastic manner. The host’s immune response (possibly aided by antiviral medications) exerts selection pressure on the pathogen, leading to the evolution of immune-resistant strains that can exacerbate the spread of disease.

[0003] Understanding and predicting evolution of pathogens could potentially contribute to prevention of communicable diseases. For instance, given the ability to predict virus evolution, vaccines can be developed that target strains that are likely to become dominant in the future. While some predictive techniques based on population-level modeling have had some success, improved techniques would be desirable.SUMMARY

[0004] Certain embodiments of the present invention relate to computer-implemented techniques that can predict evolution of a virus or other pathogen. For instance, a statistical model described herein can be used to evaluate the potential of specific mutations to become dominant in a population of hosts by leveraging a contrast between a measure of within-host fitness and between-host fitness for a given mutation. In some embodiments, the model is based on deep sequencing using next-generation sequencing techniques, in which a large number of genomic fragments of pathogens in a given host are sequenced to generate multiple “raw reads” of nucleotides at various positions in the genome of the pathogen in the host. The within-host sequence data is analyzed to determine variant frequencies within the host, where a “variant” represents a particular nucleotide or amnio acid at a particular position in the genome. Deep sequencing across multiple hosts enables a determination of between-host detection frequency. Based on discrepancies between the within-host variant frequency and the between-host detection frequency, mutations with high fitness potential can be identified (or predicted) as potential “key” mutations before they become dominant across a host population. Similar techniques can be applied to other genetic-sequence data such as alleles, amino acid sequences, or the like.

[0005] The following detailed description, together with the accompanying drawings, will provide a better understanding of the nature and advantages of the claimed invention.BRIEF DESCRIPTION OF THE DRAWINGS

[0006] FIG. 1 shows a flow diagram of a process for predicting key mutations of a pathogen according to some embodiments.

[0007] FIG. 2 graphically illustrates operating principles of a prediction process according to some embodiments.

[0008] FIG. 3 shows a flow diagram of a process for estimation of probable source region for a key mutation of a pathogen according to some embodiments.

[0009] FIGs. 4 and 5 show results from a study in which processes according to some embodiments were applied to data for SARS-CoV-2.

[0010] FIGs. 6 and 7 show results from another study in which processes according to some embodiments were used to predict mutations of SARS-CoV-2.DETAILED DESCRIPTION

[0011] The following description of exemplary embodiments of the invention is presented for the purpose of illustration and description. It is not intended to be exhaustive or to limit the claimed invention to the precise form described, and persons skilled in the art will appreciate that many modifications and variations are possible. The embodiments have been chosen and described in order to best explain the principles of the invention and its practical applications to thereby enable others skilled in the art to best make and use the invention in various embodiments and with various modifications as are suited to the particular use contemplated.

[0012] Some embodiments described herein relate to a computational framework that can predict key mutations contributing to spread of pathogenic disease. In some embodiments, by leveraging deep sequencing data of the pathogen within a single host and across multiple hosts, advantageous mutations (i.e., mutations with high likelihood of becoming dominant strains) can be identified before they become prevalent in a host population. By evaluating the contrast of fitness demonstrated between the within-host and between-host levels, the model can identify mutations that have high potential to prevail (or become predominant) in future epidemics or spread of disease. Such mutations are referred to herein as “key” mutations.

[0013] To facilitate application of the predictions, some embodiments provide a computational approach to estimating the source region of emerging variants. As illustrated in examples below, techniques described herein can effectively predict key mutations using within-host deep sequencing data prior to their dominance at population-level. The average lead-time for predictions can be on the order of months, which may be sufficient to enable development of vaccines or treatments targeting strains having the predicted key mutation (s) before such strains become predominant.

[0014] The methodological framework provided according to some embodiments can be generally applied to any pathogen that shows a contrast between micro-and macro-level fitness. Applications of methods described herein include: (1) prediction of key mutations and strains of future major variants; (2) antigen optimization for infectious disease or cancer vaccines; (3) new antigen design; (4) risk evaluation for antiviral resistance based on in-host mutation distributions; (5) risk evaluation for vaccine break-through; and (6) delineation of geographic transmission routes of infectious disease pathogens.

[0015] Existing techniques for modeling of virus evolution perform analysis at the host-population level (also referred to as “between-host” level) , with one strain of the virus being isolated per host. Modeling of the evolutionary pattern of the virus is based on comparing strains across different hosts.

[0016] In contrast, techniques described herein incorporate analysis of variants in the virus population within a single host ( “within-host” level) . The analysis is based on the hypothesis that mutations that provide a survival advantage may demonstrate early signs of evolutionary fitness in a host as the virus population within a host experiences immune selection; accordingly, a within-host analysis of the virus population may reveal early information regarding fitness of particular mutations. The within-host virus population can be profiled by deep sequencing data of virus genome of subjects, using techniques known in the art. Bridging viral fitness information across the within-host and between-host levels can be used to identify mutations with high potential to become dominant (also referred to as “key mutations” ) . As illustrated below in an example using real data, key mutations can be identified (or predicted) several months before they become dominant at population level. Early prediction of key mutations can increase the lead time for preparing vaccine antigen, targeted antivirals, or other public health responses. Overview

[0017] When a host is infected with a pathogen, such as a virus, the virions replicate and increase into a large quantity, forming a virus population in the infected subject over a few hours or days. The error-prone replication mechanism of RNA viruses generates mutations in a stochastic manner. Under host immune pressure, the mutations that can survive better grow into larger quantities and show higher fitness. Advantageous mutations, such as those enabling immune escape, may not be sufficient for viral survival -epistatic and complementary mutations are also important for the shaping of emerging variants. This means that key mutations may appear in individual hosts well before a lineage becomes established in a population.

[0018] The raw reads of next generation sequencing (NGS) data provide an opportunity to investigate mutation advantages within an individual host, since even de novo mutations can be captured as sequencing depth increases. Accordingly, some embodiments use the genetic fragments of raw reads as random samples for estimating variant frequency (VF) within a host. The VF reflects relative fitness of different mutations within an individual host.

[0019] According to some embodiments, the VF can be used in a method that contrasts the disparity of mutation fitness between the within-host level and the between-host level. Mutations that show high fitness in the within-host virus population and that have not (yet) reached dominance at the between-host level are considered as mutations having high potential to prevail in the population at a later time (sometimes referred to herein as “key” mutations or “high-potential” mutations) . The disparity between fitness at the within-host level and between-host level will eventually diminish for successful mutations, and the time that elapses between the early identification of a key mutation within-host to its reaching dominance in the host population is referred to herein as the “prediction lead-time. ”

[0020] Certain methods described herein can be used to predict high-potential mutations that may become dominant at a future time. The predicted key mutations can be used for various purposes, including: (1) informing vaccine manufacturers so that vaccines can be updated by the time a predicted key mutation becomes dominant; (2) inferring which viral strains are most likely to become prevalent at population level; and (3) informing decisions about epitopes to be included in novel types of antigens, such as recombinant virus-like particles. In any of these applications, additional filtering approaches may be applied to facilitate the selection of key mutations, epitopes, and strains; examples of such filtering approaches include examination of viral crystal structural change associated with the predicted key mutations, experimental validation of antigenic properties of predicted mutants, as well as other methods. In addition or instead, predicted mutations can be used in screening for antivirals. For example, if an antiviral agent targets a genome position that contains a predicted key mutation site, then the efficacy of the antiviral might be reduced due to evolving antiviral resistance. As yet another example, risk of vaccine break-through can be assessed based on predicted key mutations. For instance, efficacy of a vaccine against predicted key mutations can be assessed. Public health measures can be enacted when the risk of vaccine break-through is high.

[0021] While embodiments described herein focus on viral genomes at the within-host and between-host levels, the framework can be generally applied to identify mutations of high fitness potential between two levels of any living environments, such as cells, tissues, or organs, for infectious disease pathogens or cancer cells, and can be used to inform vaccines or drug design. Prediction of Key Mutations

[0022] According to some embodiments, prediction of key mutations combines an analysis of within-host deep sequencing data and an analysis of between-host fitness of particular mutations. FIG. 1 shows a flow diagram of a process 100 for prediction of key mutations according to some embodiments. Process 100 can be a computer-implemented process, using the resources of the computer to operate on large data sets that exceed human capability. Process 100 incorporates a within-host analysis (block 110) and a between-host analysis (block 120) . Process 100 can be performed for a particular time (or time period) t, within which samples of virus genome are obtained from some number (ns) of infected subjects. It is contemplated that process 100 can be repeated at different times to detect trends. For instance, the time period t can correspond to a particular week. The number of subjects (or hosts) ns can be a function of t.

[0023] Within-host analysis (block 110) can begin at block 112 with obtaining deep sequencing data for a virus genome for a given subject (denoted by index i, where 1≤i≤ns (t) ) . Deep sequencing data can be obtained using a variety of techniques, including known techniques for next generation sequencing. The deep sequencing data includes a number of raw reads, each of which represents a fragment of the virus genome; fragments are denoted by index b herein. For a given position (denoted herein by index j) in a reference genome of the virus, the number of fragments from subject i that include position j is given by: where denotes any nucleotide (or amino acid) at position j of the reference virus  genome on fragment b of the raw reads for subject i and I {·} is a Boolean function that has value 1 if the logical expression {·} is true and 0 otherwise. In some embodiments,  for a given position j can be a large number, e.g., on the order of 102 or 103 or 106.

[0024] At block 114, the variant frequency (VF) for a particular nucleotide (or amino acid) variant at position j can be computed for subject i. In notation used herein, A denotes the wild-type nucleotide (or amino acid) in the reference genome, and a denotes an alternative type (or variant) . Variant frequency of variant a at position j for a given subject i can be computed as: Eq. (2) represents the fraction of raw reads b covering position j from subject i for which  variant a is present at position j.

[0025] At block 116, variant a can be classified as present or absent in subject i, based on variant frequency For instance, a threshold (θ) can be applied to such that is counted as present in subject i if and absent otherwise. Threshold θ can be chosen to avoid false detection of a variant due to sequencing error. Optimal choice of θ may depend on the particular sequencing technique; for instance, θ can be 0.01 or a higher value.

[0026] Between-host analysis (block 120) can include combining results of the within-host analysis across different subjects i for a given time period t. At block 122, a between-host fitness for variant a at position j can be computed. For instance, the total number of subjects detected with Xj=a during time period t can be computed as: Between-host fitness for variant a at position j can be defined as the detection frequency  (DF) of variant a, which can be computed as: Eq. (4) represents the fraction of subjects detected with Xj=a during time period t and is the  detection frequency for Xj (t) =a in all sampled subjects ns (t) using deep sequencing data.

[0027] At block 124, an average (or mean) within-host fitness for variant a at position j can be computed. For instance, the total number of raw reads with Xj (t) =a among the nj (a, t, θ) samples can be computed as: and the total number of fragments in the nj (a, t, θ) samples that cover position j can be  computed as:

[0028] Taking variant frequency as a measure of within-host fitness of Xj=a for subject i, the mean within-host fitness can be defined as the average within-host VF among the subjects i at time t. In other words, the mean within-host fitness can be computed as:

[0029] At block 126, a fitness potential can be computed based on the average (or mean) within-host fitness computed at block 124 and the between-host fitness computed at block 122. In some embodiments, fitness potential can be defined based on a disparity between a high within-host fitness (e.g., high VFj (a, t) ) and a low between-host fitness (e.g., low DFj (a, t, θ) ) . For instance, fitness potential can be defined as an odds ratio (OR) , which can be computed as:

[0030] At block 130, the fitness potential can be used to identify (or predict) key mutations, e.g., mutations with high fitness potential. For instance, p-value of the fitness potential can be calculated using the Fisher’s Exact Test. The Bonferroni-corrected significance level can be used to screen the p-values and select mutations with high fitness potential.

[0031] FIG. 2 graphically illustrates some operating principles of process 100 according to some embodiments. Within-host analysis is shown at block 210. Deep sequencing is used to generate a number of raw reads 204, each of which corresponds to a fragment of a reference genome 206 at a time t. Existing techniques can be used to identify the correspondence between each fragment and reference genome 206, as illustrated by the positioning of raw reads 204. For a particular subject i, the variant frequency (VF) for a particular variant (a) at a particular position j (represented as shaded column 208) within the reference genome can be computed statistically, as described above. Between-host analysis is shown at block 220. Whether the variant a is detected or not for each subject is illustrated at 222, and detection frequency (DF) can be computed from this information. At block 230, the within-host fitness (graph 232) determined from the within-host variant frequency and the between-host fitness (graph 234) determined from the detection frequency are used to compute a fitness potential. As shown, high fitness potential occurs when there is high within-host fitness (shown in graph 232 by variant a having a higher variant frequency than the wild-type A) coupled with low between-host fitness (shown in graph 234 by variant a having a lower detection frequency than the wild-type A) .

[0032] It should be understood that process 100 is illustrative and that variations and modifications are possible. The particular metrics for within-host fitness and between-host fitness can be modified. Process 100 can be performed to investigate any number of positions j in the genetic sequence of a virus. Further, process 100 can be repeated for virus samples collected from subjects during different time periods t. The subjects during a given time period may be individuals infected with the virus during that time period (which may include different individuals for different time periods) , and the number of subjects ns (t) can be different for different time periods. Process 100 can be applied to a DNA sequence, RNA sequence, or amino acid sequence of a virus or other pathogen where the subjects are different hosts (e.g., different human or animal subjects infected with the virus) . Variants can be defined with reference to any type of genomic data, including individual nucleotides, codons, alleles, or amino acids. In other embodiments, process 100 can be applied to identify (or predict) mutations having high fitness potential between two levels of any living environments, such as cells, tissues, or organs. For instance, process 100 can be applied to cancer cells (treated as a pathogen) . Estimation of Prediction Lead Time

[0033] It is contemplated that process 100 can be used in real time (e.g., during a viral outbreak) to predict key mutations before they spread through the population. It is therefore of some interest to understand how far in advance process 100 can predict a key mutation. As used herein, the “prediction lead-time” refers to the duration from the first time of identifying a high-potential mutation (e.g., using process 100) to the time when the mutation is first observed to be dominant at population scale. As described below, efficacy of an embodiment of process 100 has been studied using data from a real-world viral outbreak, and prediction lead time has been evaluated by determining when process 100 first identified (or predicted) a given mutation as a key mutation and when that mutation became dominant in a population.

[0034] For instance, the Global Initiative on Sharing All Influenza Data (GISAID) provides a comprehensive dataset to study global genetic epidemiology of infectious disease viruses. GISAID data can be used to generate predictions and to evaluate the predictions by determining whether (and when) a predicted key mutation truly becomes dominant at population level in the GISAID data. Genetic sequences recorded in the GISAID database are isolated and recorded as one strain per subject. That is, for each subject, a single consensus sequence of the within-host viral genome is obtained and reported for that subject. This unit of viral genetic epidemiology analysis is referred to herein as the “subject-wise viral sequence” and denoted as Thus, the population-level mutation prevalence of Xj=a can be estimated as: where N (t) is the total number of subject-wise sequences downloaded from GISAID for time  period t.

[0035] The earliest time period at which fitness potential (e.g., as given by Eq. (8) ) reaches significance using the within-host data can be denoted as t0 and computed as: where T is the investigation period and α is the significance threshold in a real data analysis,  e.g., α=3.92×10-5 after Bonferroni correction. The time when mutation Xj=a is observed to reach dominance at population-level in GISAID can be denoted as td and computed as: Letting index k represent the predicted high-potential (or key) mutations, the lead time of a  predicted key mutation k can be denoted as δk and computed as: Estimating Source Region for a Key Mutation

[0036] For pandemic disease viruses like SARS-CoV-2, it may be advantageous to conduct the within-host analysis using data acquired from the source region of emerging variants. To facilitate the analysis, some embodiments provide a method to delineate the probable source region for globally circulating viruses. The basic approach is to estimate the probability of a particular geographical region to appear in a given rank order among a number of different regions, by considering the spatial-temporal distribution of the prevalence of key mutations. One such method is referred to herein as the Estimation of Probable Source region ( “EPS” ) . Key-mutation prediction process 100 described above and EPS are standalone processes; however, joint application of the two processes may improve effectiveness of predictions.

[0037] FIG. 3 shows a flow diagram of a process 300 for EPS according to some embodiments. Process 300 can be a computer-implemented process operating on a large number of data samples. At block 302, a data set is defined. The data set can include viral sequence data for a number of subjects (hosts) collected at different time periods (e.g., different weeks or months) and in different geographic regions (e.g., different countries) . The data set can be binned into space-time strata (e.g., data samples per region per month) . In some embodiments, limits (maximum and / or minimum) on the number of samples per space-time stratum can be applied to avoid statistical artifacts due to extreme disparities in number of samples per stratum.

[0038] At block 304, key mutations are identified (retrospectively) within the data set. For instance, mutations can be identified by comparing sequence data from particular subjects to a reference sequence for the virus. The reference sequence can represent the original (e.g., earliest isolated) form of the virus, and a mutation can be identified as any instance where a residue in a given position differs from the reference sequence. Key mutations can be identified as mutations that become predominant (or sufficiently prevalent) in the population of subjects in multiple geographic regions. It should be understood that a key mutation can reach predominance in different geographic regions at different times. Retrospective identification of key mutations can be based on actual data and predictions according to process 100 need not be considered.

[0039] At block 306, geographic transmission routes for some or all of the key mutations can be determined. For instance, for a given key mutation, the temporal order in which the key mutation becomes prevalent in different regions can indicate the geographic transmission route.

[0040] At block 308, a probability that a given geographic region is the source of key mutations can be determined, e.g., based on analysis of the geographic transmission routes of different key mutations.

[0041] To illustrate the application of EPS, process 300 was applied to profiling of the global circulation of SARS-CoV-2. At block 302, a data set was defined as follows. Thirteen geographic regions from six continents were selected, including Asia (India, Israel, Japan, Singapore) , Europe (Germany, Russia, United Kingdom) , Africa (South Africa) , North America (California, Mexico, New York) , South America (Brazil) , and Oceania (Australia) . The time range considered spanned the three years of COVID-19 pandemic from March 2020 to June 2023. Monthly intervals were adopted as the time binning unit for analyzing the serial cross-sectional data. Thus, the total number of space-time strata (or bins) was 13 regions × 39 months, which gives 468 sampling bins (or strata) . Subject-wise genetic sequences were downloaded from the GISAID EpiCoV database. A total number of 6,056,352 full-length human SARS-CoV-2 strains isolated in the selected geographic regions and time periods were retrieved, accounting for 43.7%of all available SARS-CoV-2 sequences worldwide in the database. Strains with duplicated names or unclear collection time were removed. It was estimated that a sample size of ~1, 500 would be sufficient to reliably detect mutation of 1%prevalence with a type I error rate of 5%and a margin of error of 0.5%. Thus, to ensure equitable statistical power for comparing across regions and time periods, the maximum sample size of each space-time stratum was set at 1, 500. When a space-time stratum had sample size below this threshold, all of the available samples were drawn for that stratum. In instances where a region was not an obvious source region of a given mutation and the sample size was less than 1500 in the previous stratum before the mutation began continuous circulation, the region was not counted towards the ranking of regions along the geographic transmission route for that mutation, to avoid under-estimation of its priority when comparing to the other regions. Multiple sequence alignment was performed using MAFFT (version 7) .

[0042] At block 304, identification of key mutations that reached global prevalence was performed as follows. Let Xk denote an amino acid mutation, where 1≤k≤K, and K= J×19, where J is the length of the amino acid sequence being considered, and 19 is the number of possible amino acid outcomes at a given genome position. The initial SARS-CoV-2 sequence, ‘Wuhan-Hu-1’ (GISAID EPI_ISL_402125) , was defined as the reference sequence, and a mutation was defined as any instance where the residue at a given position was different from the corresponding residue in the reference sequence. Letting pkl (t) denote prevalence of mutation Xk in region l at time t, where L is the total number of geographical regions being considered (l =1, …, L) , and T is the number of time intervals in the study period (t = 1, …, T) , a key mutation was defined as a mutation that satisfied the following criteria: (1) Fitness. The mutation demonstrated the ability to reach predominance in any one of the  regions, where predominance was defined as:  (2) Coverage. The mutation had good coverage globally among the geographical regions.  “Good coverage” was defined as the mutation being observed with prevalence over a certain threshold c (e.g., 20%) in more than half of the L regions (7 regions in this study) . Let Lk denote the number of regions observed with pkl (t) ≥ c, l∈L. Thus, 0≤Lk≤L. The coverage criterion was met if

[0043] Denoting the set of key mutations by W, the set of key mutations can be defined as: where k∈K, t∈T, and l∈L.

[0044] At block 306, determining the geographic transmission route of a single mutation was based on determining the mutation’s sequential order of appearing in continuous circulation among the L regions. For a mutation Xk,  was defined as the first time point when the mutation begins continuous circulation in region l, namely,  and Δt=0, 1, 2, …. In other words,  represents the time when lth region is first observed with continuous circulation of mutation Xk, and l=1, …, Lk.

[0045] The source region for mutation Xk, denoted as was identified as the region with the earliest observation time of its continuous circulation out of all regions being evaluated, that is: Similarly, the sink region for mutation Xk, denoted as was identified as the region with the  latest observation time of its continuous circulation out of all regions being evaluated, that is: The lead-time of observing mutation Xk in its source region compared to observing it in its  sink region can be defined as the source-to-sink lead-time (SSL) and computed as:

[0046] Ranks Rkl were assigned the source regions based on ordering the regions according to increasing When a rank tie occurred between two regions l and l′, the average of the two consecutive ranks was taken, such that Similar adjustments were applied when a rank tie occurred between three or more regions. When Lk<L, the rank Rkl was first calculated among the Lk regions and then normalized to such that range to allow joint analysis with the other mutations.

[0047] At block 308, probability of a particular geographic region being the source of key mutations was determined as follows. Let Rl denote the random variable of region l’s ranking along the global transmission pathway. The probability distribution of a region l to appear in the rth position along the global transmission pathway can be estimated by:

[0048] Accordingly, Pr {Rl=1} gave the probability of region l to serve as the source place of key mutations. In this manner, regions with high probability of being the source region for key mutations can be identified.

[0049] It should be appreciated that this example of process 300 is illustrative and that similar analysis can be applied to any transmissible virus or pathogen for which sequence data is available for different time periods and geographic regions. The size of a time bin and the size of a geographic region can be modified as desired. Process 300 or similar processes can be applied in a variety of contexts where understanding the geographic transmission route (s) of pathogenic strains is desired. In some embodiments, process 300 can be used to identify one or more likely source regions for key mutations, and process 100 can be applied to deep sequencing data obtained for subjects in the likely source region (s) . Examples

[0050] To illustrate the performance of an implementation of process 100, SARS-CoV-2 genetic data were used. Raw reads of SARS-CoV-2 data were downloaded from the European Nucleotide Archive (ENA) , while reference genetic sequences were downloaded from the GISAID EpiCoV database. Quality control for the raw reads data was conducted using Trimmomatic (Bolger, A.M., Lohse, M., & Usadel, B. (2014) , Trimmomatic: a flexible trimmer for Illumina sequence data, Bioinformatics, 30 (15) , 2114-2120) , followed by sequence alignment by BWA (Li, H., & Durbin, R. (2009) , Fast and accurate short read alignment with Burrows-Wheeler transform, Bioinformatics, 25 (14) , 1754-1760) and transform to the BAM format using SAMtools (Heng Li, Bob Handsaker, Alec Wysoker, Tim Fennell, Jue Ruan, Nils Homer, Gabor Marth, Goncalo Abecasis, Richard Durbin, 1000 Genome Project Data Processing Subgroup (2009) The Sequence Alignment / Map format and SAMtools, Bioinformatics, 25 (16) , 2078-2079) . Variant calling was performed using LoFreq (Andreas Wilm, Pauline Poh Kim Aw, Denis Bertrand, Grace Hui Ting Yeo, Swee Hoe Ong, Chang Hua Wong, Chiea Chuen Khor, Rosemary Petric, Martin Lloyd Hibberd, Niranjan Nagarajan (2012) LoFreq: a sequence-quality aware, ultra-sensitive variant caller for uncovering cell-population heterogeneity from high-throughput sequencing datasets, Nucleic Acids Research, 40 (22) , 11189-11201) , annotation by snpEff (Cingolani P, Platts A, Wang le L, Coon M, Nguyen T, Wang L, Land SJ, Lu X, Ruden DM (2012) A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3, Fly, 6 (2) , 80-92) . The GISAID sequences were aligned using MAFFT (Katoh, K., Misawa, K., Kuma, K., & Miyata, T. (2002) , MAFFT: a novel method for rapid multiple sequence alignment based on fast Fourier transform, Nucleic acids research, 30 (14) , 3059-3066) . The investigation period ranged from March 2020 to August 2023. Based on the EPS method (process 300) described above, the top source regions for key mutations were identified as South Africa, India, and California. Accordingly, process 100 was applied to the raw-reads data of these three regions (or other appropriate source regions) in each time interval to generate a set of predicted key mutations. Validation of the prediction was evaluated using data from four continents covering 13 geographical regions, including Asia (India, Israel, Japan, Singapore) , Europe (Germany, Russia, United Kingdom) , Africa (South Africa) , North America (California, Mexico, New York) , South America (Brazil) and Oceania (Australia) . The ground-truth set of key mutations consisted of 91 mutations from 77 residue sites that were observed to be dominant in any one of the 13 geographical regions. A prediction of a key mutation using process 100 was labeled as correct if the predicted mutation became dominant after the prediction time in any one of the regions.

[0051] Results are shown in FIGs. 4 and 5. FIG. 4 shows receiver operating characteristic graphs for predictions of key mutations on the Receptor Binding Domain (RBD) , the N-Terminal Domain (NTD8) , and the whole-length spike protein (S) . As shown, the area-under-the-curve (AUC) for prediction of key mutations on the RBD was 0.748, on the NTD was 0.772, and on the S protein was 0.769.

[0052] FIG. 5 shows graphs of prediction lead time computed using the method described above with reference to Eqs. (9) - (12) . The average prediction lead-times were 8.1 months (SD 7.0) , 8.3 months (SD 8.1) , and 8.1 months (SD 7.3) for key mutations on the RBD, NTD and S protein, respectively. These results demonstrate that some embodiments of process 100 can accurately and timely predict key mutations of a virus such as SARS-CoV-2 by leveraging within-host deep sequencing data of the virus genome.

[0053] In another illustrative example, a total number of 27 predicted mutations were identified using methods described above based on genetic data from samples obtained in June 2023. These predicted mutations were genetically engineered to create pseudoviruses of SARS-CoV-2, and the pseudoviruses were evaluated for their binding affinity to the ACE2 protein using enzyme-linked immunosorbent assay (ELISA) experiments. Out of the 27 mutations, eight showed significant ACE2 binding ability. FIG. 6 shows a graph of the binding affinity for each mutation, quantified using AUC (as commonly used in ELISA) . Asterisks (*or **) indicate the eight mutations with significant ACE2 binding ability.

[0054] The 27 predicted mutations were further evaluated to assess their ability to evade the immune systems of human subjects who had been vaccinated with three does of COVID-19 vaccination. Antisera samples were obtained from the subjects for ELISA analysis. FIG. 7 shows a graph of antibody escape of the pseudovirus with each mutation, quantified using the 50%neutralization titer (NT50) commonly used in ELISA. The results showed that 11 of the predicted mutations (indicated by horizontal bars 702) would not be neutralized by human antisera, suggesting that they could potentially escape the immune response of the vaccinated population. Accordingly, these predicted mutations could be considered for vaccine antigen design against future emerging variants of SARS-CoV-2. Additional Embodiments

[0055] While the invention has been described with reference to specific embodiments, those skilled in the art will appreciate that variations and modifications are possible. For instance, all processes described above are illustrative and may be modified. Processing operations described as separate blocks may be combined, order of operations can be modified to the extent logic permits, processing operations described above can be altered or omitted, and additional processing operations not specifically described may be added. Particular definitions and data formats can be modified as desired.

[0056] Processes described herein can be applied to sequence and epidemic data for a specific region, to global data, or to a mathematical combination of regional and global data. Predictions can be specific to a particular region (e.g., country, continent, or hemisphere) or global. The investigation period and duration of each time bin can be as long or short as desired, depending on availability of data and likely mutation rates. In some embodiments, the virus samples and population-level data can be localized to a particular area (e.g., a country, a state or region, a city) , allowing for modeling of geographic variations in virus activity. In some embodiments, EPS (e.g., process 300) can be used to identify one or more likely source regions for key mutations, and predictions of future key mutations can be made by applying process 100 (or other similar processes) to data from the likely source region (s) .

[0057] Processes described herein can be applied to genetic data (e.g., DNA, RNA, and / or amino acid sequence data) of any virus or other pathogen, including influenza viruses, coronaviruses such as SARS-CoV-2 viruses, or new emerging pathogens that may cause epidemics or pandemics. Similar processes can also be applied in other contexts in which early detection or prediction of mutations that confer a survival advantage is desirable, including cells, tissues, or organs (e.g., with cancer cells regarded as the pathogen) .

[0058] Predictions of future key mutations can be used to inform public health planning activities such as vaccine design, development of antivirals or other treatments, or the like. For instance, a vaccine can be designed or updated to include one or more strains that have the predicted key mutation (s) . Vaccines designed in this manner can be manufactured in quantity and administered to the general population or to people at particular risk for infection (e.g., in a particular geographic area that has been identified as a likely source for key mutations) using existing techniques or other techniques that may be developed in the future.

[0059] The sequencing data employed in analysis of the kind described herein can be obtained using any available sequencing technologies, including but not limited to first-generation sequencing (Sanger) , next-generation sequencing (e.g., Illumina platform) , or third-generation sequencing (e.g., PacBio platform or Nanopore platform) .

[0060] Data analysis and computational operations of the kind described herein can be implemented in computer systems that may be of generally conventional design, such as a desktop computer, laptop computer, tablet computer, mobile device (e.g., smart phone) , or the like. Computing clusters and / or cloud-based computing systems may be used for increased computational power. Such systems may include one or more processors to execute program code (e.g., general-purpose microprocessors usable as a central processing unit (CPU) and / or special-purpose processors such as graphics processors (GPUs) that may provide enhanced parallel-processing capability) ; memory and other storage devices to store program code and data; user input devices (e.g., keyboards, pointing devices such as a mouse or touchpad, microphones) ; user output devices (e.g., display devices, speakers, printers) ; combined input / output devices (e.g., touchscreen displays) ; signal input / output ports; network communication interfaces (e.g., wired network interfaces such as Ethernet interfaces and / or wireless network communication interfaces such as Wi-Fi) ; and so on.

[0061] Computer programs incorporating features of the present invention that can be implemented using program code may be encoded and stored on various computer readable storage media; suitable media include magnetic disk or tape, optical storage media such as compact disk (CD) or DVD (digital versatile disk) , flash memory, and other non-transitory media. (It is understood that “storage” of data is distinct from propagation of data using transitory media such as carrier waves. ) Computer readable media encoded with the program code may include an internal storage medium of a compatible electronic device and / or external storage media readable by the electronic device that can execute the code. In some instances, program code can be supplied to the electronic device via Internet download or other transmission paths.

[0062] Accordingly, although the invention has been described with respect to specific embodiments, it will be appreciated that the invention is intended to cover all modifications and equivalents within the scope of the following claims.

Claims

1.A computer-implemented method for predicting evolution of a pathogen, the method comprising, for a particular time bin:obtaining deep sequencing data for the pathogen for each of a plurality of individual subjects infected with the pathogen;computing a within-host fitness measure for each of a plurality of variants of the pathogen in each of the individual subjects;computing a between-host fitness measure for each of the plurality of variants of the pathogen across the individual subjects;computing, based on the within-host fitness measure and the between-host fitness measure, a fitness potential for each of the plurality of variants of the pathogen; andpredicting one or more key mutations of the pathogen based on the fitness potentials computed for the plurality of variants.2.The method of claim 1 wherein computing the within-host fitness measure includes computing a variant frequency for a particular variant of the plurality of variants, wherein the variant frequency represents a fraction of raw reads in the deep sequencing data for a particular subject of the plurality of individual subjects that include the particular variant.3.The method of claim 2 wherein computing the between-host fitness measure for the particular variant of the plurality of variants includes:computing, based on the variant frequency for each of the plurality of individual subjects, a detection frequency for the particular variant, the detection frequency corresponding to a fraction of the individual subjects in which the particular variant is detected.4.The method of claim 3 wherein computing the fitness potential for the particular variant of the plurality of variants includes:computing an average variant frequency across the plurality of individual subjectscomputing an odds ratio from the average variant frequency and the detection frequency.5.The method of claim 4 wherein identifying key mutations includes identifying, as a key mutation, a mutation with a high odds ratio.6.The method of claim 1 further comprising:repeating the acts of obtaining deep sequencing data, computing a within-host fitness measure, computing a between-host fitness measure, computing a fitness potential, and identifying key mutations for each of a plurality of time bins.7.The method of claim 1 further comprising:repeating the acts of obtaining deep sequencing data, computing a within-host fitness measure, computing a between-host fitness measure, computing a fitness potential, and identifying key mutations for each of a plurality of geographical regions.8.The method of claim 1 wherein the deep sequencing data corresponds to nucleotide sequences.9.The method of claim 1 wherein the deep sequencing data corresponds to amino acid sequences.10.The method of claim 1 wherein the plurality of variants includes mutations at different locations within a genetic sequence of a pathogen.11.The method of claim 10 wherein the pathogen is a virus.12.The method of claim 1 further comprising:determining a likely source region for new key mutations,wherein obtaining the deep sequencing data includes obtaining deep sequencing data for individual subjects in the likely source region.13.The method of claim 1 further comprising:selecting one or more strains of the pathogen to include in a vaccine, wherein at least one of the one or more strains includes one or more of the predicted key mutation.14.The method of claim 13 further comprising:manufacturing the vaccine; andadministering the vaccine to at least one subject.15.A computer system comprising:a memory to store data; anda processor coupled to the memory and configured to perform the method according to any one of claims 1 to 13.16.A computer-readable storage medium having stored therein instructions that, when executed by a processor of a computer system, cause the processor to perform the method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • System and method for detecting population variation from nucleic acid sequencing data

    WO2014152990A1

  • Disease prediction system using open source data

    WO2015127065A1

  • Measurement and prediction of virus genetic mutation patterns

    WO2019242597A1

  • Methods and systems for reporting clinically-actionable potential germline pathogenic variant sequences

    WO2023096658A1