Machine learning approach to predict virus or cancer mutations for vaccine production or drug design
The time-varying transformer-based method addresses the limitations of existing genomic mutation prediction by accurately forecasting top-k likely mutations, enhancing vaccine and treatment strategies with confidence estimation.
Patent Information
- Application Number
- PCT/IB2025/050405
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-14
- Filing Date
- 2025-01-14
- Publication Date
- 2025-08-21
AI Technical Summary
Existing methods for predicting genomic mutations are computationally intensive and have limited predictive power, often relying on observed mutations and failing to anticipate novel mutations, which poses challenges in vaccine development and drug design.
A time-varying transformer-based method that encodes molecular sequences into latent vectors, using a time-varying mutation model to predict the top-k most likely mutated sequences with confidence, incorporating biologically valid mutations and evolutionary constraints.
Enables accurate and efficient prediction of future genomic mutations, facilitating proactive vaccine design and personalized cancer treatment by simulating potential evolutionary pathways with uncertainty estimation.
Smart Images

Figure IB2025050405_21082025_PF_FP_ABST
Abstract
Description
MACHINE LEARNING APPROACH TO PREDICT VIRUS OR CANCER MUTATIONS FOR VACCINE PRODUCTION OR DRUG DESIGNCROSS REFERENCE TO RELATED APPLICATIONS
[0001] This application claims benefit to European Patent Application No. EP 24157633.9, filed on February 14, 2024, which is hereby incorporated by reference herein.FIELD
[0002] The present disclosure relates to Artificial Intelligence (Al) and machine learning (ML), and in particular to a method, system, data structure, computer program product and computer-readable medium for predicting a top-k most likely mutated molecular sequences. BACKGROUND
[0003] Viral pathogens continually evolve to survive and adopt in changing environments. This ability to mutate allows them to evade the human immune response, which poses big challenges in vaccine development and drug design. The mutations can cause previously designed substances to lose their targets or diminish their efficacy. For instance, the COVID-19 pandemic exemplified how mutations in SARS-CoV-2 may increase virus transmissibility and impact the efficacy of previously used vaccines. As another example, the Influenza virus requires an updated vaccine each season because it is constantly changing, both through recombination and mutations in surface proteins.
[0004] Genomic mutations are deemed unpredictable because of multiple known and unknown factors, and the occurrence of a mutation is usually modelled as a random event. Several methods have emerged in the literature aiming to forecast specific regions of viral genomes that are prone to mutate, however, they normally operate with mutation events that have been already observed and can’t score mutations that are not yet in population. This highlights the need for predictive methods that can anticipate novel mutations and foreseen emerging virus strains, enabling the adaptation of preventive vaccines from one season to the next.
[0005] Classical approaches predict the mutability along the viral proteome using conservation profiles built from multiple sequence alignments (MSA) of the related viruses (see Nagar, A., Hahsler, M, “Fast Discovery and Visualization of conserved regions in DNA sequences using quasi-alignment,” BMC Bioinformatics 14 (Suppl. 11), S2 (2013) , which is hereby incorporated by reference herein). On top of the MSA approach other classical approaches calculate genome conservation features, such as site-based Shannon’s entropy, or identify genome loci subject to strong selective constraints. The resulting models from these classical approaches normally require the alignment of an extensive number of sequencescollected over time and have few parameters, but have limited predictive power. A method as described in Rodriguez-Rivas J., et al., “Epistatic models predict mutable sites in SARS-CoV-2 proteins and epitopes,” PNAS 119(4) : e2113118119, (2022), which is hereby incorporated by reference, uses data from coronaviruses, preexisting to SARS-CoV-2, to build alignments of homologous sequences and train statistical sequence models to predict the mutability of each position. This method also tries to account for complex patterns resulting from epistasis while assigning mutability scores. However, the inference of the model parameters used in this method and considers high-order epistatic interactions is computationally hard.
[0006] There have also been attempts to identify viral mutations that have potential to spread in the population based on features from viral epidemiology, evolution, immunology, and neural network-based protein sequence modeling. However, it is not obvious from these attempts which type of features or combination of features will have the best predictive power. For example, the method described in Yan, S., Wu, G., “Application of neural network to predict mutations in proteins from influenza A viruses - A review of our approaches with implication for predicting mutations in coronaviruses,” J. Phys.: Conf. Ser. 1682 012019 (2020), which is hereby incorporated by reference, adopts the ideas of statistical physics to calculate mutationbased features from MSA of viral proteins and apply a neural network to predict the probabilities of mutations in influenza A virus genes. As another example, the highest predictive power of the model described in Maher et al., “Predicting the mutational drivers of future SARS-CoV-2 variants of concern,” Sci. Transl. Med. 14(633), (2022), which is hereby incorporated by reference, was obtained from an epidemiological feature, namely, the exponentially weighted mean ranking across epidemiological variables (mutation frequency, fraction of unique haplotypes in which the mutation occurs, and the number of countries in which it occurs), which was validated by predicting driver mutations in emerging SARS-CoV-2 variants of concern. Another approach includes a cutting-edge PyRo model that predicts next emerging SARS-CoV-2 strain by employing a hierarchical Bayesian multinomial logistic regression (see Fritz Obermeyer, Martin Jankowiak, Nikolaos Barkas, Stephen F Schaffner, Jesse D Pyle, Leonid Yurkovetskiy, Matteo Bosso, Daniel J Park, Mehrtash Babadi, Bronwyn L Maclnnis, Jeremy Luban, Pardis C Sabeti, Jacob E Lemieux, “Analysis of 6.4 million SARS- CoV-2 genomes identifies mutations associated with fitness,” Science 376, 1327-1332, doi: 10. 1126 / science.abml208 (2022), which is hereby incorporated by reference herein). This model determines which mutations are becoming more spread in a population and estimates how quickly each mutation could cause the lineages to spread.
[0007] Other approaches include extending text-mining and natural language processing to viral genomes to aid in estimating the mutability of genomic segments. These methods oftenevolve assessing the importance of genomic segments based on their spatial distribution and frequency over the whole genomes (see Darooneh et al., (2021), which is hereby incorporated by reference herein). Another approach involves analyzing genomic loci though embeddings derived from protein language models (see Brian Hie, Ellen D. Zhong, Bonnie Berger, Bryan Bryson, “Learning the language of viral evolution and escape,” Science 371, 284-288, (2021), which is hereby incorporated by reference herein).
[0008] Compared to standard text-mining and natural language processing methods or previously provided neural network-based protein sequence modeling, predicting genomic mutations have the following technical challenges: 1) it usually involves using large data sets that employ computationally intensive operations to properly analyze; and 2) have lower accuracy in generated predictions as previously described methods work with mutation events that have been already observed and their models can’t score mutations that are not yet in the population. These are technical problems to be overcome for generating accurate mutation sequence predictions using techniques that are less computationally intensive.SUMMARY
[0009] In an embodiment, the present disclosure provides a computer-implemented, machine learning method for predicting atop-k most likely mutated molecular sequences. A molecular sequence is encoded at a first time into a first vector in a latent space. A second vector is generated, by a time-varying mutation models, in the latent space using as input the first vector. The second vector indicates a time-varying influence of the molecular sequence on a mutated version of the molecular sequence at a subsequent time. The second vector is decoded to generate a prediction of the top-k most likely mutated molecular sequences for the molecular sequence at the subsequent time. The method has applications including, but not limited to, use cases in computational biology and medical Al and healthcare for optimizing vaccine design or supporting decision making in diagnosis and treatment of patients.BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Embodiments of the present disclosure will be described in even greater detail below based on the exemplary figures. The present disclosure is not limited to the exemplary embodiments. All features described and / or illustrated herein can be used alone or combined in different combinations in embodiments of the present disclosure. The features and advantages of various embodiments of the present disclosure will become apparent by reading the following detailed description with reference to the attached drawings which illustrate the following:
[0011] FIG. 1 illustrates an overview of the proposed system and method for mutation prediction according to an embodiment of the present disclosure;
[0012] FIG. 2 is an example of an encoder architecture that maps a virus at time t to a vector in a latent space;
[0013] FIG. 3 illustrates a time-varying layer;
[0014] FIG. 4 is an example architecture of the decoder that predicts the amino acid sequence;
[0015] FIG. 5 illustrates an overview of mutation prediction;
[0016] FIG. 6 illustrates a simulation of top-k most likely mutated sequences based on beam search; and
[0017] FIG. 7 is a block diagram of an exemplary processing system, which can be configured to perform any and all operations disclosed herein.DETAILED DESCRIPTION
[0018] Embodiments of the present disclosure provide solutions to the technical challenges in predicting most likely mutated sequences (genome sequences) discussed above by using a time-varying transformer-based method to make the predictions.
[0019] In a first aspect, the present disclosure provides a computer-implemented method for predicting a top-k most likely mutated molecular sequences. A molecular sequence is encoded at a first time into a first vector in a latent space. A second vector is generated, by a time-varying mutation model, in the latent space using as input the first vector. The second vector indicates a time-varying influence of the molecular sequence on a mutated version of the molecular sequence at a subsequent time. The second vector is decoded to generate a prediction of the top- k most likely mutated molecular sequences for the molecular sequence at the subsequent time.
[0020] In a second aspect, the present disclosure provides the method according to the first aspect, wherein the time-varying mutation model has been trained by minimizing a loss of training data provided as input to the time-varying mutation model, the training data including a training molecular sequence at the first time and a mutated version of the training molecular sequence at the subsequent time.
[0021] In a third aspect, the present disclosure provides the method according to the first aspect or the second aspect, wherein further time intervals are unevenly distributed with a time interval between the first time and the subsequent time, and wherein minimizing the loss includes tuning a coefficient of a loss function with a cross-validation procedure.
[0022] In a fourth aspect, the present disclosure provides the method according to any of the first to third aspects, wherein minimizing the loss includes using a term indicating a training objective of the time-varying mutation model that enforces biologically valid mutations in generated molecular sequences of the predicted top-k most likely mutated molecular sequences,and wherein enforcing the biologically valid mutations includes mapping a primary sequence of the training molecular sequence with structural properties of the primary sequence.
[0023] In a fifth aspect, the present disclosure provides the method according to any of the first to fourth aspects, wherein the decoding is performed by a transformer decoder that has been trained to compute a probability of a next molecular component at each subsequent position of the mutated version of the molecular sequence at the subsequent time, and wherein the top-k most likely mutated molecular sequences are based on the probabilities.
[0024] In a sixth aspect, the present disclosure provides the method according to any of the first to fifth aspects, wherein computing the probabilities includes using rules as evolutionary constraints and restrictions based on mutation rates, wherein the restrictions based on the mutation rates includes using molecular component substitution matrices to rescale the probabilities for the next molecular component at each subsequent position, and wherein the rules as the evolutionary constraints prohibit predictions that correspond to incompatible mutations for the molecular sequence.
[0025] In a seventh aspect, the present disclosure provides the method according to any of the first to sixth aspects, wherein generating, by the time-varying mutation model, the second vector in the latent space includes using a user specified time interval to determine the subsequent time, and further includes using a plurality of time-varying mechanisms, the plurality of time-varying mechanisms including a linear time-varying mechanism, a convex time-varying mechanism, and a concave time-varying mechanism.
[0026] In an eighth aspect, the present disclosure provides the method according to any of the first to seventh aspects, wherein each of the top-k most likely mutated molecular sequences for the molecular sequence at the subsequent time includes a value that indicates a confidence of the prediction for a corresponding molecular sequence mutation.
[0027] In a ninth aspect, the present disclosure provides the method according to any of the first to eighth aspects, wherein generating the prediction of the top-k most likely mutated molecular sequences for the molecular sequence at the subsequent time includes using a beamsearch method that integrates biological constraints through a pruning strategy by computing, at each position, a probability of a particular molecular component using the second vector as a condition, and selecting a top-k molecular components with largest probabilities and joint probabilities with other positions of the molecular sequence until an end is predicted.
[0028] In a tenth aspect, the present disclosure provides the method according to any of the first to ninth aspects, wherein the molecular sequence is a genome sequence for a pathogen, and wherein the method further comprises obtaining a user specified time interval, decoding the second vector in the latent space to generate the prediction of the top-k most likely mutatedmolecular sequences for the molecular sequence at the subsequent time using the user specified time interval, and identifying specific amino acids that correspond to the predicted top-k most likely mutated molecular sequences for the molecular sequence at the subsequent time as a basis for a vaccine to the pathogen.
[0029] In an eleventh aspect, the present disclosure provides the method according to any of the first to tenth aspects, wherein the molecular sequence is a genome sequence for a cancer sample, and wherein the method further comprises obtaining a user specified time interval, decoding the second vector in the latent space to generate the prediction of the top-k most likely mutated molecular sequences for the molecular sequence at the subsequent time using the user specified time interval, and identifying specific amino acids that correspond to the predicted top- k most likely mutated molecular sequences for the molecular sequence at the subsequent time to identify a treatment strategy for a cancer type associated with the cancer sample.
[0030] In a twelfth aspect, the present disclosure provides a machine learning model stored on a tangible, non-transitory computer-readable medium, the machine learning model comprising a time-varying mutation layer that learns a time-varying influence of a molecular sequence at a first time on a mutated version of the molecular sequence at a subsequent time, the time-varying mutation layer including three linear layers, an activation layer, an aggregation layer, and a batch normalization layer, the linear layers including a linear, concave, and convex time-varying mechanism.
[0031] In a thirteenth aspect, the present disclosure provides the model according to the twelfth aspect, wherein the machine learning model further comprises a transformer encoder, the transformer encoder including a multi-head attention mechanism layer, one or more normalization layers with residual addition, and a position-wise feed forward network layer, and wherein the transformer encoder has a sparse attention mechanism in the multi-head attention mechanism layer to map the molecular sequence at the first time into a vector in a latent space.
[0032] In a fourteenth aspect, the present disclosure provides the model according to the twelfth or thirteenth aspects, wherein the machine learning model further comprises a transformer decoder, the transformer decoder including a masked multi-head attention layer configured to mask subsequent positions of the molecular sequence to perform a prediction of a probability of a next molecular component at each subsequent position of the mutated version of the molecular sequence at the subsequent time in order to provide a prediction of a top-k most likely mutated molecular sequences.
[0033] In a fifteenth aspect, the present disclosure provides a tangible, non-transitory computer-readable medium containing instructions for predicting a top-k most likely mutated molecular sequences using a machine learning method, the instructions, upon being executed byone or more processors, alone or in combination, provide for execution of a machine learning method according to any of the first to eleventh aspects.
[0034] The present disclosure addresses the technical challenges of predicting mutations in molecular sequences such as protein sequences of a viral pathogen or genome sequences of a cancer tumor to characterize evolutionary patterns. This can, for example, assist in predicting future variants of concern of a pathogen, facilitate vaccine discovery, and assess the durability of vaccine targets. The present disclosure describes a time -varying transformer-based method to predict top-k mutated sequences of a pathogen or genome that are most likely to arise in the specified time interval, based on historical pathogen evolution data. The features of the present disclosure also provide a confidence estimation of the predicted mutations. The predicted mutations and associated confidence estimations for the predicted mutations can be used to more accurately identify the regions in a genome where mutations are most likely to occur, as well as to generate potential evolutionary trajectories.
[0035] The predicted mutations provided by the features of the present disclosure can be used to predict the evolution of emerging infectious diseases for vaccine production. The same features of the present disclosure can also be used to predict the evolution of cancer cells which can be used to tailor personalized treatment options for patients. Suitable vaccines can be designed proactively against emerging infectious diseases in advance using the predicted mutations generated by the features of the present disclosure. The present disclosure also enables to save time and resources in the design of vaccines, diagnosis and treatment.
[0036] Embodiments of the present disclosure enable to predict top-k de novo virus variants that have the potential to arise in a specified time interval in the human population instead of reacting after emergence of new variants. The time-varying transformer of embodiments of the present disclosure captures the patterns behind the dynamics of pathogen’s evolution at unevenly distributed time intervals. The method according to an embodiment of the present disclosure simulates the possible pathways of virus mutations, e.g. how and when it may happen with estimated uncertainty. The prior art methods in the technical field of sequence mutation predictions are typically limited to finding a single mutated sequence whereas the present disclosure provides embodiments for predicting top-k most likely mutated molecular sequences. Further, the systems and methods according to embodiments of the present disclosure provide for more accurate predictions for the mutated molecular sequences by using a time-varying transformer, and which are provided with confidence values that represent a certainty or uncertainty of the generated prediction. Previously provided methods in the art do not utilize a specific time frame in which top-k mutated sequences may emerge. As described above, the accurate prediction of mutated molecular sequences can be used to develop vaccines or designwet lab experiments with improved efficacy. The predicted most likely mutated molecular sequences generated by embodiments of the present disclosure can also be used to design treatments for cancer patients or aid clinicians in diagnosing certain cancers or diseases in a patient. For example, clinicians and doctors may utilize different chemotherapy treatments or other treatment options (e.g. surgery) based on the predicted mutation of a cancer for a patient. The ability to quickly and accurately predict mutations represents an important improvement in the technical field of sequence mutation predictions, where acting quickly to design vaccines, diagnose or provide particular treatments is time-critical and often lifesaving.
[0037] FIG. 1 provides an overview of the system and method of the present disclosure. FIG. 1 includes a molecular sequence Seq^ 100 of a pathogen at time slice t being provided as input to a transformer encoder 102. As described below, the transformer encoder may map the molecular sequence Seq^ 100 into a vector Vec^ Seq) 104 in a latent space. A time-varying layer 106 may be used to learn the time-varying influence of the molecular sequence at time t (Seq^) on its mutated version Seq(~t+&t^ at time t + At, where At 108 can by any time interval, such as 1 day, 1 week, 1 month, etc. The At 108 can be user specified. The time-varying layer 106 uses the vector representation Vec^fSeq) 104 to intimate the initial influence of the molecular sequence on its mutations in the future. The time-varying layer outputs the vector Vec(L+M>(Seq) 110 that represents the influence of the pathogen’s current state on its mutations at time t + At. The transformer decoder 112 can predict top-k mutated sequences (114) for the user-specified time interval At 108 and estimate the uncertainty of the prediction for an entire molecular sequence at time t + At 116. The predicted mutations 114 generated by the transformer decoder 112 are condition on the time-varying influence represented as Fec^t+AtSeq) 110.
[0038] The details of each component of the system and method of FIG. 1 are described in the following sections. Although the following sections may describe scenarios in which virus molecular sequences are used and analyzed the embodiments included herein are not limited to such molecular sequences, and also include genome sequences of cancer tumors or samples. Encode the virus at time t
[0039] A transformer-like encoder 200 is used to map protein sequence Seq^ 202 of a pathogen at time slice t into a vector in a latent space, Vec^ Seq) 204. Any suitable transformer encoder architecture may be used as the transformer-like encoder 200. FIG. 2 depicts an exemplary implementation of such a transformer-like encoder 200. It should be noted that the embedding layer 206, which includes an amino acid embedding 208 and amino acid position embedding 210, of the transformer-like encoder 200 maps an amino acid (such as A, R,N, G, M, F, V) to a vector. For example, embedding layer 206 may represent an embedding layer typically found in transformer encoders. In some embodiments the amino acid embedding can include using a token embedding layer that includes a fully connected linear layer to embed amino-acids as tokens. The acid position embedding can include using sinusoidal positional encoding or learned positional encoding methods, which are also found in transformer-based models. As there are around 20 amino acids in total the number of parameters for the embedding layer 206 is finite. This provides the advantage of accelerating the training of the time-varying transformer-based model as it will reduce the complexity of the disclosed method. The transformer-like encoder 200 may include a multi -head attention mechanism layer 212, one or more normalization layers with residual addition 214, and a position-wise feed forward network layer 216.
[0040] In some examples, length L of the pathogen’s protein sequence that is being analyzed often exceeds 2000 amino acids. The influence of amino acid on its close neighbors is often higher than the others. Considering the limited size of the available data, exemplary embodiments of the present disclosure use sparse attention in the multi-head attention block (multi-head attention mechanism layer 212, e.g. see Rewon Child, Scott Gray, Alec Radford, Ilya Sutskever, “Generating long sequences with sparse transformers:, arXiv preprint arXiv: 104.10509, (2019), which is hereby incorporated by reference herein, to capture sequence properties and reduce the complexity of the model. Within each multi -head attention block 212 for an amino acid at the position i, only K previous amino acids are considered to attend to the j'th amino acids, where K is significantly smaller than L (e.g. K is set to be K=50). Thus, the size of the attention matrix is decreased from LxL to KxL, which greatly reduce the number of parameters to be learned thereby providing advantages over conventional mutation sequence prediction systems by reducing training time of the disclosed model and improving accuracy of the generated predictions. The transformer-like encoder 200 of the present disclosure provides additional improvements over conventional mutation sequence prediction methods by utilizing N layers of multi -head attention blocks 212 so that the influence of far-distant residues can be captured.Time-varying layer
[0041] The time-varying layer 106 of FIG. 1 learns the time-varying influence of the pathogen’s sequence at time t (Seq(L>) on its mutated version (Se<t+At)) at time t + At. At can be any time interval, such as (n = 1, 2, ... days). FIG. 3 illustrates the proposed time-varying layer. The time-varying layer obtains or receives, as input, the vector representation Vec^ Seq) 300 of the pathogen’s sequence at time t, which is the output of the transformer encoder component 200 of FIG. 2. This vector 300 intimates the initial influence of the pathogen on itsmutations in future. However, it should be noted that the influence of the pathogen on its mutations in the future will vary with time. For example, the longer the time interval At is, the less the current state impacts its future. An exemplary embodiment of the present disclosure uses three types of varying mechanisms: linear (At) 302, convex (exp(— At)) 304, and concave (At“, a e (0,1)) 306. The use of three types of varying mechanisms ensures that the time-related dependencies are correctly captured by the time-varying layer 106. Although FIG. 3 depicts the use of three types of varying mechanisms the embodiments disclosed herein are not limited to these three and other functions may be used to approximate time-related effects. The time- related changes from the varying mechanisms 302, 304, and 306 are aggregated 308 to simulate the comprehensive effect of the time interval At. The aggregation 308 of the time-related changes from the varying mechanisms 302, 304, and 306 may be a weighted sum with trained or predefined weights, so that each type of effect may have its own contribution to the outcome. The time-varying layer of FIG. 3 also includes an activation layer 310 and a batch normalization layer 312. The activation layer 310 may include any suitable activation function used in deep neural networks such as Rectified Linear Unit (ReLU) activation function, Gaussian Error Linear Unit (GeLU) activation function, or Swish function / Swish activation function. The aggregation function may be flexible. Two examples of the aggregation function may include:Example aggregation function 1 : f(_z z2, z3) = (z, + z2+ z3)Example aggregation function 2: / '(z1, z2, z3) = [z±, z2, z3] * w where z1;z2, and z3denote the linear- 302, convex- 304, and concave-306 varying effects, respectively. [z1;z2, z3] represents the concatenation of the vectors, w is a trainable parameter matrix, which will be learned during the training period. The linear layers of FIG. 3 may include a fully -connected layer that connects every input neuron to every output neuron of a neural network. Mathematically it is formalized with a matrix of trained parameters Each value in the matrix is multiplied by a function (e.g., from linear- 302, convex- 304, and concave- 306, depending on the considered time-varying mechanism). The input 300 is multiplied by the corresponding matrix and the activation is applied to the elements of the resulting vector. Example function 1 can be viewed as a special case of Example 2, where the parameter matrix w is pre-defined without the need to train. The time-varying layer of FIG. 3 outputs the vector yec(t+At) (^Seq) 314 that represents the influence of the pathogen’s current state on its mutations at time t + At. The time-varying layer of FIG. 3 takes a hidden representation of a sequence at time t calculated by the encoder (transformer encoder component 200 of FIG. 2) and applies a transformations inside the time-varying layer to produce hidden representations at time t + At, that will be used by the transformer decoder. The transformer decoder may additionally perform cross-attention over the output of the time-varying layer, similar to the transformer sequence-to-sequence models. The transformer decoder may use the output of the time-varying layer to learn which positions / tokens / parts of a sequence in the input should stay stable over time and is therefore not influenced by mutations, and which positions / tokens / parts are likely to change and how they are affected by their initial states at time t.Decode a virus mutation at time t + At
[0042] A transformer-like decoder, such as that described in Ashish Vaswani, Noam Shazeeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aiden N. Gomez, Lukasz Kaiser, Illia Polosukhin, “Attention is All you Need,” NIPS, (2017), which is hereby incorporated by reference herein, is used to generate the pathogen’s mutations conditioned on the time-varying influence, represented with Vec(L+M>(Seq) 400 (i.e., the output of the time-varying layer depicted and described above with reference to FIG. 3). FIG. 4 depicts an example of a transformer-like decoder but any suitable transformer decoder architecture may be used with embodiments of the present disclosure.
[0043] It should be noted that as the mutated sequence is known during training, the input to the decoder is the entire sequence 402. However, the amino acids after the predicted one will be masked (see the masked multi-head attention layer 404 of the decoder in FIG. 4). Masking the amino acids after the predicted one via the masked multi-head attention layer 404 is known as teacher forcing, which provides technical improvements to predicting mutation sequences as this will accelerate the convergence of the training of the model(s) described herein.
[0044] For multi-head block 406 in the decoder, we also employ sparse attention, as described above with reference to the transformer encoder 200 of FIG. 2, to further reduce the complexity of the model. As discussed above, this will largely reduce the parameters to be learned and thus decrease the complexity of the model for efficient training with the limited size of the available data. The decoder of FIG. 4 also includes one or more normalization layers with residual addition 408, a position-wise feed forward network layer 410, and a feed forward and Softmax layer 412. The decoder of FIG. 4 also includes an embedding layer 414 with an amino acid embedding 416 and amino acid position embedding 418 similar to the transformer encoder 200 of FIG. 2. The output 420 of the decoder of FIG. 4 includes top-k mutated sequences, rather than just one mutated sequence, for a user-specified time interval t along with a value that represents a probability for each of the top-k mutated sequences. In embodiments, the transformer-like decoder uses the second vector (e.g. output of the time-varying layer) as one of the inputs at every iteration while it is generating a next sequence token for the output 420. The output 420 is generated by the transformer-like decoder auto-regressively, such as in Bayesian Additive Regression Trees (BART) and Generative Pre-Trained Transformer (GPT) models with respect to the second vector. For example, the transformer-like decoder copies informationfrom the output of the decoder to its input, and the transformer-like decoder generates the next sequence token (next amino acid symbol in this scenario) with respect to the second vector. The updated sequence from the output of the decoder is then copied again to its input, and this is repeated until it generates the full sequence (output 420).Training the parameters of the model to capture mutation patterns of the virus
[0045] The time-varying virus mutation model can be learned from training data, i.e., learning the parameters of the time-varying transformer model by minimizing the loss of the training data. The training data should include allowed symbols (e.g. amino acid symbols for protein sequences) and be relevant for the application domain (e.g. belong to relevant protein families or pathogen genomes). In an exemplary embodiment the training data is a set of pairs of sequences (sequence Seqtat time t, its mutated version Seqtlat the next observed time slice t’) of the analyzed genomic feature such as viral protein, where the time interval At = t' — t is not evenly distributed. The next observed time slice t’ may represent whichever next time slice is associated with the training data. For each pair of sequence Seqtand its mutated version SeqL+&Linthe training data, the loss is computed as
[0046] In the first term, -f denotes the position of amino acid in the sequence, which conditions on the amino acids Seqt+&tin previous positions and the pathogen’s sequence Seqtat time t. The predictive probability is the output of the decode layer, i.e., the results of the Softmax layer 412 in FIG. 4. The second term T) is an auxiliary training objective to guarantee the generated mutated sequence is valid from biological point of view. An example can be T) = log(f (SeqL+M)) that accounts for structural constraints, where the function f( eqL+M) is decided by the mapping between the primary sequence and its structural properties (e.g., Frobenius norm of the residue couplings or sequence “energy” in Potts model as described in Allan Haldane, Ronald M Levy, “Mi3-GPU: MCMC-based Inverse Ising Inference on GPUs for protein covariation analysis,” Comput. Phys. Commun., 260: 107312, doi: 10.1016 / j.cpc.2020.107312, (March 2021), which is hereby incorporated by reference herein). For example, g() above from the second term can include energy functions that reflect the ability of a sequence to fold (in this embodiment, a lower value is better). The coefficient T is a hyper-parameter that defines how to tradeoff between mutation generation and validation. The coefficient T reflects the ‘weight’ or the influence of the auxiliary loss to the total loss objective (the less weight corresponds to less of an influence of L(g())). It can be predefined or tuned using any suitable hyper parameter optimization method from such as grid search or Bayesian optimization. The coefficient will be tuned with a cross-validation procedure. The cross-validation procedure may be used to evaluate the performance of the model depending on the value of the hyper-parameter T.
[0047] Given the loss function, the virus evolution model can be trained by minimizing the loss using any suitable optimization algorithm, e.g., stochastic gradient descent.Predict top-k mutated sequences for a given time interval with uncertainty estimation and mutation pruning strategy
[0048] The system of the present disclosure can predict top-k mutated sequences (rather than one) of a virus, cancer sample, or cancer tumor for a user-specified time interval At* and estimate the uncertainty of the prediction. The system of the present disclosure can implement the inference phase on edge devices using a trained model according to the description provided herein thereby utilizing less computationally expensive computer resources. The system of the present disclosure provides users with additional flexibility of specifying any time intervals that are interesting, instead of a pre-defined time interval that constraints the users. Moreover, embodiments of the present disclosure may output a probability of each predicted mutated sequence such that users may be provided with information on how confident the predictions could be - which increases the transparency of the Al -driven results. The predictions generated by the features of the present disclosure also integrate virus mutation constraints as rules to remove unrealistic mutations. FIG. 5 depicts an overview of the mutation prediction. As described above, the system of FIG. 5 uses the time-varying context vector Vec^t+&tSeq~) 500, as a condition, and the transformer decoder 502 computes the probabilities of the next top-k amino acids at each subsequent position based on the previously predicted. For example, FIG. 5 shows that the transformer decoder 502 predicts which amino acid 504 will be the next one, given the currently predicted sequence MFVF 506. Then the predicted amino acid 504 will be added to the sequence for the prediction of the next position. Similar to the description provided above with reference to FIG. 1, FIG. 5 also includes Seq^ 508 being provided as input to the transformer encoder 510 for generating vector Vec^fSeq) 512. The time-varying layer 514, which uses the user-specified time interval At 516 to generate V ec(L+M>(Seq) 500.
[0049] To identify the top-k most likely mutated sequences of a virus, the present disclosure uses a beam-search method illustrated in FIG. 6, similar to that as described in Ashwin K Vijayakumar, Michael Cogswell, Ramprasath R. Selvaraju, Qing Sun, Stefan Lee, David Crandall, Dhruv Batra, “Diverse Beam Search: Decoding Diverse Solutions from Neural Sequence Models,” arXiv preprint ArXiv: 1610.02424, (2018), which is hereby incorporated by reference herein. The beam-search method integrates biologically relevant constraints though a pruning strategy. For example, at a very first position of a protein sequence 602, the system of the present disclosure computes the probability of each possible amino acid with the time-varying context vector Vec(L+M>(Seq) as a condition. Then, the system selects top-k amino acids with the largest predictive probabilities. At the second position 604, for each of the top-k selected amino acids at the first position 602, the system computes the probabilities of all possible amino acids, computes the joint probability of both positions 602 and 604, and selects top-k amino acids at position 604 that deliver the highest joint probabilities of the positions. The system repeats this strategy with position 606, where top-k amino acids are selected to provide the largest joint probabilities of amino acids at positions 602, 604 and 606. The same strategy is repeated in an iterative way until sequence position 608 is reached. Biological restrictions may be used in the loss component, described above, and applied at training time to the full-length generated sequences. The pruning strategy continues with the selection procedure until the system of the present disclosure predicts the [end] symbol 610, where the [end] symbol 610 represents a special token that is used in Natural Language Processing models. Finally, the top-k most likely mutated sequences (predicted sequences) are generated, together with their predictive probabilities 612. FIG. 6 also depicts the transformer decoder 614, the time-varying context vector Vec(L+&L>(Seq) 616, and the prediction of which amino acid 618 will be the next amino acid, given the currently predicted sequence MFVF 620.
[0050] The pruning strategy or beam-search method of the present disclosure may also impose biological restrictions on the mutations during the pruning strategy, such as:Using scores from amino acid substitution matrices, such as described in Rakesh Trived, Hampapathalu Adimurthy Nagarajaram, “Substitution scoring matrices for proteins - An overview,” Protein Sci., 29(11): 2150-2163, doi: 10.1002 / pro / 3954, (November 2020), which is hereby incorporated by reference herein. The amino acid substitution matrices may include standard matrices such as a Point Accepted Mutation (PAM) matrix, a BLOcks SUbstiution Matrix (BLOSUM), or variable time maximum likelihood (VTML) substitution matrix, which may be examples of 20 x 20 matrices in which the individual elements encapsulate the rates at which each of the 20 amino acid residues in proteins are substituted by other amino acid residues over time. In some implementations, the transformer decoder of the present disclosure may use a deterministic function h score) of the elements of the amino acid substitution matrices as weights to rescale the probability of the next amino acid. For example, assuming a scenario where a distribution of frequencies of amino acids 5) met at each sequence position k, it can be assumed that: h(scorejk) = freqjk. To calculate the final probabilities of amino acid j at some position k of a sequence: pk' (Sj- which represents the probabilitythat an amino acid j at position k if amino acid i at the preceding position (k-1), where p'Q represents the updated probability.- An example of a biological rule may be that in some scenarios a given amino acid site is not allowed to mutate because it would result in damages to an expressed protein using that predicted amino acid sequence. In such scenarios the system of the present disclosure would only allow one branch to be generated in the beam search.
[0051] Embodiments of the present disclosure thus provide for general improvements to computers in machine learning to enable to predict top-k most likely mutated molecular sequences, with improved accuracy and computational efficiency. Moreover, embodiments of the present disclosure can be practically applied to use cases to effect further improvements in a number of technical fields including, but not limited to, medical and healthcare (e.g., digital medicine, personalized healthcare, Al-assisted drug or vaccine development, diagnosis therapy, treatment, wet lab experiments etc.), and in particular sequence mutation prediction, and also allows to support interactive decision making by clinicians or physicians such as to support diagnosis and / or treatment recommendation and planning, in automated and semi-automated manners.
[0052] One embodiment can be practically applied in the medical domain for vaccine design against emerging infectious diseases, for example, through predicting and selecting antigen candidates for antiviral vaccines that account for potential / putative future virus variants of concern. Here, a use case is that, for certain pathogens or viruses, such as SARS-CoV-2, vaccine developers typically need to spend a long period (in some cases years) to identify the correct antigen candidates for antiviral vaccines. Moreover, vaccine development for certain pathogens is extremely complex due to the quick and varied mutation rates for these pathogens.Embodiments of the present disclosure enable the transformer-based method to learn the virus mutation patterns from the training data and predict the top-k most likely variants of a virus, such as SARS-CoV-2. This would allow vaccine developers to make a ‘mutant-proof vaccine against a selected virus and account for potential new high-risk variants as an enhancement of existing epitope selection and vaccine optimization pipelines. The data source and input for this use case includes measurements and genome sequences of a virus identified by an expert at different sampling dates, e.g., the SARS-CoV-2, Influenza sequences with the time stamps of acquisition. The output can be, for a specified time interval, a prediction of top-k most likely virus variants with the confidence estimation of the mutants. In this way, the present disclosure enables to provide accurate simulations of virus evolution based on real -world measurements. The identified confidence specifies how likely the predicted mutations can happen for the specified time interval. Vaccine developers, doctors, or clinicians who provide the specified timeinterval can use the predicted mutations to proactively design vaccines against emerging infectious diseases. The predicted mutations can also be used to operate wet lab equipment (e.g., autoclaves) or medical testing equipment, provide automated suggestions or interactive guidance for vaccines or to design vaccines in an automated or semi-automated manner, provide automated diagnoses or treatments, to reserve space in hospitals, to design molecular diagnostic tools or molecular diagnostic tests (e.g. silico aptamer design), or other automated medical planning.
[0053] Another embodiment can be practically applied in the medical domain and treatment recommendation and planning, for example, through predicting cancer evolution for a patient to develop personalized cancer treatment options. Here, a use case includes modeling possible scenarios of cancer tumor evolution based on their genome sequence data. Identifying cancer tumor evolutions can typically be very difficult due to the many varied types of cancer and the particularities of a patient suffering from the type of cancer (e.g., environmental factors or other biological factors). Moreover, predictions generated when dealing with cancer tumor evolutions must be accurate to drive treatment decisions given the fatal nature of cancer. Understanding how tumor clones may develop in terms of anticipated mutations will provide valuable feedback to optimize cancer vaccine design or tailor a personalized treatment strategy. Embodiments of the present disclosure enable the transformer-based method to learn the mutation patterns in specified genes from the training data and predict the top-k most likely mutants in the line of cancer cells. The predicted mutations may be used by doctors to anticipate treatment for a patient. The data source and input for this use case includes genome sequences of cancer tumor samples sequenced at time intervals. The output can be, for a specified time interval, such as a time interval specified by a doctor treating a patient, a prediction of top-k most likely mutated cancer sequences of a patient with an estimated uncertainty of the prediction. In this way, the present disclosure enables to provide accurate simulations of cancer evolution based on real- world measurements. The predicted mutations and estimated uncertainty may provide valuable information for the doctors to use to determine possible treatment strategies and how confident they can be trusted for the patient. The predicted mutated cancer sequences can also be used to operate wet lab equipment (e.g., autoclaves) or medical testing equipment, provide automated suggestions or interactive guidance for vaccines or to design vaccines in an automated or semiautomated manner, provide for automated diagnoses or treatments, or to design molecular diagnostic tools or molecular diagnostic tests (e.g. silico aptamer design).
[0054] In an embodiment, the present disclosure provides a method to predict top-k most likely mutated molecular sequences (e.g. most likely mutated pathogen’s protein sequences or most likely mutated genome sequences), the method is comprising the steps of:1. Encoding a pathogen sequence at time t;2. Learning time-varying influence of the pathogen at time t+At;3. Decoding pathogen mutations at time t+At;4. Training the parameters of the model to capture mutation patterns of the pathogen;5. Predicting top-k mutated sequences for a given time interval with uncertainty estimation and mutation pruning strategy.
[0055] Embodiments of the present disclosure provide for the following improvements and technical advantages over existing technology:1) Adding a new time-varying layer to the transformer, by which a time-varying context vector of a genome (or protein) sequence is learned to integrate the time factor into mutation prediction. This new layer can clearly learn diverse influence changing mechanisms of mutations, including linear, convex and concave, to predict mutations precisely.2) A method to predict top-k possible mutated sequences, rather than a single sequence, for a given time interval, in which: a. Beam-search is used by computing the probabilities of k most likely amino acids for each subsequent position. b. Evolution constraints (domain knowledge in the molecular biology) are encoded as part of the prediction objectives to prune unrealistic mutations by updating the probability of the next amino acid. c. Restrictions based on mutation rates are integrated into a pruning strategy by manipulating the predictive probabilities with respect to the amino acid substitution scores (amino acid substitution matrices).
[0056] Current methods primarily focus on comparing pathogen genomes or prioritizing known mutations to identify conserved regions in the genomes, which assumes prior knowledge of observed mutations. In contrast, the present disclosure allows to predict the top-k mutants expected to emerge within a specific time frame which may not have appeared yet.Embodiments of the present disclosure address this computationally challenging task with an Al method that explores a complex genotype space and generates plausible mutation pathways that a pathogen would adopt thereby changing the host-induces fitness landscape.
[0057] In contrast to Fritz Obermeyer, Martin Jankowiak, Nikolaos Barkas, Stephen F Schaffner, Jesse D Pyle, Leonid Yurkovetskiy, Matteo Bosso, Daniel J Park, Mehrtash Babadi, Bronwyn L Maclnnis, Jeremy Luban, Pardis C Sabeti, Jacob E Lemieux, “Analysis of 6.4 million SARS-CoV-2 genomes identifies mutations associated with fitness,” Science 376, 1327— 1332, doi: 10.1126 / science.abml208 (2022), which is a state-of-the-art model that forecasts the forthcoming SARS-CoV-2 variants and identifies potential mutation spread scenarios, is limitedto mutations observed in a population - embodiments of the present disclosure predict mutations based on the learned patterns of a pathogen observed in training data (e.g. historical data on SARS-CoV-2 or Influenza virus), but is not limited to mutations seen in the past.
[0058] The approach detailed in Brian Hie, Ellen D. Zhong, Bonnie Berger, Bryan Bryson, “Learning the language of viral evolution and escape,” Science 371, 284-288, (2021) estimates mutations within the viral genome regarding their potential for viral escape, utilizing a protein language model to generate embeddings. This approach searches for mutations in a viral sequence that are biologically viable while being antigenically different due to their impact on biological function. In contrast to the present disclosure, the previously described method doesn’t use the predictive capabilities of the model for generating mutated sequences. Instead, it relies on a Bidirectional Long Short-Term Memory (BiLSTM) recurrent neural network, which is less demanding to data compared to transformer models, but struggles to efficiently learn short-term and long-term dependencies and doesn’t incorporate time stamps in its analysis.Embodiments of the present disclosure integrate sparse attention to reduce the complexity of the model which mitigates the data issue of conventional methods and is flexible to learning any dependencies between amino acids while also considering the time-varying effects of mutations.
[0059] Referring to EIG. 7, a processing system 700 can include one or more processors 702, memory 704, one or more input / output devices 706, one or more sensors 708, one or more user interfaces 710, and one or more actuators 712. Processing system 700 can be representative of each computing system disclosed herein.
[0060] Processors 702 can include one or more distinct processors, each having one or more cores. Each of the distinct processors can have the same or different structure. Processors 702 can include one or more central processing units (CPUs), one or more graphics processing units (GPUs), circuitry (e.g., application specific integrated circuits (ASICs)), digital signal processors (DSPs), and the like. Processors 702 can be mounted to a common substrate or to multiple different substrates.
[0061] Processors 702 are configured to perform a certain function, method, or operation (e.g., are configured to provide for performance of a function, method, or operation) at least when one of the one or more of the distinct processors is capable of performing operations embodying the function, method, or operation. Processors 702 can perform operations embodying the function, method, or operation by, for example, executing code (e.g., interpreting scripts) stored on memory 704 and / or trafficking data through one or more ASICs. Processors 702, and thus processing system 700, can be configured to perform, automatically, any and all functions, methods, and operations disclosed herein. Therefore, processing system 700 can beconfigured to implement any of (e.g., all of) the protocols, devices, mechanisms, systems, and methods described herein.
[0062] For example, when the present disclosure states that a method or device performs task “X” (or that task “X” is performed), such a statement should be understood to disclose that processing system 700 can be configured to perform task “X”. Processing system 700 is configured to perform a function, method, or operation at least when processors 702 are configured to do the same.
[0063] Memory 704 can include volatile memory, non-volatile memory, and any other medium capable of storing data. Each of the volatile memory, non-volatile memory, and any other type of memory can include multiple different memory devices, located at multiple distinct locations and each having a different structure. Memory 704 can include remotely hosted (e.g., cloud) storage.
[0064] Examples of memory 704 include a non-transitory computer-readable media such as RAM, ROM, flash memory, EEPROM, any kind of optical storage disk such as a DVD, a Blu- Ray® disc, magnetic storage, holographic storage, a HDD, a SSD, any medium that can be used to store program code in the form of instructions or data structures, and the like. Any and all of the methods, functions, and operations described herein can be fully embodied in the form of tangible and / or non-transitory machine-readable code (e.g., interpretable scripts) saved in memory 704.
[0065] Input-output devices 706 can include any component for trafficking data such as ports, antennas (i.e., transceivers), printed conductive paths, and the like. Input-output devices 706 can enable wired communication via USB®, DisplayPort®, HDMI®, Ethernet, and the like. Input-output devices 706 can enable electronic, optical, magnetic, and holographic, communication with suitable memory 706. Input-output devices 706 can enable wireless communication via WiFi®, Bluetooth®, cellular (e.g., LTE®, CDMA®, GSM®, WiMax®, NFC®), GPS, and the like. Input-output devices 706 can include wired and / or wireless communication pathways.
[0066] Sensors 708 can capture physical measurements of environment and report the same to processors 702. User interface 710 can include displays, physical buttons, speakers, microphones, keyboards, and the like. Actuators 712 can enable processors 702 to control mechanical forces.
[0067] Processing system 700 can be distributed. For example, some components of processing system 700 can reside in a remote hosted network service (e.g., a cloud computing environment) while other components of processing system 700 can reside in a local computing system. Processing system 700 can have a modular design where certain modules include aplurality of the features / functions shown in FIG. 7. For example, I / O modules can include volatile memory and one or more processors. As another example, individual processor modules can include read-only-memory and / or local caches.
[0068] The following references were described or incorporated by reference above, including: Fritz Obermeyer, Martin Jankowiak, Nikolaos Barkas, Stephen F Schaffner, Jesse D Pyle, Leonid Yurkovetskiy, Matteo Bosso, Daniel J Park, Mehrtash Babadi, Bronwyn L Maclnnis, Jeremy Luban, Pardis C Sabeti, Jacob E Lemieux, “Analysis of 6.4 million SARS- CoV-2 genomes identifies mutations associated with fitness,” Science 376, 1327-1332, doi: 10.1126 / science.abml208 (2022); Brian Hie, Ellen D. Zhong, Bonnie Berger, Bryan Bryson, “Learning the language of viral evolution and escape,” Science 371, 284-288, (2021); Ashish Vaswani, Noam Shazeeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aiden N. Gomez, Lukasz Kaiser, Illia Polosukhin, “Attention is All you Need,” NIPS, (2017); Rewon Child, Scott Gray, Alec Radford, Ilya Sutskever, “Generating long sequences with sparse transformers:, arXiv preprint arXiv: 104.10509, (2019); Allan Haldane, Ronald M Levy, “Mi3-GPU: MCMC-based Inverse Ising Inference on GPUs for protein covariation analysis,” Comput. Phys. Commun., 260: 107312, doi: 10.1016 / j.cpc.2020. 107312, (March 2021); Ashwin K Vijayakumar, Michael Cogswell, Ramprasath R. Selvaraju, Qing Sun, Stefan Lee, David Crandall, Dhruv Batra, “Diverse Beam Search: Decoding Diverse Solutions from Neural Sequence Models,” arXiv preprint ArXiv: 1610.02424, (2018); and Rakesh Trived, Hampapathalu Adimurthy Nagarajaram, “Substitution scoring matrices for proteins - An overview,” Protein Sci., 29(11): 2150-2163, doi: 10.1002 / pro / 3954, (November 2020). In contrast to the existing approaches described in the above references the present disclosure predicts the next mutations and generates sequence candidates by allowing the model to capture historical patterns of molecular evolution and accounting for the properties of the sequences. Most conventional methods analyze a static picture and calculate statistics on top of this static picture in an attempt to forecast where and which mutations are possible.
[0069] While embodiments of the disclosure have been illustrated and described in detail in the drawings and foregoing description, such illustration and description are to be considered illustrative or exemplary and not restrictive. It will be understood that changes and modifications may be made by those of ordinary skill within the scope of embodiments of the present disclosure. In particular, the present disclosure covers further embodiments with any combination of features from different embodiments described herein. Additionally, statements made herein characterizing the invention or disclosure refer to an embodiment of the invention or disclosure and not necessarily all embodiments.
[0070] The terms used in the claims should be construed to have the broadest reasonable interpretation consistent with the foregoing description. For example, the use of the article “a” or “the” in introducing an element should not be interpreted as being exclusive of a plurality of elements. Likewise, the recitation of “or” should be interpreted as being inclusive, such that the recitation of “A or B” is not exclusive of “A and B,” unless it is clear from the context or the foregoing description that only one of A and B is intended. Further, the recitation of “at least one of A, B and C” should be interpreted as one or more of a group of elements consisting of A, B and C, and should not be interpreted as requiring at least one of each of the listed elements A, B and C, regardless of whether A, B and C are related as categories or otherwise. Moreover, the recitation of “A, B and / or C” or “at least one of A, B or C” should be interpreted as including any singular entity from the listed elements, e.g., A, any subset from the listed elements, e.g., A and B, or the entire list of elements A, B and C.
Claims
CLAIMSWhat is claimed is:
1. A computer-implemented, machine learning method for predicting a top-k most likely mutated molecular sequences, the method comprising: encoding a molecular sequence at a first time into a first vector in a latent space; generating, by a time-varying mutation model, a second vector in the latent space using as input the first vector, the second vector indicating a time-varying influence of the molecular sequence on a mutated version of the molecular sequence at a subsequent time; and decoding the second vector to generate a prediction of the top-k most likely mutated molecular sequences for the molecular sequence at the subsequent time.
2. The method according to claim 1, wherein the time-varying mutation model has been trained by minimizing a loss of training data provided as input to the time-varying mutation model, the training data including a training molecular sequence at the first time and a mutated version of the training molecular sequence at the subsequent time.
3. The method according to claim 2, wherein further time intervals are unevenly distributed with a time interval between the first time and the subsequent time, and wherein minimizing the loss includes tuning a coefficient of a loss function with a cross-validation procedure.
4. The method according to claim 2, wherein minimizing the loss includes using a term indicating a training objective of the time-varying mutation model that enforces biologically valid mutations in generated molecular sequences of the predicted top-k most likely mutated molecular sequences, and wherein enforcing the biologically valid mutations includes mapping a primary sequence of the training molecular sequence with structural properties of the primary sequence.
5. The method according to any of the preceding claims, wherein the decoding is performed by a transformer decoder that has been trained to compute a probability of a next molecular component at each subsequent position of the mutated version of the molecular sequence at the subsequent time, and wherein the top-k most likely mutated molecular sequences are based on the probabilities.
6. The method according to claim 5, wherein computing the probabilities includes using rules as evolutionary constraints and restrictions based on mutation rates, wherein the restrictions based on the mutation rates includes using molecular component substitutionmatrices to rescale the probabilities for the next molecular component at each subsequent position, and wherein the rules as the evolutionary constraints prohibit predictions that correspond to incompatible mutations for the molecular sequence.
7. The method according to any of the preceding claims, wherein generating, by the timevarying mutation model, the second vector in the latent space includes using a user specified time interval to determine the subsequent time, and further includes using a plurality of timevarying mechanisms, the plurality of time-varying mechanisms including a linear time-varying mechanism, a convex time-varying mechanism, and a concave time-varying mechanism.
8. The method according to any of the preceding claims, wherein each of the top-k most likely mutated molecular sequences for the molecular sequence at the subsequent time includes a value that indicates a confidence of the prediction for a corresponding molecular sequence mutation.
9. The method according to any of the preceding claims, wherein generating the prediction of the top-k most likely mutated molecular sequences for the molecular sequence at the subsequent time includes using a beam-search method that integrates biological constraints through a pruning strategy by computing, at each position, a probability of a particular molecular component using the second vector as a condition, and selecting a top-k molecular components with largest probabilities and joint probabilities with other positions of the molecular sequence until an end is predicted.
10. The method according to any of the preceding claims, wherein the molecular sequence is a genome sequence for a pathogen, and wherein the method further comprises: obtaining a user specified time interval; decoding the second vector in the latent space to generate the prediction of the top-k most likely mutated molecular sequences for the molecular sequence at the subsequent time using the user specified time interval; and identifying specific amino acids that correspond to the predicted top-k most likely mutated molecular sequences for the molecular sequence at the subsequent time as a basis for a vaccine to the pathogen.
11. The method according to any of the preceding claims, wherein the molecular sequence is a genome sequence for a cancer sample, and wherein the method further comprises: obtaining a user specified time interval;decoding the second vector in the latent space to generate the prediction of the top-k most likely mutated molecular sequences for the molecular sequence at the subsequent time using the user specified time interval; and identifying specific amino acids that correspond to the predicted top-k most likely mutated molecular sequences for the molecular sequence at the subsequent time to identify a treatment strategy for a cancer type associated with the cancer sample.
12. A machine learning model stored on a tangible, non-transitory computer readable medium, the machine learning model comprising: a time-varying mutation layer that learns a time-varying influence of a molecular sequence at a first time on a mutated version of the molecular sequence at a subsequent time, the time-varying mutation layer including three linear layers, an activation layer, an aggregation layer, and a batch normalization layer, the linear layers including a linear, concave, and convex time-varying mechanism.
13. The machine learning model according to claim 12, wherein the machine learning model further comprises a transformer encoder, the transformer encoder including a multi-head attention mechanism layer, one or more normalization layers with residual addition, and a position-wise feed forward network layer, and wherein the transformer encoder has a sparse attention mechanism in the multi-head attention mechanism layer to map the molecular sequence at the first time into a vector in a latent space.
14. The machine learning model according to claim 12, wherein the machine learning model further comprises a transformer decoder, the transformer decoder including a masked multi-head attention layer configured to mask subsequent positions of the molecular sequence to perform a prediction of a probability of a next molecular component at each subsequent position of the mutated version of the molecular sequence at the subsequent time in order to provide a prediction of a top-k most likely mutated molecular sequences.
15. A tangible, non-transitory computer-readable medium containing instructions for predicting a top-k most likely mutated molecular sequences using a machine learning method, the instructions, upon being executed by one or more processors, alone or in combination, provide for execution of the following steps: encoding a molecular sequence at a first time into a first vector in a latent space; generating, by a time varying mutation model, a second vector in the latent space using as input the first vector, the second vector indicating a time-varying influence of the molecular sequence on a mutated version of the molecular sequence at a subsequent time; anddecoding the second vector to generate a prediction of the top-k most likely mutated molecular sequences for the molecular sequence at the subsequent time.