AA2DNA: pre-trained amino acid-to-DNA sequence mapping enabling efficient expression and yield optimization

The AA2DNA framework addresses the challenges of gene optimization by using deep learning to map amino acid sequences into optimized codon sequences, enhancing protein yield and functionality, and improving the efficacy and accessibility of therapeutic proteins.

WO2025128647A1PCT designated stage expired Publication Date: 2025-06-19PROTEINEA INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/US2024/059489
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-10
Filing Date
2024-12-11
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Current methods for gene optimization, particularly in therapeutic applications, face challenges such as inadequate protein yield, misfolding, and aggregation due to data scarcity and high computational costs associated with deep learning techniques.

Method used

The AA2DNA framework uses a deep learning-based approach to map amino acid sequences into optimized codon nucleotide sequences, employing contextual and dynamic codon mapping that accounts for evolutionary features and ensures robustness to long sequences and distant token relationships.

Benefits of technology

The AA2DNA framework enhances protein yield, improves functionality, and increases the efficacy of gene therapy, while reducing costs and improving accessibility of therapeutic proteins by generating optimized codon sequences that improve protein expression and yield.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2024059489_19062025_PF_FP_ABST
    Figure US2024059489_19062025_PF_FP_ABST
Patent Text Reader

Abstract

Systems, methods, and apparatus for generating a DNA / gene sequence output, given a protein sequence input. In one aspect, the protein-to-DNA mapping system includes relative positional embedding to a multi-head transformer, that encodes relative positions of input token pairs and their feature similarities. Relative positional embedding applies to protein or nucleotide sequences and accommodates for longer (inference) sequences, structural distinctions in sequence folding, and corresponding scaled parameters. In another aspect, an optimized codon sequence is generated from an input protein sequence by confining the codon search / sampling to a subset of codons. A third aspect includes a cluster-by-cluster, high-yield codon sequence generator that produces a high yield and / or high-expression codon sequence in response to a protein input sequence. The systems, methods, and apparatus increase protein yield / expression, improve protein functionality, enhance gene therapy efficacy, streamline synthetic biology design, and / or broaden functional genomics studies, all leading to reduced costs, better performance, and increased accessibility.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No. PRTN1008WO01 AA2DNA: Pre-trained Amino Acid-to-DNA Sequence Mapping Enabling Efficient Expression and Yield Optimization PRIORITY DATA

[0001] This application claims the benefit of (priority to) US Provisional Application 63 / 616,896filed on January 2, 2024, entitled “AA2DNA: Amino Acid-to-DNA Sequence Mapping Enabling Efficient Protein Expression and Yield Optimization” (Attorney Docket No. PRTN1008USP01), US Application 18 / 769,001 filed on July 10, 2024, entitled “AA2CDS: Amino Acid-to-DNA Sequence Mapping Enabling Efficient Protein Expression and Yield Optimization” (Attorney Docket No. PRTN1003USN01), and US Provisional Application 63 / 608,546 filed on December 11, 2023, entitled “GeneEval: Enabling High- Confidence Gene Sequence Optimization via In-Silico Assessment” (Attorney Docket No. PRTN1007USP01). BACKGROUND

[0002] Optimization techniques are at the forefront of scientific advancements in the fields ofbiotechnology and genetic engineering. These techniques are designed to enhance gene expression and protein production in specific host organisms through precise modifications of DNA sequences. Despite notable advancements in this domain of optimization, especially via deep learning, significant challenges still exist in the heterologous expression of genes from one organism to another. These challenges lead to restricted protein yield, misfolding, and aggregation.

[0003] Given their data driven nature, deep learning methods are hindered by the scarcity of large-scale host / protein-specific data, or the high computational cost associated with multi-host / multi-protein data. Furthermore, deep learning methods are hindered by their limitation in covering long context lengths while retaining the relevance of distance tokens, which are of special importance in protein sequences. Finally, deep learning methods, in the best cases, produce a viable mapping of the protein sequence rather than an optimum one.

[0004] Gene optimization is a process of particular importance in improving the expression and yieldof recombinant proteins within specific host organisms. This optimization process targets specific regions within the plasmid - the extrachromosomal genetic vector that is being used to express the recombinant protein. These specific regions include the promoter, RBS (Ribosome Binding Site), CDSs (DNA Coding Sequences), signal peptides, and transcriptional terminators. Each of these regions plays a pivotal role in regulating gene expression and protein production.

[0005] The potency of gene optimization becomes particularly pronounced when consideringrecombinant protein production for therapeutic applications. Within this specialized arena, addressing the challenges associated with cross-species gene expression is paramount. Such challenges include grappling with issues of inadequate protein yield, misfolding, and aggregation – factors that can significantly impede the efficacy of therapeutic interventions. By optimizing the promoter region, efficient transcriptionAttorney Docket No. PRTN1008WO01 initiation can be ensured. Optimization of the RBS facilitates proper ribosome binding and translation initiation. Additionally, codon optimization of CDSs and signal peptides involves adapting the codons of the heterologous sequence to match the preferred codon usage of the host organism, leading to improved protein expression levels. Furthermore, optimizing the signal peptide can ensure correct targeting and localization of the protein, affecting its stability, functionality, and ultimately its expression in the context of its intended cellular environment. Finally, proper optimization of the terminator region ensures accurate transcription termination, preventing unintended read-through or interference with neighboring genes. By meticulously tailoring the DNA sequence to align with the host organism, researchers can overcome challenges associated with heterologous gene expression thereby enhancing both the quality and quantity of the produced protein.

[0006] The significance of yield improvement within this therapeutic framework extends far beyondthe laboratory bench. Indeed, the ability to control protein expression for yield quantity and quality improvement purposes has direct implications for the affordability and accessibility of critical / high-impact drugs. Optimized protein production stands as a potential cornerstone for driving down the costs of manufacturing, thereby increasing the accessibility of the biology in question. This facet resonates particularly and profoundly in the therapeutic landscape, where equitable access to effective treatments is of paramount importance.

[0007] The existing methods for gene optimization are classified as rule-based and machine learning-based. Rule-based methods hinge on predetermined criteria derived from factors such as codon usage frequency, mRNA secondary structure and stability, GC content, and restriction sites, often sourced from literature studies. However, these methods have significant limitations. Firstly, changes to synonymous codons can adversely impact the protein's structure, including its conformation, folding, stability, and post- translational modification sites. Additionally, relying solely on straightforward criteria like codon usage frequency and mRNA folding energy overlooks key biological factors influencing protein expression, such as sequence context, tRNA availability, ribosome pausing, mRNA degradation, and translation kinetics. Lastly, the optimal codons vary based on the host organism and protein type, making the generalization of rule-based techniques a source of inconsistency.

[0008] The generation of optimum gene sequences (i.e., sequences that drastically and reliablyimprove the yield of the protein) cannot occur without the presence of high-confidence in-silico evaluation enabling the high-throughput computational characterization of the wide search space of viable sequences. The presence of such means of evaluation can effectively bridge the gap between design and testing but preventing the timely cycles involved with testing non-optimum variants through early filtration. However, pursuing such means of evaluation is hindered by the scarcity of experimentally validated and labeled large-scale gene sequences satisfying therapeutic-prone hosts and protein classes. Such a significant obstacle impedes the effective deployment of traditional deep learning techniques in addressing realistic and valuable research problems in therapeutic design / engineering, thereby relegating their achievements to theoretical realms. The current research trend suggests that model performance may be improved by eitherAttorney Docket No. PRTN1008WO01 increasing the number of parameters or improving the architecture of the model. This direction has reached a limit where the model capacities surpass the sizes of the data on which they are trained. In the case of therapeutic application, the data sizes are orders of magnitude smaller than those in traditional applications. Therefore, the challenge of producing high-confidence gene sequence computational assessment given a low-scale host-specific / protein-specific dataset and thereby providing reliable and effective yield optimization remains unsolved.

[0009] Despite the transformative impact of machine learning-based techniques, particularly thoseharnessing the power of deep neural networks and specialized language models due to their ability to understand the underlying features in coding sequences and the relationship between host / protein-specific optimization, their practical success is often hindered by the prevalent issue of data scarcity, an issue that is heightened when the application of interest is yield optimization for therapeutics. For instance, if the modeling focus was the coding sequence (CDS), utilizing a sequence-to-sequence (Seq2Seq) language model to produce a host-specific codon sequence (output) given an amino acid sequence (input) requires large-scale host / protein-specific data. In fact, performing the pre-training of the designated language model from scratch entails the fulfillment of several requirements revolving around data diversity, comprehensiveness, generality, and large scale. Such adaptation of Seq2Seq language models is rather hindered due to the increased sequence length (given each amino acid is encoded by three nucleotides, the sequence length triples). Moreover, the direct adaptation of the Seq2Seq framework does not guarantee the generation of viable codon sequences (i.e., sequences that do map back to the original protein). The alternative approach of utilizing multi-host / multi-protein data requires scalability of the modeling architecture, resulting in a massive computational prerequisite cost.

[0010] Providing these challenges were mitigated via a supervised machine translation or unsupervised conditional generation frameworks, this adaptation only allows the model to produce a “valid” translation rather than an optimized translation given the inadequacy of high-quality amino acid sequence-to-codon sequence mappings (i.e., producing a codon sequence that corresponds to higher expression while retaining / improving its functional attributes). Therefore, the problem of producing a high- quality codon sequence given a low-scale host-specific / protein-specific dataset remains unsolved. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] FIG.1 shows one example of a relative positional embedding to a multi-head attention transformer system.

[0011] FIG.2 is a flowchart of one example of providing relative positional embedding to a multi- head attention transformer.

[0012] FIG.3 shows one example of a confined search of a codon element within a sequence of codon elements.

[0013] FIG.4 is a flowchart of one example of generating an output sequence of codon elements by looking back to confine the search.Attorney Docket No. PRTN1008WO01

[0014] FIG.5 shows training the cluster-by-cluster high-yield DNA sequence generator via supervised learning.

[0015] FIG.6 shows generating the high-yield DNA sequence output via the trained cluster-by- cluster high-yield DNA sequence generator.

[0016] FIG.7 shows the supervised machine translation technique to train the high-yield codon sequence generator.

[0017] FIG.8 shows one example of a high-yield codon sequence generator system.

[0018] FIG.9 is a flowchart of one example of generating higher yield codon sequences.

[0019] FIG.10 is a simplified block diagram of a computer system that can be used to implement the technology disclosed.

[0020] FIG.11 includes a chart that shows the fold-change in yield / expression for five proteins that were each expressed in the HEK293 cell line.

[0021] FIG.12 includes a chart that shows the fold-change in yield / expression for three proteins that were each expressed in the Yeast Pichia cell line.

[0022] FIG.13 includes a chart that shows the fold-change in yield / expression for two proteins that were each expressed in the HEK293 cell line.

[0023] FIG.14 includes a table that shows the fold-change in yield / expression for two protein classes, each with several corresponding proteins, that were each expressed in the CHO cell line.

[0024] FIG.15A includes a table that shows the fold-change in yield / expression for the first of two protein classes, each with several corresponding protein indices, that were each expressed in the HEK293 cell line.

[0025] FIG.15B includes a table that shows the fold-change in yield / expression for the second of two protein classes, each with several corresponding proteins, that were expressed in the HEK293 cell line.

[0026] FIG.16 is a diagram showing one example of multi-task learning and ensemble-based fine tuning of a sequence evaluation system.

[0027] FIG.17 is a table showing sequence evaluation performance compared to fold-change in experimental yields across five mAb proteins. DETAILED DESCRIPTION

[0028] Disclosed is a deep learning-based framework that maps amino acid sequences into optimized codon nucleotide sequences given a low-scale host-specific / protein-specific dataset, referred to as “AA2DNA.” The AA2DNA framework provides contextual and dynamic codon mapping. This mapping accounts for the evolutionary features of the protein which in turn opts for more conservative alterations in terms of the resulting structural, functional, and biophysical attributes. The AA2DNA framework ensures the generation of a valid DNA / gene sequence with robustness to long sequences and their distant token relationships. Finally, the AA2DNA framework enables the generation of optimized codon sequences with improved yield.Attorney Docket No. PRTN1008WO01

[0029] The utility of the AA2DNA framework translates to increased protein yield, improved functionality, enhanced gene therapy efficacy, streamlined synthetic biology design, and broadened functional genomics studies, all leading to reduced costs, better performance, and increased accessibility.

[0030] AA2DNA is a framework customized for mapping an amino acid sequence into a prominent DNA / gene sequence satisfying pre-determined constraints such as, but not limited to, host organism or protein type. AA2DNA supports application on a subset / full plasmid map, with a bias for CDS optimization for yield improvement.

[0031] AA2DNA seeks to evade the challenges introduced via the utilization of language models for the codon optimization problem. Firstly, AA2DNA adopts a novel positional embedding that provides accurate extrapolation for longer sequences while accommodating the impact of sequence folding on the relevance between distant tokens. Secondly, AA2DNA constrains the sampling of codons per amino acid that correspond to a valid sequence mapping and provide a highly efficient mapping. Thirdly, AA2DNA allows for generating a high yield codon sequence by enabling the use of a cluster-by-cluster high yield DNA sequence generator on an amino acid sequence basis.

[0032] Such source models are often pre-trained using tens to hundreds of millions of unsupervised sequences and high-complexity architectures encompassing a massive number of trainable parameters and therefore, requiring a massive computational cost. Given the existence of such source / base models in both the protein and DNA / gene domain, pre-training a machine translation network from scratch would not be needed for the default use of the framework. Utilization of embeddings of such source / base models can be viewed to introduce higher-dimensional dense information to the amino acid and codon sequences with dramatically less computational resources, making the task of inferring predictions out of their mappings more approachable. Therefore, and finally, AA2DNA supports the generation of optimum CDS sequences (as opposed to just viable CDS sequences) by enabling the mapping of an arbitrary CDS into another of higher expression and yield.

[0033] One example of the AA2DNA is a neural network system. In one implementation, the neural network system processes the input representations as input and generates the output representations as output. In some implementations, the neural network system is at least one of a language model neural network, a sequence-to-sequence neural network, an encoder-decoder neural network, an autoencoder neural network, a variational autoencoder neural network, a generative adversarial neural network, a diffusion neural network, a Transformer neural network, a recurrent neural network, a long-short term memory neural network, an autoregressive neural network, an energy-based neural network, and a flow- based neural network.

[0034] Some implementations of the technology disclosed relate to using a Transformer model to provide the AA2DNA framework. In particular, the technology disclosed proposes a parallel input, parallel output (PIPO) the AA2DNA framework based on the Transformer architecture. The Transformer model relies on a self-attention mechanism to compute a series of context-informed vector-space representations of elements in the input sequence and the output sequence, which are then used to predict distributions overAttorney Docket No. PRTN1008WO01 subsequent elements as the model predicts the output sequence element-by-element. Not only is this mechanism straightforward to parallelize, but as each input’s representation is also directly informed by all other inputs’ representations, this results in an effectively global receptive field across the whole input sequence. This stands in contrast to, e.g., convolutional architectures which typically only have a limited receptive field.

[0035] In one implementation, the disclosed AA2DNA framework is a multilayer perceptron (MLP). In another implementation, the disclosed AA2DNA framework is a feedforward neural network. In yet another implementation, the disclosed AA2DNA framework is a fully connected neural network. In a further implementation, the disclosed AA2DNA framework is a fully convolution neural network. In a yet further implementation, the disclosed AA2DNA framework is a semantic segmentation neural network. In yet another further implementation, the disclosed AA2DNA framework is a generative adversarial network (GAN) (e.g., CycleGAN, StyleGAN, pixelRNN, text-2-image, DiscoGAN, IsGAN). In yet another implementation, the disclosed AA2DNA framework includes self-attention mechanisms like Transformer, Vision Transformer (ViT), Bidirectional Transformer (BERT), Detection Transformer (DETR), Deformable DETR, UP-DETR, DeiT, Swin, GPT, iGPT, GPT- 2, GPT-3, various ChatGPT versions, various LLaMA versions, BERT, SpanBERT, RoBERTa, XLNet, ELECTRA, UniLM, BART, T5, ERNIE (THU), KnowBERT, DeiT-Ti, DeiT-S, DeiT-B, T2T-ViT-14, T2T-ViT-19, T2T-ViT-24, PVT- Small, PVT-Medium, PVT-Large, TNT-S, TNT-B, CPVT-S, CPVT-S-GAP, CPVT-B, Swin-T, Swin-S, Swin-B, Twins-SVT-S, Twins-SVT-B, Twins-SVT-L, Shuffle-T, Shuffle-S, Shuffle-B, XCiT-S12 / 16, CMT-S, CMT-B, VOLO-D1, VOLO-D2, VOLO-D3, VOLO-D4, MoCo v3, ACT, TSP, Max-DeepLab, VisTR, SETR, Hand-Transformer, HOT-Net, METRO, Image Transformer, Taming Transformer, TransGAN, IPT, TTSR, STTN, Masked Transformer, CLIP, DALL-E, Cogview, UniT, ASH, TinyBert, FullyQT, ConvBert, FCOS, Faster R-CNN + FPN, DETR-DC5, TSP-FCOS, TSP-RCNN, ACT+MKDD (L=32), ACT+MKDD (L=16), SMCA, Efficient DETR, UP-DETR, UP-DETR, ViTB / 16-FRCNN, ViT-B / 16-FRCNN, PVT- Small+RetinaNet, Swin-T+RetinaNet, Swin-T+ATSS, PVT-Small+DETR, TNT-S+DETR, YOLOS-Ti, YOLOS-S, and YOLOS-B.

[0036] In one implementation, the disclosed AA2DNA framework is a convolution neural network (CNN) with a plurality of convolution layers. In another implementation, the disclosed AA2DNA framework is a recurrent neural network (RNN) such as a long short-term memory network (LSTM), bi- directional LSTM (Bi- LSTM), or a gated recurrent unit (GRU). In yet another implementation, the disclosed AA2DNA framework includes both a CNN and an RNN.

[0037] In yet other implementations, the disclosed AA2DNA framework can use 1D convolutions, 2D convolutions, 3D convolutions, 4D convolutions, 5D convolutions, dilated or atrous convolutions, transpose convolutions, depth-wise separable convolutions, pointwise convolutions, 1 x 1 convolutions, group convolutions, flattened convolutions, spatial and cross-channel convolutions, shuffled grouped convolutions, spatial separable convolutions, and deconvolutions.Attorney Docket No. PRTN1008WO01

[0038] The disclosed AA2DNA framework can use one or more loss functions such as logistic regression / log loss, multi-class cross-entropy / softmax loss, binary cross-entropy loss, mean-squared error loss, L1 loss, L2 loss, smooth L1 loss, and Huber loss. The disclosed AA2DNA framework can use any parallelism, efficiency, and compression schemes such TFRecords, compressed encoding (e.g., PNG), sharding, parallel calls for map transformation, batching, prefetching, model parallelism, data parallelism, and synchronous / asynchronous stochastic gradient descent (SGD). The disclosed AA2DNA framework can include upsampling layers, downsampling layers, recurrent connections, gates and gated memory units (like an LSTM or GRU), residual blocks, residual connections, highway connections, skip connections, peephole connections, activation functions (e.g., non-linear transformation functions like rectifying linear unit (ReLU), leaky ReLU, exponential liner unit (ELU), sigmoid and hyperbolic tangent (tanh)), batch normalization layers, regularization layers, dropout, pooling layers (e.g., max or average pooling), global average pooling layers, and attention mechanisms.

[0039] The disclosed AA2DNA framework can be a linear regression model, a logistic regression model, an Elastic Net model, a support vector machine (SVM), a random forest (RF), a decision tree, and a boosted decision tree (e.g., XGBoost), or some other tree-based logic (e.g., metric trees, kd-trees, R-trees, universal B-trees, X-trees, ball trees, locality sensitive hashes, and inverted indexes). The disclosed AA2DNA framework can be an ensemble of multiple models, in some implementations.

[0040] In some implementations, the disclosed AA2DNA framework can be trained using backpropagation-based gradient update techniques. Example gradient descent techniques that can be used for training the disclosed AA2DNA framework include stochastic gradient descent, batch gradient descent, and mini-batch gradient descent. Some examples of gradient descent optimization algorithms that can be used to train the disclosed AA2DNA framework are Momentum, Nesterov accelerated gradient, Adagrad, Adadelta, RMSprop, Adam, AdaMax, Nadam, and AMSGrad.

[0041] In some implementations, AA2DNA can include training a Seq2Seq model from scratch, using pre-trained embeddings for a Seq2Seq model, generating codon sequences under constrained sampling, using scaled rotary positional embeddings, using DPO, generating higher-yield / expression output sequences from lower-yield / expression input sequences, and / or a subset / combination of the aforementioned. Transformer Logic

[0042] Machine learning is the use and development of computer systems that can learn and adapt without following explicit instructions, by using algorithms and statistical models to analyze and draw inferences from patterns in data. Some of the state-of-the-art models use Transformers, a more powerful and faster model than neural networks alone. Transformers originate from the field of natural language processing (NLP), but Transformers can be used in computer vision and many other fields. Neural networks process input in series and weight relationships by distance in the series. Transformers can process input in parallel and do not necessarily weigh by distance. For example, in natural language processing, neural networks process a sentence from beginning to end with the weights of words close toAttorney Docket No. PRTN1008WO01 each other being higher than those further apart. This leaves the end of the sentence very disconnected from the beginning causing an effect called the vanishing gradient problem. Transformers look at each word in parallel and determine weights for the relationships to each of the other words in the sentence. These relationships are called hidden states because they are later condensed for use into one vector called the context vector. Transformers can be used in addition to neural networks. This architecture is described in the following sections. Encoder-Decoder Architecture

[0043] An encoder-decoder architecture is often used for NLP and has two main building blocks. The first building block is the encoder that encodes an input into a fixed-size vector. In the system we describe here, the encoder is based on a recurrent neural network (RNN). At each time step, t, a hidden state of time step, t-1, is combined with the input value at time step t to compute the hidden state at timestep t. The hidden state at the last time step, encoded in a context vector, contains relationships encoded at all previous time steps. For NLP, each step corresponds to a word. Then the context vector contains information about the grammar and the sentence structure. The context vector can be considered a low-dimensional representation of the entire input space. For NLP, the input space is a sentence, and a training set consists of many sentences.

[0044] The context vector is then passed to the second building block, the decoder. For translation, the decoder has been trained on a second language. Conditioned on the input context vector, the decoder generates an output sequence. At each time step, t, the decoder is fed the hidden state of time step, t-1, and the output generated at time step, t-1. The first hidden state in the decoder is the context vector, generated by the encoder. The context vector is used by the decoder to perform the translation.

[0045] The whole model is optimized end-to-end by using backpropagation, a method of training a neural network in which the initial system output is compared to the desired output and the system is adjusted until the difference is minimized. In backpropagation, the encoder is trained to extract the right information from the input sequence, the decoder is trained to capture the grammar and vocabulary of the output language. This results in a fluent model that uses context and generalizes well. When training an encoder-decoder model, the real output sequence is used to train the model to prevent mistakes from stacking. When testing the model, the previously predicted output value is used to predict the next one.

[0046] When performing a translation task using the encoder-decoder architecture, all information about the input sequence is forced into one vector, the context vector. Information connecting the beginning of the sentence with the end is lost, the vanishing gradient problem. Also, different parts of the input sequence are important for different parts of the output sequence, information that cannot be learned using only RNNs in an encoder-decoder architecture. Attention Mechanism

[0047] Attention mechanisms distinguish Transformers from other machine learning models. The attention mechanism provides a solution for the vanishing gradient problem. For example, an attention mechanism can be added onto an RNN encoder-decoder architecture. At every step in this example, theAttorney Docket No. PRTN1008WO01 decoder is given an attention score, e, for each encoder hidden state. In other words, the decoder is given weights for each relationship between words in a sentence. The decoder uses the attention score concatenated with the context vector during decoding. The output of the decoder at time step t is based on all encoder hidden states and the attention outputs. The attention output captures the relevant context for time step t from the original sentence. Thus, words at the end of a sentence may now have a strong relationship with words at the beginning of the sentence. In the sentence “The quick brown fox, upon arriving at the doghouse, jumped over the lazy dog,” fox and dog can be closely related despite being far apart in this complex sentence.

[0048] To weight encoder hidden states, a dot product between the decoder hidden state of the current time step, and all encoder hidden states, is calculated. This results in an attention score for every encoder hidden state. The attention scores are higher for those encoder hidden states that are similar to the decoder hidden state of the current time step. Higher values for the dot product indicate the vectors are pointing more closely in the same direction. The attention scores are converted to fractions that sum to one using the SoftMax function.

[0049] The SoftMax scores provide an attention distribution. The x-axis of the distribution is position in a sentence. The y-axis is attention weight. The scores show which encoder hidden states are most closely related. The SoftMax scores specify which encoder hidden states are the most relevant for the decoder hidden state of the current time step.

[0050] The elements of the attention distribution are used as weights to calculate a weighted sum over the different encoder hidden states. The outcome of the weighted sum is called the attention output. The attention output is used to predict the output, often in combination (concatenation) with the decoder hidden states. Thus, both information about the inputs, as well as the already generated outputs, can be used to predict the next outputs.

[0051] By making it possible to focus on specific parts of the input in every decoder step, the attention mechanism solves the vanishing gradient problem. By using attention, information flows more directly to the decoder. It does not pass through many hidden states. Interpreting the attention step can give insights into the data. Attention can be thought of as a soft alignment. The words in the input sequence with a high attention score align with the current target word. Attention describes long-range dependencies better than RNN alone. This enables analysis of longer, more complex sentences.

[0052] The attention mechanism can be generalized as: given a set of vector values and a vector query, attention is a technique to compute a weighted sum of the vector values, dependent on the vector query. The vector values are the encoder hidden states, and the vector query is the decoder hidden state at the current time step.

[0053] The weighted sum can be considered a selective summary of the information present in the vector values. The vector query determines on which of the vector values to focus. Thus, a fixed-size representation of the vector values can be created that depends upon the vector query.Attorney Docket No. PRTN1008WO01

[0054] The attention scores can be calculated by the dot product, or by weighing the different values (multiplicative attention). Embeddings

[0055] For most machine learning models, the input to the model needs to be numerical. The input to a translation model is a sentence, and words are not numerical. multiple methods exist for the conversion of words into numerical vectors. These numerical vectors are called the embeddings of the words. Embeddings can be used to convert any type of symbolic representation into a numerical one.

[0056] Embeddings can be created by using one-hot encoding. The one-hot vector representing the symbols has the same length as the total number of possible different symbols. Each position in the one-hot vector corresponds to a specific symbol. For example, when converting colors to a numerical vector, the length of the one-hot vector would be the total number of different colors present in the dataset. For each input, the location corresponding to the color of that value is one, whereas all the other locations are valued at zero. This works well for working with images. For NLP, this becomes problematic because the number of words in a language is very large. This results in enormous models and the need for a lot of computational power. Furthermore, no specific information is captured with one-hot encoding. From the numerical representation, it is not clear that orange and red are more similar than orange and green. For this reason, other methods exist.

[0057] A second way of creating embeddings is by creating feature vectors. Every symbol has its specific vector representation, based on features. With colors, a vector of three elements could be used, where the elements represent the amount of yellow, red, and / or blue needed to create the color. Thus, all colors can be represented by only using a vector of three elements. Also, similar colors have similar representation vectors.

[0058] For NLP, embeddings based on context, as opposed to words, are small and can be trained. The reasoning behind this concept is that words with similar meanings occur in similar contexts. Different methods take the context of words into account. Some methods, like GloVe, base their context embedding on co-occurrence statistics from corpora (large texts) such as Wikipedia. Words with similar co-occurrence statistics have similar word embeddings. Other methods use neural networks to train the embeddings. For example, they train their embeddings to predict the word based on the context (Common Bag of Words), and / or to predict the context based on the word (Skip-Gram). Training these contextual embeddings is time intensive. For this reason, pre-trained libraries exist. Other deep learning methods can be used to create embeddings. For example, the latent space of a variational autoencoder (VAE) can be used as the embedding of the input. Another method is to use 1D convolutions to create embeddings. This causes a sparse, high-dimensional input space to be converted to a denser, low-dimensional feature space.

[0059] Transformer models are based on the principle of self-attention. Self-attention allows each element of the input sequence to look at all other elements in the input sequence and search for clues that can help it to create a more meaningful encoding. It is a way to look at which other sequence elements areAttorney Docket No. PRTN1008WO01 relevant for the current element. The Transformer can grab context from both before and after the currently processed element.

[0060] When performing self-attention, three vectors need to be created for each element of the encoder input: the query vector (Q), the key vector (K), and the value vector (V). These vectors are created by performing matrix multiplications between the input embedding vectors using three unique weight matrices.

[0061] After this, self-attention scores are calculated. When calculating self-attention scores for a given element, the dot products between the query vector of this element and the key vectors of all other input elements are calculated. To make the model mathematically more stable, these self-attention scores are divided by the root of the size of the vectors. This has the effect of reducing the importance of the scalar thus emphasizing the importance of the direction of the vector. Just as before, these scores are normalized with a SoftMax layer. This attention distribution is then used to calculate a weighted sum of the value vectors, resulting in a vector z for every input element. In the attention principle explained above, the vector to calculate attention scores and to perform the weighted sum was the same, in self-attention two different vectors are created and used. As self-attention needs to be calculated for all elements (thus a query for every element), one formula can be created to calculate a Z matrix. The rows of this Z matrix are the z vectors for every sequence input element, giving the matrix a size length sequence dimension QKV.

[0062] Multi-headed attention is executed in the Transformer. For example, consider the calculation of self-attention using one attention head. For every attention head, different weight matrices are trained to calculate Q, K, and V. Every attention head outputs a matrix, Z. Different attention heads can capture different types of information. The different Z matrices of the different attention heads are concatenated. This matrix can become large when multiple attention heads are used. To reduce dimensionality, an extra weight matrix W is trained to condense the different attention heads into a matrix with the same size as one Z matrix. This way, the amount of data given to the next step does not increase every time self-attention is performed.

[0063] When performing self-attention, information about the order of the different elements within the sequence is lost. To address this problem, positional encodings are added to the embedding vectors. Every position has its unique positional encoding vector. These vectors follow a specific pattern, which the Transformer model can learn to recognize. This way, the model can consider distances between the different elements.

[0064] As discussed above, in the core of self-attention are three objects: queries (Q), keys (K), and values (V). Each of these objects has an inner semantic meaning of their purpose. One can think of these as analogous to databases. We have a user-defined query of what the user wants to know. Then we have the relations in the database, i.e., the values which are the weights. More advanced database management systems create some apt representation of its relations to retrieve values more efficiently from the relations. This can be achieved by using indexes, which represent information about what is stored in the database. In the context of attention, indexes can be thought of as keys. So instead of running the query against valuesAttorney Docket No. PRTN1008WO01 directly, the query is first executed on the indexes to retrieve where the relevant values or weights are stored. Lastly, these weights are run against the original values to retrieve data that is most relevant to the initial query.

[0065] In another example, several attention heads can be in a Transformer block. In this example, the outputs of queries and keys dot products in different attention heads enable the multi-head attention to focus on different aspects of the input and to aggregate the obtained information by multiplying the input with different attention weights.

[0066] Examples of attention calculation include scaled dot-product attention and additive attention. There are several reasons why scaled dot-product attention is used in the Transformers. Firstly, the scaled dot-product attention is relatively fast to compute, since its main parts are matrix operations that can be run on modern hardware accelerators. Secondly, it performs similarly well for smaller dimensions of the K matrix, dk, as the additive attention. For larger dk, the scaled dot-product attention performs a bit worse because dot products can cause the vanishing gradient problem. This is compensated via the scaling factor.

[0067] As discussed above, the attention function takes as input three objects: key, value, and query. In the context of Transformers, these objects are matrices of shapes (n, d), where n is the number of elements in the input sequence and d is the hidden representation of each element (also called the hidden vector). Attention is then computed as:

[0068] Attention (Q, K, V) = SoftMax

[0069] where Q, K, V are computed as:

[0070] X is the input matrix and WQ, WK, WV are learned weights to project the input matrix into the representations. The dot products appearing in the attention function are exploited for their geometrical interpretation where higher values of their results mean that the inputs are more similar, i.e., pointing in the geometrical space in the same direction. Since the attention function now works with matrices, the dot product becomes matrix multiplication. The SoftMax function is used to normalize the attention weights into the value of 1 prior to being multiplied by the values matrix. The resulting matrix is used either as input into another layer of attention or becomes the output of the Transformer.

[0071] Transformers become even more powerful when multi-head attention is used. Queries, keys, and values are computed the same way as above, though they are now projected into h different representations of smaller dimensions using a set of h learned weights. Each representation is passed into a different scaled dot-product attention block called a head. The head then computes its output using the same procedure as described above.

[0072] Formally, the multi-head attention is defined as: MultiHeadAttention (Q, K, V) = [head1, ..., headh]W0 where headi = Attention ( ).

[0073] The outputs of all heads are concatenated together and projected again using the learned weights matrix W0 to match the dimensions expected by the next block of heads or the output of theAttorney Docket No. PRTN1008WO01 Transformer. Using multi-head attention, instead of the simpler scaled dot-product attention, enables Transformers to jointly attend to information from different representation subspaces at different positions.

[0074] One can use multiple workers to compute the multi-head attention in parallel, as the respective heads compute their outputs independently of one another. Parallel processing is one of the advantages of Transformers over RNNs. Assuming the naive matrix multiplication algorithm has a complexity of:

[0075] For matrices of shape (a, b) and (c, d), to obtain values Q, K, V, the following operations should be computed: X WQ, X WK, X WV

[0076] The matrix X is of shape (n, d) where n is the number of patches and d is the hidden vector dimension. The weights WQ, WK, WV all have a shape of (d, d). Omitting the constant factor 3, the resulting complexity is: n d2.

[0077] Next, the complexity of the attention function itself should be computed, i.e., of SoftMax . The matrices Q and K are both of shape (n, d). The transposition operation does not influence the asymptotic complexity of computing the dot product of matrices of shapes (n, d) (d, n), therefore its complexity is: n2 d.

[0078] Scaling by a constant factor, where dk is the dimension of the keys vector, as well as applyingthe SoftMax function, both have the complexity of a b for a matrix of shape (a, b), hence they do notinfluence the asymptotic complexity. Lastly the dot product SoftMax is between matrices of shapes (n, n) and (n, d) and so its complexity is: n2 d.

[0079] The final asymptotic complexity of scaled dot-product attention is obtained by summing the complexities of computing Q, K, V, and of the following attention function:

[0080] n d2 + n2 d.

[0081] The asymptotic complexity of multi-head attention is the same since the original input matrix X is projected into h matrices of shapes (n, ), where h is the number of heads. From the point of view of asymptotic complexity, h is constant, therefore we would arrive at the same estimate of asymptotic complexity using a similar approach as for the scaled dot-product attention.

[0082] Transformer models often have the encoder-decoder architecture, although this is not necessarily the case. The encoder is built out of different encoder layers which are all constructed in the same way. The positional encodings are added to the embedding vectors. Afterward, self-attention is performed. Encoder Block of Transformer

[0083] In one example of a single encoder layer of a Transformer network, every self-attention layer can be surrounded by a residual connection, summing up the output and input of the self-attention. This sum is normalized, and the normalized vectors are fed to a feed-forward layer. Every z vector is fed separately to this feed-forward layer. The feed-forward layer is wrapped in a residual connection and the outcome is normalized too. Often, numerous encoder layers are piled to form the encoder. The output of the encoder is a fixed-size vector for every element of the input sequence.Attorney Docket No. PRTN1008WO01

[0084] Just like the encoder, the decoder is built from different decoder layers. In the decoder, a modified version of self-attention takes place. The query vector is only compared to the keys of previous output sequence elements. The elements further in the sequence are not known yet, as they still must be predicted. No information about these output elements may be used. Encoder-Decoder Blocks of Transformer

[0085] In another example, a Transformer model having encoder-decoder layers is described. Next to a self-attention layer, a layer of encoder-decoder attention is present in the decoder, in which the decoder can examine the last Z vectors of the encoder, providing fluent information transmission. The ultimate decoder layer is a feed-forward layer. All layers are packed in a residual connection. This allows the decoder to examine all previously predicted outputs and all encoded input vectors to predict the next output. Thus, information from the encoder is provided to the decoder, which could improve the predictive capacity. The output vectors of the last decoder layer need to be processed to form the output of the entire system. This is done by a combination of a feed-forward layer and a SoftMax function. The output corresponding to the highest probability is the predicted output value for a subject time step.

[0086] For some tasks other than translation, only an encoder is needed. This is true for both document classification and name entity recognition. In these cases, the encoded input vectors are the input of the feed-forward layer and the SoftMax layer. Transformer models have been extensively applied in different NLP fields, such as translation, document summarization, speech recognition, and named entity recognition. These models have applications in the field of biology as well for predicting protein structure and function and labeling DNA sequences. Vision Transformer

[0087] There are extensive applications of Transformers in vision including popular recognition tasks (e.g., image classification, object detection, action recognition, and segmentation), generative modeling, multi-modal tasks (e.g., visual-question answering, visual reasoning, and visual grounding), video processing (e.g., activity recognition, video forecasting), low-level vision (e.g., image super-resolution, image enhancement, and colorization) and 3D analysis (e.g., point cloud classification and segmentation).

[0088] Transformers were originally developed for NLP and worked with sequences of words. In image classification, we often have a single input image in which the pixels are in a sequence. To reduce the computation required, Vision Transformers (ViTs) cut the input image into a set of fixed-sized patches of pixels. The patches are often 16 x 16 pixels. They are treated much like words in NLP Transformers. Unfortunately, important positional information is lost because image sets are position-invariant. This problem is solved by adding a learned positional encoding into the image patches.

[0089] The computations of the ViT architecture can be summarized as follows. The first layer of a ViT extracts a fixed number of patches from an input image. The patches are then projected to linear embeddings. A special class token vector is added to the sequence of embedding vectors to include all representative information of all tokens through the multi-layer encoding procedure. The class vector is unique to each image. Vectors containing positional information are combined with the embeddings and theAttorney Docket No. PRTN1008WO01 class token. The sequence of embedding vectors is passed into the Transformer blocks. The class token vector is extracted from the output of the last Transformer block and is passed into a multilayer perceptron (MLP) head whose output is the final classification. The perceptron takes the normalized input and places the output in categories. It classifies the images.

[0090] When the input image is split into patches, a fixed patch size is specified before instantiating a ViT. Given the quadratic complexity of attention, patch size has a large effect on the length of training and inference time. A single Transformer block comprises several layers. The first layer implements Layer Normalization, followed by the multi-head attention that is responsible for the performance of ViTs. In a Transformer block, including skip connection data can simplify the output and improve the results. The output of the multi-head attention is followed again by Layer Normalization. And finally, the output layer is an MLP (Multi-Layer Perceptron) with the GELU (Gaussian Error Linear Unit) activation function.

[0091] ViTs can be pre-trained and fine-tuned. Pretraining is generally done on a large dataset. Fine- tuning is done on a domain specific dataset.

[0092] Domain-specific architectures, like convolutional neural networks (CNNs) or long short-term memory networks (LSTMs), have been derived from the usual architecture of MLPs and suffer from so- called inductive biases that predispose the networks towards a certain output. ViTs stepped in the opposite direction of CNNs and LSTMs and became more general architectures by eliminating inductive biases. A ViT can be seen as a generalization of MLPs because MLPs, after being trained, do not change their weights for different inputs. On the other hand, ViTs compute their attention weights at runtime based on the particular input.

[0093] The following detailed description is made with reference to the figures. Example implementations are described to illustrate the technology disclosed, not to limit its scope, which is defined by the claims. Those of ordinary skill in the art will recognize a variety of equivalent variations on the description that follows.

[0094] FIGS.1-10 depict at least one example of a high-yield codon sequence generator. Also shown in the figures is at least one example of a confined search, as well as relative positional embedding to a multi-head transformer. Depicted implementations also demonstrate at least one example of training a high- yield codon sequence generator. IMPLEMENTATIONS

[0095] In this disclosure, the example embodiments may use various machine learning models for the codon sequence optimization and the high-yield codon generation problems described above. As will be described in more detail, the machine learning models may require sample data (also referred to as training data) to make predictions or decisions. In the description that follows, various implementations of the disclosed technology are described with reference to the following figures.

[0096] Transformer models process input tokes and generate query, key, and value vectors. Transformer models use self-attention to turn these vectors into attention scores. Positional encodings like rotary relative position embedding (RoPE) and Attention with Linear Biases (ALiBi) are used to negate anyAttorney Docket No. PRTN1008WO01 absolute positional information of the input tokens and to only retain information about the relative angles between every pair of word embeddings (e.g., amino acids) in a sequence (e.g., protein). Positional encodings exist that can vary the relative angles between amino acids. However, this variation is currently restricted to the confines of a given attention head. An opportunity arises to further vary the relative angles between the amino acids across different attention heads of a Transformer architecture.

[0097] The technology disclosed extends positional encodings like RoPE and ALiBI by applying a series of rotation matrices to the query and key vectors at different scaled frequencies that vary both by absolute position of the query and key vectors and by the different attention heads. For example, positional encodings like RoPE and ALiBI scale the query and key vectors by using absolute positions of the query and key vectors as scaling parameters. These absolute positions are expressed by position indices. The technology disclosed adds an additional degree of variability to the positional encodings by further multiplying the position indices with an additional scaling factor, referred to herein as “head-specific scaling parameter.” This additional scaling factor is specific to a given attention head and varies across the attention heads of the Transformer architecture. Furthermore, from one attention head to the next, this additional scaling factor can vary in a pattern, for example, change in multiples.

[0098] FIG.1 shows one example of relative positional embedding provided to a multi-head attention Transformer system. Providing relative positional embedding to a multi-head Transformer model may include providing input tokens 110, to respective attention heads, such as attention head_2 (150), of the multi-head attention Transformer model. The input tokens 110 may be used to generate respective query (Q), key (K), and value (V) vectors to be provided to respective attention. For example, input tokens 110 may be used to generate Q2 vector 120, K2 vector 122, and V2 vector 124.

[0099] Query and key vectors may be converted into position-encoded query and key vectors by applying a series of rotation matrices. As depicted in FIG.1, Q2 vector 120 and K2 vector 122 may be converted into position-encoded query_2 vector 140 and position-encoded key_2 vector 142, respectively, via application of rotation matrix Q2130 and rotation matrix K2132. Rotation matrices may be applied to the query and key vectors at different scaled frequencies that vary by absolute positions of the query and key vectors and by the respective attention heads.

[0100] For example, application of the series of rotation matrices may include rotating pairs of feature dimensions in the query and key vectors by an angle in multiples of a scaled position index of a corresponding query or key vector. The scaled position index may be scaled by a head-specific scaling parameter that varies across the respective attention heads. In some implementations, the head-specific scaling parameter may be a head-specific scaling scalar.

[0101] The position-encoded query and key vectors may be used for execution of self-attention by the respective attention heads to generate pairwise attention scores that depend on relative positions of input token pairs and on their feature similarity. As depicted in FIG.1, position-encoded query_2 vector 140 and position-encoded key vector 142 may be used for execution of self-attention_2160 by respective attention head_2150 to generate pairwise attention scores.Attorney Docket No. PRTN1008WO01

[0102] Pairwise attention scores that are generated may depend on relative positions of the input token pairs and on their feature similarity. For example, the pairwise attention scores may be penalized based on how far the position-encoded query and key vectors are located from one another. In some implementations, when a position-encoded query vector and a position-encoded key vector are close by, then the attention score penalty may be very low. In some implementations, when a position-encoded query vector and a position-encoded key vector are far away, then the attention score penalty may be very high.

[0103] In some implementations, the input token pairs may be amino acid token pairs. In other implementations, the input token pairs can be nucleotide token pairs.

[0104] One having skill in the art will recognize the relative positional embedding to a multi-head attention Transformer model can overcome the limitations arising from the longer (inference) sequence length, the nature of the sequence, resulting folding, and the corresponding scaled number of parameters. Specifically, different scaling may be performed on each attention head that is proportional to its context length.

[0105] In another example, given a vector, Q, of length R, where Q is a vector of positions (e.g. [1, 2, 3, 4, 5, 6, … R]) where R is the sequence length. By multiplying vector Q with a certain scalar (referred to herein as a “scaling parameter” or “head-specific scaling parameter”) such that the scalar changes on each head, then the amount of rotation changes for each word and the amount of rotation changes for each of the other words in each of the other attention heads.

[0106] To illustrate the above, vector Q, with length R, may be Q = [1, 2, 3, 4, 5, 6, …, R]. If given S attention heads, then it may follow that:

[0107] At the 1st head: Q will be scaled by 0.5 (e.g. Q * 0.5)

[0108] At the 2nd head: Q will be scaled by 0.25 (e.g. Q * 0.25)

[0109] At the 3rd head: Q will be scaled by 0.125 (e.g. Q * 0.125)

[0110] And so on, until the Sth attention head.

[0111] One having skill in the art will recognize that relative positional embedding to a multi-head Transformer model may be especially beneficial for predicting longer amino acid sequence embeddings that surpass the relative positional embedding dimension (i.e., efficiently extrapolating sequences longer than sequences in the pre-training data set). The relative positional embedding to a multi-head Transformer may extend the context length to sustain relationships between distant tokens of an input sequence, which is of high importance in protein sequence modeling.

[0112] FIG.2 is a flowchart of one example of a method of providing relative positional embedding to a multi-head attention Transformer 200. The exemplary method of FIG.19 depicts using attention heads of a multi-head attention Transformer model to generate query (Q), key (K), and value (V) vectors 210 for input tokens. Another method step may include converting query (Q) vectors and key (K) vectors into position-encoded query (Q) and key (K) vectors 220. Forming position-encoded (Q) and (K) vectors may be accomplished by applying a series of rotation matrices to query (Q) and key (K) vectors at different scaled frequencies 230. For example, the different scaled frequencies may vary by absolute positions of theAttorney Docket No. PRTN1008WO01 query (Q) and key (K) vectors and by the respective attention heads. In some embodiments, application of the series of rotation matrices may include rotating pairs of feature dimensions in the query (Q) and key (K) vectors by an angle in multiples of a scaled position index of a corresponding query or key vector.

[0113] The method of FIG.2 also depicts using position-encoded query (Q) and key (K) vectors for execution of self-attention by the respective attention heads to generate pairwise attention scores that depend on relative positions of input token pairs and on their feature similarity 240. One having skill in the art will recognize that relative positional embedding to a multi-head Transformer model may be especially beneficial for predicting longer amino acid sequence embeddings that surpass the relative positional embedding dimension.

[0114] FIG.3 shows one example of a confined search 300 during inference for an output sequence of codon elements 395 from a given amino acid element 320 within a sequence of amino acid elements310. A confined search for a given amino acid element within a sequence of amino acid elements 300 mayinclude processing an input sequence of amino acid elements 310 and obtaining a given amino acid element320. For given amino acid element 320, a corresponding codon element in a vocabulary of codon elements372 may be identified. A confined search 390 of the corresponding codon element in a vocabulary of codon elements 372 to a subset of codon elements 374 known to translate to the given amino acid element, may be performed. After processing, an output sequence of codon elements 395 may be generated. A person of skill in the art will recognize that generating a prediction by looking back to confine the search to a subset of corresponding codons (e.g., a subset may include 1- 6 codons per amino acid, according to the genetic code), instead of all 64 possible codons can provide an enormous speedup at runtime.

[0115] In one implementation, the confined search 390 works after the output layer 380. The confined search 390 gets the probabilities for all codons from the output layer 380. The probabilities are then passed to the confined search block 390 that suppresses (i.e., sets to -infinity) any codon that is not relevant to the current amino acid that is being currently processed by the neural network 330.

[0116] The confined search for a given amino acid element within a sequence of amino acid elements 300 may be combined with other systems or methods of the present invention. For example, the relative positional embedding to a multi-head attention transformer system 100 may be combined with a protein-to- codon sequence mapping system, that may be configured to include a confined search for a given amino acid element within a sequence of amino acid elements 300.

[0117] A given amino acid element 320 may be provided as input to a neural network 330. Neural network 330 can process the input sequence of amino acid elements 310 and generate the output sequence of codon elements 395.

[0118] As depicted in FIG.3, neural network 330 may be configured with an input layer 340 to receive the input sequence of amino acid elements 310. Input layer 340 may convert the input sequence of amino acid elements 310 from symbols into a numerical vector representation (tokenize, one-hot encode, etc.). Input sequence of amino acid elements 310 may be further processed by layers of neural network 330Attorney Docket No. PRTN1008WO01 to generate an output sequence of codon elements 395 that corresponds to a protein with optimized expression.

[0119] In some implementations, the given codon / nucleotide sequence sampling often takes place in a manner that's unrestricted by the protein sequence but rather guided with it. Therefore, it is probable that the codons / nucleotides sampled by the decoder may not encode back to the original amino acids.

[0120] Neural network 330 may be a pre-trained neural network, or alternatively, neural network 330 may be trained from scratch.

[0121] Neural network 330 may perform a search in order to predict a codon sequence element for the corresponding given amino acid element 320. In some implementations, prior to the search, a neural network component may look back (e.g., via lookup table 370) at the input sequence of amino acid elements 310 (i.e., the inference protein sequence) to determine an identity of the given amino acid element320. Then, the neural network component may use the determined identity of the given amino acid element320 (i.e., the given amino acid element 320 may be a single amino acid within the inference protein sequence) to confine the search 390 of the corresponding codon element to the subset of codon elements 374 that are valid codons known to translate to the given amino acid element 320.

[0122] For example, as depicted in FIG.3, a given amino acid element (Leucine) 320, may be received by neural network 330 during inference to predict a corresponding codon sequence element. During processing and before performing a search, the decoder 360, for example, may access the vocabulary of codon elements 372 (i.e., mapping / vocabulary of codon to amino acid elements) of lookup table 370 to determine the identity of the given amino acid (Leucine) 320 at that position of the input sequence of amino acid elements 310. By determining the identity of the given amino acid (Leucine) 320 and then determining a corresponding subset of codon elements (UUA, UUG, CUU, CUC, CUA, CUG) 374, the search may be confined to the subset of 6 codons, instead of the set of all 64 codons (i.e., decoder 360 may be constrained to sampling a subset of 6 corresponding codons and SoftMax may generate a confined subset of 6 probabilities, instead of decoder 360 sampling across the entire target codon vocabulary of 64 target codons and SoftMax generating 64 probabilities). One having skill in the art will appreciate that reducing the computation required for each position within the input sequence of amino acid elements 310 will result in a very large savings in computational resources and time.

[0123] Neural network 330 may use any suitable architecture and / or any suitable network layers to accomplish the intended goals. For example, neural network 330 may be a sequence to sequence (Seq2Seq) neural network. In some implementations, neural network 330 is an encoder 350 – decoder 360 neural network (e.g., Transformer).

[0124] As depicted in FIG.3, input sequence of amino acid elements 310 may be received by an encoder 350 neural network. In some implementations, the encoder 350 neural network may generate respective embedded tokens for respective amino acid elements in the input sequence of amino acid elements 310. In some implementations, the encoder 350 neural network applies attention between theAttorney Docket No. PRTN1008WO01 respective embedded tokens on an amino acid element-by-amino acid element. The context representation of the input sequence of amino acid elements 310 may be provided to the decoder 360 neural network.

[0125] The decoder 360 neural network may receive the context representation of the input sequence of amino acid elements 310. In some implementations of the confined search for a given amino acid 300, the decoder 360 neural network may receive results of the attention from the encoder 350 neural network and may use the results of the attention to generate the output sequence of codon elements 395. In other implementations of the confined search for a given amino acid 300, the decoder 360 neural network may look back (via lookup table 370 or another approach) at the input sequence of amino acid elements 310 to determine the identity of the given amino acid element 320, and uses the determined identity of the given amino acid element to confine the search of the corresponding codon element to the subset of codon elements 374 known to translate to the given amino acid element 320.

[0126] FIG.4 is a flowchart of one example of generating an output sequence of codon elements by looking back to confine the search 400. As depicted, the generating an output sequence of codon elements by looking back to confine the search method 400 includes processing an input sequence of amino acid elements 410, obtaining a given amino acid element 420, looking back at the input sequence of amino acid elements 430, determining the identity of the given amino acid element to confine the search 440, confining the search to a subset of codon elements known to translate to the given amino acid element 450, and generating an output sequence of codon elements 460. One having skill in the art will appreciate that generating an output sequence of codon elements by looking back to confine the search method 400 can provide an enormous speedup at runtime by placing a constraint on the number of codons for which a prediction will be made.

[0127] Processing an input sequence of amino acid elements 410 may include the input layer of a neural network to receive the input sequence of amino acid elements. The symbolic text may be converted to a numerical representation of the input sequence of amino acid elements (e.g., via tokenizer, one-hot encoding, etc.). The neural network may generate a context vector or an embedding for predicting a codon sequence element. The neural network may comprise any architecture suitable for the task of generating optimized codon sequences.

[0128] Obtaining a given amino acid element 420 may include a neural network receiving the given amino acid element 420 for processing.

[0129] Looking back at the input sequence of amino acid elements 430 may include a neural network component that determines the identity of the given amino acid element from the input sequence of amino acid elements. The term “looking back” may refer to, for example, accessing information about the input sequence of amino acid elements via a lookup table.

[0130] Determining the identity of the given amino acid element to confine the search 440 may include a neural network component that searches and finds the correct position along the input sequence of amino acid elements.Attorney Docket No. PRTN1008WO01

[0131] Confining the search to a subset of codon elements known to translate to the given amino acid element 450 may include searching the genetic code within a lookup table for a subset of codon elements that corresponds to the given amino acid element at the specified position of the input sequence of amino acid elements. Retrieving the subset of codon elements (for example, a range of 1 – 6 codons may correspond to a given amino acid element) from the lookup table for additional processing within the neural network. Based on the subset of corresponding codon elements, calculating a prediction probability for the subset of corresponding codon elements (i.e., for 1 – 6 codons) rather than a prediction probability for the entire set of all possible 64 codons.

[0132] Generating an output sequence of codon elements 460 may include a neural network providing a predicted output sequence of codon elements that is an optimized codon sequence.

[0133] FIG.5 shows a cluster-by-cluster, high-yield codon (DNA) sequence generator trained via supervised learning 500. First, a lower-to-higher yield codon training dataset is created, which is then used to train the cluster-by-cluster high yield DNA generator 590 on a cluster-basis. Training a cluster-by-cluster high-yield DNA sequence generator 590 includes creating clusters of DNA sequences on an amino acid sequence-basis via protein-DNA sequence clustering 510. Each protein-DNA sequence cluster may be sent to a cluster-by-cluster yield sorter 550, wherein each protein-DNA sequence cluster remains in-tact. All DNA sequences within a single protein-DNA sequence cluster can be sorted by expression yield to generate lower-to-higher yield codon training dataset (for example, cluster B 570). Then, the cluster-by- cluster high-yield DNA sequence generator LLM 590 may be trained via supervised learning.

[0134] A large set of DNA sequences are obtained to build the lower-to-higher yield codon (DNA) training dataset (e.g., cluster B 570). For example, DNA sequences based on labeled data, where labels may be obtained from transcriptomic labeling, can identify different DNA (codon) sequences that translate to the same amino acid sequence. In addition, clinical data can provide transcriptomic information. In some implementations, experimentally characterized data supplied information that may be used to generate at least some of the clusters of codon sequences. In other implementations, at least some of the clusters of codon sequences are generated based on “labeled data” that can use transcriptomic labeling and proteomic labeling. In another example, a protein-to-codon translator and / or a protein-to-codon sequence mapping system of AA2DNA may provide source DNA (codon) sequences.

[0135] Protein-DNA sequence clustering 510 may cluster DNA (codon) sequences based on their single corresponding protein sequence. Individual clusters of protein-DNA sequences can be formed, for example as depicted in FIG.5, cluster A 520 includes one protein sequence PS_A and multiple DNA sequences, CS_A1 through CS_A56. Each of the DNA sequences, CS_A1 through CS_A56, have a different DNA sequence that differs by one or more nucleotides (codons) from each of the other DNA sequences. Each of the DNA sequences, CS_A1 through CS_A56, corresponds to a single protein, PS_A. FIG.5 is meant for illustrative purposes only, and it is not meant to limit the number of DNA sequences that may be clustered, sorted, or contained within a single cluster.Attorney Docket No. PRTN1008WO01

[0136] Cluster-by-cluster yield sorter 550 receives each of the protein-DNA sequence clusters, such that each of the clusters remains in-tact and separated from other clusters. For example, as depicted in FIG. 5, cluster A 520 is provided to cluster-by-cluster yield sorter 550, and after being received by cluster-by- cluster yield sorter 550, cluster A 560 remains in-tact and separated from cluster B 570 and cluster C 580. All DNA sequences within a single cluster may be sorted by expression yield, low-yield to high-yield, and sorted relative to one another, to generate a lower-to-higher yield codon training dataset.

[0137] Training the cluster-by-cluster high-yield DNA sequence generator 590 may be accomplished on a cluster-by-cluster basis. For example, as depicted in FIG.5, cluster B 570 has a group of high-yield (yield / expression) DNA sequences 572 that are sorted in descending order by yield / expression and a group of low-yield DNA sequences 574 that are sorted in descending order by yield / expression. The low-yield (yield / expression) DNA sequences 574 may be input during training, while the high-yield DNA sequences 572 may be target (or ground truth). The cluster-by-cluster high-yield DNA sequence generator 590 receives input and target during supervised learning. The cluster-by-cluster high-yield DNA sequence generator 590 is only trained on input / target DNA sequences corresponding to a single cluster (i.e., cluster B 570) at a time.

[0138] After training, the (trained) cluster-by-cluster high-yield DNA sequence generator (for example, 620) may accept an inference input DNA sequence of any yield / expression and generate an inference output DNA sequence (i) that is of higher yield / expression than the input and (ii) that belongs to the same protein cluster. One having skill in the art will appreciate generating higher-yield / expression codon sequences that correspond to the same protein and therefore have retained and / or improved protein functionality.

[0139] The above description is intended to be an example and not limiting. For example, in some implementations, a lower-to-higher yield (yield / expression) codon training dataset may comprise protein embeddings and codon embeddings instead of sequences. Next, the cluster-by-cluster high-yield DNA sequence generator 590 may be trained with a lower-to-higher yield (yield / expression) codon embedding training dataset. One having skill in the art will recognize that replacing sequences with embeddings to create a lower-to-higher yield (yield / expression) codon embedding training dataset provides more robustness for a smaller dataset. In other implementations, an appropriately trained cluster-by-cluster high- yield DNA sequence generator (for example, 620) may receive an input protein sequence and may generate a high-yield (yield / expression) codon sequence. One having skill in the art will recognize the flexibility of cluster-by-cluster high-yield DNA sequence generator 620. Clustering Logic

[0140] The technology disclosed creates clusters of codon sequences on an amino acid sequence- basis. In some implementations, a particular cluster of codon sequences created for a particular amino acid sequence includes different codon sequences that translate to the particular amino acid sequence but have varying yields (i.e., yield / expression). In other implementations, the particular cluster of codon sequences isAttorney Docket No. PRTN1008WO01 created by clustering based on one or more codon sequence attributes. In one implementation, the codon sequence attributes correspond to biological constraints of the different codon sequences that are to be clustered or sub-clustered. In some implementations, the biological constraints include identity similarity of the different codon sequences that are to be clustered or sub-clustered, homology of the different codon sequences that are to be clustered or sub-clustered, structural similarity of the different codon sequences that are to be clustered or sub-clustered, size of the different codon sequences that are to be clustered or sub-clustered, length of the different codon sequences that are to be clustered or sub-clustered, distribution of the different codon sequences that are to be clustered or sub-clustered, and rarity of the different codon sequences that are to be clustered or sub-clustered.

[0141] In some implementations, the particular cluster of codon sequences is created by clustering those codon sequences in a same cluster that have an identity score for at least one codon sequence identity higher than a similarity threshold. In one implementation, the codon sequence identity includes homology overlap between the codon sequences.

[0142] In another implementation, the codon sequences are embedded in an embedding space. The codon sequence identity includes embedding distances between the codon sequences in the embedding space. An embedding space in which the codon sequences are embedded, for example, to group / cluster / subcluster similar codon sequences in a latent space. A “latent space,” for example, in deep learning is a reduced-dimensionality vector space of a hidden layer. A hidden layer of a neural network compresses an input and forms a new low-dimensional codon sequence with interesting properties that are distance-wise correlated in the latent space.

[0143] A distance is identified between each pair of the instances in the embedding space corresponding to a predetermined measure of similarity between the pair of the instances. The “embedding space,” into which the instances are embedded, for example, by an embedding module (not shown), can be a geometric space within which the instances are represented. In one implementation, the embedding space can be a vector space (or tensor space), and in another implementation the embedding space can be a metric space. In a vector space, the features of an instance define its “position” in the vector space relative to an origin. The position is typically represented as a vector from the origin to the instance’s position, and the space has a number of dimensions based on the number of coordinates in the vector. Vector spaces deal with vectors and the operations that may be performed on those vectors.

[0144] When the embedding space is a metric space, the embedding space does not have a concept of position, dimensions, or an origin. Distances among instances in a metric space are maintained relative to each other, rather than relative to any particular origin, as in a vector space. Metric spaces deal with codon sequences combined with a distance between those codon sequences and the operations that may be performed on those codon sequences.

[0145] For purposes of the present disclosure, these codon sequences are significant in that many efficient algorithms exist that operate on vector spaces and metric spaces. For example, metric trees may be used to rapidly identify codon sequences that are “close” to each other. Codon sequences can be embeddedAttorney Docket No. PRTN1008WO01 into vector spaces and / or metric spaces. In the context of a vector space, this means that a function can be defined that maps codon sequences to vectors in some vector space. In the context of a metric space, this means that it is possible to define a metric (or distance) between those codon sequences, which allows the set of all such codon sequences to be treated as a metric space. Vector spaces allow the use of a variety of standard measures of distance / divergence (e.g., the Euclidean distance). Other implementations can use other types of embedding spaces.

[0146] As used herein, “an embedding” is a map that maps instances into an embedding space. An embedding is a function that takes, as inputs, a potentially large number of characteristics of the instance to be embedded. For some embeddings, the mapping can be created and understood by a human, whereas for other embeddings the mapping can be very complex and non-intuitive. In many implementations, the latter type of mapping is developed by a machine learning algorithm based on training examples, rather than being programmed explicitly.

[0147] In order to embed an instance in a vector space, each instance must be associated with a vector. A distance between two instances in such a space is then determined using standard measures of distance using vectors.

[0148] A goal of embedding instances in a vector space is to place intuitively similar instances close to each other. One way of embedding text instances is to use a bag-of-words model. The bag of words model maintains a dictionary. Each word in the dictionary is given an integer index, for example, the word aardvark may be given the index 1, and the word zebra may be given the index 60,000. Each instance is processed by counting the number of occurrences of each dictionary word in that instance. A vector is created where the value at the ith index is the count for the ith dictionary word. Variants of this codon sequence normalize the counts in various ways. Such an embedding captures information about the content and therefore the meaning of the instances. Text instances with similar word distributions are close to each other in this embedded space.

[0149] Images may be processed to identify commonly occurring features using, e.g., scale invariant feature transforms (SIFT), which are then binned and used in a codon sequence similar to the bag-of-words embedding described above. Further, embeddings can be created using deep neural networks, or other deep learning techniques. For example, a neural network can learn an appropriate embedding by performing gradient descent against a measure of dimensionality reduction on a large set of training data. As another example, a kernel can be learned based on data and derive a distance based on that kernel. Likewise, distances may be learned directly.

[0150] These approaches generally use large neural networks to map instances, words, or images to high dimensional vectors (for example see: A brief introduction to kernel classifiers, Mark Johnson, Brown University 2009, http: / / cs.brown.edu / courses / cs195-5 / fall2009 / docs / lecture_10-27.pdf “Using Confidence Bounds for Exploitation-Exploration Trade-offs, incorporated herein by reference; and Kernel Method for General Pattern Analysis, Nello Cristianini, University of California, Davis, accessed October 2016, http: / / www.kernel-methods.net / tutorials / KMtalk.pdf). In another example, image patches can beAttorney Docket No. PRTN1008WO01 represented as deep embeddings. As an image is passed through a deep neural network model, the output after each hidden layer is an embedding in a latent space. These deep embeddings provide hints for the model to distinguish different images. In some implementations, the embeddings can be chosen from a low- dimensional layer as the latent codon sequence.

[0151] In other implementations, an embedding can be learned using examples with algorithms such as Multi-Dimensional Scaling, or Stochastic Neighbor Embedding. An embedding into a vector space may also be defined implicitly via a kernel. In this case, the explicit vectors may never be generated or used, rather the operations in the vector space are carried out by performing kernel operations in the original space.

[0152] Other types of embeddings of particular interest capture date and time information regarding the instance, e.g., the date and time when a photograph was taken. In such cases, a kernel may be used that positions images closer if they were taken on the same day of the week in different weeks, or in the same month but different years. For example, photographs taken around Christmas may be considered similar even though they were taken in different years and so have a large absolute difference in their timestamps. In general, such kernels may capture information beyond that available by simply looking at the difference between timestamps.

[0153] Similarly, embeddings capturing geographic information may be of interest. Such embeddings may consider geographic metadata associated with instances, e.g., the geo-tag associated with a photograph. In these cases, a kernel or embedding may be used that captures more information than simply the difference in miles between two locations. For example, it may capture whether the photographs were taken in the same city, the same building, or the same country.

[0154] Often embeddings will consider instances in multiple ways. For example, a product may be embedded in terms of the metadata associated with that product, the image of that product, and the textual content of reviews for that product. Such an embedding may be achieved by developing kernels for each aspect of the instance and combining those kernels in some way, e.g., via a linear combination.

[0155] In many cases a very high dimensional space would be required to capture the intuitive relationships between instances. In some of these cases, the required dimensionality may be reduced by choosing to embed the instances on a manifold (curved surface) in the space rather than to arbitrary locations.

[0156] Different embeddings may be appropriate on different subsets of the instance catalog. For example, it may be most effective to re-embed the candidate result sets at each iteration of the search procedure. In this way, the subset may be re-embedded to capture the most important axes of variation or of interest in that subset.

[0157] To embed an instance in a metric space requires associating that catalog with a distance (or metric).

[0158] A “distance” between two instances in an embedding space corresponds to a predetermined measurement (measure) of similarity among instances. Preferably, it is a monotonic function of theAttorney Docket No. PRTN1008WO01 measurement of similarity (or dissimilarity). Typically, the distance equals the measurement of similarity. Example distances include the Manhattan distance, the Euclidean distance, the Hamming distance, and the Mahalanobis distance.

[0159] Given the distance (similarity measure) between instances to be searched, or the embedding of those instances into a vector space, a metric space or a manifold, there are a variety of data structures that may be used to index the instance catalog and hence allow for rapid search. Such data structures include metric trees, kd-trees, R-trees, universal B-trees, X- trees, ball trees, locality sensitive hashes, and inverted indexes. The technology disclosed can use a combination of such data structures to identify a next set of candidate results based on a refined query. An advantage of using geometric constraints is that they may be used with such efficient data structures to identify the next results in time that is sub-linear in the size of the catalog.

[0160] There are a wide variety of ways to measure the distance (or similarity) between instances, and these may be combined to produce new measures of distance. An important concept is that the intuitive relationships between digital instances may be captured via such a similarity or distance measure. For example, some useful distance measures place images containing the same person in the same place close to each other. Likewise, some useful measures place instances discussing the same topic close to each other. Of course, there are many axes along which digital instances may be intuitively related, so that the set of all instances close (with respect to that distance) to a given instance may be quite diverse. For example, a historical text describing the relationship between Anthony and Cleopatra may be similar to other historical texts, texts about Egypt, texts about Rome, movies about Anthony and Cleopatra, and love stories. Each of these types of differences constitutes a different axis relative to the original historical text.

[0161] Such distances may be defined in a variety of ways. One typical way is via embeddings into a vector space. Other ways include encoding the similarity via a kernel. By associating a set of instances with a distance, we are effectively embedding those instances into a metric space. Instances that are intuitively similar will be close in this metric space while those that are intuitively dissimilar will be far apart. Note further that kernels and distance functions may be learned. In fact, it may be useful to learn new distance functions on subsets of the instances at each iteration of the search procedure.

[0162] Note that wherever a distance is used to measure the similarity between instances a kernel may be used to measure the similarity between instances instead, and vice-versa. However, kernels may be used directly instead without the need to transform them into distances.

[0163] Kernels and distances may be combined in a variety of ways. In this way, multiple kernels or distances may be leveraged. Each kernel may capture different information about an instance, e.g., one kernel captures visual information about a piece of jewelry, while another captures price, and another captures brand.

[0164] Also note that embeddings may be specific to a given domain, such as a given catalog of products or type of content. For example, it may be appropriate to learn or develop an embedding specificAttorney Docket No. PRTN1008WO01 to men’s shoes. Such an embedding would capture the similarity between men’s shoes but would be uninformative with regards to men’s shirts.

[0165] In other implementations, instead of a distance function, a similarity function can be used, for example, to group / cluster / subcluster visually similar images in a latent space. The similarity function, which is used to determine a measure of similarity, can be any function having kernel properties, such as but not limited to a dot product function, a linear function, a polynomial function, a Gaussian function, an exponential function, a Laplacian function, an analysis of variants (ANOVA) function, a hyperbolic tangent function, a rational quadratic function, a multi-quadratic function, an inverse multi-quadratic function, a circular function, a wave function, a power function, a log function, a spline function, a B-spline function, a Bessel function, a Cauchy function, a chi-square function, a histogram intersection function, a generalized histogram intersection function, a generalized T-student function, a Bayesian function, and a wavelet function.

[0166] In the above-described context, using similarity functions, as opposed to using distance functions, is better because neural networks are often trained with regularizers, which add an ever- increasing cost in order to reach the training objective as the weights of the neural network get larger. These regularizers are added to prevent overfitting, where the network pays undue attention to details in the training data, instead of identifying broad trends. Further, these regularizers may be viewed as applying pressure toward a default behavior, which must be overcome by the training data. When used for learning embeddings, standard regularizers have an effect of pushing the embeddings toward an origin, which tends to push them closer together. If one uses a goal to achieve large distances when items are dissimilar, then this sort of regularization pushes towards a default that items will be similar. However, if a goal is set to have the embeddings have a large dot product when the items are similar (as in the case of the above- described similarity function), then the regularizer applies pressure towards a default that items are dissimilar. It will often be the case that a typical random pair of instances should be regarded as dissimilar. An overall more accurate and efficient visual image discovery results.

[0167] FIG.6 shows generating the output higher-yield DNA sequence 640 via the trained cluster- by-cluster high-yield DNA sequence generator 620. During inference, a trained cluster-by-cluster high- yield DNA sequence generator 620 may receive an input DNA sequence 610 and, after processing, generate an output higher-yield (i.e., yield / expression) DNA sequence 640 with the same corresponding protein sequence.

[0168] Trained cluster-by-cluster high-yield DNA sequence generator 620 was trained using the Protein-DNA sequence cluster 630 (labeled “Cluster C”), such that the model learned high-yield (i.e., yield / expression) DNA sequence features and low-yield (i.e., yield / expression) DNA sequence features that all correspond to the same protein (“PS_C”). At inference, trained cluster-by-cluster high-yield DNA sequence generator 620 receives input DNA sequence 610 (labeled as “CS_CCC”) that is known to code for the same protein (“PS_C”). At inference, one or more output higher-yield codon sequences 640 may be generated, that all correspond to the same protein sequence (PS_C) as the input sequence and all haveAttorney Docket No. PRTN1008WO01 higher expression yield than input DNA sequence 610. One having skill in the art will appreciate the ease of accessing one or more valid DNA sequences having a higher yield / expression and known to code for the same protein of interest.

[0169] FIG.7 shows the supervised machine high-yield codon generation technique 700 to train the high-yield codon sequence generator 740. First, to generate high-yield codon sequences, a series of protein and codon sequences may be received by clustering logic 710 that creates clusters of codon sequences on an amino acid sequence-basis. The clusters of codon sequences (for example, Cluster C 716) may be used to build a lower-to-higher yield codon training dataset (for example, Cluster C 726). The lower-to-higher yield codon training dataset (for example, Cluster C 726) can be used to train a high-yield codon sequence generator 740 to map the lower-yield input codon sequences (e.g., 726d and 726e) to the higher-yield target codon sequences (e.g., 726b and 726c).

[0170] Input protein sequences and codon sequences may be retrieved from protein database 702 and codon database 704, respectively, and then provided to a clustering logic 710. Input protein and codon sequences may be derived from several sources, such as clinical data, databases, or others. In other implementations, at least some of the clusters of codon sequences are generated based on “labeled data” that can use transcriptomic labeling and proteomic labeling. In addition, input protein and codon sequences may be processed via a pre-trained protein language model and pre-trained DNA language model, respectively, to generate high dimensional protein embeddings and codon embeddings as input.

[0171] Clustering logic 710 processes the protein and codon sequences to generate protein-codon sequence clusters (e.g., Cluster A 712, Custer B 714, Cluster C 716). Each cluster includes only one protein sequence 716a and more than one codon sequence 716b that differs by at least one codon (for example, Cluster A 712 includes PS_A and multiple codon sequences CS_A1 to CS_A56, with no identical codon sequences in the set of sequences CS_A1 to CS_A56). Each codon sequence 716b has an associated expression yield 716c. In some implementations, a particular cluster of codon sequences 716b created for a particular amino acid sequence 716a includes different codon sequences that translate to the particular amino acid sequence but have varying yields / expression 716c. Moreover, at least some of the clusters of codon sequences are generated based on clinical data that identify different codon sequences that translate to the same amino acid sequence.

[0172] Some of the clusters of codon sequences are generated based on an amino acid sequence-to- codon sequence generator generating different output codon sequences for a same input amino acid sequence. In some implementations, the amino acid sequence-to-codon sequence generator is a neural network. In other implementations, the amino acid sequence-to-codon sequence generator that is a neural network, may be a pre-trained neural network. In some implementations, the amino acid sequence-to-codon sequence generator that is a neural network, may be an untrained neural network. In other implementations, the amino acid sequence-to-codon sequence generator that is a neural network, may be a sequence to sequence (seq2seq) neural network. Still in other implementations, the amino acid sequence-to-codonAttorney Docket No. PRTN1008WO01 sequence generator that is a neural network, may be an encoder-decoder neural network (e.g., Transformer).

[0173] The protein-codon sequence clusters may be received by training data generation logic 720. Protein-codon sequence clusters remain in-tact throughout downstream logic and when moving between downstream logic units. In addition, clusters are processed by logic on a cluster-by-cluster basis. For example, as shown in FIG.7, after Cluster A 712 is formed by clustering logic 710, then the cluster remains in-tact when received and processed by training data generation logic 720 into Cluster A 722.

[0174] Training data generation logic 720 sorts the codon sequences within each of the clusters of codon sequences by yield / expression. For example, in Cluster C 726, all codon sequences are sorted based on yield / expression into a higher-yield (i.e., yield / expression) codon 726b grouping and a lower-yield (i.e., yield / expression) codon 726d grouping.

[0175] A single protein-codon sequence cluster 732 may be received by training logic (on a cluster- basis) 730, in order to train high-yield codon sequence generator 740 on a cluster-basis. By processing the lower-to-higher yield codon training dataset 732, training logic 730 can link lower-yield input 736 codon sequences to higher-yield target 734 codon sequences. Specifically, on a cluster-basis, codon sequences in which one codon sequence has a lower-yield / expression and another codon sequence has a higher- yield / expression can be generated into pairs of codon sequences. With reference to FIG.7 for example, a lower yield / expression codon sequence, such as CS_A52 (with yield / expression of Low 27), can be paired with a higher-yield / expression codon sequence, such as CS_A2 (with yield / expression of High 3).

[0176] After pairing the lower-yield (expression) and higher-yield (expression) codon sequences, then input 736 and target 734 (ground truth) groups may be created for training the high-yield codon sequence generator 740 via supervised learning. From the paired codon sequences, the codon sequence with the lower yield / expression in the lower-to-higher yield codon training dataset 732 may be included as a lower-yield input 736 codon sequence. Similarly, from the paired codon sequences, the codon sequence with the higher-yield / expression in the lower-to-higher yield codon training dataset 732 may be included as a higher-yield input 734 codon sequence.

[0177] High-yield codon sequence generator 740 is a neural network or any suitable architecture. In some implementations, high-yield codon sequence generator 740 is a sequence to sequence (seq2seq) neural network. In other implementations, high-yield codon sequence generator 740 is an encoder-decoder neural network (e.g., Transformer). As depicted in FIG.7, high-yield codon sequence generator 740 has an encoder 750 network and a decoder 760 network.

[0178] In some implementations, high-yield codon sequence generator 740 is an untrained neural network. In other implementations, high-yield codon sequence generator 740 is a pre-trained neural network.

[0179] High-yield codon sequence generator 740 may be trained via a supervised learning technique with the lower-yield input 736 codon sequences and the higher yield target sequences. Training high-yieldAttorney Docket No. PRTN1008WO01 codon sequence generator 740 includes training with a single lower-to-higher yield codon training dataset 732, corresponding to one cluster (Cluster A), at a time.

[0180] During inference, the trained high-yield codon sequence generator may generate a higher- yield (i.e., yield / expression) codon sequence, for example, that corresponds to the protein of Cluster A 732.

[0181] FIG.8 shows one example of a high-yield codon sequence generator system 800. During inference, for example, input space 810 receives inference lower-yield (i.e., yield / expression) codon sequence 802 for processing by high-yield codon sequence generator 820. After processing, high-yield codon sequence generator 820 may generate an inference higher-yield (i.e., yield / expression) codon sequence 804 in output space 850. High-yield codon sequence generator 820 may have any suitable architecture, including an encoder 830 network and a decoder 840 network.

[0182] For example, the high-yield codon sequence generator 820 may be trained via supervised machine high-yield codon generation technique 700. During inference, a trained high-yield codon sequence generator 820 may receive an inference lower-yield (i.e., yield / expression) codon sequence 802 and, after processing, generate inference higher-yield (i.e., yield / expression) codon sequence 804 corresponding to the same protein sequence as inference lower-yield (i.e., yield / expression) codon sequence 802 (i.e., if inference lower-yield / expression and higher-yield / expression codon sequences are expressed, then expression of the same protein would result). One having skill in the art will appreciate the ease of accessing one or more valid higher-yield / expression DNA sequences

[0183] FIG.9 is a flowchart of one example of generating higher yield codon sequences 900. As depicted, the generating higher-yield (i.e., yield / expression) codon sequences method 900 includes creating clusters of codon sequences on an amino acid sequence-basis 910 and using the clusters of codon sequences to build a lower-to-higher yield codon training dataset 920. Building the lower-to-higher yield codon training dataset, method 900 can include sorting codon sequences within each cluster by yield 930, generating on a cluster-basis pairs of codon sequences 940, including lower yield codon sequence as input 950, and including higher yield codon sequence as target 950. After building the lower-to-higher yield codon training dataset, method 900 may include training a high-yield codon sequence generator 960. After training, method 900 may include providing low yield codon sequences to generate high-yield codon sequences 970. Computer System

[0184] FIG.10 shows an example computer system 1000 that can be used to implement the technology disclosed. Computer system 1000 includes at least one central processing unit (CPU) 1042 that communicates with a number of peripheral devices via bus subsystem 1026. These peripheral devices can include a storage subsystem 1002 including, for example, memory devices and a file storage subsystem 1026, user interface input devices 1028, user interface output devices 1046, and a network interface subsystem 1044. The input and output devices allow user interaction with computer system 1000. NetworkAttorney Docket No. PRTN1008WO01 interface subsystem 1044 provides an interface to outside networks, including an interface to corresponding interface devices in other computer systems.

[0185] In one implementation, the deep neural network like the large language models disclosed here is communicably linked to the storage subsystem 1002 and the user interface input devices 1028.

[0186] User interface input devices 1028 can include a keyboard; pointing devices such as a mouse, trackball, touchpad, or graphics tablet; a scanner; a touch screen incorporated into the display; audio input devices such as voice recognition systems and microphones; and other types of input devices. In general, use of the term “input device” is intended to include all possible types of devices and ways to input information into computer system 1000.

[0187] User interface output devices 1046 can include a display subsystem, a printer, a fax machine, or non-visual displays such as audio output devices. The display subsystem can include an LED display, a cathode ray tube (CRT), a flat-panel device such as a liquid crystal display (LCD), a projection device, or some other mechanism for creating a visible image. The display subsystem can also provide a non-visual display such as audio output devices. In general, use of the term “output device” is intended to include all possible types of devices and ways to output information from computer system 1000 to the user or to another machine or computer system.

[0188] Storage subsystem 1002 stores programming and data constructs that provide the functionality of some or all of the modules and methods described herein. These software modules are generally executed by processors 1048.

[0189] Processors 1048 can be graphics processing units (GPUs), field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), and / or coarse-grained reconfigurable architectures (CGRAs). Processors 1048 can be hosted by a deep learning cloud platform such as Google Cloud Platform™, Xilinx™, and Cirrascale™. Examples of processors 1048 include Google’s Tensor Processing Unit (TPU)™, rackmount solutions like GX4 Rackmount Series™, GX20 Rackmount Series™, NVIDIA DGX-1™, Microsoft' Stratix V FPGA™, Graphcore's Intelligent Processor Unit (IPU)™, Qualcomm’s Zeroth Platform™ with Snapdragon processors™, NVIDIA’s Volta™, NVIDIA’s DRIVE PX™, NVIDIA’s JETSON TX1 / TX2 MODULE™, Intel’s Nirvana™, Movidius VPU™, Fujitsu DPI™, ARM’s DynamicIQ™, IBM TrueNorth™, Lambda GPU Server with Testa V100s™, and others.

[0190] Memory subsystem 1012 used in the storage subsystem 1002 can include a number of memories including a main random access memory (RAM) 1022 for storage of instructions and data during program execution and a read only memory (ROM) 1024 in which fixed instructions are stored. A file storage subsystem 1026 can provide persistent storage for program and data files, and can include a hard disk drive, a floppy disk drive along with associated removable media, a CD-ROM drive, an optical drive, or removable media cartridges. The modules implementing the functionality of certain implementations can be stored by file storage subsystem 1026 in the storage subsystem 1002, or in other machines accessible by the processor.Attorney Docket No. PRTN1008WO01

[0191] Bus subsystem 1036 provides a mechanism for letting the various components and subsystems of computer system 1000 communicate with each other as intended. Although bus subsystem 1036 is shown schematically as a single bus, alternative implementations of the bus subsystem can use multiple busses.

[0192] Computer system 1000 itself can be of varying types including a personal computer, a portable computer, a workstation, a computer terminal, a network computer, a television, a mainframe, a server farm, a widely-distributed set of loosely networked computers, or any other data processing system or user device. Due to the ever-changing nature of computers and networks, the description of computer system 1000 depicted in FIG.10 is intended only as a specific example for purposes of illustrating the preferred implementations of the present invention. Many other configurations of computer system 1000 are possible having more or less components than the computer system depicted in FIG.10.

[0193] In various implementations, a learning system is provided. In some implementations, a feature vector is provided to a learning system. Based on the input features, the learning system generates one or more outputs. In some implementations, the output of the learning system is a feature vector. In some implementations, the learning system comprises an SVM. In other implementations, the learning system comprises an artificial neural network. In some implementations, the learning system is pre-trained using training data. In some implementations training data is retrospective data. In some implementations, the retrospective data is stored in a data store. In some implementations, the learning system may be additionally trained through manual curation of previously generated outputs.

[0194] In some implementations, a sequence generator described herein may be a trained classifier. In some implementations, the trained classifier is a random decision forest. However, it will be appreciated that a variety of other classifiers are suitable for use according to the present disclosure, including linear classifiers, support vector machines (SVM), or neural networks such as recurrent neural networks (RNN).

[0195] Suitable artificial neural networks include but are not limited to a feedforward neural network, a radial basis function network, a self-organizing map, learning vector quantization, a recurrent neural network, a Hopfield network, a Boltzmann machine, an echo state network, long short term memory, a bi- directional recurrent neural network, a hierarchical recurrent neural network, a stochastic neural network, a modular neural network, an associative neural network, a deep neural network, a deep belief network, a convolutional neural networks, a convolutional deep belief network, a large memory storage and retrieval neural network, a deep Boltzmann machine, a deep stacking network, a tensor deep stacking network, a spike and slab restricted Boltzmann machine, a compound hierarchical-deep model, a deep coding network, a multilayer kernel machine, or a deep Q-network.

[0196] The present disclosure may be embodied as a system, a method, and / or a computer program product. The computer program product may include a computer readable storage medium (or media) having computer readable program instructions thereon for causing a processor to carry out aspects of the present disclosure.Attorney Docket No. PRTN1008WO01

[0197] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium may be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. A non-exhaustive list of more specific examples of the computer readable storage medium includes the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read- only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or raised structures in a groove having instructions recorded thereon, and any suitable combination of the foregoing. A computer readable storage medium, as used herein, is not to be construed as being transitory signals per se, such as radio waves or other freely propagating electromagnetic waves, electromagnetic waves propagating through a waveguide or other transmission media (e.g., light pulses passing through a fiber- optic cable), or electrical signals transmitted through a wire.

[0198] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network may comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.

[0199] FIG.10 is a schematic of an exemplary computing node. Computing node 2000 is only one example of a suitable computing node and is not intended to suggest any limitation as to the scope of use or functionality of embodiments described herein. Regardless, computing node 1000 is capable of being implemented and / or performing any of the functionality set forth hereinabove.

[0200] In computing node 1000 there is a computer system / server, which is operational with numerous other general purpose or special purpose computing system environments or configurations. Examples of well-known computing systems, environments, and / or configurations that may be suitable for use with computer system / server include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, multiprocessor systems, microprocessor-based systems, set top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed computing environments that include any of the above systems or devices, and the like.

[0201] Computer system / server may be described in the general context of computer system- executable instructions, such as program modules, being executed by a computer system. Generally, program modules may include routines, programs, objects, components, logic, data structures, and so onAttorney Docket No. PRTN1008WO01 that perform particular tasks or implement particular abstract data types. Computer system / server may be practiced in distributed computing environments where tasks are performed by remote processing devices that are linked through a communications network. In a distributed computing environment, program modules may be located in both local and remote computer system storage media including memory storage devices.

[0202] As shown in FIG.10, computer system / server in computing node 1000 is shown in the form of a general-purpose computing device. The components of computer system / server may include, but are not limited to, one or more processors or processing units, a system memory, and a bus that couples various system components including system memory to processor.

[0203] The Bus represents one or more of any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, and a processor or local bus using any of a variety of bus architectures. By way of example, and not limitation, such architectures include Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MCA) bus, Enhanced ISA (EISA) bus, Video Electronics Standards Association (VESA) local bus, Peripheral Component Interconnect (PCI) bus, Peripheral Component Interconnect Express (PCIe), and Advanced Microcontroller Bus Architecture (AMBA).

[0204] Computer system / server typically includes a variety of computer system readable media. Such media may be any available media that is accessible by computer system / server, and it includes both volatile and non-volatile media, removable and non-removable media.

[0205] System memory can include computer system readable media in the form of volatile memory, such as random access memory (RAM) and / or cache memory. Algorithm Computer system / server may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system can be provided for reading from and writing to a non-removable, non-volatile magnetic media (not shown and typically called a "hard drive"). Although not shown, a magnetic disk drive for reading from and writing to a removable, non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading from or writing to a removable, non-volatile optical disk such as a CD-ROM, DVD-ROM or other optical media can be provided. In such instances, each can be connected to bus by one or more data media interfaces. As will be further depicted and described below, memory may include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of embodiments of the disclosure.

[0206] Program / utility, having a set (at least one) of program modules, may be stored in memory by way of example, and not limitation, as well as an operating system, one or more application programs, other program modules, and program data. Each of the operating system, one or more application programs, other program modules, and program data or some combination thereof, may include an implementation of a networking environment. Program modules generally carry out the functions and / or methodologies of embodiments as described herein.Attorney Docket No. PRTN1008WO01

[0207] Computer readable program instructions for carrying out operations of the present disclosure may be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language such as Smalltalk, C++ or the like, and conventional procedural programming languages, such as the “C” programming language or similar programming languages. The computer readable program instructions may execute entirely on the user’s computer, partly on the user’s computer, as a stand-alone software package, partly on the user’s computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer may be connected to the user’s computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection may be made to an external computer (for example, through the Internet using an Internet Service Provider). In some implementations, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) may execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.

[0208] Aspects of the present disclosure are described herein with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to implementations of the disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer readable program instructions.

[0209] These computer readable program instructions may be provided to a processor of a general- purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions may also be stored in a computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.

[0210] The computer readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable apparatus or other device to produce a computer implemented process, such that the instructions which execute on the computer, other programmable apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.Attorney Docket No. PRTN1008WO01

[0211] The flowchart and block diagrams in the FIG.s illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the FIG.s. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hardware and computer instructions. Reinforcement Learning with Human Feedback (RLHF)

[0212] RLHF incorporates human input into the RL process to improve learning efficiency, adaptability, and safety. In basic RL frameworks, an agent learns to make decisions by interacting with an environment. The agent receives feedback in the form of rewards or penalties based on its actions, guiding the agent towards optimal behavior.

[0213] Basic RL faces several challenges. First, some tasks have sparse or delayed rewards that can make RL learning slow or difficult. Second, the trade-offs between RL exploration and RL exploitation can be difficult. More specifically, it may be difficult to achieve effective RL learning, as it can be difficult to balance exploration of new strategies with exploitation of known good strategies.

[0214] To address these challenges, human feedback is incorporated into RL in several ways. First, in the form of reward shaping, wherein humans can provide additional reward signals or modify existing ones to guide the agent more effectively. Reward shaping can speed up basic RL learning by providing more informative feedback. Second, human feedback can be incorporated into basic RL by imitation learning. Imitation learning refers to the concept that humans can demonstrate desired behaviors, and the agent learns by imitating these demonstrations. This reduces exploration, compared to basic RL, in complex or dangerous environments. Third, human feedback can be incorporated into basic RL by feedback on policies. More specifically, humans can provide feedback directly to the agent, regarding the agent's policies or decision-making processes. In turn, direct feedback on policies helps the agent learn faster and avoid costly mistakes.

[0215] Human feedback can include several types. For example, human feedback may include Explicit Rewards, wherein humans assign rewards or penalties based on the agent's actions. In other examples, human feedback can include Demonstrations, such that humans demonstrate desired behaviors, and the agent learns from these examples. In some implementations, human feedback entails Preferences, such that humans express preferences or rankings over different actions or outcomes that guide the agent'sAttorney Docket No. PRTN1008WO01 decision-making. Still in other implementations, human feedback can include Critiques, wherein humans provide feedback on the agent's decisions, in which humans point out the agent’s errors in decision-making or suggest improvements to these decisions.

[0216] There are multiple ways this human integration into basic RL can be implemented. For example, human integration into basic RL may be implemented via Reward Augmentation, wherein human-provided rewards are combined with intrinsic rewards from the environment to create a more informative signal. In other examples, Inverse Reinforcement Learning (IRL) may be implemented by having the agent infer the underlying reward function from human demonstrations, that allows it to learn complex behaviors. In further examples, human integration may be implemented by Interactive Learning, wherein the agent interacts with humans in real-time, receiving feedback during training episodes.

[0217] RLHF may be implemented according to several phases. The first phase, Supervised Fine- Tuning (SFT), provides that RLHF begins with a pre-trained language model that is then fine-tuned on high-quality datasets for specific applications. The next phase, Preference Sampling and Reward Learning, entails collecting human preferences between pairs of language model outputs and using these preferences to learn a reward function, typically employing the Bradley-Terry model. The final phase, Reinforcement Learning Optimization, uses the learned reward function to further fine-tune the language model, focusing on maximizing the reward for the outputs while maintaining proximity to its original training.

[0218] The above RLHF phases may utilize, for example, any of the following language models as appropriate: cluster-by-cluster high yield DNA sequence generator 660, high-yield codon sequence generator 820, or others. In addition, the above RLHF may utilize fine-tuning components for any pre- trained models described herein, as appropriate. Direct

[0219] The language models of the present invention may also be compatible with direct performance optimization (DPO), a parameterization method of the reward model in RLHF, that enables the extraction of the corresponding optimal policy in a closed form. The DPO approach simplifies the RLHF problem to a simple classification loss, making the algorithm stable, performant, and computationally lightweight.

[0220] In the present invention, DPO may combine the reward function and language model into a single transformer network. This combined single transformer network, DPO, only requires the language model to be trained, which more directly and efficiently aligns the combined single transformer with human preferences. The combined single transformer network based on DPO, can deduce which reward function the language model can best maximize, thereby streamlining the entire process.

[0221] The above DPO approach may utilize, for example, any of the following language models as appropriate: cluster-by-cluster high yield DNA sequence generator 660, high-yield codon sequence generator 820, or others.Attorney Docket No. PRTN1008WO01 Performance Results as Objective Indicia of Inventiveness

[0222] FIG.11 includes a chart that shows the fold-change in yield / expression for five proteins that were each expressed in the HEK293 cell line. As shown in FIG.11, each of the five proteins has at least two variants with a 1.5- to 2-fold or more increase in protein expression relative to the wild type / reference sequence, demonstrating the successful performance of the protein-to-codon sequence mapping system. In fact, two of the proteins had variants with at least a 3- to 4-fold increase in protein expression.

[0223] FIG.12 includes a chart that shows the fold-change in yield / expression for three proteins that were each expressed in the Yeast Pichia cell line. As shown in FIG.12, two of the three proteins have at least two variants with a 1.5- to 2-fold or more increase in protein expression relative to the wild type / reference sequence, demonstrating the successful performance of the protein-to-codon sequence mapping system. In fact, two of the proteins had variants with at least a 6- to 7-fold increase in protein expression.

[0224] FIG.13 includes a chart that shows the fold-change in yield / expression for two proteins that were each expressed in the HEK293 cell line. As shown in FIG.13, both proteins have at least two variants with a 2-fold or more increase in protein expression relative to the wild type / reference sequence, demonstrating the successful performance of the protein-to-codon sequence mapping system. In fact, one protein had three of the six variants exhibit at least a 3- to 4-fold increase in protein expression.

[0225] FIG.14 includes a table that shows the fold-change in yield / expression for two protein classes, each with several corresponding proteins, that were each expressed in the CHO cell line. As shown in the table of FIG.14, three of the five mAb proteins have at least one variant with a 7-fold or more increase in protein expression relative to the wild type / reference sequence, demonstrating the hugely successful performance of the protein-to-codon sequence mapping system. In fact, two proteins had two of their three variants exhibit at least a 26-fold increase in protein expression, with one variant even reaching a 65-fold increase in protein expression. As shown in the table of FIG.14, three of the four VHH4 proteins have at least one variant with a 1.3-fold or more increase in protein expression relative to the wild type / reference sequence, demonstrating the successful performance of the protein-to-codon sequence mapping system

[0226] FIG.15A includes a table that shows the fold-change in yield / expression for the first of two protein classes, each with several corresponding protein indices, that were each expressed in the HEK293 cell line. As shown in the table of FIG.15A, four of the five VHH4 proteins have at least one variant with a 2-fold or more increase in protein expression relative to the wild type / reference sequence, demonstrating the successful performance of the protein-to-codon sequence mapping system.

[0227] FIG.15B includes a table that shows the fold-change in yield / expression for the second of two protein classes, each with several corresponding proteins, that were each expressed in the HEK293 cell line. As shown in the table of FIG.15B, all of the eight mAb proteins have at least one variant with a 2-fold or more increase in protein expression relative to the wild type / reference sequence, demonstrating the hugely successful performance of the protein-to-codon sequence mapping system. In fact, five of the eightAttorney Docket No. PRTN1008WO01 proteins had one or more variants exhibit at least a 10-fold increase in protein expression, with three proteins having one or more variants with at least a 20-fold increase in protein expression, and one particularly striking protein had five variants reach at least a 100-fold increase in protein expression.

[0228] Disclosed is a deep learning-based framework for sequence evaluation, referred to as “GeneEval.” GeneEval counters data sparsity limitations disabling high-confidence in-silico prediction of protein yield, allowing filtration of top generated variants and enabling faster validation of protein-based products. GeneEval is a sequence evaluation framework that enables the utilization of up to three types of hand-crafted input features to maximize the prediction correlation using a small-scale dataset. GeneEval is built on the assumption that protein yield can be correlated by utilizing one or more features of one or more of the three feature types that are directly / indirectly relevant to the yield. The three types of input features can include, for example, predicted features (i.e., via learning-based models), computed features (i.e., via reference dataset / metric equation), and annotation-based features (i.e., existing annotating in public reference databases, or known information on the protein / plasmid’s map). By incorporating the aforementioned features into a multi-task learning setting that enables the utilization of abundant training data per predictable features, GeneEval bridges the gap towards the small-scale downstream data and increases the information content of the input vector to enable the prediction of the protein yield. GeneEval is a multi-step framework that includes single-task or multi-task learning, and ensemble-based fine-tuning, which are described below in turn.involves training individual single-task models, each tailored to predict one specific metric that correlates with DNA sequence expression levels and serves as reliable expression proxies (e.g., yield-proxy metric). These models utilize a dedicated prediction head and a unified backbone architecture to focus exclusively on optimizing performance for its respective proxy metric. By aligning with the distinct characteristics of each metric, these models ensure high precision and task-specific specificity.

[0230] To further enhance predictive robustness, an ensemble of these single-task models can be constructed. By training multiple models independently on their respective proxy metrics and subsequently combining their outputs, the ensemble leverages the diverse learned representations and decision strategies of each model. This approach mitigates the risk of overfitting to any single proxy metric or dataset, yielding a more accurate and generalizable final ranking for expression levels.

[0231] Instead of single-task learning, multi-task learning can be adopted. Multi-task learning involves training a unified model on multiple distinct tasks concurrently, utilizing a separate prediction head for each task while employing a shared backbone architecture across all tasks. This multi-task training framework enhances the model's capability to rank DNA sequence expression levels without disproportionately favoring any single proxy metric employed during training. By mandating that theAttorney Docket No. PRTN1008WO01 backbone learns a shared representation applicable across all prediction heads, the model potentially develops a more comprehensive and robust intermediate representation.

[0232] In some examples, the yield-proxy metrics can include one or more protein abundance metrics. For example, the yield-proxy metric can include a protein-per-million (PPM) metric. PPM is a metric used to quantify the abundance of a protein. In another example, the yield-proxy metric can include a protein titer metric. Protein titers refer to the concentration of a protein in a solution, but can also be a ratio of volume-to-volume, weight-to-volume, or weight-to-weight. In another example, the yield-proxy metric can include a protein-molecules-per-cell metric. Protein-molecules-per-cell refers to the total number of individual protein molecules present within a single cell. Additionally, it is expressly contemplated that other protein abundance metrics can be utilized as well.

[0233] In some examples, the yield-proxy metrics can also include one or more translation efficiency metrics. For example, the yield-proxy metric can include a protein-per-transcript (PPT) metric. PPT is a metric used to quantify the abundance of a protein relative to the level of its corresponding mRNA transcript. The PPT value is calculated by dividing the protein abundance (measured, for example, by mass spectrometry) by the corresponding mRNA abundance (measured, for example, by RNA sequencing). PPT is a valuable metric as it provides insights into the efficiency of translation from mRNA to protein as it maps the relation between mRNA expression levels and protein levels, determining the protein turnover of a mRNA transcript. In another example, the yield proxy metric can include a protein-to-mRNA-ratio metric. In addition to quantifying the abundance of a protein relative to the level of its corresponding mRNA transcript (such as that with PPT), PTR accounts for both the protein and mRNA half-life. In another example, the yield-proxy metric can include a ribosomal profiling metric. Ribosomal profiling measures ribosome occupancy on mRNA to assess translation rates. This technique provides a snapshot of ribosome positions on mRNA, giving insights into translation rates and ribosome density. Ribosomal profiling involves sequencing ribosome-protected mRNA fragments to determine which mRNAs are being actively translated, and how efficiently. Additionally, it is expressly contemplated that other translation efficiency metrics can be utilized as well.

[0234] In some examples, the yield-proxy metrics can also include one or more protein stability metrics. For example, the yield-proxy metric can include a protein half-life metric. Protein half-life refers to the time it takes for half of the protein molecules in a cell to be degraded. Protein half-life can be measured using protein turnover assays. In another example, the yield-proxy metric can include a protein degradation rate metric. Protein degradation rate refers to the rate at which protein molecules are broken down, often expressed as a degradation constant (k). Protein degradation rate is determined by analyzing protein decay kinetics over time. Additionally, it is expressly contemplated that other protein stability metrics can be utilized as well.

[0235] In some examples, the yield-proxy metrics can also include one or more mRNA abundance metrics. For example, the yield-proxy metric can include a reads per kilobase million (RPKM) metric. RPKM is a measure of gene expression commonly used in RNA sequencing, and represents the number ofAttorney Docket No. PRTN1008WO01 reads mapped to a gene normalized by gene length and sequencing depth. RPKM takes into account both gene length and the total number of reads in the experiment. In another example, the yield-proxy metric can include a transcripts per million (TPM) metric. TPM is another measure of gene expression commonly used in RNA-seq analysis, and represents the relative abundance of a transcript in a sample by normalizing the number of reads mapped to a gene by the total number of mapped reads in the sample. TPM provides a standardized way to compare gene expression levels between different samples and arguably, becomes the most accurate method particularly for comparing data from different samples. In another example, the yield-proxy metric can include a fragments per kilobase million (FPKM) metric. FPKM is similar to RPKM but is used when the sequencing technology generates short fragments (fragments rather than full- length reads, hence FPKM is used for normalizing paired-end reading, right and left, of the fragment). FPKM is also used in RNA-seq analysis and provides a normalized measure of gene expression, considering fragment length and sequencing depth. Additionally, it is expressly contemplated that other mRNA abundance metrics can be utilized as well.

[0236] In some examples, the yield-proxy metrics can also include one or more mRNA stability metrics. For example, the yield-proxy metric can include an mRNA half-life metric. mRNA half-life refers to the time it takes for half of the mRNA molecules in a cell to be degraded. mRNA half-life is an important determinant of expression levels. Short-lived mRNAs are rapidly degraded, leading to lower protein production, while long-lived mRNAs can persist for a more extended period, resulting in higher protein levels. The stability of mRNA molecules is influenced by various factors, including the presence of specific sequences in the mRNA (e.g., AU-rich elements) and the action of RNA-binding proteins and microRNAs that can promote or inhibit mRNA degradation. In another example, the yield-proxy metric can include an mRNA degradation rate metric. mRNA degradation rate refers to the rate at which mRNA molecules are broken down, often expressed as a degradation constant (k). mRNA degradation rate can be determined by analyzing the decay kinetics of mRNA over time. Additionally, it is expressly contemplated that other mRNA stability metrics can be utilized as well.

[0237] In some examples, the yield proxy metrics can also include one or more functional / structural attributes. For example, the yield-proxy metric can include a solubility / monomer percentage. Solubility / monomer percentage refers to the amount of protein that exists in a soluble, non-aggregated form. In another example, the yield-proxy metric can include localization. Localization refers to the end location of the protein to perform its function. For instance, some proteins may be cytoplasmic proteins, whereas other proteins might be nuclear proteins, membrane proteins, etc. In another example, the yield- proxy metric can include protein secondary structure. Protein secondary structure refers to the local folding patterns of a polypeptide chain between backbone atoms of amino acids, resulting in structures like alpha helices and beta sheets. In another example, the yield-proxy metric can include mRNA secondary structure. mRNA secondary structure refers to the base pairing between the nucleotides of the mRNA. Additionally, it is expressly contemplated that other functional / structural attributes can be utilized as well.Attorney Docket No. PRTN1008WO01

[0238] Additionally, domain-specific features, such as computed features i.e. codon metrics derived from the input sequence, can be incorporated into each model to provide additional biological context and improve predictive performance. These features can be integrated into the model architecture through two primary methods: (1) appending them to the intermediate representation vector (embeddings) generated by the backbone, or (2) concatenating them with the input sequence prior to inputting into the backbone, with necessary architectural adjustments to accommodate the augmented input. Following the embedding extraction of the input sequences and the pooling of the embeddings, the features are concatenated to the pooled embedding vector. Examples of the embedding vector are shown below, respectively.

[0239] Embedding Vector Before Concatenation can be: [x1, x2, x3, ……., xn].

[0240] Embedding Vector After Concatenation can be: [Feature_1, Feature_2, …, Feature_m, x1, x2, x3, ……., xn].

[0241] Some examples of computed features / codon metrics can include RNA folding, and / or codon indices / metrics (e.g., CpG, GC Content, GC3 Content, Directional Relative Rodon Bias Score, Effective Number of Codons, Effective Number of Codons with Background Correction, Relative Codon Bias Score, Relative Synonymous Codon Usage, tRNA Adaptation Index, Codon Pair Bias, and / or Frequency of Optical Codons). Additionally, it is expressly contemplated that other computed features / codon metrics can be utilized as well.

[0242] Additionally, annotations that span different categories of tags can be incorporated into each model to provide additional biological context and improve predictive performance. These features can be integrated into the model architecture through two primary methods: (1) appending them to the intermediate representation vector (embeddings) generated by the backbone, or (2) concatenating them with the input sequence prior to inputting into the backbone, with necessary architectural adjustments to accommodate the augmented input. Examples of the embedding vector are shown below, respectively.

[0243] Embedding Vector Before Concatenation: [x1, x2, x3, ……., xn].

[0244] Embedding Vector After Concatenation: [Feature_1, Feature_2, …, Feature_m, x1, x2, x3, ……., xn].

[0245] Some examples of retrieved annotations can include a Plasmid Map (e.g., Promoter Sequence, Signal Peptide, Coding Sequence (CDS), Purification Tags, Terminator Sequence, Ribosome Binding Site (RBS), and / or Regulator / Enhancer Sequences). In another example, the retrieved annotations can include taxonomic tags, such as Domain (e.g., "Bacteria," "Eukaryota"), Kingdom (e.g., "Animalia," "Fungi"), Phylum (e.g., "Chordata," "Proteobacteria"), Class (e.g., "Mammalia," "Gammaproteobacteria"), and / or Order Specific classifications (e.g., family, species, strains, etc.). In another example, the retrieved annotations can include protein CATH domains. Additionally, it is expressly contemplated that other retrieved annotations can be utilized as well.

[0246] The utilized backbone model to perform the learning can span a variety of model types (pre- trained / trained from scratch) as well as fine-tuning modules. Example models can include, but are notAttorney Docket No. PRTN1008WO01 limited to, DNA BERT, DNA BERT2, The Nucleotide Transformer, The Enformer, and CodonTransformer. Ensemble-Based Fine-Tuning

[0247] Optionally, ensemble-based fine-tuning can additionally be performed. In this step, another training round is done that utilizes the same model used in the above component. However, it is expressly contemplated that a new model can be utilized as well. Unlike the previous round of training, the label in this component is the final prediction target that is protein yield or the relevant gene optimization objective function. Example fine-tuning modules can include, but are not limited to, DNA BERT, DNA BERT2, The Nucleotide Transformer, The Enformer, CodonTransformer, Full Back Propagation, and Low Rank Adaptation (LoRA).

[0248] FIG.16 is a diagram showing one example of multi-task learning and ensemble-based fine tuning of a sequence evaluation system 1600. To make use of the abundant training data that exists for each metric, multi-task learning 1602 takes place in the first component of GeneEval to utilize the relationship between different tasks in enriching the performance of each task as well as enabling the learning of a share embedding vector through a single backbone language model 1604. In the second component, the output embedding vector is then pre-processed by concatenating the three types of features to the embedding vector. In the third component, a downstream model 1606 is used to fine-tune the backbone model 1604 and pass the input to an external model to train on the ground-truth small-scale data.

[0249] FIG.17 is a table showing sequence evaluation performance compared to fold-change in experimental yields across five mAb proteins. The performance of GeneEval was evaluated by comparing the ranking done by GeneEval compared to actual experimental yields. As shown, the comparison demonstrates that the ranking of the variants according to GeneEval follows the same pattern of ranking according to the experimental yields for the five mAb proteins. In the example shown in FIG.17, the sequence evaluation of GeneEval was determined by utilizing a PTR metric as the yield-proxy metric (e.g., by utilizing models trained on the PTR metric). However, it is expressly contemplated that the utilization of the PTR metric is only by way of example, and other yield-proxy metrics could alternatively or additionally be utilized as well, such as any of those described above. Clauses

[0250] The technology disclosed can be practiced as a system, method, or article of manufacture. One or more features of an implementation can be combined with the base implementation. Implementations that are not mutually exclusive are taught to be combinable. One or more features of an implementation can be combined with other implementations. This disclosure periodically reminds the user of these options. Omission from some implementations of recitations that repeat these options should not be taken as limiting the combinations taught in the preceding sections – these recitations are hereby incorporated forward by reference into each of the following implementations.Attorney Docket No. PRTN1008WO01

[0251] One or more implementations and clauses of the technology disclosed, or elements thereof can be implemented in the form of a computer product, including a non-transitory computer readable storage medium with computer usable program code for performing the method steps indicated. Furthermore, one or more implementations and clauses of the technology disclosed, or elements thereof can be implemented in the form of an apparatus including a memory and at least one processor that is coupled to the memory and operative to perform exemplary method steps. Yet further, in another aspect, one or more implementations and clauses of the technology disclosed or elements thereof can be implemented in the form of means for carrying out one or more of the method steps described herein; the means can include (i) hardware module(s), (ii) software module(s) executing on one or more hardware processors, or (iii) acombination of hardware and software modules; any of (i)-(iii) implement the specific techniques set forth herein, and the software modules are stored in a computer readable storage medium (or multiple such media).

[0252] The clauses described in this section can be combined as features. In the interest of conciseness, the combinations of features are not individually enumerated and are not repeated with each base set of features. The reader will understand how features identified in the clauses described in this section can readily be combined with sets of base features identified as implementations in other sections of this application. These clauses are not meant to be mutually exclusive, exhaustive, or restrictive; and the technology disclosed is not limited to these clauses but rather encompasses all possible combinations, modifications, and variations within the scope of the claimed technology and its equivalents.

[0253] Other implementations of the clauses described in this section can include a non-transitory computer readable storage medium storing instructions executable by a processor to perform any of the clauses described in this section. Yet another implementation of the clauses described in this section can include a system including memory and one or more processors operable to execute instructions, stored in the memory, to perform any of the clauses described in this section.

[0254] We disclose the following clauses: Clause Set 11. A computer-implemented method of providing a relative positional embedding to a multi-head attentiontransformer model, including: using respective attention heads of the multi-head attention transformer model to generate query, key, and value vectors for inputs tokens; converting the query and key vectors into position-encoded query and key vectors by applying a series of rotation matrices to the query and key vectors at different scaled frequencies that vary by absolute positions of the query and key vectors and by the respective attention heads, wherein the application of the series of rotation matrices includes: rotating pairs of feature dimensions in the query and key vectors by an angle in multiples of aAttorney Docket No. PRTN1008WO01 scaled position index of a corresponding query or key vector, wherein the scaled position index is scaled by a head-specific scaling parameter that varies across the respective attention heads; and using the position-encoded query and key vectors for execution of self-attention by the respective attention heads to generate pairwise attention scores that depend on relative positions of input token pairs and on their feature similarity.2. The computer-implemented method of clause 1, wherein the head-specific scaling parameter is a head-specific scaling scalar.3. The computer-implemented method of clause 1, wherein the pairwise attention scores are penalized basedon how far the position-encoded query and key vectors are.4. The computer-implemented method of clause 3, wherein when a position-encoded query vector and aposition-encoded key vector are close by, the penalty is very low.5. The computer-implemented method of clause 3, wherein when a position-encoded query vector and aposition-encoded key vector are far away, the penalty is very high.6. The computer-implemented method of clause 1, wherein the input token pairs are amino acid token pairs.7. The computer-implemented method of clause 1, wherein the input token pairs are nucleotide token pairs.8. A system, comprising:using respective attention heads of the multi-head attention transformer model to generate query, key, and value vectors for inputs tokens; converting the query and key vectors into position-encoded query and key vectors by applying a series of rotation matrices to the query and key vectors at different scaled frequencies that vary by absolute positions of the query and key vectors and by the respective attention heads, wherein the application of the series of rotation matrices includes: rotating pairs of feature dimensions in the query and key vectors by an angle in multiples of a scaled position index of a corresponding query or key vector, wherein the scaled position index is scaled by a head-specific scaling parameter that varies across the respective attention heads; and using the position-encoded query and key vectors for execution of self-attention by the respective attention heads to generate pairwise attention scores that depend on relative positions of input token pairs and on their feature similarity.9. The system of clause 8, wherein the head-specific scaling parameter is a head-specific scaling scalar.Attorney Docket No. PRTN1008WO0110. The system of clause 8, wherein the pairwise attention scores are penalized based on how far the position-encoded query and key vectors are.11. The system of clause 10, wherein when a position-encoded query vector and a position-encoded key vectorare close by, the penalty is very low.12. The system of clause 10, wherein when a position-encoded query vector and a position-encoded key vectorare far away, the penalty is very high.13. The system of clause 8, wherein the input token pairs are amino acid token pairs.14. The system of clause 8, wherein the input token pairs are nucleotide token pairs.15. A non-transitory computer readable medium impressed with computer program instructions, theinstructions, when executed on a processor, implement a method of providing a relative positional embedding to a multi-head attention transformer model, including: using respective attention heads of the multi-head attention transformer model to generate query, key, and value vectors for inputs tokens; converting the query and key vectors into position-encoded query and key vectors by applying a series of rotation matrices to the query and key vectors at different scaled frequencies that vary by absolute positions of the query and key vectors and by the respective attention heads, wherein the application of the series of rotation matrices includes: rotating pairs of feature dimensions in the query and key vectors by an angle in multiples of a scaled position index of a corresponding query or key vector, wherein the scaled position index is scaled by a head-specific scaling parameter that varies across the respective attention heads; and using the position-encoded query and key vectors for execution of self-attention by the respective attention heads to generate pairwise attention scores that depend on relative positions of input token pairs and on their feature similarity.16. The non-transitory computer readable storage medium of clause 15, wherein the head-specific scalingparameter is a head-specific scaling scalar.17. The non-transitory computer readable storage medium of clause 15, wherein the pairwise attention scoresare penalized based on how far the position-encoded query and key vectors are.18. The non-transitory computer readable storage medium of clause 17, wherein when a position-encodedquery vector and a position-encoded key vector are close by, the penalty is very low.Attorney Docket No. PRTN1008WO0119. The non-transitory computer readable storage medium of clause 17, wherein when a position-encodedquery vector and a position-encoded key vector are far away, the penalty is very high.20. The non-transitory computer readable storage medium of clause 15, wherein the input token pairs areamino acid token pairs.21. The non-transitory computer readable storage medium of clause 15, wherein the input token pairs arenucleotide token pairs.22. A computer-implemented method of generating optimized codon sequences, including:processing an input sequence of amino acid elements; and based on the processing, generating an output sequence of codon elements, including: for a given amino acid element in the input sequence of amino acid elements, confining search of a corresponding codon element in a vocabulary of codon elements to a subset of codon elements known to translate to the given amino acid element. In some implementations, given the codon / nucleotide sequence sampling often takes place in a manner that's unrestricted by the protein sequence but rather guided with it, it is probable that the codons / nucleotides sampled by the decoder may not encode back to the original amino acids.23. The computer-implemented method of clause 22, further including, prior to the search, evaluating at theinput sequence of amino acid elements to determine an identity of the given amino acid element, and using the determined identity of the given amino acid element to confine the search of the corresponding codon element to the subset of codon elements known to translate to the given amino acid element.24. The computer-implemented method of clause 22, wherein a neural network (e.g., AA2DNA) processes theinput sequence of amino acid elements, and generates the output sequence of codon elements.25. The computer-implemented method of clause 22, wherein the neural network is a pre-trained neuralnetwork.26. The computer-implemented method of clause 22, wherein the neural network is trained from scratch.27. The computer-implemented method of clause 24, wherein the neural network is a sequence to sequence(seq2seq) neural network (e.g., RNNs like LSTMs and GRUs, and also Transformers).28. The computer-implemented method of clause 24, wherein the neural network is an encoder-decoder neuralnetwork (e.g., Transformer).Attorney Docket No. PRTN1008WO0129. The computer-implemented method of clause 28, wherein an encoder neural network generates respectiveembedded tokens for respective amino acid elements in the input sequence of amino acid elements.30. The computer-implemented method of clause 29, wherein the encoder neural network applies attentionbetween the respective embedded tokens on an amino acid element-by-amino acid element.31. The computer-implemented method of clause 30, wherein a decoder neural network receives results of theattention from the encoder neural network, and uses the results of the attention to generate the output sequence of codon elements.32. The computer-implemented method of clause 31, wherein the decoder neural network looks back at theinput sequence of amino acid elements to determine the identity of the given amino acid element, and uses the determined identity of the given amino acid element to confine the search of the corresponding codon element to the subset of codon elements known to translate to the given amino acid element.33. A system comprising:processing an input sequence of amino acid elements; and based on the processing, generating an output sequence of codon elements, including: for a given amino acid element in the input sequence of amino acid elements, confining search of a corresponding codon element in a vocabulary of codon elements to a subset of codon elements known to translate to the given amino acid element. In some implementations, given the codon / nucleotide sequence sampling often takes place in a manner that's unrestricted by the protein sequence but rather guided with it, it is probable that the codons / nucleotides sampled by the decoder may not encode back to the original amino acids.34. The system of clause 33, further comprising, prior to the search, evaluating at the input sequence of aminoacid elements to determine an identity of the given amino acid element, and using the determined identity of the given amino acid element to confine the search of the corresponding codon element to the subset of codon elements known to translate to the given amino acid element.35. The system of clause 33, wherein a neural network (e.g., AA2DNA) processes the input sequence of aminoacid elements, and generates the output sequence of codon elements.36. The system of clause 33, wherein the neural network is a pre-trained neural network.37. The system of clause 33, wherein the neural network is trained from scratch.38. The system of clause 35, wherein the neural network is a sequence to sequence (seq2seq) neural network(e.g., RNNs like LSTMs and GRUs, and also Transformers).Attorney Docket No. PRTN1008WO0139. The system of clause 35, wherein the neural network is an encoder-decoder neural network (e.g.,Transformer).40. The system of clause 39, wherein an encoder neural network generates respective embedded tokens forrespective amino acid elements in the input sequence of amino acid elements.41. The system of clause 40, wherein the encoder neural network applies attention between the respectiveembedded tokens on an amino acid element-by-amino acid element.42. The system of clause 41, wherein a decoder neural network receives results of the attention from theencoder neural network, and uses the results of the attention to generate the output sequence of codon elements.43. The system of clause 42, wherein the decoder neural network looks back at the input sequence of aminoacid elements to determine the identity of the given amino acid element, and uses the determined identity of the given amino acid element to confine the search of the corresponding codon element to the subset of codon elements known to translate to the given amino acid element.44. A non-transitory computer readable medium impressed with computer program instructions, theinstructions, when executed on a processor, implement a method of generating optimized codon sequences, including: processing an input sequence of amino acid elements; and based on the processing, generating an output sequence of codon elements, including: for a given amino acid element in the input sequence of amino acid elements, confining search of a corresponding codon element in a vocabulary of codon elements to a subset of codon elements known to translate to the given amino acid element. In some implementations, given the codon / nucleotide sequence sampling often takes place in a manner that's unrestricted by the protein sequence but rather guided with it, it is probable that the codons / nucleotides sampled by the decoder may not encode back to the original amino acids.45. The non-transitory computer readable storage medium of clause 44, further including, prior to the search,evaluating at the input sequence of amino acid elements to determine an identity of the given amino acid element, and using the determined identity of the given amino acid element to confine the search of the corresponding codon element to the subset of codon elements known to translate to the given amino acid element.Attorney Docket No. PRTN1008WO0146. The non-transitory computer readable storage medium of clause 44, wherein a neural network (e.g.,AA2DNA) processes the input sequence of amino acid elements, and generates the output sequence of codon elements.47. The non-transitory computer readable storage medium of clause 44, wherein the neural network is a pre-trained neural network.48. The non-transitory computer readable storage medium of clause 44, wherein the neural network is trainedfrom scratch.49. The non-transitory computer readable storage medium of clause 46, wherein the neural network is asequence to sequence (seq2seq) neural network (e.g., RNNs like LSTMs and GRUs, and also Transformers).50. The non-transitory computer readable storage medium of clause 46, wherein the neural network is anencoder-decoder neural network (e.g., Transformer).51. The non-transitory computer readable storage medium of clause 50, wherein an encoder neural networkgenerates respective embedded tokens for respective amino acid elements in the input sequence of amino acid elements.52. The non-transitory computer readable storage medium of clause 51, wherein the encoder neural networkapplies attention between the respective embedded tokens on an amino acid element-by-amino acid element.53. The non-transitory computer readable storage medium of clause 52, wherein a decoder neural networkreceives results of the attention from the encoder neural network, and uses the results of the attention to generate the output sequence of codon elements.54. The non-transitory computer readable storage medium of clause 53, wherein the decoder neural networklooks back at the input sequence of amino acid elements to determine the identity of the given amino acid element, and uses the determined identity of the given amino acid element to confine the search of the corresponding codon element to the subset of codon elements known to translate to the given amino acid element.55. A computer-implemented method of generating high-yield codon sequences, including:creating clusters of codon sequences on an amino acid sequence-basis, wherein a particular cluster of codon sequences created for a particular amino acid sequence includes different codon sequences that translate to the particular amino acid sequence but have varying yields / expression; and using the clusters of codon sequences to build a lower-to-higher yield codon training dataset that links lower-yield input codon sequences to higher-yield target codon sequences by:Attorney Docket No. PRTN1008WO01 sorting codon sequences within each of the clusters of codon sequences by yield; based on the sorting, generating, on a cluster-basis, pairs of codon sequences in which one codon sequence has a lower yield and another codon sequence has a higher yield; including the codon sequence with the lower yield in the lower-to-higher yield codon training dataset as a lower-yield input codon sequence; and including the codon sequence with the higher yield in the lower-to-higher yield codon training dataset as a higher-yield target codon sequence for the codon sequence with the lower yield.56. The computer-implemented method of clause 55, further including using the lower-to-higher yield codontraining dataset to train a high-yield codon sequence generator to map the lower-yield input codon sequences to the higher-yield target codon sequences.57. The computer-implemented method of clause 56, wherein the high-yield codon sequence generator is aneural network.58. The computer-implemented method of clause 57, wherein the neural network is a pre-trained neuralnetwork.59. The computer-implemented method of clause 57, wherein the neural network is trained from scratch.60. The computer-implemented method of clause 57, wherein the neural network is a sequence to sequence(seq2seq) neural network (e.g., RNNs like LSTMs and GRUs, and also Transformers).61. The computer-implemented method of clause 60, wherein the neural network is an encoder-decoder neuralnetwork (e.g., Transformer).62. The computer-implemented method of clause 65, wherein at least some of the clusters of codon sequencesare generated based on clinical data that identifies different codons sequences that translate to a same amino acid sequence. In other implementations, at least some of the clusters of codon sequences are generated based on “labeled data” that can use transcriptomic labeling and proteomic labeling.63. The computer-implemented method of clause 65, wherein at least some of the clusters of codon sequencesare generated based on an amino acid sequence-to-codon sequence generator generating different output codon sequences for a same input amino acid sequence.64. The computer-implemented method of clause 63, wherein the amino acid sequence-to-codon sequencegenerator is a neural network.Attorney Docket No. PRTN1008WO0165. The computer-implemented method of clause 64, wherein the neural network is a pre-trained neuralnetwork.66. The computer-implemented method of clause 64, wherein the neural network is an untrained neuralnetwork.67. The computer-implemented method of clause 64, wherein the neural network is a sequence to sequence(seq2seq) neural network.68. The computer-implemented method of clause 67, wherein the neural network is an encoder-decoder neuralnetwork (e.g., Transformer).69. A non-transitory computer readable medium impressed with computer program instructions, theinstructions, when executed on a processor, implement a method of generating high-yield codon sequences, including: creating clusters of codon sequences on an amino acid sequence-basis, wherein a particular cluster of codon sequences created for a particular amino acid sequence includes different codon sequences that translate to the particular amino acid sequence but have varying yields / expression; and using the clusters of codon sequences to build a lower-to-higher yield codon training dataset that links lower-yield input codon sequences to higher-yield target codon sequences by: sorting codon sequences within each of the clusters of codon sequences by yield; based on the sorting, generating, on a cluster-basis, pairs of codon sequences in which one codon sequence has a lower yield and another codon sequence has a higher yield; including the codon sequence with the lower yield in the lower-to-higher yield codon training dataset as a lower-yield input codon sequence; and including the codon sequence with the higher yield in the lower-to-higher yield codon training dataset as a higher-yield target codon sequence for the codon sequence with the lower yield.70. The non-transitory computer readable storage medium of clause 69, further including using the lower-to-higher yield codon training dataset to train a high-yield codon sequence generator to map the lower-yield input codon sequences to the higher-yield target codon sequences.71. The non-transitory computer readable storage medium of clause 70, wherein the high-yield codon sequencegenerator is a neural network.72. The non-transitory computer readable storage medium of clause 71, wherein the neural network is a pre-trained neural network.Attorney Docket No. PRTN1008WO0173. The non-transitory computer readable storage medium of clause 71, wherein the neural network is trainedfrom scratch.74. The non-transitory computer readable storage medium of clause 71, wherein the neural network is asequence to sequence (seq2seq) neural network (e.g., RNNs like LSTMs and GRUs, and also Transformers).75. The non-transitory computer readable storage medium of clause 74, wherein the neural network is anencoder-decoder neural network (e.g., Transformer).76. The non-transitory computer readable storage medium of clause 69, wherein at least some of the clusters ofcodon sequences are generated based on clinical data that identifies different codons sequences that translate to a same amino acid sequence. In other implementations, at least some of the clusters of codon sequences are generated based on “labeled data” that can use transcriptomic labeling a proteomic labeling.77. The non-transitory computer readable storage medium of clause 69, wherein at least some of the clusters ofcodon sequences are generated based on an amino acid sequence-to-codon sequence generator generating different output codon sequences for a same input amino acid sequence.78. The non-transitory computer readable storage medium of clause 77, wherein the amino acid sequence-to-codon sequence generator is a neural network.79. The non-transitory computer readable storage medium of clause 78, wherein the neural network is a pre-trained neural network.80. The non-transitory computer readable storage medium of clause 78, wherein the neural network is anuntrained neural network.81. The non-transitory computer readable storage medium of clause 78, wherein the neural network is asequence to sequence (seq2seq) neural network.82. The non-transitory computer readable storage medium of clause 81, wherein the neural network is anencoder-decoder neural network (e.g., Transformer).83. A system, comprising:creating clusters of codon sequences on an amino acid sequence-basis, wherein a particular cluster of codon sequences created for a particular amino acid sequence includes different codon sequences that translate to the particular amino acid sequence but have varying yields / expression; and using the clusters of codon sequences to build a lower-to-higher yield codon training dataset that linksAttorney Docket No. PRTN1008WO01 lower-yield input codon sequences to higher-yield target codon sequences by: sorting codon sequences within each of the clusters of codon sequences by yield; based on the sorting, generating, on a cluster-basis, pairs of codon sequences in which one codon sequence has a lower yield and another codon sequence has a higher yield; including the codon sequence with the lower yield in the lower-to-higher yield codon training dataset as a lower-yield input codon sequence; and including the codon sequence with the higher yield in the lower-to-higher yield codon training dataset as a higher-yield target codon sequence for the codon sequence with the lower yield.84. The system of clause 83, further comprising using the lower-to-higher yield codon training dataset to traina high-yield codon sequence generator to map the lower-yield input codon sequences to the higher-yield target codon sequences.85. The system of clause 84, wherein the high-yield codon sequence generator is a neural network.86. The system of clause 85, wherein the neural network is a pre-trained neural network.87. The system of clause 85, wherein the neural network is trained from scratch.88. The system of clause 85, wherein the neural network is a sequence to sequence (seq2seq) neural network(e.g., RNNs like LSTMs and GRUs, and also Transformers).89. The system of clause 88, wherein the neural network is an encoder-decoder neural network (e.g.,Transformer).90. The system of clause 83, wherein at least some of the clusters of codon sequences are generated based onclinical data that identifies different codons sequences that translate to a same amino acid sequence. In other implementations, at least some of the clusters of codon sequences are generated based on “labeled data” that can use transcriptomic labeling a proteomic labeling.91. The system of clause 83, wherein at least some of the clusters of codon sequences are generated based onan amino acid sequence-to-codon sequence generator generating different output codon sequences for a same input amino acid sequence.92. The system of clause 91, wherein the amino acid sequence-to-codon sequence generator is a neuralnetwork.93. The system of clause 92, wherein the neural network is a pre-trained neural network.Attorney Docket No. PRTN1008WO0194. The system of clause 92, wherein the neural network is an untrained neural network.95. The system of clause 92, wherein the neural network is a sequence to sequence (seq2seq) neural network.96. The system of clause 95, wherein the neural network is an encoder-decoder neural network (e.g.,Transformer).97. A computer-implemented method of generating optimized codon sequences, including:processing an input sequence of amino acid elements; and based on the processing, generating an output sequence of codon elements, including confining sampling of the codon elements to map back to the amino acid elements.98. A system, comprising:processing an input sequence of amino acid elements; and based on the processing, generating an output sequence of codon elements, including confining sampling of the codon elements to map back to the amino acid elements.99. A non-transitory computer readable medium impressed with computer program instructions, theinstructions, when executed on a processor, implement a method of generating optimized codon sequences, including: processing an input sequence of amino acid elements; and based on the processing, generating an output sequence of codon elements, including confining sampling of the codon elements to map back to the amino acid elements.100. A computer-implemented method of generating high-yield codon sequences, including:a high-yield codon sequence generator trained on a lower-to-higher yield codon training dataset, and configured to generate output codon sequences in response to processing input codon sequences, wherein the output codon sequences have yields higher than that of the input codon sequences.101. A system, comprising:a high-yield codon sequence generator trained on a lower-to-higher yield codon training dataset, and configured to generate output codon sequences in response to processing input codon sequences, wherein the output codon sequences have yields higher than that of the input codon sequences.102. A non-transitory computer readable medium impressed with computer program instructions, theinstructions, when executed on a processor, implement a method of generating high-yield codon sequences, including: a high-yield codon sequence generator trained on a lower-to-higher yield codon training dataset, and configured to generate output codon sequences in response to processing input codon sequences, wherein theAttorney Docket No. PRTN1008WO01 output codon sequences have yields higher than that of the input codon sequences.103. A computer-implemented method of generating high-yield codon sequences, including:using a pre-trained protein language model to process amino acid sequences and generate codon sequence embeddings; using a pre-trained DNA language model to process codon sequences and generate amino acid sequence embeddings; using the amino acid sequence embeddings as inputs and the codons sequence embeddings as targets to train an amino acid sequence-to-codon sequence generator to generate clusters of variant codon sequences that translate to a same amino acid sequence; sorting the variant codon sequences by yield and using the sorted variant codon sequences to build a lower-to-higher yield codon training dataset that links lower-yield input codon sequences to higher-yield target codon sequences; and using the lower-to-higher yield codon training dataset to train a high-yield codon sequence generator to map the lower-yield input codon sequences to the higher-yield target codon sequences.104. A system, comprising:using a pre-trained protein language model to process amino acid sequences and generate codon sequence embeddings; using a pre-trained DNA language model to process codon sequences and generate amino acid sequence embeddings; using the amino acid sequence embeddings as inputs and the codons sequence embeddings as targets to train an amino acid sequence-to-codon sequence generator to generate clusters of variant codon sequences that translate to a same amino acid sequence; sorting the variant codon sequences by yield and using the sorted variant codon sequences to build a lower-to-higher yield codon training dataset that links lower-yield input codon sequences to higher-yield target codon sequences; and using the lower-to-higher yield codon training dataset to train a high-yield codon sequence generator to map the lower-yield input codon sequences to the higher-yield target codon sequences.105. A non-transitory computer readable medium impressed with computer program instructions, theinstructions, when executed on a processor, implement a method of generating high-yield codon sequences, including: using a pre-trained protein language model to process amino acid sequences and generate codon sequence embeddings;Attorney Docket No. PRTN1008WO01 using a pre-trained DNA language model to process codon sequences and generate amino acid sequence embeddings; using the amino acid sequence embeddings as inputs and the codons sequence embeddings as targets to train an amino acid sequence-to-codon sequence generator to generate clusters of variant codon sequences that translate to a same amino acid sequence; sorting the variant codon sequences by yield and using the sorted variant codon sequences to build a lower-to-higher yield codon training dataset that links lower-yield input codon sequences to higher-yield target codon sequences; and using the lower-to-higher yield codon training dataset to train a high-yield codon sequence generator to map the lower-yield input codon sequences to the higher-yield target codon sequences.

[0255] What is claimed is:

Claims

Attorney Docket No. PRTN1008WO01 CLAIMS 1. A computer-implemented method of providing a relative positional embedding to a multi-headattention transformer model, including: using respective attention heads of the multi-head attention transformer model to generate query, key, and value vectors for inputs tokens; converting the query and key vectors into position-encoded query and key vectors by applying a series of rotation matrices to the query and key vectors at different scaled frequencies that vary by absolute positions of the query and key vectors and by the respective attention heads, wherein the application of the series of rotation matrices includes: rotating pairs of feature dimensions in the query and key vectors by an angle in multiples of a scaled position index of a corresponding query or key vector, wherein the scaled position index is scaled by a head- specific scaling parameter that varies across the respective attention heads; and using the position-encoded query and key vectors for execution of self-attention by the respective attention heads to generate pairwise attention scores that depend on relative positions of input token pairs and on their feature similarity.

2. The computer-implemented method of claim 1, wherein the head-specific scaling parameter is a head-specific scaling scalar.

3. The computer-implemented method of claim 1, wherein the pairwise attention scores are penalizedbased on how far the position-encoded query and key vectors are.

4. The computer-implemented method of claim 3, wherein when a position-encoded query vector and aposition-encoded key vector are close by, the penalty is very low.

5. The computer-implemented method of claim 3, wherein when a position-encoded query vector and aposition-encoded key vector are far away, the penalty is very high.

6. The computer-implemented method of claim 1, wherein the input token pairs are amino acid tokenpairs.

7. The computer-implemented method of claim 1, wherein the input token pairs are nucleotide tokenpairs.

8. A computer-implemented method of generating optimized codon sequences, including:processing an input sequence of amino acid elements; and based on the processing, generating an output sequence of codon elements, including:Attorney Docket No. PRTN1008WO01 for a given amino acid element in the input sequence of amino acid elements, confining search of a corresponding codon element in a vocabulary of codon elements to a subset of codon elements known to translate to the given amino acid element.

9. The computer-implemented method of claim 8, further including, prior to the search, evaluating at theinput sequence of amino acid elements to determine an identity of the given amino acid element, and using the determined identity of the given amino acid element to confine the search of the corresponding codon element to the subset of codon elements known to translate to the given amino acid element.

10. The computer-implemented method of claim 8, wherein a neural network (e.g., AA2DNA) processesthe input sequence of amino acid elements, and generates the output sequence of codon elements.

11. The computer-implemented method of claim 8, wherein the neural network is a pre-trained neuralnetwork.

12. The computer-implemented method of claim 8, wherein the neural network is trained from scratch.

13. The computer-implemented method of claim 10, wherein the neural network is a sequence tosequence (seq2seq) neural network (e.g., RNNs like LSTMs and GRUs, and also Transformers).

14. The computer-implemented method of claim 10, wherein the neural network is an encoder-decoderneural network (e.g., Transformer).

15. The computer-implemented method of claim 14, wherein an encoder neural network generatesrespective embedded tokens for respective amino acid elements in the input sequence of amino acid elements.

16. The computer-implemented method of claim 15, wherein the encoder neural network appliesattention between the respective embedded tokens on an amino acid element-by-amino acid element.

17. The computer-implemented method of claim 16, wherein a decoder neural network receives results ofthe attention from the encoder neural network, and uses the results of the attention to generate the output sequence of codon elements.

18. The computer-implemented method of claim 17, wherein the decoder neural network looks back atthe input sequence of amino acid elements to determine the identity of the given amino acid element, and uses the determined identity of the given amino acid element to confine the search of the corresponding codon element to the subset of codon elements known to translate to the given amino acid element.

19. A computer-implemented method of generating high-yield codon sequences, including:creating clusters of codon sequences on an amino acid sequence-basis, wherein a particular cluster of codon sequences created for a particular amino acid sequence includes different codon sequences that translateAttorney Docket No. PRTN1008WO01 to the particular amino acid sequence but have varying yields / expression; and using the clusters of codon sequences to build a lower-to-higher yield codon training dataset that links lower-yield input codon sequences to higher-yield target codon sequences by: sorting codon sequences within each of the clusters of codon sequences by yield; based on the sorting, generating, on a cluster-basis, pairs of codon sequences in which one codon sequence has a lower yield and another codon sequence has a higher yield; including the codon sequence with the lower yield in the lower-to-higher yield codon training dataset as a lower-yield input codon sequence; and including the codon sequence with the higher yield in the lower-to-higher yield codon training dataset as a higher-yield target codon sequence for the codon sequence with the lower yield.

20. The computer-implemented method of claim 19, further including using the lower-to-higher yieldcodon training dataset to train a high-yield codon sequence generator to map the lower-yield input codon sequences to the higher-yield target codon sequences.

Citation Information

Patent Citations

  • Transform-based two-training image classification algorithm

    CN114528928A

  • A method, apparatus, computer device, and storage medium for optimizing codon sequences.

    CN115440300B