Machine-based methods for determining the toxicity of novel epitope payloads

By expressing multiple recombinant peptides in host cells and using machine learning to analyze toxicity metrics, the problem of novel epitope toxicity in the production of recombinant therapeutic viruses has been solved, enabling technical means to improve viral titer and host cell safety, and providing a personalized method for the production of therapeutic viruses.

CN115103917BActive Publication Date: 2026-03-10NANTOMICS LLC +1
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-24
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively predict and avoid the toxicity of new epitopes during the production of recombinant therapeutic viruses, leading to toxicity issues during virus production that affect viral titers and patient safety.

Method used

By expressing multiple recombinant peptides in host cells and combining machine learning methods, toxicity measures and sequence parameters are analyzed to optimize viral expression vectors to reduce toxicity. This includes using viral expression vectors, machine learning classifiers, and autoencoders for high-throughput sequencing and biomarker detection.

Benefits of technology

This enables the effective reduction or avoidance of novel epitope toxicity during recombinant virus production, improves viral titer and host cell safety, and provides a personalized therapeutic virus production method.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115103917B_ABST
    Figure CN115103917B_ABST
Patent Text Reader

Abstract

Systems and methods are provided that allow for the determination and prediction of payload toxicity in therapeutic viruses. This document discloses methods for determining the payload toxicity of peptides expressed in cells, comprising: generating or obtaining multiple expression vectors, each containing a different recombinant nucleic acid sequence encoding a corresponding recombinant peptide; expressing the recombinant nucleic acid sequence in multiple host cells while culturing these host cells; sequencing the multiple expression vectors after culturing the host cells; and correlating at least a portion of the recombinant nucleic acid sequence with a toxicity metric.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to our co-pending U.S. provisional patent application, filed August 9, 2019, serial number 62 / 885,089, which is incorporated herein by reference in its entirety.

[0002] sequence list

[0003] The contents of an ASCII text file containing a sequence list named 102402.0071PCT_ST25, which is 2KB in size, were created on July 23, 2019, and were submitted electronically with this application via EFS-Web, and are incorporated herein by reference in their entirety. Technical Field

[0004] This disclosure relates to various systems and methods for determining and / or avoiding the toxicity of recombinant viral payloads in host organisms, particularly when they involve the toxicity of novel epitopes in host cells used to produce therapeutic viruses. Background Technology

[0005] The background description includes information that may be useful in understanding this disclosure. No information provided herein is acknowledged to be prior art or related to the currently claimed invention, nor is any publication specifically or implicitly referenced considered prior art.

[0006] All publications and patent applications herein are incorporated by reference to the same extent that each individual publication or patent application is specifically and individually indicated to be incorporated by reference. Where a definition or usage of a term in an incorporated reference is inconsistent with or contrary to the definition of that term provided herein, the definition provided herein shall apply, and not the definition in that reference.

[0007] The production of recombinant therapeutic viral vaccines has become an increasingly attractive strategy for treating a variety of diseases, particularly for viruses used in the preparation of cancer vaccines. Unfortunately, despite rapid progress in the identification and selection of potentially immunogenic novel epitope sequences, the toxicity of one or more expressed novel epitopes often only becomes apparent after the production of a therapeutic recombinant virus and the commencement of large-scale virus production. To avoid at least some of the drawbacks associated with potential payload toxicity, the expression of the recombinant payload can be suppressed in production cells in various ways, as described in PCT / US 2018 / 054982. This approach would advantageously achieve appropriately high viral titers in the production setting. However, once patient cells are infected with the recombinant therapeutic virus, the payload toxicity in patient cells can reduce the expression of one or more novel epitopes. Among other known methods, protein toxicity can be determined using predictive algorithms that identify potentially toxic sequences in a protein based on the known toxicity of known proteins (see PLoS ONE 8(9): e73957). While conceptually appealing, this approach is based on naturally occurring peptides and is generally not suitable for artificial sequence constructs (e.g., sequences encoding multiple novel epitopes linked by adapter sequences and optionally containing transport signals).

[0008] Therefore, while various methods for reducing the toxicity of recombinant viral payloads are known in the art, all or almost all of them have various drawbacks. Therefore, there is a need to provide improved compositions and methods that allow the production of recombinant therapeutic viruses with reduced toxicity. Summary of the Invention

[0009] Various systems and methods are provided that allow for the determination of payload toxicity in recombinant therapeutic viruses. In one aspect of the subject matter, the inventors have contemplated a method for determining the payload toxicity of a polypeptide expressed in cells, the method comprising the steps of generating or obtaining a plurality of expression vectors, each expression vector containing a different recombinant nucleic acid sequence encoding a corresponding recombinant polypeptide; expressing the recombinant nucleic acid sequence in a plurality of host cells while culturing the host cells; sequencing the plurality of expression vectors after culturing the host cells; and correlating at least a portion of the recombinant nucleic acid sequence with a toxicity metric.

[0010] In at least some embodiments, the expression vector is a viral expression vector, particularly a recombinant genome of a corresponding therapeutic virus. Further consideration is given to the recombinant polypeptide being a polytope containing multiple neoantigens, typically separated by linker peptides. Preferably, these neoantigens have a length of 8 to 50 amino acids, and / or the polytope contains at least 200 amino acids.

[0011] It should also be understood that the recombinant nucleic acid sequence can be expressed monoclonally or polyclonally in the multiple host cells. Therefore, the multiple expression vectors can be sequenced individually or in a mixture of expression vectors. In another aspect of the considered approach, virulence measures (e.g., as cell death, cell stress, reduced cell division, and / or decreased viral yield) are observed in the host cells, while in other aspects, virulence measures (e.g., as nonsense mutations, missense mutations, and / or deletions) are observed in the recombinant nucleic acid sequence of the virus.

[0012] Additionally, the association step utilizes machine learning, which can employ various classifiers such as linear classifiers, NMF-based classifiers, graph-based classifiers, tree-based classifiers, Bayesian-based classifiers, rule-based classifiers, network-based classifiers, or kNN classifiers. Alternatively, the machine learning can also use an autoencoder. If desired, the machine learning can further utilize secondary aspects of the recombinant peptide, such as the peptide's folding pattern, secondary structure, polar domains, charged domains, hydrophobic domains, hydrophilic domains, and / or the peptide's aggregation.

[0013] Various objects, features, aspects and advantages will become more apparent from the following detailed description of preferred embodiments and from the accompanying drawings, in which the same numerals denote the same components. Attached Figure Description

[0014] Figure 1 Exemplary assays of cellular stress induced by various payloads, measured by qPCR, are depicted.

[0015] Figure 2 Exemplary measurements of cellular stress induced by various payloads, determined by XBP1 cleavage, are depicted.

[0016] Figure 3 Exemplary assays of cellular stress induced by various payloads, measured by Western blotting, are presented. Detailed Implementation

[0017] The inventors have now discovered that a rationale-based method for determining payload toxicity can be employed, wherein multiple payload sequences of the corresponding virus are associated with one or more toxicity measures in the host cell that produced the virus, preferably using machine learning methods.

[0018] Therefore, and in a more general aspect of the subject matter of this invention, the inventors have considered expressing multiple viral payloads in the same host cell line (and its corresponding culture) to at least partially generate viral progeny. The cell and / or viral cultures are then analyzed based on the type of virulence metric (e.g., cellular stress, apoptosis, host cell growth arrest, mutations in the payload (e.g., nonsense, missense), decrease in viral titer at a predetermined culture time, increase in production time to the target titer, etc.). It should be understood, of course, that the analysis can be performed on an individual / clonal basis or in large-scale parallel use using a mixed (virus and / or host cell) clonal population. The results of the payload sequence analysis are then processed using machine learning that correlates one or more virulence metrics with one or more payload sequence parameters (e.g., charge and / or hydrophobicity patterns, specific amino acid usage or patterns, structural motifs or folding patterns, etc.). Most typically, payload sequence parameters are analyzed on more than one novel epitope (such as multiple epitopes or a single translational unit) within a single payload.

[0019] The acquisition or generation of clonal diversity of multiple viruses with corresponding payloads can be based on a variety of materials, and in particular includes novel epitope sequences from patients obtained from a variety of publicly available sources (e.g., Genomics Proteomics & Bioinformatics 16(2018)276-282; or WO 2016 / 172722), or de novo determined novel antigen sequences derived from previously unpublished patient or TCGA data using a variety of methods known in the art (e.g., Science. 2015; 348:69-74; or J Clin Invest. 2015; 125:3413-3421; or R Soc Open Sci. 2017; 4:170050; or R Soc Open Sci. 2017; 4:170050). Various bioinformatics tools can be used to further refine such data to predict MHC binding, with NetMHC 4.0 being a particularly well-known tool.

[0020] Most typically, the neoantigen (also known as a neoepitope) in the considered method is arranged in a recombinant multiepitope sequence, preferably having an intervening flexible linker sequence. Furthermore, the considered multiepitope sequence may further include a transport sequence to direct the recombinant protein to a specific subcellular location (e.g., cytoplasm, lysosome, endosome, etc.). Ubiquitination signals may also be included if desired. Exemplary suitable sequence arrangements are described in WO 2017 / 222619. In this context, it should be understood that when a neoantigen is present and expressed in multiple epitopes, toxicity metric may involve a single neoantigen or a polypeptide containing more than one neoantigen. From a different perspective, it is possible to consider two or more otherwise non-toxic neoantigens that may be toxic to host cells when such neoantigens form multiple epitopes. Such compound toxicity is undetectable when analyzing the individual antigens themselves. As can be easily understood, a neoantigen, and more preferably a multi-epitope containing the neoantigen, will be expressed from an expression vector, which may further include additional functionalities (e.g., co-stimulatory molecules, cytokines, ALT-803, TxM-type molecules, checkpoint inhibitors, etc.).

[0021] While most expression vectors are considered suitable for this document, it is particularly preferred that neoantigens or multiple epitopes be expressed from a recombinant viral genome using suitable control elements known in the art. Using such recombinant viruses in the methods provided herein offers at least two advantages, including downstream use of such viruses in the production of therapeutic viruses and assessment of potential toxicity in the event of viral replication. Therefore, the host cells used to assess toxicity will have a suitable conformation allowing viral infection. For example, when the recombinant virus is an AdV adenovirus lacking the E2b protein, the host cells considered will express CXADR (Coxsackievirus and Adenovirus Receptor) (naturally or from the recombinant nucleic acid). Exemplary host cells for adenovirus-based systems include E.C7 cells (commercially available from Etubics) and those described in WO 2009 / 006479 and WO2017 / 136748. Viruses further considered suitable as recombinant expression vectors for therapeutic antigens include various adenoviruses, adeno-associated viruses, alphaviruses, herpesviruses, lentiviruses, etc. However, adenoviruses are particularly preferred. Furthermore, it is even more preferred that the virus is a replication-defective, non-immunogenic virus, typically achieved by targeting a deleted viral protein (e.g., E1, E3 proteins). Such desired properties can be further enhanced by deleting the function of the E2b gene, and, as recently reported (e.g., J Virol. [Journal of Virology] 1998 Feb; 72(2): 926-933), high-titer recombinant viruses can be obtained using genetically modified human 293 cells.

[0022] Regarding payload toxicity, it should be understood that toxicity can affect host cells (i.e., infected or otherwise transfected) and the virus in a variety of ways. For example, expressed multiepitopes or portions thereof (e.g., one or more neoantigens or neoantigen-connector portions) may be directly toxic to cells and interfere with metabolism, cell division, or cell signaling. On the other hand, expressed multiepitopes or portions thereof can also be indirectly toxic and can affect a variety of intracellular processes and structures, such as transcription, translation, protein conversion, energy production, and the membrane integrity, nuclear and / or mitochondrial stability of various organelles. Furthermore, it should be noted that expressed multiepitopes or portions thereof can exert unfavorable selective pressure on cells and can therefore indirectly lead to mutations in the nucleic acids encoding the expressed multiepitopes or portions thereof. Thus, toxicity may also lead to the production of mutated recombinant (viral) nucleic acids, in which the mutated nucleic acids will have premature stop codons and / or missense mutations that reduce unfavorable selective pressure. Therefore, from different perspectives, toxicity can lead to cell death (usually through apoptosis or necrosis), reduced or otherwise impaired cell division, cellular stress (and usually associated reductions in metabolism and (viral) replication), mutations in recombinant payload, decreases in viral titer at predetermined culture time, and / or increased production time to a predetermined target titer.

[0023] In further consideration, toxicity can be measured in vivo using various alternative measures that can be directly or indirectly observed in host cells. For example, one or more biomarkers in host cells associated with apoptosis or cellular stress can be quantified. As shown in more detail below, the upregulation of ER stress markers (e.g., BiP / Grp78, XBP-1 cleavage) and the inhibition of CHOP-induced apoptosis associated with host cell survival can be measured. Furthermore, it should be recognized that cellular stress can also be identified and even quantified using computational omics methods, where stress-related transcription factors (e.g., XBP-1) activate the expression of recombinant biomarker molecules (e.g., GFP).

[0024] Therefore, depending on the observed type of toxicity, payload expression in host cells can be carried out either monoclonally or in mixed cultures. For example, when the payload is a multi-epitope containing actual patient neoantigens and the payload is already present in a therapeutic virus, payload expression is typically carried out monoclonally (i.e., infecting host cells with a single clone (genotype) of the therapeutic virus and culturing the so-called infected cells to the desired cell density and / or viral titer). On the other hand, when the payload is an exploratory payload (i.e., not for therapeutic use), multiple recombinant viruses with diverse libraries based on the same multi-epitopes can be used to transfect multiple host cells in polyclonal virus cultures, as described in more detail below.

[0025] Regardless of the virulence type of the payload, sequence analysis of recombinant viral nucleic acids (or other expression vectors) can be performed in a variety of ways well known in the art, and the type of payload and / or observed virulence will at least partially determine the type of sequencing used. For example, in cases where the payload is present in a therapeutic virus and the virus replicates in a monoclonal manner, sequence analysis can be performed from the viral isolate. On the other hand, in cases where multiple viruses replicate in a polyclonal virus culture, sequencing can be performed using the collective nucleic acid as a whole without prior clonal selection of individual viruses. Of course, it should be understood that all sequencing methods are preferably automated sequencing methods that allow high data throughput, such as NextGen / Illumina sequencing and other massively parallel sequencing methods. In this context, it should be recognized that in the case of sequencing mixed viral nucleic acids (e.g., such as those obtained from polyclonal virus cultures), sequence analysis will employ methods that provide "allele fraction" or "purity / mutant fraction" of specific base positions in the nucleic acid encoding novel antigens and / or novel epitopes. Exemplary suitable methods are described in our co-pending U.S. provisional applications, serial numbers 62 / 714,570 (PANBAM: BAMBAM Across Multiple Organisms InParallel) and 62 / 681,800 (Difference-Based Genomic Identity Scores), both of which are incorporated by reference.

[0026] Furthermore, it should be noted that sequence analysis can be performed at multiple points in the cell culture process, thereby helping to identify the occurrence and fraction of mutations (in one or all viral genomes) over time. Therefore, it should be understood that sequence analysis will provide not only qualitative information about mutations in a virus or viral population, but also quantitative and temporal information about mutations in a virus or viral population. For example, in the case where cell cultures are used to propagate monoclonal viral populations (e.g., for therapeutic viruses), viral samples can be removed at predetermined intervals to reveal the occurrence and fraction of viral mutants over time after sequencing. On the other hand, in the case where cell cultures are used to propagate polyclonal viral populations (e.g., libraries based on mutant sequences), viral samples can be removed at predetermined intervals to reveal the dynamic opportunities of selected viral mutants over time after sequencing.

[0027] Depending on the observed toxicity measure and mutation type, various machine learning algorithms can be employed to correlate one or more motifs in the payload sequence (e.g., domains, one or more amino acids at a specific position, sequence length, amino acid composition, predicted folding, etc.) with the observed toxicity. It will be readily understood that a variety of classifier types can be chosen, and suitable classifiers include one or more of the following: linear classifiers, NMF-based classifiers, graph-based classifiers, tree-based classifiers, Bayesian classifiers, rule-based classifiers, network-based classifiers, kNN classifiers, or other types of classifiers. More specific examples include NMF predictors (linear), SVMlight (linear), SVMlight first-order polynomial kernel (d-degree polynomial), SVMlight second-order polynomial kernel (d-degree polynomial), WEKA SMO (linear), WEKA i48 tree (tree-based), WEKA hyperpipe (distribution-based), WEKA random forest (tree-based), WEKA Naive Bayes (probabilistic / Bayesian), WEKA JRip (rule-based), glmnet lasso (sparse linear), glmnet ridge regression (sparse linear), glmnet elastic net (sparse linear), artificial neural networks (e.g., ANN, RNN, CNN, etc.), and so on. Other sources for the prediction model template 140 include Microsoft's CNTK (see URL github.com / Microsoft / cntk), TensorFlow (see URL www.tensorflow.com), PyBrain (see URL pybrain.org), or other sources.

[0028] Alternatively, particularly when the number of available toxicity examples is relatively small, the inventors considered using an encoder trained on the MHC-peptide binding problem to obtain representations of example novel epitopes, and from there training a toxicity classifier specific to the production cell line. While this approach may not generalize well and may produce errors, at least initially, human supervision can be employed to label instances where predicted toxicities prove incorrect and add them to the training set. With such intervention, the system accuracy should improve rapidly and eventually generalize well.

[0029] Once a data threshold is reached, machine learning can also utilize methods employing autoencoders (see, for example, arXiv: 1610.02415v3), which allow the transformation of polyepitopes into a continuous latent space and then back to the polyepitopes from the latent space. To constrain the structure of the latent representation, predictors for various molecular properties can be jointly trained. One advantage of any encoder / decoder pair is the ability to perturb points in the latent space or interpolate between points, subsequently passing a new representation through the decoder, in this case, sampling all possible polyepitopes. However, because the latent representation is jointly learned with the task of predicting polyepitope properties, it should be noted that the latent space also becomes more suitable for optimizing polyepitopes for desired properties. In other words, points in the latent space can be shifted using gradients from the trained predictors, such that they will lead to more or less the desired properties.

[0030] Regarding toxicity and MHC binding, it should be noted that if one of the co-trained properties predicted from the potential representation of the peptide is toxic to a particular production cell, then once a candidate epitope is available, the gradient in the latent space can be followed to minimize toxicity while attempting to maintain fidelity to the original candidate. If binding to different MHC alleles is also predicted from the same latent space representing the peptide, then theoretically, it would be possible to perform parallel optimization to maximize predicted binding in the allele of interest and minimize toxicity to select a production method.

[0031] Note that multiple parallel predictions can be made for toxicity across multiple cell lines or production processes (assuming sufficient data for training on each). Furthermore, when optimizing for toxicity, the toxicity of one or more production methods can be used as a constraint. Additionally, it should be noted that the types of peptide modifications allowed are limited only by the design choices of the models used for the encoder / decoder. Therefore, models capable of handling variable lengths in both input and output (such as fully convolutional networks or RNNs) can allow for variations in peptide length and amino acid substitutions.

[0032] Therefore, it should be understood that toxicity parameters (especially toxicity thresholds) can be determined based on observed toxicity and knowledge of the payload sequence. Once established, known payloads can be eliminated or reconstructed to reduce or completely avoid toxicity to host cells.

[0033] Example

[0034] Determination of viral payload toxicity and related observed mutationsIn the following embodiments, a payload was constructed and cloned into an AdV virus lacking the E2b gene, and the virus was propagated in E.C7 cells. Virulence was observed and genetic changes in the viral payload were detected as described. The length of the multiepithels varied between approximately 1.1 Kb and 11.2 Kb, and further included ubiquitination signals, co-stimulatory molecules, and transport signals, as shown in the table below.

[0035]

[0036]

[0037] As can be seen from the table, virulence payloads lead to deletions, point mutations, and nonsense mutations in the viral payload, as well as slower production of viral particles to the predetermined titer. Furthermore, it should be noted that virulence can be associated with the payload sequence and accompanying changes within the payload sequence.

[0038] Model biomarkers for detecting payload toxicity: In this exemplary system, E.C7 cells were treated with 1 μM thiapsigargin or transfected with pShuttle plasmid using Lipofectamine 3000. Reverse transcription and cDNA synthesis were performed according to the manufacturer's protocol using RNeasy (Qiagen) and a high-capacity cDNA synthesis kit (Applied Biosystems). Relative mRNA expression was calculated by normalizing the samples to the internal control RPL19. Expression was quantified by qPCR following rtPCR using the following primers:

[0039]

[0040] Figure 1-3 Exemplary results of this model system are described. More specifically, Figure 1 The selected biomarkers (top) and the expression vector carrying the indicated payload are shown in the figure below when cells are treated with toxic beta-carotene as a positive control. Figure 2 Exemplary results of XBP1 cutting are depicted, and Figure 3The results of Western blotting are described. E.C7 cells were treated with 1 μg / mL tunicamycin or transfected with pShuttle plasmid using Lipofectamine 3000. Total protein lysates were extracted using RIPA buffer supplemented with protease inhibitors (20 mM Tris-HCl pH 7.5, 150 mM NaCl, 1 mM Na2EDTA, 1 mM EGTA, 1% NP-40, 1% sodium deoxycholate). Lysates were probed at 1:1000 dilution with BiP (CST#3177), CHOP (CST#2895), and GAPDH (CST#2118) antibodies.

[0041] Polyclonal virus culture and sequencing: Starting with a single clone of a therapeutic virus encoding multiple epitopes of 20 neoantigens separated by flexible spacers, a diverse library is constructed in the Adv virus, where each clone will have at least one random mutation at at least one amino acid position. The first sample of the preserved library is used for sequencing. The viral expression library is then propagated in E.C7 cells, and viral samples are retrieved at different time points (e.g., 6 hours, 12 hours, 18 hours, 24 hours, etc.), and the final viral sample is retrieved at the end of viral production. Nucleic acids are then isolated from each sample, resulting in a mixed nucleic acid population representing the library members. The prepared nucleic acids are then sequenced, and the sequencing data are analyzed, preferably using simultaneous incremental alignment, for example, as described in our co-pending patent applications with publication numbers WO 2020 / 028862 (PANBAM: BAMBAM Across Multiple Organisms In Parallel) and WO 2019 / 236842 (Difference-Based Genomic Identity Scores). The base call score for each base position is then determined, and population variations can be identified. For example, in the case of a single viral clone having a lower replication rate due to a specific base at a particular position, the allele score of that base will decrease over time. Similarly, in the case of a single viral clone having a higher replication rate due to a specific base at a particular position (e.g., resulting in reduced virulence), the allele score of that base will increase over time.

[0042] Of course, it should be understood that the analysis is not necessarily limited to the observation of specific bases and direct toxicity, but can also include secondary analyses. For example, changes in a single amino acid can lead to different spatial conformations (folding), changes in net charge, changes in secondary structure, changes in lipophilicity, etc., and all such changes can be included in any machine learning algorithm. Therefore, from different perspectives, it should be understood that one or more toxicity parameters (e.g., reduced host cell growth, increased stress response in host cells, host cell death, reduced or slowed viral production in host cells, mutations in viral nucleic acids, especially mutations in recombinant payloads (e.g., deletions, nonsense, or missense mutations), decreased viral titers, etc.) can be associated not only with the linear peptide sequence, but also with secondary aspects of that linear peptide sequence. Most typically, such secondary aspects include the folding pattern and / or misfolding of the expressed peptide, the specific secondary structure of the expressed peptide, the polar domains of the expressed peptide, charge, hydrophobicity, hydrophilicity and / or aggregation, the specific length of the expressed peptide, etc.

[0043] As used herein, the term "administration" of a pharmaceutical composition or drug refers to both direct and indirect administration of the pharmaceutical composition or drug, wherein direct administration of the pharmaceutical composition or drug is typically performed by a healthcare professional (e.g., physician, nurse, etc.), and wherein indirect administration includes the steps of providing or making available to a healthcare professional the pharmaceutical composition or drug for direct administration (e.g., via injection, infusion, oral delivery, local delivery, etc.). Most preferably, cells or exosomes are administered via subcutaneous or subdermal injection. However, in other considerations, administration may also be intravenous injection. Alternatively or additionally, antigen-presenting cells may be isolated from or grown from the patient's cells, infected in vitro, and then infused into the patient. Therefore, it should be understood that the systems and methods considered can be viewed as a complete drug discovery system (e.g., drug discovery, treatment regimens, validation, etc.) for highly personalized cancer treatment.

[0044] The description of ranges of values ​​herein is intended only as a shorthand method of individually referring to each individual value falling within that range. Unless otherwise stated herein, each individual value is incorporated into the specification as if it were individually referenced herein. All methods described herein can be performed in any suitable order unless otherwise indicated herein or otherwise obviously contradicted by the context. The application of any and all instances or exemplary language (such as “e.g.”) provided with respect to certain embodiments herein is intended only to better illustrate the full scope of this disclosure and not to limit the scope of the otherwise claimed invention. No language in this specification should be construed as indicating that any unclaimed element is necessary for practicing the claimed invention.

[0045] Those skilled in the art will understand that many other modifications can be made beyond those already described without departing from the full scope of the concepts disclosed herein. Therefore, the subject matter of this disclosure is limited only to the scope of the appended claims. Furthermore, in interpreting this specification and the claims, all terms should be interpreted in the broadest possible way, consistent with the context. In particular, the terms “comprises” and “comprising” should be interpreted as referring to an element, component, or step in a non-exclusive manner, indicating that the mentioned element, component, or step may be present, used, or combined with other elements, components, or steps not expressly mentioned. Where a claim in this specification refers to at least one of the items selected from the group consisting of A, B, C, ..., and N, the text should be interpreted as requiring only one element from that group, rather than A plus N or B plus N, etc.

Claims

1. A method of determining viral payload toxicity of a polypeptide expressed in a cell, wherein the toxicity results from a viral payload, the method comprising: producing or obtaining a plurality of expression vectors, each expression vector comprising a different recombinant nucleic acid sequence encoding a corresponding recombinant polypeptide; expressing the recombinant nucleic acid sequence in a plurality of corresponding host cells while culturing the plurality of corresponding host cells; sequencing the plurality of expression vectors after culturing the host cells; associating at least a portion of the recombinant nucleic acid sequence with a measure of toxicity; and wherein the measure of toxicity is cell death, cell stress, reduced cell division, or reduced viral yield of the host cell.

2. The method of claim 1, wherein the expression vectors are viral expression vectors.

3. The method of claim 1, wherein the expression vectors are recombinant genomes of corresponding therapeutic viruses.

4. The method of claim 1, wherein the recombinant polypeptide is a polytope comprising a plurality of neoantigens.

5. The method of claim 4, wherein at least two of the neoantigens are separated by a linker peptide.

6. The method of claim 4, wherein the neoantigens have a length of 8 to 50 amino acids.

7. The method of claim 4, wherein the polytope has at least 200 amino acids.

8. The method of claim 1, wherein the recombinant nucleic acid sequence is monoclonally expressed in the plurality of host cells.

9. The method of claim 1, wherein the recombinant nucleic acid sequence is polyclonally expressed in the plurality of host cells.

10. The method of claim 1, wherein the plurality of expression vectors are sequenced individually.

11. The method of claim 1, wherein the plurality of expression vectors are sequenced in a mixture of expression vectors.

12. The method of claim 1, wherein the measure of toxicity is observed in the recombinant nucleic acid sequence of the virus.

13. The method of claim 12, wherein the measure of toxicity in the recombinant nucleic acid sequence of the virus is a nonsense mutation, a missense mutation, and a deletion.

14. The method of claim 1, wherein the associating step uses machine learning.

15. The method of claim 14, wherein the machine learning uses a classifier selected from the group consisting of a linear classifier, an NMF-based classifier, a graph-based classifier, a tree-based classifier, a Bayesian-based classifier, a rule-based classifier, a web-based classifier, and a kNN classifier.

16. The method of claim 14, wherein the machine learning uses an autoencoder.

17. The method of claim 14, wherein the machine learning uses a secondary aspect of the recombinant polypeptide.

18. The method of claim 17, wherein the secondary aspect is a folding pattern of the polypeptide, a secondary structure of the polypeptide, a polar domain, a charged domain, a hydrophobic domain, a hydrophilic domain, and / or aggregation of the polypeptide.

Citation Information

Patent Citations

  • Methods and compositions for producing an adenovirus vector for use with multiple vaccinations

    WO2009006479A2

  • Cancer neoepitopes

    WO2016172722A1

  • Compositions and methods for recombinant cxadr expression

    WO2017136748A1

  • Sequence arrangements and sequences for neoepitope presentation

    WO2017222619A2

  • Difference-based genomic identity scores

    WO2019236842A1