Systems and methods for engineering synthetic antigens to promote adaptive immune responses
In silico designed engineered antigens disrupt conserved regions to reduce memory immune responses, enhancing vaccination efficacy against mutating pathogens by promoting the production of novel antibodies.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-23
- Publication Date
- 2026-03-10
AI Technical Summary
Existing vaccination techniques struggle to keep up with constantly evolving circulating pathogen variants and newly emerging diseases, leading to ineffective immunogenic compositions that fail to reduce infection risk and disease severity.
In silico design of custom engineered antigens that disrupt conserved regions of reference antigens to reduce memory immune responses, promoting the production of novel antibodies tailored to specific epitopes through amino acid modifications.
The engineered antigens enhance immunogenic compositions by reducing memory immune responses and promoting the production of new neutralizing antibodies, offering improved efficacy against mutating viral agents.
Smart Images

Figure 2026508268000001_ABST
Abstract
Description
[Technical Field]
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS This application claims priority to and the benefit of U.S. Provisional Application Nos. 63 / 448,217 (filed February 24, 2023), 63 / 448,215 (filed February 24, 2023), 63 / 448,987 (filed February 28, 2023), 63 / 449,031 (filed February 28, 2023), 63 / 449,936 (filed March 3, 2023), 63 / 452,989 (filed March 17, 2023), 63 / 452,987 (filed March 17, 2023), 63 / 514,242 (filed July 18, 2023), and 63 / 514221 (filed July 18, 2023), the contents of each of which are incorporated herein by reference in their entirety. [Background technology]
[0002] Vaccination can play a vital role in managing and ensuring public health. When vaccinated with a sufficiently effective immunogenic composition against a particular infectious agent, individuals experience a reduced risk of infection and / or a reduced severity of disease if infected. Therefore, the development of highly effective vaccination techniques that can keep up with constantly evolving circulating pathogen variants and newly emerging diseases is a significant challenge. Summary of the Invention [Means for solving the problem]
[0003] Presented herein is a technology directed to the in silico design of custom engineered antigens. In particular, in certain embodiments, the disclosed methods and systems provide for engineering antigens to reduce activation of memory immune responses (such as B cell and / or T cell-based responses) when introduced into a subject. By engineering antigens in this manner, their performance as immunogenic compositions, for example, for vaccination purposes, can be improved. For example, without wishing to be bound by any particular theory, it is believed that, among other things, reducing the extent to which memory immune responses are elicited can lead to improved production of novel antibodies that are selectively adapted by the subject's immune system to neutralize specific (e.g., evolved) epitopes of the reference antigen.
[0004] For example, in certain embodiments, the disclosed systems and methods identify conserved regions within a computer representation of a reference antigen that are similar to regions in other (e.g., previously circulating) variants of the reference antigen and therefore likely to elicit a memory response. The approaches described herein then disrupt these conserved region(s), for example, by introducing amino acid modifications into or throughout them. In this manner, engineered antigens can be generated that retain certain portions of the reference antigen, such as specific target epitopes, but replace the conserved regions with disrupted versions. While not wishing to be bound by any particular theory, it is believed that engineered antigens with disrupted conserved region(s), when produced and introduced into a subject, are less likely to elicit a memory immune response (e.g., from memory B cells or T cells) and instead promote a naive response, thereby facilitating the production of new neutralizing antibodies specifically tailored to the retained target epitopes. Therefore, such engineered antigens may offer improved efficacy when used as immunogenic compositions, particularly against viral infectious agents prone to mutation.
[0005] In one aspect, the present disclosure provides a method for the in silico design of an engineered antigen (e.g., to elicit an immune response directed to one or more target epitopes of a reference antigen of an infectious agent, while the engineered antigen reduces (e.g., characterized by) activation of a memory immune response (e.g., of B cells and / or T cells) against the reference antigen (e.g., compared to the reference antigen). The method includes: (a) receiving and / or accessing, by a processor of a computing device, a polypeptide model representing the reference antigen of the infectious agent (e.g., as its sequence and / or a 3D structural model); (b) identifying, by the processor, within the polypeptide model, one or more memory-inducing conserved region(s) representing conserved portions of the reference antigen determined to be likely to elicit a memory immune response (e.g., portions of the reference antigen determined to (i) correspond to known epitopes and / or potential epitopes and / or (ii) not include a mutation characteristic of the reference antigen); and (c) identifying, by the processor, one or more (memory-inducing) conserved regions representing the one or more (memory-inducing) conserved regions. (d) generating one or more amino acid modifications within at least a portion of the conserved region(s) (e.g., the one or more amino acid modifications are one or more point modifications, insertions, and / or deletions), thereby creating a disrupted polypeptide model representing the engineered antigen (e.g., representing an artificially engineered version of the reference antigen in which mutations have been introduced to disrupt conserved portions of the reference antigen (e.g., at least a portion of one or more memory-inducing conserved regions, e.g., portions determined / determined to be likely to induce a memory immune response)); and (d) displaying and / or storing and / or providing, by a processor, the disrupted polypeptide model for further processing.
[0006] In some embodiments, the reference antigen is or comprises at least a portion of a naturally occurring variant of a viral protein (e.g., the infectious agent is a variant of a particular virus (e.g., influenza virus, coronavirus, respiratory syncytial virus, filovirus) and the reference antigen is or comprises at least a portion of that protein).
[0007] In some embodiments, the reference antigen (e.g., and / or viral protein) is a SARS-Cov2 spike polypeptide or comprises at least a portion thereof (e.g., a receptor binding domain (RBD), e.g., an N-terminal region, e.g., substantially all (e.g., the entire) of the spike protein) (e.g., a portion selected to focus on minimally relevant vaccine antigens, e.g., to facilitate removal of as many conserved epitopes as possible, without resorting to introducing point mutations (e.g., thereby limiting the number of epitopes into which point mutations are introduced)).
[0008] In some embodiments, the reference antigen is, or comprises at least a portion of, a particular SARS-CoV-2 variant spike polypeptide.
[0009] In some embodiments, the particular SARS-CoV-2 variant is a member of the Omicron and / or XBB phylogenetic classification (e.g., according to classifications such as WHO, Pango, Nextstrain, etc.) (e.g., the particular SARS-CoV-2 variant is XBB.1.5; e.g., the particular SARS-CoV-2 variant is JN.1).
[0010] In some embodiments, the infectious agent is or comprises (e.g., a particular variant thereof) an RNA virus (e.g., a virus that encodes genetic information in RNA), and the reference antigen is or comprises at least a portion of a protein thereof.
[0011] In some embodiments, the reference antigen is or comprises a bacterial protein (e.g., the infectious agent is a bacterium and the reference antigen is or comprises at least a portion of a particular protein thereof).
[0012] In some embodiments, the reference antigen is or comprises an antigen (e.g., a surface antigen, e.g., a protein) of a parasite (e.g., the infectious agent is a parasite and the reference antigen is or comprises at least a portion of a particular protein thereof) (e.g., the infectious agent is a malaria parasite and the reference antigen is an antigen thereof).
[0013] In some embodiments, the one or more memory-inducing conserved region(s) represent a portion(s) of the reference antigen that is substantially similar to (i) one or more (e.g., pre-existing) variants thereof and / or (ii) the initial / wild-type strain (e.g., the originally observed strain) (e.g., the one or more memory-inducing conserved region(s) represent a portion(s) of the reference antigen that has sufficient sequence similarity (e.g., at least 80% (e.g., including at least 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more) identical, e.g., identical) to the corresponding portion of (i) one or more (e.g., pre-existing) variants thereof and / or (ii) the initial / wild-type strain (e.g., the originally observed strain)).
[0014] In some embodiments, the reference antigen is a particular target SARS-CoV-2 variant (e.g., XBB.1.5; e.g., JN.1) S polypeptide or portion thereof (e.g., the reference antigen is the RBD of the target SARS-CoV-2 S protein, and the one or more memory-inducing conserved region(s) represent a portion(s) of the reference antigen that is substantially similar to a corresponding portion(s) of (i) one or more other (e.g., pre-existing) SARS-CoV-2 variant polypeptides and / or (ii) a Wuhan SARS-CoV-2 polypeptide (e.g., the one or more memory-inducing conserved region(s) represent a non-mutated portion(s) of the reference antigen that is common to the reference antigen and a corresponding portion(s) of (i) one or more (e.g., pre-existing) SARS-CoV-2 variant polypeptide(s) and / or (ii) a Wuhan SARS-CoV-2 polypeptide).
[0015] In some embodiments, the one or more memory-inducing conserved region(s) is or includes a set of conserved epitope regions that represent known epitopes that are present and that have not mutated on a reference antigen (e.g., known epitopes that do not have any characteristic mutations) (e.g., the reference antigen is a specific subregion (e.g., RBD) of a target SARS-CoV-2 variant S protein (e.g., XBB.1.5; e.g., JN.1), and the set of conserved epitope regions represents known epitopes that are present and that have not mutated on a specific subregion of the target SARS-CoV-2 variant S protein compared to corresponding subregions of one or more pre-existing variants and / or Wuhan strain SARS-CoV-2 S proteins).
[0016] In some embodiments, step (b) comprises obtaining, by a processor, data corresponding to the set of known epitopes and identifying, within the reference antigen, each of one or more specific known epitopes of the set; obtaining, by the processor, an identification of a set of characteristic mutations of the reference antigen; and identifying, by the processor, the specific known epitopes corresponding to portions of the reference antigen that do not have any characteristic mutations as the set of conserved epitope regions.
[0017] In some embodiments, the set of known epitopes includes one or more of the epitopes listed in Table 2A.
[0018] In some embodiments, the set of known epitopes includes one or more of the epitopes listed in Table 2B.
[0019] In some embodiments, the one or more memory-inducing conserved region(s) is or includes a conserved surface that represents a (e.g., contiguous, e.g., adjacent), unmutated (e.g., lacking any characteristic mutations) surface of a reference antigen (e.g., the reference antigen is SARS-Cov-2 The conserved surface of the XBB.1.5 variant of the RBD of the S protein is or contains: L335, E340, A348, S349, Y351, A352, N354, R355, K356, R357, S359, N360, V362, D364, S366, Y369, N370, A372, F377, K378, Y380, G381, S383, P384, T385, K386, N388, D389, L390, C391, F392, T393, N394, Y396, P412, G413, Q414, at least a portion (e.g., up to all of) of positions T415, K424, P426, D427, D428, T430, K444, N450, L452, R457, K458, S459, K462, P463, F464, E465, R466, D467, I468, S469, T470, E471, I472, Y473, Q474, P479, N481, G482, V483, E516, L517, L518, H519, A520, P521, T523, C525, G526, P527).
[0020] In some embodiments, the provided methods include identifying and / or accessing the identity of a set of one or more characteristic mutations (e.g., their locations) of a reference antigen (e.g., the reference antigen is a protein of a particular viral variant class, and the set of characteristic mutations includes mutations of the reference antigen that are present at or above a certain threshold rate (e.g., appear in sequences identified as belonging to the viral variant class at or above a threshold rate (e.g., proportion, percentage, etc.)).
[0021] In some embodiments, the reference antigen is a SARS-CoV2 XBB.1.5 spike protein (e.g., a SARS-CoV-2 S protein with characteristic mutations of XBB.1.5 (e.g., T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445K). P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K mutations).
[0022] In some embodiments, step (c) comprises selecting, by the processor, one or more amino acid modifications from a set of allowed mutations (e.g., selecting a position and / or a particular modification (e.g., a substitution, deletion, insertion) from the set of allowed mutations).
[0023] In some embodiments, the set of allowed mutations is or includes (e.g., a list, table, etc. representing) multiple mutations that are observed to occur (e.g., are present at or above a certain ratio) within a set of related antigens.
[0024] In some embodiments, the reference antigen is a SARS-CoV-2 (e.g., spike) protein of a particular variant, and the set of related antigens includes corresponding proteins of other related variants (e.g., within a particular set of lineages and / or sub-lineages that include the reference antigen) (e.g., within a particular set of lineages and / or sub-lineages that include the reference antigen).
[0025] In some embodiments, the reference antigen is a member of the Omicron lineage and the set of related antigens comprises observed variants belonging to the Omicron lineage.
[0026] In some embodiments, the set of allowed mutations includes at least a portion (e.g., a subset, e.g., all) of the mutations listed in Table 3 (e.g., excluding those removed, e.g., as marked with a strikethrough in Table 3) (e.g., and, optionally, one or more additional mutations, e.g., not any other mutations).
[0027] In some embodiments, the set of allowed mutations includes at least a portion (eg, a subset) of the mutations listed in Table 3, excluding one or both of L335F and L390R.
[0028] In some embodiments, the reference antigen is a SARS-CoV2 (e.g., spike) protein and the set of related antigens includes corresponding proteins of other (e.g., SARS-CoV 1, MERS) coronaviruses.
[0029] In some embodiments, the one or more memory-inducing conserved regions are or comprise a set of conserved epitope regions, and step (c) comprises introducing at least one amino acid modification into each of at least a portion (e.g., all, e.g., a particular subset) of the conserved epitope regions.
[0030] In some embodiments, the one or more memory-inducing conserved regions are or comprise a conserved surface, and step (c) comprises generating one or more amino acid modifications at positions that are distributed throughout / across the conserved surface (e.g., approximately evenly, as opposed to clustered) (e.g., dispersed in a manner that is substantially equidistant across the 3D representation of the conserved surface (e.g., the 3D linear and / or geodesic distances (across the 3D conserved surface) between each of the plurality of amino acid modifications are substantially equal and / or distributed according to a particular predefined statistical distribution (e.g., a normal distribution))).
[0031] In some embodiments, provided methods include iteratively performing steps (b) and (c) to generate a plurality of candidate polypeptide models, each of which represents a candidate engineered variant (e.g., each candidate polypeptide model includes a different combination of amino acid modifications and represents a unique engineered version of the reference antigen).
[0032] In some embodiments, a provided method includes determining, by a processor, one or more performance score values for each of the candidate polypeptide models, and selecting a subset of the candidate polypeptide models based at least in part on the determined performance score values.
[0033] In some embodiments, the one or more performance scores include one or both of: (a) an immune escape score, which indicates the likelihood and / or relative ability of a particular candidate engineered variant to be recognized and neutralized by an antibody; and (b) a fitness score, which indicates the likelihood and / or viability of a particular candidate engineered variant.
[0034] In some embodiments, determining one or both of (a) the immune escape score and (b) the fitness score includes using a machine learning model (e.g., a language model that receives as input a representation of the amino acid sequence of a particular candidate engineered variant (e.g., the input does not include a reference 3D structural representation)) (e.g., to calculate a likelihood score that represents the predicted likelihood that a particular candidate variant will occur, e.g., to calculate a semantic change score that indicates the distance between the embedded representation of (i) a particular candidate variant (e.g., generated from one or more hidden layers of the machine learning model) and (ii) one or more reference variants (e.g., a WT variant, e.g., a variant of a reference antigen to which the subject has previously been infected, e.g., via vaccination and / or natural exposure, e.g., the reference antigen)).
[0035] In some embodiments, determining one or both of (a) the immune escape score and (b) the fitness score comprises using a 3D structural model of at least a portion of the particular candidate variant (e.g., calculating a viral polypeptide receptor binding score (e.g., an ACE2 binding score), e.g., calculating an epitope alteration score).
[0036] In some embodiments, the one or more performance scores include a positional diffusion score, which measures the degree to which amino acid modifications are evenly distributed across the surface (e.g., conserved surface) of a candidate engineered variant.
[0037] In some embodiments, the one or more performance scores include a mutation co-occurrence score (e.g., measuring the degree to which one or more amino acid modifications of a particular candidate variant match their natural co-occurrence rate).
[0038] In some embodiments, provided methods include identifying, by a processor, within a polypeptide model, one or more target regions that represent portions of a reference antigen (e.g., target epitopes) that should be retained, and excluding the one or more target regions from one or more memory-inducing conserved region(s) (e.g., the reference antigen is a SARS-CoV-2 S protein and / or portions thereof, and the one or more target regions are or include an ACE2 binding interface, e.g., thereby preserving the ACE2 binding interface region).
[0039] In some embodiments, a provided method includes causing, by a processor, a rendering of the collapsed polypeptide model for graphical display.
[0040] In some embodiments, the methods provided include generating a corresponding RNA sequence from the disrupted polypeptide model.
[0041] In some embodiments, the methods provided include producing a composition (e.g., as a vaccine) that includes a polypeptide based on the disrupted polypeptide model (e.g., having substantially the same amino acid sequence as that represented by the disrupted polypeptide model).
[0042] In some embodiments, the methods provided include assessing the biological activity of the engineered antigen in vitro.
[0043] In some embodiments, the biological activity of the engineered antigen is characterized by the engineered antigen being properly expressed and folded (e.g., based on a binding assay, such as a flow cytometry-based binding assay (e.g., with hACE2)); and / or the engineered antigen does not bind to antibodies that bind the reference antigen (e.g., exhibits reduced / abolished binding (compared to the reference antigen) to a panel of one or more neutralizing antibodies (e.g., antibodies that bind to a different epitope class, e.g., one or more of classes A, B, C, D, E, and F (e.g., each of)); and / or a pseudovirus packaged with the engineered antigen is able to enter cells; and / or the engineered antigen is immunogenic and / or the engineered antigen reduces the engineered antigen's activation of B cell memory immune responses to the reference antigen.
[0044] In some embodiments, the methods provided include producing a composition (e.g., as a vaccine) that includes a nucleic acid encoding an amino acid sequence represented by the disrupted polypeptide model.
[0045] In some aspects, the present disclosure provides vaccine compositions comprising the polypeptides and / or nucleic acids of one or more aspects or embodiments described herein (e.g., as described in the paragraphs above).
[0046] In some aspects, the present disclosure provides methods of vaccination comprising administering to a subject or population of subjects a vaccine according to one or more aspects or embodiments described herein (e.g., described in the paragraphs above).
[0047] In some aspects, the present disclosure provides a system including a processor of a computing device and a memory having instructions stored thereon, the instructions, when executed by the processor, causing the processor to perform a method of one or more aspects or embodiments described herein (e.g., as described in the paragraphs above).
[0048] In some aspects, the present disclosure provides methods of making immunogenic compositions, comprising comparing sequences of viral proteins from different variants of an infectious disease agent (e.g., influenza, e.g., SARS-CoV-2) to identify (remaining) conserved sites in an antigen of interest (e.g., by one or more systems and / or methods, including any of those described in the preceding claims), replacing at least one or more of the remaining conserved sites with sequences characterized by characteristics / properties of such amino acid substitutions (examples of which may provide new immunogenic epitopes from the various variants) to generate new sequences (e.g., by one or more systems and / or methods, including any of those described in the preceding claims), and producing a vaccine that delivers at least a portion of the new sequences that include at least one of the replaced remaining conserved sites.
[0049] In some aspects, the disclosure provides RNA comprising a nucleotide sequence encoding an engineered antigen (e.g., for eliciting an immune response directed against one or more target epitopes of a reference antigen of an infectious agent, while the engineered antigen reduces (e.g., characterized by) activation of a memory immune response (e.g., of B cells and / or T cells) against the reference antigen), where the engineered antigen corresponds to (e.g., has the sequence of) a particular reference antigen, and the reference antigen has been altered (e.g., artificially) to introduce one or more amino acid modifications within at least a portion of one or more memory-inducing conserved regions identified as portions of the reference antigen determined to be likely to elicit a memory immune response (e.g., portions determined to (i) correspond to known and / or potential epitopes and / or (ii) not contain a characteristic mutation of the reference antigen) (such that, for example, the engineered antigen is an artificially engineered version of the reference antigen in which mutations have been introduced to disrupt the conserved portions of the reference antigen (e.g., portions determined / determined to be likely to elicit a memory immune response)).
[0050] In some embodiments, the one or more memory-inducing conserved region(s) (i) represent a portion(s) of a reference antigen that is substantially similar (e.g., has sufficient sequence similarity, e.g., is in common) with one or more (e.g., pre-existing) variants thereof.
[0051] In some embodiments, the one or more memory-inducing conserved region(s) is or includes a set of conserved epitope regions that represent known epitopes that are present on the reference antigen and that have not been mutated (e.g., known epitopes that do not have any characteristic mutations).
[0052] In some embodiments, the set of known epitopes includes one or more of the epitopes listed in Table 2A and / or Table 2B.
[0053] In some embodiments, the one or more memory-inducing conserved region(s) is or includes a conserved surface representing a (e.g., contiguous, e.g., adjacent), unmutated (e.g., lacking any characteristic mutations) surface of a reference antigen (e.g., the target polypeptide is or includes the XBB.1.5 variant of the RBD of the SARS-CoV-2 spike (S) protein, and the conserved surface is L335, E340, A348, S349, Y351, A352, N354, R355, K356, R357, S359, N360, V362, D364, S366, Y369, N370, A372, F377, K378, Y380, , G381, S383, P384, T385, K386, N388, D389, L390, C391, F392, T393, N394, Y396, P412, G4 13, Q414, T415, K424, P426, D427, D428, T430, K444, N450, L452, R457, K458, S459, K462, (At or including positions P463, F464, E465, R466, D467, I468, S469, T470, E471, I472, Y473, Q474, P479, N481, G482, V483, E516, L517, L518, H519, A520, P521, T523, C525, G526, P527).
[0054] In some embodiments, the reference antigen is the XBB.1.5 variant of the SARS-Cov2 spike protein or comprises at least a portion thereof (e.g., the RBD thereof) (e.g., hallmark mutations include T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371H, S372H ... and including at least some (e.g., all) of the following mutations: F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K).
[0055] In some embodiments, the one or more amino acid modifications are selected from a set of allowed mutations (e.g., selecting a position and / or a particular modification (e.g., substitution, deletion, insertion) from a set of allowed mutations).
[0056] In some embodiments, the set of allowed mutations is or includes a plurality of mutations that are observed to occur (e.g., are present at or above a certain ratio) within a set of related polypeptides.
[0057] In some embodiments, the set of related polypeptides includes corresponding polypeptides of other related SARS-CoV 2 variants (e.g., within a particular set of lineages and / or sub-lineages that include a particular SARS-CoV 2 variant).
[0058] In some embodiments, the target (SARS-CoV 2 variant) polypeptide is a member of the Omicron phylogeny, and the set of related polypeptides includes corresponding polypeptides (e.g., spike protein sequences) of observed variants belonging to the Omicron phylogeny.
[0059] In some embodiments, the one or more memory-inducing conserved regions are or include a set of conserved epitope regions, and the engineered antigen has at least one amino acid modification within each of at least a portion (e.g., all, e.g., a particular subset) of the conserved epitope regions.
[0060] In some embodiments, the one or more memory-inducing conserved regions are or comprise a conserved surface, and the one or more amino acid modifications occur at positions that are distributed (e.g., approximately evenly, as opposed to clustered) throughout / across the conserved surface (e.g., are dispersed in a manner that is substantially equidistant across a 3D representation of the conserved surface (e.g., the 3D linear and / or geodesic distances (across the 3D conserved surface) between each of the multiple amino acid modifications are substantially equal and / or are distributed according to a particular predefined statistical distribution (e.g., a normal distribution))).
[0061] In some embodiments, the engineered antigen of the provided methods, systems, vaccine compositions, manufacturing methods, or RNA comprises one or more of the mutation combinations listed in Table 5A (e.g., the second or third column of each row) (e.g., the engineered antigen comprises the RBD (e.g., residues 327-528 of SEQ ID NO: 1) of the SARS-CoV-2 S protein (e.g., as set forth in any one of SEQ ID NOs: 4 or 38) with one or more of the mutation combinations listed in the third column of Table 5A) (e.g., the engineered antigen comprises the RBD (e.g., residues 327-528 of SEQ ID NO: 1) of the SARS-CoV-2 S protein (e.g., as set forth in any one of SEQ ID NOs: 4 or 38) with one or more of the additional mutation combinations listed in the second column of Table 5A). and an RBD of an XBB.1.5 variant S protein (e.g., as set forth in any one of SEQ ID NOs: 35, 36, or 37) (e.g., residues 327-528 of SEQ ID NO: 1 having at least a portion of the characteristic mutations of XBB.1.5 identified in Table 1B).
[0062] In some embodiments, the engineered antigen of the provided methods, systems, vaccine compositions, manufacturing methods, or RNA comprises one or more of the sequences listed in Table 5B (e.g., each row) (e.g., the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or portion thereof (e.g., the reference antigen is or includes the RBD of the S protein)) (e.g., the engineered antigen is or includes the RBD of the SARS-CoV-2 S protein having any one of the sequences listed in Table 5B (e.g., any one of SEQ ID NOs: 7-34)).
[0063] In some embodiments, the engineered antigen of the provided methods, systems, vaccine compositions, manufacturing methods, or RNA comprises one or more of the combinations of mutations listed in Table 7A and / or Table 7B (e.g., the first or second column of each row) (e.g., the engineered antigen comprises the RBD of the SARS-CoV-2 XBB.1.5 variant S protein (e.g., as set forth in any one of SEQ ID NOS: 35, 36, or 37) having one or more of the combinations of additional mutations listed in the second column of Table 7A and / or Table 7B (e.g., identified as L1, L2, L3, and L4) (e.g., residues 327-528 of SEQ ID NO: 1 having at least a portion of the hallmark mutations of XBB.1.5 identified in Table 1B)).
[0064] In some embodiments, the engineered antigen of the provided methods, systems, vaccine compositions, manufacturing methods, or RNA comprises one or more of the mutation combinations listed in Table 8A (e.g., the second or third column of each row) (e.g., the engineered antigen comprises the RBD (e.g., residues 327-528 of SEQ ID NO: 1) of the SARS-CoV-2 S protein (e.g., as set forth in any one of SEQ ID NOs: 4 or 38) with one or more of the mutation combinations listed in the third column of Table 8A) (e.g., the engineered antigen comprises the RBD (e.g., residues 327-528 of SEQ ID NO: 1) of the SARS-CoV-2 S protein (e.g., as set forth in any one of SEQ ID NOs: 4 or 38) with one or more of the additional mutation combinations listed in the second column of Table 8A). and an RBD of an XBB.1.5 variant S protein (e.g., as set forth in any one of SEQ ID NOs: 35, 36, or 37) (e.g., residues 327-528 of SEQ ID NO: 1 having at least a portion of the characteristic mutations of XBB.1.5 identified in Table 1B).
[0065] In some embodiments, the engineered antigen of the provided methods, systems, vaccine compositions, manufacturing methods, or RNA comprises one or more of the sequences listed in Table 8B (e.g., each row) (e.g., the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or portion thereof (e.g., the reference antigen is or includes the RBD of the S protein)) (e.g., the engineered antigen is or includes the RBD of the SARS-CoV-2 S protein having any one of the sequences listed in Table 8B (e.g., any one of SEQ ID NOs: 39-64 and 100)).
[0066] In some embodiments, the engineered antigen of the provided methods, systems, vaccine compositions, manufacturing methods, or RNA comprises one or more of the mutation combinations listed in Table 9A (e.g., the second column of each row) (e.g., the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or portion thereof (e.g., the reference antigen is or includes the RBD of the S protein)). (e.g., the engineered antigen comprises the RBD of the SARS-CoV-2 XBB.1.5 variant S protein (e.g., as set forth in any one of SEQ ID NOs: 35, 36, or 37) with one or more of the additional mutation combinations listed in the second column of Table 9A (e.g., residues 327-528 of SEQ ID NO: 1 with at least a portion of the XBB.1.5 hallmark mutations identified in Table 1B).
[0067] In some embodiments, the engineered antigen of a provided method, system, vaccine composition, manufacturing method, or RNA comprises a combination of mutations identified as S43 in Table 9A (e.g., the engineered antigen is or comprises the RBD of an XBB.1.5 S protein with additional mutations according to S43 in Table 9A) (e.g., the engineered antigen comprises the RBD of a SARS-CoV-2 XBB.1.5 variant S protein (e.g., as set forth in any one of SEQ ID NOs: 35, 36, or 37) with additional mutations N360D, P384S, L390R, T430I, F464Y, H519N (e.g., residues 327-528 of SEQ ID NO: 1 with at least a portion of the hallmark mutations of XBB.1.5 identified in Table 1B)).
[0068] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens are or include at least a portion (e.g., a particular domain (e.g., RBD)) of the SARS-CoV-2 S protein (e.g., as set forth in SEQ ID NO: 1) having at least a portion (e.g., all) of the following mutations: N360D, P384S, L390R, T430I, F464Y, H519N (e.g., one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., SARS-CoV-2 In addition to mutations within the RBD of the S protein (e.g., T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417 One or more characteristic mutations of XBB.1.5 such as N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K).
[0069] In some embodiments, the engineered antigen of a provided method, system, vaccine composition, manufacturing method, or RNA comprises a combination of mutations identified as S48 in Table 9A (e.g., the engineered antigen is or comprises the RBD of XBB.1.5 with additional mutations according to S48 in Table 9A) (e.g., the engineered antigen comprises the RBD of a SARS-CoV-2 XBB.1.5 variant (e.g., as set forth in any one of SEQ ID NOs: 35, 36, or 37) with additional mutations L335F, K356T, P384S, L390R, T430I, F464Y, and H519N (e.g., residues 327-528 of SEQ ID NO: 1 with the hallmark mutations of XBB.1.5 identified in Table 1B)).
[0070] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens are or include at least a portion (e.g., a particular domain (e.g., RBD)) of the SARS-CoV-2 S protein (e.g., as set forth in SEQ ID NO: 1) having at least a portion (e.g., all) of the following mutations: L335F, K356T, P384S, L390R, T430I, F464Y, and H519N (e.g., one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., SARS-CoV-2 In addition to mutations within the RBD of the S protein (e.g., T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417 One or more characteristic mutations of XBB.1.5 such as N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K).
[0071] In some embodiments, (e.g., the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or portions thereof (e.g., the reference antigen is or includes the RBD of the S protein)), provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens include at least a portion of the SARS-CoV-2 S protein (e.g., the RBD) having the following mutations: P384S, L390R, T430I, F464Y, H519N.
[0072] In some embodiments, (e.g., the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or portion thereof (e.g., the reference antigen is or includes the RBD of the S protein)), provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens include at least a portion of the SARS-CoV-2 S protein (e.g., the RBD) having one or more (e.g., a subset thereof, e.g., all of) the following mutations: K356T, P384S, L390R, T430I, F464Y, H519N.
[0073] In some embodiments, (e.g., the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or portions thereof (e.g., the reference antigen is or includes the RBD of the S protein)), the provided methods, systems, vaccine compositions, manufacturing methods, or RNA do not comprise one or both of the L335F and L390R mutations.
[0074] In some embodiments, (e.g., the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or portion thereof (e.g., the reference antigen is or includes the RBD of the S protein)), the provided methods, systems, vaccine compositions, manufacturing methods, or RNAs include one or more of the combinations of mutations listed in (e.g., each row of) Table 15A and / or Table 15B (e.g., the engineered antigen includes the RBD (e.g., residues 327-528 of SEQ ID NO: 1) of the SARS-CoV-2 S protein (e.g., as set forth in any one of SEQ ID NOs: 4 or 38) with one or more of the combinations of mutations listed in Table 15A) (e.g., the engineered antigen includes the RBD (e.g., residues 327-528 of SEQ ID NO: 1) of the SARS-CoV-2 S protein (e.g., as set forth in any one of SEQ ID NOs: 4 or 38) with one or more of the combinations of additional mutations listed in column 2 of Table 15B). and an RBD of an XBB.1.5 variant (e.g., as set forth in any one of SEQ ID NOs: 35, 36, or 37) (e.g., residues 327-528 of SEQ ID NO: 1 having at least a portion of the characteristic mutations of XBB.1.5 identified in Table 1B).
[0075] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods or RNA engineered antigens comprise a combination of mutations identified as S123 in Table 15A and / or Table 15B.
[0076] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens are or include at least a portion (e.g., a particular domain (e.g., RBD)) of the SARS-CoV-2 S protein (e.g., as set forth in SEQ ID NO: 1) having at least a portion (e.g., all) of the following mutations: I332V, L335F, K356T, P384S, T430I, L452Q, F464Y, H519N (e.g., one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., SARS-CoV-2 In addition to mutations within the RBD of the S protein (e.g., T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417 One or more characteristic mutations of XBB.1.5 such as N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K).
[0077] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens include at least a portion (e.g., up to all) of the following mutations: I332V, L335F, G339H, R346T, K356T, L368I, S371F, S373P, S375F, T376A, P348S, D405N, R408S, K417N, T430I, N440K, V445P, G446S, L452Q, N460K, F464Y, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, E516Q, H519N.
[0078] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods or RNA engineered antigens comprise a combination of mutations identified as S122 in Table 15A and / or Table 15B.
[0079] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens are or include at least a portion (e.g., a particular domain (e.g., RBD)) of the SARS-CoV-2 S protein (e.g., as set forth in SEQ ID NO: 1) having at least a portion (e.g., all) of the following mutations: L335F, K356T, P384S, T430I, L452R, F464Y, H519N (e.g., one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., SARS-CoV-2 In addition to mutations within the RBD of the S protein (e.g., T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417 One or more characteristic mutations of XBB.1.5 such as N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K).
[0080] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens include at least a portion (e.g., up to all) of the following mutations: L335F, G339H, R346T, K356T, L368I, S371F, S373P, S375F, T376A, P348S, D405N, R408S, K417N, T430I, N440K, V445P, G446S, L452R, N460K, F464Y, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, E516Q, H519N, T523S.
[0081] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods or RNA engineered antigens comprise a combination of mutations identified as S109 in Table 15A and / or Table 15B.
[0082] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens are or include at least a portion (e.g., a particular domain (e.g., RBD)) of the SARS-CoV-2 S protein (e.g., as set forth in SEQ ID NO: 1) having at least a portion (e.g., all) of the following mutations: K356T, N360S, P384S, N388K, T430I, N450D, F464Y, H519N (e.g., one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., SARS-CoV-2 In addition to mutations within the RBD of the S protein (e.g., T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417 One or more characteristic mutations of XBB.1.5 such as N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K).
[0083] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens include at least a portion (e.g., up to all) of the following mutations: G339H, R346T, K356T, L368I, S371F, S373P, S375F, T376A, P348S, N388K, D405N, R408S, K417N, T430I, N440K, V445P, G446S, N450D, N460K, F464Y, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, H519N, T523S.
[0084] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods or RNA engineered antigens comprise a combination of mutations identified as S129 in Table 15A and / or Table 15B.
[0085] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens are or include at least a portion (e.g., a particular domain (e.g., RBD)) of the SARS-CoV-2 S protein (e.g., as set forth in SEQ ID NO: 1) having at least a portion (e.g., all) of the following mutations: K356T, L335F, P384S, D389G, T430I, N450D, F464Y, H519N (e.g., one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., SARS-CoV-2 In addition to mutations within the RBD of the S protein (e.g., T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417 One or more characteristic mutations of XBB.1.5 such as N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K).
[0086] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens include at least a portion (e.g., up to all) of the following mutations: L335F, G339H, R346T, K356T, L368I, S371F, S373P, S375F, T376A, P348S, D389G, D405N, R408S, K417N, T430I, N440K, V445P, G446S, N450D, L452R, N460K, F464Y, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, H519N.
[0087] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods or RNA engineered antigens comprise a combination of mutations identified as S156 in Table 15A and / or Table 15B.
[0088] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens are or include at least a portion (e.g., a particular domain (e.g., RBD)) of the SARS-CoV-2 S protein (e.g., as set forth in SEQ ID NO: 1) having at least a portion (e.g., all) of the following mutations: K356T, P384S, L390R, T430I, N450D, F464Y, I472V, H519N (e.g., one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., SARS-CoV-2 In addition to mutations within the RBD of the S protein (e.g., T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417 One or more characteristic mutations of XBB.1.5 such as N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K).
[0089] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens include at least a portion (e.g., up to all) of the following mutations: G339H, R346T, K356T, L368I, S371F, S373P, S375F, T376A, P348S, L390R, D405N, R408S, K417N, T430I, N440K, V445P, G446S, N450D, L452R, N460K, F464Y, I472V, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, H519N.
[0090] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods or RNA engineered antigens comprise a combination of mutations identified as S112 in Table 15A and / or Table 15B.
[0091] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens are or include at least a portion (e.g., a particular domain (e.g., RBD)) of a SARS-CoV-2 S protein (e.g., as set forth in SEQ ID NO: 1) having at least a portion (e.g., all) of the following mutations: K356T, P384S, D389G, T430I, N450D, F464Y, I468V, H519N (e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., mutations within the RBD of the SARS-CoV-2 S protein) (e.g., , T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G 339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445 one or more characteristic mutations of XBB.1.5 such as P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K).
[0092] In some embodiments, the engineered antigens of the provided methods, systems, vaccine compositions, manufacturing methods, or RNAs include at least a portion (e.g., up to all) of the following mutations: G339H, R346T, K356T, L368I, S371F, S373P, S375F, T376A, P348S, D389G, D405N, R408S, K417N, T430I, N440K, V445P, G446S, N450D, N460K, F464Y, I468V, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, H519N.
[0093] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods or RNA engineered antigens comprise a combination of mutations identified as S125 in Table 15A and / or Table 15B.
[0094] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens are or include at least a portion (e.g., a particular domain (e.g., RBD)) of a SARS-CoV-2 S protein (e.g., as set forth in SEQ ID NO: 1) having at least a portion (e.g., all) of the following mutations: K356T, L335F, P384S, T430I, F464Y, I468V, H519N (e.g., in addition to one or more characteristic mutations of a particular SARS-CoV-2 variant (e.g., mutations within the RBD of the SARS-CoV-2 S protein) (e.g., T1 9I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G33 9H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P , G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K)).
[0095] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or RNA engineered antigens include at least a portion (e.g., up to all) of the following mutations: L335F, G339H, R346T, K356T, L368I, S371F, S373P, S375F, T376A, P348S, D405N, R408S, K417N, T430I, N440K, V445P, G446S, N460K, F464Y, I468V, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, E516Q, H519N, T523S.
[0096] In some aspects, the disclosure provides methods for producing RNA comprising a nucleotide sequence encoding an engineered antigen that corresponds to an engineered version of a reference antigen, the method comprising producing RNA, wherein the nucleotide sequence of the RNA, when compared to (e.g., the nucleotide sequence of) the reference antigen, exhibits difference(s) relative to the reference antigen (e.g., comprises nucleotides encoding one or more mutations) in one or more memory-inducing conserved regions (e.g., one or more conserved epitopes and / or one or more conserved surface(s)) common to the reference antigen and (i) one or more pre-existing variants of the reference antigen, and / or (ii) a wild-type strain of the reference antigen.
[0097] In some embodiments, the reference antigen is or includes the RBD of a specific target variant SARS-CoV-2 S protein (e.g., the RBD of XBB.1.5 (e.g., having an amino acid sequence set forth in any one of SEQ ID NOS: 35, 36, and 37) (e.g., residues 327-528 of SEQ ID NO: 1 with the characteristic mutations of XBB.1.5 identified in Table 1B)).
[0098] In some embodiments, the RNA produced according to the provided methods (e.g., as described above) is or comprises RNA of any one of the aspects and embodiments described herein (e.g., the differences compared to the reference antigen correspond to / include nucleotide differences encoding any of the combinations of mutations of the various aspects and embodiments described herein (e.g., as described in the paragraphs above)).
[0099] In some embodiments, the methods of production provided involve producing RNA by in vitro transcription (IVT).
[0100] Features of embodiments described with respect to one aspect of the invention may also be applied with respect to other aspects of the invention.
[0101] The foregoing and other objects, aspects, features and advantages of the present disclosure will become more apparent and will be better understood by referring to the following description taken in conjunction with the accompanying drawings. [Brief explanation of the drawings]
[0102] [Figure 1] FIG. 1 is a block flow diagram and schematic diagram illustrating an exemplary process for the design of engineered antigens, according to an exemplary embodiment.
[0103] [Figure 2] FIG. 1 is a block flow diagram of an exemplary process for inserting amino acid modifications, according to an exemplary embodiment.
[0104] [Figure 3] FIG. 1 is a block flow diagram of an exemplary process for generating and scoring multiple candidate antigen designs, according to an exemplary embodiment.
[0105] [Figure 4A] 1 is a set of illustrations of a 3D model of the receptor binding domain (RBD) of the SARS-Cov2 spike (S) protein of an XBB variant, color-coded to indicate the characteristic mutations of XBB and the identified conserved regions, according to an exemplary embodiment.
[0106] [Figure 4B] 6B is another set of diagrams of the XBB RBD S protein model shown in FIG. 6A , color-coded to indicate XBB signature mutations, identified conserved regions, and the ACE2 binding interface, according to an exemplary embodiment.
[0107] [Figure 4C] 6C is another set of diagrams of the XBB RBD S protein model shown in FIGS. 6A and 6B, color-coded to indicate the relative frequency of epitope involvement at various amino acid sites within the protein, according to an exemplary embodiment.
[0108] [Figure 4D] FIG. 1 is a set of illustrations of a 3D model of the first engineered antigen design generated using the XBB RBD S protein as a reference antigen.
[0109] [Figure 4E] FIG. 10 is a set of illustrations of a 3D model of the second engineered antigen design generated using the XBB RBD S protein as a reference antigen.
[0110] [Figure 5A] Schematic diagram of SARS CoV 2 and associated proteins, adapted from Jain et al., Vaccines 2020, 8(4), 649 (www.mdpi.com / 2076-393X / 8 / 4 / 649.).
[0111] [Figure 5B] 1 is a set of illustrations of a 3D model of the receptor binding domain (RBD) of the SARS-Cov2 spike (S) protein of the XBB.1.5 variant, color-coded to indicate characteristic mutations and identified conserved regions of XBB, according to an exemplary embodiment.
[0112] [Figure 5C] 1 is a set of illustrations of a 3D model of the XBB.1.5 spike protein (complete spike protein), according to an exemplary embodiment.
[0113] [Figure 5D] 7C is another set of diagrams of the XBB RBD S protein model shown in FIG. 7B, color-coded to indicate XBB signature mutations, identified conserved regions, and the ACE2 binding interface, according to an exemplary embodiment.
[0114] [Figure 5E]5B and 5D are another set of diagrams of the XBB RBD S protein model shown in Figures 5B and 5D, color-coded to indicate the relative frequency of epitope involvement at various amino acid sites within the protein, according to an exemplary embodiment.
[0115] [Figure 5F] FIG. 5B is another set of diagrams of the XBB RBD S protein model shown in FIGS. 5B and 5D, color-coded to indicate the relative frequency of epitope involvement at various amino acid sites within the conserved surface of the protein, according to an exemplary embodiment.
[0116] [Figure 5G] FIG. 1 is a block flow diagram of an exemplary process for generating candidate variant designs, according to an exemplary embodiment.
[0117] [Figure 6A]
[0033] Figure 6A shows a set of illustrations of 3D models of candidate engineered antigen designs generated using the XBB.1.5 RBD S protein as a reference antigen, according to an exemplary embodiment. The color coding / shading in Figure 6A indicates the signature mutations of XBB.1.5, the identified conserved regions, and the amino acid sites that were mutated based on the various in silico antigen design approaches described herein.
[0118] [Figure 6B] A set of illustrations of a 3D model of the SARS-CoV-2 spike protein bearing the characteristic mutations of XBB.1.5 and additional mutations corresponding to those highlighted in Figure 6A are shown.
[0119] [Figure 6C] 1 shows a set of illustrations of a 3D model of another candidate engineered antigen design based on (e.g., using) the XBB.1.5 S protein RBD as a reference antigen, according to an exemplary embodiment.
[0120] [Figure 6D]1 shows a set of illustrations of a 3D model of another candidate engineered antigen design based on (e.g., using) the XBB.1.5 S protein RBD as a reference antigen, according to an exemplary embodiment.
[0121] [Figure 6E] 1 shows a set of illustrations of a 3D model of another candidate engineered antigen design based on (e.g., using) the XBB.1.5 S protein RBD as a reference antigen, according to an exemplary embodiment.
[0122] [Figure 6F] 1 shows a set of illustrations of a 3D model of another candidate engineered antigen design based on (e.g., using) the XBB.1.5 S protein RBD as a reference antigen, according to an exemplary embodiment.
[0123] [Figure 7] 1 shows a radar plot of specific performance scores of candidate engineered antigen designs, according to an exemplary embodiment.
[0124] [Figure 8] 1 shows a plot of predicted versus experimental ACE2 binding changes, according to an exemplary embodiment.
[0125] [Figure 9A] FIG. 10 shows a diagram of a 3D model of an engineered RBD construct, in accordance with an exemplary embodiment.
[0126] [Figure 9B] FIG. 10 shows a diagram of a 3D model of an engineered RBD construct, in accordance with an exemplary embodiment.
[0127] [Figure 9C]Schematic diagram of the SARS-CoV-2 trimer. This figure is adapted from Starr et al., SARS-CoV-2 RBD antibodies that maximize breadth and resistance to escape. Nature, 597, 97-102 (2021). https: / / doi.org / 10.1038 / s41586-021-03807-6.
[0128] [Figure 9D] FIG. 1 shows a diagram of a 3D model of an engineered construct designed to incorporate one or more SARS-CoV-1 epitopes onto an XBB.1.5 backbone, according to an exemplary embodiment.
[0129] [Figure 9E] FIG. 1 shows a diagram of a 3D model of an engineered construct designed to incorporate one or more SARS-CoV-1 epitopes onto an XBB.1.5 backbone, according to an exemplary embodiment.
[0130] [Figure 10A] FIG. 1 is a diagram of an exemplary procedure for testing engineered antigens, according to an exemplary embodiment.
[0131] [Figure 10B] FIG. 1 is a block flow diagram of an exemplary protocol for generating and evaluating engineered constructs described herein, according to an exemplary embodiment.
[0132] [Figure 10C-1] 1 is a set of graphs assessing antibody binding to WT and XBB.1.5S proteins for various epitope classes, according to an exemplary embodiment. [Figure 10C-2] 1 is a set of graphs assessing antibody binding to WT and XBB.1.5S proteins for various epitope classes, according to an exemplary embodiment.
[0133] [Figure 10D] Graph of a dilution series of hACE2-mFc binding agents.
[0134] [Figure 10E] 16 is a graph of a dilution series of human BNT162b23 (triple vaccine) polyclonal serum.
[0135] [Figure 11A] 1 shows binding dose-response curves for certain engineered SARS-CoV 2 antigens described herein.
[0136] [Figure 11B] 1 shows binding dose-response curves for certain engineered SARS-CoV 2 antigens described herein.
[0137] [Figure 11C] 1 shows binding dose-response curves for certain engineered SARS-CoV 2 antigens described herein.
[0138] [Figure 12A] 1 shows binding dose-response curves for certain engineered SARS-CoV 2 antigens described herein.
[0139] [Figure 12B] 1 shows binding dose-response curves for certain engineered SARS-CoV 2 antigens described herein.
[0140] [Figure 13A] 1 shows binding dose-response curves for certain engineered SARS-CoV 2 antigens described herein.
[0141] [Figure 13B] 1 shows binding dose-response curves for certain engineered SARS-CoV 2 antigens described herein.
[0142] [Figure 13C]1 shows binding dose-response curves for certain engineered SARS-CoV 2 antigens described herein.
[0143] [Figure 13D] 1 shows binding dose-response curves for certain engineered SARS-CoV 2 antigens described herein.
[0144] [Figure 14A] 1 shows binding dose-response curves for certain engineered SARS-CoV 2 antigens described herein.
[0145] [Figure 14B] 1 shows binding dose-response curves for certain engineered SARS-CoV 2 antigens described herein.
[0146] [Figure 14C] 1 shows binding dose-response curves for certain engineered SARS-CoV 2 antigens described herein.
[0147] [Figure 14D] 1 shows binding dose-response curves for certain engineered SARS-CoV 2 antigens described herein.
[0148] [Figure 15A] 1 is a schematic diagram showing an exemplary group and dosage / sampling schedule for assessing immune imprinting and performance of candidate vaccines in vaccinated mice. According to an exemplary embodiment, yellow-filled cells indicate days for collecting serum samples, gray-filled cells indicate days for administering vaccine, and green-filled cells indicate days for sacrificing mice and collecting final samples.
[0149] [Figure 15B]1 is a schematic diagram showing an exemplary group and dosage / sampling schedule for evaluating the performance of candidate vaccines in vaccine-naive mice, where yellow-filled cells indicate days when serum samples are collected, gray-filled cells indicate days when the vaccine is administered, and green-filled cells indicate days when mice are sacrificed and final samples are collected, according to an exemplary embodiment.
[0150] [Figure 16A] 1 is a schematic diagram showing a FACS-based assay according to an exemplary embodiment, e.g., using the approach described in Quandt and Muik et al., 2022, the contents of which are hereby incorporated by reference in their entirety.
[0151] [Figure 16B] 1 is a schematic diagram showing a depletion assay according to an exemplary embodiment, e.g., using the approach described in Quandt and Muik et al., 2022, the contents of which are hereby incorporated by reference in their entirety.
[0152] [Figure 17] 1 is a schematic diagram illustrating an approach for the collection and analysis of spleen samples, lymph nodes, and also the analysis of blood samples, according to certain embodiments.
[0153] [Figure 18] FIG. 1 shows the mutations included in the design of various engineered antigen constructs, according to an exemplary embodiment.
[0154] [Figure 19A] FIG. 1 is a schematic diagram showing an exemplary construct comprising a membrane-anchored RBD+fibritin domain with a viral signal peptide, according to an exemplary embodiment.
[0155] [Figure 19B]1 is a graph showing the results of an ACE2 binding assay for a set of engineered construct designs, according to an exemplary embodiment.
[0156] [Figure 19C] 1 is a graph showing the results of serum (from vaccinated patients) binding assays for a set of engineered construct designs, according to an exemplary embodiment.
[0157] [Figure 20A] 1 is a set of graphs showing the results of binding assays for a set of monoclonal antibodies, according to an exemplary embodiment.
[0158] [Figure 20B] 1 is a set of graphs showing the results of binding assays for a set of monoclonal antibodies, according to an exemplary embodiment.
[0159] [Figure 20C] 1 is a set of graphs showing the results of binding assays for a set of monoclonal antibodies, according to an exemplary embodiment.
[0160] [Figure 21] FIG. 1 is a block diagram of an exemplary cloud computing environment for use in certain embodiments.
[0161] [Figure 22] FIG. 1 illustrates an exemplary computing device and an exemplary mobile computing device for use in certain embodiments. DETAILED DESCRIPTION OF THE INVENTION
[0162] The features and advantages of the present disclosure will become more apparent from the following detailed description when read in conjunction with the drawings, in which like reference characters identify corresponding elements throughout. In the drawings, like reference numbers generally indicate identical, functionally similar, and / or structurally similar elements.
[0163] Specific Definitions About or Approximately: The term "about" or "approximately" used herein in reference to a value refers to a value similar to the referenced value. Generally, a person of ordinary skill in the art familiar with the context will understand the relevant degree of variation encompassed by "about" or "approximately" in that context. For example, in some embodiments, the term "about" or "approximately" can encompass a range of values that is 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the referenced value.
[0164] Administration: As used herein, the term "administration" generally refers to the administration of a composition to a subject or system. Those of skill in the art will recognize various routes that may be utilized for administration to a subject, e.g., a human, under appropriate circumstances. For example, in some embodiments, administration may be ocular, oral, parenteral, topical, etc. In some particular embodiments, administration may be bronchial (e.g., via bronchial instillation), oral, cutaneous (e.g., may be or include one or more of topical to the dermis, intradermal, interdermal, transdermal, etc.), intestinal, intra-arterial, intradermal, intragastric, intramedullary, intramuscular, intranasal, intraperitoneal, intrathecal, intravenous, intraventricular, intraspecific organ (e.g., intrahepatic), mucosal, nasal, oral, rectal, subcutaneous, sublingual, topical, tracheal (e.g., via intratracheal instillation), vaginal, vitreous, etc. In some embodiments, administration may involve intermittent dosing (e.g., multiple doses spaced apart in time) and / or periodic dosing (e.g., individual doses separated by a common period of time). In some embodiments, administration may involve continuous dosing (e.g., perfusion) over at least a selected period of time.
[0165] Adult: As used herein, the term "adult" refers to a human over the age of 18. In some embodiments, an adult human weighs between about 90 pounds and about 250 pounds.
[0166] Affinity: As known in the art, "affinity" is a measure of the strength with which two or more binding partners bind to one another. Those skilled in the art will be aware of various assays that can be used to assess affinity, and will also be aware of appropriate controls for such assays. In some embodiments, affinity is assessed in a quantitative assay. In some embodiments, affinity is assessed across multiple concentrations (e.g., one binding partner at a time). In some embodiments, affinity is assessed in the presence of one or more potentially competing entities (e.g., that may be present in a relevant (e.g., physiological) setting). In some embodiments, affinity is assessed relative to a reference (e.g., a reference with a known affinity above a certain threshold (see "positive control") or a reference with a known affinity below a certain threshold (see "negative control")). In some embodiments, affinity may be assessed relative to a contemporaneous reference. Also, in some embodiments, affinity may be assessed relative to a historical reference. Typically, when affinity is assessed relative to a reference, it is assessed under comparable conditions.
[0167] Agent: Generally, as used herein, the term "agent" is used to refer to an entity (e.g., a lipid, metal, nucleic acid, polypeptide, polysaccharide, small molecule, etc., or a complex, combination, mixture, or system thereof (e.g., a cell, tissue, organism)), or a phenomenon (e.g., heat, an electric current or field, a magnetic force or field, etc.). In appropriate circumstances, as will be clear to one of skill in the art from the context, the term may be utilized to refer to an entity that is or includes a cell or organism, or a fragment, extract, or component thereof. Alternatively, or additionally, as will be clear from the context, the term may be used to refer to a natural product found in and / or obtained from nature. In some cases, also as will be clear from the context, the term may be used to refer to one or more entities that are man-made (in that they have been designed, engineered, and / or produced by the action of the human hand and / or are not found in nature). In some embodiments, the agent may be utilized in isolated or pure form, and in some embodiments, the agent may be utilized in crude form. In some embodiments, potential agents may be provided as a collection or library that can be screened, for example, to identify or characterize active agents within. In some cases, the term "agent" may refer to a compound or entity that is or includes a polymer, and in some cases, the term may refer to a compound or entity that includes one or more polymer moieties. In some embodiments, the term "agent" may refer to a compound or entity that is not a polymer and / or that is substantially free of any polymer and / or that does not include one or more specific polymer moieties. In some embodiments, the term may refer to a compound or entity that lacks or is substantially free of any polymer moieties.
[0168] Amelioration: As used herein, the term "amelioration" refers to the prevention, reduction, or alleviation of a condition, or an improvement in a subject's condition. Amelioration includes, but does not require, complete recovery or complete prevention of a disease, disorder, or condition (such as radiation damage).
[0169] Amino Acid: As used herein, the term "amino acid" in its broadest sense refers to compounds and / or substances that can be, are, or have been incorporated into a polypeptide chain, for example, through the formation of one or more peptide bonds. In some embodiments, an amino acid has the general structure HN-C(H)(R)-COOH. In some embodiments, an amino acid is a natural amino acid. In some embodiments, an amino acid is a non-natural amino acid. In some embodiments, an amino acid is a D-amino acid. In some embodiments, an amino acid is an L-amino acid. A "standard amino acid" refers to any of the 20 standard L-amino acids commonly found in naturally occurring peptides. A "non-standard amino acid" refers to any amino acid other than the standard amino acids, whether prepared synthetically or obtained from a natural source. In some embodiments, amino acids, including the carboxy-terminal amino acid and / or amino-terminal amino acid in a polypeptide, can contain structural modifications compared to the general structures above. For example, in some embodiments, an amino acid may be modified relative to the general structure by methylation, amidation, acetylation, PEGylation, glycosylation, phosphorylation, and / or substitution (e.g., of an amino group, a carboxylic acid group, one or more protons, and / or a hydroxyl group). In some embodiments, such modifications may, for example, alter the circulating half-life of a polypeptide containing the modified amino acid compared to one containing an otherwise identical, unmodified amino acid. In some embodiments, such modifications do not significantly alter the relevant activity of a polypeptide containing the modified amino acid compared to one containing an otherwise identical, unmodified amino acid. As will be clear from the context, in some embodiments, the term "amino acid" may be used to refer to a free amino acid. In some embodiments, the term may be used to refer to an amino acid residue of a polypeptide.
[0170] Animal: As used herein, the term "animal" refers to any member of the animal kingdom. In some embodiments, "animal" refers to humans of both genders and at all stages of development. In some embodiments, "animal" refers to non-human animals at all stages of development. In particular embodiments, the non-human animal is a mammal (e.g., a rodent, mouse, rat, rabbit, monkey, dog, cat, sheep, cow, primate, and / or pig). In some embodiments, animals include, but are not limited to, mammals, birds, reptiles, amphibians, fish, insects, and / or worms. In some embodiments, animals may be transgenic animals, genetically engineered animals, and / or clones.
[0171] Antibody: As used herein, the term "antibody" refers to a polypeptide containing sufficient standard immunoglobulin sequence elements to confer specific binding to a particular target antigen. As is known in the art, intact antibodies, as produced in nature, are approximately 150 kD tetrameric agents composed of two identical heavy chain polypeptides (approximately 50 kD each) and two identical light chain polypeptides (approximately 25 kD each), which associate with each other into what is commonly referred to as a "Y-shaped" structure. Each heavy chain consists of at least four domains (each approximately 110 amino acids long): an amino-terminal variable (VH) domain (located at each end of the Y structure) followed by three constant domains (i.e., a CHI, a CH2, and a carboxy-terminal CH3 (located at the base of the stem of the Y)). A short region known as the "switch" connects the heavy chain variable and constant regions. A "hinge" connects the CH2 and CH3 domains to the rest of the antibody. Two disulfide bonds in this hinge region connect the two heavy chain polypeptides to each other in intact antibodies. Each light chain consists of two domains: an amino-terminal variable (VL) domain followed by a carboxy-terminal constant (CL) domain, separated from each other by another "switch." An intact antibody tetramer consists of two heavy-light chain dimers, with the heavy and light chains linked to each other by a single disulfide bond. Two other disulfide bonds connect the hinge regions of the heavy chains, connecting the dimers to form a tetramer. Naturally occurring antibodies are typically glycosylated on the CH2 domain. Each domain in a natural antibody has a structure characterized by an "immunoglobulin fold" formed by two beta sheets (e.g., three-, four-, or five-stranded sheets) packed together into a compressed antiparallel beta barrel. Each variable domain contains three hypervariable loops known as "complement determining regions" (CDR1, CDR2 and CDR3) and four somewhat invariant "framework" regions (FR1, FR2, FR3 and FR4).When a native antibody folds, the FR regions form a beta sheet that provides a structural framework for the domain, and the CDR loop regions from both the heavy and light chains are gathered together in three-dimensional space to create a single hypervariable antigen-binding site located at the tip of a Y-structure. The Fc region of a native antibody binds to elements of the complement system and also to receptors on effector cells, including, for example, effector cells that mediate cytotoxicity. As is known in the art, the affinity of the Fc region for the Fc receptor and / or other binding attributes can be modulated by glycosylation or other modifications. In some embodiments, antibodies produced and / or utilized in accordance with the present disclosure comprise a glycosylated Fc domain, including Fc domains in which such glycosylation has been modified or engineered. For the purposes of this disclosure, in certain embodiments, any polypeptide or complex of polypeptides that contains sufficient immunoglobulin domain sequences similar to those found in a natural antibody can be referred to and / or used as an "antibody," regardless of whether such polypeptide is produced naturally (e.g., produced by an organism in response to an antigen) or produced by recombinant engineering, chemical synthesis, or other artificial system or methodology. In some embodiments, an antibody is a polyclonal antibody, and in some embodiments, an antibody is a monoclonal antibody. In some embodiments, an antibody has constant region sequences characteristic of murine, rabbit, primate, or human antibodies. In some embodiments, antibody sequence elements are humanized, primatized, chimerized, etc., as known in the art. Moreover, as used herein, the term "antibody" can, in appropriate embodiments (unless otherwise specified or apparent from the context), refer to any construct or format known or developed in the art for utilizing the structural and functional characteristics of antibodies in alternative presentations.
[0172] For example, in some embodiments, antibodies utilized in accordance with the present disclosure are in a format selected from, but not limited to, intact IgA, IgG, IgE, or IgM antibodies; bispecific or multispecific antibodies (e.g., Zybodies®, etc.); antibody fragments such as Fab fragments, Fab' fragments, F(ab')2 fragments, Fd' fragments, Fd fragments, and isolated CDRs or sets thereof; single-chain Fvs; polypeptide-Fc fusions; single domain antibodies, alternative scaffolds, or antibody mimetics (e.g., anticalins, FN3 monobodies, DARPins, Affibodies, Affilins, Affimers, Affitins, Alphabodies, Avimers, Fynomers, Im7, VLR, VNAR, Trimab, CrossMab, Trident); nanobodies, bi-nanobodies, F(ab')2, Fab', di-sdFv, single domain antibodies, trifunctional antibodies, diabodies, and minibodies. In some embodiments, relevant formats may be or include: Adnectins®; Affibodies®; Affilins®; Anticalins®; Avimers®; BiTEs®; cameloid antibodies; Centyrins®; ankyrin repeat proteins or DARPINs®; dual affinity retargeting (DART) agents; Fynomers®; shark single domain antibodies such as IgNARs; immunomobilizing anti-cancer monoclonal T cell receptors (ImmTACs); KALBITOR®; microproteins; Nanobodies® minibodies; masked antibodies (e.g., Probodies®); Small Modular ImmunoPharmaceuticals ("SMIPs™"); single chain or tandem diabodies (TandAb®); TCR-like antibodies; Trans-bodies®; TrimerX®; VHHs. In some embodiments, the antibody may lack covalent modifications (eg, glycan attachment) that it has when produced in nature.In some embodiments, the antibody may contain a covalent modification (e.g., attachment of a glycan), a payload (e.g., a detectable moiety, a therapeutic moiety, a catalytic moiety, etc.), or other pendant group (e.g., polyethylene glycol, etc.).
[0173] Antigen: As used herein, the term "antigen" refers to an agent that elicits an immune response and / or (ii) an agent that binds to a T cell receptor (e.g., when presented by an MHC molecule) or an antibody. In some embodiments, an antigen elicits a humoral response (e.g., including the production of antigen-specific antibodies), and in some embodiments, an antigen elicits a cellular response (e.g., involving T cells having receptors that specifically interact with the antigen). In some embodiments, an antigen binds to an antibody, which may or may not elicit a specific physiological response in an organism. Generally, an antigen can be or include any chemical entity, such as, for example, a small molecule, nucleic acid, polypeptide, carbohydrate, lipid, polymer (in some embodiments, other than a biological polymer (e.g., other than a nucleic acid or amino acid polymer)), etc. In some embodiments, an antigen is or includes a polypeptide. In some embodiments, an antigen is or includes a glycan. Those skilled in the art will appreciate that, generally, antigens may be provided in isolated or pure form, or alternatively, may be provided in crude form (e.g., together with other substances, e.g., an extract such as a cell extract or other relatively crude preparation of an antigen-containing source). In some embodiments, antigens utilized in accordance with the present invention are provided in crude form. In some embodiments, the antigen is a recombinant antigen.
[0174] Emerging epitope: As used herein, the term "emerging epitope(s)" is used to refer to epitopes (e.g., portions of an antigen) that have not previously been encountered by a subject's immune system. For example, as circulating viral pathogens mutate, mutations in the various epitopes found on viral proteins may introduce one or more mutations, resulting in the emergence of epitopes that are versions of epitopes present on previously circulating variant proteins.
[0175] Composition: Those skilled in the art will understand that the term "composition" may be used to refer to a discrete physical entity that includes one or more specified components. Generally, unless otherwise specified, a composition may be in any form, examples of which include a gas, a gel, a liquid, a solid, etc.
[0176] Comprising: A composition or method described herein as "comprising" one or more specified elements or steps is open-ended, meaning that the specified elements or steps are essential, but that other elements or steps may be added within the composition or method. To avoid redundancy, any composition or method described as "comprising" (or "comprises") one or more specified elements or steps is also understood to describe a corresponding, more limited composition or method that "consistes essentially of" (or "consists essentially of") the same specified elements or steps. That is, the composition or method includes the specified essential elements or steps, and may also include additional elements or steps that do not materially affect the basic and novel characteristic(s) of the composition or method. It is also understood that any composition or method described herein as "comprising" or "consisting essentially of" one or more specified elements or steps also describes a corresponding, more specific, closed-ended composition or method that "consisting of" (or "consists of") the specified elements or steps, to the exclusion of any other unspecified elements or steps. In any composition or method disclosed herein, known or disclosed equivalents of any specified essential element or step may be substituted for that element or step.
[0177] Determining: In some embodiments, the methodologies described herein include a "determining" step. Those skilled in the art will understand, upon reading this specification, that such "determining" can utilize or be accomplished by using any of a variety of techniques available to those skilled in the art, including, for example, the specific techniques explicitly mentioned herein. In some embodiments, determining involves physical manipulation of the sample. In some embodiments, determining involves consideration and / or manipulation of data or information, for example, using a computer or other processing unit adapted to perform the relevant analysis. In some embodiments, determining involves receiving relevant information and / or material from a source. In some embodiments, determining involves comparing one or more characteristics of the sample or entity to a comparable reference.
[0178] Engineered antigen: As used herein, the term "engineered antigen" refers to an antigen that is artificially created and that is or is intended to be intentionally introduced into a subject (e.g., by vaccination), for example, to generate an immune response, as opposed to a naturally occurring antigen that has evolved through natural processes. For example, as described in further detail herein, in certain embodiments, an engineered antigen is or includes a polypeptide. An engineered polypeptide antigen may be designed to match, resemble, or be based on another reference antigen, which may itself be an engineered antigen or a naturally occurring antigen. For example, in certain embodiments, an engineered antigen is created by starting with the structure of a reference antigen and introducing one or more amino acid modifications therein to achieve a desired behavior / outcome upon deliberate introduction into a subject. In certain embodiments, an engineered antigen may be designed in silico (i.e., via computer-implemented systems and methods) using polypeptide models and other computational representations that represent various physical antigens. In certain embodiments, the engineered antigen may be encoded by ribonucleic acid (RNA), which may then be used to manufacture the engineered antigen (e.g., in vitro) or may be administered directly to a subject, for example, as an RNA vaccine.
[0179] Epitope: As used herein, the term "epitope" is used to refer to and includes any moiety that is specifically recognized by an immunoglobulin (e.g., antibody or receptor) binding component. In some embodiments, an epitope is composed of multiple chemical atoms or groups on an antigen. In some embodiments, such chemical atoms or groups are surface-exposed when the antigen adopts a suitable three-dimensional conformation. In some embodiments, such chemical atoms or groups are physically close to each other in space when the antigen adopts such a conformation. In some embodiments, at least some such chemical atoms or groups are physically separated from each other when the antigen adopts an alternative conformation (e.g., linearized).
[0180] Epitope Alteration Score: The terms "epitope alteration score" and "epitope score," used interchangeably herein, both refer to a measure of alteration of a viral polypeptide at an epitope location. In some embodiments, such alteration can be characterized by the effect of mutation(s) in one or more epitopes of a viral variant on recognition by antibodies (e.g., neutralizing antibodies). For example, in some embodiments, such alteration can be characterized by determining the number of potentially escaped antibodies. In some embodiments, the antibody being characterized is isolated from a patient who has been vaccinated against a disease or who has previously been infected with a disease (e.g., SARS-CoV-2). In some embodiments, the antibody being characterized has previously been shown to bind a reference sequence. In some embodiments, the epitope alteration score can be determined by comparing mutations in a candidate variant to one or more regions of a reference sequence previously shown to bind antibodies (e.g., by structural data). In some embodiments, an epitope alteration score can be determined by enumerating the number of unique epitopes that contain an altered position measured across one or more known antibody-viral polypeptide complex structures (e.g., all known antibody-viral polypeptide complex structures).
[0181] In some embodiments, the epitope alteration score is a measure of the number of different epitopes avoided by a candidate variant compared to a reference sequence (e.g., compared to a wild-type sequence). In some embodiments, the epitope alteration score is calculated based on known binding sites of antibodies, e.g., as reported in a protein data bank. In some embodiments, the epitope alteration score may change over time as new epitope locations are identified and / or epitope-binding antibodies are discovered. In some embodiments, the epitope alteration score can be used to characterize the degree of alteration of the SARS-CoV-2 spike polypeptide at an epitope location, for example, by counting the number or percentage of potentially escaped antibodies in some embodiments. In various embodiments described herein, the epitope alteration score can be normalized to rank from 0 to 100%.
[0182] Growth Score: As used interchangeably herein, the terms "growth," "growth metric," or "growth score" refer to a measure of the rate at which a given variant is growing in a subject population (e.g., at a given time). In some embodiments, the growth score refers to lineage-level growth. For example, in some embodiments, the growth score of a given variant can be determined by reference to the growth of a parent species, or a known variant of substantially the same lineage, or a known variant with a similar sequence (e.g., a sequence at least 90% identical to the given variant). In some embodiments, the growth of a given variant is a function of the change in the number of subjects in a subject population reported to be infected with the given variant over a given period relative to a baseline infection rate (e.g., a baseline infection rate determined over a defined period). In some embodiments, the growth of a given variant is a function of the change in the proportion of a subject population infected with the given variant over a given period relative to a baseline infection rate (e.g., a baseline infection rate determined over a defined period). In some embodiments, the growth score for a given variant can be empirically determined by considering sequences associated with the given variant (e.g., including, in some embodiments, sequences associated with a lineage) observed within a defined time period and calculating the ratio among all observed sequences at a given time point relative to a reference level (e.g., a ratio determined over a defined time period). For example, in some embodiments, for each lineage, the ratio of that sequence among all observed sequences is calculated for a long time period (e.g., an 8-week window) and a recent time window (e.g., the past 24 hours, the past 48 hours, the past 72 hours, the past 4 days, the past 5 days, the past 6 days, or the past week), represented as rextended and rlast, respectively. The growth of the lineage is defined by the ratio rextended / rlast, measuring the change in the ratio. In various embodiments described herein, growth scores can be normalized to rank from 0 to 100%.
[0183] Human: In some embodiments, the human is an embryo, fetus, infant, child, teenager, adult, or elderly.
[0184] As used herein, the terms "infectivity score" or "fitness prior score," used interchangeably with infectivity score or "fitness prior score," are a measure of the evolutionary fitness of a viral variant and are a function of how efficiently the virus replicates and / or infects host cells. In some embodiments, calculating the fitness prior score includes determining one or more of a log-likelihood score, a viral polypeptide receptor binding score, and / or a growth score. In some embodiments, the fitness prior score is determined by reference to each of the log-likelihood score, the viral polypeptide receptor binding score, and the growth score.
[0185] "Improve," "Increase," "Inhibit," or "Reduce": As used herein, the terms "improve," "increase," "inhibit," "reduce," or their grammatical equivalents refer to a value relative to a baseline or other reference measurement. In some embodiments, a suitable reference measurement may be or include a measurement in a particular system (e.g., a single individual) under otherwise comparable conditions in the absence (e.g., before and / or after) of a particular agent or treatment, or in the presence of an appropriate comparable reference agent. In some embodiments, a suitable reference measurement may be or include a measurement in a comparable system known or expected to respond in a particular way in the presence of the relevant agent or treatment.
[0186] Log-Likelihood: As used herein, the term "log-likelihood" refers to a measure of the probability of existence of a variant polypeptide sequence determined using a natural language processing algorithm. In some embodiments, the log-likelihood can be determined using a Transformer model. In some embodiments, the log-likelihood can be determined without a reference sequence. In some embodiments, the log-likelihood is a Transformer-derived log-likelihood without a reference. The higher the log-likelihood of a variant, the more likely that variant occurs from the perspective of the language model. In various embodiments described herein, the log-likelihood can be normalized to rank from 0 to 100%. In some embodiments, the log-likelihood measures how the log-likelihood of a variant polypeptide sequence compares to the entire population of known variants. In some embodiments, the log-likelihood measures how the log-likelihood of a variant polypeptide sequence compares to other variants with similar mutational loads ("conditional log-likelihood"). Such conditional log-likelihoods are particularly useful for assessing variants with a large number of mutations (e.g., at least 30 or more, including, for example, at least 40, at least 50, at least 60, at least 70, or more mutations).
[0187] Machine Learning Module, Machine Learning Model: As used herein, the terms “machine learning module” and “machine learning model” are used interchangeably and refer to a computer-implemented process (e.g., software function) that implements one or more specific machine learning algorithms (e.g., artificial neural network (ANN), random forest, decision tree, support vector machine, etc.) to determine one or more output values for a given input. In certain embodiments, the machine learning model is a deep learning model or deep neural network (i.e., ANN) that includes an input layer and an output layer as well as one or more hidden layers (e.g., layers in between). Examples of deep learning models include, but are not limited to, recurrent neural networks (RNNs) (e.g., long short-term memory networks (LSTMs), bidirectional LSTMs (biLSTMs)), attention-based networks (e.g., Transformer models), and convolutional neural networks (CNNs). In some embodiments, the machine learning module that implements the machine learning techniques is trained in a supervised manner, e.g., using curated and / or manually annotated datasets. In certain embodiments, the machine learning model may be trained in an unsupervised manner using unlabeled data. In certain embodiments, the machine learning model may be trained using a reinforcement approach, e.g., a reward / penalty system may be used to train the machine learning model to learn strategies for accomplishing a specified task. Training the machine learning model may be used to determine various parameters of the model, an example of which is the weights associated with layers of a neural network. In some embodiments, once the machine learning module is trained to accomplish a specific task, such as predicting the type of hidden amino acid in a polypeptide sequence based on its context, the determined parameter values are fixed, and the (e.g., immutable and static) machine learning module is used to process new data (e.g., different from the training data), such as new amino acid sequences, referred to as inference.In some embodiments, the machine learning module may receive feedback, e.g., based on user accuracy reviews, and may use such feedback as additional training data, e.g., to dynamically update the machine learning module. In some embodiments, the trained machine learning module is a classification algorithm (e.g., a random forest classifier) with adjustable and / or fixed (e.g., locked) parameters. In some embodiments, two or more machine learning modules may be combined and implemented as a single module and / or a single software application. In some embodiments, two or more machine learning modules may be implemented separately, e.g., as separate software applications. The machine learning modules may be software and / or hardware. For example, the machine learning module may be implemented entirely as software, or certain functions of the ANN module may be performed by specialized hardware (e.g., by an application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), etc.).
[0188] Model: As used herein, the term "model" is used to identify a computer representation of a particular physical object or quantity, examples of which include a physical object or quantity accessed by, displayed by, used as input to, generated as output from, etc., computer-implemented methods and systems, and / or one or more processes and / or modules or functions thereof. For example, as described in more detail herein, various computer-implemented systems and methods may operate on, process, and generate polypeptide models that represent physical polypeptides (e.g., particular proteins or portions thereof). A computer representation of a polypeptide (e.g., a polypeptide model) may be embodied in a variety of formats, examples of which include a string (e.g., letters each representing a particular type of amino acid) representing an amino acid sequence (e.g., a FASTA file), or a 3D structural model containing information about the 3D arrangement (e.g., relative) of amino acids and / or their atoms, examples of which include the Protein Data Bank (PDB) file format, which may be used to describe the 3D structure of a particular protein and contains, among other things, the atomic coordinates of the atoms of a particular protein.
[0189] Nucleic Acid: As used herein, the term "nucleic acid" in its broadest sense refers to any compound and / or substance that is or can be incorporated into an oligonucleotide chain. In some embodiments, a nucleic acid is a compound and / or substance that is or can be incorporated into an oligonucleotide chain via a phosphodiester bond. As is clear from the context, in some embodiments, "nucleic acid" refers to individual nucleic acid residues (e.g., nucleotides and / or nucleosides), and in some embodiments, "nucleic acid" refers to an oligonucleotide chain comprising individual nucleic acid residues. In some embodiments, "nucleic acid" is or includes RNA, and in some embodiments, "nucleic acid" is or includes DNA. In some embodiments, a nucleic acid is, comprises, or consists of one or more naturally occurring nucleic acid residues. In some embodiments, a nucleic acid is, comprises, or consists of one or more nucleic acid analogs. In some embodiments, a nucleic acid analog differs from a nucleic acid in that it does not utilize a phosphodiester backbone. For example, in some embodiments, the nucleic acid is, comprises, or consists of one or more "peptide nucleic acids," which are known in the art and have peptide bonds instead of phosphodiester bonds in the backbone, and are considered within the scope of the present disclosure. Alternatively or additionally, in some embodiments, the nucleic acid has one or more phosphorothioate and / or 5'-N-phosphoramidite linkages rather than phosphodiester linkages. In some embodiments, the nucleic acid is, comprises, or consists of one or more naturally occurring nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine).In some embodiments, the nucleic acid is, comprises, or consists of one or more nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, C-5-propynylcytidine, C-5-propynyluridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyluridine, C5-propynylcytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazaadenosine, 7-deazaguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, 2-thiocytidine, methylated bases, intercalated bases, and combinations thereof). In some embodiments, the nucleic acid comprises one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) compared to naturally occurring nucleic acids. In some embodiments, the nucleic acid has a nucleotide sequence that encodes a functional gene product (such as RNA or a protein). In some embodiments, the nucleic acid comprises one or more introns. In some embodiments, the nucleic acid is prepared by one or more of isolation from a natural source, enzymatic synthesis by polymerization based on a complementary template (in vivo or in vitro), replication in a recombinant cell or system, and chemical synthesis. In some embodiments, the nucleic acid is at least 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, 180, 190, 20, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000 or more residues in length. In some embodiments, the nucleic acid is partially or entirely single-stranded, and in some embodiments, the nucleic acid is partially or entirely double-stranded.In some embodiments, the nucleic acid has a nucleotide sequence comprising at least one element that encodes a polypeptide or is the complement of a sequence that encodes it. In some embodiments, the nucleic acid has enzymatic activity.
[0190] Pareto score: As used herein, the term "Pareto score" refers to a measure of a variant's performance relative to or assessed by one or more scoring metrics, examples of which include the various scoring metrics described herein and / or combinations thereof. These may include, but are not necessarily limited to, scores that assess fitness and ability to evade an immune response. In some embodiments, a Pareto score comprises a combination of an immune evasion score (e.g., as described herein) and a fitness prior score (e.g., as described herein). In some embodiments, a Pareto score captures the relative evolutionary advantage of a given strain. In some embodiments, such a Pareto score can be determined as described in the Examples. In some embodiments, a Pareto score is an optimality score, e.g., in some embodiments, it ranks variants relative to other sequences (e.g., sequences observed in a population). For a specific lineage, a high Pareto score at a given time point indicates that there are fewer variants with higher fitness prior scores and higher immune evasion scores at that time point. The Pareto score in some embodiments is a ranking system, and the Pareto score for a given variant may change over time because the fitness prior score and immune escape score incorporated therein may change as new data is acquired. As used herein, in some embodiments, Pareto optimality is defined over a set of lineages. In some embodiments, a lineage is Pareto optimal within a set if no lineage in the set has a higher immune escape score and a higher fitness prior score. In some embodiments, the Pareto score is a measure of the degree of Pareto optimality. For example, in some embodiments, the lineage with the highest Pareto score is Pareto optimal, and if the Pareto optimal lineage is removed from the set, etc., the lineage with the second-best Pareto score becomes Pareto optimal.
[0191] Patient: As used herein, the term "patient" refers to any organism to which provided compositions are or may be administered, e.g., for experimental, diagnostic, prophylactic, cosmetic, and / or therapeutic purposes. Typical patients include animals (e.g., mammals such as mice, rats, rabbits, non-human primates, and / or humans). In some embodiments, the patient is human. In some embodiments, the patient is suffering from or susceptible to one or more disorders or conditions. In some embodiments, the patient exhibits one or more symptoms of the disorder or condition. In some embodiments, the patient has been diagnosed with one or more disorders or conditions. In some embodiments, the disorder or condition is or includes a viral infection (e.g., SARS-CoV-2 infection). In some embodiments, the patient is undergoing or has undergone a specific therapy to diagnose and / or treat the disease, disorder, or condition.
[0192] Peptide: As used herein, the term "peptide" refers to a polypeptide that is typically relatively short, examples of which include those having a length of less than about 100 amino acids, less than about 50 amino acids, less than about 40 amino acids, less than about 30 amino acids, less than about 25 amino acids, less than about 20 amino acids, less than about 15 amino acids, or less than 10 amino acids.
[0193] Pharmaceutical composition: As used herein, the term "pharmaceutical composition" refers to an active agent formulated together with one or more pharmaceutically acceptable carriers. In some embodiments, the active agent is present in a unit dose suitable for administration in a treatment regimen that exhibits a statistically significant probability of achieving a predetermined therapeutic effect when administered to a relevant population. In some embodiments, pharmaceutical compositions may be specially formulated for administration in solid or liquid form, including those adapted for oral administration (e.g., drenches (aqueous or non-aqueous solutions or suspensions), tablets (e.g., buccal, sublingual, and those targeted for systemic absorption), boluses, powders, granules, and pastes for application to the tongue); parenteral administration (e.g., as a sterile solution or suspension, or as a sustained-release formulation, e.g., by subcutaneous, intramuscular, intravenous, or epidural injection); topical application (e.g., as a cream, ointment, or controlled-release patch or spray applied to the skin, lungs, or oral cavity); vaginal or rectal (e.g., as a pessary, cream, or foam); sublingual; ophthalmic; transdermal; or nasal, pulmonary, and other mucosal surfaces.
[0194] Pharmaceutically acceptable: As used herein, the term “pharmaceutically acceptable” as applied to a carrier, diluent, or excipient used to formulate a composition disclosed herein means that the carrier, diluent, or excipient must be compatible with the other ingredients of the composition and not deleterious to the recipient thereof.
[0195] Pharmaceutically acceptable carrier: As used herein, the term "pharmaceutically acceptable carrier" means a pharmaceutically acceptable material, composition, or vehicle, examples of which include liquid or solid fillers, diluents, excipients, or solvent encapsulating materials, which are involved in carrying or transporting a compound of interest from one organ or body part to another. Each carrier must be "acceptable" in the sense of being compatible with the other ingredients of the formulation and not injurious to the patient. Some examples of materials that can serve as pharmaceutically acceptable carriers include sugars such as lactose, glucose, and sucrose; starches such as corn starch and potato starch; cellulose and its derivatives, for example, sodium carboxymethylcellulose, ethylcellulose, and cellulose acetate; powdered tragacanth; malt; gelatin; talc; excipients such as cocoa butter and suppository waxes; oils such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; glycols such as propylene glycol; polyols such as glycerin, sorbitol, mannitol, and polyethylene glycol; esters such as ethyl oleate and ethyl laurate; agar; buffering agents such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline; Ringer's solution; ethyl alcohol; pH buffer solutions; polyesters, polycarbonates, and / or polyanhydrides; and other non-toxic, compatible substances employed in pharmaceutical formulations.
[0196] Pharmaceutical Grade: As used herein, the term "pharmaceutical grade" refers to the standards established by generally recognized national or regional pharmacopoeias (e.g., the United States Pharmacopeia and National Formulary (USP-NF)) for chemical and biological drug substances, drug products, dosage forms, compounded preparations, excipients, medical devices, and dietary supplements.
[0197] Polypeptide: As used herein, refers to a polymeric chain of amino acids. In some embodiments, a polypeptide has a naturally occurring amino acid sequence. In some embodiments, a polypeptide has a non-naturally occurring amino acid sequence. In some embodiments, a polypeptide has an engineered amino acid sequence, in that it has been designed and / or produced by the act of the human hand. In some embodiments, a polypeptide may comprise or be composed of natural amino acids, unnatural amino acids, or both. In some embodiments, a polypeptide may comprise or be composed of only natural amino acids or only unnatural amino acids. In some embodiments, a polypeptide may comprise D-amino acids, L-amino acids, or both. In some embodiments, a polypeptide may comprise only D-amino acids. In some embodiments, a polypeptide may comprise only L-amino acids. In some embodiments, a polypeptide may comprise one or more pendant groups or other modifications (e.g., modification of or attachment to one or more amino acid side chains) at the N-terminus of the polypeptide, the C-terminus of the polypeptide, or any combination thereof. In some embodiments, such pendant groups or modifications may be selected from the group consisting of acetylation, amidation, lipidation, methylation, PEGylation, etc., and combinations thereof. In some embodiments, a polypeptide may be cyclic and / or include a cyclic portion. In some embodiments, a polypeptide is not cyclic and / or does not include a cyclic portion. In some embodiments, a polypeptide is linear. In some embodiments, a polypeptide may be or include a stapled polypeptide. In some embodiments, the term "polypeptide" may be appended to the name of a reference polypeptide, activity, or structure. In such instances, it is used herein to refer to polypeptides that share a related activity or structure and therefore can be considered members of the same class or family of polypeptides.For each such class, the present specification provides, and / or one of skill in the art will recognize, exemplary polypeptides within the class whose amino acid sequence and / or function are known. In some embodiments, such exemplary polypeptides are reference polypeptides of a class or family of polypeptides. In some embodiments, members of a class or family of polypeptides exhibit significant sequence homology or identity with the reference polypeptide of the class (and, in some embodiments, with all polypeptides in the class), share common sequence motifs (e.g., characteristic sequence elements), and / or share a common activity (in some embodiments, at a similar level or within a specified range) with the reference polypeptide of the class (and, in some embodiments, with all polypeptides in the class). For example, in some embodiments, a member polypeptide exhibits an overall degree of sequence homology or identity with a reference polypeptide of at least about 30-40%, often greater than about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99%, or more, and / or comprises at least one region (e.g., a conserved region which in some embodiments may be or include a distinctive sequence element) exhibiting very high sequence identity, often greater than 90%, or even greater than 95%, 96%, 97%, 98%, or 99%. Such conserved regions typically encompass at least 3-4, and often up to 20 or more amino acids, and in some embodiments, the conserved region encompasses at least one stretch of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, or more consecutive amino acids. In some embodiments, the related polypeptide may comprise or consist of a fragment of a parent polypeptide.In some embodiments, a useful polypeptide may comprise or consist of multiple fragments, each fragment being found in the same parent polypeptide in a different spatial arrangement relative to that found in the polypeptide of interest (e.g., fragments that are directly linked in the parent may be spatially separated in the polypeptide of interest, or vice versa, and / or fragments may be present in a different order in the polypeptide of interest than in the parent), such that the polypeptide of interest is a derivative of that parent polypeptide.
[0198] Prevention or prevention: As used herein when used in relation to the occurrence of a disease, disorder, and / or condition, refers to reducing the risk of developing a disease, disorder, and / or condition and / or delaying the onset of one or more characteristics or symptoms of a disease, disorder, or condition. Prevention may be considered complete if the onset of the disease, disorder, or condition has been delayed for a predetermined period of time.
[0199] Ribonucleotide: As used herein, the term "ribonucleotide" encompasses unmodified and modified ribonucleotides. For example, unmodified ribonucleotides include the purine bases adenine (A) and guanine (G) and the pyrimidine bases cytosine (C) and uracil (U). Modified ribonucleotides may contain one or more modifications, including, but not limited to, (a) terminal modifications, including 5'-terminal modifications (e.g., phosphorylation, dephosphorylation, conjugation, reverse linkage, etc.) and 3'-terminal modifications (e.g., conjugation, reverse linkage, etc.); (b) base modifications, including substitution with modified bases, stabilizing bases, destabilizing bases, or bases that base-pair with an expanded repertoire of partners, or conjugated bases; (c) sugar modifications (e.g., at the 2' or 4' position) or sugar substitutions; and (d) internucleoside linkage modifications, including modifications or substitutions of phosphodiester bonds. The term "ribonucleotide" also encompasses ribonucleotide triphosphates, including modified and unmodified ribonucleotide triphosphates.
[0200] Ribonucleic acid (RNA): As used herein, the term "RNA" refers to a polymer of ribonucleotides. In some embodiments, the RNA is single-stranded. In some embodiments, the RNA is double-stranded. In some embodiments, the RNA includes both single-stranded and double-stranded portions. In some embodiments, the RNA can include a backbone structure described in the definition of "nucleic acid / polynucleotide" above. The RNA can be a regulatory RNA (e.g., siRNA, microRNA, etc.) or a messenger RNA (mRNA). In some embodiments, the RNA is an mRNA. In some embodiments, the RNA typically includes a poly(A) region at its 3' end. In some embodiments, the RNA is an mRNA, the RNA typically includes an art-recognized cap structure at its 5' end, for example, to recognize and attach the mRNA to a ribosome to initiate translation. In some embodiments, the RNA is synthetic RNA. Synthetic RNA includes RNA synthesized in vitro (e.g., by enzymatic and / or chemical synthesis). In some embodiments, the RNA is single-stranded RNA. In some embodiments, the single-stranded RNA may contain self-complementary elements and / or establish secondary and / or tertiary structure. Those skilled in the art will understand that when a single-stranded RNA is referred to as "encoding," it means that it either contains the encoding nucleic acid sequence itself or that it contains the complement of the encoding nucleic acid sequence. In some embodiments, the single-stranded RNA can be self-amplifying RNA (also known as self-replicating RNA).
[0201] Semantic change: As used herein, the term "semantic change" refers to a measure of functional change, from the perspective of a language model, of a variant viral polypeptide (e.g., in some embodiments, viral polypeptides that interact with a host cell receptor and / or are otherwise involved in host cell entry) relative to at least one or more (e.g., at least two, at least three, at least four or more) reference viral polypeptide(s) (e.g., in some embodiments, reference viral polypeptides of a wild-type species and / or known variants, e.g., reference viral polypeptides of the same lineage). In some embodiments, semantic change is a measure of functional change, from the perspective of a language model, of a variant viral polypeptide (e.g., in some embodiments, viral polypeptides that interact with a host cell receptor and / or are otherwise involved in host cell entry) relative to multiple (e.g., at least two or more) reference viral polypeptide(s) (e.g., in some embodiments, reference viral polypeptides of a wild-type species and / or known variants, e.g., reference viral polypeptides of the same lineage). In some embodiments, the associated language model can include transformer-derived embedded differences (e.g., as described herein) for at least one or more (e.g., at least two, at least three, at least four, or more) reference viral polypeptide(s) (e.g., in some embodiments, reference viral polypeptides of a wild-type species or known variants, e.g., reference viral polypeptides of the same lineage). In some embodiments, the semantic change score can be calculated using the LI norm. In some embodiments, the semantic change score can be calculated using the L2 norm (also called the Euclidean norm). In some embodiments, the semantic change describes how different the variants are with respect to an underlying statistical model (e.g., in some embodiments, a large-scale machine learning model fine-tuned based on viral protein sequences observed up to a given time point).In some embodiments, the semantic variation score depends on the observed sequence, and therefore may change over time as the underlying model is trained with new variant and / or reference sequences. In some embodiments, a semantic variation score for a variant spike polypeptide from SARS-CoV-2 is determined as described herein. In various embodiments described herein, the semantic variation score can be normalized to rank from 0 to 100%.
[0202] Subject: As used herein, the term "subject" refers to an organism, typically a mammal (e.g., a human, including in some embodiments prenatal human forms). In some embodiments, the subject is afflicted with the relevant disease, disorder, or condition. In some embodiments, the subject is predisposed to the disease, disorder, or condition. In some embodiments, the subject exhibits one or more symptoms or characteristics of the disease, disorder, or condition. In some embodiments, the subject does not exhibit any symptoms or characteristics of the disease, disorder, or condition. In some embodiments, the subject is a person possessing one or more characteristics characteristic of susceptibility or risk for a disease, disorder, or condition. In some embodiments, the subject is a patient. In some embodiments, the subject is an individual for whom and / or to whom diagnosis and / or treatment is administered.
[0203] Substantially: As used herein, the term "substantially" refers to a qualitative state of exhibiting a total or near-total extent or degree of a property or characteristic of interest. Those skilled in the art of biology will understand that biological and chemical phenomena rarely proceed to completion and / or perfection, or achieve or avoid absolute results. Thus, the term "substantially" is used herein to capture the potential lack of perfection inherent in many biological and chemical phenomena.
[0204] Variant: As used herein, in the context of a molecule (e.g., a nucleic acid, protein, or small molecule), the term "variant" refers to a molecule that exhibits significant structural identity with a reference molecule but differs structurally from the reference molecule, e.g., the presence, absence, or level of one or more chemical moieties compared to the reference entity. In some embodiments, a variant also differs functionally from its reference molecule. Typically, whether a particular molecule is properly considered a "variant" of a reference molecule is based on its degree of structural identity with the reference molecule. As will be understood by those skilled in the art, every biological or chemical reference molecule possesses certain characteristic structural elements. By definition, a variant is a distinct molecule that shares one or more such characteristic structural elements but differs in at least one aspect from the reference molecule. In some embodiments, a variant polypeptide or nucleic acid differs from a reference polypeptide or nucleic acid as a result of one or more differences in amino acid or nucleotide sequence and / or differences in chemical moieties (e.g., carbohydrates, lipids, phosphate groups) that are covalent components of the polypeptide or nucleic acid (e.g., attached to the backbone of the polypeptide or nucleic acid). In some embodiments, a variant polypeptide or nucleic acid exhibits at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99% overall sequence identity with a reference polypeptide or nucleic acid. In some embodiments, a variant polypeptide or nucleic acid does not share at least one characteristic sequence element with a reference polypeptide or nucleic acid. In some embodiments, a reference polypeptide or nucleic acid has one or more biological activities. In some embodiments, a variant polypeptide or nucleic acid shares one or more of the biological activities of a reference polypeptide or nucleic acid. In some embodiments, a variant polypeptide or nucleic acid lacks one or more of the biological activities of a reference polypeptide or nucleic acid.In some embodiments, a variant polypeptide or nucleic acid has a reduced level of one or more biological activities compared to a reference polypeptide or nucleic acid. In some embodiments, when a polypeptide or nucleic acid of interest is considered a "variant" of a reference polypeptide or nucleic acid, it has the same amino acid sequence or nucleotide sequence as the reference, but has a small number of sequence changes at specific positions. Typically, a variant has about 20%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, or about 2% of residues substituted, inserted, or deleted compared to the reference. In some embodiments, a variant polypeptide or nucleic acid has about 10, about 9, about 8, about 7, about 6, about 5, about 4, about 3, about 2, or about 1 substituted residue compared to the reference. Often, a variant polypeptide or nucleic acid has a very small number (e.g., less than about 5, about 4, about 3, about 2, or about 1) of functional residues (i.e., residues involved in a particular biological activity) substituted, inserted, or deleted relative to the reference. In some embodiments, the variant polypeptide or nucleic acid contains no more than about 5, no more than about 4, no more than about 3, no more than about 2, or no more than about 1 additions or deletions relative to the reference. In some embodiments, the variant polypeptide or nucleic acid contains no more than about 25, no more than about 20, no more than about 19, no more than about 18, no more than about 17, no more than about 16, no more than about 15, no more than about 14, no more than about 13, no more than about 10, no more than about 9, no more than about 8, no more than about 7, no more than about 6, and typically no more than about 5, no more than about 4, no more than about 3, or no more than about 2 additions or deletions relative to the reference. In some embodiments, the reference polypeptide or nucleic acid is one found in nature.
[0205] Vaccination: As used herein, the term "vaccination" refers to the administration of a composition intended to generate an immune response against, for example, a disease (e.g., a viral epitope). In some embodiments, vaccination can be performed before, during, and / or after the onset of disease. In some embodiments, vaccination involves multiple administrations of the vaccinating composition, appropriately spaced in time.
[0206] Viral Polypeptide Receptor Binding Score: As used herein, the term "viral polypeptide receptor binding score" refers to a measure of binding affinity between a viral polypeptide that plays a role in host recognition and / or host cell entry and a corresponding host protein with which the viral polypeptide interacts to recognize and / or enter host cells. In some embodiments, the viral polypeptide receptor binding score is determined in silico. In some embodiments, a conformational sampling algorithm can be used to determine the viral polypeptide receptor binding score. In some embodiments, the viral polypeptide receptor binding score can be determined using structures optimized using a stochastic optimization algorithm (e.g., a variant of simulated annealing, which aims to overcome local energy barriers and follow kinematically accessible paths toward achievable deep energy minima with respect to knowledge-based, protein-directed potentials). In some embodiments, the viral polypeptide receptor binding score can be calculated using the change in solvent-accessible surface area (SASA) of the viral polypeptide in complex (e.g., bound) and uncomplexed (e.g., unbound) states. In some embodiments, the viral polypeptide receptor binding score can be determined by calculating the energy change between a complexed (e.g., bound) and uncomplexed (e.g., unbound) structure of the viral polypeptide and its cognate host receptor. In some embodiments, the change in binding energy can be estimated by the difference in Gibbs free energy between the bound and unbound states. In various embodiments described herein, the viral polypeptide receptor binding score can be normalized to rank from 0 to 100%. In some embodiments, the viral polypeptide receptor binding score can be calculated in silico, for example, by calculating the change in Gibbs free energy or the change in solvent-accessible surface area between the bound and unbound states.In some embodiments, the viral polypeptide receptor binding score can be calculated using in vitro binding data (e.g., using the dissociation constant, KD, or association rate, k0n). In some embodiments, such in vitro binding data can be determined by methods known in the art, including, but not limited to, biolayer interferometry (BLI) and / or surface plasmon resonance (SPR).
[0207] ACE2 Binding Score: As used herein, the term "ACE2 binding score" refers to the viral polypeptide receptor binding score (described herein) when the viral polypeptide receptor is angiotensin-converting enzyme 2 (ACE2). The "ACE2 binding score" is a measure of the binding affinity between the S protein of a coronavirus (e.g., SARS-CoV-2), or an immunogenic fragment of the S protein (e.g., the RBD domain), and the ACE2 protein. In some embodiments, the ACE2 binding score can be calculated in silico, for example, by calculating the change in Gibbs free energy or the change in solvent-accessible surface area between the bound and unbound states. In some embodiments, the ACE2 binding score is calculated using in vitro binding data (e.g., the dissociation constant, KD, or the association rate, k). on In some embodiments, such in vitro binding data can be determined by methods known in the art, including, but not limited to, biolayer interferometry (BLI) and / or surface plasmon resonance (SPR).
[0208] Wild-type: As used herein, the term "wild-type" has its art-understood meaning and refers to an entity having a structure and / or activity found in nature in a "normal" (as opposed to mutated, diseased, altered, etc.) state or situation. Those of skill in the art will understand that wild-type genes and polypeptides often exist in multiple alternative forms (e.g., alleles). In some embodiments, in the context of SARS-CoV-2, "wild-type" refers to the Wuhan variant.
[0209] Detailed Description The systems, architectures, devices, methods, and processes of the claimed inventions are intended to encompass variations and adaptations developed using information from the embodiments described herein. Adaptations and / or modifications of the systems, architectures, devices, methods, and processes described herein may be implemented as contemplated by this description.
[0210] Throughout the description, when articles, devices, systems and architectures are described as having, including, or comprising specific components, or processes and methods are described as having, including, or comprising specific steps, it is additionally contemplated that there are articles, devices, systems and architectures of the invention that consist essentially of, or consist of, the recited components, and that there are processes and methods of the invention that consist essentially of, or consist of, the recited process steps.
[0211] It should be understood that the order of steps or order for performing certain actions is immaterial so long as the invention remains operable. Moreover, two or more steps or actions may be conducted simultaneously.
[0212] The citation of any publication herein, for example in the Background section, is not an admission that the publication serves as prior art with respect to any of the claims presented herein. The Background section is presented for purposes of clarity and is not intended to be a description of prior art with respect to any claim.
[0213] Documents are incorporated herein by reference as indicated. In the event of any discrepancy in the meaning of a particular term, the meaning provided in the Definitions section above shall prevail.
[0214] Headings are provided for the convenience of the reader, and the presence and / or placement of a heading is not intended to limit the scope of the subject matter described herein.
[0215] Among other things, the present disclosure provides systems and methods for designing engineered antigens with tailored immunological characteristics. In certain embodiments, the engineered antigens generated by the techniques described herein are computer-engineered versions of reference antigens specifically designed to, for example, interact with a subject's immune system in a desired manner.
[0216] For example, in certain embodiments, the reference antigen may be a naturally occurring protein, an example of which is a variant of a particular viral protein, or a portion thereof. Antigen engineering techniques, as described in further detail herein, may be used to design custom-adapted versions of such reference antigens to promote the production of new antibodies and / or reduce the likelihood and / or extent of eliciting a memory immune response, which may result in the generation of antibodies from memory B cells resulting from previous exposure to the reference antigen variant. Thus, new antibodies produced in this manner may be adapted to, for example, a specific version (e.g., mutation) of an evolved epitope present in the reference antigen. In contrast, antibodies generated from a memory response may effectively target a similar epitope on a previous variant, but may be avoided by a newly evolved epitope on the specific reference antigen.
[0217] Without wishing to be bound by any particular theory, we propose that it may be desirable to vaccinate with antigenic polypeptides designed to promote an immune response, and in particular an antibody response (e.g., a neutralizing antibody response), against epitopes occurring in the variant polypeptides. In particular, the present disclosure provides the insight that, particularly in the case of circulating infectious diseases (e.g., where variants are expected to emerge), it may be particularly desirable to promote an immune response that specifically includes an antibody response (e.g., a neutralizing response) against the evolved epitopes.
[0218] Among other things, the present disclosure provides techniques for engineering antigens with reduced risk of eliciting memory immune response(s) as described herein by identifying and mutating (e.g., introducing amino acid modifications) portions of a reference antigen (e.g., sequence, set of specific amino acid sites determined to be members of a discriminatory surface region) determined to be likely to contain one or more shared epitope(s) and / or (e.g., contiguous) conserved surface. As described in further detail herein, the particular approaches described herein include the recognition and establishment of various criteria related to a specific distribution of amino acid modifications, sources and rules for generating amino acid modifications, and approaches for evaluating the expected performance of engineered antigen designs aimed at sufficiently disrupting the memory-eliciting portions of the input reference antigen while simultaneously preserving its characteristics (e.g., stability, 3D structure (e.g., folding), etc.).
[0219] The present disclosure illustrates certain aspects of the provided technology through approaches for designing engineered versions of recently evolved XBB SARS-CoV2 spike protein variants. While some examples provided herein are described with reference to particular proteins and viral variants, those skilled in the art will understand, upon reading this disclosure, that the approaches described herein may be applied and employed, for example, to other pathogens (e.g., types of viruses), proteins, subregions, etc. to promote the production of a desired immune response.
[0220] A. Natural Evolution of Infectious Agents Immune imprinting is a phenomenon in which early exposure to a particular antigen can limit (e.g., subsequent) development of an immune response to epitopes specific to new variants of that antigen. In particular, upon exposure to a new, previously unencountered infectious agent (e.g., a virus), the immune system responds by producing highly specific antibodies that, among other things, bind to and neutralize portions of the agent's antigen(s). The immune system then retains a "memory" of the antigen(s) as well as the ability to produce specific antibodies that target them in the form of memory B cells and memory T cells.
[0221] On the one hand, after initial exposure to a particular agent, this immune memory allows the body to rapidly recognize and defend against that agent on subsequent encounters. On the other hand, processes such as natural mutation and evolution can result in variants of the agent that are similar enough to the initially encountered strain that they are recognized and trigger a memory response, prompting the production of antibodies designed to protect against the original strain rather than antibodies specifically adapted to the new variant. The effectiveness of this memory response can be reduced if the new variant contains sufficient mutations in critical regions targeted by these antibodies (e.g., regions involved in host cell infection, viral replication, etc.). Therefore, immune imprinting may be of particular concern for pathogens with a high concentration of mutations in neutralization-sensitive epitopes.
[0222] In the context of viral infection and vaccination, this immune imprinting phenomenon can increase the rate of reinfection by mutated variants and limit the effectiveness of vaccination in individuals who have been initially exposed to previous strains (e.g., due to natural infection or previous vaccination). Therefore, immune imprinting can be particularly problematic for vaccinations against viruses with higher mutation rates (e.g., RNA viruses). These include, but are not limited to, influenza, coronaviruses (e.g., severe acute respiratory syndrome-associated coronavirus), human immunodeficiency virus (HIV), respiratory syncytial virus (RSV), and the like.
[0223] For example, in the context of the recent SARS-CoV 2 pandemic, mutations in circulating viruses have given rise to tens of thousands of viral variants, some of which (such as Omicron and the recently emerged XBB (e.g., XBB.1.5)) are characterized by immune escape capabilities. In particular, these variants contain several mutations that enable them to evade pre-existing immune responses (e.g., memory immune responses) developed by individuals as a result of previous exposure (either by vaccination and / or natural infection) to previous strains (the original wild-type (WT) Wuhan variant).
[0224] Among other things, the present disclosure recognizes that, for example, XBB contains mutations across several epitopes. While many of these mutations are shared with other earlier variants, a subset are unique to XBB. When these shared, pre-existing epitopes trigger a memory immune response, antibodies adapted to the previous variants to which an individual was exposed are produced, instead of new antibodies explicitly designed to neutralize XBB.
[0225] Thus, without wishing to be bound by any particular theory, the systems and methods described herein include approaches for designing engineered antigens that reduce the extent to which memory responses are elicited by shared epitopes of the new antigen and encourage the immune system to generate novel responses to specific target epitopes unique to the new antigen.
[0226] B. Engineered Antigens 1, the present disclosure provides, among other things, systems and methods for the in silico design of engineered antigens designed to promote an immune response to specific epitopes of a reference antigen of an infectious agent. While approaches for engineering antigens are often described herein with reference to viruses and their proteins, they may also be utilized to design engineered antigens based on reference antigens derived from or associated with other types of infectious agents (e.g., bacteria, parasites (malaria), etc.).
[0227] i. Reference antigens and their computer representations The reference antigen may be a protein and / or portion thereof of (e.g., or produced by) the infectious agent. For example, in certain embodiments, the reference antigen is a specific viral protein and / or portion thereof, such as a surface protein. In certain embodiments, the reference antigen is a protein or portion of a protein that has been determined to be utilized by the virus to infect host cells (e.g., involved in binding to a specific host cell protein and / or facilitating fusion with the host cell membrane). In certain embodiments, the reference antigen is or includes a specific portion (e.g., a subunit, domain, etc.) of a viral protein.
[0228] For example, in certain embodiments, the infectious agent is or includes a coronavirus (such as SARS-CoV 2), and the reference antigen is the SARS-CoV 2 spike protein or a portion of the SARS-CoV 2 spike protein. For example, the reference antigen can be the entire spike protein, or a specific portion such as the N-terminal region (e.g., the N-terminal domain (NTD)) or the receptor-binding domain (RBD). In certain embodiments, the specific portion is selected to focus on the minimal relevant vaccine antigen, e.g., to facilitate removal of as many conserved epitopes as possible, e.g., without resorting to introducing point mutations (e.g., thereby limiting the number of epitopes into which point mutations are introduced).
[0229] As described herein, in certain embodiments, antigen engineering techniques of the present disclosure are directed to designing an engineered version of a reference antigen that, when introduced (e.g., administered) to a subject, prompts their immune system to generate new antibodies specifically tailored to specific epitopes of the reference antigen. In certain embodiments, this entails reducing the likelihood and / or extent to which the engineered version of the reference antigen will elicit a memory immune response, thereby ameliorating certain obstacles that the immune imprinting phenomenon may cause, e.g., with respect to vaccination, as described herein.
[0230] Thus, in certain embodiments, the systems and methods described herein identify remaining portions of a particular reference antigen that are likely to elicit a memory immune response and generate in silico one or more engineered variants in which these memory-eliciting subregion(s) are disrupted.
[0231] 1 illustrates an exemplary process 100 for generating an engineered antigen according to certain embodiments. In certain embodiments, the exemplary process 100 begins by accessing, generating, or otherwise obtaining a computer representation of at least a portion of a reference antigen of interest, referred to herein as a polypeptide model 102.
[0232] A variety of formats (e.g., data structures) may be used to represent a reference antigen (e.g., a particular protein and / or portion thereof) in a computer. For example, a polypeptide model 102 may be or include an amino acid sequence (e.g., an ordered string of characters), each representing a particular amino acid (e.g., a FASTA file). In certain embodiments, a polypeptide model may be or include a structural model, an example of which is a 3D model that includes a representation of the placement of each amino acid (or atom thereof) of a reference antigen in 3D space. For example, Protein Data Bank (PDB) format files may be used and / or accessed to obtain 3D structural information about a particular protein (e.g., based on a derived crystal structure).
[0233] In certain embodiments, the reference antigen may be a protein of a particular target viral variant, an example of which is the spike (S) protein or a portion thereof of a particular target SARS-CoV-2 variant. As used herein, a full-length SARS-CoV-2 S protein, including the "wild-type" or "Wuhan" sequence, has a sequence corresponding to that of the first detected SARS-CoV-2 strain, which consists of 1273 amino acids and has the amino acid sequence according to SEQ ID NO: 1 below. (SEQ ID NO: 1)
[0234] Unless otherwise indicated, position numbering of the SARS-CoV-2 S protein and / or portions thereof set forth herein is relative to the amino acid sequence of SEQ ID NO: 1. One of skill in the art, upon reading this disclosure, will be able to understand and determine the corresponding position in a SARS-CoV-2 S protein variant sequence from the placement of the provided position relative to the amino acid sequence of SEQ ID NO: 1 (i.e., one of skill in the art, given a position relative to SEQ ID NO: 1 or another variant, will be able to determine the corresponding position in the S protein sequence of another SARS-CoV-2 variant or fragment thereof). Where a portion of the SARS-CoV-2 S protein is described as having a particular mutation, it should be understood that the position numbering of the mutation is given to identify its placement in relation to the amino acid sequence of SEQ ID NO: 1, unless otherwise indicated.
[0235] In specific embodiments, the spike (S) proteins described herein can be modified to stabilize the prototype prefusion conformation. Certain mutations that stabilize the prefusion conformation are known in the art and are disclosed, for example, in WO 2021243122 A2 and Hsieh, Ching-Lin, et al. (“Structure-based design of prefusion-stabilized SARS-CoV-2 spikes,” Science 369.6510 (2020): 1501-1505), the contents of each of which are incorporated herein by reference in their entireties. In some embodiments, the SARS-CoV-2 S protein may be stabilized by introducing one or more proline mutations. In some embodiments, the SARS-CoV-2 S protein comprises a proline substitution at a position corresponding to residues 986 and / or 987 of SEQ ID NO:1. In some embodiments, the SARS-CoV-2 S protein comprises a proline substitution at one or more positions corresponding to residues 817, 892, 899, and 942 of SEQ ID NO:1. In some embodiments, the SARS-CoV-2 S protein comprises a proline substitution at a position corresponding to each of residues 817, 892, 899, and 942 of SEQ ID NO: 1. In some embodiments, the SARS-CoV-2 S protein comprises a proline substitution at a position corresponding to each of residues 817, 892, 899, 942, 986, and 987 of SEQ ID NO: 1.
[0236] In some embodiments, stabilization of the prototype pre-fusion conformation of the SARS-CoV-2 S protein may be achieved by introducing two consecutive proline substitutions at residues 986 and 987. Specifically, a spike (S) protein stabilized protein variant is achieved by replacing the amino acid residue at position 986 with a proline and also replacing the amino acid residue at position 987 with a proline. In one embodiment, the SARS-CoV-2 S protein variant in which the prototype pre-fusion conformation is stabilized comprises the amino acid sequence set forth below in SEQ ID NO: 2. (SEQ ID NO: 2).
[0237] Those skilled in the art will be aware of various SARS-CoV-2 spike variants and / or resources documenting them. For example, the following strains, their SARS-CoV-2 S protein amino acid sequences, and their modifications, particularly relative to the wild-type SARS-CoV-2 S protein amino acid sequence (e.g., relative to SEQ ID NO: 1), are useful herein:
[0238] B.1.1.7 ("Variant of Concern 202012 / 01" (VOC-202012 / 01))
[0239] B.1.1.7 ("alpha variant") is a SARS-CoV-2 variant first detected in the UK in October 2020 from samples collected the previous month and began spreading rapidly by mid-December. This correlates with a significant increase in COVID-19 infection rates, an increase that is thought to be at least in part due to an N501Y mutation within the receptor-binding domain of the spike glycoprotein, required for binding to ACE2 in human cells. B.1.1.7 is characterized by 23 mutations, of which 13 are nonsynonymous, 4 are deletions, and 6 are synonymous (i.e., there are 17 mutations that alter the protein and 6 that do not). Spike protein changes in B.1.1.7 include deletions 69-70, deletion 144, N501Y, A570D, D614G, P681H, T716I, S982A, and D1118H.
[0240] B.1.351(501.V2)
[0241] The B.1.351 lineage ("beta variant"), commonly known as the South African COVID-19 variant, has increased transmissibility compared to the original Wuhan strain. The B.1.351 variant is characterized by multiple spike protein changes, including L18F, D80A, D215G, deletions 242-244, R246I, K417N, E484K, N501Y, D614G, and A701V. The spike region of the B.1.351 genome contains three mutations of particular interest: K417N, E484K, and N501Y.
[0242] B.1.1.298 (Cluster 5)
[0243] B.1.1.298 was discovered in the Danish region of North Denmark and is believed to have spread from minks to humans via mink farms. Several different mutations have been identified in the spike protein of this virus. Specific mutations include deletions 69-70, Y453F, D614G, I692V, M1229I, and possibly S1147L.
[0244] P.1(B.1.1.248)
[0245] The lineage B.1.1.248 ("gamma variant"), known as the Brazilian (Brazilian) variant, is one of the SARS-CoV-2 variants designated the P.1 lineage. P.1 has several S protein mutations (L18F, T20N, P26S, D138Y, R190S, K417T, E484K, N501Y, D614G, H655Y, T1027I, V1176F) and is similar to the South African variant B.1.351 at certain key RBD positions (K417, E484, N501).
[0246] B.1.427 / B.1.429(CAL.20C)
[0247] Lineage B.1.427 / B.1.429 ("epsilon variant"), also known as CAL.20C, is characterized by the following modifications in the S protein: S13I, W152C, L452R, and D614G. Of these, the L452R modification is of particular concern. The CDC has listed B.1.427 / B.1.429 as a "variant of concern."
[0248] B.1.525
[0249] B.1.525 (the "eta variant") has the same E484K mutation found in the P.1 and B.1.351 variants, and the same ΔH69 / ΔV70 deletion found in B.1.1.7 and B.1.1.298. It also has the mutations D614G, Q677H, and F888L.
[0250] B.1.526
[0251] B.1.526 ("iota variant") was detected as a new lineage of virus isolates in New York and has the same mutations as previously reported variants. The common set of spike mutations in this lineage are L5F, T95I, D253G, E484K, D614G, and A701V.
[0252] B.1.1.529
[0253] B.1.529 (the "Omicron variant") was first detected in South Africa in November 2021. Omicron grows approximately 70 times faster than the Delta variant and quickly became the dominant strain of SARS-CoV-2 worldwide. Since its initial detection, multiple Omicron sublineages have emerged. Below, current Omicron variants of concern are listed, along with the specific characteristic mutations associated with their respective S proteins. The S proteins of BA.4 and BA.5 share the same set of characteristic mutations and are therefore listed in a single row in the tables below as "BA.4 or BA.5," and therefore, this disclosure refers to them in some embodiments as "BA.4 / 5" S proteins. Similarly, the S proteins of BA.4.6 and BF.7 Omicron variants share the same set of characteristic mutations and are therefore listed in a single row in the tables below as "BA.4.6 or BF.7." [Table 1A-1] [Table 1A-2] [Table 1A-3] [Table 1A-4]
[0254] In some embodiments, the SARS-CoV-2 S protein described herein includes one or more mutations (including, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more) characteristic of a particular Omicron variant (e.g., one or more mutations of an Omicron variant listed in Table 1A, e.g., each mutation associated with a given XBB variant in Table 1A above).
[0255] As noted elsewhere in this disclosure, in some embodiments, particular immunogenic portions (e.g., subregions) of a full-length coronavirus S protein (e.g., SARS-CoV-2 S protein) may be used as reference antigens for generating engineered antigens as described herein.
[0256] In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks a particular characteristic of the full-length polypeptide (e.g., a characteristic that has been shown or predicted to prevent the elicitation of a naive immune response). For example, in some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks a region that (i) has a low number or density of B-cell neutralizing epitopes and / or (ii) has a high number or density of B-cell epitopes not associated with neutralization. For example, in some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks the entire S2 domain. In some embodiments, a coronavirus (e.g., SARS-CoV-2) S protein that lacks the entire S2 domain lacks a region of S2 that (i) has a low number or density of B-cell epitopes associated with neutralization, or (ii) has a high number of B-cell epitopes not associated with neutralization, but retains other portions of S2. In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks the entire S2 domain. In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks the entire S2 domain but includes specific sequences that can improve the immunogenicity and / or stability of the immunogenic portion (e.g., in some embodiments, the immunogenic portion lacks the entire S2 domain but retains the TM sequence).
[0257] One of skill in the art, reading this disclosure, will be able to identify B cell epitopes in the coronavirus (e.g., SARS-CoV-2) S protein and determine which epitopes are or are not relevant for neutralization. For example, one of skill in the art will recognize numerous studies that have identified such regions using antibody binding studies (e.g., studies characterizing antibodies produced in subjects infected with or vaccinated against SARS-CoV-2).
[0258] In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein comprises a particular region determined to have a high number or density of neutralizing epitopes, and optionally a high mutation rate. In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein comprises the N-terminal domain (NTD) of the S protein. In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) protein comprises the receptor binding domain (RBD) of the S protein. In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein comprises the S1 domain of the S protein.
[0259] In some embodiments, an immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein comprises the RBD and NTD, and omits other features of the S1 domain.
[0260] The coronavirus (e.g., SARS-CoV-2) S protein has been well characterized, and one of skill in the art can determine which portions of the S protein sequence correspond to the immunogenic portions discussed herein (e.g., which portions of the S protein sequence correspond to the NTD, RBD, S1, and S2 domains). In some embodiments, the RBD of a coronavirus (e.g., SARS-CoV-2) S protein comprises residues 327-528 of SEQ ID NO:1, or a corresponding region.
[0261] In some embodiments, the RBD of a coronavirus (e.g., SARS-CoV-2) S protein has the amino acid sequence: VRFPNITNLCPFHEVFNATTFASVYAWNRKRISNCVADYSVIYNFAPFFAFKCYGVSPTKLNDLCFTNVYADSFVIRGNEVSQIAPGQTGNIADYNYKLPDDFTGCVIAWNSNKLDSKPSGNYNYLYRLFRKSKLKPFERDISTEIYQAGNKPCNGVAGPNCYSPLQSYGFRPTYGVGHQPYRVVVLSFELLHAPATVCGPK (SEQ ID NO: 3), or the corresponding region.
[0262] In some embodiments, the RBD of a coronavirus (e.g., SARS-CoV-2) S protein has the amino acid sequence: VRFPNITNLCPFGEVFNATRFASVYAWNRKRISNCVADYSVLYNSASFSTFKCYGVSPTKLNDLCFTNVYADSFVIRGDEVRQIAPGQTGKIADYNYKLPDDFTGCVIAWNSNNLDSKVGGNYNYLYRLFRKSNLKPFERDISTEIYQAGSTPCNGVEGFNCYFPLQSYGFQPTNGVGYQPYRVVVLSFELLHAPATVCGPK (SEQ ID NO: 4), or the corresponding region.
[0263] In some embodiments, the S1 domain of a coronavirus (e.g., SARS-CoV-2) S protein comprises amino acids 1-678 of SEQ ID NO: 1, or a corresponding region in the S protein of a SARS-CoV-2 variant. In some embodiments, the S1 domain of a SARS-CoV-2 S protein comprises amino acids 1-683 of SEQ ID NO: 1, or a corresponding region in the S protein of a SARS-CoV-2 variant. In some embodiments, the S1 domain of a SARS-CoV-2 S protein comprises amino acids 1-685 of SEQ ID NO: 1, or a corresponding region in the S protein of a SARS-CoV-2 variant.
[0264] In some embodiments, the S1 domain of the SARS-CoV-2 S protein comprises the amino acid sequence: (SEQ ID NO:5).
[0265] In some embodiments, the S1 domain of the SARS-CoV-2 S protein comprises the amino acid sequence: (SEQ ID NO:6).
[0266] In some embodiments, the S2 domain of the SARS-CoV-2 S protein comprises amino acids 679-1273 of SEQ ID NO: 1, or a corresponding region in the S protein of a SARS-CoV-2 variant. In some embodiments, the S1 domain of the SARS-CoV-2 S protein comprises amino acids 684-1273 of SEQ ID NO: 1, or a corresponding region in the S protein of a SARS-CoV-2 variant. In some embodiments, the S1 domain of the SARS-CoV-2 S protein comprises amino acids 686-1273 of SEQ ID NO: 1, or a corresponding region in the S protein of a SARS-CoV-2 variant.
[0267] In some embodiments, the compositions described herein deliver an immunogenic portion of the S protein of a SARS-CoV-2 variant. In some embodiments, the variant is a variant of concern (e.g., a variant predicted and / or shown to spread rapidly in the relevant jurisdiction, e.g., a variant identified by a specific public health agency, e.g., the Centers for Disease Control and Prevention (CDC), Public Health England and the COVID-19 Genomics UK Consortium, the Canadian COVID Genomics Network (CanCOGeN), and / or the World Health Organization (WHO)). In some embodiments, the variant is predicted to have a high likelihood of becoming a variant of concern (e.g., predicted using sequence-based algorithms that predict a variant's ability to escape a previously developed immune response and / or measure the "fitness" of a given variant, examples of which include those described in, e.g., WO2022 / 235847 and WO2022 / 235853, the contents of each of which are incorporated herein by reference in their entireties).
[0268] In some embodiments, the RBD comprises a mutation associated with a variant described herein. One skilled in the art can identify which portion of a given variant corresponds to an immunogenic portion described herein.
[0269] In some embodiments, the polypeptide comprises two or more SARS-CoV-2 subdomains (e.g., two or more S1 domains or RBDs). In some embodiments, the polypeptide comprises two or more receptor-binding domains linked in tandem, examples of which include those described in Dai, Lianpan, et al. "A universal design of betacoronavirus vaccines against COVID-19, MERS, and SARS," Cell 182.3 (2020): 722-733 and Han, Yuxuan, et al. "mRNA vaccines expressing homo-prototype / Omicron and hetero-chimeric RBD-dimers against SARS-CoV-2," Cell Research 32.11 (2022): 1022-1025, the contents of each of which are incorporated herein by reference in their entireties. In some embodiments, the two or more subdomains are derived from the same SARS-CoV-2 variant (e.g., a variant described herein). In some embodiments, at least two of the two or more subdomains are derived from different SARS-CoV-2 variants (e.g., from different variants of concern, different omicron variants, an omicron variant and a non-omicron variant, or a Wuhan strain and an omicron variant).
[0270] ii. Characteristic mutations In certain embodiments, the approaches described herein utilize information describing the unique characteristics of a particular reference antigen (e.g., identifying characteristic mutations). A characteristic mutation is a mutation that serves to distinguish a particular reference antigen (e.g., a protein) from other similar antigens. In certain embodiments, for example, the reference antigen is a specific protein or portion thereof of a particular viral variant, and a characteristic mutation is a mutation that serves to distinguish a particular viral variant, or even closely related variants, from other, e.g., variants. For example,
[0271] For example, in the context of SARS-CoV 2, characteristic mutations may be identified based on one or more classification schemes, examples of which include the World Health Organization (WHO) classification(s), GISAID, PANGO lineage, Nextstrain clade, etc.
[0272] For example, in certain embodiments, the reference antigen may be the SARS CoV 2 spike protein of a particular XBB variant. A characteristic mutation may then be identified as one that occurs above a certain characteristic threshold percentage of all sequences (e.g., within a particular database) classified as members of the XBB lineage, as designated by the WHO. In certain embodiments, the characteristic threshold is 50 percent or greater. In certain embodiments, the characteristic threshold is 60 percent or greater. In certain embodiments, the characteristic threshold is 70 percent or greater. In certain embodiments, the characteristic threshold is 75 percent or greater.
[0273] Exemplary XBB.1.5 signature mutations identified by the WHO as mutations occurring in more than 50% of the sequences of variants classified as XBB.1.5 (e.g., mutations compared to the Wuhan strain, e.g., mutations identified by SEQ ID NO: 1 and / or SEQ ID NO: 2) are shown in Table 1B below. [Table 1B-5]
[0274] iii. Memory-inducing storage area(s) In certain embodiments, one or more memory-inducing conserved region(s) 122 of polypeptide model 102 are identified 120. The identified (memory-inducing) conserved region(s) 122 represent specific portions of the reference antigen (represented by polypeptide model 102) that have been determined to have a high likelihood of inducing a memory immune response. The memory-inducing conserved regions may represent these specific portions of the reference antigen in various ways. For example, the memory-inducing conserved regions may be or include a list of amino acid positions, or may additionally or alternatively be or include identified 3D regions on the surface of a 3D polypeptide model.
[0275] A particular portion of a reference antigen may be determined to be likely to elicit a memory immune response by several approaches, based on a variety of criteria.
[0276] In certain embodiments, for example, but not limited to, previously known or determined data (epitope lists, structural models, etc.) may be combined with experimental techniques (e.g., binding assays) and computational methods, used individually or in combination, to identify and determine which memory-inducing conserved regions are involved in various ways in inducing memory immune response(s).
[0277] Conserved epitopes In certain embodiments, the memory-inducing portion of a reference antigen is determined by identifying a set of conserved epitopes within the reference antigen that are likely to induce a memory immune response. The set of shared epitopes of the reference antigen may be identified by comparing an initial set of known epitopes with mutations present in the reference antigen. For example, in certain embodiments, the reference antigen is a version of a particular protein, such as a version of a protein from a particular viral variant. In certain embodiments, the set of known epitopes includes epitopes that meet certain criteria, examples of which may include epitopes identified as targets for antibody binding (e.g., binding epitopes) and / or neutralizing antibodies (e.g., neutralizing epitopes).
[0278] A known epitope may be, for example, a previously characterized epitope of a particular virus of which the reference antigen is a variant, an example of which is an epitope determined to bind to an antibody. In certain embodiments, a known epitope is a site to which a neutralizing epitope binds (e.g., a neutralizing B cell epitope). A set of known epitopes may include binding epitopes and / or neutralizing epitopes. Data identifying the set of known epitopes may be or include a list of amino acid positions for each known epitope in the set. Such data may be obtained from sources, examples of which include results of targeting experiments, proprietary datasets, and public information (such as literature and public datasets).
[0279] For example, data on epitopes of SARS-CoV 2 spike protein epitopes can be found via CoV-AbDab, IEDB, etc.
[0280] The known epitope dataset may then be screened and analyzed to determine which known epitopes are mutated in the reference antigen and which are not mutated and therefore conserved. For example, data identifying the set of known epitopes can be compared to a list of characteristic mutations for a particular reference antigen. Epitopes in the set of known epitopes that include / overlap with one or more characteristic mutations of the reference antigen may be identified as mutated epitopes, while other epitopes (e.g., sets that do not include characteristic mutations) may be identified as conserved epitopes. In certain embodiments, the set of known epitopes may be filtered using additional criteria, examples of which include identifying and removing epitopes that are subsets of other known epitopes or that are or are not located within specific regions (e.g., within the context of SARS-CoV2, the RBD, and / or the N-terminal region).
[0281] Preserved surface In certain embodiments, identifying memory-inducing conserved regions includes identifying conserved surfaces of a polypeptide model representing a particular reference antigen. The conserved surfaces represent portions of the reference antigen that are capable of and / or likely to bind to antibodies generated from a memory immune response (e.g., memory B cells), but may not necessarily have been previously identified as, for example, a known epitope. For example, the conserved surfaces may represent portions (e.g., amino acid sites) of a particular reference antigen that are determined to be (i) sufficiently surface-accessible and (ii) sufficiently similar to other (e.g., related) antigens and / or unaffected by mutations present in the particular reference antigen, such that they can be binding targets for antibodies associated with a memory immune response.
[0282] For example, in certain embodiments, amino acid positions determined to be located on a surface (e.g., sufficiently surface-accessible) may be identified (e.g., using the polypeptide model 102). In certain embodiments, the set of surface amino acids may be compared to the signature mutations of a reference antigen to exclude (e.g., remove from the set) amino acid positions that are sites of signature mutations and / or that are within a certain distance (e.g., in 3D space, a linear distance or a distance across a 3D surface (e.g., a geodesic distance)) of one or more signature mutations. The remaining set of amino acid sites may then be used to define a conserved surface. In certain embodiments, the conserved surface may be further refined by evaluation of the spatial distribution of amino acid sites to, for example, eliminate (from the conserved surface) portions determined to be too small for antibody binding. For example, isolated patches below a certain threshold size (e.g., surface area) may be removed from the conserved surface.
[0283] Memory-inducing storage area In certain embodiments, one or more memory-inducing conserved regions are or include a set of conserved epitopes. In certain embodiments, one or more memory-inducing subregions are or include a conserved surface. In certain embodiments, both a set of conserved epitopes and a conserved surface are included / identified as one or more memory-inducing subregions.
[0284] iv. Destruction of storage area(s) 1 , following identification of one or more memory-inducing conserved region(s) of a reference antigen, the systems and methods of the present disclosure aim to disrupt these regions, e.g., to mitigate or avoid (e.g., reduce the likelihood and / or extent of) eliciting a memory immune response. In particular, in certain embodiments, the approaches described herein introduce (e.g., in silico) amino acid modifications (e.g., substitutions, insertions, deletions, etc.) into at least a portion of one or more identified memory-inducing conserved regions of the reference antigen.
[0285] distribution criteria In certain embodiments, the approach of introducing amino acid modifications into conserved regions aims to distribute the amino acid modifications throughout the entire conserved region. For example, in certain embodiments, one or more amino acid modifications are introduced into at least a portion of one or more conserved epitope regions. In certain embodiments, one or more amino acid modifications are introduced into each of one or more conserved epitope regions (several), for example, to disrupt each of the identified conserved epitopes. This can be achieved, for example, by considering at least a portion (e.g., all) of the conserved epitope regions separately and introducing one or more amino acid modifications into each, for example, in a stepwise manner, or additionally or alternatively, by repeatedly introducing amino acid modifications randomly / according to various statistical selection criteria until each conserved epitope region is modified by at least one amino acid modification or until other stopping criteria are met.
[0286] In certain embodiments, one or more amino acid modifications are introduced at a conserved surface (e.g., across the entire conserved surface). In certain embodiments, multiple amino acid modifications are generated. In certain embodiments, the amount and placement of amino acid modifications are introduced to generate a particular desired spatial distribution of mutations. Design criteria may include, but are not limited to, density across the 3D surface representing the conserved regions, minimum and / or maximum separation between amino acid modifications spatially across the 3D conserved surface. In certain embodiments, the desired distribution of mutations across the conserved surface may be achieved by calculating a positional diffusion score that measures the extent to which the placement of amino acid modifications is, for example, evenly distributed across the conserved surface, and evaluating the design, for example, using at least part of the calculated positional diffusion score. In certain embodiments, criteria such as density, minimum / maximum separation, and positional diffusion score are enhanced using the introduced amino acid modifications (e.g., to ensure that the introduced amino acid modifications are introduced in a manner that is distributed across the conserved regions). In certain embodiments, criteria such as density, min / max separation, positional diffusion score, etc. are enhanced using the introduced amino acid modifications and existing, e.g., characteristic mutations of a particular antigen, such as a viral protein variant (e.g., ensuring that the introduced amino acid modifications are introduced in a manner that is distributed throughout conserved regions, and avoiding mutation of existing unique epitopes that are desirable to retain, e.g., to facilitate the generation of antibodies directed against (e.g., with affinity for) these regions).
[0287] Generation of amino acid modifications In certain embodiments, the amino acid modifications introduced into the memory-inducing subregion are selected from a set of permissible mutations. In certain embodiments, the set of permissible mutations includes a set of one or more mutations identified as permissible for at least a portion of each amino acid position within a particular sequence. Permissible mutations may be identified using sequence data (e.g., nucleotide sequence and / or amino acid sequence) of an antigen related to a reference antigen. For example, if the reference antigen is a particular protein of a viral variant (e.g., a SARS-CoV 2 spike protein (e.g., an omicron spike protein, an XBB spike protein, etc.) of a particular SARS-CoV 2 variant), permissible mutations may be determined by first selecting a set of related lineage(s) and obtaining sequences of versions of the particular protein of various viral variants classified as belonging to the selected set of related lineages, e.g., from proprietary sequencing data and / or public data (e.g., GSAID). Mutations occurring within each of the related variants of the particular protein can be identified and included in the set of permissible mutations. In certain embodiments, mutations included in the set of allowed mutations are filtered to include only mutations that are observed at a sufficiently high rate or frequency (e.g., above a certain threshold frequency). For example, in certain embodiments, only mutations that are observed within at least a certain threshold percentage of related variant sequences are included in the set of allowed mutations.
[0288] In certain embodiments, the related variant sequence may be or include the sequence of a protein within a similar family and / or a protein known to perform a similar function as the reference antigen. For example, the reference antigen may be a specific SARS-CoV2 protein, such as the spike protein. In certain embodiments, the reference sequence may be the spike protein of another coronavirus, or a subset thereof, such as the spike protein of a related SARS virus (e.g., one known to infect humans, e.g., SARS-CoV1, MERS), or a spike protein known to reside / originate in a particular region. In this manner, in certain embodiments, amino acid modifications may be selected based on and / or to take advantage of changes present in other related polypeptides.
[0289] For example, in certain embodiments, the reference sequence may be or include the sequence of a variant of the reference antigen (such as an ancestor or member of another branch along the evolutionary tree). The reference sequence may be limited to variants with a particular level of prevalence and / or immune escape. In certain embodiments, the reference sequence may be obtained from one or more public and / or proprietary databases (such as GSAIDs).
[0290] In certain embodiments, one or more additional criteria are used to select amino acid modifications that provide a sufficient level of disruption and / or that still maintain polypeptide chain stability. For example, in certain embodiments, amino acid modifications that replace existing amino acids with specific equivalent / similar amino acids are excluded, e.g., because they do not provide sufficient disruption. For example, in certain embodiments, amino acid modifications that disrupt cysteine bridges are prohibited and / or penalized. In certain embodiments, amino acid modifications that cause charge reversal and / or large changes in charge are disallowed and / or penalized. Similarity / dissimilarity criteria and structural preservation criteria as described herein may be implemented via rules (e.g., encoded rules), such as, for example, conditional logic, lookup tables, etc.
[0291] 2 illustrates an exemplary process by which the sequence of a relevant antigen is accessed / obtained and analyzed to identify and generate a set of allowed mutations that can be utilized to create amino acid modifications in and disrupt conserved subregions of the reference antigen. In certain embodiments, the steps of selecting and inserting, and optionally filtering, amino acid modifications may be repeated, e.g., iteratively, to generate multiple candidate engineered variants based on a single initial conserved region.
[0292] Scoring and selection of candidate variants In certain embodiments, as conserved subregion(s) are disrupted by the introduction of amino acid modifications, the resulting in silico engineered variants of the reference antigen may be scored to assess, for example, their viability, ability to infect host cells, and relative similarity / dissimilarity to existing variants. For example, as shown in FIG. 3 , in certain embodiments, one or more candidate disrupted polypeptide models 342a, 342b, ... 342n are generated (340) from an initial polypeptide model 302 representing a particular reference antigen (e.g., by identifying (320) and disrupting (330) conserved region(s)). Each candidate disrupted polypeptide model corresponds to the initial polypeptide model, but its conserved subregion has been disrupted by the introduction of one or more amino acid modifications.
[0293] Candidate disrupted polypeptide models may be evaluated and scored (340) in a variety of ways, including by one or more of the approaches described in PCT Publication WO 2022 / 235847 A1 (entitled "TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST," published November 10, 2022) and PCT Publication WO 2022 / 235853 A1 (entitled "IMMUNOGEN SELECTION," also published November 10, 2022), the contents of each of which are incorporated herein by reference in their entirety.
[0294] For example, in certain embodiments, candidate polypeptide models may be scored using a machine learning model such as a language model. In certain embodiments, the language model is trained or has been trained to generate a predicted probability for each of one or more amino acids in each of one or more positions of an input sequence. Once trained, the language model can be used to predict the overall likelihood (log-likelihood, conditional log-likelihood, etc.) of a particular variant represented by a particular input amino acid sequence.
[0295] In certain embodiments, a language model trained in this manner may also be used to determine an embedding vector representation of an input amino acid sequence. The embedding vector representation of an amino acid sequence is an internal representation generated by a machine learning model and may be extracted as output from one or more hidden layers of the machine learning model. The embedding layer representation of the extracted variants may be used by itself and / or may be used to generate a feature vector representing a particular variant.
[0296] For example, in certain embodiments, a polypeptide sequence is provided as input to a machine learning model, which then generates an internal representation, i.e., an embedding, of the input polypeptide sequence, where the embedding may represent the input sequence as a high-dimensional matrix or vector (e.g., a numerical matrix or vector). For example, in certain embodiments, a machine learning model, such as certain language models described in further detail herein, may receive an amino acid sequence as input and generate an internal (e.g., embedding) representation. In certain embodiments, the embedding is or includes a vector (e.g., zi) of, e.g., length D, for each amino acid position in the sequence, where D is the dimensionality of the vector. Thus, an initial embedding that directly corresponds to the internal representation output by a particular layer (e.g., the final embedding layer of a recurrent neural network, e.g., the feature map from the final Transformer layer of a Transformer-based model) may be a matrix of size n×D [e.g., or (n+1)×D], where n is the length of the input amino acid sequence and D is the dimensionality [e.g., in certain embodiments, additional class tokens may be used to add to the network, resulting in a matrix of size (n+1)×D]. However, this will depend on the particular machine learning model used. In certain embodiments, the feature vector may then be determined as an average of at least a portion of the sequence positions (e.g., all of the sequence positions (e.g., excluding the first position (e.g., class token) that does not represent or correspond to an amino acid in the sequence)) to determine the feature vector for, e.g., an amino acid sequence.
number
[0297] In this manner, in certain embodiments, a machine learning model may be used to generate a corresponding feature vector (e.g., based on an embedding of the machine learning model) for each of a plurality of viral variant polypeptide sequences. This approach can be used to associate each variant sequence with a location within a higher-dimensional feature vector / embedding space.
[0298] In certain embodiments, feature vectors and / or embeddings representing particular sequences (e.g., viral variants) may be compared to each other and / or to particular versions (e.g., wild-type versions of proteins and / or other particular variants of interest). In certain embodiments, distances within the embedding space (e.g., L1 distance, L2 distance, etc.) may be calculated and used to assess the similarity and / or dissimilarity between a particular candidate variant and one or more other variants for comparison. In certain embodiments, the distance in the embedding space may be referred to as a semantic variation score and is believed to indicate, at least in part, the likelihood that a particular variant will be recognized by a host system previously exposed to one or more reference antigens.
[0299] The machine learning model used to generate the feature vector may utilize and implement various machine learning techniques. For example, the machine learning model may be a deep learning model (e.g., an artificial neural network with one or more, e.g., multiple, hidden layers) such as a language model or a large-scale language model (LLM). In particular embodiments, the machine learning model is or includes one or more recurrent models (e.g., long short-term memory (LSTM)), implemented alone or in combination (e.g., bidirectional LSTM (bi-LSTM)). In particular embodiments, the machine learning model may be or include one or more Transformer models. Examples of machine learning models include, but are not limited to, Evolutionary Scale Models (ESM), Bidirectional Encoder Representations from Transformers (BERT), etc.
[0300] Machine learning models may be trained on protein sequence data available from public and / or private repositories, examples of which include UniRef (see, e.g., Suzek et al. "UniRef: comprehensive and non-redundant UniProt reference clusters" 2007) and the National Center for Biotechnology Information's Virus Pathogen Resource (ViPR) database. Training may be achieved by providing the machine learning model with incomplete and / or partially masked input sequences and tasking the model with predicting the omitted and / or masked data. In this manner, machine learning models may be trained in an unsupervised manner. Additional details of machine learning models and approaches for training them (such as specific techniques for training recurrent and / or transformer models) are described, for example, in PCT Publications WO 2022 / 235847 and WO / 2022 / 235853, the contents of each of which are incorporated herein by reference in their entirety.
[0301] In certain embodiments, structural modeling techniques may be used to score candidate polypeptide models. For example, in certain embodiments, viral polypeptide receptor binding scores (such as ACE2 binding scores) may be determined for one or more candidate synthetic variants. In certain embodiments, epitope alteration scores may be determined for one or more candidate variants.
[0302] In certain embodiments, the amount of mutation may be assessed and used as a score to, for example, select candidate polypeptide models having a particular range in terms of amount of mutation. In certain embodiments, a mutation co-occurrence score may be used to assess whether artificially introduced mutations in a particular candidate variant match the co-occurrence observed in real-world, naturally evolved variants.
[0303] In this manner, each potential candidate may be associated with and characterized by a set of performance scores 344a, 344b, . . . 344n.
[0304] In certain embodiments, one or more candidate polypeptide models may be selected (360), e.g., based on such performance scores, e.g., as representations of engineered antigens to be used as immunogenic compositions. In certain embodiments, two or more sets of scores described herein may be combined, e.g., as a (e.g., weighted) combined score. In certain embodiments, two or more sets of scores described herein may be combined to identify a Pareto front and / or a set of Pareto-optimal solutions. In certain embodiments, a Pareto score based on a combination of two scores may further be determined, examples of which include those described in PCT Publication WO 2022 / 235847 A1 (entitled "TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST," published November 10, 2022) and PCT Publication WO 2022 / 235853 A1 (entitled "IMMUNOGEN SELECTION," also published November 10, 2022), the contents of each of which are incorporated herein by reference in their entirety.
[0305] V. Processing of Engineered Antigens In certain embodiments, disrupted polypeptide models created in silico by the systems and methods described herein may be stored and / or provided, e.g., for display and / or further processing. For example, in certain embodiments, the disrupted polypeptide models may be used to generate corresponding RNA sequences that encode the engineered antigen represented by the disrupted polypeptide models. In certain embodiments, engineered antigens may be produced based on the disrupted polypeptide models and their biological activity may be evaluated, e.g., in vitro. In this manner, multiple initial candidate engineered antigens may be generated, e.g., by the in silico design techniques described herein, and then produced and screened in vitro.
[0306] For example, in certain embodiments, assays may be used to confirm expression and folding. For example, engineered variants designed via one or more approaches described herein may be produced and tested for viability in terms of expression and folding. In certain embodiments, one or more antibodies or binding agents may be used to analyze intracellular and surface expression of the encoded antigen. In certain embodiments, for example, an ACE2 binding agent may be used to test engineered SARS-CoV 2 variants. In certain embodiments, a panel of reference (e.g., known) neutralizing antibodies may be used to assess the ability of an engineered construct to evade pre-existing neutralizing antibodies. In certain embodiments, the reference panel of neutralizing antibodies includes one or more antibodies binding to different epitope classes (e.g., A, B, C, D, E, F). In certain embodiments, the analysis may be extended to immune serum analysis from vaccinated or convalescent individuals. In certain embodiments, a pseudoviral assay is performed to confirm loss of nAb titer.
[0307] In certain embodiments, one or more (e.g., multiple) testing procedures are used, arranged in a decision tree / hierarchical manner, e.g., as described in Example 8. For example, in certain embodiments, all or a subset of steps may be used to progressively filter the engineered construct designs until a subset is obtained that passes each step, with each step / test either retaining the construct and passing it on to the next step or discarding it. Such steps may include any of the following: 1. Confirmation of expression and folding 2. Inactivation / reduction of binding of a panel of reference antibodies; 3. Inactivation / reduction of binding of complex immune sera; 4. Invalidation of neutralization of each pseudovirus; 5. Dedicated immunogenicity studies. Examples of approaches used to design engineered SARS-CoV-2 antigens are described in further detail in Example 8.
[0308] In certain embodiments, the results of the assays described herein may be combined with in silico design procedures, e.g., in an iterative manner, to inform and / or add constraints to subsequent in silico design rounds.
[0309] C. Delivery of Engineered Antigens Those skilled in the art will appreciate that effective vaccination can be achieved by delivering engineered antigens to a subject such that the antigens are exposed to the subject's B cells.
[0310] The engineered antigens of the present disclosure may comprise a full-length viral protein and / or particular portions thereof (e.g., an RBD altered to disrupt a conserved region, etc.) according to various embodiments described herein (e.g., in Section B above, and / or various examples described below). Constructs comprising the engineered antigens of the present disclosure may combine the engineered viral protein and / or portions thereof with one or more additional elements, examples of which include, for example, one or more of secretion signal(s), linker, multimerization region (e.g., fibritin domain), and membrane-associated portion(s) (e.g., transmembrane region).
[0311] In some embodiments, delivery of engineered antigens and / or constructs comprising them can be achieved by administration of such antigen(s) (e.g., polypeptide antigens). In some embodiments, such delivery can be achieved by administering a composition after which the antigen is produced (e.g., in and / or by the recipient). In some embodiments, delivery of polypeptide antigens is achieved by administration of a composition comprising a polynucleotide encoding the polypeptide antigen. In some such embodiments, such polynucleotides may be or comprise DNA. In some such embodiments, such polynucleotides may be or comprise RNA. Furthermore, those skilled in the art will recognize that non-naturally occurring residues, linkages, and / or other elements are frequently utilized in therapeutic nucleic acids administered to a subject. Examples of various polypeptide and / or nucleic acid sequences suitable for providing engineered antigens of the present disclosure are described in the Examples provided herein (e.g., Examples 10 and 11).
[0312] Of particular interest in certain embodiments of the present disclosure is the administration of a composition comprising RNA encoding an engineered antigen as described herein.
[0313] In many embodiments, provided pharmaceutical compositions (e.g., immunogenic compositions, e.g., vaccines) deliver antigens (e.g., engineered antigens) described herein by delivering a nucleic acid construct, which in many embodiments is an RNA that encodes one or more engineered antigens described herein and that is expressed in a subject upon administration of the pharmaceutical composition (e.g., immunogenic composition, e.g., vaccine).
[0314] Among other things, the present disclosure encompasses the recognition that administration of nucleic acids, and particularly RNA, to achieve delivery (e.g., by expression) of encoded antigens can offer various advantages over other strategies for immunizing against infectious diseases (such as SARS-CoV-2 infection).
[0315] Among other things, the present disclosure provides insight that RNA may be particularly useful and / or effective as an active agent in pharmaceutical compositions (e.g., immunogenic compositions, e.g., SARS-CoV-2 vaccines) for a variety of reasons, including, specifically, that RNA may have inherent adjuvant properties, such as the ability to induce very high antibody titers against SARS-CoV-2 proteins (e.g., SARS-CoV-2 antigens, particularly those associated with variants of concern that have high immune evasion potential), as described herein.
[0316] Furthermore, it has become clear that RNA activity can also induce diverse and significant T cell responses, which, especially when combined with a strong antibody response, represent a combination of immune properties that may maximize the probability of protection, as demonstrated by experience with SARS-CoV-2 vaccines.
[0317] The provided polyribonucleotides may be delivered for the therapeutic uses described herein using any suitable method known in the art, including, for example, as naked RNA or mediated by viral and / or non-viral vectors, polymer-based vectors, lipid compositions, nanoparticles (e.g., lipid nanoparticles, polymer nanoparticles, lipid-polymer hybrid nanoparticles, etc.), and / or peptide-based vectors. For information regarding various approaches potentially useful for the delivery of the polyribonucleotides described herein, see, for example, Wadhwa et al. "Opportunities and Challenges in the Delivery of mRNA-Based Vaccines" Pharmaceutics (2020) 102 (27 pages), the contents of which are incorporated herein by reference.
[0318] In some embodiments, one or more polyribonucleotides can be formulated with lipid nanoparticles for delivery (eg, administration).
[0319] In some embodiments, lipid nanoparticles can be designed to protect polyribonucleotides from extracellular RNases and / or can be engineered for systemic delivery of RNA to target cells. In some embodiments, such lipid nanoparticles can be particularly useful for delivering polyribonucleotides when the polyribonucleotides are administered intravenously or intramuscularly to a subject.
[0320] Additional details regarding approaches suitable for use in delivering the engineered antigens described herein are described in PCT Publication WO 2022 / 235847 A1 (entitled "TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST," published November 10, 2022), the contents of which are incorporated herein by reference in their entirety.
[0321] D. Targets and Indications In certain embodiments, the systems and methods described herein may be used to design engineered antigens suitable for eliciting improved immune responses. In certain embodiments, the approaches described herein may be used to create engineered antigens for use in vaccinating specific subjects (e.g., specific individuals and / or specific populations). For example, in certain embodiments, the antigen engineering techniques described herein may be used to engineer antigens for delivery to subjects who have previously been exposed to specific variants of a reference antigen (e.g., previously naturally circulating variants). For example, information related to specific epitopes of specific variants to which a subject has previously been exposed may, in certain embodiments, be used to identify specific conserved epitopes to be disrupted. Such approaches may be used to reduce memory responses and promote naive immune responses in subjects who have previously been exposed by infection and / or by previous vaccination.
[0322] In certain embodiments, the engineered antigen may be tailored to the subject whose memory B cells have been assessed (e.g., may be explicitly created or selected from a set of pre-existing options). In certain embodiments, the methods of the present disclosure comprise administering to the subject whose memory B cells have been assessed a composition that delivers an engineered antigen as described herein. For example, in certain embodiments, the approaches described herein may be used to create engineered antigen-modified memory epitope(s) that are targeted by the subject's memory B cells (at least if non-neutralizing). In certain embodiments, such engineered antigens may be administered to the subject.
[0323] E. Specific Illustrative Applications In particular, in certain embodiments, the in silico design approach of engineered antigens described herein may be applied for the design of immunogenic compositions. Such immunogenic compounds may, for example, have improved performance and / or may be specifically tailored to individual subjects and / or populations (e.g., depending on the antigen to which different individuals and / or population groups were initially exposed), e.g., taking into account particular populations of immune memory responses.
[0324] Thus, in certain embodiments, the approaches described herein may be used to generate improved immunogenic compositions.
[0325] The techniques described herein, including the various approaches specifically illustrated in the Examples relating to SARS-CoV 2 antigen engineering, may be applied to various types of antigens and pathogens.
[0326] In some embodiments, the infectious agent is a virus, bacterium, or eukaryotic cell (eg, Plasmodium).
[0327] In some embodiments, the infectious agent is a respiratory virus. In some embodiments, the infectious agent is an RNA virus. In some embodiments, the infectious agent is a coronavirus (e.g., MERS, SARS, or SARS-CoV-2). In some embodiments, the infectious agent is HIV. In some embodiments, the infectious agent is HSV (e.g., HSV-1 or HSV-2). In some embodiments, the infectious agent is RSV. In some embodiments, the infectious agent is a norovirus. In some embodiments, the infectious agent is an influenza virus. In some embodiments, the infectious agent is Plasmodium falciparum. In some embodiments, the infectious agent is an orthopoxvirus (e.g., monkeypox).
[0328] In some embodiments, the infectious agent is a bacterium. In some embodiments, the bacterium is a Mycobacterium. In some embodiments, the bacterium is selected from Haemophilus influenzae, Chlamydophila pneumoniae, Mycoplasma pneumoniae, Staphylococcus aureus, Moraxella catarrhalis, Legionella pneumophila, and Streptococcus pneumoniae. In some embodiments, the bacterium is Streptococcus pneumoniae.
[0329] In some embodiments, the infectious agent is an RNA virus. The compositions provided herein may offer particular advantages in providing an immune response to RNA viruses that have relatively high mutation rates (high compared to other infectious agents).
[0330] In some embodiments, the infectious agent comprises multiple strains, variants, or lineages, hi some embodiments, the infectious agent has a relatively high mutation rate (e.g., compared to other infectious agents).
[0331] In some embodiments, the infectious agent is prone to immune escape.
[0332] In some embodiments, the infectious agent is one for which seasonal, variant-adapted booster shots are provided periodically. In some embodiments, the infectious agent antigen is solvent exposed on the surface of the infectious agent. In some embodiments, the infectious agent antigen is a glycoprotein. In some embodiments, the infectious agent antigen is involved in host cell recognition. In some embodiments, the infectious agent antigen is involved in host cell entry. In some embodiments, the infectious agent antigen comprises one or more B cell epitopes (e.g., one or more neutralizing epitopes). [Example]
[0333] F. Working Example i. Example 1: XBB.1.5 Conserved Epitopes and Characteristic Mutations This example describes conserved regions and characteristic mutations in XBB.1.5, as well as other (eg, ACE2) regions of interest.
[0334] To promote novel B cell responses to the XBB.1.5 variants, the design approach in this example began by collecting 1,004 binding and neutralizing B cell epitopes from the CoV-AbDab and IEDB. These epitopes were compared to XBB.1.5 to identify 107 of the initial 1,004 epitopes that did not mutate when XBB.1.5 was considered. This set of conserved epitopes was further curated by manual inspection, including removing epitopes that were not located on the RBD surface and / or were subsets of other epitopes. This process resulted in a final set of 26 distinct conserved epitopes covering 140 positions within the spike protein RBD. Table 2A below lists these conserved epitopes along with their respective sources. [Table 2A-1] [Table 2A-2] [Table 2A-3] [Table 2A-4]
[0335] Certain approaches (e.g., identification of target regions) described herein utilize identification of the ACE2 interface of the SARS-CoV-2 spike (S) protein, which, in certain embodiments, is defined in Lan et al., “Structure of the SARS-CoV-2 spike receptor-binding domain bound to the ACE2 receptor,” Nature, 581:215-220 (2020) and may include the positions listed below. ACE2 interface: K417, G446, Y449, Y453, L455, F456, A475, F486, N487, Y489, Q493, G496, Q498, T500, N501, G502, Y505
[0336] ii. Example 2: Engineering an adapted XBB-based antigen This example describes the design of an engineered antigen based on (e.g., using as a reference antigen) the XBB variant of the SARS-CoV2 spike protein. The engineered antigen described in this example aims to reduce and / or avoid activation of memory immune responses originating from memory B cells and / or T cells in a subject (e.g., an individual receiving a vaccine) in order to promote the production of new neutralizing antibodies targeting specific epitopes associated with the ACE2 binding interface and containing the characteristic mutations in XBB.
[0337] Referring to Figure 4A, the XBB RBD portion was used as a reference antigen for the design of engineered antigens. The approach described in this example starts with the XBB RBD sequence. The PDB structure (PED ID 7EAM) containing RBD positions 325-527 was used as the polypeptide model. The hallmark mutations of XBB listed in Table 1B are shown in red in Figure 4A, while the unmutated regions are shown in gray.
[0338] Figure 4B shows the XBB RBD model with ACE2 interface identities shown in green, as defined in Lan et al., "Structure of the SARS-VoV-2 spike receptor-binding domain bound to the ACE2 receptor," Nature, 581:215-220 (2020), and including the positions listed in Example 1 above.
[0339] 4A and 4B also show the identified conserved surfaces colored purple. The conserved surfaces were identified as contiguous, non-variant surfaces. The amino acid positions identified as belonging to the conserved regions in this example are listed below. ·Storage area: L335, E340, A348, S349, Y351, A352, N354, R355, K356, R357, S359, N360, V362, D364, S366, Y369, N370, A372, F377, K378, Y380, G381, S383, P384, T385, K386, N388, D389, L390, C391, F392, T393, N394, Y396, P412, G413, Q414, T 415, K424, P426, D427, D428, T430, K444, N450, L452, R457, K458, S459, K462, P463, F464, E465, R466, D467, I468, S4 69, T470, E471, I472, Y473, Q474, P479, N481, G482, V483, E516, L517, L518, H519, A520, P521, T523, C525, G526, P527
[0340] Figure 4C shows a colorized version of the XBB structural model. The coloring indicates residues (amino acid positions) belonging to known neutralizing epitopes associated with (e.g., targeted by) neutralizing antibodies. Based on data from the RCSB Protein Databank (PBD) and the Immune Epitope Database and Analysis Resource (IEDB), 113 known neutralizing epitopes were identified. Each residue was scored according to the number of neutralizing epitopes in which it appears. The color coding in Figure 4C, ranging from yellow to green to blue, indicates the relative frequency of positions belonging to epitopes. The highest count is 70 and is blue, intermediate counts are green, and low counts are yellow (i.e., sites that do not or rarely belong to an epitope). While Figure 4C and this example specifically consider neutralizing epitopes, non-neutralizing epitopes are also considered in subsequent designs.
[0341] In this example, we aimed to generate engineered versions of XBB that would limit or avoid eliciting memory immune responses and instead induce the generation of new antibodies tailored to the signature mutations of XBB. Therefore, we compared the 113 neutralizing epitopes with the signature mutations of XBB and identified a subset of the remaining 26 epitopes that were not targeted by the signature mutations of XBB. These conserved, non-mutated epitopes are listed in Table 2B below.
[0342] Amino acid modifications were introduced at various positions to disrupt invariant epitopes and conserved surfaces of the XBB RBD polypeptide model. Amino acid modifications were generated by selecting from a set of permissible mutations known to occur in XBB, XBB sublineages, and omicrons. Additional criteria were used to restrict the amino acid modifications introduced: (i) the modification required significant changes (e.g., Asp to Glu, Arg to Lys, Ile to Leu, and Asn to Gln were not considered), and (ii) modifications that disrupt cysteine bonds were prohibited (e.g., excluded). For example, modifications to positions C391 or C525 were not allowed.
[0343] Based on this criteria, 200,000 randomly generated variants (each a version of XBB but with disrupted conserved regions) were evaluated against various design criteria. Specifically, candidate designs were required to modify amino acids located within each of the 26 non-mutated epitopes, for example, to maximize immune escape. Second, candidate designs were evaluated to ensure that amino acid modifications were well-distributed across the entire conserved surface, rather than clustered together. Finally, each design was scored using a combination of in silico structural modeling and machine learning-based language models. The structural modeling was used to calculate an ACE2 binding score. This is described, for example, in PCT Publication WO 2022 / 235847 A1 (titled "TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST," published November 10, 2022). A machine learning-based language model similar to that described in PCT Publication WO 2022 / 235847 A1 (entitled "TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST") was used to determine the log-likelihood and semantic change scores for each candidate synthetic variant, allowing candidate designs to be evaluated based on these criteria.
[0344] Figures 4D and 4E show two example designs: Design 1 (Figure 4D) and Design 2 (Figure 4E). Both designs introduce amino acid modifications within each non-mutated epitope, and the modifications are well-spread across the conserved surface. ACE2 binding scores, log-likelihood scores, and semantic variation scores were also calculated for each design. Design 1 (Figure 4D) was identified as having a higher ACE2 binding score, while Design 2 (Figure 4E) was identified as having stronger immune evasion (e.g., based on a higher semantic variation score).
[0345] Cyan coloring in Figures 4D and 4E identifies positions where amino acid modifications were introduced in conserved regions for Design 1 and Design 2, respectively. The positions of the modified amino acids for each design are listed below. Design 1 (higher ACE2 binding) modified positions: L335F Y351F N354D S359N A372R L390R Q414K T430I P463S E471Q N481K H519Q; Design 2 (Stronger Immune Escape) Modified positions: L335F A352V A372R N388K L390R T415I T430I P463L T470N Y473S P479S G482R L518V A520E T523A [Table 2B-1] [Table 2B-2] [Table 2B-3] [Table 2B-4]
[0346] iii. Example 3: Engineering an adapted XBB.1.5-based antigen This example describes another exemplary design approach for creating engineered antigens based on (e.g., using as a reference antigen) XBB variants of the SARS-CoV2 spike protein (specifically XBB.1.5). Similar to Example 2, the engineered antigens described in this example are intended to reduce and / or avoid activation of memory immune responses in subjects (e.g., vaccinated individuals) due to the de novo B cell immune response that starts with XBB.1.5.
[0347] Figure 5A shows the SARS-CoV 2 virus and its components, including the spike (S) protein. This spike (S) protein is shown in more detail on the right side of the figure, along with the ACE2 host receptor to which it binds. Figure 5B shows a 3D structural representation of the XBB.1.5 RBD portion of the SARS-CoV-2 S protein, encompassing positions 325-527 (PED ID 7EAM), which was used as a polypeptide model. The characteristic mutations in XBB.1.5 listed below are shown in red in Figure 5B, while the unmutated region is shown in gray. Characteristic mutations of XBB.1.5: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R 408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K
[0348] Figure 5C shows the 3D structure of the complete spike protein, with the characteristic mutations of XBB.1.5 shown in red.
[0349] To promote novel B cell responses to XBB.1.5 variants, the design approach in this example began by collecting 1,004 binding and neutralizing B cell epitopes from the CoV-AbDab and IEDB. These epitopes were compared to XBB.1.5 to identify 107 of the initial 1,004 epitopes that did not mutate when XBB.1.5 was considered. This set of conserved epitopes was further curated by manual inspection, including removing epitopes that were not located on the RBD surface and / or were subsets of other epitopes. This process resulted in a final set of 26 distinct conserved epitopes covering 140 positions within the spike protein RBD, as listed in Table 1 (Example 1) above.
[0350] A contiguous conserved surface on the XBB.1.5 RBD was also identified, which is shown in purple in Figure 5B. The amino acid positions identified in this example as belonging to the conserved region are listed below. Uninterrupted preserved surface: L335, E340, A348, S349, Y351, A352, N354, R355, K356, R357, S359, N360, V362, D364, S366, Y369, N370, A372, F377, K378, Y380, G381, S383, P384, T385, K386, N388, D389, L390, C391, F392, T393, N394, Y396, P412, G413, Q4 14, T415, K424, P426, D427, D428, T430, K444, N450, L452, R457, K458, S459, K462, P463, F464, E465, R466, D467, I468, S469, T470, E471, I472, Y473, Q474, P479, N481, G482, V483, E516, L517, L518, H519, A520, P521, T523, C525, G526, P527
[0351] Figure 5D shows the XBB RBD model with identities of the ACE2 interface shown in green, as defined by Lan et al., “Structure of the SARS-VoV-2 spike receptor-binding domain bound to the ACE2 receptor,” Nature, 581:215-220 (2020), and includes the positions listed below. ACE2 Interface: K417, G446, Y449, Y453, L455, F456, A475, F486, N487, Y489, Q493, G496, Q498, T500, N501, G502, Y505
[0352] Figure 5E shows a colorized version of the XBB structural model. The coloring indicates residues (amino acid positions) that belong to neutralizing epitopes associated with (e.g., targeted by) neutralizing antibodies. Each residue was scored according to the number of neutralizing and non-neutralizing epitopes in which it appears. The color coding in Figure 5E, ranging from yellow to green to blue, indicates the relative frequency of positions belonging to epitopes. The highest count is 287 and is blue, intermediate counts are green, and low counts are yellow (i.e., sites that do not or rarely belong to an epitope).
[0353] Figure 5F shows the epitope density of the final set of 26 conserved epitopes across the entire conserved surface. Color coding, again ranging from yellow to green to blue, indicates the relative frequency of (amino acid) positions belonging to one of the 26 conserved epitopes. Two positions (428 and 518) were hit (i.e., present) in a maximum of 10 of the 26 conserved epitopes. Positions located outside the conserved surface are displayed in gray.
[0354] To disrupt the conserved subregions of the XBB.1.5 RBD, amino acid modifications were introduced at positions across the entire conserved surface (including 76 overall positions), resulting in at least one amino acid modification in each conserved epitope. The amino acid modifications were selected from a set of allowed mutations using the sequences of related variants identified as belonging to XBB, BA, and their sublineages. Mutations present at a frequency of 1% or greater within these related lineages (i.e., XBB, BA, and their sublineages) were included in the set of allowed mutations. Additional filtering criteria were applied to exclude (i) mutations from Cys that disrupt cysteine bonds, (ii) mutations between Asp-Glu, Arg-Lys, Ile-Leu, and Asn-Gln, and (iii) a set of five mutations due to their significant impact on ACE2 binding, as demonstrated by deep mutation scanning (DMS). The final set of allowed mutations (86 in total) used in this example is listed in Table 3, along with the lineage identification of each mutation.
[0355] While the number of possible mutation positions and the absolute number of allowable mutations are relatively small, the number of possible combinations is enormous, approximately 10. For example, consider the following epitope (511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525). The allowable mutations are E516Q, L517F / H, L518V, H519L / Q / R / Y, A520E / S / T, P521L / Q / S, and T523A / S, resulting in 1 x 2 x 1 x 4 x 3 x 3 x 2 = 144 combinations for this epitope. Thus, the number of possible combinations of amino acid modifications within a single conserved epitope can be staggeringly large, and the 26 overall epitopes result in the large search landscape described above. [Table 3]
[0356] To sample the large resulting search space, the design technique in this example used process 500, shown in Figure 5G. Process 500 divides the creation of candidate engineered variants by introducing amino acid modifications into two steps. In the first step, specific (amino acid) positions to be modified are selected (510), generating multiple sets of position combinations to be mutated. In this example, each set of position combinations was generated by selecting positions until a position was selected within each of the 26 conserved epitopes. This process resulted in 200,000 sets of position combinations. As shown in Figure 5G, the position selection approach also included a filtering step 520, which calculated a position diffusion score for each set of position combinations and filtered the results of step 510 based on the position diffusion score and the number of mutations. After filtering, 2,000 combinations remained and were grouped into 20 clusters.
[0357] After arriving at a final set of position combinations, all feasible mutations were generated for each position combination in the set (530) (using the options available at each position from the set of allowed mutations). Modifications were filtered to ensure that no charge reversals were introduced and to minimize charge changes. This process produced a total of 190,945 candidate variants.
[0358] Four scores were then calculated for each candidate variant, as listed below (540). · ACE2 binding score; Likelihood (log likelihood) (RBD and complete spikes); Semantic changes for XBB.1.5 (RBD and full spike); and · Mutation co-occurrence;
[0359] The specific approaches used to calculate the ACE2 binding score and mutation co-occurrence score are described in more detail below. Methods for calculating likelihood and semantic change scores are described in detail in PCT Publication WO 2022 / 235847 A1 (entitled "TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST," published November 10, 2022) and PCT Publication WO 2022 / 235853 A1 (entitled "IMMUNOGEN SELECTION," also published November 10, 2022), the contents of each of which are incorporated herein by reference in their entirety. A final design was selected (550) using the Pareto front of all score and cluster information, followed by a final manual inspection.
[0360] Figures 6A-6F show the 3D structures of several engineered antigen designs. Tables 4A-4C list in each row a specific design and its ID, the mutations added, and the various scores and selection criteria described herein (identified in bold text in the first column of the table). Figures 6A-6E show the signature mutations of XBB.1.5 in red, positions where mutations were introduced in yellow, positions where no mutations were considered in gray (except for Figure 6B), and conserved surfaces in purple. Figures 6A and 6B show design S4_1, selected to maximize ACE2 binding. Figure 6C shows design S1_3, selected to minimize the number of mutations. Figure 6D shows design S14, selected based on the surface distribution of mutations. Figure 6E shows design S13_3, selected based on a high log-likelihood score, and Figure 6F shows design S11, which offers a diverse set of mutations. Tables 4A-4C list each of the final engineered antigen designs, showing the mutations added to XBB.1.5, the total number of mutations, the value of each of the four scores calculated for the design, and the selection criteria. [Table 4A-1] [Table 4A-2] [Table 4B-1] [Table 4B-2] [Table 4C-1] [Table 4C-2]
[0361] Figure 7 shows a radar plot of the scores for each of the final engineered antigen designs, showing good diversity.
[0362] Location Spread Score RBD candidates were scored using a location spread score, which estimates how spread out the location is across the (conserved) surface. In this example, the location spread score was calculated as follows:
number
[0363] Location Clustering A clustering algorithm based on that described in Cao et al. 2022 was used to cluster the designs according to the following steps: Given N sets of position combinations, 1. Each array was first represented in a binary encoding such that the embedding dimension = the number of positions. 2. Each sequence was again represented by the Pearson correlation between it and all other sequences, such that the embedding dimension = number of sequences. 3. Multidimensional scaling (MDS) was used to reduce the embedding space to 256.
[0364] Variant co-occurrence score A mutation co-occurrence score was calculated for each pair of mutations (mi, mj) by estimating the conditional frequency as follows:
number
[0365] In this example, for a given RBD with N mutations (relative to WT), the mutation co-occurrence score was calculated as the average log frequency between all pairs of mutations.
number
[0366] ACE2 binding calculation Finally, sequences were clustered using K-means with MDS embedding. In this example, ACE2 binding scores were calculated using deep mutation scan (DMS) data for ACE2 binding from Starr et al., 2022. This data includes DMS results for RBD:ACE2 binding at positions 331-531 for eight variants (Alpha, Beta, Delta, Eta, Omicron BA.1, Omicron BA.2, and two versions of wild type). ACE2 binding scores were calculated as the "delta binding" (log 10 The binding changes for any RBD variant are estimated by summing the ACE2 binding scores (changes in KD) for each RBD variant. When using non-omicron data, the resulting ACE2 binding scores showed a strong correlation with the BA.1 and BA.2 data, with Spearman r = 91.4%, Pearson r = 90.7%, and r 2 =83.5%. The predicted versus experimental change in binding is shown in Figure 8.
[0367] Designed sequence The approach in this example generated a list of engineered SARS-CoV 2 antigen mutations and sequences.
[0368] Tables 5A and 5B below list the engineered antigen designs generated by the approach described in this example. Table 5A lists the RBD mutations from this example design. Table 5B shows the RBD sequence for each design. As described herein, each sequence is an engineered version of the SARS-CoV 2 spike protein RBD generated using the XBB.1.5 variant RBD as a starting point. Figure 5C shows the performance metrics calculated for the sequences in Tables 5A and 5B below. For reference, Table 6 below lists the (native) XBB.1.5 RBD and Wuhan RBD sequences. Three versions of the XBB.1.5 RBD sequence are shown in Table 6, allowing for slight variations in certain portions (or boundaries) of the spike protein corresponding to the RBD region. [Table 5A-1] [Table 5A-2] [Table 5A-3] [Table 5A-4] [Table 5A-5] [Table 5A-6] [Table 5A-7] [Table 5A-8] [Table 5B-1] [Table 5B-2] [Table 5B-3] [Table 5B-4] [Table 5B-5] [Table 5B-6] [Table 5C-1] [Table 5C-2] [Table 5C-3] [Table 6]
[0369] iv. Example 4: Additional Engineered Variants Figures 9A-9B and 9D-9E show four manually designed RBD-engineered antigens. Figures 9A-9B show two constructs designed with different mutation spreads, focusing on the distance from the existing XBB.1.5 mutations. Table 7A lists the additional mutations added to XBB.1.5. Figure 9C is a schematic diagram of the SSARS-CoV-2 trimer. This figure is adapted from Starr et al., SARS-CoV-2 RBD antibodies that maximize breadth and resistance to escape. Nature, 597, 97-102 (2021). https: / / doi.org / 10.1038 / s41586-021-03807-6. Figures 9D-9E show two constructs designed to broadly promote conserved epitopes (class 4 and class 5 antibody sites) by disrupting epitopes identified as substantially conserved between SARS CoV 1 and SARS CoV 2 with mutations selected from the SARS CoV 1 sequence. Table 7B lists the additional mutations added to XBB.1.5. Tables 7C and 7D show the various scoring metrics calculated for each sequence. [Table 7A] [Table 7B] [Table 7C] [Table 7D]
[0370] Example 5: Distributing introduced amino acid modifications around signature mutations This example illustrates an embodiment of the antigen engineering techniques described herein, in which a positional diffusion score is used to promote the insertion of new amino acid modifications in a manner distributed across the conserved surface (as opposed to, e.g., clustered), taking into account not only other introduced amino acid modifications but also proximity to existing signature mutations (e.g., XBB signature mutations). Notably, the antigen engineering method used in this example proceeded similarly to that in Example 3 above, but also included signature mutations in the positional diffusion score (e.g., as described in Example 3 above), which was used to evaluate candidate engineered antigen designs. Notably, it is believed that this approach ensures that additional mutations are avoided from being placed immediately adjacent to signature mutations. In this manner, the embodiment described in this example maintains unique (e.g., omicron) epitopes in an unaltered form and preferentially mutates conserved epitopes (e.g., only epitopes). For example, in certain cases, approaches in which characteristic mutations are not explicitly considered in this manner may create a certain likelihood of altering the omicron epitope in such a way that the elicited antibody binds with lower affinity to the true desired target epitope.
[0371] Tables 8A and 8B below list the engineered antigen designs generated by the approach described in this example (i.e., including XBB signature mutations in the position diffusion score). Table 8A lists the RBD mutations from the designs in this example. Table 8B lists the RBD sequence for each design. Table 8C lists the various scoring metrics calculated for each sequence. [Table 8A-1] [Table 8A-2] [Table 8A-3] [Table 8A-4] [Table 8A-5] [Table 8A-6] [Table 8A-7] [Table 8B-1] [Table 8B-2] [Table 8B-3] [Table 8B-4] [Table 8B-5] [Table 8B-6] [Table 8B-7] [Table 8C-1] [Table 8C-2] [Table 8C-3]
[0372] vi. Example 6: Evolutionary Algorithm and Epitope Variation Score Version This example describes a particular approach for mutagenesis, scoring, and evaluation that may be used in certain embodiments in addition to or as an alternative to the various approaches described herein.
[0373] In certain embodiments, mutations may be determined using an evolutionary algorithm (e.g., in addition to or alternatively to, or in conjunction with, the two-stage position and type selection technique described herein). For example, each solution may be represented as a list of 26 mutations (in the case of the SARS-CoV-2 RBD) corresponding to the 26 conserved epitopes described in the examples above. In certain embodiments, a pool of solutions may be maintained and updated by swapping mutations on an epitope-by-epitopes basis. Solutions may be scored using position diffusion, ACE2 binding, and machine learning-based (ML) scores (e.g., log-likelihood and semantic change).
[0374] In certain embodiments, different versions of the epitope alteration score may be used. For example, current versions of the epitope alteration score (e.g., as described in PCT Publication WO 2022 / 235847 A1 (entitled "TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST," published November 10, 2022) and PCT Publication WO 2022 / 235853 A1 (entitled "IMMUNOGEN SELECTION," also published November 10, 2022), the contents of each of which are incorporated herein by reference in their entireties) consider mutations at any single position within an epitope sufficient to "escape" the corresponding antibody. In certain embodiments, a more stringent version of the epitope alteration score may be used, for example, that considers the positional placement and distance between the mutation and the CDR loops of the antibody to provide increased granularity and precision.
[0375] In certain embodiments, in silico structural modeling may be used to assess complex kinetics.
[0376] vii. Example 7: Updated Design with Integrity Risk Balancing This example describes a particular approach for mutation generation, scoring, and evaluation that may be used in certain embodiments in addition to or as an alternative to the various approaches described herein. In particular, this example shows sequences generated to balance the risk if a particular mutation adversely affects protein integrity. For example, it was found that a particular in silico sequence generation procedure may over-represent mutations at positions 384, 430, and 463 within the proposed design set. Thus, the design presented in this example sought to balance the risk (e.g., with respect to mutational diversity across the entire design set) if a particular in silico suggested mutation adversely affects RBD integrity.
[0377] Table 9A below lists RBD mutations (e.g., by any one of the three XBB.1.5 RBDs shown in Table 6) related to (e.g., added to) the XBB.1.5 RBD portion of the SARS-CoV-2 S protein according to the design in this example. As described herein, the mutation position identifies the position in the context of (i.e., with reference to) the full-length SARS-CoV-2 S protein.
[0378] Table 9B shows the scoring metrics calculated for each sequence shown in Table 9A. [Table 9A] [Table 9B-1] [Table 9B-2]
[0379] viii. Example 8: Exemplary in silico designed testing procedure This example describes an exemplary experimental procedure for testing engineered synthetic variants designed by the various approaches described herein. The procedure in this example uses three backbones / constructs to evaluate the engineered variants, as follows:
[0380] pcDNA.3.1-SARS-CoV-2-XBB.1.5-CΔ19 construct: A mammalian expression plasmid encoding the XBB.1.5 spike protein with a truncated cytoplasmic tail (harboring the respective mutated XBB.1.5 RBD). Use of this construct is expected to enable (i) confirmation of spike protein expression after transfection into HEK293T cells, (ii) production of pseudoviruses, and (iii) analysis of immune escape parameters (e.g., complete escape with no detectable titer).
[0381] pST4-v.3.0.1-AGA-hAG-SP19-XBB.1.5-RBD-Foldon-TM construct: A template for transcription of mRNA encoding a transmembrane-anchored RBD-based vaccine antigen, such as BNT162b3. This construct is expected to enable (i) confirmation of RBD expression in HEK293T cells, (ii) evaluation of antibody binding using a reference panel of various RBD epitope class binders and / or composite immune sera, and (iii) immunogenicity studies.
[0382] pcDNA3.4-SARS-CoV-2-XBB.1.5-RBD-his-avi construct: A template for RBD protein production with a HIS tag (to allow for purification) and a BAP / avi tag (to allow for assay development). This construct may be used in enzyme-linked immunosorbent assays (ELISAs), biolayer interferometry / surface plasmon resonance (BLI / SPR), and / or protein-protein interaction assays to assess, for example, whether a known antibody can bind to the generated construct.
[0383] An exemplary procedure for testing a design may include one or more (e.g., up to all) of the steps listed below and / or shown in FIG. 10A, arranging the testing steps in a hierarchical manner.
[0384] Confirmation of Expression and Folding: In step 1001, the expression and folding of various antigens (e.g., each designed antigen) is assessed using a flow cytometry-based approach. An example of a FACS staining protocol that can be used in connection with this approach is provided below.
[0385] As shown in Figure 10B, after transfection, intracellular and surface expression of the encoded antigen can be analyzed using flow cytometry with hACE2-mFc as the primary binding agent. In this approach, the full-length XBB S protein may be used as a reference. The binding affinity of the RBD protein may also be assessed, along with its intracellular and / or extracellular surface expression. The results of this process may be compared to predictions from, for example, machine learning algorithms used in the in silico design of engineered antigens and fed back into the in silico design algorithm (e.g., to improve it).
[0386] Antibody escape - monoclonal and polyclonal: Antibody escape is assessed using a flow cytometry approach using a panel of selected reference antibodies that bind to different epitope classes on the RBD (e.g., similar to the assay used to confirm expression and folding, but using reference antibodies as binding agents instead of hACE2-mFc).
[0387] Various reference antibodies may be used and / or selected based on the particular epitope class to which they bind, and additionally or alternatively, based on their level of affinity for a particular SARS-CoV 2 variant (e.g., reference antigen) that will be used as a starting point (i.e., mutated) to design the engineered variant. For example, a panel of reference antibodies may include multiple antibodies that bind to different epitope classes (e.g., A, B, C, D, E, and F). Antibodies that bind to a particular SARS-CoV 2 variant (used as a reference antigen to design an engineered version thereof (e.g., with disrupted conserved regions)) may be included along with antibodies that do not bind to that particular SARS-CoV 2 variant (but that bind to other variants).
[0388] For example, Cao et al., Nature, 2021 (doi.org / 10.1038 / s41586-021-04385-3) (Supplementary Table 1) provides a list of 247 neutralizing antibodies from which a panel can be selected. For example, Table 10A shows a subset of 14 antibodies from the table by Cao et al. that can be used as reference antibodies. Figure 10C shows example binding assay data for some of the antibodies listed in Table 10A. Antibodies that demonstrate XBB.1.5 binding are shown in Table 10B, along with IC50 values for wild-type and BA.1 variants (* indicates data provided in Cao et al.). Table 10C shows the antibodies from Table 10A along with XBB.1.5 binding data, with antibodies that bind to XBB.1.5 highlighted in green text. Without wishing to be bound by any particular theory, it is expected that certain classes of antibodies (such as antibodies that bind to class A-D epitopes) will not bind to the BA.1 S protein and therefore will not bind to the XBB S protein. Certain antibodies that bind to class E and class F are confirmed for binding to the XBB S protein. The abolition / reduction of binding of the reference antibody is determined by titrating EC in step 1002. 50 The value can be used to evaluate. [Table 10A] [Table 10B] [Table 10C]
[0389] In addition to or alternatively to the various classes of antibodies listed in Tables 10A-10C, the antibody panel may include various other exemplary antibody clones, including (but not limited to) the antibodies listed below and / or similar antibodies. WRAIR-2057 (without wishing to be bound by any particular theory, this antibody is thought to bind at a previously highly conserved site on the flank of the RBD and to have high affinity for omicron variants). WRAIR-2057 is described in further detail, for example, in Dussupt, V. et al., Low-dose in vivo protection and neutralization across SARS-CoV-2 variants by monoclonal antibody combinations. Nat Immunol 22, 1503-1514 (2021). https: / / doi.org / 10.1038 / s41590-021-01068-z. COVOX-222 is included in Supplementary Table 1 of Cao et al., Nature, 2021 (doi.org / 10.1038 / s41586-021-04385-3). COVOX-45 (without wishing to be bound by any particular theory, this antibody is thought to bind at a previously highly conserved site on the flank of the RBD and to have high affinity for omicron variants). COVOX-45 is described in further detail, for example, in: Dejnirattisai, W. et al., The antigenic anatomy of SARS-CoV-2 receptor binding domain, Cell, 184 (8), 2183-2200.e22, 2021, https: / / doi.org / 10.1016 / j.cell.2021.02.032. S2H97 (without wishing to be bound by any particular theory, this antibody is thought to bind at a previously highly conserved site on the flank of the RBD and to have high affinity for omicron variants). S2H97 is described in further detail, for example, in: Starr, TN, et al. SARS-CoV-2 RBD antibodies that maximize breadth and resistance to escape. Nature 597, 97-102 (2021). https: / / doi.org / 10.1038 / s41586-021-03807-6. 553-49 is described in further detail, for example, in: Zhan W et al., Structural Study of SARS-CoV-2 Antibodies Identifies a Broad-Spectrum Antibody That Neutralizes the Omicron Variant by Disassembling the Spike Trimer. J Virol. 96(16):e0048022. 2022. doi: 10.1128 / jvi.00480-22.
[0390] Antibody escape may additionally or alternatively be assessed using immune sera from vaccinated or convalescent individuals. The abrogation / reduction of binding of the combined immune sera is assessed by titrating EC50 The value can be used to evaluate the XBB1.5 value and compared with it.
[0391] A third approach to assess escape may employ a pseudovirus generation protocol and pVNT assay setup to confirm loss of nAb titer, with the goal of confirming abrogated neutralization of each pseudovirus in step 1004.
[0392] The final step may include, for example, a dedicated immunogenicity study to evaluate whether serum from immunized animals neutralizes XBB1.5 and / or one or more other variants, as shown, for example, in Figure 10A. Vaccine compositions based on engineered SARS-CoV 2 antigens according to various embodiments described herein may be administered to vaccine-naive animals (e.g., mice) (e.g., to evaluate the immune response and its magnitude) and to vaccine-experienced animals (e.g., mice) to evaluate, for example, their ability to overcome immune imprinting.
[0393] Figure 10B is an exemplary protocol showing the steps of transfection and flow cytometry. As shown in Figure 10B, pcDNA3.1-SARS-CoV-2-Swt-CΔ19 and / or pcDNA3.1-SARS-CoV-2-SXBB.1.5-CΔ19 may be transfected into HEK293T / 17 cells, and hACE2 binding and (neutralizing) antibody binding may be assessed using the binding assays described above, e.g., implementing the various steps in the decision tree shown in Figure 10A.
[0394] The concentration range of the dilution series of monoclonal antibodies, hACE2, and human BNT162b23 (triple vaccine) polyclonal serum that results in a sigmoidal binding curve is appropriately identified. Based on this criterion, for example, as shown in Figure 10C, based on previous apparent IC50 values for the WT and BA.1 spikes, the mAb can be tested at a concentration range of 10 μg / mL to 0.128 ng / mL (eight 5-fold dilution series). As shown in Figure 10D, hACE2-mFc binding may be tested using a concentration range of 25 μg / mL to 0.32 ng / mL. Figure 10D also shows that the XBB.1.5 spike exhibits approximately 6.7-fold higher apparent affinity for hACE2 binding compared to the wild-type spike. As shown in Figure 10E, the triple vaccine 1M post-boost polyclonal serum pool can be tested at a dilution series ranging from 1:20 to 1:1.562.500.
[0395] Table 11 lists some variants of concern (VOC) mutations and their numerical evaluation indices, which may be used as a reference. [Table 11-1] [Table 11-2]
[0396] An exemplary FACS staining protocol is shown below: Tables 12A and 12B below show the master mix used for the secondary antibody labeling step. Staining with antibodies, hACE2-mFc or human serum pool Separating cells Add PBS and centrifuge the cells Counting cells: Distribute viable cells / well into a 96 round-bottom well plate according to the plate layout. Centrifuge the plate containing the cells Discard the supernatant and vortex the plate to single the cells. Add FACS buffer and centrifuge the plate Add primary antibody, ACE2-mFc or human serum pool and gently tap the plate to mix the cells. Incubate in the dark at 4°C for approximately 15 minutes. Add FACS buffer and centrifuge the plate Discard the supernatant and vortex the plate to single the cells. Wash twice with FACS buffer and centrifuge the plate Add the master mix (MM) containing the secondary antibody and gently tap the plate to mix the cells. Incubate in the dark at 4°C for approximately 15 minutes. Add 200 μL / well of FACS buffer (DPBS + 2% FBS hi. + 2 mM EDTA) and centrifuge the plate: 5 min, 460 × g, room temperature Discard the supernatant and vortex the plate to single the cells. Wash twice with FACS buffer and centrifuge the plate: Add Histofix (in PBS) to the cells, resuspend, and incubate at 2-8°C for 15 minutes. Centrifuge the cells (5 min, 450 x g) Discard the supernatant and vortex the plate to single the cells. Wash twice with FACS buffer and centrifuge the plate Resuspend the cells in FACS buffer Store the plate in the dark at 4°C until measurement. [Table 12A] [Table 12B]
[0397] ix. Example 9: Evaluation of in silico designed constructs by binding assays This example describes the results of various assays performed on certain constructs described herein, following the exemplary testing procedures described in Example 8 above.
[0398] The first round of testing was performed on engineered XBB.1.5 variants expressed in the context of the full-length SARS-CoV-2 spike (S) protein. Specifically, the specific constructs described herein were incorporated into the pcDNA.3.1-SARS-CoV-2-XBB.1.5-CΔ19 construct, which harbors the respective mutated XBB.1.5 RBD. HEK293T / 17 cells were seeded into flasks and incubated at 37°C and 7.5% CO for 2 days. The constructs were transfected into HEK293T / 17 cells and incubated overnight at 37°C and 7.5% CO. Flow cytometry was used to (i) assess expression levels via anti-S2 fragment antibodies, (ii) confirm preserved ACE2 binding capacity (via hACE2 binding), and (iii) assess abrogation of RBD-targeted antibody binding using a panel of five monoclonal antibodies (mAbs) with demonstrated binding to XBB.1.5, as shown in Table 10B.
[0399] Referring to Figures 11A-11C, binding tests were performed on specific variant designs shown in Table 9A expressed in the context of the full-length spike (S) protein. In particular, ACE-2 binding was assessed using the hACE2-mFc antibody according to the FACS protocol described in Example 8 above. Binding responses to each of the five mAbs listed in Table 10B were also assessed by the protocol described in Example 8 above. Figures 11A-11C show results for the S43-engineered XBB.1.5 variant design (Figure 11B) and the S48-engineered XBB.1.5 variant design (Figure 11C), as well as the (parent / original) XBB.1.5 reference (Figure 11A). Notably, it was found that both the S43 and S48 variants abolished binding to all but one or two of the five mAbs, with the S43 variant maintaining hACE-2 binding.
[0400] Referring to Figures 12-14, the variant designs shown in Table 9A were also expressed in the context of a BNT162b3-like (trimerized transmembrane-anchored RBD-based) vaccine antigen. HEK293T / 17 cells were transfected with RNA for BNT162b3-XBB.1.5 and each of the variants. ACE-2 binding was then assessed using flow cytometry, as before, via hACE2-mFc binding and binding to a panel of five mAb antibodies (listed in Table 10B). Binding curves for polyclonal vaccine sera (triple BNT162b2 vaccine) were also evaluated to assess whether there was a greater impact on the polyclonal serum dose response compared to the hACE-2 dose response of the engineered variants. Without wishing to be bound by any particular theory, the stronger impact on the polyclonal serum dose response compared to the hACE-2 dose response for a particular variant when compared to the parent XBB.1.5 may suggest successful "masking / mutation" of conserved epitopes.
[0401] Figures 12A-B show dose-response curves for the XBB.1.5 reference (Figure 12A) and the S43 variant design (Figure 12B). As shown in Figure 12B, design S43 expressed in the context of a trimerized TM-anchored RBD exhibits a nearly unchanged hACE-2 binding dose response when compared to XBB.1.5. Importantly, binding of Class A Antibody 2, Class F Antibody 1, and Class B Antibody 1 is abolished, while only binding of Class F Antibody 2 and Class E Antibody 2 is preserved.
[0402] Figures 13A and 13B show results that replicate those of Figures 12A and 12B (Figure 13A shows the reference dose-response curve, and Figure 13B shows the engineered antigen design S43 dose-response curve) and confirm the results shown in Figure 12B: hACE-2 binding and mAb binding of Class F Antibody 2 and Class E Antibody 2 are preserved, while binding of Class A Antibody 2, Class F Antibody 1, and Class B Antibody 1 is abolished. Polyclonal serum binding was also compared to hACE2 binding, as shown in Figures 13C and 13D. The shift in polyclonal serum binding curves (engineered XBB.1.5 antigen versus parental XBB.1.5) was similar to, but not greater than, hACE-2 binding, and the abolished mAb binding observed in multiple sets of results shown in Figures 12A and 12B, and Figures 13A and 13B, provides evidence that the mutations introduced into the engineered S43 design successfully disrupted the conserved epitope.
[0403] Figures 14A and 14B show another set of dose-response curves for the XBB.1.5 reference and for the S48 engineered antigen design from Table 9A, expressed in the context of a trimerized TM-anchored RBD. S48, like S43, shows conserved hACE-2 and Class F Antibody 2 binding (S48 retains 5 / 7 mutations also found in S43). Class E Antibody 2 shows minimal residual binding, while binding of Class A Antibody 2, Class F Antibody 1, and Class B Antibody 1 is completely abolished.
[0404] Referring to Figures 14C and 14D, additional data are presented suggesting that disruption of a conserved epitope in S48 may be successful. Figures 14C and 14D compare (i) the shift in hACE2 binding curve for the S48 variant relative to the parental XBB.1.5 reference (Figure 14C) and (ii) the shift in the binding curve of polyclonal sera (Figure 14D). The shift in ACE2 binding (approximately 4-fold) is smaller than the shift observed in serum binding (>10-fold) when comparing the parental XBB.1.5 with S48, suggesting a progressive abrogation of the binding antibody response.
[0405] Table 13 summarizes the results of the screening data described in this example in tabular form. The first two rows of the table show the binding levels of XBB.1.5 and the ideal, desired engineered antigen. The columns representing binding data are organized into two groups corresponding to the different expression contexts analyzed: (i) full-length XBB.1.5 spike (pcDNA.3.1-SARS-CoV-2-XBB.1.5-CΔ19); and (ii) trimerized transmembrane (TM)-anchored RBD design (pST4-v.3.0.1-AGA-hAG-SP19-XBB.1.5-RBD-Foldon-TM). As shown in Table 13, the reference antigen XBB.1.5 exhibits strong ACE-2 binding and binds to all five mAbs in the panel (selected to be mAbs with known XBB.1.5 binding). The ideal engineered variant would need to evade all five mAbs known to bind to XBB.1.5 while retaining ACE-2 binding, avoiding the induction of a memory immune response. As shown in the table, constructs S43 and S48 exhibit behavior similar to the ideal constructs. S43 exhibits conserved ACE-2 binding when expressed as a full-length S protein and even in a trimerized TM-anchored RBD format, but abolishes binding to all but one or two of the mAb panel, depending on the expression context. Construct S48 performed particularly well when expressed as a trimerized TM-anchored RBD, exhibiting conserved ACE-2 binding and binding to only one of the five mAbs. [Table 13]
[0406] x. Example 10: Prophetic Mouse Immunization Studies This example describes the designed procedures and expected results of an experiment to evaluate whether RNA compositions encoding engineered antigens comprising the SARS-CoV-2 S protein and / or portions thereof (e.g., the RBD domain) described herein induce an immune response characterized by increased naive B cell activation and / or decreased memory B cell activation in vaccinated subjects (in this example, mice).
[0407] With reference to Figure 15A, vaccine candidates containing RNA encoding engineered antigens designed and screened according to the approaches described herein are administered to mice previously exposed to the full-length SARS-CoV-2 S protein. Specifically, as shown in Figure 15A, the vaccine candidates are tested in mice previously administered two doses of RNA encoding the full-length SARS-CoV-2 S protein, with each engineered antigen vaccine candidate being administered as a third dose (booster).
[0408] Mice are separated into groups, each containing approximately the same number (x) of members, and a dosing regimen is administered as shown in Figure 15A. The total number of groups depends on the number of engineered antigen candidates to be tested, as well as their different expression formats. For example, Figure 15 shows an exemplary immunization study in which engineered antigen designs S43 and S48, having the sequences listed in Table 9A, are evaluated in the context of (i) a full-length S protein encoding a specific (e.g., S43 or S48) engineered XBB.1.5 variant similar to BNT162b2, and (ii) an mRNA encoding a trimerizing TM-anchored RBD domain. Other candidate antigens and / or expression contexts may be included as well. In one example, the number of mice in each group will be approximately 7-8. As shown in Figure 15A, mice in each group are first administered two doses of a monovalent composition containing RNA encoding the SARS-CoV-2 S protein of the Wuhan strain (BNT162b2). The first and second doses of the monovalent vaccine are administered 21 days apart. Five weeks after the first dose, the mice are grouped based on their neutralization titer against the Wuhan strain (e.g., by assigning mice to each group so that the average neutralization titer for each group is approximately the same). In some embodiments, mice can be assigned to each group based on their pseudovirus neutralization titer (e.g., as shown in Figure 15A).
[0409] Each candidate vaccine is then administered as a third dose 18 weeks after administration of the first dose of vaccine (i.e., day 126 as shown in FIG. 15A). In certain embodiments, a version of the immunization study approach shown in FIG. 15A may be implemented with a third (candidate vaccine) dose administered on a shorter timeline after the first two doses. Without wishing to be bound by any particular theory, the third candidate dose may be administered quickly after the second dose, as long as sufficient time is allowed for the immune response (e.g., B cell generation) to be completed in response to the initial (e.g., Wuhan) antigen. In certain embodiments, 28 days (4 weeks) or more may be sufficient, e.g., the first dose (of BNT162b2) may be administered on day 0, the second dose (of BNT162b2) on day 21, and the vaccine candidate may be administered as a third (e.g., booster) dose 4 weeks later, e.g., on day 49 (week 7) or later.
[0410] Various engineered antigen designs and expression context formats may be administered as a third dose for comparison. For example, as shown in FIG. 15A, candidates S43 and S48 may be administered as a third dose in full spike (S) format (BNT162b2(S43) and BNT162b2(S48)), as well as in trimerized transmembrane-anchored RBD format (RBD-TM (S43) and RBD-TM (S48)). In certain embodiments, multiple control and / or reference vaccines may also be used. For example, as shown in FIG. 15A, a third dose of BNT162b2 may be administered. Additionally or alternatively, full spike (S) and RBD-TM versions of the parent XBB.1.5 variant may be administered for comparison with the engineered variant. In certain embodiments, one test group may not receive any third dose. Table 14A below lists an exemplary set of vaccine candidates (selected based on screening data described herein (e.g., previous examples)), some of which are also listed in Figure 15A. Some options for certain other candidate vaccines that may additionally or alternatively be used are listed in Table 14B. Tables 14A and 14B provide descriptions of vaccine candidate formats and references to exemplary sequences contained in Tables 14C-14I. Other vaccine candidates containing engineered antigens (e.g., variants of other VOCs) and reference compositions (vaccine doses) can be prepared and evaluated in a manner similar to that described herein for XBB.1.5. [Table 14A] [Table 14B] [Table 14C-1] [Table 14C-2] [Table 14C-3] Table 14C-4 Table 14D-1 Table 14D-2
Table 14D-3
Table 14D-5
Table 14D-6
Table 14D-7
Table 14D-8
Table 14D-9
Table 14D-10
Table 14D-11
Table 14F-1
Table 14F-2
Table 14F-3
Table 14F-4
Table 14F-5
Table 14F-6
Table 14F-7
Table 14F-8
Table 14G-1
Table 14G-2
Table 14G-3
Table 14G-4
Table 14G-5
Table 14G-6
Table 14G-7
Table 14G-8
Table 14G-9
Table 14G-10
Table 14G-11
Table 14G-12
Table 14G-13
Table 14H-2
Table 14I-1
Table 14I-2
Table 14I-3
Table 14I-5
Table 14I-6
Table 14I-7
Table 14I-8
[0411] Blood samples are collected immediately prior to administration of the first dose of RNA and at 3, 5, 9, 13, 17, 18, 19, 22, 26, 30, 34, 35, 37, and 39 weeks after administration. 39 weeks after administration of the first dose of RNA, mice are sacrificed and terminal blood, lymph node, and spleen samples are collected for analysis.
[0412] Blood sample analysis Blood samples are screened for titers of antibodies that bind and neutralize various SARS-CoV-2 strains and variants (e.g., using the ELISA and pseudoviral assays described herein). Each spleen sample can be analyzed individually. For lymph node samples, samples from two mice can be combined and analyzed.
[0413] Spleen and lymph node sample analysis B cells are isolated from the spleen and lymph nodes and phenotyped to determine B cell type (e.g., naive, memory, or plasma) and binding specificity (e.g., specificity for the S protein of various SARS-CoV-2 strains and variants). Phenotypic analysis can be performed using FACS-based and depletion assays similar to those shown in Figures 16A and 16B. Further details are described in Quandt and Muik et al., Science Immunol., 7(75) eabq2427 (2022) (doi / 10.1126 / sciimmunol.abq2427), the contents of which are incorporated herein by reference in their entirety. Pseudovirus neutralization titers for each serum sample are also collected. BCR repertoire analysis is also performed on B cells isolated from spleen samples.
[0414] For RBD binding assays, samples from all mice can be grouped together to generate enough samples to perform the method, and each sample can be screened for negative Wuhan-specific XBB.1.5-only binders and for cross-reactivity between Wuhan and XBB.1.5 binding.
[0415] In particular, the protocol described in this example can be used to characterize the binding specificity of B cells from subjects who have received a booster vaccine delivering an antigen of a variant of concern (here, the XBB.1.5 S protein or an immunogenic portion thereof). Specifically, this experimental protocol can be used to determine the relative number of XBB.1.5-specific B cells and / or which portion of the XBB.1.5 S protein those B cells recognize. These results can be used to assess the impact of immune imprinting on various vaccine candidates and the ability of a particular vaccine candidate to circumvent immune imprinting and elicit a de novo immune response.
[0416] Vaccine candidates that are less susceptible to immune imprinting (i.e., more likely to generate a de novo response) can be characterized by one or more of the following: (i) an increased proportion of B cells specific for the variant of concern (i.e., XBB.1.5 or an immunogenic portion thereof) delivered by the vaccine candidate, (ii) an increased neutralizing titer against the variant of concern encoded by the vaccine candidate, and / or (iii) an increased number of B cell receptors that recognize epitopes unique to the antigen encoded by the vaccine candidate (i.e., increased B cell breadth).
[0417] Additional analytical techniques include analysis of spleen samples, analysis of lymph nodes, and characterization of blood samples, as shown in Figure 17. Analysis of spleen samples involves the collection of 5 x 10 cells per spleen after isolation and red blood cell (RBC) depletion (80-90% cell viability). 7This may involve the preparation of a single cell suspension that yields more than 1.5 x 10 leukocytes per spleen. 7 Leukocytes may be used for FACS phenotypic analysis for immunogenicity testing. This analysis may involve, for example, naive cells, memory cells, and plasma cells specific for different S proteins. Approximately 1.5 × 10 leukocytes per spleen may be tagged and pooled for subsequent magnetic-activated cell sorting (MACS). Subsequent steps may include FACS sorting and staining for memory and plasma cell populations. Another subsequent step may include BCR repertoire analysis. The remaining leukocytes extracted from the spleen may be used for enzyme-linked immunosorbent spot (ELISpot) assays, and the remaining samples may be frozen, or frozen and ELISpot. Lymph node analysis may involve approximately 1 × 10 leukocytes per inguinal (iLN) and pelvic (pLN) lymph node. 6 This may include preparation of a single cell suspension, yielding 100 white blood cells. This analysis may include FACS phenotyping for immunogenicity testing. This analysis may involve, for example, naive, memory, and plasma cells specific for different S proteins. Finally, characterization of blood samples may include pseudotyped virus neutralization test (pVNT) ELISA.
[0418] Immune responses in vaccine-naive mice Referring to Figure 15B, in certain embodiments, vaccine candidates may be tested in vaccine-naive mice along with a reference. As shown in Figure 15B, testing in vaccine-naive mice may be performed on a shorter timescale and can be used to evaluate and / or confirm whether vaccine candidates based on various engineered antigenic compounds can elicit a meaningful immunogenic response against XBB.1.5 alone. Figure 15B illustrates an exemplary protocol for administering the vaccine candidates and reference described herein to vaccine-naive mice in a two-dose format, with the first dose administered on day 0 and the second dose administered three weeks later on day 21.
[0419] Blood samples can be collected immediately before the first dose of RNA is administered, and 2, 3, 4, and 7 weeks after administration. Seven weeks after the first dose of RNA is administered, the mice are sacrificed, and terminal blood, lymph node, and spleen samples are collected for analysis. Neutralizing antibody production can be tested by pVNT assay using the XBB.1.5 pseudovirus. Binding antibodies can be tested by ELISA using the XBB.1.5 RBD as the target. Ideal candidates successfully maintain neutralizing antibody titers, e.g., at levels comparable to those of the parent XBB.1.5 (e.g., RBD-TM(XBB.1.5)), while simultaneously exhibiting reduced binding antibody titers in ELISA tests, e.g., due to disruption of conserved regions.
[0420] xi. Example 11: Evaluation and selection of individual engineered RBD mutations for engineered antigen design This example describes the evaluation and selection of RBD mutations introduced into certain engineered antigens described herein. In particular, mutations present in various constructs were individually evaluated and a subset selected to generate further fine-tuned engineered antigens.
[0421] In particular, analysis of various in silico designed constructs provided herein identified specific mutations introduced into construct S48 as having an effect on expression. As described herein, construct S48 corresponds to the SARS-CoV-2 S protein RBD containing (i) the hallmark mutations of XBB.1.5, and also (ii) an additional set of introduced mutations engineered to disrupt memory-inducing conserved regions of the baseline, reference, and XBB.1.5 reference antigens. As shown in Table 9A, the additional set of introduced mutations in construct S48 included the following mutations: Additional mutations in S48: L335F, K356T, P384S, L390R, T430I, F464Y, and H519N.
[0422] Of these S48-added mutations, a subset was determined to be beneficial for expression of constructs containing the S48 RBD design, while another subset was selected for further evaluation, e.g., as potentially suboptimal and / or potentially causing reduced expression. Mutations determined to be beneficial included, but were not limited to, K356T, P384S, T43OI, F464Y, and H519N. Mutations selected for further evaluation included L335F and L390R.
[0423] Based on these identified mutations, additional rounds of construct design based on the XBB.1.5 RBD reference antigen (i.e., including the hallmark mutations of XBB.1.5) were created, in which (i) the set of mutations determined to be beneficial was retained, (ii) two mutations selected for further evaluation were removed from the majority of constructs (none contained both the L335F and L390R mutations), and (iii) an additional set of new mutations was introduced.
[0424] New mutations were introduced following a similar approach as described in Example 3, with reference to operation 500 in Figure 5G.
[0425] Figure 18 shows this additional set of adapted constructs. Construct design indicators are displayed vertically, with each row corresponding to a specific construct design, and a list of potential mutations is displayed horizontally. The signature mutation of XBB.1.5 is identified by a black asterisk ("*") symbol, mutations determined to be beneficial are identified by a green star, two mutations selected for further evaluation are identified by a red downward-pointing triangle, and new mutations are identified by a purple diamond. Shading (dark green) identifies mutations present in a particular construct. As shown in Figure 18, all constructs contained the signature mutation of XBB.1.5 as well as five additional mutations determined to be beneficial. Constructs S122-S127 and S128 contained the L335F mutation, and constructs S145 and S156 contained the L390R mutation. New mutations were introduced into the various constructs, as shown in Figure 18. Tables 15A and 15B below list the RBD mutations (including the XBB.1.5 signature mutations) and additional engineered RBD mutations (i.e., in addition to the reference XBB.1.5 mutations) for each construct design, respectively.
[0426] The construct designs shown in Figure 18 and identified in Tables 15A and 15B below were then tested by incorporating each construct design's set of additional mutations (i.e., those listed in Table 15 below) along with the XBB.1.5 signature mutations into BNT162b3 mRNA for in vitro testing. In particular, the BNT162b3 construct shown in Figure 21A is a 1397 base pair mRNA encoding a membrane-anchored RBD and fibritin domain (F) along with a viral signal peptide. As shown in Figure 19A, this construct contains a secretion signal ("sec"), an RBD domain ("RBD"), a fibritin domain ("F"), and a transmembrane anchor ("TM"). For each construct design, the RBD-encoding domain was modified to encode the XBB.1.5 signature mutations and each construct design's specific set of additional mutations.
[0427] To assess expression, ACE-2 binding was assessed using flow cytometry with an hACE2-mFc binding assay. This assay was performed on the 22 construct designs shown in Figure 18 and listed in Table 15, as well as the XBB.1.5 and S48 constructs as reference points. The results are shown in Figure 19B and Table 16A below. To assess immune escape (e.g., the constructs' ability to evade antibodies generated by a prior Wuhan-induced immune response and potentially induce a de novo immune response), binding to polyclonal vaccine sera from patients receiving triple BNT162b2 vaccinations was measured. Results for the 22 construct designs, as well as the XBB.1.5 and S48 constructs, are shown in Figure 19C and Table 16B below. The values in Tables 16A and 16B were obtained from duplicate experiments and represent the results in terms of area under the curve (AUC) values and flow cytometry intensity (FC).
[0428] As shown in Figure 19B and Table 16A, all 22 construct designs from Figure 18, Table 15A, and Table 15B exhibited higher binding to ACE2 and higher expression than the S48 construct design. In terms of immune escape, the S48 construct design exhibited the most significant reduction in binding to serum from patients triple-vaccinated with BNT162b2. Several construct designs listed in Tables 15A and 15B exhibited significant reductions in (BNT162 serum) binding. In particular, seven constructs exhibited a greater than 1.5-fold reduction in serum binding, while none exhibited a greater than 2-fold reduction in ACE2. These construct designs are identified by an asterisk (*) in Tables 16A and 16B below.
[0429] Figures 20A-20C show the binding affinity of XBB.1.5, S48, and the seven construct designs shown in Figure 18 to various SARS-CoV-2 monoclonal antibodies. Figure 20A shows the results for five antibodies for which the seven construct designs showed complete restoration of binding, Figure 20B shows the results for five antibodies for which the seven construct designs showed partial restoration of binding, and Figure 20C shows the results for two antibodies for which the seven constructs showed little to no binding. The results appear consistent with the polyclonal serum data, although S48 still shows the most significant reduction in binding with the monoclonal antibodies studied. [Table 15A-1] [Table 15A-2] [Table 15A-3] [Table 15A-4] [Table 15B-1] [Table 15B-2] [Table 16A] [Table 16B]
[0430] G. Computer System and Network Environment As shown in FIG. 21 , an implementation of a network environment 2100 for use in providing the systems and methods described herein is shown and described. Briefly, referring now to FIG. 21 , a block diagram of an exemplary cloud computing environment 2100 is shown and described. The cloud computing environment 2100 may include one or more resource providers 2102 a, 2102 b, and 2102 c (collectively 2102). Each resource provider 2102 may include computing resources. In some implementations, computing resources may include any hardware and / or software used to process data. For example, computing resources may include hardware and / or software capable of executing algorithms, computer programs, and / or computer applications. In some implementations, exemplary computing resources may include application servers and / or databases with storage and retrieval capabilities. Each resource provider 2102 may be connected to any other resource providers 2102 within the cloud computing environment 2100. In some embodiments, the resource providers 2102 may be connected through a computer network 2108. Each resource provider 2102 may be connected to one or more computing devices 2104 a , 2104 b , 2104 c (collectively 2104 ) through a computer network 2108 .
[0431] The cloud computing environment 2100 may include a resource manager 2106. The resource manager 2106 may be connected to the resource providers 2102 and the computing devices 2104 through a computer network 2108. In some implementations, the resource manager 2106 may facilitate the provision of computing resources by one or more resource providers 2102 to one or more computing devices 2104. The resource manager 2106 may receive a request for a computing resource from a particular computing device 2104. The resource manager 2106 may identify one or more resource providers 2102 that can provide the computing resource requested by the computing device 2104. The resource manager 2106 may select a resource provider 2102 to provide the computing resource. The resource manager 2106 may facilitate a connection between a resource provider 2102 and a particular computing device 2104. In some implementations, the resource manager 2106 may establish a connection between a particular resource provider 2102 and a particular computing device 2104. In some implementations, the resource manager 2106 may redirect a particular computing device 2104 that has the requested computing resource to a particular resource provider 2102 .
[0432] 22 illustrates examples of a computing device 2200 and a mobile computing device 2250 that can be used to implement the techniques described in this disclosure. The computing device 2200 is intended to represent various types of digital computers, examples of which include laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. The mobile computing device 2250 is intended to represent various types of mobile devices, examples of which include personal digital assistants, mobile phones, smartphones, and other similar computing devices. The components, their connections and relationships, and their functions illustrated herein are intended to be exemplary only and not limiting.
[0433] Computing device 2200 includes processor 2202, memory 2204, storage device 2206, high-speed interface 2208 connecting memory 2204 and multiple high-speed expansion ports 2210, and low-speed interface 2212 connecting low-speed expansion port 2214 and storage device 2206. Each of processor 2202, memory 2204, storage device 2206, high-speed interface 2208, high-speed expansion port 2210, and low-speed interface 2212 may be interconnected using various buses and mounted on a common motherboard or in other manners as desired. Processor 2202 is capable of processing instructions for execution within computing device 2200, including instructions stored in memory 2204 or storage device 2206 for displaying graphical information for a GUI on an external input / output device (such as a display 2216 coupled to high-speed interface 2208). Other implementations may use multiple processors and / or multiple buses as desired, along with multiple memories and multiple memory types. Also, multiple computing devices may be connected, with each device providing a portion of the required operations (e.g., as a bank of servers, a group of blade servers, or a multiprocessor system). Thus, as used herein, when functions are described as being performed by a "processor," this encompasses embodiments in which the functions are performed by any number of processor(s) in any number of computing device(s). Furthermore, when functions are described as being performed by a "processor," this encompasses embodiments in which the functions are performed by any number of processor(s) in any number of computing device(s) (e.g., in a distributed computing system).
[0434] The memory 2204 stores information within the computing device 2200. In some embodiments, the memory 2204 is a volatile memory unit(s). In some embodiments, the memory 2204 is a non-volatile memory unit(s). The memory 2204 may also be another form of computer-readable medium, examples of which include a magnetic disk or an optical disk.
[0435] Storage device 2206 can provide mass storage for computing device 2200. In some embodiments, storage device 2206 can be or contain a computer-readable medium, examples of which include a floppy disk device, a hard disk device, an optical disk or tape device, a flash memory or other similar solid-state memory device, or an array of devices, including devices in a storage area network or other configuration. Instructions can be stored on an information carrier. When executed by one or more processing devices (e.g., processor 2202), the instructions perform one or more methods, such as those described above. Instructions can also be stored by one or more storage devices, examples of which include a computer- or machine-readable medium (e.g., memory 2204, storage device 2206, or memory on processor 2202).
[0436] High-speed interface 2208 manages bandwidth-intensive operations for computing device 2200, while low-speed interface 2212 manages less bandwidth-intensive operations. This allocation of functionality is merely exemplary. In some embodiments, high-speed interface 2208 is coupled to memory 2204, display 2216 (e.g., via a graphics processor or accelerator), and high-speed expansion port 2210, which may accommodate various expansion cards (not shown). In this implementation, low-speed interface 2212 is coupled to storage device 2206 and low-speed expansion port 2214. Low-speed expansion port 2214, which may include various communication ports (e.g., USB, Bluetooth, Ethernet, wireless Ethernet), may be coupled to one or more input / output devices, examples of which include a keyboard, pointing device, scanner, or network device, such as a switch or router via a network adapter.
[0437] Computing device 2200 may be implemented in several different forms, as shown. For example, it may be implemented as a standard server 2220, or multiple times within a group of such servers. It may also be implemented within a personal computer (such as laptop computer 2222). It may also be implemented as part of a rack server system 2224. Alternatively, components of computing device 2200 may be combined with other components in a mobile device (not shown) (such as mobile computing device 2250). Each such device may contain one or more of computing device 2200 and mobile computing device 2250, and the entire system may be made up of multiple computing devices in communication with each other.
[0438] The mobile computing device 2250 includes, among other components, a processor 2252, a memory 2264, an input / output device (such as a display 2254), a communication interface 2266, and a transceiver 2268. The mobile computing device 2250 may also be provided with a storage device (such as a microdrive or other device) to provide additional storage. Each of the processor 2252, memory 2264, display 2254, communication interface 2266, and transceiver 2268 may be interconnected using various buses, and some of the components may be mounted on a common motherboard or in other manners as desired.
[0439] The processor 2252 can execute instructions within the mobile computing device 2250, including instructions stored in the memory 2264. The processor 2252 may be implemented as a chipset of chips including multiple individual analog and digital processors. The processor 2252 may provide, for example, coordination of other components of the mobile computing device 2250, including control of a user interface, applications run by the mobile computing device 2250, and wireless communication by the mobile computing device 2250.
[0440] The processor 2252 may communicate with a user via a display interface 2256 coupled to the control interface 558 and the display 2254. The display 2254 may be, for example, a TFT (thin film transistor liquid crystal display) display or an OLED (organic light emitting diode) display, or other suitable display technology. The display interface 2256 may include appropriate circuitry for driving the display 2254 to present graphical and other information to the user. The control interface 2258 may receive commands from the user and convert them for submission to the processor 2252. Additionally, an external interface 2262 may provide communication with the processor 2252 to enable near-area communication between the mobile computing device 2250 and other devices. The external interface 2262 may, for example, provide wired communication in some embodiments or wireless communication in other embodiments, and multiple interfaces may be used.
[0441] Memory 2264 stores information within mobile computing device 2250. Memory 2264 may be embodied as one or more of a computer-readable medium(s), a volatile memory unit(s), or a non-volatile memory unit(s). Expansion memory 2274 may also be provided and connected to mobile computing device 2250 via expansion interface 2272. This expansion interface 2272 may include, for example, a SIMM (Single In-Line Memory Module) card interface. Expansion memory 2274 may provide extra storage space for mobile computing device 2250 or may store applications or other information for mobile computing device 2250. Specifically, expansion memory 2274 may include instructions for performing or supplementing the above-described processes and may also include secure information. Thus, for example, expansion memory 2274 may be provided as a security module for mobile computing device 2250 and may be programmed with instructions that enable secure use of mobile computing device 2250. Additionally, secure applications may be provided via SIMM cards with additional information, an example of which is placing identifying information on the SIMM card in an unhackable manner.
[0442] As discussed below, the memory may include, for example, flash memory and / or NVRAM memory (non-volatile random access memory). In some implementations, the instructions are stored on an information carrier. When executed by one or more processing devices (e.g., processor 2252), the instructions perform one or more methods, such as those described above. The instructions may also be stored by one or more storage devices, examples of which include one or more computer- or machine-readable media (e.g., memory 2264, expansion memory 2274, or memory on processor 2252). In some embodiments, the instructions may be received in a propagated signal, for example, through transceiver 2268 or external interface 2262.
[0443] Mobile computing device 2250 may communicate wirelessly via communication interface 2266, which may include digital signal processing circuitry as needed. Communication interface 2266 may provide for communication under various modes or protocols, examples of which include GSM (Global System for Mobile communications) voice, SMS (Short Message Service), EMS (Enhanced Messaging Service), or MMS (Multimedia Messaging Service), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service), among others. Such communication may be performed, for example, by transceiver 2268 using radio frequencies. Additionally, short-range communication may be performed using, for example, Bluetooth, Wi-Fi, or other such transceivers (not shown). Additionally, a Global Positioning System (GPS) receiver module 2270 may provide additional navigation and location related wireless data to the mobile computing device 2250, which may be used as needed by applications running on the mobile computing device 2250.
[0444] The mobile computing device 2250 may also perform voice communications using an audio codec 2260, which may receive voice information from a user and convert it into usable digital information. The audio codec 2260 may similarly generate audible sounds for the user, such as through a speaker in the handset of the mobile computing device 2250. Such sounds may include sounds from a voice call, recorded sounds (e.g., voice messages, music files, etc.), and sounds generated by applications running on the mobile computing device 2250.
[0445] The mobile computing device 2250 may be implemented in several different forms, as shown, such as a mobile phone 2280, or as part of a smartphone 2282, personal digital assistant, or other similar mobile device.
[0446] Various implementations of the systems and techniques described herein can be realized in digital electronic circuitry, integrated circuits, specially designed ASICs (application-specific integrated circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs executable and / or interpretable on a programmable system including at least one programmable processor, which may be special purpose or general purpose, coupled to receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0447] These computer programs (also known as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in a high-level procedural and / or object-oriented programming language, and / or assembly / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device (e.g., magnetic disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0448] To provide for user interaction, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, as well as a keyboard and pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other types of devices can be used to provide for user interaction as well. For example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, haptic feedback, etc.), and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0449] The systems and techniques described herein can be implemented in a computing system that includes a back-end component (e.g., as a data server), or a computing system that includes a middleware component (e.g., an application server), or a computing system that includes a front-end component (e.g., a client computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or any combination of such back-end, middleware, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0450] A computing system may include clients and servers. Clients and servers are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.
[0451] In some embodiments, the various modules described herein may be separated, combined, or incorporated into a single module or combined modules. The modules shown in the figures are not intended to limit the systems described herein to the software architectures shown therein.
[0452] equivalent Elements of various implementations described herein may be combined to form other implementations not specifically described above. Elements may be omitted from the processes, computer programs, databases, etc. described herein without adversely affecting operation. Furthermore, the logic flows depicted in the figures do not require the particular order shown, or sequential order, to achieve desirable results. Various individual elements may be combined into one or more individual elements to perform the functions described herein.
[0453] Throughout the description, where devices and systems are described as having, including, or comprising specific components, or processes and methods are described as having, including, or comprising specific steps, it is additionally contemplated that there are devices and systems of the invention that consist essentially of, or consist of, the recited components, and that there are processes and methods of the invention that consist essentially of, or consist of, the recited process steps.
[0454] It should be understood that the order of steps or order for performing certain actions is immaterial so long as the invention remains operable. Moreover, two or more steps or actions may be conducted simultaneously.
[0455] While the present invention has been particularly shown and described with reference to specific preferred embodiments, it should be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope of the invention as defined by the appended claims.
Claims
1. 1. A method for the in silico design of engineered antigens, comprising: (a) receiving and / or accessing, by a processor of a computing device, a polypeptide model representing a reference antigen of an infectious agent; (b) identifying, by the processor, within the polypeptide model, one or more memory-inducing conserved region(s) representing conserved portions of the reference antigen determined to be likely to elicit a memory immune response; (c) generating, by said processor, one or more amino acid modifications within at least a portion of said one or more conserved region(s), thereby creating a disrupted polypeptide model representative of said engineered antigen; (d) by the processor, displaying and / or storing and / or providing the collapsed polypeptide model for further processing.
2. 10. The method of claim 1, wherein the reference antigen is a naturally occurring variant of a viral protein or comprises at least a portion thereof.
3. 3. The method of claim 2, wherein the reference antigen is or comprises at least a portion of a SARS-Cov2 spike polypeptide.
4. 4. The method of claim 3, wherein the reference antigen is or comprises at least a portion of a specific SARS-CoV-2 variant spike polypeptide.
5. 5. The method of claim 4, wherein the specific SARS-CoV-2 variant is a member of the Omicron and / or XBB phylogenetic tree.
6. 3. The method of claim 1 or 2, wherein the infectious agent is or comprises an RNA virus and the reference antigen is or comprises at least a portion of a protein thereof.
7. The method of claim 1 , wherein the reference antigen is or comprises a bacterial protein.
8. 10. The method of claim 1, wherein the reference antigen is or comprises an antigen of a parasite.
9. 10. The method of any one of the preceding claims, wherein the one or more memory-inducing conserved region(s) represent portions of the reference antigen that are substantially similar to (i) one or more variants thereof and / or (ii) an initial / wild-type strain.
10. 10. The method of any one of the preceding claims, wherein the reference antigen is a specific target SARS-CoV-2 variant S polypeptide or portion thereof, and the one or more memory-inducing conserved region(s) represent portion(s) of the reference antigen that are substantially similar to corresponding portion(s) of (i) one or more other SARS-CoV-2 variant polypeptides and / or (ii) a Wuhan SARS-CoV-2 polypeptide.
11. 10. The method of any one of the preceding claims, wherein the one or more memory-inducing conserved region(s) is or comprises a set of conserved epitope regions that represent known epitopes that are present on the reference antigen and that have not been mutated.
12. Step (b) is obtaining, by the processor, data corresponding to a set of known epitopes and identifying, within the reference antigen, each of one or more specific known epitopes of the set; obtaining, by the processor, an identification of a set of characteristic mutations of the reference antigen; and identifying, by the processor, specific known epitopes corresponding to portions of the reference antigen that do not have any characteristic mutations as the set of conserved epitope regions.
13. 13. The method of claim 12, wherein the set of known epitopes comprises one or more of the epitopes listed in Table 2A.
14. 13. The method of claim 12, wherein the set of known epitopes comprises one or more of the epitopes listed in Table 2B.
15. 10. The method of any one of the preceding claims, wherein the one or more memory-inducing conserved region(s) is or comprises a conserved surface representing a non-mutated surface of the reference antigen.
16. 10. The method of any one of the preceding claims, comprising identifying and / or accessing the identity of a set of one or more characteristic mutations of said reference antigen.
17. 17. The method of claim 16, wherein the reference antigen is or comprises the SARS-Cov2 XBB.1.5 spike protein.
18. 10. The method of any one of the preceding claims, wherein step (c) comprises selecting, by the processor, the one or more amino acid modifications from a set of allowed mutations.
19. 19. The method of claim 18, wherein the set of allowed mutations is or comprises a plurality of mutations observed as occurring within a set of related antigens.
20. 20. The method of claim 19, wherein the reference antigen is a SARS-CoV-2 protein of a particular variant and the set of related antigens includes corresponding proteins of other related variants.
21. 21. The method of claim 20, wherein the reference antigen is a member of the Omicron lineage and the set of related antigens comprises observed variants belonging to the Omicron lineage.
22. 22. The method of any one of claims 19 to 21, wherein the set of allowed mutations comprises at least a portion of the mutations listed in Table 3.
23. 23. The method of claim 22, wherein the set of allowed mutations includes at least a portion of the mutations listed in Table 3, excluding one or both of L335F and L390R.
24. 24. The method of any one of claims 19 to 23, wherein the reference antigen is a SARS-CoV2 protein and the set of related antigens comprises corresponding proteins of other coronaviruses.
25. 10. The method of any one of the preceding claims, wherein the one or more memory-inducing conserved regions are or comprise a set of conserved epitope regions, and step (c) comprises introducing at least one amino acid modification into at least a portion of each of the conserved epitope regions.
26. 10. The method of any one of the preceding claims, wherein the one or more memory-inducing conserved regions are or comprise a conserved surface, and step (c) comprises generating the one or more amino acid modifications at positions distributed throughout / across the conserved surface.
27. 10. The method of any one of the preceding claims, comprising iteratively performing steps (b) and (c) to generate a plurality of candidate polypeptide models, each of the plurality of candidate polypeptide models representing a candidate engineered variant.
28. 28. The method of claim 27, comprising: determining, by the processor, one or more performance score values for each of the candidate polypeptide models; and selecting a subset of the candidate polypeptide models based at least in part on the determined performance score values.
29. The one or more performance scores may be: (a) an immune escape score indicating the likelihood and / or relative ability of a particular candidate engineered variant to be recognized and neutralized by an antibody; and 29. The method of claim 28, comprising one or both of (b) a fitness score indicating the likelihood and / or viability of a particular candidate engineered variant.
30. 30. The method of claim 29, wherein determining one or both of (a) the immune escape score and (b) the fitness score comprises using a machine learning model.
31. 31. The method of claim 29 or 30, wherein determining one or both of (a) the immune escape score and (b) the fitness score comprises using a 3D structural model of at least a portion of the identified candidate variant.
32. 32. The method of any one of claims 27-31, wherein said one or more performance scores comprise a positional diffusion score that measures the degree to which amino acid modifications are evenly distributed across the surface of said candidate engineered variant.
33. The method of any one of claims 27 to 32, wherein the one or more performance scores comprise a variant co-occurrence score.
34. 10. The method of any one of the preceding claims, comprising identifying, by the processor, one or more target regions within the polypeptide model that represent portions of the reference antigen to be retained; and excluding the one or more target regions from the one or more memory-inducing conserved region(s).
35. 10. A method according to any one of the preceding claims, comprising causing the processor to render the collapsed polypeptide model for graphical display.
36. 10. A method according to any one of the preceding claims, comprising generating a corresponding RNA sequence from the disrupted polypeptide model.
37. 10. A method according to any one of the preceding claims, comprising producing a composition comprising a polypeptide based on said disrupted polypeptide model.
38. 10. The method of any one of the preceding claims, comprising assessing the biological activity of the engineered antigen in vitro.
39. The biological activity of the engineered antigen is: the engineered antigen is properly expressed and folded; and / or the engineered antigen does not bind to an antibody that binds to the reference antigen; and / or The engineered antigen-packed pseudovirus is capable of entering cells; and / or the engineered antigen is immunogenic, and / or 39. The method of claim 38, wherein the engineered antigen reduces activation of a B cell memory immune response to the reference antigen.
40. 10. A method according to any one of the preceding claims, comprising producing a composition comprising a nucleic acid encoding an amino acid sequence represented by said disrupted polypeptide model.
41. A vaccine composition comprising a polypeptide and / or nucleic acid according to claim 36 or 37.
42. 42. A method of vaccination comprising administering the vaccine of claim 41 to a subject or population of subjects.
43. 37. A system comprising a processor of a computing device and a memory having stored thereon instructions that, when executed by the processor, cause the processor to perform a method according to any one of claims 1 to 36.
44. 1. A method for producing an immunogenic composition, comprising: comparing sequences of viral proteins from different variants of an infectious disease agent to identify conserved sites remaining in the antigen of interest; replacing at least one or more of the remaining conserved sites with a sequence that generates a new sequence; and producing a vaccine that delivers at least a portion of said new sequence that includes at least one of said remaining conserved sites substituted.
45. 1. An RNA comprising a nucleotide sequence encoding an engineered antigen, wherein the engineered antigen corresponds to a particular reference antigen, wherein the reference antigen has been altered to introduce one or more amino acid modifications within at least a portion of one or more memory-inducing conserved regions identified as portions of the reference antigen determined to have a high likelihood of inducing a memory immune response.
46. 46. The RNA of claim 45, wherein said one or more memory-inducing conserved region(s) represent portions of said reference antigen(s), said portions being substantially similar to one or more variants thereof.
47. 47. The RNA of claim 45 or 46, wherein the one or more memory-inducing conserved region(s) is or comprises a set of conserved epitope regions that represent known epitopes that are present on the reference antigen and that have not been mutated.
48. 48. The RNA of claim 47, wherein the set of known epitopes comprises one or more of the epitopes listed in Table 2A and / or Table 2B.
49. 49. The RNA of any one of claims 45 to 48, wherein the one or more memory-inducing conserved region(s) is or comprises a conserved surface representing a non-mutated surface of the reference antigen.
50. 50. The RNA of claim 49, wherein the reference antigen is or comprises at least a portion of the XBB.1.5 variant of the SARS-Cov2 spike protein.
51. 51. The RNA of any one of claims 45 to 50, wherein the one or more amino acid modifications are selected from a set of allowed mutations.
52. 52. The RNA of claim 51, wherein the set of allowed mutations is or includes a plurality of mutations observed as occurring within a set of related polypeptides.
53. 53. The RNA of claim 52, wherein the set of related polypeptides includes corresponding polypeptides of other related SARS-CoV 2 variants.
54. 54. The RNA of claim 53, wherein the target polypeptide is a member of the Omicron lineage and the set of related polypeptides includes polypeptides corresponding to observed variants belonging to the Omicron lineage.
55. 55. The RNA of any one of claims 53-54, wherein the one or more memory-inducing conserved regions are or comprise a set of conserved epitope regions, and the engineered antigen has at least one amino acid modification within at least a portion of each of the conserved epitope regions.
56. 56. The RNA of any one of claims 45 to 55, wherein the one or more memory-inducing conserved regions are or comprise a conserved surface, and wherein the one or more amino acid modifications occur at positions distributed throughout / across the conserved surface.
57. 10. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen comprises one or more of the combinations of mutations listed in Table 5A.
58. 10. The method, system, vaccine composition, method of manufacture or RNA of any one of the preceding claims, wherein said engineered antigen comprises one or more of the sequences listed in Table 5B.
59. 10. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen comprises one or more of the combinations of mutations listed in Table 7A and / or Table 7B.
60. 8. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen comprises one or more of the combinations of mutations listed in Table 8A.
61. 10. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen comprises one or more of the sequences listed in Table 8B.
62. 9. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen comprises one or more of the combinations of mutations listed in Table 9A.
63. 10. The method, system, vaccine composition, method of manufacture or RNA of any one of the preceding claims, wherein said engineered antigen comprises the combination of mutations identified as S43 in Table 9A.
64. 10. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein having at least a portion of the following mutations: N360D, P384S, L390R, T430I, F464Y and H519N.
65. 10. The method, system, vaccine composition, method of manufacture or RNA of any one of the preceding claims, wherein said engineered antigen comprises the combination of mutations identified as S48 in Table 9A.
66. 10. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein having at least a portion of the following mutations: L335F, K356T, P384S, L390R, T430I, F464Y and H519N.
67. 10. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen comprises at least a portion of SARS-CoV-2 S protein having the following mutations: P384S, L390R, T430I, F464Y and H519N.
68. 10. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen comprises at least a portion of a SARS-CoV-2 S protein having one or more of the following mutations: K356T, P384S, L390R, T430I, F464Y and H519N.
69. 10. The method, system, vaccine composition, method of manufacture or RNA of any one of the preceding claims, wherein said engineered antigen does not comprise one or both of the mutations L335F and L390R.
70. 13. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen comprises one or more of the combinations of mutations listed in Table 15A and / or Table 15B.
71. 10. The method, system, vaccine composition, method of manufacture or RNA of any one of the preceding claims, wherein the engineered antigen comprises a combination of mutations identified as S123 in Table 15A and / or Table 15B.
72. 10. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein having at least a portion of the following mutations: I332V, L335F, K356T, P384S, T430I, L452Q, F464Y, H519N.
73. 73. The method, system, vaccine composition, method of manufacture or RNA of claim 72, wherein the engineered antigen comprises at least a portion of the following mutations: I332V, L335F, G339H, R346T, K356T, L368I, S371F, S373P, S375F, T376A, P348S, D405N, R408S, K417N, T430I, N440K, V445P, G446S, L452Q, N460K, F464Y, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, E516Q, H519N.
74. 10. The method, system, vaccine composition, RNA or method of manufacture of any one of the preceding claims, wherein said engineered antigen comprises a combination of mutations identified as S122 in Table 15A and / or Table 15B.
75. 10. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein having at least a portion of the following mutations: L335F, K356T, P384S, T430I, L452R, F464Y, H519N.
76. 76. The method, system, vaccine composition, method of manufacture or RNA of claim 75, wherein the engineered antigen comprises at least a portion of the following mutations: L335F, G339H, R346T, K356T, L368I, S371F, S373P, S375F, T376A, P348S, D405N, R408S, K417N, T430I, N440K, V445P, G446S, L452R, N460K, F464Y, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, E516Q, H519N, T523S.
77. 10. The method, system, vaccine composition, RNA or method of manufacture of any one of the preceding claims, wherein said engineered antigen comprises a combination of mutations identified as S109 in Table 15A and / or Table 15B.
78. 10. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein having at least a portion of the following mutations: K356T, N360S, P384S, N388K, T430I, N450D, F464Y, H519N.
79. 79. The method, system, vaccine composition, method of manufacture or RNA of claim 78, wherein the engineered antigen comprises at least a portion of the following mutations: G339H, R346T, K356T, L368I, S371F, S373P, S375F, T376A, P348S, N388K, D405N, R408S, K417N, T430I, N440K, V445P, G446S, N450D, N460K, F464Y, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, H519N, T523S.
80. 10. The method, system, vaccine composition, RNA or method of manufacture of any one of the preceding claims, wherein said engineered antigen comprises a combination of mutations identified as S129 in Table 15A and / or Table 15B.
81. 10. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein having at least a portion of the following mutations: K356T, L335F, P384S, D389G, T430I, N450D, F464Y, H519N.
82. 82. The method, system, vaccine composition, method of manufacture or RNA of claim 81 , wherein the engineered antigen comprises at least a portion of the following mutations: L335F, G339H, R346T, K356T, L368I, S371F, S373P, S375F, T376A, P348S, D389G, D405N, R408S, K417N, T430I, N440K, V445P, G446S, N450D, L452R, N460K, F464Y, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, H519N.
83. 10. The method, system, vaccine composition, RNA or method of manufacture of any one of the preceding claims, wherein said engineered antigen comprises a combination of mutations identified as S156 in Table 15A and / or Table 15B.
84. 10. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein having at least a portion of the following mutations: K356T, P384S, L390R, T430I, N450D, F464Y, I472V, H519N.
85. 85. The method, system, vaccine composition, method of manufacture or RNA of claim 84, wherein the engineered antigen comprises at least a portion of the following mutations: G339H, R346T, K356T, L368I, S371F, S373P, S375F, T376A, P348S, L390R, D405N, R408S, K417N, T430I, N440K, V445P, G446S, N450D, L452R, N460K, F464Y, I472V, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, H519N.
86. 10. The method, system, vaccine composition, RNA or method of manufacture of any one of the preceding claims, wherein said engineered antigen comprises a combination of mutations identified as S112 in Table 15A and / or Table 15B.
87. 10. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein having at least a portion of the following mutations: K356T, P384S, D389G, T430I, N450D, F464Y, I468V, H519N.
88. 88. The method, system, vaccine composition, method of manufacture or RNA of claim 87, wherein the engineered antigen comprises at least a portion of the following mutations: G339H, R346T, K356T, L368I, S371F, S373P, S375F, T376A, P348S, D389G, D405N, R408S, K417N, T430I, N440K, V445P, G446S, N450D, N460K, F464Y, I468V, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, H519N.
89. 10. The method, system, vaccine composition, RNA or method of manufacture of any one of the preceding claims, wherein said engineered antigen comprises a combination of mutations identified as S125 in Table 15A and / or Table 15B.
90. 10. The method, system, vaccine composition, manufacturing method or RNA of any one of the preceding claims, wherein said engineered antigen is or comprises at least a portion of a SARS-CoV-2 S protein having at least a portion of the following mutations: K356T, L335F, P384S, T430I, F464Y, I468V, H519N.
91. 91. The method, system, vaccine composition, method of manufacture or RNA of claim 90, wherein the engineered antigen comprises at least a portion of the following mutations: L335F, G339H, R346T, K356T, L368I, S371F, S373P, S375F, T376A, P348S, D405N, R408S, K417N, T430I, N440K, V445P, G446S, N460K, F464Y, I468V, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, E516Q, H519N, T523S.
92. 1. A method for producing RNA comprising a nucleotide sequence encoding an engineered antigen that corresponds to an engineered version of a reference antigen, the method comprising producing RNA, wherein the nucleotide sequence of the RNA, when compared to the reference antigen, exhibits difference(s) relative to the reference antigen in one or more memory-inducing conserved regions common to the reference antigen and (i) one or more pre-existing variants of the reference antigen, and / or (ii) a wild-type strain of the reference antigen.
93. 93. The method of claim 92, wherein the reference antigen is or comprises the RBD of a specific target variant SARS-CoV-2 S protein.
94. The method of claim 92 or 93, wherein the RNA is or comprises the RNA of any one of claims 45 to 91.
95. 95. The method of any one of claims 92 to 94, comprising producing the RNA by in vitro transcription (IVT).