Systems and methods for engineering synthetic antigens to facilitate tailored immune responses
By using computer simulation design and amino acid modification, engineered antigens are generated, which solves the problem that existing vaccination technologies are unable to cope with variant pathogens, and improves the immunogenicity of vaccines and the efficiency of customized antibody production.
Patent Information
- Application Number
- CN202480024042.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-18
- Filing Date
- 2024-02-23
- Publication Date
- 2026-01-13
AI Technical Summary
Existing vaccination technologies are ill-suited to effectively address evolving pathogen variants and emerging diseases, and are prone to triggering memory immune responses, resulting in a lack of customization in the production of new antibodies.
By using computer simulation design, conserved regions of reference antigens are identified and modified with amino acids to generate engineered antigens, thereby reducing the activation of memory immune responses, promoting primary responses, and generating new customized antibodies.
It improved the immunogenicity of vaccination, reduced the triggering of memory immune responses, and enhanced the immune response against volatile viruses.
Smart Images

Figure CN121336262A_ABST
Abstract
Description
[0001] Cross-reference to related applications
[0002] This application claims priority and benefit to U.S. Provisional Applications Nos. 63 / 448,217 (filed February 24, 2023), 63 / 448,215 (filed February 24, 2023), 63 / 448,987 (filed February 28, 2023), 63 / 449,031 (filed February 28, 2023), 63 / 449,936 (filed March 3, 2023), 63 / 452,989 (filed March 17, 2023), 63 / 452,987 (filed March 17, 2023), 63 / 514,242 (filed July 18, 2023), and 63 / 514221 (filed July 18, 2023), the contents of which are incorporated herein by reference in their entirety. Background Technology
[0003] Vaccination plays a crucial role in managing and protecting public health. When administered with a sufficiently effective immunogenic composition against a specific infectious pathogen, individuals experience a reduced risk of infection and / or a reduction in disease severity. Therefore, developing highly efficient vaccination technologies that can keep pace with evolving circulating pathogen variants and emerging diseases is a critical challenge. Summary of the Invention
[0004] This document presents techniques for computer-simulated design of customized, engineered antigens. Specifically, in some embodiments, the methods and systems of this disclosure provide for engineering antigens to reduce their activation of memory immune responses (such as B-cell and / or T-cell-based responses) upon introduction into a subject. For example, designing antigens in this manner can improve their performance as immunogenic compositions for vaccination. For example, without being bound by any particular theory, it is considered that, among other things, reducing the degree to which a memory immune response is triggered can lead to an increase in the production of novel antibodies that are selectively tailored by the subject's immune system to neutralize specific (e.g., present) epitopes of a reference antigen.
[0005] For example, in some embodiments, the systems and methods of this disclosure identify conserved regions in the computer representation of a reference antigen that are similar to regions of other variants of the reference antigen (e.g., previously circulating variants), and thus may trigger a memory response. The methods described herein then disrupt these conserved regions, for example, by introducing amino acid modifications therein / across them. In this way, an engineered antigen can be generated that retains certain portions of the reference antigen, such as specific target epitopes, but replaces the conserved regions with its disrupted version. Without wishing to be bound by any particular theory, it is thought that, when manufactured and introduced into a subject, an engineered antigen with disrupted conserved regions is less likely to trigger a memory immune response (e.g., from memory B or T cells), but rather promotes a primary response and thus produces novel neutralizing antibodies specifically tailored to the retained target epitopes. Therefore, such engineered antigens may provide improved efficacy when used as immunogenic compositions, particularly for easily mutated viral infectious pathogens.
[0006] In one aspect, this disclosure provides a method for computer-simulated design of engineered antigens [e.g., for (e.g., characterized by) inducing an immune response against one or more target epitopes of a reference antigen of an infectious pathogen, while reducing the activation of a memory immune response of the engineered antigen against the reference antigen (e.g., relative to) B cell and / or T cell], the method comprising: (a) receiving and / or accessing a peptide model representing a reference antigen of an infectious pathogen (e.g., as a sequence and / or 3D structural model) via a processor of a computing device; and (b) identifying, via the processor, one or more conserved regions within the peptide model that trigger memory, the conserved regions representing conserved portions of the reference antigen, the conserved portions being determined to be likely to trigger a memory immune response [e.g., determined to be] (i) a portion of a reference antigen corresponding to a known epitope and / or a potential epitope and / or (ii) a portion of a reference antigen that does not contain a signature mutation of the reference antigen; (c) generating one or more amino acid modifications within at least a portion of one or more conserved regions (triggering memory) by a processor [e.g., one or more amino acid modifications are one or more point modifications, insertions, and / or deletions], thereby creating a peptide model representing the disruption of an engineered antigen [e.g., representing an artificially engineered version of a reference antigen in which mutations are introduced to disrupt its conserved portions (e.g., at least a portion of one or more conserved regions that trigger memory, e.g., identified as potentially triggering a memory immune response)]; and (d) storing and / or providing the disrupted peptide model by a processor for display and / or further processing.
[0007] In some implementations, the reference antigen is or contains at least a portion of a naturally occurring variant of a viral protein [e.g., where the infectious pathogen is a variant of a particular virus (e.g., influenza virus, coronavirus, respiratory syncytial virus, filovirus) and where the reference antigen is or contains at least a portion of its protein].
[0008] In some implementations, the reference antigen (e.g. and / or viral protein) is or contains at least a portion of the SARS-Cov2 spike polypeptide [e.g., receptor-binding domain (RBD); e.g., N-terminal region; e.g., substantially all (e.g., the entire) spike protein] [e.g., selected to focus on the portion of the least relevant vaccine antigen, e.g. to facilitate the removal of as many conserved epitopes as possible without, for example, introducing point mutations (e.g., thus limiting the number of epitopes to be introduced with point mutations)].
[0009] In some implementations, the reference antigen is or contains at least a portion of a specific SARS-CoV-2 variant spike polypeptide.
[0010] In some implementations, a particular SARS-CoV-2 variant is a member of the Omicron and / or XBB lineage classification (e.g., according to WHO, Pango, Nextstrain, etc.) (e.g., where the particular SARS-CoV-2 variant is XBB.1.5; e.g., where the particular SARS-CoV-2 variant is JN.1).
[0011] In some embodiments, the infectious pathogen is or contains an RNA virus (e.g., a specific variant thereof) (e.g., a virus whose genetic information is encoded by RNA), and the reference antigen is or contains at least a portion of its protein.
[0012] In some implementations, the reference antigen is or contains a bacterial protein (e.g., where the infectious pathogen is a bacterium and the reference antigen is or contains at least a portion of its specific protein).
[0013] In some embodiments, the reference antigen is or contains an antigen of the parasite (e.g., a surface antigen; e.g., a protein) (e.g., where the infectious pathogen is a parasite and the reference antigen is or contains at least a portion of its specific protein) (e.g., where the infectious pathogen is the malaria parasite and the reference antigen is its antigen).
[0014] In some embodiments, one or more conserved regions that trigger memory represent portions of the reference antigen that are substantially similar to (i) one or more (e.g., pre-existing) variants and / or (ii) initial / wild-type strains (e.g., first-observed strains) {e.g., wherein the one or more conserved regions that trigger memory represent portions of the reference antigen that have sufficient sequence similarity [e.g., at least 80% (including, for example, at least 85%, 90%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or higher]; e.g., identical]}.
[0015] In some implementations, the reference antigen is a specific target SARS-CoV-2 variant (e.g., XBB.1.5; e.g., JN.1) S polypeptide or a portion thereof (e.g., where the reference antigen is the RBD of the target SARS-CoV-2 S protein), and where one or more conserved regions that trigger memory represent portions of the reference antigen that are substantially similar to (i) one or more (e.g., pre-existing) other SARS-CoV-2 variant polypeptides and / or (ii) wild-type SARS-CoV-2 polypeptides [e.g., where one or more conserved regions that trigger memory represent non-mutated portions of the reference antigen that are common to the reference antigen and (i) one or more (e.g., pre-existing) SARS-CoV-2 variant polypeptides and / or (ii) wild-type SARS-CoV-2 polypeptides].
[0016] In some embodiments, one or more conserved regions that trigger memory are or comprise a set of conserved epitope regions that represent known epitopes present on a reference antigen and not mutated (e.g., known epitopes without any signature mutations) [e.g., where the reference antigen is a specific subregion of the target SARS-CoV-2 variant S protein (e.g., RBD) (e.g., XBB.1.5; e.g., JN.1), and wherein the set of conserved epitope regions represents a known epitope present on a specific subregion of the target SARS-CoV-2 variant S protein and not mutated, relative to a corresponding subregion of the SARS-CoV-2 S protein of one or more pre-existing variants and / or wild-type strains].
[0017] In some implementations, step (b) includes: acquiring data corresponding to a set of known epitopes via a processor, and identifying each of one or more specific known epitopes in the set within a reference antigen; acquiring an identifier of a set of signature mutations of the reference antigen via a processor; and identifying those specific known epitopes corresponding to portions of the reference antigen that do not have any signature mutations as conserved epitope groups via a processor.
[0018] In some implementations, a set of known tabletops includes one or more of the tabletops listed in Table 2A.
[0019] In some implementations, a set of known tabletops includes one or more of the tabletops listed in Table 2B.
[0020] In some implementations, one or more conserved regions that trigger memory are or contain conserved surfaces of non-mutated (e.g., lacking any signature mutations) surfaces representing a reference antigen (e.g., continuous; e.g., adjacent) [e.g., where the reference antigen is or contains SARS-CoV-2]. The XBB.1.5 variant of the RBD of the S protein, and the conserved surface is or contains at least a portion (e.g., at most all) of the following sites: L335, E340, A348, S349, Y351, A352, N354, R355, K356, R357, S359, N360, V362, D364, S366, Y369, N370, A372, F377, K378, Y380, G381, S383, P384, T385, K386, N388, D389, L390, C391, F392, T393, N394, Y3 96. P412, G413, Q414, T415, K424, P426, D427, D428, T430, K444, N450, L452, R457, K458, S459, K462, P463, F464, E465, R466, D467 , I468, S469, T470, E471, I472, Y473, Q474, P479, N481, G482, V483, E516, L517, L518, H519, A520, P521, T523, C525, G526, P527].
[0021] In some embodiments, the provided method includes identifying and / or obtaining the identifier (e.g., location) of one or more signature mutations of a reference antigen [e.g., where the reference antigen is a protein of a specific viral variant classification, and the signature mutations of the group include those mutations of the reference antigen that appear at or above a specific threshold rate (e.g., appearing at a rate equal to or above a threshold rate (e.g., fraction, percentage, etc.) in sequences identified as belonging to a viral variant classification)].
[0022] In some implementations, the reference antigen is or contains the SARS-CoV-2 XBB.1.5 spike protein {e.g., the SARS-CoV-2S protein with the XBB.1.5 signature mutation} [e.g., containing at least some (e.g., all) of the following mutations: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S3...]. 73P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F4 86P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K]}.
[0023] In some implementations, step (c) includes the processor selecting one or more amino acid modifications from a set of permissible mutations [e.g., selecting positional and / or specific modifications (e.g., substitution, deletion, insertion) from the set of permissible mutations].
[0024] In some implementations, a set of permitted mutations is or includes (e.g., a list, table, etc.) multiple mutations observed to occur (e.g., at a specific rate or higher) in a set of related antigens.
[0025] In some implementations, the reference antigen is a specific variant of the SARS-CoV-2 protein (e.g., spike protein), and the group-related antigen includes corresponding proteins of other related variants (e.g., within a specific group lineage and / or sublineage that includes the reference antigen).
[0026] In some implementations, the reference antigen is a member of the Omicron lineage, and the group-associated antigen includes the observed variants belonging to the Omicron lineage.
[0027] In some implementations, a set of permitted mutations includes at least a portion (e.g., a subgroup; e.g., all) of the mutations listed in Table 3 (e.g., excluding those that have been deleted, as marked with strikethrough in Table 3) (e.g., and optionally, one or more additional mutations; e.g., not any additional mutations).
[0028] In some implementations, a set of permitted mutations includes at least a portion (e.g., a subgroup) of the mutations listed in Table 3, excluding one or both of L335F and L390R.
[0029] In some implementations, the reference antigen is the SARS-CoV2 (e.g., spike) protein, and the group-associated antigen includes corresponding proteins of other coronaviruses (e.g., SARS-CoV 1, MERS).
[0030] In some embodiments, one or more conserved regions that trigger memory are or contain a set of conserved epitope regions, and step (c) includes introducing at least one amino acid modification into each of at least a portion (e.g., all; e.g., a specific subgroup) of the conserved epitope regions.
[0031] In some implementations, one or more conserved regions that trigger memory are or contain a conserved surface, and step (c) includes generating one or more amino acid modifications at locations distributed throughout / across the conserved surface (e.g., approximately uniform rather than aggregated) [e.g., distributed substantially equidistantly on a 3D representation of the conserved surface (e.g., where the 3D linear distance and / or geodesic distance (across the 3D conserved surface) between each of the plurality of amino acid modifications is substantially equal and / or distributed according to a particular predefined statistical distribution (e.g., a normal distribution)].
[0032] In some implementations, the provided method includes repeating steps (b) and (c) to generate multiple candidate peptide models, each candidate peptide model representing a candidate engineered variant (e.g., each candidate peptide model contains a different combination of amino acid modifications and represents a unique artificially engineered version of a reference antigen).
[0033] In some implementations, the provided method includes determining one or more performance scores for each candidate peptide model via a processor, and selecting a subgroup of candidate peptide models based at least in part on the determined performance score values.
[0034] In some implementations, one or more performance scores include one or both of the following: (a) an immune escape score, which indicates the likelihood and / or relative ability of a particular candidate engineered variant to be recognized and neutralized by an antibody, and (b) a fitness score, which indicates the likelihood and / or activity of a particular candidate engineered variant.
[0035] In some implementations, determining one or both of (a) the immune escape score and (b) the fitness score includes using a machine learning model [e.g., a language model that receives an amino acid sequence representation of a particular candidate engineered variant as input (e.g., where the input does not include a reference 3D structural representation)] [e.g., calculating a probability score indicating the predicted likelihood of a particular candidate variant occurring; e.g., calculating a semantic change score indicating the distance between (i) the embedding representation of a particular candidate variant (e.g., generated by one or more hidden layers of a machine learning model) and (ii) the embedding representation of one or more reference variants (e.g., WT variants; e.g., variants of a reference antigen that the subject has previously been infected with, e.g., through vaccination and / or natural exposure; e.g., reference antigens)].
[0036] In some implementations, determining one or both of (a) the immune escape score and (b) the fitness score includes using a 3D structural model of at least a portion of a particular candidate variant [e.g., calculating a viral peptide receptor binding score (e.g., ACE2 binding score); e.g., calculating an epitope alteration score].
[0037] In some implementations, one or more performance scores include a positional expansion score, which measures the degree to which amino acid modifications are uniformly distributed on the surface (e.g., a conserved surface) of the candidate engineered variant.
[0038] In some implementations, one or more performance scores include a mutation co-occurrence score (e.g., measuring the degree to which one or more amino acid modifications of a particular candidate variant are compared to the natural co-occurrence rate).
[0039] In some implementations, the provided method includes identifying one or more target regions within a peptide model by a processor that represent a reference antigen portion (e.g., a target epitope) to be retained, and excluding one or more target regions from one or more conserved regions that trigger memory (e.g., where the reference antigen is a SARS-CoV-2S protein and / or a portion thereof, and one or more target regions are or contain an ACE2 binding interface, thereby retaining the ACE2 binding interface region).
[0040] In some implementations, the provided method includes the presentation of a peptide model that has been disrupted by a processor for graphical display.
[0041] In some implementations, the provided method includes generating a corresponding RNA sequence from a disrupted peptide model.
[0042] In some embodiments, the provided method includes producing (e.g., as a vaccine) a composition comprising a polypeptide based on a disrupted polypeptide model (e.g., having substantially the same amino acid sequence as that it represents).
[0043] In some implementations, the provided methods include in vitro assessment of the bioactivity of engineered antigens.
[0044] In some embodiments, the bioactivity of the engineered antigen is characterized by: proper expression and folding of the engineered antigen [e.g., based on binding assays, such as flow cytometry-based binding assays (e.g., with hACE2)]; and / or the engineered antigen does not bind to antibodies that bind to a reference antigen [e.g., showing reduced / eliminated binding (compared to the reference antigen) with a group of one or more neutralizing antibodies (e.g., antibodies that bind to different epitope classes; e.g., one or more of classes A, B, C, D, E, and F (e.g., each))]; and / or pseudoviruses loaded with the engineered antigen are able to enter cells; and / or the engineered antigen is immunogenic; and / or the engineered antigen reduces the activation of B-cell memory immune responses to the reference antigen.
[0045] In some embodiments, the provided method includes producing (e.g., as a vaccine) a composition comprising a nucleic acid encoding an amino acid sequence represented by a disrupted polypeptide model.
[0046] In some respects, this disclosure provides vaccine compositions comprising one or more aspects or embodiments described herein (e.g., in the preceding paragraphs).
[0047] In some aspects, this disclosure provides a method of vaccination, which includes administering to a subject or group of subjects a vaccine provided according to one or more aspects or implementation schemes described herein (e.g., in the preceding paragraphs).
[0048] In some aspects, this disclosure provides a system comprising a processor of a computing device and a memory thereon storing instructions, wherein when executed by the processor, the instructions cause the processor to perform one or more aspects or embodiments of the methods described herein (e.g., in the paragraphs above).
[0049] In some aspects, this disclosure provides a method for manufacturing an immunogenic composition, the method comprising: comparing viral protein sequences from different variants of infectious disease pathogens (e.g., influenza; e.g., SARS-CoV-2) to identify (residual) conserved sites in a target antigen (e.g., by one or more systems and / or methods, including any one of the preceding claims); replacing one or more of the residual conserved sites with a sequence characterized by [characteristics / features of such amino acid substitutions, e.g., providing new immunogenic epitopes from different variants] to generate a new sequence (e.g., by one or more systems and / or methods, including any one of the preceding claims); and producing a vaccine that delivers at least a portion of the new sequence, including at least one replaced residual conserved site.
[0050] In some aspects, this disclosure provides RNA comprising a nucleotide sequence encoding an engineered antigen [e.g., for (e.g., characterized in that) inducing an immune response against one or more target epitopes of a reference antigen against an infectious pathogen, while reducing activation of a memory immune response (e.g., B cell and / or T cell) against the reference antigen) by the engineered antigen], wherein the engineered antigen corresponds to a specific reference antigen (e.g., having a sequence of the specific reference antigen), wherein the reference antigen has been (e.g., artificially) altered to introduce one or more amino acid modifications within at least a portion of one or more conserved regions that trigger memory, the conserved regions being identified as portions of the reference antigen, the portions being determined to be likely to trigger a memory immune response [e.g., determined to be (i) a portion corresponding to a known epitope and / or a potential epitope and / or (ii) a portion not containing a signature mutation of the reference antigen] [e.g., such that the engineered antigen is an artificially engineered version of the reference antigen, wherein mutations are introduced to disrupt its conserved portions (e.g., determined to be likely to trigger a memory immune response)].
[0051] In some implementations, one or more conserved regions that trigger memory represent portions of the reference antigen that are substantially similar (e.g., have sufficient sequence similarity; e.g., common) to one or more (e.g., pre-existing) variants of the reference antigen.
[0052] In some implementations, one or more conserved regions that trigger memory are or contain a set of conserved epitope regions that represent known epitopes that are present on the reference antigen and are not mutated (e.g., known epitopes without any signature mutations).
[0053] In some implementations, a set of known tabletops includes one or more of the tabletops listed in Table 2A and / or Table 2B.
[0054] In some implementations, one or more conserved regions that trigger memory are or contain conserved surfaces of a non-mutated (e.g., lacking any signature mutations) surface representing a reference antigen (e.g., continuous; e.g., adjacent). [e.g., where the target polypeptide is or contains an XBB.1.5 variant of the RBD of the SARS-CoV-2 spike (S) protein, and the conserved surface is or contains the following locations: L335, E340, A348, S349, Y351, A352, N354, R355, K356, R357, S359, N360, V362, D364, S366, Y369, N370, A372, F377, K378, Y380, G381, S383, P384, T385, K...] 386, N388, D389, L390, C391, F392, T393, N394, Y396, P412, G413, Q414, T415 , K424, P426, D427, D428, T430, K444, N450, L452, R457, K458, S459, K462, P46 3. F464, E465, R466, D467, I468, S469, T470, E471, I472, Y473, Q474, P479, N4 81, G482, V483, E516, L517, L518, H519, A520, P521, T523, C525, G526, P527].
[0055] In some embodiments, the reference antigen is or comprises at least a portion (e.g., its RBD) of the XBB.1.5 variant of the SARS-CoV-2 spike protein [e.g., and the marker mutation includes at least a portion (e.g., all) of the following mutations: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L3 68I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E 484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K].
[0056] In some embodiments, one or more amino acid modifications are selected from a set of permissible mutations [e.g., selecting positional and / or specific modifications (e.g., substitution, deletion, insertion) from the set of permissible mutations].
[0057] In some implementations, a set of permitted mutations are or include multiple mutations observed to occur (e.g., at a specific rate or higher) in a set of related peptides.
[0058] In some implementations, a group of related peptides includes corresponding peptides of other related SARS-CoV 2 variants (e.g., within a lineage and / or sublineage that includes a particular group of SARS-CoV 2 variants).
[0059] In some implementations, the target (SARS-CoV 2 variant) peptide is a member of the Omicron lineage, and the group-associated peptide includes the corresponding peptide (e.g., spike protein sequence) of the observed variant belonging to the Omicron lineage.
[0060] In some embodiments, one or more conserved regions that trigger memory are or comprise a set of conserved epitope regions, and the engineered antigen has at least one amino acid modification in each of at least a portion (e.g., all; e.g., a specific subgroup) of the conserved epitope regions.
[0061] In some implementations, one or more conserved regions that trigger memory are or contain a conserved surface, and one or more amino acid modifications occur at locations that are distributed throughout / across the conserved surface (e.g., approximately uniform rather than clustered) [e.g., distributed substantially equidistantly on the 3D representation of the conserved surface (e.g., where the 3D linear distance and / or geodesic distance (across the 3D conserved surface) between each of the multiple amino acid modifications is substantially equal and / or distributed according to a specific predefined statistical distribution (e.g., a normal distribution)].
[0062] In some embodiments, [for example, where the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or a portion thereof (for example, where the reference antigen is or contains an S protein RBD) and] the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA contain one or more combinations of mutations listed in Table 5A (for example, in each row, second column, or third column) [for example, where the engineered antigen contains an RBD of the SARS-CoV-2 S protein having one or more combinations of mutations listed in the third column of Table 5A (for example, according to any one of SEQ ID NO:4 or 38) (for example, residues 327 to 528 of SEQ ID NO:1)] [for example, where the engineered antigen contains an RBD of the SARS-CoV-2 XBB.1.5 variant S protein (for example, according to any one of SEQ ID NO:35, 36, or 37) (for example, residues 327 to 528 of SEQ ID NO:1, having at least a portion of the XBB.1.5 signature mutation identified in Table 1B, and) has one or more additional combinations of mutations listed in the second column of Table 5A].
[0063] In some implementations, [for example, where the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or a portion thereof (e.g., where the reference antigen is or contains the S protein RBD) and] the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA contain one or more sequences listed in Table 5B (e.g., each row) [e.g., where the engineered antigen is or contains the SARS-CoV-2 S protein RBD having any of the sequences listed in Table 5B (e.g., any of SEQ ID NO: 7 to 34)].
[0064] In some embodiments, [for example, where the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or a portion thereof (e.g., where the reference antigen is or contains the S protein RBD) and] the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA contain one or more of the mutation combinations listed in Table 7A and / or Table 7B (e.g., in each row, first column, or second column) [for example, where the engineered antigen contains the SARS-CoV-2 XBB.1.5 variant S protein RBD (e.g., according to any one of SEQ ID NO: 35, 36, or 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 having at least a portion of the XBB.1.5 signature mutation identified in Table 1B, and) have one or more of the additional mutation combinations listed in Table 7A and / or Table 7B second column (e.g., identified as L1, L2, L3, and L4)].
[0065] In some embodiments, [for example, where the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or a portion thereof (for example, where the reference antigen is or contains an S protein RBD) and] the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA contain one or more combinations of mutations listed in Table 8A (for example, in each row, second column, or third column) [for example, where the engineered antigen contains an RBD of the SARS-CoV-2 S protein having one or more combinations of mutations listed in the third column of Table 8A (for example, according to any one of SEQ ID NO:4 or 38) (for example, residues 327 to 528 of SEQ ID NO:1)] [for example, where the engineered antigen contains an RBD of the SARS-CoV-2 XBB.1.5 variant S protein (for example, according to any one of SEQ ID NO:35, 36, or 37) (for example, residues 327 to 528 of SEQ ID NO:1, having at least a portion of the XBB.1.5 signature mutation identified in Table 1B, and) has one or more additional combinations of mutations listed in the second column of Table 8A].
[0066] In some implementations, [for example, where the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or a portion thereof (e.g., where the reference antigen is or contains the S protein RBD) and] the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA contain one or more sequences listed in Table 8B (e.g., each row) [e.g., where the engineered antigen is or contains the SARS-CoV-2 S protein RBD having any of the sequences listed in Table 8B (e.g., any of SEQ ID NO: 39 to 64 and 100)].
[0067] In some embodiments, [for example, where the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or a portion thereof (e.g., where the reference antigen is or contains the S protein RBD) and] the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA contain one or more combinations of mutations listed in Table 9A (e.g., in each row, second column) [e.g., where the engineered antigen contains the SARS-CoV-2 XBB.1.5 variant S protein RBD (e.g., according to any one of SEQ ID NO: 35, 36, or 37) (e.g., residues 327 to 528 of SEQ ID NO: 1 having at least a portion of the XBB.1.5 signature mutation identified in Table 1B, and) have one or more additional combinations of mutations listed in the second column of Table 9A].
[0068] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA comprise a combination of mutations identified as S43 in Table 9A (e.g., wherein the engineered antigen is or comprises an XBB.1.5 S protein RBD with an additional mutation according to S43 in Table 9A) [e.g., wherein the engineered antigen comprises a SARS-CoV-2 XBB.1.5 variant S protein RBD (e.g., according to any one of SEQ ID NO: 35, 36, or 37) (e.g., residues 327 to 528 of SEQ ID NO: 1, having at least a portion of the XBB.1.5 signature mutation identified in Table 1B, and) having an additional mutation N360DP384SL390R T430IF464Y H519N].
[0069] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA are or contain at least a portion of the SARS-CoV-2S protein (e.g., according to SEQ ID NO:1) [e.g., a specific domain (e.g., RBD)] having at least a portion (e.g., all) of the following mutations: N360D, P384S, L390R, T430I, F464Y, H519N {e.g., in addition to one or more characteristic mutations of a specific SARS-CoV-2 variant (e.g., those within the RBD of the SARS-CoV-2S protein) [e.g., one or more XBB.1.5 signature mutations as follows: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, ...} H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N 460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K]}.
[0070] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA comprise a combination of mutations identified as S48 in Table 9A (e.g., wherein the engineered antigen is or comprises an XBB.1.5 RBD with an additional mutation according to S48 in Table 9A) [e.g., wherein the engineered antigen comprises a SARS-CoV-2 XBB.1.5 variant RBD (e.g., according to any one of SEQ ID NO: 35, 36, or 37) (e.g., residues 327 to 528 of SEQ ID NO: 1, having the XBB.1.5 signature mutation identified in Table 1B, and) having the additional mutation L335F K356T P384S L390RT430IF464YH519N].
[0071] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA are or contain at least a portion of the SARS-CoV-2S protein (e.g., according to SEQ ID NO:1) [e.g., a specific domain (e.g., RBD)] having at least a portion (e.g., all) of the following mutations: L335F, K356T, P384S, L390R, T430I, F464Y, and H519N {e.g., in addition to one or more characteristic mutations of a specific SARS-CoV-2 variant (e.g., those within the RBD of the SARS-CoV-2S protein) [e.g., one or more XBB.1.5 signature mutations as follows: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y14] 4-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S , N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K]}.
[0072] In some embodiments, [for example, where the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or a portion thereof (e.g., where the reference antigen is or contains the S protein RBD) and] the provided methods, systems, vaccine compositions, manufacturing methods, or engineered RNA antigens contain at least a portion of the SARS-CoV-2 S protein (e.g., RBD) having the mutation P384S L390R T430I F464Y H519N.
[0073] In some embodiments, [for example, where the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or a portion thereof (e.g., where the reference antigen is or contains the S protein RBD) and] the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA contain at least a portion of one or more SARS-CoV-2 S proteins (e.g., RBDs) having (e.g., a subgroup; e.g., all) the mutations K356T P384SL390RT430I F464Y H519N.
[0074] In some implementations, [for example, where the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or a portion thereof (e.g., where the reference antigen is or contains the S protein RBD) and] the provided methods, systems, vaccine compositions, manufacturing methods, or RNA do not include one or both of the mutations L335F and L390R.
[0075] In some embodiments, [for example, where the reference antigen is a specific target variant of the SARS-CoV-2 S protein and / or a portion thereof (for example, where the reference antigen is or contains the S protein RBD) and] the provided methods, systems, vaccine compositions, manufacturing methods, or RNA contain one or more combinations of mutations listed in Table 15A and / or Table 15B (for example, in each row) [for example, where the engineered antigen contains the RBD of the SARS-CoV-2 S protein having one or more combinations of mutations listed in Table 15A (for example, according to any one of SEQ ID NO:4 or 38) (for example, residues 327 to 528 of SEQ ID NO:1)] [for example, where the engineered antigen contains the SARS-CoV-2 XBB.1.5 variant RBD (for example, according to any one of SEQ ID NO:35, 36, or 37) (for example, residues 327 to 528 of SEQ ID NO:1, having at least a portion of the XBB.1.5 signature mutation identified in Table 1B, and) has one or more additional combinations of mutations listed in the second column of Table 15B].
[0076] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA include combinations of mutations identified as S123 in Tables 15A and / or 15B.
[0077] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA are or contain at least a portion of the SARS-CoV-2S protein (e.g., according to SEQ ID NO:1) [e.g., a specific domain (e.g., RBD)] having at least a portion (e.g., all) of the following mutations: I332V L335F K356T P384ST430IL452QF464Y H519N {e.g., in addition to one or more characteristic mutations of a specific SARS-CoV-2 variant (e.g., those within the RBD of the SARS-CoV-2S protein) [e.g., where one or more XBB.1.5 signature mutations are as follows: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L 368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E 484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K]}.
[0078] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA contain at least some (e.g., at most all) of the following mutations: I332V L335F G339H R346T K356T L368IS371F S373P S375F T376AP348SD405N R408S K417N T430I N440K V445P G446SL452QN460K F464Y S477N T478K E484A F486P F490S Q498R N501YY505H E516Q H519N.
[0079] In some implementations, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA include combinations of mutations identified as S122 in Tables 15A and / or 15B.
[0080] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA are or contain at least a portion of the SARS-CoV-2S protein (e.g., according to SEQ ID NO:1) [e.g., a specific domain (e.g., RBD)] having at least a portion (e.g., all) of the following mutations: L335F K356T P384S T430IL452RF464YH519N {e.g., in addition to one or more characteristic mutations of a specific SARS-CoV-2 variant (e.g., those within the RBD of the SARS-CoV-2S protein) [e.g., where one or more XBB.1.5 signature mutations are as follows: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478 K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K]}.
[0081] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA contain at least some (e.g., at most all) of the following mutations: L335F G339H R346T K356T L368I S371FS373P S375F T376A P348SD405N R408S K417N T430I N440K V445P G446S L452RN460KF464Y S477N T478K E484A F486P F490S Q498R N501Y Y505HE516Q H519N T523S.
[0082] In some implementations, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA include combinations of mutations identified as S109 in Tables 15A and / or 15B.
[0083] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA are or contain at least a portion of the SARS-CoV-2S protein (e.g., according to SEQ ID NO:1) [e.g., a specific domain (e.g., RBD)] having at least a portion (e.g., all) of the following mutations: K356T N360S P384S N388KT430IN450D F464Y H519N {e.g., in addition to one or more characteristic mutations of a specific SARS-CoV-2 variant (e.g., those within the RBD of the SARS-CoV-2S protein) [e.g., where one or more XBB.1.5 signature mutations are as follows: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L 368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E 484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K]}.
[0084] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA contain at least some (e.g., at most all) of the following mutations: G339H R346T K356T L368I S371F S373PS375F T376A P348S N388KD405N R408S K417N T430I N440K V445P G446S N450DN460KF464Y S477N T478K E484A F486P F490S Q498R N501Y Y505HH519N T523S.
[0085] In some implementations, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA include combinations of mutations identified as S129 in Tables 15A and / or 15B.
[0086] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA are or contain at least a portion of the SARS-CoV-2S protein (e.g., according to SEQ ID NO:1) [e.g., a specific domain (e.g., RBD)] having at least a portion (e.g., all) of the following mutations: K356T L335F P384SD389GT430IN450D F464Y H519N {e.g., in addition to one or more characteristic mutations of a specific SARS-CoV-2 variant (e.g., those within the RBD of the SARS-CoV-2S protein) [e.g., where one or more XBB.1.5 signature mutations are as follows: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L 368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E 484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K]}.
[0087] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA contain at least some (e.g., at most all) of the following mutations: L335F G339H R346T K356T L368I S371FS373P S375F T376A P348SD389G D405N R408S K417N T430I N440K V445P G446SN450DL452R N460K F464Y S477N T478K E484A F486P F490S Q498RN501Y Y505H H519N.
[0088] In some implementations, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA include combinations of mutations identified as S156 in Tables 15A and / or 15B.
[0089] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA are or contain at least a portion of the SARS-CoV-2S protein (e.g., according to SEQ ID NO:1) [e.g., a specific domain (e.g., RBD)] having at least a portion (e.g., all) of the following mutations: K356T P384S L390RT430IN450DF464Y I472V H519N {e.g., in addition to one or more characteristic mutations of a specific SARS-CoV-2 variant (e.g., those within the RBD of the SARS-CoV-2S protein) [e.g., where one or more XBB.1.5 signature mutations are as follows: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L 368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E 484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K]}.
[0090] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA contain at least some (e.g., at most all) of the following mutations: G339H R346T K356T L368I S371F S373PS375F T376A P348S L390RD405N R408S K417N T430I N440K V445P G446S N450DL452RN460K F464Y I472V S477N T478K E484A F486P F490S Q498RN501Y Y505H H519N.
[0091] In some implementations, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA include combinations of mutations identified as S112 in Tables 15A and / or 15B.
[0092] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA are or contain at least a portion of the SARS-CoV-2S protein (e.g., according to SEQ ID NO:1) [e.g., a specific domain (e.g., RBD)] having at least a portion (e.g., all) of the following mutations: K356T P384SD389G T430IN450DF464YI468V H519N {e.g., in addition to one or more characteristic mutations of a specific SARS-CoV-2 variant (e.g., those within the RBD of the SARS-CoV-2S protein) [e.g., where one or more XBB.1.5 signature mutations are as follows: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L 368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E 484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K]}.
[0093] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA contain at least some (e.g., at most all) of the following mutations: G339H R346T K356T L368I S371F S373PS375F T376A P348SD389GD405N R408S K417N T430I N440K V445P G446S N450DN460KF464Y I468V S477N T478K E484A F486P F490S Q498R N501YY505H H519N.
[0094] In some implementations, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA include combinations of mutations identified as S125 in Tables 15A and / or 15B.
[0095] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA are or contain at least a portion of the SARS-CoV-2S protein (e.g., according to SEQ ID NO:1) [e.g., a specific domain (e.g., RBD)] having at least a portion (e.g., all) of the following mutations: K356T L335F P384ST430IF464YI468VH519N {e.g., in addition to one or more characteristic mutations of a specific SARS-CoV-2 variant (e.g., those within the RBD of the SARS-CoV-2S protein) [e.g., where one or more XBB.1.5 signature mutations are as follows: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G33 9H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R408S, K417N, N440K, V445P, G446S, N460K, S477N, T4 78K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K]}.
[0096] In some embodiments, the provided methods, systems, vaccine compositions, manufacturing methods, or engineered antigens of RNA contain at least some (e.g., at most all) of the following mutations: L335F G339H R346T K356T L368I S371FS373P S375F T376A P348SD405N R408S K417N T430I N440K V445P G446S N460KF464YI468V S477N T478K E484A F486P F490S Q498R N501Y Y505HE516Q H519N T523S.
[0097] In some aspects, this disclosure provides a method for manufacturing RNA comprising a nucleotide sequence encoding an engineered antigen corresponding to an engineered version of a reference antigen, the method comprising producing such RNA having a nucleotide sequence that, when compared with a reference antigen (e.g., the nucleotide sequence of the reference antigen), shows differences relative to the reference antigen in one or more memory-triggering conserved regions [e.g., one or more conserved epitopes and / or one or more conserved surfaces] (e.g., comprising nucleotides encoding one or more mutations), the conserved regions being common to the reference antigen and (i) one or more pre-existing variants of the reference antigen and / or (ii) wild-type strains of the reference antigen.
[0098] In some implementations, the reference antigen is or contains a specific target variant of the SARS-CoV-2 S protein RBD [e.g., XBB.1.5 RBD (e.g., having an amino acid sequence according to any one of SEQ ID NO:35, 36 and 37) (e.g., residues 327 to 528 of SEQ ID NO:1 with the XBB.1.5 signature mutation identified in Table 1B)].
[0099] In some embodiments, the RNA produced according to the provided methods (e.g., as described above) is or contains any of the aspects and embodiments described herein (e.g., where the difference compared to the reference antigen corresponds to / includes nucleotide differences encoding any combination of mutations in the aspects and embodiments described herein, e.g., in the paragraph above).
[0100] In some implementations, the manufacturing method provided includes producing RNA via in vitro transcription (IVT).
[0101] Features of the embodiments described with respect to one aspect of the invention may be applied with respect to another aspect of the invention. Brief description of the attached diagram
[0103] The foregoing and other objects, aspects, features, and advantages of this disclosure will become more apparent and better understood from the following description taken in conjunction with the accompanying drawings, wherein:
[0104] Figure 1 These are block diagrams and schematic diagrams illustrating an example process for designing engineered antigens, based on an illustrative implementation scheme.
[0105] Figure 2 This is a flowchart illustrating an example process of inserting amino acid modifications according to an illustrative implementation scheme.
[0106] Figure 3 This is a block diagram of an example method for generating and scoring multiple candidate antigen designs, according to an illustrative implementation scheme.
[0107] Figure 4A This is a set of views of a 3D model of the receptor-binding domain (RBD) of the XBB variant of the SARS-Cov2 spike (S) protein, according to an illustrative implementation scheme, which is colored to show the signature mutations of XBB and the identified conserved regions.
[0108] Figure 4B It is based on the illustrative implementation plan. Figure 6A Another set of views of the XBB RBD S protein model shown is colored to show the signature XBB mutation, the identified conserved region, and the ACE2 binding interface.
[0109] Figure 4C It is based on the illustrative implementation plan. Figure 6A and 6B Another set of views of the XBB RBD S protein model shown is colored to show the relative frequencies of epitope involvement at individual amino acid sites within the protein.
[0110] Figure 4D This is a set of views of the first engineered antigen design using the XBB RBD S protein as a reference antigen.
[0111] Figure 4E This is a set of views of a 3D model of a second engineered antigen design generated using the XBB RBD S protein as a reference antigen.
[0112] Figure 5A This is a schematic diagram of SARS-CoV 2 and its related proteins, adapted from Jain et al., Vaccines 2020, 8(4), 649 (www.mdpi.com / 2076-393X / 8 / 4 / 649.).
[0113] Figure 5B This is a set of views of a 3D model of the receptor-binding domain (RBD) of the SARS-Cov2 spike (S) protein of the XBB.1.5 variant, according to an illustrative embodiment, which is colored to show the signature mutations of XBB and the identified conserved regions.
[0114] Figure 5C This is a set of views of a 3D model of the XBB.1.5 spike protein (the whole spike protein) according to the illustrative implementation scheme.
[0115] Figure 5D It is based on the illustrative implementation plan. Figure 7 Another set of views of the XBB RBD S protein model shown in B is colored to show the XBB signature mutation, the identified conserved region, and the ACE2 binding interface.
[0116] Figure 5EIt is based on the illustrative implementation plan. Figure 5B and 5D Another set of views of the XBB RBD S protein model shown is colored to show the relative frequencies of epitope involvement at individual amino acid sites within the protein.
[0117] Figure 5F It is based on the illustrative implementation plan. Figure 5B and 5D Another set of views of the XBB RBD S protein model shown is colored to show the relative frequencies of epitope involvement at individual amino acid sites within the conserved surface of the protein.
[0118] Figure 5G This is a block flowchart illustrating an example method for generating candidate variant designs, based on an illustrative implementation scheme.
[0119] Figure 6A A set of views is shown of a 3D model of a candidate engineered antigen design generated using the XBB.1.5RBD S protein as a reference antigen, according to an illustrative embodiment. Figure 6A The coloring / shading in the images identified the signature mutations of XBB.1.5, the identified conserved regions, and the amino acid sites mutated based on the various computer-simulated antigen design methods described herein.
[0120] Figure 6B A set of views showing a 3D model of the SARS-CoV-2 spike protein, highlighting the XBB.1.5 signature mutation and its relationship with... Figure 6A The additional mutations corresponding to those mutations shown are shown.
[0121] Figure 6C A set of views is shown of a 3D model designed based on (e.g., using) the XBB.1.5S protein RBD as a reference antigen, according to an illustrative embodiment.
[0122] Figure 6D A set of views is shown of a 3D model designed based on (e.g., using) the XBB.1.5S protein RBD as a reference antigen, according to an illustrative embodiment.
[0123] Figure 6E A set of views is shown of a 3D model designed based on (e.g., using) the XBB.1.5S protein RBD as a reference antigen, according to an illustrative embodiment.
[0124] Figure 6FA set of views is shown of a 3D model designed based on (e.g., using) the XBB.1.5S protein RBD as a reference antigen, according to an illustrative embodiment.
[0125] Figure 7 A radar chart showing some performance scores of candidate engineered antigen designs according to the illustrative implementation is presented.
[0126] Figure 8 A graph showing the predicted changes in ACE2 binding compared to the experimental implementation is shown, based on the illustrative implementation scheme.
[0127] Figure 9A A 3D model view of the engineered RBD construct according to the illustrative implementation is shown.
[0128] Figure 9B A 3D model view of the engineered RBD construct according to the illustrative implementation is shown.
[0129] Figure 9C This is a schematic diagram of the SARS-CoV-2 trimer, adapted from Starr et al., SARS-CoV-2 RBD antibodies that maximize breadth and resistance to escape. Nature, 597, 97–102 (2021). https: / / doi.org / 10.1038 / s41586-021-03807-6.
[0130] Figure 9D A 3D model view of an engineered construct designed to integrate one or more SARS-CoV-1 epitopes on the XBB.1.5 main chain, according to an illustrative implementation, is shown.
[0131] Figure 9E A 3D model view of an engineered construct designed to integrate one or more SARS-CoV-1 epitopes on the XBB.1.5 main chain, according to an illustrative implementation, is shown.
[0132] Figure 10A This is an example procedure for testing engineered antigens, based on an illustrative implementation scheme.
[0133] Figure 10B It is a block diagram of an example scheme used to generate and evaluate the engineered building blocks described herein, based on an illustrative implementation scheme.
[0134] Figure 10C It is a set of graphs evaluating the binding of antibodies to WT and XBB.1.5S proteins for different epitope classes, according to the illustrative implementation scheme.
[0135] Figure 10D This is a graph showing the dilution series of hACE2-mFc binders.
[0136] Figure 10E It is human BNT162b2 3 (Triple vaccination) Dilution series of polyclonal serum.
[0137] Figure 11A The binding dose-response curves for some of the engineered SARS-CoV 2 antigens described herein are shown.
[0138] Figure 11B The binding dose-response curves for some of the engineered SARS-CoV 2 antigens described herein are shown.
[0139] Figure 11C The binding dose-response curves for some of the engineered SARS-CoV 2 antigens described herein are shown.
[0140] Figure 12A The binding dose-response curves for some of the engineered SARS-CoV 2 antigens described herein are shown.
[0141] Figure 12B The binding dose-response curves for some of the engineered SARS-CoV 2 antigens described herein are shown.
[0142] Figure 13A The binding dose-response curves for some of the engineered SARS-CoV 2 antigens described herein are shown.
[0143] Figure 13B The binding dose-response curves for some of the engineered SARS-CoV 2 antigens described herein are shown.
[0144] Figure 13C The binding dose-response curves for some of the engineered SARS-CoV 2 antigens described herein are shown.
[0145] Figure 13D The binding dose-response curves for some of the engineered SARS-CoV 2 antigens described herein are shown.
[0146] Figure 14A The binding dose-response curves for some of the engineered SARS-CoV 2 antigens described herein are shown.
[0147] Figure 14B The binding dose-response curves for some of the engineered SARS-CoV 2 antigens described herein are shown.
[0148] Figure 14CThe binding dose-response curves for some of the engineered SARS-CoV 2 antigens described herein are shown.
[0149] Figure 14D The binding dose-response curves for some of the engineered SARS-CoV 2 antigens described herein are shown.
[0150] Figure 15A This is a schematic diagram illustrating example groups and a dose / sampling schedule used to evaluate the immunoblotting and performance of a candidate vaccine in vaccine-vaccinated mice. According to the illustrative implementation, yellow-filled cells indicate the date on which serum samples will be collected, gray-filled cells indicate the date on which the vaccine will be administered, and green-filled cells indicate the date on which mice will be euthanized and the final samples collected.
[0151] Figure 15B This is a schematic diagram illustrating the example groups and dosage / sampling schedule used to evaluate the performance of the candidate vaccine in vaccine-unvaccinated mice. According to the illustrative implementation, yellow-filled cells indicate the date on which serum samples will be collected, gray-filled cells indicate the date on which the vaccine will be administered, and green-filled cells indicate the date on which mice will be euthanized and the final samples will be collected.
[0152] Figure 16A This is an illustrative implementation of a FACS-based measurement, for example using the method described in Quandt and Muik et al., 2022, the contents of which are incorporated herein by reference in their entirety.
[0153] Figure 16B This is an illustrative embodiment illustrating a consumption measurement, for example using the method described in Quandt and Muik et al., 2022, the contents of which are incorporated herein by reference in their entirety.
[0154] Figure 17 This is a schematic diagram illustrating, according to certain implementation schemes, methods for collecting and analyzing spleen samples, lymph node samples, and blood samples.
[0155] Figure 18 It is a diagram illustrating the mutations included in various engineered antigen construct designs, based on an illustrative implementation scheme.
[0156] Figure 19A This is a schematic diagram illustrating an exemplary construct comprising a membrane-anchored RBD+ fibrous substitute protein domain having a viral signal peptide, according to an illustrative embodiment.
[0157] Figure 19B Based on the illustrative implementation scheme, a graph showing the ACE2 binding assay results of a set of engineered construct designs is presented.
[0158] Figure 19C This is a chart showing the results of a set of engineered constructs for serum (from vaccinated patients) binding assays, based on an illustrative implementation scheme.
[0159] Figure 20A It is a set of graphs showing the results of binding assays for a group of monoclonal antibodies, according to the illustrative implementation scheme.
[0160] Figure 20B It is a set of graphs showing the results of binding assays for a group of monoclonal antibodies, according to the illustrative implementation scheme.
[0161] Figure 20C It is a set of graphs showing the results of binding assays for a group of monoclonal antibodies, according to the illustrative implementation scheme.
[0162] Figure 21 This is a block diagram of an exemplary cloud computing environment for certain implementation schemes.
[0163] Figure 22 These are diagrams of example computing devices and example mobile computing devices used in certain implementations.
[0164] The features and advantages of this disclosure will become more apparent from the detailed description set forth in conjunction with the accompanying drawings, in which the same reference characters consistently identify corresponding elements. In the drawings, similar reference numerals generally indicate the same, functionally similar, and / or structurally similar elements.
[0165] Some definitions
[0166] About or approximately: The term “about” or “approximately”, when used herein to refer to a value, means a value similar to a reference value. Generally, those skilled in the art will understand the extent of variation covered by “about” or “approximately” in that context. For example, in some embodiments, the term “about” or “approximately” may cover a range of values within the range of 25%, 20%, 19%, 18%, 17%, 16%, 15%, 14%, 13%, 12%, 11%, 10%, 9%, 8%, 7%, 6%, 5%, 4%, 3%, 2%, 1%, or less of the mentioned value.
[0167] Administration: As used herein, the term "administration" generally refers to the application of the composition to a subject or system. Those skilled in the art will recognize a variety of routes of administration that may be used to administer the composition to a subject (e.g., a person) where appropriate. For example, in some embodiments, administration may be via the eye, oral, parenteral, topical, etc. In some specific embodiments, administration may be via the bronchus (e.g., by bronchial infusion), buccal, dermal (which may be one or more of, for example, topical application to the dermis, intradermis, intradermal, transdermal, etc.), intestine, intra-arterial, intradermal, gastric, intramedullary, intramuscular, intranasal, intraperitoneal, intrathecal, intravenous, intravenous, intracardiac, in a specific organ (e.g., intrahepatic), mucosa, nose, oral, rectal, subcutaneous, sublingual, topical, trachea (e.g., by tracheal infusion), vagina, vitreous, etc. In some embodiments, administration may involve intermittent (e.g., multiple doses separated by time) and / or periodic (e.g., individual doses separated by a common time period) administration. In some implementations, administration may involve continuous administration (e.g., infusion) for at least a selected period of time.
[0168] Adult: As used herein, the term “adult” refers to a human being eighteen years of age or older. In some implementations, an adult’s weight is in the range of approximately 90 pounds to approximately 250 pounds.
[0169] Affinity: As is known in the art, “affinity” is a measure of the tightness with which two or more binding couples bind to each other. Those skilled in the art are aware of various assays that can be used to assess affinity, and also of appropriate controls for such assays. In some embodiments, affinity is assessed by quantitative assay. In some embodiments, affinity is assessed at multiple concentrations (e.g., one binding couple at a time). In some embodiments, affinity is assessed in the presence of one or more potential competing entities (e.g., possibly present in a relevant environment (e.g., a physiological environment)). In some embodiments, affinity is assessed relative to a reference (e.g., a known affinity above a specific threshold [“positive control” reference”] or a known affinity below a specific threshold [“negative control” reference”). In some embodiments, affinity may be assessed relative to a contemporaneous reference; in some embodiments, affinity may be assessed relative to a historical reference. Generally, when affinity is assessed relative to a reference, it is assessed under comparable conditions.
[0170] Agent: Generally, as used herein, the term "agent" refers to an entity (lipids, metals, nucleic acids, polypeptides, polysaccharides, small molecules, etc., or complexes, combinations, mixtures, or systems thereof [e.g., cells, tissues, organisms]) or phenomenon (e.g., heat, electric current or electric field, magnetic force or magnetic field, etc.). Where appropriate, as will be clear from the context to those skilled in the art, the term may be used to refer to an entity that is or includes cells or organisms, or fractions, extracts, or components thereof. Alternatively or additionally, as will be clear from the context, the term may be used to refer to natural products found in and / or obtained from nature. In some cases, as will be clear from the context, the term may be used to refer to one or more such entities that are artificial because they are designed, engineered, and / or produced by human hand and / or not found in nature. In some embodiments, the agent may be used in isolated or pure form; in some embodiments, the agent may be used in crude form. In some embodiments, potential agents may be provided as a collection or library, for example, that can be screened to identify or characterize the active agents therein. In some instances, the term "agent" may refer to a polymer or a compound or entity comprising a polymer; in other instances, the term may refer to a compound or entity comprising one or more polymer portions. In some embodiments, the term "agent" may refer to a compound or entity that is not a polymer and / or substantially free of any polymer and / or one or more specific polymer portions. In some embodiments, the term may refer to a compound or entity lacking or substantially free of any polymer portion.
[0171] Improvement: As used herein, the term “improvement” means the prevention, reduction, or alleviation of a state, or the improvement of a subject’s condition. Improvement includes, but is not limited to, the complete restoration or complete prevention of a disease, ailment, or disorder (e.g., radiation damage).
[0172] Amino acids: As used herein, the term "amino acid" in its broadest sense refers to compounds and / or substances that can, are, or have been incorporated into a polypeptide chain, for example, by forming one or more peptide bonds. In some embodiments, amino acids have the general structure H₂N-C(H)(R)-COOH. In some embodiments, amino acids are naturally occurring amino acids. In some embodiments, amino acids are non-natural amino acids; in some embodiments, amino acids are D-amino acids; in some embodiments, amino acids are L-amino acids. "Standard amino acid" refers to any of the twenty standard L-amino acids commonly found in naturally occurring peptides. "Non-standard amino acid" refers to any amino acid other than a standard amino acid, whether it is synthetically prepared or obtained from a natural source. In some embodiments, the amino acids in the polypeptide (including carboxyl and / or amino-terminal amino acids) may contain structural modifications compared to the general structure described above. For example, in some embodiments, amino acids may be modified compared to the general structure by methylation, amidation, acetylation, PEGylation, glycosylation, phosphorylation, and / or substitution (e.g., of amino, carboxylic acid groups, one or more protons and / or hydroxyl groups). In some embodiments, such modifications may, for example, alter the cycling half-life of a peptide containing modified amino acids compared to a peptide containing otherwise identical unmodified amino acids. In some embodiments, such modifications do not significantly alter the relevant activity of a peptide containing modified amino acids compared to a peptide containing otherwise identical unmodified amino acids. As will be clear from the context, in some embodiments, the term "amino acid" may be used to refer to a free amino acid; in some embodiments, it may be used to refer to the amino acid residues of a peptide.
[0173] Animal: As used herein, the term "animal" means any member of the animal kingdom. In some embodiments, "animal" means a human being of any sex and at any developmental stage. In some embodiments, "animal" means a non-human animal at any developmental stage. In some embodiments, a non-human animal is a mammal (e.g., rodents, mice, rats, rabbits, monkeys, dogs, cats, sheep, cattle, primates, and / or pigs). In some embodiments, animals include, but are not limited to, mammals, birds, reptiles, amphibians, fish, insects, and / or worms. In some embodiments, an animal may be a transgenic animal, a genetically engineered animal, and / or a clone.
[0174] Antibody Agent: As used herein, the term "antibody" refers to any polypeptide or polypeptide complex comprising immunoglobulin sequence elements sufficient to confer specific binding to a particular antigen. As is known in the art, a naturally occurring complete antibody is a tetramer of approximately 150 kDa, consisting of two identical heavy-chain polypeptides (approximately 50 kDa each) and two identical light-chain polypeptides (approximately 25 kDa each), which bind together to form what is commonly referred to as a "Y-shaped" structure. Each heavy chain consists of at least four domains (each approximately 110 amino acids long) – an amino-terminal variable (VH) domain (located at the top of the Y-shaped structure), followed by three constant domains: CHI, CH2, and a carboxyl-terminal CH3 (located at the bottom of the Y-shaped stem). A short region called a "switch" connects the heavy-chain variable and constant regions. A "hinge" links the CH2 and CH3 domains to the rest of the antibody. Two disulfide bonds in this hinge region link the two heavy-chain polypeptides in the complete antibody to each other. Each light chain consists of two domains—an amino-terminal variable (VL) domain followed by a carboxyl-terminal constant (CL) domain, separated from each other by another “switch.” A complete antibody tetramer consists of two heavy-chain-light-chain dimers, where the heavy and light chains are linked to each other by a single disulfide bond; two other disulfide bonds link the heavy-chain hinge regions to each other, thus linking the dimers together to form a tetramer. Naturally occurring antibodies are also glycosylated, typically at the CH2 domain. Each domain in a natural antibody has a structure characterized by an “immunoglobulin fold,” formed by two β-sheets (e.g., 3-, 4-, or 5-chain folds) compressed together in a compressed, antiparallel β-barrel. Each variable domain contains three hypervariable loops (CDR1, CDR2, and CDR3) called “complement-determining regions” and four slightly invariant “framework” regions (FR1, FR2, FR3, and FR4). When a natural antibody folds, the FR region forms a β-sheet that provides a structural framework for the domain, and the CDR loop regions of the heavy and light chains aggregate in three-dimensional space, thereby creating a single hypervariable antigen-binding site located at the top of the Y-shaped structure. The Fc region of a naturally occurring antibody binds to elements of the complement system and also to receptors on effector cells (including, for example, effector cells mediating cytotoxicity). As is known in the art, the affinity and / or other binding properties of the Fc region to Fc receptors can be modulated by glycosylation or other modifications. In some embodiments, antibodies generated and / or utilized according to this disclosure include glycosylated Fc domains, including such glycosylated Fc domains with modifications or engineering. For the purposes of this disclosure, in some embodiments, any polypeptide or polypeptide complex containing sufficient immunoglobulin domain sequences as found in natural antibodies may be called and / or used as an "antibody," whether such polypeptide is naturally generated (e.g., produced by an organism reacting with an antigen) or generated by recombinant engineering, chemical synthesis, or other artificial systems or methods.In some embodiments, the antibody is polyclonal; in other embodiments, the antibody is monoclonal. In some embodiments, the antibody has a constant region sequence specific to mouse, rabbit, primate, or human antibodies. In some embodiments, the antibody sequence element is humanized, primate-derived, chimeric, etc., as known in the art. Additionally, the term "antibody," as used herein, may refer to any construct or form known or developed in the art (unless otherwise specified or obvious from the context) for utilizing antibody structural and functional characteristics in alternative presentations.
[0175] For example, in some embodiments, the antibodies used according to this disclosure are selected from, but not limited to, intact IgA, IgG, IgE, or IgM antibodies; bispecific or multispecific antibodies (e.g., (etc.); antibody fragments, such as Fab fragments, Fab' fragments, F(ab')2 fragments, Fd' fragments, Fd fragments, and isolated CDRs or groups thereof; single-chain Fvs; peptide-Fc fusions; single-domain antibodies, alternative scaffolds, or antibody mimics (e.g., anti-carrier proteins, FN3 single-domain antibodies, DARPin, affinity molecules, Affilin, Affimer, Affitin, alpha bodies, Avimer, Fynomer, Im7, VLR, VNAR, Trimab, CrossMab, Trident); nanobodies, dual nanobodies, F(ab')2, Fab', di-sdFv, single-domain antibodies, trifunctional antibodies, biantibodies, and microantibodies, etc. In some embodiments, the relevant form may be or includes: Antibodies against camelids; Anchor protein repeat sequence or Biparental and highly targeted (DART) agents; Shark single-domain antibodies, such as IgNAR; anti-cancer immune mobilization monoclonal T-cell receptor (ImmTAC); Microproteins; Microantibodies; masking antibodies (e.g., Small modular immunotherapies (“SMIP™”); single-chain or tandem dual antibodies TCR-like antibodies; VHH. In some embodiments, the antibody may lack the covalent modifications (e.g., glycan linkages) it would have in naturally occurring form. In some embodiments, the antibody may contain covalent modifications (e.g., glycans, payloads [e.g., detectable portions, therapeutic portions, catalytic portions, etc.] or other side groups [e.g., polyethylene glycol, etc.] linkages).
[0176] Antigen: As used herein, the term "antigen" refers to an agent that elicits an immune response; and / or (ii) an agent that binds to T cell receptors (e.g., when presented by MHC molecules) or to antibodies. In some embodiments, an antigen elicits a humoral response (e.g., including the production of antigen-specific antibodies); in some embodiments, an antigen elicits a cellular response (e.g., involving T cells whose receptors specifically interact with the antigen). In some embodiments, the antigen binds to an antibody and may or may not induce a specific physiological response in the organism. Generally, an antigen can be or comprises any chemical entity, such as, for example, small molecules, nucleic acids, peptides, carbohydrates, lipids, polymers (in some embodiments, other than biopolymers [e.g., other than nucleic acid or amino acid polymers]), etc. In some embodiments, the antigen is or comprises a peptide. In some embodiments, the antigen is or comprises a glycan. Those skilled in the art will appreciate that, generally, antigens can be provided in an isolated or pure form, or alternatively, in a crude form (e.g., together with other materials, such as extracts of cell extracts or other relatively crude formulations containing antigen sources). In some embodiments, the antigen used according to the invention is provided in a crude form. In some implementations, the antigen is a recombinant antigen.
[0177] Emerging epitopes: As used herein, the term “emerging epitope” refers to an epitope that the subject’s immune system has not previously encountered, such as a part of an antigen. For example, as circulating viral pathogens mutate, mutations in various epitopes found on viral proteins can lead to the emergence of epitopes that are of the type of epitopes present on proteins of previously circulating variants, with one or more mutations introduced.
[0178] Composition: Those skilled in the art will understand that the term "composition" can be used to refer to a discrete physical entity comprising one or more specified components. Generally, unless otherwise specified, a composition can be in any form, such as a gas, gel, liquid, solid, etc.
[0179] The term "including" is open-ended, meaning that a composition or method described herein as "including" one or more specified elements or steps is essential, but additional elements or steps may be added to the composition or method. To avoid verbosity, it should also be understood that any composition or method described as "including" (or "comprising") one or more specified elements or steps also describes a corresponding, more limited composition or method that is "substantially composed of (or "substantially composed of") the same specified elements or steps, meaning that the composition or method includes the specified essential elements or steps and may also include additional elements or steps that do not materially affect the fundamental and novel characteristics of the composition or method. It should also be understood that any composition or method described herein as "including" or "substantially composed of one or more specified elements or steps" also describes a corresponding, more limited and closed composition or method that is "composed of" (or "composed of") the specified elements or steps, excluding any other unspecified elements or steps. In any composition or method disclosed herein, any known or disclosed equivalent of any specified essential element or step may replace that element or step.
[0180] Determination: In some embodiments, the methods described herein include a "determination" step. Those skilled in the art who read this specification will understand that such "determination" can be achieved using any of the various techniques available to those skilled in the art (including, for example, specific techniques explicitly mentioned herein) or by using any of the various techniques available to those skilled in the art (including, for example, specific techniques explicitly mentioned herein). In some embodiments, determination involves manipulating a physical sample. In some embodiments, determination involves considering and / or manipulating data or information, for example, using a computer or other processing unit suitable for performing relevant analysis. In some embodiments, determination involves receiving relevant information and / or materials from a source. In some embodiments, determination involves comparing one or more features of a sample or entity with a comparable reference.
[0181] Engineered Antigens: As used herein, the term "engineered antigen" refers to an antigen that is artificially created or intended to be artificially created and intentionally introduced into a subject, for example, to elicit an immune response (e.g., through vaccination), rather than a naturally evolved antigen that has evolved through natural processes. For example, as described further in detail herein, in some embodiments, an engineered antigen is or comprises a polypeptide. Engineered polypeptide antigens may be designed to match, resemble, or be based on other reference antigens, which may themselves be engineered antigens or may be naturally occurring antigens. For example, in some embodiments, an engineered antigen is created by starting with the structure of a reference antigen and introducing one or more amino acid modifications therein, for example, to achieve the desired behavior / outcome when planned for introduction into a subject. In some embodiments, engineered antigens may be designed via computer simulation—i.e., systems and methods implemented via computers, using polypeptide models and other computer representations representing various physical antigens. In some embodiments, engineered antigens may be encoded by ribonucleic acid (RNA), which can then be used to manufacture the engineered antigen (e.g., in vitro) or may be administered directly to a subject, for example, as an RNA vaccine.
[0182] Epitope: As used herein, the term "epitaxe" is used to refer to any portion specifically recognized by an immunoglobulin (e.g., antibody or receptor) binding component. In some embodiments, an epitope consists of a plurality of chemical atoms or groups on an antigen. In some embodiments, such chemical atoms or groups are surface-exposed when the antigen takes an associated three-dimensional conformation. In some embodiments, such chemical atoms or groups are physically close to each other in space when the antigen takes such a conformation. In some embodiments, at least some of such chemical atoms or groups are physically spaced apart when the antigen takes an alternative conformation (e.g., is linearized).
[0183] Epitope Change Score: Used interchangeably herein, the terms “eptope change score” and “eptop score” refer to a measure of a change in the position of a viral peptide at an epitope. In some embodiments, such a change can be characterized by the effect of mutations in one or more epitopes of a viral variant on antibody (e.g., neutralizing antibody) recognition. For example, in some embodiments, such a change can be characterized by determining the number of antibodies that may escape. In some embodiments, the antibody used for characterization is isolated from a patient who has been vaccinated against a disease or has previously been infected with a disease (e.g., SARS-CoV-2). In some embodiments, the antibody used for characterization has previously been shown to bind to a reference sequence. In some embodiments, the epitope change score can be determined by comparing mutations in a variant candidate to one or more regions of a reference sequence that have previously shown to bind antibodies (e.g., via structural data). In some embodiments, the epitope change score can be determined by counting the number of unique epitopes involved in the changed position, as measured in one or more known antibody-viral peptide complex structures (e.g., all known antibody-viral peptide complex structures).
[0184] In some embodiments, the epitope alteration score is a measure of the number of different epitopes evaded by a variant candidate compared to a reference sequence (e.g., compared to a wild-type sequence). In some embodiments, the epitope alteration score is calculated based on known antibody binding sites, such as those reported in protein databases. In some embodiments, the epitope alteration score may vary over time as new epitope locations are identified and / or epitope-binding antibodies are discovered. In some embodiments, the epitope alteration score can be used to characterize the degree of alteration of the SARS-CoV-2 spike peptide at epitope locations, for example, by counting the number or percentage of antibodies that may escape. In the various embodiments described herein, the epitope alteration score may be normalized such that it is ranked between 0 and 100%.
[0185] Growth Score: As used interchangeably herein, the terms “growth,” “growth index,” or “growth score” refer to a measure of the rate of growth of a given variant in a subject population (e.g., at a specific time). In some embodiments, the growth score refers to pedigree-level growth. For example, in some embodiments, the growth score of a given variant can be determined by the growth of a known variant of a reference parent species or a substantially identical pedigree, or a known variant having a similar sequence (e.g., a sequence at least 90% identical to the given variant). In some embodiments, the growth of a given variant is a function of the change in the number of subjects in a subject population reported to be infected with the given variant over a given time period relative to a reference infection rate (e.g., a reference infection rate determined over a specified time period). In some embodiments, the growth of a given variant is a function of the change in the proportion of a subject population infected with the given variant over a given time period relative to a reference infection rate (e.g., a reference infection rate determined over a specified time period). In some embodiments, the growth score of a given variant can be determined empirically by considering sequences observed within a specified time period that are associated with the given variant (e.g., in some embodiments, including sequences associated with a lineage) and calculating its proportion among all observed sequences relative to a reference level at a given time (e.g., a proportion determined within a specified time period). For example, in some embodiments, for each lineage, its proportion among all observed sequences within an extended time period (e.g., an eight-week window) and a more recent time window (e.g., the past 24 hours, past 48 hours, past 72 hours, past 4 days, past 5 days, past 6 days, or last week) is calculated, denoted by r_extended and r_duration, respectively. The growth of a lineage is defined by its ratio r_extended / r_duration, measuring the change in proportion. In the various embodiments described herein, the growth score can be normalized such that its ranking is between 0 and 100%.
[0186] Human: In some implementations, human refers to an embryo, fetus, infant, child, adolescent, adult, or elderly person.
[0187] Infectivity score or “fitness prior score” is used interchangeably herein. The term “infectivity score” or “fitness prior score” is a measure of the evolutionary fitness of a viral variant and is a function of viral replication efficiency and / or viral efficiency in infecting host cells. In some embodiments, the calculation of the fitness prior score includes determining one or more of a log-likelihood score, a viral peptide receptor binding score, and / or a growth score. In some embodiments, the fitness prior score is determined by referring to each of the log-likelihood score, the viral peptide receptor binding score, and the growth score.
[0188] “Improvement,” “Increase,” “Inhibit,” or “Reduction”: As used herein, the terms “improvement,” “increase,” “inhibit,” “reduction,” or their grammatical equivalents indicate a value relative to a baseline or other reference measurement. In some embodiments, an appropriate reference measurement may be or be included in a measurement performed in a particular system (e.g., in a single individual) under other comparable conditions in the absence of a particular agent or treatment (e.g., before and / or after) or in the presence of an appropriate comparable reference agent. In some embodiments, an appropriate reference measurement may be or be included in a measurement in a comparable system in which a particular agent or treatment is known or expected to respond in a particular manner in the presence of the relevant agent or treatment.
[0189] Log-Likelihood: As used herein, the term "log-likelihood" refers to a measure of the probability of the existence of a variant polypeptide sequence, determined using a natural language processing algorithm. In some embodiments, a transformer model may be used to determine the log-likelihood. In some embodiments, the log-likelihood may be determined without a reference sequence. In some embodiments, the log-likelihood is a reference-free transformer-derived log-likelihood. From a language model perspective, the higher the log-likelihood of a variant, the greater the likelihood of the variant's occurrence. In the various embodiments described herein, the log-likelihood score may be normalized such that it is ranked between 0 and 100%. In some embodiments, the log-likelihood measures how the log-likelihood of a variant polypeptide sequence compares to the entire population of known variants. In some embodiments, the log-likelihood measures how the log-likelihood of a variant polypeptide sequence compares to other variants with similar mutational loads ("conditional log-likelihood"). This conditional log-likelihood is particularly useful for evaluating variants with high mutation counts (e.g., at least 30 or more, including, for example, at least 40, at least 50, at least 60, at least 70 or more mutation counts).
[0190] Machine learning module, machine learning model: As used herein, the terms “machine learning module” and “machine learning model” are used interchangeably and refer to a computer-implemented process (e.g., software function) that implements one or more specific machine learning algorithms, such as artificial neural networks (ANNs), random forests, decision trees, support vector machines, etc., to determine one or more output values for a given input. In some implementations, the machine learning model is a deep learning model or a deep neural network—an ANN—that includes one or more hidden layers (e.g., in the middle) in addition to input and output layers. Examples of deep learning models include, but are not limited to, recurrent neural networks (RNNs) (e.g., long short-term memory networks (LSTMs), bidirectional LSTMs (biLSTMs)), attention-based networks (such as transformer models), and convolutional neural networks (CNNs). In some implementations, the machine learning module implementing the machine learning techniques is trained in a supervised manner, such as using a scrambled and / or manually annotated dataset. In some implementations, the machine learning model can be trained in an unsupervised manner using unlabeled data. In some implementations, the machine learning model can be trained using reinforcement methods, such as using a reward / penalty system to train the machine learning model to learn a strategy to perform a specified task. Training the machine learning model can be used to determine various parameters of the model, such as the weights associated with the layers in the neural network. In some implementations, once the machine learning module is trained, for example, to perform a specific task, such as predicting the type of hidden amino acids within a polypeptide sequence based on its context, the determined parameter values are fixed, and the machine learning module is used to process new data (e.g., different from the training data), such as new amino acid sequences, referred to as inference. In some implementations, the machine learning module may receive feedback, such as user comments on accuracy, and this feedback may be used as additional training data, for example, to dynamically update the machine learning module. In some implementations, the trained machine learning module is a classification algorithm with tunable and / or fixed (e.g., locked) parameters, such as a random forest classifier. In some implementations, two or more machine learning modules may be combined and implemented as a single module and / or a single software application. In some implementations, two or more machine learning modules may also be implemented separately, for example as separate software applications. The machine learning module can be software and / or hardware. For example, the machine learning module can be implemented entirely as software, or certain functions of the ANN module can be implemented using dedicated hardware (e.g., via application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), etc.).
[0191] Model: As used herein, the term "model" is used to identify a particular physical object or number of computer representations, such as methods and systems implemented by a computer and / or one or more steps and / or modules or functions thereof that access, display, serve as input, generate as output, etc. For example, as described further in detail herein, various computer-implemented systems and methods can operate, process, and generate peptide models representing physical peptides (such as a particular protein or a portion thereof). Computer representations of peptides, such as peptide models, can be implemented in a variety of formats, such as strings representing amino acid sequences (e.g., letters each representing a particular type of amino acid) (e.g., FASTA files), or 3D structural models that include (e.g., relative) 3D positional information about amino acids and / or their atoms, such as protein database (PDB) file formats, which can be used to describe the 3D structure of a particular protein and, in particular, include the atomic coordinates of the atoms of the particular protein.
[0192] Nucleic Acids: As used herein, the term "nucleic acid" in its broadest sense refers to any compound and / or substance incorporated into or potentially incorporated into an oligonucleotide chain. In some embodiments, nucleic acids are compounds and / or substances incorporated into or potentially incorporated into an oligonucleotide chain via phosphodiester bonds. As will become clear from the context, in some embodiments, "nucleic acid" refers to a single nucleic acid residue (e.g., a nucleotide and / or nucleoside); in some embodiments, "nucleic acid" refers to an oligonucleotide chain containing a single nucleic acid residue. In some embodiments, "nucleic acid" is or comprises RNA; in some embodiments, "nucleic acid" is or comprises DNA. In some embodiments, a nucleic acid is one or more native nucleic acid residues, comprises one or more native nucleic acid residues, or is composed of one or more native nucleic acid residues. In some embodiments, a nucleic acid is one or more nucleic acid analogs, comprises or is composed of said analogs. In some embodiments, a nucleic acid analog differs from a nucleic acid in that the nucleic acid analog does not utilize a phosphodiester backbone. For example, in some embodiments, a nucleic acid is one or more "peptide nucleic acids," comprises or is composed of said peptide nucleic acids, said peptide nucleic acids being known in the art and having peptide bonds rather than phosphodiester bonds in the backbone, and is considered to be within the scope of this disclosure. Alternatively or additionally, in some embodiments, the nucleic acid has one or more thiophosphate and / or 5'-N-phosphamide bonds instead of phosphodiester bonds. In some embodiments, the nucleic acid is one or more natural nucleosides (e.g., adenosine, thymidine, guanosine, cytidine, uridine, deoxyadenosine, deoxythymidine, deoxyguanosine, and deoxycytidine), comprising or composed of one or more natural nucleosides. In some embodiments, the nucleic acid is one or more nucleoside analogs (e.g., 2-aminoadenosine, 2-thiothymidine, inosine, pyrrolopyrimidine, 3-methyladenosine, 5-methylcytidine, C-5-propynylcytidine, C-5-propynyluridine, 2-aminoadenosine, C5-bromouridine, C5-fluorouridine, C5-iodouridine, C5-propynyluridine, C5-propynylcytidine, C5-methylcytidine, 2-aminoadenosine, 7-deazoadenosine, 7-deazoguanosine, 8-oxoadenosine, 8-oxoguanosine, O(6)-methylguanine, 2-thiocytidine, methylated bases, intercalated bases, and combinations thereof), comprising or consisting of one or more nucleoside analogs. In some embodiments, the nucleic acid comprises one or more modified sugars (e.g., 2'-fluororibose, ribose, 2'-deoxyribose, arabinose, and hexose) compared to the sugars in natural nucleic acids. In some embodiments, the nucleic acid has a nucleotide sequence encoding a functional gene product (such as RNA or protein). In some embodiments, the nucleic acid contains one or more introns.In some embodiments, nucleic acids are prepared by one or more of the following methods: isolation from natural sources, enzymatic synthesis (in vivo or in vitro) via complementary template-based polymerization, replication in recombinant cells or systems, and chemical synthesis. In some embodiments, the length of the nucleic acid is at least 3, 4, 5, 6, 7, 8, 9, 10, 15, 20, 25, 30, 35, 40, 45, 50, 55, 60, 65, 70, 75, 80, 85, 90, 95, 100, 110, 120, 130, 140, 150, 160, 170, or 18. 0, 190, 20, 225, 250, 275, 300, 325, 350, 375, 400, 425, 450, 475, 500, 600, 700, 800, 900, 1000, 1500, 2000, 2500, 3000, 3500, 4000, 4500, 5000, or more residues. In some embodiments, the nucleic acid is partially or entirely single-stranded; in some embodiments, the nucleic acid is partially or entirely double-stranded. In some embodiments, the nucleic acid has a nucleotide sequence comprising at least one element that encodes a polypeptide or is complementary to a sequence encoding a polypeptide. In some embodiments, the nucleic acid has enzymatic activity.
[0193] Pareto score: As used herein, the term "Pareto score" refers to a measure of variant performance evaluated relative to / as described herein as one or more scoring metrics (such as various scoring metrics and / or combinations thereof). These may include, but are not limited to, scores evaluating fitness and the ability to evade immune responses. In some embodiments, the Pareto score includes a combination of immune evasion scores (e.g., as described herein) and fitness prior scores (e.g., as described herein). In some embodiments, the Pareto score reflects the relative evolutionary advantage of a given strain. In some embodiments, this Pareto score may be determined as described in the examples. In some embodiments, the Pareto score is an optimal score, which, for example, in some embodiments, ordinates the variant relative to other sequences (e.g., sequences observed in the population). A higher Pareto score for a particular lineage at a particular time indicates that fewer variants at that time have higher fitness prior and immune evasion scores. Because the Pareto score in some embodiments is an ordination system, and the fitness prior and immune evasion scores included therein change as new data becomes available, the Pareto score for a given variant will change over time. As used herein, in some embodiments, Pareto optimality is defined over a set of lineages. In some embodiments, a lineage is Pareto optimal within a set if no lineage in that set has a higher immune escape and a higher fitness prior score. In some embodiments, the Pareto score is a measure of the degree of Pareto optimality. For example, in some embodiments, the lineage with the highest Pareto score is Pareto optimal; and if the Pareto optimal lineage is removed from the set, the lineage with the second-best Pareto score will be Pareto optimal, and so on.
[0194] Patient: As used herein, the term "patient" means any organism to which the provided composition is applied or may be applied, for example, for experimental, diagnostic, preventive, cosmetic, and / or therapeutic purposes. Typical patients include animals (e.g., mammals such as mice, rats, rabbits, non-human primates, and / or humans). In some embodiments, the patient is a human. In some embodiments, the patient has or is susceptible to one or more diseases or disorders. In some embodiments, the patient exhibits one or more symptoms of a disease or disorder. In some embodiments, the patient has been diagnosed with one or more diseases or disorders. In some embodiments, the disease or disorder is or comprises a viral infection (e.g., SARS-CoV-2 infection). In some embodiments, the patient is receiving or has received some form of therapy to diagnose and / or treat a disease, condition, or disorder.
[0195] Peptide: As used in this article, the term “peptide” refers to polypeptides that are generally relatively short, such as those with a length of less than about 100 amino acids, less than about 50 amino acids, less than about 40 amino acids, less than about 30 amino acids, less than about 25 amino acids, less than about 20 amino acids, less than about 15 amino acids, or less than about 10 amino acids.
[0196] Pharmaceutical composition: As used herein, the term "pharmaceutical composition" refers to an active agent formulated together with one or more pharmaceutically acceptable carriers. In some embodiments, the active agent is present in an amount suitable for a unit dose administered in a treatment regimen that, when administered to a relevant population, shows a statistically significant probability of achieving the intended therapeutic effect. In some embodiments, the pharmaceutical composition may be specifically formulated for administration in solid or liquid form, including pharmaceutical compositions suitable for: oral administration, such as a drench (aqueous or non-aqueous solution or suspension), tablets (e.g., those for buccal, sublingual, and systemic absorption), pills, powders, granules, or pastes for application to the tongue; parenteral administration, such as by subcutaneous, intramuscular, intravenous, or epidural injection, as, for example, a sterile solution or suspension, or a sustained-release formulation; topical application, such as as a cream, ointment, controlled-release patch, or spray to the skin, lungs, or mouth; intravaginal or rectal administration, such as as a vaginal suppository, cream, or foam; sublingual; ocular; transdermal; or via the nose, lungs, and other mucosal surfaces.
[0197] Pharmaceutically acceptable: As used herein, the term “pharmaceutically acceptable” applies to carriers, diluents, or excipients used to formulate the compositions disclosed herein, meaning that the carrier, diluent, or excipient must be compatible with the other components of the composition and harmless to its recipients.
[0198] Pharmaceutically acceptable carriers: As used herein, the term "pharmaceutically acceptable carrier" means a pharmaceutically acceptable material, composition, or medium, such as a liquid or solid filler, diluent, excipient, or solvent encapsulating material, that participates in carrying or transporting the subject compound from one organ or part of the body to another. Each carrier must be "acceptable" in the sense that it is compatible with other components of the formulation and harmless to the patient. Some examples of materials that can be used as pharmaceutically acceptable carriers include: sugars, such as lactose, glucose, and sucrose; starches, such as corn starch and potato starch; cellulose and its derivatives, such as sodium carboxymethyl cellulose, ethyl cellulose, and cellulose acetate; powdered tragacanth gum; malt; gelatin; talc; excipients, such as cocoa butter and suppository waxes; oils, such as peanut oil, cottonseed oil, safflower oil, sesame oil, olive oil, corn oil, and soybean oil; diols, such as propylene glycol; polyols, such as glycerol, sorbitol, mannitol, and polyethylene glycol; esters, such as ethyl oleate and ethyl laurate; agar; buffers, such as magnesium hydroxide and aluminum hydroxide; alginic acid; pyrogen-free water; isotonic saline; Ringer's solution; ethanol; pH buffer solutions; polyesters, polycarbonates, and / or polyanhydrides; and other non-toxic and compatible substances used in pharmaceutical formulations.
[0199] Pharmaceutical grade: As used herein, “pharmaceutical grade” refers to the standards established by recognized national or regional pharmacopoeias (e.g., the United States Pharmacopeia and Formulary (USP–NF)) for chemical and biological pharmaceutical substances, pharmaceutical products, dosage forms, compound preparations, excipients, medical devices, and dietary supplements.
[0200] Polypeptide: As used herein, refers to a polymeric chain of amino acids. In some embodiments, the polypeptide has a naturally occurring amino acid sequence. In some embodiments, the polypeptide has an amino acid sequence that is not naturally occurring. In some embodiments, the polypeptide has an engineered amino acid sequence because the sequence is designed and / or generated artificially. In some embodiments, the polypeptide may contain natural amino acids, non-natural amino acids, or both, or consist of them. In some embodiments, the polypeptide may contain only natural amino acids or only non-natural amino acids, or consist of them. In some embodiments, the polypeptide may contain D-amino acids, L-amino acids, or both. In some embodiments, the polypeptide may contain only D-amino acids. In some embodiments, the polypeptide may contain only L-amino acids. In some embodiments, the polypeptide may include one or more side groups or other modifications, such as modifications or attachments to one or more amino acid side chains at the N-terminus, C-terminus, or any combination thereof. In some embodiments, such side groups or modifications may be selected from the group consisting of: acetylation, amidation, esterification, methylation, polyethylene glycolation, etc., including combinations thereof. In some embodiments, the polypeptide may be cyclic, and / or may contain a cyclic moiety. In some embodiments, the polypeptide is not cyclic and / or does not contain any cyclic portion. In some embodiments, the polypeptide is linear. In some embodiments, the polypeptide may be or comprise a bound polypeptide. In some embodiments, the term "polypeptide" may be attached to the name, activity, or structure of a reference polypeptide; in such cases, it is used herein to refer to polypeptides that share a common related activity or structure and can therefore be considered members of the same class or family of polypeptides. For each such class, this specification provides and / or those skilled in the art will know of exemplary polypeptides with known amino acid sequences and / or functions within the class; in some embodiments, such exemplary polypeptides are reference polypeptides of the polypeptide class or family. In some embodiments, members of a polypeptide class or family exhibit significant sequence homology or identity with a reference polypeptide of the class, share common sequence motifs (e.g., characteristic sequence elements) with a reference polypeptide of the class, and / or share common activity (in some embodiments at comparable levels or within a specified range) with a reference polypeptide of the class; in some embodiments, this applies to all polypeptides within the class. For example, in some embodiments, the member polypeptide exhibits at least about 30-40% homology or identity with the overall sequence of the reference polypeptide, and typically greater than about 50%, 60%, 70%, 80%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, 98%, 99% or more, and / or includes at least one region (e.g., in some embodiments, this may be or contain a conserved region of characteristic sequence elements) that exhibits very high sequence identity, typically greater than 90% or even 95%, 96%, 97%, 98% or 99%.Such conserved regions typically encompass at least 3-4, and usually up to 20 or more, amino acids; in some embodiments, the conserved region encompasses at least one segment of at least 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15 or more consecutive amino acids. In some embodiments, the relevant polypeptide may comprise or consist of fragments of the parent polypeptide. In some embodiments, the useful polypeptide may comprise or consist of multiple fragments, each of which is found in the same parent polypeptide in a spatial arrangement relative to each other, rather than in the polypeptide of interest (e.g., fragments directly linked in the parent may be spatially separated in the polypeptide of interest, and vice versa, and / or the order in which the fragments are present in the polypeptide of interest may differ from that in the parent), such that the polypeptide of interest is a derivative of its parent polypeptide.
[0201] Prevention (or prevention): As used herein, when used in conjunction with the occurrence of a disease, condition, and / or disorder, refers to reducing the risk of developing a disease, condition, and / or disorder and / or delaying the onset of one or more characteristics or symptoms of a disease, condition, or disorder. Prevention is considered complete when the onset of a disease, condition, or disorder has been delayed by a predetermined period of time.
[0202] Ribonucleotides: As used herein, the term "ribonucleotide" encompasses both unmodified and modified ribonucleotides. For example, unmodified ribonucleotides include purine bases adenine (A) and guanine (G) and pyrimidine bases cytosine (C) and uracil (U). Modified ribonucleotides may include one or more modifications, including but not limited to, for example, (a) end modifications, such as 5' end modifications (e.g., phosphorylation, dephosphorylation, conjugation, reverse bonding, etc.), 3' end modifications (e.g., conjugation, reverse bonding, etc.), (b) base modifications, such as substitution with a modified base, a stabilizing base, a destabilizing base, or a base paired with or conjugated with an extended partner library base, (c) sugar modifications (e.g., at the 2' or 4' position) or sugar substitutions, and (d) internucleotide linking modifications, including modifications or substitutions of phosphodiester links. The term "ribonucleotide" also covers ribonucleotide triphosphates, including both modified and unmodified ribonucleotide triphosphates.
[0203] Ribonucleic acid (RNA): As used herein, the term "RNA" refers to a polymer of ribonucleotides. In some embodiments, the RNA is single-stranded. In some embodiments, the RNA is double-stranded. In some embodiments, the RNA comprises both single-stranded and double-stranded portions. In some embodiments, the RNA may comprise a backbone structure as described in the definition of "nucleic acid / polynucleotide" above. The RNA may be regulatory RNA (e.g., siRNA, microRNA, etc.) or messenger RNA (mRNA). In some embodiments where the RNA is mRNA. In some embodiments where the RNA is mRNA, the RNA typically includes a poly(A) region at its 3' end. In some embodiments where the RNA is mRNA, the RNA typically includes a cap structure recognized in the art at its 5' end, for example, for recognizing the mRNA and linking it to a ribosome to initiate translation. In some embodiments, the RNA is synthetic RNA. Synthetic RNA includes RNA synthesized in vitro (e.g., by enzymatic synthesis and / or by chemical synthesis). In some embodiments, the RNA is single-stranded RNA. In some embodiments, the single-stranded RNA may contain self-complementary elements and / or be capable of establishing secondary and / or tertiary structures. Those skilled in the art will understand that when a single-stranded RNA is referred to as "coding," it can mean that the single-stranded RNA contains a nucleic acid sequence that it encodes or that the single-stranded RNA contains a complementary sequence to a encoded nucleic acid sequence. In some embodiments, the single-stranded RNA may be self-amplifying RNA (also known as self-replicating RNA).
[0204] Semantic variation: As used herein, the term "semantic variation" refers to a measure of functional variation of a viral peptide variant (e.g., in some embodiments, a viral peptide that interacts with host cell receptors and / or otherwise participates in host cell entry) from a linguistic model perspective relative to at least one or more (e.g., at least two, at least three, at least four or more) reference viral peptides (e.g., in some embodiments, wild-type species and / or known variants, e.g., reference viral peptides of the same lineage). In some embodiments, semantic variation is a measure of functional variation of a viral peptide variant (e.g., in some embodiments, a viral peptide that interacts with host cell receptors and / or otherwise participates in host cell entry) from a linguistic model perspective relative to multiple (e.g., at least two or more) reference viral peptides (e.g., in some embodiments, wild-type species and / or known variants, e.g., reference viral peptides of the same lineage). In some embodiments, the relevant language model may include transformer-derived embedding differences (e.g., as described herein) for at least one or more (e.g., at least two, at least three, at least four or more) reference viral peptides (e.g., in some embodiments, wild-type species or known variants, e.g., reference viral peptides of the same lineage). In some embodiments, the LI norm may be used to compute the semantic change score. In some embodiments, the L2 norm may be used to compute the semantic change score (also known as the Euclidean norm). In some embodiments, the semantic change describes the difference of the variant relative to the underlying statistical model (e.g., in some embodiments, a large machine learning model fine-tuned on the observed viral protein sequence at a given time point). In some embodiments, the semantic change score depends on the observed sequence, and therefore the semantic change score may change over time as the base model is trained on new variant sequences and / or reference sequences. In some embodiments, the semantic change score is determined for variant spike peptides of SARS-Co-V-2, as described herein. In the various embodiments described herein, the semantic change score may be normalized such that its order is between 0 and 100%.
[0205] Subject: As used herein, the term "subject" refers to an organism, typically a mammal (e.g., a human, including prenatal human forms in some embodiments). In some embodiments, the subject suffers from a relevant disease, condition, or disorder. In some embodiments, the subject is susceptible to a disease, condition, or disorder. In some embodiments, the subject exhibits one or more symptoms or characteristics of a disease, condition, or disorder. In some embodiments, the subject does not exhibit any symptoms or characteristics of a disease, condition, or disorder. In some embodiments, the subject is a person who has one or more characteristics that characterize susceptibility or risk to a disease, condition, or disorder. In some embodiments, the subject is a patient. In some embodiments, the subject is an individual who has received and / or has received a diagnosis and / or treatment.
[0206] Essentially: As used herein, the term “essentially” refers to a qualitative condition that exhibits all or nearly all of the target characteristics or properties of a particular range or degree. Those skilled in the art of biology will understand that biological and chemical phenomena rarely (if at all) complete and / or proceed to completeness or achieve or avoid absolute results. Therefore, the term “essentially” is used herein to reflect the inherent potential lack of completeness in many biological and chemical phenomena.
[0207] Variants: As used herein in the context of molecules (e.g., nucleic acids, proteins, or small molecules), the term "variant" refers to a molecule that exhibits significant structural identity with a reference molecule but is structurally different from it, for example, by the presence or absence of one or more chemical parts or their different levels compared to the reference entity. In some embodiments, variants are also functionally different from their reference molecule. Generally, whether a particular molecule is properly considered a "variant" of a reference molecule is based on the degree of its structural identity with the reference molecule. As those skilled in the art will understand, any biological or chemical reference molecule has certain characteristic structural elements. A variant is defined as a different molecule that shares one or more of these characteristic structural elements with a reference molecule but is different in at least one respect. In some embodiments, a variant polypeptide or nucleic acid may differ from a reference polypeptide or nucleic acid due to one or more differences in amino acid or nucleotide sequences and / or one or more differences in chemical parts (e.g., carbohydrates, lipids, phosphate groups) that are covalent components of a polypeptide or nucleic acid (e.g., linked to the polypeptide or nucleic acid backbone). In some embodiments, the variant peptide or nucleic acid exhibits at least 85%, 86%, 87%, 88%, 89%, 90%, 91%, 92%, 93%, 94%, 95%, 96%, 97%, or 99% overall sequence identity with the reference peptide or nucleic acid. In some embodiments, the variant peptide or nucleic acid does not share at least one characteristic sequence element with the reference peptide or nucleic acid. In some embodiments, the reference peptide or nucleic acid has one or more biological activities. In some embodiments, the variant peptide or nucleic acid shares one or more biological activities with the reference peptide or nucleic acid. In some embodiments, the variant peptide or nucleic acid lacks one or more biological activities of the reference peptide or nucleic acid. In some embodiments, the variant peptide or nucleic acid exhibits a reduced level of one or more biological activities compared to the reference peptide or nucleic acid. In some embodiments, a peptide or nucleic acid of interest is considered a “variant” of the reference peptide or nucleic acid if it has the same amino acid or nucleotide sequence as the reference, except for minor sequence changes at specific positions. Typically, compared to the reference, the variant contains fewer than about 20%, about 15%, about 10%, about 9%, about 8%, about 7%, about 6%, about 5%, about 4%, about 3%, or about 2% of residues that are substituted, inserted, or deleted. In some embodiments, the variant polypeptide or nucleic acid contains about 10, about 9, about 8, about 7, about 6, about 5, about 4, about 3, about 2, or about 1 substituted residue compared to the reference. Typically, the variant polypeptide or nucleic acid contains a very small number (e.g., fewer than about 5, about 4, about 3, about 2, or about 1) of substituted, inserted, or deleted functional residues (i.e., residues involved in a specific biological activity) relative to the reference. In some embodiments, the variant polypeptide or nucleic acid contains no more than about 5, about 4, about 3, about 2, or about 1 added or deleted residues compared to the reference, and in some embodiments, no added or deleted residues are included.In some embodiments, the variant polypeptide or nucleic acid contains fewer than about 25, about 20, about 19, about 18, about 17, about 16, about 15, about 14, about 13, about 10, about 9, about 8, about 7, about 6, and typically fewer than about 5, about 4, about 3, or about 2 additions or deletions compared to the reference. In some embodiments, the reference polypeptide or nucleic acid is naturally occurring.
[0208] Vaccination: As used herein, the term "vaccination" refers to the administration of a composition designed to produce an immune response, for example, against a disease (e.g., a viral epitope). In some embodiments, vaccination may be administered before, during, and / or after the development of a disease. In some embodiments, vaccination comprises administering the vaccination composition multiple times at appropriate time intervals.
[0209] Viral peptide receptor binding score: As used herein, the term "viral peptide receptor binding score" refers to a measure of the binding affinity between a viral peptide that plays a role in host recognition and / or host cell entry and a corresponding host protein that interacts with the viral peptide to recognize and / or enter the host cell. In some embodiments, the viral peptide receptor binding score is determined by computer simulation. In some embodiments, the viral peptide receptor binding score can be determined using a conformational sampling algorithm. In some embodiments, the viral peptide receptor binding score can be determined using a structure optimized using a probabilistic optimization algorithm (e.g., in some embodiments, a variant of simulated annealing is designed to overcome local energy barriers and follow kinetically accessible paths toward a knowledge-based, protein-directed potential to reach an achievable depth energy minimum). In some embodiments, the viral peptide receptor binding score can be used to calculate the change in the solvent accessible surface area (SASA) of the viral peptide in a complexed state (e.g., bound state) and a non-complexed state (e.g., unbound state). In some embodiments, the viral peptide receptor binding score can be determined by calculating the energy changes of the complexed (e.g., bound) and non-complexed (e.g., unbound) structures of the viral peptide and its homologous host receptor. In some embodiments, the change in binding energy can be estimated by the difference in Gibbs free energy between the bound and unbound states. In the various embodiments described herein, the viral peptide receptor binding score can be normalized such that it is ranked between 0 and 100%. In some embodiments, the viral peptide receptor binding score can be calculated by computer simulation, for example, by calculating the change in Gibbs free energy or the change in solvent-accessible surface area between the bound and unbound states. In some embodiments, the viral peptide receptor binding score can be calculated using in vitro binding data (e.g., using the dissociation constant KD or binding rate k0n). In some embodiments, such in vitro binding data can be determined by methods known in the art, including but not limited to, biolayer interferometry (BLI) and / or surface plasmon resonance (SPR).
[0210] ACE2 Binding Score: As used herein, the term "ACE2 binding score" refers to a viral peptide receptor binding score (as described herein), where the viral peptide receptor is angiotensin-converting enzyme 2 (ACE2). The "ACE2 binding score" is a measure of the binding affinity between the S protein of a coronavirus (e.g., SARS-CoV-2) or an immunogenic fragment of the S protein (e.g., the RBD domain) and the ACE2 protein. In some embodiments, the ACE2 binding score can be calculated using computer simulations, for example, by calculating changes in Gibbs free energy or changes in solvent-accessible surface area in bound and unbound states. In some embodiments, the ACE2 binding score can use in vitro binding data (e.g., using the dissociation constant KD or binding rate k).on The data can be calculated using methods known in the art, including but not limited to biolayer interferometry (BLI) and / or surface plasmon resonance (SPR). In some embodiments, such in vitro binding data can be determined using methods known in the art, including but not limited to biolayer interferometry (BLI) and / or surface plasmon resonance (SPR).
[0211] Wild-type: As used herein, the term "wild-type" has its meaning as understood in the art, referring to an entity having the structure and / or activity found in nature in a "normal" state or context (as opposed to mutation, disease, alteration, etc.). Those skilled in the art will understand that wild-type genes and polypeptides often exist in many different forms (e.g., alleles). Detailed Implementation
[0212] It is conceivable that the systems and methods for which protection is claimed encompass variations and modifications developed using information from the embodiments described herein. Adjustments and / or modifications can be made to the systems, architectures, apparatuses, methods, and processes described herein, in accordance with the intent of this specification.
[0213] Throughout this specification, when articles, apparatuses, systems, and architectures are described as having, including, or containing specific components, or when processes and methods are described as having, including, or containing specific steps, it is also contemplated that there are articles, apparatuses, systems, and architectures of the invention that are substantially composed of or comprised of the listed components, and that there are processes and methods of the invention that are substantially composed of or comprised of the listed processing steps.
[0214] It should be understood that the order of steps or the sequence of certain operations is not important as long as the invention remains operational. Furthermore, two or more steps or actions may be performed simultaneously.
[0215] Any publications mentioned herein (e.g., in the background section) do not acknowledge that such publications are prior art to any of the claims set forth herein. The background section is provided for clarity and is not intended to describe prior art to any of the claims.
[0216] As stated above, this document is incorporated herein by reference. In the event of any discrepancy in the meaning of a term, the meaning provided in the definitions section above shall prevail.
[0217] A title is provided for the convenience of the reader—the existence and / or placement of a title is not intended to limit the scope of the subject matter described herein.
[0218] Among other things, this disclosure provides systems and methods for designing engineered antigens with customized immunological characteristics. In some embodiments, the engineered antigen created by the techniques described herein is a computer-engineered version of a reference antigen designed to interact with the immune system of a subject in a specific manner (e.g., a desired manner).
[0219] For example, in some embodiments, the reference antigen may be a naturally occurring protein, such as a variant of a specific viral protein, or a portion thereof. The antigen engineering techniques described further in detail herein can be used to design customized versions of such reference antigens to promote the generation of novel antibodies and reduce the likelihood and / or extent of triggering a memory immune response, which may result in the generation of antibodies from memory B cells that have been previously exposed to a variant of the reference antigen. Thus, novel antibodies generated in this manner may be tailored to a specific version of an epitope (e.g., a mutation) present in, for example, the reference antigen, while antibodies generated by a memory response may effectively target similar epitopes on previous variants but be evaded by newly emerging epitopes on a specific reference antigen.
[0220] Without being bound by any particular theory, we propose that vaccination may require antigenic peptides designed to promote an immune response, particularly an antibody response (e.g., a neutralizing antibody response), to epitopes appearing in variant peptides. Among other things, this disclosure provides the insight that, particularly for circulating infectious diseases (e.g., those where variants are expected), there may be a particular need to promote an immune response, specifically including an antibody response (e.g., a neutralizing response) to appearing epitopes.
[0221] Among other things, this disclosure provides techniques for engineering antigens by identifying and mutating (e.g., introducing amino acid modifications) portions of a reference antigen (e.g., sequences, specific sets of amino acid sites identified as members of a recognized surface region), said antigens having a reduced risk of triggering memory immune responses, as described herein, said reference antigens being identified as potentially comprising one or more shared epitopes and / or (e.g., continuous) conserved surfaces. As further described herein in detail, the specific methods described herein include identifying and establishing various criteria relating to specific distributions of amino acid modifications, sources and rules for generating amino acid modifications, and methods for evaluating the expected performance of engineered antigen designs designed to sufficiently disrupt the memory-triggering portions of the input reference antigen while retaining characteristics such as stability, 3D structure (e.g., folding), etc.
[0222] This disclosure illustrates certain aspects of the provided technology by way of designing engineered versions of recently evolved XBB SARS-Cov2 spike protein variants. While some examples provided herein are described with reference to specific proteins and viral variants, those skilled in the art will understand upon reading this disclosure that the methods described herein can be applied and employed, for example, to other pathogens (e.g., virus types), proteins, subregions, etc., to promote the generation of a desired immune response.
[0223] A. Natural evolution of infectious pathogens
[0224] Immunoblotting is a phenomenon in which initial exposure to a specific antigen can limit (e.g., subsequently) the development of an immune response specific to epitopes of a novel variant of that antigen. Specifically, when exposed to a new, previously unencountered infectious pathogen (such as a virus), the immune system responds, in particular, by producing antibodies that bind with and neutralize portions of the pathogen's antigen with high specificity. The immune system then retains a "memory" of the antigen in the form of memory B cells and T cells, along with the ability to produce specific antibodies targeting that antigen.
[0225] On the one hand, after initial exposure to a specific pathogen, this immune memory enables the body to rapidly recognize and defend against it upon subsequent encounters. On the other hand, processes such as natural mutation and evolution can produce variants sufficiently similar to the initially encountered strain to be recognized and trigger a memory response, leading to the production of antibodies designed to defend against the original strain, rather than antibodies specifically tailored to the new variant. The effectiveness of this memory response is diminished if the new variant contains sufficient mutations in key regions targeted by these antibodies (e.g., those related to host cell infection, viral replication, etc.). Therefore, immunoblotting is particularly concerning for pathogens with high concentrations of mutations on neutralizing sensitive epitopes.
[0226] In the context of viral infection and vaccination, this immunoblotting phenomenon can lead to increased infection rates or reinfection with variant strains and limit the effectiveness of vaccination after initial exposure to early strains (e.g., whether due to natural infection or early vaccine doses). Therefore, immunoblotting is particularly problematic for vaccination against viruses with high mutation rates, such as RNA viruses. These include, but are not limited to, influenza, coronaviruses (e.g., SARS-CoV-2), human immunodeficiency virus (HIV), and respiratory syncytial virus (RSV).
[0227] For example, in the context of the recent SARS-CoV-2 pandemic, mutations in circulating viruses have generated tens of thousands of viral variants, some of which, such as Omicron and the more recent XBB (e.g., XBB.1.5), are characterized by their immune evasion potential. Specifically, these variants contain a variety of mutations that enable them to evade existing (e.g., memory) immune responses formed in individuals due to prior exposure (through vaccination and / or natural infection) to previous strains (such as the original wild-type (WT) variant).
[0228] Among other things, this disclosure recognizes that, for example, XBB comprises mutations across multiple epitopes. Many of these mutations are the same as those of other earlier variants, but there is a subgroup that is unique to XBB. When these shared, pre-existing epitopes trigger a memory immune response, antibodies tailored to the earlier variants to which an individual has been exposed are generated instead of novel antibodies specifically designed to neutralize XBB.
[0229] Therefore, without being bound by any particular theory, the systems and methods described herein include methods for designing engineered antigens that can reduce the degree of memory responses triggered by shared epitopes of neoantigens and promote novel responses of the immune system to specific target epitopes unique to the neoantigen.
[0230] B. Engineered antigens
[0231] Go to Figure 1 Among other things, this disclosure provides systems and methods for the computer-simulated design of engineered antigens designed to promote an immune response to specific epitopes of a reference antigen of an infectious pathogen. Methods for engineered antigens are generally described herein with reference to viruses and their proteins, but engineered antigens may also be designed based on reference antigens derived from / related to other types of infectious pathogens, such as bacteria, parasites (e.g., malaria), etc.
[0232] i. Reference antigen and its computer representation
[0233] A reference antigen can be a protein of an infectious pathogen (e.g., a protein of an infectious pathogen or a protein produced by an infectious pathogen) and / or a portion thereof. For example, in some embodiments, the reference antigen is a specific viral protein and / or a portion thereof, such as a surface protein. In some embodiments, the reference antigen is a protein or a portion of a protein that has been identified as being used by the virus to infect host cells, for example, involving binding to specific host cell proteins and / or promoting fusion with the host cell membrane. In some embodiments, the reference antigen is or contains a specific portion—e.g., a subunit, domain, etc., of a viral protein.
[0234] For example, in some embodiments, the infectious pathogen is or comprises a coronavirus, such as SARS-CoV 2, and its reference antigen is the SARS-CoV 2 spike protein or a portion thereof. For example, the reference antigen may be the entire spike protein or a specific portion, such as an N-terminal region (e.g., an N-terminal domain (NTD)) or a receptor-binding domain (RBD). In some embodiments, specific portions are selected to focus on the least relevant vaccine antigen, for example, to facilitate the removal of as many conserved epitopes as possible without, for example, introducing point mutations (e.g., thereby limiting the number of epitopes to which point mutations may be introduced).
[0235] As described herein, in some embodiments, the antigen engineering techniques of this disclosure are designed to engineer a reference antigen that, when introduced (e.g., administered) to a subject, promotes the production of novel antibodies specifically tailored to a particular epitope of the reference antigen by the subject's immune system. In some embodiments, this involves reducing the likelihood and / or extent to which the engineered version of the reference antigen triggers a memory immune response, thereby mitigating certain obstacles that immunoblotting may (as described herein) pose in, for example, vaccination.
[0236] Therefore, in some implementations, the systems and methods described herein identify those remaining portions of a particular reference antigen that may trigger a memory immune response, and computer simulations generate one or more engineered variants in which these memory-triggering subregions are disrupted.
[0237] Figure 1 An example process 100 for generating engineered antigens according to certain embodiments is illustrated. In some embodiments, example process 100 begins by accessing, generating, or otherwise acquiring a computer representation of at least a portion of a reference antigen of interest, referred to herein as peptide model 102.
[0238] Various formats (e.g., data structures) can be used for computer representations of reference antigens (such as specific proteins and / or portions thereof). For example, peptide model 102 is or includes amino acid sequences, such as ordered strings of letters, each representing a specific amino acid (e.g., a FASTA file). In some embodiments, the peptide model may be or include a structural model, such as a 3D model including indications of the position of each amino acid (or atom thereof) of the reference antigen in 3D space. For example, protein database (PDB) format files can be used and / or accessed to obtain 3D structural information about a specific protein, such as based on derived crystallographic structures.
[0239] In some embodiments, the reference antigen may be a protein of a specific target viral variant, such as the spike (S) protein or a portion thereof of a specific target SARS-CoV-2 variant. As used herein, the full-length SARS-CoV-2 S protein containing the “wild-type” sequence has a sequence corresponding to the sequence of the first detected SARS-CoV-2 strain, consists of 1273 amino acids, and has the amino acid sequence according to SEQ ID NO:1:
[0240]
[0241] (SEQ ID NO:1)
[0242] Unless otherwise indicated, the position numbers given herein for the SARS-CoV-2S protein and / or portions thereof are relative to the amino acid sequence of SEQ ID NO:1. Those skilled in the art, upon reading this disclosure, will understand and be able to determine the corresponding position in a SARS-CoV-2S protein variant sequence based on the location provided relative to the amino acid sequence of SEQ ID NO:1 (i.e., those skilled in the art will be able to determine the corresponding position in the S protein sequence of another SARS-CoV-2 variant or fragment thereof based on the position provided relative to SEQ ID NO:1 or another variant). When a portion of the SARS-CoV-2S protein is described as having certain mutations, unless otherwise indicated, it should be understood that position numbers for the mutations are given, and the position is identified relative to the amino acid sequence of SEQ ID NO:1.
[0243] In specific embodiments, the spike (S) protein described herein can be modified in a manner that stabilizes the prefusion-stabilized SARS-CoV-2 spikes. Certain mutations that stabilize the prefusion-stabilized SARS-CoV-2 spikes are known in the art, for example, as disclosed in WO2021243122 A2 and Hsieh, Ching-Lin et al. ("Structure-based design of prefusion-stabilized SARS-CoV-2 spikes," Science 369.6510(2020):1501-1505), the contents of which are incorporated herein by reference in their entirety. In some embodiments, the SARS-CoV-2S protein can be stabilized by introducing one or more proline mutations. In some embodiments, the SARS-CoV-2S protein contains a proline substitution at positions corresponding to residues 986 and / or 987 of SEQ ID NO:1. In some embodiments, the SARS-CoV-2S protein contains a proline substitution at one or more positions corresponding to residues 817, 892, 899, and 942 of SEQ ID NO:1. In some embodiments, the SARS-CoV-2S protein contains a proline substitution at each of the positions corresponding to residues 817, 892, 899, and 942 of SEQ ID NO:1. In some embodiments, the SARS-CoV-2S protein contains a proline substitution at each of the positions corresponding to residues 817, 892, 899, 942, 986, and 987 of SEQ ID NO:1.
[0244] In some embodiments, stabilization of the prototype pre-fusion conformation of the SARS-CoV-2S protein can be achieved by introducing two consecutive proline substitutions at residues 986 and 987. Specifically, the spike (S) protein stabilized variant is obtained by exchanging the amino acid residue at position 986 for proline and also exchanging the amino acid residue at position 987 for proline. In one embodiment, the SARS-CoV-2S protein variant whose prototype pre-fusion conformation is stabilized comprises the amino acid sequence shown in SEQ ID NO:2:
[0245]
[0246] (SEQ ID NO:2)
[0247] Those skilled in the art are familiar with various SARS-CoV-2 spike mutants and / or resources documenting them. For example, the following strains, their SARS-CoV-2S protein amino acid sequences, and in particular their modifications compared to the wild-type SARS-CoV-2S protein amino acid sequence (e.g., compared to SEQ ID NO:1) may be used herein.
[0248] B.1.1.7 (“Variant of Concern 202012 / 01” (VOC-202012 / 01))
[0249] B.1.1.7 (“α variant”) is a SARS-CoV-2 variant first detected in the UK in October 2020 from samples collected last month and rapidly spreading by mid-December. It is associated with a significant increase in COVID-19 infection rates; this increase is thought to be at least partly due to changes in the N501Y domain within the receptor-binding domain of the spike glycoprotein, which is essential for binding to ACE2 in human cells. B.1.1.7 is defined by 23 mutations: 13 non-synonymous mutations, 4 deletions, and 6 synonymous mutations (i.e., 17 mutations alter the protein, while six do not). Spike protein variations in B.1.1.7 include deletions 69-70, deletion 144, N501Y, A570D, D614G, P681H, T716I, S982A, and D1118H.
[0250] B.1.351(501.V2)
[0251] The B.1.351 lineage (“β variant”), commonly known as the South African COVID-19 variant, exhibits increased transmissibility compared to the original wild-type strain. The B.1.351 variant is defined by several spike protein variations, including: L18F, D80A, D215G, deletions 242-244, R246I, K417N, E484K, N501Y, D614G, and A701V. Three mutations of particular concern are located in the spike regions of the B.1.351 genome: K417N, E484K, and N501Y.
[0252] B.1.1.298 (Cluster 5)
[0253] Virus B.1.1.298 was discovered in North Jutland, Denmark, and is believed to have been transmitted from mink to humans through mink farms. Several different mutations in the virus's spike protein have been identified. Specific mutations include deletions of 69-70, Y453F, D614G, I692V, M1229I, and the optional S1147L.
[0254] P.1(B.1.1.248)
[0255] Lineage B.1.1.248 (“γ variant”), known as the Brazilian variant, is one of the SARS-CoV-2 variants and is named the P.1 lineage. P.1 has multiple S protein modifications (L18F, T20N, P26S, D138Y, R190S, K417T, E484K, N501Y, D614G, H655Y, T1027I, V1176F) and is similar to the variant B.1.351 from South Africa at certain key RBD positions (K417, E484, N501).
[0256] B.1.427 / B.1.429(CAL.20C)
[0257] The lineage B.1.427 / B.1.429 (“ε variant”), also known as CAL.20C, is defined by the following modifications in the S protein: S13I, W152C, L452R, and D614G, with the L452R modification being of particular concern. The CDC has listed B.1.427 / B.1.429 as a “variant of concern”.
[0258] B.1.525
[0259] B.1.525 (“η variant”) carries the same E484K modification as in variants P.1 and B.1.351, and also carries the same ΔH69 / ΔV70 deletion as in variants B.1.1.7 and B.1.1.298. It also carries the modifications D614G, Q677H, and F888L.
[0260] B.1.526
[0261] B.1.526 (“ι variant”) was identified as a newly emerging viral isolate lineage in the New York area, sharing mutations with previously reported variants. The most common spike mutations in this lineage are L5F, T95I, D253G, E484K, D614G, and A701V.
[0262] B.1.1.529
[0263] B.1.529 (“Omicron variant”) was first discovered in South Africa in November 2021. Omicron replicates approximately 70 times faster than the delta variant and quickly became the dominant SARS-CoV-2 strain globally. Since its initial discovery, numerous Omicron sublineages have emerged. The Omicron variants of current interest are listed below, along with certain characteristic mutations associated with the S protein of each variant. The S proteins of BA.4 and BA.5 share the same group of characteristic mutations, which is why the table below has a single row “BA.4 or BA.5”, and why this disclosure refers to the “BA.4 / 5” S protein in some embodiments. Similarly, the S proteins of the BA.4.6 and BF.7 Omicron variants share the same group of characteristic mutations, which is why the table below has a single row “BA.4.6 or BF.7”.
[0264] Table 1A: Specific Omicron variants of interest and their characteristic mutations.
[0265]
[0266]
[0267] In some implementations, the SARS-CoV-2S protein described herein contains one or more mutations specific to a certain Omicron variant (including, for example, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20 or more) (e.g., one or more mutations of the Omicron variants listed in Table 1A, such as each of the mutations associated with a given XBB variant in Table 1A above).
[0268] As described elsewhere in this disclosure, in some embodiments, specific immunogenic portions (e.g., subregions) of the full-length coronavirus S protein (e.g., SARS-CoV-2 S protein) can be used as reference antigens for creating engineered antigens as described herein.
[0269] In some embodiments, the immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks certain features of the full-length polypeptide (e.g., features shown or predicted to interfere with inducing a primary immune response). For example, in some embodiments, the immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks regions having: (i) a low number or density of B-cell neutralizing epitopes and / or (ii) a high number or density of B-cell epitopes unrelated to neutralization. For example, in some embodiments, the immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks the complete S2 domain. In some embodiments, coronavirus (e.g., SARS-CoV-2) S proteins lacking the complete S2 domain lack an S2 region having: (i) a low number or density of neutralizing-related B-cell epitopes or (ii) a high number of B-cell epitopes unrelated to neutralization but retaining other parts of the S2 domain. In some embodiments, the immunogenic portion of a coronavirus (e.g., SARS-CoV-2) S protein lacks the entire S2 domain. In some embodiments, the immunogenic portion of the coronavirus (e.g., SARS-CoV-2) S protein lacks the complete S2 domain but contains certain sequences that can improve the immunogenicity and / or stability of the immunogenic portion (e.g., in some embodiments, the immunogenic portion lacks the complete S2 domain but retains the TM sequence).
[0270] Those skilled in the art who read this disclosure will be able to identify B-cell epitopes in the S protein of coronaviruses (e.g., SARS-CoV-2) and determine which epitopes are associated with or unrelated to neutralization. For example, those skilled in the art will recognize that many studies have already identified such regions using antibody binding studies (e.g., studies characterizing antibodies produced in subjects infected with SARS-CoV-2 or vaccinated against SARS-CoV-2).
[0271] In some embodiments, the immunogenic portion of the coronavirus (e.g., SARS-CoV-2) S protein includes certain regions identified as having a high number or density of neutralizing epitopes and optionally a high mutation rate. In some embodiments, the immunogenic portion of the coronavirus (e.g., SARS-CoV-2) S protein includes the N-terminal domain (NTD) of the S protein. In some embodiments, the immunogenic portion of the coronavirus (e.g., SARS-CoV-2) protein includes the receptor-binding domain (RBD) of the S protein. In some embodiments, the immunogenic portion of the coronavirus (e.g., SARS-CoV-2) S protein includes the S1 domain of the S protein.
[0272] In some implementations, the immunogenic portion of the coronavirus (e.g., SARS-CoV-2) S protein includes RBD and NTD, and other features of the S1 domain are omitted.
[0273] The coronavirus (e.g., SARS-CoV-2) S protein has been well characterized, and those skilled in the art will be able to determine which portions of the S protein sequence correspond to the immunogenic portions discussed herein (e.g., which portions of the S protein sequence correspond to the NTD, RBD, S1, and S2 domains). In some embodiments, the RBD of the coronavirus (e.g., SARS-CoV-2) S protein comprises residues 327 to 528 of SEQ ID NO:1 or the corresponding region.
[0274] In some implementations, the RBD of the coronavirus (e.g., SARS-CoV-2) S protein contains an amino acid sequence:
[0275] VRFPNITNLCPFHEVFNATTFASVYAWNRKRISNCVADYSVIYNFAPFFAFKCYGVSPTKLNDLCFTNVYADSFVIRGNEVSQIAPGQTGNIADYNYKLPDDFTGCVIAWNSNKLDSKPSGNYNYLYRLFRKSKLKPFERDISTEIYQAGNKPCNGVAGPNCYSPLQSYGFRPTYGVGHQPYRVVVLSFELLHAPATVCGPK(SEQ ID NO:3),
[0276] Or the corresponding area.
[0277] In some implementations, the RBD of the coronavirus (e.g., SARS-CoV-2) S protein contains an amino acid sequence:
[0278] VRFPNITNLCPFGEVFNATRFASVYAWNRKRISNCVADYSVLYNSASFSTFKCYGVSPTKLNDLCFTNVYADSFVIRGDEVRQIAPGQTGKIADYNYKLPDDFTGCVIAWNSNNLDSKVGGNYNYLYRLFRKSNLKPFERDISTEIYQAGSTPCNGVEGFNCYFPLQSYGFQPTNGVGYQPYRVVVLSFELLHAPATVCGPK(SEQ ID NO:4),
[0279] Or the corresponding area.
[0280] In some embodiments, the S1 domain of the coronavirus (e.g., SARS-CoV-2) S protein comprises amino acids 1 to 678 of SEQ ID NO:1, or the corresponding region in the S protein of a SARS-CoV-2 variant. In some embodiments, the S1 domain of the SARS-CoV-2 S protein comprises amino acids 1 to 683 of SEQ ID NO:1, or the corresponding region in the S protein of a SARS-CoV-2 variant. In some embodiments, the S1 domain of the SARS-CoV-2 S protein comprises amino acids 1 to 685 of SEQ ID NO:1, or the corresponding region in the S protein of a SARS-CoV-2 variant.
[0281] In some implementations, the S1 domain of the SARS-CoV-2S protein contains the amino acid sequence: MFVFLVLLPLVSSQCVNLITRTQSYTNSFTRGVYYPDKVF (SEQ ID NO:5).
[0282] In some implementations, the S1 domain of the SARS-CoV-2S protein contains the amino acid sequence: MFVFLVLLPLVSSQCVNLTTRTQLPPAYTNSFTRGVYYP (SEQ ID NO:6).
[0283] In some embodiments, the S2 domain of the SARS-CoV-2 S protein comprises amino acids 679 to 1273 of SEQ ID NO:1, or the corresponding region in the S protein of a SARS-CoV-2 variant. In some embodiments, the S1 domain of the SARS-CoV-2 S protein comprises amino acids 684 to 1273 of SEQ ID NO:1, or the corresponding region in the S protein of a SARS-CoV-2 variant. In some embodiments, the S1 domain of the SARS-CoV-2 S protein comprises amino acids 686 to 1273 of SEQ ID NO:1, or the corresponding region in the S protein of a SARS-CoV-2 variant.
[0284] In some embodiments, the compositions described herein deliver the immunogenic portion of the S protein of a SARS-CoV-2 variant. In some embodiments, the variant is the variant of interest (e.g., a variant that has been predicted and / or shown to be rapidly spreading within the relevant jurisdiction, e.g., as identified by certain public health agencies, such as the Centre for Disease Control and Prevention (CDC) of the UK, Public Health England and the UK COVID-19 Genomics Consortium, the Canadian COVID Genomics Network (CanCOGeN), and / or the World Health Organization (WHO)). In some embodiments, a variant has been predicted to have a high probability of becoming the variant of interest (e.g., using the ability of predicted variants to evade previously developed immune responses and / or sequence-based algorithms that measure the “fitness” of a given variant, such as those described in WO2022 / 235847 and WO2022 / 235853, the contents of which are incorporated herein by reference in their entirety).
[0285] In some implementations, the RBD contains mutations associated with the variants described herein. Those skilled in the art will be able to identify which parts of a given variant correspond to the immunogenic parts described herein.
[0286] In some embodiments, the peptide comprises two or more SARS-CoV-2 subdomains (e.g., two or more S1 domains or RBDs). In some embodiments, the peptide comprises two or more tandemly linked receptor-binding domains, as described, for example, in Dai, Lianpan et al., "A universal design of betacoronavirus vaccines against COVID-19, MERS, and SARS," Cell 182.3(2020):722-733; and Han, Yuxuan et al., "mRNA vaccines expressing homo-prototype / Omicron and hetero-chimeric RBD-dimers against SARS-CoV-2," Cell Research 32.11(2022):1022-1025, the contents of which are incorporated herein by reference in their entirety. In some embodiments, the two or more subdomains originate from the same SARS-CoV-2 variant (e.g., the variant described herein). In some implementations, at least two of the two or more subdomains come from different SARS-CoV-2 variants (e.g., from different variants of interest, different Omicron variants, Omicron variants and non-Omicron variants, or wild-type strains and Omicron variants).
[0287] ii. Significant mutations
[0288] In some embodiments, the methods described herein utilize information describing the unique characteristics of a particular reference antigen, such as the identification of signature mutations. Signature mutations are those used to distinguish a particular reference antigen (e.g., a protein) from other similar antigens. In some embodiments, for example, the reference antigen is a specific protein or part thereof of a particular viral variant, and signature mutations are those used to distinguish the particular viral variant and closely related variants from other variants, such as [other variants].
[0289] For example, in the context of SARS-CoV 2, signature mutations can be identified based on one or more classification schemes, such as the World Health Organization (WHO) classification, GISAID, Pango lineage, Nextstrain clade, etc.
[0290] For example, in some embodiments, the reference antigen may be the SARS-CoV 2 spike protein of a specific XBB variant. A marker mutation can then be identified as one that occurs at a percentage greater than a specific marker threshold across all sequences (e.g., within a specific database) classified as members of the XBB lineage according to WHO specifications. In some embodiments, the marker threshold is 50 percent or higher. In some embodiments, the marker threshold is 60 percent or higher. In some embodiments, the marker threshold is 70 percent or higher. In some embodiments, the marker threshold is 75 percent or higher.
[0291] Exemplary XBB.1.5 hallmark mutations (e.g., mutations relative to wild-type strains, identified, for example, by SEQ ID NO:1 and / or SEQ ID NO.2) that are identified as occurring in more than 50% of the sequences in variants classified by WHO as XBB.1.5 are shown in Table 1B below:
[0292] Table 1B.XBB.1.5 hallmark mutations.
[0293]
[0294] iii. Conservative areas that trigger memories
[0295] In some embodiments, one or more conserved regions 122 of the polypeptide model 102 that trigger memory are identified. The identified (memory-triggered) conserved regions 122 represent specific portions of a reference antigen (represented by the polypeptide model 102) that are identified as potentially triggering a memory immune response. The memory-triggered conserved regions can represent these specific portions of the reference antigen in a variety of ways. For example, the memory-triggered conserved regions may be or contain a list of amino acid positions, or alternatively, may be or contain identified 3D regions on the surface of the 3D polypeptide model.
[0296] Several methods, based on various criteria, can be used to determine whether a specific portion of the reference antigen is likely to trigger a memory immune response.
[0297] In some implementations, conserved regions that trigger memories can be identified and determined in a variety of ways, such as, but not limited to, using prior known or established data, such as epitope lists, structural models, combined with experimental techniques (e.g., combined assays) and computational methods, used alone or in combination.
[0298] conservative epitope
[0299] In some embodiments, the memory-triggering portion of the reference antigen is determined by identifying a set of conserved epitopes within the reference antigen that may trigger a memory immune response. A shared set of epitopes of the reference antigen can be identified by comparing an initial set of known epitopes with mutations present in the reference antigen. For example, in some embodiments, the reference antigen is a form of a specific protein, such as a protein of a specific viral variant. In some embodiments, a set of known epitopes includes epitopes that meet specific criteria, such as those identified as targets for antibody binding (e.g., binding epitopes) and / or neutralizing antibodies (e.g., neutralizing epitopes).
[0300] Known epitopes can be, for example, epitopes of a specific virus previously characterized (with a variant of the reference antigen), or epitopes to which antibodies have been identified as binding. In some embodiments, a known epitope is a binding site for a neutralizing epitope (e.g., a neutralizing B-cell epitope). A set of known epitopes may include binding epitopes and / or neutralizing epitopes. Data identifying a set of known epitopes may be, or may include, a list of amino acid positions for each known epitope in the set. Such data may be available from sources such as results of targeted experiments, proprietary datasets, and public information such as literature and publicly available datasets.
[0301] For example, data on epitopes of the SARS-CoV 2 spike protein can be found using CoV-AbDab, IEDB, and other databases.
[0302] The known epitope dataset can then be filtered and analyzed to determine which known epitopes in the reference antigen are mutated and which are not mutated and are therefore conserved. For example, data identifying a set of known epitopes can be compared to a list of signature mutations for a specific reference antigen. Epitopes in a group of known epitopes that contain / overlap one or more signature mutations of the reference antigen can be identified as mutated epitopes, and other epitopes in that group (e.g., epitopes that do not contain signature mutations) can be identified as conserved epitopes. In some implementations, additional criteria can be used to filter a set of known epitopes, such as identifying and removing epitopes that are subgroups of other known epitopes or known epitopes located or not located in specific regions (e.g., SARS-CoV2, RBD, and / or N-terminal regions).
[0303] Conservative surface
[0304] In some embodiments, identifying conserved regions that trigger memory includes recognizing conserved surfaces of peptide models representing a specific reference antigen. A conserved surface represents a portion of the reference antigen that is available and / or potentially binds to antibodies created by a memory immune response (e.g., memory B cells), but may not necessarily have been previously identified as a known epitope. For example, a conserved surface may represent those portions (e.g., amino acid sites) of a specific reference antigen that (i) are sufficiently surface-accessible and (ii) are determined to be sufficiently similar to other (e.g., related) antigens and / or unaffected by mutations present in the specific reference antigen, such that they can become binding targets for antibodies associated with a memory immune response.
[0305] For example, in some embodiments, amino acid sites located on a surface (e.g., sufficiently surface-accessible) can be identified (e.g., using peptide model 102). In some embodiments, a set of surface amino acids can be compared to a signature mutation of a reference antigen to filter out (e.g., remove from the set) amino acid sites that are sites of signature mutations and / or within a specific distance [e.g., in 3D space, a "straight-line distance"; or across one or more signature mutations on a 3D surface (e.g., geodesic distance)]. The remaining amino acid sites in the set can then be used to define a conserved surface. In some embodiments, the conserved surface can be further refined by evaluating the spatial distribution of the amino acid sites, for example, to exclude portions (from the conserved surface) determined to be too small for an antibody to bind to. For example, isolated plaques below a certain threshold size (e.g., surface area) can be removed from the conserved surface.
[0306] Conservative areas that trigger memories
[0307] In some embodiments, one or more memory-triggered conservative regions are or contain a set of conservative tabletops. In some embodiments, one or more memory-triggered subregions are or contain conservative surfaces. In some embodiments, a set of conservative tabletops and a conservative surface are included / identified as one or more memory-triggered subregions.
[0308] iv. Destroy the conservative area
[0309] Turn to Figure 1 After identifying one or more conserved regions of a reference antigen that trigger memory, the systems and methods of this disclosure are designed to disrupt these regions, for example, to mitigate or avoid (e.g., reduce the likelihood and / or extent of a trigger memory immune response). Specifically, in some embodiments, the methods described herein introduce (e.g., computer-simulated) amino acid modifications (such as substitution, insertion, deletion, etc.) into at least a portion of one or more identified conserved regions of a reference antigen that trigger memory.
[0310] Distribution Standard
[0311] In some embodiments, the method of introducing amino acid modifications into conserved regions is intended to distribute the amino acid modifications throughout the conserved region. For example, in some embodiments, one or more amino acid modifications are introduced into at least a portion of one or more conserved epitope regions. In some embodiments, one or more amino acid modifications are introduced into each of the one or more conserved epitope regions, for example, to disrupt each identified conserved epitope. This can be achieved, for example, by considering at least a portion (e.g., all) of the conserved epitope regions separately and introducing one or more amino acid modifications into each region (e.g., in a stepwise manner), or alternatively, by repeatedly introducing amino acid modifications randomly / according to various statistical selection criteria until each conserved epitope region is modified by at least one amino acid modification, or meets other stopping criteria.
[0312] In some embodiments, one or more amino acid modifications are introduced (e.g., across) a conserved surface. In some embodiments, multiple amino acid modifications are generated. In some embodiments, a certain number and location of amino acid modifications are introduced to create a specific desired spatial distribution of mutations. Design criteria may include, but are not limited to, density on a 3D surface representing a conserved region, minimum and / or maximum separation between amino acid modifications, and spatial separation on the 3D conserved surface. In some embodiments, the ideal distribution of mutations on a conserved surface can be achieved by calculating a positional expansion score, which measures the degree to which the positions of amino acid modifications are distributed on the conserved surface, e.g., uniform distribution, and the design is evaluated, for example, using at least a partially calculated positional expansion score. In some embodiments, introduced amino acid modifications are used to improve criteria such as density, minimum / maximum separation, and positional expansion score (e.g., to ensure that the introduced amino acid modifications are introduced in a manner that allows them to be distributed within the conserved region). In some implementations, introduced amino acid modifications and signature mutations of existing, such as specific antigens (e.g., viral protein variants), are used to enhance criteria such as density, minimum / maximum separation, positional expansion scores, etc. (e.g., to ensure that the introduced amino acid modifications are introduced in a way that distributes them in conserved regions, and also to avoid mutating existing unique epitopes, for example, those that are desired to be preserved, for example, to promote the production of antibodies against these regions (e.g., with affinity)).
[0313] Amino acid modification
[0314] In some embodiments, the amino acid modifications that introduce the memory-triggered subregion are selected from a set of permitted mutations. In some embodiments, the set of permitted mutations includes one or more mutations identified as permitted for each of at least a portion of the amino acid positions within a specific sequence. Permissible mutations can be identified using sequence data (e.g., nucleotide and / or amino acid sequences) of an antigen associated with a reference antigen. For example, where the reference antigen is a specific protein of a viral variant, such as the SARS-CoV 2 spike protein of a specific SARS-CoV 2 variant (e.g., Omicron spike protein, XBB spike protein, etc.), permissible mutations can be determined by first selecting a set of relevant lineages and obtaining, for example, proprietary sequencing data and / or public data (e.g., GSAID) for sequences of specific protein versions of various viral variants classified as belonging to the selected relevant lineage group. Mutations occurring in each of the relevant variants of the specific protein can be identified and included in the set of permitted mutations. In some embodiments, mutations included in the set of permitted mutations are filtered to include only those mutations observed at a sufficiently high rate or frequency, such as at or above a specific threshold frequency. For example, in some implementations, only those mutations observed within at least a certain threshold percentage of the relevant variant sequence are included in the group of permitted mutations.
[0315] In some embodiments, the relevant variant sequence may be or comprise a protein sequence within a similar family and / or a protein sequence known to perform a function similar to that of the reference antigen. For example, the reference antigen may be a specific SARS-Cov2 protein, such as the spike protein. In some embodiments, the reference sequence may be a spike protein of another coronavirus, or a subgroup thereof, such as the spike protein of a related SARS virus (e.g., viruses known to infect humans; e.g., SARS-Cov1, MERS), or a virus known to exist / originate from a specific region. In this way, in some embodiments, amino acid modifications may be selected and / or these modifications may be utilized based on variations present in other relevant polypeptides.
[0316] For example, in some embodiments, the reference sequence may be a sequence containing a variant of the reference antigen, such as a precursor or member of another branch on a phylogenetic tree. The reference sequence may be limited to variants with a specific level of prevalence and / or immune escape. In some embodiments, the reference sequence may be available from one or more public and / or proprietary databases, such as GSAID.
[0317] In some embodiments, one or more additional criteria are used to select amino acid modifications that provide a sufficient level of disruption and / or still maintain polypeptide chain stability. For example, in some embodiments, amino acid modifications that replace existing amino acids with certain comparable / similar amino acids are excluded, for example, because they do not provide sufficient disruption. For example, in some embodiments, amino acid modifications that disrupt cysteine bridges are not permitted and / or are penalized. In some embodiments, amino acid modifications that cause charge flipping and / or produce large charge changes are not permitted and / or are penalized. For example, the similarity / dissimilarity criteria and structure preservation criteria described herein can be implemented through (e.g., encoded) rules, such as in conditional logic, lookup tables, etc.
[0318] Figure 2 An example process is illustrated in which the sequence of a relevant antigen is accessed / acquired and analyzed to identify and generate a set of permissible mutations that can be used to generate amino acid modifications and disrupt conserved subregions of a reference antigen. In some embodiments, the steps of selection and insertion, and optionally filtering of amino acid modifications, can be repeated (e.g., iteratively) to generate multiple candidate engineered variants based on a single initial conserved region.
[0319] Scoring and selecting candidate variants
[0320] In some embodiments, since conserved subregions are disrupted by the introduction of amino acid modifications, the resulting computer-simulated engineered variants of the reference antigen can be scored, for example, to assess their viability, ability to infect host cells, and relative similarity / dissimilarity to existing variants. For example, in some embodiments, such as... Figure 3 As shown, one or more candidate disrupted peptide models 342a, 342b...342n are generated from an initial peptide model 302 representing a specific reference antigen (e.g., by recognizing and disrupting conserved regions 320 and 330). Each candidate disrupted peptide model corresponds to the initial peptide model, but its conserved subregions are disrupted by introducing one or more amino acid modifications.
[0321] Candidate disrupted peptide models can be evaluated and scored in a variety of ways, such as by one or more methods described in PCT Publication WO 2022 / 235847A1 entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST” and PCT Publication WO 2022 / 235853A1 entitled “IMMUNOGENSELECTION”, both published on November 10, 2022, the contents of each of which are incorporated herein by reference in their entirety.
[0322] For example, in some implementations, machine learning models (such as language models) can be used to score candidate peptide models. In some implementations, the language model is trained or has been trained to generate predicted probabilities for each of one or more amino acids at each of one or more positions in an input sequence. Once trained, the language model can be used to predict the overall probability of a particular variant represented by a specific input amino acid sequence, such as log-likelihood, conditional log-likelihood, etc.
[0323] In some implementations, the language model trained in this manner can also be used to determine the embedding vector representation of the input amino acid sequence. The embedding vector representation of the amino acid sequence is an internal representation generated by the machine learning model and can be extracted as output from one or more hidden layers of the machine learning model. The extracted embedding layer representation of the variant can be used independently and / or used to generate feature vectors representing specific variants.
[0324] For example, in some embodiments, a polypeptide sequence is provided as input to a machine learning model, which in turn generates an internal representation (i.e., embedding) of the input polypeptide sequence, which can represent the input sequence as a high-dimensional matrix or vector (e.g., a numerical matrix or vector). For example, in some embodiments, a machine learning model (such as certain language models described further in detail herein) may receive an amino acid sequence as input and generate an internal (e.g., embedding) representation. In some embodiments, the embedding is or contains a vector (e.g., zi) of, for example, length D, where D is the dimension of the vector at each amino acid position in the sequence, such that the initial embedding directly corresponding to the internal representation output of a particular layer (e.g., the final embedding layer of a recurrent neural network, e.g., the feature map of the final converter layer of a converter-based model) can be a matrix of size n×D [e.g., or (n+1)×D], where n is the length of the input amino acid sequence and D is the dimension [e.g., in some embodiments, additional class tokens may be used and appended to the network, such that the matrix size is (n+1)×D depending on the particular machine learning model used]. In some implementations, the feature vector can then be determined as the average of at least a subset of sequence positions {e.g., all sequence positions, [e.g., excluding the first (e.g., class token) position, which does not represent or correspond to an amino acid in the sequence]}, to determine, for example, the feature vector of the amino acid sequence. Further details of the embedding vectors and the methods for determining the feature vectors based thereon are provided in PCT Publications WO 2022 / 235847 and WO / 2022 / 235853, the contents of each of which are incorporated herein by reference in their entirety.
[0325] In this way, in some implementations, the machine learning model can be used to generate a corresponding feature vector (e.g., machine learning model-based embedding) for each of the multiple viral variant peptide sequences. This method can be used to associate each variant sequence with a location in a higher-dimensional feature vector / embedding space.
[0326] In some embodiments, feature vectors and / or embeddings representing a specific sequence (e.g., a viral variant) can be compared with each other and / or with a specific version, such as with a wild-type version of a protein and / or other specific variants of interest. In some embodiments, distances in the embedding space (e.g., L1 distance, L2 distance, etc.) can be calculated and used to evaluate the similarity and / or dissimilarity between a specific candidate variant and one or more other variants for comparison. In some embodiments, distances in the embedding space may be referred to as semantic variation scores and are considered to at least partially indicate the potential of a specific variant to be recognized by a host system previously exposed to one or more reference antigens.
[0327] Machine learning models used to generate feature vectors can utilize and implement a variety of machine learning techniques. For example, a machine learning model can be a deep learning model (e.g., an artificial neural network with one or more (e.g., multiple) hidden layers), such as a language or large language model (LLM). In some implementations, the machine learning model is or includes one or more recurrent models, such as long short-term memory (LSTM), implemented alone or in combination, for example in a bidirectional LSTM (bi-LSTM). In some implementations, the machine learning model can be or includes one or more transformer models. Examples of machine learning models include, but are not limited to, evolutionary scaling models (ESM), bidirectional encoder representations from transformers (BERT), etc.
[0328] Machine learning models can be trained, for example, on protein sequence data available from public and / or private repositories, such as UniRef (see, for example, Suzek et al., “UniRef: comprehensive and non-redundant UniProt reference clusters”, 2007), the Viral Pathogens Resource (ViPR) database of the National Center for Biotechnology Information, etc. Training can be accomplished by providing the machine learning model with incomplete and / or partially masked input sequences and asking it to predict the missing and / or masked data. In this way, the machine learning model can be trained in an unsupervised manner. Further details regarding machine learning models and their training methods, such as specific techniques for training recurrent and / or transformative models, can be found, for example, in PCT Publications WO 2022 / 235847 and WO / 2022 / 235853, the contents of which are incorporated herein by reference in their entirety.
[0329] In some embodiments, structural modeling techniques can be used to score candidate peptide models. For example, in some embodiments, viral peptide receptor binding scores, such as ACE2 binding scores, can be determined for one or more candidate synthetic variants. In some embodiments, epitope alteration scores can be determined for one or more candidate variants.
[0330] In some embodiments, the number of mutations can be evaluated and used as a score, for example, to select candidate peptide models that have a specific range in terms of the number of mutations. In some embodiments, a mutation co-occurrence score can be used to evaluate whether artificially introduced mutations in a particular candidate variant match co-occurrences observed in naturally evolving variants in the real world.
[0331] In this way, each candidate can be associated with and characterized by a set of performance scores 344a, 344b...344n.
[0332] In some embodiments, one or more candidate peptide models may be selected, for example, based on such performance scores, as representatives of the engineered antigen to be used, e.g., as an immunogenic composition. In some embodiments, a set of two or more scores described herein may be combined, for example, as a (e.g., weighted) combined score. In some embodiments, a set of two or more scores described herein may be combined to identify a group of Pareto fronts and / or Pareto optimal solutions. In some embodiments, a Pareto score may be determined based on a combination of two or more scores, such as PCT Publication WO 2022 / 235847 A1 entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST”, published on November 10, 2022, and PCT Publication WO 2022 / 235853A1 entitled “IMMUNOGEN SELECTION”, the contents of each of which are incorporated herein by reference in their entirety.
[0333] V. Processing engineered antigens
[0334] In some embodiments, disrupted peptide models created by computer simulation using the systems and methods described herein can be stored and / or provided for, for example, display and / or further processing. For example, in some embodiments, the disrupted peptide model can be used to generate corresponding RNA sequences encoding engineered antigens represented by the disrupted peptide model. In some embodiments, engineered antigens based on disrupted peptide models can be fabricated and their bioactivity evaluated, e.g., in vitro. For example, in this manner, multiple initial candidate engineered antigens can be created using the computer simulation design techniques described herein, followed by in vitro fabrication and screening.
[0335] For example, in some embodiments, the assay can be used to examine expression and folding. For example, engineered variants designed using one or more methods described herein can be produced and their activity in terms of expression and folding can be tested. In some embodiments, one or more antibodies or binding agents can be used to analyze intracellular and surface expression of an antigen. In some embodiments, for example, to test engineered SARS-CoV 2 variants, an ACE2 binding agent can be used. In some embodiments, a set of reference (e.g., known) neutralizing antibodies can be used to evaluate the ability of engineered constructs to evade existing neutralizing antibodies. In some embodiments, the reference set of neutralizing antibodies includes one or more antibodies that bind to different epitope classes (e.g., A, B, C, D, E, F). In some embodiments, the assay can be extended to the analysis of immune serological data from vaccinated / recovered individuals. In some embodiments, a sham virus assay is performed to examine for loss of nAb titers.
[0336] In some implementations, one or more (e.g., multiple) test procedures are used and arranged in a decision tree / hierarchical manner, such as as described in Example 8. For example, in some implementations, all or a subgroup of steps may be used, such that in each step / test, the build is either retained and passed to the next step or discarded, in order to progressively filter the engineered build design until a subgroup passes each step. These steps may include any of the following:
[0337] 1. Expression and folding check
[0338] 2. Eliminate / reduce the binding of a set of reference antibodies;
[0339] 3. Eliminate / reduce the binding of complex immune sera;
[0340] 4. Neutralize and eliminate corresponding pseudoviruses;
[0341] 5. Specific immunogenicity studies.
[0342] Example 8 describes in more detail an example method for the design of engineered SARS-CoV-2 antigens.
[0343] In some implementations, the measurement results, as described herein, can be combined with computer simulation design procedures, for example, in an iterative manner, to provide information and / or add constraints for subsequent rounds of computer simulation design.
[0344] C. Delivery of engineered antigens
[0345] Those skilled in the art will understand that effective vaccination can be achieved by delivering engineered antigens to a subject, thereby exposing the antigens to the subject's B cells.
[0346] According to the various embodiments described herein (e.g., in Section B above, and / or the various examples described below), the engineered antigens of this disclosure may comprise a full-length viral protein and / or specific portions thereof, such as a modified RBD, for example, to disrupt conserved regions. Constructs comprising the engineered antigens of this disclosure may combine the engineered viral protein and / or portions thereof with one or more additional elements, such as, for example, one or more of the following: secretion signals, adapters, multimerized regions (e.g., fibrous substitute protein domains), and membrane-associated portions (e.g., transmembrane regions).
[0347] In some embodiments, the delivery of engineered antigens and / or constructs containing them can be achieved by administering such antigens (e.g., peptide antigens). In some embodiments, such delivery can be achieved by administering a composition that subsequently generates the antigen (e.g., in and / or by the recipient). In some embodiments, delivery of peptide antigens is achieved by administering a composition containing a polynucleotide encoding the peptide antigen. In some such embodiments, the polynucleotide may be or contains DNA; in some such embodiments, the polynucleotide may be or contains RNA. Furthermore, those skilled in the art will recognize that non-natural residues, bonds, and / or other elements are frequently used in therapeutic nucleic acids administered to a subject. Various exemplary peptide and / or nucleic acid sequences suitable for providing engineered antigens of this disclosure are described in the examples provided herein (e.g., Examples 10 and 11).
[0348] Of particular interest in certain embodiments of this disclosure is the administration of a composition comprising RNA encoding an engineered antigen as described herein.
[0349] In many embodiments, the provided pharmaceutical composition (e.g., an immunogenic composition, such as a vaccine) delivers an antigen (e.g., an engineered antigen) as described herein by delivering a nucleic acid construct, such as, in many embodiments, RNA encoding one or more engineered antigens as described herein and expressed in a subject after administration of the pharmaceutical composition (e.g., an immunogenic composition, such as a vaccine).
[0350] Among other things, this disclosure covers the understanding that administering nucleic acids, and in particular RNA, to deliver encoded antigens (e.g., by expression) can provide a variety of benefits relative to other strategies for immunizing against infections, such as SARS-CoV-2 infection.
[0351] Among other things, this disclosure provides the following insight: RNA may be particularly useful and / or effective as an active agent in pharmaceutical compositions (e.g., immunogenic compositions, such as SARS-CoV-2 vaccines) for a variety of reasons, including the potential for RNA to possess intrinsic adjuvant properties. As described herein, RNA can induce very high antibody titers against SARS-CoV-2 proteins, for example, particularly SARS-CoV-2 antigens associated with variants of interest that have a high potential for immune escape.
[0352] Furthermore, experience with SARS-CoV-2 vaccines has demonstrated that RNA active ingredients can also elicit significant and diverse T-cell responses, particularly when combined with strong antibody responses, representing a combination of immune features considered to be the most likely to maximize protection.
[0353] The provided polynucleotides can be delivered for the therapeutic applications described herein using any suitable method known in the art, including, for example, delivery as naked RNA, or delivery mediated by viral and / or non-viral vectors, polymer-based vectors, lipid compositions, nanoparticles (e.g., lipid nanoparticles, polymer nanoparticles, lipid-polymer hybrid nanoparticles, etc.) and / or peptide-based vectors. See, for example, Wadhwa et al., “Opportunities and Challenges in the Delivery of mRNA-Based Vaccines”, Pharmaceutics (2020) 102 (27 pages), where information regarding various methods that can be used to deliver the polynucleotides described herein is incorporated herein by reference.
[0354] In some implementations, one or more polynucleotides may be formulated together with lipid nanoparticles for delivery (e.g., application).
[0355] In some embodiments, lipid nanoparticles can be designed to protect polynucleotides from extracellular RNases and / or engineered to deliver RNA systemically to target cells. In some embodiments, these lipid nanoparticles are particularly useful for delivering polynucleotides when administered intravenously or intramuscularly to a subject.
[0356] Further details regarding methods applicable to the delivery of the engineered antigens described herein can be found in PCT Publication WO 2022 / 235847A1 entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST”, published on November 10, 2022, the contents of which are incorporated herein by reference in their entirety.
[0357] D. Subjects and indications
[0358] In some embodiments, the systems and methods described herein can be used to design engineered antigens suitable for evoking improved immune responses. In some embodiments, the methods described herein can be used to create engineered antigens for vaccination of specific subjects, such as specific individuals and / or specific populations. For example, in some embodiments, the antigen engineering techniques described herein can be used to engineer antigens for delivery to subjects who have previously been exposed to a specific variant of a reference antigen (e.g., a previously naturally circulating variant). For example, in some embodiments, specific conserved epitopes to be destroyed can be identified using information associated with, for example, a specific epitope of a specific variant previously exposed to by the subject. For subjects previously exposed through infection and / or previous vaccination, such methods can be used to reduce memory responses and promote primary immune responses.
[0359] In some embodiments, engineered antigens can be customized for subjects whose memory B cells have been evaluated (e.g., explicitly created or selected from a set of pre-existing options). In some embodiments, the methods of this disclosure include administering a composition delivering an engineered antigen as described herein to a subject whose memory B cells have been evaluated. For example, in some embodiments, the methods described herein can be used to create engineered antigens that modify memory epitopes and target them (at least in a non-neutralized state) with memory B cells of a subject. In some embodiments, such engineered antigens can be administered to a subject.
[0360] E. Some example applications
[0361] Among other things, in some embodiments, the computer simulation design method for engineered antigens described herein can be applied to the design of immunogenic compositions. Such immunogenic compounds may have, for example, improved performance and / or be specifically tailored to individual subjects and / or populations, such as taking into account specific populations with regard to immune memory responses (e.g., depending on the antigen first exposed to various individuals and / or populations).
[0362] Therefore, in some embodiments, the methods described herein can be used to create improved immunogenic compositions.
[0363] The techniques described herein, including the various methods specifically described in the embodiments for the engineering of SARS-CoV 2 antigens, can be applied to a variety of types of antigens and pathogens.
[0364] In some implementations, the infectious pathogen is a virus, bacteria, or eukaryotic cell (e.g., Plasmodium).
[0365] In some embodiments, the infectious pathogen is a respiratory virus. In some embodiments, the infectious pathogen is an RNA virus. In some embodiments, the infectious pathogen is a coronavirus (e.g., MERS, SARS, or SARS-CoV-2). In some embodiments, the infectious pathogen is HIV. In some embodiments, the infectious pathogen is HSV (e.g., HSV-1 or HSV-2). In some embodiments, the infectious pathogen is RSV. In some embodiments, the infectious pathogen is norovirus. In some embodiments, the infectious pathogen is influenza virus. In some embodiments, the infectious pathogen is Plasmodium falciparum. In some embodiments, the infectious pathogen is orthopoxvirus (e.g., monkeypox).
[0366] In some embodiments, the infectious pathogen is bacteria. In some embodiments, the bacteria are Mycobacterium. In some embodiments, the bacteria are selected from Haemophilus influenzae, Chlamydophila pneumoniae, Mycoplasma pneumoniae, Staphylococcus aureus, Moraxella catarrhalis, Legionella pneumophila, and Streptococcus pneumoniae. In some embodiments, the bacteria are Streptococcus pneumoniae.
[0367] In some embodiments, the infectious pathogen is an RNA virus. The compositions provided herein may have specific advantages in providing an immune response against RNA viruses with a relatively high mutation rate (relative to other infectious pathogens).
[0368] In some embodiments, the infectious pathogen comprises many strains, variants, or lineages. In some embodiments, the infectious pathogen has a relatively high mutation rate (e.g., relative to other infectious pathogens).
[0369] In some implementations, infectious pathogens tend to escape the immune system.
[0370] In some implementations, the infectious pathogen is an infectious pathogen that requires regular injections of a seasonal, adaptive variant of the booster.
[0371] In some embodiments, the infectious pathogen antigen is a solvent exposed on the surface of the infectious pathogen. In some embodiments, the infectious pathogen antigen is a glycoprotein. In some embodiments, the infectious pathogen antigen is involved in host cell recognition. In some embodiments, the infectious pathogen antigen is involved in host cell entry. In some embodiments, the infectious pathogen antigen contains one or more B cell epitopes (e.g., one or more neutralizing epitopes).
[0372] F. Example
[0373] i. Example 1: XBB.1.5 conserved epitopes and marker mutations
[0374] This embodiment describes conserved regions and signature mutations in XBB.1.5 and other regions of interest, such as ACE2.
[0375] To facilitate novel B-cell responses to the XBB.1.5 variant, the design method in this embodiment first collected 1,004 binding and neutralizing B-cell epitopes from CoV-AbDab and IEDB. These epitopes were compared to XBB.1.5 to identify 107 of the initial 1,004 epitopes that were not mutated when considering XBB.1.5. This set of conserved epitopes was further refined through manual examination, including the removal of epitopes located not on the RBD surface and / or in other epitope subgroups. This process yielded a final set of 26 distinct conserved epitopes covering 140 sites in the spike protein RBD. Table 2A below lists these conserved epitopes and their respective sources.
[0376] Table 2A. Twenty-six non-mutated binding / neutralizing epitopes.
[0377]
[0378]
[0379]
[0380] Some of the methods described herein (e.g., identifying target regions) utilize the recognition of the ACE2 interface of the SARS-CoV-2 spike (S) protein, which in some embodiments may be defined as in Lan et al., “Structure of the SARS-VoV-2 spike receptor-binding domain bound to the ACE2 receptor,” Nature, 581:215-220 (2020), and include the following locations:
[0381] • ACE2 interface: K417, G446, Y449, Y453, L455, F456, A475, F486, N487, Y489, Q493, G496, Q498, T500, N501, G502, Y505ii. Example 2: Engineered custom-designed XBB-based antigen
[0382] This embodiment describes the design of engineered antigens based on XBB variants of the SARS-Cov2 spike protein (e.g., used as reference antigens). The engineered antigens described in this embodiment are designed to mitigate and / or avoid activation of memory immune responses induced by memory B cells and / or T cells in subjects (such as vaccinated individuals) to promote the production of novel neutralizing antibodies against specific epitopes associated with the ACE2 binding interface and containing XBB signature mutations.
[0383] Go to Figure 4A The XBB RBD portion was used as a reference antigen for the design of engineered antigens. The method described in this embodiment begins with the XBB RBD sequence. A PDB structure including RBD positions 325 to 527 (PED ID 7EAM) was used as a peptide model. XBB signature mutations (as listed in Table 1B) were... Figure 4A The regions shown are in red, while non-mutated regions are shown in gray.
[0384] Figure 4B The XBB RBD model is shown, in which the ACE2 interface is identified in green, as defined by Lan et al., “Structure of the SARS-VoV-2spike receptor-binding domain bound to the ACE2 receptor,” Nature, 581:215-220 (2020), and includes the locations listed in Example 1 above:
[0385] Figure 4A and 4B The identified purple conserved surface is also shown. The conserved surface was identified as a continuous, non-mutated surface. The amino acid sites identified as belonging to the conserved region in this embodiment are listed below:
[0386] ·Conserved regions: L335, E340, A348, S349, Y351, A352, N354, R355, K356, R357, S359, N360, V362, D364, S366, Y369, N370, A372, F 377, K378, Y380, G381, S383, P384, T385, K386, N388, D389, L390, C391, F392, T393, N394, Y396, P412, G413, Q414, T4 15. K424, P426, D427, D428, T430, K444, N450, L452, R457, K458, S459, K462, P463, F464, E465, R466, D467, I468, S46 9. T470, E471, I472, Y473, Q474, P479, N481, G482, V483, E516, L517, L518, H519, A520, P521, T523, C525, G526, P527
[0387] Figure 4C A color version of the XBB structural model is shown. Color changes indicate which residues (amino acid sites) belong to known neutralizing epitopes associated with (e.g., targeting) neutralizing antibodies. One hundred and thirteen (113) known neutralizing epitopes were identified based on data from the RCSB Protein Database (PBD) and the Immunoeptope Database and Analysis Resource (IEDB). They were scored according to the number of neutralizing epitopes appearing at each residue. Figure 4C The color-changing range in the image is from yellow to green to blue to indicate the relative frequency of locations belonging to epitopes, with the highest count being 70 and the color being blue, the middle count being green, and the low count being yellow (i.e., sites that do not belong to or are not frequently associated with epitopes). Figure 4C This embodiment specifically considers neutralization epitopes, but subsequent designs also consider non-neutralization epitopes.
[0388] This embodiment aims to generate an engineered version of XBB that limits or avoids triggering a memory immune response, instead producing novel antibodies tailored to the XBB signature mutation. Therefore, 113 neutralizing epitopes were compared with the XBB signature mutation to identify a subgroup of twenty-six (26) remaining epitopes not targeted by the XBB signature mutation. These conserved, non-mutated epitopes are listed in Table 2B below.
[0389] To disrupt the non-mutant epitopes and conserved surfaces of the XBB RBD peptide model, amino acid modifications were introduced at various positions. These modifications were generated through selection from a set of permissible mutations known to occur in the XBB, XBB sublineages, and Omicron. Additional criteria were used to restrict the introduced amino acid modifications to (i) require that the modifications cause significant changes—e.g., changes from Asp to Glu, from Arg to Lys, from Ile to Leu, and from Asn to Gln are not considered; and (ii) prohibit (e.g., exclude) modifications that disrupt cysteine bonds—e.g., modifications at positions C391 or C525 are not permitted.
[0390] According to this standard, 200,000 variants were randomly generated, each a version of XBB, but with conserved regions being disruptive. Each candidate variant was evaluated against various design criteria. Specifically, first, the candidate design needed to modify amino acids located in each of the 26 non-mutant epitopes, for example, to maximize immune evasion. Second, the candidate designs were evaluated to ensure that the amino acid modifications were sufficiently extended across the conserved surface, rather than clustered together. Finally, each design was scored using a combination of computer-simulated structural modeling and a machine learning-based language model. Structural modeling was used to calculate the ACE2 binding score, as described, for example, in PCT Publication WO 2022 / 235847 A1, entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST”, published on November 10, 2022. A machine learning-based language model, similar to that described in PCT Publication WO 2022 / 235847 A1 entitled "TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST", is used to determine the log-likelihood and semantic variation score for each candidate synthetic variant. The candidate designs can then be evaluated based on these criteria.
[0391] Figure 4D and 4E Two design examples are shown, namely Design 1 ( Figure 4D ) and Design 2 ( Figure 4E Both designs introduced amino acid modifications within each non-mutant epitope and extended the modifications sufficiently to conserved surfaces. ACE2 binding score, log-likelihood score, and semantic change score were also calculated for each design. Design 1 ( Figure 4D ) were identified as having higher ACE2 binding scores, while design 2 ( Figure 4E () were identified as more immune escapes (e.g., based on higher semantic change scores).
[0392] Figure 4D and 4E The cyan color change identified the sites of amino acid modifications introduced in the conserved regions of Designs 1 and 2, respectively. The modified amino acid sites for each design are listed below:
[0393] • Design 1 (Higher ACE2 Combination) Modification Locations: L335F Y351F N354DS359N A372R L390RQ414K T430IP463S E471Q N481K H519Q;
[0394] • Design 2 (More Immune Escape) Modification Locations: L335F A352V A372RN388K L390R T415IT430IP463L T470N Y473S P479S G482R L518VA520E T523A
[0395] Table 2B. Twenty-six non-mutated binding / neutralizing epitopes.
[0396]
[0397]
[0398]
[0399]
[0400] iii. Example 3: Engineered Custom-Made Antigen Based on XBB.1.5
[0401] This embodiment describes another example design approach for an XBB variant based on the SARS-Cov2 spike protein, specifically, XBB.1.5 (e.g., used as a reference antigen) to generate an engineered antigen. Similar to Example 2, the engineered antigen described in this embodiment is designed to mitigate and / or avoid activating the memory immune response in subjects (such as individuals receiving a vaccine) in order to generate a novel B-cell immune response starting from XBB.1.5.
[0402] Figure 5A The image shows the SARS-CoV 2 virus and its components, including the spike (S) protein (shown in more detail on the right side of the image) and the ACE2 host receptor that binds to it. Figure 5B The 3D structural representation of the XBB.1.5 RBD portion of the SARS-CoV-2S protein, including positions 325 to 527 (PED ID 7EAM), is shown as a peptide model. The following XBB.1.5 signature mutations are... Figure 5BThe regions shown are in red, while non-mutated regions are shown in gray.
[0403] • Marker mutations of XBB.1.5: T19I, L24-, P25-, P26-, A27S, V83A, G142D, Y144-, H146Q, Q183E, V213E, G252V, G339H, R346T, L368I, S371F, S373P, S375F, T376A, D405N, R 408S, K417N, N440K, V445P, G446S, N460K, S477N, T478K, E484A, F486P, F490S, Q498R, N501Y, Y505H, D614G, H655Y, N679K, P681H, N764K, D796Y, Q954H, N969K
[0404] Figure 5C The complete 3D structure of the spike protein is shown, with the signature XBB.1.5 mutation indicated in red.
[0405] To facilitate novel B-cell responses to the XBB.1.5 variant, the design method in this embodiment first collected 1,004 binding and neutralizing B-cell epitopes from CoV-AbDab and IEDB. These epitopes were compared to XBB.1.5 to identify 107 of the initial 1,004 epitopes that were not mutated when considering XBB.1.5. This set of conserved epitopes was further refined through manual examination, including the removal of epitopes located not on the RBD surface and / or in other epitope subgroups. This process yielded a final set of 26 distinct conserved epitopes covering 140 sites in the spike protein RBD, listed in Table 1 above (Example 1).
[0406] Continuous conservative surfaces on XBB.1.5RBD were also identified, and in Figure 5B The amino acid sites identified as belonging to the conserved region in this embodiment are shown in purple.
[0407] • Continuously conservative surfaces: L335, E340, A348, S349, Y351, A352, N354, R355, K356, R357, S359, N360, V362, D364, S366, Y369, N370, A372, F377, K378, Y380, G381, S383, P384, T385, K386, N388, D389, L390, C391, F392, T393, N394, Y396, P412, G413, Q414 T415, K424, P426, D427, D428, T430, K444, N450, L452, R457, K458, S459, K462, P463, F464, E465, R466, D467, I468, S4 69. T470, E471, I472, Y473, Q474, P479, N481, G482, V483, E516, L517, L518, H519, A520, P521, T523, C525, G526, P527
[0408] Figure 5D The XBB RBD model is shown, in which the ACE2 interface is identified in green, as defined by Lan et al., “Structure of the SARS-VoV-2spike receptor-binding domain bound to the ACE2 receptor,” Nature, 581:215-220 (2020), and includes the following locations:
[0409] • ACE2 interface: K417, G446, Y449, Y453, L455, F456, A475, F486, N487, Y489, Q493, G496, Q498, T500, N501, G502, Y505
[0410] Figure 5E A color-coded version of the XBB structural model is shown. The color changes indicate which residues (amino acid sites) belong to neutralizing epitopes associated with neutralizing antibodies (e.g., targeting). The model is scored based on the number of 1,004 neutralizing and non-neutralizing epitopes appearing at each residue. Figure 5E The color-changing range in the image is from yellow to green to blue to indicate the relative frequency of locations belonging to epitopes, with the highest count being 287 and the color being blue, the middle count being green, and the low count being yellow (i.e., sites that do not belong to or are not frequently associated with epitopes).
[0411] Figure 5FThe epitope density of the final group of 26 conserved epitopes on a continuously conserved surface is shown. The color change range is again from yellow to green to blue, indicating the relative frequency of the (amino acid) position belonging to one of the 26 conserved epitopes. Among the 26 conserved epitopes, two positions (428 and 518) reach (i.e., are present) a maximum value of 10. Positions located outside the continuously conserved surface are colored gray.
[0412] To disrupt the conserved subregions of XBB.1.5RBD, amino acid modifications were introduced at various sites on the conserved surface, totaling 76 sites, such that each conserved epitope was modified with at least one amino acid. Amino acid modifications to be introduced were selected from a set of permissible mutations using sequences identified as belonging to relevant variants of XBB, BA, and their sublineages. Mutations in these relevant lineages (i.e., XBB, BA, and their sublineages) with a frequency greater than or equal to 1% were included in the permissible mutations of this set. Additional filtering criteria were applied to exclude (i) Cys mutations that disrupt cysteine bonds, (ii) mutations between Asp-Glu, Arg-Lys, Ile-Leu, and Asn-Gln, and (iii) a set of five mutations that, as indicated by deep mutational analysis (DMS), significantly affect ACE2 binding. The final set of permissible mutations used in this embodiment (86 in total) is listed in Table 3, along with the lineage identification for each mutation.
[0413] Although the possible mutation sites and the absolute number of allowed mutations are relatively low, the number of possible combinations they generate is very large—approximately 10^34. For example, consider the following epitopes: (511, 512, 513, 514, 515, 516, 517, 518, 519, 520, 521, 522, 523, 524, 525). Allowed mutations are: E516Q, L517F / H, L518V, H519L / Q / R / Y, A520E / S / T, P521L / Q / S, T523A / S. Therefore, this epitope has 1 × 2 × 1 × 4 × 3 × 3 × 2 = 144 combinations. Thus, the number of possible combinations of amino acid modifications within a single conserved epitope can be very large, with a total of twenty-six epitopes, leading to the large search range mentioned earlier.
[0414]
[0415] In order to sample the larger resulting search space, the design technique in this embodiment uses Figure 5GThe process 500 is illustrated. Process 500 divides the creation of candidate engineered variants by introducing amino acid modifications into two steps. In the first step, specific (amino acid) positions 510 to be modified are selected to generate multiple sets of position combinations to be mutated. In this embodiment, each set of position combinations is generated by selecting positions until a position is selected within each of 26 conserved epitopes. This process produces 200,000 sets of position combinations. Figure 5G As shown, the location selection method also includes a filtering step 520, in which a location expansion score is calculated for each location combination, and the results of step 510 are filtered based on the location expansion score and the number of mutations. After filtering, 2,000 combinations remain and are grouped into 20 clusters.
[0416] After determining the final group of positional combinations, for each positional combination within the group, all feasible mutations 530 are generated (using the available options at each position in the mutation, allowing for mutations within the group). Modifications are filtered to ensure they do not introduce charge flips and to minimize charge changes. This process yields a total of 190,945 candidate variants.
[0417] Then, for each candidate variant, calculate the following four scores: 540:
[0418] • ACE2 combined scoring;
[0419] • Likelihood (log-likelihood) (RBD and full spikes);
[0420] • Semantic changes related to XBB.1.5 (RBD and full spikes); and
[0421] • Mutation co-occurrence;
[0422] The specific methods used to calculate the ACE2 binding score and mutation co-occurrence score will be described in detail below. The methods used to calculate the likelihood and semantic change scores are described in detail in PCT Publication WO 2022 / 235847A1, entitled "TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST," published on November 10, 2022, and PCT Publication WO 2022 / 235853A1, also published on November 10, 2022, each of which is incorporated herein by reference in its entirety. The final design 550 was selected using the Pareto front from all scores and clustering information, followed by a final manual review.
[0423] Figures 6A-6FThe 3D structures of several engineered antigen designs are shown. Tables 4A-4C list a specific design and its ID, the added mutations, and the various scoring and selection criteria described herein (identified in bold text in the first column of the table) in each row. Figures 6A-6E Each of these is shown in red as the signature XBB.1.5 mutation, and in yellow as the location where the mutation was introduced, and ( Figure 6B (Except for) locations where mutations are not considered are shown in gray, and conserved surfaces are shown in purple. Figure 6A and 6B Design S4_1 is shown, which was selected to maximize ACE2 binding. Figure 6C Design S1_3 is shown, which was chosen to minimize the number of mutations. Figure 6D Design S14 is shown, which was selected based on the aberrant surface distribution. Figure 6E The design S13_3, selected based on a high log-likelihood score, is shown. Figure 6F Design S11, which provides different groups of mutations, is shown. Tables 4A-4C list each final engineered antigen design, indicating the mutations added to XBB.1.5, the total number of mutations, the values of the four scores calculated for the design, and the selection criteria.
[0424]
[0425]
[0426]
[0427]
[0428] Figure 7 A radar chart showing the scores for each final engineered antigen design is presented, demonstrating good diversity.
[0429] Location expansion rating
[0430] RBD candidates are scored using a location expansion score, which estimates how the location expands across a (conservative) surface. In this embodiment, the location expansion score is calculated as follows:
[0431]
[0432] Where N is the number of mutations in the variant, and r i,j It is the C in 3D coordinate space between residue numbers i and j. ɑ -C ɑDistance. The coordinates are extracted from structure 7EAM. The score can later be optimized using the distance on the surface mesh instead of the direct 3D distance, since two residues may be close in 3D space but located on different faces of the RBD.
[0433] Location clustering
[0434] The design is clustered using a clustering algorithm described by Cao et al. (2022) according to the following steps:
[0435] Given N combinations of positions,
[0436] 1. Each sequence is first represented using binary encoding, such that the embedding dimension equals the number of positions.
[0437] 2. Each sequence is again represented by its Pearson correlation with every other sequence, such that the embedding dimension equals the number of sequences.
[0438] 3. Use Multidimensional Scaling (MDS) to reduce the embedding space to 256.
[0439] Mutation co-occurrence score
[0440] For each pair of mutations (mi, mj), the mutation co-occurrence score is calculated by approximating the following conditional frequencies:
[0441]
[0442] For a given RBD with N mutations (relative to WT), the mutation co-occurrence score in this embodiment is calculated as the average log frequency among all mutation pairs.
[0443]
[0444] ACE2 combined calculation
[0445] Finally, K-means with MDS embeddings were used to cluster the sequences. In this example, ACE2 binding scores were calculated using Deep Mutation Scan (DMS) data on ACE2 binding from Starr et al., 2022. This data includes DMS results for ACE2 binding positions 331-531 of the RBD for eight variants: Alpha, Beta, Delta, Eta, OmicronBA.1, OmicronBA.2, and two wild-type versions. The ACE2 binding score represents the “Δ binding” (log) of the mutation and variant. 10The changes in (KD) were summed to estimate the binding variation of any RBD variant. Using non-Omicron data, the ACE2 binding score obtained strong correlations on the BA.1 and BA.2 data, where Spearman r = 91.4% and Pearson r = 90.7%, and r 2 =83.5%. Figure 8 The combined change of the predicted values compared to the experimental values is shown.
[0446] Design sequence
[0447] The method described in this embodiment creates a list of mutations and sequences of engineered SARS-CoV 2 antigens.
[0448] Tables 5A and 5B below list the engineered antigen designs created using the methods described in this embodiment. Table 5A lists the RBD mutations according to the designs in this embodiment. Table 5B provides the RBD sequence for each design. As described herein, each sequence is an engineered version of the SARS-CoV 2 spike protein RBD created using the XBB.1.5 variant RBD as a starting point. Figure 5C The metrics calculated for the sequences in Tables 5A and 5B below are shown. For reference, Table 6 below lists the (natural) XBB.1.5 RBD and wild-type RBD sequences. Three versions of XBB.1.5 are shown. The RBD sequences shown in Table 6 allow for minor variations within specific portions (boundaries) of the spike protein corresponding to the RBD region.
[0449] Table 5A. RBD mutations used in engineered antigen design.
[0450]
[0451]
[0452]
[0453]
[0454] Table 5B. RBD sequences for engineered antigen design.
[0455]
[0456]
[0457]
[0458]
[0459]
[0460]
[0461]
[0462]
[0463] Table 6. Example XBB.1.5 RBD and wild-type RBD portions.
[0464]
[0465] iv. Example 4: Additional engineered variants
[0466] Figures 9A-9B 9D-9E shows four artificially designed RBD engineered antigens. Figures 9A-9B Two constructs with different mutation distributions are shown, focusing on their distance from existing mutations in XBB.1.5, with Table 7A listing the additional mutations added to XBB.1.5. Figure 9C This is a schematic diagram of the SARS-CoV-2 trimer, adapted from Starr et al., SARS-CoV-2RBD antibodies that maximize breadth and resistance to escape. Nature, 597, 97–102 (2021). https: / / doi.org / 10.1038 / s41586-021-03807-6. Figures 9D-9E Two constructs are shown, designed to enhance broadly conserved epitopes (class 4 and class 5 antibody sites) by disrupting epitopes identified as largely conserved between SARS CoV 1 and SARS CoV 2 (selected from mutations in the SARS CoV 1 sequence). Table 7B lists the additional mutations added to XBB.1.5.
[0467] Tables 7C and 7D show the various scoring metrics calculated for each sequence.
[0468] Table 7A. Two manual designs focusing on mutation propagation.
[0469]
[0470] Table 7B. Two manual designs for strengthening broadly conservative tabletops.
[0471]
[0472] Design parameters in Tables 7C and 7A.
[0473]
[0474] Design parameters in Tables 7D and 7B.
[0475]
[0476]
[0477] v. Example 5: Distributing the introduced amino acid modifications around the signature mutation
[0478] This embodiment describes an implementation of the antigen engineering technique described herein, wherein a positional expansion score is used to facilitate the insertion of new amino acid modifications into conserved surfaces in a distributed manner (e.g., as opposed to aggregation), taking into account not only other introduced amino acid modifications but also proximity to existing marker mutations (e.g., XBB marker mutations). Specifically, the antigen engineering method used in this embodiment is similar to those in Embodiment 3 above, but also includes a marker mutation in the positional expansion score (e.g., as described in Embodiment 3 above), which is used to evaluate candidate engineered antigen designs. Among other things, this method is intended to ensure that additional mutations are avoided from being placed in positions directly adjacent to the marker mutation. In this way, the implementation described in this embodiment maintains the unaltered form of the unique (e.g., Omicron) epitope and preferentially (e.g., only) mutates conserved epitopes. For example, in some cases where the marker mutation is not explicitly interpreted in this way, there is a possibility that the Omicron epitope may be modified in some way such that the induced antibody has a lower affinity for the actual, desired target epitope.
[0479] Tables 8A and 8B below list the engineered antigen designs created using the methods described in this embodiment (i.e., including XBB signature mutations in the positional expansion score). Table 8A lists the RBD mutations for the designs according to this embodiment. Table 8B provides the RBD sequences for each design. Table 8C shows the various scoring metrics calculated for each sequence.
[0480] Table 8A. RBD mutations used in engineered antigen design.
[0481]
[0482]
[0483]
[0484]
[0485]
[0486] Table 8B. RBD sequences for engineered antigen design.
[0487]
[0488]
[0489]
[0490]
[0491]
[0492]
[0493]
[0494] vi. Example 6: Evolutionary Algorithm and Epitope Change Scoring Version
[0495] This embodiment describes certain methods for mutation generation, scoring, and evaluation, which may additionally or alternatively be used in some implementations for the various methods described herein.
[0496] In some implementations, evolutionary algorithms (e.g., additionally or alternatively, or in combination with the two-step position and type selection techniques described herein) can be used to determine mutations. For example, each solution can be represented as a list of 26 mutations (in the case of SARS-CoV-2RBD), corresponding to the 26 conserved epitopes described in the examples above. In some implementations, a set of solutions can be maintained and updated by swapping the mutations of each epitope. Solutions can be scored using position expansion, ACE2 binding, and machine learning (ML)-based scoring such as log-likelihood and semantic change.
[0497] In some implementations, another version of the epitope change score may be used. For example, the current version of the epitope change score is described in, for instance, PCT Publication WO 2022 / 235847 A1 entitled “TECHNOLOGIES FOR EARLY DETECTION OF VARIANTS OF INTEREST”, published November 10, 2022, and PCT Publication WO 2022 / 235853A1 entitled “IMMUNOGEN SELECTION”, also published November 10, 2022, the contents of each of which are incorporated herein by reference in their entirety, considering that a mutation at any single position in the epitope is sufficient to “evade” the corresponding antibody. In some implementations, a more stringent version of the epitope change score may be used, for example, considering the position and distance between the mutation and the antibody CDR loop, to provide increased granularity and accuracy.
[0498] In some implementations, computer simulation structural modeling can be used to evaluate complex dynamics.
[0499] vii. Example 7: Updated Design Balancing Integrity Risk
[0500] This embodiment describes certain methods for mutation generation, scoring, and evaluation, which may additionally or alternatively be used in some embodiments of the various methods described herein. Specifically, this embodiment provides sequences generated to balance the risk of certain mutations being detrimental to protein integrity. For example, studies have found that some computer-simulated sequence generation programs may overrepresent mutations at positions 384, 430, and 463 in a set of suggested designs. Therefore, the designs proposed in this embodiment attempt to balance the risk (e.g., regarding mutation diversity across the designs) in case some computer-simulated mutations are detrimental to RBD integrity.
[0501] Table 9A below lists the RBD mutations according to the design of this embodiment concerning (e.g., added to) the XBB.1.5RBD portion of the SARS-CoV-2S protein (e.g., according to any of the three XBB.1.5RBDs provided in Table 6). As explained herein, the mutation location is identified in the context of the full-length SARS-CoV-2S protein (i.e., with reference to it).
[0502] Table 9B shows the scoring metrics calculated for each sequence shown in Table 9A.
[0503] Table 9A. RBD mutations used in engineered antigen design.
[0504]
[0505]
[0506]
[0507] viii. Example 8: Example Computer Simulation Design Test Program
[0508] This embodiment describes an example experimental procedure for testing engineered synthetic variants designed using the various methods described herein. The procedure in this embodiment uses three main chains / constructs to evaluate the engineered variants, as shown below.
[0509] pcDNA.3.1-SARS-CoV-2-XBB.1.5-CΔ19 construct: A mammalian expression plasmid encoding the XBB.1.5 spike protein with a truncated cytoplasmic tail and a corresponding mutant XBB.1.5RBD. The use of this construct is expected to allow for (i) examination of spike protein expression after transfection into HEK293T cells, (ii) pseudovirus generation, and (iii) analysis of immune escape parameters (e.g., complete escape with no detectable titer).
[0510] pST4-v.3.0.1-AGA-hAG-SP19-XBB.1.5-RBD-Foldon-TM construct: a template for transcribing mRNA encoding BNT162b3-like transmembrane anchored RBD-based vaccine antigens. The intended use of this construct is to allow for (i) examination of RBD expression in HEK293T cells, (ii) assessment of antibody binding using reference panels with different RBD epitope classes of binders and / or complex immune sera, and (iii) immunogenicity studies.
[0511] pcDNA3.4-SARS-CoV-2-XBB.1.5-RBD-his-avi construct: A template for RBD protein production, featuring an HIS tag (allowing for purification) and a BAP / avi tag (allowing for assay development). This construct can be used for enzyme-linked immunosorbent assays (ELISA), biolayer interferometry / surface plasmon resonance (BLI / SPR), and / or protein-protein interaction assays to evaluate, for example, whether a known antibody can bind to the created construct.
[0512] Example procedures for test design may include the steps listed below and / or Figure 10A One or more (e.g., at most all) of the steps shown are arranged in a hierarchical manner for the test steps:
[0513] Expression and folding checks: In step 1001, a flow cytometry-based method will be used to evaluate the expression and folding of various antigens (e.g., antigens of each design). An example FACS staining protocol that can be used in conjunction with this method is shown below.
[0514] like Figure 10BAs shown, after transfection, flow cytometry using hACE2-mFc as the primary binding agent can be used to analyze the intracellular and surface expression of the encoded antigen. The full-length XBBS protein can be used as a reference. The binding affinity of the RBD protein, as well as its intracellular and / or extracellular surface expression, can also be evaluated. The results of this step can be compared with predictions from machine learning algorithms, for example, used in computer simulation design for engineered antigens, and fed back to (e.g., to improve) the computer simulation design algorithm.
[0515] Antibody escape – monoclonal and polyclonal: Antibody escape will be evaluated using flow cytometry, with a set of selected reference antibodies binding to different epitope classes on the RBD (e.g., similar to the assay used for expression and folding checks, but using reference antibodies as binders instead of hACE2-mFc).
[0516] Multiple reference antibodies can be used and / or selected based on the specific epitope class to which the antibody binds and, alternatively, the affinity level to a specific SARS-CoV 2 variant (e.g., a reference antigen) used as a starting point for designing engineered variants (i.e., mutated). For example, a set of reference antibodies may include multiple antibodies that bind to different epitope classes (e.g., A, B, C, D, E, and F). Antibodies may be included that bind to a specific SARS-CoV 2 variant used as a reference antigen for designing its engineered version (e.g., with a disrupted conserved region), as well as antibodies that do not bind to that specific SARS-CoV 2 variant (but bind to other variants).
[0517] For example, Cao et al., Nature, 2021 (doi.org / 10.1038 / s41586-021-04385-3) (in Supplementary Table 1) provide a list of 247 neutralizing antibodies from which one can select an antibody group. For example, Table 10A shows a subset of 14 antibodies from Cao et al.'s table, which can be used as reference antibodies. Figure 10C Example binding assay data for several antibodies listed in Table 10A are shown. Table 10B shows the antibodies exhibiting XBB.1.5 binding, along with IC50 values for wild-type and BA.1 variants (* indicates data provided by Cao et al.). Table 10C shows the antibodies in Table 10A and their XBB.1.5 binding data, with those binding to XBB.1.5 highlighted in green text. Not wanting to be bound by any particular theory, some types of antibodies (such as those binding to AD class epitopes) do not bind to the BA.1S protein and are therefore not expected to bind to the XBB.S protein. Some antibodies binding to class E and F will be examined for binding to the XBB.S protein. The EC titration in step 1002 can be used. 50 The value is used to evaluate the elimination / reduction of the reference antibody binding.
[0518] Table 10A. Example panel of reference antibodies.
[0519]
[0520] Table 10B shows antibodies that bind to XBB.1.5.
[0521]
[0522] Table 10C. Antibodies from Table 10A, with XBB.1.5 binding / assay data.
[0523]
[0524]
[0525] In addition to or as alternatives to the various types of antibodies listed in Tables 10A-C, the antibody group may also include examples of various other antibody clones, such as (but not limited to) the following antibodies and / or similar antibodies:
[0526] • WRAIR-2057 (not wanting to be bound by any particular theory, this antibody is thought to bind to a site on the flanking side of the RBD that is so far highly conserved and has high affinity for the Omicron variant). For example, a more detailed description of WRAIR-2057 can be found in Dussupt, V. et al., Low-dose in vivo protection and neutralization across SARS-CoV-2 variants by monoclonal antibody combinations. Nat Immunol 22, 1503–1514 (2021). https: / / doi.org / 10.1038 / s41590-021-01068-z.
[0527] • COVOX-222, included in Supplementary Table 1 of Cao et al., Nature, 2021 (doi.org / 10.1038 / s41586-021-04385-3);
[0528] • COVOX-45 (not wanting to be bound by any particular theory, this antibody is thought to bind to a site on the flanking side of the RBD that is so far highly conserved and has high affinity for the Omicron variant). For example, COVOX-45 is described in more detail in Dejnirattisa i, W. et al., The antigenic anatomy of SARS-CoV-2 receptor binding domain, Cell, 184(8), 2183-2200.e22, 2021, https: / / doi.org / 10.1016 / j.cell.2021.02.032.
[0529] • S2H97 (not wanting to be bound by any particular theory, this antibody is thought to bind to a site on the flanking side of the RBD that is so far highly conserved and has high affinity for the Omicron variant). For example, S2H97 is described in more detail; see Starr, TN et al. SARS-CoV-2RBD antibodies that maximize breadth and resistance to escape. Nature 597, 97–102 (2021). https: / / doi.org / 10.1038 / s41586-021-03807-6.
[0530] ·553-49, for example, see Zhan W et al., Structural Study of SARS-CoV-2 Antibodies Identifies a Broad-Spectrum Antibody That Neutralizes the OmicronVariant by Disassembling the Spike Trimer. JVirol. 96(16):e0048022.2022.doi:10.1128 / jvi.00480-22.
[0531] Alternatively, immune serum from vaccinated / recovered individuals can be used to evaluate antibody escape. EC titration can be used. 50 The value is compared with the XBB1.5 value in step 1003 to evaluate whether the binding of the compound immune serum is eliminated / reduced.
[0532] A third method for assessing escape can be achieved by using a pseudovirus generation protocol and a pVNT assay setting to examine the loss of nAb titer. This method aims to check whether the neutralization of each pseudovirus in step 1004 has been eliminated.
[0533] For example, Figure 10A As shown, the final step may include dedicated immunogenicity studies, such as to evaluate whether the serum of immunized animals neutralizes XBB1.5 and / or one or more other variants. Vaccine compositions based on engineered SARS-CoV 2 antigens according to various embodiments described herein may be administered to unvaccinated animals (e.g., mice) (e.g., to evaluate the immune response and its extent) and vaccinated animals (e.g., mice), for example, to evaluate the ability to overcome the immunoblot.
[0534] Figure 10B An example protocol is shown, illustrating the transfection and flow cytometry steps. (See attached image.) Figure 10B As shown, pcDNA3.1-SARS-CoV-2-Swt-CΔ19 and / or pcDNA3.1-SARS-CoV-2-SXBB.1.5-CΔ19 can be transfected into HEK293T / 17 cells, and the binding assays described above can be used to evaluate hACE2 binding and (neutralizing) antibody binding, for example, for... Figure 10A The steps in the decision tree shown.
[0535] Concentration ranges for monoclonal antibody, hACE2, and human BNT162b23 (triple vaccination) polyclonal serum dilution series that produce S-shaped binding curves were identified as appropriate. Based on this criterion, mAbs can be tested in concentration ranges of 10 μg / mL–0.128 ng / mL (8-step 5-fold dilution series) based on prior apparent IC50 values of WT and BA.1 spikes, for example... Figure 10C As shown. Figure 10D As shown, hACE2-mFc binding can be tested using a concentration range from 25 μg / mL to 0.32 ng / mL. Figure 10D It was also shown that the XBB.1.5 spike exhibited approximately 6.7 times higher apparent affinity for hACE2 binding compared to the wild-type spike. For example... Figure 10E As shown, the triple vaccine 1M booster polyclonal serum confluence can be tested in a dilution series ranging from 1:20 to 1:1.56-2.500.
[0536] Table 11 shows several noteworthy variant (VOC) mutations and their numerical indicators, which can be used as a reference.
[0537] Table 11. Variant of interest (VOC) mutations and related scores.
[0538]
[0539]
[0540] The following shows an example FACS staining protocol, and Tables 12A and 12B below show the master mixture used for the secondary antibody labeling step.
[0541] Staining with a mixture of antibody, hACE2-mFc, or human serum.
[0542] • Cell isolation
[0543] Add PBS and centrifuge the cells.
[0544] • Cell counting:
[0545] • Distribute live cells / wells into 96-well round-bottom plates according to the plate layout.
[0546] • Centrifuge plates containing cells
[0547] • Discard the supernatant and vortex the plate to isolate the cells.
[0548] • Add FACS buffer and centrifuge plates
[0549] • Add the first antibody, ACE2-mFc, or human serum conjugate solution, and gently tap the plate to mix the cells.
[0550] • Incubate in the dark at 4°C for approximately 15 minutes.
[0551] • Add FACS buffer and centrifuge plates
[0552] • Discard the supernatant and vortex the plate to isolate the cells.
[0553] Wash twice with FACS buffer and centrifuge the plate.
[0554] • Add Mastermix (MM) containing secondary antibody and gently tap the plate to mix the cells.
[0555] • Incubate in the dark at 4°C for approximately 15 minutes.
[0556] • Add 200 μL / well FACS buffer (DPBS + 2% FBS hi. + 2 mM EDTA) and centrifuge the plate: 5 min, 460 × g, RT
[0557] • Discard the supernatant and vortex the plate to isolate the cells.
[0558] • Wash twice with FACS buffer and centrifuge the plate:
[0559] • Add Histofix (in PBS) to the cells, resuspend, and incubate at 2–8°C for 15 min.
[0560] • Centrifuge cells (5 min, 450 x g)
[0561] • Discard the supernatant and vortex the plate to isolate the cells.
[0562] Wash twice with FACS buffer and centrifuge the plate.
[0563] • Resuspend the cells in FACS buffer.
[0564] Store the plates in the dark at 4°C until they are to be measured.
[0565] Table 12A. Second antibody Mastermix I (used for primary antibody and human serum).
[0566]
[0567] Table 12B. Second antibody Mastermix II (for hACE2-mFc).
[0568]
[0569] ix. Example 9: Evaluation of computer-simulated design constructs by combining measurement
[0570] This embodiment describes the results of various measurements performed on certain constructs described herein according to the example test procedure described in Embodiment 8 above.
[0571] The first round of testing was conducted on engineered XBB.1.5 variants expressed in the context of the full-length SARS-CoV-2 spike (S) protein. Specifically, certain constructs described herein were integrated into the pcDNA.3.1-SARS-CoV-2-XBB.1.5-CΔ19 construct containing the corresponding mutated XBB.1.5 RBD. HEK293T / 17 cells were seeded in flasks and incubated at 37°C and 7.5% CO2 for 2 days. The constructs were transfected into HEK293T / 17 cells and incubated overnight at 37°C and 7.5% CO2. Flow cytometry was used to (i) evaluate expression levels by anti-S2 fragment antibody, (ii) examine conserved ACE2 binding capacity (by hACE2 binding), and (iii) assess elimination of RBD-targeting antibody binding using five (5) monoclonal antibodies (mAbs) proven to bind to XBB.1.5, as shown in Table 10B.
[0572] Go to Figure 11A-11C Binding tests were designed for certain variants expressed in the context of the full-length spike (S) protein as shown in Table 9A. Specifically, ACE-2 binding was evaluated using an hACE2-mFc antibody according to the FACS protocol described in Example 8 above. The binding response of each of the five mAbs listed in Table 10B was also evaluated using the protocol described in Example 8 above. Figure 11A-11C The S43 engineered XBB.1.5 variant design is shown. Figure 11B ) and S48 engineered XBB.1.5 variant design ( Figure 11C ) and (parent / original) XBB.1.5 reference ( Figure 11A The results showed that, among other things, variants S43 and S48 exhibited elimination binding to all five mAbs except for 1-2, and variant S43 maintained hACE-2 binding.
[0573] Turning to Figures 12-14, the variant designs shown in Table 9A are also expressed in the context of the BNT162b3-like (based on trimerized transmembrane anchored RBD) vaccine antigen. HEK293T / 17 cells were transfected with RNA from BNT162b3-XBB.1.5 and its respective variants. ACE-2 binding was then assessed by flow cytometry via hACE2-mFc binding and binding to five mAb antibody groups (listed in Table 10B), as previously described. Binding curves for polyclonal vaccine serum (triple BNT162b2 vaccination) were also evaluated to assess whether the effect of polyclonal serum dose response was greater compared to the hACE-2 dose response of engineered variants. Not wishing to be bound by any particular theory, it is considered that a stronger effect of polyclonal serum dose response on hACE-2 dose response compared to the parental XBB.1.5 of a specific variant indicates successful “masking / mutation” of conserved epitopes.
[0574] Figure 12A -B indicates the XBB.1.5 reference ( Figure 12A ) and S43 variant design ( Figure 12B The dose-response curve of (e.g.) Figure 12B As shown, compared to XBB.1.5, design S43, expressed in a trimerized™-anchored RBD environment, exhibited near-unchanged hACE-2 binding dose-response. Importantly, binding to class A antibody 2, class F antibody 1, and class B antibody 1 was eliminated, while binding to only class F antibody 2 and class E antibody 2 was retained.
[0575] Figure 13A and 13B It shows repetition Figure 12A The result of B ( Figure 13A Show the reference dose-response curve. Figure 13B The dose-response curve of the engineered antigen design S43 was displayed, confirming that... Figure 12B The results shown indicate that hACE-2 binding and binding to class F antibody 2 and class E antibody 2 mAbs are retained, but binding to class A antibody 2, class F antibody 1, and class B antibody 1 is eliminated. Figure 13C and 13DAs shown, polyclonal serum binding was also compared with hACE2 binding. Although the changes in the polyclonal serum binding curves (engineered XBB.1.5 antigen compared to parental XBB.1.5) were similar to (but not greater than) those in hACE-2 binding, in Figure 12A and 12B The elimination of mAb binding observed in multiple sets of results shown in 13A and 13B provides evidence that the mutations introduced into the engineered S43 design successfully disrupted the conserved epitope.
[0576] Figure 14A and 14B The XBB.1.5 reference and another set of dose-response curves for the S48 engineered antigen design expressed in the context of trimer™ anchored RBD are shown in Table 9A. S48 shows conserved binding to hACE-2 and class F antibody 2, as does S43 (S48 exhibits the 5 / 7 mutation also found in S43); class E antibody 2 shows minimal residual binding, while binding to class A antibody 2, class F antibody 1, and class B antibody 1 is completely eliminated.
[0577] Go to Figure 14C and 14D Additional data were provided to demonstrate the possibility of successfully disrupting conservative epitopes in S48. Figure 14C and 14D The changes in hACE2 binding curves of (i) the S48 variant related to the parental XBB.1.5 reference were compared. Figure 14C (ii) Changes in the binding curves of polyclonal sera () and (ii) polyclonal sera Figure 14D When comparing parental XBB.1.5 with S48, the change in ACE2 binding (approximately 4-fold) was smaller than the change observed in serum binding (>10-fold), suggesting a gradual elimination of the binding antibody response.
[0578] Table 13 summarizes the results of the screening data described in this embodiment in tabular form. The first two rows of the table show the binding levels of XBB.1.5 and the ideal, desired engineered antigen. The columns representing the binding data are arranged in two groups, corresponding to the different expression environments analyzed - (i) full-length XBB.1.5 spike (pcDNA.3.1-SARS-CoV-2-XBB.1.5-CΔ19), (ii) trimerized transmembrane (TM) anchored RBD design (pST4-v.3.0.1-AGA-hAG-SP19-XBB.1.5-RBD-Foldon-TM). As shown in Table 13, the reference antigen XBB.1.5 exhibited strong ACE-2 binding affinity and bound to all five mAbs in the group (selected as mAbs known to bind to XBB.1.5). The ideal, desired engineered variant should maintain ACE-2 binding but escape all five mAbs known to bind to XBB.1.5 to avoid triggering a memory immune response. As shown in the table, constructs S43 and S48 exhibit near-ideal construct behavior. When expressed as the full-length S protein and in the form of a trimerized™ anchored RBD, S43 shows conserved ACE-2 binding, while eliminating binding to all mAb groups except for 1-2, depending on the expression environment. Construct S48 performs particularly well when expressed as a trimerized™ anchored RBD, demonstrating conserved ACE-2 binding and binding to only one of the five mAbs.
[0579] Table 13. Summary of some test data.
[0580]
[0581] x. Example 10: Predictive Mouse Immunity Study
[0582] This embodiment describes the planning procedures and expected results of an experiment for evaluating whether an RNA composition encoding an engineered antigen comprising the SARS-CoV-2S protein and / or a portion thereof (e.g., the RBD domain) described herein induces an immune response characterized by increased primary B cell activation and / or decreased memory B cell activation in vaccinated subjects (mice in this embodiment).
[0583] Go to Figure 15A Vaccine candidates containing RNA encoding engineered antigens, designed and screened according to the methods described herein, will be administered to mice previously exposed to the full-length SARS-CoV-2S protein. Specifically, as Figure 15A As shown, the vaccine candidates were tested in mice that had previously been given two doses of RNA encoding the full-length SARS-CoV-2S protein, and each engineered antigen vaccine candidate was administered as a third dose (booster).
[0584] The mice were divided into several groups, each containing approximately the same number of members (x), and then... Figure 15A The dosing regimen is shown. The total number of groups depends on the number of engineered antigen candidates to be tested and their different expression forms. For example, Figure 15 shows an exemplary immunoassay in which engineered antigen designs with sequences described in Table 9A, S43 and S48, are evaluated in the context of mRNA encoding (i) a full-length S protein (similar to BNT162b2, but encoding a specific (e.g., S43 or S48) engineered XBB.1.5 variant) and (ii) a trimerized™-anchored RBD domain. Similarly, other candidate antigens and / or expression backgrounds may be included. In one example, the number of mice in each group is approximately 7-8. Figure 15A As shown, each group of mice was first administered two doses of a monovalent composition containing RNA encoding the SARS-CoV-2S protein of the wild-type strain (BNT162b2). The first and second doses of the monovalent vaccine were administered 21 days apart. Five (5) weeks after the first dose, the mice were grouped based on their neutralizing titer against the wild-type strain (e.g., the mice were divided into several groups such that the average neutralizing titer was approximately the same in each group). In some embodiments, mice may be assigned to groups based on pseudovirus neutralizing titers (e.g., as shown in [examples of other methods]). Figure 15A (As shown).
[0585] Then, 18 weeks after the first dose of vaccine is administered (i.e., on day 126), Figure 15A As shown), each candidate vaccine is administered as a third dose. In some implementations, Figure 15A Versions of the immunization study method shown may administer a third dose of the candidate vaccine shortly after the first two doses. Not bound by any particular theory, a third dose of the candidate vaccine may be administered rapidly after the second dose, provided sufficient time is allowed for an immune response to occur, such as the generation of B cells against the initial antigen (e.g., wild-type antigen). In some implementations, 28 days (4 weeks) or longer may be sufficient; for example, it may be permissible to administer the first dose (BNT162b2) on day 0, the second dose (BNT162b2) on day 21, and the vaccine candidate as a third dose (e.g., a booster) four weeks later (e.g., on day 49 (week 7) or later).
[0586] Various engineered antigen designs and expression environment formats can be administered as a third agent and compared. For example, such as Figure 15AAs shown, candidate drugs S43 and S48 can be administered as a third dose, in the whole spike (S) form - BNT162b2 (S43) and BNT162b2 (S48), and in the trimerized transmembrane anchored RBD form - RBD-TM (S43) and RBD-TM (S48). In some embodiments, numerous control and / or reference vaccines may also be used. For example, such as Figure 15A As shown, a third dose of BNT162b2 can be administered. Alternatively or concurrently, the full spike (S) and RBD-TM versions of the parental XBB.1.5 variant can be administered for comparison with its engineered variant. In some implementations, a test group may not receive any third dose. Table 14A below lists a set of example vaccine candidates (selected based on screening data described herein, such as in the preceding examples), some of which are also... Figure 15A The following are listed. Table 14B lists several alternative candidate vaccines that can be used separately or as alternatives. Tables 14A and 14B provide descriptions of the vaccine candidate formats and references to exemplary sequences included in Tables 14C-14I. Other vaccine candidates comprising engineered antigens (e.g., variants of other VOCs) and a reference composition (vaccine dose) can be prepared and evaluated in a manner similar to that described herein with respect to XBB.1.5.
[0587] Table 14A. Vaccine candidates administered to vaccinated mice.
[0588]
[0589]
[0590] Table 14B: Additional vaccine candidate options.
[0591]
[0592]
[0593]
[0594]
[0595]
[0596]
[0597]
[0598]
[0599]
[0600]
[0601]
[0602]
[0603]
[0604]
[0605]
[0606]
[0607]
[0608]
[0609]
[0610]
[0611]
[0612]
[0613]
[0614]
[0615]
[0616]
[0617]
[0618]
[0619]
[0620]
[0621]
[0622]
[0623]
[0624]
[0625]
[0626]
[0627]
[0628]
[0629]
[0630]
[0631]
[0632]
[0633]
[0634]
[0635]
[0636]
[0637]
[0638]
[0639]
[0640]
[0641]
[0642]
[0643]
[0644]
[0645]
[0646]
[0647]
[0648]
[0649]
[0650]
[0651] Blood samples will be collected before administration of the first dose of RNA and at 3, 5, 9, 13, 17, 18, 19, 22, 26, 30, 34, 35, 37, and 39 weeks after administration of the first dose of RNA. Mice will be sacrificed at 39 weeks after administration of the first dose of RNA, and final blood, lymph node, and spleen samples will be collected for analysis.
[0652] Blood sample analysis
[0653] Blood samples will be screened for antibody titers that bind to and neutralize various SARS-CoV-2 strains and variants (e.g., using the ELISA and pseudovirus assays described herein). Each spleen sample can be analyzed individually. For lymph node samples, samples from two mice can be pooled for analysis.
[0654] Analysis of spleen and lymph node samples
[0655] B cells were isolated from the spleen and lymph nodes and subjected to phenotypic analysis to determine B cell type (e.g., primary, memory, or plasma) and binding specificity (e.g., specificity for the S protein of various SARS-CoV-2 strains and variants). Phenotypic analysis can be performed using... Figure 16A and 16B The similar methods described in the literature, based on FACS and consumption assays, are further described in detail in Quandt and Muik et al., Science Immunol., 7(75)eabq2427(2022) (doi / 10.1126 / sciimmunol.abq2427), the contents of which are incorporated herein by reference in their entirety. The pseudovirus neutralizing titer for each serum sample will also be collected. BCR library analysis will also be performed on B cells isolated from spleen samples.
[0656] For RBD binding assays, samples from all mice can be combined to generate enough samples to perform these methods. Each sample can be screened for negative, wild-type specific, XBB.1.5-specific binders, as well as cross-reactivity between wild-type and XBB.1.5 binding.
[0657] Among other things, the protocol described in this embodiment can be used to characterize the B-cell binding specificity of subjects receiving a booster vaccine, the vaccine delivering an antigen of interest—here, the XBB.1.5S protein or an immunogenic portion thereof. Specifically, this experimental protocol can be used to determine the relative number of XBB.1.5-specific B cells and / or which portions of the XBB.1.5S protein these B cells recognize. These results can be used to assess the impact of immunoblotting on various vaccine candidates, and the ability of certain vaccine candidates to evade immunoblotting and induce de novo immunization responses.
[0658] Vaccine candidates that are less susceptible to immunoblotting (i.e. are more likely to produce de novo reactions) may have one or more of the following characteristics: (i) an increased proportion of B cells against the target variant delivered by the vaccine candidate (i.e., XBB.1.5 or its immunogenic portion), (ii) an increased neutralizing titer against the target variant encoded by the vaccine candidate, and / or (iii) an increased number of B cell receptors that recognize epitopes unique to the antigen encoded by the vaccine candidate (i.e., increased B cell breadth).
[0659] Additional analytical techniques include spleen sample analysis, lymph node analysis, and blood sample characterization, such as... Figure 17 As shown. Spleen sample analysis may involve the preparation of single-cell suspensions, with results showing >5 × 10⁵ cells after separation. 7 The spleen and red blood cells (RBCs) were consumed (cell viability was 80-90%). Approximately 1.5 × 10⁻⁶ cells were isolated from these cells. 7 One leukocyte / spleen can be used for FACS phenotypic analysis for immunogenicity testing. This analysis may involve, for example, primary, memory, and plasma cells specific to different S proteins. Approximately 1.5 × 10⁷ leukocytes / spleen can be used for labeling and pooling in the spleen isolate, followed by magnetically activated cell sorting (MACS). Subsequent steps may involve FACS sorting and staining of memory and plasma cell populations. Another subsequent step may include BCR library analysis. Residual leukocytes extracted from the spleen can be used for enzyme-linked immunosorbent assay (ELISpot) and freezing of the remaining sample or freezing and ELISpotting. Lymph node analysis may involve the preparation of single-cell suspensions, approximately 1 × 10⁷ per inguinal (iLN) and pelvic (pLN) lymph node. 6 The analysis may involve FACS phenotypic analysis for immunogenicity testing. This analysis may involve, for example, primary, memory, and plasma cells specific for different S proteins. Finally, blood sample characterization may involve a pseudovirus neutralization assay (pVNT) ELISA.
[0660] Immune response in unvaccinated mice
[0661] Go to Figure 15B In some implementations, the vaccine candidate and the reference vaccine can be tested in unvaccinated mice. For example... Figure 15B As shown, it can be tested on vaccine-naïve mice on a shorter timescale and can be used to evaluate and / or confirm that vaccine candidates based on various engineered antigen compounds can spontaneously elicit a meaningful immunogenic response to XBB.1.5. Figure 15BAn example protocol is shown in which the vaccine candidate described herein and the reference vaccine are administered in two doses to unvaccinated mice, with the first dose administered on day zero and the second dose administered three weeks later, on day 21.
[0662] Blood samples can be collected before administration of the first dose of RNA and at 2, 3, 4, and 7 weeks after administration of the first dose of RNA. Seven (7) weeks after administration of the first dose of RNA, mice are sacrificed, and final blood, lymph node, and spleen samples are collected for analysis. The generation of neutralizing antibodies can be tested using the XBB.1.5 pseudovirus via pVNT assay, while binding antibodies can be tested using XBB.1.5RBD as a target via ELISA. An ideal candidate will successfully maintain neutralizing antibody titers, for example, showing levels comparable to parental XBB.1.5 (e.g., RBD-TM(XBB.1.5)), but simultaneously exhibiting reduced binding antibody titers in ELISA assays, for example, due to disruption of conserved regions.
[0663] xi. Example 11: Evaluation and selection of a single engineered RBD mutation designed for engineered antigens
[0664] This embodiment describes the evaluation and selection of RBD mutations introduced into certain engineered antigens described herein. Specifically, mutations present in various constructs are individually evaluated, and subgroups are selected to create additional fine-tuned engineered antigens.
[0665] Specifically, analysis of the various computer-simulated constructs presented in this paper led to the identification of certain mutations introduced in construct S48 that affect expression. As explained in this paper, construct S48 corresponds to the SARS-CoV-2S protein RBD, which has (i) the XBB.1.5 signature mutation and (ii) an additional group of introduced mutations engineered to disrupt conserved regions of the baseline, reference, and XBB.1.5 reference antigen trigger memory. As shown in Table 9A, the additional group of introduced mutations in construct S48 includes the following mutations:
[0666] • Additional S48 mutations: L335F, K356T, P384S, L390R, T430I, F464Y, and H519N.
[0667] Of these additional S48 mutations, one subgroup was identified as beneficial for expression in constructs containing the S48 RBD design, while another subgroup was selected for further evaluation, e.g., mutations that may not be optimal and / or lead to reduced expression. Mutations identified as beneficial include, but are not limited to, K356T, P384S, T43OI, F464Y, and H519N. Mutations selected for further evaluation include L335F and L390R.
[0668] Based on these identified mutations, an additional round of construct designs was created based on the XBB.1.5RBD reference antigen (i.e., including the XBB.1.5 signature mutation), wherein (i) mutations identified as beneficial to the group were retained, (ii) two mutations selected for further evaluation were excluded from most constructs (and none of them contained both L335F and L390R mutations), and (iii) new mutations of the additional group were introduced.
[0669] Following a method similar to that described in Example 3 and referring to Figure 5G The process 500 introduces new mutations.
[0670] Figure 18 The custom constructs for this additional group are shown. Construct design labels are displayed vertically, with each row corresponding to a specific construct design, and a list of possible mutations is displayed horizontally. XBB.1.5 signature mutations are marked with a black asterisk ("*"), mutations identified as beneficial are marked with a green asterisk, two mutations selected for further evaluation are marked with a red downward-pointing triangle, and new mutations are marked with a purple diamond. Shading (dark green) identifies mutations present in a specific construct. Figure 18 As shown, all constructs contain the XBB.1.5 signature mutation and five mutations identified as beneficial. Constructs S122-S127 and S128 contain the L335F mutation, and constructs S145 and S156 contain the L390R mutation. Figure 18 As shown, new mutations were introduced in various constructs. Tables 15A and 15B below list the RBD mutations (including the XBB.1.5 signature mutation) and additional engineered RBD mutations (i.e., in addition to the reference XBB.1.5 mutation) for each construct design.
[0671] Then, by incorporating additional mutations (i.e., as listed in Table 15 below) and the XBB.1.5 signature mutation into the BNT162b3 mRNA for in vitro testing, the results were evaluated. Figure 18 The construct designs shown below and identified in Tables 15A and 15B were tested. Specifically, Figure 21 The BNT162b3 construct shown in A is a 1397-base-pair mRNA that encodes the membrane-anchored RBD and fibrous substitute protein domain (F), as well as a viral signal peptide. Figure 19A As shown, the construct includes a secretion signal (“sec”), an RBD domain (“RBD”), a fibrous alternative protein domain (“F”), and a transmembrane anchor (“TM”). For each construct design, the RBD coding domain is modified to encode the XBB.1.5 signature mutation and additional mutations specific to each construct design.
[0672] To evaluate expression, ACE-2 binding was assessed using flow cytometry via an hACE2-mFc binding assay, which targets ACE-2 expression. Figure 18 The 22 construct designs shown in the diagram and listed in Table 15, as well as the XBB.1.5 and S48 constructs, are used as reference points. Results are shown in... Figure 19B And Table 16A below. To evaluate immune escape (e.g., the ability of a construct to avoid antibodies generated by a previously wild-type induced immune response and thus potentially trigger a de novo immune response), binding to polyclonal vaccine serum from patients receiving triple BNT162b2 vaccination was measured. Results for 22 construct designs, as well as for the XBB.1.5 and S48 constructs, are shown below. Figure 19C And in Table 16B. The values in Tables 16A and 16B were obtained by repeating the experiment twice and the results are presented as area under the curve (AUC) values and flow cytometry intensity (FC).
[0673] like Figure 19B As shown in Table 16A, Figure 18 All 22 construct designs in Tables 15A and 15B showed higher ACE2 binding than the S48 construct design, indicating higher expression. Regarding immune escape, the S48 construct design showed the most significant decrease in binding to serum from patients receiving the BNT162b2 triple vaccine. Several construct designs listed in Tables 15A and 15B showed significantly reduced binding to (BNT162 serum). Specifically, seven constructs showed a decrease in serum binding of more than 1.5-fold, while ACE2 binding decreased by no more than two (2)-fold. These construct designs are identified by an asterisk (*) in Tables 16A and 16B below.
[0674] Figure 20A -C indicates XBB.1.5, S48 and Figure 18 The seven construct designs shown exhibit binding affinity to various SARS-CoV-2 monoclonal antibodies. Figure 20A Results for five antibodies are shown, with seven construct designs demonstrating complete restoration of binding. Figure 20B Results for five antibodies are shown, with seven construct designs demonstrating partial restoration of binding. Figure 20C Results for two antibodies are shown, with seven constructs exhibiting little to no binding. The results are consistent with polyclonal serum data, with S48 showing the most significant decrease in binding to the studied monoclonal antibody.
[0675] Table 15A. RBD mutations used in engineered antigen design.
[0676]
[0677]
[0678] Table 15B. RBD mutations used in engineered antigen design.
[0679]
[0680] Table 16A. Results were evaluated using hACE2-mFC binding assays. Area under the curve (AUC) and FC intensity (relative to XBB.1.5 = 1.00) values are provided.
[0681]
[0682]
[0683] Table 16B. Immune escape results, evaluated by binding assays in the serum of patients triple-vaccinated with BNT162b3. Area under the curve (AUC) and FC intensity (relative to XBB.1.5 = 1.00) values are provided.
[0684]
[0685] G. Computer systems and network environment
[0686] like Figure 21 As shown, an implementation of the network environment 2100 used to provide the systems and methods described herein is illustrated and described. A brief overview is provided below, now referring to... Figure 21 A block diagram of an exemplary cloud computing environment 2100 is shown and described. The cloud computing environment 2100 may include one or more resource providers 2102a, 2102b, 2102c (collectively referred to as 2102). Each resource provider 2102 may include computing resources. In some implementations, computing resources may include any hardware and / or software for processing data. For example, computing resources may include hardware and / or software capable of executing algorithms, computer programs, and / or computer applications. In some implementations, exemplary computing resources may include application servers and / or databases with storage and retrieval capabilities. Each resource provider 2102 may connect to any other resource provider 2102 in the cloud computing environment 2100. In some implementations, resource providers 2102 may be connected via a computer network 2108. Each resource provider 2102 may be connected via the computer network 2108 to one or more computing devices 2104a, 2104b, 2104c (collectively referred to as 2104).
[0687] The cloud computing environment 2100 may include a resource manager 2106. The resource manager 2106 can be connected to resource providers 2102 and computing devices 2104 via a computer network 2108. In some implementations, the resource manager 2106 may facilitate one or more resource providers 2102 to provide computing resources to one or more computing devices 2104. The resource manager 2106 may receive requests for computing resources from a particular computing device 2104. The resource manager 2106 may identify one or more resource providers 2102 capable of providing the computing resources requested by the computing device 2104. The resource manager 2106 may select a resource provider 2102 to provide computing resources. The resource manager 2106 may facilitate a connection between the resource provider 2102 and the particular computing device 2104. In some implementations, the resource manager 2106 may establish a connection between the particular resource provider 2102 and the particular computing device 2104. In some implementations, the resource manager 2106 may redirect a particular computing device 2104 to a particular resource provider 2102 that has the requested computing resources.
[0688] Figure 22 Examples of computing devices 2200 and mobile computing devices 2250 that can be used to implement the techniques described in this disclosure are shown. Computing device 2200 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Mobile computing device 2250 is intended to represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are intended only as examples and are not intended to be limiting.
[0689] The computing device 2200 includes a processor 2202, a memory 2204, a storage device 2206, a high-speed interface 2208 connecting the memory 2204 and multiple high-speed expansion ports 2210, and a low-speed interface 2212 connecting a low-speed expansion port 2214 and the storage device 2206. Each of the processor 2202, memory 2204, storage device 2206, high-speed interface 2208, high-speed expansion port 2210, and low-speed interface 2212 is interconnected using various buses and can be mounted on a common motherboard or otherwise suitably mounted. The processor 2202 can process instructions executed within the computing device 2200, including instructions stored in the memory 2204 or storage device 2206, to display graphical information of a GUI on an external input / output device, such as a display 2216 coupled to the high-speed interface 2208. In other implementations, multiple processors and / or multiple buses, as well as multiple memories and various types of memory, may be used as appropriate. Furthermore, multiple computing devices can be connected together, with each device providing a portion of the necessary operation (e.g., as a server group, a set of blade servers, or a multiprocessor system). Therefore, as used herein, the description of multiple functions as implemented by a “processor” covers implementations where multiple functions are implemented by any number of processors (one or more) of any number of computing devices (one or more). Moreover, when a function is described as implemented by a “processor,” this covers implementations where said function is implemented by any number of processors (one or more) of any number of computing devices (one or more) (e.g., in a distributed computing system).
[0690] Memory 2204 stores information within computing device 2200. In some implementations, memory 2204 is one or more volatile memory cells. In some implementations, memory 2204 is one or more non-volatile memory cells. Memory 2204 can also be another form of computer-readable medium, such as a magnetic disk or optical disk.
[0691] Storage device 2206 provides large-capacity storage for computing device 2200. In some implementations, storage device 2206 may be or contain computer-readable media such as floppy disk devices, hard disk devices, optical disk devices, or magnetic tape devices, flash memory, or other similar solid-state storage devices or device arrays, including those in storage area networks or other configurations. Instructions may be stored in an information carrier. When executed by one or more processing devices (e.g., processor 2202), the instructions implement one or more methods, such as those described above. Instructions may also be stored by one or more storage devices, such as computer or machine-readable media (e.g., memory 2204, storage device 2206, or memory on processor 2202).
[0692] High-speed interface 2208 manages bandwidth-intensive operations of computing device 2200, while low-speed interface 2212 manages lower bandwidth-intensive operations. This functional allocation is merely illustrative. In some implementations, high-speed interface 2208 is coupled to memory 2204, display 2216 (e.g., via a graphics processor or accelerator), and high-speed expansion port 2210 which accepts various expansion cards (not shown). In this implementation, low-speed interface 2212 is coupled to storage device 2206 and low-speed expansion port 2214. Various communication ports (e.g., USB, Bluetooth) may be included. The low-speed expansion port 2214 (such as Ethernet, wireless Ethernet) can be coupled to one or more input / output devices (such as keyboards, pointing devices, scanners), or, for example, to network devices such as switches or routers via a network adapter.
[0693] The computing device 2200 can be implemented in a variety of different forms as shown in the figure. For example, it can be implemented as a standard server 2220, or multiple times in a group of such servers. Alternatively, it can be implemented in a personal computer such as a laptop computer 2222. It can also be implemented as part of a rack server system 2224. Alternatively, components of the computing device 2200 can be integrated with other components (such as a mobile computing device 2250) in a mobile device (not shown). Each of such devices may include one or more of the computing device 2200 and the mobile computing device 2250, and the entire system may consist of multiple computing devices communicating with each other.
[0694] Mobile computing device 2250 includes processor 2252, memory 2264, input / output devices (such as display 2254), communication interface 2266, and transceiver 2268, as well as other components. Mobile computing device 2250 may also be equipped with storage devices (such as microdrives or other devices) to provide additional storage. Each of the processor 2252, memory 2264, display 2254, communication interface 2266, and transceiver 2268 is interconnected using various buses, and multiple components may be mounted on a common motherboard or otherwise suitably mounted.
[0695] Processor 2252 can execute instructions within mobile computing device 2250, including instructions stored in memory 2264. Processor 2252 can be implemented as a chipset including multiple separate analog and digital processors. Processor 2252 can provide, for example, coordination of other components of mobile computing device 2250, such as control of user interface, applications running on mobile computing device 2250, and wireless communication of mobile computing device 2250.
[0696] Processor 2252 can communicate with the user via control interface 558 and display interface 2256 coupled to display 2254. Display 2254 can be, for example, a TFT (Thin Film Transistor Liquid Crystal Display) or OLED (Organic Light Emitting Diode) display, or other suitable display technologies. Display interface 2256 may include suitable circuitry for driving display 2254 to present graphics and other information to the user. Control interface 2258 can receive commands from the user and translate the commands for submission to processor 2252. Furthermore, external interface 2262 can provide communication with processor 2252 to enable near-field communication between mobile computing device 2250 and other devices. External interface 2262 may provide wired communication in some implementations, or wireless communication in others, and multiple interfaces may be used.
[0697] Memory 2264 stores information within mobile computing device 2250. Memory 2264 may be implemented as one or more computer-readable media, one or more volatile memory cells, or one or more non-volatile memory cells. Extended memory 2274 may also be provided and connected to mobile computing device 2250 via extended interface 2272, which may include, for example, a SIMM (Single In-line Memory Module) card interface. Extended memory 2274 may provide additional storage space for mobile computing device 2250, or it may store applications or other information of mobile computing device 2250. Specifically, extended memory 2274 may include instructions for performing or supplementing the above processes, and may also include security information. Therefore, for example, extended memory 2274 may be provided as a security module of mobile computing device 2250 and may be programmed using instructions that allow secure use of mobile computing device 2250. Additionally, secure applications and additional information, such as identification information placed on the SIMM card in a manner resistant to hacking, may be provided via a SIMM card.
[0698] The memory may include, for example, flash memory and / or NVRAM (non-volatile random access memory), as discussed below. In some implementations, instructions are stored in an information carrier. When executed by one or more processing devices (e.g., processor 2252), the instructions implement one or more methods, such as those described above. Instructions may also be stored by one or more storage devices, such as one or more computer or machine-readable media (e.g., memory 2264, extended memory 2274, or memory on processor 2252). In some implementations, instructions may be received by propagation signals, for example, via transceiver 2268 or external interface 2262.
[0699] Mobile computing device 2250 can wirelessly communicate via communication interface 2266, which may include digital signal processing circuitry if necessary. Communication interface 2266 can provide communication under various modes or protocols, such as GSM voice calls (Global System for Mobile Communications), SMS (Short Message Service), EMS (Enhanced Messaging Service) or MMS messages (Multimedia Messaging Service), CDMA (Code Division Multiple Access), TDMA (Time Division Multiple Access), PDC (Personal Digital Cellular), WCDMA (Wideband Code Division Multiple Access), CDMA2000, or GPRS (General Packet Radio Service). For example, such communication can be performed using a radio frequency transceiver 2268. Alternatively, Bluetooth or similar technologies can be used. Wi-Fi TM Or other such transceivers (not shown) for short-range communication. Furthermore, the GPS (Global Positioning System) receiver module 2270 can provide additional navigation and location-related wireless data to the mobile computing device 2250, which can be used as needed by applications running on the mobile computing device 2250.
[0700] Mobile computing device 2250 can also communicate audibly using audio codec 2260, which receives spoken information from a user and converts it into usable digital information. Audio codec 2260 can also generate audible sounds for the user, such as through a speaker in, for example, the handset of mobile computing device 2250. Such sounds may include sounds from voice telephone calls, recorded sounds (e.g., voice messages, music files, etc.), and sounds generated by applications operating on mobile computing device 2250.
[0701] The mobile computing device 2250 can be implemented in a variety of different forms as shown in the figure. For example, it can be implemented as a cellular phone 2280. It can also be implemented as a smartphone 2282, a personal digital assistant, or part of other similar mobile devices.
[0702] Various implementations of the systems and techniques described herein can be implemented using digital electronic circuits, integrated circuits, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be dedicated or general-purpose and coupled to receive and transmit data and instructions from a storage system, at least one input device, and at least one output device.
[0703] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and may be implemented in high-level programming languages and / or object-oriented programming languages, and / or in assembly language / machine language. As used herein, the terms machine-readable medium and computer-readable medium refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term machine-readable signal refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0704] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and pointing device (e.g., a mouse or trackball) by which the user can provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form, including sound, speech, or tactile input.
[0705] The systems and technologies described herein can be implemented in computing systems that include back-end components (e.g., as data servers), middleware components (e.g., application servers), or front-end components (e.g., client computers having a graphical user interface or web browser through which users can interact with implementations of the systems and technologies described herein), or any combination of such back-end, middleware, or front-end components. Components of the system can be interconnected via any form or medium of digital data communication (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0706] A computing system may include clients and servers. Clients and servers are typically geographically separated and usually interact through a communication network. The client-server relationship is established by computer programs running on their respective computers that create a client-server relationship between them.
[0707] In some implementations, the various modules described herein may be separated, combined, or incorporated into a single or merged module. The modules depicted in the figures are not intended to limit the system described herein to the software architecture shown therein.
[0708] Equivalent solution
[0709] Elements from the different embodiments described herein can be combined to form other embodiments not specifically described above. Elements can be omitted from the processes, computer programs, databases, etc., described herein without adversely affecting their operation. Furthermore, the logical flows depicted in the figures do not require the desired results to be achieved in the specific or sequential order shown. Various individual elements can be combined into one or more individual elements to implement the functions described herein.
[0710] Throughout this specification, where devices and systems are described as having, including, or comprising specific components, or where methods are described as having, including, or comprising specific steps, it is envisioned that additional devices and systems substantially composed of or comprised of said components, as well as methods substantially composed of or comprised of said processing steps, as covered by the subject matter of this invention, exist.
[0711] It should be understood that the order of steps or the sequence of certain operations is not important as long as the invention remains operational. Furthermore, two or more steps or actions may be performed simultaneously.
[0712] While the invention has been specifically shown and described with reference to preferred embodiments, those skilled in the art will understand that various changes in form and detail may be made therein without departing from the spirit and scope of the invention as covered by the appended claims.
Claims
1. A computer simulation design method for engineered antigens, the method comprising: (a) Receive and / or access a polypeptide model representing a reference antigen of an infectious pathogen via the processor of a computing device; (b) The processor identifies one or more conserved regions in the polypeptide model that trigger memory, the conserved regions representing conserved portions of the reference antigen that are identified as potentially triggering a memory immune response. (c) Generating one or more amino acid modifications within at least a portion of the one or more conserved regions by the processor, thereby producing a polypeptide model representing the disruption of the engineered antigen; and (d) The processor stores and / or provides the disrupted peptide model for display and / or further processing.
2. The method of claim 1, wherein the reference antigen is or comprises at least a portion of a naturally occurring variant of a viral protein.
3. The method of claim 2, wherein the reference antigen is or comprises at least a portion of the SARS-Cov2 spike polypeptide.
4. The method of claim 3, wherein the reference antigen is or comprises at least a portion of a specific SARS-CoV-2 variant spike polypeptide.
5. The method of claim 4, wherein the specific SARS-CoV-2 variant is a member of the Omicron and / or XBB lineage classification.
6. The method of claim 1 or 2, wherein the infectious pathogen is or comprises an RNA virus, and the reference antigen is or comprises at least a portion of its protein.
7. The method of claim 1, wherein the reference antigen is or comprises a bacterial protein.
8. The method of claim 1, wherein the reference antigen is or contains an antigen of a parasite.
9. The method as described in any of the preceding claims, wherein the one or more conserved regions that trigger memory represent portions of the reference antigen that are substantially similar to (i) one or more variants thereof and / or (ii) an initial / wild-type strain.
10. The method of any of the preceding claims, wherein the reference antigen is a specific target SARS-CoV-2 variant S polypeptide or a portion thereof, and wherein the one or more conserved regions that trigger memory represent portions of the reference antigen that are substantially similar to (i) one or more other SARS-CoV-2 variant polypeptides and / or (ii) corresponding portions of wild-type SARS-CoV-2 polypeptides.
11. The method of any of the preceding claims, wherein the one or more conserved regions that trigger memory are or comprise a set of conserved epitope regions representing known, non-mutated epitopes present on the reference antigen.
12. The method of claim 11, wherein step (b) comprises: The processor acquires data corresponding to a set of known epitopes and identifies each of one or more specific known epitopes in the set within the reference antigen; The processor acquires the identification of a set of signature mutations of the reference antigen; as well as The processor identifies specific known epitopes that correspond to portions of the reference antigen and do not contain any signature mutations as the group of conserved epitope regions.
13. The method of claim 12, wherein the group of known table positions includes one or more of the positions listed in Table 2A.
14. The method of claim 12, wherein the group of known table positions includes one or more of the table positions listed in Table 2B.
15. The method of any of the preceding claims, wherein the one or more conserved regions that trigger memory are or contain a conserved surface representing a non-mutated surface of the reference antigen.
16. The method as described in any of the preceding claims, the method comprising identifying one or more signature mutations of the reference antigen and / or accessing its identification information.
17. The method of claim 16, wherein the reference antigen is or comprises the aSARS-Cov2XBB.1.5 spike protein.
18. The method as claimed in any of the preceding claims, wherein step (c) comprises selecting the one or more amino acid modifications from a set of permissible mutations by means of the processor.
19. The method of claim 18, wherein the group of permitted mutations is or is included in a group of related antigens when multiple mutations are observed.
20. The method of claim 19, wherein the reference antigen is a specific variant of the SARS-CoV-2 protein, and the group-related antigen includes corresponding proteins of other related variants.
21. The method of claim 20, wherein the reference antigen is a member of the Omicron lineage, and the group-associated antigen includes the observed variant belonging to the Omicron lineage.
22. The method of any one of claims 19-21, wherein the group allows mutations to include at least a portion of the mutations listed in Table 3.
23. The method of claim 22, wherein the group allows mutations to include at least a portion of the mutations listed in Table 3, excluding one or both of L335F and L390R.
24. The method of any one of claims 19-23, wherein the reference antigen is a SARS-CoV-2 protein, and the group-related antigen includes corresponding proteins of other related coronaviruses.
25. The method of any of the preceding claims, wherein the one or more conserved regions that trigger memory are or comprise a set of conserved epitope regions, and step (c) comprises introducing at least one amino acid modification into each of at least a portion of the conserved epitope regions.
26. The method of any of the preceding claims, wherein the one or more conserved regions that trigger memory are or comprise a conserved surface, and step (c) includes generating the one or more amino acid modifications at locations distributed throughout / across the conserved surface.
27. The method as claimed in any of the preceding claims, the method comprising repeatedly performing steps (b) and (c) to generate a plurality of candidate peptide models, each candidate peptide model representing a candidate engineered variant.
28. The method of claim 27, the method comprising determining, by the processor, values of one or more performance scores for each of the candidate peptide models, and selecting a subgroup of candidate peptide models based at least in part on the determined performance score values.
29. The method of claim 28, wherein the one or more performance scores include one or both of the following: (a) Immune escape score, indicating the likelihood and / or relative ability of a particular candidate engineered variant to be recognized and neutralized by the antibody, and (b) Fitness score, indicating the likelihood and / or viability of a particular candidate engineered variant.
30. The method of claim 29, wherein determining one or both of (a) the immune escape score and (b) the fitness score comprises using a machine learning model.
31. The method of claim 29 or 30, wherein determining one or both of (a) the immune escape score and (b) the fitness score comprises a 3D structural model using at least a portion of the particular candidate variant.
32. The method of any one of claims 27 to 31, wherein the one or more performance scores include a positional expansion score, the score measuring the degree to which the amino acid modification is uniformly distributed on the surface of the candidate engineered variant.
33. The method of any one of claims 27 to 32, wherein the one or more performance scores include a mutation co-occurrence score.
34. The method of any of the preceding claims, the method comprising identifying, by means of the processor, one or more target regions representing the reference antigen portion to be retained within the polypeptide model, and excluding the one or more target regions from the one or more conserved regions that trigger memory.
35. The method of any of the preceding claims, the method comprising presenting a polypeptide model of said destruction caused by said processor for graphical display.
36. The method as described in any of the preceding claims, wherein the method comprises generating a corresponding RNA sequence from the disrupted polypeptide model.
37. The method as claimed in any of the preceding claims, the method comprising producing a composition comprising a polypeptide based on the disrupted polypeptide model.
38. The method as described in any of the preceding claims, wherein the method comprises evaluating the bioactivity of the engineered antigen in vitro.
39. The method of claim 38, wherein the bioactivity of the engineered antigen is characterized by: The engineered antigen is correctly expressed and folded; and / or The engineered antigen does not bind to antibodies that bind to the reference antigen; and / or The pseudovirus loaded with the engineered antigen can enter cells; and / or The engineered antigen is immunogenic, and / or The engineered antigen reduces the activation of the B-cell memory immune response to the reference antigen by the engineered antigen.
40. The method of any of the preceding claims, the method comprising producing a composition comprising a nucleic acid encoding the amino acid sequence represented by the disrupted polypeptide model.
41. A vaccine composition comprising the polypeptide and / or nucleic acid as described in claim 36 or 37.
42. A method of vaccination comprising administering the vaccine of claim 41 to a subject or a group of subjects.
43. A system comprising a processor of a computing device and a memory storing instructions thereon, wherein when executed by the processor, the instructions cause the processor to perform the method of any one of claims 1 to 36.
44. A method for manufacturing an immunogenic composition, the method comprising: Compare the sequences of viral proteins from different variants of infectious disease pathogens to identify residual conserved sites in target antigens; Replacing at least one or more of the remaining conserved sites with a sequence characterized by generating a new sequence; and Producing a vaccine that delivers at least a portion of the new sequence, including at least one replaced residual conserved site.
45. An RNA comprising a nucleotide sequence encoding an engineered antigen, wherein the engineered antigen corresponds to a specific reference antigen that has been altered to introduce one or more amino acid modifications into at least a portion of one or more conserved regions that trigger memory, the conserved regions that trigger memory being identified as portions of the reference antigen, the reference antigen being determined to potentially trigger a memory immune response.
46. The RNA of claim 45, wherein the one or more conserved regions that trigger memory represent portions of the reference antigen substantially similar to one or more variants thereof.
47. The RNA of claim 45 or 46, wherein the one or more conserved regions that trigger memory are or comprise a set of conserved epitope regions representing known, non-mutated epitopes present on the reference antigen.
48. The RNA of claim 47, wherein the set of known epitopes includes one or more of the epitopes listed in Tables 2A and / or 2B.
49. The RNA of any one of claims 45 to 48, wherein the one or more conserved regions that trigger memory are or contain a conserved surface representing a non-mutated surface of the reference antigen.
50. The RNA of claim 49, wherein the reference antigen is or comprises at least a portion of an XBB.1.5 variant of the SARS-Cov2 spike protein.
51. The RNA of any one of claims 45 to 50, wherein the one or more amino acid modifications are selected from a group of permissible mutations.
52. The RNA of claim 51, wherein the group of permitted mutations is or is included in a group of related polypeptides when multiple mutations are observed.
53. The RNA of claim 52, wherein the group-associated polypeptide includes corresponding polypeptides of other related SARS-CoV 2 variants.
54. The RNA of claim 53, wherein the target polypeptide is a member of the Omicron lineage, and the associated polypeptides of the group include the corresponding polypeptides of the observed variants belonging to the Omicron lineage.
55. The RNA of any one of claims 53 to 54, wherein the one or more conserved regions that trigger memory are or comprise a set of conserved epitope regions, and the engineered antigen has at least one amino acid modification in each of at least a portion of the conserved epitope regions.
56. The RNA of any one of claims 45 to 55, wherein the one or more conserved regions that trigger memory are or comprise a conserved surface, and the one or more amino acid modifications occurring at locations distributed throughout / across the conserved surface.
57. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen comprises one or more combinations of mutations listed in Table 5A.
58. The method, system, vaccine composition, manufacturing method, or RNA as described in any of the preceding claims, wherein the engineered antigen comprises one or more sequences listed in Table 5B.
59. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen comprises one or more of the mutation combinations listed in Table 7A and / or Table 7B.
60. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen comprises one or more of the mutation combinations listed in Table 8A.
61. The method, system, vaccine composition, manufacturing method, or RNA as described in any of the preceding claims, wherein the engineered antigen comprises one or more sequences listed in Table 8B.
62. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen comprises one or more combinations of mutations listed in Table 9A.
63. The method, system, vaccine composition, manufacturing method, or RNA as described in any of the preceding claims, wherein the engineered antigen comprises a combination of mutations identified as S43 in Table 9A.
64. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen is or comprises at least a portion of the SARS-CoV-2S protein having at least a portion of the following mutations: N360D, P384S, L390R, T430I, F464Y, and H519N.
65. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen comprises a combination of mutations identified as S48 in Table 9A.
66. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen is or comprises at least a portion of the SARS-CoV-2S protein having at least a portion of the following mutations: L335F, K356T, P384S, L390R, T430I, F464Y, and H519N.
67. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen comprises at least a portion of the SARS-CoV-2S protein having mutant P384S, L390R, T430I, F464Y, and H519N.
68. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen comprises at least a portion of one or more SARS-CoV-2S proteins having mutant K356T, P384S, L390R, T430I, F464Y, and H519N.
69. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen does not include one or both of the mutants L335F and L390R.
70. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen comprises one or more combinations of mutations listed in Tables 15A and / or 15B.
71. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen comprises a combination of mutations identified as S123 in Table 15A and / or Table 15B.
72. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen is or comprises at least a portion of the SARS-CoV-2S protein having at least a portion of the following mutations: I332V L335F K356TP384S T430IL452Q F464Y H519N.
73. The method, system, vaccine composition, manufacturing method, or RNA of claim 72, wherein the engineered antigen comprises at least a portion of the following mutations: I332VL335F G339H R346T K356T L368I S371F S373PS375F T376AP348SD405N R408S K417N T430I N440K V445P G446S L452Q N460KF464YS477N T478K E484AF486P F490S Q498R N501Y Y505HE516Q H519N.
74. The method, system, vaccine composition, RNA, or manufacturing method as claimed in any of the preceding claims, wherein the engineered antigen comprises a combination of mutations identified as S122 in Tables 15A and / or 15B.
75. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen is or comprises at least a portion of the SARS-CoV-2S protein having at least a portion of the following mutations: L335F K356T P384ST430I L452R F464Y H519N.
76. The method, system, vaccine composition, manufacturing method, or RNA of claim 75, wherein the engineered antigen comprises at least a portion of the following mutations: L335FG339H R346T K356T L368I S371F S373P S375FT376A P348SD405NR408S K417N T430I N440K V445P G446S L452R N460K F464YS477NT478K E484A F486P F490S Q498R N501Y Y505H E516QH519N T523S.
77. The method, system, vaccine composition, RNA, or manufacturing method as claimed in any of the preceding claims, wherein the engineered antigen comprises a combination of mutations identified as S109 in Table 15A and / or Table 15B.
78. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen is or comprises at least a portion of the SARS-CoV-2S protein having at least a portion of the following mutations: K356T N360S P384SN388K T430IN450D F464Y H519N.
79. The method, system, vaccine composition, manufacturing method, or RNA of claim 78, wherein the engineered antigen comprises at least a portion of the following mutations: G339HR346T K356T L368I S371F S373P S375F T376AP348S N388K D405NR408S K417N T430I N440K V445P G446S N450D N460K F464YS477NT478K E484A F486P F490S Q498R N501Y Y505H H519NT523S.
80. The method, system, vaccine composition, RNA, or manufacturing method as claimed in any of the preceding claims, wherein the engineered antigen comprises a combination of mutations identified as S129 in Table 15A and / or Table 15B.
81. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen is or comprises at least a portion of the SARS-CoV-2S protein having at least a portion of the following mutations: K356T L335F P384SD389G T430IN450D F464Y H519N.
82. The method, system, vaccine composition, manufacturing method, or RNA of claim 81, wherein the engineered antigen comprises at least a portion of the following mutations: L335FG339H R346T K356T L368I S371F S373P S375FT376A P348SD389GD405N R408S K417N T430I N440K V445P G446S N450D L452RN460KF464Y S477N T478K E484A F486P F490S Q498R N501YY505H H519N.
83. The method, system, vaccine composition, RNA, or manufacturing method as described in any of the preceding claims, wherein the engineered antigen comprises a combination of mutations identified as S156 in Table 15A and / or Table 15B.
84. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen is or comprises at least a portion of the SARS-CoV-2S protein having at least a portion of the following mutations: K356T P384S L390RT430I N450D F464Y I472V H519N.
85. The method, system, vaccine composition, manufacturing method, or RNA of claim 84, wherein the engineered antigen comprises at least a portion of the following mutations: G339HR346T K356T L368I S371F S373P S375F T376AP348S L390R D405NR408S K417N T430I N440K V445P G446S N450D L452R N460KF464YI472V S477N T478K E484A F486P F490S Q498R N501YY505H H519N.
86. The method, system, vaccine composition, RNA, or manufacturing method as claimed in any of the preceding claims, wherein the engineered antigen comprises a combination of mutations identified as S112 in Tables 15A and / or 15B.
87. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen is or comprises at least a portion of the SARS-CoV-2S protein having at least a portion of the following mutations: K356T P384SD389GT430I N450D F464Y I468V H519N.
88. The method, system, vaccine composition, manufacturing method, or RNA of claim 87, wherein the engineered antigen comprises at least a portion of the following mutations: G339HR346T K356T L368I S371F S373P S375F T376AP348SD389G D405NR408S K417N T430I N440K V445P G446S N450D N460K F464YI468VS477N T478K E484A F486P F490S Q498R N501Y Y505HH519N.
89. The method, system, vaccine composition, RNA, or manufacturing method as claimed in any of the preceding claims, wherein the engineered antigen comprises a combination of mutations identified as S125 in Table 15A and / or Table 15B.
90. The method, system, vaccine composition, manufacturing method, or RNA as claimed in any of the preceding claims, wherein the engineered antigen is or comprises at least a portion of the SARS-CoV-2S protein having at least a portion of the following mutations: K356T L335F P384ST430I F464Y I468V H519N.
91. The method, system, vaccine composition, manufacturing method, or RNA of claim 90, wherein the engineered antigen comprises at least a portion of the following mutations: L335FG339H R346T K356T L368I S371F S373P S375FT376A P348SD405NR408S K417N T430I N440K V445P G446S N460K F464Y I468VS477NT478K E484A F486P F490S Q498R N501Y Y505H E516QH519N T523S.
92. A method of producing RNA comprising a nucleotide sequence encoding an engineered antigen corresponding to an engineered version of a reference antigen, the method comprising producing such RNA having a nucleotide sequence that, when compared with the reference antigen, shows a difference relative to the reference antigen in one or more conserved regions that trigger memory, the conserved regions being common to the reference antigen and (i) one or more pre-existing variants of the reference antigen and / or (ii) wild-type strains of the reference antigen.
93. The method of claim 92, wherein the reference antigen is or comprises a specific target variant of the SARS-CoV-2S protein RBD.
94. The method of claim 92 or 93, wherein the RNA is or comprises any one of claims 45 to 91.
95. The method of any one of claims 92 to 94, wherein the method comprises generating the RNA by in vitro transcription (IVT).
Citation Information
Patent Citations
Engineered coronavirus spike (s) protein and methods of use thereof
WO2021243122A2
Technologies for early detection of variants of interest
WO2022235847A1
Immunogen selection
WO2022235853A1