Method and system for identifying one or more candidate regions of one or more source proteins predicted to induce an immunogenic response, and method for producing a vaccine.

A computational method identifies candidate regions in source proteins to induce broad T-cell responses across diverse HLA types, addressing vaccine effectiveness and safety issues by predicting epitopes using statistical models and machine learning, enabling universal or personalized immunogenic responses.

JP2026065000APending Publication Date: 2026-04-14NEC ONCOIMMUNITY AS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-17
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing coronavirus vaccines face challenges in inducing a broad T-cell immune response due to human leukocyte antigen (HLA) polymorphism, leading to variable vaccine effectiveness across different populations, and there are concerns about antibody-dependent enhancement and safety issues with S protein-centric approaches.

Method used

A computer-aided method to identify candidate regions of source proteins that can induce an adaptive immunogenic response across a variety of HLA types by analyzing amino acid sequences and applying statistical models to predict epitopes, using tools like Immunoepitope Database and machine learning algorithms to ensure broad T-cell response.

Benefits of technology

This method allows for the identification of regions in source proteins that can stimulate a wide range of adaptive immune responses, potentially leading to universal vaccines or personalized treatments, while avoiding adverse reactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026065000000001_ABST
    Figure 2026065000000001_ABST
Patent Text Reader

Abstract

This invention provides a method, system, and vaccine production method for identifying candidate regions of source proteins that are predicted to induce adaptive immunogenic responses across multiple human leukocyte antigen (HLA) types. [Solution] The method implemented by the computer includes: step S201 accessing the amino acid sequence of a source protein; step S205 predicting the immunogenicity potential of multiple candidate epitopes within the amino acid sequence for each set of HLA types; step S211 generating a regional metric for each of multiple amino acid subsequences that indicates the predicted ability of the amino acid subsequence to induce an immunogenic response across the set of HLA types, based on the predicted immunogenicity potential; and step S213 applying a statistical model to identify amino acid subsequences having a statistically significant regional metric.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Introduction Vaccines have been established as an effective form of epidemiological control and have achieved great success in helping to reduce infections and the fatality rates associated with viral infections such as smallpox and polio. However, infections caused by other coronaviruses, such as severe acute respiratory syndrome coronavirus (SARS-CoV), SARS-CoV-2, and Middle East respiratory syndrome coronavirus (MERS-CoV), have proven more difficult to prevent with vaccination.

Background Art

[0002] Much of the worldwide effort to develop coronavirus vaccines to date has mainly focused on stimulating an antibody response against the exposed spike glycoprotein (S protein), which functions as the most exposed structural protein on the virus. However, the response against the S protein of SARS-CoV has been shown to confer short-term protection in mice (Yang et al. 2004, Nature 428(6982):561-4), but the neutralizing antibody response against the same structure in recovered patients is typically of low titer and short-lived (Channappanavar et al. 2014, Immunol Res 88(19):11034-44) (Yang et al. 2006, Clin Immunol 120(2)171-8). Furthermore, the induction of an antibody response against the S protein of SARS- CoV is associated with adverse effects in some animal models and may raise concerns about safety. For example, in the macaque model, it has been observed that anti-S protein antibodies are associated with severe acute lung injury (Liu et al. 2019 JCI Insight 4(4) ), while sera from SARS-CoV patients have also revealed an increase in anti-S protein antibodies in patients who died from this disease.

[0003] Antibody-dependent enhancement (ADE), a biological phenomenon in which antibodies promote viral entry into host cells and enhance viral infectivity (Tirado & Yoon 2003, Viral Immunol 16(1)69-86) When considering the possibilities, further concerns arise regarding S protein-centric approaches. It has been demonstrated that neutralizing antibodies can bind to the coronavirus S protein and trigger conformational changes that facilitate viral entry (Wan et al. J Virol 2020, 94(5)).

[0004] Therefore, due to these issues, it is desirable to develop additional vaccine design strategies, such as the use of T-cell antigens designed to induce a broad T-cell immune response in recipients.

[0005] However, when considering vaccines designed to induce a broad T-cell response, there is a further problem of human leukocyte antigen (HLA) constraint within individuals and broader populations. The HLA system is a gene complex that codes for human major histocompatibility complex (MHC) proteins, which are responsible for regulating an individual's immune system and for specifically presenting epitopes on the surface of infected cells, thereby eliciting immune responses to epitopes from intracellular pathogens and epitopes delivered to the individual in the form of vaccines (Marsh et al. 2010 Tissue Antigens 75(4):291-455).

[0006] The high polymorphism of HLA alleles and the subsequent inter-individual diversity of the immune system result in a diverse range of "HLA types" across populations. Adding to the complexity, such HLA types can significantly impact the effectiveness of viral vaccine compositions with potential for prevention among different individuals. Therefore, the design and production of epitope-based vaccines compatible with specific subsets of HLA types may prove ineffective for a significant proportion of the global population, including individuals with different HLA types. [Overview of the project] [Problems that the invention aims to solve]

[0007] Therefore, there is a need to develop methods for designing and producing vaccines that have the potential to stimulate a wide range of adaptive immune responses across a large proportion of the world's population. [Means for solving the problem]

[0008] Summary of the Invention According to a first aspect of the present invention, a computer-aided method for identifying one or more candidate regions of one or more source proteins that are predicted to induce an adaptive immunogenic response across a plurality of human leukocyte antigen HLA types, wherein one or more source proteins have an amino acid sequence, and the method comprises (a) accessing the amino acid sequence of one or more source proteins, (b) accessing a set of HLA types, (c) predicting the immune potential of a plurality of candidate epitopes within the amino acid sequence for each of the set of HLA types, (d) dividing the amino acid sequence into a plurality of amino acid subsequences, and (e) for each of the plurality of amino acid subsequences (f) generating a regional metric that indicates the predicted ability of an amino acid subsequence to induce an immunogenic response across a set of HLA types, wherein for each set of HLA types, the regional metric is based on the predicted immunogenicity potential of several candidate epitopes; and (f) applying a statistical model to identify whether any of the generated regional metrics are statistically significant, wherein an amino acid subsequence identified as having a statistically significant regional metric corresponds to a candidate region of an amino acid sequence predicted to induce an immunogenic response across at least a subset of the set of HLA types.

[0009] The method of the present invention is advantageous in that it uses a statistical model to quantitatively analyze the predicted immunogenicity potential of one or more candidate epitopes within amino acid subsequences across different sets of HLA types—in other words, the predicted ability of one or more candidate epitopes to induce an immunogenic response. Candidate regions of amino acid sequences (or “hotspots”) identified by quantitative statistical analysis may represent one or more regions (e.g., areas) of source proteins that are most likely to be promising vaccine targets and can be used in vaccine design and production. In particular, identified candidate regions are likely to contain one or more promising T cell epitopes (“predicted epitopes”) that can induce a broad T cell immune response across populations with different sets of HLA types.

[0010] As used in the present invention, the term “epitope” refers to any portion of an antigen recognized by any antibody, B cell, or T cell. “Antigen” refers to a molecule that can be bound by an antibody, B cell, or T cell, and may consist of one or more epitopes. Therefore, the terms epitope and antigen may be used synonymously herein. An epitope may also be referred to by the molecule to which it binds, such as a “T cell epitope,” or more specifically, an “MHC class I epitope” or an “MHC class II epitope.”

[0011] The human leukocyte antigen (HLA) system is a complex of genes that encode human MHC proteins. The high polymorphism of HLA genes, where the term "polymorphism" refers to the high diversity of different alleles, means that the precise MHC proteins encoded by various HLA genes in each individual can differ in fine-tuning their adaptive immune system. Hundreds of different alleles are recognized by HLA molecules. The terms HLA type and HLA allele may be used synonymously herein.

[0012] The amino acid subsequence region metric indicates the predicted immunogenicity potential of one or more candidate epitopes within amino acid subsequences across a set of HLA-typed tests. However, A "relatively good" region metric indicates that one or more candidate epitopes within that amino acid subsequence are collectively predicted to induce an immunogenic response across a large population of HLA types. A "relatively poor" region metric indicates that one or more candidate epitopes within that amino acid subsequence are not collectively predicted to induce an immunogenic response across a large population of HLA types in the analysis.

[0013] A statistical model is applied to identify amino acid subsequences that have statistically significant regional metrics. In particular, the statistical model is applied to identify any regional metrics that are better than would be expected by chance. As those skilled in the art will understand, the significance threshold of the statistical model may be selected accordingly, for example, based on the perceived accuracy of the predicted immunogenicity potential of a candidate epitope.

[0014] A candidate region may contain a single candidate epitope ("live" or "predicted" epitope) that is predicted to induce an immunogenic response across multiple HLA types. Such an epitope may be said to "overlap" with several HLA types. More typically, however, a candidate region contains multiple candidate epitopes that are predicted to induce an immunogenic response and, collectively, overlap with a large population of HLA types being analyzed. For example, one promising epitope within a candidate region may overlap with n HLA types, and different promising epitopes within that candidate region may overlap with m HLA types, thereby predicting that the candidate region will induce an immunogenic response across (m+n) HLA types.

[0015] Predicted epitopes can have different lengths from each other and may overlap. For example, a candidate region may include a predicted epitope the length of 8 amino acids, as well as a further predicted epitope the length of 25 amino acids. This 25-amino acid epitope may overlap with a portion of the 8-amino acid epitope, or it may completely encompass the entirety of the 8-amino acid epitope.

[0016] Typically, the method may further include the step of assigning an epitope score to each amino acid for each set of HLA types, where the epitope score is based on the predicted immunogenicity potential of one or more candidate epitopes containing that amino acid for that HLA type, and each of the regional metrics is generated across the set of HLA types based on the epitope scores of the amino acids within each of the amino acid subsequences.

[0017] Therefore, by generating a regional metric based on the epitope score of the amino acids within each amino acid subsequence (indicating the immunogenicity potential of the corresponding candidate epitopes), each regional metric represents the ability of an amino acid subsequence to induce an immunogenic response across a set of HLA types.

[0018] The regional metric may be the average of the amino acid epitope scores within each amino acid subsequence across a set of HLA types.

[0019] In each embodiment, at least a subset of epitope scores includes: (i) identifying a first set of candidate epitopes having a first (typically fixed) length across an amino acid sequence; (ii) for each set of HLA types, generating an epitope score for each of the first set of candidate epitopes that indicates the predicted immunogenicity potential of each candidate epitope of that HLA type; (iii) identifying a second set of candidate epitopes having a second (typically fixed) length across an amino acid sequence; (iv) for each set of HLA types, generating an epitope score for each of the second set of candidate epitopes that indicates the predicted immunogenicity potential of each candidate epitope of that HLA type; and (v) for each set of HLA types, for each amino acid in the amino acid sequence, for that HLA type This can be assigned by assigning an epitope score to the candidate epitope that is predicted to have the best immunogenicity potential among all of the first and second candidate epitopes containing the amino acids.

[0020] First, a plurality of first candidate epitopes are identified across the amino acid sequence, preferably in a "sliding window" of amino acids of a fixed length. In such a "sliding window" approach, the step size between consecutive candidate epitopes is less than the length of the candidate epitope so that consecutive candidate epitopes overlap. Typically, the step size is one amino acid. This is performed for each HLA type. For each of the plurality of first candidate epitopes, an epitope score indicative of the immunogenic potential of that candidate epitope for each HLA type is generated. How these epitope scores are generated will be considered in more detail later.

[0021] Subsequently, a second plurality of candidate epitopes are successively identified across the amino acid sequence for each HLA type. Again, this is preferably performed using a "sliding window approach". For each of the second epitopes, an epitope score indicative of the immunogenic potential of that epitope for each HLA type is assigned.

[0022] Next, to each amino acid, for each HLA type, an epitope score of the candidate epitope predicted to have the best immunogenic potential among all candidate epitopes containing that amino acid is assigned. Thus, for a particular HLA type, if both candidate epitope "A" and candidate epitope "B" contain a particular amino acid "X", then to amino acid "X" is assigned the epitope score of either candidate epitope "A" or "B" that is predicted to have the best immunogenic potential. In other words, for a given HLA type, the epitope score assigned to an amino acid corresponds to the best score obtained by the candidate epitopes that overlap this amino acid.

[0023] The first plurality of candidate epitopes and the second plurality of candidate epitopes have different lengths.

[0024] This method is typically extended to similarly identify multiple candidate epitopes of three or more. For example, when considering class I HLA types, candidate epitopes of lengths 8, 9, 10, 11, and 12 amino acids can be identified and scored based on their associated predicted immunogenic potential. Thus, in embodiments, multiple 8-mer candidate epitopes across an amino acid sequence can be identified and scored, and then multiple 9-mers, multiple 10-mers, multiple 11-mers, and 12-mers can be identified and scored. Then, each amino acid can be assigned an epitope score corresponding to the best score obtained by one of the identified candidate epitopes that contains that amino acid.

[0025] Preferably, the candidate epitope has a length of at least 8 amino acids, and preferably, the candidate epitope has a length of 8, 9, 10, 11, 12, or 15 amino acids. Typically, in class I HLA types, candidate epitopes of lengths from 8 to 12 amino acids are identified, and in class II HLA types, candidate epitopes of length 15 amino acids are identified, but other lengths may be used.

[0026] In a preferred embodiment, the predicted immunogenic potential of a candidate epitope for a particular HLA type is based on one or more of the predicted binding affinities and predicted processing of the identified candidate epitope.

[0027] Preferably, the predicted immunogenic potential (or "immunogenicity") of a candidate epitope is based on both the predicted binding affinity and processing of the candidate epitope. The combination of predicted binding affinity and predicted processing can be referred to as the predicted presentation of the candidate epitope. However, good results can still be obtained if the predicted immunogenic potential is based on one of these metrics (e.g., in class II HLA types, good results have been obtained when candidate epitopes are predicted for percentile rank binding affinity scores).

[0028] Such predictions can be performed using antigen presentation or binding affinity prediction algorithms, experimental data, or both. Examples of publicly available databases and tools that can be used for such predictions include the Immunoepitope Database (IEDB) (https: / / www.iedb.org / ). Other tools include the NetMHC prediction tool (http: / / www.cbs.dtu.dk / services / NetMHC / ), the TepiTool prediction tool (http: / / tools.iedb.org / tepitool / ), the MHCflurry prediction tool, the NetChop prediction tool (http: / / www.cbs.dtu.dk / services / NetChop / ), and the MHC-NP prediction tool (http: / / tools.immuneepitope.org / mhcnp / ). Other techniques are described in International Publication No. 2020 / 070307 and the same. This is disclosed in document number 2017 / 186959.

[0029] In a particularly preferred embodiment, antigen presentation is predicted from a machine learning model integrated with an ensemble machine learning layer of information from several HLA binding predictors (e.g., trained on ic50nm binding affinity data) and multiple different predictors for antigen processing (e.g., trained on mass spectrometry data).

[0030] Immunogenic potential may be based on alternative means of measuring heterogeneity or the ability of a candidate epitope to stimulate an immune response. Such examples may include comparing candidate epitopes to pathogen databases to identify the degree of similarity, or using predictive models that attempt to learn the physicochemical differences between immunogenic and non-immunogenic epitopes.

[0031] In some embodiments, the immunogenicity potential of a candidate epitope may be further based on its similarity to human proteins. Therefore, a candidate epitope may be penalized (e.g., assigned a lower score) if it is similar to a human protein.

[0032] An advantageous feature of the present invention is that this method not only identifies candidate regions containing epitopes capable of binding to HLA molecules, but also identifies CD8 epitopes that are naturally processed by the cell's antigen processing mechanism and presented on the surface of infected host cells.

[0033] The method may further include digitizing ("binarining") the assigned epitope scores, where each epitope score meeting a predetermined criterion is converted to "1" and each epitope score not meeting the criterion is converted to "0". The region metric for the amino acid subsequence can then be calculated, typically as the average of the number of amino acids in the subsequence assigned a value of "1" across the set of HLA types.

[0034] After the digitization process, amino acids assigned an epitope score of "1" can be considered components of promising epitopes predicted to induce an immunogenic response. Therefore, regions of amino acids with an assigned score of "1" may contain one or more (possibly overlapping) candidate epitopes predicted to bind to multiple HLA types.

[0035] Preferably, the set of HLA types includes HLA types of major histocompatibility complex (MHC) class I and HLA types of MHC class II. Thus, the method advantageously allows for the prediction of candidate regions expected to induce a broad T cell response across CD8+ and CD4+ T cell types. However, useful results can also be obtained when the set of HLA types includes only HLA types of MHC class I or only HLA types of MHC class II.

[0036] A set of HLA types may include HLA types that represent exactly one human population. A population is, This could be a racial group (e.g., Caucasian, African, Asian) or a geographical group (e.g., Lombardy, Wuhan). Therefore, the present invention can be used to identify candidate regions for a particular group. Thus, identified candidate regions common to several different groups are particularly advantageous for use in vaccine production.

[0037] In embodiments, the set of HLA types may include HLA types representing different human population groups. Thus, the method of the present invention can be usefully used to identify candidate regions that are expected to provide an immunogenic response across large populations of human populations.

[0038] In a preferred embodiment, the set of HLA types includes HLA types representing a human population. Thus, candidate regions predicted to induce an immunogenic response across most (or all) of the HLA types in such a set of HLA types may be promising candidates for a “universal” vaccine.

[0039] The set of HLA types may include the top N most frequent HLA types within a human population or group of human populations, preferably N is at least 5, more preferably at least 50, and even more preferably N=100. The statistical model of the present invention is particularly advantageous because it allows for the identification of a large number (e.g., 100) of candidate HLA type regions. Thus, the present invention can be used to design and produce vaccines with the potential to stimulate a broad adaptive immune response across large populations worldwide.

[0040] The present invention is particularly useful for identifying candidate regions that are expected to provide an immunogenic response across large populations of humans, but it can also be used to generate personalized vaccines for individuals (e.g., cancer treatment vaccines in the field of neoantigens). Thus, in embodiments, a set of HLA types may represent a given individual.

[0041] It will be understood that the method of the present invention can identify different candidate regions based on the set of HLA types used.

[0042] Statistical models can generally be based on one or more parametric distributions (e.g., binomial, Poisson, or hypergeometric distributions) or sampling methods to identify statistically significant amino acid subsequences. In a particularly preferred embodiment, applying a statistical model involves applying Monte Carlo simulations to estimate p-values ​​for each generated region metric. The estimated p-values ​​are then used to identify statistically significant amino acid subsequences and, consequently, candidate regions. The use of Monte Carlo algorithms is particularly advantageous because it allows the complexity in generating epitope scores to be reflected in the null model.

[0043] In statistical modeling, the null model is typically defined as the generative model of a set of epitope scores for each HLA type, assuming they are generated by chance. A set of epitope scores for a particular HLA type can be called an "HLA track." Using Monte Carlo simulations, a randomized set of HLA tracks and several related simulation domain metrics can be iteratively generated, from which p-values ​​of the domain metrics and, consequently, their statistical significance can be estimated.

[0044] Preferably, the null model reflects the complexity behind the epitope scores. Therefore, it is preferable to apply Monte Carlo simulation to (i) place the epitope scores into multiple epitope segments and epitope gaps based on the distribution of epitope scores for each HLA type, and (ii) repeatedly randomize the epitope segments and epitope gaps for each HLA type. This includes generating.

[0045] The placement of epitope scores for each HLA type into multiple epitope segments and epitope gaps (the placement of each HLA track) reflects, based on the assigned score, whether or not that amino acid was part of a candidate epitope predicted to have good immunogenic potential. Thus, an epitope segment is a sequence of epitopes (typically at least eight) assigned to amino acids within an epitope predicted to have good immunogenic potential. Such an epitope segment, consisting of the sequence “epitope amino acids”, can be considered as an amino acid region containing one or more predicted epitopes, which may or may not overlap with each other. An epitope gap is one or more consecutive scores assigned to amino acids that are not part of such predicted epitopes. By iteratively randomizing epitope segments and epitope gaps rather than individual amino acid epitope scores, the null model more faithfully reflects the methodology behind the region metrics, thereby providing more reliable results.

[0046] This method may further include applying a false detection rate (FDR) procedure to the results of a statistical model, preferably the Benjamin-Hochberg procedure or the Benjamin-Jektieri procedure.

[0047] In some embodiments, epitope scores may be weighted according to the human population frequency of each HLA type within a set of HLA types. Thus, candidate epitopes predicted to most frequently induce immunogenic responses across HLA types may be given a preference weight reflected in their amino acid epitope scores.

[0048] Statistically significant amino acid subsequences are identified as candidate regions that are likely to be promising vaccine targets. Therefore, the size of the amino acid subsequence is typically selected based on the intended vaccine platform. Preferably, each amino acid subsequence has the same length. For example, in step (b) of this method, the amino acid sequence may be divided into multiple amino acid subsequences of 20 to 50 amino acids in length toward a peptide vaccine platform capable of synthesizing the identified candidate region. Longer amino acid subsequences (e.g., 50 to 150 amino acids) may be used in vaccine platforms based on encoding the candidate region into a corresponding DNA or RNA sequence. Protein domains identified as having a large T-cell epitope population may also be used in vaccines. Such domains may provide a conformational antibody response.

[0049] Particularly preferred amino acid subsequence sizes are those of 27 amino acids, 50 amino acids, or 100 amino acids.

[0050] Amino acid subsequences are typically chosen to have the same length, but they can also be chosen to have different lengths. Amino acid subsequences can overlap each other, as described above using the "moving window" technique. However, to reduce the computational resources required to run the statistical model, amino acid subsequences can also be chosen to not overlap, for example, by arranging them consecutively across the amino acid sequence.

[0051] The candidate regions identified by the methods described above are predicted to contain promising T cell epitopes capable of inducing a broad T cell immune response across populations with different sets of HLA types. In preferred embodiments, each region metric may further indicate the predicted B cell response potential of each amino acid subsequence. In other words, the region metric may indicate the presence of any B cell epitope within the amino acid subsequence. In some embodiments, allocation The resulting epitope score may be further based on the predicted B cell response potential of each amino acid (e.g., within a predicted B cell epitope).

[0052] Additionally or alternatively, this method may further include analyzing each candidate region of one or more source proteins for the presence of B cell epitopes.

[0053] B cell response prediction can be based on B cell binding prediction algorithms, experimental data, or both. An example of a prediction tool that can be used in such embodiments is the BepiPred prediction tool (http: / / www.cbs.dtu.dk / services / BepiPred / ).

[0054] In embodiments, the method may further include comparing each identified candidate region with at least one human protein sequence to determine similarity, and ranking, extracting, or discarding candidate regions based on whether their similarity to at least one human protein is greater than a predetermined threshold.

[0055] These techniques advantageously allow for the comparison of the similarity of identified candidate regions to the expression profiles of proteins expressed in different major organs, thereby avoiding adverse reactions to vaccines based on such candidate regions. Different predetermined thresholds may be used. For example, if a candidate region contains one or more epitopes that closely match those of a human protein, that candidate region may be discarded.

[0056] This method may involve modifying candidate regions based on one or more adjacent amino acid subsequences. For example, if a candidate region is identified, but adjacent amino acid subsequences are known to have a predicted T cell epitope near the boundary between the two subsequences, the amino acid sequence of the candidate region may be extended to include further epitopes. It will also be understood that identified candidate regions may be joined together. For example, two candidate regions of 50 amino acids each may be joined to form a candidate region of 100 amino acids to be used in a vaccine.

[0057] The one or more source proteins are preferably one or more proteins of a virus, bacterium, parasite, tumor, or fragment thereof. The one or more source proteins may include a nascent antigen. For example, the one or more source proteins may be a spike (S) protein, a nucleoprotein (N), a membrane (M) protein, an envelope (E) protein, and one or more open reading frames such as ORF10, ORF1AB, ORF3A, ORF6, ORF7A, ORF8. Thus, the method of the present invention can be applied to the entire viral proteome. This is particularly useful for identifying candidate regions for vaccine design. In embodiments, the source proteins may be one or more proteins of a coronavirus, preferably the SARS-CoV-2 virus.

[0058] One or more source proteins may be multiple variations of one or more source proteins or may include multiple variations (and / or the method may be applicable to multiple variations of one or more source proteins). Each variation may be, for example, a mutation in a viral protein. Thus, the method of the present invention can be advantageously used to analyze the immunogenicity of all non-synonymous variations across multiple different protein sequences (e.g., of viruses). The method can advantageously include filtering one or more candidate regions to select one or more candidate regions in a conserved area (i.e., an area less likely to present mutations) of one or more proteins. The conserved area can be identified using techniques known in the art.

[0059] The amino acid sequences of one or more source proteins are determined by oligonucleotide hybridization, nucleic acid amplification-based methods (including, but not limited to, polymerase chain reaction-based methods), automated prediction based on DNA or RNA sequencing, or de novo peptides. The amino acid sequence can be obtained by sequencing, Edman sequencing, or mass spectrometry. The amino acid sequence can be obtained from a bioinformatics repository such as UniProt (http: / / www.uniprot.org). It can be downloaded.

[0060] The method may further include synthesizing one or more identified candidate regions and / or one or more predicted ("promising") epitopes within one or more identified candidate regions.

[0061] The method may further include encoding one or more identified candidate regions and / or one or more predicted ("feasible") epitopes within one or more identified candidate regions into corresponding DNA or RNA sequences. Such DNA or RNA sequences may be incorporated into a delivery system used in a vaccine (e.g., using naked or encapsulated DNA or encapsulated RNA). The method may also include incorporating the DNA or RNA sequences into the genome of a bacterial or viral delivery system to produce a vaccine.

[0062] Accordingly, a second aspect of the present invention provides a method for producing a vaccine, which includes identifying at least one candidate region of at least one source protein by any of the methods of the first aspect disclosed earlier, and synthesizing or encoding the at least one candidate region and / or at least one predicted epitope within the at least one candidate region into a DNA or RNA sequence corresponding to the at least one candidate region and / or at least one predicted epitope within the at least one candidate region. Such DNA or RNA sequences may be delivered in a naked or encapsulated form, or incorporated into the genome of a bacterial or viral delivery system to produce a vaccine. In addition, the DNA may be delivered to a vaccine-vaccinated host cell using a bacterial vector. In the case of peptide vaccines, the candidate region and / or epitope may typically be synthesized as an amino acid sequence or a "string".

[0063] A third aspect of the present invention provides a system for identifying one or more candidate regions of one or more source proteins that are predicted to induce an immunogenic response across a plurality of human leukocyte HLA allele types, wherein the one or more source proteins have an amino acid sequence, and the system comprises at least one processor communicating with at least one memory device, the at least one memory device storing instructions causing the at least one processor to perform any of the methods of the first aspect disclosed earlier.

[0064] According to a fourth aspect of the present invention, a computer-readable medium is provided which stores computer-executable instructions that implement any of the methods of the first aspect disclosed earlier.

[0065] In a further aspect of the present invention, a method is provided for constructing a diagnostic assay to determine whether a patient is infected with a pathogen or has been infected in the past (and has developed, for example, a protective immune response), the diagnostic assay being performed on a biological sample obtained from a subject and comprising identifying at least one candidate region of at least one source protein of a pathogen using any of the methods of the first aspect disclosed earlier, the diagnostic assay comprising utilizing or identifying at least one identified candidate region and / or at least one predicted epitope within at least one candidate region in the biological sample.

[0066] Thus, the present invention can be advantageously used to prepare rapid diagnostic tests or assays. Candidate regions and epitopes within the candidate regions can be further analyzed in laboratory tests to create such diagnostic tests or assays, thereby significantly reducing the time required for test development compared to conventional laboratory methods.

[0067] When used herein, the term "utilization" is intended to mean that at least one identified region and / or at least one predictive epitope within at least one identified region is used in the assay to identify a patient's (e.g., protective) immune response. In this context, the identified region and / or epitope within the identified region are components of the assay, not targets of the assay.

[0068] An in vitro diagnostic assay may involve the identification of an immune system component in a biological sample that recognizes at least one identified candidate region and / or at least one predicted epitope within the at least one candidate region. Thus, the diagnostic assay may utilize at least one identified candidate region and / or at least one predicted epitope. Typically, the diagnostic assay may include at least one identified candidate region and / or a predicted epitope (e.g., synthesized). In a preferred embodiment, the immune system component may be a T cell, and therefore the diagnostic assay may include a T cell assay. In another preferred embodiment, the immune system component may be a B cell. For example, the assay may include the identification of an antibody or B cell that recognizes a predicted B cell epitope within the at least one candidate region.

[0069] As an example of such diagnostic use, a sample taken from a patient, preferably a blood sample, can be analyzed for the presence of T cells, B cells, or antibodies in the biological sample that recognize and bind to epitopes within candidate regions identified as part of the present invention and included in the assay. The T cell epitopes identified as part of the present invention are expected to be presented by HLA molecules and therefore recognizable by T cells. Such a diagnostic response (e.g., T cell) indicates to those skilled in the art whether the patient has been exposed to an infection by a pathogen and whether they have developed a protective immune response, such infection resulting in an observable level of cellular immunity and / or immunological memory.

[0070] Suitable diagnostic assays will be understood by those skilled in the art, but may include enzyme-linked immunosorbent spot (ELISPOT) assays, enzyme-linked immunosorbent assays (ELISA), cytokine capture assays, intracellular staining assays, tetramer staining assays, or limiting dilution culture assays.

[0071] In a method for creating a diagnostic test, the amino acid sequences of one or more source proteins (from which at least one candidate region can be identified) may be selected based on the desired response to be tested. For example, one or more source proteins may be one or more source proteins of a coronavirus (or fragment thereof), such as the SARS-CoV-2 virus. In such a case, the present invention can be used to create a diagnostic test that determines whether or not a patient is infected with the SARS-CoV-2 virus or has been infected in the past. However, as will be understood by those skilled in the art, one or more source proteins may be from any pathogen (e.g., a virus or a bacterium).

[0072] Further disclosed herein are diagnostic assays for determining whether a patient is or has been previously infected with a pathogen, the diagnostic assays being performed on a biological sample obtained from a subject and comprising utilizing or identifying in the biological sample at least one candidate region of at least one source protein of a pathogen identified using any of the methods of the first embodiment described above, and / or at least one predicted epitope within said at least one candidate region. The diagnostic assays may also include identifying immune system components (e.g., T cells or B cells) in the biological sample that recognize at least one identified candidate region and / or at least one predicted epitope within said candidate region.

[0073] Brief explanation of the drawing Embodiments will be described in more detail with reference to the attached diagrams, which are merely examples. [Brief explanation of the drawing]

[0074] [Figure 1A]The epitope maps of the SARS-CoV-2 virus S protein across the most frequent HLA-A, HLA-B, and HLA-DRB alleles in the human population are shown. In these epitope maps, the data are converted so that CD8 positivity results are associated with 0.7 or higher, and less than 10% (represented as 0.1 in the figure) are related to class II, demonstrating broad coverage of CD8 and CD4 with overlapping B cell antibody carriers. [Figure 1B] The epitope maps of the SARS-CoV-2 virus S protein across the most frequent HLA-A, HLA-B, and HLA-DRB alleles in the human population are shown. In these epitope maps, the data are converted so that CD8 positivity results are associated with 0.7 or higher, and less than 10% (represented as 0.1 in the figure) are related to class II, demonstrating broad coverage of CD8 and CD4 with overlapping B cell antibody carriers. [Figure 2] This shows hierarchical clustering of the binary conversion of epitope maps of class ICD8 epitopes in the HLA-A and HLA-B alleles of the SARS-CoV-2 virus S protein. [Figure 3] Using preservation and human self-peptide filtering procedures, we show epitope hotspots captured across the entire viral proteome of the SARS-CoV-2 virus from Monte Carlo analysis. [Figure 4] This is a scatter plot showing mutant AP scores compared to wild-type AP score protein variants. [Figure 5] This demonstrates the application of Monte Carlo epitope hotspot prediction to 10 variant viral sequences at different geographical locations. [Figure 6] A scatter plot showing the distribution of protein hotspot conservation scores in the viral genome is shown. [Figure 7] This is a flowchart illustrating the steps of a preferred embodiment of the method. [Figure 8] This is an example of a system suitable for implementing an embodiment of the method. [Figure 9] This is an example of a suitable server. [Modes for carrying out the invention]

[0075] Detailed description of the drawing According to certain embodiments described herein, methods and systems are proposed for identifying one or more candidate regions of one or more source proteins that are predicted to induce adaptive immunogenic responses across multiple HLA types. Such candidate regions may be referred to as “hotspots,” and the terms “candidate region” and “hotspot” can be used synonymously herein. In embodiments, identified hotspots and / or epitopes identified within hotspots can be used in vaccine design and production.

[0076] Next, preferred embodiments for identifying such hotspots will be described. The following description refers to the analysis of the entire proteome of the SARS-CoV-2 virus, but it will be understood that the present invention can be used for the analysis of different viruses, tumors, bacteria, parasites, or fragments thereof such as nascent antigens.

[0077] Generation of global epitope maps and amino acid scores For a given HLA allele, the score assigned to an amino acid corresponds to the best score obtained by predicting the epitope overlapping with that amino acid. For class IHLA alleles, the epitope lengths are preferably 8, 9, 10, 11, and 12, and antigen presentation (AP) or immune presentation (IP) of the viral peptide to the surface of infected host cells is predicted. Various methods and tools are available for AP prediction, such as the publicly available NETCHop and NETMHC prediction tools and those discussed in the summary section of this specification. These Class I scores range from 0 to 1, with 1 being the best score (i.e., more likely to be spontaneously presented on the cell surface). In this embodiment, for Class II HLA alleles, prediction was performed at 15mer. Class II prediction is a percentile-rank binding affinity score (not antigen presentation), and therefore a lower score is better (the score ranges from 0 to 100, with 0 being the best score).

[0078] A statistical framework for detecting epitope hotspots and epitope regions in different HLA populations Input data The datasets input into the statistical framework are epitope maps generated for each amino acid position in one or more source proteins (e.g., all proteins in the SARS-CoV-2 proteome) for all studied (e.g., 100 HLA alleles). The score for any given amino acid was determined as the maximum AP or IP score held in the epitope map by the peptide (candidate epitope) overlapping that amino acid. All peptide lengths of 8 to 11 amino acids in class I and all peptide lengths of 15 amino acids in class II were processed to generate one HLA dataset per viral protein. Each row in the dataset represents the predicted amino acid epitope score for one HLA type.

[0079] Statistical framework The central question that the statistical framework attempts to answer is, "Are specific regions in a given viral protein that score highly immunogenic compared to a given set of HLA types higher than would be expected by chance?"

[0080] HLA Trucks The raw input dataset (e.g., AP or percentile-rank binding affinity scores) is first converted to a binary track. For each class IHLA dataset, the epitope scores are converted to binary (0 or 1) values ​​such that amino acid positions with predicted epitope scores greater than 0.7 (for AP) and greater than 0.5 (for IP) are assigned a value of 1 (positive predicted epitopes), and the rest are assigned a value of 0. Similarly, for class IIHLA datasets, amino acid positions with predicted epitope scores less than 10 are assigned a value of 1, and all others are assigned a value of 0. These thresholds are relatively conservative, and it should be understood that other thresholds may be chosen based on the techniques and confidence levels used in generating the raw data. Each binary track can be efficiently presented as a list of consecutive intervals of 1—segments—and consecutive intervals of 0 between segments or gaps.

[0081] Test statistic For a group of k HLA binary tracks, the test statistic ("region metric") Si is calculated for each bin bi of a given size m, dividing the protein into n bins (for example, for larger proteins, m = 100 amino acids). For a single HLA track, the test statistic is s i This is calculated for each bin bi:

number

number

[0082] Null model An efficient method for estimating the statistical significance of observed HLA tracks is Monte Carlo-based simulation. If HLA tracks are generated randomly, a null model is defined as the generation model for HLA tracks. From the model, through sampling, the test statistic Si of the null distribution is obtained. The null model must reflect the complexity due to the properties of the HLA tracks. Epitope amino acids in a single HLA track always generate a continuum of at least 8 (the minimum peptide size used in the prediction framework) in length. Similarly, amino acids with low epitope scores are also clustered together.

[0083] p-value estimation To sample from the null model, each of the k HLA tracks is divided into segments and gaps, which are then shuffled to generate randomized HLA tracks. In this embodiment, this is repeated 10,000 times to generate Si statistics for 10,000 samples in each bin. For each bin, the p-value is estimated as the proportion of samples above the truly observed environment. Furthermore, the generated p-values ​​are adjusted for multiple testing using the Benjamin-Jektieri procedure to a false detection rate (FDR) of 0.05. It will be understood that other multiple testing procedures (e.g., Benjamin-Hochberg) may be used. Different false detection rates may be implemented.

[0084] Epitope Hotspot Preservation Score An example of generating a measure of conservation is described below. For each protein in the viral genome, a set of unique amino acid sequences was compiled from all lineages available in the GISAID database as of March 29, 2020 (Shu, Y. and J. McCauley, GISAID: Global initiative on sharing all influenza data-from vision to reality. Euro Surveill, 2017.22(13)). These sets were processed individually using the Clustal Omega (v1.2.4) software (Sievers, F. and D. Ghiggins, Clustal Omega for making accurate alignments of many protein sequences. Protein Sci, 2018.27(1):p.135-145.) via a command-line interface with default parameter settings. The software outputs a consensus sequence containing conservation information for each amino acid in the protein sequence. Therefore, at position i in the consensus sequence, * The amino acid indicated as '' can be rephrased as meaning that this amino acid is conserved at position i in all input sequences (Sievers, F. and D. Ghiggins, Clustal Omega for making accurate alignments of many protein sequences. Protein Sci, 2018. 27(1): p. 135-145.).

[0085] Next, the hotspot offset was used to extract each consensus sequence. For each hotspot, the " within the consensus subsequence relative to the total length of the subsequence" * The preservation score was calculated as a ratio of "[ ]". Therefore, each hotspot was assigned a preservation score from 0 to 1, where 1 represents perfect preservation across all available lines.

[0086] 1,000 hotspots, equal to the size of the hotspots from the entire consensus sequence of the protein. The median preserved score was calculated by sampling a subarray. A preserved score was assigned to each sample, and the median was calculated from all 1,000 preserved scores. The minimum preserved score was calculated using a sliding window method where the window size was equal to the hotspot size. For each increment, a preserved score was calculated and the generated minimum preserved score was retained.

[0087] Next, we will describe an example of applying the method of the present invention to the SARS-CoV-2 viral proteome. However, as discussed earlier, the method can be applied to several different source proteins, such as different viruses, bacteria, tumors, or parasites. The method can also be applied to newly synthesized antigens.

[0088] The immunogenicity landscape of SARS-CoV-2 reveals diversity among different HLA groups in human populations. Epitope mapping of the entire SARS-CoV-2 viral proteome was performed. Antigen presentation (AP) was predicted from a machine learning model integrated into an ensemble machine learning layer of information from several HLA-binding predictors (in this case, three separate HLA-binding predictors trained on ic50nm binding affinity data) and 13 different antigen-processing predictors (all trained on mass spectrometry data). The output AP scores ranged from 0 to 1 and were used as input to calculate immune presentation (IP) across the epitope map. The IP score penalizes presentation peptides with a high degree of "human similarity" compared to human proteins and rewards peptides with low similarity. The IP score represents the HLA-presenting peptide that is most likely to be recognized by circulating T cells, i.e., T cells that are not deficient or anergized, in the periphery, and therefore most likely to be immunogenic.

[0089] Both AP and IP epitope prediction are "pan" HLA or HLA-independent and can be performed on any allele in the human population; however, for the purposes of this study, the analysis was limited to the 100 most frequent HLA-A, HLA-B, and HLA-DR alleles in the human population. Class II HLA binding prediction was also incorporated into a large-scale epitope screening from the IEDB consensus tool (Dhanda, SK, et al., IEDB-AR: immune epitope database-analysis resource in 2019. Nucleic Acids Res, 2019. 47(W1): p. W502-W506.), and B cell epitope prediction was performed using BepiPred (Dhanda, SK, et al., IEDB-AR: immune epitope database-analysis resource in 2019. Nucleic Acids Res, 2019. 47(W1): p. W502-W506.). The resulting epitope map allowed us to identify regions in the viral proteome most likely to be presented by infected host cells using the most frequent HLA-A, HLA-B, and HLA-DR alleles in human populations worldwide.

[0090] All epitope maps of the viral protein were created, and an example based on the IP score of the S protein is shown in Figure 1A, and an example for AP is shown in Figure 1B, showing distinct regions of the S protein containing candidate CD8 and CD4 epitopes for the 100 most frequent human HLA-A, HLA-B, and HLA-DR alleles. This set of HLA types is shown in Figure 1A. Interestingly, predicted B cell epitopes often map to regions of the protein containing high-density predicted T cell epitopes, and therefore the heatmap provides a comprehensive overview of the most relevant regions of the SARS-CoV-2 virus that can be used for vaccine development. From Figure 1, it is clear that different HLA alleles have different class IAP and class II binding affinities. This strongly suggests, as expected, that the SARS-CoV-2 antigen presentation landscape clusters into distinct populations across a range of different human HLA alleles. This trend is further illustrated in the hierarchical clustering map presented in Figure 2, after the AP scores have been binaryized. Figure 2 clearly shows that some allele clusters present many viral targets to the human immune system, while others present only a few targets, and some are unable to present any targets at all. Figure 2 shows shuffleable epitope segments and epitope gaps for each HLA type in Monte Carlo simulations. This suggests that different groups within human populations with different HLAs respond differently to T cell-driven vaccines composed of viral peptides. Therefore, to design optimal vaccines that leverage the benefits of T cell immunity across a broad human population, it is desirable to predict “epitope hotspots” in the viral proteome. These hotspots are viral regions rich in overlapping epitopes and / or spatially close epitopes that can be recognized by multiple HLA types across human populations.

[0091] Prior to the discovery of such epitope hotspots with the broadest coverage in the human population, we confirmed, to the extent possible, that T-cell-based AP and IP scores predict promising targets from a limited number of validated SARS-CoV virus epitopes. We identified class I epitopes from the original SARS-CoV virus (first appearing in Guangdong Province, China in 2002) that shared more than 90% sequence identity with current SARS-CoV-2. Unfortunately, many of the publicly available epitopes were identified using ELISPOT against PBMCs from convalescent patients and / or healthy donors (or humanized mouse models), where restriction HLA was not explicitly deconvoluted. To mitigate this problem, we used tetramers to identify a subset of five epitopes with minimal epitopes and HLA restriction (Grifoni, A., et al., A Sequence Homology and Bioinformatic Approach Can Predict Candidate Targets for Immune Responses to SARS-CoV-2. Cell Host Microbe, 2020).

[0092] Of the five epitopes tested, four were identified as positive, i.e., had an IP score greater than 0.5 (see Table 1), showing an accuracy of 80%. Although this was a very small test dataset, the NEC Immune Profiler prediction pipeline showed promising immunogenicity candidates. This allows for precise identification and provides a degree of confidence that the epitope hotspots identified by this analysis and subsequent analyses represent targets of interest for vaccine development.

[0093] [Table 1]

[0094] Robust statistical analysis identifies epitope hotspots across a wide range of T cell responses. To identify potential epitope hotspots that are promising immunogenic targets for the majority of the human population, we first performed a Monte Carlo random sampling procedure on the previously generated epitope map (of the Wuhan reference sequence exemplified in Figure 1 for the S protein) to identify specific areas of the SARS-CoV-2 proteome with the highest probability of being epitope hotspots using the method described above. Three bin sizes were examined for potential epitope hotspots: 27, 50, and 100. Statistics were calculated for each defined subset region (bin) of protein from a set of 100 HLAs. Then, using Monte Carlo simulation, the p-value for each bin was estimated, thereby indicating that each bin represented a candidate epitope hotspot. Statistically significant bins that emerged from the simulations represented epitope hotspots or regions of interest for each protein analyzed.

[0095] Epitope hotspots are constructed for each individual epitope score, epitope length, and each amino acid contained within the epitope hotspot. These scores are generated for each amino acid in the hotspots of all 100 most frequent HLA alleles in the human population. Based on Monte Carlo analysis, significant hotspots have a false detection rate (FDR) of less than 5% and represent the regions most likely to contain promising T-cell-driven vaccine targets recognizable by multiple HLA types across the human population. A summary of epitope hotspots identified across the entire range of the virus is shown in Figure 3, revealing that most immunogenic regions of the virus targeting the most frequent human HLA alleles in the global population are found in several viral proteins, in addition to antibody-exposure structural proteins such as the S protein.

[0096] Conservation analysis identifies robust epitope hotspots in SARS-CoV-2. A universal vaccine blueprint should ideally also be capable of protecting populations from different branching groups from which the SARS-CoV-2 virus emerges. Therefore, the AP potentials of 3400 viral sequences in the GISAID database were compared to the AP potentials of the Wuhan Genbank reference sequence. The results of this comparison are shown in Figure 4, suggesting a trend in which SARS-CoV-2 mutations appear to decrease their potential to be presented and therefore detected by the host immune system. Similar trends have been observed in chronic infections such as HPV and HIV.

[0097] To assess whether these epitope hotspots are sufficiently robust across all sequenced and mutant lineages of SARS-CoV-2, we then used the epitope hotspot Monte Carlo statistical framework to analyze 10 viral sequences from 10 of the most mutated viral sequences from different geographical regions (Shu, Y. and J. McCauley, GISAID: Global initiative on sharing all influenza data-from vision to reality. Euro Surveill, 2017. 22(13)). The vast majority of hotspots were present in all sequenced viruses, but occasionally, hotspots disappeared and / or new hotspots appeared in these diverse lineages. This is shown in Figure 5. Figure 5 shows the application of the Monte Carlo epitope hotspot prediction method to 10 mutant viral sequences from different geographical locations. Hotspots in the 10 mutant sequences compared to the Wuhan reference sequence are on the x-axis, and the frequency of epitope hotspots is on the y-axis. Frequencies are shown for three different hotspot bin lengths: 27 (left), 50 (center), and 100 (right). While epitope hotspots are robust across variant sequences, it is evident that novel epitope hotspots occasionally appear in several sequences at different geographical locations.

[0098] While the identified hotspots appeared robust across different viral lineages, epitope hotspots were subjected to sequence conservation analysis to design the most robust vaccine blueprint, hopefully providing broad protection from newly emerging branches of the SARS-CoV-2 virus. The goal of this analysis was to identify hotspots that appeared resistant to mutation across thousands of viral sequences. Conservation scores for each hotspot were calculated based on the consensus sequence of the protein using the techniques discussed earlier. Figure 6 shows the conservation scores of the hotspots identified based on IP using different bin sizes. Only epitope hotspots exhibiting conservation scores higher than the median conservation score were retained for further analysis. This allowed for filtering out approximately half of the hotspots in 50 and 100 amino acid bin sizes, and over 70% in 27 amino acid bin sizes. In addition, bins containing tight sequence matches with proteins in the human proteome were removed to reduce the potential for non-targeted autoimmune responses against host tissues.

[0099] Variant immunogenicity potential across SARS-CoV-2 mutation sequences As of March 31, 2020, all available strains in the GISAID database were downloaded (Shu, Y. and J. McCauley, GISAID: Global initiative on sharing all influenza data-from vision to reality. Euro Surveill, 2017. 22(13)), and executed using default parameters through the Nextstrain / Augur software suite (Hadfield, J., et al., Nextstrain: real-time tracking of pathogen evolution. Bioinformatics, 2018. 34(23): p.4121-4123). The generated phylogenetic trees were analyzed to obtain all protein variants. For each, HLA-A *The wild-type score and variant antigen presentation (AP) score for 02:01 were calculated. The variant score is the highest AP score among the nine possible 9mer peptides, including the variant. The wild-type score is the highest AP score of the 9mer at the same position in the reference (Wuhan) lineage.

[0100] Figure 7 is a flowchart summarizing the steps of a preferred embodiment of the present invention, which have been described in more detail above.

[0101] In step S201, the amino acid sequences of one or more source proteins are obtained. These may be, for example, one or more source proteins of a virus, bacteria, parasite, or tumor.

[0102] In step S203, multiple candidate epitopes are identified within the amino acid sequence. These candidate epitopes may have a length of 8, 9, 10, 11, 12, or 15 amino acids and can be identified, for example, by a “movement window” technique.

[0103] In step S205, the immune response potential of each candidate epitope is predicted for each set of HLA types (e.g., representing a human population). The immune response potential can be either an antigen presentation (AP) score or an immune presentation (IP) score, as discussed earlier.

[0104] In step S207, an epitope score is assigned to each amino acid for each HLA type based on the overlapping candidate epitope with the best predicted immunogenicity potential for that HLA type. The epitope score may be, for example, an AP value or an IP value.

[0105] In step S208, the epitope score is digitized into epitope segments and epitope gaps based on predetermined thresholds. The epitope segments indicate promising epitopes for HLA types.

[0106] In step S209, the amino acid sequence is divided into multiple amino acid subsequences or "bins." These may have varying lengths, for example, depending on the intended vaccine platform.

[0107] In step S211, a region metric for each amino acid subsequence is calculated based on the assigned epitope score within the amino acid subsequence.

[0108] In step S213, a statistical model (such as Monte Carlo simulation) is used to identify candidate regions (or "hot spots") that have statistically significant regional metrics.

[0109] In step S215, the identified candidate regions may be filtered to prioritize those occurring in the conserved region. For example, different viral sequences may be analyzed, and candidate regions identified in the conserved region across different analyses may be prioritized.

[0110] This document provides a clear application of this method in vaccine design. However, it will be understood that the technique described herein can also be equally applied to the design of T cells that recognize epitopes in identified candidate regions ("hotspots"). Similarly, this technique can be used to identify neo-antigen loadings in tumors, i.e., to predict the response to treatment, where it is used as a biomarker.

[0111] Referring now to Figure 8, an example of a system suitable for implementing an embodiment of this method is shown. System 1100 includes at least one server 1110 that communicates with a reference data store 1120. The server can also communicate with an automated peptide synthesis device 1130, for example, via a communication network 1140.

[0112] In certain embodiments, the server may retrieve, for example, the amino acid sequences of one or more source proteins from a reference data store, along with data associated with a set of HLA types. The server may then use the steps described above to identify one or more candidate hotspots in the amino acid sequences.

[0113] Candidate regions (or one or more predicted epitopes within the candidate regions) are sent to an automated peptide synthesis device 1130 to synthesize the candidate regions or epitopes. Such peptide synthesis is particularly suitable for candidate regions or epitopes up to 30 amino acids in length. Techniques for automated peptide synthesis are well known in the art, and it will be understood that any known technique can be used. Typically, candidate regions or epitopes are synthesized using standard solid-phase peptide synthesis chemistry, purified using reverse-phase high-performance liquid chromatography, and then formulated into aqueous solutions. When used for vaccination, prior to administration, the peptide solution is usually mixed with an adjuvant before being administered to the patient.

[0114] Peptide synthesis technology has existed for over 20 years, but in recent years, it has undergone rapid improvements, bringing the synthesis process to a point where it now takes only a few minutes on commercial machines. For brevity, such machines will not be described in detail, but their operation will be understood by those skilled in the art, and such conventional machines can be adapted to receive candidate regions or candidate epitopes from a server.

[0115] The server, including the functions described above, can identify candidate regions on an amino acid sequence. Naturally, it will be understood that these functions may be subdivided across different processing entities and different processing modules communicating with each other within the computer network.

[0116] Techniques for identifying candidate regions are useful for developing a broader ecosystem of customized vaccines. It can be integrated into a system (for example, using the method of the present invention for an individual's HLA type). Examples of vaccine development ecosystems are well known in the art and are described at a high level in relation to the situation, but for the sake of brevity, the ecosystem will not be described in detail.

[0117] In one example of the ecosystem, the first sample step may involve separating DNA from a tumor biopsy and matching healthy tissue control. In the second sequencing step, the data is sequenced and variants, such as mutations, are identified. In the immunoprofiler step, the relevant mutant peptides are identified.<in silico> >It can be generated with

[0118] Using the relevant mutant peptides and the techniques described herein, candidate regions can be predicted and selected, and target epitopes can be identified for vaccine design. Specifically, candidate peptide sequences selected based on their predicted binding affinity were determined using the techniques described herein.

[0119] As described above, the target epitope is then synthesized and produced using conventional techniques. Prior to administration, the peptide solution is usually mixed with an adjuvant before being administered to the patient (vaccinated). Alternatively, as with any conventional vaccine, the target epitope can be incorporated into DNA or RNA, or into the genome of bacteria or viruses.

[0120] The candidate regions predicted by the methods described herein can also be used to produce other types of vaccines besides peptide-based vaccines. For example, the candidate region (or predicted epitope within the candidate region) can be encoded into a corresponding DNA or RNA sequence and used for patient vaccination. Note that DNA is typically inserted into a plasmid construct. Alternatively, the DNA can be incorporated into the genome of a bacterial or viral delivery system (it can also be RNA, depending on the viral delivery system) – thus producing a vaccine in a genetically modified virus or bacterium that manufactures the target after vaccination in the patient, i.e., in vivo.

[0121] An example of a suitable server 1110 is shown in Figure 9. In this example, the server includes at least one microprocessor 1200 interconnected via bus 1204 as shown, memory 1201, optional input / output devices 1202 such as a keyboard and / or display, and an external interface 1203. In this example, the external interface 1203 can be used to connect the server 1110 to peripherals such as a communication network 1140, a reference data store 1120, and other storage devices. Although a single external interface 1203 is shown, this is for illustrative purposes only, and in practice, multiple interfaces (e.g., Ethernet, serial, USB, wireless, etc.) can be provided using various methods.

[0122] In use, the microprocessor 1200 executes instructions in the form of application software stored in memory 1201 to enable the execution of necessary processes, including communication with the reference data store 1120 for receiving and processing input data and / or receiving sequence data of one or more source proteins, and communication with a client device for generating potential predictions (including, for example, predicted binding affinity and processing). The application software may include one or more software modules and may run in a suitable execution environment such as an operating system environment.

[0123] Therefore, it will be understood that the server 1200 can be formed from any suitable processing system such as appropriately programmed client devices, PCs, web servers, network servers, etc. In a particular example, the server 1200 is an Intel architecture running software applications stored on non-volatile (e.g., hard disk) storage. A standard processing system, such as a base processing system, is used, but this is not mandatory. However, it should be understood that the processing system can be any processing device, such as a microprocessor, microchip processor, logic gate configuration, optionally FPGA (Field Programmable Gate Array) or related firmware for logic implementation, or any other electronic device, system, or mechanism. Therefore, the term "server" is used, but only as an example and not intended to be limiting.

[0124] Although Server 1200 is presented as a single entity, it will be understood that Server 1200 can be distributed across several geographically separate locations by using, for example, a processing system and / or database 1201 provided as part of a cloud-based environment. Therefore, the above-described deployment is not mandatory, and other suitable configurations can also be used.

[0125] As discussed earlier, this method is used in vaccine design. The method can also be used to design and prepare in vitro diagnostic tests or assays. For example, such a diagnostic assay may be used to identify T cells or B cells in a biological sample that recognize and bind to “hot spots” and / or epitopes contained within the assay, which have been identified using the technique of the present invention. A diagnostic response to such a diagnostic assay may indicate to those skilled in the art whether a patient has been exposed to infection by a pathogen of interest (e.g., SARS-CoV-2 virus) and whether the patient has developed protective immunity.

Claims

1. A computer-aided method for identifying one or more candidate regions of one or more source proteins predicted to induce adaptive immunogenic responses across multiple human leukocyte antigen (HLA) types, wherein the one or more source proteins have an amino acid sequence, and the method is: (a) Accessing the amino acid sequence of one or more source proteins, (b) Accessing a set of HLA types, (c) For each of the sets of HLA types, predict the immune potential of multiple candidate epitopes within the amino acid sequence, (d) Dividing the amino acid sequence into multiple amino acid subsequences, (e) For each of the plurality of amino acid subsequences, to generate a regional metric indicating the predicted ability of the amino acid subsequence to induce an immunogenic response across the set of HLA types, wherein for each of the set of HLA types, the regional metric is based on the predicted immunogenic potential of the plurality of candidate epitopes. (f) Applying a statistical model to determine whether any of the generated region metrics are statistically significant, wherein an amino acid subsequence identified as having a statistically significant region metric corresponds to a candidate region of the amino acid sequence that is predicted to induce an immunogenic response across at least a subset of the set of HLA types, A computer implementation method, including

2. The process further includes assigning an epitope score to each amino acid for each of the sets of HLA types, wherein the epitope score is based on the predicted immunogenicity potential of one or more of the candidate epitopes containing that amino acid for that HLA type. The computer implementation method according to claim 1, wherein each of the region metrics is generated based on the epitope score of the amino acid in each of the amino acid subsequences across the set of HLA types.

3. At least a subset of the epitope scores are (i) Identifying a first group of candidate epitopes having a first length across the amino acid sequence, (ii) For each of the sets of HLA types, generate an epitope score for each of the first plurality of candidate epitopes that represent the predicted immunogenicity potential of each candidate epitope of that HLA type, (iii) Identifying a second group of candidate epitopes having a second length across the amino acid sequence, (iv) For each of the sets of HLA types, generate an epitope score for each of the second plurality of candidate epitopes that represent the predicted immunogenicity potential of each candidate epitope of that HLA type, (v) For each of the sets of HLA types, assign the epitope score to each amino acid in the amino acid sequence of the candidate epitope that is predicted to have the best immunogenicity potential among all of the first and second candidate epitopes containing that amino acid in that HLA type, A computer implementation method according to claim 1 or 2, which is assigned by performing the following.

4. The computer-aided method according to any one of claims 1 to 3, wherein the candidate epitope has a length of at least eight amino acids, and preferably the candidate epitope has a length of eight, nine, ten, eleven, twelve, or fifteen amino acids.

5. The computer-aided method according to any one of claims 1 to 4, wherein the predicted immunogenicity potential of a specific HLA type candidate epitope is based on one or more predicted binding affinities and prediction processes of the identified candidate epitope.

6. The computer-aided method according to any one of claims 1 to 5, wherein the immunogenicity potential of the candidate epitope is further based on the similarity of the candidate epitope to human proteins.

7. The computer implementation method according to any one of claims 2 to 6, further comprising digitizing the assigned epitope scores, wherein each epitope score that meets a predetermined criterion is converted to "1", and each epitope score that does not meet the predetermined criterion is converted to "0".

8. The computer-aided method according to any one of claims 1 to 7, wherein the set of HLA types includes HLA types of major histocompatibility complex (MHC) class I and HLA types of MHC class II.

9. The computer implementation method according to any one of claims 1 to 8, wherein the set of HLA types includes HLA types representing at least one human population group, and preferably the set of HLA types represents the human population.

10. The computer implementation method according to any one of claims 1 to 9, wherein the set of HLA types includes the top N most frequent HLA types within the human population or group of human populations, preferably N is at least 5, more preferably at least 50, and more preferably at least 100.

11. The computer implementation method according to any one of claims 1 to 8, wherein the set of HLA types represents a given individual.

12. The computer implementation method according to any one of claims 1 to 11, wherein applying the statistical model includes applying a Monte Carlo simulation to estimate the p-values ​​for each of the generated domain metrics.

13. If at least dependent on claim 2, applying the Monte Carlo simulation means (i) For each HLA type, the epitope scores are arranged in multiple epitope segments and epitope gaps based on the distribution of the epitope scores, (ii) For each HLA type, repeatedly generate the random arrangement of the epitope segment and the epitope gap, The computer implementation method according to claim 12, including the method described in claim 12.

14. The computer implementation method according to any one of claims 1 to 13, further comprising applying an FDR procedure, which is a false detection rate procedure, to the results of the statistical model, wherein the FDR procedure is the Benjaminie-Hochberg procedure or the Benjaminie-Jektieri procedure.

15. The computer implementation method according to any one of claims 2 to 14, further comprising weighting the epitope score according to the human population frequency of each HLA type within the set of HLA types.

16. Each amino acid subsequence consists of at least 8 amino acids, preferably 20 to 50 amino acids. A computer-aided method according to any one of claims 1 to 15, comprising an acid, more preferably 50 to 150 amino acids.

17. The computer-aided method according to any one of claims 1 to 16, wherein each of the region metrics further indicates the predicted B-cell response potential of each of the amino acid subsequences.

18. The computer-aided method according to claim 17, wherein, when dependent on claim 2, each assigned epitope score is further based on the predicted B-cell response potential of each of the amino acids.

19. The computer-aided method according to any one of claims 1 to 18, further comprising analyzing each candidate region of the one or more source proteins for the presence of a B cell epitope.

20. To determine the similarity, each identified candidate region is compared to at least one human protein sequence, The candidate regions are ranked or discarded based on whether the similarity to at least one of the human proteins is greater than a predetermined threshold. A computer implementation method according to any one of claims 1 to 19, further comprising:

21. A computer-aided method according to any one of claims 1 to 20, further comprising adjusting a candidate region based on one or more adjacent amino acid subsequences.

22. The computer-aided method according to any one of claims 1 to 21, wherein the one or more source proteins are one or more proteins of a virus, tumor, bacterium, parasite, or fragment thereof containing a nascent antigen.

23. The computer-aided method according to any one of claims 1 to 22, wherein the one or more source proteins are one or more proteins of a coronavirus, preferably a SARS-CoV-2 virus.

24. The computer-aided method according to any one of claims 1 to 23, wherein the one or more source proteins include multiple variations of one or more proteins.

25. The computer implementation method according to claim 24, further comprising filtering one or more candidate regions in a storage area in order to select one or more candidate regions.

26. A method for producing a vaccine, Identifying at least one candidate region of at least one source protein by the method according to any one of claims 1 to 25, Synthesizing the at least one candidate region and / or at least one predicted epitope within the at least one candidate region, or encoding the at least one candidate region and / or at least one predicted epitope within the at least one candidate region to form a corresponding DNA sequence or RNA sequence. Methods that include...

27. A system for identifying one or more candidate regions of one or more source proteins that are predicted to induce an immunogenic response across multiple human leukocyte antigen (HLA) allele types, The system comprises one or more source proteins having an amino acid sequence, and the system comprises at least one processor communicating with at least one memory device, the at least one memory device storing instructions causing the at least one processor to perform the method according to any one of claims 1 to 25.

28. A computer-readable medium storing computer-executable instructions for implementing the method described in any one of claims 1 to 25.

29. A method for producing a diagnostic assay to determine whether a patient is infected with or has been infected with a pathogen, the diagnostic assay being performed on a biological sample obtained from a subject and comprising identifying at least one candidate region of at least one source protein of the pathogen using the method according to any one of claims 1 to 25, The diagnostic assay is a method comprising utilizing or identifying the at least one identified candidate region and / or at least one predicted epitope within the at least one candidate region in the biological sample.

30. A diagnostic assay for determining whether a patient is infected with or has been infected with a pathogen, the diagnostic assay being performed on a biological sample obtained from a subject, and comprising utilizing or identifying in the biological sample at least one candidate region of at least one source protein of the pathogen identified using the method according to any one of claims 1 to 25 and / or at least one predicted epitope within the at least one candidate region.

31. The method according to claim 29, wherein the diagnostic assay includes identifying an immune system component in the biological sample that recognizes the at least one identified candidate region and / or at least one predicted epitope within the at least one candidate region.

32. The diagnostic assay according to claim 30, comprising identifying an immune system component in the biological sample that recognizes the at least one identified candidate region and / or at least one predicted epitope within the at least one candidate region.