Method and computer program for predicting neoantigens by processing the lengths of peptides of various lengths through folding

By folding peptide sequences of non-uniform lengths into a unified unit length and using a learned model to predict binding strength with HLA alleles, the method addresses the challenges of neoantigen prediction, achieving accurate determination of neoantigens in cancer tissues.

JP2025519177APending Publication Date: 2025-06-24THERAGEN BIO CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024570337
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-11-25
Filing Date
2023-08-08
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

Existing methods for predicting neoantigen potential struggle with the diversity of neoantigen sequences and the challenge of non-uniform peptide lengths, leading to data insufficiency and information loss.

Method used

A method that processes peptide sequences of non-uniform lengths by folding them into a unified unit length, using a model learned from such data to predict binding strength with HLA alleles and determine neoantigens based on predicted binding strength and immunogenicity.

Benefits of technology

This approach effectively predicts the immunity and binding affinity of peptides, enabling accurate determination of neoantigens in cancer tissues, thereby overcoming the limitations of existing prediction models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025519177000001_ABST
    Figure 2025519177000001_ABST
Patent Text Reader

Abstract

The present disclosure discloses a method for processing a peptide sequence exceeding a unit length in the folding process when predicting neoantigens using peptide sequences and HLA class I and / or II allele sequences. According to this, peptide sequences contained in cancer tissues can determine neoantigens within the cancer tissues regardless of the diversity of lengths. Through this, it is possible to overcome the imbalance and lack of information of learning data by length and predict a more accurate binding force to determine neoantigens within cancer tissues.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present disclosure relate to a method and a computer program for predicting the neoantigen potential of peptides extracted from cancer tissues and cancer-derived DNA against HLA class I and / or II alleles.

Background Art

[0002] Cancer is one of the most common causes of death worldwide. Approximately ten million new cases occur each year, accounting for about 12% of all deaths and ranking as the third most common cause of death. In recent years, with the development of new anti-cancer drugs, through first-generation anti-cancer drugs, chemical anti-cancer drugs, and second-generation targeted anti-cancer drugs, in recent years, third-generation immune anti-cancer drugs have been in the spotlight. In particular, in the case of third-generation immune anti-cancer drugs, unlike previous anti-cancer drugs, it is a treatment strategy that utilizes the patient's own immune system, so it has the advantage of significantly lower side effects. However, despite such advantages, only patients showing the expression of biomarker genes such as PD-L1 and microsatellite instability (MSI-H) have limitations in establishing a treatment strategy using immune anti-cancer drugs. Due to such limitations, it is necessary to establish a strategy for treating patients for whom it is difficult to administer existing anti-cancer drugs, and one of the alternatives presented is precisely a cancer vaccine that utilizes neoantigens.

[0003] Our immune system is trained not to attack normal cells by removing T cells that recognize self in the thymus (thymic selection), which is called central tolerance. Cancer is a disease that occurs while mutations accumulate in the genome. Solid cancer cells such as gastric cancer, colorectal cancer, and breast cancer have an average of about 60 mutations, while lung cancer and melanoma caused by carcinogens have more, about 150 mutations.

[0004] If a mutation occurs, the immune system will produce abnormal proteins that have not been seen before. In this case, since the immune system has not come into contact with the mutant protein during the thymic selection process, immune central tolerance does not act on the protein with the mutation. Therefore, all mutant proteins are potentially considered to have immunogenicity. However, most mutations involve intracellular proteins, which cannot reach the surface of tumor cells and thus cannot directly contact immune cells. Mutations can only contact immune cells through a method of presenting epitopes, which are short peptide fragments 8 to 10 in length, on the surface of cancer cells through the class 1 Major Histocompatibility Complex (class 1 MHC, also called Human Leukocyte Antigen (HLA) in humans). At this time, the mutant peptide bound to class 1 MHC and presented on the surface of tumor cells is called a neoepitope or neoantigen, and the complex of the neoepitope and MHC is called pMHC (or pHLA in the case of humans). Such neoepitopes can be said to be a kind of label that helps CD8 cytotoxic T cells recognize cancer cells as external attackers.

[0005] However, for the immune response against the mutations possessed by cancer cells to be activated, the cancer cell-derived neoantigen pMHC must be cross-presented to T-immune cells by antigen-presenting cells (APCs) that are not cancer cells, particularly dendritic cells, together with co-proteins that are only expressed in dendritic cells. Unlike other somatic cells, APCs present longer peptides as neoantigens through MHC class II, which activates CD4 T helper cells, which are accessory cells, and ultimately all of class 1 pMHC, class 2 pMHC, and co-stimulatory molecules act to generate CD8 cytotoxic T cells with strong immunity against neoantigens.

[0006] The T-immune cells activated through such a process attempt to attack cancer cells, but cancer cells express a shield called an immune checkpoint to paralyze the activated T-cells and suppress the immune response (T cell exhaustion). Currently, most immunotherapeutic agents in clinical application are immune checkpoint inhibitors. Immune checkpoint inhibitors are effective only in some cancers where there are many already-activated and paralyzed T-cells. However, in many cases, a state of immunological ignorance occurs where T-cells cannot be activated in the first place against neoantigens, and an immune response cannot occur. In recent years, it has been revealed that during neoantigen-targeted cancer vaccine treatment, the phenomenon of activation of the immune response also occurs against neoantigens that had immunological ignorance. That is, neoantigen-targeted cancer vaccines are expected to be applicable to cancers that were not the target of existing immune checkpoint inhibitors, and a greater therapeutic effect can be expected during combination therapy even in the group of patients targeted by immune checkpoint inhibitors.

[0007] Previous cancer vaccines targeted "tumor-associated antigens," which are "normal" proteins that are abnormally highly expressed in cancer cells compared to normal cells. By using such commonly appearing targets, off-the-shelf vaccines can be made, but there are drawbacks such as low antigenicity due to central tolerance, and there is still no vaccine that has proven a definite clinical effect. Different from such "tumor-associated antigens," neoantigens are not affected by central tolerance and can thus be the target of an ideal cancer vaccine.

[0008] In contrast, the HLA gene that encodes MHC proteins has the highest polymorphism of more than 13,000. Such polymorphism causes differences in the peptide binding groove sequences of MHC proteins, and thus MHC proteins come to have various affinities with peptides. Therefore, even for peptides with the same mutation, they can be presented as neoantigens only in patients with specific HLA types that have a high affinity with some MHC proteins. Due to such MHC restriction, only less than 10% of the peptides with mutations are presented by MHC, and most neoantigens appear differently in each patient. For this reason, cancer vaccines have to be made customized for each patient.

[0009] One of the various problems that occur in the development process of cancer vaccines is caused by the diverse lengths of neoantigens that bind to HLA. Such diverse lengths come to limit the establishment of prediction model based on artificial intelligence including deep learning. Conventional methods have taken strategies such as independently developing models for each length or developing a model after correcting to a specific length in order to solve such problems. However, such preprocessing processes have drawbacks such as inducing insufficient learning data and loss of information. Therefore, the present invention attempts to solve the diversity of neoantigen sequences while minimizing data insufficiency and information loss. SUMMARY OF THE INVENTION

Problems to be Solved by the Invention

[0010] The present invention is based on the above-mentioned necessity, and predicts the immunity of a peptide by using a model learned from data in which peptide sequences of non-uniform lengths are each folded to be unified into a unit length, predicts the binding strength which is the binding strength with an HLA allele, and determines neoantigens in a cancer tissue based on the predicted binding strength and immunity.

Means for Solving the Problems

[0011] The method according to an embodiment of the present disclosure includes steps of: inputting, by a neoantigen prediction device, a peptide sequence extracted from a target cancer tissue, and determining whether the length of the peptide sequence exceeds a predetermined unit length; when the peptide sequence exceeds the unit length, determining, among the peptide sequence, a remaining region excluding a predetermined fixed region, and processing the value of the region exceeding the unit length for the remaining region by folding to reduce the peptide sequence to the unit length; inputting, by the neoantigen prediction device, a sequence of a complex of an HLA class I or II allele; extracting, by the neoantigen prediction device, a feature value from input data including sequence values included in the peptide sequence processed by the neoantigen prediction device, and outputting an immunity prediction value for the feature value; extracting, by the neoantigen prediction device, a peptide feature value from data processed to reduce the peptide sequence to the unit length; one-hot processing, by the neoantigen prediction device, an α-chain sequence or a β-chain sequence of the HLA class I or II allele, and extracting an HLA feature value corresponding to the processed chain sequence; outputting, by the neoantigen prediction device, a binding prediction value with a combination of the peptide and the HLA feature value; and determining, by the neoantigen prediction device, whether the peptide sequence is a neoantigen based on the immunity prediction value or the binding prediction value.

[0012] Reducing the peptide to the unit length is characterized by processing in a folding process such that, within the remaining region of the peptide, a first value at a first position and a second value at a second position are folded and represented by the sum of the first value and the second value, and the length of the peptide is reduced by 1.

[0013] The fixed region of the peptide includes a first value, a second value, and a last value, and the values included in the fixed region may not be folded. The folding process can be repeated until the peptide reaches the unit length.

[0014] The step of outputting the binding prediction value outputs the binding prediction value using a binding prediction model, and the binding prediction model takes as input a peptide feature value extracted from a peptide sequence, an HLA-α binding feature value extracted from an α-chain sequence of an HLA class I or II allele, and an HLA-β feature value extracted from a β-chain sequence of an HLA class II allele, and is a model learned to output a binding prediction value corresponding to the binding strength.

[0015] The binding prediction model further includes a peptide processing model for extracting peptide feature values, an HLAα processing model for extracting HLA-α feature values, and an HLAβ processing model for extracting HLA-β feature values, and the peptide, HLAα, and HLAβ binding processing models are in the form of encoders and may be different encoders.

[0016] The step of outputting the immunogenicity prediction value outputs the immunogenicity prediction value using an immunogenicity prediction model, and the immunogenicity prediction model takes as input a feature value extracted from a peptide sequence and is a model learned to output an immunogenicity prediction value corresponding to the immunogenicity of the peptide sequence.

[0017] The target cancer tissue can include cells engineered to express a single HLA class I and / or class II allele. The target cancer tissue can contain mutant cells obtained from or derived from a plurality of living organisms.

[0018] The target cancer tissue can contain fresh or frozen tumor cells obtained from a plurality of living organisms. The target cancer tissue can contain fresh or frozen tissue cells obtained from a plurality of living organisms.

Advantages of the Invention

[0019] According to one embodiment of the present invention made as described above, the immunity of a peptide is predicted using a model learned with data in which peptide sequences of non-uniform lengths are each folded and unified to a unit length, the binding affinity with an HLA allele is predicted, and neoantigens in cancer tissue can be determined based on the predicted binding affinity and immunity.

Brief Description of the Drawings

[0020]

Figure 1

Figure 2

Figure 3a

Figure 3b

Figure 4a

Figure 4b

Figure 4c

Figure 5

Figure 6

Mode for Carrying Out the Invention

[0021] Hereinafter, the configuration and operation of the present invention will be described in detail with reference to embodiments of the present invention shown in the accompanying drawings. The present invention can be subjected to various conversions and can have various embodiments. Specific embodiments are illustrated in the drawings and will be described in detail in the detailed description. The effects and features of the present invention, and the methods for achieving them, will become clear by referring to the embodiments described in detail later together with the drawings. However, the present invention is not limited to the embodiments disclosed below and can be embodied in various forms.

[0022] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings. When describing with reference to the drawings, the same or corresponding components are given the same reference numerals, and duplicate descriptions thereof are omitted.

[0023] In the following embodiments, terms such as first, second, etc. are not used in a limiting sense and are used for the purpose of distinguishing one component from another. In the following embodiments, singular expressions include plural expressions unless the context clearly indicates otherwise.

[0024] In the following embodiments, terms such as "including" or "having" mean that the features or components described in the specification exist, and do not preclude the possibility of adding one or more other features or components in advance.

[0025] In the drawings, for convenience of explanation, the size of components can be exaggerated or reduced. For example, the size and thickness of each configuration shown in the drawings are arbitrarily shown for convenience of explanation, so the present invention is not necessarily limited to what is shown.

[0026] When a certain embodiment can be implemented differently, the specific process order may be carried out differently from the order described. For example, two processes described consecutively may be carried out substantially simultaneously and may proceed in an order reverse to the described order.

[0027] Here, the target cancer tissue means the tissue to be the subject of the experiment. For example, the target cancer tissue may be a cancer tissue attempting to detect an antigen capable of causing an immune reaction. Preferably, the target cancer tissue may be an aggregate of tumor cells or cancer cells.

[0028] Here, a mutation means all phenomena in which the sequence of the bases A (adenine), T (thymine), G (guanine), and C (cytosine) of the gene containing the genetic information in each living body is altered so as to be different from the original genetic information of the corresponding species. Such mutations induce structural mutations on a small or large scale. Small-scale mutations include point mutations where a single base sequence is converted and appears, and there may also be mutations where a base sequence is further inserted or deleted. Large-scale mutations that occur and affect the structure include gene duplication, gene deletion, chromosomal inversion, interstitial deletion, chromosomal translocation, loss of heterozygosity, etc.

[0029] Mutations are broadly classified into germ cell mutations and somatic mutations according to the type of cells in which they occur. Somatic mutations are gene mutations that occur in somatic cells, also called somatic cell mutations or somatic mutations, and may be caused by gene mutations or chromosomal abnormalities.

[0030] Due to the occurrence of such mutations, changes may occur in the function of the protein produced by the corresponding gene, and a specific function may be lost or activated to another function. Such changes in protein function cause or accelerate cancer development, so such mutations may be directly or indirectly deeply related to cancer development and progression.

[0031] The base sequence in DNA containing the genetic information of a living organism consists of A, T, G, and C. When such base sequences gather in groups of three in a row, they form a code that forms one specific amino acid, and when several such codes gather, they can be converted into one protein. Amino acids consist of alanine (Ala), cysteine (Cys), aspartic acid (Asp), glutamic acid (Glu), phenylalanine (Phe), glycine (Gly), histidine (His), isoleucine (Ile), lysine (Lys), leucine (Leu), methionine (Met), asparagine (Asn), pyrrolysine (Pyl), proline (Pro), glutamine (Gln), arginine (Arg), serine (Ser), threonine (Thr), selenocysteine (Sec), valine (Val), tryptophan (Trp), and tyrosine (Tyr).

[0032] Peptide can mean a peptide or polypeptide formed by an amino acid sequence. In a living organism, there is an immune system for removing external substances not derived from the genetic information within various types. In particular, there are immunogenic peptides among externally derived peptides that can cause an immune reaction. Mutations that occur differently from the original genetic information during the cancer development process also generate such immunogenic peptides, and the immunogenic peptides thus generated can be capable of binding to HLA class I or HLA class II proteins through a series of processes within the immune system. Furthermore, the immunogenic peptide can have a mutated amino acid sequence, and the length of the amino acid can be 8 - 12 for a neoantigen that binds to HLA class I and 13 - 20 for a neoantigen that binds to HLA class II, but is not limited to this and can be of various lengths.

[0033] Here, machine learning refers to using a model composed of a large number of parameters and optimizing the parameters with the given data. Machine learning can be classified into supervised learning, unsupervised learning, and reinforcement learning according to the form of the learning problem. Supervised learning is to learn the mapping between input and output, and it can be applied when the input and output pairs are given by data.

[0034] In this specification, a neoantigen may be a peptide that elicits an immune response, that is, an immunogenic peptide. A neoantigen can be induced by specific mutations of tumor cells and can be represented by epitopes of tumor cells. In the following, for the sake of simplicity of explanation, the immunogenic peptide will be named and explained as a neoantigen.

[0035] The neoantigen prediction device according to an embodiment of the present invention can analyze the peptide sequence of a target cancer tissue and the HLA class I allele sequence of a living body to confirm whether a specific peptide of the target cancer tissue used for the treatment of the target cancer tissue can be determined as a neoantigen. The length of the peptide sequence of the target cancer tissue does not have to be constant.

[0036] Here, HLA (human Leukocyte antigen) is a protein present on the surface of human cells and may have various types of gene sequences for each corresponding living body. Also, the protein sequence of HLA is different for each living body. Based on this, each living body has different antigenicity.

[0037] FIG. 1 is a block diagram of a neoantigen prediction device 100 according to an embodiment of the present disclosure. The neoantigen prediction device 100 may include a first data input unit 111, a second data input unit 112, an immunogenicity prediction unit 120, a binding prediction unit 130, and a neoantigen prediction unit 140.

[0038] The neoantigen prediction device 100 can determine neoantigens, which are immunogenic peptide sequences for the treatment of cancer tissues and cancer cells, based on the peptide sequences of cancer tissues, cancer cells, etc. and the HLA class I and / or II sequences of the living body. At this time, the neoantigen prediction device 100 can determine neoantigens among the peptide sequences based on the immunogenicity of the peptide sequences and / or the binding affinity between the peptide sequences and HLA class I and / or II.

[0039] The neoantigen prediction device 100 can predict the binding affinity between the mutant-derived peptide and the specific HLA class I and / or II alleles of the living body in order to utilize the peptide derived from the mutation of the cancer tissue of the living body as a neoantigen so that the immune system of the corresponding living body can recognize and attack the neoantigen of the peptide.

[0040] The neoantigen prediction device 100 can determine the neoantigen by preprocessing the peptide sequence even for mutant-derived peptides having various lengths. The neoantigen prediction device 100 can fold the sequence values of some regions in the peptide sequence to unify the length of the peptide to a defined unit length. In the case of a peptide, folding can occur naturally, but it is possible to determine the regions where folding occurs and the regions where it does not occur even within the peptide region. More specifically, it is known that folding occurs naturally in the middle part of the peptide. The neoantigen prediction device 100 can be fixed so as not to perform folding processing on the regions where folding occurs with a relatively low probability in the peptide sequence. The neoantigen prediction device 100 can determine the fixed regions that are not subjected to folding processing in the peptide sequence based on the data on the folding of each peptide sequence.

[0041] The array value of the region to be folded by the peptide array can be calculated by applying a weighted value in a defined manner. The weighted value can be determined based on the number of folding times, the number of values combined by the folding process, etc. For example, when performing folding four times, a weighted value of 4 can be applied to the value of the folded region. When combining five values into one value, a weighted value of 5 can be applied to the value of the folded region.

[0042] The neoantigen prediction device 100 can predict the binding property corresponding to the binding force with the peptide sequence by preprocessing the complex of specific HLA class I and II alleles possessed by the living body. According to one embodiment, the neoantigen prediction device 100 can infer the binding property corresponding to the binding force with the peptide sequence after performing one-hot processing on the HLA class I and II alleles.

[0043] The neoantigen prediction device 100 can calculate an immunity prediction value by using an immunity prediction model learned with the data obtained by folding the peptide sequence as an input. The neoantigen prediction device 100 can output a binding property prediction value of the peptide sequence by using a binding property prediction model learned with the data obtained by folding the peptide sequence as an input.

[0044] Also, the neoantigen prediction device 100 can generate a binding property prediction model learned with the data obtained by encoding the HLA class I and / or II alleles by a method such as one-hot processing as an input. The neoantigen prediction device 100 can input the HLA class I and / or II alleles into the binding property prediction model and output a binding property prediction value of the HLA class I and / or II alleles. The neoantigen prediction device 100 can output an immunity prediction value by using the immunity prediction model, output a binding property prediction value by using the binding property prediction model, and output data about the neoantigen by using the neoantigen prediction model for the immunity prediction value and the binding property prediction value.

[0045] The first data input unit 111 inputs a peptide sequence extracted from a target cancer tissue or cancer cells. Among a plurality of peptides extracted from the target cancer tissue or cancer cells, a selected peptide sequence is input.

[0046] The first data input unit 111 can generate processed data by performing one-hot encoding on the input peptide sequence and then performing folding processing. More specifically, the first data input unit 111 can determine whether the length of the input peptide sequence exceeds a defined unit length. The first data input unit 111 can perform folding processing on the peptide sequence that exceeds the unit length to normalize it to the unit length.

[0047] Regarding the folding process, as shown in FIG. 3, the sequence values CD1 from the 3rd to the 11th, excluding P1, P2, and P12 fixed in the 12-mer peptide sequence PD1, can be folded to process the 11-mer peptide sequence PD2. To reduce the 11th sequence value in the 12-mer peptide sequence PD1, the 3rd and 4th sequence values of PD1 can be combined and converted into the 3rd sequence value. More specifically, the sum of the 3rd and 4th sequence values of PD1, or the 3rd or 4th sequence value of PD1, can be converted into the 3rd sequence value. A predetermined weighting value can be applied to the combined value of the 3rd and 4th sequence values of PD1 to convert it into the 3rd sequence value.

[0048] The 4th and 5th sequence values of PD1 can be combined and converted into the 4th sequence value. The 5th and 6th sequence values of PD1 can be combined and converted into the 5th sequence value. The 6th and 7th sequence values of PD1 can be combined and converted into the 6th sequence value. The 7th and 8th sequence values of PD1 can be combined and converted into the 7th sequence value. The 8th and 9th sequence values of PD1 can be combined and converted into the 8th sequence value. The 9th and 10th sequence values of PD1 can be combined and converted into the 9th sequence value. The 10th and 11th sequence values of PD1 can be combined and converted into the 10th sequence value.

[0049] The 11-mer peptide sequence PD2 can be folded to convert it into a 10-mer peptide sequence PD3. The 10-mer peptide sequence PD3 can be folded to convert it into a 9-mer peptide sequence PD4. The 9-mer peptide sequence PD4 can be folded to convert it into an 8-mer peptide sequence PD5.

[0050] The second data input unit 112 inputs the sequences of HLA class I and / or II for checking the neoantigen against the peptide. The second data input unit 112 can encode the sequences of HLA class I and / or II using a one-hot method or the like.

[0051] The immunogenicity prediction unit 120 can output an immunogenicity prediction value of the peptide sequence based on the data obtained by folding the peptide sequence. The immunogenicity prediction unit 120 can obtain the immunogenicity values for one or more amino acid sequences included in the folded peptide sequence from an external database. The immunogenicity prediction value of the peptide sequence for the combination of immunogenicity values can be output to an immunogenicity prediction model learned by machine learning.

[0052] The immunogenicity prediction model can be machine-learned using, as input data, one or more amino acid sequences included in the folded peptide sequence and, as output data, values corresponding to the immunogenicity of the peptide sequence. Alternatively, the immunogenicity prediction model can be machine-learned using, as input data, a combination of immunogenicity values for one or more amino acid sequences included in the folded peptide sequence and, as output data, values corresponding to the immunogenicity of the peptide sequence.

[0053] The immunogenicity prediction model can be learned by supervised learning in particular, but is not limited thereto, and can be learned by unsupervised learning or reinforcement learning. The immunogenicity prediction model can learn the input data using one or more neural networks and output an immunogenicity prediction value for the peptide sequence.

[0054] The immunity prediction model can be embodied to output an immunity prediction value corresponding to a peptide sequence based on the learned data even when there is no learning data for one or more amino acid sequences included in the folded peptide sequence. Using the immunity data for each amino acid sequence included in the peptide sequence as input, the immunity prediction value for an unlearned peptide sequence can be inferred. The immunity prediction model can infer the immunity prediction value for an unlearned peptide sequence based on the immunity values of each of the plurality of amino acid sequences included in the peptide sequence.

[0055] More specifically, the immunity prediction model can separate a peptide sequence into amino acid sequence units, convert it into values according to the presence or absence of each amino acid sequence, and fold the encoded data. For example, a peptide sequence can be converted into values (0 or 1) for the first to nth amino acid sequences. The converted values can become input data for the immunity prediction model.

[0056] The immunity prediction model is learned with input data corresponding to the folded peptide sequence and output data including the immunity value of the peptide sequence. At this time, the immunity value of the peptide sequence can be determined by the immunity values for one or more amino acid sequences included in the peptide sequence. The immunity value for one or more amino acid sequences included in the peptide sequence or the immunity value of the peptide sequence can be obtained from pre-stored data.

[0057] The immunity value for a novel peptide sequence can be further used to input to the immunity prediction model and learn the immunity prediction model. By inputting a novel peptide sequence and the immunity value therefor, the immunity prediction model can be continuously upgraded.

[0058] The binding prediction unit 130 can input the peptide sequence obtained from the first data input unit 111 and the HLA class I and / or II allele sequences obtained from the second data input unit 112, and output a binding prediction value corresponding to the binding force between the peptide sequence and the HLA class I and / or II allele sequences.

[0059] The binding prediction unit 130 can output a binding prediction value corresponding to the binding force between the folded peptide sequence and the one-hot processed HLA class I and / or II allele sequences.

[0060] A binding prediction model can be used to output the binding prediction value. The binding prediction model uses the processed data obtained by folding the peptide sequence and the data on the one-hot processed HLA allele sequences as input data, extracts the characteristic values corresponding to the input data, and can output the binding prediction value based on the characteristic values.

[0061] The one-hot encoding method may convert one or more amino acid sequences included in the peptide sequence or HLA allele sequence into their respective corresponding values. Further, according to the embodiments of the present disclosure, the peptide sequence can be one-hot encoded and the one-hot encoded peptide sequence can be folded. Values can be predefined for all amino acid sequences included in the peptide sequence experimentally or empirically. For example, a plurality of peptide sequences can be multi-aligned positionally, the peptide sequences can be arranged on the x-axis and y-axis, and based on the similarity between the peptide sequence on the x-axis and the peptide sequence on the y-axis, positions with high similarity and positions with low similarity can be extracted in the peptide sequence. At this time, it can be determined based on a reference similarity as a reference. Furthermore, a processing method for unifying the length of the peptide can be performed. An example of such a processing method may be a folding processing method.

[0062] It can be said that the folding processing method reduces the length of the peptide sequence by binding the sequence values of the region defined by the one-hot processed peptide sequence according to a defined rule. The binding prediction model can input the data obtained by processing the peptide sequence by the folding processing method and the data on the HLA allele sequence, and output the binding prediction value.

[0063] The binding prediction model can extract one or more feature values that can be utilized for the input data corresponding to the peptide sequence, HLA allele sequence, and the output of the binding between the peptide sequence and the HLA allele sequence. The types of feature values to be extracted can be determined by the result of machine learning using, as input data, combinations of binding values for the amino acid sequences included in the peptide sequence, and, as output data, values corresponding to the binding between the peptide sequence and the HLA allele sequence. The binding prediction model can train the input data using one or more neural networks to extract main feature values, and output a binding prediction value based on the main feature values.

[0064] Even when there is no learning data for the input peptide sequence, the binding prediction model can output a binding prediction value for a new peptide sequence based on data learned with other peptide sequences.

[0065] The binding prediction model can output a binding feature value corresponding to the folded peptide sequence after folding the peptide sequence according to a defined rule for the length of the peptide sequence. The binding prediction model can encode and fold a peptide to output a binding feature value for the peptide sequence.

[0066] The binding prediction model can use, as input data, values corresponding to the encoded HLA class I and / or II allele sequences, and can output, as output data, binding feature values corresponding to the HLA class I and / or II allele sequences.

[0067] The binding prediction unit 130 can take, as input, data on the peptide sequence and the HLA class I and / or II alleles, and output a binding prediction value corresponding to the binding force between the peptide sequence and the HLA class I and / or II alleles.

[0068] In a further embodiment, the binding prediction unit 130 can output a binding prediction value using a binding prediction model. The binding prediction model is learned by machine learning. As a result of the learning, the types of binding characteristic values for the peptide sequence and the binding characteristic values for the HLA class alleles can also be determined. In the binding prediction model, the binding prediction value utilized as output data may be a value empirically or experimentally obtained for the binding force between the peptide sequence and the HLA class I and / or II alleles, and may be a value obtained from an external database.

[0069] The neoantigen prediction unit 140 can output neoantigen data including the possibility of the presence of a neoantigen in the peptide sequence, what the neoantigen is, etc., based on the immunity prediction value and / or the binding prediction value. The neoantigen prediction unit 140 can output neoantigen data using a neoantigen prediction model learned with the immunity prediction value and the binding prediction value as inputs.

[0070] The neoantigen prediction model is learned with the immunity prediction value and / or the binding prediction value as input data, and can be learned with the neoantigen data of the peptide sequence for the HLA class I and / or II alleles as output data, but is not limited thereto, and may be learned by unsupervised learning or reinforcement learning. The neoantigen data can include the presence or absence of the peptide sequence of the neoantigen, the position of the neoantigen, the specificity of the neoantigen, the possibility of neoantigen by position, etc.

[0071] The immunity prediction value and the binding prediction value can include one or more 0s or 1s, but are not limited thereto, and can include variously defined values between 0 and 1. The neoantigen prediction device 100 according to the embodiment of the present disclosure can output neoantigen data including the presence or absence of the peptide sequence of the neoantigen, the main position within the neoantigen, the specificity of the neoantigen, etc., for the peptide sequence derived from the mutation of the target cancer tissue. In particular, the neoantigen prediction device 100 is meaningful in terms of determining whether there is a neoantigen in the peptide sequence for the HLA class I and / or II alleles.

[0072] FIG. 2 is an exemplary diagram of a learning model 200 according to an embodiment of the present disclosure. The learning model 200 can include a neoantigen prediction model M that takes a peptide sequence i1 and HLA class I and / or II alleles i2 as inputs and outputs immunogenicity data of the peptide.

[0073] The peptide sequence i1 can be input into the neoantigen prediction model M after undergoing a one-hot processing step P11 and then a folding processing step P12. The one-hot processed peptide sequence is as shown in FIG. 4a. The folding processing step P12 can convert the combined value into a first value calculated according to a defined rule, and apply a weighting value to the calculated first value to calculate a final second value. Converting the combined value into a first value according to a defined rule is as shown in FIG. 4b. The defined rule may be to sum the combined values, but is not limited thereto, and may be calculated as a weighted average value, an average value, etc. In the learning model 200, the peptide sequence i1 can be folded in units of two values at a time until it reaches the unit length (see FIG. 3a). Applying a weighting value to the first value to calculate a final second value is as shown in FIG. 4c. The weighting value applied at this time may be a value proportional to the number of combined values or the number of folding times, but is not limited thereto, and may be determined by combining various values.

[0074] The HLA class I and / or II alleles i2 can be input into the neoantigen prediction model M after undergoing a one-hot processing step P21. The neoantigen prediction model M can output data on the immunogenicity of the peptide sequence i1 with the peptide sequence i1 processed by P11 and P12 and the HLA class I and / or II allele i2 processed by P21 as inputs. At this time, the length of the peptide sequence i1 may be folded into a predetermined unit length, for example, 8mer. The peptide sequence i1 may go through a process of being folded into a unit length for the region excluding the fixed region after one-hot processing. The fixed region that is not folded can include, but is not limited to, the first, second, and last values, and can include important regions. The HLA class I allele i2 may be one-hot processed based on a defined amino acid sequence.

[0075] The neoantigen prediction model M can be designed to include layers of neural networks NN(256), NN(128), and NN(64). As a result of inputting a training dataset into the neoantigen prediction model M and learning, the layers of NN(256), NN(128), and NN(64) can be completed. Here, the number of neurons in NN and the structure of the layers are not limited to the above examples and may be in other forms.

[0076] The neural network may be a set of algorithms that utilize the results of statistical machine learning to extract various attribute information of the input data and identify and / or judge the objects in the image based on the extracted attribute information. The neural network can be embodied in software or an engine for executing the set of algorithms. The neural network embodied in software or an engine can be executed by a processor in a device or a processor of a server.

[0077] The neoantigen prediction model M can include an immunogenicity prediction model for inferring the immunogenicity of the peptide sequence and a binding prediction model. The binding prediction model can infer the binding between the peptide sequence and the HLA class I and / or II alleles.

[0078] The neoantigen prediction model M according to one embodiment can output an output value, which is data about neoantigens, by using a model learned with the output values of an immunity prediction model and a binding prediction model between HLA class I and / or II and peptides as inputs.

[0079] The learning model 200 can be embodied in software or hardware, or can be embodied by the combination of software and hardware. The learning model 200 can be provided inside the neoantigen prediction device 100, or can be provided in a device separate from the neoantigen prediction device 100. The learning model 200 can be connected to a plurality of neoantigen prediction devices through a network or electrically connected. The learning model 200 can receive data necessary for learning from an external device. The learning model 200 can transmit the learned prediction model to the neoantigen prediction device so as to output neoantigen data for the peptide sequence.

[0080] FIG. 3a is an exemplary diagram of a peptide sequence processed by the method according to an embodiment of the present disclosure. FIG. 3b is an exemplary diagram of input data processed by a conventional method. The one-hot processed peptide sequence PD1 has a length exceeding 20×12 and unit length, and folding processing is required. The values of P1, P2, and P12, which are fixed regions determined in the peptide sequence PD1, can be fixed without folding processing. The fixed region can be set differently for each peptide sequence. The fixed region is determined based on the probability of folding processing occurring in each peptide sequence, and can be set in a region where the probability of folding processing occurring is low. By performing folding processing after fixing some regions, the immunity and binding properties of the peptide sequence can be predicted more accurately.

[0081] If the region CD1 excluding the fixed region in the peptide sequence PD1 is folded once, it is as shown in PD2. The values of the third and fourth positions are folded to obtain the third position value, and the values from the third position to the tenth position CD2 are calculated. The length of the peptide sequence folded once can be changed to 20×11.

[0082] If the region CD2 excluding the fixed region in the peptide sequence PD2 is folded once again, it is as shown in PD3. The 9th value CD3 is calculated from the 3rd value in such a way that the 3rd and 4th values after the first folding process are folded to calculate the 3rd value. The length of the peptide sequence after two folding processes can be changed to 20×10.

[0083] If the region CD3 excluding the fixed region in the peptide sequence PD3 is folded once again, it is as shown in PD4. The 8th value CD4 is calculated from the 3rd value in such a way that the 3rd and 4th values after the second folding process are folded to calculate the 3rd value. The length of the peptide sequence after three folding processes can be changed to 20×9.

[0084] If the region CD4 excluding the fixed region in the peptide sequence PD4 is folded once again, it is as shown in PD5. The 7th value CD5 is calculated from the 3rd value in such a way that the 3rd and 4th values after the third folding process are folded to calculate the 3rd value. The length of the peptide sequence after four folding processes can be changed to 20×8.

[0085] In FIG. 3a, the fixed regions are P1, P2, and P12. However, in order to increase the prediction probability of immunogenicity, the fixed regions can be changed according to the target cancer, HLA class I, or the characteristics of the peptide sequence.

[0086] The folding method can also be changed in order to increase the prediction probability of immunogenicity. When applying the method of extracting the array that gives the maximum value of the diagonal sum after taking the inner product of the PPSM matrix of the conventional peptide sequence and MHC data, there may arise a problem that one of the information of the two anchors will inevitably be lost. Also, when using the PSSM matrix, there was a problem that data had to be preprocessed manually every time new reference data was updated. As shown in Fig. 3b, when applying the method of extracting the array that gives the maximum value of the diagonal sum after taking the inner product of the PPSM matrix of the peptide sequence and MHC data, it can be seen that the 10th value, the 1st value, and the 1st to 3rd values are lost (see OD).

[0087] Fig. 4a is an exemplary diagram of a one-hot encoded peptide sequence according to an embodiment of the present disclosure. Fig. 4b is an exemplary diagram of a folded peptide sequence according to an embodiment of the present disclosure.

[0088] Fig. 4c is an exemplary diagram of a peptide sequence to which weighted values are applied according to an embodiment of the present disclosure. With the peptide sequence P41 in FIG. 4a, the remaining region D41 excluding the defined fixed region can be changed to a combined value D42 according to a defined rule to generate P42 in FIG. 4b. At this time, in order to combine a 12×20 matrix into an 8×20 matrix, 5 values must be combined into 1 value. As shown in FIG. 4b, the third array value can be determined by combining the third to seventh array values according to a defined rule. The defined rule may be one of the sum value, weighted average value, and average value. Optionally, after being combined according to a defined rule, respective weighting values can be applied to each array value to generate P43. As shown in FIG. 4c, data P43 and D43 processed by applying a multiple of 5 to each combined value of the peptide sequence can be generated. The multiple of 5 is only one example, and other multiples can be applied. In FIGS. 4a to 4c, a 12×20 peptide sequence is changed to an 8×20 peptide sequence with a unit length of 8, but peptide sequences of various sizes such as 11×20, 10×20, 13×20, etc. can be changed to a defined unit length.

[0089] FIG. 5 is a flowchart of an immunogenicity prediction method according to an embodiment of the present disclosure. In S110, the neoantigen prediction device 100 inputs a peptide sequence extracted from a target cancer tissue. The neoantigen prediction device 100 can determine whether the length of the peptide sequence exceeds a defined unit length.

[0090] In S120, when the length of the peptide sequence exceeds a defined unit length, the neoantigen prediction device 100 sets the remaining region excluding the already determined fixed region in the peptide sequence, and can process the value of the region exceeding the unit length for the remaining region by folding to reduce the peptide sequence to the unit length (S130).

[0091] In S140, the neoantigen prediction device 100 inputs the sequence of the HLA class I allele complex, and can process the sequence of the complex by a one-hot processing method. In S150, the neoantigen prediction device 100 can extract feature values from input data including sequence values contained in a peptide sequence and output an immunity prediction value for the feature values. At this time, the neoantigen prediction device 100 can output an immunity prediction value using a learned immunity prediction model.

[0092] In S160, the neoantigen prediction device 100 can perform one-hot processing on data obtained by folding a peptide sequence and HLA class I and / or II allele sequences, input the processed sequences into a binding prediction model, and output a binding prediction value.

[0093] In S170, the neoantigen prediction device 100 can output neoantigen data of a peptide sequence for an immunity prediction value and / or a binding prediction value using an immunogenicity prediction model. FIG. 6 is a diagram comparing the neoantigen prediction model according to an embodiment of the present disclosure with a previous neoantigen prediction model.

[0094] Experimentally derived that the AUC of the result predicted with data obtained by folding the peptide sequence as input is 0.822, which is significantly better than the AUC value of 0.720 of the result predicted with data obtained by processing the peptide sequence by other processing methods.

[0095] The neoantigen prediction device 100 can output an immunity prediction value of a peptide sequence from a combination of one or more amino acid sequences contained in the folded peptide sequence and immunity values corresponding to the one or more amino acid sequences.

[0096] The neoantigen prediction device 100 can obtain immunity values for one or more amino acid sequences contained in a peptide sequence from an external database. The immunity prediction value of the peptide sequence for the combination of immunity values can be output to an immunity prediction model learned by machine learning.

[0097] An immunity prediction model can be machine - learned using, as input data, one or more amino acid sequences contained in a peptide sequence and, as output data, values corresponding to the immunity of the peptide sequence. Alternatively, the immunity prediction model can be machine - learned using, as input data, combinations of immunity values for one or more amino acid sequences contained in a peptide sequence and, as output data, values corresponding to the immunity of the peptide sequence. In particular, it can be learned by supervised learning, but is not limited thereto and can be learned by unsupervised learning or reinforcement learning. The immunity prediction model can output an immunity prediction value by having the input data learned by one or more neural networks to extract main feature values. The immunity prediction value can include one or more immunity prediction values.

[0098] Even when there is no learning data for one or more amino acid sequences contained in a peptide sequence, the immunity prediction model can output an immunity prediction value corresponding to the peptide sequence based on the learned data.

[0099] The immunity prediction model can perform an embedding process that separates a peptide sequence into amino acid sequence units and converts it into values based on the presence or absence of each amino acid sequence. For example, a peptide sequence can be converted into values (0 or 1) for the first to the nth amino acid sequences. The converted values can be input data for the immunity prediction model.

[0100] The immunity prediction model is learned using input data corresponding to a plurality of peptide sequences and output data including the immunity values of the peptide sequences. At this time, the immunity value of the peptide sequence can be determined by the immunity values for one or more amino acid sequences contained in the peptide sequence. The immunity value of the peptide sequence can be obtained from pre - stored data. The immunity value for a novel peptide sequence can be further input into the immunity prediction model and utilized to learn the immunity prediction model. The immunity prediction model can be updated through the novel peptide sequence and its corresponding immunity value.

[0101] The binding prediction model can use the folded peptide sequence as input data and extract the characteristic values corresponding to the input data. The types of characteristic values extracted by the binding prediction model can be determined by the result of machine learning with the combination of binding values for the amino acid sequences contained in the peptide sequence as input data and the values corresponding to the binding of the peptide sequence as output data.

[0102] The binding prediction model can input the data obtained by processing the peptide sequence by a folding method and the values corresponding to the encoded HLA class I and / or II allele sequences, and output a binding prediction value.

[0103] The binding prediction model can train the input data with one or more neural networks to extract the main characteristic values, and output a binding prediction value based on such main characteristic values. Even when there is no learning data for the peptide sequence, the binding prediction model can output a binding prediction value by outputting the binding characteristic values for the peptide sequence based on the learned data.

[0104] The binding processing model can fold the peptide sequence to unify the length of the peptide sequence to a unit length, then calculate the amino acid usage characteristics of the bound / non-bound peptides for each length, and extract the corresponding peptide characteristic values. In other embodiments, the binding prediction model can convert the peptide sequence into a vector by one-hot encoding, and output a binding prediction value through a plurality of neural networks.

[0105] In an alternative embodiment, the binding prediction model includes one or more prediction models, and each processing model can use, as input data, either an HLA class I allele or the α-chain or β-chain of an HLA class II allele, respectively. Each prediction model can output a binding characteristic value corresponding to an HLA class I and / or class II allele. The binding prediction model can cause input data to be learned by one or more neural networks to extract principal characteristic values, and can output a binding prediction value between a peptide corresponding to the principal characteristic values and an HLA allele.

[0106] The binding prediction model can obtain an α-chain sequence of an HLA class I or II allele extracted from a peptide sequence, and input the α-chain sequence of the HLA class I or II allele into an HLAα processing model to extract an HLA-α binding characteristic value. The binding prediction model can input the β-chain sequence of an HLA class II allele into an HLAβ processing model to extract an HLA-β binding characteristic value.

[0107] The binding prediction model may be a model learned with, as input, a peptide characteristic value extracted from a peptide sequence, an HLA-α binding characteristic value extracted from an α-chain sequence of an HLA class I or II allele, and an HLA-β binding characteristic value extracted from a β-chain sequence of an HLA class II allele, and outputting, as output, a binding prediction value corresponding to binding strength.

[0108] The binding prediction model can extract a peptide characteristic value using a peptide processing model from data processed to reduce a peptide sequence to a unit length. The binding prediction model can one-hot process an α-chain sequence or β-chain sequence of an HLA class I or II allele, extract an HLA characteristic value corresponding to the processed chain sequence, and output a binding prediction value as a combination of the peptide and the HLA characteristic value.

[0109] The neoantigen prediction device 100 can output data regarding immunogenicity by combining an immunity prediction value and a binding prediction value for a peptide and an HLA. The neoantigen prediction device 100 can output neoantigen data including the possibility of the presence of neoantigens in a peptide sequence, the positions of neoantigens, the possibility of neoantigens for each position, etc., based on at least one of the immunity prediction value and the binding prediction value of a peptide to HLA.

[0110] The neoantigen prediction device 100 can output neoantigen data by using a neoantigen prediction model learned with the immunity prediction value and / or the binding prediction value of a peptide to HLA as input.

[0111] The neoantigen prediction model is learned with the immunity prediction value and / or the binding prediction value of a peptide to HLA as input data, and can be learned with neoantigen data of a peptide sequence for HLA class I and / or II alleles as output data, but is not limited thereto, and may be learned by unsupervised learning or reinforcement learning.

[0112] The immunity prediction value and the binding prediction value of a peptide to HLA can include one or more 0s or 1s, but are not limited thereto, and can include variously defined values. The apparatuses described above may be embodied in hardware components, software components, and / or combinations of hardware components and software components. For example, the apparatuses and components described in the embodiments may be embodied using one or more general-purpose computers or special-purpose computers, such as, for example, a processor, a controller, an ALU (arithmetic logic unit), a digital signal processor, a microcomputer, an FPGA (field programmable gate array), a PLU (programmable logic unit), a microprocessor, or some other device capable of executing and responding to instructions. The processing device may execute an operating system (OS) and one or more software applications executed on the operating system. Further, the processing device may also access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device may sometimes be described as if one were used, but those of ordinary skill in the art will understand that the processing device may include a plurality of processing elements and / or multiple types of processing elements. For example, the processing device may include a plurality of processors or one processor and one controller. Also, other processing configurations, such as a parallel processor, are possible.

[0113] Software can include a computer program, code, instruction, or a combination of one or more of these, and can configure a processing device to operate as desired or can instruct the processing device independently or collectively. Software and / or data can be permanently or temporarily embodied in a certain type of machine, component, physical device, virtual equipment, computer storage medium or device, or signal wave being transmitted, in order to be analyzed by the processing device or to provide instructions or data to the processing device. Software can be distributed on a computer system connected by a network and stored or executed in a distributed manner. Software and data can be stored in one or more computer-readable writeable media.

[0114] The method according to the embodiment can be embodied in the form of program instructions to be executed through various computer means and can be written on a computer-readable medium. The computer-readable medium can include program instructions, data files, data structures, etc. alone or in combination. The program instructions written on the medium can be those specially designed and configured for the embodiment or those known to and usable by those skilled in the art of computer software. Examples of computer-readable writeable media include magnetic media such as hard disks, floppy (registered trademark) disks, and magnetic tapes, optical media such as CD-ROMs, DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program instructions such as ROMs, RAMs, flash memories, etc. Examples of program instructions include not only machine language codes such as those made by compilers but also high-level language codes that can be executed by a computer using an interpreter or the like. The above-described hardware device can be configured to operate as one or more software modules to perform the operations of the embodiment, and vice versa.

[0115] As described above, the embodiment has been described with reference to the limited embodiments and drawings. However, various modifications and variations are possible for those having ordinary knowledge in the relevant technical field from the above description. For example, whether the described technology is executed in a different order from the described method, and / or whether the components of the described system, structure, device, circuit, etc. are combined or assembled in a different form from the described method, or replaced or substituted by other components or equivalents, appropriate results can be achieved.

[0116] Therefore, other embodiments, other implementations, and those equivalent to the claims also fall within the scope of the claims described hereinafter.

Claims

1. A step in which a neoantigen prediction device inputs a peptide sequence extracted from a target cancer tissue and determines whether the length of the peptide sequence exceeds a predetermined unit length; When the peptide sequence exceeds the unit length, in the peptide sequence, a remaining region excluding a predetermined fixed region is determined, and a value of a region exceeding the unit length in the remaining region is subjected to a folding process so as to reduce the peptide sequence to the unit length; A step in which the neoantigen prediction device inputs a sequence of a complex of HLA class I or II alleles; A step in which the neoantigen prediction device extracts a feature value from input data including a sequence value included in the processed peptide sequence and outputs an immunity prediction value for the feature value; A step in which the neoantigen prediction device extracts a peptide feature value from data processed so as to reduce the peptide sequence to the unit length; A step in which the neoantigen prediction device performs one-hot processing on an α-chain sequence or a β-chain sequence of the HLA class I or II allele and extracts an HLA feature value corresponding to the processed chain sequence; A step in which the neoantigen prediction device outputs a binding prediction value based on a combination of the peptide feature value and the HLA feature value; and A step in which the neoantigen prediction device determines whether a peptide sequence is a neoantigen based on the immunity prediction value or the binding prediction value; A method for predicting a neoantigen by processing the lengths of peptides of various lengths by folding.

2. Reducing the peptide to the unit length In the remaining region of the peptide, the first value at the first position and the second value at the second position are folded and represented by the sum of the first value and the second value, and the length of the peptide is reduced by 1 in the folding process; A method for predicting a neoantigen by processing the lengths of peptides of various lengths by folding according to claim 1.

3. The fixed region of the peptide Includes the first value, the second value, and the last value, The values included in the fixed region are not folded. A method for predicting a neoantigen by processing the lengths of peptides of various lengths by folding according to claim 1.

4. The method for predicting a neoantigen by processing the lengths of peptides of various lengths by folding according to claim 2, wherein the folding process is repeated until the peptide reaches the unit length.

5. The step of outputting the binding prediction value Utilize a binding prediction model to output a binding prediction value, The binding prediction model is a model learned by taking as input peptide feature values extracted from a peptide sequence, HLA-α feature values extracted from the α-chain sequence of an HLA class I or II allele, and HLA-β feature values extracted from the β-chain sequence of an HLA class II allele, and outputting a binding prediction value corresponding to the binding affinity. The method for predicting neoantigens by processing the lengths of peptides of various lengths in a folded manner according to claim 1.

6. The binding prediction model is further includes a peptide processing model for extracting peptide feature values, an HLAα processing model for extracting HLA-α feature values, and an HLAβ processing model for extracting HLA-β feature values, The peptide, HLAα, and HLAβ binding processing models are in the form of encoders, and are different encoders respectively. The method for predicting neoantigens by processing the lengths of peptides of various lengths in a folded manner according to claim 5.

7. The step of outputting the immunogenicity prediction value is using an immunogenicity prediction model to output an immunogenicity prediction value, The immunogenicity prediction model is a model learned by taking as input feature values extracted from a peptide sequence and outputting an immunogenicity prediction value corresponding to the immunogenicity of the peptide sequence. The method for predicting neoantigens by processing the lengths of peptides of various lengths in a folded manner according to claim 1.

8. The target cancer tissue is cells engineered to express a single HLA class I or class II allele. The method for predicting neoantigens by processing the lengths of peptides of various lengths in a folded manner according to claim 1.

9. The target cancer tissue is mutant cells obtained from or derived from a plurality of living organisms. The method for predicting neoantigens by processing the lengths of peptides of various lengths in a folded manner according to claim 1.

10. The target cancer tissue is fresh or frozen tumor cells obtained from a plurality of living organisms. The method for predicting neoantigens by processing the lengths of peptides of various lengths in a folded manner according to claim 1.

11. The target cancer tissue is fresh or frozen tissue cells obtained from a plurality of living organisms. The method for predicting neoantigens by processing the lengths of peptides of various lengths in a folded manner according to claim 1.

12. A computer program stored in a computer-readable storage medium for causing a computer to execute the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Method for predicting neoantigen using a peptide sequence and HLA class ii allele sequence and computer program

    KR102278727B1

  • Method and computer program to predict neoantigen for pan-HLA class ii allele sequences and pan-length of peptides

    KR102330099B1