Methods and systems for discovering cancer-rich motifs from splicing mutations in tumors
A machine learning model identifies cancer-enriched motifs from splicing mutations in tumors, enhancing neoantigen detection and immunotherapy by classifying alternative splicing events, addressing the limitations of existing methods in scalability and efficiency.
Patent Information
- Application Number
- JP2024574601
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2022-08-12
- Publication Date
- 2025-07-30
AI Technical Summary
Existing methods for identifying neoantigens from splicing mutations in tumors are time-consuming and difficult to generalize with new patient data, focusing mainly on somatic mutations and neglecting the potential of splicing mutations as a source of neoantigens.
A computer-implemented method using a machine learning model to identify cancer-enriched motifs from alternative splicing events by training on nucleotide sequences with labels, classifying healthy and disease samples, and selecting high attribution regions to generate DNA sequence motifs, which can be used to predict cancer presence and develop immunotherapy candidates.
The method efficiently identifies unique cancer-specific motifs, improving specificity and scalability in neoantigen detection, enabling personalized cancer vaccines and immunotherapy by leveraging deep learning for pan-cancer analysis.
Smart Images

Figure 2025524430000001_ABST
Abstract
Description
Technical Field
[0001] The present application relates to methods and systems for discovering cancer-rich motifs from splicing mutations in tumors.
Background Art
[0002] Tumors can alter the human transcriptome and generate neoantigens (i.e., foreign proteins) not present in normal tissues. Cancer neoantigens can be detected by the immune system and are thus used in immunotherapy to enhance the reactivity of T cells against tumor cells, i.e., T cells can eliminate affected tumor cells when activated by neoantigens. Previous approaches for identifying neoantigen candidates have focused on neoantigens derived from somatic mutations, but recent studies have shown that cancer-specific alternative splicing (AS) events represent an additional source of neoantigen candidates. Through AS, a single gene can generate multiple transcripts that join exons in alternative ways. These transcripts can be translated into different proteins. Similarly, cancer-specific AS events can alter gene transcripts, enrich specific (DNA) sequence motifs, and generate neoantigens.
[0003] In state-of-the-art research, cancer motifs are being investigated to predict transcription factor binding sites. In other approaches, cancer motifs are investigated from the perspective of mutation patterns widely found in the cancer genome. However, they have mainly addressed cancer mutation signatures, i.e., the nucleotide sequence context upstream and / or downstream of specific cancer mutations.
[0004] In recent years, various studies have focused on neoantigens generated from changes in somatic DNA such as non-synonymous point mutations, insertions-deletions, and frameshift mutations, but neoantigens derived from splicing mutations in tumors have received less attention.
[0005] For example, several studies, such as Kahles et al., "Comprehensive analysis of alternative splicing across tumors from 8,705 patients", Cancer cell 34.2 (2018): 211-224, have addressed the problem of identifying AS-based neoantigens. This study utilized RNA and whole-exome sequence data from tumor and healthy samples from the Cancer Genome Atlas (TCGA) and the Genotype-Tissue Expression (GTEx) portal. Their workflows identified significant changes in AS in tumors compared to normal tissues and enumerated novel splicing junctions (i.e., neo-junctions) in TCGA samples that did not naturally occur in GTEx normal samples. In addition, they investigated whether the discovered neo-junctions were translated into proteins. A subset of peptides derived from neo-junctions could be experimentally validated as potential cancer neoantigens.
[0006] ASNEO is another recent computational pipeline for identifying individualized AS-based neoantigens from RNA-seq data, as described in Zhang, Zhanbing, et al., "ASNEO: identification of personalized alternative splicing based neoantigens with RNA-seq", Aging (Albany NY) 12.14 (2020): 14633. ASNEO identifies novel gene isoforms based on novel splicing junctions that exist only in tumors. The identified novel isoforms are then translated into novel proteins. From the set of novel proteins, ASNEO generates a set of peptides as potential sources of cancer neoantigens. The ASNEO pipeline has been validated in two published immunotherapy treatment cohorts, and two important findings are that (i) AS-based neoantigens have higher immune scores compared to neoantigens identified from somatic DNA mutations, and (ii) AS-based neoantigens have the potential to predict patient survival patterns.
[0007] The two methods described above rely only on identifying novel splicing junctions within cancer samples as potential sources of neo-junctions. Identifying novel splicing junctions within tumor samples and verifying that such junctions do not exist in normal tissues can be a time-consuming step. Furthermore, using conventional approaches to identify and verify novel cancer junctions can be difficult to generalize when new patient data becomes available.
[0008] Smart, Alicia C., et al., "Intron retention is a source of neoepitopes in cancer," Nature biotechnology 36.11 (2018): 1056-1058, proposed an in silico approach to identify neoepitopes derived from intron retention events in tumors. In vivo validation was performed, and mass spectrometry was used to show that neoepitopes derived from intron retention events are presented on MHC I of cancer cells.
Prior art documents
Non-patent documents
[0009]
Non-patent document 1
Non-patent document 2
Non-patent document 3
Non-patent document 4
[0010] Therefore, providing a comprehensive approach for extracting motifs rich in cancer remains an important goal in the industry. [Means for Solving the Problems]
[0011] According to one aspect of the present invention, there is provided a computer-implemented method for training a machine learning model for use in identifying sequence motifs indicative of a disease from alternative splicing events, the method comprising: obtaining a list of nucleotide sequences, each sequence representing a step; obtaining a label for each sequence, the label indicating whether the sequence is associated with a healthy sample or a disease sample; and training a machine learning model with the list and each respective label to classify alternative splicing events from those sequences based on whether the sequences occur in healthy samples or disease samples.
[0012] Preferably, each array represents the DNA sequence of an exon of a mature messenger RNA (mRNA), or the constitutive exons and alternative exons of an alternative splicing event, and such sequences are exon sequences. Preferably, the nucleotide sequence is a DNA sequence such that the sequence motif is a DNA sequence motif, and by training a machine learning model with a list and their respective labels, based on whether the sequences occur in healthy samples or in diseased samples, alternative splicing events are classified from those DNA sequences. Preferably, an AS event is considered cancerous if it uniquely occurs in one or more cancer samples and does not occur in any other healthy samples. To obtain a list of cancer-specific AS events, all alternative splicing events within tumor samples can be analyzed and then a subset of events common to healthy and cancerous tissues can be excluded.
[0013] Preferably, each label indicates whether the sequence is associated with a sample of healthy tissue or a sample of cancerous tissue, and the machine learning model is trained to classify alternative splicing events from those sequences based on whether the sequences occur in healthy tissue or in cancerous tissue.
[0014] This method may further include identifying features of the sequences that contribute to indicating whether the sequences occur in healthy samples or in diseased samples by interpreting the trained model, selecting one or more high attribution regions of one or more sequences based on the identified features, where the high attribution regions indicate regions of the sequences that have features that contribute positively, and outputting a sequence motif based on the one or more high attribution regions. The DNA sequence motif can be considered a high attribution sub-sequence of the nucleotide sequence that contributes to whether the DNA sequence is associated with alternative splicing events that occur in diseased samples.
[0015] According to the concepts described herein, it is possible to identify unique cancer-enriched motifs from tumor-specific splicing mutations, and these motifs can provide a valuable source for cancer diagnosis and can function as a potential source of neoantigens.
[0016] The methods and systems described herein can be used to identify cancer-enriched motifs from splicing mutations in tumors and are not limited to so-called mutation signatures. Mutations in tumors can inhibit alternative splicing and lead to new splicing mutations. Therefore, it is proposed that by considering DNA mutations present in tumor samples, a comprehensive understanding of splicing mutations and the accompanying cancer-enriched motifs can be obtained depending on the implementation form. Therefore, the proposed systems and methods provide a more comprehensive perspective compared to the latest techniques for eliciting cancer-enriched motifs. Motifs based on splicing mutations in tumors have not been investigated so far.
[0017] The examples shown herein can identify motifs significantly enriched in different types of alternative splicing events. Alternative splicing events can be selected from the group consisting of exon skipping, alternative donor sites, alternative acceptor sites, intron retention, mutually exclusive exons, and other more complex splicing patterns. The proposal shown herein is not limited to specific alternative splicing events such as intron retention events, and the computational pipeline can identify motifs significantly enriched in any type of alternative splicing event.
[0018] Preferably, each nucleotide sequence represents the exon sequence of an alternative splicing event.
[0019] "Obtaining" means that the sequences and labels are retrieved from a data store or, if not, are generated and assigned by this method in preprocessing steps, etc. For example, the method may include obtaining a list of nucleotide sequences from a store and assigning a label to each sequence. The label can be a ground truth label. The method may include retrieving a list of nucleotide sequences and enumerating AS events based on sequencing read evidence.
[0020] Each region can have a length such that it is a contiguous sub-sequence of nucleotides present within a DNA sequence. Preferably, the length of the region exceeds a minimum threshold value and preferably the length can be 5 or more.
[0021] In certain implementations, each feature can be a nucleotide of the sequence. In this way, the trained model learns which nucleotides positively contribute to the occurrence of alternative splicing events in a disease sample. Alternatively, the features can be a group of nucleotides or can be based on the analysis of nucleotide sequences.
[0022] This method may include interpreting the predictions of a model trained using a saliency approach. Thus, the steps of selection may include generating attribution scores for each nucleotide in each sequence based on the contribution of that nucleotide to the classification, and selecting regions of the sequence based on the scores of each nucleotide. The attribution scores may be assigned based on Integrated Gradients, i.e., applying the Integrated Gradients interpretation approach to the trained model. Alternatively, the trained model can also be interpreted using Guided Backpropagation and occlusion maps. Other suitable interpretation methods can be used to assign scores to each feature for subsequent motif identification based on the relationship between the features and the input and output of the model.
[0023] The step of selecting regions of the sequence based on the scores of each nucleotide may include comparing the score of each nucleotide in the candidate region with a threshold, calculating the average score of the nucleotides in the candidate region, and calculating the average score of the sequence scores and comparing the score of each nucleotide in the candidate region with the average score of the sequence scores.
[0024] According to one aspect of the present invention, a method for identifying sequence motifs indicative of a disease from alternative splicing events may be provided, the method comprising extracting features of the sequence that positively contribute to indicating whether a nucleotide sequence representing the sequence of the alternative splicing event appears in a healthy sample or in a disease sample, extracting an attribution score for each feature, comparing the nucleotide sequence with the features, and identifying motifs within the sequence based on the scores for each feature within the nucleotide sequence.
[0025] According to one aspect of the present invention, a method for identifying an array motif indicating a disease from alternative splicing events may be provided. The method includes extracting the attribution score of each nucleotide in the alternative splicing array to indicate whether the nucleotide and its adjacent regions represent subsequences likely to occur in a healthy sample or a disease sample, and identifying a motif in the array based on the scores of consecutive regions in the nucleotide sequence that meet various predefined thresholds. Such regions may be referred to as motifs rich in disease.
[0026] The step of selection may include obtaining significant regions present in the sequence associated with the disease sample label and applying a hypergeometric test to one or more high-attribution regions to exclude regions common to the nucleotide sequence associated with the disease sample label and the nucleotide sequence associated with the healthy sample label.
[0027] This method may further include the step of merging a plurality of selected one or more regions using pairwise sequence alignment. By merging the regions, the total number of motifs is reduced, a list of representative motifs is provided, and the statistical likelihood that the motifs are related is improved. The merging step may reduce redundancy. For example, two motifs may have the same sequence in the center but one is shifted to the right or left, or in another example, one motif is a subset of another motif, so they are aligned and merged to obtain a list of representative motifs.
[0028] This method may further include filtering one or more selected regions based on the number of occurrences of each region in the sequence associated with the disease sample label. By filtering in this way, the statistical likelihood that the region is related is improved, which helps to eliminate artifacts and non-general motifs.
[0029] In a preferred implementation form, the DNA sequence motif can be output together with the corresponding genomic coordinates and a list of genes and transcripts in which the motif appears.
[0030] The DNA sequence motif can be utilized for cancer diagnosis and can serve as a potential source of cancer neoantigens, that is, it can function as a cancer treatment or immunotherapy.
[0031] The machine learning model can be any suitable machine learning model or statistical model. By way of example, the machine learning model can be a derivable parametric model. The machine learning model is a neural network, preferably a convolutional neural network with multi-kernel sizes for 1D signals. The machine learning model may also be referred to as a deep learning model.
[0032] Therefore, in a preferred implementation form, the method disclosed herein uses a deep learning-based approach to classify alternative splicing events in cancerous and normal tissues, in contrast to existing approaches that rely on enumerating all novel splicing in tumors by a one-to-one comparison with all splicing events that occur in normal samples. Using a deep learning-based model improves the specificity in identifying common cancer-specific splicing patterns in pan-cancer analysis.
[0033] In the benchmark analysis, it has been shown that the proposed convolutional neural network with multi-kernel sizes for 1D signals provides good performance in terms of training time, accuracy, F1, and Matthews correlation coefficient (MCC) score compared to state-of-the-art models.
[0034] This method may further include applying the trained model to one or more unknown nucleotide sequences to identify whether the nucleotide sequence corresponds to a healthy alternative splicing event or a disease-specific alternative splicing event. For example, normal or cancer-specific alternative splicing events and the like. In this way, the approach shown herein generalizes to new unknown alternative splicing events and does not require repeated manual comparison between tumor samples and healthy samples. This model is trained once on a large number of samples and can distinguish normal AS events from cancer-specific AS events by having an F1 score and a recall score exceeding 90%.
[0035] Preferably, this method may further include identifying neoantigen candidates or neoepitope candidates for immunotherapy based on sequence motifs.
[0036] An alternative method relies on identifying cancer neoantigens based on either cancer-specific splicing junctions or somatic DNA changes. On the other hand, this system identifies neoantigens based on cancer-enriched motifs from novel splicing mutations in tumors. The alternative method does not easily scale with additional datasets and requires more computational resources to assist in the comparison between alternative splicing events in normal and tumor samples, while this system includes a trained model that can efficiently classify cancerous and healthy AS events from those DNA sequences.
[0037] In addition, according to one aspect of the present invention, a method of identifying neoantigen candidates or neoepitope candidates for immunotherapy based on the sequence motifs output according to any of the above aspects or implementations may be provided.
[0038] Obtaining a list of candidates for neoepitopes or neoantigens derived from alternative splicing may involve steps of extracting k-mer peptides (where k≥9) that cover each of the cancer-rich motifs, excluding peptides present in proteomics datasets other than cancer, confirming potential neoepitopes or neoantigens in protein mass spectrometry (MS) databases obtained from various tumor types (such as MS data of the Clinical Proteomic Tumor Analysis Consortium (CPTAC)), predicting whether the neoepitopes or neoantigens bind to the major histocompatibility complex (MHC), and running existing tools to check their immunogenicity. Neoepitopes or neoantigens that elicit an immune response can be utilized in cancer immunotherapy. These validation steps are described in the literature (for example, Kahles et al., "Comprehensive analysis of alternative splicing across tumors from 8,705 patients", Cancer cell 34.2 (2018): 211-224). Additionally, peptides derived from newly identified alternative splicing events, which have not been detected previously and thus do not exist in standard MS reference databases, can be utilized when identifying candidates for novel neoepitopes or neoantigens in MS data from eluted peptide-MHC complexes. Finally, the candidate neoepitopes or neoantigens are utilized in vaccines for cancer immunotherapy, helping the immune system to recognize such epitopes or antigens and eliminate the cancer cells that produce them.
[0039] According to one aspect of the present invention, a method of producing a vaccine can be provided, the method comprising selecting one or more predicted immunogenic candidate amino acid sequences that cover the array motifs output according to any of the above aspects or implementations for inclusion in the vaccine; and synthesizing one or more amino acid sequences, or encoding one or more amino acid sequences into corresponding DNA or RNA sequences, and / or incorporating the DNA or RNA sequences into the genome of a bacterial or viral delivery system to produce the vaccine.
[0040] Furthermore, according to one aspect of the present invention, a method of predicting the likelihood that a tissue sample is cancerous can be provided, the method comprising retrieving the sequence of the tissue sample; analyzing the sequence of the tissue sample for similarity to the array motifs output according to any of the above aspects or implementations; and predicting the likelihood that the tissue sample is cancerous based on the analysis.
[0041] According to a further implementation of the present invention, each label indicates whether the sequence is positively associated with a genetic disorder, and the machine learning model can be trained to classify alternative splicing events from those sequences based on whether they are associated with a genetic disorder. The genetic disorder can be, for example, an autism spectrum disorder or a Mendelian disorder. Identifying significant motifs from alternative splicing events in genetic disorders enables the development of potential therapies for such disorders and can deepen the understanding of their origin, cause, and diagnosis. The concepts presented herein can be used to identify significant motifs from novel splicing mutations in genetic disorders.
[0042] According to one aspect of the present invention, there is provided a computer-implemented method for identifying a DNA sequence motif indicative of a disease from alternative splicing events, the method comprising: obtaining a list of nucleotide DNA sequences, each sequence representing the sequence of an alternative splicing event; obtaining a label for each sequence, the label indicating whether the sequence is positively associated with a healthy sample or a disease sample; training a machine learning model with the list and respective labels to classify alternative splicing events from those DNA sequences based on whether the sequences occur in healthy samples or disease samples; identifying, by interpreting the trained model, features of the sequences that positively contribute to indicating whether the sequences occur in healthy samples or disease samples; selecting, based on the identified features, one or more high attribution regions of one or more sequences, the high attribution regions indicating regions of sequences having features that positively contribute; generating a DNA sequence motif based on the one or more high attribution regions; identifying a neoantigen candidate or neoepitope candidate for immunotherapy based on the DNA sequence motif, or comparing the motif to a DNA sequence to predict the presence of cancer in a tissue based on the alternative splicing event.
[0043] According to one aspect of the present invention, a method for identifying unique cancer-enriched motifs from alternative splicing events can be provided, the method comprising the following steps: a) collecting DNA sequences of AS events obtained from healthy samples and tumor samples, and assigning to each sequence, i.e., a cancerous or healthy label; b) training a derivable parametric model to classify AS events based on whether the sequence appears in normal tissue or cancer tissue, using a deep learning-based approach; c) using a saliency approach to assign attribution values, i.e., positive values, to nucleotides that positively contribute to the cancerous class, and vice versa, to interpret the predictions of the selected model; d) selecting high-attribution regions where all nucleotides within the region have an attribution score greater than a threshold; e) applying a hypergeometric test to retain only motifs significantly enriched in cancer-specific AS sequences; f) merging similar motifs and filtering out low-frequency motifs; g) the output is a list of cancer-enriched motifs that can be used for cancer diagnosis to identify cancer neoantigens.
[0044] According to one aspect of the present invention, a computer-readable medium storing computer-executable instructions for implementing the method of any of the above aspects of the present invention can be provided.
[0045] According to one aspect of the present invention, a system can be provided, the system comprising at least one processor communicating with at least one memory device, and the at least one memory device stores instructions for causing the at least one processor to execute the method according to any of the above aspects.
[0046] Next, embodiments will be described in detail by way of example only with reference to the accompanying drawings.
Brief Description of the Drawings
[0047]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Mode for Carrying Out the Invention
[0048] In the past decade, the use of sequencing technology has facilitated the analysis of tumor genomes and opened the way to a deeper understanding of tumor evolution. However, identifying neoantigens for targeted cancer immunotherapy remains an ongoing challenge in cancer genomics.
[0049] The stated objectives of the methods and systems provided herein are to identify unique cancer-rich motifs from tumor-specific splicing mutations, which can provide valuable sources for cancer diagnosis and function as potential sources of neoantigens. Methods and systems are presented for identifying a set of DNA regions of overexpressed motifs from dysregulation of alternative splicing in the tumor transcriptome.
[0050] Alternative splicing (AS) plays an important role in cancer development and progression. Recent studies have shown that there are approximately 30% more alternative splicing events in tumors compared to normal tissues. Therefore, splicing mutations and the resulting overexpressed motifs provide a valuable source for cancer diagnosis and treatment. Neoantigens derived from tumor splicing mutations can be utilized in the development of novel immunotherapies such as personalized cancer vaccines.
[0051] Throughout this book, the terms neoantigen and neoepitope may be used interchangeably to refer to changes in the amino acid sequence of proteins within tumor cells. Neoantigens or neoepitopes can be recognized by the immune system and can trigger an immune response against cancer.
[0052] Figure 1 shows a high-level schematic of the steps of the proposed pipeline. In the first step, DNA sequences are prepared. Next, a machine learning model is trained on the prepared sequences before a motif discovery algorithm is employed. The motif discovery algorithm aims to interpret the trained model to identify regions of the input DNA sequences that contribute to the classification of healthy and cancerous. These regions are then considered cancer-enriched motifs as they occur in DNA sequences associated with alternative splicing events related to cancerous tissue.
[0053] First, the steps of data collection and model training are performed. In step 101, sequencing data of healthy samples and tumor samples are collected. Based on the evidence of sequencing reads, the workflow prepares a list of DNA sequences of various AS events (e.g., exon skipping, alternative acceptor sites, alternative donor sites, intron retention, etc.) and assigns a ground truth label, i.e., cancerous or healthy, to each sequence. An appropriate table showing the list of multiple sequences and their respective labels is shown in Figure 1.
[0054] Next, a derivable parametric model is trained to classify AS events based on whether the AS event appears in normal tissue or cancer tissue. Figure 1 shows a model as a convolutional neural network with a specific configuration. The details of the model configuration and training method will be described below. However, it is important to understand that, in contrast to existing approaches that rely on enumerating all novel splicing in tumors by a one-to-one comparison with all splicing events that appear in normal samples, the model is trained to classify RNA splicing in cancerous and normal tissues. The model is suitable for learning how each input DNA sequence is related to its respective label and needs to be interpretable later.
[0055] As shown, in the last step, motifs are discovered from the trained model by interpreting the predictions of the selected model. The motif discovery algorithm uses a saliency approach to assign attribution values to each input nucleotide of cancer-specific AS events based on the model predictions.
[0056] In the example, for each cancerous AS event, the workflow looks for a DNA region of length ≧M where each nucleotide in the region has an attribution value higher than a predefined threshold. Here, these DNA regions are called sequence motifs.
[0057] Next, the workflow applies a hypergeometric test to retain only motifs significantly enriched in cancer-specific AS sequences, merges similar motifs using an algorithm efficient for pairwise alignment of sequences, and filters out low-frequency motifs.
[0058] The final motif list can be called a cancer-enriched motif list. This list can be utilized for further downstream analysis to identify cancer neoantigens, as will be described below.
[0059] Figure 2 breaks down the three steps of Figure 1 into more detailed technical blocks and shows the schematic workflow of the proposed system's pipeline in more detail.
[0060] In step 201, a set of alternative splicing event arrays is collected. The input to the system includes DNA sequences of alternative events obtained from healthy samples and tumor samples. Based on the evidence of sequencing reads, various AS events are enumerated, such as exon skipping, alternative acceptor sites, alternative donor sites, intron retention, etc. To represent the exon sequences of AS events, a list X of nucleotide sequences is created. For each sequence x ∈ X, a label or value y ∈ Y is assigned. In the case of a binary classification problem, the label is assigned based on whether an event appears in a healthy sample or in a tumor sample, i.e., y ∈ {healthy, cancerous}. In the case of a regression problem, for example, to predict the exon inclusion ratio (or equivalent splicing rate) in healthy and cancerous AS events, the system can support a continuous labeling scheme y ∈ [0, 1]. The exon inclusion ratio can be quantified from sequencing reads.
[0061] Alternative splicing occurs at the RNA level, but preferably, the input to the system is the corresponding DNA sequence obtained by substituting uracil (U) nucleotides in the RNA with thymine (T) nucleotides.
[0062] The input array can already be retrieved from a preprocessed data store or can be processed by a workflow. For example, the array can be retrieved from the store and labels can be assigned based on the array read data, or the array may already be stored associated with appropriate labels. An example of a suitable dataset available for this concept is ExonSkipDB, which is a resource for the cancer and drug research communities to identify exon skipping events that could be targets of treatment and contains exon skipping events from 14,272 genes based on evidence from RNA-seq and whole exome sequencing. Additional resources such as The Cancer Genome Atlas (TCGA) and Genotype-Tissue Expression (GTEx) may also be used.
[0063] In step 202, a derivable parametric model is trained. In this step, the workflow uses the data provided in the previous step, i.e., the input array and labels shown in FIG. 1, to train a derivable parametric model. The derivable parametric model models the probability P θ (y|x), where θ represents the model parameters. The model can be, but is not limited to, a neural network. Training is performed using, for example, stochastic gradient descent to minimize a predefined error metric between the predicted values and the ground truth values. For example, training stops when user-defined criteria are met, such as when the validation recall exceeds a specific threshold after a predefined number of iterations (or epochs).
[0064] In the present disclosure, a deep neural network model, hereinafter referred to as DNA-Inception, is provided. DNA-Inception is merely a label used to refer to a specific neural network configuration proposed according to an exemplary implementation of the present disclosure. The system is not limited to using DNA-Inception to perform classification or regression tasks, but the use of any other derivable parametric function is contemplated. However, the specific configuration of DNA-Inception in the initial benchmark analysis has been shown to achieve good performance compared to two other state-of-the-art approaches. Further details regarding the preferred neural network configuration and the contemplated alternative machine learning models will be described hereinafter in the context of FIGS. 3 and 4.
[0065] Once the model is trained, as described above in the context of FIG. 1, the next step is to interpret the trained model to discover cancer-rich motifs, which is described as a model explanation method and shown as steps 203 to 207.
[0066] To interpret the predictions of the selected model, a saliency approach such as integrated gradients is proposed, and a feature attribution score a i ∈A is assigned to each nucleotide in the input sequence of the cancer-specific alternative splicing event, i.e., a i →x i ∈x, where i ∈ {1, 2,..., L}, x ∈ X, L = len(x) and y = cancerous. Positive values are assigned to nucleotides that positively contribute to the cancerous class, and vice versa.
[0067] In this exemplary implementation, integrated gradients are used, but other model explanation methods that explain the relationship between input and output based on what the model has learned, such as guided backpropagation, occlusion maps, etc., can also be utilized.
[0068] These and other suitable methods of interpretation will be understood by those of ordinary skill in the art, but the important thing is that the method of interpretation can identify the characteristics of the input that lead to predictions from the trained network model. In other words, which portions of the DNA sequence are associated with the cancerous alternative splicing event, i.e., which features were learned to accurately predict the class from the DNA sequence.
[0069] Once the model is interpreted and it is determined how the input positively contributes to the classification, the workflow uses that interpretation to select the high attribution regions (step 204).
[0070] For each input sequence x ∈ X, for y = cancerous, the system selects a DNA region of length ≧ M where all nucleotides in the region have an attribution score greater than the threshold t. For example, the threshold t can be defined based on the minimum positive attribution score or the average attribution score of the input sequence.
[0071] Other mechanisms for grouping features are also contemplated. For example, the scores may be compared to the scores of adjacent features, to a metric of the entire sequence, to a metric of the region, or to a metric calculated for the region.
[0072] Similarly, while here it has been described that the features utilized may be the individual nucleotides of the sequence, it will be understood that the model can also be trained using additional features such as groups of nucleotides or adjacent nucleotide information. The important thing is that the model learns to classify cancerous or healthy related DNA sequences based on the features of the sequence, and that those features are interpretable in a way that leads to the discovery of motifs.
[0073] A series of optional filtering steps follows to ensure that the list of high attribution regions is statistically relevant and likely to be associated with alternative splicing events corresponding to cancerous tumors. It will be understood that these inference and correction steps may be performed in a different order, or replaced or supplemented by additional steps that are likely to be essentially statistical.
[0074] Preferably, in step 205, a hypergeometric test is applied for the application of statistical inference and test correction.
[0075] In the hypergeometric test, it is assumed that the number of positive (i.e., cancerous) sequences containing motifs rich in cancer follows a hypergeometric distribution X~Hypergeometric(N,K,n). The probability mass function (pmf) of the random variable X following the hypergeometric distribution is as follows.
[0076]
Equation
[0077] Where N is the total number of sequences (i.e., the size of the population), K is the total number of cancerous sequences, n is the number of sequences having a specific motif, and k is the number of cancerous sequences containing that specific motif.
[0078] The hypergeometric test is performed to obtain motifs significantly enriched in cancer, and motifs common to both healthy and cancerous alternative splicing event sequences are excluded.
[0079] For each motif, the hypergeometric test first calculates the probability P(X≥k) of extracting the array containing that motif without replacement under the null hypothesis, that is, such a motif may appear equally in healthy AS events and cancerous AS events. For this purpose, the survival function sf is utilized, that is, the reciprocal of the cumulative distribution function cdf, that is, sf = 1 - cdf = 1 - P(X < k). After adjusting for multiple testing, if the p-value (i.e., the output probability) is low enough, the null hypothesis is likely not to hold and is rejected. That motif is considered to be significantly enriched in cancer events and is added to the motif output list.
[0080] For the correction of multiple testing, various methods can be used, such as the Bonferroni correction, the false discovery rate of Benjamini and Hochberg, etc.
[0081] When a preliminary motif output list corresponding to the high attribution region is obtained, in step 206, the workflow optionally merges similar motifs. The merging of similar motifs can be performed using, for example, the implementation of the PairwiseAligner() class in the biopython library. That is, motifs can be merged using appropriate pairwise sequence alignment techniques. Sequence alignment is a process of arranging two or more sequences (of DNA, RNA, or protein sequences) in a specific order to identify the similar regions between them. As is well understood by those skilled in the art, pairwise sequence alignment compares two sequences at a time and provides the best possible sequence alignment. Pairwise is easy to understand, and it is very easy to infer from the obtained sequence alignment.
[0082] For each query motif, if it can be aligned to an existing motif, the alignment with the maximum score where internal gaps are prohibited to maintain motif consistency is selected. If a motif cannot be aligned in this approach, that motif is added to the output list. Other alignment techniques or implementations can be utilized to execute the task.
[0083] Then, in step 207, the list of merged motifs is filtered to remove low-frequency motifs. The system filters motifs that occur in the cancer sequences below a user-defined number, i.e., less than min_occurrence. In this way, it is possible to be more confident that each motif is statistically significant and representative of the DNA sequences corresponding to cancerous alternative splicing events.
[0084] In step 208, the output of the system is a list of cancer-enriched motifs. In a further exemplary implementation, the output may also include the genomic coordinates of the motifs and a list of the genes and transcripts in which those motifs occur.
[0085] This final output motif list can be used for cancer diagnosis. Specifically, when obtaining a new RNA-seq sample, all alternative splicing event sequences are listed and the trained model is run to predict, for example, whether each event belongs to cancer tissue or healthy tissue. If a nucleotide sequence likely to appear in a cancer sample is found, the sample is diagnosed as cancerous. Additionally, the motif discovery algorithm can be run to list all significant cancer-enriched motifs and compare new motifs with previously discovered motifs to obtain neoepitopes from all previously validated alternative splicing for which immunogenicity has been verified. If no nucleotide sequences likely to appear in cancer are found, a list of candidate neoepitopes or neoantigens from alternative splicing can be obtained by extracting unique k-mer peptides (where k≧9) that cover cancer-enriched motifs and excluding peptides present in proteomics datasets other than cancer. Next, protein mass spectrometry (MS) databases obtained from various tumor types (such as the MS data of the Clinical Proteomic Tumor Analysis Consortium (CPTAC)) need to be used to confirm potential neoepitopes or neoantigens. The immunogenicity of the resulting neoepitopes or neoantigens can be predicted using existing tools to extract candidate neoepitopes or neoantigens that can be utilized in cancer immunotherapy. An exemplary workflow is shown in Figure 6.
[0086] Motifs can be used not only for cancer diagnosis but also to identify neoantigens based on cancer-enriched motifs.
[0087] Splicing mutations in cancer and the resulting motifs are a rich source of neoepitope or neoantigen candidates. AS-based neoantigens can be utilized in the development of novel immunotherapies such as personalized cancer vaccines. To utilize cancer-rich motifs for therapeutic purposes, additional validation experiments may be used to check the immunogenicity of the peptides resulting from covering those motifs. The goal is to elicit a T cell response in patient samples when pulsed with peptides covering the motifs.
[0088] Obtaining a list of candidates for neoepitopes or neoantigens derived from alternative splicing involves the steps of extracting k-mer peptides (where k≥9) that cover each of the cancer-rich motifs, excluding peptides present in proteomics datasets other than cancer, identifying potential neoepitopes or neoantigens in protein mass spectrometry (MS) databases obtained from various tumor types (such as MS data of the Clinical Proteomic Tumor Analysis Consortium (CPTAC)), and running existing tools to predict whether the neoepitopes or neoantigens bind to the major histocompatibility complex (MHC) and check their immunogenicity. Neoepitopes or neoantigens that trigger an immune response can be utilized in cancer immunotherapy. These verification steps are well-described in the literature (e.g., Kahles et al., "Comprehensive analysis of alternative splicing across tumors from 8,705 patients", Cancer cell 34.2 (2018): 211-224). Additionally, peptides derived from newly identified alternative splicing events, which have not been detected previously and thus do not exist in standard MS reference databases, can be utilized when identifying candidates for novel neoepitopes or neoantigens in MS data from eluted peptide-MHC complexes. Finally, candidate neoepitopes (or neoantigens) are utilized in cancer immunotherapy vaccines to help the immune system recognize such antigens or epitopes and eliminate the cancer cells that produce them.
[0089] Samples of the immunopeptidome and peripheral blood mononuclear cells (PBMCs) are available and can be used to complete this step. Such verification experiments are well-known to those skilled in the art.
[0090] Furthermore, the algorithms and workflows can be modified to identify significant motifs from novel splicing mutations in genetic diseases such as autism spectrum disorder, which may facilitate the development of potential therapies.
[0091] Figure 3 shows a model specially developed to address specific classification tasks required in the proposed workflow. As described above, here we refer to the specially configured and trained model as DNA-Inception.
[0092] DNA-Inception is an example of a derivable parametric model that can be utilized in the workflow. This model was developed to prove the effectiveness of the proposed system.
[0093] The DNA-Inception model is a multi-kernel size convolutional network for 1-dimensional (1D) signals, inspired by Inception v1. Inception was developed to solve pattern recognition problems in computer vision and is described in Szegedy, Christian et al., "Going deeper with convolutions", Proceedings of the IEEE conference on computer vision and pattern recognition, 2015. Inception enables interpretability while providing high performance.
[0094] DNA-Inception, i.e., the exemplary implementation proposed herein, is a matrix
[0095]
Number
[0096] including an embedding layer represented by, where n is the vocabulary size and d is the customized embedding dimension. Considering the input array x of nucleotides 301 of length L, the array is first encoded using a tokenizer encoding method. Tokenization replaces the elements within the array with tokens to facilitate analysis. As understood, any appropriate encoding technique can be used to facilitate the calculation of vectors or matrices.
[0097] Next, the embedding layer 302 performs matrix multiplication xW, mapping each nucleotide into
[0098]
Number
[0099] space, and the shape of the output matrix is xd. The embedding layer is followed by n 1D convolutional blocks 303, where n ∈ {1,..., N}. Each convolutional block consists of (i) multiple 1D convolutional layers 304 with a kernel size equal to 1, which are added to improve computational efficiency and reduce the number of parameters, and then (ii) another set of 1D convolutions 305 with variable kernel sizes to capture the dynamics at different scales of the input signal, and then (iii) a 1D max-pooling layer 306. All output tensors from the final convolutional block are concatenated at 307 and passed through a series of fully connected layers 308.
[0100] The proposed deep learning-based approach is used to classify RNA splicing in cancerous and normal tissues, in contrast to existing approaches that rely on enumerating all novel splicing in tumors by a one-to-one comparison with all splicing events that occur in normal samples. Using a deep learning-based model improves the specificity in identifying common cancer-specific splicing patterns in pan-cancer analysis.
[0101] As described above, when the trained model can be interpreted to identify features that contribute positively (i.e., the contribution of each nucleotide to the classification), multiple appropriate machine learning techniques can be used.
[0102] To demonstrate this, multiple comparisons were made between the constructed machine learning models specifically disclosed herein, namely DNA-Inception, the above-described pattern recognition CNN, and other models.
[0103] First, DNA-Inception, the model developed in the present invention, is compared with a baseline model developed using state-of-the-art DNABERT and Bi-LSTM. DNABERT is described in Ji, Yanrong, et al., "DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome", Bioinformatics 37.15 (2021): 2112-2120. Next, the proposed system is compared with existing methods and pipelines for identifying alternative splicing-based cancer neoantigens. Third, the system is compared with existing methods for identifying cancer motifs.
[0104] DNABERT is a Bidirectional Encoder Representations from Transformers (BERT) model pre-trained for the DNA language within the genome. DNABERT is a general-purpose model that can be trained and fine-tuned for any classification task given a DNA sequence as input. The authors provided a pre-trained model that utilized the entire human genome in the pre-training step. The authors stated that the pre-training of DNABERT took approximately 25 days using eight NVIDIA 2080Ti GPUs. In an initial benchmark analysis using the ExonSkipDB dataset, DNABERT-XL, which was developed to process DNA sequences longer than 512 nucleotides, was trained and evaluated. DNABERT uses k-mer sequences as input. The authors reported that the highest performance was achieved when k = 6, and thus this value was used for k in the benchmark analysis.
[0105] Table 1 and Figure 4 below show the performance evaluation over 10 runs, where in each run, the dataset is randomly split into 80% for training and validation and 20% for testing. Figure 4 shows the violin plots of the performance of DNABERT-XL, Bi-LSTM, and DNA-Inception over 10 runs. The three methods were benchmarked to classify cancerous and healthy AS sequences from the ExonSkipDB dataset.
[0106] DNABERT-XL consistently showed lower average performance values and higher variability compared to DNA-Inception. Additionally, DNABERT-XL requires 24GB of GPU RAM during training. DNA-Inception outperformed DNABERT-XL in terms of performance, model complexity in terms of the total number of trainable parameters, training time, and GPU requirements.
[0107] The Bi-LSTM is a baseline model consisting of an embedding layer, a bidirectional long short-term memory network, and then a final fully connected layer. Bi-LSTM is a common approach for handling genomic sequences. In this baseline approach, k-mer sequences with k = 6 were used as input. During the initial benchmarking, it was confirmed that using k-mer sequences instead of a character-level approach significantly improved performance. Empirically, setting k = 6 yielded the best results compared to k = 3, 4, or 5. Again, the disclosed DNA-Inception outperformed Bi-LSTM in the initial benchmark analysis.
[0108] Finally, when benchmarking the three models, the same evaluation criteria and stopping metrics were set.
[0109] Table 1 below shows the evaluation of the three models in terms of training time, complexity from the perspective of the total number of trainable parameters, and GPU models. The average and standard deviation of the training time over 10 different runs are shown.
[0110] [Table 1]
[0111] According to the proposed exemplary implementation, a deep learning system for identifying cancer-enriched sequence motifs from novel splicing mutations in tumors is described. For this purpose, a neural network model based on one-dimensional (1D) convolution was developed. This model classifies cancerous AS events and normal AS events. Then, the algorithm highlights important input nucleotides in cancerous events based on model predictions using the integrated gradient saliency approach. Subsequently, the algorithm applies a hypergeometric test and additional merging and filtering techniques to identify a set of DNA regions of motifs that are significantly represented in cancer-specific AS events. The set of cancer-enriched motifs can be utilized in further downstream analysis to obtain AS-based neoantigen candidates.
[0112] Furthermore, since the model is trained once on a large number of samples and can distinguish normal AS events from cancer-specific AS events with F1 score and recall score exceeding 90%, this approach can be generalized to new unknown AS events and does not require repeated manual comparison between tumor samples and healthy samples.
[0113] The proposed system identifies cancer-enriched motifs from splicing mutations in tumors and is not limited to so-called mutation signatures. Tumor mutations can inhibit AS and lead to novel splicing mutations. Therefore, DNA mutations present in tumor samples are considered as inputs to the system to obtain a comprehensive understanding of splicing mutations and the accompanying cancer-enriched motifs. Thus, the system provides a more comprehensive perspective for extracting cancer-enriched motifs. Motifs based on tumor splicing mutations have not been investigated so far.
[0114] The dataset providing AS annotation for tumor samples and healthy samples has been collected based on various sequencing technologies, such as whole-genome sequencing and whole-exome sequencing. This system may become more powerful when applied to overlapping reads in both the coding and non-coding regions of DNA, i.e., the AS dataset collected based on whole-genome sequencing. When applied only to the coding exome, i.e., the AS dataset based on whole-exome sequencing, this system may not be as powerful because there is a possibility of missing cancer-rich motifs from non-coding regions (i.e., introns). Nevertheless, this system still shows usefulness in this regard.
[0115] Various datasets have been generated using different sequencing depths, which may introduce unwanted biases in the evidence of AS reads. This can be avoided by ensuring that the initial dataset meets specific quality criteria. Otherwise, variations in the results are expected due to differences in the execution, depth, and quality of sequencing.
[0116] For the sake of completeness, the present disclosure provides a computer-implemented method for identifying DNA sequence motifs indicative of a disease from alternative splicing events. FIG. 5 shows the steps of an exemplary implementation in the form of a flowchart. The method includes, at step 501, obtaining a list of nucleotide DNA sequences, each sequence representing the sequence of an alternative splicing event. At step 502, the method obtains a label for each sequence, the label indicating whether the sequence is associated with a healthy sample or a diseased sample. Next, at step 503, the implementation trains a machine learning model with the list and respective labels to classify alternative splicing events from those DNA sequences based on whether the sequences occur in healthy samples or diseased samples. From the trained model, the implementation can then generate cancer-enriched motifs for use as neoantigens for immunotherapy and diagnosis. At step 504, the process identifies features of the sequences that positively contribute to indicating whether the sequences occur in healthy samples or diseased samples by interpreting the trained model. Next, at step 505, based on the identified features, one or more high attribution regions of one or more sequences are selected, the high attribution regions indicating regions of the sequences having features that positively contribute. At step 506, the process generates a DNA sequence motif based on the one or more high attribution regions. Optionally, at step 507, the motif is used to identify neoantigen candidates or neoepitope candidates for immunotherapy based on the DNA sequence motif, or at step 508, the motif is compared to a DNA sequence to predict the presence of cancer in a tissue based on alternative splicing events.
[0117] The above described how DNA sequence motifs can be utilized for the formation of neoantigens for vaccine development or as a diagnosis.
[0118] Using various bioinformatics approaches, the DNA sequences of motifs can be converted into corresponding peptide sequences, enabling the prediction of neoantigens or neoepitopes. Machine learning-based software solutions are available for in-silico prediction and identification of optimal immunogenic neoantigens or neoepitopes for personalized cancer immunotherapy. For example, neoantigen prediction systems (such as NeoAntigen Quest (NAQ) and NEC Immune Profiler (NIP)) combine transcriptomics and proteomics data to identify immunogenic neoantigen candidates or neoepitope candidates from DNA sequence information, thereby determining whether the neoantigen or neoepitope is naturally processed and presented on the surface of tumor cells and, thus, obtaining the ability to predict with high accuracy the immunogenic potential (or "immunogenicity") of said neoantigen candidates or neoepitope candidates.
[0119] Obtaining a list of candidates for neoepitopes or neoantigens derived from alternative splicing involves extracting k-mer peptides (where k ≥ 9) that cover each of the cancer-rich motifs, excluding peptides present in proteomic datasets other than cancer, identifying potential neoepitopes or neoantigens in protein mass spectrometry (MS) databases obtained from various tumor types (such as MS data of the Clinical Proteomic Tumor Analysis Consortium (CPTAC)), and running existing tools to predict whether the neoepitopes or neoantigens bind to the major histocompatibility complex (MHC) and check their immunogenicity. Neoepitopes or neoantigens that trigger an immune response can be utilized in cancer immunotherapy. These validation steps are well described in the literature (e.g., Kahles et al., "Comprehensive analysis of alternative splicing across tumors from 8,705 patients", Cancer cell 34.2 (2018): 211 - 224). Additionally, peptides derived from newly identified alternative splicing events, which have not been detected previously and thus do not exist in standard MS reference databases, can be utilized in identifying candidates for novel neoepitopes or neoantigens in MS data from eluted peptide-MHC complexes. Finally, the candidate neoepitopes (or neoantigens) are utilized in vaccines for cancer immunotherapy, helping the immune system to recognize such antigens or epitopes and eliminate the cancer cells that produce them.
[0120] FIG. 6 shows, in the form of a flowchart including steps 601 to 605, the plurality of steps by which a list of candidates for neoepitopes or neoantigens derived from alternative splicing is obtained. The flowchart shows the workflow of neoantigens derived from alternative splicing. The process starts with tumor-specific alternative splicing (step 601). In step 602, cancer-enriched motifs generated as a result of tumor-specific alternative splicing events are identified using the methods for identifying DNA sequence motifs described herein. In step 603, k-mer peptides (where k≥9) covering the cancer-enriched motifs are extracted. Peptides present in proteomics datasets other than cancer are excluded to obtain unique k-mer peptides derived from alternative splicing. In step 604, potential neoepitopes or neoantigens are identified using protein mass spectrometry (MS) databases obtained from various tumor types (such as MS data of the Clinical Proteomic Tumor Analysis Consortium (CPTAC)). Finally, step 605 involves predicting the immunogenicity of the resulting neoepitopes or neoantigens. Neoepitopes or neoantigens that bind to MHC are predicted to trigger an immune response and can thus be utilized in cancer immunotherapy.
[0121] Neoantigen candidates or neoepitope candidates identified by such quantitative statistical analysis may represent effective vaccine targets that have the potential to elicit a broad T cell immune response and can be used in vaccine design and production. Once the sequence of a neoantigen or neoepitope is obtained (e.g., by mass spectrometry), the peptide can be synthesized, i.e., by in vitro solid-phase or liquid-phase peptide synthesis methods, and purified, i.e., by classical separation-based methods. These methods are well known to those skilled in the art.
[0122] As used herein, the term "neoepitope" refers to any part of a neoantigen that is recognized by any antibody, B cell, or T cell. A "neoantigen" refers to a molecule that can be bound by an antibody, B cell, or T cell and can be composed of one or more neoepitopes. Thus, the terms "neoepitope" and "neoantigen" may be used interchangeably herein.
[0123] Such an approach is validated against clinically relevant neoantigens or neoepitopes, checks the immunogenicity of the resulting peptides, and ensures that the selected neoantigen or neoepitope successfully elicits a T cell response in patient samples. This can be performed using an immunogenicity assay (e.g., a quantitative enzyme-linked immunosorbent assay (ELISA)) pulsed with peptides covering the motif in patient samples.
[0124] Alternatively, additional immunopeptidome analysis can be used to identify which peptides are at least presented on the tumor cell surface from alternative splicing events. Neoantigens or neoepitopes are presented on the tumor cell surface by class I and class II major histocompatibility complex (MHC) molecules. Peptides associated with and presented by MHC molecules are collectively referred to as the immunopeptidome. The immunopeptidome can be extracted from cell or tissue samples using various separation techniques known in the art, followed by immunoaffinity purification of MHC molecules and release of peptides from the separated MHC molecules. The resulting purified peptides can then be analyzed by mass spectrometry. Peptides from mass spectrometry can be mapped to the healthy immunopeptidome and to novel splicing mutations in cancer.
[0125] The term "vaccine" relates to a biological preparation that provides active acquired immunity against a specific disease, in this case cancer or a tumor. Usually, a vaccine contains a substance similar to a neoantigen or neoepitope from the surface of cancer or tumor cells, i.e., an "external" substance. Such an external substance is recognized by the immune system of the vaccinated person, whereby in turn the substance is destroyed, "memory" against the cancer or tumor is formed, and a permanent level of protection against future diseases caused by recurrence of the same cancer or tumor is induced. Through the vaccination route involving the vaccine produced by the method of the present invention, when a vaccinated subject is confronted again with the same cancer or tumor as the one vaccinated against, it is envisaged that the individual's immune system will recognize the cancer or tumor and elicit a more effective immune defense.
[0126] The induced active acquired immunity can be humoral and / or cellular. Humoral immunity refers to a reaction involving B cells that produce antibodies that specifically bind to a neoantigen or neoepitope, or any future neoantigen or neoepitope corresponding to those in the administered vaccine. B cells each express a unique B cell receptor (BCR) and recognize neoantigens or neoepitopes in their native form. Through this recognition and further interaction with other cells of the immune system, activated B cells can differentiate into plasma cells specialized to secrete antibodies against the encountered neoantigen or neoepitope. The term "antibody" refers to immunoglobulins (Ig) used by the immune system to specifically identify and neutralize external antigens. A subset of these B cell-derived plasma cells become long-lived antigen-specific memory B cells, as is well understood by those skilled in the art.
[0127] On the one hand, cellular immunity can be divided into two different aspects. The first involves helper T cells, or CD4+ T cells, which produce cytokines and regulate the activities of other immune cells in the immune response. The second involves killer T cells, also known as cytotoxic T lymphocytes (CTLs) or CD8+ T cells, which are cells that can recognize neoantigens or neoepitopes presented on the surface of cancer or tumor cells and eliminate cancer or tumor cells. In contrast to B cells, T cells recognize only neoantigens or neoepitopes that have been processed into peptides, loaded onto MHC molecules, and presented on the cell surface. CD4+ T cells interact with MHC class-II molecules and play roles in regulating the immune response, recognizing foreign antigens (i.e., neoantigens or neoepitopes), activating various parts of the immune system, and activating B cells and CD8+ T cells. CD8+ T cells interact with MHC class I receptors on the surface of antigen-presenting cells (APCs) and target cells that display antigen peptide fragments generated by proteasomal degradation. As will be understood by those skilled in the art, after an immune response, subsets of both CD8+ T cells and CD4+ T cells can contribute to acquired adaptive immunity and persist as memory T cells that enable a more rapid and robust response to any future recognition of the same neoantigen or neoepitope.
[0128] The vaccine produced by the method of the present invention may be an epitope-based (i.e., neoepitope-based) vaccine, in other words, it is assumed to be composed of one or more epitopes. Epitope-based vaccines (EVs) utilize short antigen-derived peptides corresponding to immune epitopes and are administered to induce protective humoral and / or cellular immune responses. EVs may be able to precisely control the activation of the immune response by focusing on the most relevant (i.e., immunogenic and conserved) antigen regions. Since experimental screening of large sets of peptides is time-consuming and costly, in silico methods that facilitate T cell epitope mapping of protein antigens are extremely important in EV development. Prediction of T cell epitopes focuses on the presentation of peptides on the surface of cancer or tumor cells by proteins encoded by MHC.
[0129] The neoantigen or neoepitope of the present invention may interact with MHC class-I and / or MHC class-II molecules to induce responses of CD8+ T cells and / or CD4+ T cells, respectively. There may be at least one neoantigen or neoepitope that interacts with MHC class-I and at least one neoantigen or neoepitope that interacts with MHC class-II.
[0130] Any or all steps of the method can be performed on local, remote, or cloud computing devices. The trained model and its associated parameters may be stored centrally, i.e., on the cloud. The methods and processes described herein can be embodied as code (e.g., software code) and / or data. Models, methods, and algorithms can be implemented in hardware or software, as is well known in the field of machine learning. For example, hardware acceleration using a specially programmed graphics processing unit (GPU) or a specially designed field programmable gate array (FPGA) may provide certain efficiencies. For completeness, such code and data can be stored on one or more computer-readable media that may include any device or medium capable of storing code and / or data for use by a computer system. When the computer system reads and executes the code and / or data stored on the computer-readable medium, it performs the methods and processes embodied as data structures and code stored within the computer-readable storage medium. In certain embodiments, one or more steps of the methods and processes described herein can be performed by a processor (e.g., the processor of a computer system or data storage system).
[0131] In general, any of the functions described in this document or shown in the figures can be implemented using software, firmware (e.g., fixed logic circuitry), programmable or non-programmable hardware, or a combination of these implementation forms. The terms "component" or "function" as used herein generally represent software, firmware, hardware, or a combination thereof. For example, in the case of a software implementation, the terms "component" or "function" may refer to program code that executes a specified task when executed on a processing device(s). The separation of the shown components and functions into distinct units may reflect the actual or conceptual physical grouping and allocation of such software and / or hardware with the tasks.
Description of Symbols
[0132] 301 Nucleotide 302 Embedded layer 303 1D Convolution Block 305 1D Convolution 306 1D Max Pooling Layer 308 Fully Connected Layer
Claims
1. A computer-implemented method for training a machine learning model for use in identifying an array motif indicative of a disease from alternative splicing events, comprising: obtaining a list of nucleotide sequences, each sequence representing an array of alternative splicing events; obtaining a label for each sequence, the label indicating whether the sequence is associated with a healthy sample or a disease sample; training a machine learning model with the list and respective labels to classify alternative splicing events from those sequences based on whether the sequences occur in healthy samples or disease samples; and a method comprising the steps.
2. The method according to claim 1, wherein each label indicates whether the sequence is associated with a sample of healthy tissue or a sample of cancerous tissue, and the machine learning model is trained to classify alternative splicing events from those sequences based on whether the sequences occur in healthy tissue or cancerous tissue.
3. The method according to claim 1, wherein each label indicates whether the sequence is associated with a genetic disease, and the machine learning model is trained to classify alternative splicing events from those sequences based on whether the sequences are associated with a genetic disease.
4. The method according to any one of claims 1 to 3, wherein the machine learning model is a neural network, preferably a convolutional neural network with a multi-kernel size for one-dimensional signals.
5. The method further comprises: identifying features of the sequences that positively contribute to indicating whether the sequences occur in healthy samples or disease samples by interpreting the trained model; selecting one or more high attribution regions of one or more of the sequences based on the identified features, the high attribution regions indicating regions of the sequences having features that positively contribute; and outputting an array motif based on the one or more high attribution regions. The method according to any one of claims 1 to 4, further comprising the steps.
6. The method according to claim 5, wherein each feature is a nucleotide of said sequence and said sequence is a DNA sequence.
7. The step of selecting comprises generating a feature attribution score for each nucleotide in each sequence based on the contribution of that nucleotide to said classification; and selecting a region of said sequence based on said score for each nucleotide. The method according to claim 6, comprising the steps of
8. The step of selecting a region of said sequence based on said score for each nucleotide comprises comparing the score of each nucleotide in a candidate region to a threshold; calculating an average of the scores of the nucleotides in the candidate region; and calculating an average of the scores of said sequence and comparing the score of each nucleotide in the candidate region to said average of the scores of said DNA sequence. The method according to claim 7, comprising the steps of
9. The step of selecting comprises obtaining significant regions present in said sequence associated with a disease sample label and applying a hypergeometric test to said one or more high attribution regions to exclude regions common to said sequence associated with a disease sample label and said sequence associated with a healthy sample label. The method according to any one of claims 5 to 8, comprising the steps of
10. The method according to any one of claims 5 to 9, further comprising merging a plurality of said selected one or more regions using pairwise sequence alignment.
11. The method according to any one of claims 5 to 10, further comprising filtering said selected one or more regions based on the abundance of each region in said DNA sequence associated with a disease sample label.
12. The method according to any one of claims 5 to 11, further comprising identifying neoantigen candidates or neoepitope candidates for immunotherapy based on said sequence motif.
13. A computer-readable medium storing computer-executable instructions for performing the method according to any one of claims 1 to 12.
14. Selecting one or more predicted immunogenic candidate amino acid sequences covering the sequence motif output according to the method according to any one of claims 5 to 12 for inclusion in a vaccine. To produce a vaccine, the step of synthesizing the one or more amino acid sequences, or the step of encoding the one or more amino acid sequences into the corresponding DNA or RNA sequences, and / or the step of integrating the DNA or RNA sequences into the genome of a bacterial or viral delivery system, and A method for producing a vaccine, comprising:
15. A method for predicting the likelihood that a tissue sample is cancerous, comprising: Retrieving the sequence of the tissue sample, and Analyzing the sequence of the tissue sample for similarity to the sequence motif output according to the method according to any one of claims 5 to 12, and Predicting the likelihood that the tissue sample is cancerous based on the analysis, and A method, comprising:
Citation Information
Patent Citations
Cancer-related variable splicing database system of long non-coding RNA
CN111508563A
Cancer-specific molecules and methods of use thereof
JP2022526424A
Systems and methods for analysis of alternative splicing
WO2019226804A1
Neoantigens, methods and detection of use thereof
WO2022047242A2