Tumor tissue splicing recognition and neoantigen peptide screening method based on ONT sequencing platform

By identifying alternative splicing events in tumor tissue using the ONT sequencing platform and SUPPA2 software, and verifying them with MHC binding affinity and mass spectrometry, 9mer peptides are extracted. This solves the problem of unsystematic neoantigen screening in existing technologies, achieving highly accurate and reliable neoantigen peptide screening, which is suitable for tumor immunotherapy and personalized treatment.

CN121331233APending Publication Date: 2026-01-13SOUTHEAST UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511530553.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-24
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing technologies lack the ability to identify and predict the function of novel transcripts in tumor tissues based on the ONT sequencing platform. Furthermore, the prediction of peptide affinity with MHC in splice boundary regions, expression filtering, and mass spectrometry verification are not effectively combined, resulting in unsystematic and unreliable neoantigen screening.

Method used

RNA sequencing was performed using the ONT sequencing platform. Alternative splicing events were identified using SUPPA2 software. Neoantigen peptides were screened by combining TPM expression levels and MHC binding affinity, and mass spectrometry was used for verification. 9-mer peptides across splicing sites were extracted, achieving data integration and standardized screening throughout the entire process.

Benefits of technology

It significantly improves the accuracy of alternative splicing event identification and the biological relevance of neoantigen peptides, enhances the immune response potential and actual expression reliability of candidate neoantigens, and is applicable to tumor mechanism analysis, personalized treatment and vaccine design.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121331233A_ABST
    Figure CN121331233A_ABST
Patent Text Reader

Abstract

The invention discloses a tumor tissue splicing recognition and neoantigen peptide screening method based on an ONT sequencing platform, and belongs to the technical field of tumor immunology. Comprising the following steps: extracting RNA from a tumor tissue, separating mRNA, constructing a sequencing library, and sequencing through an ONT sequencing platform to obtain original data; performing basic group identification on the original data to obtain fastq format data; comparing the data to a reference genome reconstruction transcript, identifying a variable splicing event based on a comparison result, and performing difference analysis; screening out variable splicing events with significant differences, and extracting novol transcript coding sequences related to the events; extracting a 9mer peptide fragment crossing a splicing site in a splicing boundary region of the transcript; performing MHC binding affinity scoring on the peptide fragments, and screening out candidate new antigen peptide fragments; and carrying out mass spectrum verification on the candidate new antigen peptide fragment, and screening the actually expressed peptide fragment. The method can be widely applied to tumor immunomarker screening, new antigen vaccine design and individualized treatment strategy development.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of tumor immunology technology, specifically relating to a method for tumor tissue splicing recognition and neoantigen peptide screening based on the ONT sequencing platform. Background Technology

[0002] Alternative splicing (AS) is a crucial regulatory mechanism in the post-transcriptional modification of RNA in eukaryotic cells. It selectively splices exons or introns in various ways, generating a wide variety of mRNA transcripts and significantly expanding protein diversity. In tumor tissues, aberrant splicing events are closely associated with cell proliferation, apoptosis escape, and immune escape, and have been shown to serve as diagnostic markers and therapeutic targets. Furthermore, the novel isoform transcripts generated by splicing can produce "neoantigen peptides" not expressed in normal tissues, making them important candidates for tumor vaccines and immunotherapies.

[0003] Third-generation sequencing technology, especially the Oxford Nanopore Technologies (ONT) platform, provides the ability to directly sequence RNA molecules, obtaining raw RNA sequences and epigenetic modification information without reverse transcription or amplification. It is particularly suitable for detecting single-base levels of m6A modifications. Compared to traditional second-generation sequencing, which has short read lengths, difficulty in covering full-length transcripts, insufficient ability to identify complex splicing events, and incomplete reconstruction of exon boundary regions, ONT sequencing offers significant advantages such as long read lengths and the ability to capture full-length transcripts, making it ideal for full-length transcript reconstruction and the analysis of complex splicing events.

[0004] Although existing studies have used the ONT platform to identify splicing in the tumor transcriptome, the technical methods for systematically screening splice-driven neoantigen peptides remain incomplete. On the one hand, current methods mostly focus on the analysis of known annotated transcripts, lacking the identification and functional prediction of novel transcripts; on the other hand, few workflows can organically combine peptides in splice boundary regions with MHC affinity prediction, expression filtering, and mass spectrometry verification to form a standardized and reproducible candidate neoantigen screening system.

[0005] Therefore, there is an urgent need to develop an analytical method based on the ONT sequencing platform that can automatically identify differential splicing events, extract potential neoantigen peptides, and screen them by combining affinity and expression levels. This method can enhance our understanding of tumor transcriptional isoforms and provide precise target resources for tumor immunotherapy. Summary of the Invention

[0006] To address the shortcomings of existing technologies, the present invention aims to provide a method for tumor tissue splicing recognition and neoantigen peptide screening based on the ONT sequencing platform, thereby solving the problems in the prior art.

[0007] The objective of this invention can be achieved through the following technical solutions: A method for tumor tissue splicing recognition and neoantigen peptide screening based on the ONT sequencing platform includes the following steps: RNA was extracted from tumor tissue, and mRNA was isolated. Sequencing libraries were then constructed and sequenced using the ONT sequencing platform to obtain raw long read data. Base identification is performed on the raw long-read data to obtain fastq format data; FastQ format data were aligned to a reference genome to reconstruct transcripts, and alternative splicing events were identified and differential analyses were performed based on the alignment results. Significantly different alternative splicing events were screened out, and the coding sequences of novel transcripts associated with these events were extracted. The 9mer peptide across the splice boundary region of the novel transcript was extracted; The MHC binding affinity of the 9mer peptide was scored, and candidate neoantigen peptides were screened. Mass spectrometry was used to validate candidate neoantigen peptides and screen for peptides that were actually expressed.

[0008] Furthermore, the process of identifying alternative splicing events is as follows: using SUPPA2 software, combined with annotation files and TPM expression levels, the types of alternative splicing events present in each sample are identified, including: exon skipping, intron retention, variable 5' splicing sites, variable 3' splicing sites, and mutually exclusive exons.

[0009] Furthermore, the process of differential analysis of alternative splicing events is as follows: using the DiffSplice module of SUPPA2, the PSI values ​​of splicing events are compared between groups, ΔPSI is calculated and significance is tested, thereby identifying alternative splicing events with significant differences.

[0010] Furthermore, when screening for significantly different alternative splicing events, FDR < 0.05 and |ΔPSI| > 0.1 were used as significance screening thresholds.

[0011] Furthermore, the process of extracting the 9mer peptide is as follows: using a sliding window approach, extract a continuous 9-amino acid peptide spanning at least two different exon regions from the novel transcript corresponding to the alternative splicing event.

[0012] Furthermore, the criteria for screening candidate neoantigen peptides are: MHC binding affinity prediction value less than 500 nM, %Rank less than 2, and source transcript TPM greater than 1.

[0013] Furthermore, the mass spectrometry verification process involves comparing the mass spectrometry data of the candidate neoantigen peptide with those of the tumor tissue to confirm its expression at the protein level.

[0014] A tumor tissue splicing recognition and neoantigen peptide screening system based on the ONT sequencing platform includes: Sequencing module: RNA is extracted from tumor tissue and mRNA is isolated. Then, a sequencing library is constructed and sequenced using the ONT sequencing platform to obtain raw long read data. Base recognition module: performs base recognition on the raw long read data to obtain fastq format data; Event identification and analysis module: Aligns fastq format data to the reference genome to reconstruct transcripts, and based on the alignment results, identifies alternative splicing events and performs differential analysis; Transcript extraction module: Screens for significantly different alternative splicing events and extracts the coding sequences of novel transcripts related to these events; Peptide generation module: Extracts 9mer peptides across splice sites in the splice boundary region of the novel transcript; Peptide screening module: MHC binding affinity score is performed on 9mer peptides, and candidate neoantigen peptides are screened out; Peptide validation module: Performs mass spectrometry validation on candidate neoantigen peptides to screen for peptides that are actually expressed.

[0015] A computer storage medium storing a readable program that, when executed, instructs a computing device to perform the tumor tissue splicing recognition and neoantigen peptide screening method based on the ONT sequencing platform described above.

[0016] An electronic device, characterized in that it comprises: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, which causes the processor to perform operations corresponding to the tumor tissue splicing recognition and neoantigen peptide screening method based on the ONT sequencing platform described above.

[0017] The beneficial effects of this invention are: 1. This invention obtains full-length transcript information based on a third-generation sequencing platform, which significantly improves the accuracy of identifying alternative splicing events. It also proposes a complete workflow method from RNA extraction, transcript reconstruction, differential splicing identification, peptide extraction, immune prediction to validation and evaluation, realizing data integration and functional integration of key links, reducing errors caused by manual intervention and step-by-step operation, and contributing to the standardization and large-scale promotion of results.

[0018] 2. This invention uses splicing events as the source of antigen generation, proposing to extract 9mer peptides from the splicing boundary region of novel transcripts, and combining expression level (TPM) with MHC binding affinity (IC50, %Rank) for joint screening, thereby improving the biological relevance and immune response potential of candidate neoantigens from the source. Furthermore, a mass spectrometry validation strategy is introduced to directly verify the predicted peptides at the protein expression level, enhancing the reliability of the actual expression of candidate peptides. This invention is applicable to tumor mechanism analysis and epitranscriptome feature mining in basic research, and can also serve multiple directions such as target screening for personalized treatment, vaccine design, and research on immune escape mechanisms, possessing extremely high scientific research value and industrial transformation potential.

[0019] 3. This invention uses SUPPA2 software in conjunction with annotation files and TPM expression levels to identify various variable splicing events, such as exon skipping, intron retention, variable 5' splicing sites, variable 3' splicing sites, and mutually exclusive exons. It can achieve standardized and systematic identification of splicing events across the entire transcriptome. Compared with manual or low-throughput tools, it improves the accuracy and comprehensiveness of splicing event identification and provides high-quality input data for subsequent differential analysis.

[0020] 4. This invention compares the PSI values ​​of splicing events among different samples or groups using the DiffSplice module of SUPPA2, calculates ΔPSI and performs significance tests, which can quantitatively assess the differences in splicing patterns and exclude random noise or intra-sample variation, thereby improving the statistical reliability of screening for differentially variable splicing events and laying a solid foundation for the subsequent functional prediction of candidate events.

[0021] 5. When screening for significantly different alternative splicing events, this invention uses significance thresholds of FDR < 0.05 and |ΔPSI| > 0.1, which can effectively control the false positive rate and ensure that the differential events have sufficient biological effect size, so that the screened events are both statistically significant and practically meaningful, thereby improving the reliability and reproducibility of downstream analysis results.

[0022] 6. This invention extracts at least 9 consecutive amino acid peptides spanning two different exon regions from novel transcripts corresponding to alternative splicing events using a sliding window approach. This ensures that the obtained peptides truly cover the splicing boundaries and maintain the correct reading frame, making it more likely to capture neoantigens generated by tumor-specific splicing, thereby improving the specificity and effectiveness of subsequent immune prediction.

[0023] 7. This invention uses multiple criteria to screen candidate neoantigen peptides, including MHC binding affinity prediction value of less than 500 nM, %Rank of less than 2, and source transcript TPM of greater than 1. By considering two key indicators, affinity and expression level, peptides with weak affinity or no actual expression in tissues are effectively eliminated, significantly improving the immune response potential and biological relevance of candidate neoantigen peptides.

[0024] 8. This invention confirms the existence of candidate neoantigen peptides at the protein level by comparing them with tumor tissue mass spectrometry data, thereby achieving direct experimental verification of predicted peptides. This significantly improves the reliability and translational value of candidate peptides as targets for tumor immunotherapy and avoids false positives caused by relying solely on computational predictions. Attached Figure Description

[0025] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is the overall flowchart of the tumor tissue splicing event identification and neoantigen prediction screening of the present invention; Figure 2 This is a statistical chart showing the distribution of splicing event types in tumor samples provided in Example 2. Detailed Implementation

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] Example 1 like Figure 1 As shown, the method for tumor tissue splicing recognition and neoantigen peptide screening based on the ONT sequencing platform includes the following steps: S1. Total RNA was extracted from tumor tissue and mRNA was isolated. Then, a sequencing library was constructed and sequenced using a third-generation RNA (ONT) sequencing platform to obtain raw long read data. Total RNA was extracted from tumor tissue, mRNA was isolated using poly(A)+ RNA capture technology, a sequencing library was constructed, and sequencing was performed on a third-generation RNA sequencing platform to obtain raw long-read data. Sample preparation and RNA extraction: Tumor tissue and its paired adjacent normal tissue samples, flash-frozen in liquid nitrogen, were used as starting materials. Total RNA was extracted using the Novizan FastPure® Cell / Tissue Total RNA Isolation Kit V2, and poly(A)+ mRNA was enriched using VAHTS mRNA Capture Beads 2.0 magnetic beads to ensure the acquisition of high-purity, intact mRNA molecules, providing high-quality input for subsequent library construction and sequencing.

[0029] Direct RNA sequencing: The enriched mRNA is directly sequenced using the Oxford Nanopore platform to generate pod5 or fast5 format data containing the original signal, obtaining long read data covering the full-length transcript.

[0030] S2, perform base identification on the raw long read data to obtain fastq format data; Data quality control and base identification: Dorado (v0.8.1) software was used for base calling to convert the fast5 / pod5 format signal data into recognizable base sequences in fastq format. During this process, data were classified based on the average quality value of each read. Data with a Q value greater than or equal to 10 were defined as pass reads, considered clean reads for subsequent analysis, while data below this threshold were considered fail reads and excluded, thus improving the accuracy and reliability of the analysis results.

[0031] S3. FastQ format data is aligned to the reference genome to reconstruct transcripts, and alternative splicing events are identified and differential analyses are performed based on the alignment results. Specifically: Alignment analysis and expression quantification: Clean reads were aligned to a reference genome (e.g., GRCh38) using minimap2, and then the alignment results were statistically analyzed and filtered using samtools to obtain high-confidence alignment files (.bam format). Full-length transcripts were reconstructed using annotation information. Furthermore, appropriate expression quantification tools (e.g., featureCounts, salmon) were used to estimate the TPM expression level of each transcript.

[0032] The process of identifying alternative splicing events is as follows: using SUPPA2 software, combined with annotation files (GTF) and TPM expression levels, the types of alternative splicing events present in each sample are identified, including but not limited to: exon skipping (SE), intron retention (RI), variable 5' splice site (A5SS), variable 3' splice site (A3SS), and mutually exclusive exons (MXE).

[0033] The process of variable splicing event differential analysis is as follows: By using the DiffSplice module of SUPPA2, the percentage of spliced ​​in (PSI) values ​​of splicing events are compared between groups, ΔPSI is calculated and a significance test (p-value) is performed to identify significantly different alternative splicing events. S4. Screen for significantly different alternative splicing events and extract the coding sequences of novel transcripts related to the events; When screening for significantly different alternative splicing events, the significance screening thresholds are set as follows: False Discovery Rate (FDR) < 0.05 and |ΔPSI| > 0.1.

[0034] The novel transcript extraction process is as follows: the coding sequence (CDS) is extracted from novel transcripts that have undergone splicing variations (significantly different alternative splicing events). The 'novel transcript' refers to a transcript that has a new splicing structure, start site, or termination site compared with the transcripts recorded in the reference annotation database, and is labeled as 'i', 'j', 'k', 'm', 'n', 'u', or 'x' by a transcript alignment tool (such as Cuffcompare).

[0035] S5, extract the 9mer peptide across the splice site in the splice boundary region of the novel transcript; The 9mer peptide spanning the splice site refers to a continuous sequence of 9 amino acids containing a variable splice boundary position, spanning at least two different exon regions.

[0036] The peptide generation strategy is as follows: using a sliding window approach, a continuous 9-amino acid peptide (9mer peptide) spanning at least two different exon regions is extracted from the novel transcript corresponding to the alternative splicing event for subsequent neoantigen candidate screening. When extracting the 9mer peptide, its source region is limited to coherent coding segments within the same reading frame to avoid meaningless translation caused by frame shifts.

[0037] S6. Based on the prediction model, the MHC binding affinity of the 9mer peptide was predicted, and candidate neoantigen peptides were screened out. The screening criteria for candidate neoantigen peptides are: MHC binding affinity prediction value (IC50 value) less than 500 nM, %Rank (percentile ranking) less than 2, and source transcript TPM (number of transcripts per million) greater than 1; Specifically, the screening process for candidate neoantigen peptides involves using MHC binding prediction tools such as NetMHCpan to assess and predict the affinity of the aforementioned peptides. Peptides with high affinity (IC50 < 500 nM, %Rank < 2) and whose source transcript TPM expression level is > 1 are selected as candidate neoantigen peptides. Regarding HLA subtype selection, known HLA-I typing information from patient samples is prioritized, or common subtypes with high population frequency (such as HLA-A02:01, HLA-B07:02) are selected for prediction.

[0038] S7, perform mass spectrometry verification on candidate neoantigen peptides to screen for peptides that are actually expressed; Mass spectrometry verification is used to confirm whether candidate neoantigen peptides are actually translated in tumor tissues and presented via MHC-I. If so, it indicates that the peptide has been actually expressed and can be further used for tumor immunotherapy target discovery or personalized vaccine design.

[0039] The mass spectrometry validation process involves comparing the mass spectrometry data of candidate neoantigen peptides with those of tumor tissue to confirm their expression at the protein level. Validated peptides can then be used for immunotherapy target discovery, personalized vaccine design, or functional pathway annotation analysis.

[0040] Based on a similar inventive concept, embodiments of the present invention also provide a computer storage medium storing a readable program that, when run by a processor, can execute the above-described method for tumor tissue splicing recognition and neoantigen peptide screening based on the ONT sequencing platform.

[0041] Based on a similar inventive concept, this invention provides an electronic device, including: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other through the communication bus; The memory is used to store at least one executable instruction, which causes the processor to perform the operations corresponding to the tumor tissue splicing recognition and neoantigen peptide screening method based on the ONT sequencing platform described above.

[0042] Based on a similar inventive concept, embodiments of the present invention also provide a computer program product, including computer instructions, which instruct a computing device to perform the operations corresponding to the above-described tumor tissue splicing recognition and neoantigen peptide screening method based on the ONT sequencing platform.

[0043] Example 2 In this embodiment, a specific example is used to illustrate the technical solution of the present invention; 1. RNA extraction and poly(A)+ RNA enrichment from tumor tissue samples In this embodiment, colorectal cancer (CRC) tumor tissue and adjacent normal tissue were used as sample materials. All tissue samples were flash-frozen in liquid nitrogen. Flash-frozen tumor tissue samples should be immediately transferred to a -80°C ultra-low temperature freezer for long-term storage to avoid repeated freeze-thaw cycles. Sterile cryovials are recommended for tissue aliquoting, and each sample should be labeled with its number and date. This preservation method effectively maintains RNA integrity (RIN≥8), meeting the Agilent 2100 Bioanalyzer standards for clinical sample processing.

[0044] Approximately 30–50 mg of tissue was ground in liquid nitrogen and total RNA was extracted immediately using the Novizan FastPure® Cell / Tissue TotalRNA Isolation Kit V2. Following the instructions, a combination of protease digestion and silica membrane adsorption was used to extract high-purity, well-preserved total RNA.

[0045] RNA concentration and integrity (RIN value) were measured using Nanodrop and Agilent 2100. A RIN value ≥8 and an OD260 / 280 ratio between 1.8 and 2.0 were required.

[0046] VAHTS mRNA Capture Beads 2.0 was used to enrich poly(A)+ RNA from total RNA, ensuring the acquisition of mature mRNA with complete poly(A) tails, removing a large amount of non-coding RNA, and optimizing sequencing efficiency and result quality.

[0047] 2. Nanopore direct RNA sequencing and basecalling Direct RNA sequencing of enriched mRNA samples was performed using the MinION platform from Oxford Nanopore Technologies. Sequencing was conducted using a MinION flow cell (R9.4.1 chemical version, catalog number FLO-MIN004RA) and the accompanying kit SQK-RNA004 (RNA004). After library construction, sequencing was controlled using MinKNOW software (v22.05+), with a typical run time of 48-72 hours. A single flow cell could generate approximately 300-600 GB of raw data in pod5 format. Electrical signal stability and data output rate were monitored in real-time during the run to ensure sequencing quality met standards.

[0048] The data format after shutdown is pod5 / fast5, which contains all raw current signals.

[0049] Use Dorado (v0.8.1) to perform base calling operations and convert pod5 / fast5 data to fastq format.

[0050] To ensure data quality, the following strategy is used to filter raw reads: Reads are filtered based on the average quality value (Q). Reads with Q ≥ 10 are passed and used as clean reads for downstream analysis, while reads with Q < 10 are classified as fail and excluded.

[0051] Statistical analysis of the length and quality distribution of FASTQ sequences was performed to ensure that the sequencing sample quality met the standards.

[0052] 3. Variable splicing event identification and differential analysis Clean reads were aligned to the human reference genome (GRCh38 / hg38) using minimap2 software (parameters: -ax splice -uf -k14). The alignment results were output in BAM format. samtools was used to sort the alignment results, remove low-quality alignments, and calculate the alignment rate. The featureCounts tool was used to quantify expression levels based on the alignment results, and TPM values ​​for each transcript were output for subsequent use.

[0053] The following is the procedure for using SUPPA2 software to identify alternative splicing events (AS) and analyze differences between groups: Download and prepare a GTF file containing gene annotation information for the target species; Using SUPPA2's generateEvents module, all standard variable splice event types are extracted from GTF, including: SE (Skipped Exon): Exon skipping; RI (Retained Intron): Intron retention; MXE (Mutually Exclusive Exons): Mutually exclusive exons; A5SS / A3SS: Variable 5' / 3' splicing site; AF / AL (Alternative First / Last Exon): Alternative first / last exon.

[0054] The statistical distribution of splicing event types in this tumor sample is shown in the figure below. Figure 2 As shown, Figure 2 (a) in the figure is a distribution diagram of alternative splicing event types in adjacent normal tissue (N1). Figure 2Figure (b) shows the distribution of alternative splicing event types in tumor tissue (T1). The proportions of events such as exon skipping (SE), intron retention (RI), variable 5′ splice site (A5SS), variable 3′ splice site (A3SS), mutually exclusive exons (MXE), variable first exon (AF), and variable tail exon (AL) are statistically compared. Figure 2 (a) and Figure 2 As can be seen in (b), the proportions of exon skipping and mutually exclusive exons in tumor tissues are significantly different from those in adjacent normal tissues, suggesting that alternative splicing patterns are remodeled during tumorigenesis, providing a basis for subsequent screening of neoantigen peptides.

[0055] Using the psiPerEvent module, combined with TPM expression data, the PSI (Percent Spliced ​​In) value of each splicing event in different samples was calculated.

[0056] Use the diffSplice module to perform between-group comparisons and output a list of differential splicing events, including: dPSI (difference splicing intensity), p-value (significance test), and FDR (multiple hypothesis correction result).

[0057] Commonly used screening criteria are: |dPSI| ≥ 0.1 and p-value < 0.05.

[0058] 4. Neoantigen peptide prediction based on differential splicing events Novel transcripts involved in differential splicing events were extracted. Using annotation information, the open reading frame (ORF) region of each novel transcript was located, and its complete CDS sequence was extracted. To accurately capture the protein sequence changes caused by splicing variations, the analysis focused on the new connection regions formed across splice sites.

[0059] In the splicing variation region of the novel transcript (i.e., the boundary region spanning two exons), continuous short peptides (9mers) of 9 amino acids in length are extracted using a sliding window method. The extraction principles are as follows: Each 9mer must cover at least one splice event boundary; If multiple splicing structural variations exist, peptides that occur in exon skipping or mutually exclusive exon regions are preferentially retained. All extracted peptides must have their source transcript ID, splicing type, amino acid sequence, and the start and end positions of their respective CDS segments recorded.

[0060] The final result is a set of 9mer peptides containing tens of thousands of segments covering splice boundary regions.

[0061] The NetMHCpan 4.1 online prediction tool was used to predict the affinity of each peptide for human HLA class I subtypes. Input included the 9-mer sequence and HLA subtype (e.g., HLA-A02:01, HLA-B07:02, etc.), and output included predicted IC50 value, %Rank, and bind level (Strong binder / Weak binder, etc.). The selection criteria were as follows: IC50 < 500 nM; %Rank<2; The TPM expression level of the source transcript is >1.

[0062] Only peptides that simultaneously meet all three criteria above will be retained as candidate neoantigen peptides.

[0063] Candidate neoantigen peptides were compared with mass spectrometry data from tumor tissue proteomics. MS / MS peptide identification results were used to confirm the presence of peptides at the protein level. Peptides confirmed by mass spectrometry were defined as high-confidence neoantigens.

[0064] The final output is a structured neoantigen candidate table, which includes: 9mer peptide sequences generated across splicing event regions; The type of splicing event it belongs to (such as SE, RI, A3SS, etc.); Source: Novel transcript number; The matched HLA-I subtype (such as HLA-C) 07:01); Predictive affinity metrics: IC50 value and %Rank_EL; TPM expression levels of source transcripts; Whether it has been verified by mass spectrometry data.

[0065] See Table 1 below for an example of the prediction results for candidate neoantigen peptides.

[0066] Table 1. Partial candidate neoantigen peptides and their predicted affinity for a certain splicing variant transcript. Note: The table shows the MHC affinity prediction results of a portion of the 9mer peptide extracted from a splice variant transcript (ID:peptide44) in a colorectal cancer sample according to the present invention.

[0067] This table is an exemplary sample, showing only a portion of representative data to illustrate the applicability and analytical capabilities of the method. The method of this invention can be extended to multiple samples and multiple transcripts, and the candidate neoantigen peptides obtained through screening can be further used for mass spectrometry validation and functional evaluation.

[0068] The method described in this embodiment can systematically identify neoantigen peptides in tumor tissue caused by alternative splicing, and screen them by combining affinity prediction and expression level information. The introduction of mass spectrometry validation further enhances the biological reliability of candidate peptides, providing effective support for subsequent vaccine design, immunotherapy target development, and personalized treatment strategies.

[0069] Example 3 In this embodiment, a tumor tissue splicing recognition and neoantigen peptide screening system based on the ONT sequencing platform is proposed, specifically including: Sequencing module: RNA is extracted from tumor tissue and mRNA is isolated. Then, a sequencing library is constructed and sequenced using the ONT sequencing platform to obtain raw long read data. Base recognition module: performs base recognition on the raw long read data to obtain fastq format data; Event identification and analysis module: Aligns fastq format data to the reference genome to reconstruct transcripts, and based on the alignment results, identifies alternative splicing events and performs differential analysis; Transcript extraction module: Screens for significantly different alternative splicing events and extracts event-related novel transcript coding sequences; Peptide generation module: Extracts 9mer peptides across splice sites in the splice boundary region of the novel transcript; Peptide screening module: Based on a prediction model, the MHC binding affinity of 9mer peptides is scored, and candidate neoantigen peptides are screened out. Peptide validation module: Performs mass spectrometry validation on candidate neoantigen peptides to screen for peptides that are actually expressed.

[0070] The methods of the present invention can be implemented in hardware, firmware, or as software or computer code that can be stored in a recording medium (such as a CD-ROM, RAM, floppy disk, hard disk, or magneto-optical disk), or as computer code originally stored on a remote recording medium or a non-transitory machine-readable medium and subsequently stored on a local recording medium, downloaded via a network. Thus, the methods described herein can be processed by software stored on a recording medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware (such as an ASIC or FPGA). It is understood that the computer, processor, microprocessor controller, or programmable hardware includes storage components (e.g., RAM, ROM, flash memory, etc.) capable of storing or receiving software or computer code that, when accessed and executed by the computer, processor, or hardware, implements the methods described herein. Furthermore, when a general-purpose computer accesses the code used to implement the methods shown herein, the execution of the code transforms the general-purpose computer into a dedicated computer for performing the methods shown herein.

[0071] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the claimed invention.

Claims

1. A method for tumor tissue splicing recognition and neoantigen peptide screening based on the ONT sequencing platform, characterized in that, Includes the following steps: RNA was extracted from tumor tissue, and mRNA was isolated. Sequencing libraries were then constructed and sequenced using the ONT sequencing platform to obtain raw long read data. Base identification is performed on the raw long-read data to obtain fastq format data; FastQ format data were aligned to a reference genome to reconstruct transcripts, and alternative splicing events were identified and differential analyses were performed based on the alignment results. Significantly different alternative splicing events were screened out, and the coding sequences of novel transcripts associated with these events were extracted. The 9mer peptide across the splice boundary region of the novel transcript was extracted; The MHC binding affinity of the 9mer peptide was scored, and candidate neoantigen peptides were screened. Mass spectrometry was used to validate candidate neoantigen peptides and screen for peptides that were actually expressed.

2. The method for tumor tissue splicing recognition and neoantigen peptide screening based on the ONT sequencing platform according to claim 1, characterized in that, The process of identifying alternative splicing events is as follows: using SUPPA2 software, combined with annotation files and TPM expression levels, the types of alternative splicing events present in each sample are identified, including: exon skipping, intron retention, variable 5' splicing sites, variable 3' splicing sites, and mutually exclusive exons.

3. The method for tumor tissue splicing recognition and neoantigen peptide screening based on the ONT sequencing platform according to claim 2, characterized in that, The process of differential analysis of alternative splicing events is as follows: using the DiffSplice module of SUPPA2, the PSI values ​​of splicing events are compared between groups, ΔPSI is calculated and significance is tested, thereby identifying alternative splicing events with significant differences.

4. The method for tumor tissue splicing recognition and neoantigen peptide screening based on the ONT sequencing platform according to claim 3, characterized in that, When screening for significantly different alternative splicing events, FDR < 0.05 and |ΔPSI| > 0.1 were used as significance screening thresholds.

5. The method for tumor tissue splicing recognition and neoantigen peptide screening based on the ONT sequencing platform according to claim 1, characterized in that, The process of extracting the 9mer peptide is as follows: using a sliding window method, extract a continuous 9-amino acid peptide spanning at least two different exon regions from the novel transcript corresponding to the alternative splicing event.

6. The method for tumor tissue splicing recognition and neoantigen peptide screening based on the ONT sequencing platform according to claim 1, characterized in that, The criteria for screening candidate neoantigen peptides are: MHC binding affinity prediction value less than 500 nM, %Rank less than 2, and source transcript TPM greater than 1.

7. The method for tumor tissue splicing recognition and neoantigen peptide screening based on the ONT sequencing platform according to claim 1, characterized in that, The mass spectrometry verification process involves comparing the mass spectrometry data of the candidate neoantigen peptide with those of the tumor tissue to confirm its expression at the protein level.

8. A tumor tissue splicing recognition and neoantigen peptide screening system based on the ONT sequencing platform, characterized in that, include: Sequencing module: RNA is extracted from tumor tissue and mRNA is isolated. Then, a sequencing library is constructed and sequenced using the ONT sequencing platform to obtain raw long read data. Base recognition module: performs base recognition on the raw long read data to obtain fastq format data; Event identification and analysis module: Aligns fastq format data to the reference genome to reconstruct transcripts, and based on the alignment results, identifies alternative splicing events and performs differential analysis; Transcript extraction module: Screens for significantly different alternative splicing events and extracts the coding sequences of novel transcripts related to these events; Peptide generation module: Extracts 9mer peptides across splice sites in the splice boundary region of the novel transcript; Peptide screening module: MHC binding affinity score is performed on 9mer peptides, and candidate neoantigen peptides are screened out; Peptide validation module: Performs mass spectrometry validation on candidate neoantigen peptides to screen for peptides that are actually expressed.

9. A computer storage medium storing a readable program, characterized in that, When the program is running, it can instruct the computing device to perform the tumor tissue splicing recognition and neoantigen peptide screening method based on the ONT sequencing platform as described in any one of claims 1-7.

10. An electronic device, characterized in that, include: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction, which causes the processor to perform the operation corresponding to the tumor tissue splicing recognition and neoantigen peptide screening method based on the ONT sequencing platform as described in any one of claims 1-7.