A method for finding novel alternative splicing events for third generation full-length transcript sequencing data

By using TAGET software and filtering steps to identify novel alternative splicing events in third-generation full-length transcript sequencing data, this approach solves the detection challenges in existing technologies, achieving efficient and accurate alternative splicing event identification and isoform analysis, while saving resources and time.

CN121506243BActive Publication Date: 2026-04-17BEIJING VIEWSOLIDBIOTECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING VIEWSOLIDBIOTECH
Filing Date
2026-01-09
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively identify and detect the occurrence and types of novel alternative splicing events in third-generation full-length transcript sequencing data.

Method used

Data mapping and annotation were performed using TAGET software to screen NNC and NIC transcripts. Alternative splicing events, including intron retention, exon skipping, and splicing site changes, were identified through a filtering process. Intron splicing was verified using the GT-AG rule, and the results of alternative splicing events were output.

Benefits of technology

It improves the accuracy and efficiency of detecting novel alternative splicing events, reduces the time and economic cost of repeated sequencing, can directly identify the generation of novel alternative splicing isoforms, and is suitable for high-quality sequencing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121506243B_ABST
    Figure CN121506243B_ABST
Patent Text Reader

Abstract

The present application relates to a method for finding new variable splicing events in three-generation full-length transcript sequencing data in the field of bioinformatics. The method for detecting new variable splicing events in three-generation full-length transcript sequencing data of a test sample of the present application comprises the following steps: 1) obtaining three-generation full-length transcript sequencing data of a test sample; 2) performing gene function annotation on the three-generation full-length transcript sequencing data of step 1) to obtain transcript set 1, wherein the transcript set 1 comprises NNC and NIC; 3) screening whether variable splicing events occur in the transcript set 1 of step 2) to obtain transcript set 2 with variable splicing events; 4) filtering and screening the transcript set 2; 5) after step 4) is completed, screening new splicing isoforms formed by known variable splicing events and new splicing isoforms formed by multiple variable splicing events; 6) after step 5) is completed, outputting a variable splicing event result information file.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method in the field of bioinformatics for finding novel alternative splicing events in third-generation full-length transcript sequencing data. Background Technology

[0002] In full-length transcript sequencing, PacBio technology utilizes the Iso-Seq (Isoform Sequencing) method. Iso-Seq technology can directly sequence complete mRNA molecules without fragmenting them, thus obtaining the full-length transcript sequence. This method is particularly suitable for studying the complexity of gene expression, including alternative splicing and transcript isoforms.

[0003] Alternative splicing is one of the core mechanisms regulating gene expression in eukaryotes. It involves differentially splicing the precursor mRNA of a single gene to generate mature mRNA isoforms with distinct structures and functions. Partially spliced ​​isoforms can competitively bind to targets (such as receptors, signaling proteins, or domains), interfering with the activity of functional isoforms and thus inhibiting their normal biological functions. This offers more possibilities for disease research, especially targeted cancer therapy. Summary of the Invention

[0004] The main problem this invention aims to solve is how to identify or detect the occurrence and type of novel alternative splicing events in third-generation full-length transcript sequencing data of a test sample.

[0005] To address the aforementioned issues, this invention provides a method for detecting novel alternative splicing events in third-generation full-length transcript sequencing data of a sample.

[0006] The method for detecting novel alternative splicing events in third-generation full-length transcript sequencing data of a test sample provided by this invention includes the following steps:

[0007] 1) Obtain the third-generation full-length transcript sequencing data of the sample to be tested;

[0008] 2) Gene function annotation is performed on the third-generation full-length transcript sequencing data described in step 1), and transcript set 1 is obtained by screening. Transcript set 1 includes NNC and NIC; NNC represents new transcripts containing new exons, and NIC represents new transcripts with known exon combinations.

[0009] 3) Screen the transcripts in transcript set 1 in step 2) one by one to see if alternative splicing events have occurred, and obtain transcripts 2 that have undergone alternative splicing events;

[0010] 4) Filter transcript set 2, filtering for the following events in sequence:

[0011] a. Filter the intron retention events in the transcript set 2 to remove intron retention events that exist only in the left and right edge exons of a transcript; the removal of IR (intron retention) events that exist only in the left and right edge exons of a transcript is to reduce the impact of intron excision errors on the left and right sides of the transcript in third-generation sequencing on the judgment of alternative splicing events;

[0012] b. Filter for events of recessive exon insertion, selective 3' splicing site, and selective 5' splicing site, verify whether the newly generated introns of the splicing events conform to the GT-AG rule, and remove splicing events where the newly generated introns do not conform to the GT-AG rule at both ends;

[0013] The GT-AG rule refers to the core rule of highly conserved sequences at the splicing boundaries of introns in eukaryotic nuclear genes. That is, at the DNA level, the intron start segment (5' end) must be GT and the intron end segment (3' end) must be AG. It also includes the low-probability GC-AG and AT-AC combinations.

[0014] c. Filter exon skipping events, verify whether small exon skipping events actually exist, and delete small exon skipping events that actually exist but are not annotated due to mapping issues;

[0015] After the above filtering, the target transcript of the alternative splicing event is obtained;

[0016] 5) After completing step 4), the new splice isomers formed by known alternative splicing events in the target transcript of the alternative splicing event are marked as Novel Junction Combination; the new splice isomers formed by multiple alternative splicing events obtained by screening are marked as Multiple Events;

[0017] 6) After completing step 5), output the variable splicing event result information file.

[0018] Further, in step 2), the TAGET software is used for mapping and annotation.

[0019] Furthermore, before annotation in step 2), the sequencing data from step 1) is processed to obtain a FASTA data file, wherein the processing is performed using isoseq3 (version 3.3.0) software.

[0020] The specific operation is as follows: use TAGET software to map the fasta file to the reference genome database to obtain transcript mapping results and annotation results.

[0021] Further, step 3) of the screening includes the following steps:

[0022] m. Based on the gene annotated with the measured transcripts, count all exon regions that may be covered by the transcripts of the gene;

[0023] n. If one or more exons in all the exons of the measured transcript set 1 are located in regions that belong to the reference transcript but not to regions covered by known exons of the reference transcript, then the generation of the transcript set 1 can be considered to be related to a recessive exon insertion event.

[0024] o. Obtain a set of reference transcripts based on the annotated genes of the experimental transcripts. Filter out reference transcripts whose number of exons is more than one less than the number of exons in the experimental transcripts (number of exons in reference transcripts - number of exons in experimental transcripts < -1). Cyclicly match the remaining reference transcripts and experimental transcripts. Determine the type of alternative splicing events in the experimental transcripts based on the number of alternative splicing events, the degree of difference between the experimental transcripts and the reference transcripts, and other conditions.

[0025] Furthermore, the variable splicing event types described in step o include splicing events involving changes in the selection of 5' or 3' splicing sites, exon skipping, and intron retention.

[0026] Furthermore, the method for determining the type of alternative splicing events occurring in transcript set 1 is as follows:

[0027] 1) Determining changes in splice site selection:

[0028] Compare the introns of the transcript set 1 with the introns of the reference transcript to determine whether a selective 3' splicing event has occurred. If the introns of the measured transcripts have a different start position but the same end position compared to the corresponding introns in the reference transcripts, then a selective 3' splicing event has occurred; otherwise, a selective 3' splicing event has not occurred.

[0029] Compare the introns of the transcript set 1 with the introns of the reference transcript to determine whether a selective 5' splicing event has occurred; if the introns of the tested transcripts have the same start position but a different end position compared to the corresponding introns in the reference transcripts, then a selective 5' splicing event has occurred, otherwise, a selective 5' splicing event has not occurred.

[0030] 2) Determining splicing exon skipping events:

[0031] First, determine the exon positions mapped by the exons in transcript set 1 to the reference transcript. If there are exons that cannot be mapped, then transcript set 1 has not experienced an ES (exon skipping) event relative to the current reference transcript; otherwise, continue with subsequent analysis.

[0032] If all exon mapping positions of transcript set 1 are continuous, then transcript set 1 does not have an exon skipping event relative to the current reference transcript; otherwise, continue with subsequent analysis.

[0033] Then it is verified whether the missing exons in transcript set 1 are due to intron retention events and merging with other exons. If they are true missing exons, then the transcripts in transcript set 1 have an ES (exon skipping) event relative to the current reference transcript.

[0034] 3) Intron retention event determination: Compare the introns of the transcript set 1 with the introns of the reference transcript to determine whether an intron retention event has occurred; if the introns of the transcript can be matched one by one with the introns of the reference transcript, and there are no cases where they cannot be matched or the differences are large, but one or more introns are missing compared with the introns of the reference transcript, and the missing introns are not at the beginning or end of the matched intron set, then an intron retention event has occurred; otherwise, an intron retention event has not occurred.

[0035] 4) Count the number of alternative splicing events in the experimental transcripts relative to different reference transcripts, and retain the set of reference transcripts with the fewest events;

[0036] 5) Compare the difference between the transcript and the reference transcript within the exon mapping range, i.e., the sum of the number of bases deleted or added in the transcript due to alternative splicing events, retain the reference transcript with the lowest difference from the transcript, and the alternative splicing events that occurred in the transcript relative to the reference transcript.

[0037] This invention also provides a novel alternative splicing event classification system based on third-generation full-length transcript sequencing data, the system comprising:

[0038] 1) Acquisition Unit: Used to acquire third-generation full-length transcript sequencing data of the sample to be analyzed;

[0039] 2) Alternative splicing event acquisition unit: used to perform gene function annotation on the full-length genome to obtain transcripts; includes transcript set 1 screening unit and transcript 2 screening unit;

[0040] Transcript Set 1 Screening Unit: Used to screen the full-length genome to obtain Transcript Set 1. The Transcript Set 1 Screening Unit includes NNC and NIC; NNC represents new transcripts containing new exons, and NIC represents new transcripts with known exon combinations.

[0041] The criteria for screening the full-length genome included: based on TAGET annotation results. Transcripts tagged with NIC and NNC were retained during the screening.

[0042] Transcript 2 screening unit: used to further screen the transcript set 1 to obtain transcript 2;

[0043] The criteria for further screening of the transcript set 1 include: alternative splicing is defined as the type of alternative splicing event including changes in splice site selection, splice exon skipping, splice intron retention, and mutually exclusive splicing events;

[0044] 3) Alternative splicing event filtering unit: used to filter and select based on transcript 2 to obtain target transcripts of alternative splicing events;

[0045] 4) Alternative splicing event reading and output unit: used for classifying the alternative splicing event types based on the target transcript and reading and outputting the alternative splicing event type results.

[0046] The present invention also provides a novel alternative splicing event identification device based on third-generation full-length transcript sequencing data, the device comprising: a memory and a processor;

[0047] The memory is used to store program instructions;

[0048] The processor is used to invoke program instructions, which, when executed, are used to perform any of the methods described above for detecting variable splicing events.

[0049] The present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the methods described above.

[0050] The computer program product may be a software product whose solution is primarily implemented through a computer program.

[0051] The computer-readable storage medium refers to a carrier for storing data, which may be magnetic tape, disk, floppy disk, optical disk, magneto-optical disk, ROM, PROM, VCD, DVD, hard disk, flash memory, USB flash drive, CF card, SD card, MMC card, SM card, Memory Stick, or xD card, etc.

[0052] The method for detecting novel alternative splicing in third-generation full-length transcripts established in this invention takes into account the advantage of complete transcript mapping inherent in third-generation full-length transcripts. Compared with second-generation sequencing, which requires short reads to piece together long sequence isoforms using algorithms, this method more accurately and quickly determines whether a transcript is a novel splicing isoform generated by an alternative splicing event. It fills the gap in the detection of novel alternative splicing in third-generation full-length transcripts and improves detection accuracy.

[0053] The method established in this invention ensures that novel alternative splicing events can be effectively detected in samples with only third-generation full-length transcript data, saving the time and economic cost of re-sequencing to identify novel alternative splicing. This invention has significant value in detecting novel alternative splicing genotyping in third-generation full-length transcripts, and has high application prospects, especially when the sequencing sample quality is sufficiently good. Attached Figure Description

[0054] Figure 1 This is a flowchart of a novel alternative splicing detection method for third-generation full-length transcripts. Detailed Implementation

[0055] The present invention will now be described in further detail with reference to specific embodiments. The given embodiments are merely illustrative of the invention and not intended to limit its scope. The embodiments provided below can serve as a guide for further improvements by those skilled in the art and do not constitute a limitation on the invention in any way.

[0056] Unless otherwise specified, the experimental methods used in the following examples are conventional methods, performed according to the techniques or conditions described in the literature in this field or according to the product instructions. Unless otherwise specified, the materials and reagents used in the following examples are commercially available.

[0057] Unless otherwise specified, the quantitative experiments in the following examples are all repeated three times, and the results are averaged.

[0058] The eight samples in the following examples are described in: Xia, Y., Jin, Z., Zhang, C., Ouyang, L., Dong, Y., & Li, J., et al. (2023 “TAGET: a toolkit for analyzing full-length transcripts from long-read sequencing”, Nature Communications, Vol. 14, p. 5935). This biological material is publicly available from the applicant and is intended solely for the replication of experiments of this invention and may not be used for any other purpose.

[0059] Example 1: Establishment of a novel alternative splicing detection method for third-generation full-length transcripts

[0060] This invention, through extensive experimentation, establishes a novel method for detecting variable splicing of third-generation full-length transcripts. Figure 1 The specific steps are as follows:

[0061] 1. Using TAGET software, the full-field transcript data of the three generations were mapped to the reference genome database and annotated to obtain the mapping information and annotation information of the transcripts;

[0062] 2. After completing step 1, based on the annotation results of TAGET, the transcripts marked as NNC and NIC (referred to as transcript set 1) are selected according to the annotation information to determine whether they are splice isoforms; where NNC represents a new transcript containing a new exon, and NIC represents a new transcript with a known exon combination;

[0063] 3. After completing step 2, for each labeled transcript, determine whether an alternative splicing event has occurred; specifically as follows:

[0064] 1) Based on the gene annotated with transcripts, count the exon regions contained in all known transcripts of that gene;

[0065] 2) If one or more exons in a measured transcript belong to the gene but are not covered by the known exons of the gene, the generation of the transcript can be considered to be related to a CEI (recessive exon insertion) event.

[0066] 3) Based on the annotated genes of the experimental transcripts, obtain a set of reference transcripts. Filter out reference transcripts whose number of exons is more than one less than the number of exons in the experimental transcripts (number of exons in reference transcripts - number of exons in experimental transcripts < -1). Cyclicly match the remaining reference transcripts with the experimental transcripts. Based on the number of alternative splicing events, the degree of difference between the experimental transcripts and the reference transcripts, determine which alternative splicing events are related to the generation of the experimental transcripts, and identify the most relevant reference transcripts. The specific steps are as follows:

[0067] A. To determine whether the tested transcript produced an ES (exon skipping) event, the specific steps are as follows:

[0068] I. Determine the exon positions that the tested transcripts map to on the reference transcripts. If there are tested exons that cannot be mapped, then the tested transcripts do not have an ES (exon skipping) event relative to the current reference transcripts; otherwise, continue with the subsequent analysis.

[0069] II. If all exon mapping positions of the measured transcript are continuous, then there is no ES (exon skipping) event relative to the current reference transcript; otherwise, continue with subsequent analysis.

[0070] III. Verify whether the missing exons in the experimental transcript have merged with other exons due to an IR (intron retention) event. If it is a true deletion, then the experimental transcript has an ES (exon skipping) event relative to the current reference transcript.

[0071] B. Compare the introns of the experimental transcript with the introns of the reference transcript to determine whether IR (intron retention) events, A3SS (selective 3' splicing site) events, and A5SS (selective 5' splicing site) events have occurred;

[0072] C. Compare the difference between the experimental transcript and the reference transcript within the exon mapping range, i.e. the sum of the number of bases deleted or added in the experimental transcript due to alternative splicing events, retain the reference transcript with the lowest difference from the experimental transcript, and the alternative splicing events that occurred in the experimental transcript relative to the reference transcript.

[0073] Transcript 2 was obtained.

[0074] 4. After completing step 3, filter for IR (intron retention) events and remove IR (intron retention) events that only exist in the left and right edge exons of a transcript. Removing IR (intron retention) events that only exist in the left and right edge exons of a transcript is to reduce the impact of intron excision errors on the left and right sides of the transcript in third-generation sequencing on the judgment of alternative splicing events.

[0075] 5. After completing step 4, filter for CEI (recessive exon insertion), A3SS (selective 3' splicing site), and A5SS (selective 5' splicing site) events to verify whether the newly generated introns conform to the GT-AG rule. Remove splicing events where the ends of the newly generated introns do not conform to the GT-AG rule. The GT-AG rule refers to the core rule of highly conserved sequences at the splicing boundaries of introns in eukaryotic nuclear genes. That is, at the DNA level, the intron start segment (5' end) must be GT, and the intron termination segment (3' end) must be AG. It also includes a small probability of GC-AG and AT-AC combinations.

[0076] 6. After completing step 5, filter the ES (exon skip) events, verify whether the small fragments of ES (exon skip) events actually exist, and delete the small fragments of ES (exon skip) events that actually exist but are not annotated due to mapping issues.

[0077] 7. After completing step 6, select new splice isomers formed by known variable splicing events and label them as NJC (Novel Junction Combination); select new splice isomers formed by multiple variable splicing events and label them as ME (Multiple Events).

[0078] 8. After completing step 8, output the result information file.

[0079] Example 2: Using the method established in Example 1, novel alternative splicing events were identified in eight full-length third-generation transcripts.

[0080] The novel alternative splicing was detected in eight experimental samples using the method established in Example 1. The specific steps are as follows:

[0081] 1. Obtained third-generation full-length transcript sequencing data from 8 samples;

[0082] Eight samples were sequenced using the PacBio platform to obtain corresponding third-generation full-length transcript data;

[0083] 2. After completing step 1, rename the full-length transcript sequencing data of the 8 samples as Sample1, Sample2, Sample3, Sample4, Sample5, Sample6, Sample7, and Sample8. Then, use TAGET software to map the sequencing sequences onto the reference genome and annotate the transcripts to obtain the mapping and annotation information of the transcripts.

[0084] 3. Based on the annotation results of TAGET, transcripts marked as NNC and NIC are selected to determine whether they are splice isoforms; NNC indicates a new transcript containing a new exon, and NIC indicates a new transcript with a known exon combination.

[0085] 4. After completing step 3, for each tagged transcript, determine whether an alternative splicing event has occurred.

[0086] 5. After completing step 4, filter for IR (intron retention) events to remove IR (intron retention) events that only exist in the left and right edge exons of a transcript;

[0087] 6. After completing step 5, filter the CEI (recessive exon insertion), A3SS (selective 3' splicing site), and A5SS (selective 5' splicing site) events, verify whether the newly generated introns of the splicing events conform to the GT-AG rule, and remove splicing events whose newly generated introns do not conform to the GT-AG rule at both ends.

[0088] 7. After completing step 6, filter the ES (exon skip) events, verify whether the small fragments of ES (exon skip) events actually exist, and delete the small fragments of ES (exon skip) events that actually exist but are not annotated due to mapping issues.

[0089] 8. After completing step 7, select new splice isomers formed by known variable splice events and label them as NJC (Novel Junction Combination); select new splice isomers formed by multiple variable splice events and label them as ME (Multiple Events).

[0090] 9. After completing step 8, output the result information file.

[0091] Information on the amount of alternative splicing in the three generations of full-length transcripts from the eight samples is shown in Table 1.

[0092]

[0093] The above results demonstrate that the method established in Example 1 can detect novel alternative splicing in three generations of full-length transcripts.

[0094] In summary, the method for detecting novel variable splicing established in this study has the following beneficial effects:

[0095] (1) The isomers have good integrity and high accuracy, which is convenient for protein level studies.

[0096] For the eight experimental samples in Example 2, after obtaining the novel alternative splicing events using the method provided by this invention, it is possible to directly know which isoforms were produced by these novel alternative splicing events without splicing short read sequences. The accuracy is high, and the changes in the splicing isoforms produced by the novel alternative splicing events at the protein level can be intuitively obtained.

[0097] (2) It has high data utilization, saves samples, avoids multiple sequencing, and reduces time and economic costs.

[0098] The method provided by this invention can effectively identify alternative splicing events in samples using third-generation full-length transcript sequencing data, without requiring the samples to be reused or collected for other sequencing analyses, thus saving the time and economic costs of other sequencing methods.

[0099] The present invention has been described in detail above. Those skilled in the art will recognize that the invention can be practiced in a wide range of ways with equivalent parameters, concentrations, and conditions without departing from its spirit and scope, and without requiring unnecessary experiments. While specific embodiments have been provided, it should be understood that further modifications can be made to the invention. In summary, according to the principles of the invention, this application is intended to include any changes, uses, or improvements to the invention, including changes made using conventional techniques known in the art that depart from the scope disclosed herein.

Claims

1. A method for detecting novel alternative splicing events in third-generation full-length transcript sequencing data of a test sample, characterized in that, The method includes the following steps: 1) Obtain the third-generation full-length transcript sequencing data of the sample to be tested; 2) Gene function annotation is performed on the third-generation full-length transcript sequencing data described in step 1), and transcript set 1 is obtained by screening. Transcript set 1 includes NNC and NIC; NNC represents new transcripts containing new exons, and NIC represents new transcripts with known exon combinations. 3) Screen the transcripts in transcript set 1 in step 2 one by one to see if alternative splicing events have occurred, and obtain transcript set 2 in which alternative splicing events have occurred; 4) Filter transcript set 2, filtering for the following events in sequence: a. Filter the intron retention events in the transcript set 2 to remove intron retention events that only exist in the left and right edge exons of a transcript; b. Filter for events of recessive exon insertion, selective 3' splicing site, and selective 5' splicing site, verify whether the newly generated introns of the splicing events conform to the GT-AG rule, and remove splicing events where the newly generated introns do not conform to the GT-AG rule at both ends; c. Filter exon skipping events, verify whether small exon skipping events actually exist, and delete small exon skipping events that actually exist but are not annotated due to mapping issues; 5) After completing step 4), new splice isomers formed by known variable shear events are selected and labeled as Novel Junction Combination; new splice isomers formed by multiple variable shear events are selected and labeled as Multiple Events. 6) After completing step 5), output the variable splicing event result information file.

2. The method of claim 1, wherein, In step 2), the TAGET software is used for mapping and annotation.

3. The method according to claim 1 or 2, characterized in that, Step 2) before annotation also includes processing the sequencing data from step 1) to obtain a FASTA data file, wherein the processing is performed using isoseq3 software version 3.3.

0.

4. The method of claim 3, wherein, Step 3) The screening of each transcript includes the following steps: m. Based on the gene annotated with measured transcripts, count all exon regions that the gene may cover in all transcripts; n. If one or more exons in all the exons of the measured transcript are located in regions that belong to the reference transcript but are not covered by known exons of the reference transcript, then the generation of the transcript can be considered to be related to a recessive exon insertion event; o. Based on the annotated genes of the experimental transcripts, obtain a set of reference transcripts. Filter out reference transcripts that have more than one fewer exon than the experimental transcripts. Circularly match the remaining reference transcripts with the experimental transcripts. Determine the type of alternative splicing events in the experimental transcripts based on the number of alternative splicing events and the degree of difference between the experimental transcripts and the reference transcripts.

5. The method of claim 4, wherein, The variable splicing event types described in step o include splicing events involving changes in the selection of 5' or 3' splicing sites, exon skipping, and intron retention.

6. The method according to claim 4 or 5, characterized in that, The method for determining the type of alternative splicing event occurring in transcripts in transcript set 1 is as follows: 1) Determining events indicating changes in splice site selection: Compare the introns of the transcript with the introns of the reference transcript to determine whether a selective 3' splicing event has occurred. If the introns of the measured transcript have a different start position but the same termination position compared to the corresponding introns in the reference transcript, then a selective 3' splicing event has occurred; otherwise, a selective 3' splicing event has not occurred. Compare the introns of the transcript with those of the reference transcript to determine whether a selective 5' splicing event has occurred; If the introns in the measured transcript have the same start position but a different end position compared to the corresponding introns in the reference transcript, then a selective 5' splicing event has occurred; otherwise, a selective 5' splicing event has not occurred. 2) Determining exon skipping events: First, determine the exon positions that the exons in the transcript map to on the reference transcript. If there are exons that cannot be mapped, then the transcript has not experienced an exon skipping event relative to the current reference transcript; otherwise, continue with the subsequent analysis. If all exon mapping positions of the transcript are continuous, then the transcript does not have an exon skipping event relative to the current reference transcript; otherwise, continue with subsequent analysis. The missing exons in the transcript were then verified to determine whether they were due to intron retention events or merging with other exons. If they were true missing exons, then the transcript contained an exon skipping event relative to the current reference transcript. 3) Intron retention event judgment: The introns of the transcript are compared with the introns of the reference transcript to determine whether an intron retention event has occurred. If the introns of the transcript can be matched one by one with the introns of the reference transcript, and there are no cases where they cannot be matched or the differences are large, but one or more introns are missing compared with the introns of the reference transcript, and the missing introns are not at the beginning or end of the set of matched introns, then an intron retention event has occurred; otherwise, an intron retention event has not occurred. 4) Count the number of alternative splicing events in the experimental transcripts relative to different reference transcripts, and retain the set of reference transcripts with the fewest events; 5) Compare the difference between the transcript and the reference transcript within the exon mapping range, i.e., the sum of the number of bases deleted or added in the transcript due to alternative splicing events, retain the reference transcript with the lowest difference from the transcript, and the alternative splicing events that occurred in the transcript relative to the reference transcript.

7. A system for classifying novel alternative splicing events based on third generation full-length transcript sequencing data, comprising: The system includes: 1) Acquisition Unit: Used to acquire third-generation full-length transcript sequencing data of the sample to be analyzed; 2) Alternative splicing event acquisition unit: used for gene function annotation of the full-length genome to obtain transcripts; includes transcript set 1 screening unit and transcript 2 screening unit; The transcript set 1 screening unit is used to screen the full-length genome to obtain transcript set 1. The transcript set 1 screening unit includes NNC and NIC; NNC represents new transcripts containing new exons, and NIC represents new transcripts with known exon combinations. The criteria for screening the full-length genome include: selecting and retaining transcripts tagged as NIC and NNC based on TAGET annotation results; The transcript 2 screening unit is used to further screen the transcript set 1 to obtain transcript 2; The criteria for further screening of the transcript set 1 include: alternative splicing is defined as the type of alternative splicing event including changes in splice site selection, splice exon skipping, splice intron retention, and mutually exclusive splicing events; 3) Alternative splicing event filtering unit: used to filter and select based on transcript 2 to obtain target transcripts of alternative splicing events; 4) Alternative splicing event reading and output unit: used for classifying the alternative splicing event types based on the target transcript and reading and outputting the alternative splicing event type results.

8. A novel alternative splicing event identification device based on third-generation full-length transcript sequencing data, characterized in that, The device includes: a memory and a processor; The memory is used to store program instructions; The processor is used to invoke program instructions, which, when executed, are used to perform the method described in any one of claims 1-6.

9. A computer readable storage medium having stored thereon a computer program, characterized in that, When the computer program is executed by a processor, it implements the method described in any one of claims 1-6.

Citation Information

Patent Citations

  • Eucaryon alternative splicing analysis method and system based on RNA-seq data

    CN107766696A

  • Analysis method and system of parametric transcriptome sequencing data adaptive to ONT sequencing

    CN116013415A