Systems and methods for interpreting transcript expression levels from RNA sequencing data using local unique features

By extracting unique features in gene transcripts and establishing a database, the challenge of estimating gene and transcript expression in RNA sequencing data is solved, and rapid and efficient extraction and compilation of expression information is achieved.

CN112041933BActive Publication Date: 2025-05-13KONINKLIJKE PHILIPS NV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201980025788.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2018-03-14
Filing Date
2019-03-13
Publication Date
2025-05-13
Estimated Expiration
2039-03-13

AI Technical Summary

Technical Problem

Existing tools face challenges in estimating gene and transcript expression levels from RNA sequencing data, including assigning sequencing reads to corresponding transcripts, uneven distribution of read coverage, etc., and tools that rely on complete RNA sequencing reads are inefficient in analyzing complex transcriptome structures.

Method used

A unique feature database is established by extracting unique features such as unique exons, exon junctions, introns, initiation and termination locations from gene transcripts, and comparing sequencing sequences with these features to identify gene transcripts and compile expression level information.

Benefits of technology

This method can quickly and efficiently extract the expression information of gene transcripts from RNA sequencing data, reduce calculation time, and is suitable for processing low-quality RNA sequencing data and complex transcriptome structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112041933B_ABST
    Figure CN112041933B_ABST
Patent Text Reader

Abstract

A method (100) for characterizing gene transcript expression levels, comprising: (i) extracting (110) one or more unique features from each of a plurality of gene transcripts; (ii) storing (120) the extracted unique features in a unique feature database; (iii) receiving (130) a plurality of sequences sequenced from the gene transcripts, wherein at least some of the sequences contain one or more extracted unique features; (iv) comparing the plurality of sequences with the extracted unique features stored in the unique feature database by a processor (140); (v) based on a match between the sequence and the extracted unique features, identifying (150) a gene transcript and / or a gene that is generated based on the gene transcript and / or a gene; and (vi) based on the identified gene transcripts, compiling (160) information about the gene transcript expression levels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure is generally directed to methods and systems for characterizing gene transcript expression levels using unique features in gene transcripts. Background Art

[0002] RNA sequencing is an important tool for transcriptome studies. This high-throughput technology offers several advantages over previous techniques, including the ability to detect novel and low-expressed transcripts with a wide dynamic range.

[0003] Protein diversity in eukaryotes is greatly increased by alternative splicing, which greatly increases the complexity of the transcriptome. For example, it is estimated that more than 90% of multi-exon human genes undergo alternative splicing, many of which are revealed by RNA sequencing data. The expression of these transcript variants is highly modulated and differentially expressed in different tissues or developmental stages and in tumors or diseases. As a result, estimating gene and transcript expression from RNA sequencing data is a key element of basic and clinical bioinformatics research.

[0004] However, estimating gene and transcript expression from RNA sequencing data is challenging. For example, since many genes express more than one transcript, assigning sequencing reads to the transcripts from which they originate is a major problem that any transcript expression estimation program must address. Other challenges include, for example, uneven distribution of read coverage, etc.

[0005] Current tools attempt to resolve the structure of different expressed isoforms and estimate their expression levels based on RNA sequencing data. For example, some software can assemble RNA sequencing reads to a minimum number of transcripts to try to identify all fragments, and then use the generated statistical models to estimate transcript abundance. Other analysis software maps reads directly to the transcriptome rather than the genome, and then uses models to assign reads to different isoforms.

[0006] However, these current tools cannot address all challenges faced when analyzing RNA sequencing data. For example, tools typically examine the entire RNA sequencing read from the transcript start site to the transcript stop site, which is both time-consuming and computationally inefficient. In addition, as the complexity of parsing transcriptome structure increases (e.g., small modulated RNAs or low-quality RNA sequencing data), the effectiveness of tools that rely on complete RNA sequencing reads decreases. Summary of the invention

[0007] There remains a need for tools to effectively and efficiently determine gene transcript expression levels from RNA sequencing data.

[0008] The present disclosure is directed to creative methods and systems for characterizing gene transcript expression levels based on RNA sequencing data. Various embodiments and implementations herein are directed to a system for extracting unique features from gene transcripts, including but not limited to unique exons, unique exon connections, unique introns, unique start positions, and / or unique end positions, etc. The system receives or sequences gene transcripts and compares the sequences with the extracted unique features, which are stored in a unique feature database. Based on the matching between these sequences and the extracted unique features, the system identifies gene transcripts and compiles information about gene transcript expression levels.

[0009] Generally, in one aspect, a method for characterizing the expression level of a gene transcript is provided. The method comprises: (i) extracting one or more unique features from each of a plurality of gene transcripts; (ii) storing the extracted unique features in a unique feature database; (iii) receiving a plurality of sequences sequenced from the gene transcripts, wherein at least some of the sequences contain one or more extracted unique features; (iv) comparing the plurality of sequences with the extracted unique features stored in the unique feature database by a processor; (v) based on a match between the sequence and the extracted unique features, identifying the gene transcript as follows: the sequence is generated from the gene transcript; and (vi) based on the identified gene transcript, compiling information about the transcript expression level.

[0010] According to one embodiment, the unique features include one or more of the following: a unique exon, a unique exon junction, a unique intron, a unique start position and / or a unique end position.

[0011] According to one embodiment, comparing comprises aligning each of the plurality of sequences sequenced from gene transcripts to one or more unique features.

[0012] According to one embodiment, the method further comprises the step of providing a sample for RNA sequencing.

[0013] According to one embodiment, the method further comprises the step of sequencing gene transcripts from one or more cells to generate the plurality of sequences.

[0014] According to one embodiment, the method further comprises the step of associating at least some of the extracted unique features with annotation information in the unique feature database.

[0015] According to one embodiment, the unique signature database includes extracted unique signatures instead of complete gene transcripts.

[0016] According to one embodiment, the identifying step comprises the likelihood that the identified gene transcript is the transcript from which the sequence was generated.

[0017] According to one embodiment, the sequence matches unique features extracted from two different genes, and the identifying step comprises identifying two or more gene transcripts from which the sequence is generated, or may have been generated.

[0018] According to one aspect, a system for characterizing the expression level of a gene transcript is provided. The system includes: a database of unique features extracted from each of a plurality of gene transcripts; a comparison module configured to: (i) compare a plurality of sequences sequenced from the gene transcripts with the extracted unique features stored in the unique feature database; and (ii) based on the match between the sequence and the extracted unique features, identify the following gene transcript, the sequence is generated according to the gene transcript; and a compilation module configured to compile information about the expression level of the gene transcript based on the identified gene transcripts.

[0019] According to one embodiment, the system further comprises a feature extraction module, the feature extraction module being configured to extract the unique features from the plurality of gene transcripts. According to one embodiment, the feature extraction module is further configured to associate at least some of the extracted unique features with annotation information.

[0020] In various embodiments, a processor or controller may be associated with one or more storage media (generally referred to herein as "memory", such as volatile and non-volatile computer memory, such as RAM, PROM, EPROM and EEPROM, compact disks, optical disks, optical disks, magnetic tapes, etc.). In some implementations, the storage media may be encoded with one or more programs that, when executed on one or more processors and / or controllers, perform at least some of the functions discussed herein. The various storage media may be fixed within the processor or controller, or may be transferable so that one or more programs stored thereon may be loaded into a processor or controller in order to implement various aspects of the various embodiments discussed herein. The term "program" or "computer program" is used herein in a general sense to refer to any type of computer code (e.g., software or microcode) that can be employed to program one or more processors or controllers.

[0021] It should be understood that all combinations of the above-described concepts and the additional concepts discussed in more detail below (assuming such concepts are not mutually inconsistent) are contemplated as being part of the inventive subject matter disclosed herein. In particular, all combinations of the subject matter of the claims are contemplated as being part of the inventive subject matter disclosed herein. It should also be understood that terms explicitly employed herein, which may also appear in any disclosure incorporated by reference, should be given the meaning that best matches the specific concepts disclosed herein.

[0022] These and other aspects of the embodiments will be apparent from and elucidated with reference to the embodiment(s) described hereinafter. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In the drawings, like reference numerals generally refer to the same parts throughout the different views. Also, the drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the various embodiments.

[0024] Figure 1 is a flow chart of a method for characterizing gene expression levels according to one embodiment.

[0025] Figure 2 is a schematic diagram of transcript expression estimation using unique features of gene transcripts according to one embodiment.

[0026] Figure 3 is a schematic diagram of a system and method for characterizing the expression level of a gene or gene transcript according to one embodiment.

[0027] Figure 4 is a schematic representation of a system for characterizing gene expression levels according to one embodiment. DETAILED DESCRIPTION

[0028] The present disclosure describes various embodiments of the system and method for compiling information about gene transcript expression levels using unique features extracted from gene transcripts. More generally, the applicant has recognized and appreciated that it would be beneficial to provide a system that can use RNA sequencing data to quickly and effectively characterize gene transcript expression levels. The system includes a unique feature database, and the unique feature database stores the unique features extracted from the gene transcript, including but not limited to unique exons, unique exon connections, unique introns, unique start positions and / or unique end positions and many other unique features. The system receives the gene transcript or performs sequencing on it, and compares the sequence with the unique features extracted and stored in the unique feature database. If at least a portion of the sequence matches with one or more unique features extracted, the gene transcript from which the sequence is generated is identified. Like this, the system can compile information about gene transcript expression levels from the source of RNA sequencing data.

[0029] refer to Figure 1 , in one embodiment, is a flow chart of a method 100 for characterizing the expression level of gene transcripts using RNA sequencing data. In step 110 of the method, unique features are extracted from the gene transcripts. According to one embodiment, for most or all transcripts in the target or studied transcriptome, the system can scan the transcripts obtained by sequencing and / or identified based on genetic analysis, and these transcripts can be compared to identify unique features. The system can obtain results from the transcription and / or alternative splicing of a single gene using only the unique features found based on this comparison. Alternatively, the system can use the discovered unique features that are generated by the transcription and / or alternative splicing of two or more genes. For example, there may be a threshold value for determining how many genes or alternative splicing can be found before and / or after which the feature will be identified or will not be identified as a sufficiently unique feature for the methods described or contemplated herein.

[0030] Unique features are parameters of RNA sequences, and the parameters are produced by the splicing of the gene transcribed by RNA. In many cases, the parameters come from the alternative splicing of the gene transcribed by RNA. For example, the unique features of a gene transcript can be produced by unique exons, and the unique exons can be exons unique to a subset of the transcript of a gene. The unique features of a gene transcript may come from unique exon connections, which may be exon connections unique to a subset of the transcript of a gene, such as from skipping exons in other processes. The unique features of a gene transcript may be caused by unique intron retention events, and intron retention events may be caused by one or more introns remaining in the transcript. The unique features of a gene transcript may be produced by unique transcription initiation and / or termination sites, because different transcripts from a gene can start and / or terminate at different positions along the gene.

[0031] As described herein, quantifying these unique identifiers can effectively solve the deconvolution problem usually caused by RNA sequencing data. For example, even if the degraded RNA is sequenced, as long as the unique features are still covered by enough reads, the expression of the transcript can still be evaluated accordingly. In addition, the extracted unique features can only include a part of the total information found in the entire transcriptome of the organism from which the RNA sequencing data is obtained. This further solves many problems faced by existing systems and significantly reduces the computing time. It can also quickly screen a large amount of RNA sequencing data in a short time.

[0032] In step 120 of the method, the extracted unique features are stored in a unique feature database. The unique feature database may be part of the system, or may be located away from the system. For example, the unique feature database may be a database or memory associated with a processor or other component of the system. Alternatively, the unique feature database may be a database or memory away from the system that uses unique features to characterize RNA sequencing data. For example, the generated unique feature database may be utilized by one or more systems, some or all of which may be decentralized relative to the database or memory to perform the analysis described herein or otherwise contemplated. Therefore, the system may include or communicate with a wired and / or wireless communication system to facilitate communication between the system and a remote database or memory. The extracted unique features may be stored in a unique feature database for retrieval and downstream use, or may be stored in a unique feature database in a format that enables rapid search of RNA sequencing data and / or comparison or alignment of RNA sequencing data with the extracted unique features. According to one embodiment, the unique feature database includes extracted unique features instead of complete gene transcripts, which helps to quickly identify genes and / or gene transcripts.

[0033] In step 122 of the method, one or more unique features in the unique feature database are associated with annotation information. For example, the unique features can be tagged, marked, annotated, or otherwise associated with information about the gene from which the tag was extracted and / or the gene from which the transcript was extracted in a memory. The annotation information can include information about the location of the unique feature or the associated transcript in the genome, information about the organism from which the unique feature was extracted, information about alternative splicing of the gene from which the unique feature was extracted, and / or other information about the source of the unique feature, the location of the unique feature.

[0034] In step 130 of the method, RNA is sequenced or RNA sequencing data is obtained. For example, RNA can be sequenced from a sample that contains or may contain ribonucleic acid. Therefore, according to one embodiment, in step 128 of the method, a sample is provided for nucleic acid extraction and analysis. The sample can be composed of ribonucleic acid from one or more cells of one or more microorganisms (e.g., bacteria, viruses, fungi) and / or from plants or animals and many other sources. The sample can include ribonucleic acid molecules from one or more organisms. Samples can be obtained in clinical settings, environments, indoor or outdoor surfaces or any other source. It should be recognized that there is no restriction on the source of the sample or the ribonucleic acid in the sample. Any preparation method can be used to prepare samples and / or ribonucleic acid therein for sequencing, and the method can depend at least in part on the sequencing platform. According to one embodiment, in many other preparations or disposals, ribonucleic acid can be extracted, purified and / or amplified.

[0035] The system may include a sequencing platform configured to sequence at least a portion of the RNA from the sample. Any method and / or platform for sequencing RNA can be used to obtain RNA sequencing data. Therefore, the sequencing platform can be any sequencing platform, including but not limited to any system described or envisioned herein. According to one embodiment, the sequencing platform may include a controller or other analysis module for downstream analysis and characterization. According to another embodiment, the RNA sequencing data generated by the sequencing platform is transmitted to a local or remote controller or other analysis module in real time or at certain time points for downstream analysis and characterization.

[0036] Alternatively, the system can retrieve or otherwise receive RNA sequencing data from a remote sequencing platform or from a database or memory comprising stored RNA sequencing data. For example, the system can communicate with a local and / or remote database or memory comprising stored RNA sequencing data, or can receive an upload or other delivery of RNA sequencing data. Thus, the analyses described or envisioned herein can be obtained while the RNA sequencing data is being obtained and / or can be obtained after the RNA sequencing data is obtained.

[0037] At step 140 of the method, the system compares the sequenced or obtained sequence with the extracted unique features stored in the unique feature database. For example, the system may include a processor or other computing component configured or programmed to compare the sequenced or obtained sequence with the extracted unique features stored in the unique feature database. The comparison may be performed, for example, by aligning the sequenced or obtained sequence with one or more of the extracted unique features in the unique feature database or in a memory or processor.

[0038] According to one embodiment, the system can utilize an algorithm to compare the sequence that has been sorted or obtained with the unique features extracted. For example, a splicing quantification algorithm, such as SpliceTrap, which quantifies exon inclusion levels using paired-end RNA sequencing data, or MISO (mixture of isomers) that identifies differentially regulated isomers or exons in a sample, can be optionally modified for use. For example, a splicing quantification algorithm can quantify known or novel alternative splicing events from RNA sequencing reads. These are suitable for quantifying unique features and can be used and / or modified to estimate the ratio and expression of unique features. Reads about exon junctions and unique regions may be important, and the algorithm can be used to find the best solution. According to one embodiment, cassette exons can be skipped in certain transcripts, and their inclusion rate and expression level can be studied by examining the reads of (one or more) intermediate exons and / or exon junctions.

[0039] In step 150 of the method, the following gene transcripts are identified and / or quantified based on the match between the sequence and the extracted unique features, and the sequence is generated from the gene transcripts. According to one embodiment, there may be a threshold or probability requirement for the positive identification of the gene transcript, which can optionally be based on the quality of the unique features identified, the quantity of unique features and / or other parameters. According to one embodiment, the system quantifies the gene transcripts while identifying the gene transcripts, or in addition to identifying them. For example, the system counts, tracks, records or otherwise quantifies the identified gene transcripts, which helps to promote information about the expression of the gene transcripts based on the relative expression measured from the unique features. For example, a splicing quantitative algorithm can be used for quantitative gene transcripts.

[0040] According to one embodiment, the sequence is matched to one or more unique features extracted from two or more different gene transcripts. For example, in some embodiments, a short sequence may contain unique features found in several different gene transcripts, but lacks additional sequence information that may distinguish between complete transcripts. Therefore, the identification step 150 may include identifying two or more transcripts from which the sequence was generated or may have been generated. The system can be configured to report only transcripts that can be clearly defined, or can report sequences that potentially identify multiple transcripts.

[0041] refer to Figure 2 , which is a schematic diagram 200 of transcript expression estimation using unique features of gene transcripts, in one embodiment. A gene 10 contains at least three different transcripts (n1, n2, and n3), each of which contains a different set of exons 20. According to one embodiment, the three different transcripts of the gene can be distinguished by two unique features 30, a skipped exon 50, and an alternative splice site 60. For example, the presence of unique feature 50 in comparison 42 enables the read to be identified as n2 to n1 or n3. As another example, the presence of unique feature 60 in comparison 44 enables the read to be identified as n3 to n1 or n2. The expression of transcripts n1, n2, and n3 can be resolved by looking at each feature separately and then combining the observations.

[0042] In step 160 of the method, the system compiles information about gene transcripts and / or gene expression levels based on gene transcripts and / or genes identified from the analyzed RNA sequences. According to one embodiment, when each sequence is identified in step 150 of the method, the system can track, record, store or otherwise count specific gene transcripts or genes. Transcript expression levels can be summarized in any format, including standard formats, such as FPKM values ​​and many other formats. Collect and summarize feature quantification to explain transcript expression based on the relationship between features and transcripts. In complex cases, linear models can be used to solve matrices. When there is a conflict between the results summarized from different features due to uneven distribution of transcription results over the entire transcript, some representative values, such as mean or maximum values, can be used. According to one embodiment, the compilation includes annotation information from a unique feature database. According to one embodiment, the system can report transcript expression levels as probability information or with probability information, and the probability information includes the probability that the identified transcript is a transcript as follows: the sequence is generated from the transcript.

[0043] As described herein, the extracted unique features can be used as markers for certain gene transcripts and / or gene expression profiles. One advantage of using unique features is that they can combine views from the gene level and the splicing level. In addition, the quantification of the unique features from a gene can be used to simulate the expression pattern of the transcripts from the gene. In fact, this operation can be performed even if the actual expression value of the transcript is not known.

[0044] refer to Figure 3 300 is a schematic diagram of a system and method for characterizing gene transcript expression levels as described herein or otherwise contemplated. The system includes a unique feature database 320, which includes unique features 322, which are extracted from a gene structure 310, as described herein or otherwise contemplated. The unique feature database 320 may also include one or more feature annotations 324 associated with the extracted unique features 322. A plurality of RNA sequencing reads 330 are obtained by sequencing or by receiving sequencing data, and compared to the unique features 322 extracted from the unique feature database 320 at 340. Genes and / or gene transcripts are compiled, summarized, or otherwise characterized using the feature annotations in the unique feature database 320, thereby obtaining transcript expression levels 350.

[0045] refer to Figure 4, which is a schematic diagram of a system 400 for characterizing gene transcript expression levels in one embodiment. System 400 includes one or more of a processor 420, a memory 426, a user interface 440, a communication interface 450, and a memory 460 interconnected via one or more system buses 410. In some embodiments, such as those in which the system includes or implements a sequencer or sequencing platform, the hardware may include additional sequencing hardware 415, which may be any sequencer or sequencing platform. It should be understood that Figure 4 Some aspects are abstracted, and the actual organization of components of system 400 may be different and more complex than shown.

[0046] According to one embodiment, the system 400 includes a processor 420 capable of executing instructions stored in a memory 426 or a storage device 460 or otherwise processing data. The processor 420 performs one or more steps of the method and may include one or more modules described or otherwise contemplated herein. The processor 420 may be formed of one or more modules and may include, for example, a memory 426. The processor 420 may take any suitable form, including but not limited to a microprocessor, a microcontroller, multiple microcontrollers, a circuit, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a single processor, or multiple processors.

[0047] Memory 426 can take any suitable form, including non-volatile memory and / or RAM. Memory 426 can include various memories, such as cache or system memory. In this way, memory 426 can include static random access memory (SRAM), dynamic RAM (DRAM), flash memory, read-only memory (ROM) or other similar memory devices. The memory can store operating systems, etc. The processor uses RAM to temporarily store data. According to one embodiment, the operating system can include code that controls the operation of one or more components of system 400 when executed by the processor. It is obvious that in an embodiment where the processor implements one or more functions described herein in hardware, the software described in other embodiments as corresponding to such functions can be omitted.

[0048] User interface 440 may include one or more devices for realizing communication with users such as administrators. User interface may be any device or system that allows to convey and / or receive information, and may include a display, mouse and / or keyboard for receiving user commands. In certain embodiments, user interface 440 may include a command line interface or a graphical user interface that may be presented to a remote terminal via a communication interface. The user interface may be located together with one or more other components of the system, or may be located at a position away from the system and communicate via a wired and / or wireless communication network.

[0049] The communication interface 450 may include one or more devices for implementing communication with other hardware devices. For example, the communication interface 450 may include a network interface card (NIC) configured to communicate according to the Ethernet protocol. In addition, the communication interface 450 may implement a TCP / IP stack for communicating according to the TCP / IP protocol. Various alternative or additional hardware or configurations for the communication interface 450 will be apparent.

[0050] The storage device 460 may include one or more machine-readable storage media, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium, an optical storage medium, a flash memory device, or a similar storage medium. In various embodiments, the storage device 460 may store instructions for execution by the processor 420 or data that the processor 420 may operate on. For example, the storage device 460 may store an operating system 461 for controlling various operations of the system 400. In the case where the system 400 implements a sequencer and includes sequencing hardware 415, the storage device 460 may include sequencing instructions 462 for operating the sequencing hardware 415. According to one embodiment, the storage device 460 may include a unique feature database 464 that has been extracted according to the methods described or otherwise contemplated herein.

[0051] It will be apparent that various information stored in memory 460 may be stored in addition or alternatively in memory 426. In this regard, memory 426 may also be considered to constitute a storage device, and storage device 460 may be considered to be a memory. Various other arrangements will be apparent. In addition, both memory 426 and memory 460 may be considered to be non-transitory machine-readable media. As used herein, the term non-transitory will be understood to exclude transient signals but include all forms of storage devices, including volatile and non-volatile memory.

[0052] Although system 400 is shown as including one of each described component, in various embodiments, various components may be multiple. For example, processor 420 may include multiple microprocessors configured to independently perform the methods described herein, or configured to perform the steps or subroutines of the methods described herein, so that multiple processors cooperate to implement the functions described herein. In addition, in the case of implementing system 400 in a cloud computing system, various hardware components may belong to separate physical systems. For example, processor 420 may include a first processor in a first server and a second processor in a second server. Many other variations and configurations are possible.

[0053] According to one embodiment, the processor 420 includes one or more modules to perform one or more functions or steps of the methods described or otherwise contemplated herein. For example, the processor 420 may include a feature extraction module 422, a comparison module 424, and / or an assembly module 428. According to one embodiment, the feature extraction module 422 analyzes genes and / or gene transcripts to identify one or more parameters of RNA sequences, the parameters being generated by splicing of genes that transcribe RNA, including but not limited to alternative splicing of genes from which RNA is transcribed. Any method for feature identification from genes and / or gene transcripts may be used to extract the unique features. According to one embodiment, the system may only utilize unique features that are found to be due to transcription and / or alternative splicing from a single gene. Alternatively, the system may utilize unique features that are found to be generated by transcription and / or alternative splicing of two or more genes. For example, there may be a threshold value for determining how many genes or alternative splicing can be found before and / or after which a feature will be identified or will not be identified as a sufficiently unique feature for the methods described or contemplated herein. The extracted unique features may be the result of unique exon junctions, unique intron retention events, unique transcription start and / or stop sites, etc., among many other features. Once extracted, the unique features may be stored in a unique feature database 464 or other memory. In some embodiments, the unique features are stored remotely from one or more other components of the system.

[0054] According to one embodiment, the processor 420 includes a comparison module 424. According to one embodiment, the comparison module 424 compares the sequence ordered or obtained with the extracted unique features stored in the unique feature database 464. The comparison can be performed, for example, by comparing the RNA sequence with one or more of the extracted unique features in the unique feature database or in the memory or processor. According to one embodiment, the system can use an algorithm to compare the sequence ordered or obtained with the extracted unique features. The comparison module 424 can identify the gene transcript from which the sequence is generated based on the match between the sequence and the extracted unique features, and / or can identify the gene from which the gene transcript is transcribed. According to one embodiment, there may be a threshold or probability requirement for the positive identification of the gene transcript and / or gene, which can be optionally based on the quality of the identified unique features, the number of unique features and / or other parameters. The comparison module 424 can count, track, record or otherwise quantify the gene transcript, which helps to promote information about the expression of the gene transcript based on the relative expression measured from the unique features. Among other methods, the comparison module 424 can use a splicing quantification algorithm to quantify the gene transcript.

[0055] According to one embodiment, processor 420 includes compilation module 428. According to one embodiment, compilation module 428 compiles or summarizes information about gene transcripts and / or gene expression levels based on identified gene transcripts and / or identified genes generated or transcribed sequences. According to one embodiment, when analyzing each sequence, the system can track, record, store or otherwise count specific gene transcripts or genes. Transcript expression levels can be summarized in any format, including standard formats, such as FPKM values ​​and many other formats. According to one embodiment, compilation module 428 retrieves, compiles and / or summarizes annotation information from a unique feature database associated with identified gene transcripts and / or identified genes.

[0056] According to one embodiment, the system described herein or otherwise contemplated has significant functional advantages over existing systems in both efficiency and accuracy. For example, by improving the identification of gene transcripts, the system provides significant computational efficiency over existing systems. By using information only in small regions, rather than reading all information from the transcripts, gene expression estimation can be simplified to quantifying local key elements. This enables the system to perform improved high-throughput screening of RNA sequencing data.

[0057] According to another embodiment, the systems described herein or otherwise contemplated improve existing systems by enabling determination of transcript expression levels from incomplete RNA, which is common in low-quality RNA sequencing data and scRNA sequencing data. The methods described herein avoid bias from regions where transcription is very high or very low.

[0058] According to another embodiment, the system described or otherwise contemplated herein improves existing systems in which unique features are associated with phenotypes. Quantification of these features provides higher resolution than gene expression. It may also be more robust because unique features may be able to capture the effects of unknown transcript variants because more detailed patterns can be revealed through these local measurements. Similarly, unique features can be used as additional evidence for clustering RNA sequencing samples, such as for subpopulation inference on scRNA sequencing data in other processes.

[0059] All definitions, as defined and used herein, should be understood to control over dictionary definitions, definitions in documents incorporated by reference, and / or ordinary meanings of the defined terms.

[0060] As used herein in the specification and claims, the terms "a" and "an" should be understood to mean "at least one" unless expressly stated otherwise.

[0061] The phrase "and / or" as used in the specification and claims herein should be understood to mean "one or both" of the elements so combined, i.e., the elements are present in combination in some cases, and separately in other cases. Multiple elements listed with "and / or" should be interpreted in the same manner, i.e., "one or more" elements so connected. Optionally, other elements may be present in addition to the elements specifically identified by the "and / or" clause, whether related or unrelated to those elements specifically identified.

[0062] As used herein in the specification and claims, "or" should be understood to have the same meaning as "and / or" defined above. For example, when separating items in a list, "or" or "and / or" should be interpreted as inclusive, that is, including at least one of several elements or a list of elements, but also including more than one, and optionally, additional unlisted items. Only when the opposite term is clearly indicated, such as "only one" or "exactly one", or when "consisting of..." is used in the claims, it will refer to an exact element in a list of several elements or elements. In general, the term "or" used in this article should be interpreted as indicating exclusive alternatives only when it is preceded by an exclusive term (i.e., "one or the other but not both"), such as "either", "one of", "only one of" or "exactly one of".

[0063] As used herein in the specification and claims, the phrase "at least one", in reference to a list of one or more elements, should be understood to mean at least one element selected from one or more of the elements in the listed elements, but not necessarily including at least one of each and every element specifically listed in the list of elements, and not excluding any combination of elements in the listed elements. This definition also allows for the optional presence of elements other than the elements specifically identified in the list of elements to which the phrase "at least one" refers, whether related or unrelated to the specifically identified elements.

[0064] It should also be understood that in any method claimed herein that includes more than one step or action, the order of the method steps or actions is not necessarily limited to the order in which the method steps or actions are described unless explicitly stated to the contrary.

[0065] In the claims and the above description, all transitional phrases such as "include", "comprising", "carrying", "having", "containing", "involving", "holding" shall be understood as open-ended, i.e., meaning including but not limited to the above. Only the transitional phrases "consisting of" and "consisting essentially of" shall be closed or semi-closed transitional phrases, respectively.

[0066] Although several innovative embodiments have been described and illustrated herein, a person skilled in the art will readily envision a variety of other ways and / or structures for performing the functions and / or obtaining the results and / or one or more of the advantages described herein, and each of such variations and / or modifications is shown as being within the scope of the innovative embodiments described herein. More generally, a person skilled in the art will readily recognize that all parameters, dimensions, materials, and configurations described herein are intended to be exemplary, and that the actual parameters, dimensions, materials, and / or configurations will depend on the specific application or application to which the innovative teachings are used. A person skilled in the art will recognize or be able to determine many equivalents of the specific inventive embodiments described herein using no more than routine experimentation. Therefore, it should be understood that the foregoing embodiments are presented only by way of example, and within the scope of the appended claims and their equivalents, the inventive embodiments may be practiced in a manner different from that specifically described and claimed. The innovative embodiments of the present disclosure relate to each individual feature, system, article, material, complete set of equipment, and / or method described herein. In addition, any combination of two or more such features, systems, articles, materials, kits and / or methods, if such features, systems, articles, materials, kits and / or methods are not mutually inconsistent, are included within the innovative scope of the present disclosure.

Claims

1. A method for characterizing the expression level of a gene transcript, comprising: extracting one or more unique features from each gene transcript of a plurality of gene transcripts, wherein the unique features result from alternative splicing of the gene and include unique exons, the unique exons including skipped exons, wherein the unique exons are unique to a subset of transcripts from a gene; storing the extracted unique features in a unique feature database; receiving a plurality of sequences generated from gene transcripts sequenced from a cell, wherein at least some of the sequences include one or more of the extracted unique features; comparing, by a processor, the plurality of sequences to the extracted unique signatures stored in the unique signature database; identifying, based on a match between the sequence and the extracted unique features, a gene transcript from which the sequence was generated; and Based on the identified gene transcripts, information on the expression levels of the gene transcripts is compiled.

2. The method according to claim 1, wherein: The unique features also include one or more of the following: unique introns, unique transcription start positions, and / or unique transcription stop positions.

3. The method according to claim 1, wherein: Comparing includes aligning each sequence in the plurality of sequences to one or more unique features.

4. The method of claim 1, further comprising the step of quantifying the identified gene transcripts.

5. The method of claim 1, further comprising the step of sequencing gene transcripts from one or more cells to generate the plurality of sequences.

6. The method according to claim 1, further comprising the steps of: In the unique feature database, at least some of the extracted unique features are associated with annotation information.

7. The method according to claim 1, wherein: The unique signature database includes extracted unique signatures rather than complete gene transcripts.

8. The method according to claim 1, wherein: The identifying step includes the likelihood that the identified gene transcript is the transcript from which the sequence was generated.

9. The method according to claim 1, wherein: The sequence matches unique features extracted from two different gene transcripts, and the identifying step includes identifying two or more gene transcripts from which the sequence is generated, or could have been generated.

10. A system for characterizing the expression level of a gene transcript, comprising: a feature extraction module configured to extract a unique feature from each gene transcript of a plurality of gene transcripts generated by sequencing gene transcripts from a cell, wherein the unique feature is generated by alternative splicing of a gene and comprises unique exons, the unique exons comprising skipped exons, wherein the unique exons are unique to a subset of transcripts from a gene; a database of unique features extracted from each gene transcript in a plurality of gene transcripts; a comparison module configured to: (i) compare a plurality of sequences sequenced from gene transcripts with the extracted unique features stored in the unique feature database; and (ii) identify, based on a match between the sequences and the extracted unique features, a gene transcript and / or a gene from which the sequence is generated; and A compilation module is configured to compile information about the expression levels of gene transcripts based on the identified gene transcripts.

11. The system according to claim 10, wherein: The feature extraction module is further configured to associate at least some of the extracted unique features with annotation information.

12. The system according to claim 10, wherein: The unique features stored in the unique feature database further include one or more of the following: unique introns, unique transcription start positions and / or unique transcription end positions.

13. The system according to claim 10, wherein: Comparing includes aligning each sequence in the plurality of sequences to one or more unique features.

14. The system according to claim 10, wherein: The sequence matches unique features extracted from two different gene transcripts, and the identifying step includes identifying two or more gene transcripts from which the sequence is generated, or could have been generated.

Citation Information

Patent Citations

  • Methods and systems for annotating biomolecular sequences

    US20040142325A1

  • Fusion transcript detection methods and fusion transcripts identified thereby

    US20160078168A1